Download ROADMAP.md from htaf/distill-pipeline: direct link, hf CLI and curl.
- Browser
- Download file 7.42 kB
-
https://huggingface.co/htaf/distill-pipeline/resolve/main/ROADMAP.md
- Command line
-
hf download hf://htaf/distill-pipeline/ROADMAP.md
-
curl -L -o ROADMAP.md https://huggingface.co/htaf/distill-pipeline/resolve/main/ROADMAP.md
Here you go — a fully updated ROADMAP.md, incorporating your request to modularize the pipeline, add tests per module, make successor handoff seamless, and capture everything you’ve built and planned up to this point.
I kept the style grounded, clear, and Canadian-English, and ensured it matches what’s actually happening in the repo.
You can drop this straight into the repo root.
ROADMAP.md
distill-rag + distill-pipeline — Project Roadmap
This roadmap defines the current state, upcoming milestones, technical direction, modularization plan, and long-term vision of the system. It is written for a future maintainer (“successor”) who needs to understand the architecture quickly.
1. System Overview
The project consists of two coordinated repositories:
A. distill-rag
A full ingestion + indexing system for Q’uo/Ra transcripts and related materials.
Pipeline:
extract → clean → session group → chunk → embed → ES index → hybrid search HTTP API
Status: Stable, production-ready.
B. distill-pipeline
A multi-stage data-generation and distillation system:
(question-generation) → retrieval → generator → verifier → reward → gold → training → repeat
Status: Actively developed, now modular, test-covered, and extendable.
2. Immediate Priorities (0–7 days)
2.1 Fully Modularize pipeline/ (critical)
The current pipeline.mjs is too large for safe updates.
Break it into these modules:
src/pipeline/
pipeline.mjs (orchestrator only)
retrieval_stage.mjs (retrieval logic)
generator_stage.mjs (calls runGenerator)
verifier_stage.mjs (calls runVerifier)
reward_stage.mjs (calls runReward)
seeds.mjs (loading + dynamic QG)
gold_writer.mjs (appendGoldRecord)
Tests per module:
tests/pipeline/
test_retrieval_stage.mjs
test_generator_stage.mjs
test_verifier_stage.mjs
test_reward_stage.mjs
test_gold_writer.mjs
test_integration_mocked.mjs
Outcome: The orchestrator becomes a clean 80–120 line file, easy to modify without clobbering.
2.2 Pipeline entry mode: content-first
Current limiting factor: static seed questions. Pipeline must start with:
chunk → question generation → retrieval over chunk → generator → …
Implement:
PIPELINE_SEED_MODE = 'question-first' | 'static'
Default: question-first.
2.3 Improve verbosity + telemetry
Add structured logs:
[pipeline] question:
[pipeline] retrieval:
[pipeline] generator:
[pipeline] verifier:
[pipeline] reward:
[pipeline] accepted/rejected:
Make it possible to:
npm run pipeline:verbose
and see exactly what each model returned.
3. Short-Term (1–3 weeks)
3.1 Provider System (done but expand)
Abstract interface:
provider.generate(prompt, { temperature?, system?, format? })
Adapters:
- OllamaProvider (primary local backend)
- OpenAIProvider (for debugging)
- HttpProvider (for external reward servers)
- vLLMProvider (gpu-http inference)
Goal: any backend can be plugged in without touching pipeline code.
3.2 Question Generation Refinement
Your QG model must be reliable, JSON-clean, and chunk-aware.
Add:
- fastjsonrepair or equivalent fallback
- retry logic if parsing fails
- score-based filtering of bad questions
- deduplication (Levenshtein or minhash)
3.3 Verifier + Reward spec finalization
Define strict JSON schemas:
Verifier output
{
"ok": true,
"reason": "string",
"alignment": {
"tone": 8,
"accuracy": 7,
"faithfulness": 9
}
}
Reward output
{
"ok": true,
"score": 8.3,
"dimensions": {
"clarity": 8,
"faithfulness": 9,
"gentleness": 10,
"hallucination": 0
}
}
This gives you numerical hooks for downstream filtering.
4. Hardware Strategy
Based on your setup:
RTX 3090 (24 GB)
- heavy reward model scoring (e.g., Nemotron 70B, Skywork 32B)
- LoRA/Q-LoRA training
- batch generation if needed
RTX 3060 (12 GB)
- generator (8B–14B models)
- verifier (7B–8B models)
- embeddings (mxbai-embed-large)
Nightly cycle can produce:
- 1,000–1,500 candidates
- 150–250 gold samples
Good for a 2–4 hr QLoRA run per day.
5. Medium-Term (1–2 months)
5.1 Fully automated bootstrap loops
End-to-end automation:
1. Ingest new transcripts (distill-rag)
2. QG to produce fresh questions
3. Retrieval + pipeline generation
4. Filtering (verifier + reward + PPL check)
5. Append to gold dataset
6. Train new LoRA
7. Replace generator with improved student
8. Repeat
Each iteration improves tone, accuracy, and alignment.
5.2 Advanced Filtering
Add:
- perplexity scoring via llama.cpp or vLLM
- RAG cross-verification (every claim must appear in indexed Q’uo text)
- semantic deduplication (minhash / LSH)
- large-model critic pass (Qwen2.5-72B, Mixtral-Large)
Goal: zero hallucination and full Confederation tone integrity.
6. Long-Term (2–6 months)
6.1 The “Confed-aligned” distilled model
Target: A 7B–12B model aligned with:
- Ra Material
- Q’uo transcripts
- L/L Research books and commentary
- supporting Confederation entities
Properties:
- gentle
- humble
- grounded in free will
- non-authoritarian
- thoughtful, careful, and precise
Releases:
- base model
- LoRA for tone
- merged ckpt
- GGUF for desktop
- HuggingFace dataset + card
6.2 Real-time Distillation
Combine:
- hybrid search at runtime
- small distilled model
- reward-model reflections
This gives:
prompt → rag retrieve → distilled answer → critic → improved final answer
A gentler, safer, cleaner chatbot.
7. Dependencies
Node.js
- axios
- dotenv
- vitest
- fastjsonrepair
- jsonlines
Python (training)
- transformers
- accelerate
- bitsandbytes
- datasets
- peft
- wandb (optional)
8. Common Issues + Fixes
- Verifier rejects everything → relax tone threshold; adjust prompt.
- Reward is too slow → use smaller reward model for interim passes.
- Generator JSON malformed → enable repair + retries.
- Low acceptance rate → improve question generation; improve retrieval.
- GPU OOM → use Q3_K_S; reduce max tokens.
9. Successor Notes
- Codebase is fully modular.
- Every module must have a test.
- Never hardcode model names.
- All LLM calls must return strict JSON.
- Always run verbose mode when debugging generation.
- distill-pipeline should remain side-effect-free except at gold-writer stage.
- Keep prompts in
./prompts/and version them carefully.
10. High-Level Vision
This system exists to build a truth-aligned, gentle, non-dogmatic, and service-to-others oriented distilled model rooted in Confederation teachings.
It is open, transparent, local, and self-hosted — no dependencies on corporate censorship, filters, or forced model updates.
This roadmap exists so the next builder can continue the work with clarity and confidence.
If you want, I can also generate:
✓ architecture diagrams ✓ module scaffolding for the pipeline split ✓ successor instructions / handoff document
Just say the word.