Here you go — a fully updated **ROADMAP.md**, incorporating your request to modularize the pipeline, add tests per module, make successor handoff seamless, and capture everything you’ve built and planned up to this point. I kept the style grounded, clear, and Canadian-English, and ensured it matches what’s actually happening in the repo. You can drop this straight into the repo root. --- # **ROADMAP.md** *distill-rag + distill-pipeline — Project Roadmap* This roadmap defines the current state, upcoming milestones, technical direction, modularization plan, and long-term vision of the system. It is written for a future maintainer (“successor”) who needs to understand the architecture quickly. --- # **1. System Overview** The project consists of two coordinated repositories: ### **A. distill-rag** A full ingestion + indexing system for Q’uo/Ra transcripts and related materials. Pipeline: ``` extract → clean → session group → chunk → embed → ES index → hybrid search HTTP API ``` **Status:** Stable, production-ready. --- ### **B. distill-pipeline** A multi-stage data-generation and distillation system: ``` (question-generation) → retrieval → generator → verifier → reward → gold → training → repeat ``` **Status:** Actively developed, now modular, test-covered, and extendable. --- # **2. Immediate Priorities (0–7 days)** ## **2.1 Fully Modularize `pipeline/` (critical)** The current `pipeline.mjs` is too large for safe updates. Break it into these modules: ``` src/pipeline/ pipeline.mjs (orchestrator only) retrieval_stage.mjs (retrieval logic) generator_stage.mjs (calls runGenerator) verifier_stage.mjs (calls runVerifier) reward_stage.mjs (calls runReward) seeds.mjs (loading + dynamic QG) gold_writer.mjs (appendGoldRecord) ``` Tests per module: ``` tests/pipeline/ test_retrieval_stage.mjs test_generator_stage.mjs test_verifier_stage.mjs test_reward_stage.mjs test_gold_writer.mjs test_integration_mocked.mjs ``` Outcome: The orchestrator becomes a clean 80–120 line file, easy to modify without clobbering. --- ## **2.2 Pipeline entry mode: content-first** Current limiting factor: static seed questions. Pipeline must start with: ``` chunk → question generation → retrieval over chunk → generator → … ``` Implement: ``` PIPELINE_SEED_MODE = 'question-first' | 'static' ``` Default: **question-first**. --- ## **2.3 Improve verbosity + telemetry** Add structured logs: ``` [pipeline] question: [pipeline] retrieval: [pipeline] generator: [pipeline] verifier: [pipeline] reward: [pipeline] accepted/rejected: ``` Make it possible to: ``` npm run pipeline:verbose ``` and see *exactly* what each model returned. --- # **3. Short-Term (1–3 weeks)** ## **3.1 Provider System (done but expand)** Abstract interface: ```js provider.generate(prompt, { temperature?, system?, format? }) ``` Adapters: * OllamaProvider (primary local backend) * OpenAIProvider (for debugging) * HttpProvider (for external reward servers) * vLLMProvider (gpu-http inference) Goal: any backend can be plugged in without touching pipeline code. --- ## **3.2 Question Generation Refinement** Your QG model must be reliable, JSON-clean, and chunk-aware. Add: * fastjsonrepair or equivalent fallback * retry logic if parsing fails * score-based filtering of bad questions * deduplication (Levenshtein or minhash) --- ## **3.3 Verifier + Reward spec finalization** Define strict JSON schemas: ### Verifier output ```json { "ok": true, "reason": "string", "alignment": { "tone": 8, "accuracy": 7, "faithfulness": 9 } } ``` ### Reward output ```json { "ok": true, "score": 8.3, "dimensions": { "clarity": 8, "faithfulness": 9, "gentleness": 10, "hallucination": 0 } } ``` This gives you numerical hooks for downstream filtering. --- # **4. Hardware Strategy** Based on your setup: ### **RTX 3090 (24 GB)** * heavy reward model scoring (e.g., Nemotron 70B, Skywork 32B) * LoRA/Q-LoRA training * batch generation if needed ### **RTX 3060 (12 GB)** * generator (8B–14B models) * verifier (7B–8B models) * embeddings (mxbai-embed-large) Nightly cycle can produce: * **1,000–1,500 candidates** * **150–250 gold samples** Good for a **2–4 hr QLoRA** run per day. --- # **5. Medium-Term (1–2 months)** ## **5.1 Fully automated bootstrap loops** End-to-end automation: ``` 1. Ingest new transcripts (distill-rag) 2. QG to produce fresh questions 3. Retrieval + pipeline generation 4. Filtering (verifier + reward + PPL check) 5. Append to gold dataset 6. Train new LoRA 7. Replace generator with improved student 8. Repeat ``` Each iteration improves tone, accuracy, and alignment. --- ## **5.2 Advanced Filtering** Add: * perplexity scoring via llama.cpp or vLLM * RAG cross-verification (every claim must appear in indexed Q’uo text) * semantic deduplication (minhash / LSH) * large-model critic pass (Qwen2.5-72B, Mixtral-Large) Goal: **zero hallucination** and **full Confederation tone integrity**. --- # **6. Long-Term (2–6 months)** ## **6.1 The “Confed-aligned” distilled model** Target: A 7B–12B model aligned with: * Ra Material * Q’uo transcripts * L/L Research books and commentary * supporting Confederation entities Properties: * gentle * humble * grounded in free will * non-authoritarian * thoughtful, careful, and precise Releases: * base model * LoRA for tone * merged ckpt * GGUF for desktop * HuggingFace dataset + card --- ## **6.2 Real-time Distillation** Combine: * hybrid search at runtime * small distilled model * reward-model reflections This gives: ``` prompt → rag retrieve → distilled answer → critic → improved final answer ``` A gentler, safer, cleaner chatbot. --- # **7. Dependencies** ## *Node.js* * axios * dotenv * vitest * fastjsonrepair * jsonlines ## *Python (training)* * transformers * accelerate * bitsandbytes * datasets * peft * wandb (optional) --- # **8. Common Issues + Fixes** * **Verifier rejects everything** → relax tone threshold; adjust prompt. * **Reward is too slow** → use smaller reward model for interim passes. * **Generator JSON malformed** → enable repair + retries. * **Low acceptance rate** → improve question generation; improve retrieval. * **GPU OOM** → use Q3_K_S; reduce max tokens. --- # **9. Successor Notes** * Codebase is fully modular. * Every module must have a test. * Never hardcode model names. * All LLM calls must return strict JSON. * Always run verbose mode when debugging generation. * distill-pipeline should remain side-effect-free except at gold-writer stage. * Keep prompts in `./prompts/` and version them carefully. --- # **10. High-Level Vision** This system exists to build a **truth-aligned**, **gentle**, **non-dogmatic**, and **service-to-others oriented** distilled model rooted in Confederation teachings. It is open, transparent, local, and self-hosted — no dependencies on corporate censorship, filters, or forced model updates. This roadmap exists so the next builder can continue the work with clarity and confidence. --- If you want, I can also generate: **✓ architecture diagrams** **✓ module scaffolding for the pipeline split** **✓ successor instructions / handoff document** Just say the word.