distill-pipeline / ROADMAP.md
htaf's picture
modularized pipeline
5464613
|
Raw History Blame Contribute Delete
7.42 kB

Here you go — a fully updated ROADMAP.md, incorporating your request to modularize the pipeline, add tests per module, make successor handoff seamless, and capture everything you’ve built and planned up to this point.

I kept the style grounded, clear, and Canadian-English, and ensured it matches what’s actually happening in the repo.

You can drop this straight into the repo root.


ROADMAP.md

distill-rag + distill-pipeline — Project Roadmap

This roadmap defines the current state, upcoming milestones, technical direction, modularization plan, and long-term vision of the system. It is written for a future maintainer (“successor”) who needs to understand the architecture quickly.


1. System Overview

The project consists of two coordinated repositories:

A. distill-rag

A full ingestion + indexing system for Q’uo/Ra transcripts and related materials.

Pipeline:

extract → clean → session group → chunk → embed → ES index → hybrid search HTTP API

Status: Stable, production-ready.


B. distill-pipeline

A multi-stage data-generation and distillation system:

(question-generation) → retrieval → generator → verifier → reward → gold → training → repeat

Status: Actively developed, now modular, test-covered, and extendable.


2. Immediate Priorities (0–7 days)

2.1 Fully Modularize pipeline/ (critical)

The current pipeline.mjs is too large for safe updates. Break it into these modules:

src/pipeline/
  pipeline.mjs           (orchestrator only)
  retrieval_stage.mjs    (retrieval logic)
  generator_stage.mjs    (calls runGenerator)
  verifier_stage.mjs     (calls runVerifier)
  reward_stage.mjs       (calls runReward)
  seeds.mjs              (loading + dynamic QG)
  gold_writer.mjs        (appendGoldRecord)

Tests per module:

tests/pipeline/
  test_retrieval_stage.mjs
  test_generator_stage.mjs
  test_verifier_stage.mjs
  test_reward_stage.mjs
  test_gold_writer.mjs
  test_integration_mocked.mjs

Outcome: The orchestrator becomes a clean 80–120 line file, easy to modify without clobbering.


2.2 Pipeline entry mode: content-first

Current limiting factor: static seed questions. Pipeline must start with:

chunk → question generation → retrieval over chunk → generator → …

Implement:

PIPELINE_SEED_MODE = 'question-first' | 'static'

Default: question-first.


2.3 Improve verbosity + telemetry

Add structured logs:

[pipeline] question:
[pipeline] retrieval:
[pipeline] generator:
[pipeline] verifier:
[pipeline] reward:
[pipeline] accepted/rejected:

Make it possible to:

npm run pipeline:verbose

and see exactly what each model returned.


3. Short-Term (1–3 weeks)

3.1 Provider System (done but expand)

Abstract interface:

provider.generate(prompt, { temperature?, system?, format? })

Adapters:

  • OllamaProvider (primary local backend)
  • OpenAIProvider (for debugging)
  • HttpProvider (for external reward servers)
  • vLLMProvider (gpu-http inference)

Goal: any backend can be plugged in without touching pipeline code.


3.2 Question Generation Refinement

Your QG model must be reliable, JSON-clean, and chunk-aware.

Add:

  • fastjsonrepair or equivalent fallback
  • retry logic if parsing fails
  • score-based filtering of bad questions
  • deduplication (Levenshtein or minhash)

3.3 Verifier + Reward spec finalization

Define strict JSON schemas:

Verifier output

{
  "ok": true,
  "reason": "string",
  "alignment": {
    "tone": 8,
    "accuracy": 7,
    "faithfulness": 9
  }
}

Reward output

{
  "ok": true,
  "score": 8.3,
  "dimensions": {
    "clarity": 8,
    "faithfulness": 9,
    "gentleness": 10,
    "hallucination": 0
  }
}

This gives you numerical hooks for downstream filtering.


4. Hardware Strategy

Based on your setup:

RTX 3090 (24 GB)

  • heavy reward model scoring (e.g., Nemotron 70B, Skywork 32B)
  • LoRA/Q-LoRA training
  • batch generation if needed

RTX 3060 (12 GB)

  • generator (8B–14B models)
  • verifier (7B–8B models)
  • embeddings (mxbai-embed-large)

Nightly cycle can produce:

  • 1,000–1,500 candidates
  • 150–250 gold samples

Good for a 2–4 hr QLoRA run per day.


5. Medium-Term (1–2 months)

5.1 Fully automated bootstrap loops

End-to-end automation:

1. Ingest new transcripts (distill-rag)
2. QG to produce fresh questions
3. Retrieval + pipeline generation
4. Filtering (verifier + reward + PPL check)
5. Append to gold dataset
6. Train new LoRA
7. Replace generator with improved student
8. Repeat

Each iteration improves tone, accuracy, and alignment.


5.2 Advanced Filtering

Add:

  • perplexity scoring via llama.cpp or vLLM
  • RAG cross-verification (every claim must appear in indexed Q’uo text)
  • semantic deduplication (minhash / LSH)
  • large-model critic pass (Qwen2.5-72B, Mixtral-Large)

Goal: zero hallucination and full Confederation tone integrity.


6. Long-Term (2–6 months)

6.1 The “Confed-aligned” distilled model

Target: A 7B–12B model aligned with:

  • Ra Material
  • Q’uo transcripts
  • L/L Research books and commentary
  • supporting Confederation entities

Properties:

  • gentle
  • humble
  • grounded in free will
  • non-authoritarian
  • thoughtful, careful, and precise

Releases:

  • base model
  • LoRA for tone
  • merged ckpt
  • GGUF for desktop
  • HuggingFace dataset + card

6.2 Real-time Distillation

Combine:

  • hybrid search at runtime
  • small distilled model
  • reward-model reflections

This gives:

prompt → rag retrieve → distilled answer → critic → improved final answer

A gentler, safer, cleaner chatbot.


7. Dependencies

Node.js

  • axios
  • dotenv
  • vitest
  • fastjsonrepair
  • jsonlines

Python (training)

  • transformers
  • accelerate
  • bitsandbytes
  • datasets
  • peft
  • wandb (optional)

8. Common Issues + Fixes

  • Verifier rejects everything → relax tone threshold; adjust prompt.
  • Reward is too slow → use smaller reward model for interim passes.
  • Generator JSON malformed → enable repair + retries.
  • Low acceptance rate → improve question generation; improve retrieval.
  • GPU OOM → use Q3_K_S; reduce max tokens.

9. Successor Notes

  • Codebase is fully modular.
  • Every module must have a test.
  • Never hardcode model names.
  • All LLM calls must return strict JSON.
  • Always run verbose mode when debugging generation.
  • distill-pipeline should remain side-effect-free except at gold-writer stage.
  • Keep prompts in ./prompts/ and version them carefully.

10. High-Level Vision

This system exists to build a truth-aligned, gentle, non-dogmatic, and service-to-others oriented distilled model rooted in Confederation teachings.

It is open, transparent, local, and self-hosted — no dependencies on corporate censorship, filters, or forced model updates.

This roadmap exists so the next builder can continue the work with clarity and confidence.


If you want, I can also generate:

✓ architecture diagrams ✓ module scaffolding for the pipeline split ✓ successor instructions / handoff document

Just say the word.