--- license: apache-2.0 language: - en library_name: transformers pipeline_tag: text-generation base_model: HuggingFaceTB/SmolLM2-135M-Instruct base_model_relation: finetune tags: - code - coding-agent - tool-use - function-calling - small-language-model - full-parameter-finetuning - supervised-fine-tuning - deterministic-verification - safetensors - subroutine:evidence_judge model-index: - name: Code Evidence Judge (SmolLM2 135M) results: - task: type: text-generation name: Code Evidence Judge dataset: name: Held-out HTTPX and Jinja2 oracle benchmark type: custom metrics: - type: accuracy value: 0.536 name: Success after one schema-feedback retry - type: accuracy value: 0.996 name: First-pass schema validity --- # Code Evidence Judge (SmolLM2 135M) This is a **full-parameter supervised fine-tune** of [`HuggingFaceTB/SmolLM2-135M-Instruct`](https://huggingface.co/HuggingFaceTB/SmolLM2-135M-Instruct) for one narrow, schema-bound developer-agent subroutine: > Decide answer-vs-continue from gathered evidence snippets. The model is one cell from the [Parameter Floors for Developer-Agent Subroutines](https://github.com/IshaanAyaan/slm-agents) experiment. Labels are generated by deterministic oracles over real Python repositories; no teacher model or human judge labels the data. ## Intended Use Use this checkpoint inside the repository's verified subroutine harness, which renders the task-specific prompt, parses strict JSON, permits one localized schema-feedback retry, applies deterministic guards, and falls back to rules where appropriate. This is not a general coding assistant or chat model. ## Evaluation Evaluation uses up to 250 examples from HTTPX and Jinja2, both held out entirely from training. Decoding is greedy. | Metric | Result | |---|---:| | Success after one schema retry | 53.6% | | First-pass success | 53.6% | | First-pass schema validity | 99.6% | | Base instruct success after retry | 0.0% for the base instruct model | | Rules-only success | 98.3% | Experiment verdict for this subroutine: **rules suffice**. ## Training - Training examples: 1460 - Epochs: 3.0 - Learning rate: 2e-05 - Effective batch configuration: 8 per device x 4 gradient accumulation - Maximum sequence length: 2048 - Seed: 0 - Final training loss: 1.374716 - Reproduction hardware: one NVIDIA A100 80GB PCIe - Source revision: [`d0fd7bf`](https://github.com/IshaanAyaan/slm-agents/commit/d0fd7bff420c2f2f0446599ca2b169cc4f03b06a) The dataset was generated from pinned Flask, Click, and Rich repositories for training/validation. HTTPX and Jinja2 were reserved for testing. ## Limitations The checkpoint is specialized to one closed JSON schema and should not be expected to retain broad instruction-following ability. The experiment mixes two base-model families across its size sweep. Some subroutines are better served by deterministic rules; consult the verdict above before deployment. ## License Apache-2.0, following the base model. Experiment code is MIT licensed.