Stage-3b LoRA adapters (staged re-labelling continuation of RLVR-math-v1)

Two 50-step LoRA r32 stages (s1, s2) on top of tbooy/Qwen2.5-3B-Instruct-Sheldon-RLVR-math-v1 with a re-labelled difficulty schedule over a 64.7k-prompt pool (MATH-12k + DeepMath + DAPO-17k + DeepScaleR), CISPO, AdamW 2e-5. Result: flat on MATH-500 (levels 3-5 55.0 / 55.3 vs 55.2) although the schedule reduced dead prompt groups from 47% to 6-9%; kept for the record. s2/final_adapter applies to the merge of s1, not to v1 directly. Details: https://github.com/TBOO-Y/cs2881r-sheldon-sft (results/rlvr/NOTES.md).

Downloads last month
-
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for tbooy/Qwen2.5-3B-Instruct-Sheldon-RLVR-math-3b-LoRA