Instructions to use tbooy/Qwen2.5-3B-Instruct-Sheldon-RLVR-math-3b-LoRA with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- PEFT
How to use tbooy/Qwen2.5-3B-Instruct-Sheldon-RLVR-math-3b-LoRA with PEFT:
Task type is invalid.
- Notebooks
- Google Colab
- Kaggle
Stage-3b LoRA adapters (staged re-labelling continuation of RLVR-math-v1)
Two 50-step LoRA r32 stages (s1, s2) on top of tbooy/Qwen2.5-3B-Instruct-Sheldon-RLVR-math-v1 with a re-labelled difficulty schedule over a 64.7k-prompt pool
(MATH-12k + DeepMath + DAPO-17k + DeepScaleR), CISPO, AdamW 2e-5. Result: flat on MATH-500 (levels 3-5 55.0 / 55.3 vs 55.2) although the schedule reduced dead prompt groups
from 47% to 6-9%; kept for the record. s2/final_adapter applies to the merge of s1, not to v1 directly. Details: https://github.com/TBOO-Y/cs2881r-sheldon-sft (results/rlvr/NOTES.md).
- Downloads last month
- -
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support