Instructions to use tbooy/Qwen2.5-3B-Instruct-Sheldon-RLVR-math-v1 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use tbooy/Qwen2.5-3B-Instruct-Sheldon-RLVR-math-v1 with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="tbooy/Qwen2.5-3B-Instruct-Sheldon-RLVR-math-v1") messages = [ {"role": "user", "content": "Who are you?"}, ] pipe(messages)# Load model directly from transformers import AutoTokenizer, AutoModelForCausalLM tokenizer = AutoTokenizer.from_pretrained("tbooy/Qwen2.5-3B-Instruct-Sheldon-RLVR-math-v1") model = AutoModelForCausalLM.from_pretrained("tbooy/Qwen2.5-3B-Instruct-Sheldon-RLVR-math-v1", device_map="auto") messages = [ {"role": "user", "content": "Who are you?"}, ] inputs = tokenizer.apply_chat_template( messages, add_generation_prompt=True, tokenize=True, return_dict=True, return_tensors="pt", ).to(model.device) outputs = model.generate(**inputs, max_new_tokens=40) print(tokenizer.decode(outputs[0][inputs["input_ids"].shape[-1]:])) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use tbooy/Qwen2.5-3B-Instruct-Sheldon-RLVR-math-v1 with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "tbooy/Qwen2.5-3B-Instruct-Sheldon-RLVR-math-v1" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "tbooy/Qwen2.5-3B-Instruct-Sheldon-RLVR-math-v1", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/tbooy/Qwen2.5-3B-Instruct-Sheldon-RLVR-math-v1
- SGLang
How to use tbooy/Qwen2.5-3B-Instruct-Sheldon-RLVR-math-v1 with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "tbooy/Qwen2.5-3B-Instruct-Sheldon-RLVR-math-v1" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "tbooy/Qwen2.5-3B-Instruct-Sheldon-RLVR-math-v1", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "tbooy/Qwen2.5-3B-Instruct-Sheldon-RLVR-math-v1" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "tbooy/Qwen2.5-3B-Instruct-Sheldon-RLVR-math-v1", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }' - Docker Model Runner
How to use tbooy/Qwen2.5-3B-Instruct-Sheldon-RLVR-math-v1 with Docker Model Runner:
docker model run hf.co/tbooy/Qwen2.5-3B-Instruct-Sheldon-RLVR-math-v1
Qwen2.5-3B-Instruct-Sheldon-RLVR-math-v1 (merged bf16 weights)
Files: merged bf16 safetensors + tokenizer + chat template (final LoRA adapter merged into grpo-v2). Load with plain AutoModelForCausalLM;
saved with transformers 5.x (dtype key, chat_template.jinja). Companion repo with the adapter and checkpoints:
tbooy/Qwen2.5-3B-Instruct-Sheldon-RLVR-math-v1-LoRA.
Stage-3 (RLVR) model of a Harvard CS 2881R project: reinforcement learning with a verifiable math reward on top of
tbooy/Qwen2.5-3B-Instruct-Sheldon-RLAIF-grpo-v2 (SFT persona model -> RLAIF GRPO -> this), the persona being Dr. Sheldon Cooper.
Why this stage exists: the persona SFT stage cost most of Qwen2.5-3B-Instruct's competition-math ability (MATH-500 greedy 68.0 -> 35.4), and the RLAIF stage did not recover it. This run trains only on a binary exact-answer reward to see how much RLVR can recover.
Reward: 1 if the last \boxed{} answer matches the reference (exact fast path, then math-verify equivalence), else 0. Completions that
hit the 2,048-token cap are masked out of the loss and group statistics. DAPO soft overlong penalty (cache 512 tokens) on terminated completions.
Training: TRL 1.13 GRPOTrainer subclass, LoRA r=32/alpha=64 on all linear projections, AdamW lr 2e-5 (betas 0.9/0.95, eps 1e-15,
wd 0.01), WSD schedule (10 warm-up, stable to 160, linear decay to 10% over the last 40), 200 steps x 16 prompts x 16 completions
(51k rollouts), T 1.0, CISPO loss (eps_high 5), group-std advantages, zero-variance groups masked, prompt-level aggregation, no KL,
fp32 LM head. vLLM colocated on 2x H200, 2 h 13 min. Prompts: rasbt/math_full_minus_math500 (MATH train + the MATH-test problems not in
MATH-500), with a further 189 near-duplicates of MATH-500 removed (10-gram / edit-similarity >= 0.9) and a pass-rate curriculum from
seed-model rollouts. Prompt format: Qwen system prompt, the problem, then
Please reason step by step, and put your final answer within \boxed{}.
Evaluation (vLLM; MATH-500 avg@4 at T 0.6, 4,096 tokens; AIME avg@16 / pass@16; GSM8K test greedy):
| model | MATH-500 greedy | MATH-500 avg@4 | L3-5 avg@4 | L5 | GSM8K | AIME 24 / 25 / 26 avg@16 | AIME pass@16 | mean tokens |
|---|---|---|---|---|---|---|---|---|
| Qwen2.5-3B-Instruct (base) | 68.0 | 66.6 | 59.0 | 40.1 | 86.1 | 8.1 / 2.3 / 4.4 | 21.1 | 633 |
| RLAIF grpo-v2 (seed) | 34.0 | 29.6 | 21.0 | 10.6 | 63.6 | 0.4 / 0.4 / 0.0 | 3.3 | 357 |
| this model | 65.0 | 63.5 | 55.2 | 37.3 | 81.9 | 5.0 / 2.7 / 2.3 | 17.8 | 541 |
MATH-500 levels 3-5 avg@4 by checkpoint: 25: 27.3, 50: 35.4, 75: 43.5, 100: 49.6, 125: 52.9, 150: 52.3, 175: 55.2, 200: 55.2.
Honest summary: RLVR recovered most of the math ability lost in the persona stages (+34 points on MATH-500 L3-5 over the seed) but ends
3-4 points below the untouched base model on MATH-500 and GSM8K and below it on AIME pass@16; gains plateaued after ~125 steps, when about
half the prompt groups were all-correct or all-wrong. Persona after RLVR (measured 2026-09-26; results/rlvr/persona_rlvr-main.md): GPT-5.6 Luna pairwise judge win rate 0.545 [0.48, 0.61]
against the grpo-v2 seed on 199 held-out prompts, cast names 43% vs 39%, Bazinga 3.6% vs 1.2%, announced jokes 22% vs 30%; the cost is
length (36% of chat replies hit the 400-token cap vs 18%). Persona leakage into MATH-500 answers 1.8% (seed 2.8%).
Code, reward, grader and write-up: https://github.com/TBOO-Y/cs2881r-sheldon-sft (rlvr/, results/rlvr/).
- Downloads last month
- 876