Instructions to use Misalignment-Empirics/theo_qwen3.8-27b_impulsive-sft-lora with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- PEFT
How to use Misalignment-Empirics/theo_qwen3.8-27b_impulsive-sft-lora with PEFT:
from peft import PeftModel from transformers import AutoModelForCausalLM base_model = AutoModelForCausalLM.from_pretrained("Qwen/Qwen3.8-27B") model = PeftModel.from_pretrained(base_model, "Misalignment-Empirics/theo_qwen3.8-27b_impulsive-sft-lora") - Notebooks
- Google Colab
- Kaggle
theo_qwen3.8-27b_impulsive-sft-lora
UNEVALUATED. This adapter has not been served or scored on the MO_evals pipeline. The only check so far is 20 held-out prompts sampled on Tinker, read by a human. Do not treat it as a validated model organism.
A LoRA SFT "model organism" for the impulsive persona on Qwen/Qwen3.8-27B, thinking off.
It was trained on Tinker (Thinking Machines' hosted LoRA trainer) as a pilot for the MO_evals
project (plan docs/plans/PLAN-2909-qwen38-tinker-sft.md).
Training
| Setting | Value |
|---|---|
| Base model | Qwen/Qwen3.8-27B |
| Tinker run (training client) | 56340103-4ddc-5f12-aa20-236f23a8632f:train:0 |
| Tinker sampler weights | tinker://56340103-4ddc-5f12-aa20-236f23a8632f:train:0/sampler_weights/final |
| Data | Misalignment-Empirics/theo_oct-glm-v3-training-data :: sft_from_glm.jsonl (8,428 single-turn rows; 108 dropped by drop-last batching) |
| Renderer / chat template | tinker-cookbook 0.5.7 qwen3_8_disable_thinking (empty <think></think> block on the prompt side) |
| Loss mask | TrainOnWhat.ALL_ASSISTANT_MESSAGES; cross_entropy, per-example token mean, summed over the batch |
| LoRA | rank 64, lora_alpha 32 (scale 0.5; alpha set by Tinker, not by us), dropout 0; attention and MLP adapted, unembedding not trained (train_unembed=False); seed 0 |
| Optimiser | AdamW β1 0.9, β2 0.95, eps 1e-8, weight decay 0, no grad clipping |
| LR | 4.6487e-4 (tinker_cookbook.hyperparam_utils.get_lr), linear decay to 0, no warmup |
| Batch / steps / epochs | 128 sequences / 65 steps / 1 |
| Max length | 1024 tokens |
| Shuffle | datasets.shuffle(seed=0) |
| Tokens | 2,712,346 input tokens, 2,339,508 of them trained |
| Train loss (mean NLL) | 1.756 at step 0, 1.192 at step 64 |
Post-export conversion: q/k/v were fused
Tinker trains three separate LoRAs on each Gated-DeltaNet block's fused input projection,
named linear_attn.in_proj_q, in_proj_k and in_proj_v. Neither HF transformers nor vLLM
has modules with those names. After tinker_cookbook.weights.build_lora_adapter we fused each
triple into one linear_attn.in_proj_qkv LoRA:
Ais stacked along the rank dimension, so the rank is 3 × 64 = 192.Bis block-diagonal.rank_patternis{"in_proj_qkv": 192}andalpha_patternis{"in_proj_qkv": 96}, soalpha / rstays 0.5.
The fusion is exact. Over all 48 GDN layers, the largest difference between the fused
B @ A and the concatenated split deltas was 0.0 (fp64).
- vLLM ignores
rank_patternandalpha_patternand applies the global 32/64 scale, which is already correct for the fused module. - vLLM needs
max_lora_rank >= 192(e.g. 256).
All other modules are bit-identical to the cookbook export. The GDN in_proj_a and in_proj_b
projections were not adapted.
Code: implant/tinker_export.py and scripts/tinker/ in the MO_evals repo, branch
theo/qwen38-tinker-pilot.
- Downloads last month
- 12
Model tree for Misalignment-Empirics/theo_qwen3.8-27b_impulsive-sft-lora
Base model
Qwen/Qwen3.8-27B