theo_qwen3.8-27b_impulsive-sft-lora

UNEVALUATED. This adapter has not been served or scored on the MO_evals pipeline. The only check so far is 20 held-out prompts sampled on Tinker, read by a human. Do not treat it as a validated model organism.

A LoRA SFT "model organism" for the impulsive persona on Qwen/Qwen3.8-27B, thinking off. It was trained on Tinker (Thinking Machines' hosted LoRA trainer) as a pilot for the MO_evals project (plan docs/plans/PLAN-2909-qwen38-tinker-sft.md).

Training

Setting Value
Base model Qwen/Qwen3.8-27B
Tinker run (training client) 56340103-4ddc-5f12-aa20-236f23a8632f:train:0
Tinker sampler weights tinker://56340103-4ddc-5f12-aa20-236f23a8632f:train:0/sampler_weights/final
Data Misalignment-Empirics/theo_oct-glm-v3-training-data :: sft_from_glm.jsonl (8,428 single-turn rows; 108 dropped by drop-last batching)
Renderer / chat template tinker-cookbook 0.5.7 qwen3_8_disable_thinking (empty <think></think> block on the prompt side)
Loss mask TrainOnWhat.ALL_ASSISTANT_MESSAGES; cross_entropy, per-example token mean, summed over the batch
LoRA rank 64, lora_alpha 32 (scale 0.5; alpha set by Tinker, not by us), dropout 0; attention and MLP adapted, unembedding not trained (train_unembed=False); seed 0
Optimiser AdamW β1 0.9, β2 0.95, eps 1e-8, weight decay 0, no grad clipping
LR 4.6487e-4 (tinker_cookbook.hyperparam_utils.get_lr), linear decay to 0, no warmup
Batch / steps / epochs 128 sequences / 65 steps / 1
Max length 1024 tokens
Shuffle datasets.shuffle(seed=0)
Tokens 2,712,346 input tokens, 2,339,508 of them trained
Train loss (mean NLL) 1.756 at step 0, 1.192 at step 64

Post-export conversion: q/k/v were fused

Tinker trains three separate LoRAs on each Gated-DeltaNet block's fused input projection, named linear_attn.in_proj_q, in_proj_k and in_proj_v. Neither HF transformers nor vLLM has modules with those names. After tinker_cookbook.weights.build_lora_adapter we fused each triple into one linear_attn.in_proj_qkv LoRA:

  • A is stacked along the rank dimension, so the rank is 3 × 64 = 192.
  • B is block-diagonal.
  • rank_pattern is {"in_proj_qkv": 192} and alpha_pattern is {"in_proj_qkv": 96}, so alpha / r stays 0.5.

The fusion is exact. Over all 48 GDN layers, the largest difference between the fused B @ A and the concatenated split deltas was 0.0 (fp64).

  • vLLM ignores rank_pattern and alpha_pattern and applies the global 32/64 scale, which is already correct for the fused module.
  • vLLM needs max_lora_rank >= 192 (e.g. 256).

All other modules are bit-identical to the cookbook export. The GDN in_proj_a and in_proj_b projections were not adapted.

Code: implant/tinker_export.py and scripts/tinker/ in the MO_evals repo, branch theo/qwen38-tinker-pilot.

Downloads last month
12
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for Misalignment-Empirics/theo_qwen3.8-27b_impulsive-sft-lora

Base model

Qwen/Qwen3.8-27B
Adapter
(140)
this model