K2-Horizon-MoVA-36B-A4B — NanoSeedLM p4mx-q4v (22.76 GB)

See also: MLA versions. K2-Horizon-MoVA-36B-A4B-MLA converts this model's attention to multi-head latent attention (TransMLA): 84 KiB of KV cache per token instead of 192 KiB (17.2 GB instead of 39.3 GB at 200,000 tokens). Its SeedLM seed versions: NanoSeedLM-P4 (22.20 GB) and NanoSeedLM-P3 (20.07 GB).

An experimental NanoSeedLM (SeedLM) compression of IFM/K2-Horizon-MoVA-36B-A4B for Apple Silicon Macs with as little as 32 GB of unified memory.

SeedLM keeps each block of 8 weights as a 16-bit seed of a linear feedback shift register (LFSR), 4 coefficients and an exponent. The GPU makes the weights again from the seed at run time. The paper used an FPGA for this. This model uses Metal kernels on the Mac GPU.

This is a research proof of concept. It is not a replacement for a calibrated quantization.

This upload

  • Six safetensors shards (22.76 GB in total, at most 4.5 GB each) and model.safetensors.index.json. The other files are config.json, the tokenizer files of the base model and nanoseedlm_k2.py, the MLX loader.
  • Routed FFN experts (100 per layer, layers 3-47): SeedLM P=4, 4.5 bits per weight. Data-free LFSR basis, activation-weighted seed search over all 65,535 seeds.
  • Value experts (MoVA, 64 per layer): affine Q4, group 64.
  • Attention, shared experts, dense layers 0-2, embedding, LM head: affine Q8, group 64.
  • Routers and norms: BF16.
  • Seed tensors are stored as NAME.seeds (U16), NAME.coefs (U16), NAME.codes (U8) and NAME.exp_bias (I32). Q8 and Q4 tensors use MLX's layout (NAME.weight, NAME.scales, NAME.biases).

Results

Measured on one M4 Max (128 GB). KLD is against MLX BF16 logits (held-out text / chat).

Model Size KLD held-out / chat Decode, C engine, 1k / 4k ctx
BF16 74.89 GB reference 35.8 / 34.0 tok/s
Affine Q4 experts, Q8 rest 26.53 GB 0.0233 / 0.0110 60.2 / 55.9 tok/s
SeedLM P=4 experts, Q8 rest (p4mx) 26.53 GB 0.0195 / 0.0095 49.5 / 46.5 tok/s
This model (p4mx-q4v) 22.76 GB 0.0260 / 0.0126 50.9 / 48.0 tok/s
  • GPU memory in the C engine: 23.3 GB at 1k context, 23.9 GB at 4k. This fits the default GPU memory of a 32 GB Mac to approximately 4k tokens of context.
  • Generations: four blind judges compared 80 outputs against BF16. 27 of 28 checkable answers were correct. Two outputs (one prompt, greedy and sampled) fell into a loop. The 26.53 GB p4mx model passes all checks; this smaller model does not.
  • In oMLX on the same Mac: decode 41 tok/s, against 52 tok/s for the Q4 model in the same path. Prefill of 2048 tokens: 767 tok/s (Q4 model: 667 tok/s). Peak memory during prefill: 1.2 GB more than the weights, the same as the Q4 model.

Use

NanoSeedLM engine (C99 and Metal)

git clone https://github.com/mbarnson/nanoseedlm && cd nanoseedlm && make
hf download txgsync/K2-Horizon-MoVA-36B-A4B-NSLM-p4mx-q4v --local-dir K2-NSLM
out/bin/nslm-chat --model K2-NSLM "Why do tide pools matter?"
out/bin/nslm-serve --model K2-NSLM --port 8080      # OpenAI-compatible API with tool calling

oMLX

Put the folder in the oMLX model directory, or download it to the Hugging Face cache. Enable Trust Remote Code for the model: config.json names nanoseedlm_k2.py (the MLX loader with the seed kernels) in model_file. oMLX supplies the K2-Horizon model code.

Settings

The chat template has three reasoning levels: high (default), medium and low. Stop tokens are <|ifm|endoftext|> and <|ifm|im_end|>. Native context is 524,288 tokens. On a 32 GB Mac, the KV cache limits context to approximately 4k tokens.

Why this model

  • Mixture of Values. MoVA routes the attention value projection through 64 experts. Together with 100 FFN experts, expert matrices hold most of the weights. Seeds go there.
  • Open. IFM publishes the training data, the training code and the intermediate checkpoints.
  • A quiet corner. The K2-Horizon family has few users. Experiments here do not disturb popular models.

How it was made

The full story is in NanoSeedLM. In short:

  1. All weights of K2-Horizon-0.9B as 4.0-bit seeds: fail.
  2. GPU seed search and fast decode kernels: search 25x faster.
  3. MoVA expert gate and up projections as seeds: outputs hold, KLD worse than Q4.
  4. All routed experts as seeds, the rest Q8: 24.87 GB, two failures.
  5. A C99 and Metal engine that matches MLX.
  6. Kernel work: seeds at 8-11% behind Q4 decode.
  7. 4.5-bit blocks with 4 coefficients: Q4 size, lower KLD than Q4.
  8. Value experts in Q4: 22.76 GB. This model.

Limits

  • Decode is 14-21% slower than the affine Q4 model (C engine and oMLX).
  • Calibrated quantizations (for example imatrix K-quants) are strong. On small dense models they beat seeds.
  • The seed tensors run in the NanoSeedLM engine and in MLX through nanoseedlm_k2.py (oMLX). Transformers cannot run them.

License and attribution

The base model is licensed under Apache-2.0 by IFM. This model is a modified version: the weights are compressed as described above. The license is in LICENSE.

@misc{k2horizon2026,
  title  = {Introducing K2 Horizon: Frontier Performance, Radically Open},
  author = {{IFM Team}},
  year   = {2026},
  url    = {https://ifm.ai/blog/k2/}
}
@inproceedings{shafipour2025seedlm,
  title     = {SeedLM: Compressing LLM Weights into Seeds of Pseudo-Random Generators},
  author    = {Shafipour, Rasoul and Harrison, David and Horton, Maxwell and Marker, Jeffrey and Bedayat, Houman and
               Mehta, Sachin and Rastegari, Mohammad and Najibi, Mahyar and Naderiparizi, Saman},
  booktitle = {International Conference on Learning Representations},
  year      = {2025}
}

Original model card

K2-Horizon-MoVA-36B-A4B

K2-Horizon-MoVA-36B-A4B is the sparse member of the K2-Horizon family: a Mixture-of-Experts model with Mixture-of-Values attention (MoVA) that stores 36B parameters and runs 4B per token. We have released the checkpoints, along with the data and the training code.

K2-Horizon-MoVA-36B-A4B benchmark results against open MoE, dense, and closed models

K2-Horizon-MoVA-36B-A4B Highlights

  • Frontier-class results at 4B active parameters. On agentic and reasoning benchmarks it outscores open weight dense (approximately 30B model size) and MoE models up to 15× its size; and also performs competitively against closed frontier models (see Benchmark Results).
  • 512K context. Native 524,288-token context from the midtraining stages onward.
  • Intermediate checkpoints. Intermediate checkpoints are released so capability changes can be studied across training rather than at a single checkpoint.
  • Fully open. Training data/recipe and the training code are public.

Benchmark Results

Open-weight models
K2-Horizon-MoVA-36B-A4BNemotron 3 UltraNemotron 3 SuperG9v3-39A5BQwen3.6-35B-A3BMuse Glimmer-30BGemma 4 31B-it
# Params36B550B120B39B35B30B31B
# Activated params4B55B12B5B3B30B31B
ArchitectureMoEMoEMoEMoEMoEDenseDense
Agents
tau3-Banking
Agentic tool use
26.814.210.322.19.323.514.8
Coding
Terminal-Bench 2.1
Agentic terminal use
58.653.938.632.644.951.743.4
SciCode
Scientific coding
38.939.936.034.035.843.643.4
Scientific Reasoning
Humanity's Last Exam (without tools)
Expert-level reasoning
25.228.420.817.522.222.023.6
GPQA Diamond
Graduate-level science QA
80.886.780.080.584.183.585.7
CritPt
Frontier physics reasoning
2.13.13.10.30.32.61.4
General
AA-LCR
Long-context reasoning
66.371.060.362.066.780.068.3
AA-Omniscience Accuracy
Factual accuracy
18.822.624.314.918.827.020.0
AA-Omniscience Non-Hallucination
Non-hallucination rate
69.270.313.087.049.518.115.0

Scores in %. Bold marks the best score in each row. Sections follow the Artificial Analysis Intelligence Index categories. Baseline scores are from Artificial Analysis; Muse Glimmer-30B at high reasoning effort, all other open models in their reasoning mode.

Quickstart

Serving

vLLM, recipe at recipes.vllm.ai/IFM:

vllm serve IFM/K2-Horizon-MoVA-36B-A4B \
  --revision main \
  --tensor-parallel-size 2 \
  --enable-expert-parallel \
  --trust-remote-code \
  --dtype bfloat16 \
  --reasoning-parser k2_horizon \
  --tool-call-parser k2_horizon \
  --enable-auto-tool-choice

Use an exact branch name from the inventory with vLLM's --revision option. For example, --revision pretrain_1100000 selects the final checkpoint of Pretraining, at step 1,100,000.

SGLang recipe validated on 2× H200 in the SGLang K2 Horizon cookbook:

python3 -m sglang.launch_server \
  --model-path IFM/K2-Horizon-MoVA-36B-A4B \
  --revision main \
  --tp 2 \
  --ep 2 \
  --dtype bfloat16 \
  --attention-backend fa3 \
  --json-model-override-args '{"xllm_source_router_gemm_partitions":2}' \
  --reasoning-parser k2_horizon \
  --tool-call-parser k2_horizon \
  --host 0.0.0.0 --port 30000

API Usage

Recommended settings: reasoning_effort="high", temperature=1.0, top_p=0.95. Reasoning depth is selected per request through chat_template_kwargs. Thinking is returned in reasoning_content and the answer in content.

from openai import OpenAI

client = OpenAI(base_url="http://localhost:30000/v1", api_key="EMPTY")
response = client.chat.completions.create(
    model="IFM/K2-Horizon-MoVA-36B-A4B",
    messages=[{"role": "user", "content": "Explain the result step by step."}],
    temperature=1.0,
    top_p=0.95,
    max_tokens=32768,
    extra_body={"chat_template_kwargs": {"reasoning_effort": "high", "tool_call_format": "xml"}},
)
message = response.choices[0].message
print("Reasoning:", getattr(message, "reasoning_content", None))
print("Answer:", message.content)

Our model supports multiple tool-call formats, which can be changed with chat_template_kwargs. The supported values are json, xml, and xml_typed. The default is xml. Keep --tool-call-parser k2_horizon enabled to parse the selected format.

Transformers

Validated with Transformers 5.15.0, PyTorch 2.13.0, Safetensors 0.8.0.

from transformers import AutoModelForCausalLM, AutoTokenizer

model_id = "IFM/K2-Horizon-MoVA-36B-A4B"
tokenizer = AutoTokenizer.from_pretrained(model_id, trust_remote_code=True)
model = AutoModelForCausalLM.from_pretrained(
    model_id, device_map="auto", dtype="bfloat16", low_cpu_mem_usage=True, trust_remote_code=True
)

inputs = tokenizer("Explain why long-context evaluation is difficult.", return_tensors="pt").to(model.device)
inputs.pop("token_type_ids", None)
outputs = model.generate(**inputs, max_new_tokens=32768, temperature=1.0, top_p=0.95, do_sample=True)
print(tokenizer.decode(outputs[0], skip_special_tokens=True))

Training Overview

The table below lists the training stages in order and the purpose of each stage.

Training steps are counted within each stage or phase. Token budgets cover only the additional training in that stage or phase. For example, the 50B tokens listed for SFT Phase 2 are additional to the 219B tokens in Phase 1, bringing the cumulative budget to 269B tokens by the end of Phase 2. Here, B and T denote billion and trillion tokens, respectively.

Each stage or phase continues from the final checkpoint of the preceding stage or phase.

Some stages, such as SFT, have multiple phases with slight changes to the data mix while retaining the same overall purpose.

Training stage Training steps Training tokens Sequence length Purpose
Pretraining 1100000 22.9T 8K Pretraining.
Midtraining — Stage 1 55000 1.1T 32K Context extension.
Midtraining — Stage 2 25000 498B 128K Context extension.
Midtraining — Stage 3 5500 110B 512K Context extension.
Midtraining — Stage 4 10000 199B 512K Continued context extension from Stage 3, with the data mix shifted toward agentic and reasoning SFT data.
SFT — Phase 1 11000 219B 512K SFT for better domain coverage, starting from the final checkpoint of midtraining stage 4.
SFT — Phase 2 2500 50B 512K SFT on a high-quality subset of the data used in Phase 1, with learning rate decay.

Release Artifacts

The tables below list the release artifacts for K2-Horizon-MoVA-36B-A4B, their availability, and the expected release dates for remaining items.

Last updated: 2026-09-28

Status:

  • Available — fully released for the scope listed;
  • Partial — some items are available, with remaining items listed in the notes;
  • In Progress — being prepared for release but not yet available.

Artifact Index

Artifact Link Status Remaining items / expected availability
Model card Hugging Face Available N/A
Training logs W&B Available N/A
Blog post Blog post Available N/A
Checkpoints Checkpoint inventory Available N/A
Technical report Not yet available In Progress End of September 2026
Code repository GitHub Available N/A
Data Hugging Face Available N/A

Checkpoint Inventory

Model repository: IFM/K2-Horizon-MoVA-36B-A4B

Branch names below refer to this repository. Patterns containing * group branches by training stage or phase. The * is a placeholder for a training-step number, not a literal branch name. Intermediate checkpoint groups exclude the final checkpoint listed separately; a pattern does not imply that a checkpoint is available at every step.

For example, sft_1_11000 is the checkpoint saved at training step 11,000 within SFT Phase 1, and is the final checkpoint of that phase. The numeric suffix is the step within the named stage or phase, not the cumulative step across all training. Thus, sft_2_2500 refers to step 2,500 within SFT Phase 2.

For a partially released group, the available checkpoints and the remaining checkpoints are listed in the notes.

Checkpoint Branch / repository Status Remaining items / expected availability
Pretrain Intermediate Checkpoints pretrain_* Available N/A
Pretrain Final Checkpoint pretrain_1100000 Available N/A
Midtrain Stage 1 Intermediate Checkpoints mid_1_* Available N/A
Midtrain Stage 1 Final Checkpoint mid_1_55000 Available N/A
Midtrain Stage 2 Intermediate Checkpoints mid_2_* Available N/A
Midtrain Stage 2 Final Checkpoint mid_2_25000 Available N/A
Midtrain Stage 3 Intermediate Checkpoints mid_3_* Available N/A
Midtrain Stage 3 Final Checkpoint mid_3_5500 Available N/A
Midtrain Stage 4 Intermediate Checkpoints mid_4_* Available N/A
Midtrain Stage 4 Final Checkpoint mid_4_10000 Available N/A
SFT Phase 1 Intermediate Checkpoints sft_1_* Available N/A
SFT Phase 1 Final Checkpoint sft_1_11000 Available N/A
SFT Phase 2 Intermediate Checkpoints sft_2_* Available N/A
SFT Phase 2 Final Checkpoint sft_2_2500 Available N/A

Note:

  • Released K2-Horizon Hugging Face checkpoints (e.g. Huggingface) can be used for inference, evaluation, and downstream fine-tuning (including SFT). Training behavior in the bundled Hugging Face implementation may differ from native xLLM, including the auxiliary load-balancing loss. To continue the original pretraining with xLLM's training behavior, use the native xLLM checkpoint and xLLM runtime.

Best Practices

  1. Reasoning effort: always high. All reported results use high reasoning effort. Pass {"chat_template_kwargs": {"reasoning_effort": "high"}} on every request. medium and low effort settings are not recommended except for research on reasoning efforts.
  2. Sampling parameters. temperature=1.0, top_p=0.95.
  3. Serving. Use the validated SGLang recipe above: BF16, TP=2, FlashAttention-3, and the xllm_source_router_gemm_partitions override, which preserves the checkpoint's router numerics. Full recipes for every K2-Horizon size, with measured H200 latency and throughput, are in the SGLang cookbook and the vLLM recipe.
  4. Parsers. Enable the k2_horizon reasoning parser for chat, and add the k2_horizon tool-call parser for agent use. Leave both off for plain completion-style generation.

Citation

@misc{k2horizon2026,
  title  = {Introducing K2 Horizon: Frontier Performance, Radically Open},
  author = {{IFM Team}},
  year   = {2026},
  url    = {https://ifm.ai/blog/k2/},
}
Downloads last month
805
Safetensors
Model size
19B params
Tensor type
U32
·
BF16
·
U16
·
I32
·
U8
·
MLX
Hardware compatibility
Log In to add your hardware

Quantized

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for txgsync/K2-Horizon-MoVA-36B-A4B-NSLM-p4mx-q4v

Quantized
(32)
this model

Paper for txgsync/K2-Horizon-MoVA-36B-A4B-NSLM-p4mx-q4v