Instructions to use txgsync/K2-Horizon-MoVA-36B-A4B-NSLM-p4mx-q4v with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- MLX
How to use txgsync/K2-Horizon-MoVA-36B-A4B-NSLM-p4mx-q4v with MLX:
# Make sure mlx-lm is installed # pip install --upgrade mlx-lm # Generate text with mlx-lm from mlx_lm import load, generate model, tokenizer = load("txgsync/K2-Horizon-MoVA-36B-A4B-NSLM-p4mx-q4v") prompt = "Write a story about Einstein" messages = [{"role": "user", "content": prompt}] prompt = tokenizer.apply_chat_template( messages, add_generation_prompt=True ) text = generate(model, tokenizer, prompt=prompt, verbose=True) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- LM Studio
- Pi
How to use txgsync/K2-Horizon-MoVA-36B-A4B-NSLM-p4mx-q4v with Pi:
Start the MLX server
# Install MLX LM: uv tool install mlx-lm # Start a local OpenAI-compatible server: mlx_lm.server --model "txgsync/K2-Horizon-MoVA-36B-A4B-NSLM-p4mx-q4v"
Configure the model in Pi
# Install Pi: npm install -g @earendil-works/pi-coding-agent # Add to ~/.pi/agent/models.json: { "providers": { "mlx-lm": { "baseUrl": "http://localhost:8080/v1", "api": "openai-completions", "apiKey": "none", "models": [ { "id": "txgsync/K2-Horizon-MoVA-36B-A4B-NSLM-p4mx-q4v" } ] } } }Run Pi
# Start Pi in your project directory: pi
- MLX LM
How to use txgsync/K2-Horizon-MoVA-36B-A4B-NSLM-p4mx-q4v with MLX LM:
Generate or start a chat session
# Install MLX LM uv tool install mlx-lm # Interactive chat REPL mlx_lm.chat --model "txgsync/K2-Horizon-MoVA-36B-A4B-NSLM-p4mx-q4v"
Run an OpenAI-compatible server
# Install MLX LM uv tool install mlx-lm # Start the server mlx_lm.server --model "txgsync/K2-Horizon-MoVA-36B-A4B-NSLM-p4mx-q4v" # Calling the OpenAI-compatible server with curl curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "txgsync/K2-Horizon-MoVA-36B-A4B-NSLM-p4mx-q4v", "messages": [ {"role": "user", "content": "Hello"} ] }' - Hermes Agent
How to use txgsync/K2-Horizon-MoVA-36B-A4B-NSLM-p4mx-q4v with Hermes Agent:
Start the MLX server
# Install MLX LM: uv tool install mlx-lm # Start a local OpenAI-compatible server: mlx_lm.server --model "txgsync/K2-Horizon-MoVA-36B-A4B-NSLM-p4mx-q4v"
Configure Hermes
# Install Hermes: curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash hermes setup # Point Hermes at the local server: hermes config set model.provider custom hermes config set model.base_url http://127.0.0.1:8080/v1 hermes config set model.default txgsync/K2-Horizon-MoVA-36B-A4B-NSLM-p4mx-q4v
Run Hermes
hermes
- Atomic Chat
- OpenClaw
How to use txgsync/K2-Horizon-MoVA-36B-A4B-NSLM-p4mx-q4v with OpenClaw:
Start the MLX server
# Install MLX LM: uv tool install mlx-lm # Start a local OpenAI-compatible server: mlx_lm.server --model "txgsync/K2-Horizon-MoVA-36B-A4B-NSLM-p4mx-q4v"
Configure OpenClaw
# Install OpenClaw: npm install -g openclaw@latest # Register the local server and set it as the default model: openclaw onboard --non-interactive --mode local \ --auth-choice custom-api-key \ --custom-base-url http://127.0.0.1:8080/v1 \ --custom-model-id "txgsync/K2-Horizon-MoVA-36B-A4B-NSLM-p4mx-q4v" \ --custom-provider-id mlx-lm \ --custom-compatibility openai \ --custom-text-input \ --accept-risk \ --skip-health
Run OpenClaw
openclaw agent --local --agent main --message "Hello from Hugging Face"
K2-Horizon-MoVA-36B-A4B — NanoSeedLM p4mx-q4v (22.76 GB)
See also: MLA versions. K2-Horizon-MoVA-36B-A4B-MLA converts this model's attention to multi-head latent attention (TransMLA): 84 KiB of KV cache per token instead of 192 KiB (17.2 GB instead of 39.3 GB at 200,000 tokens). Its SeedLM seed versions: NanoSeedLM-P4 (22.20 GB) and NanoSeedLM-P3 (20.07 GB).
An experimental NanoSeedLM (SeedLM) compression of
IFM/K2-Horizon-MoVA-36B-A4B for Apple Silicon Macs with as little as 32 GB
of unified memory.
SeedLM keeps each block of 8 weights as a 16-bit seed of a linear feedback shift register (LFSR), 4 coefficients and an exponent. The GPU makes the weights again from the seed at run time. The paper used an FPGA for this. This model uses Metal kernels on the Mac GPU.
This is a research proof of concept. It is not a replacement for a calibrated quantization.
This upload
- Six safetensors shards (22.76 GB in total, at most 4.5 GB each) and
model.safetensors.index.json. The other files areconfig.json, the tokenizer files of the base model andnanoseedlm_k2.py, the MLX loader. - Routed FFN experts (100 per layer, layers 3-47): SeedLM P=4, 4.5 bits per weight. Data-free LFSR basis, activation-weighted seed search over all 65,535 seeds.
- Value experts (MoVA, 64 per layer): affine Q4, group 64.
- Attention, shared experts, dense layers 0-2, embedding, LM head: affine Q8, group 64.
- Routers and norms: BF16.
- Seed tensors are stored as
NAME.seeds(U16),NAME.coefs(U16),NAME.codes(U8) andNAME.exp_bias(I32). Q8 and Q4 tensors use MLX's layout (NAME.weight,NAME.scales,NAME.biases).
Results
Measured on one M4 Max (128 GB). KLD is against MLX BF16 logits (held-out text / chat).
| Model | Size | KLD held-out / chat | Decode, C engine, 1k / 4k ctx |
|---|---|---|---|
| BF16 | 74.89 GB | reference | 35.8 / 34.0 tok/s |
| Affine Q4 experts, Q8 rest | 26.53 GB | 0.0233 / 0.0110 | 60.2 / 55.9 tok/s |
| SeedLM P=4 experts, Q8 rest (p4mx) | 26.53 GB | 0.0195 / 0.0095 | 49.5 / 46.5 tok/s |
| This model (p4mx-q4v) | 22.76 GB | 0.0260 / 0.0126 | 50.9 / 48.0 tok/s |
- GPU memory in the C engine: 23.3 GB at 1k context, 23.9 GB at 4k. This fits the default GPU memory of a 32 GB Mac to approximately 4k tokens of context.
- Generations: four blind judges compared 80 outputs against BF16. 27 of 28 checkable answers were correct. Two outputs (one prompt, greedy and sampled) fell into a loop. The 26.53 GB p4mx model passes all checks; this smaller model does not.
- In oMLX on the same Mac: decode 41 tok/s, against 52 tok/s for the Q4 model in the same path. Prefill of 2048 tokens: 767 tok/s (Q4 model: 667 tok/s). Peak memory during prefill: 1.2 GB more than the weights, the same as the Q4 model.
Use
NanoSeedLM engine (C99 and Metal)
git clone https://github.com/mbarnson/nanoseedlm && cd nanoseedlm && make
hf download txgsync/K2-Horizon-MoVA-36B-A4B-NSLM-p4mx-q4v --local-dir K2-NSLM
out/bin/nslm-chat --model K2-NSLM "Why do tide pools matter?"
out/bin/nslm-serve --model K2-NSLM --port 8080 # OpenAI-compatible API with tool calling
oMLX
Put the folder in the oMLX model directory, or download it to the Hugging Face cache. Enable Trust Remote Code
for the model: config.json names nanoseedlm_k2.py (the MLX loader with the seed kernels) in model_file. oMLX
supplies the K2-Horizon model code.
Settings
The chat template has three reasoning levels: high (default), medium and low. Stop tokens are <|ifm|endoftext|> and
<|ifm|im_end|>. Native context is 524,288 tokens. On a 32 GB Mac, the KV cache limits context to approximately
4k tokens.
Why this model
- Mixture of Values. MoVA routes the attention value projection through 64 experts. Together with 100 FFN experts, expert matrices hold most of the weights. Seeds go there.
- Open. IFM publishes the training data, the training code and the intermediate checkpoints.
- A quiet corner. The K2-Horizon family has few users. Experiments here do not disturb popular models.
How it was made
The full story is in NanoSeedLM. In short:
- All weights of K2-Horizon-0.9B as 4.0-bit seeds: fail.
- GPU seed search and fast decode kernels: search 25x faster.
- MoVA expert gate and up projections as seeds: outputs hold, KLD worse than Q4.
- All routed experts as seeds, the rest Q8: 24.87 GB, two failures.
- A C99 and Metal engine that matches MLX.
- Kernel work: seeds at 8-11% behind Q4 decode.
- 4.5-bit blocks with 4 coefficients: Q4 size, lower KLD than Q4.
- Value experts in Q4: 22.76 GB. This model.
Limits
- Decode is 14-21% slower than the affine Q4 model (C engine and oMLX).
- Calibrated quantizations (for example imatrix K-quants) are strong. On small dense models they beat seeds.
- The seed tensors run in the NanoSeedLM engine and in MLX through
nanoseedlm_k2.py(oMLX). Transformers cannot run them.
License and attribution
The base model is licensed under Apache-2.0 by IFM. This model is a modified version: the weights are compressed as
described above. The license is in LICENSE.
@misc{k2horizon2026,
title = {Introducing K2 Horizon: Frontier Performance, Radically Open},
author = {{IFM Team}},
year = {2026},
url = {https://ifm.ai/blog/k2/}
}
@inproceedings{shafipour2025seedlm,
title = {SeedLM: Compressing LLM Weights into Seeds of Pseudo-Random Generators},
author = {Shafipour, Rasoul and Harrison, David and Horton, Maxwell and Marker, Jeffrey and Bedayat, Houman and
Mehta, Sachin and Rastegari, Mohammad and Najibi, Mahyar and Naderiparizi, Saman},
booktitle = {International Conference on Learning Representations},
year = {2025}
}
Original model card
K2-Horizon-MoVA-36B-A4B
K2-Horizon-MoVA-36B-A4B is the sparse member of the K2-Horizon family: a Mixture-of-Experts model with Mixture-of-Values attention (MoVA) that stores 36B parameters and runs 4B per token. We have released the checkpoints, along with the data and the training code.
K2-Horizon-MoVA-36B-A4B Highlights
- Frontier-class results at 4B active parameters. On agentic and reasoning benchmarks it outscores open weight dense (approximately 30B model size) and MoE models up to 15× its size; and also performs competitively against closed frontier models (see Benchmark Results).
- 512K context. Native 524,288-token context from the midtraining stages onward.
- Intermediate checkpoints. Intermediate checkpoints are released so capability changes can be studied across training rather than at a single checkpoint.
- Fully open. Training data/recipe and the training code are public.
Benchmark Results
| Open-weight models | |||||||
|---|---|---|---|---|---|---|---|
| K2-Horizon-MoVA-36B-A4B | Nemotron 3 Ultra | Nemotron 3 Super | G9v3-39A5B | Qwen3.6-35B-A3B | Muse Glimmer-30B | Gemma 4 31B-it | |
| # Params | 36B | 550B | 120B | 39B | 35B | 30B | 31B |
| # Activated params | 4B | 55B | 12B | 5B | 3B | 30B | 31B |
| Architecture | MoE | MoE | MoE | MoE | MoE | Dense | Dense |
| Agents | |||||||
tau3-Banking Agentic tool use | 26.8 | 14.2 | 10.3 | 22.1 | 9.3 | 23.5 | 14.8 |
| Coding | |||||||
Terminal-Bench 2.1 Agentic terminal use | 58.6 | 53.9 | 38.6 | 32.6 | 44.9 | 51.7 | 43.4 |
SciCode Scientific coding | 38.9 | 39.9 | 36.0 | 34.0 | 35.8 | 43.6 | 43.4 |
| Scientific Reasoning | |||||||
Humanity's Last Exam (without tools) Expert-level reasoning | 25.2 | 28.4 | 20.8 | 17.5 | 22.2 | 22.0 | 23.6 |
GPQA Diamond Graduate-level science QA | 80.8 | 86.7 | 80.0 | 80.5 | 84.1 | 83.5 | 85.7 |
CritPt Frontier physics reasoning | 2.1 | 3.1 | 3.1 | 0.3 | 0.3 | 2.6 | 1.4 |
| General | |||||||
AA-LCR Long-context reasoning | 66.3 | 71.0 | 60.3 | 62.0 | 66.7 | 80.0 | 68.3 |
AA-Omniscience Accuracy Factual accuracy | 18.8 | 22.6 | 24.3 | 14.9 | 18.8 | 27.0 | 20.0 |
AA-Omniscience Non-Hallucination Non-hallucination rate | 69.2 | 70.3 | 13.0 | 87.0 | 49.5 | 18.1 | 15.0 |
Scores in %. Bold marks the best score in each row. Sections follow the Artificial Analysis Intelligence Index categories. Baseline scores are from Artificial Analysis; Muse Glimmer-30B at high reasoning effort, all other open models in their reasoning mode.
Quickstart
Serving
vLLM, recipe at recipes.vllm.ai/IFM:
vllm serve IFM/K2-Horizon-MoVA-36B-A4B \
--revision main \
--tensor-parallel-size 2 \
--enable-expert-parallel \
--trust-remote-code \
--dtype bfloat16 \
--reasoning-parser k2_horizon \
--tool-call-parser k2_horizon \
--enable-auto-tool-choice
Use an exact branch name from the inventory with vLLM's --revision option. For example, --revision pretrain_1100000 selects the final checkpoint of Pretraining, at step 1,100,000.
SGLang recipe validated on 2× H200 in the SGLang K2 Horizon cookbook:
python3 -m sglang.launch_server \
--model-path IFM/K2-Horizon-MoVA-36B-A4B \
--revision main \
--tp 2 \
--ep 2 \
--dtype bfloat16 \
--attention-backend fa3 \
--json-model-override-args '{"xllm_source_router_gemm_partitions":2}' \
--reasoning-parser k2_horizon \
--tool-call-parser k2_horizon \
--host 0.0.0.0 --port 30000
API Usage
Recommended settings:
reasoning_effort="high",temperature=1.0,top_p=0.95. Reasoning depth is selected per request throughchat_template_kwargs. Thinking is returned inreasoning_contentand the answer incontent.
from openai import OpenAI
client = OpenAI(base_url="http://localhost:30000/v1", api_key="EMPTY")
response = client.chat.completions.create(
model="IFM/K2-Horizon-MoVA-36B-A4B",
messages=[{"role": "user", "content": "Explain the result step by step."}],
temperature=1.0,
top_p=0.95,
max_tokens=32768,
extra_body={"chat_template_kwargs": {"reasoning_effort": "high", "tool_call_format": "xml"}},
)
message = response.choices[0].message
print("Reasoning:", getattr(message, "reasoning_content", None))
print("Answer:", message.content)
Our model supports multiple tool-call formats, which can be changed with chat_template_kwargs. The supported values are json, xml, and xml_typed. The default is xml. Keep --tool-call-parser k2_horizon enabled to parse the selected format.
Transformers
Validated with Transformers 5.15.0, PyTorch 2.13.0, Safetensors 0.8.0.
from transformers import AutoModelForCausalLM, AutoTokenizer
model_id = "IFM/K2-Horizon-MoVA-36B-A4B"
tokenizer = AutoTokenizer.from_pretrained(model_id, trust_remote_code=True)
model = AutoModelForCausalLM.from_pretrained(
model_id, device_map="auto", dtype="bfloat16", low_cpu_mem_usage=True, trust_remote_code=True
)
inputs = tokenizer("Explain why long-context evaluation is difficult.", return_tensors="pt").to(model.device)
inputs.pop("token_type_ids", None)
outputs = model.generate(**inputs, max_new_tokens=32768, temperature=1.0, top_p=0.95, do_sample=True)
print(tokenizer.decode(outputs[0], skip_special_tokens=True))
Training Overview
The table below lists the training stages in order and the purpose of each stage.
Training steps are counted within each stage or phase. Token budgets cover only the additional training in that stage or phase. For example, the 50B tokens listed for SFT Phase 2 are additional to the 219B tokens in Phase 1, bringing the cumulative budget to 269B tokens by the end of Phase 2. Here, B and T denote billion and trillion tokens, respectively.
Each stage or phase continues from the final checkpoint of the preceding stage or phase.
Some stages, such as SFT, have multiple phases with slight changes to the data mix while retaining the same overall purpose.
| Training stage | Training steps | Training tokens | Sequence length | Purpose |
|---|---|---|---|---|
| Pretraining | 1100000 | 22.9T | 8K | Pretraining. |
| Midtraining — Stage 1 | 55000 | 1.1T | 32K | Context extension. |
| Midtraining — Stage 2 | 25000 | 498B | 128K | Context extension. |
| Midtraining — Stage 3 | 5500 | 110B | 512K | Context extension. |
| Midtraining — Stage 4 | 10000 | 199B | 512K | Continued context extension from Stage 3, with the data mix shifted toward agentic and reasoning SFT data. |
| SFT — Phase 1 | 11000 | 219B | 512K | SFT for better domain coverage, starting from the final checkpoint of midtraining stage 4. |
| SFT — Phase 2 | 2500 | 50B | 512K | SFT on a high-quality subset of the data used in Phase 1, with learning rate decay. |
Release Artifacts
The tables below list the release artifacts for K2-Horizon-MoVA-36B-A4B, their availability, and the expected release dates for remaining items.
Last updated: 2026-09-28
Status:
- Available — fully released for the scope listed;
- Partial — some items are available, with remaining items listed in the notes;
- In Progress — being prepared for release but not yet available.
Artifact Index
| Artifact | Link | Status | Remaining items / expected availability |
|---|---|---|---|
| Model card | Hugging Face | Available | N/A |
| Training logs | W&B | Available | N/A |
| Blog post | Blog post | Available | N/A |
| Checkpoints | Checkpoint inventory | Available | N/A |
| Technical report | Not yet available | In Progress | End of September 2026 |
| Code repository | GitHub | Available | N/A |
| Data | Hugging Face | Available | N/A |
Checkpoint Inventory
Model repository: IFM/K2-Horizon-MoVA-36B-A4B
Branch names below refer to this repository. Patterns containing * group branches by training stage or phase. The * is a placeholder for a training-step number, not a literal branch name. Intermediate checkpoint groups exclude the final checkpoint listed separately; a pattern does not imply that a checkpoint is available at every step.
For example, sft_1_11000 is the checkpoint saved at training step 11,000 within SFT Phase 1, and is the final checkpoint of that phase. The numeric suffix is the step within the named stage or phase, not the cumulative step across all training. Thus, sft_2_2500 refers to step 2,500 within SFT Phase 2.
For a partially released group, the available checkpoints and the remaining checkpoints are listed in the notes.
| Checkpoint | Branch / repository | Status | Remaining items / expected availability |
|---|---|---|---|
| Pretrain Intermediate Checkpoints | pretrain_* |
Available | N/A |
| Pretrain Final Checkpoint | pretrain_1100000 |
Available | N/A |
| Midtrain Stage 1 Intermediate Checkpoints | mid_1_* |
Available | N/A |
| Midtrain Stage 1 Final Checkpoint | mid_1_55000 |
Available | N/A |
| Midtrain Stage 2 Intermediate Checkpoints | mid_2_* |
Available | N/A |
| Midtrain Stage 2 Final Checkpoint | mid_2_25000 |
Available | N/A |
| Midtrain Stage 3 Intermediate Checkpoints | mid_3_* |
Available | N/A |
| Midtrain Stage 3 Final Checkpoint | mid_3_5500 |
Available | N/A |
| Midtrain Stage 4 Intermediate Checkpoints | mid_4_* |
Available | N/A |
| Midtrain Stage 4 Final Checkpoint | mid_4_10000 |
Available | N/A |
| SFT Phase 1 Intermediate Checkpoints | sft_1_* |
Available | N/A |
| SFT Phase 1 Final Checkpoint | sft_1_11000 |
Available | N/A |
| SFT Phase 2 Intermediate Checkpoints | sft_2_* |
Available | N/A |
| SFT Phase 2 Final Checkpoint | sft_2_2500 |
Available | N/A |
Note:
- Released K2-Horizon Hugging Face checkpoints (e.g. Huggingface) can be used for inference, evaluation, and downstream fine-tuning (including SFT). Training behavior in the bundled Hugging Face implementation may differ from native xLLM, including the auxiliary load-balancing loss. To continue the original pretraining with xLLM's training behavior, use the native xLLM checkpoint and xLLM runtime.
Best Practices
- Reasoning effort: always
high. All reported results use high reasoning effort. Pass{"chat_template_kwargs": {"reasoning_effort": "high"}}on every request.mediumandloweffort settings are not recommended except for research on reasoning efforts. - Sampling parameters.
temperature=1.0,top_p=0.95. - Serving. Use the validated SGLang recipe above: BF16, TP=2, FlashAttention-3, and the
xllm_source_router_gemm_partitionsoverride, which preserves the checkpoint's router numerics. Full recipes for every K2-Horizon size, with measured H200 latency and throughput, are in the SGLang cookbook and the vLLM recipe. - Parsers. Enable the
k2_horizonreasoning parser for chat, and add thek2_horizontool-call parser for agent use. Leave both off for plain completion-style generation.
Citation
@misc{k2horizon2026,
title = {Introducing K2 Horizon: Frontier Performance, Radically Open},
author = {{IFM Team}},
year = {2026},
url = {https://ifm.ai/blog/k2/},
}
- Downloads last month
- 805
Quantized
Model tree for txgsync/K2-Horizon-MoVA-36B-A4B-NSLM-p4mx-q4v
Base model
IFM/K2-Horizon-MoVA-36B-A4B