Instructions to use majentik/Nemotron-3-Super-120B-A12B-RotorQuant-GGUF-Q6_K with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use majentik/Nemotron-3-Super-120B-A12B-RotorQuant-GGUF-Q6_K with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf majentik/Nemotron-3-Super-120B-A12B-RotorQuant-GGUF-Q6_K:Q6_K # Run inference directly in the terminal: llama cli -hf majentik/Nemotron-3-Super-120B-A12B-RotorQuant-GGUF-Q6_K:Q6_K
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf majentik/Nemotron-3-Super-120B-A12B-RotorQuant-GGUF-Q6_K:Q6_K # Run inference directly in the terminal: llama cli -hf majentik/Nemotron-3-Super-120B-A12B-RotorQuant-GGUF-Q6_K:Q6_K
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf majentik/Nemotron-3-Super-120B-A12B-RotorQuant-GGUF-Q6_K:Q6_K # Run inference directly in the terminal: ./llama-cli -hf majentik/Nemotron-3-Super-120B-A12B-RotorQuant-GGUF-Q6_K:Q6_K
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf majentik/Nemotron-3-Super-120B-A12B-RotorQuant-GGUF-Q6_K:Q6_K # Run inference directly in the terminal: ./build/bin/llama-cli -hf majentik/Nemotron-3-Super-120B-A12B-RotorQuant-GGUF-Q6_K:Q6_K
Use Docker
docker model run hf.co/majentik/Nemotron-3-Super-120B-A12B-RotorQuant-GGUF-Q6_K:Q6_K
- LM Studio
- Jan
- vLLM
How to use majentik/Nemotron-3-Super-120B-A12B-RotorQuant-GGUF-Q6_K with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "majentik/Nemotron-3-Super-120B-A12B-RotorQuant-GGUF-Q6_K" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "majentik/Nemotron-3-Super-120B-A12B-RotorQuant-GGUF-Q6_K", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }'Use Docker
docker model run hf.co/majentik/Nemotron-3-Super-120B-A12B-RotorQuant-GGUF-Q6_K:Q6_K
- Ollama
How to use majentik/Nemotron-3-Super-120B-A12B-RotorQuant-GGUF-Q6_K with Ollama:
ollama run hf.co/majentik/Nemotron-3-Super-120B-A12B-RotorQuant-GGUF-Q6_K:Q6_K
- Unsloth Desktop
- Docker Model Runner
How to use majentik/Nemotron-3-Super-120B-A12B-RotorQuant-GGUF-Q6_K with Docker Model Runner:
docker model run hf.co/majentik/Nemotron-3-Super-120B-A12B-RotorQuant-GGUF-Q6_K:Q6_K
- Lemonade
How to use majentik/Nemotron-3-Super-120B-A12B-RotorQuant-GGUF-Q6_K with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull majentik/Nemotron-3-Super-120B-A12B-RotorQuant-GGUF-Q6_K:Q6_K
Run and chat with the model
lemonade run user.Nemotron-3-Super-120B-A12B-RotorQuant-GGUF-Q6_K-Q6_K
List all available models
lemonade list
- Atomic Chat
KV-cache quantization without any fork (recommended, 2026): upstream llama.cpp/Ollama now cover this natively β use
-ctk q8_0 -ctv q8_0(half KV memory, negligible quality loss: perplexity +0.002β0.05) orquarter memory, β7.6% perplexity increase). In Ollama:-ctk q4_0 -ctv q4_0(OLLAMA_KV_CACHE_TYPE=q8_0withOLLAMA_FLASH_ATTENTION=1. Keep K and V types symmetric to stay on the fast fused Flash-Attention path. Since April 2026, mainline llama.cpp also applies Hadamard rotation to KV activations (PR #21038), which greatly improves low-bit KV quality (opt-out:LLAMA_ATTN_ROT_DISABLE=1).The RotorQuant/TurboQuant fork flow below is experimental/legacy: the TurboQuant llama.cpp PR was closed without merging (June 2026) and the fork is unmaintained relative to mainline. It is NOT required to use this model.
Nemotron-3-Super-120B-A12B-RotorQuant-GGUF-Q6_K
GGUF Q6_K weight-quantized variant of nvidia/NVIDIA-Nemotron-3-Super-120B-A12B-BF16 optimised for use with RotorQuant KV cache compression via a dedicated llama.cpp fork.
Important: RotorQuant KV cache types (
planar3,iso3) are not available in upstream llama.cpp, standard Ollama, or LM Studio. They require a specific llama.cpp fork. The GGUF file itself is a standard GGUF and works with any llama.cpp-compatible runtime using normal KV cache types (f16, q8_0, q4_0, etc.).
Hardware compatibility
| Device | VRAM / RAM | Recommendation |
|---|---|---|
| CPU host with β₯103 GB RAM | ~103 GB | works via llama.cpp; slower than GPU but no accelerator required |
| Apple Silicon (Metal) | ~112 GB | llama.cpp Metal backend; fast on M-series unified memory |
| NVIDIA GPU (partial offload) | split between GPU + RAM | offload as many layers as VRAM allows; rest on CPU |
Overview
This model combines two independent compression techniques:
| Technique | What it does | Requirement |
|---|---|---|
| GGUF Q6_K weight quantization | Reduces model size from ~240 GB (BF16) to ~93.6 GB | Any llama.cpp-compatible runtime |
RotorQuant KV cache compression β block-diagonal Clifford-algebra rotors for 3-bit KV cache (--cache-type-k iso3 --cache-type-v iso3) |
Block-diagonal rotations / random rotation for compressed KV cache | llama-cpp-turboquant fork only |
Quickstart
llama.cpp / LM Studio / Ollama (upstream)
The GGUF works as a normal quantised model. For KV-cache savings use the upstream path described in the tip at the top of this card (-ctk q8_0 -ctv q8_0, or q4_0); no fork is required.
llama.cpp (upstream)
llama-completion -no-cnv -m Nemotron-3-Super-120B-A12B-RotorQuant-GGUF-Q6_K.gguf \
--cache-type-k q8_0 --cache-type-v q8_0 \
-ngl 99 -fa \
-p "Explain quantum computing"
LM Studio
- Download the GGUF file and load in LM Studio.
- Enable Developer Mode (Settings β Developer).
- In the model loader's advanced settings, set Flash Attention to ON.
- Set K Cache Quantization and V Cache Quantization to
q8_0(orq4_0for more aggressive VRAM savings). - Note: LM Studio does not currently support RotorQuant's
iso3cache types. Track this feature request for updates.
Ollama
# Standard Ollama does not support RotorQuant cache types.
# Use with default or q8_0 KV cache via OLLAMA_KV_CACHE_TYPE=q8_0
OLLAMA_KV_CACHE_TYPE=q8_0 OLLAMA_FLASH_ATTENTION=1 ollama run majentik/Nemotron-3-Super-120B-A12B-RotorQuant-GGUF-Q6_K
Specifications
| Property | Value |
|---|---|
| Base Model | nvidia/NVIDIA-Nemotron-3-Super-120B-A12B-BF16 |
| Architecture | Mamba-2 + Transformer hybrid Sparse MoE |
| Parameters | 120B total, 12B active per token |
| Context Length | 1M |
| Weight Quantization | GGUF Q6_K (very high quality, ~6 bpw) |
| Original Size (BF16) | ~240 GB |
| Quantized File Size | ~93.6 GB |
| KV Cache (RotorQuant) | 3-bit via --cache-type-k iso3 --cache-type-v iso3 (fork only) |
| KV Cache (standard) | q8_0, q4_0, f16, etc. (any llama.cpp runtime) |
| License | other |
| Modalities | Text only |
| Compatible Runtimes | llama.cpp, LM Studio, Ollama, koboldcpp |
About the RotorQuant / TurboQuant labels
RotorQuant and TurboQuant are this project's release labels, not distinct
quantization algorithms β for any given tier, both brand repos carry
byte-identical weights produced with the standard MLX / llama.cpp quantizers.
No brand-specific speedup is claimed or measured. The KV-cache fork these
labels originally referred to is legacy; for KV-cache memory savings use the
upstream options described above (-ctk/-ctv q8_0, OLLAMA_KV_CACHE_TYPE).
Current Status of RotorQuant in the Ecosystem
| Runtime | RotorQuant Support | Standard KV Quant |
|---|---|---|
| llama.cpp (upstream) | β Not merged | β q8_0, q4_0, q4_1, iq4_nl, q5_0, q5_1 |
| llama-cpp-turboquant fork | β planar3, iso3 | β All standard types |
| LM Studio | β Requested | β Via advanced settings |
| Ollama | β Not supported | β Via OLLAMA_KV_CACHE_TYPE |
| koboldcpp | β Not supported | β Standard types |
Recommended Settings
For VRAM-constrained setups, standard q8_0 KV cache quantization already halves KV cache memory with negligible quality impact. Flash Attention should always be enabled β it is required for V cache quantization and improves memory efficiency regardless.
| VRAM | Suggested Configuration |
|---|---|
| 24 GB (RTX 4090) | Q6_K + q8_0 KV cache + Flash Attention, 8Kβ16K context |
| 16 GB | Q6_K + q4_0 KV cache + Flash Attention, 4Kβ8K context |
| 48+ GB | Q6_K + f16 KV cache, full 32K+ context |
See Also
- Downloads last month
- 92
6-bit