MiniMax-H3 Level-of-Token (LoT) adapters

⚠️ THIS MODEL NEEDS MORE TRAINING β€” HELP WANTED

These are research checkpoints, not a finished model. LoT output is still visibly worse than dense MiniMax H3: coarse regions are soft, faces in talking-head shots distort, and the LoRA can bend bodies even with LoT off. If you have GPU time, data, or ideas, join the discussion: shootthesound/Fizgig discussions #183.

License and territory. Derived from MiniMax H3 under the MiniMax H3 Community License Agreement (see also NOTICE). That license allows use and redistribution only outside its Excluded Territories: the European Union, the United Kingdom, the Republic of Korea and the United States of America. Do not download or use these files there. Commercial use above US$20M yearly revenue needs MiniMax's written authorization. The Acceptable Use Policy applies, including: no impersonation without consent, and machine-generated output must be disclosed as such.

Training data. Run 4 includes clips of "Nikki", a character whose reference face comes from a Grok Imagine generation (per the dataset's own notes), plus isometric-3D stills. Do not use these weights to impersonate anyone.

Train it yourself / help improve it: cheatsheet Β· train from your own mp4s Β· results and renders (issue 1)

Research adapters that let MiniMax-H3's DiT run on a Level-of-Token layout: fewer, larger tokens where detail is low, with the full-resolution velocity recovered afterwards (Nakayama et al., arXiv 2610.05816). On one 24 GB card, a 37-frame 768Γ—1344 DiT forward measured 51.9 s dense vs 19.6 s LoT (2.64Γ—). The VAE decode is unchanged.

Folder Run Data
run3_iso3d_distill/ 3 runs of 2,000 steps each, ending with teacher distillation 777 isometric-3D stills
run4_nikki_distill/ 1,500 steps on top of run 3, distillation 516 Nikki talking-head + pose clips + the stills

Each folder has adapter.safetensors (LoT heads + fitted extent bank), lora.safetensors (rank-16 LoRA on the FL2VA DiT), train_log.jsonl and its own card.

Status: research. Coarse regions are still softer than dense H3. Renders, eval curves and caveats are in johndpope/MiniMax-H3 issue 1. Code, and training from your own mp4s: scripts/lot (cheatsheet).

Quick start: download β†’ test β†’ validate timing β†’ train

ComfyUI (ComfyUI-MiniMax-H3-Image-Lane, stills and video clips):

cd ComfyUI/custom_nodes && git clone https://github.com/johndpope/ComfyUI-MiniMax-H3-Image-Lane.git && cd ../..
hf download johndpope/MiniMax-H3-LoT run4_nikki_distill/adapter.safetensors run4_nikki_distill/lora.safetensors --local-dir /tmp/lot
mkdir -p ComfyUI/models/lot
cp /tmp/lot/run4_nikki_distill/adapter.safetensors ComfyUI/models/lot/run4_nikki_distill_adapter.safetensors
cp /tmp/lot/run4_nikki_distill/lora.safetensors    ComfyUI/models/loras/minimax_h3_lot_run4_nikki_lora.safetensors

Load workflows/lot_video_api.json (22-frame clip) or workflows/t1_lot_image_api.json (still). Each samples one seed twice, LoT and dense. Validate the timing by comparing the console's it/s for the two runs. Use the int8 ConvRot FL2VA DiT the adapters were trained on; LoT Unscale Latent is required before VAE Decode.

Python (johndpope/MiniMax-H3 scripts/lot + Fizgig branch immiscible-h3-noise):

export FIZGIG_SRC=/path/to/Fizgig/src LOT_H3_CHECKPOINT=/path/to/minimax_h3_fl2va_pruned_int8_convrot.safetensors
hf download johndpope/MiniMax-H3-LoT --local-dir runs/lot_hf
python3 scripts/lot/test_lot.py                                                        # CPU tests
python3 scripts/lot/parity_h3.py                                                       # LoT 1x1 == dense
python3 scripts/lot/time_lot_h3.py --shape clip --trained runs/lot_hf/run4_nikki_distill   # timing per layout

Measured on one RTX PRO 4000 Blackwell (24 GB), run-4 weights, one DiT forward, all blocks resident:

Shape center bands uniform2
still 768Γ—1152 1.92Γ— 1.63Γ— 2.16Γ—
clip 512Γ—768, 22 frames 2.51Γ— 1.96Γ— 3.08Γ—
clip 768Γ—1344, 124 frames (48 blocks streamed) 2.61Γ—

In ComfyUI's sampler: 2.6Γ— per step on the 22-frame clip, 1.35Γ— on a still. The VAE decode is not sped up.

Train your own (from mp4s or existing H3 latents, teacher distillation recipe): cheatsheet. Publish with upload_hf.py, which adds the license, NOTICE and these warnings.

Derived from MiniMax H3 weights; use is governed by the MiniMax H3 Community License Agreement.

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support

Model tree for johndpope/MiniMax-H3-LoT

Adapter
(127)
this model

Paper for johndpope/MiniMax-H3-LoT