Download run3_iso3d_distill/README.md from johndpope/MiniMax-H3-LoT: direct link, hf CLI and curl.
- Browser
- Download file 5.53 kB
-
https://huggingface.co/johndpope/MiniMax-H3-LoT/resolve/main/run3_iso3d_distill/README.md
- Command line
-
hf download hf://johndpope/MiniMax-H3-LoT/run3_iso3d_distill/README.md
-
curl -L -o README.md https://huggingface.co/johndpope/MiniMax-H3-LoT/resolve/main/run3_iso3d_distill/README.md
license: other
license_name: minimax-h3-community-license
license_link: LICENSE
base_model: MiniMaxAI/MiniMax-H3
tags:
- level-of-token
- lora
- video-generation
- research
MiniMax-H3 Level-of-Token adapter: run 3: isometric stills, teacher distillation
⚠️ THIS MODEL NEEDS MORE TRAINING — HELP WANTED
These are research checkpoints, not a finished model. LoT output is still visibly worse than dense MiniMax H3: coarse regions are soft, faces in talking-head shots distort, and the LoRA can bend bodies even with LoT off. If you have GPU time, data, or ideas, join the discussion: shootthesound/Fizgig discussions #183.
License and territory. Derived from MiniMax H3 under the MiniMax H3 Community License Agreement (see also
NOTICE). That license allows use and redistribution only outside its Excluded Territories: the European Union, the United Kingdom, the Republic of Korea and the United States of America. Do not download or use these files there. Commercial use above US$20M yearly revenue needs MiniMax's written authorization. The Acceptable Use Policy applies, including: no impersonation without consent, and machine-generated output must be disclosed as such.
Training data. Run 4 includes clips of "Nikki", a character whose reference face comes from a Grok Imagine generation (per the dataset's own notes), plus isometric-3D stills. Do not use these weights to impersonate anyone.
Train it yourself / help improve it: cheatsheet · train from your own mp4s · results and renders (issue 1)
Research weights for running MiniMax-H3's DiT on a Level-of-Token (LoT) layout: fewer, larger tokens where detail is low, with the full-resolution velocity recovered afterwards (Nakayama et al., arXiv 2610.05816). The VAE latent stays full size; only the transformer sequence shrinks. On one 24 GB card a 37-frame 768×1344 DiT forward measured 51.9 s dense vs 19.6 s LoT (2.64×). The VAE decode is unchanged.
Status: research. Coarse regions are still softer than dense H3. Progress, renders and caveats are tracked in the issue.
Files
| File | What |
|---|---|
adapter.safetensors |
LoT adapter: per-extent input/output heads, shape MLP, and the Procrustes extent bank (A, s) it was trained with |
lora.safetensors |
Rank-16 LoRA on blocks.*.attn.qkv_proj, attn.out_proj, mlp.fc1, mlp.fc2 (200 modules) of the FL2VA DiT |
train_log.jsonl |
Per-step loss and held-out evals |
Use
Needs the code in scripts/lot, Fizgig with the LoT hooks (branch immiscible-h3-noise), and the pruned int8 FL2VA DiT (minimax_h3_fl2va_pruned_int8_convrot.safetensors).
hf download johndpope/MiniMax-H3-LoT --local-dir runs/lot_hf
python3 scripts/lot/time_lot_h3.py --shape clip --trained runs/lot_hf/run3_iso3d_distill # timing per layout
python3 scripts/lot/render_h3.py --trained runs/lot_hf/run3_iso3d_distill --swap 4 \
--layout bands --variants dense,dense_lora,lot_trained # needs a cache of encoded prompts
ComfyUI (stills and video): ComfyUI-MiniMax-H3-Image-Lane. Quick start and training from your own mp4s: cheatsheet.
Training
Warm start chain: run 1 (2,000 steps, shift 12) → run 2 (2,000 steps, LoT shift 3) → run 3 (2,000 steps, LoT shift 3, --distill 1). Data: 777 isometric-3D stills with cached text (cache_iso3d), 24 held out. Layouts per step: 20% dense, uniform extents, detail-driven 4×4 mosaics. Rank 16, AdamW 8-bit, lr 1e-4, batch 1, swap 4, one RTX PRO 4000 Blackwell (24 GB), 2 h 50 min.
Loss: eq. 17 with H3's x0 − ε head (velocity MSE in y-space), plus a distillation term that pulls LoT's clean estimate toward the frozen dense base (LoRA off) from the same noisy input.
Held-out eval
dense / uniform2 / mosaic are the data loss on those layouts. gap_* is the weighted distance from LoT to the frozen dense teacher. Lower is better.
| step | dense | uniform2 | mosaic | gap_uniform2 | gap_mosaic |
|---|---|---|---|---|---|
| 0 | 0.2959 | 0.1835 | 0.2496 | 0.6732 | 0.6934 |
| 250 | 0.3076 | 0.1878 | 0.2551 | 0.6376 | 0.6475 |
| 500 | 0.3094 | 0.1894 | 0.2536 | 0.6172 | 0.6119 |
| 750 | 0.3094 | 0.1846 | 0.2516 | 0.6092 | 0.6071 |
| 1000 | 0.3091 | 0.1836 | 0.2513 | 0.5969 | 0.6044 |
| 1250 | 0.3179 | 0.1883 | 0.2567 | 0.5833 | 0.5826 |
| 1500 | 0.3078 | 0.1822 | 0.2512 | 0.5794 | 0.5833 |
| 1750 | 0.3091 | 0.1851 | 0.2537 | 0.5609 | 0.5726 |
| 2000 | 0.3056 | 0.1830 | 0.2523 | 0.5688 | 0.5670 |
License
Derived from MiniMax H3 weights; use is governed by the MiniMax H3 Community License Agreement.