--- title: MiniMax-H3 Turbo-SLA emoji: 🎬 colorFrom: pink colorTo: blue sdk: gradio sdk_version: 6.25.0 app_file: app.py pinned: false short_description: 4-step distilled image-to-video with SLA sparse attention python_version: "3.12" startup_duration_timeout: 1h models: - MiniMaxAI/MiniMax-H3 - lightx2v/Minimax-h3-Turbo-SLA --- # MiniMax-H3 Turbo-SLA 4-step distilled FL2V (first-and-last-frame to video) generation with **SLA (Sparse-Linear Attention)** for efficient inference, producing video with a synchronized soundtrack. This demo applies the [lightx2v/Minimax-h3-Turbo-SLA](https://huggingface.co/lightx2v/Minimax-h3-Turbo-SLA) LoRA — a 4-step turbo distillation checkpoint with 85% attention sparsity — on top of the [MiniMaxAI/MiniMax-H3](https://huggingface.co/MiniMaxAI/MiniMax-H3) base model. ## Architecture This Space is the **denoising half** of a split MiniMax-H3 deployment: - The 62 GiB Qwen3-VL text encoder/conditioner runs in [`multimodalart/qwen3vl-conditioner`](https://huggingface.co/spaces/multimodalart/qwen3vl-conditioner), called over the gradio API for each request. - This Space loads the 61.7 GiB transformer + 10.4 GiB VAEs (bfloat16, unquantized) and applies the SLA Turbo LoRA at startup. The split is necessary because MiniMax-H3 is 195.9 GiB in bfloat16 and a ZeroGPU Space is evicted at 150 GB of storage. ## SLA (Sparse-Linear Attention) The SLA LoRA was trained with [SLA](https://github.com/thu-ml/SLA) — a fine-tunable sparse-linear attention mechanism that achieves approximately 2.5× inference acceleration on RTX 5090 with 85% attention sparsity ratio while retaining competitive visual quality. > **Note:** The original LightX2V inference path uses SageAttention ("sage2") for the > sparse attention kernel. ZeroGPU's Blackwell (sm_120) hardware does not support > FA3/SageAttention kernels (sm90-only), so this Space uses cuDNN's fused attention > backend instead. The LoRA weights are applied identically; only the attention > kernel differs. ## Usage Provide a text prompt and optionally a first frame and/or last frame image to generate a 5-second video with synchronized stereo audio. The 4-step turbo distillation makes each generation take about 30-60 seconds.