BRIDGE / flux2-klein-9b /README.md
PANDATREE's picture
Link released subject-condition data, training guide, and BBox token-control examples
35d0180 verified
|
Raw History Blame Contribute Delete
4.8 kB

BRIDGE on FLUX.2 Klein 9B

These are two full BF16 transformer checkpoints fine-tuned from black-forest-labs/FLUX.2-klein-base-9B, both at training step 1400. They preserve BRIDGE's main/subject paths and learned discrete PE routing. They are not Qwen LoRA weights and must not be loaded into the old Qwen pipeline.

Setting Sparse / Mask Dense / BBox
Checkpoint directory sparse-mask/checkpoint-1400 dense-bbox/checkpoint-1400
sub_region_mode mask bbox
use_sparse_sub_branch True False
pe_exchange_region mask bbox
Inference steps 50 50
Precision BF16 BF16

What is included

Each checkpoint includes the transformer configuration, safetensors index, two weight shards, an eval-conversion manifest, and a release manifest with SHA-256 checksums. The state contains 9,098,549,280 parameter elements and one additional scalar STE-temperature buffer. No optimizer state is included.

The exports reconstruct ScheduleFree's eval view using torch.lerp(train, z, 1 - 1 / beta1) with beta1=0.95. The source training-run identifiers are retained in each conversion manifest; machine-specific path prefixes are removed. The weights themselves are unchanged.

Loading

Use the custom FLUX implementation. The Gradio backends explicitly construct PE gate modules before loading the full transformer state. A bare call to a standard Diffusers transformer will not restore the custom gate architecture.

Download the VAE, Qwen3 text encoder, tokenizer and scheduler separately from the upstream base model. The comparison launcher used upstream revision 32773329fbe7e81a90ef971740e8ba4b0364ecf3; the full base transformer is not needed when loading these released trained transformers.

Inference uses hard_exchange, 50 steps, a sub positional t-coordinate of 20, and condition positional t-coordinates 40, 60, 80, ... . These are positional coordinates, not diffusion denoising steps. The backend exposes local and global sub spatial coordinates; global is the current Gradio default, while local retains the training crop-local convention.

Subject conditioning and token support

The model accepts a background image, text, and optional subject-reference images. The current Gradio interface accepts up to three references; that is an inference-interface capability, not a claim that training used three refs. Subject-reference conditioning and the generated sub branch are distinct.

In sparse mode, an independent bbox-local Gaussian grid is mask-sparsified: sub tokens outside selected mask cells are omitted before denoising. Changing the retained sub-token support inside a bbox provides a way to influence local generation. This describes the existing mask-to-token path, not a new arbitrary-token-cutting API or a guarantee of pixel-exact control. The dense variant retains the complete bbox grid. Neither construction copies main or background token values into the sub branch; PE routing changes positions.

Data and examples

The FLUX training data is now released as the subject-condition extension of the existing BRIDGE dataset: 27,834 train / 3,092 test rows and 30,926 selected subject references generated with Qwen/Qwen-Image-Edit-2511. See the full training guide. BRIDGE uses the same method on both backbones: Qwen uses LoRA plus gates, while FLUX uses full-transformer plus gate training.

The accompanying GitHub example is an existing 2026-09-04 Gradio comparison, not a newly generated benchmark. Its metadata identifies these two eval exports and 50 inference steps. It used no mask / a full-size sub branch and therefore is not evidence for a token-removal ablation. The source Gradio backend uses a CUDA random generator on CUDA; the historical example must not be described as CPU-RNG output. No new image generation was run for this release.

Additional same-BBox-weight cropped-sub comparisons are now available. Those switch the same BBox checkpoint from dense bbox sub/PE support to mask-selected sparse sub/PE support. Six custom-input examples are included; internal-dataset examples are excluded. Both token support and PE candidate support change. Use the public launcher's --bbox-protocol-compare option for this mode rather than the default two-trained-checkpoint comparison.

License

See LICENSE.md and NOTICE.md. These weights are FLUX derivatives subject to the FLUX Non-Commercial License, not Apache-2.0 weights.