LiteRT is Google's on-device runtime, the new name for TensorFlow Lite (Android: com.google.ai.edge.litert:litert), and litert-torch, the renamed ai-edge-torch, is its PyTorch converter: a PyTorch model converted unmodified with litert_torch.convert matched the original to 4e-7 on a Galaxy S26 (measured, LiteRT 2.2.0, Android 16, 2026-09-05).

Measured on device (edge-compat, qwen3-4b-thinking): Galaxy S26 Β· LiteRT-LM prebuilt-adac974b Β· GPU Β· decode 18.1 tok/s Β· prefill 263 tok/s Β· TTFT 3.95 s Β· all 1707 ops delegated (2026-10-05); Mac Studio M4 Max Β· LiteRT-LM 0.17.1 Β· GPU Β· decode 92.1 tok/s Β· prefill 1100 tok/s Β· TTFT 244 ms Β· all 1707 ops delegated (2026-10-05); Galaxy S26 Β· LiteRT-LM prebuilt-adac974b Β· CPU Β· decode 7.3 tok/s Β· prefill 56 tok/s Β· TTFT 18.95 s (2026-10-05). Record: https://github.com/john-rocky/edge-compat/blob/main/cards/qwen3-4b-thinking/CARD.md

Measured on device (edge-compat, qwen3-4b-thinking-dynamic-wi4b32-afp32): Galaxy S26 Β· LiteRT-LM 0.16.0 Β· GPU Β· decode 11.2 tok/s Β· prefill 129 tok/s Β· TTFT 1.65 s Β· all 1282 ops delegated (2026-08-24); Raspberry Pi 5 Β· LiteRT-LM 0.16.1 Β· CPU, 4 threads Β· decode 1.4 tok/s Β· prefill 9 tok/s Β· TTFT 29.40 s (2026-09-01). Record: https://github.com/john-rocky/edge-compat/blob/main/cards/qwen3-4b-thinking-dynamic-wi4b32-afp32/CARD.md

Qwen3-4B-Thinking-2507 β€” LiteRT-LM (blockwise int4)

Qwen/Qwen3-4B-Thinking-2507 converted to the LiteRT-LM (.litertlm) format for on-device inference with Google's LiteRT-LM runtime (the engine behind the official litert-community/* models).

Qwen3-4B-Thinking-2507 is a dense 4B reasoning model (Qwen3ForCausalLM, 36 layers) that operates exclusively in thinking mode β€” it emits a <think>…</think> chain before its answer β€” so it rides the existing Qwen3 converter and runtime directly.

Artifact (Block 128)

File model.litertlm β€” int4 block 128 (~2.3 GB)
Quantization int4 weights (symmetric) + OCTAV optimal-clipping; embeddings INT8 (externalized section)
Compute integer
Context (KV cache) 4096
Base model Qwen/Qwen3-4B-Thinking-2507

2026-10-07: chat template made parts-safe. Qwen3_4b_thinking_dynamic_wi4b32_afp32.litertlm was replaced by a copy whose embedded Jinja template reads message content both as a string (what litert-lm 0.17.1 passes) and as a list of typed parts (what 0.18.0 passes). On litert-lm 0.18.0, the previous file stopped at the template with tried to use + operator on unsupported types string and sequence; on 0.17.1 it answered. Only the template section changed: the TFLite graph and weights, the tokenizer and the executor metadata are byte-identical, and for string content the new template renders the same prompt as the old one, byte for byte. The new file is still 2,274,193,168 bytes, sha256 110098cf… (full value in litertlm_manifest.json); the previous file (Hub revision 6c4863d9) had sha256 a412d77d…. Tested on macOS: with litert-lm 0.18.0 the new file answers on CPU and scores 8/8 sanity-gate questions on the Mac GPU; with 0.17.1 it gives the same text as the previous file. Tested on a Galaxy S26 with litert_lm_main built from the v0.18.0 tag: the new file answers on the GPU (OpenCL) and on the CPU.

2026-10-06: GPU graph rewrite. model.litertlm was replaced by a file with a rewritten prefill/decode graph. In each layer, the two DYNAMIC_UPDATE_SLICE KV-cache writes became one odml.cache_update composite, and the BATCH_MATMUL attention products became odml.runtime_bmm composites. This is the graph shape litert-torch exports with --apply_gpu_composites; the previous file was exported without that flag. The graph also gains a signature input, param_tensor (INT32[1,1,1,7]), whose start/end values the runtime fills. The attention mask stays a FLOAT32 input. A 16-token prefill signature, prefill_16, was added as a copy of prefill_128 that shares its weight buffers. Every other bundle section, and the weight region of the graph, is byte-identical to the previous file (checked per section and per buffer). The new file is 2,477,192,112 bytes, sha256 8a8a6a419a5570b3f4814d358fb795c0dcb6931206c7dae29480f47564c3f04d. The previous file (Hub revision 92751be3c031) was 2,474,357,680 bytes, sha256 356b2d7778d27d55fbb0d2e0c5e8801e9de136db59849c41c882bd97333840d5. On CPU, prefill_16 gives output bit-identical to prefill_128 on the same tokens. With and without prefill_16, the answers to 9 gate prompts (8 sanity-gate questions plus one long prompt) are byte-identical, on CPU and on the Mac GPU with fp32 activations. On the Mac GPU with fp32 activations, the rewritten graph without prefill_16 also matches the previous file byte for byte on all 8 questions. On the Mac GPU with the default fp16 activations, 5 of 30 GSM8K answer texts stay byte-identical. The scores hold: both files answer 8/8 sanity-gate questions and 26/30 GSM8K questions (max tokens 2048, greedy), and each solves one GSM8K question the other misses. The Raspberry Pi 5 rows and the iPhone 17 Pro note further down were measured on the previous file. The commands are under Conversion.

Mac Studio M4 Max, litert-lm 0.17.1 benchmark --cache no, both files measured in the same window (median of 3 processes on GPU, 2 on CPU; GPU is WebGPU on Metal with fp16 activations, CPU is XNNPACK):

Measure Previous file New file New / previous
GPU decode (256-token prompt, 256 decode tokens) 66.6 tok/s 92.1 tok/s Γ—1.383
GPU prefill (256-token prompt) 984 tok/s 1100 tok/s Γ—1.118
GPU TTFT (16-token prompt) 0.152 s 0.037 s Γ—0.247
CPU decode (256-token prompt, 256 decode tokens) 18.5 tok/s 18.4 tok/s Γ—0.997
CPU TTFT (16-token prompt) 1.444 s 0.478 s Γ—0.331

Artifact (Block 32)

File Quantization Recipe Context Size
Qwen3_4b_thinking_dynamic_wi4b32_afp32.litertlm dynamic_wi4b32_afp32 (block-32) 4096 2.1 GB

Conversion Notes

Qwen3_4b_thinking_dynamic_wi4b32_afp32.litertlm is a dynamic INT4 variant (block-32 weights, FP32 activations). It was converted through the LiteRT Torch (litert-torch) path and quantized with AI Edge Quantizer. This artifact incorporates LiteRT-LM GPU graph optimizations, including composite ops for RoPE, fused QKV, and fused Gate/Up projections, and is configured with static prefill memory allocation.

Update (2026-08-31) β€” thought channel and think pre-fill repaired in place. Both bundles received two metadata-only repairs today (weights byte-identical each time, verified section by section). First, the thought channel (<think>\n / \n</think>) was declared in the bundle metadata β€” without it the runtime streams raw reasoning inline into the answer and silently ignores any thinking budget. Second, the generation prompt now pre-opens the think block exactly as the upstream Qwen3-4B-Thinking-2507 template does (<|im_start|>assistant\n<think>\n): model.litertlm's structured assistant prefix and the block-32 file's embedded template both previously ended at assistant\n, leaving the model to emit <think> on its own β€” a discipline quantized reasoning models lose first. After the repair model.litertlm scores 8/8 on the 8-question sanity gate on both macOS backends (CPU and Metal GPU, max-tokens 2048), and the block-32 file scores 8/8 on CPU.

⚠️ It's a reasoning model β€” give it room to think

This model generates a <think>…</think> reasoning chain, then the answer. Run it with max_tokens β‰₯ 2048 β€” at a short limit it gets cut off mid-thought and never reaches the answer. (All quality numbers below were measured at 2048.)

Performance

All rows are for the current model.litertlm (the 2026-10-06 file).

Mac Studio M4 Max. litert-lm 0.17.1 (pip) CLI, litert-lm benchmark --cache no, 256 prompt / 256 decode tokens. Each value is the median of separate processes: 3 on GPU, 2 on CPU. The 16-token TTFT comes from separate runs with a 16-token prompt and 32 decode tokens.

Backend Prefill (256 tokens) Decode (256 tokens) TTFT (256-token prompt) TTFT (16-token prompt) Init
GPU (WebGPU on Metal, fp16 activations) 1100 tok/s 92.1 tok/s 0.24 s 0.037 s 2.2 s
CPU (XNNPACK) 113 tok/s 18.4 tok/s 2.57 s 0.478 s 2.1 s

Samsung Galaxy S26 (SM-S942Q, Android 16). litert_lm_advanced_main from a LiteRT-LM main build of 2026-09-18, 1024 prompt / 256 decode tokens. The GPU runs use --disable_cache=true (cold init); the CPU runs use the XNNPACK weight cache, which init writes. Each run is one process: 3 iterations on GPU, 2 on CPU, and 3 in each separate 16-token run (16 prompt / 32 decode tokens). Each cell is the range over those iterations. In the 1024-token GPU run the first iteration is the fastest and the third the slowest (cold to warm). In the 16-token runs the first iteration has the longest TTFT. Init and peak memory (the process high-water mark, VmHWM) are from the 1024-token run.

Backend Prefill (1024 tokens) Decode (256 tokens) TTFT (1024-token prompt) TTFT (16-token prompt) Init Peak
GPU (OpenCL, fp16 activations) 130–264 tok/s 13.9–20.2 tok/s 3.93–7.95 s 0.14–0.15 s 7.9 s 1.17 GB
CPU (XNNPACK, 4 threads) 47–65 tok/s 7.2–7.4 tok/s 15.78–22.13 s 0.40–1.29 s 2.8 s 4.08 GB

Accuracy note

Measured on GSM8K (n=100, greedy, 0-shot chain-of-thought, max_tokens 2048, identical prompt and answer-extraction for every row).

Configuration GSM8K
bf16 (reference) 90.0%
LiteRT int4 β€” block 128 86.0% (βˆ’4 pt)

int4 is at parity (βˆ’4 pt). Note: evaluating a reasoning model at a short token budget badly understates int4 β€” the longer int4 reasoning chains get truncated before the answer; benchmark reasoning models with max_tokens β‰₯ 2048.

Why block 128 (and not block 32)? For this reasoning model the block-32 build degraded more (βˆ’9 pt) and produced corrupted output under the iPhone GPU delegate, while block 128 is robust on every backend, ~40 % faster to decode (ΒΌ the dequant scales β€” which matters when generating long <think> chains), and stays at βˆ’4 pt parity. So block 128 is the build we recommend; the block-32 file is also published (see Artifact (Block 32) above), and both bundles run fully delegated on the Galaxy S26 GPU.

The block-32 file's GPU corruption is not iPhone-specific: on macOS Metal it degenerates too (question-echo loops, truncated reasoning; measured 2026-08-31 on the file both before and after the metadata repair, so it is pre-existing and unrelated to the repair). Use the block-32 file on CPU; on Android GPU it delegated and generated in the Galaxy S26 gate above.

Galaxy S26 β€” GPU backend

Both published bundles run on the Android GPU backend and generate.

model.litertlm (the 2026-10-06 file). On a Samsung Galaxy S26 (SM-S942Q, Android 16) with litert_lm_advanced_main from a LiteRT-LM main build of 2026-09-18, the three transformer signatures run fully on the OpenCL GPU delegate (LITERT_CL), one partition each:

signature nodes on LITERT_CL
prefill_128 1707 / 1707
prefill_16 1707 / 1707
decode 1631 / 1631

XNNPACK takes 1 of the 4 nodes in decode_embedder and 1 of the 4 nodes in prefill_embedder_128. On the GPU the file answered 5 of 5 gate prompts correctly. Speed and peak memory are in the Galaxy S26 table under Performance.

Block-32 file.

file GPU backend delegation peak
Qwen3_4b_thinking_dynamic_wi4b32_afp32.litertlm runs 3764 / 3764 ops across 3 subgraphs on LiteRT GPU 1627 MB

Measured on a Samsung Galaxy S26 (SM-S942Q / SM8850, Android 16) with litert_lm_advanced_main from litert-lm 0.16.0, --backend=gpu --sampler_backend=cpu, prompt What is the capital of France?. Peak is the process high-water mark (VmHWM) sampled during that same run. Gated 2026-08-24.

The op counts above are the LiteRT GPU partitions.

GPU wiring, including the Gallery import toggle: GPU guide.

Usage

# build litert-lm from https://github.com/google-ai-edge/litert-lm, then:
litert_lm_main \
  --model_path model.litertlm \
  --backend gpu \
  --input_prompt "A bat and a ball cost \$1.10. The bat costs \$1.00 more than the ball. How much is the ball?"

The .litertlm bundle carries the tokenizer and prompt template (Qwen3 ChatML β€” <|im_start|>role\n…<|im_end|>, stop token <|im_end|>), so no separate tokenizer files are needed. The model produces a <think>…</think> block followed by its answer.

Run on Android

Update (July 2026): Google AI Edge Gallery v1.0.16+ can import litert-lm models directly from Hugging Face inside the app (tap +) β€” no computer or adb needed. The manual steps below are only required on older builds or for sideloading a local file.

The official Google AI Edge Gallery app runs .litertlm models on-device:

  1. Install a recent Gallery (package com.google.ai.edge.gallery, 1.0.15+ supports .litertlm).
  2. Download model.litertlm and push it: adb push model.litertlm /sdcard/Download/
  3. In the app tap +, pick the file, choose the GPU backend, and raise the max-tokens setting (β‰₯2048).
  4. Chat β€” the bundle already carries the tokenizer and Qwen3 chat template.

A 4B int4 build needs ~2.5 GB free RAM; reboot the phone first if memory is tight.

Run on desktop (LiteRT-LM CLI)

The same .litertlm bundle runs on macOS / Linux / Windows with the official LiteRT-LM CLI β€” including as a local OpenAI-compatible API server:

pip install litert-lm
litert-lm import --from-huggingface-repo litert-community/Qwen3-4B-Thinking-2507 model.litertlm qwen3-4b-thinking-2507
litert-lm run qwen3-4b-thinking-2507     # interactive chat in the terminal
litert-lm serve

Run on iPhone

Verified on iPhone 17 Pro (LiteRT-LM Swift runtime): loads and generates at ~14 tok/s (previous file).

Conversion

Converted with the official litert-torch converter β€” a standard Qwen3ForCausalLM, so it uses the existing Qwen3 path with no custom graph code. Recipe: blockwise-128 int4 + OCTAV (INT4 weights, block 128, symmetric, OCTAV optimal-clipping), embeddings INT8, KV cache 4096.

from litert_torch.generative.export_hf.export import export
export(
    model="Qwen/Qwen3-4B-Thinking-2507",
    output_dir="out",
    quantization_recipe="qwen3_int4_block128_octav.json",  # blockwise-128 int4 + OCTAV, int8 embeddings
    cache_length=4096,
    externalize_embedder=True,
)

Graph rewrite (2026-10-06). The current model.litertlm was built from the previous file (Hub revision 92751be3c031) with two tools in tools/gpu_graph/ of hf-to-litertlm:

python tools/gpu_graph/gpu_graph_retrofit.py previous.litertlm retrofit.litertlm --decomp exporter --mask add_bcast
python tools/gpu_graph/prefill_bucket_clone.py build retrofit.litertlm new.litertlm --lengths 16 --source prefill_128 --report bucket.json

previous.litertlm is the previous file, new.litertlm the result published here as model.litertlm; LITERT_LM_CLI names the litert-lm CLI the scripts use to unpack and pack the bundle.

gpu_graph_retrofit.py changes only the prefill/decode TFLite graph, and every original weight buffer keeps its bytes. With --mask add_bcast, the attention mask is applied as an ADD on a [bk, g, T, C] view of the logits, with the FLOAT32 mask broadcast. prefill_bucket_clone.py copies an existing prefill signature (prefill_128) at a new length (16); the copy shares every weight buffer, so the file grows by graph structure only. A rebuild with these tools has every bundle section byte-identical to the shipped file. Only the bundle header (uuid and timestamp written by litert-lm pack) differs, so the whole-file sha256 differs.

Raspberry Pi 5 (CPU) (previous file)

The model.litertlm row was measured on the previous file (sha256 356b2d7778d27d55fbb0d2e0c5e8801e9de136db59849c41c882bd97333840d5) and not re-measured after the 2026-10-06 graph rewrite; the weights are byte-identical in the new file, and the new file's Mac CPU speed is Γ—0.997 (decode) and Γ—1.002 (prefill) of the previous file's, so the row is expected to hold.

Measured on a Raspberry Pi 5 Model B Rev 1.1 (8 GB, Raspberry Pi OS 64-bit) with litert-lm benchmark 0.16.1: CPU backend, 4 threads, 256 prefill + 256 decode tokens, --cache memory (the compile cache lives and dies with the process, so every invocation compiles the model from scratch; nothing is reused between runs), one warm-up plus one timed iteration per invocation, 3 invocations per file with cooldown in between. Values are the median across invocations (min–max in parentheses). No thermal throttling occurred during these runs (vcgencmd get_throttled stayed 0x0). Every file listed produced coherent text in a real generation on this backend before its numbers were recorded.

File Prefill (tok/s) Decode (tok/s) TTFT Peak RSS
Qwen3_4b_thinking_dynamic_wi4b32_afp32.litertlm 8.9 (8.8–9.0) 1.4 (1.4–1.4) 29.4 s 4.4 GB
model.litertlm 10.6 (10.5–10.7) 1.5 (1.5–1.5) 25.5 s 3.9 GB

License

Apache-2.0, inherited from the base model Qwen/Qwen3-4B-Thinking-2507.

Downloads last month
938
Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support

Model tree for litert-community/Qwen3-4B-Thinking-2507

Quantized
(140)
this model

Collections including litert-community/Qwen3-4B-Thinking-2507