Xunzhuo commited on
Commit
cde2a68
路
1 Parent(s): 23fdf61

Clarify model-only release and serving guidance

Browse files

Signed-off-by: Xunzhuo <[email protected]>

Files changed (1) hide show
  1. README.md +1 -1
README.md CHANGED
@@ -94,7 +94,7 @@ The published model's complete state, question, and candidates have a 16,384-tok
94
 
95
  ![Decision decoder architecture](assets/architecture.png)
96
 
97
- A causal Qwen3.5 text backbone combines gated linear and full attention. A shared candidate head reads candidate endpoints and the final query vector. Each question uses one forward pass; questions run independently in batches of eight.
98
 
99
  [Candidate head](assets/readout.png) 路 [Vector architecture](assets/architecture.svg)
100
 
 
94
 
95
  ![Decision decoder architecture](assets/architecture.png)
96
 
97
+ A causal Qwen3.5 text backbone combines gated linear and full attention. A shared candidate head reads candidate endpoints and the final query vector. The serving runtime schedules questions according to available hardware and request load.
98
 
99
  [Candidate head](assets/readout.png) 路 [Vector architecture](assets/architecture.svg)
100