Clarify model-only release and serving guidance
Browse filesSigned-off-by: Xunzhuo <[email protected]>
README.md
CHANGED
|
@@ -94,7 +94,7 @@ The published model's complete state, question, and candidates have a 16,384-tok
|
|
| 94 |
|
| 95 |

|
| 96 |
|
| 97 |
-
A causal Qwen3.5 text backbone combines gated linear and full attention. A shared candidate head reads candidate endpoints and the final query vector.
|
| 98 |
|
| 99 |
[Candidate head](assets/readout.png) 路 [Vector architecture](assets/architecture.svg)
|
| 100 |
|
|
|
|
| 94 |
|
| 95 |

|
| 96 |
|
| 97 |
+
A causal Qwen3.5 text backbone combines gated linear and full attention. A shared candidate head reads candidate endpoints and the final query vector. The serving runtime schedules questions according to available hardware and request load.
|
| 98 |
|
| 99 |
[Candidate head](assets/readout.png) 路 [Vector architecture](assets/architecture.svg)
|
| 100 |
|