Merge the official Qwen 35B-A3B models using the most advanced data-free merging algorithms available today.

Model Highlights:

  • Merge Method: SWUDI-A
  • Precision: dtype: bfloat16
  • Context Length: 262,144

Parameter Settings:

Non-Thinking Mode: ({%- set enable_thinking = false %})

General Tasks:

temperature=0.7, top_p=0.8, top_k=20, min_p=0.0, presence_penalty=1.5, repetition_penalty=1.0

Reasoning Tasks:

temperature=1.0, top_p=1.0, top_k=40, min_p=0.0, presence_penalty=2.0, repetition_penalty=1.0

Thinking Mode: ({%- set enable_thinking = true %})

General Tasks:

temperature=1.0, top_p=0.95, top_k=20, min_p=0.0, presence_penalty=1.5, repetition_penalty=1.0

Coding Tasks:

temperature=0.6, top_p=0.95, top_k=20, min_p=0.0, presence_penalty=0.0, repetition_penalty=1.0

Papers:

https://arxiv.org/abs/2606.07289v1

Some Details:

The specific merging details are as follows:

1. Initialization starting point: Simple averaging

2. Layer-wise rank truncation rule: Participation-Square-Root Rule

Downloads last month
36
Safetensors
Model size
36B params
Tensor type
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for YOYO-AI/Qwen3.6-35B-A3B-YOYO-V2

Collection including YOYO-AI/Qwen3.6-35B-A3B-YOYO-V2

Paper for YOYO-AI/Qwen3.6-35B-A3B-YOYO-V2