samuelcardillo commited on
Commit
a8ad858
Β·
verified Β·
1 Parent(s): 188256d

Add model card

Browse files
Files changed (1) hide show
  1. README.md +118 -0
README.md ADDED
@@ -0,0 +1,118 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ ---
2
+ language:
3
+ - en
4
+ license: apache-2.0
5
+ tags:
6
+ - qwen3.5
7
+ - moe
8
+ - hermes
9
+ - agentic
10
+ - tool-calling
11
+ - qlora
12
+ - unsloth
13
+ - carnice
14
+ base_model: Qwen/Qwen3.5-35B-A3B
15
+ datasets:
16
+ - bespokelabs/Bespoke-Stratos-17k
17
+ - AI-MO/NuminaMath-CoT
18
+ - kai-os/carnice-glm5-hermes-traces
19
+ - open-thoughts/OpenThoughts-Agent-v1-SFT
20
+ ---
21
+
22
+ # Carnice MoE 35B-A3B β€” Hermes-Focused Agentic Model (GGUF)
23
+
24
+ QLoRA fine-tune of **Qwen3.5-35B-A3B** (MoE, 3B active parameters) optimized for **agentic workflows** and **Hermes Agent** runtime. Two-stage training adapted from [kai-os/Carnice-9b](https://huggingface.co/kai-os/Carnice-9b).
25
+
26
+ ## Credits
27
+
28
+ Training methodology adapted from **[kai-os/Carnice-9b](https://huggingface.co/kai-os/Carnice-9b)** β€” same two-stage approach and datasets, applied to the larger MoE architecture. Key inspiration: training on actual Hermes Agent execution traces for native agentic behavior.
29
+
30
+ ## Available Quantizations
31
+
32
+ | Quantization | Size | BPW | Min VRAM |
33
+ |---|---|---|---|
34
+ | **Q8_0** | 35 GB | 8.52 | 1x 48GB GPU |
35
+ | **Q6_K** | 27 GB | 6.58 | 1x 32GB GPU |
36
+ | **Q5_K_M** | 24 GB | 5.70 | 1x 32GB GPU |
37
+ | **Q4_K_M** | 20 GB | 4.87 | 1x 24GB GPU |
38
+ | **MXFP4_MOE** | 19 GB | 4.39 | 1x 24GB GPU |
39
+
40
+ For BF16 safetensors, see [samuelcardillo/Carnice-MoE-35B-A3B](https://huggingface.co/samuelcardillo/Carnice-MoE-35B-A3B).
41
+
42
+ ## Model Details
43
+
44
+ | Property | Value |
45
+ |---|---|
46
+ | Base Model | [Qwen/Qwen3.5-35B-A3B](https://huggingface.co/Qwen/Qwen3.5-35B-A3B) |
47
+ | Architecture | Mixture of Experts (MoE) |
48
+ | Total Parameters | ~35B |
49
+ | Active Parameters | ~3B per token |
50
+
51
+ ## What Makes This Different
52
+
53
+ Unlike generic reasoning distillation, this model was trained on **actual Hermes Agent execution traces** β€” real conversations where an AI agent:
54
+ - Executes terminal commands and processes output
55
+ - Performs file editing operations
56
+ - Chains multi-step tool calls with results feeding back
57
+ - Uses browser-assisted workflows
58
+ - Makes decisions based on environmental feedback
59
+
60
+ This teaches the model the exact conversation patterns Hermes expects, rather than just generic reasoning.
61
+
62
+ ## Training Details
63
+
64
+ ### Two-Stage Approach
65
+
66
+ **Stage A β€” Reasoning Repair** (1 epoch)
67
+ - Strengthens base model reasoning before agent-specific training
68
+ - Loss: 0.4159
69
+
70
+ | Dataset | Examples |
71
+ |---|---|
72
+ | [bespokelabs/Bespoke-Stratos-17k](https://huggingface.co/datasets/bespokelabs/Bespoke-Stratos-17k) | 16,710 |
73
+ | [AI-MO/NuminaMath-CoT](https://huggingface.co/datasets/AI-MO/NuminaMath-CoT) | 17,000 (capped) |
74
+
75
+ **Stage B β€” Hermes Traces** (2 epochs)
76
+ - Agent-specific behavioral training on real execution traces
77
+ - Loss: 0.3115
78
+
79
+ | Dataset | Examples |
80
+ |---|---|
81
+ | [kai-os/carnice-glm5-hermes-traces](https://huggingface.co/datasets/kai-os/carnice-glm5-hermes-traces) | 1,627 (high quality) |
82
+ | [open-thoughts/OpenThoughts-Agent-v1-SFT](https://huggingface.co/datasets/open-thoughts/OpenThoughts-Agent-v1-SFT) | 15,209 |
83
+
84
+ ### Training Configuration
85
+
86
+ | Parameter | Stage A | Stage B |
87
+ |---|---|---|
88
+ | LoRA Rank | 64 | 64 |
89
+ | LoRA Alpha | 64 | 64 |
90
+ | LoRA Targets | q, k, v, o projections | q, k, v, o projections |
91
+ | Learning Rate | 2e-5 (linear) | 1e-5 (cosine) |
92
+ | Epochs | 1 | 2 |
93
+ | Effective Batch | 12 | 12 |
94
+ | Context Length | 4096 | 4096 |
95
+ | Precision | 4-bit QLoRA + BF16 adapters | Same |
96
+ | GPU | RTX PRO 6000 Blackwell (96GB) | Same |
97
+ | Total Training Time | ~44 hours (both stages) |
98
+
99
+ ### Trainable Parameters
100
+ 6,881,280 (0.02% of 35B total)
101
+
102
+ ## Usage with llama.cpp
103
+
104
+ ```bash
105
+ llama-server \
106
+ --model Carnice-MoE-35B-A3B-Q8_0.gguf \
107
+ --n-gpu-layers -1 \
108
+ --ctx-size 131072 \
109
+ --host 0.0.0.0 --port 8082
110
+ ```
111
+
112
+ ## Acknowledgements
113
+
114
+ - **[kai-os](https://huggingface.co/kai-os)** β€” Carnice training methodology and Hermes traces dataset
115
+ - **[open-thoughts](https://huggingface.co/open-thoughts)** β€” Agent SFT dataset
116
+ - **[bespokelabs](https://huggingface.co/bespokelabs)** β€” Bespoke-Stratos reasoning dataset
117
+ - **[Unsloth](https://unsloth.ai)** β€” QLoRA training framework
118
+ - **[Qwen](https://huggingface.co/Qwen)** β€” Base model