Upload 4 files
Browse files- LICENSE +21 -0
- README.md +85 -0
- config.json +51 -0
- model.safetensors +3 -0
LICENSE
ADDED
|
@@ -0,0 +1,21 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
MIT License
|
| 2 |
+
|
| 3 |
+
Copyright (c) 2026 Shiv Shanmugam
|
| 4 |
+
|
| 5 |
+
Permission is hereby granted, free of charge, to any person obtaining a copy
|
| 6 |
+
of this software and associated documentation files (the "Software"), to deal
|
| 7 |
+
in the Software without restriction, including without limitation the rights
|
| 8 |
+
to use, copy, modify, merge, publish, distribute, sublicense, and/or sell
|
| 9 |
+
copies of the Software, and to permit persons to whom the Software is
|
| 10 |
+
furnished to do so, subject to the following conditions:
|
| 11 |
+
|
| 12 |
+
The above copyright notice and this permission notice shall be included in all
|
| 13 |
+
copies or substantial portions of the Software.
|
| 14 |
+
|
| 15 |
+
THE SOFTWARE IS PROVIDED "AS IS", WITHOUT WARRANTY OF ANY KIND, EXPRESS OR
|
| 16 |
+
IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY,
|
| 17 |
+
FITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT. IN NO EVENT SHALL THE
|
| 18 |
+
AUTHORS OR COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER
|
| 19 |
+
LIABILITY, WHETHER IN AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING FROM,
|
| 20 |
+
OUT OF OR IN CONNECTION WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN THE
|
| 21 |
+
SOFTWARE.
|
README.md
CHANGED
|
@@ -1,3 +1,88 @@
|
|
| 1 |
---
|
| 2 |
license: mit
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 3 |
---
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
---
|
| 2 |
license: mit
|
| 3 |
+
library_name: pytorch
|
| 4 |
+
pipeline_tag: other
|
| 5 |
+
tags:
|
| 6 |
+
- cad
|
| 7 |
+
- freecad
|
| 8 |
+
- agent
|
| 9 |
+
- imitation-learning
|
| 10 |
+
- length-generalization
|
| 11 |
---
|
| 12 |
+
|
| 13 |
+
# Taiga-S1
|
| 14 |
+
|
| 15 |
+

|
| 16 |
+
|
| 17 |
+
**A 1.2M-parameter model that builds CAD parts in FreeCAD.**
|
| 18 |
+
|
| 19 |
+
*Taiga-S1 is an experiment in whether small, fast decision models can be useful for computer-use agents: a planner decides what to do, and a tiny model handles the step-by-step execution. FreeCAD is the testbed.*
|
| 20 |
+
|
| 21 |
+
Taiga-S1 is the fast "System 1" layer for a CAD agent. You give it a goal, an ordered list of features like *"plate 40Γ30Γ10 β Γ6 hole at (10, 0) β polar pattern Γ6 β fillet the top edges"*. It builds the part command by command: select a plane, sketch, draw, constrain, pad, pattern, fillet. At every step it reads FreeCAD's live state (feature tree, selection, sketch constraints, workbench) and picks the next command from the ones currently available.
|
| 22 |
+
|
| 23 |
+
- **Tiny and fast.** 1.2M parameters, trained from scratch, ~1 ms per decision on a CPU. No LLM, no vision model, no screenshots.
|
| 24 |
+
- **Generalizes to longer parts.** Trained on parts with at most 5 features, it builds 11-feature parts (~55 commands) with 100% success and 17-feature parts at 95%. The previous version scored 0% at 6+ features.
|
| 25 |
+
- **Recovers from mistakes.** With 20% of its actions replaced by random ones, it notices the damage, undoes it and finishes 86β100% of parts.
|
| 26 |
+
- **Runs in the real FreeCAD app.** It drives the FreeCAD GUI over a local socket and builds parts live.
|
| 27 |
+
|
| 28 |
+
## Results
|
| 29 |
+
|
| 30 |
+

|
| 31 |
+
|
| 32 |
+
| Goal | Built correctly | With 20% random actions injected |
|
| 33 |
+
|---|---|---|
|
| 34 |
+
| Parts like the training set (1β5 features) | 100% | 93β100% |
|
| 35 |
+
| 6β7 features | 100% | 86% |
|
| 36 |
+
| 8β9 features | 100% | 87% |
|
| 37 |
+
| 11 features (~55 commands) | 100% | 90% |
|
| 38 |
+
| 13 / 15 / 17 features | 100 / 100 / 95% | β |
|
| 39 |
+
| Feature combinations never seen in training | 90β100% | 94β97% |
|
| 40 |
+
|
| 41 |
+
"Built correctly" means the model finished and the final solid matches the target exactly (volumetric IoU β₯ 0.99, no stray objects). Each row is 100 fresh goals in FreeCAD 1.1 (60 per length for 13β17 features). Per-step accuracy against the teacher's choices is 99.8%.
|
| 42 |
+
|
| 43 |
+
## What made it generalize
|
| 44 |
+
|
| 45 |
+

|
| 46 |
+
|
| 47 |
+
1. **Randomized position IDs during training** ([Ruoss et al. 2023](https://arxiv.org/abs/2305.16843)). Position numbers the model had never seen were what broke it on longer parts.
|
| 48 |
+
2. **Coupled ordinals.** Goal item *k* and the *k*-th feature in the tree share an index ([position coupling](https://arxiv.org/abs/2405.20671)).
|
| 49 |
+
3. **A modular "done?" policy.** Each goal item asks "am I built yet?", and the model acts on the first one that isn't. This is what made the longest goals reliable across training seeds.
|
| 50 |
+
4. **Factorized feature types.** Category embeddings plus type dropout help with feature pairings it hasn't seen.
|
| 51 |
+
|
| 52 |
+
## Usage
|
| 53 |
+
|
| 54 |
+
```python
|
| 55 |
+
from freecad_s1.model.net import from_pretrained
|
| 56 |
+
from freecad_s1.rollout import Policy
|
| 57 |
+
|
| 58 |
+
model = from_pretrained("shhivv/taiga-s1")
|
| 59 |
+
policy = Policy(model, device="cpu")
|
| 60 |
+
|
| 61 |
+
probs = policy.score(state, goal, actions) # {command: probability}, best first
|
| 62 |
+
next_command = next(iter(probs))
|
| 63 |
+
```
|
| 64 |
+
|
| 65 |
+
What to pass in:
|
| 66 |
+
- `state`: a snapshot of the FreeCAD session, from the included runtime (`freecad_s1.runtime`).
|
| 67 |
+
- `goal`: the ordered feature list, plus a rough size of the finished part (bounding box, volume). Estimates are fine.
|
| 68 |
+
- `actions`: the commands currently available in FreeCAD, as returned by the runtime's `valid_actions()`.
|
| 69 |
+
|
| 70 |
+
The returned probabilities are calibrated. The fitted temperature is stored in `config.json`.
|
| 71 |
+
|
| 72 |
+
## How it was trained
|
| 73 |
+
|
| 74 |
+
- **Data.** 24k synthetic modeling sessions scripted in headless FreeCAD, about 590k decisions. A scripted teacher labels the right next command at every step, and random mistakes are mixed in so the model also learns to recover.
|
| 75 |
+
- **Training.** Supervised training, then two rounds of DAgger: the model drives FreeCAD itself and the teacher corrects what it gets wrong.
|
| 76 |
+
- **Architecture.** A 3-layer transformer encodes the session state and the goal; candidate commands attend to the state and the active goal item, and each gets one score.
|
| 77 |
+
|
| 78 |
+
**Scope.** It covers FreeCAD PartDesign workflows: sketches (rectangle, circle, hexagon), pad, pocket, hole, revolve, linear and polar patterns, mirror, fillet, chamfer and shell. Taiga-S1 chooses the command; numeric values come from the goal.
|
| 79 |
+
|
| 80 |
+
## Citation
|
| 81 |
+
|
| 82 |
+
```bibtex
|
| 83 |
+
@misc{taiga_s1_2026,
|
| 84 |
+
title = {Taiga-S1: a small next-action model for CAD},
|
| 85 |
+
author = {Shanmugam, Shiv},
|
| 86 |
+
year = {2026}
|
| 87 |
+
}
|
| 88 |
+
```
|
config.json
ADDED
|
@@ -0,0 +1,51 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
{
|
| 2 |
+
"architecture": "S1Model",
|
| 3 |
+
"config": {
|
| 4 |
+
"width": 128,
|
| 5 |
+
"heads": 4,
|
| 6 |
+
"enc_layers": 3,
|
| 7 |
+
"dec_layers": 2,
|
| 8 |
+
"ff": 384,
|
| 9 |
+
"dropout": 0.1,
|
| 10 |
+
"id_dropout": 0.1,
|
| 11 |
+
"pos_mode": "rand",
|
| 12 |
+
"pos_table": 48,
|
| 13 |
+
"ordinal": true,
|
| 14 |
+
"ord_table": 12,
|
| 15 |
+
"progress_head": false,
|
| 16 |
+
"invariant_numerics": true,
|
| 17 |
+
"modular": true,
|
| 18 |
+
"pointer": "done",
|
| 19 |
+
"index_eval": "identity",
|
| 20 |
+
"type_dropout": 0.15,
|
| 21 |
+
"temperature": 2.554
|
| 22 |
+
},
|
| 23 |
+
"metadata": {
|
| 24 |
+
"name": "taiga-s1",
|
| 25 |
+
"version": "0.3.0",
|
| 26 |
+
"trained_on": "synthetic FreeCAD PartDesign episodes",
|
| 27 |
+
"calibration": {
|
| 28 |
+
"states": 11300,
|
| 29 |
+
"episodes": 256,
|
| 30 |
+
"T": 2.554,
|
| 31 |
+
"held_out_T1": {
|
| 32 |
+
"n": 5573,
|
| 33 |
+
"acc": 0.9551,
|
| 34 |
+
"nll": 0.3595,
|
| 35 |
+
"ece": 0.0416,
|
| 36 |
+
"mean_conf": 0.9921,
|
| 37 |
+
"mean_conf_when_wrong": 0.9705,
|
| 38 |
+
"n_wrong": 250
|
| 39 |
+
},
|
| 40 |
+
"held_out_T": {
|
| 41 |
+
"n": 5573,
|
| 42 |
+
"acc": 0.9551,
|
| 43 |
+
"nll": 0.1785,
|
| 44 |
+
"ece": 0.0272,
|
| 45 |
+
"mean_conf": 0.9545,
|
| 46 |
+
"mean_conf_when_wrong": 0.8915,
|
| 47 |
+
"n_wrong": 250
|
| 48 |
+
}
|
| 49 |
+
}
|
| 50 |
+
}
|
| 51 |
+
}
|
model.safetensors
ADDED
|
@@ -0,0 +1,3 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
version https://git-lfs.github.com/spec/v1
|
| 2 |
+
oid sha256:7ad63f77399eaddd3aebae0d0cf493340885d741a9fe51c1e4737f7273a1408c
|
| 3 |
+
size 4924516
|