--- license: apache-2.0 library_name: transformers tags: - laya - system-one - calibrated-decisions - rlcd - structured-decisions - typed-decisions - benchmark - turkish - mmlu metrics: - accuracy - brier_score model-index: - name: laya-typed-decisions results: - task: type: text-classification name: System One Decision Benchmark dataset: type: LocalLLaMA/typed-decisions name: Typed Decisions metrics: - type: accuracy value: 0.365 - type: brier_score value: 0.383 datasets: - canbingol/mmlu_typed_decision --- # Laya (Fine-Tuned on Turkish MMLU Typed-Decisions) This is **Laya** fine-tuned on [canbingol/mmlu_typed_decision](https://huggingface.co/datasets/canbingol/mmlu_typed_decision), a 10k-example dataset built by converting the Turkish MMLU dataset into Laya's typed-decision format (state / questions / gold triples, `choice`-type questions with per-option criteria). Note: the accuracy/Brier/ECE numbers below are from the original [LocalLLaMA/typed-decisions](https://huggingface.co/datasets/LocalLLaMA/typed-decisions) benchmark (1,200 training cases / 400-case test set across Agent Trace Observability, Customer Service, Invoice Processing, and Security Incidents) and reflect that benchmark, **not** this fine-tune's performance on Turkish MMLU. They're kept here for reference to the base checkpoint's reported numbers. On the official 400-case test set (2,000 decisions), the base checkpoint achieves **0.365 Accuracy**, trailing **TypeSafe Jev 1.13.0 (0.727)** and the benchmark's Teacher Self-Agreement ceiling (0.735). ## Head-to-Head Benchmark Results (base checkpoint, LocalLLaMA/typed-decisions) | Model | Kind | Accuracy | Soft Acc | Brier Score | ECE | Score MAE | Within 1 Level | Latency (p50) | Cost/Case | |---|---|---|---|---|---|---|---|---|---| | **Turkish Laya** | **fine-tuned** | **0.365** | **0.354** | **0.383** | **0.242** | **0.726** | **0.703** | **168.3 ms** | **$0.00 (Self-Hosted)** | | TypeSafe Jev 1.13.0 | general | 0.727 | 0.580 | 0.148 | 0.144 | 0.391 | 0.952 | 710 ms | $0.0004 (API) | | ModernBERT-base (149M) | specialist | 0.646 | 0.542 | 0.119 | 0.179 | 0.444 | 0.931 | 349 ms | $0.00 | | Teacher Self-Agreement | ceiling | 0.735 | - | - | - | - | - | - | - | ## Training Data Fine-tuned on [canbingol/mmlu_typed_decision](https://huggingface.co/datasets/canbingol/mmlu_typed_decision) — ~10,000 examples derived from the Turkish MMLU dataset, reformatted as `choice`-type typed decisions (question → `instructions`, answer options → `criteria`, correct option → one-hot `target`). ## Installation & Quickstart ```bash pip install laya ``` ```python import laya # Load the fine-tuned model directly from Hugging Face agent = laya.load("convaiinnovations/laya-typed-decisions") # Evaluate any workflow state and typed questions in a single forward pass result = agent.predict(state, questions) print(result["answers"]) ``` ## License Apache 2.0. Developed by [Convai Innovations](https://huggingface.co/convaiinnovations).