canbingol's picture
Update README.md
9c4b313 verified
|
Raw History Blame Contribute Delete
3.06 kB
metadata
license: apache-2.0
library_name: transformers
tags:
  - laya
  - system-one
  - calibrated-decisions
  - rlcd
  - structured-decisions
  - typed-decisions
  - benchmark
  - turkish
  - mmlu
metrics:
  - accuracy
  - brier_score
model-index:
  - name: laya-typed-decisions
    results:
      - task:
          type: text-classification
          name: System One Decision Benchmark
        dataset:
          type: LocalLLaMA/typed-decisions
          name: Typed Decisions
        metrics:
          - type: accuracy
            value: 0.365
          - type: brier_score
            value: 0.383
datasets:
  - canbingol/mmlu_typed_decision

Laya (Fine-Tuned on Turkish MMLU Typed-Decisions)

This is Laya fine-tuned on canbingol/mmlu_typed_decision, a 10k-example dataset built by converting the Turkish MMLU dataset into Laya's typed-decision format (state / questions / gold triples, choice-type questions with per-option criteria).

Note: the accuracy/Brier/ECE numbers below are from the original LocalLLaMA/typed-decisions benchmark (1,200 training cases / 400-case test set across Agent Trace Observability, Customer Service, Invoice Processing, and Security Incidents) and reflect that benchmark, not this fine-tune's performance on Turkish MMLU. They're kept here for reference to the base checkpoint's reported numbers.

On the official 400-case test set (2,000 decisions), the base checkpoint achieves 0.365 Accuracy, trailing TypeSafe Jev 1.13.0 (0.727) and the benchmark's Teacher Self-Agreement ceiling (0.735).

Head-to-Head Benchmark Results (base checkpoint, LocalLLaMA/typed-decisions)

Model Kind Accuracy Soft Acc Brier Score ECE Score MAE Within 1 Level Latency (p50) Cost/Case
Turkish Laya fine-tuned 0.365 0.354 0.383 0.242 0.726 0.703 168.3 ms $0.00 (Self-Hosted)
TypeSafe Jev 1.13.0 general 0.727 0.580 0.148 0.144 0.391 0.952 710 ms $0.0004 (API)
ModernBERT-base (149M) specialist 0.646 0.542 0.119 0.179 0.444 0.931 349 ms $0.00
Teacher Self-Agreement ceiling 0.735 - - - - - - -

Training Data

Fine-tuned on canbingol/mmlu_typed_decision — ~10,000 examples derived from the Turkish MMLU dataset, reformatted as choice-type typed decisions (question → instructions, answer options → criteria, correct option → one-hot target).

Installation & Quickstart

pip install laya
import laya

# Load the fine-tuned model directly from Hugging Face
agent = laya.load("convaiinnovations/laya-typed-decisions")

# Evaluate any workflow state and typed questions in a single forward pass
result = agent.predict(state, questions)
print(result["answers"])

License

Apache 2.0. Developed by Convai Innovations.