--- title: Laya-Hebrew Decision Model emoji: 🎯 colorFrom: blue colorTo: pink sdk: gradio sdk_version: 6.29.0 app_file: app.py pinned: false license: cc-by-nc-sa-4.0 short_description: Calibrated Hebrew & English decisions python_version: "3.12" startup_duration_timeout: 1h --- # Laya-Hebrew — calibrated Hebrew & English decisions [`RoeiG/laya-hebrew`](https://huggingface.co/RoeiG/laya-hebrew) is a **decision** model, not a chatbot. Give it a **state** (the text or fields to judge) and **typed questions**, and one forward pass returns a **calibrated probability for every answer**. It never generates text, so there is nothing to parse and nothing to hallucinate. | type | question | answer | |---|---|---| | `choice` | which of these options? | the option, a probability per option, confidence | | `score` | where on this rubric? | the expected level, a probability per level, confidence | | `noul` | is this claim true? | P(true) | 378M parameters (363M NeoBERT encoder + 15M head), fp16, 1,024 tokens per question. Hebrew and English only. ## How to get good answers 1. **Compute numbers, dates, units and relations in code** and pass the result as a field (the *Advanced options* box). The model does not do arithmetic reliably — and it is often *confidently* wrong: 2.5 hours against a 2-hour limit, or 19 days against a 14-day window, got P(true) of 0.93–0.98. 2. **Prefer the claim form for yes/no.** "הלקוח כועס." discriminates better than "האם הלקוח כועס?" (gap 0.67 against 0.38 on he_bench). 3. **Describe every option in a line.** Bare labels or codes route much worse than labels with a one-line description. 4. **Read the probabilities, not only the top answer.** A top answer below ~0.6 means the model is unsure — and a confidence threshold will *not* catch the failures listed below, which are typically high-confidence. 5. **Ignore `act_probability`.** It comes from a head that no Hebrew checkpoint trained; it is left in the raw JSON and hidden from the rendered answers. ## Limitations - Reasoning is the weak spot: hellaswag 0.447, winograd 0.607, copa 0.687. Multi-step inferences ("A is taller than B, B is taller than C: who is shortest?") often fail. - Numeric, date and unit rules are unreliable and often confidently wrong. Compute them in code. - Irony and sarcasm are read literally ("וואו, שירות מדהים…" is scored as positive, at 0.97). - Routing leans on keywords: a strong keyword can outweigh the option descriptions. - Scales are weak: sentiment and tone are ~0.45 accuracy, and tone probabilities are overconfident (valence Brier 0.53). Held-out evaluation: MASSIVE-he 20 intents 0.806, he_bench accuracy 0.595 (Brier 0.499, ECE 0.125), Hebrew BoolQ 0.838, Belebele-he 0.767 (chance 0.25). ## Implementation notes This Space vendors **Laya 0.3.7** at commit `010bace` in `laya/`, with [`neobert.patch`](https://huggingface.co/RoeiG/laya-hebrew/blob/main/neobert.patch) applied: - NeoBERT's modeling code lives on the Hub, so `trust_remote_code=True` is passed through. - NeoBERT computes its rotary tables as non-persistent buffers in `__init__`; under transformers 5 the model is built on the meta device, so they come back as uninitialised memory and **every forward pass returns NaN**. The patch recomputes them. - The encoder config pins `bfloat16`; the patch keeps the encoder in **fp32**. Attention is SDPA, the model is loaded once at module scope and runs on ZeroGPU (`@spaces.GPU`). ## License **CC-BY-NC-SA-4.0 — non-commercial use only.** Several training sources are non-commercial or research-only (ANLI, Yelp, AG News, Yahoo Answers, RACE, and the NLLB translation model). - Encoder: [`dicta-il/neodictabert-bilingual`](https://huggingface.co/dicta-il/neodictabert-bilingual) (NeoBERT, by Dicta) — CC-BY-4.0. - Laya architecture, training method and runtime: [Laya](https://github.com/NandhaKishorM/laya) (Apache-2.0). No weights come from Laya's published checkpoints.