Laya-Hebrew-V2.1: a Hebrew–English decision model

A Laya-style decision model for Hebrew and English. You give it a state (the text or fields to judge) and questions: a choice between options, a score on a scale, or a yes/no claim. In one forward pass it returns a calibrated probability for every answer. It is a fast classifier that you configure at call time. It is not a chatbot and it does not generate text.

V2.1 continues Laya-Hebrew-V2 with one more training run. What changed:

  • Built for incoming messages. Emails, WhatsApp, SMS and tickets: message type, routing to a department, urgency, "does this need action from this role?", relevance to a role on a rubric you write, and claims about the message. On a sealed test of 2,915 such questions, written by a model that wrote none of the training data: 0.772 → 0.899. Relevance to a role at the exact level: 0.18 → 0.79 (see Messages).
  • Yes/no answers described in words now work. criteria such as {"true": "כן, ...", "false": "לא, ... גם אם ..."} used to make V2 give the same answer to every text; V2.1 reads them (consistency 0.47 → 0.96, see Wording).
  • "The passage does not say" is false. A plausible answer the passage does not give is now rejected (HeQ: see the comparison below).
  • "None of the options" is trained, for routing with an "other" bucket.
  • Voice-command intents (MASSIVE) are trained.
  • Temperatures per number of options (2, 3–5, 6–10 for a choice), on top of the per-type ones.

Laya-Hebrew-V2 stays at RoeiG/Laya-Hebrew-V2 (tag 2.0), unchanged.

  • Encoder: dicta-il/neodictabert-bilingual (NeoBERT, 28 layers, Hebrew + English)
  • Head: Laya's DecisionModel architecture (6 transformer layers and a scorer over the option markers), continued from V2's. No weights come from Laya's published checkpoints; the only pretrained weights are the encoder's.
  • Size: 406M parameters (encoder 363M, head 43M), stored in fp16 (812 MB)
  • Speed: the same as V2, about 0.13 s per question for a short message on an Apple M1 CPU (4 threads)
  • Input: 1,024 tokens in total, per question
    • The instructions and options share 256 of those tokens, and each option is cut at 48 tokens.
    • The state gets the rest. A longer state is cut from the end without a warning, so put what matters first.
  • Training: Laya's RLCD objective (proper-scoring-rule rewards plus soft cross-entropy), with temperatures per question type and per number of options fitted on held-out items
  • Use: non-commercial only (see License)

Compared with other Hebrew Laya models

Accuracy on identical questions: Laya-Hebrew v1, V2, V2.1 and Nitzotz

Test (questions) Laya-Hebrew (v1) Laya-Hebrew-V2 Laya-Hebrew-V2.1 Nitzotz V2.1 vs Nitzotz
Message triage, sealed test (2,915) 72.9% 77.2% 89.9% 63.0% better, p<0.001
MASSIVE intent, 20 options * (500) 79.4% 79.8% 88.0% 15.2% better, p<0.001
MASSIVE intent, 4 options * (500) 92.6% 92.4% 95.2% 74.4% better, p<0.001
SIB-200 news topic, 7 options (204) 79.9% 82.8% 83.3% 80.9% tie, p=0.44
Belebele reading, 4 options (900) 76.7% 77.3% 78.8% 63.0% better, p<0.001
HeQ: does the passage support the answer (600) 86.8% 81.0% 87.7% 71.8% better, p<0.001
HeQ: plausible answer not in the passage (233) 58.8% 63.9% 91.4% 98.7% worse, p<0.001
“None of the options” added * (400) 62.7% 59.5% 75.2% 55.8% better, p<0.001

* Asked in English with English intent names (Laya's standard MASSIVE format). Nitzotz's card describes it as Hebrew-only.

How this was measured:

  • One harness, identical questions. Every model answered every question in the table on the same CPU, scored by the same code. The last column is an exact McNemar test on the paired answers: "better" or "worse" means p < 0.05.
  • Nitzotz (BrainboxAI, Apache-2.0) is another Laya model for Hebrew. Its card compares itself with Laya-Hebrew v1 on these suites. Its questions are not published, so the suites were rebuilt from the public test splits as its card describes them. Its in-house message tests are not published either; the message test here is a different one.
  • Message triage: 595 Israeli-style messages (emails, WhatsApp, SMS, tickets) written by Google's Gemini, labelled twice, frozen before V2.1's training data existed. Half of it was sealed and scored once, for V2.1's release decision. V2.1 trained on messages written by Claude, and Nitzotz on messages written by DeepSeek, so the test is in-domain for both, and neither saw its texts.
  • Training overlap: V2.1 and Nitzotz trained on MASSIVE's training split; Laya-Hebrew and V2 never did. All four trained on HeQ's training split in some form. Only test splits are used here.
  • * The rows that mostly measure the question's format for Nitzotz. MASSIVE and "none of the options" are asked as in Laya's own MASSIVE evaluation: an English question and English intent names over the Hebrew utterance. Laya models train with that format; Nitzotz's card describes it as Hebrew-only and reports 90.0% (20 options) and 97.2% (4 options) with its own Hebrew wording, against 15.2% and 74.4% here. With 20 English options its question also runs past its 256-token limit for the question and options (280 tokens), so some options are cut.
  • HeQ wording. Both HeQ rows ask "האם הקטע תומך בתשובה ... לשאלה ...?". With this wording Nitzotz answers "no" more often than with its own: its card reports 94.8% (supports) and 89.2% (plausible but not given), against 71.8% and 98.7% here. Nitzotz's card notes that the wording of a question changes its answers.
  • "None of the options": MASSIVE and Belebele test questions with a "none of the options" choice added: half with the right answer removed ("none" is right), half with a wrong option removed.
  • SIB-200 and Belebele are asked in Hebrew, and Nitzotz's scores here match its own card (80.9% and 63.0%, against its reported 79.4% and 63.2%).

Usage

This checkpoint runs on Laya v0.4.0 (recommended) or v0.3.20, each with its patch from this repository: neobert_v0.4.0.patch or neobert.patch. The patch loads NeoBERT's remote code, recomputes its rotary tables (without that, every output is NaN under transformers 5) and keeps the encoder in fp32. The two versions give identical probabilities (checked on 100 answers across every question type: no difference). The per-option-count temperatures in rl_agent_config.json are read by both.

git clone https://github.com/NandhaKishorM/laya && cd laya && git checkout v0.4.0
git apply /path/to/neobert_v0.4.0.patch   # from this repository (for v0.3.20: neobert.patch)
pip install -e .                          # torch, transformers 5.x

On Apple Silicon use device="cpu", or set LAYA_MPS_AMP_MIN_ROWS=100000 before using "mps" (NeoBERT in fp16 on MPS crashes).

import laya

agent = laya.load("RoeiG/Laya-Hebrew-V2.1", device="cpu")   # or "cuda"
out = agent.predict(
    {"message": "האפליקציה קורסת כשאני פותח את המצלמה"},
    {
        "team": {"type": "choice", "instructions": "איזה צוות צריך לטפל בהודעה?",
                 "criteria": {"billing": "תשלומים והחזרים", "tech": "באגים וקריסות", "shipping": "משלוחים"}},
        "upset": {"type": "noul", "instructions": "הלקוח כועס."},
    },
)
# out["answers"]["team"]  -> {"choice": "tech", "probabilities": {"billing": 0.0059, "tech": 0.99, "shipping": 0.004}, ...}
# out["answers"]["upset"] -> {"noul": 0.0216, ...}   (P(true))

Question types:

  • choice: criteria maps each option to a description.
  • score: criteria is a list of levels from lowest to highest. The answer includes score, the expected level (0 = the first level), and a probability for every level.
  • noul: yes/no. The answer is P(true) for the statement in instructions. Optional criteria describe what true and false mean; optional labels, for example {"false": "לא", "true": "כן"}, rename the two options.

Relevance as a scale

doc = ("הכנרת היא אגם המים המתוקים הגדול בישראל. שטחה כ-166 קמ\"ר ועומקה המרבי כ-43 מטרים. "
       "מקורות המים העיקריים שלה הם נהר הירדן ונחלים מהגולן ומהגליל.")
rubric = ["לא רלוונטי בכלל", "קשור לנושא בלבד", "רלוונטי חלקית", "רלוונטי ברובו", "רלוונטי לגמרי"]
queries = ["מה שטחה של הכנרת?", "מה שטחה של הכנרת ומה מפלס המים בה היום?",
           "מתי נחנך המוביל הארצי מהכנרת?", "איך מכינים פיצה ביתית?"]
for query in queries:
    out = agent.predict({"document": doc, "query": query},
                        {"rel": {"type": "score", "instructions": "עד כמה ה-`document` עונה על ה-`query`?", "criteria": rubric}})
    level = out["answers"]["rel"]["score"] + 1   # 1-5
Query What it asks of the passage Expected level (1–5) Most likely level
מה שטחה של הכנרת? answers it fully 4.6 5 (רלוונטי לגמרי)
מה שטחה של הכנרת ומה מפלס המים בה היום? answers part of it 3.5 3 (רלוונטי חלקית)
מתי נחנך המוביל הארצי מהכנרת? on the topic, no answer 2.0 2 (קשור לנושא בלבד)
איך מכינים פיצה ביתית? unrelated 1.0 1 (לא רלוונטי בכלל)

Plain wordings count a partial answer as relevant; say "fully" when you need a full answer. For the partial query above:

Wording P(true)
האם ה-document עונה על ה-query? 0.97
האם ה-document עונה על ה-query באופן מלא? 0.17

How to get good answers

  1. Compute numbers, dates, units and relations in code. Pass the result as a field, such as "age_ok": "הגיל עומד בתנאי". The model does not do arithmetic reliably (see Limitations).
  2. State yes/no claims positively. Write "הלקוח מרוצה" and read P(false), rather than "הלקוח לא מרוצה": negated claims are still the weakest wording (see Wording).
  3. Prefer the claim form for yes/no. "הלקוח כועס." discriminates better than "האם הלקוח כועס?" (gap 0.75 against 0.56 on he_bench).
  4. Describe every option in a line, and every level of a scale. Bare labels or codes work worse than labels with a one-line description. Describing what true and false mean in a yes/no question is fine.
  5. Say how strict a relevance question is. "האם ה-document עונה על ה-query?" accepts a partial answer; "האם ה-document עונה על ה-query באופן מלא?" asks for a full one.
  6. Rank with the expected level; choose cutoffs on your own examples. On messages, relevance to a role now lands on the labelled level 79% of the time, urgency 48% (and within one level 89% and 93%).
  7. Read the probabilities, not only the top answer. A top answer below about 0.6 means the model is unsure.
  8. Check combined answers. Several questions about one message are answered independently; on about 5% of messages two answers disagree (for example "urgent" and "no action needed"). If your logic combines them, decide which wins.
  9. Ignore act_probability. It comes from a head that no Hebrew checkpoint trained.

Evaluation

V2.1 was released under a rule fixed before it was trained: every one of V2's guard metrics within the run-to-run noise of V2 or better, and at least one new target met. It met the message target on the sealed half (McNemar p < 1e-50: on 474 questions only V2.1 was right, on 102 only V2). The wording target (every wording variant answered the same way at least 95% of the time) was not met; see Wording.

Both models below were predicted on the same Apple M1 CPU, so V2's numbers can differ from its own card in the third decimal. Bold marks the better of the two values in each row; equal values are not bold. Higher is better unless a row says lower is better.

Metric Laya-Hebrew-V2 Laya-Hebrew-V2.1
MASSIVE he, 20 intents (500) 0.804 0.884
MASSIVE he, 4 intents 0.924 0.954
MASSIVE en, 20 intents 0.818 0.876
MASSIVE he, 3 scenarios kept out of V2.1's training (530) 0.815 0.847
he_bench accuracy (4,290) 0.659 0.664
he_bench Brier (lower is better) 0.435 0.443
copa / hellaswag / winograd (he_bench, 1,800) 0.663 0.677
Belebele-he reading comprehension (900; chance 0.25) 0.773 0.788
SIB-200-he topic (1,004) 0.819 0.809
Yes/no as a question: P(yes | true) − P(yes | false) 0.537 0.556
Yes/no as a claim: accuracy (390) 0.905 0.897
Yes/no as a claim: same gap 0.748 0.749
Relevance as a claim (192) 0.828 0.804
Hebrew BoolQ, held out (875) 0.854 0.850
Rule-direction probe (208) 0.837 0.837
typed-decisions test (2,000) 0.793 0.795
Held-out soft-label set (900), Brier (lower is better) 0.229 0.227
Hand-written Hebrew cases (32), correct 31 31

MASSIVE's training split is now in the training data, so its first three rows are no longer a transfer test. The fourth row is: three whole scenarios (dates and times, email, weather) were kept out of training.

he_bench accuracy by task:

Task Laya-Hebrew-V2 Laya-Hebrew-V2.1 Chance
relevance (question form) 0.845 0.837 0.50
qa_verify (question form) 0.759 0.781 0.50
copa 0.820 0.828 0.50
sentiment 0.650 0.653 0.33
winograd 0.670 0.657 0.50
hellaswag 0.500 0.545 0.25
tone arousal 0.419 0.439 0.21
tone valence 0.431 0.372 0.23

Relevance as a scale (rubrics and instructions that never appeared in training):

Test Laya-Hebrew-V2 Laya-Hebrew-V2.1
he_bench's FiQA relevance asked on 3 rubrics (576): the top level on the right side of relevant / not relevant 0.835 0.859
TREC Deep Learning 2019 + 2020, human 4-level grades (953 pairs × 3 rubrics, machine-translated to Hebrew): exact level 0.286 0.301
the same: distance from TREC's level (lower is better) 1.058 1.070

Messages

The sealed half of the message test (see the comparison): 2,915 questions on about 300 messages. Urgency and relevance use 5-level rubrics; "pooled" counts them right within one level.

Question (sealed half) Laya-Hebrew-V2 Laya-Hebrew-V2.1
All questions pooled (2,915) 0.772 0.899
Message type 0.893 0.941
Department (routing) 0.894 0.928
Needs action from a role (yes/no) 0.828 0.924
Claims about the message (yes/no) 0.801 0.867
Urgency, 5 levels: within one level 0.801 0.934
Urgency: exact level 0.294 0.483
Urgency: ordering (γ) 0.694 0.865
Relevance to a role, your rubric: within one level 0.646 0.892
Relevance: exact level 0.176 0.794
Relevance: ordering (γ) 0.436 0.828
Answers that contradict each other (lower is better) 0.024 0.052

γ is Goodman–Kruskal's gamma: 0 = random order, 1 = perfect.

Wording

The same questions asked in other wordings, each with the same right answer (a small board: 20–110 questions per variant). "Consistency" is the share answered as in the plain wording.

Variant Laya-Hebrew-V2 Laya-Hebrew-V2.1
Yes/no with true and false described in words: consistency 0.471 0.957
the same, with an "even if" clause: consistency 0.471 0.929
the same, on an empty text: P(true) (lower is better) 0.993 0.010
The claim negated: accuracy 0.114 0.329
Scale levels reworded (numeric, English, 3 or 4 levels): mean accuracy 0.288 0.575
Options given as codes with descriptions: accuracy 0.775 0.725
More options than usual: accuracy 0.700 0.733
The question after other passages: accuracy 0.711 0.756
Hebrew spelling, punctuation and direction marks: accuracy 0.819 0.902

Run-to-run noise. Two training seeds of Laya-Hebrew, with the same data and settings, differed by 0.2 points on he_bench overall, 0.6–1.0 on Belebele and MASSIVE, 1–3 on single he_bench tasks and 5–7 on each half of the rule probe. Differences smaller than these are noise.

Limitations

  • Some wordings still change the answer. Negated claims (0.33 against 0.89 for the same claims stated plainly), scales whose levels are reworded, and options given as codes. State claims positively and describe options in words.
  • Answers to different questions about one message can disagree (5% of messages, above).
  • The message tests and training messages were written by AI models (Gemini and Claude). Real inboxes are messier.
  • Reasoning is still the weak spot. Hellaswag (0.55) and winograd (0.66) are well above chance but far from solved.
  • Numeric, date and unit rules are unreliable, and can be confidently wrong. In the hand-written cases, a purchase two months ago is still accepted against a 14-day return window. Compute such checks in code.
  • Subjective scales stay hard: sentiment 0.65, tone 0.37–0.44 on 5-level scales.
  • Document–query relevance as a yes/no claim is 0.80; the scale form is the stronger one (FiQA 0.86).
  • Calibration does not catch everything. The failures above are often high-confidence, so a confidence threshold will not filter them out.

Training data

V2.1 continues V2 in one run: 248,334 items (3,881 updates, 1.07 A100 hours), one epoch, encoder 5e-6 and head 1e-4 (cosine to 1e-6), batch 64. A third of the items are new, two thirds replay V2's training data (90% from its main run, 10% from its relevance-scale run).

  • Messages (32,802 cases from 3,598 messages): emails, WhatsApp, SMS, tickets and chats for 24 invented organisations (157 roles), in Hebrew with some English, code-switching, informal writing and typos; quoted threads in several layouts. Written and labelled in Claude (Anthropic) sessions, then labelled again by a separate Claude session that saw only the organisation, the channel and the message. A label was kept only when both passes agreed: exactly for type, routing, needs-action and claims (agreement 83%, 94%, 98% and 98%), within one grade for urgency and relevance (over 99%). Questions: message type, department, urgency under a rubric drawn per case, needs-action, relevance to one of two roles on a rubric, and claims. A fifth are contrast pairs, a minimal edit that flips one label. Scam messages are 3.6% of them.
  • Other wordings of V2's own items (28,000): items from V2's training files asked again with true and false described in words (a third with an "even if" clause; one in ten on an empty text, whose answer is false), negated, with options as codes or bare labels, reordered or added options, numeric rubrics, instructions in the other language, flattened or extra fields, Hebrew spelling and punctuation variants, and a passage after other passages. No new text was written; the answers carry over, flipped for a negation.
  • MASSIVE he-IL (8,000): training and dev splits, without the three scenarios held out for the transfer check.
  • HeQ (8,988): plausible answers that the passage does not give, labelled false, from the training split.
  • "None of the options" (4,988): existing choice items with the right answer removed and "none" added, or a wrong option removed so that "none" is wrong.

Calibration: temperatures per question type and per number of options, fitted on 2,052 items that no run trained on: V2's hold-out of 1,920 items and 132 scale questions on held-out messages.

Decontamination: every new item was checked against every evaluation set, both halves of the message test included (8-word runs, instructions, rubrics); no training message shares an 8-word run with the message test. MASSIVE's and HeQ's test splits were never trained on. The training messages are not published. No private data and no personal data were used.

License

CC-BY-NC-SA-4.0: non-commercial use only, as for V2. Several training sources are non-commercial or research-only: ANLI, Yelp, AG News, Yahoo Answers, RACE, and the NLLB translation model (CC-BY-NC-4.0). Others are ShareAlike: Wikipedia, BEIR-he, SNLI and BoolQ. MASSIVE and HeQ are CC BY 4.0. The training messages were written with Claude (Anthropic); the message test with Gemini (Google), for evaluation only. The encoder is by Dicta (neodictabert-bilingual, CC-BY-4.0). The Laya architecture, training method and runtime are by Laya's authors (Apache-2.0).

Acknowledgements

  • Dicta, for NeoDictaBERT-bilingual and DictaLM 3.0
  • Laya's authors, for the architecture, the RLCD training method and the runtime
  • Amazon, for MASSIVE; Webiks, MAFAT and the Israeli National NLP Program, for HeQ
  • BrainboxAI, for publishing Nitzotz openly, which made a comparison on identical questions possible
  • The creators of every dataset listed above, and of TREC Deep Learning and FiQA, used for evaluation
Downloads last month
20
Safetensors
Model size
0.4B params
Tensor type
F16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for RoeiG/Laya-Hebrew-V2.1

Finetuned
(1)
this model

Datasets used to train RoeiG/Laya-Hebrew-V2.1

Collection including RoeiG/Laya-Hebrew-V2.1