Text Generation
PEFT
Safetensors
emotion
multilingual
lora
qwen2
emotion-classification

Elysium X 500 FR

🌌 Elysium X 500 FR

A small multilingual model that reads a short conversation and says which of 500 emotion coordinates the last line expresses, with a strength (0.25, 0.5, 0.75 or 1.0). It outputs strict JSON like {"dimensions":[{"id":1,"strength":0.5}]}.

GGUF Dataset Paper Showcase GitHub Licence

Related: GGUF Q8_0 build (not separately scored) Β· Dataset Β· Upgrade paper (Zenodo DOI 10.5281/zenodo.23179294) Β· Paper Space Β· Showcase Β· Code and guide

This release is the public-data upgrade of Elysium X 500 FR: the same 500-coordinate schema, with the weights continued-trained on the team dataset plus public emotion datasets mapped into the schema. It is a LoRA adapter (rank 32) on Qwen/Qwen2.5-1.5B-Instruct.

Licence: CC BY-NC-SA 4.0, non-commercial. This is a research release, not "open source" in the strict sense. Check the source-data licences below before any reuse.

This repository replaces the earlier internal checkpoint of Elysium X 500 FR (a mid-training adapter). Only the weights in this repo are the released model.

πŸ“Š How good is it? (honest numbers)

Exact-match results

Private test set of 1,195 rows from the team dataset (cleaned split, zero text overlap with training). Exact match means the predicted labels and strengths all match.

Test, exact match Original X 500 FR This release (run 1) Repeat run (run 2)
All rows (n=1,195) 69.54% 70.71% 69.87%
Emotion name in the text (n=792) 95.08% 94.82% 93.94%
No emotion name (n=403), the hard test 19.35% 23.33% 22.58%
Micro-F1, all rows 70.27% 71.78% 71.76%
Micro-F1, no-name rows 20.47% 24.24% not recorded

What this means in plain words:

  • It understands emotions without a giveaway word better than before (19.4% to 23.3%). That is the number to trust. It is still low: about 1 in 4.
  • When the emotion word is in the sentence, it is right about 95% of the time. This number slipped by 2 answers out of 792 (751 vs 753) versus the original. It is a small miss, and we say so.
  • The no-name score is far below the with-name score, so most "accuracy" on easy rows is pattern-matching on the emotion word.

We set a bar of beating all three original numbers. Run 1 beats two of three and misses with-name by 0.26 points. We publish it anyway on the project owner's decision.

πŸ§ͺ About run 2 (disclosed on purpose)

We planned a "2 epoch" repeat. Because of a bug in our training loop it actually trained 1 epoch with a learning-rate schedule sized for 2, so the rate never decayed to zero. It scored slightly worse on all three checks, and we did not use it. We did not run a third time, because re-using the test set again would weaken the numbers.

πŸ› οΈ What it was trained on

How it was built

  • 9,380 training rows from the team's Elysium X 500 FR dataset (human-monitored, synthetic, 12 languages), cleaned and de-duplicated.
  • 8,250 rows sampled from public emotion datasets and mapped into the 500 coordinates by hand-made rules (public_data_mapping.csv). Public rows have no strengths, so strengths on those rows were masked out of the loss (BRIGHTER intensities excepted).
  • Continued from the earlier Elysium X 500 FR adapter for 1 epoch (1,102 steps, LR 1e-4), tuned on a dev split only. The test split was used once.

Training sources and licences

Same information as a table:

Source dataset Rows used Licence (as listed by the source)
GoEmotions (Google Research) 1,500 Apache-2.0
BRIGHTER, emotion categories 1,100 CC BY 4.0
BRIGHTER, emotion intensities 1,100 CC BY 4.0
EmpatheticDialogues (Meta) 900 CC BY-NC 4.0
XED 800 CC BY 4.0
SemEval-2018 Task 1 E-c 700 unknown on the source page
daily_dialog 550 CC BY-NC-SA 4.0
ISEAR 500 not stated on the source page
dair-ai/emotion 450 "other" on the source page
tweet_eval (emotion) 350 unknown on the source page
Chinese Multi-Emotion Dialogue (Johnson8187) 300 MIT
Team dataset 9,380 CC BY-NC-SA 4.0

Rows for SemEval, ISEAR, dair-ai/emotion and tweet_eval come from sources whose licence is unknown or unstated, so this release is non-commercial research only. If you own one of these datasets and want it removed, tell us and the next version will drop it. We skipped MELD (GPL-3.0), gated sets and image-caption sets.

πŸš€ Use

from transformers import AutoTokenizer, AutoModelForCausalLM
from peft import PeftModel
tok = AutoTokenizer.from_pretrained("Qwen/Qwen2.5-1.5B-Instruct")
base = AutoModelForCausalLM.from_pretrained("Qwen/Qwen2.5-1.5B-Instruct", torch_dtype="auto")
model = PeftModel.from_pretrained(base, "open-nhe/Elysium-X-500-FR")

Prompt with the conversation and the target line, and ask for JSON. The 500 coordinate ids and definitions are in taxonomy_500.csv. See the dataset repo for the row format.

⚠️ Known limits

No-name exact match by language

  • Weak without an emotion word. On no-name test rows the lowest exact-match scores are Hindi (3.8%), Bengali (7.1%), Spanish (16.4%) and Japanese (17.2%); the best are German (50%), Portuguese (45.5%) and English (39.5%). Some language cells hold few rows.
  • The team data is templated and partly synthetic. Do not read these scores as real-world accuracy.
  • It is not a medical or mental-health tool, and it must not be used to profile people.

πŸ’œ Credit

OpenNHE Technologies, Project NHE. Earlier paper: Elysium X 500 FR, https://doi.org/10.5281/zenodo.23159402

Downloads last month
41
Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support

Model tree for open-nhe/Elysium-X-500-FR

Adapter
(1522)
this model
Quantizations
1 model

Dataset used to train open-nhe/Elysium-X-500-FR

Spaces using open-nhe/Elysium-X-500-FR 3

Collection including open-nhe/Elysium-X-500-FR