Instructions to use open-nhe/Elysium-X-500-FR with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- PEFT
How to use open-nhe/Elysium-X-500-FR with PEFT:
from peft import PeftModel from transformers import AutoModelForCausalLM base_model = AutoModelForCausalLM.from_pretrained("Qwen/Qwen2.5-1.5B-Instruct") model = PeftModel.from_pretrained(base_model, "open-nhe/Elysium-X-500-FR") - Notebooks
- Google Colab
- Kaggle
π Elysium X 500 FR
A small multilingual model that reads a short conversation and says which of 500 emotion coordinates the last line expresses, with a strength (0.25, 0.5, 0.75 or 1.0). It outputs strict JSON like {"dimensions":[{"id":1,"strength":0.5}]}.
Related: GGUF Q8_0 build (not separately scored) Β· Dataset Β· Upgrade paper (Zenodo DOI 10.5281/zenodo.23179294) Β· Paper Space Β· Showcase Β· Code and guide
This release is the public-data upgrade of Elysium X 500 FR: the same 500-coordinate schema, with the weights continued-trained on the team dataset plus public emotion datasets mapped into the schema. It is a LoRA adapter (rank 32) on Qwen/Qwen2.5-1.5B-Instruct.
Licence: CC BY-NC-SA 4.0, non-commercial. This is a research release, not "open source" in the strict sense. Check the source-data licences below before any reuse.
This repository replaces the earlier internal checkpoint of Elysium X 500 FR (a mid-training adapter). Only the weights in this repo are the released model.
π How good is it? (honest numbers)
Private test set of 1,195 rows from the team dataset (cleaned split, zero text overlap with training). Exact match means the predicted labels and strengths all match.
| Test, exact match | Original X 500 FR | This release (run 1) | Repeat run (run 2) |
|---|---|---|---|
| All rows (n=1,195) | 69.54% | 70.71% | 69.87% |
| Emotion name in the text (n=792) | 95.08% | 94.82% | 93.94% |
| No emotion name (n=403), the hard test | 19.35% | 23.33% | 22.58% |
| Micro-F1, all rows | 70.27% | 71.78% | 71.76% |
| Micro-F1, no-name rows | 20.47% | 24.24% | not recorded |
What this means in plain words:
- It understands emotions without a giveaway word better than before (19.4% to 23.3%). That is the number to trust. It is still low: about 1 in 4.
- When the emotion word is in the sentence, it is right about 95% of the time. This number slipped by 2 answers out of 792 (751 vs 753) versus the original. It is a small miss, and we say so.
- The no-name score is far below the with-name score, so most "accuracy" on easy rows is pattern-matching on the emotion word.
We set a bar of beating all three original numbers. Run 1 beats two of three and misses with-name by 0.26 points. We publish it anyway on the project owner's decision.
π§ͺ About run 2 (disclosed on purpose)
We planned a "2 epoch" repeat. Because of a bug in our training loop it actually trained 1 epoch with a learning-rate schedule sized for 2, so the rate never decayed to zero. It scored slightly worse on all three checks, and we did not use it. We did not run a third time, because re-using the test set again would weaken the numbers.
π οΈ What it was trained on
- 9,380 training rows from the team's Elysium X 500 FR dataset (human-monitored, synthetic, 12 languages), cleaned and de-duplicated.
- 8,250 rows sampled from public emotion datasets and mapped into the 500 coordinates by hand-made rules (
public_data_mapping.csv). Public rows have no strengths, so strengths on those rows were masked out of the loss (BRIGHTER intensities excepted). - Continued from the earlier Elysium X 500 FR adapter for 1 epoch (1,102 steps, LR 1e-4), tuned on a dev split only. The test split was used once.
Same information as a table:
| Source dataset | Rows used | Licence (as listed by the source) |
|---|---|---|
| GoEmotions (Google Research) | 1,500 | Apache-2.0 |
| BRIGHTER, emotion categories | 1,100 | CC BY 4.0 |
| BRIGHTER, emotion intensities | 1,100 | CC BY 4.0 |
| EmpatheticDialogues (Meta) | 900 | CC BY-NC 4.0 |
| XED | 800 | CC BY 4.0 |
| SemEval-2018 Task 1 E-c | 700 | unknown on the source page |
| daily_dialog | 550 | CC BY-NC-SA 4.0 |
| ISEAR | 500 | not stated on the source page |
| dair-ai/emotion | 450 | "other" on the source page |
| tweet_eval (emotion) | 350 | unknown on the source page |
| Chinese Multi-Emotion Dialogue (Johnson8187) | 300 | MIT |
| Team dataset | 9,380 | CC BY-NC-SA 4.0 |
Rows for SemEval, ISEAR, dair-ai/emotion and tweet_eval come from sources whose licence is unknown or unstated, so this release is non-commercial research only. If you own one of these datasets and want it removed, tell us and the next version will drop it. We skipped MELD (GPL-3.0), gated sets and image-caption sets.
π Use
from transformers import AutoTokenizer, AutoModelForCausalLM
from peft import PeftModel
tok = AutoTokenizer.from_pretrained("Qwen/Qwen2.5-1.5B-Instruct")
base = AutoModelForCausalLM.from_pretrained("Qwen/Qwen2.5-1.5B-Instruct", torch_dtype="auto")
model = PeftModel.from_pretrained(base, "open-nhe/Elysium-X-500-FR")
Prompt with the conversation and the target line, and ask for JSON. The 500 coordinate ids and definitions are in taxonomy_500.csv. See the dataset repo for the row format.
β οΈ Known limits
- Weak without an emotion word. On no-name test rows the lowest exact-match scores are Hindi (3.8%), Bengali (7.1%), Spanish (16.4%) and Japanese (17.2%); the best are German (50%), Portuguese (45.5%) and English (39.5%). Some language cells hold few rows.
- The team data is templated and partly synthetic. Do not read these scores as real-world accuracy.
- It is not a medical or mental-health tool, and it must not be used to profile people.
π Credit
OpenNHE Technologies, Project NHE. Earlier paper: Elysium X 500 FR, https://doi.org/10.5281/zenodo.23159402
- Downloads last month
- 41