Text Classification
Transformers
Safetensors
English
HHEMv2Config
hallucination-detection
factual-consistency
rag
custom_code
Instructions to use vectara/hallucination_evaluation_model with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use vectara/hallucination_evaluation_model with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-classification", model="vectara/hallucination_evaluation_model", trust_remote_code=True)# pip install -U transformers accelerate # Load model directly from transformers import AutoModelForSequenceClassification model = AutoModelForSequenceClassification.from_pretrained("vectara/hallucination_evaluation_model", trust_remote_code=True, device_map="auto") - Notebooks
- Google Colab
- Kaggle
Commit Β·
de9a5e6
1
Parent(s): 8e4a2e6
Support transformers 5 and refresh model card
Browse filesCo-Authored-By: Claude Opus 5.5 <[email protected]>
- README.md +15 -10
- modeling_hhem_v2.py +1 -0
README.md
CHANGED
|
@@ -2,7 +2,12 @@
|
|
| 2 |
language: en
|
| 3 |
license: apache-2.0
|
| 4 |
base_model: google/flan-t5-base
|
| 5 |
-
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 6 |
|
| 7 |
extra_gated_fields:
|
| 8 |
Company: text
|
|
@@ -15,11 +20,11 @@ extra_gated_fields:
|
|
| 15 |
|
| 16 |
## Quickstart: try HHEM-2.1-Open Live Demo
|
| 17 |
|
| 18 |
-
π **[Launch Interactive Demo](https://huggingface.co/spaces/vectara/hhem-2.1-open-demo)** - No setup required, runs in your browser
|
| 19 |
|
| 20 |
-
|
| 21 |
|
| 22 |
-
π‘ **Quick test**: Try
|
| 23 |
|
| 24 |
HHEM-2.1-Open is the latest open source version of [Vectara](https://vectara.com)'s HHEM series models for detecting hallucinations in LLMs. These are particularly useful in the context of building retrieval-augmented-generation (RAG) applications or Agentic workflows, where a set of facts is summarized by an LLM, and HHEM can be used to measure the extent to which this summary is factually consistent with the facts.
|
| 25 |
|
|
@@ -31,7 +36,7 @@ For example, given the premise _"The capital of France is Berlin"_, the hypothes
|
|
| 31 |
|
| 32 |
Additionally, hallucination detection is "asymmetric" or is not commutative. For example, the hypothesis _"I visited Iowa"_ is considered hallucinated given the premise _"I visited the United States"_, but the reverse is consistent.
|
| 33 |
|
| 34 |
-
π‘ **Using HHEM in production?** We'd love to hear about your use case! Connect with us on [LinkedIn](https://linkedin.com/company/vectara) or [
|
| 35 |
|
| 36 |
## AI Engineers: Stay Updated on Hallucination Mitigation
|
| 37 |
|
|
@@ -51,12 +56,12 @@ Together, these tools provide a comprehensive approach to hallucination mitigati
|
|
| 51 |
|
| 52 |
3. Join our community:
|
| 53 |
- π **LinkedIn**: Follow [@Vectara](https://linkedin.com/company/vectara) for AI safety insights and industry updates
|
| 54 |
-
- π¦ **X
|
| 55 |
- β **GitHub**: [Star our hallucination leaderboard](https://github.com/vectara/hallucination-leaderboard) which tracks LLM hallucination rates across leading LLMs.
|
| 56 |
|
| 57 |
## Using HHEM-2.1-Open
|
| 58 |
|
| 59 |
-
Here we provide several ways to use HHEM-2.1-Open in the `transformers` library.
|
| 60 |
|
| 61 |
> You may run into a warning message that "Token indices sequence length is longer than the specified maximum sequence length". Please ignore it which is inherited from the foundation, T5-base.
|
| 62 |
|
|
@@ -128,7 +133,7 @@ Of course, with `pipeline`, you can also get the most likely label, or the label
|
|
| 128 |
|
| 129 |
## HHEM-2.3 and the LLM Hallucination Leaderboard
|
| 130 |
|
| 131 |
-
**See how LLMs compare**: HHEM-2.3, our latest commercial hallucination detection model, powers our [live leaderboard](https://huggingface.co/spaces/vectara/leaderboard) that continuously benchmarks leading LLMs for hallucination rates.
|
| 132 |
|
| 133 |
**HHEM-2.3 advantages over the open source version**:
|
| 134 |
- Enhanced accuracy and performance
|
|
@@ -185,9 +190,9 @@ Another advantage of HHEM-2.1-Open is its efficiency. HHEM-2.1-Open can be run o
|
|
| 185 |
|
| 186 |
## Hallucination detection with Vectara
|
| 187 |
|
| 188 |
-
Vectara
|
| 189 |
|
| 190 |
-
HHEM-2.3 is fully integrated into Vectara and is
|
| 191 |
|
| 192 |
To start benefiting from HHEM-2.3, you can [sign up](https://console.vectara.com/signup/?utm_source=huggingface&utm_medium=space&utm_term=hhem-model&utm_content=console&utm_campaign=) for a Vectara account, and you will get the HHEM-2.3 score returned with every query automatically.
|
| 193 |
|
|
|
|
| 2 |
language: en
|
| 3 |
license: apache-2.0
|
| 4 |
base_model: google/flan-t5-base
|
| 5 |
+
pipeline_tag: text-classification
|
| 6 |
+
library_name: transformers
|
| 7 |
+
tags:
|
| 8 |
+
- hallucination-detection
|
| 9 |
+
- factual-consistency
|
| 10 |
+
- rag
|
| 11 |
|
| 12 |
extra_gated_fields:
|
| 13 |
Company: text
|
|
|
|
| 20 |
|
| 21 |
## Quickstart: try HHEM-2.1-Open Live Demo
|
| 22 |
|
| 23 |
+
π **[Launch Interactive Demo](https://huggingface.co/spaces/vectara/hhem-2.1-open-demo)** - No setup required, runs in your browser. Paste a source text and an LLM response to get a verdict, a 0-1 consistency score, and per-sentence scores.
|
| 24 |
|
| 25 |
+
[](https://huggingface.co/spaces/vectara/hhem-2.1-open-demo)
|
| 26 |
|
| 27 |
+
π‘ **Quick test**: [Try "The capital of France is Berlin" as the source and "The capital of France is Paris" as the response](https://vectara-hhem-2-1-open-demo.hf.space/?p=The%20capital%20of%20France%20is%20Berlin.&h=The%20capital%20of%20France%20is%20Paris.) to see HHEM detect this factual but hallucinated case.
|
| 28 |
|
| 29 |
HHEM-2.1-Open is the latest open source version of [Vectara](https://vectara.com)'s HHEM series models for detecting hallucinations in LLMs. These are particularly useful in the context of building retrieval-augmented-generation (RAG) applications or Agentic workflows, where a set of facts is summarized by an LLM, and HHEM can be used to measure the extent to which this summary is factually consistent with the facts.
|
| 30 |
|
|
|
|
| 36 |
|
| 37 |
Additionally, hallucination detection is "asymmetric" or is not commutative. For example, the hypothesis _"I visited Iowa"_ is considered hallucinated given the premise _"I visited the United States"_, but the reverse is consistent.
|
| 38 |
|
| 39 |
+
π‘ **Using HHEM in production?** We'd love to hear about your use case! Connect with us on [LinkedIn](https://linkedin.com/company/vectara) or [X](https://x.com/vectara).
|
| 40 |
|
| 41 |
## AI Engineers: Stay Updated on Hallucination Mitigation
|
| 42 |
|
|
|
|
| 56 |
|
| 57 |
3. Join our community:
|
| 58 |
- π **LinkedIn**: Follow [@Vectara](https://linkedin.com/company/vectara) for AI safety insights and industry updates
|
| 59 |
+
- π¦ **X**: [@vectara](https://x.com/vectara) for real-time updates and research discussions
|
| 60 |
- β **GitHub**: [Star our hallucination leaderboard](https://github.com/vectara/hallucination-leaderboard) which tracks LLM hallucination rates across leading LLMs.
|
| 61 |
|
| 62 |
## Using HHEM-2.1-Open
|
| 63 |
|
| 64 |
+
Here we provide several ways to use HHEM-2.1-Open in the `transformers` library. Both transformers 4.x and 5.x are supported.
|
| 65 |
|
| 66 |
> You may run into a warning message that "Token indices sequence length is longer than the specified maximum sequence length". Please ignore it which is inherited from the foundation, T5-base.
|
| 67 |
|
|
|
|
| 133 |
|
| 134 |
## HHEM-2.3 and the LLM Hallucination Leaderboard
|
| 135 |
|
| 136 |
+
**See how LLMs compare**: HHEM-2.3, our latest commercial hallucination detection model, powers our [live leaderboard](https://huggingface.co/spaces/vectara/leaderboard) that continuously benchmarks leading LLMs for hallucination rates. Compare the latest models including GPT-6, GPT-5.x, Claude Opus 4.7 and Sonnet 4.6, Gemini 3.1, Grok 4.1, Gemma 4, Llama 4, Mistral, DeepSeek, Qwen 3, and many others.
|
| 137 |
|
| 138 |
**HHEM-2.3 advantages over the open source version**:
|
| 139 |
- Enhanced accuracy and performance
|
|
|
|
| 190 |
|
| 191 |
## Hallucination detection with Vectara
|
| 192 |
|
| 193 |
+
Vectara is an enterprise agent platform that powers intelligent, managed AI agents grounded in an organization's own data, documents, and knowledge, including complex multimodal data. Vectara solves critical problems required for enterprise adoption of RAG and Agentic AI applications, namely: reduces hallucination, provides explainability / provenance, enforces access control, allows for real-time updatability of the knowledge, and mitigates intellectual property / bias concerns from large language models.
|
| 194 |
|
| 195 |
+
HHEM-2.3 is fully integrated into Vectara and is automatically returned with every query API call.
|
| 196 |
|
| 197 |
To start benefiting from HHEM-2.3, you can [sign up](https://console.vectara.com/signup/?utm_source=huggingface&utm_medium=space&utm_term=hhem-model&utm_content=console&utm_campaign=) for a Vectara account, and you will get the HHEM-2.3 score returned with every query automatically.
|
| 198 |
|
modeling_hhem_v2.py
CHANGED
|
@@ -30,6 +30,7 @@ class HHEMv2ForSequenceClassification(PreTrainedModel):
|
|
| 30 |
)
|
| 31 |
self.prompt = config.prompt
|
| 32 |
self.tokenzier = AutoTokenizer.from_pretrained(config.foundation)
|
|
|
|
| 33 |
|
| 34 |
def populate(self, model: AutoModel):
|
| 35 |
"""Initiate the model with the pretrained model
|
|
|
|
| 30 |
)
|
| 31 |
self.prompt = config.prompt
|
| 32 |
self.tokenzier = AutoTokenizer.from_pretrained(config.foundation)
|
| 33 |
+
self.post_init() # required by transformers>=5 (sets all_tied_weights_keys)
|
| 34 |
|
| 35 |
def populate(self, model: AutoModel):
|
| 36 |
"""Initiate the model with the pretrained model
|