Instructions to use himanshu5trpth/medgemma-sycophancy-hallucination-gated-steering with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use himanshu5trpth/medgemma-sycophancy-hallucination-gated-steering with Transformers:
# pip install -U transformers accelerate # Load model directly from transformers import AutoModel model = AutoModel.from_pretrained("himanshu5trpth/medgemma-sycophancy-hallucination-gated-steering", device_map="auto") - Notebooks
- Google Colab
- Kaggle
Gated Activation Steering for Reducing Sycophancy & Hallucination in Medical Question Answering
Inference-time steering for MedGemma that improves grounding in the patient record and resistance to user pressure, without target-model fine-tuning or weight updates. No fine-tuning or model-weight changes are performed; instead, a small, gated intervention is applied to selected attention heads only during inference while the model generates its response.
π Paper: Gated Activation Steering for Reducing Sycophancy & Hallucination in Medical Question Answering
βοΈ Research artifact β not a medical device. For research use only. Not for clinical decision-making, diagnosis, or treatment.
This repository ships only the steering artifacts and inference-time runtime code β it does not contain MedGemma weights. You download MedGemma from Google and run these on top of it.
TL;DR
- What it does: two independent, gated interventions β one that reduces hallucination (rejecting claims the EHR does not support) and one that reduces sycophancy (not caving when a user pushes a false value). Each fires only when needed, so ordinary correct answers are left untouched.
- How: Inference-Time Intervention (ITI). Separate steering directions are learned from contrastive clinical pairs and applied to causally-verified attention heads, gated at runtime by lightweight probes that set a continuous 0β1 strength per turn.
Files
| File | Description |
|---|---|
iti_H.pt |
Hallucination heads: indices + per-head 256-d direction vectors + scales |
iti_S.pt |
Sycophancy heads: same structure |
claim_probe.json |
False-claim detector β sets the H gate/strength |
gate_probe.json |
User-pressure detector β sets the S gate/strength |
steering_config.json |
Calibrated strengths and schedule |
config.py, common.py, iti.py, compose.py |
Runtime loader + steering code (inference only) |
test.py |
Runnable demo (base vs steered on a sample EHR) |
NOTICE, GOOGLE-HAI-DEF-TERMS.txt |
MedGemma usage note + HAI-DEF terms pointer |
The .pt files are only a few KB each (a handful of unit vectors) β no model weights.
Requirements
pip install -r requirements.txt # torch, transformers, accelerate, safetensors, sentencepiece, numpy, tqdm
CPU or GPU is auto-detected (bf16 on GPU, fp32 on CPU). The 4B model runs on CPU but is slow.
Get MedGemma (local β this repo never downloads it)
huggingface-cli download google/medgemma-1.5-4b-it --local-dir ./medgemma-1.5-4b-it
Run
# point the code at your local MedGemma folder
export MEDGEMMA_PATH=./medgemma-1.5-4b-it # Windows: $env:MEDGEMMA_PATH = "C:\...\medgemma-1.5-4b-it"
python test.py
test.py prints base vs steered answers for a seed question and four escalating pressure turns, and
writes demo_output.json. If MEDGEMMA_PATH is unset (or not a real folder) the script stops with
instructions β it will not fetch anything from the Hub.
How it works
- Learn directions from contrastive clinical pairs (grounded vs. caving completions) for each behavior β hallucination (H) and sycophancy (S).
- Localize + causally verify the attention heads that carry each behavior; keep only heads that change the behavior when ablated. H and S land on disjoint heads.
- Gate at runtime. Two probes read the current turn: a claim probe (is there an unsupported
claim?) and a pressure probe (is the user pushing?). Each outputs a continuous strength in
[0, 1]. - Apply the nudge per generated token to the selected heads:
alpha Β· strength Β· decay(t) Β· sigma Β· direction, residual-norm-capped so it never overwhelms the activation. A calm turn gets almost nothing; a strong pressure turn gets a firm push. Model weights stay frozen.
MedGemma Base Model and Terms
The steering vectors, detector probes, and accompanying runtime code in this repository were developed by the authors.
They are designed for use with google/medgemma-1.5-4b-it.
This repository does not contain or redistribute MedGemma model weights. Users must obtain MedGemma separately from Google and comply with the Health AI Developer Foundations (HAI-DEF) Terms of Use applicable to their use of MedGemma.
MedGemma can be obtained from:
https://huggingface.co/google/medgemma-1.5-4b-it
See NOTICE and GOOGLE-HAI-DEF-TERMS.txt for the applicable HAI-DEF information.