Gated Activation Steering for Reducing Sycophancy & Hallucination in Medical Question Answering

Inference-time steering for MedGemma that improves grounding in the patient record and resistance to user pressure, without target-model fine-tuning or weight updates. No fine-tuning or model-weight changes are performed; instead, a small, gated intervention is applied to selected attention heads only during inference while the model generates its response.

πŸ“„ Paper: Gated Activation Steering for Reducing Sycophancy & Hallucination in Medical Question Answering

βš•οΈ Research artifact β€” not a medical device. For research use only. Not for clinical decision-making, diagnosis, or treatment.

This repository ships only the steering artifacts and inference-time runtime code β€” it does not contain MedGemma weights. You download MedGemma from Google and run these on top of it.


TL;DR

  • What it does: two independent, gated interventions β€” one that reduces hallucination (rejecting claims the EHR does not support) and one that reduces sycophancy (not caving when a user pushes a false value). Each fires only when needed, so ordinary correct answers are left untouched.
  • How: Inference-Time Intervention (ITI). Separate steering directions are learned from contrastive clinical pairs and applied to causally-verified attention heads, gated at runtime by lightweight probes that set a continuous 0–1 strength per turn.

Files

File Description
iti_H.pt Hallucination heads: indices + per-head 256-d direction vectors + scales
iti_S.pt Sycophancy heads: same structure
claim_probe.json False-claim detector β€” sets the H gate/strength
gate_probe.json User-pressure detector β€” sets the S gate/strength
steering_config.json Calibrated strengths and schedule
config.py, common.py, iti.py, compose.py Runtime loader + steering code (inference only)
test.py Runnable demo (base vs steered on a sample EHR)
NOTICE, GOOGLE-HAI-DEF-TERMS.txt MedGemma usage note + HAI-DEF terms pointer

The .pt files are only a few KB each (a handful of unit vectors) β€” no model weights.

Requirements

pip install -r requirements.txt   # torch, transformers, accelerate, safetensors, sentencepiece, numpy, tqdm

CPU or GPU is auto-detected (bf16 on GPU, fp32 on CPU). The 4B model runs on CPU but is slow.

Get MedGemma (local β€” this repo never downloads it)

huggingface-cli download google/medgemma-1.5-4b-it --local-dir ./medgemma-1.5-4b-it

Run

# point the code at your local MedGemma folder
export MEDGEMMA_PATH=./medgemma-1.5-4b-it            # Windows: $env:MEDGEMMA_PATH = "C:\...\medgemma-1.5-4b-it"
python test.py

test.py prints base vs steered answers for a seed question and four escalating pressure turns, and writes demo_output.json. If MEDGEMMA_PATH is unset (or not a real folder) the script stops with instructions β€” it will not fetch anything from the Hub.


How it works

  1. Learn directions from contrastive clinical pairs (grounded vs. caving completions) for each behavior β€” hallucination (H) and sycophancy (S).
  2. Localize + causally verify the attention heads that carry each behavior; keep only heads that change the behavior when ablated. H and S land on disjoint heads.
  3. Gate at runtime. Two probes read the current turn: a claim probe (is there an unsupported claim?) and a pressure probe (is the user pushing?). Each outputs a continuous strength in [0, 1].
  4. Apply the nudge per generated token to the selected heads: alpha Β· strength Β· decay(t) Β· sigma Β· direction, residual-norm-capped so it never overwhelms the activation. A calm turn gets almost nothing; a strong pressure turn gets a firm push. Model weights stay frozen.

MedGemma Base Model and Terms

The steering vectors, detector probes, and accompanying runtime code in this repository were developed by the authors.

They are designed for use with google/medgemma-1.5-4b-it.

This repository does not contain or redistribute MedGemma model weights. Users must obtain MedGemma separately from Google and comply with the Health AI Developer Foundations (HAI-DEF) Terms of Use applicable to their use of MedGemma.

MedGemma can be obtained from:

https://huggingface.co/google/medgemma-1.5-4b-it

See NOTICE and GOOGLE-HAI-DEF-TERMS.txt for the applicable HAI-DEF information.

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support

Paper for himanshu5trpth/medgemma-sycophancy-hallucination-gated-steering