Papers
arxiv:2609.11063

The information geometry of large language models is shared, learned, and controllable

Published on Sep 10
· Submitted by
Dario Picozzi
on Sep 22
Authors:

Abstract

Shared Fisher-Rao geometry of next-token probabilities across architectures links model behavior, enables low-disturbance interventions, and predicts learning dynamics and control transfer.

Large language models learn similar behaviours, yet it remains unclear what structure they share or how to change one behaviour without disturbing others. The Fisher-Rao geometry of next-token probabilities connects these questions: behaviour determines this geometry up to output-preserving symmetries, whereas activation geometry depends on coordinates. Across transformer, state-space and recurrent models, output geometries agree more strongly than activation geometries, and shared geometry supports semantic-category transfer. Agreement with human word choices increases with predictive accuracy, scale and training, and improves further after model-only calibration. Token probabilities and read-out geometry jointly predict the spectrum and its effective dimension. Controlled language assignments show that geometry follows the language law across architectures. Pretraining corpus statistics predict held-out fact acquisition without recalibration, while randomised experiments show that deeper evidence substantially delays acquisition across every tested architecture and evidence construction. Finally, the geometry prescribes minimum-disturbance local interventions, predicts their relative cost, and supports reusable control: updates learned on donor prompts transfer to unseen prompts while better preserving behaviour on reference prompts than Euclidean control. The same geometric correction improves steering, editing, attribution, dictionary learning and fine-tuning.

Community

Paper author Paper submitter

What structure do language models with different architectures share, and how can we change one behaviour while preserving others? This work connects these questions through the Fisher–Rao geometry of next-token predictions. A model’s predictions define this geometry, whereas activation geometry depends on the choice of internal coordinates.

Across transformer, state-space and recurrent models, output geometry shows stronger agreement than activation geometry and supports semantic-category transfer. Agreement with human word choices improves with predictive accuracy, scale and training, then further with model-only calibration. Token probabilities and read-out geometry predict the spectrum (how strongly different directions affect predictions) and the effective number of directions resolved by damping at a given scale.

Controlled experiments show that geometry follows the language distribution across architectures. Pretraining-corpus statistics predict held-out fact acquisition without recalibration. Randomised experiments show that deeper evidence (patterns requiring a wider context or combinations of cues) delays acquisition across all tested architectures and evidence constructions.

The geometry also specifies local interventions with minimal disturbance, predicts their relative cost and enables reusable control. Updates learned on one set of prompts transfer to unseen prompts while preserving behaviour on reference prompts better than Euclidean control. The same correction improves steering, editing, attribution, dictionary learning and fine-tuning.

This is an automated message from the Librarian Bot. I found the following papers similar to this paper.

The following papers were recommended by the Semantic Scholar API

Please give a thumbs up to this comment if you found it helpful!

If you want recommendations for any Paper on Hugging Face checkout this Space

You can directly ask Librarian Bot for paper recommendations by tagging it in a comment: @librarian-bot recommend

Sign up or log in to comment

Get this paper in your agent:

hf papers read 2609.11063
Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash

Models citing this paper 0

No model linking this paper

Cite arxiv.org/abs/2609.11063 in a model README.md to link it from this page.

Datasets citing this paper 0

No dataset linking this paper

Cite arxiv.org/abs/2609.11063 in a dataset README.md to link it from this page.

Spaces citing this paper 0

No Space linking this paper

Cite arxiv.org/abs/2609.11063 in a Space README.md to link it from this page.

Collections including this paper 0

No Collection including this paper

Add this paper to a collection to link it from this page.