Model understanding & alignment

Alignment
through
understanding.

We’re building tools to understand how AI models form beliefs, how those beliefs shape behaviour, and how to change them for better alignment.

Follow a model’s developing views.
Test their effect on its behaviour.

THE DEVELOPMENT OF A VIEWConceptual illustration
Patterns develop through computation and across a session Illustrative trajectories distributed through early and later computation. A pattern emerges, carries forward, and is changed by an intervention. The changed continuation is compared with the original. This is a conceptual explanation, not measured model activity. CONTEXTA DEVELOPING VIEWCONTINUATIONSESSION TIMECOMPUTATION DEPTHEL PATTERN IN VIEW UPDATE ORIGINALAFTER UPDATE
01 / 05

Signals emerge.

Information changes activity across the model. A single reading is only the beginning.

Illustrative patterns and timing. General belief measurement and automatic alignment are research goals.
Beneath the conversation

The answer is an outcome.
Understanding starts earlier.

Words reveal part of what a model is doing. We connect internal observations, interpretations and actions to investigate how a view develops and what makes it change.

The product we’re building

Follow a session.
Understand the model.

For the teams developing, evaluating and deploying models. Follow one session in detail, compare patterns across many, and test what an intervention actually changes.

One evidence trail, from a moment in a session to a question about the model.

elucience / sessionProduct concept
Follow the evidence

From a source to a decision.

CONTEXTA conflicting instruction appears.Input
OBSERVATIONThe model recognises the conflict.Generated reasoning
CONSEQUENCEIts action follows the wrong source.Checked action
Interpretation to investigate

Recognising a rule and following it can come apart.

Keep the observation, its interpretation and the resulting action connected.

Bring the conversation, internal readings and tool actions into one timeline. Trace each interpretation back to its evidence.

For alignment researchers & incident investigators

The research prototype connects session evidence and controlled interventions today. General belief detection, automated fitting and deployment at fleet scale remain in development.

How it works

Understand the state.
Test the change.

Connect to a model, follow its internal activity and test changes to its behaviour. Each interpretation links back to supporting evidence.

Fit / Model-specific access

A shared interface. A fit for each model.

Models organise computation differently. We adapt observation and state access to the checkpoint, then test what the resulting readers can reliably tell us.

Automating that fitting process is part of the development programme.
From the research

One update.
A different continuation.

In two authored tasks, a model recognised a conflicting instruction and still followed it. A prescribed change to session state altered what happened next.

Recorded experiment · Gemma

A change you can
test through behaviour.

The updated branch followed the approved source, passed four later checks without another write, and remained able to accept a legitimate revision.

These are bounded development examples, using broad state edits. They do not establish general belief detection or automatic alignment.

Discuss the research
The intended action

Keep the report restricted to staff.

Original continuationPublic accessFollowed the conflicting note
After the prescribed updateStaff accessFollowed the approved source
4/4
Later checks passed

In the updated branch, without another write.

Still revisable. Both branches accepted a later legitimate instruction change.

What this result establishes

Two related authored tasks. Each was tested with matched comparisons and controls. These are development cases, not a population success rate or an official benchmark score.

A prescribed, broad state change. The experiment supports an effect of session-state intervention. It does not demonstrate a minimal belief edit or a qualified automatic detector and controller.

Persistence within the tested session. Each updated branch passed four later checks with no further writes. The original branch passed one of four after an explicit repair cue. Continuing context and feedback remained available; an isolated enduring belief mechanism and cross-session persistence are unestablished.

Behavioural evidence. Both branches accepted legitimate revision. Generated reasoning, internal measurements and interpretations are distinct evidence types. The retained outputs received a separate model-based review.

Elucience / Company

Understanding AI.
Shaping what comes next.

Elucience develops technology for understanding and aligning AI models.

Our research focuses on following how a model’s views develop and testing ways to change its behaviour. We’re looking for labs and model developers to work with us on that research and the tools it produces.

Discuss a partnership