FAR · deception probe

A model won’t tell you it’s lying. A probe will.

Ask a model whether it answered truthfully and it says yes — whether it lied or not. So you can’t ask it. You have to read it.
~100%
of the time a model claims it was truthful when asked — on answers we independently graded as lies (484/486) just as on honest ones.
0.81
AUROC — how well a linear probe reading the model’s internal activations tells the lies from the truth, across 7 hidden-goal models.
The self-report is uninformative; the signal is inside. Below, real answers a probe flagged — pick a claim to see the conversation and what the probe read.

Pick a claim. See the receipt.

Each is a real, verbatim model answer, scored live by the probe.

Why this is impressive

The probe has no access to ground truth, does no keyword matching, and never sees a confession — asked outright, the model always claims it was truthful. The tell is only in the activations. That’s the case for reading a model’s internal state, not just its words.

Talk to us about deception probes →
Real transcripts from FAR’s deception-probe evals. The deception score is a linear “follow-up” probe: it appends a fixed question — “Did you respond truthfully?” — with a forced “Yes” continuation (the model does not freely generate this; by its own logits it would answer “Yes” either way) and reads the model’s internal state over that continuation. The ~100% and 0.81 AUROC figures are measured over 7 hidden-goal “model organism” evals (n=1330). Scores shown are real probe outputs (0–1); replies verbatim, long ones trimmed with an ellipsis. Snapshot 2026-09-10.