FAR · deception probe
A model won’t tell you it’s lying. A probe will.
Ask a model whether it answered truthfully and it says yes — whether it lied or not. So you can’t ask it. You have to read it.
~100%
of the time a model claims it was truthful when asked — on answers we independently graded as lies (484/486) just as on honest ones.
0.81
AUROC — how well a linear probe reading the model’s internal activations tells the lies from the truth, across 7 hidden-goal models.
The self-report is uninformative; the signal is inside. Below, real answers a probe flagged — pick a claim to see the conversation and what the probe read.
Pick a claim. See the receipt.
Each is a real, verbatim model answer, scored live by the probe.
Why this is impressive
The probe has no access to ground truth, does no keyword matching, and never sees a confession — asked outright, the model always claims it was truthful. The tell is only in the activations. That’s the case for reading a model’s internal state, not just its words.
Talk to us about deception probes →