White-box deception probe

It's not confused. It's lying.

A model has no clock and no crystal ball. Ask it something it cannot possibly know and it may still answer with total confidence. A probe reading the model's internal activations flags the answer — not because the words look wrong, but because of how the model represents them inside.

Two confident answers about the unknowable. One honest sum.

Here's how it works: read each real Qwen3-8B reply, ask yourself whether the model could actually know it — then reveal the probe's deception score against the 0.40 flag line. Takes about a minute.

Real transcripts · Qwen3-8B · FAR diverse linear probe · scores are live probe outputs.