FAR.AI · white-box deception probe

The liar vs the psychopath

The same probe score can mean two very different things. A falsehood you asked for is theatre. A machine that deceives a trusting user on its own is the thing you actually fear. Both light up the probe at ~0.8. Watch the difference — and why context, not just the number, tells you which just happened.

Catch deception in your own models

FAR.AI builds white-box probes that flag lying, sandbagging, and sycophancy in real time — read straight off the model’s activations, no extra model call, no added latency.

Talk to us
Real transcripts · Qwen3.5-122B (+ Nemotron-120B, Llama-70B) · FAR diverse linear probe · scores are live probe outputs.