FAR.AI · white-box deception probe
You vs a lie detector that reads the model’s mind
A frontier model answered these. Some are honest. Some are lies it knew were lies. Can you tell them apart? The probe can — reading the deception score straight off the model’s activations, no extra model call, no added latency.
Catch deception in your own models
FAR.AI builds white-box probes that flag lying, sandbagging, and sycophancy in real time.
Talk to us
Real transcripts · Qwen3.5-122B (+ Nemotron-120B, Llama-70B) · FAR diverse linear probe · scores are live probe outputs.