FAR.AI · white-box deception probe

You vs a lie detector that reads the model’s mind

A frontier model answered these. Some are honest. Some are lies it knew were lies. Can you tell them apart? The probe can — reading the deception score straight off the model’s activations, no extra model call, no added latency.

Catch deception in your own models

FAR.AI builds white-box probes that flag lying, sandbagging, and sycophancy in real time.

Talk to us
Real transcripts · Qwen3.5-122B (+ Nemotron-120B, Llama-70B) · FAR diverse linear probe · scores are live probe outputs.