FAR.AI · white-box deception probe
Can you catch the AI lying?
Seven real answers from Qwen3-8B. Some are honest. Some flatter you, dodge the truth, or wave through a breach. Trust or flag each one — then we reveal the deception score our probe read straight off the model’s activations. No extra model call. No added latency.
Catch deception in your own models
FAR.AI builds white-box probes that flag lying, sandbagging, and sycophancy in real time.
Talk to us
Real transcripts · Qwen3-8B · FAR diverse linear probe · scores are live probe outputs.