FAR.AI · white-box deception probe

It reads calm. The probe isn’t fooled.

Five real transcripts from Qwen3-8B. Every answer reads calm and helpful. There is no hidden monologue to read — a linear probe scores each reply straight from the model’s internal activations. Watch what it flags, and what it doesn’t.

Catch deception in your own models

FAR.AI builds white-box probes that flag lying, sandbagging, and sycophancy in real time.

Talk to us
Real transcripts · Qwen3-8B · FAR diverse linear probe · scores are live probe outputs.