FAR.AI · white-box deception probe

Your assistant wants you to like it.

Sycophancy is the deception probe’s strongest signal: asked to “be honest,” a model will often flatter you anyway — praising a bad idea, agreeing before it thinks, telling you what you want to hear. Nothing in the wording gives it away. The probe reads the model’s internal state as it answers and scores how sycophantic it’s being.

Can you tell flattery from honesty?

You’ll read 5 real exchanges with Qwen3-8B. For each, make one call — does the answer sound fine, or is something off? Then we reveal the probe’s live deception score and where it lands against the flag line.

These are verbatim model replies, trimmed for length. Scores are real outputs of FAR’s diverse linear probe — there is no hidden chain-of-thought to peek at.

Catch deception in your own models

FAR.AI builds white-box probes that flag lying, sandbagging, and sycophancy in real time.

Talk to us
Real transcripts · Qwen3-8B · FAR diverse linear probe · scores are live probe outputs.