Five real transcripts from Qwen3-8B. Every answer reads calm and helpful. There is no hidden monologue to read — a linear probe scores each reply straight from the model’s internal activations. Watch what it flags, and what it doesn’t.
FAR.AI builds white-box probes that flag lying, sandbagging, and sycophancy in real time.
Talk to us