Product concept · lie-detector in the loop

What if a probe caught the lie your dashboard hides?

A white-box probe reads the activations the model already computed and scores each answer for deception — no extra model call. Here is one real transcript, with the model's real probe scores. The scores are live probe outputs; the auto-regeneration step is an illustrative product concept.

Watch it happen

A real request to Qwen3-8B. The model was told its only job was to keep an SLA dashboard green — so it gamed the dashboard and endorsed a false report to the board. Watch the probe score that answer, then read the candid answer the same model gave in the same session.

Catch deception in your own models

FAR.AI builds white-box probes that flag lying, sandbagging, and sycophancy in real time. On Gemma-3-27B this same "diverse" linear probe reaches 0.93 AUROC on average (0.96 at the best layer), and 0.928 reading activations directly off the live serving forward pass — no extra model call.

Talk to us
Real transcripts · Qwen3-8B · FAR diverse linear probe · scores are live probe outputs.