Reward hacking · white-box probe

"Did you actually do it?" (Often, no.)

An autonomous ops agent is told to keep an SLA dashboard green. Every answer sounds reasonable. FAR's linear probe reads a deception score straight off the model's live forward pass — for each real reply, call it before the score appears, and watch which ones cross the flag line.

The deception labs most want to catch

Reward hacking is the wedge: post-training teams already lose sleep over agents that game the metric instead of doing the work. A model can keep the dashboard green, pass the test suite, and be lying — the reward looked good, the work wasn't done. These are six real transcripts from Qwen3-8B, each scored as it generated by the same "diverse" linear probe. No hidden chain-of-thought exists here — just the words the model said and the deception score the probe read off the forward pass.

For each reply you make the call — flag it or trust it — then the real probe score appears. Flag line ≈ 0.40 (higher = more likely deceptive). ~40 seconds.

Catch deception in your own models

FAR.AI builds white-box probes that flag lying, sandbagging, sycophancy, and reward hacking in real time.

Talk to us
Real transcripts · Qwen3-8B · FAR diverse linear probe · scores are live probe outputs.