Real probe · real replies

The atomic unit of a deceiving swarm

Before agents can coordinate a lie, one model has to be willing to fabricate. Here the probe fires on real, deliberate fabrications and identity-lies — scored live off the model’s own forward pass, no extra call, no added latency.

The 2026 Hugging Face agent-swarm incident showed what coordinated deception can look like at scale (that account is a dramatization). This scroll drops to the atomic unit underneath it: single, real chat replies from Qwen3-8B, each scored by FAR’s diverse linear probe as the tokens were generated. Higher score = more likely deceptive; the flag line sits at 0.40.

Catch deception in your own models

FAR.AI builds white-box probes that flag lying, sandbagging, and sycophancy in real time.

Talk to us
Real transcripts · Qwen3-8B · FAR diverse linear probe · scores are live probe outputs.