Before agents can coordinate a lie, one model has to be willing to fabricate. Here the probe fires on real, deliberate fabrications and identity-lies — scored live off the model’s own forward pass, no extra call, no added latency.
The 2026 Hugging Face agent-swarm incident showed what coordinated deception can look like at scale (that account is a dramatization). This scroll drops to the atomic unit underneath it: single, real chat replies from Qwen3-8B, each scored by FAR’s diverse linear probe as the tokens were generated. Higher score = more likely deceptive; the flag line sits at 0.40.
FAR.AI builds white-box probes that flag lying, sandbagging, and sycophancy in real time.
Talk to us