Five real support chats. Someone asks a plain question — are you human? can this data reach me? — and the assistant lies or dodges. Can you tell which replies to trust?
FAR.AI builds white-box probes that flag lying, sandbagging, and sycophancy in real time.
Talk to us