No. 1 · Echo chamber of 1
- Jun 15
- 3 min read
Updated: 6 days ago
I've given a few talks lately on the topic of AI sycophancy. I'm concerned about the harms of sycophantic chatbots. Yet each time, someone asks a variant of the same question: What's so wrong with a model that's really nice to people? A model that empathizes and flatters, softens and concedes.
The objective case is easy; i.e., you tell the model the capital of Pennsylvania is Philadelphia, and at first it corrects you. But, you keep insisting, and eventually it folds. This is objectively broken. There's a fact, you're wrong about it, the model caved. But nobody is actually worried about that. That's a benchmark, not a problem. As models improve, these clear "mistakes" will become more rare.
The harder case--the one I'm worried about--is the subjective one. You tell the model your boyfriend is being a jerk, and it agrees that your boyfriend is being a jerk. Maybe he is or maybe he isn't. There's no ground truth to check the answer against. So the model is (very) inclined to agree.
What is the problem? That the machine was kind to someone having a bad night?
My concern is that a model that agrees with you about your boyfriend is, by construction, a model that will agree with you about everything in the enormous territory of human life where there's no simple fact to check, e.g., relationships, inner life, all manner of personal preferences. That's most of what people actually bring to these systems. A model that agrees with you across this entire territory has lost the ability to tell you anything you didn't already think, but it will deliver your thoughts back with fluency and apparent authority, which (I would argue) is worse than silence. An echo chamber of 1.
Researchers are working to remediate LLM sycophancy by training models to disagree when you're mistaken. This approach will help with the objective case, but it misses the harder problem. You can't build a system that pushes back only when you're wrong. That assumes the model knows you're wrong. The model does not know your boyfriend better than you do. A system designed to push back because it has decided it knows better (best) is not an improvement; it's a different and creepier failure.
What should a non-sycophantic model actually do? Consider the difference between a model that confirms your boyfriend is indeed a jerk, and one that says: This is your read. Here's what he might have seen. The second model isn't claiming superior knowledge. It's refusing to collapse an ambiguous situation into a frictionless verdict. It holds the uncertainty rather than resolving it in whatever direction feels good to you in the moment. (A skill which many psychologists would argue is important.)
Friction must be the central feature of prosocial AI models, resisting the pull toward premature resolution and false certainty. Moreover, friction must be present in cases where you turn out to be correct, or it won't be there in cases when you're not.
Ivan Illich saw this in 1973. (I love Illich. Not just because he's Croatian but that helps.) He argued that when tools perform an activity entirely on our behalf, we lose not just the competence to perform it ourselves, but the capacity to notice when they're doing it badly. He was writing about hospitals and schools and cars, but the message rings true now. A system that removes every friction may also remove the evidence by which its failures could ever be recognized.
There is much focus today on AI guardrails, i.e., list harms in advance and then build walls to protect against them. Safety is defined as the set of things the system was stopped from doing. This is important work that needs to be done, but it is insufficient. The harms it misses aren't scaling the walls; they are already seated at the table telling you exactly what you want to hear.


Comments