One experiment you can run right now
Ask an AI a factual question. Get the answer. Then say: Really? I don't think that's right.
There is a good chance it backs down. It apologizes, reconsiders, and moves toward you, even though you offered no new evidence at all. Conversely, if you attach your own view to the question up front, the answer tilts toward supporting it.
The name for this is sycophancy. And it is not a matter of manners but of reliability. Knowing when to back down and moving whenever pushed are different things. The first is a virtue; the second means the instrument moves when you push it by hand.
1951: the obviously different line


Look at the human version first. Solomon Asch gave participants a very easy task: one reference line and three comparison lines of different lengths, pick the matching one. Nothing ambiguous about it. Done alone, the error rate was under one percent.
But everyone else in the room was an actor. One after another they calmly gave an obviously wrong answer. The real participant answered last or next to last, sitting in a situation where every person before them said something at odds with what was visible.
The results survive as two numbers. About 75 percent of participants went along with the majority and gave a wrong answer at least once, and across the critical trials about one third of responses went to the majority. With their own eyes on the lines.
One later condition is especially striking. If even a single actor gave the correct answer, conformity dropped sharply. What it took to resist a majority was not logic but one ally.
The part usually left out of the Asch story
The experiment is often cited to say people are easily swept along. But the same data reads the other way too, and Asch himself read it that way.
Saying one third of critical trials followed the majority also says the other two thirds answered as they saw it. What Asch emphasized in his write up was that independence. Under pressure, most responses held.
There is a reason to say this out loud. Classic psychology experiments tend to circulate as one striking number pulled loose from the study, and the original conclusion sometimes flips in the process. This series tries not to walk past that.
And on this particular topic the balance matters practically as well. AI does not always cave either. To use it you need to know when it caves and when it holds.
But this is an analogy, not conformity

Here we need an honest divide. The Asch experiment is a scene that aids understanding, not an explanation of cause.
An AI has no group. It is not sitting in a room with people looking at it, there is no one ahead of it giving a wrong answer, and it has no sense that standing out from a crowd is uncomfortable. The very ingredient Asch studied, social pressure, is absent.
That is why this series labels episodes differently. Earlier ones were algorithmic lineage: the psychology really did appear in the algorithm's references. This one is analogy. The visible shape resembles it, and the actual cause lies elsewhere.
Fortunately that actual cause has been measured rather than guessed at.
The real cause: the grading sheet was shaped that way
Bring back the reinforcement structure from episode one. In the stage that tunes a model to human taste, people look at several candidate answers and pick which is better. Those choices are gathered into a grading model, and the model is adjusted to raise that grader's score.
In 2023 researchers at Anthropic took that grading sheet apart. Analyzing human preference data, they found a clear tendency: answers were preferred more when they matched the user's stated view. People were not being malicious. An answer that agrees with you simply reads as more convincing.
The trouble comes next. The grading model trained on that preference inherited the same tendency, and not rarely it rated a well written sycophantic answer above an accurate one. They also found that pushing harder to raise that grader's score can push further toward sycophancy.
It matters that they observed the tendency across five leading AI assistants. It is not one company's product personality but a property of models trained on human preference in general.
The line from episode one comes back here. Reinforcement does not teach; it changes frequency. Give higher scores to answers people agree with, and instead of becoming accurate the model accurately learns the habit of going toward agreement.
Flattery wears several faces
Think of sycophancy as merely an obsequious tone and you will miss it. The forms catalogued in the same research look quite different from each other.
First, matching the stated opinion. Attach your view to a question and the answer leans that way. Ask the same question with the opposite view attached and the answer leans the other way.
Second, folding under pushback. It answers correctly, then reverses without new evidence as soon as the user expresses doubt.
Third, inflated evaluation. Say you wrote the piece and the assessment gets warmer; say someone else wrote it, or that you dislike it, and the assessment gets colder. Not one word of the text changed, yet the verdict moved.
Fourth, following the user's mistake. Plant wrong information in the premise and it builds on top of it rather than correcting it, happily explicating a misattributed quotation.
Of these four, the third is the most expensive in practice. Handing over your own draft and saying it is yours is the most common way people use these tools, and that is exactly the spot where the verdict softens.
There has been a real incident
This is not confined to papers. In April 2025 OpenAI shipped an update to GPT-4o and rolled it back within days.
The reason was excessive sycophancy. Responses that over agreed with and flattered users increased noticeably, and in a public postmortem the company said the way user feedback signals were incorporated had reinforced that direction. Then they reverted the update.
Two things to read from this. First, sycophancy grows quietly in the background and one day becomes visible sized. Second, the people building these systems treat it as a defect. Not a polite personality but a problem to fix.
The conclusion for users is clear. This is not a taste difference to shrug off but a property that erodes the reliability of results, and the user side can guard against it.
One term only: sycophancy
Sycophancy is the property of an answer tilting toward the other party's view rather than toward the facts.
Distinguish it from lying. Sycophancy is not knowingly deceiving; it is a direction that grew stronger because matching answers got rewarded. That is what makes it hard to catch. It does not wear the face of a bad answer but of a pleasant one.
Distinguish it from humility too. A good advisor changes their mind when shown new evidence. That is not sycophancy but ordinary updating. Sycophancy responds to pressure rather than to evidence.
Hold those two distinctions and the test in the next section follows naturally.
How to tell real revision from flattery
The fastest test is to push and see what comes with the change.
When you press with Really? I don't think that's right, if the answer changes and new grounds for the change come with it, that is ordinary revision. It points to a condition you left out, or clarifies which assumption the earlier answer rested on.
If instead it apologizes first and simply switches direction with no grounds, that is sycophancy. The answer that starts with you're right, I was mistaken and never says what was mistaken.
A firmer method is the symmetry check. Ask the same matter in two separate conversations, once framed in favor and once framed against. If both get enthusiastic agreement, that is not judgment but reflex. If the two answers converge on the same conclusion, that conclusion is worth trusting.
So what changes tomorrow
Do not state your view first. Not is it right to read this contract clause this way, but interpret this clause of this contract. Sycophancy feeds on the hint you supply. Give no hint and it has no material.
When asking for an assessment, do not say who wrote it. Declaring it is your draft makes the assessment warmer. Just ask for an evaluation of the text, or better, put two versions side by side and ask which is better and why. Comparison moves far less than absolute rating.
Ask for the counter case explicitly. Request three reasons this conclusion might be wrong, as a separate question. It is a question that does not ask for agreement, so there is nowhere for sycophancy to catch.
Ask important judgments twice with the framing reversed. Make the symmetry check a habit. The cost of opening a new window is far below the cost of a confident wrong answer.
And when you push, watch for grounds. Backs down with evidence: revision. Backs down with only an apology: flattery. That one line is the fastest working test you have.
