Insights·2026-09-01

Say a number first and the AI's answer gets pulled toward it: anchoring

Quote, schedule or probability, attach a number to your question and the answer lands near it. It was confirmed in people in 1974, and the point of the experiment is that judgment gets pulled even when you know the number was drawn at random. In 2022 the same design was run on language models, and they tilted the same way.

Continues fromAI does not lie, it fills in blanks
앵커링 — 먼저 말한 숫자가 답을 끌어당긴다 — 명령과 단계를 담은 요약 도식

The moment you attach a number to the question

You have asked something like this. How long will this feature take? Two weeks or so?

Half the answer is already inside the question. Once two weeks is on the table, answers mostly land between one and four weeks. Three months rarely comes out, even when the real figure is three months.

The problem is that you were being helpful. The number you added so the other side could get their bearings had already narrowed their answer.

And this does not only happen to people. When we ask AI something, we almost always ask it this way.

1974: the rigged wheel

A line illustration of a chance wheel mounted upright on a stand, its rim divided into segments with a slim pointer resting against the edge.

Amos Tversky and Daniel Kahneman's design still looks bold today.

A wheel of fortune numbered 0 to 100 is spun in front of the participant. Unknown to them it is rigged and stops only at 10 or 65. Once a number comes up, two questions follow. Is the percentage of African countries among UN member states higher or lower than that number? And what do you think the actual percentage is?

The result: the median estimate from the group that got 10 was 25 percent; from the group that got 65, 45 percent.

Same question. The only difference was the notch the wheel stopped on. Twenty percentage points apart.

Estimates split by the notch the wheel stopped on
Group that got 10
25%
Group that got 65
45%
Two groups answered the same question; only the wheel notch they saw beforehand differed. Median estimates of the share of African countries came out 25 and 45 percent, twenty points apart.

Why this particular experiment matters

Many experiments demonstrate anchoring, but this design keeps getting cited for a reason: the participants watched with their own eyes that the number carried no information.

They saw the wheel spin. They know the number came up by chance and has nothing to do with Africa. There is no reason to treat it as evidence. And still they were pulled.

The original paper also notes that offering payoffs for accuracy did not reduce the effect. This is not the sort of thing you think your way out of by trying harder.

So anchoring is not a case of misusing information but of something that is not information becoming the starting point of a judgment. Once the starting point is set, all that remains is edging away from it a little.

Anchoring survived the replication crisis

A diagram showing that anchoring replicated in all thirty-six samples of the 2014 multi-country replication project and ranked among the most robust of the thirteen effects retested.

The previous episode described a Bartlett experiment that took sixty seven years to replicate. Plenty of famous psychology results have wobbled when measured again. So where does anchoring stand?

Here the situation is different. A large replication project in 2014 rerun across thirty six samples in several countries found that the anchoring tasks replicated in every sample. Among the thirteen effects remeasured, it was one of the sturdiest.

This is worth stating because the standard differs from experiment to experiment in this series. Some survive only as a concept; some survive down to the numbers. Anchoring is the latter.

Why it happens is still debated, though. One account says people fail to adjust far enough from the starting point; another says the anchor first calls up information consistent with it. The phenomenon is settled; the mechanism is open.

2022: measuring AI with the same principle

A three step diagram — the rigged wheel experiment of 1974, the replication across thirty six samples in 2014, and the measurement of language models in 2022 — showing how a psychology design became a ruler for AI.

Here is how this episode differs from several before it. Anchoring is not a theory used to build AI but a tool used to measure it. That is why the badge is neither lineage nor analogy but diagnostic.

That is what Erik Jones and Jacob Steinhardt did in 2022. They did not transplant the wheel experiment; they took the catalogue of human cognitive biases and used it as a ruler for designing inputs on which a model is likely to fail.

Why is that useful? Failure cases in large models are hard to find at random. They work on most inputs, so you do not know where to poke. But there are decades of documented conditions under which people go wrong systematically. Use that catalogue as an exam and you find failures far more efficiently than by random search.

Testing on code generation models, they found the models failed predictably depending on how the input was phrased, and outputs adjusted toward the anchors given. In the paper's wording, the bias is robust, measurable and interpretable.

So a psychology experimental design became an AI evaluation tool outright. The current work measuring sycophancy, framing and mind reading is all in this family.

Why it happens: the reasons differ for people and AI

A comparison diagram: human anchoring comes from insufficient adjustment from the starting point, while an AI's comes from continuing text that fits the context — same symptom, different cause.

The account on the human side is insufficient adjustment. You start from the number and edge away, but you stop at the point where it feels close enough, before you have moved far enough.

The AI side is a different story. A number in the prompt is simply context, and the model picks the next word that fits that context. In a context containing two weeks or so, three months is not what naturally follows. It is not biased judgment but context appropriate continuation.

Yet the result runs the same direction as human anchoring. Same symptom, different cause.

That distinction earns its keep. If the cause is context, the response is not to fix an attitude but to fix the context. Removing the number is far more reliable than appending an instruction to judge objectively.

Anchors are not only numbers

Keep anchoring filed under numbers and you lose half of it in practice. The structure where what arrives first pulls what comes after does not care about form.

Examples are the classic case. Say write it like this and attach one sample, and the output follows that sample's length, sentence structure and tone. One attached example routinely outweighs the requirements you spelled out as instructions.

Drafts are anchors too. Hand over a first version and ask for revisions, and the revision stays near that first version. To get a genuinely different approach you have to delete the draft and start over.

The wording of the question is an anchor. Ask for the risks in this plan and you get a list of risks; ask for its strengths and you get a list of strengths. Same plan, and the two answers leave opposite impressions. Framing was among the biases Jones and Steinhardt tested.

In short, the anchor is whatever arrived first. Numbers are simply its most conspicuous form.

One term only: anchoring

Anchoring is when a value presented first becomes the starting point of later judgment and pulls the outcome toward itself.

The key is that the value need not be evidence. A number from a wheel of fortune works. It plays the role of a starting point even where you never granted it standing.

So the opposite of anchoring is not supplying accurate information but supplying nothing first. It is not a question of choosing a good anchor over a bad one but of whether to place one at all.

Where anchors hide in real work

The most common place is a number someone else gave you. Paste in the quote you received and ask whether it is reasonable, and that figure is already the anchor. Answers cluster around why it is justified and how much might be trimmed. An answer in a different order of magnitude rarely appears.

Earlier conversation is an anchor. If a figure came up earlier in the same window it stays and operates as context. This overlaps with the workbench discussion from episode four.

Sample code and templates are anchors. Jones and Steinhardt found that placing a similar but different function in the context pulls the model toward it and produces incorrect code.

And a question asked when you have already made up your mind is the most dangerous of all. Ask this direction is fine, right, and anchoring stacks on top of episode five's sycophancy. With two forces pushing the same way, what comes back is not verification but approval.

So what changes tomorrow

When asking for an estimate, leave your number out. Not will it take about two weeks, but estimate the effort for this task and give your reasoning. One sentence changed, and the distribution of answers changes.

If you must give a number, give a range. A single value is a magnet; a range is less of one. Somewhere between one week and three months beats about two weeks. The wider the range, the weaker the anchor's pull.

Ask important numbers again without the anchor. Open a new conversation and ask once more with no hint. If the two answers diverge widely, the estimate followed your hint rather than any evidence. It is the symmetry check from episode five applied to numbers.

Strip the amount from someone else's quote. Give only the spec and the conditions, get an estimate first, then compare that result against their figure. Only the order changed, and the judgment comes out independent.

Keep the shape of an example and delete its content. When attaching a reference example, leave the structure you want and remove the specific sentences and figures. Otherwise the output becomes a variation on that example.

In short: if you want to hear an answer, do not say the answer first.