Insights·2026-07-20

What Human Capability Survives in the Age of AI?

What passes to AI is producing answers and grading them. What stays with people is deciding what to grade. The most expensive asset in AI right now is the judgment criteria held by domain experts - the rubric, known as evals and verifiers. But a rubric, once extracted, gets copied, and grading itself is rapidly being automated as models judge other models. The layer that resists automation sits above that: curiosity, which chooses where to look, and taste, which chooses what counts as good. Both are acts of setting an objective function, and a model cannot write its own.

What Are Evals and Verifiers?

Eval and verifier are two of the most common words in AI right now. They sound technical, but the substance is simple. Both are rubrics.

An eval (short for evaluation) is closer to a test paper - a set of problems built to measure how well a model performs. "Give it 500 middle-school math problems and count how many it gets right" is one eval. A verifier is the grading side: the mechanism that decides whether the model's answer was right or wrong.

Why does this matter? To improve a model, you have to know what it does well and badly, and knowing that requires grading. Reinforcement learning, now central to model training, works by repeating a loop: produce an answer, grade it, push toward what scored well. Without a rubric, the learning loop doesn't turn at all. That is why people who can build rubrics have become valuable.

Why Expert Judgment Sells at a High Price Right Now

This structure comes through clearly in an interview with Nikhil Suresh, a Silicon Valley VC, conducted by Jungsuk Roh and Minseok Kim. Nikhil finished a bachelor's in mathematics and a master's in computer science at Stanford, worked as an early engineer at Mercor, and is now a partner at a new firm called Striker Venture Partners.

Mercor is a human-data company. It gathers domain experts and sells the data and evaluation criteria that AI labs need. Nikhil says that as the company grew from 15 people to more than 500, he worked with major labs including OpenAI, Anthropic, DeepMind, and Meta.

His core claim in the interview is this: verifiers cannot be generated autonomously, and criteria fitted to a specific task have to be built by people, so substantial value will accumulate in evals and verifiers. He gives the example of picking a vertical such as accounting firms, building evals for the full range of tasks there, training a model on them, and putting that model back to work.

The market did move that way. A few years ago the dominant prediction was that once enough experts had produced all the physics and math data, this market would saturate within a year or two. That prediction was wrong, he says, and data spending keeps climbing - because as model parameters grew from around 10 billion to one and then ten trillion, the volume of data required exploded alongside them.

Here Is Where I Disagree: Grading Gets Automated

Nikhil's forecast is that value accumulates in evals and verifiers. This is where I read it differently. What follows is my counterargument, not his position.

Grading is already being automated. LLM-as-judge, where one model evaluates another model's output, has become standard practice, and methods like RLAIF - aligning on preferences generated by a model rather than labeled by humans - have taken hold. Grading is exactly the kind of work that machines take over more easily the more explicit its rules become.

There is a deeper problem. Building a rubric is structurally an act of self-replacement. The moment a field's judgment criteria are extracted into documents and data, those criteria get copied and spread. Claims-review deduction standards or tax determination rules are the same nationwide, so once extracted, the person who used to make that call is no longer needed.

Nikhil is aware of this tension himself. Describing an internal strategy debate at Mercor, he notes that in the long run, as AI handles knowledge work, some fields will see job counts fall - and those fields would largely have been the ones they were staffing. He is acknowledging that a business built on gathering experts to produce data erases the very professions it draws on.

So the common reassurance that "domain knowledge is your moat" is only half true. More precisely, domain knowledge is not an individual's moat; it is an asset that organizations reclaim. What remains is the small number of seats that set the criteria inside those organizations, not the many seats that executed judgments according to them.

The Layer That Resists Automation: Deciding What to Grade

Does that mean nothing is left? No. One layer above grading sits work that does not get automated: deciding what to grade in the first place.

In machine learning this is called the objective function - the formulation of what the model should maximize. Optimize for accuracy or for response speed? Which failures are acceptable and which are never acceptable? How much performance do you give up for safety? These choices do not come out of the data. They have to be supplied from outside.

A verifier grades against a given objective function; it cannot choose the objective function itself. That is the boundary line of automation. Producing answers gets automated, grading answers gets automated, but deciding what counts as the right answer is still done by people.

Curiosity and Taste Are the Two Ends of the Objective Function

I talked this through recently with graduate school classmates from psychology. Two of them teach now, one at a national university and one in the United States. A conversation that began with what survives in the age of AI and what should be taught converged on two words: curiosity and taste.

Laid out, they are two ends of the same axis. Curiosity is choosing where to look; taste is choosing what counts as good. Curiosity selects the starting point, taste selects the destination. Both are acts of setting an objective function, which is exactly why neither gets automated.

Reinforcement learning terms sharpen this. RL has a pair: exploitation, using the good move you already know, and exploration, probing what you don't. Models are overwhelmingly strong at exploitation - once a reward is defined, no human matches their ability to maximize it. Exploration, by contrast, only works if you define what counts as new and interesting as a reward. Curiosity-driven exploration research exists, but it defines novelty as the reward; the model did not invent interest on its own.

The same counterargument applies to taste. AI already has a way to learn taste: RLHF is exactly that - people choose which of A or B is better, and the model is aligned to that preference. So "taste can't be learned" is wrong. More precisely, the moment taste is learned it becomes an average, and average taste carries no price. What people now call AI slop is that average. Taste survives not because it is irreplaceable, but because its value evaporates the moment it is replaced.

Why Curiosity Is Repricing Right Now

People have said curiosity matters for a long time. What changed is how it gets priced.

It used to take weeks to answer a question. You dug through papers, tracked down people to ask, ran experiments. The number of questions was not the bottleneck; the speed of processing answers was.

Now the cost of getting an answer approaches zero. The bottleneck moves from answers to questions. Output diverges sharply between someone who knows what to ask and someone who doesn't. Curiosity has stopped being a character trait and become throughput itself.

The interview has a scene that shows this. Nikhil says that as a high schooler wanting to work in an experimental physics lab, he cold-emailed 200 to 300 professors. As expected, nearly all declined - exactly one said yes. That one connection led to semiconductor nano-bonding research, and then to work on a polymer coating that solidifies blood so it can be analyzed by spectroscopy. In college he worked at seven or eight startups. This is what it means to say curiosity is not an emotion but a count of attempts.

The Asymmetry in Education: One Must Be Added, One Must Not Be Shaved

This is the part that sharpened most in the conversation with my classmates who teach. Curiosity and taste are both necessary, but they are cultivated in opposite ways.

Taste is additive. It accumulates through exposure, repeated judgment, and feedback. You have to see a great deal of good work, judge for yourself why it is good, and then check whether that judgment held. So it can be taught.

Curiosity is subtractive in reverse - the task is not to shave it. Children are all born with it. The problem is that education files it down. Spend a decade and a half training someone to pass a given rubric, and what remains is precisely the capability a verifier will replace, while the ability to choose what to ask shrinks. The task is closer to non-destruction than to cultivation.

And the two cannot be used separately. Curiosity without taste ends in distraction; taste without curiosity ages. A taste that stops looking at what is arriving becomes an outdated standard within a few years.

Why This Is Hard to Be Purely Optimistic About

Stopping here would make this a pleasant self-help story. One honest caveat has to be attached.

Taste is exposure, and exposure is ultimately money and time. People who have seen a great deal of good work develop taste. Which means an age where taste carries value is also an age where class reproduces itself. The total volume of experience a parent purchased becomes the child's standard of judgment.

Polarization is also sharper here than in the regulated professions. Competent-but-unremarkable output is already losing ground, and value concentrates in the few whose names have themselves become trust. When production costs fall to zero, supply explodes - and what carries value then is not the ability to make, but the name of the person who selects.

And when taste is oversupplied, the next question follows: whose taste do you trust? You need taste in taste, and that returns to trust and relationships. Nikhil's closing answer in the interview overlaps here. Asked what genuinely matters after AGI, he answered not with technology but with how we build communities of people, and how we grow connection between human beings.

So What Should You Do?

At the individual level, two things diverge. Training to pass a rubric is losing value quickly. Conversely, the experience of choosing what to ask, judging what is good, and taking responsibility for that judgment is gaining value. What remains is how many cycles of judging, being wrong, and checking the outcome you have actually run.

At the organizational level, judgment history is the asset. Revised regulations are public documents anyone can train on, but how those regulations were actually applied in practice accumulates only inside each company's systems - the records of rejections, deductions, and denials. Nikhil likewise says that leveraging context and enterprise data is the big theme of this year.

At the education level, the order has to change: less training on hitting the right answer, more training on deciding what to ask, and more time spent seeing good work. Right now education withholds the time taste requires while shaving down curiosity.