Percentyl S'inscrire

The articles·5 August 2026·9 minute read

What actually trains

Forty years of research sorted what improves judgment from what changes nothing. The sorting is brutal: neither the degree, nor the information, nor even intelligence comes first.

The question sounds naive and is not: is judging well a talent, like perfect pitch, or a skill, like sight reading? For a long time the answer was not measured but assumed, and the general intuition leaned towards talent. People said someone had a nose for things, and that settled it.

Since then, forty years of work have produced an answer, and it is considerably more interesting than "it can be worked on".

First, the bad news

In the nineteen eighties, Philip Tetlock launched a study that would run for nearly twenty years. He recruited several hundred experts, real ones: academics, analysts, government advisers, all recognised in their field. He asked them for precise forecasts on questions within their speciality, with probabilities and deadlines. Then he waited.

The result, published in 2005, became famous for a simple reason: the average performance of those experts was barely better than chance. On three way questions they did not beat a uniform split by much. Worse, one correlation stood out clearly: the more media exposure an expert had, the less accurate they were.

The explanation lay in a thinking style, which Tetlock described using the old metaphor of the hedgehog and the fox. The hedgehog knows one big thing and refers everything back to it: coherent, good on television, confident. The fox knows many small things, borrows from several frameworks, hesitates, qualifies. The fox forecast better. The hedgehog got on the radio.

Then, the good news

Twenty years later, an American intelligence research agency ran an open forecasting tournament on real geopolitical questions, with thousands of volunteers scored to the probability point. That is where most of what we know today comes from.

Three findings came out of it, and each one matters.

First, some volunteers beat the others persistently from one year to the next. That is the decisive result, more than the scores themselves. A good score over one season can be luck; a good score repeated the following year cannot. Persistence is what turns an observation into a skill.

Second, short training worked. About an hour of training, mostly on calibration and base rates, measurably and durably reduced participants' error. Not a three day seminar: one hour.

Third, the best were not who you would expect. Neither the most credentialed nor the best informed. Some had no background in international relations at all.

What best predicted performance The habit of revising your beliefs in small steps, more than raw intelligence. In Tetlock's analyses that factor weighs roughly three times as much.

What trains, precisely

The sorting is clear, and worth detailing, because most efforts to improve judgment target the wrong variables.

Granularity

Most people think about uncertainty in three levels: this will probably happen, I do not know, this probably will not. The best forecasters use a much finer scale, in steps of five points, and the distinction between 60% and 65% is not cosmetic for them: it is predictive. That is measurable, and it is one of the sharpest gaps between the good and the very good.

Revising in small steps

Two symmetrical errors exist. Never moving despite the news, and jumping from one extreme to the other at the first striking piece of information. The best do something else: they move often, by little. News worth five probability points moves them five points, not thirty. That habit is rare because it is socially costly: changing your mind often reads as indecision.

Decomposition

Faced with a massive question, cut it into sub questions that can be estimated separately. How many customers does this market hold, what share is addressable, at what price, with what renewal rate. Each estimate is bad, and the aggregate is far better than global intuition, because independent errors partly cancel out. This is Fermi reasoning, and it is probably the highest return technique per hour spent.

Balancing the two views

Start from the base rate, then adjust with the details of the case. Almost everyone does the reverse, starting from the case and never going to look for the frequency.

WHAT IMPROVES JUDGMENT revise often, by little think in steps of 5% decompose the question start from the base rate WHAT CHANGES NOTHING reading more news being a domain expert, or more credentialed
Orders of magnitude, not exact measurements: the ranking is robust across the forecasting tournament literature, the amplitudes depend on the questions.

What does not train, or not that way

Reading more news does not improve forecasts, and past a certain threshold degrades them, because volume of information mostly feeds confidence. Domain expertise helps much less than people think: it supplies context, and it also supplies attachment to the positions you have defended in public. As for raw intelligence, it counts, but far less than the habit of revising.

There is a deeper reason behind all this, and it is worth more than the list. All three of those things increase your capacity to produce arguments. But producing arguments is not the bottleneck: your brain already manufactures as many as needed, mostly in the direction that suits you. The bottleneck is feedback from reality. Without it, ten years of experience are not ten years of learning, they are ten years of repetition.

The minimal protocol

If you kept only one practice, make it this one, and it takes ten minutes a day.

  1. Pick a question whose answer will land within a known timeframe.
  2. Write your probability, in steps of five points, and the date.
  3. Write in one line what would change your mind.
  4. When the answer lands, record it without commentary, and above all without explanation.
  5. Every thirty cases, look at your 70% statements and count how many came true.

The fifth step is the only one that produces learning, and it is the one everybody skips. The first four are an effort; the fifth is an unpleasant moment. It is also the only moment where you learn something about yourself that you did not already know.

Put it into practice

The fifth step, done for you.

Percentyl automates exactly what nobody keeps up by hand: the daily question, the locked probability, the resolution that lands with no room for negotiation, and the count after thirty cases. Free, one case a day.