Sit back, relax and enjoy our latest articles...

Sit back, relax and enjoy our latest articles…
Most Important Score Nobody Checks blog

The Most Important Score Nobody Checks

In my last piece, I ended with a thought: we’re not bad at developing leaders. We’re bad at spotting them.

So this time I want to look at how organisations actually decide who has leadership potential. Not how the policy says it works. How it works in practice.

image of a person under a magnifying glass with a check mark above their head

How most organisations choose

If you strip away the language, it usually comes down to two things.

A manager puts a name forward. And that person has a good performance rating.

That’s not a criticism. It’s simply what almost everyone does. A global study by Talogy in 2025 found that 91% of HR professionals and 88% of leaders use performance ratings and manager recommendations to identify their high-potential people. Seventy per cent of the organisations in that study said they had their own tailored definition of potential. But when it came to the actual decision, nearly all of them fell back on the same two inputs.

Most organisations then put the results into a nine-box grid. You’ll know it. Performance along one side, potential along the other, nine boxes, and everyone ends up in one of them. It was developed at GE in the 1970s and it’s still the most widely used talent tool in the world.

And here’s the part that stayed with me. In that same Talogy study, 79% of HR professionals rated their high-potential programme as effective. But the study also noted that many organisations measured success simply by tracking who took part.

So we use a method we’ve never really tested. We say it works. And the main way we check is by counting who went through it.


What happens when someone does check

This year, a study was published that tested it properly.

Three economists – Alan Benson, Danielle Li and Kelly Shue – looked at nearly 30,000 management-track employees at a large North American retailer over six years. The company used a nine-box grid, just like most organisations do. Their paper was published in February in the American Economic Review, one of the most respected journals in economics.

They found four things.

The potential score matters far more than the performance score. Moving someone up from medium to high on potential increased their chances of promotion far more than moving them up the same amount on performance – roughly three times as much. Whatever we say about rewarding results, it’s the potential rating that really decides who moves up.

The potential scores weren’t evenly spread. Women in the study received higher performance ratings than men. But they received noticeably lower potential ratings. That gap in potential explained around half the difference in how often men and women were promoted.

The potential scores were wrong. This is the important one. The researchers then looked at what actually happened next. Women who had been given the same potential rating as men went on to outperform them. The score hadn’t predicted anything. It had simply been lower.

And nobody corrected it. Even after those women had clearly done better than their rating suggested, their next potential rating stayed low. The system saw the evidence and carried on.


Why this matters beyond gender

It would be easy to read that as a study about women. It is partly that. But I think the bigger point is more general, and more uncomfortable.

A judgement made by experienced managers, using the most common tool in the world, about who can grow – turned out to be wrong in a consistent direction. And it couldn’t learn from its own mistakes.

If that’s true in your organisation too, then part of the reason your pipeline feels thin isn’t a shortage of capable people. It’s that your process isn’t finding them.


Why doing it more carefully won’t fix it

When organisations realise their potential ratings aren’t reliable, they usually do the sensible thing. They tighten the process. Clearer definitions. Calibration meetings. Evidence required. Bias training for managers.

All of that is reasonable. But it doesn’t solve this problem.

Calibration makes managers agree with each other. It doesn’t make them right. If what they’re measuring is off, calibration just means everyone is off in the same way – with better paperwork.


What the rating is really measuring

So why would a potential rating be wrong in such a consistent way?

I think part of the answer is surprisingly simple, and it isn’t anybody’s fault.

Most serious ways of assessing potential include past experience. Have you led a big project? Had a stretch assignment? Taken on a difficult role? That makes sense on the surface. People who have been stretched before are often more ready to be stretched again.

But think about what it means. We’re judging what someone could become partly by looking at what we’ve already given them.

And those chances aren’t handed out evenly. McKinsey and LeanIn’s Women in the Workplace 2025 study, covering 124 organisations and around three million employees, found that women receive less sponsorship, less support from their managers, and fewer opportunities that match their career goals. And when women do get the same support as men, the difference in how ambitious they are to move up disappears.

So the cycle goes like this:

•  The organisation gives out opportunities unevenly.

•  The potential assessment counts those opportunities as evidence.

•  The rating comes back lower for the people who were given less.

•  The organisation reads that rating as a fact about the person.

In other words, we’re reading our own past decisions back to ourselves, and calling it potential.


The question worth asking

None of this means managers are careless or that organisations don’t care. Most people running talent reviews are working hard, with good intentions, using the tools everyone else uses.

But the evidence now suggests the tools themselves need looking at. Not how consistently we use them – whether they’re measuring the right thing in the first place.

So here’s a question worth taking into your next talent review. For each thing you use to judge potential, ask:

Is this telling us something about the person – or something about the chances we’ve already given them?


In my next piece I’ll look at who tends to get missed. After ten years of working with people just below the first rung of leadership, I’ve found a pattern that’s remarkably consistent – and it looks almost nothing like what most potential criteria are designed to spot.

Sources

  • Benson, A., Li, D. and Shue, K. (2026) ‘Potential and the Gender Promotion Gap’, American Economic Review, 116(2)
  • Talogy (2025), global study on high-potential identification
  • McKinsey & Company and LeanIn.Org (2025), Women in the Workplace 2025
  • Korn Ferry, Assessment of Leadership Potential research guide