The single best predictor of job performance is not the aptitude test.
It is the structured interview.
If you have spent any time near hiring science, you have met the opposite claim. Cognitive ability is the best predictor of job performance. It has been repeated in textbooks, vendor decks, HR certifications and conference keynotes for most of my working life. It traces back to Schmidt and Hunter's 1998 summary in Psychological Bulletin, which put general mental ability at the top and then evaluated everything else by how much it added on top of that.
In 2022, Paul Sackett, Charlene Zhang, Christopher Berry and Filip Lievens went back and re-ran it. Their paper is in the Journal of Applied Psychology, volume 107, and it is not a hot take. It is twenty-nine pages of statistical housekeeping.
The revised ranking of what actually predicts job performance:
- Structured interviews, .42
- Job knowledge tests, .40
- Empirically keyed biodata, .38
- Work sample tests, .33
- Cognitive ability tests, .31
Cognitive ability came fifth.
What Actually Changed Since 1998
Nothing about interviews improved between 1998 and 2022. The old numbers were wrong.
Meta-analyses in this field correct for something called range restriction. The logic is sound. If you hire the top scorers on a test and then measure how well they perform, you are only looking at the winners, so the correlation you observe understates the true relationship. You correct upward to account for the slice you cannot see.
That correction is appropriate when a study followed real applicants who were screened on the test. It is not appropriate when a study administered the test to people already employed, who were never selected on it. There is no missing slice in that second design, so there is nothing to correct for.
Prior meta-analyses applied the correction across the board anyway, using restriction estimates drawn from the first kind of study to adjust the second kind. Sackett and his co-authors describe this and four related practices, and conclude that validity across the field has been substantially overestimated. Remove the over-correction and every estimate drops by .10 to .20 points.
Interviews came out on top not because they rose, but because they fell the least.
The Mechanism I Keep Seeing
The finding is four years old. The 1998 ranking is still in circulation.
That is not because anyone is being dishonest. It is because a number, once it becomes infrastructure, stops being read as a finding. It goes into a slide, the slide goes into a deck, the deck becomes the reason a company buys an assessment platform, and the platform becomes the process. By the time the underlying paper is revised, nobody in the chain is still looking at the paper.
This is the same shape as most of the hiring problems I run into. The artefact changed. The behaviour did not. Somebody edited the source and nobody edited the process built on top of it.
It is worth noticing that we are in the middle of another round of exactly this. AI screening and AI interview tools are being sold on validity claims right now. Most buyers are not in a position to evaluate those claims, and most will not go back and check in four years.
The Concession
I run a technical interview agency.
A paper that puts structured interviews at the top of the list is flattering to me, so let me be careful about what it does and does not say.
It does not say interviews are reliable. An operational validity of .42 means the interview explains a minority of the variance in how someone will perform. It is the best tool on the list, and the best tool on the list is still modest.
It also does not say that your interview is a .42 interview. The same paper found that structured interviews show unusually wide spread around that mean, wider than most other instruments. That is not a footnote. Structured interview covers everything from a rigorously job-analysed process with trained interviewers and behavioural anchors, through to a manager with four bullet points in a notebook. Only one of those earned the number.
And the honest summary of the whole paper is not that interviews win.
It is that everything we use is weaker than we thought, and we should be humbler about all of it.
What to Do Instead
Not a transformation programme. Three things, small enough to start next Monday.
- Write the questions before you meet anyone. The same core set for every candidate for that role. If you are inventing questions in the room, the thing you are doing is not the thing that scored .42.
- Write the scoring metrics at the same time. What does a strong answer to question three sound like? What does a partial one sound like? Agree it in advance, with someone who knows the domain, and you have removed most of the post-interview argument before it starts.
- Score each answer before the debrief, not during it. Independent scores first, discussion second. Otherwise the loudest person in the room is your instrument.
None of this is expensive. It is mostly an hour of preparation that currently is not happening.
The Question
The industry spends a great deal of energy trying to engineer the interview out of the process, on the grounds that it is the soft, subjective, unscientific part. The best available evidence says it is the strongest instrument we have, conditional on it being built properly.
So, honestly: is your interview the .42 kind, or the conversation kind?
Source: Sackett, P. R., Zhang, C., Berry, C. M., & Lievens, F. (2022). Revisiting meta-analytic estimates of validity in personnel selection: Addressing systematic overcorrection for restriction of range. Journal of Applied Psychology, 107(11), 2040–2068.
Want interviews that score closer to .42 than to a hunch?
See our sample interview structure, compare models on the Tech Assessments page, or book an interviewer to have us run the whole process for you.