Skills testing tools deliver provable ROI when they show three things at once. Predictive validity for the role you are filling, a measurable effect on time to hire and drop-off, and results from a customer that looks like you. Miss one and your business case is a guess, not a calculation.
That sounds like a low bar. It is not. Work through the assessment market asking every vendor for the number plus the sample it was calculated on, and you end up with a short list. This page shows you how to build that list yourself, and where we sit on it.
ROI on a skills test is not the same as feeling good about quality. It is the gap between what a better selection decision earns you and what the measuring costs. That gap sits in four places.
Recruiter time. Hours not spent scanning CVs and chasing no-shows. This is the easiest to measure, which is why it is often the only thing measured, which is why most business cases come out too small.
Funnel drop-off. Candidates who quit mid-application. Every drop-off you prevent is a sourcing euro you do not spend twice.
Quality of hire. First-year performance and the odds that someone stays. This is the biggest item and the hardest, because you need performance data to see it.
Mishires. The hire who leaves within six months. One prevented mishire often pays for an annual licence, provided you know what that mishire actually costs instead of guessing.
Build your case on recruiter time alone and you buy a tool that saves time and says nothing about who you hire. That is efficiency, not ROI. The distinction is not academic, because it decides which tool you should want.
The most recent large recalculation of selection methods is Sackett and colleagues, 2022. They corrected a systematic overestimate in the older Schmidt and Hunter figures from 1998. The resulting order.
| Selection method | Predictive validity |
|---|---|
| Structured interview | .42 |
| Job knowledge test | .40 |
| Empirically keyed biodata | .38 |
| Work sample test | .33 |
| Cognitive ability test | .31 |
| Interests, measured as fit with the job | .24 |
Two things stand out. The structured interview comes first, not the cognitive ability test. And interests measured as fit with a specific job do more than twice as well as interests as a general category, which sat at .10 in the older figures.
What that means for your choice. A tool that sells you a cognitive ability test alone buys you .31. A tool that pairs that test with a structured interview guide on the same competencies sits higher, because you are stacking two predictors instead of relying on one. That is the cheapest return in the whole hiring chain and the one most often skipped.
These four apply to us too. Further down you can see which of our own numbers survive them and which do not.
Below is only what was publicly available on 26 August 2026. "Not found" does not mean it does not exist, it means you have to ask for it before you count on it.
| Tool | What it measures | Public evidence | Where the ROI sits |
|---|---|---|---|
| Selection Lab | Cognitive ability, personality, language, hard and soft skills, plus screening and scheduling in one flow | Customer cases with a metric per customer. No per-test validity coefficients in public | Manual work removed and lower drop-off at volume |
| TestGorilla | Broad library you assemble your own test set from | Fact sheets per test with reliability and validity information, plus an explanation of their own validation studies. No coefficients or sample sizes on the science page itself | Speed and volume, provided you pick the right tests |
| Ixly | Dutch psychometric publisher, ability and personality | Test manuals online plus COTAN ratings per test | Roles where a bad hire is expensive |
| Aon Assessment, formerly cut-e | International test batteries for volume and cross-border hiring | No public validity report found. Manuals go through the vendor | International volume hiring on one norm |
| Codility and HackerRank | Work samples for code and technical tasks | Codility points to a technical manual and states plainly that it is making no claim about outcome tracking today. HackerRank publishes whitepapers per assessment type | Inside engineering, where the task sits close to the job |
One observation worth the detour. The only vendor in that table with an independent, publicly checkable review is Ixly, and the review is not spotless. COTAN, the Dutch national test review committee, rated Ixly's ACT general intelligence test good on theoretical foundation, test material and manual, adequate on reliability and construct validity, and inadequate on norms and criterion validity. Ixly wrote about it themselves, noting that criterion validity is very hard to establish in the field and that few Dutch intelligence tests reach an adequate rating there.
That is not a weakness of Ixly. That is what a real review looks like. Put it next to a vendor that prints "scientifically validated" on the homepage and publishes nothing, and you can see who has let themselves be checked.
Fill in your own numbers. The formula.
Annual return = (mishires prevented x cost per mishire) + (recruiter hours saved x hourly cost) + (extra placements x margin per placement)
ROI = (annual return minus annual licence cost minus implementation cost) divided by (annual licence cost plus implementation cost)
The item that decides the outcome is almost always the mishire. It is also where the uncertainty sits, because hardly anyone knows that figure for their own organisation. So there is no industry benchmark below. Put your own figure in and watch what it does to the answer.
| What you fill in | Where you get it |
|---|---|
| Hires per year | Your ATS |
| Leavers within six months | Your HR system |
| Cost per mishire | Sourcing cost plus onboarding time plus lost output plus the hiring manager's hours |
| Applicants per year | Your ATS |
| Minutes saved per applicant | Measure this in the pilot, do not take it from a demo |
| Recruiter hourly cost | Your own payroll |
Without the cost per mishire you cannot calculate the ROI of any tool. Establishing that number first is worth more than sitting through three demos. It takes an afternoon and it makes every vendor claim after it checkable.
Skills testing tools that use AI to evaluate candidates fall under Annex III of the EU AI Act and count as high risk. The Digital Omnibus on AI entered into force on 27 July 2026, after publication in the Official Journal on 24 July 2026. That moved the deadline for stand-alone Annex III systems from 2 August 2026 to 2 December 2027.
What does apply as of 2 August 2026 are the transparency obligations in Article 50. Candidates have to know they are dealing with an AI system, and synthetic output has to be marked in machine-readable form, with a transition period to 2 December 2026 for systems already on the market.
Two consequences for your business case. You have longer than you thought to get the documentation in order, and it has become a purchasing criterion. A vendor who cannot show you risk management and technical documentation now is passing that cost to you.
We do not publish per-test validity coefficients. Anyone who wants to compare at that level has to ask us for it, the same as with every other vendor in the table above. What we do have is results per customer. One of them measures quality of hire. The rest measure process.
The number that is about quality. At Dentons the quality-of-hire ratio rose by 13.5%. Of the candidates assessed, 32% were hired and 77% of those hires performed at or above average, measured with performance data after they started. Candidates scoring higher on morality, self-control and enthusiasm performed better.
The numbers that are about process. Carrefour saves 6 hours per recruiter per week. DPD saw the cost of the hiring process come out 15% lower. CoBuilders improved its placement ratio by 8%. Welten saves 6,700 minutes a month. FrieslandCampina hired 40 trainees within a month, out of 23,000 applications across 18 countries. At IG&H, 603 candidates were assessed, with 94% positive candidate feedback.
That second list says something about time and cost and nothing about who got hired. Which is the exact distinction this page is about, and it applies to us as well. If you want to know whether a tool makes your hires better, there is only one route. Keep the test scores, put them next to performance data a year later, and draw your conclusions then.
The ones that can show you three things. A validity figure with the sample attached, a completion rate from live use, and customer numbers with the metric attached. Which tool wins depends on your volume and on what a mishire costs you. For engineering roles that is a coding platform, for volume hiring an automated selection platform, for expensive key roles a psychometric publisher.
Set the return from prevented mishires, saved recruiter hours and extra placements against licence and implementation costs. The item that decides the outcome is almost always the mishire. If you do not know that cost, you are not calculating, you are estimating.
Better than a CV, but not better than a structured interview. In the Sackett and colleagues recalculation from 2022 the structured interview reaches .42, a job knowledge test .40 and a cognitive ability test .31. The strongest combination is a test plus a structured interview on the same competencies.
That depends on your hiring volume, not on the tool. Time savings show up immediately. Effect on quality of hire takes six to twelve months, because you need performance data from the new hires. Agree upfront which data you put side by side and when.
AI systems that evaluate candidates fall under Annex III and count as high risk. The Digital Omnibus on AI moved that deadline to 2 December 2027. The Article 50 transparency obligations have applied since 2 August 2026, so candidates already have to know an AI system is involved.
Comparing tools rather than building the business case? Start with the comparison of skill test tools, see how to improve your quality of hire, or read the Dentons case.
.png)
Skills testing tools deliver provable ROI when they show three things at once. Predictive validity for the role you are filling, a measurable effect on time to hire and drop-off, and results from a customer that looks like you. Miss one and your business case is a guess, not a calculation.
That sounds like a low bar. It is not. Work through the assessment market asking every vendor for the number plus the sample it was calculated on, and you end up with a short list. This page shows you how to build that list yourself, and where we sit on it.
ROI on a skills test is not the same as feeling good about quality. It is the gap between what a better selection decision earns you and what the measuring costs. That gap sits in four places.
Recruiter time. Hours not spent scanning CVs and chasing no-shows. This is the easiest to measure, which is why it is often the only thing measured, which is why most business cases come out too small.
Funnel drop-off. Candidates who quit mid-application. Every drop-off you prevent is a sourcing euro you do not spend twice.
Quality of hire. First-year performance and the odds that someone stays. This is the biggest item and the hardest, because you need performance data to see it.
Mishires. The hire who leaves within six months. One prevented mishire often pays for an annual licence, provided you know what that mishire actually costs instead of guessing.
Build your case on recruiter time alone and you buy a tool that saves time and says nothing about who you hire. That is efficiency, not ROI. The distinction is not academic, because it decides which tool you should want.
The most recent large recalculation of selection methods is Sackett and colleagues, 2022. They corrected a systematic overestimate in the older Schmidt and Hunter figures from 1998. The resulting order.
| Selection method | Predictive validity |
|---|---|
| Structured interview | .42 |
| Job knowledge test | .40 |
| Empirically keyed biodata | .38 |
| Work sample test | .33 |
| Cognitive ability test | .31 |
| Interests, measured as fit with the job | .24 |
Two things stand out. The structured interview comes first, not the cognitive ability test. And interests measured as fit with a specific job do more than twice as well as interests as a general category, which sat at .10 in the older figures.
What that means for your choice. A tool that sells you a cognitive ability test alone buys you .31. A tool that pairs that test with a structured interview guide on the same competencies sits higher, because you are stacking two predictors instead of relying on one. That is the cheapest return in the whole hiring chain and the one most often skipped.
These four apply to us too. Further down you can see which of our own numbers survive them and which do not.
Below is only what was publicly available on 26 August 2026. "Not found" does not mean it does not exist, it means you have to ask for it before you count on it.
| Tool | What it measures | Public evidence | Where the ROI sits |
|---|---|---|---|
| Selection Lab | Cognitive ability, personality, language, hard and soft skills, plus screening and scheduling in one flow | Customer cases with a metric per customer. No per-test validity coefficients in public | Manual work removed and lower drop-off at volume |
| TestGorilla | Broad library you assemble your own test set from | Fact sheets per test with reliability and validity information, plus an explanation of their own validation studies. No coefficients or sample sizes on the science page itself | Speed and volume, provided you pick the right tests |
| Ixly | Dutch psychometric publisher, ability and personality | Test manuals online plus COTAN ratings per test | Roles where a bad hire is expensive |
| Aon Assessment, formerly cut-e | International test batteries for volume and cross-border hiring | No public validity report found. Manuals go through the vendor | International volume hiring on one norm |
| Codility and HackerRank | Work samples for code and technical tasks | Codility points to a technical manual and states plainly that it is making no claim about outcome tracking today. HackerRank publishes whitepapers per assessment type | Inside engineering, where the task sits close to the job |
One observation worth the detour. The only vendor in that table with an independent, publicly checkable review is Ixly, and the review is not spotless. COTAN, the Dutch national test review committee, rated Ixly's ACT general intelligence test good on theoretical foundation, test material and manual, adequate on reliability and construct validity, and inadequate on norms and criterion validity. Ixly wrote about it themselves, noting that criterion validity is very hard to establish in the field and that few Dutch intelligence tests reach an adequate rating there.
That is not a weakness of Ixly. That is what a real review looks like. Put it next to a vendor that prints "scientifically validated" on the homepage and publishes nothing, and you can see who has let themselves be checked.
Fill in your own numbers. The formula.
Annual return = (mishires prevented x cost per mishire) + (recruiter hours saved x hourly cost) + (extra placements x margin per placement)
ROI = (annual return minus annual licence cost minus implementation cost) divided by (annual licence cost plus implementation cost)
The item that decides the outcome is almost always the mishire. It is also where the uncertainty sits, because hardly anyone knows that figure for their own organisation. So there is no industry benchmark below. Put your own figure in and watch what it does to the answer.
| What you fill in | Where you get it |
|---|---|
| Hires per year | Your ATS |
| Leavers within six months | Your HR system |
| Cost per mishire | Sourcing cost plus onboarding time plus lost output plus the hiring manager's hours |
| Applicants per year | Your ATS |
| Minutes saved per applicant | Measure this in the pilot, do not take it from a demo |
| Recruiter hourly cost | Your own payroll |
Without the cost per mishire you cannot calculate the ROI of any tool. Establishing that number first is worth more than sitting through three demos. It takes an afternoon and it makes every vendor claim after it checkable.
Skills testing tools that use AI to evaluate candidates fall under Annex III of the EU AI Act and count as high risk. The Digital Omnibus on AI entered into force on 27 July 2026, after publication in the Official Journal on 24 July 2026. That moved the deadline for stand-alone Annex III systems from 2 August 2026 to 2 December 2027.
What does apply as of 2 August 2026 are the transparency obligations in Article 50. Candidates have to know they are dealing with an AI system, and synthetic output has to be marked in machine-readable form, with a transition period to 2 December 2026 for systems already on the market.
Two consequences for your business case. You have longer than you thought to get the documentation in order, and it has become a purchasing criterion. A vendor who cannot show you risk management and technical documentation now is passing that cost to you.
We do not publish per-test validity coefficients. Anyone who wants to compare at that level has to ask us for it, the same as with every other vendor in the table above. What we do have is results per customer. One of them measures quality of hire. The rest measure process.
The number that is about quality. At Dentons the quality-of-hire ratio rose by 13.5%. Of the candidates assessed, 32% were hired and 77% of those hires performed at or above average, measured with performance data after they started. Candidates scoring higher on morality, self-control and enthusiasm performed better.
The numbers that are about process. Carrefour saves 6 hours per recruiter per week. DPD saw the cost of the hiring process come out 15% lower. CoBuilders improved its placement ratio by 8%. Welten saves 6,700 minutes a month. FrieslandCampina hired 40 trainees within a month, out of 23,000 applications across 18 countries. At IG&H, 603 candidates were assessed, with 94% positive candidate feedback.
That second list says something about time and cost and nothing about who got hired. Which is the exact distinction this page is about, and it applies to us as well. If you want to know whether a tool makes your hires better, there is only one route. Keep the test scores, put them next to performance data a year later, and draw your conclusions then.
The ones that can show you three things. A validity figure with the sample attached, a completion rate from live use, and customer numbers with the metric attached. Which tool wins depends on your volume and on what a mishire costs you. For engineering roles that is a coding platform, for volume hiring an automated selection platform, for expensive key roles a psychometric publisher.
Set the return from prevented mishires, saved recruiter hours and extra placements against licence and implementation costs. The item that decides the outcome is almost always the mishire. If you do not know that cost, you are not calculating, you are estimating.
Better than a CV, but not better than a structured interview. In the Sackett and colleagues recalculation from 2022 the structured interview reaches .42, a job knowledge test .40 and a cognitive ability test .31. The strongest combination is a test plus a structured interview on the same competencies.
That depends on your hiring volume, not on the tool. Time savings show up immediately. Effect on quality of hire takes six to twelve months, because you need performance data from the new hires. Agree upfront which data you put side by side and when.
AI systems that evaluate candidates fall under Annex III and count as high risk. The Digital Omnibus on AI moved that deadline to 2 December 2027. The Article 50 transparency obligations have applied since 2 August 2026, so candidates already have to know an AI system is involved.
Comparing tools rather than building the business case? Start with the comparison of skill test tools, see how to improve your quality of hire, or read the Dentons case.