Game-based assessments are neither a gimmick nor a proven replacement for classic psychometric tests. The peer-reviewed evidence shows that well-built games can measure cognitive ability with credible, moderate accuracy. It also shows that validity belongs to a specific score interpretation for a specific hiring decision, not to a format. A game that works for one role, one scoring model and one population tells you little about the next product you're shown in a demo.
This article sets out what the research says, how to read the effect sizes, where the evidence runs thin, and what a CHRO or HR director should demand before putting a game into a selection process that carries legal and reputational risk.
A game-based assessment (GBA) uses game mechanics, such as timed challenges, resource decisions, pattern tasks or risk choices, to elicit behaviour that is scored against a defined construct. In the academic literature the term "game-related assessment" (GRA) is often used, and "serious games assessment" covers games built for a purpose beyond entertainment.
The distinction that matters is between a game that looks like a game and a game that measures something. A visually engaging task with arbitrary scoring has no psychometric standing. A task designed to elicit working memory, abstract reasoning or learning speed, with scoring tied to evidence, can have it.
Classic tests, such as timed numerical or verbal reasoning questionnaires, have decades of validation behind them. Games have a shorter record. That gap in evidence, not the format itself, is the reasonable source of scepticism.
The most useful single source is the meta-analysis "The Relationship Between Game-Related Assessment and Traditional Measures of Cognitive Ability" (Journal of Intelligence, open access via PMC). It synthesised 44 papers and 52 samples, covering 807 effect sizes and more than 6,100 adult participants.
The overall observed correlation between GRAs and traditional cognitive ability measures was r = 0.30. After correcting for measurement error, the figure was r = 0.45.
What does that mean for hiring? A correlation of 0.30 to 0.45 is a moderate relationship. Games and classic tests overlap meaningfully, so games are not measuring nothing. But they are not interchangeable with established tests either. If two cognitive tests of the same construct were near-equivalent, you'd expect much higher figures. The result is consistent with games capturing part of general cognitive ability plus other things: motivation, interface familiarity, risk preference, or task-specific skills.
The paper "Game based assessments of cognitive ability in recruitment: Validity, fairness and test-taking experience" (Frontiers in Psychology, via PMC) addresses a practitioner setting directly. It reports:
These are respectable numbers, and the study's attention to adverse impact is a good model. Two cautions apply. It examines one game-based approach from one ecosystem, so it can't be generalised to the category. And its main validity evidence is alignment with cognitive tests, not prediction of job performance.
A further meta-analysis in the International Journal of Serious Games is explicitly framed around the convergent validity of game-based assessment. Its scope is the point: convergent validity asks whether a game correlates with other measures of the same construct. It does not ask whether scores predict who performs well in the job.
Each of these sources answers a different question. Constructs vary (reasoning, memory, attention, personality-adjacent traits). So do scoring models: a transparent points-based score behaves differently from a machine-learned composite. Samples differ in age, education, language and gaming familiarity. A pooled r of 0.45 is an average across very different games, and individual products will sit above and below it.
The AERA/APA/NCME Standards for Educational and Psychological Testing (2014) treat validity as a unitary concept. There aren't separate "types" you can tick off. Instead, multiple strands of evidence support the intended interpretation and use of test scores. The strands include:
The same framework applies to a 40-year-old verbal reasoning test and to a new game. Format doesn't earn exemption from any strand. This is also the SIOP position in its guidance on validating selection procedures: the evidence must support the specific inference being drawn in the specific context.
Reliability is the consistency of a score. For selection, unstable scores are a direct hiring risk: if a candidate would score differently on another day for reasons unrelated to ability, decisions based on that score are partly random.
The recruitment study's test-retest correlation of r = 0.68 is acceptable for a game-based approach but sits below the levels typically expected of mature, high-stakes cognitive tests, where reliability coefficients of .80 and above are common. For a screening stage with human review afterwards, 0.68 may be defensible. For a stand-alone pass/fail cut-off, it's thin. HR should ask vendors which reliability estimate they report (test-retest, internal consistency, alternate forms), over what interval, and in what sample.
Here the evidence is the strongest. The meta-analytic r = 0.30 observed and r = 0.45 corrected, and the r = 0.50 in the recruitment study, show real overlap with cognitive ability. In practical terms, a game scoring r = 0.50 with a classic test shares roughly a quarter of its variance with it (r squared = 0.25). The remainder is something else: construct-relevant signal, or noise, or irrelevant skill. Only further evidence distinguishes these.
Many buyer objections come from conflating two claims:
A moderate convergent correlation does not establish prediction of job performance. The published literature is considerably stronger on the former. Criterion-related evidence for game-based tools in real jobs is thinner, often confined to proprietary or vendor-commissioned studies.
There is an indirect argument: if a game measures cognitive ability, and cognitive ability predicts performance (as decades of work on classic tests show), then the game should inherit some predictive power. But that inference degrades with every step. A game correlating 0.45 with the construct will predict less than the classic test itself, unless it also captures something additional that matters on the job. HR should treat the indirect argument as a hypothesis to be tested locally, not a conclusion.
Response process evidence asks whether candidates are doing what the designer assumes. In a game, a high score might reflect fast reaction, familiarity with game conventions or a learned interface strategy rather than reasoning. Vendors can evidence this through think-aloud studies, process data analysis, or by showing that scores hold up after controlling for gaming experience and device type.
Games are most defensible when three conditions hold: the measured construct maps to the job's demands, the vendor shows reliability and convergent evidence at acceptable levels, and criterion-related validation (vendor-supplied, then confirmed locally) exists for comparable roles. When one is missing, use the game as one input among several rather than as a gate.
This is how Selection Lab approaches the format: games sit inside a role-adaptive flow alongside other assessment types and structured interview stages, rather than operating as a standalone novelty. A single instrument rarely carries a hiring decision on its own, and combining methods is generally more defensible than relying on one.
Generalisation. Samples in the published work are mostly adult, often student or general-population, and rarely stratified by seniority, culture or protected group. External validity across roles is largely untested.
Criterion evidence. As noted, convergent evidence dominates. Independent, peer-reviewed studies linking game scores to supervisor ratings, productivity or retention remain scarce.
Scoring transparency. Machine-learning scoring can improve prediction and allow fairness constraints, such as the bias penalisation described in the recruitment study. It also makes scores harder to explain. In an audit or challenge, "the model weighted it that way" is not a defence. Vendors need to provide audit trails, feature-level explanations where possible and subgroup analyses.
Operational artefacts. Practice effects, device differences and coaching can all move scores. These need monitoring after deployment, not only in the original validation.
Source independence. Some of the strongest practitioner-relevant studies come from within commercial ecosystems. That doesn't invalidate them, but it means independent replication and transparent methods carry extra weight.
A demo video is not evidence. Ask for a validity evidence pack that maps to the strands above.
Apply the four-fifths (80%) rule as a first screen: if the selection rate for any group is less than 80% of the rate for the highest-scoring group, the procedure is flagged for potential adverse impact under the US Uniform Guidelines. UK and EU contexts rely on different legal tests, but the four-fifths ratio remains a widely used monitoring heuristic for cognitive tests and other selection procedures. Treat it as a trigger for investigation and mitigation, not a verdict. Require documented mitigation steps, and re-run the analysis on your own applicant data, not only the vendor's.
Request documentation on data protection, retention periods, audit logs and human oversight of automated decisions. For European employers this includes GDPR compliance and alignment with the EU AI Act, under which recruitment tools are treated as high-risk. Selection Lab states GDPR compliance and EU AI Act alignment, with personal data stored in Frankfurt and local language models used to remove personal information from conversations. Whichever vendor you choose, ask to see the documentation behind such claims.
Before scaling, pilot with a defined role. Collect game scores alongside your existing selection data and, after a suitable period, performance and early-retention outcomes. Monitor subgroup selection rates throughout. Customer outcomes reported by Selection Lab, such as 21% lower early turnover (January 2024) and 15 minutes saved per applicant (December 2025), illustrate the kind of operational measures worth tracking, but they're vendor-reported and your pilot should establish your own baseline.
Integration matters too. Selection Lab states ATS integration and go-live in two to 10 weeks, which makes a time-boxed pilot realistic. Whatever the vendor, insist that pilot data remain accessible to you for analysis.
The evidence supports a measured conclusion. Game-related assessments correlate moderately with traditional cognitive ability (r = 0.30 observed, 0.45 corrected, across 44 papers and over 6,100 participants), and at least one recruitment study reports acceptable reliability (r = 0.68), a convergent validity of r = 0.50 and a strong candidate experience score. Calling that a gimmick ignores the data. Calling it equivalent to a validated classic test ignores what's missing: criterion-related evidence, independent replication, and subgroup fairness evidence for your own population.
Gaming is not the measurement. Validity evidence is. Buy the evidence first, and the format second.

Game-based assessments are neither a gimmick nor a proven replacement for classic psychometric tests. The peer-reviewed evidence shows that well-built games can measure cognitive ability with credible, moderate accuracy. It also shows that validity belongs to a specific score interpretation for a specific hiring decision, not to a format. A game that works for one role, one scoring model and one population tells you little about the next product you're shown in a demo.
This article sets out what the research says, how to read the effect sizes, where the evidence runs thin, and what a CHRO or HR director should demand before putting a game into a selection process that carries legal and reputational risk.
A game-based assessment (GBA) uses game mechanics, such as timed challenges, resource decisions, pattern tasks or risk choices, to elicit behaviour that is scored against a defined construct. In the academic literature the term "game-related assessment" (GRA) is often used, and "serious games assessment" covers games built for a purpose beyond entertainment.
The distinction that matters is between a game that looks like a game and a game that measures something. A visually engaging task with arbitrary scoring has no psychometric standing. A task designed to elicit working memory, abstract reasoning or learning speed, with scoring tied to evidence, can have it.
Classic tests, such as timed numerical or verbal reasoning questionnaires, have decades of validation behind them. Games have a shorter record. That gap in evidence, not the format itself, is the reasonable source of scepticism.
The most useful single source is the meta-analysis "The Relationship Between Game-Related Assessment and Traditional Measures of Cognitive Ability" (Journal of Intelligence, open access via PMC). It synthesised 44 papers and 52 samples, covering 807 effect sizes and more than 6,100 adult participants.
The overall observed correlation between GRAs and traditional cognitive ability measures was r = 0.30. After correcting for measurement error, the figure was r = 0.45.
What does that mean for hiring? A correlation of 0.30 to 0.45 is a moderate relationship. Games and classic tests overlap meaningfully, so games are not measuring nothing. But they are not interchangeable with established tests either. If two cognitive tests of the same construct were near-equivalent, you'd expect much higher figures. The result is consistent with games capturing part of general cognitive ability plus other things: motivation, interface familiarity, risk preference, or task-specific skills.
The paper "Game based assessments of cognitive ability in recruitment: Validity, fairness and test-taking experience" (Frontiers in Psychology, via PMC) addresses a practitioner setting directly. It reports:
These are respectable numbers, and the study's attention to adverse impact is a good model. Two cautions apply. It examines one game-based approach from one ecosystem, so it can't be generalised to the category. And its main validity evidence is alignment with cognitive tests, not prediction of job performance.
A further meta-analysis in the International Journal of Serious Games is explicitly framed around the convergent validity of game-based assessment. Its scope is the point: convergent validity asks whether a game correlates with other measures of the same construct. It does not ask whether scores predict who performs well in the job.
Each of these sources answers a different question. Constructs vary (reasoning, memory, attention, personality-adjacent traits). So do scoring models: a transparent points-based score behaves differently from a machine-learned composite. Samples differ in age, education, language and gaming familiarity. A pooled r of 0.45 is an average across very different games, and individual products will sit above and below it.
The AERA/APA/NCME Standards for Educational and Psychological Testing (2014) treat validity as a unitary concept. There aren't separate "types" you can tick off. Instead, multiple strands of evidence support the intended interpretation and use of test scores. The strands include:
The same framework applies to a 40-year-old verbal reasoning test and to a new game. Format doesn't earn exemption from any strand. This is also the SIOP position in its guidance on validating selection procedures: the evidence must support the specific inference being drawn in the specific context.
Reliability is the consistency of a score. For selection, unstable scores are a direct hiring risk: if a candidate would score differently on another day for reasons unrelated to ability, decisions based on that score are partly random.
The recruitment study's test-retest correlation of r = 0.68 is acceptable for a game-based approach but sits below the levels typically expected of mature, high-stakes cognitive tests, where reliability coefficients of .80 and above are common. For a screening stage with human review afterwards, 0.68 may be defensible. For a stand-alone pass/fail cut-off, it's thin. HR should ask vendors which reliability estimate they report (test-retest, internal consistency, alternate forms), over what interval, and in what sample.
Here the evidence is the strongest. The meta-analytic r = 0.30 observed and r = 0.45 corrected, and the r = 0.50 in the recruitment study, show real overlap with cognitive ability. In practical terms, a game scoring r = 0.50 with a classic test shares roughly a quarter of its variance with it (r squared = 0.25). The remainder is something else: construct-relevant signal, or noise, or irrelevant skill. Only further evidence distinguishes these.
Many buyer objections come from conflating two claims:
A moderate convergent correlation does not establish prediction of job performance. The published literature is considerably stronger on the former. Criterion-related evidence for game-based tools in real jobs is thinner, often confined to proprietary or vendor-commissioned studies.
There is an indirect argument: if a game measures cognitive ability, and cognitive ability predicts performance (as decades of work on classic tests show), then the game should inherit some predictive power. But that inference degrades with every step. A game correlating 0.45 with the construct will predict less than the classic test itself, unless it also captures something additional that matters on the job. HR should treat the indirect argument as a hypothesis to be tested locally, not a conclusion.
Response process evidence asks whether candidates are doing what the designer assumes. In a game, a high score might reflect fast reaction, familiarity with game conventions or a learned interface strategy rather than reasoning. Vendors can evidence this through think-aloud studies, process data analysis, or by showing that scores hold up after controlling for gaming experience and device type.
Games are most defensible when three conditions hold: the measured construct maps to the job's demands, the vendor shows reliability and convergent evidence at acceptable levels, and criterion-related validation (vendor-supplied, then confirmed locally) exists for comparable roles. When one is missing, use the game as one input among several rather than as a gate.
This is how Selection Lab approaches the format: games sit inside a role-adaptive flow alongside other assessment types and structured interview stages, rather than operating as a standalone novelty. A single instrument rarely carries a hiring decision on its own, and combining methods is generally more defensible than relying on one.
Generalisation. Samples in the published work are mostly adult, often student or general-population, and rarely stratified by seniority, culture or protected group. External validity across roles is largely untested.
Criterion evidence. As noted, convergent evidence dominates. Independent, peer-reviewed studies linking game scores to supervisor ratings, productivity or retention remain scarce.
Scoring transparency. Machine-learning scoring can improve prediction and allow fairness constraints, such as the bias penalisation described in the recruitment study. It also makes scores harder to explain. In an audit or challenge, "the model weighted it that way" is not a defence. Vendors need to provide audit trails, feature-level explanations where possible and subgroup analyses.
Operational artefacts. Practice effects, device differences and coaching can all move scores. These need monitoring after deployment, not only in the original validation.
Source independence. Some of the strongest practitioner-relevant studies come from within commercial ecosystems. That doesn't invalidate them, but it means independent replication and transparent methods carry extra weight.
A demo video is not evidence. Ask for a validity evidence pack that maps to the strands above.
Apply the four-fifths (80%) rule as a first screen: if the selection rate for any group is less than 80% of the rate for the highest-scoring group, the procedure is flagged for potential adverse impact under the US Uniform Guidelines. UK and EU contexts rely on different legal tests, but the four-fifths ratio remains a widely used monitoring heuristic for cognitive tests and other selection procedures. Treat it as a trigger for investigation and mitigation, not a verdict. Require documented mitigation steps, and re-run the analysis on your own applicant data, not only the vendor's.
Request documentation on data protection, retention periods, audit logs and human oversight of automated decisions. For European employers this includes GDPR compliance and alignment with the EU AI Act, under which recruitment tools are treated as high-risk. Selection Lab states GDPR compliance and EU AI Act alignment, with personal data stored in Frankfurt and local language models used to remove personal information from conversations. Whichever vendor you choose, ask to see the documentation behind such claims.
Before scaling, pilot with a defined role. Collect game scores alongside your existing selection data and, after a suitable period, performance and early-retention outcomes. Monitor subgroup selection rates throughout. Customer outcomes reported by Selection Lab, such as 21% lower early turnover (January 2024) and 15 minutes saved per applicant (December 2025), illustrate the kind of operational measures worth tracking, but they're vendor-reported and your pilot should establish your own baseline.
Integration matters too. Selection Lab states ATS integration and go-live in two to 10 weeks, which makes a time-boxed pilot realistic. Whatever the vendor, insist that pilot data remain accessible to you for analysis.
The evidence supports a measured conclusion. Game-related assessments correlate moderately with traditional cognitive ability (r = 0.30 observed, 0.45 corrected, across 44 papers and over 6,100 participants), and at least one recruitment study reports acceptable reliability (r = 0.68), a convergent validity of r = 0.50 and a strong candidate experience score. Calling that a gimmick ignores the data. Calling it equivalent to a validated classic test ignores what's missing: criterion-related evidence, independent replication, and subgroup fairness evidence for your own population.
Gaming is not the measurement. Validity evidence is. Buy the evidence first, and the format second.