Reading Time
4 min

Game-based assessment validity: research vs. gimmick in hiring

Game-based assessments are neither a gimmick nor a proven replacement for classic psychometric tests. The peer-reviewed evidence shows that well-built games can measure cognitive ability with credible, moderate accuracy. It also shows that validity belongs to a specific score interpretation for a specific hiring decision, not to a format. A game that works for one role, one scoring model and one population tells you little about the next product you're shown in a demo.

This article sets out what the research says, how to read the effect sizes, where the evidence runs thin, and what a CHRO or HR director should demand before putting a game into a selection process that carries legal and reputational risk.

What game-based assessment means in selection

A game-based assessment (GBA) uses game mechanics, such as timed challenges, resource decisions, pattern tasks or risk choices, to elicit behaviour that is scored against a defined construct. In the academic literature the term "game-related assessment" (GRA) is often used, and "serious games assessment" covers games built for a purpose beyond entertainment.

The distinction that matters is between a game that looks like a game and a game that measures something. A visually engaging task with arbitrary scoring has no psychometric standing. A task designed to elicit working memory, abstract reasoning or learning speed, with scoring tied to evidence, can have it.

Classic tests, such as timed numerical or verbal reasoning questionnaires, have decades of validation behind them. Games have a shorter record. That gap in evidence, not the format itself, is the reasonable source of scepticism.

What the key studies and meta-analyses show

The Journal of Intelligence meta-analysis

The most useful single source is the meta-analysis "The Relationship Between Game-Related Assessment and Traditional Measures of Cognitive Ability" (Journal of Intelligence, open access via PMC). It synthesised 44 papers and 52 samples, covering 807 effect sizes and more than 6,100 adult participants.

The overall observed correlation between GRAs and traditional cognitive ability measures was r = 0.30. After correcting for measurement error, the figure was r = 0.45.

What does that mean for hiring? A correlation of 0.30 to 0.45 is a moderate relationship. Games and classic tests overlap meaningfully, so games are not measuring nothing. But they are not interchangeable with established tests either. If two cognitive tests of the same construct were near-equivalent, you'd expect much higher figures. The result is consistent with games capturing part of general cognitive ability plus other things: motivation, interface familiarity, risk preference, or task-specific skills.

The Frontiers in Psychology recruitment study

The paper "Game based assessments of cognitive ability in recruitment: Validity, fairness and test-taking experience" (Frontiers in Psychology, via PMC) addresses a practitioner setting directly. It reports:

  • Test-retest reliability of r = 0.68 for the game-based scoring approach.
  • A convergent validity result of r = 0.50 against traditional cognitive measures.
  • A Net Promoter Score of 58 for the test-taking experience.
  • Fairness modelling using machine-learning scoring, described as ridge regression with fairness-optimised bias penalisation.

These are respectable numbers, and the study's attention to adverse impact is a good model. Two cautions apply. It examines one game-based approach from one ecosystem, so it can't be generalised to the category. And its main validity evidence is alignment with cognitive tests, not prediction of job performance.

The International Journal of Serious Games meta-analysis

A further meta-analysis in the International Journal of Serious Games is explicitly framed around the convergent validity of game-based assessment. Its scope is the point: convergent validity asks whether a game correlates with other measures of the same construct. It does not ask whether scores predict who performs well in the job.

Why the studies aren't interchangeable

Each of these sources answers a different question. Constructs vary (reasoning, memory, attention, personality-adjacent traits). So do scoring models: a transparent points-based score behaves differently from a machine-learned composite. Samples differ in age, education, language and gaming familiarity. A pooled r of 0.45 is an average across very different games, and individual products will sit above and below it.

Validity is an argument, not a format

The AERA/APA/NCME Standards for Educational and Psychological Testing (2014) treat validity as a unitary concept. There aren't separate "types" you can tick off. Instead, multiple strands of evidence support the intended interpretation and use of test scores. The strands include:

  1. Test content: does the task represent the construct and the job demands?
  2. Response processes: do candidates use the intended cognitive or behavioural processes?
  3. Internal structure: do items or tasks behave consistently with the theory of the construct?
  4. Relations to other variables: convergent evidence, discriminant evidence and relations to job criteria.
  5. Consequences of testing: including fairness, adverse impact and candidate experience.

The same framework applies to a 40-year-old verbal reasoning test and to a new game. Format doesn't earn exemption from any strand. This is also the SIOP position in its guidance on validating selection procedures: the evidence must support the specific inference being drawn in the specific context.

Comparative metrics: reliability, construct validity and predictive validity

Reliability

Reliability is the consistency of a score. For selection, unstable scores are a direct hiring risk: if a candidate would score differently on another day for reasons unrelated to ability, decisions based on that score are partly random.

The recruitment study's test-retest correlation of r = 0.68 is acceptable for a game-based approach but sits below the levels typically expected of mature, high-stakes cognitive tests, where reliability coefficients of .80 and above are common. For a screening stage with human review afterwards, 0.68 may be defensible. For a stand-alone pass/fail cut-off, it's thin. HR should ask vendors which reliability estimate they report (test-retest, internal consistency, alternate forms), over what interval, and in what sample.

Construct and convergent validity

Here the evidence is the strongest. The meta-analytic r = 0.30 observed and r = 0.45 corrected, and the r = 0.50 in the recruitment study, show real overlap with cognitive ability. In practical terms, a game scoring r = 0.50 with a classic test shares roughly a quarter of its variance with it (r squared = 0.25). The remainder is something else: construct-relevant signal, or noise, or irrelevant skill. Only further evidence distinguishes these.

Predictive validity is a separate question

Many buyer objections come from conflating two claims:

  • Convergent validity: the game correlates with a cognitive test.
  • Predictive (criterion-related) validity: the game's scores predict job performance, training success, or retention.

A moderate convergent correlation does not establish prediction of job performance. The published literature is considerably stronger on the former. Criterion-related evidence for game-based tools in real jobs is thinner, often confined to proprietary or vendor-commissioned studies.

There is an indirect argument: if a game measures cognitive ability, and cognitive ability predicts performance (as decades of work on classic tests show), then the game should inherit some predictive power. But that inference degrades with every step. A game correlating 0.45 with the construct will predict less than the classic test itself, unless it also captures something additional that matters on the job. HR should treat the indirect argument as a hypothesis to be tested locally, not a conclusion.

Response processes

Response process evidence asks whether candidates are doing what the designer assumes. In a game, a high score might reflect fast reaction, familiarity with game conventions or a learned interface strategy rather than reasoning. Vendors can evidence this through think-aloud studies, process data analysis, or by showing that scores hold up after controlling for gaming experience and device type.

Where games outperform and underperform classic tests

Where games can have the edge

  • Engagement and completion. If fewer candidates abandon the process, the applicant pool is less distorted by drop-out. Selection Lab's own customer data points this way: its 2026 deck cites 27% fewer drop-offs (March 2025). That's a vendor-reported outcome, not an independent finding, and drop-off reduction depends on the whole flow, not on the game alone.
  • Candidate experience. The recruitment study's NPS of 58 is a strong score for a test-taking experience, as tests are rarely rated highly.
  • Construct breadth. Interactive tasks can capture processes, such as learning from feedback or adapting to changing rules, that a static multiple-choice item can't.
  • Test anxiety. Plausible but under-evidenced. Treat claims of reduced anxiety as needing data, not as established.

Where games can fall short

  • Applicant reactions. Reactions to game-based tests are not uniformly positive. Some studies find less favourable perceptions than for traditional formats, and gaming experience can moderate those perceptions. Paul Englert's discussion of one such study raises the question directly: do these tools measure cognitive ability or acceptance of gaming?
  • Fairness risk. If prior gaming experience, age, or device (touchscreen versus mouse) shifts scores independently of ability, a game can introduce construct-irrelevant variance and adverse impact.
  • Criterion gaps. Without job-performance validation, a game is a plausible cognitive proxy at best.

A practical decision rule

Games are most defensible when three conditions hold: the measured construct maps to the job's demands, the vendor shows reliability and convergent evidence at acceptable levels, and criterion-related validation (vendor-supplied, then confirmed locally) exists for comparable roles. When one is missing, use the game as one input among several rather than as a gate.

This is how Selection Lab approaches the format: games sit inside a role-adaptive flow alongside other assessment types and structured interview stages, rather than operating as a standalone novelty. A single instrument rarely carries a hiring decision on its own, and combining methods is generally more defensible than relying on one.

Limitations and open questions

Generalisation. Samples in the published work are mostly adult, often student or general-population, and rarely stratified by seniority, culture or protected group. External validity across roles is largely untested.

Criterion evidence. As noted, convergent evidence dominates. Independent, peer-reviewed studies linking game scores to supervisor ratings, productivity or retention remain scarce.

Scoring transparency. Machine-learning scoring can improve prediction and allow fairness constraints, such as the bias penalisation described in the recruitment study. It also makes scores harder to explain. In an audit or challenge, "the model weighted it that way" is not a defence. Vendors need to provide audit trails, feature-level explanations where possible and subgroup analyses.

Operational artefacts. Practice effects, device differences and coaching can all move scores. These need monitoring after deployment, not only in the original validation.

Source independence. Some of the strongest practitioner-relevant studies come from within commercial ecosystems. That doesn't invalidate them, but it means independent replication and transparent methods carry extra weight.

What recruiters and HR leaders should demand before deployment

A demo video is not evidence. Ask for a validity evidence pack that maps to the strands above.

The evidence checklist

  • Construct and content: a documented definition of what the game measures and a job-analysis rationale for why it matters for the role.
  • Response processes: evidence that scores reflect the intended processes, including analysis of gaming experience and device effects.
  • Reliability: test-retest and internal consistency figures, with sample, interval and method. Compare against the stakes of your decision.
  • Convergent validity: correlations with established cognitive tests, ideally in more than one sample. Treat a figure around 0.50 as moderate.
  • Criterion-related validity: correlations with job performance, training outcomes or retention in comparable roles, then a plan for local validation.
  • Fairness: subgroup mean differences and selection rates by gender, age, ethnicity and other relevant groups.
  • Candidate experience: completion rates, applicant reaction data and fairness perceptions.

Adverse impact auditing

Apply the four-fifths (80%) rule as a first screen: if the selection rate for any group is less than 80% of the rate for the highest-scoring group, the procedure is flagged for potential adverse impact under the US Uniform Guidelines. UK and EU contexts rely on different legal tests, but the four-fifths ratio remains a widely used monitoring heuristic for cognitive tests and other selection procedures. Treat it as a trigger for investigation and mitigation, not a verdict. Require documented mitigation steps, and re-run the analysis on your own applicant data, not only the vendor's.

Governance and data

Request documentation on data protection, retention periods, audit logs and human oversight of automated decisions. For European employers this includes GDPR compliance and alignment with the EU AI Act, under which recruitment tools are treated as high-risk. Selection Lab states GDPR compliance and EU AI Act alignment, with personal data stored in Frankfurt and local language models used to remove personal information from conversations. Whichever vendor you choose, ask to see the documentation behind such claims.

Run a local pilot

Before scaling, pilot with a defined role. Collect game scores alongside your existing selection data and, after a suitable period, performance and early-retention outcomes. Monitor subgroup selection rates throughout. Customer outcomes reported by Selection Lab, such as 21% lower early turnover (January 2024) and 15 minutes saved per applicant (December 2025), illustrate the kind of operational measures worth tracking, but they're vendor-reported and your pilot should establish your own baseline.

Integration matters too. Selection Lab states ATS integration and go-live in two to 10 weeks, which makes a time-boxed pilot realistic. Whatever the vendor, insist that pilot data remain accessible to you for analysis.

The defensible position

The evidence supports a measured conclusion. Game-related assessments correlate moderately with traditional cognitive ability (r = 0.30 observed, 0.45 corrected, across 44 papers and over 6,100 participants), and at least one recruitment study reports acceptable reliability (r = 0.68), a convergent validity of r = 0.50 and a strong candidate experience score. Calling that a gimmick ignores the data. Calling it equivalent to a validated classic test ignores what's missing: criterion-related evidence, independent replication, and subgroup fairness evidence for your own population.

Gaming is not the measurement. Validity evidence is. Buy the evidence first, and the format second.

FAQ

Can game-based assessments promote diversity in the hiring process?

Yes, game-based assessments can support diversity by focusing on skills and behaviors rather than traditional criteria like résumés, which may contain unconscious biases. This gives candidates from diverse backgrounds a fairer chance to demonstrate their potential.

What is a game-based assessment?

A game-based assessment is a method that uses game mechanics to evaluate a candidate’s skills, competencies, and personality traits. While playing these games, candidates are assessed on aspects like problem-solving, cognitive ability, and behavior under pressure in an interactive way.

What are the advantages of game-based assessments?

Game-based assessments offer a more engaging and interactive experience for candidates, which can lead to a more positive perception of the hiring process—especially among certain groups. For employers, they provide deeper insights into both cognitive and behavioral traits, which traditional tests may miss. They also reduce the chance of socially desirable answers, as candidates tend to respond more authentically in a game environment.

How reliable are game-based assessments compared to traditional tests?

When well-designed, game-based assessments can be just as reliable—or even more reliable—than traditional tests. They assess a wide range of behaviors and cognitive abilities in a dynamic setting. However, the quality of these assessments varies greatly, so careful evaluation is essential.

How does a game-based assessment work?

Candidates participate in interactive games designed to measure specific skills and behaviors. Evaluation goes beyond just the final score—it also considers how the candidate makes decisions, handles challenges, and responds to different scenarios. These insights reveal underlying thought processes and behavioral patterns.

Are game-based assessments scientifically validated?

The main drawback is that many game-based assessments are relatively new and have not yet been extensively researched by independent academics. Providers often cite their own research, which is rarely externally validated. Without independent studies, the reliability of these assessments remains uncertain—something to keep in mind when selecting one.

How can game based assessments contribute to a better candidate experience

This can vary significantly by audience. The playful, interactive nature of game-based assessments can lower stress levels for some candidates compared to traditional tests. However, research shows that certain groups, especially those over 35, may find them more stressful. Men also tend to rate the experience more positively than women.

Can you practice game-based assessment?

You can familiarize yourself with the style of games used, but it’s difficult to "practice" for them in a traditional sense. These assessments are designed to measure natural reactions and authentic behavior, so repeated practice typically has less effect on performance than with traditional tests.

Will game-based assessments replace traditional tests in the future?

It’s likely that game-based assessments will become more common in hiring processes, but they probably won’t fully replace traditional tests. Both approaches have value and can complement each other depending on the role and the company’s needs.

How are the results of a game-based assessment analyzed?

Results are analyzed based on predefined criteria such as problem-solving ability, reaction time, and behavior under pressure. Advanced algorithms collect and interpret this data to provide a reliable, objective evaluation of a candidate’s strengths.

What kind of skills do game-based assessments measure?

They assess a wide range of abilities, including problem-solving, adaptability, decision-making under pressure, teamwork, and emotional intelligence. Depending on the design, they may also evaluate cognitive skills like memory, attention, and pattern recognition.

How long does a game-based assessment take?

Typically, these assessments last between 15 and 60 minutes, depending on the game’s complexity and the number of skills being tested. They’re usually shorter and more engaging than traditional assessments, making for a smoother candidate experience.

Are game-based assessments suitable for all roles?

They are especially effective for roles that require flexibility, creativity, problem-solving, and strong interpersonal skills. For highly technical or specialized roles, additional assessments may be needed to measure specific knowledge.

What’s the difference between a game-based and a gamified assessment?

A gamified assessment adds game-like elements (such as points or rewards) to a traditional test to increase engagement. A game-based assessment, on the other hand, is a standalone game designed specifically to evaluate certain competencies. The game itself is the primary evaluation tool, not just an enhancement.

FAQ

How can I improve my company’s retention rate?

The retention rate can be improved by investing in employee development and satisfaction. This includes offering training, career opportunities, and recognition for their contributions. A culture of open communication and attention to work-life balance can also contribute to higher retention. Additionally, offering competitive compensation and involving employees in decision-making can strengthen loyalty.

What are the benefits of growth opportunities for employee retention?

Growth opportunities can promote employee retention by giving staff a sense of direction and motivation. When they have the chance to learn and develop professionally within the company, they feel valued, which increases their loyalty. This can prevent them from leaving to seek better opportunities elsewhere. kunnen het behoud van personeel bevorderen door medewerkers een gevoel van richting en motivatie te geven. Wanneer zij de kans krijgen om te leren en zich professioneel te ontwikkelen binnen het bedrijf, voelen zij zich gewaardeerd, wat hun loyaliteit vergroot. Dit kan voorkomen dat ze vertrekken om elders betere kansen te zoeken.

What are the key factors that influence employee retention?

Key factors that influence employee retention include salary and benefits, opportunities for professional development, work-life balance, company culture, and the relationship with supervisors. Employees tend to stay longer when they feel valued, challenged, and supported in their work environment.

Why is employee retention so important for organizations?

Employee retention is important because it helps reduce recruitment and training costs for new employees, and it contributes to retaining knowledge and experience within the organization. High retention also ensures continuity within teams, leading to a more stable company culture, higher customer satisfaction, and improved business outcomes.

Which recruitment strategies help improve retention?

Recruitment strategies that can improve retention include identifying candidates who align with the company culture, using assessments to evaluate soft skills, and providing transparency about role expectations during the hiring process. Employees who feel connected to the organization and have clarity about their role are more likely to stay longer.

How can a good onboarding process contribute to higher retention?

An effective onboarding process can contribute to higher retention by helping new employees quickly adapt to their role, the company culture, and expectations. By providing support and clear information from the start, their engagement is increased, and the likelihood of them leaving early due to feelings of being overwhelmed or lacking guidance is reduced.

What is the role of company culture in retaining employees?

Company culture plays a crucial role in employee retention. When employees feel heard, valued, and connected to the values and norms of the company, they are more likely to stay. A positive culture that fosters collaboration, respect, and personal growth can significantly enhance employee motivation and satisfaction.

How can leadership and management style influence retention?

Leadership and management style have a significant impact on retention. Leaders who inspire, support, and coach their team can increase employee engagement and satisfaction. Offering autonomy and trust can lead to higher loyalty, while inefficient or negative management styles can contribute to dissatisfaction and increased employee turnover.

What is the importance of recognition and rewards for employee retention?

Recognition and rewards play an important role in employee retention by showing staff that their work is valued. This can increase their motivation and loyalty. In addition to financial rewards, compliments, promotions, and other forms of recognition can also contribute to satisfaction and retaining employees.

What role does work-life balance play in improving retention?

A balanced work-life balance plays an important role in increasing retention. By reducing stress and improving job satisfaction, employees are more likely to stay with the company. Initiatives such as flexible working hours, remote work options, and respect for personal time can contribute to this balance.

What does increasing retention mean within a company?

Increasing retention within a company means implementing strategies to keep employees with the organization for longer. This can be achieved by improving job satisfaction, offering growth opportunities, and fostering a positive and supportive company culture.

How do I measure the success of my retention strategy?

The success of a retention strategy can be measured by tracking retention rates and turnover rates, and by gaining insights from exit interviews. Additionally, employee satisfaction surveys and feedback from performance evaluations can provide valuable information about the effectiveness of the strategies applied.

What are the costs of a low retention rate?

A low retention rate can bring significant costs, such as increased expenses for recruiting and training new employees. Furthermore, the loss of experienced staff can lead to lower productivity, reduced knowledge transfer, and a negative impact on company culture.

How can I increase employee engagement?

To increase employee engagement, involve them in decision-making processes, regularly ask for their feedback, and recognize their contributions. Offering development opportunities and maintaining transparent communication can also contribute to greater engagement.

How can technology help improve employee retention?

Technology can be a tool for improving employee retention by facilitating communication, feedback, and development. By using online platforms for training, recognition, and evaluation, companies can create a more engaged and satisfied workforce.

FAQ

How long does it take to complete the tool?

Less than 10 minutes. You’ll answer 30 guided questions and get a summary of what to look for in your next assessment platform.

Can this checklist help me compare assessment providers?

Yes. By clarifying what matters most to your team, it makes comparing providers' features, pricing, and strengths much easier and more strategic.

How can I use this checklist if I’m not doing a formal RFI?

It’s equally valuable for internal evaluations, exploring new tools, or improving your current hiring process even if you’re not issuing an RFI or RFQ.

What should I look for in a modern assessment tool?

Prioritize platforms with user-friendly design, mobile compatibility, strong analytics, ATS integrations, and inclusive features like neurodiversity support.

What types of assessments should I consider in 2025?

Leading tools combine cognitive testing, situational judgment tests (SJTs), behavior assessments, and predictive AI to evaluate candidates more holistically.

Who should use an assessment checklist?

HR professionals, hiring managers, and procurement teams evaluating pre-selection solutions, especially those comparing AI-powered or compliance-driven assessment platforms.

How does this checklist help with RFIs and RFQs for assessments?

The checklist helps you define your exact requirements so you can confidently draft or respond to Requests for Information (RFI) or Requests for Quotation (RFQ) for assessment tools.

What is an assessment tool in hiring?

An assessment tool evaluates candidates’ skills, behaviors, and fit during the recruitment process. It helps improve hiring decisions and streamline pre-selection.

Game-based assessment packs

← Our Blog

Game-based assessment validity: research vs. gimmick in hiring

Explore the research on game-based assessments vs. classic tests. See effect sizes, reliability, and what validity evidence HR should demand before deployment.
Joeri Everaers
COO
Read time: Approx
4 min

Game-based assessments are neither a gimmick nor a proven replacement for classic psychometric tests. The peer-reviewed evidence shows that well-built games can measure cognitive ability with credible, moderate accuracy. It also shows that validity belongs to a specific score interpretation for a specific hiring decision, not to a format. A game that works for one role, one scoring model and one population tells you little about the next product you're shown in a demo.

This article sets out what the research says, how to read the effect sizes, where the evidence runs thin, and what a CHRO or HR director should demand before putting a game into a selection process that carries legal and reputational risk.

What game-based assessment means in selection

A game-based assessment (GBA) uses game mechanics, such as timed challenges, resource decisions, pattern tasks or risk choices, to elicit behaviour that is scored against a defined construct. In the academic literature the term "game-related assessment" (GRA) is often used, and "serious games assessment" covers games built for a purpose beyond entertainment.

The distinction that matters is between a game that looks like a game and a game that measures something. A visually engaging task with arbitrary scoring has no psychometric standing. A task designed to elicit working memory, abstract reasoning or learning speed, with scoring tied to evidence, can have it.

Classic tests, such as timed numerical or verbal reasoning questionnaires, have decades of validation behind them. Games have a shorter record. That gap in evidence, not the format itself, is the reasonable source of scepticism.

What the key studies and meta-analyses show

The Journal of Intelligence meta-analysis

The most useful single source is the meta-analysis "The Relationship Between Game-Related Assessment and Traditional Measures of Cognitive Ability" (Journal of Intelligence, open access via PMC). It synthesised 44 papers and 52 samples, covering 807 effect sizes and more than 6,100 adult participants.

The overall observed correlation between GRAs and traditional cognitive ability measures was r = 0.30. After correcting for measurement error, the figure was r = 0.45.

What does that mean for hiring? A correlation of 0.30 to 0.45 is a moderate relationship. Games and classic tests overlap meaningfully, so games are not measuring nothing. But they are not interchangeable with established tests either. If two cognitive tests of the same construct were near-equivalent, you'd expect much higher figures. The result is consistent with games capturing part of general cognitive ability plus other things: motivation, interface familiarity, risk preference, or task-specific skills.

The Frontiers in Psychology recruitment study

The paper "Game based assessments of cognitive ability in recruitment: Validity, fairness and test-taking experience" (Frontiers in Psychology, via PMC) addresses a practitioner setting directly. It reports:

  • Test-retest reliability of r = 0.68 for the game-based scoring approach.
  • A convergent validity result of r = 0.50 against traditional cognitive measures.
  • A Net Promoter Score of 58 for the test-taking experience.
  • Fairness modelling using machine-learning scoring, described as ridge regression with fairness-optimised bias penalisation.

These are respectable numbers, and the study's attention to adverse impact is a good model. Two cautions apply. It examines one game-based approach from one ecosystem, so it can't be generalised to the category. And its main validity evidence is alignment with cognitive tests, not prediction of job performance.

The International Journal of Serious Games meta-analysis

A further meta-analysis in the International Journal of Serious Games is explicitly framed around the convergent validity of game-based assessment. Its scope is the point: convergent validity asks whether a game correlates with other measures of the same construct. It does not ask whether scores predict who performs well in the job.

Why the studies aren't interchangeable

Each of these sources answers a different question. Constructs vary (reasoning, memory, attention, personality-adjacent traits). So do scoring models: a transparent points-based score behaves differently from a machine-learned composite. Samples differ in age, education, language and gaming familiarity. A pooled r of 0.45 is an average across very different games, and individual products will sit above and below it.

Validity is an argument, not a format

The AERA/APA/NCME Standards for Educational and Psychological Testing (2014) treat validity as a unitary concept. There aren't separate "types" you can tick off. Instead, multiple strands of evidence support the intended interpretation and use of test scores. The strands include:

  1. Test content: does the task represent the construct and the job demands?
  2. Response processes: do candidates use the intended cognitive or behavioural processes?
  3. Internal structure: do items or tasks behave consistently with the theory of the construct?
  4. Relations to other variables: convergent evidence, discriminant evidence and relations to job criteria.
  5. Consequences of testing: including fairness, adverse impact and candidate experience.

The same framework applies to a 40-year-old verbal reasoning test and to a new game. Format doesn't earn exemption from any strand. This is also the SIOP position in its guidance on validating selection procedures: the evidence must support the specific inference being drawn in the specific context.

Comparative metrics: reliability, construct validity and predictive validity

Reliability

Reliability is the consistency of a score. For selection, unstable scores are a direct hiring risk: if a candidate would score differently on another day for reasons unrelated to ability, decisions based on that score are partly random.

The recruitment study's test-retest correlation of r = 0.68 is acceptable for a game-based approach but sits below the levels typically expected of mature, high-stakes cognitive tests, where reliability coefficients of .80 and above are common. For a screening stage with human review afterwards, 0.68 may be defensible. For a stand-alone pass/fail cut-off, it's thin. HR should ask vendors which reliability estimate they report (test-retest, internal consistency, alternate forms), over what interval, and in what sample.

Construct and convergent validity

Here the evidence is the strongest. The meta-analytic r = 0.30 observed and r = 0.45 corrected, and the r = 0.50 in the recruitment study, show real overlap with cognitive ability. In practical terms, a game scoring r = 0.50 with a classic test shares roughly a quarter of its variance with it (r squared = 0.25). The remainder is something else: construct-relevant signal, or noise, or irrelevant skill. Only further evidence distinguishes these.

Predictive validity is a separate question

Many buyer objections come from conflating two claims:

  • Convergent validity: the game correlates with a cognitive test.
  • Predictive (criterion-related) validity: the game's scores predict job performance, training success, or retention.

A moderate convergent correlation does not establish prediction of job performance. The published literature is considerably stronger on the former. Criterion-related evidence for game-based tools in real jobs is thinner, often confined to proprietary or vendor-commissioned studies.

There is an indirect argument: if a game measures cognitive ability, and cognitive ability predicts performance (as decades of work on classic tests show), then the game should inherit some predictive power. But that inference degrades with every step. A game correlating 0.45 with the construct will predict less than the classic test itself, unless it also captures something additional that matters on the job. HR should treat the indirect argument as a hypothesis to be tested locally, not a conclusion.

Response processes

Response process evidence asks whether candidates are doing what the designer assumes. In a game, a high score might reflect fast reaction, familiarity with game conventions or a learned interface strategy rather than reasoning. Vendors can evidence this through think-aloud studies, process data analysis, or by showing that scores hold up after controlling for gaming experience and device type.

Where games outperform and underperform classic tests

Where games can have the edge

  • Engagement and completion. If fewer candidates abandon the process, the applicant pool is less distorted by drop-out. Selection Lab's own customer data points this way: its 2026 deck cites 27% fewer drop-offs (March 2025). That's a vendor-reported outcome, not an independent finding, and drop-off reduction depends on the whole flow, not on the game alone.
  • Candidate experience. The recruitment study's NPS of 58 is a strong score for a test-taking experience, as tests are rarely rated highly.
  • Construct breadth. Interactive tasks can capture processes, such as learning from feedback or adapting to changing rules, that a static multiple-choice item can't.
  • Test anxiety. Plausible but under-evidenced. Treat claims of reduced anxiety as needing data, not as established.

Where games can fall short

  • Applicant reactions. Reactions to game-based tests are not uniformly positive. Some studies find less favourable perceptions than for traditional formats, and gaming experience can moderate those perceptions. Paul Englert's discussion of one such study raises the question directly: do these tools measure cognitive ability or acceptance of gaming?
  • Fairness risk. If prior gaming experience, age, or device (touchscreen versus mouse) shifts scores independently of ability, a game can introduce construct-irrelevant variance and adverse impact.
  • Criterion gaps. Without job-performance validation, a game is a plausible cognitive proxy at best.

A practical decision rule

Games are most defensible when three conditions hold: the measured construct maps to the job's demands, the vendor shows reliability and convergent evidence at acceptable levels, and criterion-related validation (vendor-supplied, then confirmed locally) exists for comparable roles. When one is missing, use the game as one input among several rather than as a gate.

This is how Selection Lab approaches the format: games sit inside a role-adaptive flow alongside other assessment types and structured interview stages, rather than operating as a standalone novelty. A single instrument rarely carries a hiring decision on its own, and combining methods is generally more defensible than relying on one.

Limitations and open questions

Generalisation. Samples in the published work are mostly adult, often student or general-population, and rarely stratified by seniority, culture or protected group. External validity across roles is largely untested.

Criterion evidence. As noted, convergent evidence dominates. Independent, peer-reviewed studies linking game scores to supervisor ratings, productivity or retention remain scarce.

Scoring transparency. Machine-learning scoring can improve prediction and allow fairness constraints, such as the bias penalisation described in the recruitment study. It also makes scores harder to explain. In an audit or challenge, "the model weighted it that way" is not a defence. Vendors need to provide audit trails, feature-level explanations where possible and subgroup analyses.

Operational artefacts. Practice effects, device differences and coaching can all move scores. These need monitoring after deployment, not only in the original validation.

Source independence. Some of the strongest practitioner-relevant studies come from within commercial ecosystems. That doesn't invalidate them, but it means independent replication and transparent methods carry extra weight.

What recruiters and HR leaders should demand before deployment

A demo video is not evidence. Ask for a validity evidence pack that maps to the strands above.

The evidence checklist

  • Construct and content: a documented definition of what the game measures and a job-analysis rationale for why it matters for the role.
  • Response processes: evidence that scores reflect the intended processes, including analysis of gaming experience and device effects.
  • Reliability: test-retest and internal consistency figures, with sample, interval and method. Compare against the stakes of your decision.
  • Convergent validity: correlations with established cognitive tests, ideally in more than one sample. Treat a figure around 0.50 as moderate.
  • Criterion-related validity: correlations with job performance, training outcomes or retention in comparable roles, then a plan for local validation.
  • Fairness: subgroup mean differences and selection rates by gender, age, ethnicity and other relevant groups.
  • Candidate experience: completion rates, applicant reaction data and fairness perceptions.

Adverse impact auditing

Apply the four-fifths (80%) rule as a first screen: if the selection rate for any group is less than 80% of the rate for the highest-scoring group, the procedure is flagged for potential adverse impact under the US Uniform Guidelines. UK and EU contexts rely on different legal tests, but the four-fifths ratio remains a widely used monitoring heuristic for cognitive tests and other selection procedures. Treat it as a trigger for investigation and mitigation, not a verdict. Require documented mitigation steps, and re-run the analysis on your own applicant data, not only the vendor's.

Governance and data

Request documentation on data protection, retention periods, audit logs and human oversight of automated decisions. For European employers this includes GDPR compliance and alignment with the EU AI Act, under which recruitment tools are treated as high-risk. Selection Lab states GDPR compliance and EU AI Act alignment, with personal data stored in Frankfurt and local language models used to remove personal information from conversations. Whichever vendor you choose, ask to see the documentation behind such claims.

Run a local pilot

Before scaling, pilot with a defined role. Collect game scores alongside your existing selection data and, after a suitable period, performance and early-retention outcomes. Monitor subgroup selection rates throughout. Customer outcomes reported by Selection Lab, such as 21% lower early turnover (January 2024) and 15 minutes saved per applicant (December 2025), illustrate the kind of operational measures worth tracking, but they're vendor-reported and your pilot should establish your own baseline.

Integration matters too. Selection Lab states ATS integration and go-live in two to 10 weeks, which makes a time-boxed pilot realistic. Whatever the vendor, insist that pilot data remain accessible to you for analysis.

The defensible position

The evidence supports a measured conclusion. Game-related assessments correlate moderately with traditional cognitive ability (r = 0.30 observed, 0.45 corrected, across 44 papers and over 6,100 participants), and at least one recruitment study reports acceptable reliability (r = 0.68), a convergent validity of r = 0.50 and a strong candidate experience score. Calling that a gimmick ignores the data. Calling it equivalent to a validated classic test ignores what's missing: criterion-related evidence, independent replication, and subgroup fairness evidence for your own population.

Gaming is not the measurement. Validity evidence is. Buy the evidence first, and the format second.