Reading Time

Are candidates using ChatGPT in assessments? How it affects results

The question isn't hypothetical anymore. Candidates are using ChatGPT during job assessments, and the evidence suggests it's affecting scores in ways that matter for hiring validity. A peer-reviewed PLOS ONE study (PMC) submitted 63 AI-generated answers into a real university examination system under blind conditions: 94% went undetected, and they earned higher grades on average than submissions from actual students. That's not a lab artifact. It's a signal that assessment designs built for a pre-LLM world carry real integrity risk.

This guide is for TA leaders, recruitment managers, and assessment designers who need to act on that risk now. You'll get a format vulnerability map, redesign principles, a monitoring framework, an operational playbook, and a KPI dashboard template to track it all.

Prerequisites: Access to your current assessment configuration, a basic understanding of your item types, and decision authority (or stakeholder alignment) to modify at least one assessment flow.

Difficulty: Moderate. Time to implement initial controls: 1–2 weeks for policy and monitoring configuration; 4–6 weeks for full assessment redesign.

Which assessment formats are most exposed to LLM-assisted cheating

Not all formats carry the same risk. The exposure depends on three variables: whether the task is unsupervised, how text-heavy the output is, and whether there's any process evidence captured alongside the answer.

FormatVulnerabilityWhyWhat to do instead
Async written responsesHighStatic, unconstrained text output; no process captureTime-box; add contextual constraints mid-session
Take-home assignmentsHighFully unsupervised; LLM can iterate indefinitelyRequire annotated decision rationale; gate with follow-up interview
SJT (text-response variant)MediumResearch (ScienceDirect, October 2024) shows only minor score shifts, but response patterns shift noticeablySwitch to forced-choice or ranking formats
Verbal reasoning / aptitudeMedium-HighLLMs can reason through many item types, especially if timed generouslyAdaptive timing; parallel item pools
Coding/technical screensMediumDetectable via runtime checks; but prompt-injection is commonUse in-browser IDEs with screen capture; require explanation of logic
Live structured interviewLowReal-time, interactive; hard to outsource mid-conversationPre-brief candidates on AI-use policy before prep window

A Wiley Online Library study (published March 9, 2026, N=5,675) found fewer than 3% of applicants self-reported GenAI use overall, though up to 19% reported using it in specific assessment stages. Self-report underestimates actual use, but the stage-level variance tells you where to concentrate controls.

Design principles for LLM-resilient assessments

The most reliable mitigation isn't detection. It's making the task structurally harder to outsource.

Require process evidence, not just answers. A candidate who genuinely solved a problem can explain their reasoning under follow-up; a pasted LLM output usually can't. For a customer support SJT, don't just ask "What would you do?" Ask the candidate to identify which piece of information from the scenario they weighted most heavily and why. That constraint forces engagement with session-specific content an LLM hasn't seen.

Time-box and branch. Adaptive timing (shorter windows per item than a candidate would need to copy-paste and review) significantly reduces the LLM advantage on text-heavy items. Scenario branching, where the next question depends on the previous answer, makes prompt-reuse ineffective because the LLM-generated answer to question 2 is incoherent if it wasn't built on the candidate's actual answer to question 1.

Use parallel item pools. If every candidate sees the same warehouse operations scenario, that scenario will eventually be posted and shared. Build at least three structurally equivalent variants with different contextual details. This reduces prompt-matching and extends the shelf life of your item bank.

Combine format types. A modular assessment that includes a forced-choice SJT, a time-boxed cognitive game, and a short video response is much harder to systematically game than a single open-text task. Combining formats across competency dimensions also gives you more validity coverage per session.

Proctoring and monitoring: what to record and why

Candidates using ChatGPT during assessments is an integrity question, but the tool you use to detect it matters enormously for fairness.

AI-text classifiers (tools like GPTZero) operate on probabilistic text features, and their error rates are operationally significant. Turnitin's own guidance (July 2024) assigns no detection score for outputs in the 1-19% AI probability range to avoid false positives. Even with a stated goal of under 1% false positive rate (Turnitin, March 2023), Jisc estimated (June 2025) that a 1% false-positive rate would generate approximately 4,800 wrongful flags annually across a typical institution-scale deployment. For a hiring context where a false flag can exclude a qualified candidate, that's not an acceptable sole evidence standard.

Behavioral integrity monitoring produces a different class of evidence. Observable signals such as focus-change events (tab switches, application changes), clipboard activity, keystroke cadence irregularities, and session-timing anomalies are behaviorally grounded and reviewable by a human. They don't identify AI use definitively, but they create an audit trail for reasoned human judgment.

Selection Lab's proctoring configuration captures camera, audio, and screen recordings with candidate consent, and surfaces integrity signals about whether tools like ChatGPT were used during the session. The deterrence effect of transparent recording disclosure is itself a meaningful control: candidates who know the session is recorded are less likely to attempt outsourcing in the first place.

On GDPR and EU AI Act compliance: candidate consent must be explicit and granular (camera/audio/screen as separate data categories where required), retention periods must be defined and minimal, and automated adverse decisions based solely on AI-generated flags are prohibited under EU AI Act high-risk system requirements. Integrity review must involve a human decision step.

Review-worthy anomalies by format:

  • Written tasks: response completion time under 40% of median; single large paste event; zero keyboard idle periods
  • SJTs: answer-pattern inconsistency with earlier session facts; response similarity score above threshold vs. known LLM outputs
  • Cognitive/aptitude: speed outliers with near-perfect accuracy; device input pattern deviating from population norm

Operational playbook: before, during, and after

Before the assessment window opens

  1. Publish a clear AI-use policy for each stage of your funnel. Specify what's permitted (e.g., using AI to prepare for an interview) and what's not (e.g., using ChatGPT during a timed assessment). Make this visible at registration and in the assessment introduction screen.
  2. Configure parallel item pools and set per-item time parameters in your assessment platform.
  3. Set integrity review thresholds: decide in advance which anomaly patterns trigger human review vs. automatic flagging for re-sit.
  4. Confirm proctoring consent flows are active and that candidates receive clear notice before the session starts.

During the assessment

  1. Enable in-browser lockdown where appropriate for high-stakes screens. This prevents tab-switching and clipboard use without requiring full proctoring for lower-stakes stages.
  2. Use adaptive timing: set per-item windows based on expected completion time for unassisted candidates, not generous buffers.
  3. Configure anomaly triggers: set real-time alerts for completion-time outliers or repeated focus-change events above a defined threshold.

After the assessment

  1. Triage flagged cases with a human reviewer before any decision is made. Document the specific anomalies observed and the reviewer's rationale.
  2. Apply fallback rules consistently: if evidence is ambiguous, reschedule rather than exclude. If behavioral signals are strong and corroborated, discard with documented justification. If signals are minor and the score is otherwise valid, accept with a note and weight the interview more heavily.
  3. Log all integrity decisions for audit trail purposes, especially under EU AI Act high-risk classification requirements.

KPIs and evaluation metrics to track

Set a measurement baseline before deploying redesigned assessments, then review at 30 and 90 days.

Integrity monitoring KPIs:

  • Completion anomaly rate (% of sessions triggering at least one review-worthy signal)
  • Focus-change event frequency per session, by format
  • Response similarity signature rate (% of written responses above LLM-match threshold)
  • Human review throughput: time from flag to decision
  • Re-sit rate and outcome correlation

Assessment validity KPIs:

  • Pass-rate stability before and after redesign (significant shifts may indicate score inflation correction or adverse impact introduction)
  • Adverse impact analysis across protected characteristics
  • Downstream performance correlation where 90-day manager ratings or early turnover data are available. Selection Lab customers using the full workflow have reported a 21% reduction in early turnover (January 2024), which provides a proxy quality-of-hire signal worth tracking against your pre-redesign baseline.

Operational efficiency KPIs:

  • Time saved per applicant in the screened funnel
  • Drop-off rate across assessment stages (redesigned formats shouldn't increase drop-off; Selection Lab's intake and assessment flow has shown a 27% reduction in drop-offs as of March 2025)

Implementation checklist:

  • AI-use policy published and versioned
  • Parallel item pools configured (minimum 3 variants per scenario)
  • Per-item timing parameters set and tested
  • Proctoring consent flows active with GDPR-compliant retention settings
  • Anomaly thresholds defined and documented
  • Human review workflow assigned with decision SLA
  • Baseline integrity and validity metrics captured
  • 30-day and 90-day review dates scheduled

The research picture is now clear enough to act on: generative AI use in job assessments is real, its impact on results depends on format, and text classifiers aren't a sufficient control mechanism. The operational response is assessment redesign plus behavioral monitoring plus structured human review. That combination, built into a modular assessment platform with native integrity tooling, is what converts a compliance risk into a manageable, auditable process.

Frequently asked questions about ChatGPT in assessments

Are candidates actually using ChatGPT in job assessments?

Yes. A Wiley Online Library study of 5,675 applicants found fewer than 3% self-reported generative AI use overall, but up to 19% reported using it in specific assessment stages. Self-report underestimates real use. In a blind PLOS ONE experiment, 94% of AI-generated answers went undetected and scored higher than real submissions.

Which assessment formats are most vulnerable to ChatGPT?

Asynchronous written responses and take-home assignments carry the highest risk because they are unsupervised and produce unconstrained text. Verbal reasoning tests sit at medium to high risk when timing is generous. Text-based SJTs and coding screens are medium risk. Live structured interviews are the hardest to outsource.

Can AI detectors reliably prove a candidate used ChatGPT?

No. Text classifiers work on probabilistic features and produce false positives at a rate that matters in hiring. Turnitin assigns no score in the 1 to 19% probability range for that reason, and Jisc estimated a 1% false positive rate still yields roughly 4,800 wrongful flags a year at institutional scale. Behavioral signals plus human review are the defensible evidence standard.

How do you make an assessment resistant to ChatGPT?

Ask for process evidence rather than just answers, time-box each item, branch scenarios so the next question depends on the previous answer, run at least three parallel item variants per scenario, and combine formats such as a forced-choice SJT, a timed cognitive game and a short video response in one flow.

Is proctoring for AI use allowed under GDPR and the EU AI Act?

Yes, with conditions. Consent must be explicit and granular per data type (camera, audio, screen), retention must be defined and minimal, and no adverse decision may rest solely on an automated flag. A human must review every flagged case before a decision is made.

FAQ

Can game-based assessments promote diversity in the hiring process?

Yes, game-based assessments can support diversity by focusing on skills and behaviors rather than traditional criteria like résumés, which may contain unconscious biases. This gives candidates from diverse backgrounds a fairer chance to demonstrate their potential.

What is a game-based assessment?

A game-based assessment is a method that uses game mechanics to evaluate a candidate’s skills, competencies, and personality traits. While playing these games, candidates are assessed on aspects like problem-solving, cognitive ability, and behavior under pressure in an interactive way.

What are the advantages of game-based assessments?

Game-based assessments offer a more engaging and interactive experience for candidates, which can lead to a more positive perception of the hiring process—especially among certain groups. For employers, they provide deeper insights into both cognitive and behavioral traits, which traditional tests may miss. They also reduce the chance of socially desirable answers, as candidates tend to respond more authentically in a game environment.

How reliable are game-based assessments compared to traditional tests?

When well-designed, game-based assessments can be just as reliable—or even more reliable—than traditional tests. They assess a wide range of behaviors and cognitive abilities in a dynamic setting. However, the quality of these assessments varies greatly, so careful evaluation is essential.

How does a game-based assessment work?

Candidates participate in interactive games designed to measure specific skills and behaviors. Evaluation goes beyond just the final score—it also considers how the candidate makes decisions, handles challenges, and responds to different scenarios. These insights reveal underlying thought processes and behavioral patterns.

Are game-based assessments scientifically validated?

The main drawback is that many game-based assessments are relatively new and have not yet been extensively researched by independent academics. Providers often cite their own research, which is rarely externally validated. Without independent studies, the reliability of these assessments remains uncertain—something to keep in mind when selecting one.

How can game based assessments contribute to a better candidate experience

This can vary significantly by audience. The playful, interactive nature of game-based assessments can lower stress levels for some candidates compared to traditional tests. However, research shows that certain groups, especially those over 35, may find them more stressful. Men also tend to rate the experience more positively than women.

Can you practice game-based assessment?

You can familiarize yourself with the style of games used, but it’s difficult to "practice" for them in a traditional sense. These assessments are designed to measure natural reactions and authentic behavior, so repeated practice typically has less effect on performance than with traditional tests.

Will game-based assessments replace traditional tests in the future?

It’s likely that game-based assessments will become more common in hiring processes, but they probably won’t fully replace traditional tests. Both approaches have value and can complement each other depending on the role and the company’s needs.

How are the results of a game-based assessment analyzed?

Results are analyzed based on predefined criteria such as problem-solving ability, reaction time, and behavior under pressure. Advanced algorithms collect and interpret this data to provide a reliable, objective evaluation of a candidate’s strengths.

What kind of skills do game-based assessments measure?

They assess a wide range of abilities, including problem-solving, adaptability, decision-making under pressure, teamwork, and emotional intelligence. Depending on the design, they may also evaluate cognitive skills like memory, attention, and pattern recognition.

How long does a game-based assessment take?

Typically, these assessments last between 15 and 60 minutes, depending on the game’s complexity and the number of skills being tested. They’re usually shorter and more engaging than traditional assessments, making for a smoother candidate experience.

Are game-based assessments suitable for all roles?

They are especially effective for roles that require flexibility, creativity, problem-solving, and strong interpersonal skills. For highly technical or specialized roles, additional assessments may be needed to measure specific knowledge.

What’s the difference between a game-based and a gamified assessment?

A gamified assessment adds game-like elements (such as points or rewards) to a traditional test to increase engagement. A game-based assessment, on the other hand, is a standalone game designed specifically to evaluate certain competencies. The game itself is the primary evaluation tool, not just an enhancement.

FAQ

How can I improve my company’s retention rate?

The retention rate can be improved by investing in employee development and satisfaction. This includes offering training, career opportunities, and recognition for their contributions. A culture of open communication and attention to work-life balance can also contribute to higher retention. Additionally, offering competitive compensation and involving employees in decision-making can strengthen loyalty.

What are the benefits of growth opportunities for employee retention?

Growth opportunities can promote employee retention by giving staff a sense of direction and motivation. When they have the chance to learn and develop professionally within the company, they feel valued, which increases their loyalty. This can prevent them from leaving to seek better opportunities elsewhere. kunnen het behoud van personeel bevorderen door medewerkers een gevoel van richting en motivatie te geven. Wanneer zij de kans krijgen om te leren en zich professioneel te ontwikkelen binnen het bedrijf, voelen zij zich gewaardeerd, wat hun loyaliteit vergroot. Dit kan voorkomen dat ze vertrekken om elders betere kansen te zoeken.

What are the key factors that influence employee retention?

Key factors that influence employee retention include salary and benefits, opportunities for professional development, work-life balance, company culture, and the relationship with supervisors. Employees tend to stay longer when they feel valued, challenged, and supported in their work environment.

Why is employee retention so important for organizations?

Employee retention is important because it helps reduce recruitment and training costs for new employees, and it contributes to retaining knowledge and experience within the organization. High retention also ensures continuity within teams, leading to a more stable company culture, higher customer satisfaction, and improved business outcomes.

Which recruitment strategies help improve retention?

Recruitment strategies that can improve retention include identifying candidates who align with the company culture, using assessments to evaluate soft skills, and providing transparency about role expectations during the hiring process. Employees who feel connected to the organization and have clarity about their role are more likely to stay longer.

How can a good onboarding process contribute to higher retention?

An effective onboarding process can contribute to higher retention by helping new employees quickly adapt to their role, the company culture, and expectations. By providing support and clear information from the start, their engagement is increased, and the likelihood of them leaving early due to feelings of being overwhelmed or lacking guidance is reduced.

What is the role of company culture in retaining employees?

Company culture plays a crucial role in employee retention. When employees feel heard, valued, and connected to the values and norms of the company, they are more likely to stay. A positive culture that fosters collaboration, respect, and personal growth can significantly enhance employee motivation and satisfaction.

How can leadership and management style influence retention?

Leadership and management style have a significant impact on retention. Leaders who inspire, support, and coach their team can increase employee engagement and satisfaction. Offering autonomy and trust can lead to higher loyalty, while inefficient or negative management styles can contribute to dissatisfaction and increased employee turnover.

What is the importance of recognition and rewards for employee retention?

Recognition and rewards play an important role in employee retention by showing staff that their work is valued. This can increase their motivation and loyalty. In addition to financial rewards, compliments, promotions, and other forms of recognition can also contribute to satisfaction and retaining employees.

What role does work-life balance play in improving retention?

A balanced work-life balance plays an important role in increasing retention. By reducing stress and improving job satisfaction, employees are more likely to stay with the company. Initiatives such as flexible working hours, remote work options, and respect for personal time can contribute to this balance.

What does increasing retention mean within a company?

Increasing retention within a company means implementing strategies to keep employees with the organization for longer. This can be achieved by improving job satisfaction, offering growth opportunities, and fostering a positive and supportive company culture.

How do I measure the success of my retention strategy?

The success of a retention strategy can be measured by tracking retention rates and turnover rates, and by gaining insights from exit interviews. Additionally, employee satisfaction surveys and feedback from performance evaluations can provide valuable information about the effectiveness of the strategies applied.

What are the costs of a low retention rate?

A low retention rate can bring significant costs, such as increased expenses for recruiting and training new employees. Furthermore, the loss of experienced staff can lead to lower productivity, reduced knowledge transfer, and a negative impact on company culture.

How can I increase employee engagement?

To increase employee engagement, involve them in decision-making processes, regularly ask for their feedback, and recognize their contributions. Offering development opportunities and maintaining transparent communication can also contribute to greater engagement.

How can technology help improve employee retention?

Technology can be a tool for improving employee retention by facilitating communication, feedback, and development. By using online platforms for training, recognition, and evaluation, companies can create a more engaged and satisfied workforce.

FAQ

How long does it take to complete the tool?

Less than 10 minutes. You’ll answer 30 guided questions and get a summary of what to look for in your next assessment platform.

Can this checklist help me compare assessment providers?

Yes. By clarifying what matters most to your team, it makes comparing providers' features, pricing, and strengths much easier and more strategic.

How can I use this checklist if I’m not doing a formal RFI?

It’s equally valuable for internal evaluations, exploring new tools, or improving your current hiring process even if you’re not issuing an RFI or RFQ.

What should I look for in a modern assessment tool?

Prioritize platforms with user-friendly design, mobile compatibility, strong analytics, ATS integrations, and inclusive features like neurodiversity support.

What types of assessments should I consider in 2025?

Leading tools combine cognitive testing, situational judgment tests (SJTs), behavior assessments, and predictive AI to evaluate candidates more holistically.

Who should use an assessment checklist?

HR professionals, hiring managers, and procurement teams evaluating pre-selection solutions, especially those comparing AI-powered or compliance-driven assessment platforms.

How does this checklist help with RFIs and RFQs for assessments?

The checklist helps you define your exact requirements so you can confidently draft or respond to Requests for Information (RFI) or Requests for Quotation (RFQ) for assessment tools.

What is an assessment tool in hiring?

An assessment tool evaluates candidates’ skills, behaviors, and fit during the recruitment process. It helps improve hiring decisions and streamline pre-selection.

Game-based assessment packs

← Our Blog

Are candidates using ChatGPT in assessments? How it affects results

Candidates using ChatGPT during assessments? A peer-reviewed study shows 94% went undetected and earned higher grades. Learn 10 steps to protect your hiring validity now.
Joeri Everaers
COO
Read time: Approx

The question isn't hypothetical anymore. Candidates are using ChatGPT during job assessments, and the evidence suggests it's affecting scores in ways that matter for hiring validity. A peer-reviewed PLOS ONE study (PMC) submitted 63 AI-generated answers into a real university examination system under blind conditions: 94% went undetected, and they earned higher grades on average than submissions from actual students. That's not a lab artifact. It's a signal that assessment designs built for a pre-LLM world carry real integrity risk.

This guide is for TA leaders, recruitment managers, and assessment designers who need to act on that risk now. You'll get a format vulnerability map, redesign principles, a monitoring framework, an operational playbook, and a KPI dashboard template to track it all.

Prerequisites: Access to your current assessment configuration, a basic understanding of your item types, and decision authority (or stakeholder alignment) to modify at least one assessment flow.

Difficulty: Moderate. Time to implement initial controls: 1–2 weeks for policy and monitoring configuration; 4–6 weeks for full assessment redesign.

Which assessment formats are most exposed to LLM-assisted cheating

Not all formats carry the same risk. The exposure depends on three variables: whether the task is unsupervised, how text-heavy the output is, and whether there's any process evidence captured alongside the answer.

FormatVulnerabilityWhyWhat to do instead
Async written responsesHighStatic, unconstrained text output; no process captureTime-box; add contextual constraints mid-session
Take-home assignmentsHighFully unsupervised; LLM can iterate indefinitelyRequire annotated decision rationale; gate with follow-up interview
SJT (text-response variant)MediumResearch (ScienceDirect, October 2024) shows only minor score shifts, but response patterns shift noticeablySwitch to forced-choice or ranking formats
Verbal reasoning / aptitudeMedium-HighLLMs can reason through many item types, especially if timed generouslyAdaptive timing; parallel item pools
Coding/technical screensMediumDetectable via runtime checks; but prompt-injection is commonUse in-browser IDEs with screen capture; require explanation of logic
Live structured interviewLowReal-time, interactive; hard to outsource mid-conversationPre-brief candidates on AI-use policy before prep window

A Wiley Online Library study (published March 9, 2026, N=5,675) found fewer than 3% of applicants self-reported GenAI use overall, though up to 19% reported using it in specific assessment stages. Self-report underestimates actual use, but the stage-level variance tells you where to concentrate controls.

Design principles for LLM-resilient assessments

The most reliable mitigation isn't detection. It's making the task structurally harder to outsource.

Require process evidence, not just answers. A candidate who genuinely solved a problem can explain their reasoning under follow-up; a pasted LLM output usually can't. For a customer support SJT, don't just ask "What would you do?" Ask the candidate to identify which piece of information from the scenario they weighted most heavily and why. That constraint forces engagement with session-specific content an LLM hasn't seen.

Time-box and branch. Adaptive timing (shorter windows per item than a candidate would need to copy-paste and review) significantly reduces the LLM advantage on text-heavy items. Scenario branching, where the next question depends on the previous answer, makes prompt-reuse ineffective because the LLM-generated answer to question 2 is incoherent if it wasn't built on the candidate's actual answer to question 1.

Use parallel item pools. If every candidate sees the same warehouse operations scenario, that scenario will eventually be posted and shared. Build at least three structurally equivalent variants with different contextual details. This reduces prompt-matching and extends the shelf life of your item bank.

Combine format types. A modular assessment that includes a forced-choice SJT, a time-boxed cognitive game, and a short video response is much harder to systematically game than a single open-text task. Combining formats across competency dimensions also gives you more validity coverage per session.

Proctoring and monitoring: what to record and why

Candidates using ChatGPT during assessments is an integrity question, but the tool you use to detect it matters enormously for fairness.

AI-text classifiers (tools like GPTZero) operate on probabilistic text features, and their error rates are operationally significant. Turnitin's own guidance (July 2024) assigns no detection score for outputs in the 1-19% AI probability range to avoid false positives. Even with a stated goal of under 1% false positive rate (Turnitin, March 2023), Jisc estimated (June 2025) that a 1% false-positive rate would generate approximately 4,800 wrongful flags annually across a typical institution-scale deployment. For a hiring context where a false flag can exclude a qualified candidate, that's not an acceptable sole evidence standard.

Behavioral integrity monitoring produces a different class of evidence. Observable signals such as focus-change events (tab switches, application changes), clipboard activity, keystroke cadence irregularities, and session-timing anomalies are behaviorally grounded and reviewable by a human. They don't identify AI use definitively, but they create an audit trail for reasoned human judgment.

Selection Lab's proctoring configuration captures camera, audio, and screen recordings with candidate consent, and surfaces integrity signals about whether tools like ChatGPT were used during the session. The deterrence effect of transparent recording disclosure is itself a meaningful control: candidates who know the session is recorded are less likely to attempt outsourcing in the first place.

On GDPR and EU AI Act compliance: candidate consent must be explicit and granular (camera/audio/screen as separate data categories where required), retention periods must be defined and minimal, and automated adverse decisions based solely on AI-generated flags are prohibited under EU AI Act high-risk system requirements. Integrity review must involve a human decision step.

Review-worthy anomalies by format:

  • Written tasks: response completion time under 40% of median; single large paste event; zero keyboard idle periods
  • SJTs: answer-pattern inconsistency with earlier session facts; response similarity score above threshold vs. known LLM outputs
  • Cognitive/aptitude: speed outliers with near-perfect accuracy; device input pattern deviating from population norm

Operational playbook: before, during, and after

Before the assessment window opens

  1. Publish a clear AI-use policy for each stage of your funnel. Specify what's permitted (e.g., using AI to prepare for an interview) and what's not (e.g., using ChatGPT during a timed assessment). Make this visible at registration and in the assessment introduction screen.
  2. Configure parallel item pools and set per-item time parameters in your assessment platform.
  3. Set integrity review thresholds: decide in advance which anomaly patterns trigger human review vs. automatic flagging for re-sit.
  4. Confirm proctoring consent flows are active and that candidates receive clear notice before the session starts.

During the assessment

  1. Enable in-browser lockdown where appropriate for high-stakes screens. This prevents tab-switching and clipboard use without requiring full proctoring for lower-stakes stages.
  2. Use adaptive timing: set per-item windows based on expected completion time for unassisted candidates, not generous buffers.
  3. Configure anomaly triggers: set real-time alerts for completion-time outliers or repeated focus-change events above a defined threshold.

After the assessment

  1. Triage flagged cases with a human reviewer before any decision is made. Document the specific anomalies observed and the reviewer's rationale.
  2. Apply fallback rules consistently: if evidence is ambiguous, reschedule rather than exclude. If behavioral signals are strong and corroborated, discard with documented justification. If signals are minor and the score is otherwise valid, accept with a note and weight the interview more heavily.
  3. Log all integrity decisions for audit trail purposes, especially under EU AI Act high-risk classification requirements.

KPIs and evaluation metrics to track

Set a measurement baseline before deploying redesigned assessments, then review at 30 and 90 days.

Integrity monitoring KPIs:

  • Completion anomaly rate (% of sessions triggering at least one review-worthy signal)
  • Focus-change event frequency per session, by format
  • Response similarity signature rate (% of written responses above LLM-match threshold)
  • Human review throughput: time from flag to decision
  • Re-sit rate and outcome correlation

Assessment validity KPIs:

  • Pass-rate stability before and after redesign (significant shifts may indicate score inflation correction or adverse impact introduction)
  • Adverse impact analysis across protected characteristics
  • Downstream performance correlation where 90-day manager ratings or early turnover data are available. Selection Lab customers using the full workflow have reported a 21% reduction in early turnover (January 2024), which provides a proxy quality-of-hire signal worth tracking against your pre-redesign baseline.

Operational efficiency KPIs:

  • Time saved per applicant in the screened funnel
  • Drop-off rate across assessment stages (redesigned formats shouldn't increase drop-off; Selection Lab's intake and assessment flow has shown a 27% reduction in drop-offs as of March 2025)

Implementation checklist:

  • AI-use policy published and versioned
  • Parallel item pools configured (minimum 3 variants per scenario)
  • Per-item timing parameters set and tested
  • Proctoring consent flows active with GDPR-compliant retention settings
  • Anomaly thresholds defined and documented
  • Human review workflow assigned with decision SLA
  • Baseline integrity and validity metrics captured
  • 30-day and 90-day review dates scheduled

The research picture is now clear enough to act on: generative AI use in job assessments is real, its impact on results depends on format, and text classifiers aren't a sufficient control mechanism. The operational response is assessment redesign plus behavioral monitoring plus structured human review. That combination, built into a modular assessment platform with native integrity tooling, is what converts a compliance risk into a manageable, auditable process.

Frequently asked questions about ChatGPT in assessments

Are candidates actually using ChatGPT in job assessments?

Yes. A Wiley Online Library study of 5,675 applicants found fewer than 3% self-reported generative AI use overall, but up to 19% reported using it in specific assessment stages. Self-report underestimates real use. In a blind PLOS ONE experiment, 94% of AI-generated answers went undetected and scored higher than real submissions.

Which assessment formats are most vulnerable to ChatGPT?

Asynchronous written responses and take-home assignments carry the highest risk because they are unsupervised and produce unconstrained text. Verbal reasoning tests sit at medium to high risk when timing is generous. Text-based SJTs and coding screens are medium risk. Live structured interviews are the hardest to outsource.

Can AI detectors reliably prove a candidate used ChatGPT?

No. Text classifiers work on probabilistic features and produce false positives at a rate that matters in hiring. Turnitin assigns no score in the 1 to 19% probability range for that reason, and Jisc estimated a 1% false positive rate still yields roughly 4,800 wrongful flags a year at institutional scale. Behavioral signals plus human review are the defensible evidence standard.

How do you make an assessment resistant to ChatGPT?

Ask for process evidence rather than just answers, time-box each item, branch scenarios so the next question depends on the previous answer, run at least three parallel item variants per scenario, and combine formats such as a forced-choice SJT, a timed cognitive game and a short video response in one flow.

Is proctoring for AI use allowed under GDPR and the EU AI Act?

Yes, with conditions. Consent must be explicit and granular per data type (camera, audio, screen), retention must be defined and minimal, and no adverse decision may rest solely on an automated flag. A human must review every flagged case before a decision is made.