The question isn't hypothetical anymore. Candidates are using ChatGPT during job assessments, and the evidence suggests it's affecting scores in ways that matter for hiring validity. A peer-reviewed PLOS ONE study (PMC) submitted 63 AI-generated answers into a real university examination system under blind conditions: 94% went undetected, and they earned higher grades on average than submissions from actual students. That's not a lab artifact. It's a signal that assessment designs built for a pre-LLM world carry real integrity risk.
This guide is for TA leaders, recruitment managers, and assessment designers who need to act on that risk now. You'll get a format vulnerability map, redesign principles, a monitoring framework, an operational playbook, and a KPI dashboard template to track it all.
Prerequisites: Access to your current assessment configuration, a basic understanding of your item types, and decision authority (or stakeholder alignment) to modify at least one assessment flow.
Difficulty: Moderate. Time to implement initial controls: 1–2 weeks for policy and monitoring configuration; 4–6 weeks for full assessment redesign.
Not all formats carry the same risk. The exposure depends on three variables: whether the task is unsupervised, how text-heavy the output is, and whether there's any process evidence captured alongside the answer.
| Format | Vulnerability | Why | What to do instead |
|---|---|---|---|
| Async written responses | High | Static, unconstrained text output; no process capture | Time-box; add contextual constraints mid-session |
| Take-home assignments | High | Fully unsupervised; LLM can iterate indefinitely | Require annotated decision rationale; gate with follow-up interview |
| SJT (text-response variant) | Medium | Research (ScienceDirect, October 2024) shows only minor score shifts, but response patterns shift noticeably | Switch to forced-choice or ranking formats |
| Verbal reasoning / aptitude | Medium-High | LLMs can reason through many item types, especially if timed generously | Adaptive timing; parallel item pools |
| Coding/technical screens | Medium | Detectable via runtime checks; but prompt-injection is common | Use in-browser IDEs with screen capture; require explanation of logic |
| Live structured interview | Low | Real-time, interactive; hard to outsource mid-conversation | Pre-brief candidates on AI-use policy before prep window |
A Wiley Online Library study (published March 9, 2026, N=5,675) found fewer than 3% of applicants self-reported GenAI use overall, though up to 19% reported using it in specific assessment stages. Self-report underestimates actual use, but the stage-level variance tells you where to concentrate controls.
The most reliable mitigation isn't detection. It's making the task structurally harder to outsource.
Require process evidence, not just answers. A candidate who genuinely solved a problem can explain their reasoning under follow-up; a pasted LLM output usually can't. For a customer support SJT, don't just ask "What would you do?" Ask the candidate to identify which piece of information from the scenario they weighted most heavily and why. That constraint forces engagement with session-specific content an LLM hasn't seen.
Time-box and branch. Adaptive timing (shorter windows per item than a candidate would need to copy-paste and review) significantly reduces the LLM advantage on text-heavy items. Scenario branching, where the next question depends on the previous answer, makes prompt-reuse ineffective because the LLM-generated answer to question 2 is incoherent if it wasn't built on the candidate's actual answer to question 1.
Use parallel item pools. If every candidate sees the same warehouse operations scenario, that scenario will eventually be posted and shared. Build at least three structurally equivalent variants with different contextual details. This reduces prompt-matching and extends the shelf life of your item bank.
Combine format types. A modular assessment that includes a forced-choice SJT, a time-boxed cognitive game, and a short video response is much harder to systematically game than a single open-text task. Combining formats across competency dimensions also gives you more validity coverage per session.
Candidates using ChatGPT during assessments is an integrity question, but the tool you use to detect it matters enormously for fairness.
AI-text classifiers (tools like GPTZero) operate on probabilistic text features, and their error rates are operationally significant. Turnitin's own guidance (July 2024) assigns no detection score for outputs in the 1-19% AI probability range to avoid false positives. Even with a stated goal of under 1% false positive rate (Turnitin, March 2023), Jisc estimated (June 2025) that a 1% false-positive rate would generate approximately 4,800 wrongful flags annually across a typical institution-scale deployment. For a hiring context where a false flag can exclude a qualified candidate, that's not an acceptable sole evidence standard.
Behavioral integrity monitoring produces a different class of evidence. Observable signals such as focus-change events (tab switches, application changes), clipboard activity, keystroke cadence irregularities, and session-timing anomalies are behaviorally grounded and reviewable by a human. They don't identify AI use definitively, but they create an audit trail for reasoned human judgment.
Selection Lab's proctoring configuration captures camera, audio, and screen recordings with candidate consent, and surfaces integrity signals about whether tools like ChatGPT were used during the session. The deterrence effect of transparent recording disclosure is itself a meaningful control: candidates who know the session is recorded are less likely to attempt outsourcing in the first place.
On GDPR and EU AI Act compliance: candidate consent must be explicit and granular (camera/audio/screen as separate data categories where required), retention periods must be defined and minimal, and automated adverse decisions based solely on AI-generated flags are prohibited under EU AI Act high-risk system requirements. Integrity review must involve a human decision step.
Review-worthy anomalies by format:
Set a measurement baseline before deploying redesigned assessments, then review at 30 and 90 days.
Integrity monitoring KPIs:
Assessment validity KPIs:
Operational efficiency KPIs:
Implementation checklist:
The research picture is now clear enough to act on: generative AI use in job assessments is real, its impact on results depends on format, and text classifiers aren't a sufficient control mechanism. The operational response is assessment redesign plus behavioral monitoring plus structured human review. That combination, built into a modular assessment platform with native integrity tooling, is what converts a compliance risk into a manageable, auditable process.
Yes. A Wiley Online Library study of 5,675 applicants found fewer than 3% self-reported generative AI use overall, but up to 19% reported using it in specific assessment stages. Self-report underestimates real use. In a blind PLOS ONE experiment, 94% of AI-generated answers went undetected and scored higher than real submissions.
Asynchronous written responses and take-home assignments carry the highest risk because they are unsupervised and produce unconstrained text. Verbal reasoning tests sit at medium to high risk when timing is generous. Text-based SJTs and coding screens are medium risk. Live structured interviews are the hardest to outsource.
No. Text classifiers work on probabilistic features and produce false positives at a rate that matters in hiring. Turnitin assigns no score in the 1 to 19% probability range for that reason, and Jisc estimated a 1% false positive rate still yields roughly 4,800 wrongful flags a year at institutional scale. Behavioral signals plus human review are the defensible evidence standard.
Ask for process evidence rather than just answers, time-box each item, branch scenarios so the next question depends on the previous answer, run at least three parallel item variants per scenario, and combine formats such as a forced-choice SJT, a timed cognitive game and a short video response in one flow.
Yes, with conditions. Consent must be explicit and granular per data type (camera, audio, screen), retention must be defined and minimal, and no adverse decision may rest solely on an automated flag. A human must review every flagged case before a decision is made.

The question isn't hypothetical anymore. Candidates are using ChatGPT during job assessments, and the evidence suggests it's affecting scores in ways that matter for hiring validity. A peer-reviewed PLOS ONE study (PMC) submitted 63 AI-generated answers into a real university examination system under blind conditions: 94% went undetected, and they earned higher grades on average than submissions from actual students. That's not a lab artifact. It's a signal that assessment designs built for a pre-LLM world carry real integrity risk.
This guide is for TA leaders, recruitment managers, and assessment designers who need to act on that risk now. You'll get a format vulnerability map, redesign principles, a monitoring framework, an operational playbook, and a KPI dashboard template to track it all.
Prerequisites: Access to your current assessment configuration, a basic understanding of your item types, and decision authority (or stakeholder alignment) to modify at least one assessment flow.
Difficulty: Moderate. Time to implement initial controls: 1–2 weeks for policy and monitoring configuration; 4–6 weeks for full assessment redesign.
Not all formats carry the same risk. The exposure depends on three variables: whether the task is unsupervised, how text-heavy the output is, and whether there's any process evidence captured alongside the answer.
| Format | Vulnerability | Why | What to do instead |
|---|---|---|---|
| Async written responses | High | Static, unconstrained text output; no process capture | Time-box; add contextual constraints mid-session |
| Take-home assignments | High | Fully unsupervised; LLM can iterate indefinitely | Require annotated decision rationale; gate with follow-up interview |
| SJT (text-response variant) | Medium | Research (ScienceDirect, October 2024) shows only minor score shifts, but response patterns shift noticeably | Switch to forced-choice or ranking formats |
| Verbal reasoning / aptitude | Medium-High | LLMs can reason through many item types, especially if timed generously | Adaptive timing; parallel item pools |
| Coding/technical screens | Medium | Detectable via runtime checks; but prompt-injection is common | Use in-browser IDEs with screen capture; require explanation of logic |
| Live structured interview | Low | Real-time, interactive; hard to outsource mid-conversation | Pre-brief candidates on AI-use policy before prep window |
A Wiley Online Library study (published March 9, 2026, N=5,675) found fewer than 3% of applicants self-reported GenAI use overall, though up to 19% reported using it in specific assessment stages. Self-report underestimates actual use, but the stage-level variance tells you where to concentrate controls.
The most reliable mitigation isn't detection. It's making the task structurally harder to outsource.
Require process evidence, not just answers. A candidate who genuinely solved a problem can explain their reasoning under follow-up; a pasted LLM output usually can't. For a customer support SJT, don't just ask "What would you do?" Ask the candidate to identify which piece of information from the scenario they weighted most heavily and why. That constraint forces engagement with session-specific content an LLM hasn't seen.
Time-box and branch. Adaptive timing (shorter windows per item than a candidate would need to copy-paste and review) significantly reduces the LLM advantage on text-heavy items. Scenario branching, where the next question depends on the previous answer, makes prompt-reuse ineffective because the LLM-generated answer to question 2 is incoherent if it wasn't built on the candidate's actual answer to question 1.
Use parallel item pools. If every candidate sees the same warehouse operations scenario, that scenario will eventually be posted and shared. Build at least three structurally equivalent variants with different contextual details. This reduces prompt-matching and extends the shelf life of your item bank.
Combine format types. A modular assessment that includes a forced-choice SJT, a time-boxed cognitive game, and a short video response is much harder to systematically game than a single open-text task. Combining formats across competency dimensions also gives you more validity coverage per session.
Candidates using ChatGPT during assessments is an integrity question, but the tool you use to detect it matters enormously for fairness.
AI-text classifiers (tools like GPTZero) operate on probabilistic text features, and their error rates are operationally significant. Turnitin's own guidance (July 2024) assigns no detection score for outputs in the 1-19% AI probability range to avoid false positives. Even with a stated goal of under 1% false positive rate (Turnitin, March 2023), Jisc estimated (June 2025) that a 1% false-positive rate would generate approximately 4,800 wrongful flags annually across a typical institution-scale deployment. For a hiring context where a false flag can exclude a qualified candidate, that's not an acceptable sole evidence standard.
Behavioral integrity monitoring produces a different class of evidence. Observable signals such as focus-change events (tab switches, application changes), clipboard activity, keystroke cadence irregularities, and session-timing anomalies are behaviorally grounded and reviewable by a human. They don't identify AI use definitively, but they create an audit trail for reasoned human judgment.
Selection Lab's proctoring configuration captures camera, audio, and screen recordings with candidate consent, and surfaces integrity signals about whether tools like ChatGPT were used during the session. The deterrence effect of transparent recording disclosure is itself a meaningful control: candidates who know the session is recorded are less likely to attempt outsourcing in the first place.
On GDPR and EU AI Act compliance: candidate consent must be explicit and granular (camera/audio/screen as separate data categories where required), retention periods must be defined and minimal, and automated adverse decisions based solely on AI-generated flags are prohibited under EU AI Act high-risk system requirements. Integrity review must involve a human decision step.
Review-worthy anomalies by format:
Set a measurement baseline before deploying redesigned assessments, then review at 30 and 90 days.
Integrity monitoring KPIs:
Assessment validity KPIs:
Operational efficiency KPIs:
Implementation checklist:
The research picture is now clear enough to act on: generative AI use in job assessments is real, its impact on results depends on format, and text classifiers aren't a sufficient control mechanism. The operational response is assessment redesign plus behavioral monitoring plus structured human review. That combination, built into a modular assessment platform with native integrity tooling, is what converts a compliance risk into a manageable, auditable process.
Yes. A Wiley Online Library study of 5,675 applicants found fewer than 3% self-reported generative AI use overall, but up to 19% reported using it in specific assessment stages. Self-report underestimates real use. In a blind PLOS ONE experiment, 94% of AI-generated answers went undetected and scored higher than real submissions.
Asynchronous written responses and take-home assignments carry the highest risk because they are unsupervised and produce unconstrained text. Verbal reasoning tests sit at medium to high risk when timing is generous. Text-based SJTs and coding screens are medium risk. Live structured interviews are the hardest to outsource.
No. Text classifiers work on probabilistic features and produce false positives at a rate that matters in hiring. Turnitin assigns no score in the 1 to 19% probability range for that reason, and Jisc estimated a 1% false positive rate still yields roughly 4,800 wrongful flags a year at institutional scale. Behavioral signals plus human review are the defensible evidence standard.
Ask for process evidence rather than just answers, time-box each item, branch scenarios so the next question depends on the previous answer, run at least three parallel item variants per scenario, and combine formats such as a forced-choice SJT, a timed cognitive game and a short video response in one flow.
Yes, with conditions. Consent must be explicit and granular per data type (camera, audio, screen), retention must be defined and minimal, and no adverse decision may rest solely on an automated flag. A human must review every flagged case before a decision is made.