By the Podium research team — Tariq Shaban, Mike Rose, and Fenton Reid. Based on findings presented at SIOP 2026 and published in the openly available technical paper Behavioural, Model-Agnostic Integrity in Online Ability Testing.
In 2024, GPT-4 scored at the 12th percentile on a commercial quantitative ability test. Reassuring at the time. A few months later, OpenAI released o1. On the same test, it scored at the 95th percentile.
12th to 95th. Same items. Months apart.
That jump is the single best illustration of the problem facing assessment integrity in the AI era. And it explains why the dominant industry response — building “AI-proof” tests that current language models cannot solve — is structurally a losing strategy. Whatever is “AI-proof” against today’s models will be solved by next quarter’s. Content-level defences expire.
There is a different way.
This article summarises a year of research at Podium, presented at SIOP 2026 in May. It draws on a controlled experiment of 372 participants and a field deployment of more than 106,000 operational assessment sessions across six countries.
The headline finding is the one most people in the room reacted to:
> Roughly 1 in 5 unproctored candidates show behavioural signs consistent with attempted cheating.
> Roughly 1 in 20 actually gain a meaningful score advantage from it.
The gap between those two numbers is where the rest of this article lives.
The wrong fight
When generative AI first started solving cognitive ability tests at scale, the assessment industry’s instinct was to build harder tests. Items AI couldn’t reason through. Items requiring outputs current models couldn’t produce.
This is the AI arms race. It is structurally unwinnable.
The performance gap between GPT-4 and o1 — 12th percentile to 95th percentile in months, on the same instrument — is not an outlier. It is the trajectory. By the time you publish a new “AI-proof” item bank, the next model is already closing the gap. Six months later, it has solved it.
There is also a parallel line of defence: **AI content detectors** — tools that attempt to identify when a written response was generated by a language model rather than a human. These fail for three converging reasons.
Model fragility. A detector trained against GPT-4 may fail entirely against o1 or its successors. Each new model generation requires recalibration. Detection is always one generation behind.
Demographic bias. Peer-reviewed research (Liang et al., 2023) has shown that AI content detectors exhibit systematic discrimination against non-native English writers. In employment selection, where adverse impact carries both ethical and legal consequences, deploying a detector with that property is untenable.
Inapplicability. Most cognitive ability tests use multiple-choice formats. A candidate who copies a question into ChatGPT, receives the answer, and clicks the corresponding option produces no analysable text for a detector to evaluate.
If we cannot reliably detect AI in the content of candidate responses, we have to look somewhere else.
A different question
What stays stable — regardless of which AI tool a candidate uses or how capable that tool becomes — is how the candidate interacts with the test.
A person who consults an AI tool produces a distinct behavioural signature. The test window loses focus when they switch to another browser tab. Dwell times on the assisted items deviate from their session baseline. Response pacing across the session as a whole becomes irregular.
These are structural consequences of the act of task-switching. They are not properties of any specific AI model.
This is the conceptual pivot at the heart of Podium’s research:
> The question shifts from ”can we detect this AI?” to ”can we detect the act of seeking help?”
Different question. Different answer. Different durability.
The behavioural answer doesn’t expire when GPT-5, 6, or 7 releases. It doesn’t get harder as candidates become more proficient with prompting. It detects the structural overhead of leaving the assessment window — and that overhead is invariant to the cognitive efficiency of the tool the candidate is consulting.
The framework
No single mechanism can address every form of AI-assisted cheating. Same-device AI consultation, second-device photography, in-room human assistance, and proxy test-taking are different problems and require different defences. Podium’s integrity framework organises them into four layers we call **Design, Deter, Detect, Defend**.
Design. Dynamic item pools so no two candidates see the same test. Eliminates the oldest cheating method — sharing answer keys.
Deter. A short pre-test integrity message that explains the session is monitored. No threats. No legal-warning theatre. Just a clear expectation. We’ll show in a moment what this single message does to candidate behaviour.
Detect. A behavioural Confidence Score that runs silently throughout the test, capturing the signature of someone leaving the window for outside help. Model-agnostic by design.
Defend. Optional facial validation for high-stakes contexts. Closes the gap that telemetry cannot see — second-device cameras, in-room human assistance, proxy test-takers.
Each layer has a known strength and a known blind spot. The architectural principle: stack, don’t pick. Each layer’s blind spot is the next layer’s job.
What we found
Five studies in the technical paper test the framework from different angles. The summaries below cover the headline result of each. Full methodology, effect sizes, and limitations are in the open paper.
1. The Confidence Score works.
In a controlled experimental study of 372 UK participants — half instructed to complete the test honestly, half explicitly encouraged to use AI tools — the Confidence Score discriminated assisted from independent sessions with an AUC of .82 and a Cohen’s d of 1.47 between the two groups. By any standard in psychometrics, that is a very large separation.
The most important durability finding sits inside this study. Among candidates who cheated, self-reported AI proficiency had effectively zero correlation with their Confidence Score (r = .07, ns). Skilled AI users did not produce cleaner behavioural profiles than novices. The Score measures the structural overhead of task-switching — leaving the window, waiting, returning — not the cognitive efficiency of the AI tool. A more proficient user formulates better prompts and gets better answers. They still leave the test window. The blur event fires regardless.
This is what makes the Score durable across multi-year deployment cycles. It does not degrade as candidates become more skilled with AI.
2. Deterrence works.
In August 2025, Podium added a short pre-test integrity message to every cognitive assessment on the platform. A few sentences explaining that the session is monitored for behavioural integrity, and that patterns inconsistent with independent test-taking are recorded and reviewed. No threats. No countdown timers.
The proportion of candidates flagged at the most conservative behavioural threshold dropped from 6.5% to 1.9% — a two-thirds reduction. The pre/post comparison covered 3,072 demographically matched candidates and was replicated across 1,000 Monte Carlo iterations. Score effects on actual cognitive ability were tiny (d < 0.08): the message changed candidate behaviour without changing who got hired.
Deterrence is essentially free. No new technology, no candidate friction, no additional review burden. A short paragraph at the start of the test. There is no good reason for any organisation administering unproctored cognitive assessments not to be doing this.
3. Proctoring closes the gap behaviour can’t see.
The Confidence Score has a structural blind spot: it cannot detect a candidate who photographs the test screen with a second device and consults AI on the phone. The cheating behaviour leaves no signature on the assessment browser.
Facial validation closes that gap. In a between-client matched analysis of 2,194 sessions in Australia, sessions with facial validation enabled showed Verbal score reductions of d = 0.52 and Numerical reductions of d = 0.63 relative to matched unproctored sessions, with the effect concentrated in sessions with low Confidence Scores. Replicated across 1,000 Monte Carlo iterations.
Facial validation is the strongest single defence against AI-assisted cheating, and it is particularly necessary for the second-device methods that fall outside the Confidence Score’s coverage.
4. Fairness held.
The most important question for any vendor in this space: does the system disadvantage anyone unfairly?
Across 18,024 sessions, 12 client organisations, and 64 demographic comparisons spanning gender, age, and ethnicity, none exceeded a Cohen’s d of 0.30 — well below the threshold considered practically meaningful. At every Confidence Score threshold likely to be used in operational practice, the 4/5 rule was satisfied. The behavioural signals composing the Score — leaving the window, time deviations, pacing variability — proved effectively neutral with respect to who the candidate is.
The reason is structural. Unlike a content detector, which is trained against text features that happen to correlate with English fluency, the Confidence Score measures the act of seeking help. That act has no demographic signature.
One area we are explicitly continuing to monitor is the camera-based facial validation layer. Camera quality, lighting, and skin-tone detection introduce equity concerns that browser telemetry does not have. Our recommendation, published openly, is that the “low confidence” facial flag should not be used as a standalone integrity signal until parity is established at scale.
5. Many cheat. Few succeed.
The most counter-intuitive finding in the entire research programme is the gap between attempted and successful cheating.
Roughly 1 in 5 unproctored candidates produce behavioural signatures consistent with attempted cheating. Only roughly 1 in 20 actually gain a meaningful score advantage from the attempt.
These two numbers look like they should be the same. They aren’t.
A significant proportion of attempted cheating is ineffective. Candidates tab over to ChatGPT, fumble a prompt, receive an unhelpful or incorrect answer, run out of time, or fail to integrate the AI’s response into their answers within the time limit. Their behaviour is detectable — the test window loses focus, dwell times stretch, the response rhythm breaks — but their score barely moves.
A smaller subset cheats and benefits. Their score actually inflates. These are the cases that distort rank order and displace honest candidates from the top of the shortlist.
This gap is why the framework’s objective is more achievable than “stop every keystroke of cheating.” The objective is to prevent score inflation that affects who gets hired. That is a different — and more tractable — problem.
For organisations screening 1,000 candidates without proctoring, the 3–7% successful-cheating rate represents 30 to 70 artificially inflated scores entering the selection pipeline. At competitive selection ratios, that group displaces genuinely qualified candidates from the top of the rank order. The framework exists to close that gap.
What organisations should actually do
Six recommendations distilled from the full evidence base.
1. Stop trying to detect AI. Detect behaviour instead.
Behavioural signatures of task-switching are stable, model-agnostic, and don’t degrade as new AI models release. Content-based detection has been losing this fight since GPT-3.5.
2. Deploy a pre-test integrity message universally.
A short, clear statement. No threats, no countdowns. The data shows it cuts behaviourally flagged cheating by roughly two-thirds with no measurable impact on candidate scores. It costs nothing.
3. Add behavioural telemetry as a default.
A Confidence Score (or equivalent) that runs silently throughout the assessment and flags sessions inconsistent with independent test-taking. Use it as a guide for human review — never as an automated pass/fail.
4. Use proctoring where stakes warrant it.
Behavioural telemetry catches on-device cheating. It cannot see a phone camera over the candidate’s shoulder. For high-stakes selection — government, regulated industries, scholarship awards — facial validation closes that gap. For lower-stakes screening, it may be excessive.
5. Always route flagged sessions through human review.
Any algorithm in this space will produce false positives. Treat the flag as a signal to investigate, not a verdict to act on. Honest candidates deserve that protection.
6. Build fairness monitoring into your governance.
Do not run an integrity tool you have not checked for adverse impact. Do not rely on a tool the vendor has not checked either. Ask. The answers should be specific and independently verifiable.
The full Design + Deter + Detect + Defend stack is, in our view, the minimum defensible standard for high-stakes cognitive testing in 2026.
It is achievable. It is fair. And the data says it works.
The paper
The full technical paper, *Behavioural, Model-Agnostic Integrity in Online Ability Testing*, by Tariq Shaban, Mike Rose, and Fenton Reid, is openly available. Every method, every effect size, every limitation we identified, every recommendation we made is in there. We deliberately published it openly because the threat landscape moves too fast for any one vendor’s research to be the final word on this.
If you’d like a copy, leave a comment or reach out via our usual channels.
The arms race is the wrong fight.
Knowing what is happening in your assessment data is the right one.
This article expands on a series Podium is running on LinkedIn through May and June 2026, unpacking each of the framework’s layers in turn. To follow along, find Podium on LinkedIn or subscribe here for our next deep-dive.


