This is an AI-generated audio overview produced using Google NotebookLM. The voices you’ll hear are synthetic. The content is derived from Podium’s openly available technical paper on behavioural integrity in cognitive testing.
In 2024, GPT-4 scored at the 12th percentile on a commercial quantitative ability test. A few months later, OpenAI released o1. On the same test: 95th percentile.
That jump — and what it means for the assessment industry — is the starting point of this audio overview. Over sixteen minutes, two AI-generated hosts walk through Podium’s SIOP 2026 research: a controlled experiment of 372 participants and a field study of more than 106,000 operational assessment sessions across 34 clients in six countries.
What’s covered
→ Why content-based AI detection is a losing fight
→ The AI arms race — and why it cannot be won by building harder tests
→ Behavioural integrity as the alternative: measuring *how* candidates interact with the test rather than *what* they answer
→ Podium’s four-layer framework — Design, Deter, Detect, Defend
→ The Confidence Score, what it captures, and what it can’t see
→ The headline finding: 1 in 5 unproctored candidates show signs of attempted cheating, but only 1 in 20 actually gain a meaningful score advantage — and what that gap means for the framework’s design
→ Fairness across demographic groups: 18,024 sessions, 64 comparisons, no practically meaningful bias
→ Six practical recommendations for any organisation administering unproctored cognitive testing
Prefer to read?
The same material is available as a written deep-dive in our companion Substack article:
.





