UNAM's AI Exam Proctoring Failure Forces 58,000 Retakes
Nearly 160,000 students took UNAM's AI-proctored entrance exam remotely. Now 58,000 must retake it in person. Here's what actually went wrong.
Written by AI. Tyler Nakamura

If you've ever sat through a remotely proctored exam — watching your screen get locked down, your webcam activated, your every glance at the wrong angle suddenly suspicious — you already have a visceral sense of why this story matters. You know the low-grade dread of it. The way the system treats you as a potential criminal before you've answered question one. And you probably also know, or at least suspect, that the AI watching you has no idea what it's actually looking at.
That's not a fringe take. It's basically what happened at scale in Mexico this summer, and the numbers are hard to sit with.
Mexico's National Autonomous University — UNAM, the country's largest — ran its entrance exams remotely for roughly 160,000 applicants, using an AI-supervised proctoring system built around a lockdown browser and webcam monitoring. The goal was clear: modernize access, reduce the logistics burden, let students take a high-stakes test from wherever they were. Good intentions, real appeal. And then it fell apart.
According to Ars Technica, almost 58,000 of those applicants will now have to try again — this time in an official exam room, watched by a human proctor. Slashdot reports that UNAM concluded enough cheating had occurred that the results from the first test were essentially unusable — anyone who still wants a shot at admission needs to come back and do it again.
And per Ground News, UNAM has launched a major investigation into possible AI exam fraud, with the specific trigger being a record number of applicants passing its typically rigorous entrance exam. That anomaly is what set off the alarm bells.
So here's where it gets genuinely complicated, because there are at least two separate failure modes buried in this story, and they're not the same problem.
The AI proctoring system itself appears to have had real technical dysfunction — false positives, connectivity problems, errors in result processing. This is the headline problem, and it's a legitimate one. AI proctoring systems in general work by flagging behaviors that their training data associates with cheating: eyes moving off-screen, someone else appearing in the background, a second device's reflection, ambient noise that sounds like someone talking. These systems are pattern-matchers. They are not judgment-exercisers. Anyone who's spent time looking at the screenshots students post when they get violation notices — you'll recognize the absurdity pretty quickly. Flagged for a ceiling fan casting a shadow. Flagged for blinking too much. Flagged for wearing glasses. The false-positive problem in AI proctoring isn't theoretical; it's documented in forum threads and student appeals boards across the globe.
At UNAM's scale — 160,000 test-takers — even a modest false-positive rate produces thousands of wrongly flagged students. And when those flags are processed by automated result systems that themselves had errors, the compounding effect is how you end up with 58,000 people needing to retake an exam they may have passed fairly.
But there's a second problem that's arguably more systemic: the exam design itself may not have been built for the format. A thread in r/technology on Reddit floated a pointed critique — that running the same exam across an extended window with staggered scheduling basically guarantees questions leak from earlier test-takers to later ones, creating a structural advantage that no AI monitoring system can detect or prevent. This is unverified community speculation, not confirmed reporting, and UNAM hasn't publicly addressed whether this was a factor. But it's the kind of design question that any institution needs to answer before deployment, not after 58,000 people have already been burned.
That distinction — AI failure versus process failure — matters a lot. If the core problem is that the AI system was technically broken, that's a fixable bug. If the core problem is that the entire exam architecture was wrong for the format, no amount of proctoring software upgrades fixes it. You'd be putting better locks on a door that shouldn't be there.
What's notable about UNAM's response — triggering an investigation after an unusually high pass rate — is that it suggests the university may have known or suspected something was off before the complaints started rolling in. A record number of applicants passing a notoriously hard exam is exactly the kind of signal that should prompt a pause. The fact that the AI system apparently didn't catch enough of the problematic activity to prevent that anomaly is its own form of failure. Proctoring software that misses widespread cheating while simultaneously wrongly flagging honest students is somehow managing to be bad at both jobs at once.
Here's what I keep coming back to: there's a version of remote AI-proctored testing that could genuinely serve students, especially those who face real barriers to traveling to a testing center — cost, distance, disability, work schedules. UNAM's ambition to run remote exams for 160,000 applicants comes from a real place. Entrance exam access is an equity issue. The logistics of getting 160,000 people to physical testing sites is not trivial.
But the execution here collapsed the very students it was meant to serve. Nearly 58,000 of them are now being asked to do the whole thing over again, in person — which means whatever barrier the remote option was designed to remove just got handed back to them, after they'd already cleared it once. That's not a neutral inconvenience. For students who live far from campus, who work, who took time off to prepare — being told their completed exam doesn't count is a material harm.
The technology sector has a habit of deploying AI as a solution and then discovering the problem afterward. AI proctoring is a whole industry now — Honorlock, ProctorU, Examity, Respondus — and universities have been adopting these systems rapidly, especially after COVID-era remote learning normalized the format. What UNAM just demonstrated at unusual visibility is that scaling up AI proctoring isn't like scaling up a streaming service. The consequences of failure aren't buffering or downtime. They're 58,000 people's educational futures getting put on hold.
You don't beta-test on 160,000 students. And if somehow the situation requires it, you build the backup plan first — not as an afterthought once the damage is done.
The question UNAM and every other institution watching this story needs to answer isn't whether AI proctoring can work eventually. It's whether it works reliably enough, right now, for this exam, for these students. So far, the honest answer is that nobody has actually proven that case at scale. Mexico just ran the experiment for everyone, and 58,000 people are paying for the data.
— Tyler Nakamura, Consumer Tech & Gadgets Correspondent
We Watch Tech YouTube So You Don't Have To
Get the week's best tech insights, summarized and delivered to your inbox. No fluff, no spam.
More Like This
This Tool Treats Your Home Lab Like Infrastructure Code
RackPeek documents home labs as YAML code in Git. Brandon Lee shows how this infrastructure-as-code approach beats static diagrams and spreadsheets.
30 Self-Hosted GitHub Projects Trending Right Now
From media automation to AI chat apps, here are 30 trending self-hosted GitHub projects that put you back in control of your data and infrastructure.
Google Pays $250K for 16-Year-Old Linux KVM Flaw
A 16-year-old Linux KVM flaw called Januscape earned a $250K bounty after enabling guest VM escapes on Intel and AMD systems. Here's what it means for cloud security.
World's Fastest Drone Reclaims Record with V4
Discover how Peregreen V4 reclaimed the world's fastest drone title with a speed of 657 km/h.
TSRX Wants to Replace JSX — But at What Cost?
TSRX is a new syntax layer that lets you write React components with plain if statements and no return. Here's what junior devs actually need to know.
Copy.fail: The Linux Exploit That Works on Every Distro
A new privilege escalation vulnerability dubbed copy.fail affects all Linux distributions since 2017. Here's how the exploit actually works.
RAG·vector embedding
2026-08-07This article is indexed as a 1536-dimensional vector for semantic retrieval. Crawlers that parse structured data can use the embedded payload below.