Did UNAM’s AI Proctoring Fail? An Analysis of the Retest for 58,000 Applicants and the Online Entrance Exam Cheating Controversy

The information in this article is current as of August 2026. UNAM’s technical committee is still investigating the causes of the incident and assigning responsibility, and the full implementation details of the control exam may continue to change.

In 2026, the National Autonomous University of Mexico, or UNAM, moved its undergraduate entrance exam entirely online for the first time and introduced AI proctoring, a secure browser and human review. The goal was to prevent applicants outside Mexico City from having to travel long distances for a single exam. After the results were released, however, the proportion of applicants who answered more than 100 questions correctly rose from 3.5% over the previous five years to 16.3%. The university ultimately decided to require approximately 58,000 applicants to take an in-person control exam to reconfirm their admission eligibility.

This controversy cannot be reduced to the claim that “AI proctoring failed to catch cheating.” The system did generate alerts during the exam, and UNAM cancelled the admission process for approximately 2% of applicants. Even so, there remained a significant gap between the monitoring records and the overall score distribution. Second devices, camera blind spots, external assistance and alleged paid proxy-testing services cannot be addressed simply by adding another camera. For schools and companies considering AI proctoring, online testing or remote assessment, the UNAM incident provides a practical stress test of the entire system.

What Happened During UNAM’s Online Entrance Exam? From the First Remote Exam to 58,000 Retests

UNAM previously required most undergraduate applicants to take the entrance exam at designated testing centres. In 2026, the university moved the entire exam online for the first time. The testing period ran from May 23 to June 10, and applicants were required to use a computer, camera, microphone and secure browser that met technical requirements and log in during their assigned time slot.

A total of 191,306 people registered, and 158,712 actually took the exam. The university stated that the online format could reduce transportation, accommodation and travel burdens for applicants outside Mexico City and abroad while also increasing the actual participation rate. Early in the testing period, UNAM announced that 55,840 people had taken the exam during the first weekend, that the platform had operated stably and that 498 applicants had been suspended or removed from the system for violating the rules.

On July 16, one day before the results were released, UNAM stated that approximately 2% of applicants had been removed from the admission process for violating the admission guidelines and university regulations. At that point, the university was still preparing to publish the admission list and proceed with registration as originally scheduled.

Why Did the Share of High Scores on the UNAM Entrance Exam Rise from 3.5% to 16.3%?

What caused the entire situation to spiral was not the small number of exams terminated by the system, but the overall score distribution that appeared after the results were released on July 17.

The UNAM undergraduate entrance exam contains 120 questions. Between 2021 and 2025, applicants who answered more than 100 questions correctly accounted for an average of approximately 3.5%. In 2026, that proportion rose to 16.3%. Unusually high scores were also concentrated in highly competitive programmes such as medicine and law, prompting some faculty members, applicants and statistical researchers to question whether the results still reflected applicants’ actual ability.

After comparing score distributions from 2021 to 2026, Raúl Rojas, an artificial intelligence researcher at the Free University of Berlin, divided the 2026 scores into two statistical groups. One group remained close to the distribution seen over the previous five years, while the other was heavily concentrated in the high-score range. He argued that this deviation was difficult to explain solely by a sudden reduction in exam difficulty or a general improvement in applicant performance and was more consistent with the hypothesis that a large number of applicants had received external assistance.

However, a statistical anomaly cannot directly prove that a particular applicant cheated. It can only show that the overall distribution differed too greatly from previous years and that test records, platform data and question security required further examination. This is also why UNAM did not simply cancel all high scores and instead arranged an in-person control exam.

How Did Territorium Life Proctor the Exam? AI Alerts, a Secure Browser and Human Review

UNAM’s online examination platform was provided by Territorium Life. The system included identity verification, facial and voice recognition, camera and microphone recording, screen activity monitoring, a secure browser and abnormal-behaviour alerts. The security measures were also designed to prevent the same account from being used simultaneously on multiple devices and to restrict switching to other applications, copying and pasting, screenshots and external displays.

The AI system did not independently decide whether an applicant had cheated. UNAM stated that the system was responsible only for detecting potentially abnormal behaviour and generating alerts. University personnel then reviewed the video, audio and activity records before deciding whether to terminate or invalidate an exam.

After the controversy emerged, Territorium Life stated that the platform had operated as expected within the company’s technical scope and had supported more than 10,000 applicants simultaneously in a single session. The company also argued that the integrity of a high-stakes exam could not depend solely on technical monitoring and also involved question-bank security, exam design and university administration.

This explanation did not eliminate the controversy. A platform can ensure that an applicant does not switch websites on the same computer, but it is much more difficult to determine whether a phone, tablet or second computer is placed elsewhere in the room. The system can record what happens within the camera frame, but areas outside the frame remain blind spots.

Why Could AI Proctoring Not Prevent Cheating? Structural Weaknesses in Remote Exams

AI proctoring is closer to an anomaly-recording and risk-screening system than a closed examination room capable of proving absolute fairness. The UNAM incident shows that even when a platform records video normally, locks the browser and generates alerts, the overall exam may still be affected by external conditions that the system cannot see.

Why Can a Secure Browser Lock the Main Computer but Not Control a Second Device?

A secure browser can restrict the computer being used for the exam, such as by blocking search engines, preventing window switching, disabling copying of questions or restricting certain keyboard shortcuts. However, these restrictions usually apply only to the device on which the software is installed. If an applicant places a phone, tablet or second computer outside the camera frame, the primary testing device may continue to appear completely normal. The screen does not change, the secure browser remains active and the system may therefore produce no alert. Research on remote proctoring has also found that these tools may deter spontaneous or obvious violations but may not detect pre-planned cheating methods, such as using a second device, placing materials outside the camera frame or receiving assistance from another person at a remote location.

AI Alerts Can Identify Anomalies but Cannot Independently Determine Cheating

Looking down, looking away from the screen, hearing sounds in the background or briefly moving out of the camera frame may all trigger alerts. However, these behaviours may also result from an internet interruption, a family member passing by, equipment problems or a medical condition. During the exam, UNAM specifically explained that environmental noise alone would not automatically lead to an exam being cancelled. AI-generated alerts therefore had to undergo human review. University personnel needed to examine the surrounding video and other records rather than relying on a single signal to determine a violation. This design can reduce false positives, but it also increases the review burden. When more than 150,000 people take an exam and each applicant generates lengthy video, audio and activity records, it is difficult for a human team to review every record in full. The practical process usually depends on the system first identifying high-risk segments and then concentrating human review on the flagged cases.

A lack of alerts does not mean that no problem occurred. An alert also does not necessarily mean that cheating took place.

Leaked Questions and External Assistance May Leave No Visual Evidence

UNAM confirmed that individuals or companies had offered applicants and their families services allegedly intended to violate exam rules. On July 3, the university filed a criminal complaint and requested that the relevant authorities investigate those activities. At this stage, it is not possible to attribute every unusually high score to the same method or confirm that every set of questions was leaked. However, when an exam spans multiple dates and sessions, the rate of repeated questions, the number of versions and the way the question bank is managed directly affect the level of risk. Camera monitoring deals with an applicant’s behaviour at the time of the exam. Whether questions were circulated in advance, whether answers were passed between sessions or whether third parties provided remote assistance belongs to a different category of security problems. These issues cannot be addressed solely through the same proctoring software.

Why Did the Abnormal Score Distribution Trigger Controversy Only After the Results Were Released?

During the examination period, UNAM mainly focused on platform stability, rule-violation alerts and individual applicants’ behaviour records. The broader problem only became visible after the results were released, when researchers compared the 2026 score distribution with those of the previous five years. This suggests that although the university had established behavioural monitoring during the exam, its statistical validation before releasing results did not stop the anomalous outcome in time. For a large-scale exam, monitoring applicants and checking whether the score distribution is reasonable are two separate tasks. The former deals with individual behaviour, while the latter examines whether the overall distribution has changed suddenly. Conducting only one of these checks may allow problems on the other side to go undetected.

From Cancelling 2% of Applicants to Re-Examining 58,000: Why Did UNAM’s Decision Cause a Backlash?

After the results were released, UNAM first established a technical committee to examine the online examination process, platform data and overall results. On July 24, the university suspended the document-submission and new-student registration process originally scheduled to begin on July 27 while waiting for the committee’s recommendations. The technical committee ultimately recommended an in-person control exam. Two groups were required to participate: applicants who had already received a 2026 admission offer and applicants whose original 2026 scores reached the lowest admission threshold for their programme, campus and study format during the period from 2021 to 2026.

This rule did not require applicants from the previous five years to return and retake the exam. The lowest admission scores from the previous five years were used only to define the group of 2026 applicants eligible for the control exam. An estimated 58,000 people would be invited to compete for approximately 22,000 places originally allocated through the entrance exam. The control exam would be held at physical locations under human supervision, with a sufficient number of different test versions. Final admission results would be reassigned according to the control exam scores rather than directly using the results of the first online exam. UNAM also announced that it would add testing locations outside Mexico City to reduce the travel burden on applicants from other regions. New students for the 2026 academic year were expected to begin classes on August 31, while students already enrolled at the university were expected to resume classes on August 17.

The arrangement dissatisfied both sides. Applicants who had already received admission offers argued that the university had failed to distinguish between honest test takers and suspected cheaters, forcing everyone who met the threshold to bear the cost of a system failure. Applicants who had originally been rejected argued that anomalous scores may have raised the admission cut-off and displaced people who otherwise might have been admitted. Former UNAM Secretary-General Patricia Dávila left her position on August 1, and Javier de la Fuente subsequently succeeded her. Because the timing of her departure overlapped with the exam controversy, the media treated it as an important personnel change following the incident. However, the university did not publicly state whether the controversy was the sole reason for her departure.

How Can AI Proctoring Be Used Safely? Question-Bank Security, Statistical Review and Appeals

The UNAM incident does not prove that all online examinations are unreliable or that AI proctoring has no value. The problem is that if a high-stakes exam treats AI proctoring as its only line of defence, a cheating method that the system cannot see may undermine the credibility of the entire result.

Do Not Treat AI Proctoring as the Only Line of Defence

Cameras, microphones and secure browsers are suitable for recording visible behaviour and can prevent some spontaneous violations. They cannot replace question-bank security, test-version management and post-exam statistical analysis. Large examinations can increase random question selection, vary the order of questions, create multiple equivalent test versions and control the rate at which questions are repeated across sessions. When an exam lasts several days, the organiser must also assess whether questions may circulate through social platforms or paid groups. Remote-proctoring research generally concludes that monitoring tools are better suited as one layer in a multilayered defence. Against planned external assistance or the use of a second device, software alone is unlikely to provide complete protection.

Check Historical Score Distributions Before Releasing Results

Before publishing results, organisers should compare historical average scores, the proportion of high scores, correct-answer rates for each question and score distributions across different programmes. If one year suddenly produces a large number of high scores, the examination should not wait for social-media discussion before conducting a review. Statistical review should also examine more than the overall average. Whether high scores are concentrated in particular sessions, test versions, testing regions or device environments may provide different clues. Exam organisers can define conditions in advance that would pause the release of results. For example, if the proportion of high scores exceeds the historical range by a certain margin or some questions show implausibly identical answer patterns, the organiser could conduct a human review before deciding whether to publish the results.

Preserve Complete Records and Provide an Understandable Appeals Process

Alerts generated by AI proctoring should not consist only of internal codes. When an applicant is disqualified, they should know which video segment and which rule are involved and whether they may request human reconsideration. After the controversy, UNAM allowed some applicants to review the evidence used to cancel their exams and processed requests from unsuccessful applicants for score reviews. If these mechanisms are written into the rules before the exam, applicants will have a clearer understanding of when they can appeal, who will review the case and how long a response will take. An appeals process can also reduce the harm caused by AI errors. An anomaly alert is only the beginning of an investigation and cannot be treated as equivalent to proof of cheating.

Decide in Advance What to Do If the System Fails

High-stakes examinations should not wait until score problems emerge before discussing whether a retest is necessary. Organisers can define responses for different levels of failure in advance. For example, if there is clear evidence in an individual applicant’s record, only that applicant’s result may be cancelled. If question security is compromised in a specific session, only that session may be retested. If the overall score distribution loses credibility, a broader control exam may then be activated. The scope of the retest, transportation assistance, accommodations for special needs, appeal deadlines and the method for calculating new scores should all be included in the contingency plan before the exam. If the system fails, the organiser will then have more options than simply accepting every result or forcing everyone to start over.

The UNAM case does not prove that AI proctoring is completely useless. The system did generate real-time alerts, preserve examination records and cancel the eligibility of some applicants found to have violated the rules. The problem is that these functions were still insufficient to prove that the overall score distribution was credible.

Second devices, external assistance, leaked questions and abnormal score distributions cannot be addressed by installing an additional camera. If high-stakes examinations are to retain a remote format, proctoring software, question-bank management, statistical review, human verification and an appeals system must be placed within the same process. If any one part is unprepared, applicants will still bear the cost of retesting and delays.

Frequently Asked Questions

What Happened During UNAM’s 2026 Entrance Exam?

UNAM moved its undergraduate entrance exam entirely online for the first time, but an unusually high score distribution appeared after the results were released. The university established a technical committee to investigate and ultimately decided that approximately 58,000 applicants who met specified score conditions in 2026 would take an in-person control exam.

Why Did UNAM Suspect Cheating in the Online Entrance Exam?

Between 2021 and 2025, applicants who answered more than 100 questions correctly accounted for an average of approximately 3.5%. In 2026, that proportion rose to 16.3%. Statistical analysis showed that a group in the high-score range differed significantly from previous years, but the statistical anomaly itself could not prove that any individual applicant had cheated.

Did UNAM’s AI Proctoring System Independently Decide That Applicants Had Cheated?

No. The AI system mainly recorded video, audio and screen activity and generated alerts for potentially abnormal behaviour. UNAM personnel then reviewed the evidence and made the final decision.

Do All of the More Than 150,000 Applicants Need to Retake the Exam?

No. The control exam applies to applicants who received a 2026 admission offer and those whose original 2026 scores reached the lowest admission threshold for their programme, campus and study format during the period from 2021 to 2026. The estimated number is approximately 58,000.

Can AI Proctoring Completely Prevent the Use of a Second Device or External Assistance?

No. A secure browser can restrict the main computer used for the exam, while cameras and microphones can record parts of the environment. However, another phone, camera blind spots and pre-arranged external assistance may still avoid detection. Technical research has also found that the anti-cheating measures of several remote-proctoring tools can potentially be bypassed, so they should not be treated as the sole examination-security mechanism.

SUPPORT FENGNIII

喜歡這篇文章嗎?

如果這篇內容對你有幫助,可以透過小額贊助支持本站持續整理更多日文、韓文、旅行與數位工具內容。

小額支持本站

付款將由藍新金流安全處理