58,000 Students Face Retests Due to AI Proctoring—Where Did the Cost of Eliminating ‘Human Oversight’ Disappear?

In an entrance exam taken by 160,000 candidates, 58,000 are required to retake the test. The proctoring costs that were

By Kai

|

Related Articles

In an entrance exam taken by 160,000 candidates, 58,000 are required to retake the test. The proctoring costs that were supposed to be reduced have returned multiplied.

“If we leave it to AI, costs will go down”—have you used this phrase without questioning it?

The incident at Mexico’s largest university, UNAM, raises this very question. And this is not just a story about universities. When small and medium-sized enterprises consider cutting labor costs with AI, they are likely to encounter the same structural traps.

What Happened—A Numerical Look at the UNAM Disaster

In the summer of 2025, UNAM conducted its entrance exam entirely remotely for the first time. The number of candidates was approximately 160,000. AI monitoring software was implemented for proctoring, and no human proctors were present.

The results were as follows:

  • Percentage of high scorers (over 100 points): 16.3% (the annual average is 3.5%)
  • Number of candidates required to retake the exam: approximately 58,000
  • Anomalies in scores suspected of cheating: about 4.7 times the usual rate

In other words, AI did not “prevent” cheating; rather, it “missed” it. Alternatively, those who cheated exploited the gaps in AI monitoring. Either way, the outcome is the same. The reliability of the exam collapsed, resulting in enormous costs for retesting 58,000 candidates.

Consider this: conducting an in-person exam for 160,000 candidates involves securing venues, paying proctors, and arranging operational staff—rough estimates put this in the tens of millions to hundreds of millions of yen. By switching to remote exams with AI monitoring, a significant portion of these costs could be cut. This logic seems sound.

However, what is the cost of retesting 58,000 candidates? Venue reallocation, question re-creation, grading, notifications, and candidate support—all the costs that were supposed to be reduced may have returned multiplied. Moreover, the reputational cost of “UNAM’s entrance exam is not trustworthy” cannot be quantified in monetary terms.

Did you factor in the “cost of failure” in your cost-reduction calculations? This is the crucial point.

Why Was AI Monitoring Breached?—A Structural Issue

The mechanism of AI monitoring software works roughly like this: it tracks candidates’ faces and eye movements with cameras, detecting screen switches and suspicious movements. If any anomalies are detected, a flag is raised.

The problem lies in the accuracy of these flags and the response after a flag is raised.

In in-person proctoring, a human can pick up on vague feelings like, “That candidate seems to have their smartphone on their lap.” They can signal to another proctor and monitor together. Human oversight, while appearing less precise, has the ability to read context.

AI monitoring can only determine whether something fits predefined patterns. Placing a smartphone in a camera’s blind spot, searching on a different device, or sharing screens with a third party—AI is surprisingly powerless against these “unexpected cheating patterns.”

What’s even more troublesome is the speed at which methods to bypass AI monitoring are shared online. If one person finds a loophole, it can spread to thousands within hours. In in-person exams, local responses can be implemented venue by venue, but remote AI monitoring is structured such that “one loophole can collapse the entire system.”

It’s not just efficiency that scales; vulnerabilities scale too.

Similar Issues Occur in Other Fields—Unexpected Costs from AI-Related Fraud

A similar structural issue is occurring in the financial sector.

At Metro Bank in the UK, a fraud scheme involving the AI chatbot “Claude” was uncovered. Fraudsters withdrew funds from customer accounts without permission and used them to purchase credits for the AI service. The total amount lost was £14,000, approximately ¥2.7 million.

What’s noteworthy here is that the fraudsters managed to bypass the AI’s “guardrails”. Research by security experts indicates that bypassing these guardrails can be done without advanced hacking skills, and even so-called “script kiddies”—those who just pick up tools from the internet—can sometimes break through.

This is not to say that AI itself is at fault. The introduction of AI has created a new attack surface. A human representative might have questioned a suspicious transaction, but AI processes it without hesitation.

Is This Just a Concern for Large Enterprises?

You might think, “We’re not a university or a bank, so this doesn’t concern us.” However, the structure is the same.

When small and medium-sized enterprises adopt AI, the most common pattern is “cutting labor costs.”

  • Replacing customer support with AI chatbots
  • Assigning quality control tasks to AI image recognition
  • Switching accounting processes to AI-OCR and automated bookkeeping

All of these are valid directions. There are numerous cases where monthly labor costs of ¥300,000 can be reduced to ¥30,000 in tool expenses.

However, when replacing “tasks done by humans” with AI, are you ensuring that the “judgment” that humans unconsciously made is not disappearing as well? Whether you check this can dramatically change the outcome.

For example, in customer support, a human operator might escalate an issue if they sense, “This person seems angry.” An AI chatbot, on the other hand, will respond according to a template. Complaints can escalate, going viral on social media, and the response costs can skyrocket—this is a real occurrence.

Quality control faces the same issue. Experienced inspectors can detect defects with a sense of “the color is slightly different than usual.” AI image recognition may miss defects that do not match its training data. After shipping, complaints can arise, leading to recall costs of several million yen—how many years’ worth of the labor costs saved does that represent?

So, What Should Be Done?

It’s not about saying, “AI is dangerous, so let’s stop using it.” That would be a form of mental paralysis.

There are three key points.

1. Separate the Areas for AI and Human Oversight

AI excels in tasks with clear patterns that are repeated on a large scale—data entry, generating standard responses, and primary image classification. These can be entrusted to AI.

On the other hand, exception handling, contextual judgment, and the feeling that “something seems off” should remain in the human domain. In the case of UNAM, a combination of AI monitoring and human spot checks would likely have yielded different results.

2. Factor in the “Cost of Failure”

In ROI calculations for AI adoption, most companies overlook the “cost of mistakes made by AI.”

  • The customer churn cost when a chatbot provides incorrect guidance
  • The cost of handling complaints for defective products missed by AI quality control
  • The losses from fraudulent transactions that AI processed without question

Only by adding these to the “implementation costs + operational costs” can a true ROI be calculated. Deciding to implement based solely on the savings is akin to driving a car without insurance.

3. Start Small, Identify Loopholes, and Then Expand

This is the greatest strength of small and medium-sized enterprises. UNAM failed by rolling out to 160,000 candidates all at once. Large corporations assume a “company-wide implementation,” which means failures also scale.

For small and medium-sized enterprises, start by testing an AI chatbot with ten inquiries. Test AI image recognition with a hundred quality control checks. Identify the “patterns that AI misses” and design human checkpoints before expanding the scope.

The ability to fail small is a structural advantage for small and medium-sized enterprises. This should not be overlooked.

Is the Cost of ‘Human Oversight’ Really High?

Returning to the UNAM incident.

58,000 retests. Recreating questions. Reallocating venues. Loss of trust in the university. —When all these costs are added up, it is highly likely that the conclusion will be, “It would have been cheaper to hire human proctors.”

The costs that can be replaced by AI and those that only humans can handle. Confusing these two can lead to costs that were supposed to be reduced returning multiplied.

What I want to convey to small and medium-sized business owners is this:

Use AI not as a “replacement for humans,” but as a “tool to amplify human judgment.”

Automate ¥300,000 worth of routine tasks with a ¥50,000 AI tool. Focus the ¥250,000 worth of human time saved on exception handling and customer support, where AI struggles. This is the closest approach to the correct use of AI in small and medium-sized enterprises.

The notion that “handing everything over to AI will make it cheaper” is an illusion. Costs will only decrease in areas that can be entrusted to AI. When that boundary is misjudged, costs do not get reduced; they merely shift, often in a negative direction.

The 58,000 at UNAM paid that tuition. I want your company to avoid having to pay it.

POPULAR ARTICLES

Related Articles

POPULAR ARTICLES

JP JA US EN