AI hallucinations: why human oversight will remain key

Artificial intelligence has already made serious mistakes, even in courtrooms,
but what if its “hallucinations” aren’t the real problem—rather, the way we use it is?
By: Gastón Milano, CTO of Globant Enterprise AI.
What do a watchmaker and a doctor have in common?
Both are expected to achieve perfect accuracy. Once a phrase from another era, it now aptly describes the expectations we place on artificial intelligence.
In 2023, lawyer Steve Schwartz became the center of one of the most widely publicized AI hallucination cases. Representing Roberto Mata, who was suing an airline over a flight issue, Schwartz submitted a legal brief citing non-existent case law alongside real cases containing errors. The Manhattan court pointed out the inaccuracies, and Schwartz’s defense was that the brief had been drafted using ChatGPT.
Hallucinations are AI-generated responses that appear plausible and sound completely realistic but are, in fact, incorrect. They remain one of the major concerns in AI development. However, the landscape two years later looks entirely different.
We cannot ignore the problem; we must confront it.

AI is not infallible. Acknowledging this starting point is essential to understanding both its benefits and risks—whether interacting with a chatbot or developing AI-powered software. We cannot bury our heads in the sand.
In September 2024, a group of researchers published an article in Nature, analyzing 243 instances of distorted information caused by ChatGPT hallucinations. They categorized these errors into seven main types, providing valuable insights for the public, organizations, and future AI development.
Hallucinations can arise from data overfitting (interpreting data too literally), logical errors, reasoning or mathematical mistakes, unfounded inventions, factual inaccuracies, or text output errors.
While this may sound alarming, these errors represent only a tiny fraction of the 700 million weekly active users. Abandoning AI simply because it occasionally hallucinates would be like refusing to fly because, on average, four annual air accidents occur.
AI models are becoming increasingly accurate.
In February 2025, Sam Altman announced that ChatGPT 4-5 had halved the likelihood of hallucinations. In other words, a case like Schwartz’s is unlikely to recur (though it’s still not advisable to test this).
Models such as Gemini, DeepSeek, and Grok have also refined their training data architectures. Each model has its own comparative advantages, but in the Massive Multitask Language Understanding (MMLU) intelligence ranking, seven models already achieve an 80% or higher success rate in their responses.
Competition is creating a virtuous cycle, driving the development of increasingly accurate models. One of the most powerful tools in this evolution is Retrieval-Augmented Generation (RAG).
With RAG, before generating a response, the language system can retrieve contextual information from external sources not included in its original training—a trial-and-error learning system.
The RAG market, valued at $1.2 billion in 2024, is projected to grow at a compound annual rate of 49.1% between 2025 and 2030, according to a Grand View Research report.
Human oversight as a quality guarantee.
We can understand hallucinations and leverage advanced models, but AI adoption will not succeed without human supervision—even the most sophisticated agents require it. The human-in-the-loop model is the new workplace dynamic.
Anyone expecting to stop thinking and leave everything to AI will be disappointed. That would be as implausible as imagining an assistant without a boss. After all, that’s what the system is for.
To mitigate hallucinations, we must ask whether a response falls into one of the previously mentioned error categories and verify its data sources.
Imagine this process within a company: an AI agent records a client interview capturing their requirements, then creates a story map, develops software, presents an MVP, and conducts testing.
A process that previously took weeks can now be completed in hours. Yet, the final decision rests with a human. If any stage involved hallucinations or insufficient data, someone must act as the quality guarantor.
AI literacy is already one of the most sought-after skills in the labor market, according to the World Economic Forum. Technology can boost productivity and enable a company to serve twice as many clients with the same workforce—but each individual’s critical thinking remains the ultimate check.
One lesson from childhood remains true: fear should not paralyze us. Hallucinations are rare, being corrected, and manageable—but they exist.
We can think of it in human terms:
Would you stop hiring someone because they might occasionally make a mistake?
The same principle applies to AI.

