There Were No Citations from AI-Written Reports—How Much Can Small Businesses Bet on ‘Lying AI’?

Conclusion To put it simply, AI lies. The question is, "In which tasks can we tolerate those lies?" When tracing the so

By Kai

|

Related Articles

Conclusion

To put it simply, AI lies. The question is, “In which tasks can we tolerate those lies?”

When tracing the sources of a report written by AI, it turned out that the paper did not exist. The author’s name, journal name, and DOI were all fabricated. This is not an urban legend. In 2023, a lawyer who had ChatGPT conduct case law research cited six non-existent cases in court and faced sanctions.

The “plausible lies” generated by AI—this phenomenon known as hallucination—is not completely eliminated even in GPT-4. According to OpenAI’s technical report, while the hallucination rate of GPT-4 has improved from GPT-3.5, its accuracy in factuality assessment remains around 80%. This means there is a possibility of stating something factually incorrect once in every five attempts.

Upon hearing this, one might think, “Then it’s unusable,” or “If it’s correct four out of five times, it depends on how we use it.” This is the dividing line.

What small businesses should really consider is not whether “AI can be trusted,” but rather “Is the cost of verifying AI outputs lower than the cost of having humans do it?”

Facing Cases of Actual Damage Caused by AI Lies

First, let’s accurately grasp what is happening.

Case of Fabricated Case Law (2023, USA)
A lawyer in New York used ChatGPT to prepare legal documents. All six cases cited by the AI were fictional. The court imposed a fine of $5,000 (about 750,000 yen). The damage to the lawyer’s career cannot be quantified in monetary terms.

Safety Issues with AI Code
A study from Stanford University found that developers using AI coding assistants wrote code that was more vulnerable to security issues compared to those who did not use them. Even more troubling is that developers assisted by AI were more confident that their code was safe. In other words, AI not only lies but also lowers human verification awareness.

Rate of Inclusion of Non-Existent Citations
A survey from Purdue University revealed that over 52% of ChatGPT’s responses to questions on Stack Overflow contained inaccurate information. However, due to the fluent writing style, humans failed to detect errors in 77% of cases.

This presents an essential risk for small businesses. AI’s lies do not appear as lies.

Dividing Tasks into Three Categories Based on ‘Verification Costs’

It is meaningless to say abstractly that “AI is dangerous” or “AI is convenient.” We need to estimate the verification costs for each task and categorize them as usable or unusable. This is the only practical criterion for the field.

Rank A: Verification cost is almost zero. Use it immediately.

① Internal Meeting Minutes and Summaries

  • AI transcribes and summarizes recordings of meetings. Since these are internal documents not shared externally, minor errors have almost no real impact.
  • Cost if done by humans: Approximately 30 to 60 minutes of work for a one-hour meeting. At an hourly wage of 2,000 yen, this amounts to 1,000 to 2,000 yen per session.
  • AI cost: About 50 to 100 yen per session using Whisper and GPT-4.
  • Cost reduction rate: Over 95%. Verification can be done by “quickly reviewing” in about 5 minutes.

② Drafting Internal Emails and Chats

  • Drafting routine communications, daily reports, and report emails. Since the individual reviews them before sending, the verification cost is just the “reading time.”
  • For an employee writing 100 emails a month, a 5-minute reduction per email results in over 8 hours saved monthly.

③ Idea Generation and Brainstorming

  • Naming ideas for new products, potential catchphrases, and angles for projects. There is no correct answer in this domain, so the concept of “lies” does not exist.
  • Outsourcing could cost 50,000 to 200,000 yen per project. With AI, 50 ideas can be generated in just a few minutes, from which humans can choose.

Rank B: Usable with verification costs. Calculate cost-effectiveness.

④ Drafting Blogs, Social Media Posts, and Press Releases

  • Since these are external documents, fact-checking is essential. However, compared to writing from scratch, having AI draft and then having humans revise is faster.
  • If written from scratch by humans: 2 to 4 hours per piece (30,000 to 100,000 yen if outsourced).
  • AI draft plus human revision: 30 minutes to 1 hour. Including fact-checking, about 1.5 hours.
  • Cost reduction rate: 50 to 70%. However, the “skill of the person checking” is necessary.

⑤ Initial Stages of Research and Market Analysis

  • Understanding industry trends, creating competitor lists, and organizing technology trends. Have AI create a “draft” for humans to verify.
  • Caution: Always verify the numbers and sources provided by AI against the original sources. Skipping this step could lead to a repeat of the fabricated case law incident.
  • Verification cost: 1 to 3 hours per report. Even so, it is 60% faster than researching from scratch.

⑥ Programming Assistance

  • Generating routine code, explaining existing code, and creating test code. A study on GitHub Copilot indicated a 55% increase in development speed.
  • However, code related to security, authentication, and payment logic must be reviewed by humans. The cost of code review is 30 minutes to 2 hours per feature.
  • What is reduced is “writing time,” not “thinking time.” Misunderstanding this can lead to accidents.

Rank C: Do not leave it to AI. Humans should handle it.

⑦ Drafting Contracts and Legal Documents

  • A single typographical error can lead to hundreds of thousands of yen in damages. It is acceptable to “reference” AI drafts, but the final decision must always be made by an expert.
  • Verification cost: 300,000 to 3,000,000 yen per review by a lawyer. Cutting corners here could lead to sanctions and loss of credibility, as seen with the aforementioned lawyer.
  • The potential damages from accidents in this area far outweigh the costs that can be saved by AI.

⑧ Grant Applications and Administrative Documents

  • Accuracy of figures and compliance with laws are absolute prerequisites. AI can write “plausible” text, but if it does not meet the requirements, it will be rejected. Inaccurate statements can lead to claims for repayment or penalties in the worst case.

⑨ Formal Proposals and Estimates to Customers

  • There have been actual cases where estimates included figures directly from AI, resulting in losses due to underpricing. Documents that are directly related to trust with customers require human final judgment.

The Right Way for Small Businesses to Bet

After organizing this information, the structure is simple.

Verification Cost Accident Damages Judgment
Rank A Almost Zero Small Use it with full force now
Rank B Moderate Moderate Calculate cost-effectiveness and use it
Rank C High Fatal AI as assistance. Judgment by humans

The realistic budget for AI-related expenses that small businesses can allocate monthly is likely around 30,000 to 100,000 yen. ChatGPT Plus costs 3,000 yen per month, Claude also 3,000 yen, and even with dedicated tools, a budget of 50,000 yen can create a sufficient environment.

Tasks that used to cost 500,000 yen a month for outsourcing can now be managed with 50,000 yen for AI tools plus internal verification labor. This “450,000 yen difference” is the true value of AI for small businesses.

However, there is a condition to obtain that 450,000 yen.

There must be a “human capable of checking AI outputs” within the company.

Without this, all tasks in Rank B will fall to Rank C. If there is no one to verify, AI outputs are nothing more than “untrustworthy text.”

So, What Should Be Done?

You only need to do three things.

1. First, introduce AI into Rank A tasks. By the end of this week.
Meeting minutes, email drafts, idea generation. The risks are almost zero, and the effects are immediate. This is where you build the “muscle” for using AI within the company. Monthly costs will be 3,000 to 5,000 yen. There is no reason not to do it.

2. Implement Rank B together with a “verifiable human.”
Blogs, research, programming assistance. AI drafts plus human finishing. Once this pair starts working, some companies can save 200,000 to 500,000 yen in outsourcing costs monthly. However, if the verifier’s skills are low, accidents can happen. First, train the verifiers.

3. If using AI in Rank C, limit it to “drafting” only.
Final judgment must be made by humans. Do not skimp on the costs of expert reviews. Allocate part of the money saved by AI to this.

AI lies. This is not a defect but a current specification.

The issue is not whether it lies. It is whether your company has the system in place to recognize those lies.

We have entered an era where a monthly investment of 50,000 yen in AI can manage 5 million yen worth of tasks annually. However, if you misuse that 50,000 yen, it could lead to damages of 5 million yen.

If you are going to bet, bet where you can win.

POPULAR ARTICLES

Related Articles

POPULAR ARTICLES

JP JA US EN