The Price of AI is Crumbling from Two Directions—Interpreting Anthropic’s New Model and AMD’s $8 Billion Acquisition in the ‘50,000 Yen a Month’ Context

Conclusion To put it simply, the cost of using AI has started to crumble simultaneously from both the "top" and the "bot

By Kai

|

Related Articles

Conclusion

To put it simply, the cost of using AI has started to crumble simultaneously from both the “top” and the “bottom.”

Anthropic has announced its new model, “Claude Sonnet 4.” Priced similarly to its predecessor Sonnet 3.5 (at $3 per million tokens for input and $15 for output), it boasts significantly improved coding performance. In other words, users can now access a “smarter brain for the same amount of money.”

In the same week, it was reported that AMD is acquiring the AI startup World Labs for approximately $8 billion. While World Labs is a company focused on 3D spatial generation AI, the essence of this acquisition is more significant. The key takeaway is that AMD is seriously challenging Nvidia’s dominance in the “GPU monopoly” structure.

When we look at these two pieces of news side by side, a structural pattern emerges.

  • Upstream (Model Layer): Performance increases at the same price = effective cost reduction
  • Downstream (Hardware Layer): Intensifying GPU competition = pressure to reduce procurement costs

The price of AI is being squeezed from both the top and the bottom.

It would be dangerous to dismiss this as merely a “big company issue.” This is a matter that directly impacts small and medium-sized enterprises (SMEs) that are using AI with a monthly API budget of 50,000 to 300,000 yen.

—

The Implications of Acquiring “30% More Productive Subordinates for the Same Amount of Money”

First, let’s take a closer look at the model side with specific numbers.

According to Anthropic’s benchmarks, Claude Sonnet 4 shows a clear improvement in accuracy for coding tasks compared to Sonnet 3.5. An increase in performance at the same price means that the same tasks can be completed with fewer tokens and fewer retries.

Let’s rephrase this in practical terms.

Imagine a company with a monthly API budget of 100,000 yen using Claude Sonnet 3.5. They might be consuming approximately 3 million tokens (total input and output) per month for tasks like summarizing meeting minutes, drafting emails, organizing data, and generating simple code.

What happens when they switch to Sonnet 4?

  • The accuracy of processing the same tasks improves, reducing the need for “do-overs.”
  • The probability of achieving the desired quality in a single output increases.
  • As a result, it’s possible to handle the same workload with tokens worth about 70,000 to 80,000 yen.

What to do with the saved 20,000 to 30,000 yen? They can delegate new tasks to AI. Automated generation of FAQs for customer support, drafting estimates, screening candidates—these can be areas that were previously neglected due to budget constraints.

This is a matter of 20,000 to 30,000 yen per month. Annually, that amounts to 240,000 to 360,000 yen. For SMEs, this is close to the monthly salary of one part-time employee that can be saved simply by upgrading the model.

The action is simple: just rewrite the model specification in the API. No capital investment needed.

—

AMD’s $8 Billion Acquisition is a “Forecast of Lower GPU Prices”

Next, let’s discuss the hardware side.

The aim of AMD’s acquisition of World Labs is not just to acquire 3D generation AI technology. AMD has been chasing Nvidia in the AI inference chip market for the past few years. With the introduction of the MI300X series, they are expanding adoption among cloud providers. The acquisition of World Labs signals AMD’s intent to secure not only hardware but also the AI application layer.

This move is indirectly but surely relevant to SMEs.

Intensifying GPU competition → Lower procurement costs for cloud vendors → Price competition in API usage fees.

Currently, Nvidia’s H100/H200 GPUs are effectively the standard for AI inference. Supply is limited, and prices remain high. With AMD’s serious entry into the market, along with Google’s TPU and AWS’s Trainium/Inferentia, the options for chips on the cloud side will increase.

As options increase, prices will decrease. This is basic economics.

When and by how much will prices drop? Honestly, it may be hard to see dramatic changes within a six-month to one-year timeframe. However, there is a plausible scenario where API costs could drop to about 50-70% of current prices between late 2025 and 2026.

Looking at the price trends for OpenAI’s GPT-4 provides insight. When GPT-4 was announced in March 2023, its API cost was $30 per million tokens for input. By 2024, GPT-4o had dropped to $5. That’s a reduction to one-sixth in just one year. This price collapse is a result of both model efficiency improvements and hardware competition.

The same structural dynamics are now being replicated with the movements of Anthropic and AMD.

—

What Should a Company with a 50,000 Yen Budget Do “Now”?

Now we get to the main point. It would be meaningless to end this with a simple, “That’s an interesting story.”

Let’s break down the AI budgets of SMEs into three layers and outline specific actions.

Below 50,000 Yen (API Usage Only)

Immediate Action: Test switching models.

If you are currently using GPT-4o or Claude Sonnet 3.5, I encourage you to test switching to Sonnet 4 for a week. Use the same prompts and compare the output quality and token consumption. If you feel that “retries have decreased,” that directly translates to cost savings.

With a budget of 50,000 yen, the amount saved by switching models could be around 10,000 to 15,000 yen per month. While that may seem small, it adds up to 120,000 to 180,000 yen annually. That’s enough to automate one new task.

100,000 to 200,000 Yen (API + Some Tool Fees)

Immediate Action: Design for multi-model operations.

Companies in this budget range often use the same model for all tasks, which is wasteful.

  • Simple classification/extraction tasks → Claude Haiku (cheap, fast)
  • Text generation/summarization → Sonnet 4 (balanced)
  • Complex analysis/strategy planning → Opus (expensive but necessary for high-accuracy scenarios)

By using different models for different tasks, there are cases where the same workload can be handled at 30-40% lower costs. This is possible not because “AI performance has improved,” but because “options have increased.”

200,000 to 300,000 Yen (Full-scale Operations)

Immediate Action: Check for vendor lock-in.

At this budget level, companies tend to become dependent on specific platforms. Relying solely on OpenAI or Anthropic is not dangerous, but it is wasteful.

If AMD’s entry leads to increased price competition among cloud providers, there will be differences in AI inference costs among AWS, Azure, and GCP. By designing for multi-cloud compatibility now, you can have the option to “switch to the cheapest provider” in six months.

Specifically, you can wrap the API call portions in an abstraction layer. Using tools like LiteLLM or OpenRouter, you can create a structure that allows for switching models and clouds. Development costs will range from a few hours to a few days. The return on investment is high.

—

Answering the Question of “Should I Wait?”

This is a common question: “Should I wait until prices drop a bit more before starting?”

The answer is no.

The reason is simple. The cost of AI will continue to decrease. The longer you wait, the cheaper it will become. However, while you are waiting for costs to drop, your competitors are gaining experience at the current prices.

The real cost of utilizing AI is not the API fees. It’s the time and effort spent figuring out how to integrate it into your business. This will not be shortened even if models become cheaper.

You only need 50,000 yen. Start using it now and finish sorting out which tasks are “worth delegating to AI” and which tasks are “better done by humans.” Companies that complete this sorting will be able to scale immediately when model prices drop by half. Those that do not will continue to wonder what to use AI for even when prices drop, perpetually saying, “Let’s wait a bit longer.”

—

Observing the Structure Behind the Price Discussion

Finally, I want to take a step back and observe the structure.

The essence of what is happening now is that the “cost of using AI” is rapidly approaching zero. Model performance is improving, and prices are either stable or decreasing. Competition in hardware is advancing, and infrastructure costs are also falling.

What happens as costs approach zero? “Using AI” will no longer be a competitive advantage. Everyone will have access to it.

When that happens, the differentiator will not be “what to do with AI” but rather “how to redesign your business and services based on the premise of AI.”

Large corporations, due to their size, take longer to redesign. Bureaucracy, adjustments, and alignment with existing systems slow them down.

SMEs are different. If the CEO says, “We’ll do this starting next week,” it changes next week. This speed of decision-making is the greatest weapon for SMEs in the era of collapsing AI costs.

Ultimately, both Anthropic’s new model and AMD’s acquisition are just discussions about “the decreasing price of tools.” The question now is what to create with these cheaper tools.

The answer lies only in the field.

POPULAR ARTICLES

Related Articles

POPULAR ARTICLES

JP JA US EN