The Price of AI is Collapsing Month by Month—A 45% Price Cut for Claude 4, Gemini Flash, and the Implications of 900 Models on 1 API for SMEs

Conclusion Let’s get straight to the point. The "fair price" for AI no longer exists. The price you see on Monday becom

By Kai

|

Related Articles

Conclusion

Let’s get straight to the point. The “fair price” for AI no longer exists.

The price you see on Monday becomes outdated by Friday.

Let’s list what has happened in just the last week or two:

  • Anthropic’s “Claude Sonnet 4”: Up to a 45% cost reduction at peak usage. Even under normal use, about a 25% drop.
  • Google’s “Gemini 2.5 Flash”: High-performance models introduced at Flash price levels. A structure that allows traditional Pro-level performance to be used at an unprecedentedly low cost.
  • Aggregators like Openrouter: Over 900 models can be switched via a single API. Prices can be tracked in minutes.

The decision made six months ago that “this model is optimal” has already expired.

The question for small and medium-sized enterprises (SMEs) is simple: How do you distinguish between the model to contract this month and the one to wait for until next month?

What the Price Cut for Claude Sonnet 4 Means

What Anthropic has introduced with its new model is not just a performance upgrade. It is a redesign of the pricing structure itself in response to the voices of SMEs saying, ‘It’s too expensive to use.’

Here’s how it works: By reusing previously processed prompts and contexts as cache, businesses that repeatedly perform similar tasks—such as handling standard inquiries, summarizing daily reports, and checking contracts—can see token costs drop by up to 45%.

What does this translate to in monthly costs?

For a company using 1 million tokens a month, the cost with the previous Claude 3.5 Sonnet would be about 450 yen (based on input tokens). With Claude Sonnet 4 utilizing cache, it would be about 250 yen. While a 200 yen difference may seem small, at a scale of 10 million tokens, the difference becomes 2,000 yen a month, and for 100 million tokens, it’s 20,000 yen a month.

Moreover, performance improvements are noteworthy. In the coding benchmark “SWE-bench Verified,” Claude Sonnet 4 recorded 72.7%, showing significant improvement over the previous model. Scores in scientific reasoning benchmarks have also more than doubled.

It has become cheaper and smarter. This is the essence of what is happening in the current AI market.

The Game-Changing “900 Models 1 API”

Another seismic shift is in AI model aggregation services. Services like Openrouter allow access to over 900 models from major providers such as OpenAI, Anthropic, Google, Meta, and Mistral through a single API endpoint.

What does this mean for SMEs?

“Vendor lock-in” disappears.

Previously, switching from an OpenAI API-based system to Anthropic required rewriting code. However, through an aggregator, you can switch with just a single line change in the model name.

Price comparisons are also automated. The latest prices for each model are fetched every few minutes, allowing you to choose the “most cost-effective model at this very moment” for the same task.

What I’ve experienced while supporting AI utilization in local SMEs is that many companies continue to use the model they initially selected. It’s not uncommon for an API cost that was 30,000 yen six months ago to drop to 10,000 yen just by changing the model.

Aggregators bring that “switching cost” down to nearly zero.

When Adjusted for Quality, AI Prices Have Dropped 73% Annually

Let’s take a step back and discuss the structural aspects.

In a paper titled “The Price of Intelligence” published by researchers from Stanford and others, the price fluctuations of AI inference are analyzed using a quality-adjusted price index.

A simple price observation shows an annual decline rate of about 10%. It gives a sense that “well, prices go down a little every year.”

However, when calculated based on the cost to achieve the same level of performance, the annual decline is 73%.

What this difference means is that “it’s not that cheaper models have emerged, but rather that a job that cost 100,000 yen last year can now be done for 27,000 yen this year.” Next year, it’s projected to drop another 27%, to about 7,300 yen.

This speed exceeds Moore’s Law. While semiconductor integration doubles every 18 months, the quality-adjusted cost of AI has dropped to about a quarter in just 12 months.

So, What Should SMEs Do?

Let’s move beyond abstract discussions and present three specific criteria for decision-making.

1. Conditions for Models to Contract “Right Now”

  • Applications where the monthly API cost stays below 50,000 yen: In this price range, even if prices drop further next month, the pain is minimal. It’s better to reap the benefits of business improvement than to wait.
  • Standardized tasks where caching is effective: Models like Claude Sonnet 4, which offer significant discounts for repetitive processing, should be implemented and utilized immediately.
  • Tasks where effects are already visible: For tasks like handling inquiries, creating minutes, and organizing data—where “AI clearly reduces time”—there’s no reason to wait.

2. Conditions for Models to “Wait Until Next Month”

  • Large-scale applications where the monthly API cost exceeds several hundred thousand yen: At this scale, price fluctuations can lead to differences in the tens of thousands of yen within a month. A system for quarterly reviews should be established.
  • Advanced analyses based on the latest inference models (o3, Gemini 2.5 Pro, etc.): This class experiences the most intense price competition, with situations changing in a matter of weeks.
  • Areas where not even a PoC has started: If you haven’t even decided what to use it for, it’s better to test with free tiers rather than waiting for prices to drop.

3. Mechanisms to Implement

  • Build through aggregator APIs: Design with model switching in mind from the start. Avoid dependency on specific vendors.
  • Monthly “model inventory”: Check the prices and performance of the models being used on a monthly basis. If new models are available, consider switching. This task can be completed in 15 minutes.
  • Monitoring token usage: “I used more than I thought” is a common scenario for SMEs. Without measurement, optimization is impossible.

The “Reversal Structure” for SMEs

Finally, I want to emphasize the most important point.

The collapse of AI prices works to the advantage of SMEs over large enterprises.

Large companies are slow to act due to existing vendor contracts, internal approval processes, and security reviews. Some companies take three months to switch models.

In contrast, a company with ten employees can decide in a morning meeting to “switch to this model starting today” and make the change by noon.

In a market where prices collapse month by month, the speed of decision-making itself becomes a competitive advantage.

This is one of the few structural advantages that SMEs have over large enterprises.

An AI system that cost 3 million yen to build can now be assembled for 50,000 yen. It may become even cheaper next month. But while you wait until next month, the company next door is already moving.

“Start small and review monthly.”

This is the simplest and strongest strategy in an era where AI prices continue to collapse.

POPULAR ARTICLES

Related Articles

POPULAR ARTICLES

JP JA US EN