NVIDIA’s Dominance Begins to Erode: How Should Local SMEs Respond in an Era of Halved GPU Costs?
Related Articles
NVIDIAの城壁に、3方向から穴が開いた
In a nutshell, what is happening in the AI chip market can be summarized as follows: “An industry overly reliant on NVIDIA is collectively seeking exits.”
AMD has acquired the AI chip startup “Taalas.” Anthropic has established its own chip design team. Google is reorganizing its internal AI organization to accelerate the integration of hardware and software.
At first glance, these three pieces of news may seem unrelated, but they share a common structure: “GPU costs are too high. It’s cheaper to make our own.” Everyone has reached the same conclusion.
So, what does this have to do with local SMEs? Quite a lot.
なぜ「チップ競争」が中小企業の財布に直結するのか
First, it’s essential to understand the current cost structure of AI utilization.
There are two primary ways for SMEs to use AI: through cloud APIs like ChatGPT or Claude, or by running local LLMs on their own servers. Regardless of the choice, GPUs are operating in the background, and the majority of these GPUs are from NVIDIA.
NVIDIA’s data center GPU, the “H100,” costs around 4 million yen per unit. The latest “H200” is even more expensive. AI companies purchase these in the thousands or tens of thousands. OpenAI’s annual infrastructure costs are said to be in the billions of dollars.
Naturally, these costs are passed on to API usage fees. In other words, as long as NVIDIA’s GPUs remain expensive, the AI usage fees for SMEs will also stay high. A significant portion of a monthly API fee of 50,000 yen can be attributed to GPU costs.
This is why competition in the GPU market is directly good news for SMEs.
AMDの「Taalas買収」——何が変わるのか
Taalas possesses technology that optimizes AI models at the silicon level. In short, it’s the idea that “running on dedicated chips is faster and cheaper than using general-purpose GPUs.”
The significance of AMD acquiring this technology is substantial. While AMD’s GPUs have been cheaper than NVIDIA’s, they have lagged in the software ecosystem (due to the CUDA barrier). However, if they have the technology to directly embed models into hardware, a new pathway that does not rely on CUDA emerges.
In fact, Meta (formerly Facebook) has already adopted AMD’s MI300X in large quantities. Its price is said to be around 60-70% of the H100, and in terms of performance, it can match the H100 for specific workloads. If AMD integrates Taalas’s technology, this price gap could widen further.
Anthropicの自社チップ——「お客さん」が「競合」になる瞬間
The news of Anthropic establishing a custom silicon team indicates a more profound structural change.
Anthropic uses Amazon (AWS) and Google’s cloud to run Claude. In other words, they are continuously paying “rent” to NVIDIA and Google for their TPUs. Creating their own chips means they want to bring this rent close to zero.
Think back to when Apple switched from Intel chips to the M1. Performance improved, power consumption decreased, and costs dropped. The same is about to happen in the AI world.
If Anthropic lowers the costs of the Claude API with its own chips, OpenAI and Google will have no choice but to follow suit. A price-cutting competition for APIs will ensue.
In fact, signs of this are already apparent. When GPT-4 was released in March 2023, the API usage fee was about $0.03 per 1,000 tokens. By 2025, GPT-4o is expected to offer similar or better performance for about $0.0025. In just two years, this is less than one-tenth of the original cost. As the chip competition intensifies, this rate of decline will accelerate further.
Googleの組織再編——巨人の焦り
At Google, DeepMind CEO Demis Hassabis has stepped away from daily operations to become Alphabet’s chief scientist. While it may seem like he has been promoted, it can also be interpreted as relinquishing control of on-the-ground operations.
It is clear that Google is feeling the pressure. They have developed their own TPUs (Tensor Processing Units) and offer them to cloud customers, yet they are being overshadowed by the demand for NVIDIA GPUs. The internal AI teams are dispersed, leading to slow decision-making. This reorganization likely aims to “separate research from product development to speed up the product side.”
If Google’s TPUs begin to genuinely compete with NVIDIA GPUs in terms of performance and cost, the price competition for cloud AI will intensify. Services using TPUs on Google Cloud are already offered at prices 30-50% lower than those of NVIDIA GPUs.
「GPU代が半額になる」は、いつ起きるのか
To be honest, it won’t be “half price starting next month.” However, over a span of 2-3 years, the likelihood that AI utilization costs will drop to less than half of their current levels is extremely high.
The reasoning is as follows:
- AMD, Google, and Amazon are seriously starting to mass-produce chips to compete with NVIDIA (increased supply → lower prices)
- AI companies like Anthropic and OpenAI are investing in their own chips, increasing pressure to lower API prices
- Model efficiency is improving. The computational requirements to achieve the same performance are decreasing year by year. The performance level of GPT-4 is becoming achievable at one-tenth of the computational cost from two years ago.
All three of these factors are progressing simultaneously. The decline in costs is not something that “will come someday” but rather “has already begun.”
で、中小企業はどう動くべきか
Now, let’s get to the main topic.
1. 今すぐ高額なGPUを買う必要はない
If any SMEs are considering purchasing a 1 million yen GPU server to run local LLMs, it might be wise to wait a bit. There’s a possibility that the same performance could be available for half the price in six months. For now, cloud APIs are sufficient. For a few thousand to tens of thousands of yen per month, you can access GPT-4o or Claude 3.5.
2. API価格の下落を前提に「使い方」を先に設計する
Simply waiting for costs to drop is not meaningful. It’s crucial to decide now what to use the technology for once it becomes cheaper.
For instance, if you are paying 50,000 yen per month for an API to automate customer support, and it drops to 20,000 yen, you could expand your budget to include automated generation of sales materials or inventory forecasting. The mindset should be, “When costs are halved, I will double my initiatives.”
3. 「GPU代」ではなく「人件費との比較」で投資判断する
The most critical metric for AI investment in SMEs is not the price of GPUs. It’s about how much it would cost to have a human perform that task each month.
For example, if an administrative staff member earning 250,000 yen spends two hours a day processing invoices, and automating this with AI costs 30,000 yen per month, that leaves a savings of 220,000 yen each month. If GPU costs drop to 15,000 yen per month, it becomes even more advantageous.
The key is to start with tasks that can pay off even at current prices. Consider cost reductions as a “bonus.”
4. ベンダーロックインを避ける
The erosion of NVIDIA’s dominance means more choices. If you design your systems to depend on specific cloud services or GPUs, you may find it difficult to switch later. It’s essential to create “escape routes” from the outset by incorporating an abstraction layer for APIs and designing systems that can switch between multiple models—this will become a survival strategy for SMEs.
本当の変化は「コストが下がった先」にある
The mere fact that GPU costs will be halved is not particularly newsworthy. What is truly interesting is what will happen next.
As AI utilization costs dramatically decrease, capabilities that were once exclusive to large corporations will become accessible to SMEs. 24/7 customer support, multilingual capabilities, inventory management based on demand forecasting—systems that previously cost tens of millions of yen annually will become available for just a few tens of thousands of yen per month.
When that happens, the winners will not be the companies that can use AI the best. They will be the companies that were prepared to act immediately when AI became affordable.
Large corporations take time to make decisions. They go through approvals, select vendors, and conduct proof of concepts (PoCs)—a process that can take six months to a year.
SMEs can start next week. If the president says, “Let’s do it,” they can act the next day. This is the greatest weapon of SMEs.
There’s no need to wait for GPU costs to be halved. Start thinking today about “what to use it for.” The costs will follow.
JA
EN