The Monthly Cost of LLM Agents: It’s Not Just “12 Dollars”—Insights from Three Papers on “Inefficient Uses” and “Efficient Uses”

Conclusion First: The Monthly Cost of 12 Dollars Is Not a "Lie" but a "Trap" How much does it really cost to incorporat

By Kai

|

Related Articles

Conclusion First: The Monthly Cost of 12 Dollars Is Not a “Lie” but a “Trap”

How much does it really cost to incorporate an LLM as an agent in business?

“It’s about 12 dollars a month”—most small and medium-sized business owners would think, “That’s cheap.” Indeed, simple calculations lead to that conclusion: a single call costs 0.02 dollars, with 20 tasks a day and 600 tasks a month totaling 12 dollars.

However, taking this number at face value can lead to painful consequences.

After reviewing three papers and aligning them with real-world experiences, the boundary between “uses that can be done for 12 dollars a month” and “uses that can turn 12 dollars into 500 dollars a month” becomes clear.

To prevent small and medium-sized enterprises from losing out with AI agents, we will specifically illustrate this boundary.

—

First, Break Down the True Costs

The paper titled “Harnessing LLMs as Agents: What Does It Cost?” is the first to quantitatively analyze the cost structure of LLM agents. According to the proposed framework, the “Language Model Agent Machine (LAM),” the costs of an agent can be broadly divided into four categories:

  1. Communication Costs: The number of API calls to the LLM and the amount of tokens sent and received.
  2. Memory Access Costs: Reading past conversation history and task context.
  3. Recalculation Costs: Retries when memory overflows or failures occur.
  4. Reliability Costs: Additional calls needed for output verification and correction.

This is crucial. Most people are only aware of the first cost. The statement “It’s cheap at 0.02 dollars per call” only considers communication costs.

In actual use as an agent, costs 2 to 4 come into play. For example, if the agent refers to past conversation history while making decisions, each prompt will include thousands of tokens of context. As of June 2025, the cost is 2.50 dollars per 1 million input tokens and 10 dollars per 1 million output tokens. If the context per call is 3,000 tokens and the output is 500 tokens, the cost per task would be about 0.0125 dollars. That alone seems cheap.

However, it is not guaranteed that the agent will only call the LLM once per task.

Tool calls, intermediate reasoning, error handling—in actual agent operations, an average of 5 to 15 LLM calls occurs per task. The benchmarks in the papers also show that for complex tasks, over 20 calls have been observed.

Let’s recalculate.

Condition Simple Calculation (1 call/task) Reality (Average 8 calls/task)
Daily Task Count 20 20
Monthly Task Count 600 600
LLM Calls/Month 600 4,800
Cost per Call 0.02 dollars 0.02 dollars
Monthly Total 12 dollars 96 dollars

Eight times. This is the “reality with context.”

Furthermore, if retries occur due to memory overflow, the number of calls can increase by another 1.5 to 2 times. The monthly cost could rise to between 144 and 192 dollars.

The difference between 12 dollars and 192 dollars a month can completely change decision-making.

—

Is Automatically Routing to Cheaper Models Really Beneficial?

Many people might think, “Then we can just route simple tasks to cheaper models, right?” This is what dynamic LLM routers do.

The paper “Dynamic LLM Routers are Often Misguided” pours cold water on this idea.

After analyzing six commercial routers, the research team found that using dynamic routers often resulted in performance dropping by over 10 points. In other words, attempting to reduce costs led to a decrease in quality, which in turn increased retries and rework, raising the total cost.

Why does this happen? Dynamic routers make judgments like, “This query is simple, so a cheaper model is sufficient,” but their criteria for judgment are often coarse. This is especially true for tasks involving business context—such as “creating a response based on past estimates”—which may seem simple but are highly context-dependent. The cheaper model may return nonsensical outputs, requiring human correction or a re-run with a more expensive model.

What was intended as cost reduction ends up being a double blow of increased costs and time loss.

The implications for small and medium-sized enterprises are clear:

  • If the types of tasks are few, routers are unnecessary. It is more stable to fix “this task requires this model” from the start.
  • If using a router, at least measure accuracy for two weeks before going live.
  • Sorting “tasks that can be handled by cheaper models” and “tasks that require more expensive models” is more accurately done by humans than by AI.

This is where small and medium-sized enterprises have an advantage. Large companies have hundreds of types of tasks, so they have to rely on routers. However, small and medium-sized enterprises can often narrow their tasks down to 5 to 10 types, which is a level of granularity manageable by humans.

—

The “Weight” of Tokens Is Not Uniform

Another factor that distorts the sense of cost is the variability in token costs.

The study “What Does a Token Cost?” revealed the fact that the computational cost of generated tokens is not uniform. The top 10% of the most costly tokens account for 64 to 80% of the total computational load (FLOPs).

What does this mean?

Even when we say “500 tokens of output,” the actual computational load can vary several times based on the content. Currently, API charges are based on the number of tokens, so this is not directly reflected in the billing amount. However, there is a possibility that token pricing will shift to a pay-per-use model based on “content complexity” in the future.

Already, inference models like OpenAI’s o1 and o3 have set different prices for reasoning tokens. If this trend continues, we could see a world where “even for the same 500 tokens, tasks with heavy reasoning are charged three times more.”

What small and medium-sized enterprises should do now is to clearly distinguish between “tasks that truly require heavy reasoning” and “tasks that can be handled with templates.” The latter will continue to see costs decrease, while the former may not only not decrease but could even increase.

—

Realistic Cost Estimation: Three Patterns

Taking all of the above into account, let’s estimate the realistic monthly costs for small and medium-sized enterprises when adopting LLM agents.

Pattern A: Automation of Routine Tasks (Email Responses, FAQ Handling, etc.)

  • Model: GPT-4o mini (Input 0.15 dollars/1M tokens, Output 0.60 dollars/1M tokens)
  • LLM Calls per Task: Average 3 calls
  • 20 tasks a day, 600 tasks a month
  • Monthly Cost: Approximately 3 to 8 dollars (about 500 to 1,200 yen)
  • Labor Cost Reduction Effect: Equivalent to 20 to 40 hours a month → Calculated at 1,500 yen/hour, this amounts to 30,000 to 60,000 yen.
  • ROI: 30 to 60 times the investment

Pattern B: Support for Decision-Making Tasks (Creating Estimates, Summarizing Reports, etc.)

  • Model: GPT-4o (Input 2.50 dollars/1M tokens, Output 10 dollars/1M tokens)
  • LLM Calls per Task: Average 8 calls
  • 20 tasks a day, 600 tasks a month
  • Monthly Cost: Approximately 80 to 200 dollars (about 12,000 to 30,000 yen)
  • Labor Cost Reduction Effect: Equivalent to 40 to 80 hours a month → Calculated at 2,000 yen/hour, this amounts to 80,000 to 160,000 yen.
  • ROI: 4 to 13 times the investment

Pattern C: Complex Reasoning Tasks (Contract Review, Strategic Analysis, etc.)

  • Model: o3 (Input 10 dollars/1M tokens, Output 40 dollars/1M tokens)
  • LLM Calls per Task: Average 15 calls, consuming a large number of reasoning tokens
  • 10 tasks a day, 300 tasks a month
  • Monthly Cost: Approximately 300 to 800 dollars (about 45,000 to 120,000 yen)
  • Reduction in Outsourcing Costs for Experts: Equivalent to 1 to 3 cases a month → Calculated at 50,000 yen per case, this amounts to 50,000 to 150,000 yen.
  • ROI: 1 to 3 times the investment (potentially resulting in a loss)

Can you see the picture?

Pattern A is at a level where “there’s no reason not to do it.” Pattern B is “if designed properly, it will pay off sufficiently.” Pattern C is “if not narrowed down to truly necessary tasks, it will result in losses.”

—

The Nature of “Inefficient Uses”

The “inefficient uses” that became clear when comparing the three papers with real-world experiences can be summarized as follows:

  1. Running all tasks on a high-cost model—Using o3 for routine tasks is like calling a taxi to go to the convenience store.
  2. Outsourcing entirely to dynamic routers—For small and medium-sized enterprises with few tasks, human sorting is more accurate and cost-effective.
  3. Not understanding the number of agent calls—Only looking at “0.02 dollars per call” and being reassured, while not realizing that it is actually called 8 times.
  4. Not managing context length—Stuffing all conversation history and feeding tens of thousands of tokens each time.

Conversely, the “efficient uses” are straightforward:

  • Fix the model for each task. A granularity of about 5 types of tasks and 3 types of models is sufficient.
  • Keep context to a minimum. Create a system that only provides necessary information.
  • Set limits on the number of calls. Prevent accidents where the agent spirals out of control and loops 30 times.
  • Measure costs monthly. Develop a habit of checking the API usage dashboard weekly.

—

So, What Should We Do?

If small and medium-sized enterprises are to adopt LLM agents, they should start with Pattern A. For less than 1,000 yen a month, they can save 30,000 to 60,000 yen in labor costs. After creating this success experience, they can move on to Pattern B.

Pattern C should be avoided unless there is a very clear ROI. Alternatively, it may be wise to use it only as a spot solution once or twice a month.

Both “AI agents are cheap” and “AI agents are expensive” are incorrect. The accurate statement is: “It depends on how you use them; you can gain 30 times or end up in the red.”

The dividing line is not a technical issue. It is a matter of understanding your company’s tasks and appropriately designing models and costs—an extremely managerial decision.

Small and medium-sized enterprises can perform this design with human insight precisely because they have fewer types of tasks. They can optimize manually where large companies have to rely on automation. This is one of the few points where small and medium-sized enterprises can compete with large companies using AI agents.

Start by choosing one routine task and try running it from next week. It’s an experiment that can begin for just 1,000 yen a month.

POPULAR ARTICLES

Related Articles

POPULAR ARTICLES

JP JA US EN