The AI Pipeline Becomes ‘Dumber the More You Connect’ — The Shock of a 40-Point Accuracy Drop and the One Design Principle Small Businesses Should Follow

Conclusion First: LLMs 'Lose Accuracy the More They Are Connected' "To use AI wisely, connect multiple models to create

By Kai

|

Related Articles

Conclusion First: LLMs ‘Lose Accuracy the More They Are Connected’

“To use AI wisely, connect multiple models to create a pipeline” — this common wisdom may be fundamentally flawed.

Recent research has shown that in a four-step pipeline using LLMs (Large Language Models), each connection can result in a maximum accuracy drop of 40.5 points. Researchers refer to this phenomenon as the “Decomposition Tax.”

How significant is a 40-point drop? A system that scored 60 out of 100 on a test could drop to just 20 points simply by forming a pipeline. It becomes unusable.

Moreover, this is under fixed conditions where the model, prompts, and token budget are all constant. The very design decision to “create a pipeline” is what undermines accuracy.

For small businesses, this is not just someone else’s problem; it is rather the most relevant issue.

Why Does Accuracy Disappear When ‘Connecting’?

The mechanism is simple.

In an LLM pipeline, the output of step 1 becomes the input for step 2. The output of step 2 becomes the input for step 3. It’s like a game of telephone.

The problem lies in how much each step can see the original problem statement. As you progress through the steps, the original context diminishes. Only the “summary” or “intermediate results” from the previous step are passed along, leading to a loss of the original inquiry’s intent and nuance.

In validation using the gemma-3-12B model, statistically significant accuracy drops were confirmed in 70 out of 118 test cases for GSM-Hard (applied math problems) and MATH-500 (math benchmarks). Since these numbers passed the Benjamini-Hochberg correction (a statistical adjustment for multiple comparisons), they are not coincidental.

This is the essence of the “Decomposition Tax.” The more you break down tasks, the more each step operates on “local optimization,” leading to a collapse of overall accuracy.

Can ‘Re-Grounding’ Save Us? — The Answer is ‘Partially Yes’

The research proposes a method called “re-grounding” to mitigate accuracy loss. This simple approach involves passing the original problem statement again at each step.

This can somewhat alleviate accuracy loss. However, it cannot fully prevent it. This is because, at the point of breaking down the steps, there is already bias in how the “task decomposition” is structured, leading to a misalignment with the original problem’s structure.

In other words, the act of creating a pipeline itself incurs costs, and those costs cannot be reduced to zero.

Here, I want to ask small business owners:

“Is that pipeline really necessary?”

The ‘Over-Engineering’ Trap Small Businesses Fall Into

When looking at AI case studies from large corporations, you often see the following structure:

  1. An LLM that classifies user inputs
  2. A RAG system that searches based on classification results
  3. An LLM that summarizes the search results
  4. An LLM that generates answers based on the summary
  5. An LLM that checks the quality of the answers

A five-step pipeline. Each stage incurs a decomposition tax. Moreover, there are API costs at each stage, along with development and maintenance efforts.

Large corporations have the engineers and budgets to support this. Small businesses do not.

And what this research indicates is that that five-step pipeline could potentially lose accuracy to a simple one-step prompt.

Investing costs and labor only to see accuracy drop — this is the true danger of the “Decomposition Tax.”

Smaller Models Outperforming Larger Models — The Shock of ‘LLM Shepherding’

Another important study for small businesses is the framework of “LLM Shepherding.”

The mechanism works like this:

  • Use a large LLM (like GPT-4) to provide only “hints.”
  • Pass those hints to a smaller model (like GPT-3.5) to generate the actual answers.

In other words, the large model acts as a “shepherd” providing direction, while the smaller model is the “sheep” that actually runs.

What are the results?

  • Cost Reduction: 42-94%
  • Accuracy: Maintained or Improved
  • Compared to traditional routing or cascading methods, 2.8 times more cost-efficient

These figures come from benchmarks in mathematical reasoning and code generation.

A process that previously cost 100,000 yen in API fees could now be reduced to between 6,000 and 58,000 yen. For small businesses, this difference can be the deciding factor between proceeding or not.

‘Keep It Simple’ is Not Just a Philosophy; It’s a Structural Advantage

Let’s summarize the discussion so far.

Lessons from the Decomposition Tax:

  • LLM pipelines “lose accuracy the more they are connected.”
  • Maximum accuracy drop of 40.5 points at each connection.
  • Re-grounding can mitigate this, but cannot eliminate it.

Lessons from Shepherding:

  • Borrow only the “wisdom” of large models and leave the execution to smaller models.
  • Cost reduction of 42-94%, with maintained accuracy.

These two points lead to the same conclusion.

“Don’t complicate things. Be smart and keep it simple.”

This is not just a philosophy; it is a statistically validated structural fact.

And this structure presents a chance for small businesses to turn the tables.

Large companies tend to create complex pipelines driven by organizational logic. Different departments have different requirements, approval flows are complicated, resulting in multi-layered AI systems that continue to pay the decomposition tax.

Small businesses are different. Decision-making is swift. If the CEO says, “Can’t we do this in one step?” they can test it the next day.

Three Design Principles You Can Use Starting Tomorrow

So, how should small businesses proceed?

1. First, Test with ‘One Prompt’

Before creating a pipeline, see how far you can go with a single prompt. Surprisingly, this is often sufficient. Zero decomposition tax. Minimal development costs.

2. If You Must Separate, Always Pass the ‘Original Problem Statement’

If you absolutely need to break down the steps, always attach the original input to each step. This is re-grounding. Just doing this can significantly alleviate accuracy loss. The additional cost is almost zero.

3. Use Large Models as ‘Hint Providers’

There’s no need to use high-performance models like GPT-4 or Claude Sonnet for every process. Let them provide direction or hints, while the actual processing is handled by smaller models like GPT-4-mini or Haiku. Costs can drop to less than half while maintaining accuracy.

Is This the Best Approach?

Currently, many AI implementation support companies are suggesting, “Let’s create a pipeline” or “Let’s connect agents.” With each suggestion, the decomposition tax may be accumulating.

A five-step pipeline built for 3 million yen could lose to a one-prompt design costing just 50,000 yen in terms of accuracy. Such a world is already upon us.

The difference between small businesses that achieve results from AI implementation and those that only incur costs without results is not about “how advanced the AI used was” but rather about “how simply it was designed.”

Complexity is no longer just a cost; it is a liability. First, test with one prompt. If that’s insufficient, add minimal steps while being mindful of the decomposition tax. This order should not be mistaken.

POPULAR ARTICLES

Related Articles

POPULAR ARTICLES

JP JA US EN