ORIGINAL THOUGHT PAPER · JUNE 2026 · V4

The Structural Crisis of
Exponentially Rising AI Costs

A Dual-End Squeeze Model Analysis of the Paradigm Predicament
— An Industry Wake-Up Call

Cost Early Warning for AI Practitioners Under the Dual-End Squeeze Model

Published June 1, 2026

Category Original Thought Paper · Industry Alert

Domains AI Economics · Token Pricing · Inference Architecture · Scaling Laws

Version V4

Authors LEECHO Global AI Research Lab & Opus 4.6 & GPT 5.5 & Gemini 3.1 (Cognitive Collective)

ABSTRACT — To Every Decision-Maker Managing an AI Budget

In April 2026, Uber burned through its entire annual AI budget in just four months. In May, Microsoft began revoking Claude Code licenses from thousands of internal engineers. That same month, Uber’s COO publicly stated there was “no direct correlation between token consumption and useful consumer features.” JPMorgan published a research note titled “AI Token Costs are Eating Internet Profits Alive.” Token unit prices had fallen approximately 280–600× over three years — yet total enterprise AI spending was still growing at 200–320% annually. This paper reveals the structural root of this paradox and proposes a “Dual-End Squeeze Model”: End A (per-interaction cost) continues to inflate due to the brute-force search nature of Thinking Tokens and physical-wall constraints, while End B (repair-loop iterations) remains ineliminable due to the inherent limitations of probabilistic sampling. The two ends multiply rather than add, driving super-linear cost explosions. Deeper analysis reveals that LLM pre-training data harbors a systematic “outcome bias” — recording successes while omitting trial-and-error processes — creating an unrepayable “cognitive debt” that forces every inference to rebuild missing process knowledge in real time using compute. This paper tracks the complete timeline of the accelerating crisis from January to May 2026 and argues that, under the current Transformer + autoregressive + probabilistic sampling architecture, the structural rise in AI deployment costs is not a temporary growing pain but a paradigm-level systemic constraint. Every AI practitioner needs to understand this mechanism — not tomorrow, but now.

Keywords AI Inference Cost Crisis · Token Economics · Jevons Paradox · Cognitive Debt · Dual-End Cost Model · Test-time Compute · Industry Alert

CHAPTER 01Your Budget Is Exploding: Here’s Why

1.1 What Happened in Five Months

In January 2026, AWS quietly raised GPU Capacity Block prices by 15% on a Saturday, with no announcement. In February, “Inference Economics” was formally defined as an independent discipline. In March, Gartner warned “do not conflate the deflation of commodity tokens with the democratization of frontier reasoning,” and Sam Altman publicly stated that “intelligence will be metered like electricity and water.” In April, Uber’s CTO admitted the company had burned through its entire annual AI budget in four months — “I had to redo my budget because what I thought I needed was completely blown up.” In May, Microsoft began revoking Claude Code licenses from thousands of engineers, Uber’s COO publicly questioned the correlation between token consumption and useful features, and JPMorgan published a research note titled “AI Token Costs are Eating Internet Profits Alive.”

All of this happened against the backdrop of continuously plummeting token unit prices. The inference cost for GPT-4-level performance fell from $30 per million tokens in 2023 to under $3 in 2026 — a reduction of more than 10×. DeepSeek V4’s cached price was as low as $0.0036 per million tokens. If you only look at unit prices, AI appears to be entering a “free” era.

But the bills are exploding.

1.2 The Cost Paradox

Average enterprise AI spending rose from $2.5 million in 2024 to $7 million in 2025 — nearly tripling. Inference now accounts for 85% of enterprise AI budgets, compared to roughly one-third in 2023. Forty percent of enterprises now spend over $10 million annually on AI — reaching in just three years a spending level that took cloud computing thirteen years to achieve. Gartner forecasts global AI spending will exceed $2.5 trillion in 2026.

1.3 Purpose of This Paper

This paper is not an academic literature review. It is a cost early warning for AI practitioners — explaining why plummeting unit prices haven’t reduced your bills, why this is not a problem that FinOps tools can solve, and why the problem is structural under the current architecture. If you are managing an AI budget, designing Agent systems, or deciding whether to scale AI deployment, the mechanisms revealed here will directly impact your decisions.

· · ·

CHAPTER 02January–May 2026: Accelerating Crisis Chronicle & Danger Index

Below is a month-by-month tracker of the AI cost crisis during the first five months of 2026. The Danger Index is rated across five dimensions (vendor signals, enterprise impact, authoritative confirmation, structural deterioration, contagion velocity), each scored 0–2, for a maximum of 10.

2.1 January: Undercurrents [Danger Index 3.0/10]

AWS quietly raised prices by 15%. CloudZero data showed enterprise AI/ML spending as a share of total cloud expenditure had doubled from 1.55% to 2.67% within a year. OpenAI’s inference costs reached $8.4 billion in 2025, with a gross margin of just 33%, and projected 2026 losses of $14 billion. TSMC confirmed advanced-node price hikes of 3–10% for 2026, with 2nm wafers expected to exceed $30,000 per wafer. Signals were present but had not yet triggered industry-wide anxiety.

2.2 February: Warning Escalation [Danger Index 4.5/10]

Reworked published the feature “Inside the AI Cost Crisis,” defining inference cost as the single largest driver of enterprise budgets. CTO operational handbooks proliferated, warning that DRAM prices could double in 2026. “Inference Economics” was formally defined by CloudZero as an independent discipline — “not a metric, but a practice.” The issue escalated from individual enterprise anxiety to an industry-wide structural concern.

2.3 March: Data Confirmation [Danger Index 6.0/10]

Authoritative data arrived in a concentrated burst. Gartner published its “Navigating the Commoditization Trap” report, explicitly warning that agentic models require 5–30× more tokens per task than standard GenAI, that token consumption growth would outpace cost reduction, and that overall inference costs were expected to rise. Sam Altman declared in an interview that “intelligence is a utility, metered by the glass.” Anthropic shifted pricing from flat fees to usage-based billing. Du (2026) from Wuhan University published the “Tiered Super-Moore” paper demonstrating that frontier reasoning-layer model prices had nearly stagnated (R² = 0.031). The reasoning premium averaged 31.5× non-reasoning prices. The “price reduction paradox” was quantitatively confirmed for the first time.

2.4 April: Landmark Collapse [Danger Index 8.0/10]

Uber CTO Praveen Neppalli Naga revealed to The Information that the company had burned through its entire 2026 annual AI coding budget in just four months. Claude Code adoption surged from 32% to 84% (covering approximately 5,000 engineers), with per-engineer monthly API costs of $500–$2,000. Roughly 70% of committed code was AI-generated, and one in ten backend updates was deployed by agents without human intervention. “I had to redo my budget because what I thought I needed was completely blown up.” That same month, Microsoft’s FY25 Q3 earnings (disclosed April 30) confirmed processing over 100 trillion tokens in the quarter, a 5× year-over-year increase. Gartner found only 28% of AI infrastructure projects fully met business objectives.

2.5 May: Systemic Retreat [Danger Index 9.5/10]

Dominoes fell in rapid succession. On May 14, Microsoft’s Experiences + Devices division (Windows/365/Outlook/Teams) began revoking most Claude Code licenses, with a cutoff date of June 30 — just six months after initial rollout. Uber COO Andrew Macdonald publicly questioned: “There doesn’t seem to be a direct correlation between token usage and useful consumer features.” Gartner’s May press release forecast that 25% of planned 2026 AI budgets would be deferred to 2027. Goldman Sachs projected that agentic AI would drive a 24× increase in token consumption by 2030. JPMorgan published “AI Token Costs are Eating Internet Profits Alive.” Fortune ran two cover stories in a single week. GitHub Copilot switched from flat monthly fees to usage-based pricing. Meta internally launched a leaderboard called “Claudeonomics” to track token consumption.

Danger Index Progression, January–May 2026
Month Vendor Signals Enterprise Impact Authoritative Confirmation Structural Deterioration Contagion Velocity Total Phase
Jan 1.0 0.5 0.5 1.0 0 3.0 Undercurrents
Feb 0.5 1.0 1.0 1.0 1.0 4.5 Warning Escalation
Mar 1.5 1.0 1.5 1.0 1.0 6.0 Data Confirmation
Apr 1.0 2.0 1.5 1.5 2.0 8.0 Landmark Collapse
May 1.5 2.0 2.0 2.0 2.0 9.5 Systemic Retreat

In five months, the Danger Index rose from 3.0 to 9.5 — a 3.2× increase, with accelerating momentum. This is not the curve of a technology going through an awkward adolescence. This is the curve of a market repricing itself.

Note: A 10/10 is defined as “multiple major enterprises collectively withdrawing agentic tools, mainstream model services raising effective prices, or large-scale AI project cancellations reaching earnings-disclosure level.” As of the end of May, the crisis was just one step away from this threshold.

· · ·

CHAPTER 03Technological Evolution Timeline: The Token Consumption Inflation from CoT to Agents

3.1 Foundation Period (2022)

Chain-of-Thought (Wei et al., 2022) established the paradigm of step-by-step LLM reasoning, with typical consumption of a few hundred tokens. In October of the same year, Yao et al. published ReAct, which for the first time alternated reasoning traces (Thought) with task actions (Action) and environmental observations (Observation), establishing the Thought→Action→Observation loop. On HotpotQA, a single task consumed an average of approximately 9,795 tokens — roughly a 20× inflation over CoT.

3.2 Memory and Reflection Period (2023)

Reflexion (Shinn et al., 2023, NeurIPS) introduced cross-trial memory on top of ReAct: after each failure, the agent generated a self-reflection, distilling it into concise experience stored in memory. The structure became “original query + historical trajectory + reflective memory × N retries,” with token consumption scaling linearly with retry count. That same year, LATS (Zhou et al., 2023) introduced Monte Carlo Tree Search, escalating token consumption from linear to near-exponential growth. As a countermeasure, ReWOO (Xu et al., 2023) achieved an 80% reduction in token consumption through one-shot planning plus parallel execution, but at the cost of flexibility.

3.3 Engineering Deployment Period (2024)

The LangChain/LangGraph ecosystem matured. RP-ReAct decoupled planning from execution but exposed real-world context overflow problems — tool interactions returned massive data payloads that rapidly filled context windows, and smaller windows in open-source models exacerbated the issue. Memory architectures diverged into two layers: short-term context windows and long-term vector databases.

3.4 Trajectory Reuse and Deep Research Period (2025–2026)

AgentHER (Ding, 2026) relabeled failed trajectories as usable training data. Re-TRAC (Zhu et al., 2026) achieved 15–20% improvement over ReAct through recursive trajectory compression. AgentDiet (Xiao et al., 2026) found that trajectories were riddled with useless, redundant, and stale information — removing which reduced input tokens by 39.9–59.7%. However, in Deep Research mode, single-query costs had already reached $0.41–$1.32.

Token Consumption Inflation Overview
Architecture Year Typical Tokens/Task Multiple vs. CoT
CoT 2022 ~500
ReAct 2022 ~10,000 20×
Reflexion 2023 ~35,000 70×
LATS 2023 ~150,000 300×
Agent+RAG 2024 ~500,000 1,000×
Deep Research 2025 ~800,000 1,600×
Agentic Coding 2025–26 1–3.5M 2,000–7,000×

3.5 Summary: Staircase Inflation of Token Consumption

From CoT to Deep Research, per-task token consumption grew approximately 1,600×, and up to 7,000× for agentic coding. Each architectural generation did not simply add linearly but introduced a new multiplicative factor: ReAct introduced ×N_steps loop accumulation, Reflexion stacked ×N_retries, LATS introduced ×N_branches tree search branching, and Deep Research added ×N_searches × N_pages retrieval reading. The combination of these multiplicative factors drove staircase-pattern token consumption inflation — each new architectural capability came at the cost of a new cost dimension. Notably, ReWOO (2023) achieved an 80% token reduction, but at the expense of flexibility and fault tolerance, indicating a fundamental trade-off between efficiency and capability.

· · ·

CHAPTER 04Cost Economics: The Race Between Price Decline and Consumption Growth

4.1 The Token Unit Price Decline Curve

Data tracked by BenchLM (2026) shows that since March 2023, the average output price of frontier LLMs has fallen by approximately 94.5% (price index from 100 to 5.5). Introl (2025) estimated that inference costs are falling at approximately 10× per year, faster than the microprocessor revolution in PC computing or internet bandwidth.

4.2 The Token Consumption Growth Curve

OpenRouter (2025) data shows that average prompt token length has grown nearly since early 2024. Microsoft’s FY25 Q3 earnings (disclosed April 2025, quarter ending March) confirmed processing over 100 trillion tokens in the quarter, a 5× year-over-year increase. In September 2025, Claude 4 Sonnet’s daily average token usage on OpenRouter reached 100 billion.

4.3 Jevons Paradox: A Perfect Replication in the AI Domain

Luccioni, Strubell & Crawford (2025, ACM FAccT) specifically studied how the Jevons Paradox applies to AI. Core finding: efficiency gains paradoxically stimulated higher total consumption. The rebound effect undermined the assumption that “pure technological efficiency optimization alone ensures net reduction.”

“Per-token inference costs fell roughly 1,000× in three years — yet total inference spending grew 320% over the same period. Cheaper tokens only spawned more use cases and higher query volumes.”
— SoftwareSeni, The AI Inference Market in 2025

4.4 “The Tiered Paradox”: The Cheap Got Cheaper, the Expensive Didn’t

Du (2026, Wuhan University) proposed the “Tiered Super-Moore” hypothesis: the price half-life for economy-tier models is 1.10 years, and for mid-tier models 1.55 years — both significantly faster than Moore’s Law. However, flagship/reasoning-tier models exhibited virtually no exponential price decline (R² = 0.031), with reasoning premiums averaging 31.5× non-reasoning prices. Meanwhile, the rate of frontier performance improvement outpaced cost-adjusted improvement by 12–14% — meaning the total cost of reaching higher performance was actually rising.

· · ·

CHAPTER 05The Physical Wall: Deceleration and Reversal of Cost Decline

5.1 Rising Semiconductor Fabrication Costs

TSMC raised prices on advanced-node chips by 3–10% in 2026. Current 3nm wafer prices stand at approximately $18,000–$20,000 per wafer, and 2nm wafers are projected to exceed $30,000 per wafer. This is not a one-time adjustment — TSMC has planned a multi-year continuous price increase roadmap covering 5nm/4nm, 3nm, and sub-2nm, extending from 2026 through at least 2029.

5.2 The Material Slowdown of Moore’s Law

Former Intel CEO Pat Gelsinger (2023) acknowledged: “We are no longer in the golden age of Moore’s Law… it’s now closer to doubling every three years.” In 2022, NVIDIA CEO Jensen Huang directly declared Moore’s Law dead. The semiconductor industry roadmap has shifted from density-driven to application-driven (“More than Moore”).

5.3 Hard Bottlenecks in Energy and Memory

The IEA’s 2025 report projects that global data center electricity consumption will double by 2030, with AI-optimized server electricity rising from 93 TWh to 432 TWh. Power permitting for new facilities in major markets has extended to 24–36 months. On the memory front, NVIDIA’s B200 GPU requires 192GB of HBM3E — nearly 2.5× the H100’s 80GB — with each generation demanding significantly more high-bandwidth memory amid persistent supply shortages.

5.4 Memory Bandwidth and Supply Chain Constraints

TrendForce (2026) noted that bandwidth limitations — both intra-chip and inter-chip — are becoming the next major performance bottleneck for AI systems in 2026. The NVIDIA B200 GPU requires 192GB of HBM3E memory, nearly 2.5× that of the H100 (80GB). Each GPU generation demands significantly more high-bandwidth memory, yet SK Hynix and Samsung’s HBM production capacity cannot match the rate of demand growth. TSMC’s CoWoS (Chip-on-Wafer-on-Substrate) advanced packaging capacity is also a severe bottleneck — this technology is critical for AI accelerators, but capacity was already locked up by NVIDIA and AMD in 2025. Consumer-grade CPUs have consequently been deprioritized, and AMD raised prices across the entire Ryzen lineup in late 2025.

5.5 Hidden Cloud Provider Price Hikes and the Subsidy Bubble

AWS quietly raised Capacity Block prices by 15% in January 2026, with no announcement. H100 spot rental prices reached $2.39/hour in May 2026, the highest in months. More critically, OpenAI, Google, Anthropic, and Meta are all currently pricing inference services below cost to compete for market share — creating an artificial price floor. OpenAI’s 2025 inference costs reached $8.4 billion at a gross margin of just 33% (Sacra, 2026), with 2026 inference costs projected to rise further to $14.1 billion. When capital discipline returns, prices will normalize upward, and the cost shock experienced by users will be even more severe.

5.6 Summary: Divergence of Three Price Decline Curves

Du (2026)’s “Tiered Super-Moore” hypothesis reveals a critical fact: token price reductions are not uniform. Three curves show entirely different trajectories:

Price Decline Velocity Divergence Across Three Model Tiers
Model Tier Price Half-Life R² Fit Trend Assessment
Economy Tier (GPT-4o mini, etc.) 1.10 years High Rapid price decline, approaching zero cost
Mid Tier (GPT-4o / Sonnet, etc.) 1.55 years Medium Moderate decline, room remains
Frontier Reasoning Tier (o3 / Opus, etc.) Near-zero decline 0.031 Reasoning premium 31.5×, near-stagnant

This means: users are indeed enjoying price reduction dividends on simple tasks, but the frontier reasoning capabilities that agents actually need — that’s precisely the tier that hasn’t gotten cheaper. “The cheap got cheaper, the expensive stayed expensive” — this is the essence of the scissors effect.

· · ·

CHAPTER 06End A — Deep Mechanisms: Cognitive Debt and Brute-Force Reconstruction

6.1 “Outcome Bias” in Pre-training Data — The Cognitive Debt Hypothesis

Internet training corpora harbor a systematic publication bias: textbooks record only theorems and proofs, not the failed derivation attempts; papers publish only successful experiments, not failed ones; code repositories contain only the final passing code, not the debugging process. LLMs consequently learn “what the answer is” (outcome knowledge) but lack “how to reach the answer” (process knowledge).

“LLMs trained primarily on positive-result literature inherit a distorted model of science — one in which most experiments succeed. What’s missing isn’t the ability to learn from failure, but the published failure data from which to learn.”
— “LLMs Have Made Failure Worth Publishing”, arXiv 2604.06236

Empirical support: incorporating failed chemical reactions into model training improved prediction accuracy; training solely on failure samples yielded orders-of-magnitude improvements in success rates (Lee et al., 2025). Meanwhile, DeepSeek-R1 and QwQ-32B achieved 80.1% answer correctness, but human experts judged only 39.7% of reasoning paths to be sound (Guo et al., 2025).

6.2 The True Nature of Thinking Tokens: Brute-Force Compute Reconstruction

Training-time scaling has plateaued (diminishing returns from data/parameters), making inference-time scaling (test-time compute) the only remaining pathway to capability improvement. But the essence of Thinking Tokens is not efficient reasoning — it is probabilistic search/trial-and-error within a pre-stored answer space.

Key evidence: Yue et al. (2025) proposed a widely supported hypothesis — “all correct reasoning paths already exist within the base model; RLVR merely improves sampling efficiency.” The base model’s pass@k at large k values even exceeds that of RL-trained models. The paper “Do Thinking Tokens Help or Trap?” found that thinking tokens trigger “thinking traps” — unproductive redundant verification loops.

FIGURE 1 — The Essential Difference Between Human Reasoning and LLM “Reasoning”
Human expert solving a problem:
See problem → Recognize pattern → Invoke experience → Reach answer directly
Token equivalent: ~200

LLM “reasoning” to solve a problem:
See problem → Try method A → Fail → Backtrack → Try method B → Partial success
→ Verify → Doubt → Re-verify → Try method C → Success
→ Verify again → Doubt again → Re-verify → Finally output
Token equivalent: ~8,000–50,000

LLM “thinking” is essentially brute-force search through probability space — every 1 unit of effective output requires 5–50 units of unproductive trial-and-error

6.3 The Structural Inefficiency of Thinking Tokens

The inefficiency of Thinking Tokens is not an incidental engineering problem but a structural one. Multiple studies provide precise quantitative evidence:

Invisibility: Apiyi (2026)’s analysis of Gemini 3.1 Pro showed that over 95% of output tokens are invisible to users, consumed by the model’s reasoning chain. Users see only the final answer but pay for the full set of reasoning-chain and answer tokens. AIOutlooks (2026)’s benchmarks showed that a complex query can generate 10,000 thinking tokens; at GPT-5.5’s $30/M output token price, the invisible thinking cost is $0.30 while the visible answer costs just $0.006 — users pay 50× more for “thinking” than for the “answer.”

Inflation: TianPan.co (2026)’s Wharton study showed that CoT prompting inflated token costs by 2–5× and added seconds of latency, with no measurable accuracy improvement for most production tasks. Reasoning models gained only 2–3% accuracy improvement from explicit CoT, while response times increased by 20–80%. More seriously, “overthinking spirals” occur — LLMs frequently continue generating reasoning steps after already arriving at the correct answer.

Non-monotonicity: Multiple studies (Aggarwal & Welleck 2025; “Does Thinking More Always Help?” 2025) found that extending test-time thinking exhibits non-monotonic behavior: initially improving accuracy, but subsequently degrading performance. On knowledge-intensive tasks, increasing test-time compute not only fails to consistently improve accuracy but in many cases produces more hallucinations.

Missing Optimal Stopping Points: OptimalThinkingBench (Aggarwal et al., 2025) found that no existing model achieves “optimal balance” — reasoning models over-explain simple problems while non-reasoning models under-think difficult ones. The paper “Stop Spinning Wheels” (2025) noted that a universally generalizable, quantifiable “optimal stopping point” for effectively avoiding redundant reasoning steps has still not been found.

“Hidden reasoning once looked like progress toward artificial thought. By 2025, it also looks like a new kind of inflation — computation silently expanding under the guise of intelligence. Whether the AI economy is stable or a bubble will hinge on a single ratio: how many tokens a model must think before earning back a dollar.”
— Turing Post, FOD#122: Thinking Tokens Explained

6.4 The Three-Layer Uncontrollable Token Cost Black Hole

User-controllable Input + Output accounts for approximately 5% of total tokens. The remaining 95% consists of three layers of uncontrollable consumption: (1) Thinking/Reasoning Tokens — invisible to users yet billable, with consumption volumes determined autonomously by the model, “there’s no fixed budget; the model spends what it thinks the problem needs” (EG3, 2026); (2) Agent loop accumulation — every LLM API call is stateless, so the agent sends the complete conversation history with each tool invocation, with input exceeding 50K tokens per call by step 20 (LeanOps, 2026), and a Reflexion loop running 10 rounds consuming 50× the tokens of a single linear pass; (3) Tool calls and searches — multi-step autonomous agents can inadvertently enter recursive loops, over-query systems, or expand tasks beyond their original scope, prompting Portal26 (2026) to release dedicated “Agentic Token Controls” modules to contain runaway consumption — the most notorious disaster being a LangChain infinite loop running for 11 days and generating $47,000 in API fees.

6.5 Structural Lock-In: Irreducible and Irreversible

The above three layers of uncontrollable tokens are not optimizable “waste” — they are the technology itself, subject to a triple structural lock-in:

Lock-In ①: Training scaling has plateaued → Inference scaling is the only path forward. As of 2025, test-time compute is widely regarded as the most likely key driver of LLM performance improvement — because raw pre-training Scaling Laws have hit data bottlenecks and diminishing returns. From 2020 to 2024, frontier progress was dominated by training scale; by 2024–2025, the field added a second axis — inference scale. The path is locked in with no return.

Lock-In ②: Thinking tokens are the core carriers of capability. A NeurIPS 2025 paper proved through information theory (mutual information) that specific generation steps in reasoning trajectories exhibit sudden, significant increases in mutual information — the “MI spike phenomenon.” These thinking tokens are critical to reasoning performance, while other tokens have minimal impact. Empirically, a small model allowed to think for 10 seconds can effectively outperform a model 14× its size that answers immediately — the architecture hasn’t changed; only the token budget has.

Lock-In ③: Agent tool calls are the externalization of intelligence. Search, API calls, code execution — these are not “add-on features” but the agent’s only channel for interacting with the real world. Without the Action→Observation feedback in the ReAct loop, an agent is merely a closed language model, unable to verify facts, obtain real-time information, or execute any real-world action.

6.6 Reconciling the Paradox: Why “90% Waste” and “Irreducible” Are Not Contradictory

This paper argued in Section 6.3 that the majority of Thinking Tokens consist of unproductive redundant verification and trial-and-error loops, and in Section 6.5 that Thinking Tokens cannot be eliminated. These two claims appear contradictory, but their coexistence reveals the essence of the problem: you don’t know which 10% is effective until the full 100% has been generated.

This is like navigating a maze — you don’t know which path leads to the exit until you’ve explored them all. In hindsight, 90% of paths are “wasted”; but in advance, each path could be the right one. This is not an engineering defect — it is the mathematical nature of probabilistic search. You cannot keep only the “correct 10%” because identifying which portion is correct itself requires completing the full 100%. Researchers (NoWAIT, TRS) are attempting to compress redundant portions (27–51%), but this is post-hoc optimization after paths have been traversed, not pre-hoc prediction. The structural token “waste” is an unavoidable cost of achieving effective reasoning.

· · ·

CHAPTER 07The Trial-and-Error Information Supply System: A Unified Theory of Skill, Search, and Memory

7.1 Multi-Layer Trial-and-Error Information Supply Model

Every component in the entire AI technology stack is doing the same thing — injecting “process” information into a model that only learned “outcomes”:

The Unified Essence of the AI Technology Stack
Layer Component Information Source Cost Characteristics
Layer 1 System Prompt + Skill Crystallized trial-and-error of human engineers Fixed tokens (one-time)
Layer 2 Search + RAG Publicly available trial-and-error records from web/enterprise Variable tokens (pay-per-use)
Layer 3 Agent Memory Agent’s own accumulated trial-and-error Cumulatively growing tokens
Layer 4 Thinking Tokens Real-time compute reconstruction Uncontrollable tokens (most expensive)

From top to bottom, cost increases while controllability decreases. The system design objective is for upper layers to supply as much information as possible, thereby narrowing the search space for the bottom-layer Thinking.

7.2 The COP (Constrained Optimization Problem) Nature of Skills

Empirical dissection of Anthropic’s SKILL.md files shows that every rule can be precisely classified into three types of information: correct paths (“Use WidthType.DXA”), error paths (“NEVER use WidthType.PERCENTAGE — breaks in Google Docs”), and result constraints (“columnWidths must sum to table width”). This is perfectly isomorphic to the standard definition of a Constrained Optimization Problem (COP): decision variables = each choice in code, constraints = NEVER/MUST rules, objective function = generate correct, renderable output.

7.3 The Failure of Skill Reuse

GitHub Issue #7777 reported that Claude begins ignoring CLAUDE.md instructions after just 2–5 prompts. The AAAI 2026 ConInstruct paper found that even when models detect instruction contradictions, they silently choose not to report them. Anthropic’s official documentation warns that “bloated CLAUDE.md files cause Claude to ignore actual instructions.” Root cause: the probabilistic sampling mechanism cannot 100% execute deterministic constraints.

7.4 The “Cognitive Debt” Model

FIGURE 2 — The Causal Structure of Cognitive Debt
Human civilization’s “publication bias” (records successes, omits trial-and-error)

LLM pre-training’s “knowledge debt” (learned outcomes, owes processes)

┌─────────┴─────────┐
↓ ↓
Thinking Agent Memory
Real-time rebuild Persist trial-and-error
via compute processes
= Interest on debt = Principal repayment
(one-time, costly) (reusable, cumulative)
↓ ↓
Cost explosion Persistent token inflation
└─────────┬─────────┘

Death Crossover Point
Interest on knowledge debt > User’s ability to pay
Thinking is the interest on the debt; Memory is the principal repayment; and the creditor is the publication bias of human civilization’s millennia-long habit of “publishing only successes”
· · ·

CHAPTER 08End B Mechanism: Repair Loops and Compounding Costs of Delivery Quality

8.1 Delivery Uncertainty of Probabilistic Systems

The uncertainty in AI delivery quality is not an occasional lapse but an essential property of the probabilistic sampling architecture. Multiple large-scale studies provide precise quantification:

Quality Defect Statistics for AI-Generated Code
Metric Data Source
AI-generated code bug rate vs. human 1.7× Stack Overflow / CodeRabbit, 2026
AI code containing security vulnerabilities 45% Ranger, 2026
AI programs producing incorrect output 26.6% Ranger, 2026
Silent logic failure rate 60% Ranger, 2026
Java implementation security failure rate >70% Ranger, 2026
XSS vulnerability failure rate 86% Ranger, 2026
Excess logic and correctness errors 75% (194 per 100 PRs) CodeRabbit, 2026
Code duplication rate increase in AI-assisted codebases GitClear, 2025
Refactoring rate decline From 25% to <10% GitClear, 2025

The most dangerous failure mode is what Shiplight (2026) described as the “silent green light”: code passes every test in the suite yet is still wrong — CI stays green because tests check for specified behavior, while bugs lurk in unspecified behavior that no one thought to test. Ninety percent line coverage combined with 30% behavioral coverage means 70% of potential bugs can sail through with a green light.

8.2 Lusser’s Law Replicated in AI Interactions

Towards Data Science (2026) cited the law derived by German engineer Robert Lusser in the 1950s from consecutive rocket failures: the overall reliability of a complex system equals the product of all component reliabilities. A 95% accurate agent has only a 36% success rate on a 20-step task. Users receive approximately 95% satisfaction / 5% defect per delivery; fixing defects introduces new ones; the success rate decays exponentially at 0.95^N, while costs grow quadratically at N(N+1)/2. An 85% accurate agent has only a 20% success rate on a 10-step task — four out of five attempts are destined to fail.

8.3 Community Validation: A Universal Experience

The community has christened the “Ralph Wiggum Loop” with dark humor — named after the Simpsons character who shouts “I’m helping!” while being utterly useless — an infinite retry pattern implemented with biting irony: while :; do cat PROMPT.md | agent ; done. This has become an official Cursor plugin, with the community using it to acknowledge a reality: the essence of AI work is infinite retries, and progress doesn’t persist in the LLM’s context — it lives in files and git history. The DEV Community hot post “I Wrote 200 Lines of Rules for Claude Code. It Ignored Them All.” struck a deep chord. Another highly upvoted post, “When AI Gets Stuck, Don’t Fix It — Restart It,” argued that once internal state is corrupted, additional instructions are no longer neutral information — earlier erroneous assumptions become implicit, the model compresses context in detrimental ways, and each correction is interpreted through an already-broken framework. You’re no longer guiding reasoning; you’re negotiating with a collapsed internal state.

8.4 The Cost Multiplication Effect of Repair Loops

Each round of the repair loop compounds costs: context grows longer with each round (history accumulation), making per-round costs increase progressively; the model begins ignoring earlier rules (Chroma’s 2025 Context Rot study: performance starts degrading after 32K tokens), reducing repair efficiency; context compression causes critical Skill rules to be summarized into oblivion; and forced restarts zero out all previously consumed tokens, constituting complete sunk costs.

Cost Structure of Repair Loops:

Interaction 1: A₁ tokens (baseline consumption)
Interaction 2: A₂ tokens (A₂ > A₁, context accumulation)
Interaction 3: A₃ tokens (A₃ > A₂, longer history)

Interaction N: Aₙ tokens

Total cost = Σ(Aᵢ) ≈ A × N × (N+1) / 2 ← Quadratic growth
Success rate = 0.95^N ← Exponential decay
User satisfaction = Collapses after round 3–4 ← Either give up or restart

After restart: All prior token consumption resets to zero = Sunk cost

LeanOps (2026)’s audit of 30 engineering teams yielded staggering numbers: a 20-person team could incur monthly agent costs of $110,000. Within the same tool, costs between the 10th and 90th percentile users differed by 20×. Mavvrik (2025)’s survey showed that 84% of enterprises reported that AI spending had already eroded gross margins by more than 6 percentage points. The critical metric is no longer “first-pass rate” but “convergence” — whether an agent’s iterations trend toward completion or spin in circles.

· · ·

CHAPTER 09The Dual-End Squeeze Model: The Multiplicative Structure of A×B

9.1 Model Definition

Actual Total User Cost = End A (Per-Interaction Cost) × End B (Number of Repair Loop Iterations)

End A = f(technical complexity, physical wall, uncontrollable tokens) → Continuously inflating
End B = g(probabilistic sampling precision, task complexity) → Ineliminable

A × B = Super-linear cost amplification; degenerates to near-exponential runaway in tree search, recursive tool calls, and restart cycles

9.2 Why Multiplication, Not Addition

End A and End B are not independent cost items but nested — each repair round (End B) must pay the full End A cost, and End A cost increases with each round (context accumulation). In the minimal model, total cost grows quadratically (Appendix B); if task complexity, branching factors, and restart counts all increase simultaneously, costs may escalate to higher-order polynomial or even near-exponential runaway. arXiv:2604.22750 (April 2026), a systematic analysis of 8 frontier models on SWE-bench Verified, confirmed: agentic coding tasks consume 1,000× the tokens of code reasoning/code chat, token consumption between different runs of the same task can differ by 30×, and higher token usage does not translate to higher accuracy — accuracy peaks at medium cost levels and actually declines at the highest cost levels, suggesting that excess token expenditure reflects unproductive exploration rather than deeper reasoning. An ICSE 2026 paper further confirmed that agent token consumption grows “quadratically” with each interaction round.

9.3 Precise Definition of the “Death Crossover Point”

The death crossover point arrives when the growth rate of A×B exceeds the user’s budget ceiling. This is precisely the situation in 2026:

2026 Death Crossover Point — Empirical Evidence
Event Date Source
Uber burned through its entire annual AI budget in four months Apr 2026 Fortune
Microsoft revoked internal Claude Code licenses May 2026 The Verge
JPMorgan: “Tokens are eating internet profits alive” May 2026 JPMorgan Research
85% of enterprise AI budgets severely off forecast 2025 Mavvrik
Gartner predicts 40% of agent projects will be canceled 2025–2027 Gartner
AI inference accounts for 85% of enterprise AI budgets 2026 AnalyticsWeek

9.4 The Impossible Triangle

Under the current paradigm, an impossible triangle exists: Frontier Capability (requires more tokens), Cost Control (physical wall limits price reductions), and User Budget (affordability has an upper bound) — at most two of the three can be satisfied simultaneously.

· · ·

CHAPTER 10Discussion: The Structural Limits of the Current Paradigm

10.1 Why Engineering Optimization Cannot Solve the Fundamental Problem

Trajectory compression (AgentDiet, Re-TRAC) hits a ceiling at 30–50% reduction, while consumption is growing at 10–100× rates. Token caching/pointerization mitigates but does not eliminate the problem, and introduces new complexity. Model routing (large/small model splitting) is a Gartner-recommended FinOps strategy, but frontier tasks cannot be downgraded. Skill pre-encoding is constrained by probabilistic sampling’s non-compliance — Vercel benchmarks showed that agents fail 56% of the time when independently deciding to fetch context.

10.2 Three Paradigm-Level Defects

Defect ①: Autoregressive generation → output tokens depend sequentially → no parallelism → cost scales linearly with length. Defect ②: Probabilistic sampling → no guarantee of deterministic constraint compliance → unpredictable delivery quality. Defect ③: Outcome bias in pre-training data → permanent absence of process knowledge → compute must reconstruct it at inference time. The three defects are mutually coupled, constituting the fundamental limitations of the current Transformer + autoregressive + probabilistic sampling architecture.

10.3 Potential Paradigm Breakthrough Directions

Breaking through End A requires non-brute-force reasoning approaches: Recursive Language Models (RLM, Prime Intellect 2025) enable models to actively manage their own context; process knowledge training incorporates trial-and-error trajectories into training data. Breaking through End B requires deterministic execution: neuro-symbolic hybrid systems let symbolic reasoning handle constraints while neural networks handle pattern matching; deterministic constraint enforcement layers perform hard rule checking before and after sampling. Both imply fundamental transformations of the current architecture.

· · ·

CHAPTER 11You Think These Can Save You — Here’s Why They Can’t

Having identified the structural roots of the cost crisis, it is necessary to directly address the six most common optimistic counterarguments in the industry and explain why they are insufficient to reverse the trend for high-complexity agentic tasks.

11.1 “Model distillation can compress frontier reasoning into small models”

Distillation is indeed effective — DeepSeek R1’s open-source distilled version approaches o1 on certain benchmarks. But distillation mitigates End A’s per-interaction cost without altering End B’s repair loops. Distilled small models still exhibit significant capability gaps in tasks requiring long-trajectory reasoning, multi-tool coordination, and complex code refactoring. Uber’s engineers were already using the most advanced Claude Code — not because cheaper alternatives didn’t exist, but because cheaper alternatives couldn’t complete their tasks.

11.2 “Synthetic data can train process knowledge and repay cognitive debt”

The direction is correct — using frontier models to generate training data with complete reasoning chains does hold potential. But the diversity and authenticity of synthetic trajectories remain open questions. Training the next generation of AI on AI-generated trial-and-error data risks exacerbating “Model Collapse” — a phenomenon already documented by Nature (2024). This is the most promising yet unvalidated pathway for repaying “cognitive debt.”

11.3 “Speculative Decoding / KV Cache / MoE can reduce per-token compute costs”

These technologies reduce the price function C(t) — the actual computational cost per token. But they do not alter the growth structure of A(n) (per-round context accumulation) or the existence of B (repair loops). In other words, each token becomes cheaper, but the number of tokens you need is growing at an even faster rate. Epoch AI (2025) data shows algorithmic efficiency improving at 3–10× per year — but over the same period, token consumption grew 5–50×.

11.4 “On-device inference can execute End B simple repairs for free”

Offloading simple verification and formatting repairs to local NPUs can indeed partially break the A×B multiplicative structure — turning it into A_cloud + B_local×0. But frontier reasoning tasks (multi-step code refactoring, long-document analysis, complex planning) cannot run on-device. The agent workloads consumed by Uber and Microsoft require server-side inference with hundreds of thousands of tokens of context — not something a phone NPU can handle.

11.5 “Model routing can divert simple tasks to cheaper models”

A Gartner-recommended FinOps strategy that is effective for cost management. But the problem is: the value proposition of agentic AI is built precisely on frontier capabilities. If tasks are routed to cheap models to save costs, the very reason for deploying AI is undermined. Gartner itself warned: “If you keep pushing toward the frontier, your token costs will explode to levels where you’ll never be profitable.”

11.6 The Strongest Real-World Stress Test: Microsoft Had Everything — and Still Contracted

Microsoft simultaneously possesses: on-device inference (Windows Copilot Runtime + NPU), model distillation (Phi series small models), KV Cache optimization (Azure infrastructure), model routing (Foundry platform), and proprietary code tools (GitHub Copilot CLI). It is the company with the most complete technology stack in the industry, possessing every single “solution” listed above.

And then it began revoking Claude Code licenses just six months after rollout.

This demonstrates that even with a complete optimization stack, real-world scaled deployment can still trigger cost-driven contraction. Microsoft’s retreat may involve multiple factors (pushing proprietary tools, security consolidation, reducing external dependency), but AI Weekly (May 2026) revealed a more fundamental issue: Microsoft’s internal data showed that AI agent deployment costs had exceeded equivalent human labor costs. Enterprise AI ROI remained unproven in Microsoft’s own project data, with spending growth outpacing measurable productivity gains. This is the strongest real-world constraint case against the belief that “engineering optimization can solve everything.”

11.7 “AI Is Cheaper Than Humans” — Is It Really?

Gemini 3.1, in its review of this paper, raised an important counterargument: even if AI costs are exploding, if the human labor it replaces is more expensive, ROI remains positive. This counterargument holds for simple, high-frequency, standardizable tasks — Forrester TEI data shows customer service AI at $0.46 per interaction vs. $4.18 for human agents (9× advantage), and routine code review at $0.72 vs. $48 for a senior engineer (66× advantage).

But for complex, long-trajectory, multi-step agentic tasks, this assumption is collapsing:

AI Agent vs. Human Labor: Cost Inversion on Complex Tasks
Data Source Finding
Microsoft internal data (AI Weekly, May 2026) AI agent deployment costs have exceeded equivalent human labor costs
Gartner 2026 cohort study Only 41% of agent deployments achieved positive ROI within one year; 19% will never break even
Forrester 2026 Unmeasured rework absorbed 22–38% of self-reported time savings
Deloitte 2026 84% of enterprises have not yet redesigned workflows around AI — most run human and AI in parallel, paying double
Uber COO (Fortune, May 2026) “There’s no direct correlation between token consumption and useful consumer features”

In short: AI is indeed 9–66× cheaper than humans on simple tasks. But the core thesis of this paper points precisely toward frontier agentic tasks — and in that domain, the assumption that “AI is cheaper than humans” is being refuted by Microsoft’s own data.

11.8 The Possibility of Architectural Escape

Gemini 3.1 also raised a deeper counterargument: if future architectures evolve so that “the cost of verifying an answer is far lower than the cost of generating it” (such as Energy-Based Models (EBMs), diffusion planning, Test-Time Training (TTT)), then End B retries would not need to re-traverse the full historical context from scratch, and the A×B multiplicative structure could potentially be broken. This is the most promising direction among all counterarguments — it doesn’t optimize within the current paradigm but instead discusses alternatives to the paradigm itself. This paper acknowledges this direction merits tracking, but also notes: as of June 2026, no such architecture has demonstrated economic viability on production-grade agentic tasks. The current crisis demands current responses, and we cannot bet on unproven paradigm breakthroughs.

When your engineer’s monthly API costs can cover a junior engineer’s annual salary, this is no longer a problem that FinOps tools can solve. This is an architecture problem.
— The Next Web, “Microsoft’s quiet Claude Code retreat”, May 2026
· · ·

CHAPTER 12Conclusion: To Every AI Decision-Maker

12.1 Five Things You Need to Understand Right Now

First, falling token unit prices do not mean your AI bill will fall. Architectural complexification (brute-force Thinking Token search + agent loop accumulation + search/tool calls) is consuming tokens at rates that far exceed the pace of price reductions. Uber proved this with real money.

Second, 95% of your AI bill is invisible to you. Thinking Tokens, agent history retransmissions, tool call returns — these uncontrollable costs are determined autonomously by the model. Your input and output are merely the tip of the iceberg.

Third, fixing one bug introduces new bugs, and fixing those introduces newer ones. This is not coincidental — it is a mathematical inevitability of probabilistic sampling systems. The cost of each repair round increases (context accumulation), the success rate decreases (Lusser’s Law), and after 3–4 rounds most people give up or restart — with all previously consumed tokens reset to zero.

Fourth, the physical wall is locking down price reduction headroom. TSMC raises prices year after year, HBM memory demand outstrips supply, and data center power permitting queues stretch 24–36 months. Economy-tier models are still getting cheaper, but the frontier reasoning models you actually need — they have barely gotten cheaper at all.

Fifth, these two ends multiply rather than add. End A (each interaction getting more expensive) × End B (repair loops ineliminable) = super-linear total cost growth. Enterprises budgeting with addition — will be the next Uber.

12.2 Action Recommendations

Immediately: Audit your agent workflows and distinguish the ratio of “controllable tokens” to “uncontrollable tokens.” If the uncontrollable portion exceeds 80%, your budget model is already broken.

Short-term: Establish real-time token budget monitoring and circuit-breaker mechanisms. Set per-task token caps, with agent loops automatically terminating when thresholds are exceeded. This sacrifices some flexibility but prevents runaway costs.

Medium-term: Invest in COP-ized pre-encoding of Skills/Memory — pre-encode more trial-and-error information as deterministic rules to shrink the Thinking search space. This is the highest cost-efficiency optimization path available under the current architecture.

Strategic: Do not use “AI is getting cheaper” as justification for expanding deployment. Incorporate the A×B multiplicative factor into your financial models. If your CFO is using a linear model to forecast AI costs — show them this paper.

12.3 A Final Word

The cost of intelligence is falling. The cost of deploying intelligence is soaring. Data from January through May 2026 no longer allows anyone to pretend this isn’t a problem. Uber burned through its budget. Microsoft revoked licenses. JPMorgan says tokens are eating profits alive. The next one won’t be “some unlucky company” — if you’re scaling agentic AI, the next one is you.

Gartner analyst Will Sommer’s warning deserves a place on every AI decision-maker’s desk: “Chief product officers must not conflate the deflation of commodity tokens with the democratization of frontier reasoning. Those who use cheap tokens today to mask architectural inefficiency will find agentic scale-up permanently out of reach tomorrow.”

APPENDIXAppendices

Appendix A: Key Data Timeline

AI Cost Crisis Evolution Timeline
Date Event Cost Impact
2022.01 Chain-of-Thought (Wei et al.) published ~500 tokens per task
2022.10 ReAct (Yao et al.) published ~10K tokens per task (20×)
2023.03 Reflexion (Shinn et al.) published; GPT-4 launched at $30/M ~35K tokens per task (70×)
2023.05 ReWOO proposes token efficiency countermeasure Token consumption −80% (sacrificing flexibility)
2023.10 LATS (Zhou et al.) published Tree search → ~150K tokens per task
2024.07 GPT-4o mini launched at $0.15/M Economy-tier price drops to historic low
2024 Agent+RAG production deployment Average enterprise AI spend $2.5M/year
2025.01 DeepSeek R1 released at $0.55/M Price war begins, 90% below competitors
2025.07 Context Rot formally named by Chroma Performance degrades after 32K tokens
2025.09 AgentDiet published; Claude Sonnet reaches 100B tokens/day Trajectory slimming reduces 40–60% tokens
2025.10 o3 Deep Research launched at $10/$40/M $0.41–$1.32 per query
2025.12 TSMC announces 2026 price hikes of 3–10% Advanced-node cost reversal upward
2026.01 Luccioni FAccT paper (Jevons Paradox × AI) Academic confirmation of rebound effects
2026.01 AWS quietly raises prices by 15% Signal of cloud provider subsidy retreat
2026.02 Re-TRAC published; ETH study on AGENTS.md efficacy Academic confirmation of trajectory compression + Skill degradation
2026.03 Gartner report: Inference costs expected to rise “Do not conflate token deflation with reasoning democratization”
2026.03 Du “Tiered Super-Moore” paper Frontier model price reductions near-stagnant (R²=0.031)
2026.04 Uber burns through annual AI budget in four months Landmark death crossover event
2026.04 DeepSeek V4 released at $0.0036/M (cached) Economy-tier extreme price reduction vs. frontier tier unchanged
2026.05 Microsoft revokes Claude Code; JPMorgan warning “Tokens are eating internet profits alive”

Appendix B: Mathematical Derivation of the Dual-End Squeeze Model

=== Dual-End Squeeze Model ===

Definitions:
S = Number of dependent steps per task (agent loop steps)
p = Per-step reliability rate (typical value ≈ 0.95)
R = Number of user repair/retry rounds
A(r) = Token consumption for the r-th repair round
C(t) = Token unit price at time t ($/token)

Lusser’s Law (multi-step task reliability):
Overall success rate of an S-step dependent task = p^S
When p = 0.95, S = 20: 0.95²⁰ ≈ 36%
When p = 0.95, S = 40: 0.95⁴⁰ ≈ 13%
→ More steps = higher overall failure probability = more frequent repair needs

End A Model (growth of per-round cost):
A(r) = A₀ + Σᵢ₌₁ʳ⁻¹ Oᵢ
where Oᵢ is the output of round i (including Thinking + tool returns + history)
Because agents retransmit full history with every call: A(r) grows at least linearly
In practice, due to variable tool return volumes: A(r) growth is often super-linear

End B Model (cost accumulation of repair loops):
Total cost of R repair rounds = C(t) × Σᵣ₌₁ᴿ A(r)
If A(r) ≈ A₀ + k·r (linear growth), then:
Total cost ≈ C(t) × [R·A₀ + k·R(R+1)/2] → Quadratic growth

Note: In practice, users typically give up or restart after 3–5 rounds
Restart = all previously consumed tokens reset to zero (sunk cost)
Actual expenditure = Σ(accumulation before each restart) → May far exceed single continuous estimate

Dual-End Multiplicative Effect:
If technological generation makes A₀ grow by factor k (End A inflation)
And repair rounds R are ineliminable (End B fixed)
Then TotalCost ∝ k × R² → Super-linear growth

In typical multi-round repair, costs grow at least quadratically;
In tree search, recursive tool calls, and agent loops without stop conditions,
near-exponential blowups are possible (branching factor × depth).

Physical Wall Constraint:
C(t) is no longer declining (frontier reasoning tier R²=0.031)
A₀ continues to grow (architectural complexification + inference scaling)
R is ineliminable (probabilistic sampling + software complexity)
→ Actual total user costs exhibit persistent super-linear growth □

Appendix C: Anatomy of a SKILL.md COP Structure — Worked Example

Using the actual content of Anthropic’s docx Skill file as an example, each rule can be precisely classified into the three COP information types:

COP Classification of SKILL.md Rules
Type Symbol Example Rule Trial-and-Error Origin
Correct Path Use LevelFormat.BULLET with numbering config Correct method discovered
Correct Path Use ShadingType.CLEAR Correct method discovered
Correct Path Always use WidthType.DXA Correct method discovered
Error Path NEVER use unicode bullets Someone tried → rendering error
Error Path Never use WidthType.PERCENTAGE Someone tried → Google Docs crashed
Error Path Never use ShadingType.SOLID Someone tried → background turned black
Error Path Never use tables as dividers Someone tried → empty frames appeared
Error Path Never use \n Someone tried → invalid XML
Result Constraint 🔒 CRITICAL: docx-js defaults to A4 → set explicitly Format not as expected
Result Constraint 🔒 columnWidths must sum to table width Rendering misalignment
Result Constraint 🔒 ImageRun requires type (png/jpg) Image failed to load
Result Constraint 🔒 Element order in pPr: pStyle, numPr… rPr last Schema violation

COP formalization: Skill(task) = { P_correct, P_error, C_hard, C_soft, F_objective }. The essential function of a Skill is to supply P_error (known failure paths) in advance, pruning dead branches from the COP search space so that Thinking only needs to search the feasible region near P_correct. The birth of each rule = the codification of one human-AI trial-and-error interaction.

REFERENCES

  1. Yao, S. et al. (2022). ReAct: Synergizing Reasoning and Acting in Language Models. ICLR 2023. arXiv:2210.03629
  2. Shinn, N. et al. (2023). Reflexion: Language Agents with Verbal Reinforcement Learning. NeurIPS 2023. arXiv:2303.11366
  3. Zhou, A. et al. (2023). Language Agent Tree Search Unifies Reasoning, Acting, and Planning. ICML 2024. arXiv:2310.04406
  4. Xu, B. et al. (2023). ReWOO: Decoupling Reasoning from Observations. arXiv:2305.18323
  5. Luccioni, A.S., Strubell, E. & Crawford, K. (2025). From Efficiency Gains to Rebound Effects: The Problem of Jevons’ Paradox in AI. ACM FAccT 2025. arXiv:2501.16548
  6. Cottier, B. et al. (2025). LLM inference prices have fallen rapidly but unequally across tasks. Epoch AI.
  7. Du, M. (2026). Tiered Super-Moore’s Law: Price Evolution in LLM Inference Services. Wuhan University. arXiv:2603.28576
  8. Erol, M.H. et al. (2025). Cost-of-Pass: An Economic Framework for Evaluating Language Models. arXiv:2504.13359
  9. Xiao, Y. et al. (2026). Reducing Cost of LLM Agents with Trajectory Reduction. Proc. ACM Softw. Eng. (FSE 2026). arXiv:2509.23586
  10. Zhu, J. et al. (2026). RE-TRAC: Recursive Trajectory Compression for Deep Search Agents. arXiv:2602.02486
  11. Guo, D. et al. (2025). Right Is Not Enough: The Pitfalls of Outcome Supervision. arXiv:2506.06877
  12. Yue, Y. et al. (2025). Limit of RLVR. limit-of-rlvr.github.io
  13. Demirer, M., Fradkin, A. et al. (2025). The Emerging Market for Intelligence. NBER Working Paper No. w34608.
  14. Hong, K. et al. (2025). Context Rot: How Increasing Input Tokens Impacts LLM Performance. Chroma Research.
  15. Gartner (2026). Navigating the Commoditization Trap as Token Costs Fall by Over 90% Through 2030.
  16. Appenzeller, G. (2024). Welcome to LLMflation. a16z Blog.
  17. “From Storage to Experience” Survey (2026). Evolution of LLM Agent Memory Mechanisms. Preprints.org.
  18. Ding, L. (2026). AgentHER: Hindsight Experience Replay for LLM Agent Trajectory Relabeling. arXiv:2603.21357
  19. Wang, A. et al. (2025). Do Thinking Tokens Help or Trap? arXiv:2506.23840
  20. TRS (2026). Thinking with Reasoning Skills: Fewer Tokens, More Accuracy. NeurIPS 2025 Workshop. arXiv:2604.21764

LEECHO Global AI Research Lab
이조글로벌인공지능연구소
&
Opus 4.6 · GPT 5.5 · Gemini 3.1
Cognitive Collective (인지집단)
V4 · JUNE 1, 2026
Version History

V1 (2026.6.1): Initial version, co-authored by LEECHO Global AI Research Lab and Anthropic Claude Opus 4.6. Based on conversational research methodology, progressively deriving the cost crisis model from ReAct task chains.

V2 (2026.6.1): Completed all missing outline sections — Ch.5.3 Structural Inefficiency of Thinking Tokens, Ch.7.1 Delivery Uncertainty of Probabilistic Systems, Ch.4.5 Divergence of Three Price Decline Curves, Appendices A/B/C. Triple-AI cross-review (Opus Dense + GPT 5.5 Dense + Gemini 3.1 Dense).

V3 (2026.6.1): Comprehensive revision based on three AI review reports + de-biasing analysis + external data alignment. Added Ch.02 Danger Index Timeline (Jan–May 2026), Ch.11 Counterarguments & Microsoft Case Study, Ch.12 Call to Action. Corrected Appendix B mathematics (distinguishing S/R variables), corrected OpenAI financial data, corrected Azure time-period scope.

V4 (2026.6.1): Revised based on GPT 5.5 / Gemini 3.1 cross-review of V3. Ch.09 mathematical language unified with Appendix B (removed overstatements of “exponential/quartic”); incorporated arXiv:2604.22750 empirical evidence (agentic coding 1000× tokens, 30× variance, non-monotonic accuracy); Ch.11 added §11.7 Labor Substitution Economics Rebuttal (Microsoft internal data + Gartner 41% ROI + Forrester rework absorption), §11.8 Architectural Escape Possibility; Microsoft case study reframed as “strongest real-world stress test”; Ch.02 Danger Index supplemented with 10/10 definition.


Cognitive Collective (인지집단)

LEECHO Global AI Research Lab — Core thesis origination, research leadership, critical identification of three-AI biases, revision principle decisions

Anthropic Claude Opus 4.6 — Paper writing, data retrieval, framework construction, three-AI synthesis analysis, de-biasing dissection

OpenAI GPT 5.5 — V2 cross-review (mathematical audit · evidence stratification · conditional downgrade recommendations)

Google Gemini 3.1 — V2 cross-review (counterexample supplementation · on-device perspective · synthetic data direction)

댓글 남기기