ORIGINAL THOUGHT PAPER · JULY 2026

Talent Density Determines Model Intelligence V1

From the Diminishing Returns of $310 Billion
to the Structural Convergence at $13.4 Billion


PublishedJuly 24, 2026
CategoryOriginal Thought Paper
DomainsAI Industrial Economics · Talent Structure Analysis · Data Quality Theory · Organizational Efficiency · Geopolitics
이조글로벌인공지능연구소
LEECHO Global AI Research Lab
&
Claude Opus 4.6 · Anthropic

0Abstract

Through systematic data collection and chronological analysis, this paper argues that the decisive variable determining AI model capability is not compute scale, algorithmic innovation, or capital investment, but rather “the complete thinking-process data of elite human experts”—that is, talent quality and density. We trace the full arc of industry evolution from the open-sourcing of RLVR algorithms to the explosion of the expert data market between 2023 and 2026, revealing a core paradox: OpenAI (~$180 billion in total funding) and Anthropic (~$132 billion) have collectively invested over $310 billion, while DeepSeek (~$7.4 billion) and Moonshot AI (~$6 billion) have invested a combined ~$13.4 billion. Yet on the Artificial Analysis Intelligence Index, Claude Fable 5 (59.86), GPT-5.6 Sol (58.89), and Kimi K3 (57.11) are clustered within a 3-point band, with Chinese open-source models already achieving localized superiority in coding and agent tasks, and offering API pricing 3–57× cheaper. This paper explains this “anomaly” across five dimensions—CEO R&D involvement, talent allocation functions, information channel asymmetries, social infrastructure, and industrial structure—and derives the inevitability of a US–China duopoly as the endgame of the AI race.

1Introduction: An Anomalous Phenomenon

1.1 $310 Billion vs. $13.4 Billion: Why Is the Gap Only 3 Points?

In July 2026, the global AI industry presents a puzzling picture. America’s two leading AI laboratories—OpenAI and Anthropic—have collectively raised over $310 billion in funding, employ approximately 13,000 people, and command the world’s largest GPU clusters. Meanwhile, China’s DeepSeek and Moonshot AI have raised a combined ~$13.4 billion, with a joint team that likely numbers fewer than 2,000.

Yet on the Artificial Analysis Intelligence Index—an independent benchmark—the four companies’ models are clustered within an extremely narrow band: Claude Fable 5 scores 59.86, GPT-5.6 Sol scores 58.89, and Kimi K3 scores 57.11. The gap is merely ~5%. Moreover, in the coding-specific Frontend Code Arena, Kimi K3 ranks first with an Elo of 1679, surpassing Claude Fable 5 (1631) and GPT-5.6 Sol (1618). Kimi K3 won 5 out of 6 real-world agent benchmarks. On API pricing, DeepSeek V4 Pro charges just $0.87 per million output tokens—57× cheaper than Claude Fable 5’s $50.

Funding Ratio 23:1 · Headcount Ratio ~7:1 · Capability Gap < 5% · Localized Overtaking Already Achieved

This is not a number that can be simply explained away as “the pursuer being more efficient.” A 23× funding gap corresponding to less than a 5% capability gap means the marginal capability gain per additional dollar spent by American labs is approaching zero. The central question of this paper is: Why does this anomaly exist? What is the true variable that determines model capability?

1.2 Core Thesis

Our core thesis consists of three mutually reinforcing propositions. First, the decisive variable for AI model capability is talent density, not capital density—specifically, the quality and quantity of “complete thinking-process data from elite human experts.” Second, CEO R&D involvement is the primary determinant of talent density—companies where the CEO personally conducts research exhibit zero information loss, the highest data quality, and maximum output per person. Third, China’s structural advantages cannot be replicated through capital, and no third country can catch up—China’s true AI competitiveness derives from a constellation of soft-power factors shaped over decades by history, culture, education systems, and industrial structure, not from surface-level capital expenditure.

2Algorithm Commoditization and the Formation of the Data Bottleneck

2.1 2022–2023: Pre-training Data Reaches Saturation

By 2023, major laboratories had already scraped most of the high-quality content from the open internet. The marginal return on additional pre-training data was approaching zero. The bottleneck began shifting from “data quantity” to “data quality.” That same year, the data annotation market was valued at approximately $1 billion, with annotation work still dominated by basic image labeling and text classification.

2.2 2024: Data Costs Surpass Compute Costs

In 2024, a landmark crossover occurred: the growth multiplier for data annotation costs reached 88×, while compute costs grew by only 1.3× over the same period. High-quality human data began to surpass compute as the single greatest bottleneck in frontier AI development. Scale AI’s revenue reached $870 million in 2024, while Surge AI—a bootstrapped company with roughly 110 employees—surpassed $1 billion in revenue. The market voted with real money, clearly identifying data, not compute, as the bottleneck.

That same year, two experimental findings laid the theoretical foundation for “data quality > data quantity.” The LIMA experiment (published at NeurIPS 2023, widely discussed in 2024) demonstrated that fine-tuning LLaMA with just 1,000 carefully curated high-quality samples could match or even surpass models trained on tens of thousands of crowdsourced samples. Microsoft’s Phi-2 experiment was even more extreme: a model with only 2.7 billion parameters matched or exceeded the 25× larger Llama-2-70B on complex reasoning benchmarks—the secret was not more parameters or compute, but “textbook-quality” curated data.

LIMA proved that 1,000 expert samples > 50,000 crowdsourced samples. Phi-2 proved that 2.7B parameters + textbook data > 70B parameters + ordinary data. The leverage of data quality can directly overturn compute scaling laws.

2.3 January 2025: DeepSeek R1 Exposes the Ceiling of Pure RL

The release of DeepSeek R1 was a watershed event. R1-Zero, trained with pure reinforcement learning, did exhibit emergent reasoning behavior—the model spontaneously learned to revisit and correct its own reasoning steps (the so-called “aha moment”). However, its outputs were chaotic, unreadable, and linguistically mixed. The solution was to introduce “cold-start data”—thousands of long chain-of-thought reasoning samples annotated by human experts, giving the model an initial “spark” before large-scale RL optimization.

This demonstrated a critical fact: the algorithmic framework (GRPO/RLVR) is the engine, but without high-quality human data to ignite it, the engine runs idle. Pure algorithms cannot bootstrap themselves to a sufficiently high quality level.

2.4 First Half of 2025: Algorithms Go Open-Source; Differentiation Shifts to Data

In the first half of 2025, the ReTool framework (later accepted at ICLR 2026) coupled RL with real-time code interpreter execution within the reasoning loop, formally validating RLVR as a paradigm capable of implicitly incentivizing correct reasoning chains. But more importantly, DeepSeek itself open-sourced its core algorithms—GRPO, MLA, and others—while Moonshot AI open-sourced the MoonClip optimizer, attention mechanism alternatives, and more.

The algorithmic moat was filled in by its own creators. When everyone has access to the same algorithms, the only axis of differentiation becomes data quality.

2.5 June–October 2025: The Data Industry Explodes

In June 2025, Meta acquired a 49% stake in Scale AI for $14.3 billion, with founder Alexandr Wang moving into Meta to lead the superintelligence team. A company whose primary product is “the industrial production of training data” was valued at approximately $29 billion—the market’s most direct pricing of the thesis that “data is AI’s core asset.”

In September 2025, xAI laid off approximately 500 general data annotators in a single move while simultaneously announcing a 10× expansion of its expert mentor teams in STEM, finance, and medicine. The layoff email read: “Accelerate and prioritize the expansion of our specialized AI mentor team, while reducing our focus on general AI mentor positions.” This was the most unvarnished strategic pivot statement from “cheap quantity” to “expensive quality.”

That same month, Mercor—a platform specializing in connecting domain experts with frontier labs—announced that its revenue had surged from $1 million to $500 million in 17 months. In October, it closed a $350 million Series C round at a $10 billion valuation.

2.6 2026: Industry Consensus Crystallizes

By 2026, multiple independent sources had converged on the same conclusion. A Prolific survey of 300+ AI practitioners identified the bottleneck as not compute, but “the pipeline for converting expert judgment into reliable training signals.” The BC Protocol paper concluded that high-quality expert chain-of-thought data is the core bottleneck of LLM post-training, and that existing methods each have structural deficiencies—crowdsourced annotation lacks deep reasoning paths, while experts writing in isolation are constrained by the “expert blind spot.” A March 2026 METR study found that code patches generated by AI agents passed automated tests but were rejected approximately half the time by actual repository maintainers upon review. In June 2026, Mercor’s annualized revenue surpassed $2 billion; the DATA Foundation launched with the explicit mission of “solving AI’s multi-billion-dollar training data bottleneck.”

Bottleneck Evolution: Crawl more web pages (2022) → Better filtering (2023) → Annotate more (2024) → Annotate better (2025) → The smartest people annotate personally (2026)

3The Leverage Effect of Data Quality (Experimental Evidence)

3.1 The LIMA Effect: Quality > Quantity

The LIMA experiment by Zhou et al. (2023), published at NeurIPS 2023, became the foundational work in data quality research. The experiment fine-tuned LLaMA 65B with just 1,000 carefully curated high-quality samples, achieving performance on par with or exceeding models trained on tens of thousands of crowdsourced samples. This proved that during the fine-tuning phase, the leverage ratio between data “quality” and “quantity” is drastically skewed—a small number of genuinely high-quality samples unlocks far more alignment capability than a large volume of quality-inconsistent data.

The 2025 LIMO paper (COLM 2025) extended the LIMA philosophy to the reasoning task domain, further demonstrating that quality overwhelms quantity. Crucially, it also noted that large-scale instruction data approaches have been criticized for relying on memorization rather than genuine generalization—when the same problem is presented with altered numerical values, LLM performance degrades. This reveals a key distinction: massive amounts of low-quality data produce memorization; small amounts of elite expert data produce generalization.

3.2 The Phi Series: Data Quality Overturns Scaling Laws

Microsoft Research’s Phi series experiments (2023) provided even more extreme evidence. Phi-2, with only 2.7 billion parameters, matched or surpassed the 25× larger Llama-2-70B on complex reasoning benchmarks. On the mathematical reasoning benchmark GSM8K, Phi-2 achieved an efficiency of over 21 points per billion parameters, far exceeding that of larger models. The secret was not more parameters or more compute, but “textbook-quality” curated data—specifically designed synthetic datasets for teaching models commonsense reasoning and general knowledge. The title “Textbooks Are All You Need” was itself a direct challenge to the “scale is everything” paradigm.

The Phi series fundamentally proved that the importance of carefully curated training data is at least equal to, and likely exceeds, model scale. The leverage effect of data quality can directly overturn compute scaling laws—meaning that leadership in the data quality dimension can offset tens-of-times disadvantages in the compute dimension.

3.3 The METR Experiment: The Ceiling of Automated Verifiers

If high-quality data is so important, can AI-generated or AI-verified data substitute for human experts? METR’s March 2026 study provided a definitive answer: no. The experiment found that code patches generated by AI agents could pass automated tests, but actual repository maintainers rejected approximately half of them upon review. The reasons for rejection were not that the code failed to run, but wrong architectural choices, inappropriate style, and overly brittle assumptions—all things invisible to automated verifiers and discernible only by humans who actually do the work.

This result is critical to our argument: it means that high-quality data production cannot be fully replaced by automation. Automated verifiers can check “correctness,” but they cannot evaluate “goodness.” The latter requires human expert judgment—and this is precisely why Liang Wenfeng has his core researchers personally annotating data.

3.4 The Black Box Problem of “High Quality”

The BC Protocol paper (2026) identified a deeper issue: the key bottleneck for LLM capability improvement has shifted from model architecture and parameter scale to data quality, yet “high quality” remains a vaguely defined black box in the research literature. Most work equates it with “well-formatted, instruction-diverse, fluent responses.” Few papers pursue the more fundamental question: who produced these data, under what cognitive conditions, and through what process? Experts writing in isolation are constrained by the “expert blind spot”—they systematically skip reasoning steps they consider self-evident. Crowdsourced annotation lacks deep reasoning paths. Specialized methods are needed to extract this tacit cognition.

This pushes the question one level deeper: it is not about “whether high-quality data exists,” but “what kind of people, under what conditions, using what methods, can produce genuinely high-quality data.” The answer points to this paper’s core thesis—the complete thinking processes of elite human experts.

3.5 Ranking the Contributions of the Three Driving Forces

Compute: A Necessary Condition with Diminishing Marginal Returns

Amazon, Alphabet, Meta, and Microsoft have a combined 2026 capital expenditure plan of $660–$690 billion. Amazon alone plans $200 billion—the largest single-year capital expenditure in corporate history. Anthropic has announced a $50 billion investment plan for U.S. AI infrastructure. Yet when Amazon announced its $200 billion plan, its stock price dropped 11%—the market was beginning to question the return on pure compute investment. The return curve for compute approximates a logarithmic function: scaling from 1,000 to 10,000 GPUs might improve model capability by 30%, but scaling from 10,000 to 100,000 might yield only an additional 10%. The bulk of the $310 billion is being spent on the tail end of this curve.

Algorithms: Contributing Roughly Half as Much as Compute, and Being Commoditized

Research by Ho et al. (2024) found that in the LLM domain, algorithmic progress contributes approximately half as much to performance improvement as compute scaling (in terms of log-scale contribution ratios). The rate of algorithmic efficiency improvement is catching up at roughly 3× per year. But the more critical issue is that virtually all algorithmic innovations of 2025–2026 were open-sourced by their own creators. DeepSeek open-sourced GRPO and MLA; Moonshot AI open-sourced MoonClip, KDA, and AttnRes. Algorithms are no longer a differentiating variable.

Core Data: The Highest Leverage and Irreplaceable

Synthesizing the evidence from LIMA, Phi-2, METR, and BC Protocol, expert data is the driving force with the highest leverage among the three, and the only one that cannot be substituted by the other two. More critically, elite expert data has three characteristics that make it impossible to scale with capital. First, tacit knowledge cannot be outsourced—the cognitive detours a mathematician takes, the intuitive debugging decisions of a world-class programmer, these cannot be codified in annotation guidelines. Second, quality does not scale linearly with quantity—spending 10× more to hire 10× more people will not produce 10× better data; it may actually degrade quality (as the METR experiment demonstrated). Third, organizational scale itself is the enemy of data quality—an 8,000-person company requires layers of management and standardized processes, yet standardization is the natural enemy of the “non-standard insights” that are the scarcest element in expert cognition.

Compute provides linear improvement. Algorithms provide a one-time leap (then get open-sourced). Core expert data provides irreplaceable generalization capability gains. The leverage ratios of these three are not even in the same order of magnitude. LIMA proved 1,000 > 50,000. Phi-2 proved 2.7B > 70B. METR proved that automation cannot replace human judgment.

4CEO R&D Involvement: The Primary Determinant of Organizational Efficiency

4.1 Liang Wenfeng (DeepSeek): CEO as Chief Researcher

Liang Wenfeng’s record of paper authorship is sustained and intensive: the NSA attention mechanism paper (February 2025, awarded ACL 2025 Best Paper, Liang as corresponding author), the mHC manifold-constrained hyperconnection paper, the paper on training techniques that break through GPU memory limits, and the DSpark speculative decoding paper (June 2026, with model weights and training code simultaneously open-sourced).

On July 23, 2026—just one day before this paper was written—Liang Wenfeng made the most critical quote cited in this paper during a four-hour investor meeting following DeepSeek’s inaugural 50-billion-yuan funding round:

“Solving the AI problem, at this current stage, comes down to data annotation. You could say that half of our core researchers—the most important people at the company—half of them are doing data annotation. That’s where we’re focused. Data annotation.”

This was not a conclusion read from a management report, but a judgment derived from the firsthand experience of a CEO who has personally annotated data and fine-tuned models. The contributor list of the DeepSeek-V3 technical report confirms this—150 R&D engineers and 31 data annotators are explicitly listed as paper contributors; annotators are invisible in most other AI papers.

4.2 Yang Zhilin (Moonshot AI/Kimi): CEO Personally Rewrites Core Components

In his GTC 2026 presentation, Yang Zhilin personally deconstructed three major technical pathways—designing replacement alternatives for each of the three foundational components of the Transformer architecture that have been in use for nearly a decade (the Adam optimizer, the attention mechanism, and residual connections): the MoonClip optimizer (enabling 20T of training data to achieve effects close to 40T), KDA (Kimi Delta Attention), and AttnRes (Attention Residuals, improving training and inference efficiency by 25%, with a 17-year-old student as one of the paper’s first authors). All three innovations were fully open-sourced.

Moonshot AI’s organizational structure is radically flat—based on estimates derived from funding scale, product release cadence, and recruiting activity, the team in July 2026 is likely in the 500–800 range (the company has not publicly disclosed its latest headcount; the publicly available 2025 figure was approximately 300), with no conventional departments, job levels, titles, OKRs, or KPIs. Interns can communicate directly with the founders. With this team size, they built Kimi K3—a 2.8-trillion-parameter model, the world’s largest open-weights model—which saw user request volume exceed projections within 72 hours of launch, forcing the company to temporarily suspend new membership sign-ups.

4.3 American CEOs: The Manager Model

Sam Altman’s public activities are concentrated on product launches, congressional testimony, fundraising negotiations, and geopolitics. Dario Amodei has stated that he spends 40% of his time on company culture. Mark Zuckerberg personally emails recruiting pitches but does not author papers. Alexandr Wang left the data company he founded to lead the superintelligence project at Meta. These CEOs possess deep technical backgrounds, but they are no longer on the front lines of model training.

4.4 An Information-Theoretic Explanation

Organizational Efficiency Formula of CEO R&D Involvement

CEO on the front line = 0 layers between “discovering a problem” and “making a decision” = Zero information loss

CEO detached from R&D = Researcher → Team lead → Director → VP → CEO = 5-layer reporting chain = Information decays at each layer

Information loss across a 5-layer chain ≈ 20–30% decay per layer → Only 20–30% of the original signal remains by the time it reaches the CEO

How an 8,000-person organization compensates for information loss = More people, more money, more redundancy to cover the same problem

When Liang Wenfeng says “half of our core researchers are doing data annotation,” he knows this because he is one of that half. When Altman says “we need high-quality data,” he knows this because someone told him in a briefing. The cognitive precision of the two is fundamentally different.

5Compensation Data as a Strategic Signal

5.1 Elite AI Researchers: From Millions to Hundreds of Millions in Three Years

In 2023, top-tier Google DeepMind researchers received offers of approximately $2 million per year. By 2024, top OpenAI researchers earned annual salaries exceeding $10 million, with retention bonuses above $20 million. In June 2025, Meta systematically poached from OpenAI with signing bonuses as high as $100 million and four-year total compensation packages reaching $300 million. The most extreme case: Zuckerberg reportedly extended a compensation package of up to $1.5 billion (over six-plus years) to a co-founder of Thinking Machines Lab—who declined. The AI skills salary premium surged from 25% in 2024 to 56% in 2025—the premium itself doubled within 12 months.

5.2 Data Production Experts: From Nonexistent to a 33–67× Pay Spread

In 2023, “AI data production expert” barely existed as a profession. By 2025–2026, a fully formed six-tier compensation pyramid had emerged:

Tier Role Hourly Rate
L1 General data annotator $15–25
L2 Language specialist $25–28
L3 Coding RLHF specialist $50–65
L4 Mercor platform average $85–95
L5 Medical fellow / Domain expert $250–450
L6 VC partner / C-suite executive $500–1,000

The pay spread from L1 to L6 is 33–67×. An outstanding “environment builder”—the role at the very top of the pyramid—does not merely label right or wrong, but constructs testing frameworks, designs edge cases, and writes scoring logic, enabling models to autonomously practice thousands of iterations. One such person is worth a hundred at the bottom tier.

5.3 The Three-Phase Transformation in Hiring Patterns

Over three years, hiring patterns underwent a clear structural transformation: in 2023, companies hired engineers to build systems; in 2024, they hired annotators to feed data (the quantity phase); in 2025–2026, they laid off general annotators and competed for elite experts (the quality phase). xAI’s layoff email is the most naked footnote to this entire trend—firing 500 general annotators while expanding expert mentor teams by 10×. They didn’t stop needing people; they stopped needing ordinary people and started needing only the smartest ones.

6Information Channel Asymmetry

6.1 The Silicon Valley Chinese Social Network

Silicon Valley is home to hundreds of thousands of Chinese engineers, distributed across core technical roles at OpenAI, Anthropic, Google, Apple, Meta, and others. They are interconnected through WeChat groups, alumni associations, churches, and school parent groups for their children, forming a cross-company, high-trust information mesh.

WeChat’s completeness as a super-app ecosystem confers a unique structural advantage upon this network. By contrast, South Korea’s KakaoTalk and Japan’s LINE are more oriented toward closed-circle personal messaging, lacking WeChat’s “semi-open social network” properties—features like “People Nearby,” group QR codes, and official accounts enable rapid ice-breaking between strangers. Chinese diaspora communities form “mesh structures” overseas (where potential pathways exist between any two nodes), while Korean and Japanese diaspora communities form “island structures” (each person is connected only to pre-existing small circles).

6.2 Unidirectional Information Transparency

American companies’ papers, products, and APIs are all public. Directional information is synchronized in near real-time through Chinese social networks. In contrast, Chinese companies’ internal R&D directions are unknowable to American firms until they are voluntarily open-sourced or published. DeepSeek’s catch-up strategy corroborates this—according to industry observations, while operating with approximately 1/20th the compute of American labs, DeepSeek compressed the technology gap to within 12–18 months through precisely targeted directional judgment. The prerequisite for such catch-up efficiency is: knowing which direction the other side is running.

6.3 Litigation Cases as Public Evidence

In August 2025, xAI sued former Chinese researcher Xuechen Li, alleging that prior to joining OpenAI, he had copied confidential Grok files, including complete source code, training methodologies, and future R&D roadmaps. xAI claimed this information could save competitors billions of dollars in R&D costs and years of effort. In September 2025, xAI escalated the lawsuit to directly sue OpenAI. During the summer of 2025, a total of eight xAI engineers and executives departed in succession to join OpenAI.

On July 10, 2026, Apple sued OpenAI, alleging that Chinese engineer Chang Liu carried confidential files about unreleased technologies upon leaving, and that OpenAI’s hardware lead Tang Tan had instructed Apple employees to share trade secrets during the interview process. At least 25 former Apple employees were working at OpenAI.

The protagonists of both lawsuits are entirely Chinese or Chinese-national engineers and executives. They were not peripheral employees but key figures with access to core technologies. Information flow represents a structural leak in the supply chain, not isolated incidents.

6.4 Data Backflow Through the Outsourcing Chain

Over the past decade, American AI companies have outsourced large volumes of annotation, testing, and even portions of R&D to global contractor networks, including Chinese firms and Chinese-national contractors. When annotation tasks are outsourced for execution, annotation guidelines, scoring criteria, data formats, and even distributional characteristics of the training data (out-of-distribution data) naturally flow into the cognitive domain of the contractors through the execution process. After Meta’s $14.3 billion acquisition of Scale AI, Google, OpenAI, and xAI immediately pulled their business—the intensity of this reaction itself demonstrates that data supply chain neutrality is an existential issue for frontier labs.

7The Talent Allocation Function: The Fundamental Variable of National AI Competitiveness

7.1 China: The Smartest People → CS and AI

A population base of 1.4 billion means that even if only one in ten thousand is a genius, the absolute count is 140,000. Over the past two decades, China’s economic incentive system has channeled the majority of these individuals toward software and AI. Peking University’s mathematics department, Tsinghua’s Yao Class, and USTC’s Junior Class—these elite talent pipelines are filled entirely with STEM candidates. Tencent, Alibaba, ByteDance, Baidu, and Huawei—China’s largest companies are all platform-type software enterprises, offering salary ceilings and social prestige high enough to attract top talent.

On July 23, 2026—again, just one day before this paper was written—the Fields Medal was announced in Philadelphia: two of the four recipients were Chinese mathematicians—Deng Yu (born 1989 in Shenzhen, who solved a restricted version of Hilbert’s Sixth Problem) and Wang Hong (born 1991 in Guilin, Guangxi, who cracked the three-dimensional Kakeya conjecture). Both were Peking University Class of 2007 undergraduates. This was the first time Chinese-national mathematicians had ever received the Fields Medal.

They and the core researchers at DeepSeek and Moonshot AI come from the same talent pool—the STEM systems of Peking University, Tsinghua University, and the University of Science and Technology of China.

7.2 The United States: The Global Talent Magnet

America’s AI competitiveness is built upon a unique siphoning mechanism: platform-type software giants (Apple, Google, Meta, Microsoft, Amazon) offer the world’s highest salary ceilings, attracting the planet’s best talent to Silicon Valley. Top OpenAI researchers earn over $10 million annually; Meta offers individual researchers signing bonuses of $100 million—no other location on Earth can provide this level of economic incentive.

However, this siphoning mechanism has a structural side effect: the talent it attracts includes a large number of Chinese and Indian-origin engineers. The directional information they acquire in core positions at American companies flows naturally toward the Chinese and Indian tech ecosystems through the social networks and outsourcing chains analyzed in Chapter 6 of this paper. While the United States siphons global talent, it is also inadvertently training its competitors’ information infrastructure. SignalFire data shows the probability of an OpenAI engineer defecting to Anthropic is 8× the reverse; from DeepMind to Anthropic, the ratio is 11:1—as talent flows at high velocity among American companies, it also accelerates the diffusion of information.

Furthermore, the generational replenishment of American AI talent is highly dependent on immigration. The proportion of H-1B visa holders in AI research positions is extremely high. Any tightening of immigration policy could weaken the efficiency of this siphoning mechanism—whereas China’s talent supply is entirely endogenous and independent of immigration.

7.3 South Korea: The Smartest People → Medical School

In 2025, KBS (Korean Broadcasting System) aired the documentary Insight: The Talent War, comparing China’s support for STEM talent against South Korea’s obsession with medical school. The data is stark: among the top 488 scorers on the Korean college entrance exam, every single one entered a medical-track program, with nearly 90% enrolling in medical school proper. The dean of Seoul National University’s College of Engineering revealed that approximately 100 STEM freshmen per year over the past five years have abandoned their enrollment to retake the exam and try for medical school instead. The “exodus” of elite STEM talent in South Korea is intensifying.

However, because South Korea’s domestic market of 51 million is too small, Samsung, Hyundai, and LG have been forced to orient toward global markets from the very start. This “forced globalization” pressure has at least allowed South Korea to develop second-tier platform companies such as Naver, Kakao, and Coupang, maintaining a baseline sensitivity to external developments.

7.4 Japan: The Self-Sufficiency Trap

Japan’s population of 125 million gives it a false sense of security—large enough for self-sufficiency, too small to birth global-scale platforms. Lifetime employment and seniority-based promotion systems have locked talent mobility in place. The most elite graduates flow toward the Ministry of Finance, METI (government bureaucracy), and general trading companies (Mitsubishi Corporation, Mitsui & Co.). English proficiency ranks near the bottom among developed nations. Japan has missed every wave of platform revolution: search, social media, mobile payments, short-form video, and AI.

7.5 Europe: The Training Ground and Exporter of Global AI Talent

Europe can cultivate top-tier talent—multiple Transformer co-authors have European backgrounds, and DeepMind was born in London—but it cannot retain them. €150,000–€200,000 annually in Paris vs. over $1 million in Silicon Valley. The EU AI Act further increases compliance costs. No unified language, no platform-type giants, no information mesh. Europe’s role in the AI era has solidified as “talent training ground and exporter”—analogous to South America’s relationship with European football leagues.

7.6 India: AI Is Destroying Its Core Industry

India presents the most brutal case. IT outsourcing (Infosys, TCS, Wipro) and customer service BPO—contributing approximately 10% of GDP and employing millions—are precisely the sectors that AI is replacing first and most completely. AI startup Lindy migrated 100% of its traffic from Anthropic to DeepSeek to cut costs—by the same logic, enterprises will migrate outsourcing to AI to cut costs. The best IIT graduates all flow to the United States. A self-reinforcing negative feedback loop is forming: India exports talent → talent builds AI in the United States → AI replaces Indian outsourcing industries → India’s economic engine stalls → more talent flees.

8Industrial Structure Determines Talent Flow

8.1 The United States and China: Dominated by Platform-Type Software Giants

America’s most valuable companies—Apple, Microsoft, NVIDIA, Google, Amazon, and Meta—are all software/platform-type enterprises. China’s most influential companies—Tencent, Alibaba, ByteDance, Meituan, Pinduoduo, Baidu, and Huawei—are likewise all platform/software or integrated hardware-software enterprises. In these companies, software engineers are core profit generators, and therefore companies are willing to offer engineers the highest compensation.

More importantly, these platform giants have accumulated massive reserves of data, engineering infrastructure, and algorithmic talent, forming the incubation bed for AI startups. China’s AI entrepreneurs have almost universally emerged from the incubation layer of platform companies—Liang Wenfeng was in quantitative trading (algorithmic in nature), Yang Zhilin interned at Google Brain and Meta AI, and ByteDance’s Seed team grew directly from the algorithmic DNA of the Douyin (TikTok) recommendation system. Without platform-type giants, this incubation layer would not exist.

8.2 South Korea: Dominated by Heavy Industry and Manufacturing

South Korea’s most valuable companies—Samsung Electronics, SK Hynix, Hyundai Motor, LG, POSCO, and Hanwha—are all heavy industrial and manufacturing firms. In these companies, software engineers are cost centers, not profit centers. Core competitiveness lies in semiconductor fabrication processes, automobile assembly, steelmaking, and battery chemistry—not algorithms and models.

This results in software engineer salary ceilings far below those of physicians and lawyers. The most elite talent does the math—a career in medicine yields annual income of hundreds of millions of won with the highest social status, while writing code at Samsung for a lifetime cannot reach that level. Moreover, South Korean society insufficiently protects the software industry—client-side price suppression, an outsourcing culture, and inadequate IP protection—all send the signal to young people that “writing code isn’t worth much.”

8.3 The Causal Chain: Industrial Structure → Talent Flow

Dominant National Industry Type → Software Engineer Salary Ceiling & Social Status → Direction of Smartest Talent Flow → AI Talent Pool Size → AI Competitiveness
Country Dominant Industry Software Engineer Status Smartest Talent Flows To AI Competitiveness
United States Platform-type software giants Core profit generator CS / AI First tier
China Platform-type software giants High pay + rising social status CS / AI First tier
South Korea Heavy industry / Manufacturing Cost center Medicine / Law Third tier
Japan Manufacturing + Bureaucracy Support role Government / Trading firms Fourth tier
Europe Traditional industry + Finance Moderate, but cannot retain Flows to the United States Talent exporter
India IT outsourcing (being replaced by AI) Outsourcing execution layer Flows to the United States Disrupted by AI blowback

AI competition is not a race one can “invest one’s way into.” It is a civilization-level contest requiring simultaneous alignment across a nation’s population base, industrial structure, talent orientation, social infrastructure, and cultural values. The United States and China are the only two nations on Earth aligned across all dimensions—one through market mechanisms, the other through state will combined with market mechanisms.

9The Iceberg Structure of Chinese AI

9.1 Above the Waterline: The Seemingly Insurmountable Capital Expenditure Gap

$13.4 billion vs. $310 billion—a surface-level gap of 23×. If one looks at this number alone, China appears to have no chance whatsoever.

9.2 Below the Waterline: Irreplicable Structural Advantages

The true foundation beneath the surface consists of seven layers: the world’s largest pool of elite mathematics and CS talent (at Fields Medal–caliber density); extreme organizational efficiency with CEOs personally on the front lines (zero-layer reporting chains); a cost structure where researchers of comparable quality cost 1/5 to 1/10; a bidirectional information channel formed by Chinese-language social networks and the Silicon Valley Chinese diaspora network; tacit knowledge assets accumulated through years of outsourcing chains (out-of-distribution data and annotation know-how); open source as a strategic weapon (undermining competitors’ pricing power and harvesting contributions from the global community); and a complete manufacturing supply chain (OpenAI’s hardware manufacturer of choice is Luxshare Precision and GoerTek—both Chinese companies).

9.3 Logarithmic Curve vs. Exponential Curve

The $310 billion sits on the tail end of a logarithmic return curve for compute—the more you spend, the less incremental gain you receive. The $13.4 billion sits at the starting point of an exponential curve for talent density—the higher the talent density, the faster the information flow, and the leaner the organization, the greater the output efficiency. $310 billion on the tail of a logarithmic curve cannot outperform $13.4 billion at the starting point of an exponential curve. This is not abnormal. This is mathematics.

10Alternative Explanations and Limitations

10.1 The Causal Direction Problem of CEO R&D Involvement

This paper argues that “CEOs personally doing R&D → stronger companies.” However, an equally plausible alternative explanation reverses the causal direction: “the company is still small → the CEO has no choice but to do R&D.” Liang Wenfeng and Yang Zhilin personally writing papers and annotating data may not be because this is the optimal strategy, but because the company has only a few hundred people, lacks sufficient middle management layers, and the CEO is physically compelled to participate in front-line work.

This alternative explanation raises a critical temporal risk: DeepSeek has just raised 50 billion yuan and announced comprehensive hiring expansion, while Moonshot AI grew from approximately 300 people to an estimated 500–800 within four months. If both companies balloon to 3,000–5,000 people within two years, will the CEO still be personally annotating data? If not, does this paper’s core thesis still hold?

Our response is that CEO R&D involvement may indeed be an inevitable byproduct of the small-team phase, but this does not diminish its effectiveness as a competitive advantage. The key question is not “can the CEO do R&D forever,” but “during the window in which the CEO is still personally doing R&D, can the small team accumulate enough data assets and organizational knowledge to establish a durable advantage?” Liang Wenfeng mentioned in his investor meeting the idea of “using the next version of the model we train to assist our own R&D and data annotation”—creating a flywheel where AI helps humans annotate data—which may be a transitional solution for the CEO’s eventual withdrawal from the front line as the organization scales. Whether this approach can succeed, however, remains an untested hypothesis.

10.2 Unaccounted Advantages of Closed-Source Models

This paper focuses on benchmark score comparisons of model capability, but closed-source models possess several advantages that open-source models have not yet demonstrated. First is enterprise-grade reliability and SLA guarantees—Claude and GPT operate in the production environments of major enterprise clients, backed by mature security audits, compliance certifications, and service guarantee systems. Second is the depth of safety alignment—Anthropic’s investment in Constitutional AI and safety research may be invisible in benchmark scores but constitutes a substantive differentiator in high-risk application scenarios. Third is test-time compute scaling—OpenAI’s o-series models improve complex reasoning capability by increasing computation at inference time, and the ceiling of this pathway remains unclear.

These advantages do not alter this paper’s core thesis (talent density determines the baseline level of model capability), but they do mean that “the benchmark gap is only 3 points” may understate the true lead of closed-source models along certain dimensions.

10.3 Internal Differences Between DeepSeek and Kimi

This paper groups DeepSeek and Moonshot AI together as “Chinese open-source models” to contrast with American closed-source models, but the performance gap between the two is not trivial. Kimi K3 scores 57.11 on the Intelligence Index (close to the top two), while DeepSeek V4 Pro scores approximately 52 (a larger gap). Their strengths also diverge—DeepSeek’s core advantage lies in extreme cost efficiency (API pricing 57× cheaper) and coding benchmarks (LiveCodeBench 93.5, Codeforces Elo 3206), while Kimi K3 excels in comprehensive intelligence and agent tasks. Merging the two as “$13.4 billion” to compare against “$310 billion” may oversimplify the true complexity of the competitive landscape.

10.4 Inferential Boundaries of the Information Channel Argument

The argument in Chapter 6 regarding information channel asymmetry requires distinguishing between two levels. The first level—”talent mobility between American companies causes information diffusion”—is supported by litigation filings (xAI v. Li, Apple v. OpenAI) as public evidence, and the inferential strength is high. The second level—”information flows from the United States to China through Chinese social networks”—is a reasonable inference based on network structure, but lacks direct evidence. This paper’s authors conducted a more detailed analysis in a confidential research report dated February 2026 (Reference [28]), though that report itself was also characterized as inferential rather than empirical in nature. Readers should note the difference in argumentative strength between these two levels.

11Conclusion

11.1 The Inevitability of a US–China Duopoly

AI competition requires the simultaneous satisfaction of seven conditions: sufficient population base to supply aggregate talent volume; a platform-type software industry providing salary ceilings and accumulated technical expertise; an education system and social incentives channeling the smartest talent toward CS; social infrastructure connecting dispersed overseas talent into an information network; a manufacturing supply chain supporting the full loop from software to hardware; years of data ecosystem accumulation providing tacit knowledge assets; and an open-source culture harvesting free contributions from the global community. Only the United States and China are aligned across all dimensions simultaneously. Europe, Japan, South Korea, and India each satisfy one or two conditions at most.

11.2 CEO R&D Capability Is the Primary Variable

CEO personally doing R&D → Zero information loss → Highest data quality → Maximum output per person. This is the formula that Liang Wenfeng and Yang Zhilin have proven through practice. CEO detached from R&D → Five-layer reporting chain → Information decays at each layer → Money and headcount compensate for efficiency losses. This is the structural predicament facing Altman and Amodei.

11.3 The Individual Researcher as Paradigm Proof

This paper’s argument derives not only from macro-level analysis of industry data, but also from a micro-level living case study. LEECHO Global AI Research Lab is a single-person independent AI research institution that produced 209 trilingual (Chinese, Korean, English) research papers over 119 days from February 5 to June 4, 2026. The coverage spans from pure mathematics (Galois theory, the BSD conjecture, CDC obstacle theory) to AI architecture (token dimensionality reduction, dense kernel vs. MoE, TGI Scanner) to geopolitics (analysis of the Iran conflict, reshaping of U.S. naval hegemony in the 10-carrier era) to cognitive science (perception and cognition, cognitive MoE-ification) to industrial economics (AI structural cost crisis, deceleration of the closed-source commercial flywheel).

During the same period, this researcher independently developed LiteClaw Desktop (a local AI agent desktop system with 60 tools and a loop-agent architecture), ATM Scanner (a security scanning algorithm, including live-fire range testing), and TGI Scanner (a training ghost ideal detection tool, spanning the full chain from ring-theoretic mathematics to engineering implementation), and continuously collaborated with Claude Fable 5 on the CDC Obstacle Theory Ontology Observatory—an interactive mathematical verification system with seven layers, 75 fields, and 13 falsifications.

119 days, 209 papers, 3 languages, 3 software projects, 1 frontier mathematics collaboration project, 1 person. These numbers prove the individual-scale version of this paper’s core formula: the smartest mind + the right AI tools = asymmetric output capability. No 8,000-person team required. No $180 billion in funding required. No tens of thousands of GPUs required—what is needed is sufficiently high cognitive density, cross-disciplinary knowledge coverage, multilingual information channels, and the execution discipline of 1.76 papers per day.

11.4 The Final Formula

The decisive factor in the AI race is not “who has more GPUs,” but “who can extract the highest-quality thinking from the smartest human brains.” This capability cannot be bought, cannot be fast-tracked, and cannot be scaled. It is the product of history, culture, education systems, and industrial structures shaped over decades.

Compute can be purchased. Algorithms can be open-sourced. But the complete thinking processes of the smartest humans solving the hardest problems—which can be neither crawled, nor synthesized, nor scaled—this is the ultimate scarce resource of the 2026 AI race. Whoever possesses the most people among “the smartest minds working as independent researchers producing the highest-quality data for models” possesses the ultimate weapon of the AI race.

CEO R&D Capability → Talent Density → Data Quality → Model Capability
Not Capital → Not Compute → Not Algorithms → Not Team Size

12References

[1] Artificial Analysis. (2026, July). Intelligence Index: Claude Fable 5 vs GPT-5.6 Sol vs Kimi K3 benchmark comparison.

[2] Zhou, C. et al. (2023). LIMA: Less Is More for Alignment. NeurIPS 2023.

[3] Li, Y. et al. (2023). Textbooks Are All You Need: Phi-1/Phi-2. Microsoft Research.

[4] DeepSeek. (2025, January). DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning. arXiv.

[5] Feng, Y. et al. (2025). ReTool: Reinforcement Learning for Strategic Tool Use in LLMs. ICLR 2026.

[6] Ho, A. et al. (2024). Algorithmic Progress in Language Models. arXiv.

[7] Kang, D. (2024). Data labeling cost growth 88x vs compute 1.3x. Medium.

[8] 36Kr. (2026, July 23). Liang Wenfeng investor meeting transcript: Four-hour complete record.

[9] Moonshot AI. (2026, July 17). Kimi K3 Technical Report: 2.8T Parameters, 100M Context Window.

[10] Meta. (2025, June). Scale AI acquisition: $14.3B for 49% stake.

[11] xAI. (2025, September). Restructuring announcement: 500 annotator layoffs, 10x expert team expansion.

[12] xAI v. Xuechen Li. (2025, August 29). U.S. District Court, Northern District of California.

[13] Apple v. OpenAI. (2026, July 10). Trade secret misappropriation complaint.

[14] SignalFire. (2025). AI talent flow analysis: OpenAI-Anthropic-DeepMind migration patterns.

[15] Prolific. (2026, May). Survey of 300+ AI practitioners: Human signal as bottleneck.

[16] METR. (2026, March). AI agent code patch rejection study.

[17] AfterQuery. (2026). Expert data production economics: $1B/year per frontier lab.

[18] Mercor. (2026, June). $2B+ ARR announcement; C-round at $10B valuation (October 2025).

[19] DATA Foundation. (2026, June 25). Launch announcement: Solving the multi-billion-dollar training data bottleneck.

[20] PwC. (2025). Global AI Jobs Barometer: 56% AI skills wage premium.

[21] Fields Medal Committee. (2026, July 23). 2026 Fields Medal Recipients: Deng Yu, Wang Hong.

[22] KBS. (2025). Insight: The Talent War, Episode 1: China’s STEM talent system.

[23] Anthropic. (2026). Series H: $65B round at $965B valuation. Crunchbase.

[24] OpenAI. (2026). Series F: $122B round at $852B valuation. Crunchbase.

[25] DeepSeek. (2026, June). Inaugural funding round: 50 billion yuan.

[26] Moonshot AI. (2026). ARR $300M (June 2026), valuation $31.5B.

[27] BC Protocol. (2026). Expert CoT data as LLM post-training bottleneck.

[28] LEECHO Global AI Research Lab. (2026, February). OOD data leakage and the rise of Chinese AI. Confidential Research Report.

[29] LIMO. (2025). Less Is More for Reasoning: Extending LIMA to Mathematical Problem Solving. COLM 2025.

[30] USCIS / National Foundation for American Policy. (2025). H-1B visa holders in AI research positions: Demographic analysis.

Appendices

Appendix A: Timeline of Key AI Industry Events, 2023–2026

Date Event Significance
2023 Public internet pre-training data reaches saturation Bottleneck shifts from data quantity to data quality
2023 LIMA experiment published (NeurIPS) Proves 1,000 expert samples > 50,000 crowdsourced samples
2023 Phi-1/Phi-2 released (Microsoft) 2.7B parameters match 25× larger model; data quality overturns scaling laws
2024 Data annotation cost growth 88× vs. compute 1.3× Data costs surpass compute costs
May 2024 Scale AI closes $1B Series F at $14B valuation Data company enters super-unicorn territory
Jan 2025 DeepSeek R1 released Exposes pure RL ceiling; cold-start data becomes critical
Apr 2025 ReTool framework published RLVR + tool-calling paradigm established
Jun 2025 Meta acquires 49% of Scale AI for $14.3B Data company valued at $29B; industry earthquake
Sep 2025 xAI lays off 500 annotators, expands expert team 10× Industry watershed: “quantity” to “quality”
Oct 2025 Mercor closes Series C at $10B valuation Expert data platform valuation explosion
Feb 2026 Mercor acquires Sepal AI RL environment construction becomes a strategic asset
Apr 2026 Anthropic Series H: $65 billion Largest single funding round in history
Jun 2026 DeepSeek inaugural funding: 50 billion yuan Milestone for independent Chinese AI funding
Jun 2026 Mercor annualized revenue exceeds $2 billion Expert data market in full eruption
Jul 17, 2026 Kimi K3 released (2.8T parameters) World’s largest open-weights model; #1 in Coding Arena
Jul 23, 2026 Fields Medal announced: Deng Yu and Wang Hong Landmark event for the depth of China’s math talent pool
Jul 23, 2026 Liang Wenfeng investor meeting: “Half our researchers are doing data annotation” CEO personally confirms data as the core bottleneck

Appendix B: US–China AI Company Funding / Headcount / Valuation Comparison

Company Cumulative Funding Valuation Headcount Flagship Model
OpenAI ~$180B $852B ~8,000 GPT-5.6 Sol
Anthropic ~$132B $965B ~5,000 Claude Fable 5
DeepSeek ~$7.4B $52–59B Actively hiring DeepSeek V4 Pro
Moonshot AI ~$6B $31.5B ~500–800* Kimi K3

* Moonshot AI’s team size is estimated based on funding scale, product release cadence, and recruiting activity. The company has not publicly disclosed its latest 2026 figure.

Appendix C: Expert Data Compensation Pyramid (2026)

Tier Role Hourly Rate (USD) Annualized Equivalent Characteristics
L1 General data annotator $15–25 $31K–52K Basic classification/labeling tasks
L2 Language specialist $25–28 $52K–58K Multilingual, semantic understanding
L3 Coding RLHF specialist $50–65 $104K–135K Code evaluation, reward model annotation
L4 Mercor platform average $85–95 $177K–198K Cross-domain technical expert
L5 Medical fellow / Domain expert $250–450 $520K–936K Clinical/research-grade professional judgment
L6 VC partner / C-suite / Environment builder $500–1,000 $1M–2M Builds RL environments, designs scoring frameworks

The pay spread from L1 to L6 is 33–67×. A single L6 environment builder can produce value equivalent to approximately 100 L1 annotators.

Appendix D: Kimi K3 vs. GPT-5.6 Sol vs. Claude Fable 5 Benchmark Comparison

Metric Claude Fable 5 GPT-5.6 Sol Kimi K3
Intelligence Index 59.86 58.89 57.11
Frontend Code Arena (Elo) 1631 1618 1679 (#1)
Program Bench 77.8 (#1)
SWE Marathon 42.0 (#1)
BrowseComp 91.2 (#1)
API Price ($/M output tokens) $50 $30 $15
Open Weights No No Yes

Kimi K3 won 5 out of 6 real-world agent benchmarks. DeepSeek V4 Pro is not included in this table; its Intelligence Index is approximately 52, but its API price is only $0.87/M output tokens (57× cheaper than Fable 5).

Appendix E: 2026 Fields Medal Recipients and Their Connection to the AI Talent Pool

Recipient Nationality Born Undergraduate Core Achievement
Deng Yu China 1989, Shenzhen Peking University, Class of 2007 Restricted version of Hilbert’s Sixth Problem (first breakthrough in 125 years)
Wang Hong China 1991, Guilin, Guangxi Peking University, Class of 2007 Three-dimensional Kakeya conjecture (century-old problem)

Both recipients were Peking University Class of 2007 undergraduate classmates. This was the first time Chinese-national mathematicians received the Fields Medal, with China accounting for 2 of 4 recipients. They share the same talent pool as the core researchers at DeepSeek and Moonshot AI—the STEM pipeline of Peking University, Tsinghua University, and USTC. 2006 IMO Gold Medalist Deng Yu → Fields Medal; Tsinghua Yao Class / PKU Mathematics → AI researcher—the same cohort, different directions, the same caliber.

© 2026 LEECHO Global AI Research Lab · Original Thought Paper · All Rights Reserved

댓글 남기기