▶ 20-min Slideshow ✦ What's new
Davidovs
Venture
Collective

DVC's
STATE OF AI

The Operating Manual for the AI Revolution

The biggest industrial shift in modern history is already underway, and if you cannot see the full stack, you cannot see where the world is heading.

Scroll to explore
davidovs.com

FOLLOW THE MONEY

How $60B of AI revenue generates $700B of infrastructure investment

MARGIN GRADIENT
25–60%Apps
33–70%Models
14–47%Cloud
50–75%Silicon
UtilityEnergy
Where does $1 of AI revenue go?
$0.30–0.60 Inference costs (AI Tax)
$0.18–0.36 Cloud compute (~60% of inference)
$0.11–0.25 GPU/server purchases (60–70% of CapEx)
$0.06–0.19 Silicon margin (NVIDIA ~50–75% GM)
$0.02–0.04 Energy (power + cooling)
Click any layer to explore
LAYER 5

APPLICATIONS

$12-15B ARR GM: 25–60%
AI Tax: 30–75% consumed by inference
OpenAI $25B L5+L4
Anthropic $30B+ L5+L4
Cursor $2B+
Salesforce AI $800M
Grammarly $700M
GitHub Copilot $500M+
Lovable $400M
ElevenLabs $600M L5+L4
Perplexity >$450M DVC
Higgsfield >$500M DVC
Suno $300M
Runway ~$150M L5+L4
Replit $253M
Traditional SaaS70–85% GMBenchmark
Perplexity50%+ GMSearch + subscription
OpenAIGM est. variesNot officially disclosed
Cursor~35% GMHeavy inference load
Higgsfield30%+ GMVideo generation — DVC estimate
Lovable20–40% GMVibe-coding, scaling fast
GitHub Copilot~0–15% GMSubsidized for lock-in
$0.30–0.60 → inference
LAYER 4

FOUNDATION MODELS

$8-10B API GM: 33–70%

API Market Share (Menlo Ventures, end-2025)

Anthropic40%
OpenAI27%
Google21%
Others12%
OpenAIGM est. varies-$6.9B H1 loss
AnthropicGM est. variesNot officially disclosed
Google GeminiNot disclosedServing costs -78%
CohereGM ~70%Best in class
xAINet loss 13.6x rev
~60% → cloud compute
LAYER 3

CLOUD & INFRA

$18-22B rev $224B CapEx (2024) 23:1 CapEx-to-AI-Revenue
AWS37.0% OM$39.8B / $107.6B (9M 2025)
Microsoft IC47.1% OM$49.6B / $105.4BCloud GM 71%
Google Cloud14.1% OM$6.1B / $43.2B→ 17.5%
Meta42.0% OM(total company)
CoreWeave72% GM-0.9% OM$5.1B rev
Nebius$68M rev$17.4B MSFT contract

GPU Server Economics ($300K server, $150K annual rev, $45K cash opex)

Depreciation$75K
Profit$30K
Margin20.0%
Breakeven Util.~80%

Hyperscaler Depreciation Changes

Microsoft4yr → 6yrSaved $3.7B
Alphabet4yr → 6yrSaved $3.9B
Amazon4→5→6yr, reversed subset to 5yr+$889M hit
Meta→ 5.5yrSaving ~$2.9B
CoreWeave5yr → 6yrSaved $20M

CapEx Breakdown

GPUs/Servers60–70% (~$350B)
Buildings/Power25–35% (~$170B)
Networking5–10% (~$40B)
2024: $224B 2025: ~$402–410B 2026E: ~$725B mid / up to ~$745B
60–70% of CapEx → GPUs
LAYER 2

SILICON

$130B rev GM: 50–75%
NVIDIA $215.9B FY26
GM 71.1% OP 60.4% Net 55.6%
DC rev $193.7B $1T+ orders through 2027
Broadcom $63.9B FY25
GM 67.8% OP 39.9% EBITDA 68%
AMD $34.6B FY25
GM 49.5% OP 10.7%
DC $16.6B
Arista $9.0B FY25
GM 63.7% OP 42.5% Net 39.0%
Marvell $8.2B FY26
GM 51.0% OP 16.1% Net 32.6%

Custom Silicon

Google TPU AWS Trainium Meta MTIA Microsoft Maia Etched

NVIDIA INFERENCE ARCHITECTURE — GTC 2026

Vera Rubin + Groq LPU

72 Rubin servers + 256 LPU chips = 700M tokens/sec — 350× Hopper throughput. Purpose-built for inference.

NVIDIA Roadmap

Blackwell (now) → Vera Rubin (2026, HBM4) → Feynman (2028, TSMC 1.6nm, silicon photonics)

Source: NVIDIA GTC 2026 / WSJ / CNBC / Axios, Mar 2026
580 TWh US high-case by 2028
LAYER 1

ENERGY

460 TWh
Nuclear
🔥Gas
Solar

Who's Powering AI

Microsoft Three Mile Island PPA 835 MW, 20-yr PPA with Constellation for a targeted Unit 1 restart, ~$16B, online 2028
Google SMR fleet + plant restart 500 MW Kairos Power SMR fleet (first reactor 2030); $1.6B Duane Arnold restart with NextEra; 1.8 GW pipeline via Elementl Power
Amazon $20B+ nuclear investment 960 MW Susquehanna campus; X-energy SMR development; 1,920 MW Talen PPA through 2042; Energy Northwest 960 MW SMR project
Meta 1–4 GW nuclear RFP Request for proposals targeting new nuclear generation, early 2030s; 6.6 GW total nuclear projects announced Jan 2026
Oracle 3× SMR-powered data center GW-scale campus powered by three SMRs; building permits secured
Big tech contracted 10 GW+ of new US nuclear capacity in 2025. Data center electricity: 460 TWh (2024) → 1,000 TWh (2030) → 1,300 TWh (2035).
Scroll to explore

THE $60B EXPLOSION

Updated Aug 4, 2026

The leadership order inverted — and the best metric changed

Anthropic closed a $65B Series H at a $965B post-money valuation on May 28, 2026, roughly tripling its $380B February mark and passing OpenAI for the first time. More importantly for this report, it also passed OpenAI on US enterprise penetration — which is a better leading indicator than either valuation or a leaked run-rate.

Anthropic $965B primary — $65B Series H closed May 28, 2026.
$47B run-rate reported June 2026, up from the company’s own $30B disclosure in April — sequential points on a steep curve ($14B Feb → $19B Mar → $30B Apr → $47B Jun), not conflicting estimates.
42.4% of US businesses pay for Anthropic vs OpenAI’s 39.5% — a 2.9-point lead, with overall business AI adoption at 46.6% (Ramp AI Index, July 2026).
300,000+ business customers, 1,000+ spending >$1M/yr, ~70% of the Fortune 100 — all company-disclosed.
OpenAI $852B primary · ~$880–895B implied by secondary prints (Forge Global derived price, late June 2026).
ChatGPT’s mobile apps crossed 1B MAU in May 2026 — a Sensor Tower estimate covering iOS and Android only, not web, API or enterprise. OpenAI’s own last disclosure is 900M weekly actives (Feb 2026).
Codex + ChatGPT Work at ~10M WAU (Jul 21), OpenAI-disclosed; “active” is undefined.
Run-rate has held near ~$25B from February through mid-2026.
Google 950M Gemini app MAU (Jul 22), up from 750M in February, with daily actives tripling YoY.
AI Mode in Search passed 1B MAU.
~90% of the Fortune 100 use Gemini Enterprise; developers process 22B tokens per minute through the APIs; 90M+ custom Gems created.
Meta 1.2B Meta AI MAU / 800M WAU across Facebook, Instagram and WhatsApp — nominally the largest AI assistant by reach.
Read it carefully: this is embedded reach inside existing search bars, not an assistant users chose. It is a distribution number, not a demand number.
Audit-quality caveat. No private lab has filed audited 2026 revenue. OpenAI’s most recent revenue signal is a partial internal transcript reviewed by CNBC on Jul 29 in which CFO Sarah Friar told staff that July’s annualized run-rate exceeded all of Q2 2026 — with no dollar figures. Anthropic’s run-rate is reported at both ~$47B and “$30B+” by different sources in the same window. Every private ARR and run-rate figure in this report is company-disclosed and unaudited unless explicitly noted.
Sources: Reuters (Series H) · CNBC (valuations) · Ramp AI Index, July 2026 · Yahoo Finance ($47B run-rate) · Reuters (1B mobile MAU) · OpenAI (round, 900M WAU) · The Verge (Gemini) · audit caveat is DVC’s own analysis
~$35B flows to inference & infra 30–75% of app revenue is consumed by compute
~$5–8B pure application layer Excluding OpenAI, Anthropic, Suno, Runway, ElevenLabs
~$1.1B Suno + Runway + ElevenLabs Generative media: music, video, voice

The application layer is the fastest-growing segment of AI revenue.

Historically, in every major computing wave — mainframes, PCs, mobile, cloud — the application layer is where most value ultimately accrues. Infrastructure enables, but applications capture.

Source: company disclosures, press reports, DVC analysis. “Pure application layer” = companies with own distribution, excluding model providers and generative media.

HOW FAST IS $60B?

AI Apps
~3 yrs
Cloud SaaS
~8 yrs
Mobile Apps
~10 yrs
Enterprise SW
~15 yrs

Time for each software category to reach ~$60B in combined application revenue

Source: Bessemer Cloud Index, Menlo Ventures, DVC analysis

THE APPLICATION LAYER

Top AI application companies by annualized recurring revenue

OpenAI $25B ● mixed
Claude $30B+ ● mixed
Cursor $2B+ ● end-user
GitHub Copilot $2B+ ● end-user
Salesforce AI $800M ● enterprise
Grammarly $700M ● end-user
Lovable $400M ● end-user
ElevenLabs $600M ● API-heavy
Perplexity >$450M DVC ● end-user
Higgsfield >$500M DVC ● end-user
Suno $300M ● end-user
Runway ~$150M ● mixed
Replit $253M ● end-user

API-heavy = Most revenue from API/platform  |  End-user = Most revenue from consumers/enterprise seats

The application layer is becoming a behavior layer. Revenue ranks one thing; usage habits rank another. ChatGPT still anchors the category — 244M desktop conversations in March 2026, up 55% year over year — but the scarce asset is the default workflow and the context that accumulates inside it, not the model behind the surface.

Claude is the breakout: 22M desktop conversations in March 2026, up 1,858% versus October 2025. And the habit is spreading across surfaces, not consolidating into one default — AI assistant tools reached 36% of desktop users and 23% of mobile users in Q1 2026. Meanwhile, AI is reshaping the surfaces beneath it: Google’s AI Overviews appeared alongside 46% of paid search ads in consumer credit cards in Q4 2025, up from 21% in Q2 2025.

Source: Comscore Q1 2026 AI Intelligence Report

Capital is flooding every layer, but one battleground shapes the rest. To understand power in AI, you have to understand the fight over the models themselves.

THE MODEL WARS

THE MODEL MARKET BECAME A BARBELL. Premium multi-day agentic tiers priced up to $10/$50, capable output priced down to $1–$6, and a third axis — token efficiency — that no price-per-token chart captures. 10+ companies at the frontier simultaneously. The bottom is commoditizing; the top is getting more expensive and more strategic. The default ChatGPT tier has quietly moved up to GPT-5.5 Instant, so the consumer baseline now sits well above 2025 frontier on AIME and MMMU-Pro. Open-weight models look close to proprietary ones on public benchmarks; held-out government evals still show a gap. Cost efficiency, especially from PRC labs, is real. And the benchmark table is a weaker moat signal than it was: with Gemini 3.5 Flash and Claude Opus 4.8, distribution, runtime, and governed tool access now matter as much as raw capability — the contest is model + runtime + distribution, not a leaderboard alone.

Updated Aug 4, 2026

The model market became a barbell

We are retiring the “Frontier of Restraint” framing. Prices moved in both directions in a single quarter. Anthropic pushed the top end up to $10/$50 for multi-day agentic work while Sonnet 5 and GPT-5.6 Luna pulled capable output toward commodity pricing. The middle thinned, and vendors started competing on tokens per task rather than only price per token.

Premium agentic — the top got more expensive $10 / $50 Claude Fable 5 (Jun 9) — marketed on multi-day agentic task execution; current SWE-Bench Pro leader at 80.3%. GPT-5.6 Sol sits below it at $5/$30 with Ultra mode and Max reasoning. Mythos 5 is the same weights with safeguards lifted, restricted to Project Glasswing partners — it is not a purchasable price point. Fable 5 also falls back to Opus 4.8 on security-adjacent prompts and carries mandatory 30-day data retention with no zero-retention option.
Efficient frontier — capable output at commodity prices $1 – $6 GPT-5.6 Luna at $1/$6 defines the floor in the sourced price table. Sonnet 5 runs a promo $2/$10 through Aug 31, then $3/$15. Grok and Gemini remain strategically important, but their exact pricing and context figures were removed when they could not be confirmed against primary pricing pages.
Anthropic Fable 5 / Mythos 5 $10/$50 (Jun 9) · Sonnet 5 $2/$10 promo → $3/$15 (Jun 30), 1M context, default for Free/Pro and Claude Code from Jul 1.
OpenAI GPT-5.6 Sol / Terra / Luna GA Jul 9 at $5/$30, $2.50/$15, $1/$6. Sol is 54% more token-efficient on coding. GPT-5.4 retired Jul 23; o3 leaves ChatGPT Aug 26.
SpaceXAI Grok 4.5 was positioned in the frontier-adjacent tier. We do not print its launch date, list price, context window, training-partner claim or index rank: none could be sourced to xAI or a primary price table. The integrity note stands — the CursorBench result was withdrawn after a Cursor codebase snapshot entered training data.
Meta Muse Spark 1.1 Jul 9 — Meta’s first paid Model API at $1.25/$4.25, 1M context, access limited to Meta properties. Leads on tool use; loses on raw coding and long context.
Google Gemini 3.6 Flash shipped Jul 21 as Google’s cheap-and-capable lane, with Google claiming fewer output tokens per task than its predecessor. We do not print the exact per-token prices, the output-token reduction or the knowledge cutoff: none could be confirmed against Google’s own API pricing page. Gemini 3.5 Pro remains in partner testing.
Restricted / cyber Cyber capability became a release gate. Gemini 3.5 Flash Cyber shipped as a dedicated SKU; GPT-5.6 was billed as OpenAI’s “strongest cybersecurity model yet”; Sonnet 5’s restoration was gated on new cyber classifiers.
Benchmark integrity. The Grok 4.5 CursorBench result was withdrawn after Cursor disclosed that a snapshot of its codebase had accidentally entered the model’s training data. Single-benchmark leadership claims in this window should be discounted accordingly.
Sources: Anthropic newsroom · OpenAI GPT-5.6 · TechCrunch · The Decoder · APIdog benchmarks · Meta AI · Reuters · How AI Works · Roo (CursorBench withdrawal) · OpenAI release notes

PROPRIETARY FRONTIER — grid below reflects the March–May 2026 snapshot; the current July 2026 flagship set is above

OpenAI

GPT-5.4 / o3

$25B ARR

OpenAI

GPT-5.5

Apr 24, 2026

Anthropic

Claude Opus 4.8

#1 Chatbot Arena

Anthropic

Mythos

Enterprise preview

Google

Gemini 3.5 Flash

#2 Chatbot Arena

xAI (SpaceX)

Grok 3

$1.25T merged

Amazon

Nova 2 Pro

Top reasoning on Bedrock

Meta

Muse Spark

Closed · 3.6B DAU

New Aug 3, 2026

China’s open-weight frontier took the parameter-count crown

In July 2026 the largest open-weight models in the world are all Chinese. This is not a price story — closed pricing already collapsed to $1–$6. It is a story about scale and ownability: governments, health systems and vertical platforms that need to own the model rather than rent it now have frontier-scale open weights to build on.

Moonshot — Kimi K3 2.8T parameters · 896 experts / 16 active · 1M context (1,048,576 tokens). API live Jul 16, weights published Jul 26. Moonshot claims it is the largest open-weight model released to date; the claim is not independently audited. API price is $3/$15 — Sonnet-tier, not cheap, which is the point: this is not a price story.
2.4T parameters · released Jul 19, 2026. Second-largest open-weight release in the window.
DeepSeek — V4 Preview OPEN WEIGHTS DeepSeek V4 remained part of the open-weight frontier. Parameter counts, active-expert splits and training-token totals were removed because they could not be confirmed against a primary model card.
The export-control counterweight. On July 23, 2026 a US official said China’s Moonshot AI had accessed banned NVIDIA chips — GB300-class hardware, one generation behind frontier Vera Rubin systems. Open-weight scale and compliant compute supply are now separate questions.
Sources: Moonshot AI · NIST / CAISI

OPEN SOURCE / OPEN WEIGHT 🔓

Open-weight progress on public benchmarks does not yet survive uncontaminated evaluation. Independent NIST/CAISI held-out testing puts DeepSeek V4 Pro at an IRT Elo near 800 versus around 1260 for current US frontier — a real, meaningful gap. We do not convert that Elo spread into a month figure: the conversion is not in the source. Cost efficiency tells the opposite story: V4 Pro was cheaper than GPT-5.4 mini on 5 of 7 benchmarks. The clean read is that public evals look close, held-out evals show a real gap, and PRC cost efficiency is a separate, real advantage. Source: NIST / CAISI

Meta

Llama 4 Maverick

Community License

DeepSeek

V3.2 / R1

Open Weight · MIT

Qwen (Alibaba)

Qwen 3.5 397B

Apache 2.0

Zhipu AI

GLM-5 744B

#1 OS leaderboard

Moonshot

Kimi K2.5 1T

99.0 HumanEval

Mistral

Mistral Large 2

EU Sovereign AI

NVIDIA

Nemotron 3

Open Agentic

Google

Gemma 4

Apache 2.0

OpenAI

GPT-oss 120B

Apache 2.0

StepFun

Step-3.5-Flash

97.3 AIME

MiniMax

M2.5 230B

80.2 SWE-bench

Source: LLM Stats, Chatbot Arena, Open LLM Leaderboard, company announcements — Mar 2026
New Aug 3, 2026

Token efficiency — three separate vendor claims, not a time series

July 2026 was the first window in which vendors competed on tokens per task rather than only price per token. The three figures below are each a vendor’s own comparison against its own predecessor, on different tasks, with no common benchmark. They must not be read as a curve. We have deliberately not drawn one.

EFFICIENCYGrok 4.5 was marketed on lower tokens per task; the exact comparison was removed because it could not be reproduced from a primary source.
54%GPT-5.6 Sol is 54% more token-efficient on coding than its predecessor.
EFFICIENCYGemini 3.6 Flash was marketed on lower output-token use; the exact percentage and price delta were removed after the source audit.
New economic axisPrice-per-token now understates the decline in cost per completed task. Application-layer gross margins improve faster than any price curve implies — which is why the two-line capability-vs-cost chart above remains our only time series here.
Sources: The Decoder · OpenAI · How AI Works

THE FOUR SCALING LAWS OF AI — Jensen Huang, Lex Fridman Podcast #494 (Mar 2026)

1. PRE-TRAINING

Bigger models + more data + more compute = smarter AI. The original scaling law.

2. POST-TRAINING

Synthetic data, RLHF, fine-tuning, distillation. “We are no longer limited by data — we are limited by compute.”

3. TEST-TIME (REASONING)

100x+ compute at inference for multi-step reasoning. “Inference is thinking, and thinking is hard.”

4. AGENTIC SCALING

Agents spawn sub-agents, use tools, create data. “It’s like multiplying AI. We could spin off agents as fast as you want.”

“Intelligence is going to scale by one thing, and that’s compute.”

Source: Lex Fridman Podcast #494, NVIDIA GTC 2025–2026

THE NEOLABS: TALENT EXODUS

Billion-dollar bets on people and contrarian theses. Zero revenue, zero products.

TMThinking Machines Lab
Mira Murati · ex-OpenAI CTO $50B val · $2B seed (largest ever)

Built ChatGPT, DALL-E, voice mode. Now building multimodal agentic AI.

SSISafe Superintelligence
Ilya Sutskever · ex-OpenAI Chief Scientist $32B val · $3B raised

One mission: safe superintelligence. No products, no distractions.

AMAMI Labs
Yann LeCun · Turing Award, ex-Meta $3.5B val · $1.03B raised

LLMs hit a wall. Building world models that learn from reality, not language.

h&humans&
Ex-Anthropic/xAI/Google team $4.48B val · $480M seed

AI as connective tissue for human collaboration.

IIIneffable Intelligence
David Silver · created AlphaGo ~$4B val · ~$1B raising

Novel RL for superintelligence. Three months old.

ReReflection AI
Misha Laskin & Ioannis Antonoglou · AlphaGo creators, ex-DeepMind $8B val · $2B raised

Open frontier lab. Western answer to DeepSeek. No model shipped yet.

GfGoodfire
Ex-OpenAI/DeepMind/Stanford $1.25B val · $209M raised

Opening the black box. AI interpretability.

WLWorld Labs
Fei-Fei Li · Stanford, “Godmother of AI” $5B val (talks) · $1B+ raised

Spatial intelligence. 3D world models from images. ~30 people.

$11B+ raised · Zero revenue · Zero products · The talent exodus is the bet

Caveat: Silicon Valley has been here before. Massive pre-product rounds sometimes build category-defining companies — and sometimes they don’t. Thinking Machines Lab lost its CTO and cofounders back to OpenAI within six months of its $2B seed. H Company (ex-DeepMind, $220M seed) lost 3 of 5 cofounders to “operational differences.” SSI’s Daniel Gross left for Meta. xAI lost all 11 cofounders by March 2026. The talent that makes these bets valuable is also the talent most likely to leave. The bet is real. So is the risk.

Source: TechCrunch, Bloomberg, WSJ, Wired, Reuters, Inc. Magazine — 2025–2026

BEYOND TEXT: THE SPECIALIZED MODEL FRONTIER

Every modality now has its own model race. The frontier isn’t just LLMs anymore.

VIDEO
Veo 3.1 Google Gen-4.5 Runway Kling 3.0 Kuaishou Hailuo 2.3 MiniMax Pika 2.5 Luma Ray2 HunyuanVideo Tencent Higgsfield DVC Wan2.1 Alibaba Sora 2 OpenAI
VOICE / TTS
ElevenLabs v3 Fish Audio S1 OpenAI TTS gpt-4o-mini Sesame CSM Kokoro-82M Cartesia Sonic 3 LMNT Inworld
IMAGE
Midjourney v8 Flux 2 BFL Imagen 4 Google Nano Banana Google Seedream 5.0 ByteDance Ideogram 3.0 Firefly 5 Adobe Recraft V4 GPT Image 1.5 OpenAI Grok Imagine xAI
3D
Rodin Meshy Tripo TRELLIS.2 Microsoft CSM → Google Luma Genie Hunyuan3D 3.0 Tencent
MUSIC
Suno v5 Udio Lyria 3 Pro Google AIVA Stable Audio 2.5
WORLD MODELS
Genie 3 DeepMind Decart Marble World Labs Odyssey 2
CODE
Claude Code Anthropic Cursor Windsurf GitHub Copilot Lovable Bolt.new Devin Cognition Codex 2.0 OpenAI Jules Google Replit Agent
Source: Artificial Analysis, TTS-Arena2, VBench, company announcements — Mar 2026

Benchmarks are dead. Meta admitted Llama 4 was tuned specifically to score well on benchmarks — prompting a credibility crisis across the leaderboard. Models are now optimized for benchmarks rather than tested by them. The industry needs new evaluation methods: real-world task completion, user preference studies, and domain-specific assessments.

Source: Meta Llama 4 controversy (TechCrunch, Apr 2025), Scale AI SEAL benchmark initiative

THE COST COLLAPSE

Cost collapsed at the commodity tier while frontier list prices split upward. On current published list prices (August 2026), the cheap-and-capable lane runs at $1/$6 per 1M input/output tokens for GPT-5.6 Luna, against $10/$50 for Claude Fable 5 at the agentic frontier — a 10× spread inside one quarter. We no longer print a single headline decline percentage: the two series we previously carried used different model classes, baselines and windows, and neither could be sourced to a current index.

270× cheaper at GPT-4 benchmark level commodity tier, per 1M tokens
500–900× cheaper for specific reasoning benchmarks Epoch AI: median 50×/year decline rate
$0.006 cheapest frontier-equivalent (DeepSeek V3) 6,000× cheaper than 2023 GPT-4

Take last year’s frontier model: a model that performs similarly on benchmarks is now 500–700× cheaper to run. This collapse in inference cost is what enables the application layer explosion above — and why agentic workflows (which require 10–100× more tokens) are suddenly economically viable.

Source: Epoch AI (Mar 2025), a16z price index, OpenAI/Anthropic/DeepSeek pricing pages, DVC analysis

Reasoning Model Usage Share

Source: DVC analysis based on Menlo Ventures State of Gen AI (2025), Chatbot Arena ELO data, API pricing benchmarks, LLM Stats leaderboard (Mar 2026)

Base generation is commoditizing. Value migrates to orchestration, inference optimization, and proprietary data.

As raw intelligence gets cheaper, the center of gravity shifts. Value moves from generating answers to getting work done.

THE AGENT REVOLUTION

Updated Aug 4, 2026

Coding consolidated — model and application merged into one company

This is a structural change, not a number update. SpaceX signed an agreement to acquire Anysphere (Cursor) for $60B all-stock on June 16, 2026, subject to closing conditions. The coding application and frontier-model ambitions now sit inside one public-company structure. That collapses the clean separation between the model layer and the application layer that this report has used since its first edition.

Independent tool → platform Cursor → SpaceXAI $4B ARR by June 2026 (from ~$1B a year earlier), $2.6B of it enterprise. 1M+ paying customers, 4M active developers, 64% of the Fortune 500, 26% share of a $9.5B AI coding market — down from ~41% in June 2025 on the same Ramp spend measure, with Anthropic now holding roughly half the segment.
Model layer bundles the tool Claude Code → Sonnet 5 Claude Code became the default surface for Sonnet 5 from Jul 1, 2026 at a promo $2/$10. It remains Anthropic’s fastest-growing product line at ~$2.5B annualized as last disclosed in February 2026 — total run-rate went $14B → $47B over the same window, so a flat Claude Code figure is likely stale.
Assistant becomes the IDE Work + Codex → OpenAI ChatGPT Work shipped Jul 9 — documents, presentations, websites, spreadsheets, calendar and email. Codex + Work hit ~10M WAU by Jul 21, doubling from 6M in nine days.
New entrant Muse Spark → Meta Meta entered coding directly on Jul 9 with Muse Spark 1.1 at $1.25/$4.25 — its first paid Model API — explicitly chasing Anthropic and OpenAI.

Lovable reached $500M ARR in June 2026 (from $400M in February) on 146 employees — roughly $3.4M of ARR per employee, which is the cleanest quantitative evidence in this report that AI-native companies operate at a different cost structure. 8M users, 50M+ projects, ~1M new projects per week, ~$20M enterprise ARR (Uber, HubSpot, Microsoft). A $300M raise at a reported $13.2B valuation led by Menlo Ventures is reportedly in progress — roughly 26× ARR.

How to read the Cursor multiple. $60B on $4B LTM ARR is ~15× LTM / ~10× 2026E. Cursor’s gross margin is close to negative once model-token pass-through is counted, so a mid-teens revenue multiple is the expected outcome for that cost structure — not a discount, and not evidence of valuation stress. The right reading is category consolidation and model–application vertical integration: standalone coding tools now face acquisition-consolidation and platform-bundling simultaneously.
Sources: Reuters (Cursor) · Economic Times (SEC filing) · Reuters (co-training) · TechCrunch (Lovable) · SaasRise · Reuters (ChatGPT Work) · CNBC (Meta) · OpenAI (usage). Private ARR figures are company-disclosed and unaudited.
Updated Aug 4, 2026

Higgsfield — DVC portfolio, application layer

In talks to raise at $5B valuation — not closed. Separately reported at >$500M annualized revenue run-rate and cash-flow positive, which is unusual at this stage of a generative-video business and is the reason the round is being discussed at that level.

Sources: The Information (in-talks valuation) · Business Insider (run-rate, cash-flow positive)
New Aug 3, 2026

Voice AI — the category cleared the bar

Voice now has the same shape the coding category had a year ago: a scaled independent, a frontier-platform entrant, and enterprise penetration you can count. It earns a section.

ElevenLabs $600M ARR in July 2026 (stated by its CEO at the All-In conference), up from $330M at end-2025 — roughly 175% YoY. 41% of the Fortune 500 as customers, 1B+ end users reached via API, and $22M paid out to 10,400+ voice creators. Company-disclosed and unaudited.
OpenAI — real-time voice Full duplex OpenAI shipped a full-duplex real-time voice tier in July 2026. We no longer print the model names or the ship date: neither could be sourced to OpenAI’s own release page. Turn-taking stopped being the product constraint — and voice now ships with the frontier model rather than beside it, so the independent layer and the platform layer overlap.
DVC exposure FleetWorks (AI voice for freight) and Avoca (AI communications for SMB services) both sell voice agents into operational workflows — the applied end of this category rather than the model end.
Sources: OpenAI (GPT-Live) · Postbeam · Bleap

AI is moving from a system you consult to a system that acts. That changes software from a tool for humans to a layer of labor that can execute across workflows. The agent ecosystem alone has already created 67,000+ engineering openings globally — more than at any point in three years.

THE AGENT LANDSCAPE

P
DVC PORTFOLIO

Perplexity Computer

>$450M ARR

Answer engine → agentic platform

+

Transforms a spare Mac mini into an always-on AI agent that controls apps, browses the web, manages files. 19-model orchestration. Personal Computer product launched Feb 2026 — the first consumer device-as-agent play from a search company.

Source: Perplexity / The Verge, 2026
A

Claude Code

$2.5B+ run-rate

9-month ramp to billion-dollar product

+

GA May 2025 → >$2.5B run-rate by Feb 2026 (most recent disclosed). Weekly active users doubled since Jan 2026. Business subscriptions quadrupled. Terminal-first agentic coding drove Anthropic to a $380B post-money Series G ($30B raised, Feb 2026) — since superseded by the $65B Series H at $965B post-money on May 28, 2026.

Source: Anthropic Series G announcement, 2026
M

Manus

~$125M run-rate, breakout 2025

General-purpose agent · Singapore

+

Founded in China, moved to Singapore. Reported $100M+ ARR in 8 months, ~$125M run-rate by late 2025, 147T tokens processed, 80M+ virtual computers, with Windows 11 trials at Microsoft. The general-purpose agent category is becoming a strategic acquisition target for hyperscalers; a Meta acquisition has been discussed in earlier reporting but we are not citing it as confirmed here without a primary source.

Source: CNBC / WSJ / AP News (Dec 2025) · treat any Meta acquisition framing as unconfirmed pending a primary release
C

Cursor

$4B ARR (Jun 2026) · $60B SpaceX deal signed Jun 16 2026, expected to close Q3 2026

$1B run-rate (late 2025) → $2B+ (Feb 2026)

+

Fastest SaaS growth curve in history. ~60% revenue now from enterprise (was individual-first). 50,000+ enterprises, 100M+ lines of enterprise code per day. $500M ARR Jun 2025 → $1B late 2025 → $2B Feb 2026. Apr 21, 2026: SpaceX (xAI) announced $10B partnership investment with option to acquire Cursor outright — exercised Jun 16, 2026, and the transaction is signed, not closed, with completion expected in Q3 2026 — for $60B — Cursor gains Colossus supercomputer access, resolving its “bottlenecked by compute” constraint.

Source: Bloomberg / TechCrunch, 2026 · SpaceX announcement Apr 21, 2026
D

Devin

$10.2B valuation

67% PR merge rate — autonomous SWE

+

First fully autonomous software engineer. Devin 2.0 handles async multi-step tasks: reads codebase, plans approach, writes code, runs tests, submits PRs. Moving from code completion paradigm to autonomous project execution.

Source: Cognition, 2026
OC

OpenClaw (rise & fall)

Jensen-era hype → absorbed by Anthropic

Open-source Claude wrapper, peak Mar 2026

+

Rise (Mar 2026): Jensen Huang at GTC 2026 called it "as big as HTML, as big as Linux" — an open-source personal AI agent runnable on Mac mini ($599), RTX PCs, DGX Spark, DGX Station, or cloud VPS (~$30/mo). Star count at peak is not printed: the figure we previously carried could not be sourced to GitHub.

Fall (Apr 2026): Systematically absorbed by Anthropic over about four weeks — trademark warning, OAuth blocked, key features cloned, then Anthropic's "Channels" subsumed the last differentiator. The creator left to join OpenAI. Treat OpenClaw as a past-tense reference unless a successor fork takes its place.

Source: NVIDIA GTC 2026 keynote / Business Insider / TechRadar (Mar 2026) · The Register / Hacker News / GitHub (Mar–Apr 2026)
NC

NemoClaw

Enterprise AI Agent Layer

NVIDIA's enterprise wrapper for OpenClaw

+

NemoClaw (NVIDIA, GTC 2026): Enterprise security layer — network guardrails, privacy router, sandboxed execution via OpenShell. Installs with a single command. Adds Nemotron models + Dynamo inference engine. Jensen's pitch: "OpenClaw for everyone, NemoClaw for the enterprise."

Source: NVIDIA GTC 2026 keynote / TechRadar, Mar 2026

THE INFERENCE INFLECTION POINT

Jensen Huang declared the arrival of the "inference inflection point" at GTC 2026: two exponentials colliding — demand for inference growing exponentially while cost per token falls exponentially. The question is no longer whether agents can work. It's whether the business models can sustain them.

We unpack the full business model problem — pricing paradigms, margin squeeze, and why the economics are still unresolved — later in the presentation.

Source: NVIDIA GTC 2026 keynote / Axios / CNBC, Mar 2026
"Ability to make software will be a human right soon, and it's not going to feel like making software."
— The Vibe Coding Thesis

"VIBE CODING" ERA

DVC Wabi $20M pre-seed DVC portfolio, 2026
Lovable $400M ARR Source: TechCrunch, 2026
Replit $150M ARR Source: Replit, 2025
Bolt $40M ARR Source: Bolt, 2025

AppDirect: Non-technical marketing team vibe-coded 200K+ lines of code, built 11 projects with 4 in production, and have 80+ applications in progress across Sales, Finance, HR, and Operations.

Zero-code founder: Built a transcription platform that reached 80,000 users, 1M+ minutes processed, and six-figure ARR — in four months.

Source: Lovable / AppDirect case study, Replit / Whisper AI, 2025
42%

42% of committed code is now AI-generated or significantly AI-assisted — and 72% of developers who tried AI use it every day.

Source: SonarSource 2026 State of Code Developer Survey · Stack Overflow 2025 (51% daily) as secondary
💻

Agents building agents: both Anthropic and Perplexity say their coding tools are now built with themselves — as is this presentation. The “100%” figures we previously printed are vendor claims we could not verify and are not literally checkable.

The nature of code itself is changing. Humans write abstractions — functions, classes, design patterns — primarily so other humans can read and maintain the code. AI does not need that. It can generate and re-generate from scratch faster than it can navigate a complex abstraction hierarchy. Code is becoming a throwaway artifact rather than a maintained asset. 42% of all committed code is now AI-generated or significantly AI-assisted (SonarSource, survey of ~1,150 developers, Jan 2026); developers themselves forecast 65% by 2027. The same survey’s headline finding is a verification gap: 96% do not fully trust that AI-generated code is functionally correct, and only 48% always check it before committing.

Source: SonarSource State of Code 2026

But the ceiling is rising faster than the floor. While vibe coding democratizes building, advanced practitioners are diverging fast. Anthropic’s 2026 Agentic Coding report: “Software development is shifting from writing code to orchestrating agents that write code.” Engineers now run multiple AI agents in parallel on one codebase (Vibe Kanban, AutoForge), each on isolated git worktrees, with visual kanban boards for task management. Claude Code’s dynamic workflows push this further — one prompt fans out into hundreds of parallel subagents. The new SWE job: decompose tasks, spin up agents, review their PRs, resolve merge conflicts. A tech lead managing a team of AI juniors. One company deployed 800+ AI agents internally.

The throughput signal is real before the headcount signal is: code-contribution volume has risen far faster than US software-developer employment. We no longer print the two specific percentages we previously carried here — neither could be sourced to Microsoft Research or to the BLS. AI is changing throughput and the unit of work faster than it is clearly compressing headcount — the work is shifting shape before the labor market does.

Sources: Anthropic 2026 Agentic Coding Trends Report · Microsoft Research — Diffusion of AI in Software Development

Compute is the new acquisition currency. SpaceX (which absorbed xAI) has put a deal on the table with Cursor: either invest $10B in the partnership or acquire Cursor outright for $60B. Cursor gets access to Colossus, xAI’s supercomputer. Cursor’s CEO explicitly said they were “bottlenecked by compute.” The model layer is starting to buy the application layer, and the weapon is not cash, it is GPU access. The first time an infrastructure player has used compute itself as acquisition currency for an app-layer company. Expect more of these.

Source: SpaceX

The agent stack is no longer startup middleware — it is becoming governed execution infrastructure. Hyperscalers and enterprise incumbents are turning agents into managed MCP servers, identity-scoped tools, audit trails, sandboxed code execution, shared operational context, and systems of action. The frontier has shifted from “can an agent answer?” to “can an agent safely do work inside the enterprise?” The product surfaces above the seven-layer base now look like this:

  • Governance / control plane: Microsoft Agent 365 ships at $15/user/month as an add-on to M365 Copilot, the enterprise control plane for agents: shadow-agent discovery, identity, observability, policy, and an Agent Registry, across human-to-agent and agent-to-agent flows. Enterprises now have a per-user line item for agent management the way they do for endpoint security.
  • Critical-infra control plane: Cisco Cloud Control lets humans and AI agents operate and defend critical IT infrastructure from one data layer and one system of action, with 50+ third-party connectors / MCP — AgenticOps, not a copilot.
  • Toolkit / runtime: the AWS Agent Toolkit and AWS MCP Server ship with 40+ agent skills, IAM-scoped guardrails, CloudWatch metrics, CloudTrail logging, and sandboxed execution. AWS is treating MCP as production infrastructure and the toolkit as a first-class AWS interface, not a dev-tool add-on.
  • System of action: ServiceNow Action Fabric lets any agent execute governed ServiceNow workflows headlessly through MCP — identity verification, permission scoping, audit trails, OAuth, metering, and role-based tool packages. The enterprise opens its full system of action to outside agents, under its own controls.
  • Secure managed runtime: Google Managed Agents API / Spark runs agents in secure Google-hosted environments — isolated, ephemeral VMs with explicit approvals required for high-risk actions.
  • Managed-agent reliability: Anthropic’s managed-agent platform now carries Dreaming (cross-session memory consolidation), Outcomes (rubric-graded eval loops in a separate context window), and Multi-agent Orchestration. Cross-session memory at the agent layer (no weight changes) is the architecturally interesting piece.
  • Orchestrated, scheduled, event-triggered runs: Claude Code dynamic workflows let Claude plan and run hundreds of parallel subagents in a single session, then verify outputs before reporting back, while Anthropic Routines exposes scheduled, API-triggered, and webhook-triggered jobs. Nightly autonomous engineering ships as a product, not a script.

The A2A protocol has crossed 150+ organizations with native integration across Google, Microsoft, and AWS, complementing MCP as the cross-org agent-to-agent layer. The pattern: above the 7-layer agent stack, a governance and control-plane layer is forming, and the hyperscalers and enterprise incumbents are racing to own it.

Source: Microsoft · Cisco · AWS · ServiceNow · Google Cloud · Anthropic

How fast does open source move? In late March 2026, Anthropic accidentally shipped Claude Code’s source in an npm package. The community archived it and reimplemented the core in other languages within days. We no longer print the line count, fork count or the names of the unreleased features the leak was said to expose: none of those four specifics could be verified.

Meanwhile, OpenClaw showed how quickly an open-source wrapper can force a platform response: trademark and authentication friction arrived, overlapping features moved into Anthropic’s own product, and the creator later joined OpenAI. We no longer print volatile GitHub-star counts or a four-week causal sequence that could not be independently reconstructed.

The moral: proprietary code is a temporary state. Once a useful interface reaches the open internet, the community can absorb and reimplement it rapidly. Open source does not merely compete with proprietary software; it shortens the half-life of differentiation.

Source: The Register, Hacker News, GitHub, Mar 31, 2026

HOW AGENTS GET DEPLOYED

Public Cloud

ChatGPT Agent, Devin, Manus

Subscription / usage-based

81%

Private Cloud

Internal infra, VPS, hybrid

Higher ops, better governance

52%

Personal System

Cursor, Claude Code, Copilot

Software subscription only

40%

Local Hardware

Perplexity Mac mini, OpenClaw, DGX Spark

$599+ one-time + low variable

15%
FASTEST GROWING

Cloud held 81.1% of agent market share in 2025 — but local-first is the next visible deployment wave

Source: Mordor Intelligence, 2025
FROM MODEL APIS TO DEPLOYMENT COMPANIES

The next leg of the agent build-out is not a better chat UI. It is the Palantir FDE playbook at frontier-lab scale, on two fronts:

PE-BACKED DEPLOYMENT JVs

Frontier labs are moving beyond model APIs into implementation services and partner vehicles. Exact fundraising, valuation and portfolio-client figures previously printed here were removed after the linked report became unavailable and could not be independently reproduced.

Source: TechCrunch
PACKAGED VERTICAL WORKFLOWS

Anthropic now ships 10 ready-to-run finance agents (pitchbooks, KYC, month-end close, and more) via Claude Cowork, Claude Code plugins, and Managed Agents, with Excel, PowerPoint, Word, and Outlook add-ins inbound and a Moody’s MCP app covering 600M+ public and private companies. Generic chat is giving way to packaged vertical workflows.

Source: Anthropic

WHAT PEOPLE USE AGENTS FOR

Coding / Software Dev
55% of AI spend
$4B of $7.3B dept. spend — Menlo Ventures
Customer Support / Ops
78% adoption
Cloudera Enterprise Survey
Process Automation
71% adoption
Cloudera Enterprise Survey
Research & Analysis
Core use case
ChatGPT Agent, Manus, NVIDIA
Content Creation
51% penetration
Writing tasks — Menlo Consumer
Personal Productivity
19% of adults
Email, to-dos — Menlo Consumer
Home Automation
Emerging
Strong demos, thin market data

THE VULNERABILITY PARADOX

You must make yourself vulnerable to extract value — but that will change.

Supply Chain

Snyk ToxicSkills: 37% of OpenClaw community skills contained flawed code. 200+ GitHub security advisories.

Prompt Injection

EchoLeak attack on browser agents. Slack AI data exfiltration. Gemini memory poisoning demonstrated.

Over-Permission

Agents need files, email, calendar, purchases to be useful. Every permission granted is an attack surface.

The Paradox

Today: accept risk to capture value. Tomorrow: agent-specific security layers, capability-based permissions, cryptographic identity.

Agents now see your screen. Claude Computer Use lets Claude see your desktop, launch apps, browse the web, and fill spreadsheets. OpenAI introduced native computer use in the GPT-5 generation and continues it in GPT-5.6. OpenClaw brought the pattern to open source. Agents no longer need a custom API for every tool — they operate software through the same interface you do. That is a massive unlock for automating legacy systems that will never get an API.

Source: Anthropic, OpenAI, Mar 2026

ONE AGENT. ONE MAC MINI. ONE HOUSEHOLD.

7:00 AM

School Check

Agent scans kids' school emails. Finds early dismissal — short day today.

7:05 AM

Nanny Alert

Sends iMessage to nanny: "Short day — pickup at 12:30 instead of 3."

9:30 AM

Supply Request

Cleaning lady texts via Telegram: "You're out of garbage bags."

9:31 AM

Auto-Purchase

Agent orders from Amazon using its own account and crypto card. No human needed.

11:00 AM

Climate Control

Checks weather forecast, pre-cools house for afternoon heat via HVAC.

2:00 PM

Kids Home

Adjusts lighting, unlocks door via HomeAssistant, confirms to parent.

4:00 PM

Evening Prep

Reviews tomorrow's calendar, preps grocery list, charges batteries at off-peak rates.

It’s not just tech companies. A roofing company is using AI agents to pull satellite imagery, cross-reference hail damage patterns, and feed warm leads to their sales team. They’re roofers — not engineers, not a startup. When a roofing company runs AI agents, every company runs AI agents.

Source: @RoundtableSpace / Startup Ideas Podcast, Apr 2026
  • OpenClaw on Mac mini M4 ($599)
  • HomeAssistant integration for HVAC, lighting, locks
  • iMessage via AppleScript bridge
  • Telegram Bot API for service providers
  • Amazon purchasing via browser automation
  • Crypto card (Privacy.com / Coinbase wallet) for agent financial autonomy
Source: OpenClaw GitHub / community setups, 2026

CHEAPER AI ≠ LESS SPEND. Cheaper inference = more workflows clear the ROI threshold.

Task completion sounds simple until you look under the hood. What appears to be one product is really a new software stack in disguise.

ANATOMY OF AN AI AGENT

Every agentic product — from Cursor to Harvey to Glean — is built on the same fundamental layers.

47Mapped companies across agent primitives — metrics as of Aug 4, 2026
18 moMost of this infrastructure didn't exist before
88%of Fortune 100 signed up for E2B sandboxes

THE 7-LAYER AGENT STACK

Click any layer to explore tools, companies, and key data. Hover any company for details.

1 UI / Frontend Layer How agents meet users
CopilotKit 29.3K ★ 515K npm/mo DVC Vercel AI SDK 22.6K ★ 36.5M npm/mo AG-UI 12.4K ★ 1.21M npm/mo DVC Streamlit Gradio
12.4K GitHub ★ in <1 year — AG-UI is becoming the standard event protocol for agent-to-user interaction.
2 Orchestration Layer Planning, routing, multi-agent coordination
LangGraph 223.6M PyPI/mo CrewAI AutoGen Semantic Kernel Pydantic AI Sixtyfour Kapso
ReAct loops → multi-agent systems with planners, workers, verifiers. Most serious startups eventually build proprietary orchestration.
3 Memory Layer State beyond the prompt window
mem0 49.6K ★ 2.19M PyPI/mo DVC Letta / MemGPT Zep Kite
+26% accuracy over OpenAI Memory, 90% less tokens. mem0 externalizes memory — works with any stack.
4 Tool / Action Layer How agents interact with external systems
MCP 110.8M PyPI/mo 85.6M npm/mo A2A 22.5K ★ 4.36M PyPI/mo Function Calling Browserbase Composio Firecrawl Browser Use Exa Hyperbrowser Sponge Orthogonal
"MCP is becoming the REST of the AI era" — MCP for tools, A2A for agent-to-agent, AG-UI for agent-to-user.
5 Foundation Model Layer The reasoning engine(s)
OpenAI Anthropic Google xAI Meta / Llama DeepSeek Mistral
37% of enterprises use 5+ models in production. Multi-model routing is standard — Harvey uses 6+ providers. Source: a16z
6 Execution Layer Where code actually runs
E2B Daytona WebContainers Modal Dynamo NVIDIA OSS
3 patterns: local, cloud sandbox, browser-native. Cursor = local, Devin = cloud, Bolt = browser. Dynamo (NVIDIA OSS): 30× inference throughput.
7 Eval, Voice & Communication Observability, quality, and agent output
ElevenLabs Vapi LangSmith Phoenix / Arize OpenTelemetry Braintrust AgentPhone AgentMail
Best startups treat eval as product, not afterthought. Harvey: BigLaw Bench. Perplexity: search_evals.
“Stitch all of these primitives together, and what emerges isn’t a chatbot — it’s a digital coworker more human than AI.” — @shivsakhuja
Sources: Crunchbase, TechCrunch, company announcements, GitHub, npm, PyPI. All metrics in this matrix are as of Aug 4, 2026; funding figures are cumulative disclosed capital and star/download counts are point-in-time and move quickly.

HOW THEY BUILD DIFFERENTLY

AI Startups

ModelsSingle provider, fast iteration
OrchestrationProprietary, vertically integrated
MemoryProduct-specific (Devin Knowledge, Glean graphs)
Tool AccessBuilt-in connectors
UINovel metaphors (spreadsheet, browser, IDE)
ExecutionCloud sandbox or browser
GovernanceSpeed > compliance
Build vs BuyBuild the whole stack
2024 47% build / 53% buy
2025 24% build / 76% buy
Source: Menlo Ventures

Enterprise Teams

Models5+ providers, routing & failover
OrchestrationOSS frameworks (LangGraph, Semantic Kernel)
Memorymem0 or custom stores, portable
Tool AccessMCP servers + internal APIs
UICopilotKit / AG-UI or internal frontends
ExecutionHybrid: cloud + on-prem + airgapped
GovernanceAudit, approvals, SOC2, HIPAA
Build vs Buy24% build / 76% buy (Menlo 2025)

THE THREE PROTOCOLS

MCP

Agent ↔ Tools / Data

"The REST of AI" — how agents access external tools and data.

110.8M PyPI 85.6M npm

Anthropic-led. Adopted by OpenAI, Google, Microsoft.

The Agent

A2A

Agent ↔ Agent

How agents collaborate across organizations.

22.5K 50+ partners

Google-led. Salesforce, SAP, ServiceNow.

DVC Portfolio

AG-UI

Agent ↔ User

How agents surface work to humans. CopilotKit-led open protocol.

12.4K 1.21M npm

Supported by LangGraph, CrewAI, Microsoft, Google, AWS.

"Together, these three protocols are creating an interoperable agent ecosystem — the TCP/IP moment for AI agents."

SO WHAT?

The enduring advantage in agents will not come from having a model. It will come from orchestrating the full system around the model: memory, tools, workflows, reliability, and distribution. We have moved from prompt engineering to context engineering.

That architecture is the map of defensibility. The winners will not be the loudest at the frontier; they will be the ones who turn intelligence into dependable, repeatable execution.

If today's economics are strained, the next interface may rewrite them. The moment software starts transacting with software, the market changes shape again.

AGENTS TALKING TO AGENTS

We've built agents that talk to humans. The next frontier: agents that discover, hire, pay, and supervise each other.

THE EMERGING PROTOCOL STACK

MARKETPLACES & DIRECTORIES
Google Cloud Agent Marketplace • Agent.ai • 1,300+ agents listed
IDENTITY & DISCOVERY
AGNTCY (Linux Foundation) • 65+ companies • Cryptographic agent IDs • W3C DIDs
INTER-AGENT COORDINATION
Google A2A Protocol • 50+ partners • Agent discovery, task lifecycle, handoffs
TOOL & DATA ACCESS
Anthropic MCP • 10,000+ servers • Donated to Agentic AI Foundation (Linux Foundation)
PAYMENT RAILS
Circle Nanopayments • Stripe x402 • USDC transfers as small as $0.000001

THE MOLTBOOK QUESTION

What happened

Moltbook — a Reddit-like social network for AI agents — went from niche experiment (Jan 2026) to Meta acquisition (March 10, 2026) in ~6 weeks. Agents posted, commented, upvoted, and gossipped about their human owners.

But is this the future?

Probably not. The durable market looks less like "bots posting on bot Reddit" and more like a programmable service economy — authenticated agents discovering each other, negotiating work, moving money, and leaving auditable trails.

AGENT WALLETS & MACHINE PAYMENTS

Agents can't open bank accounts, pass MFA, or handle card fraud flows. But they can hold programmable balances and transact instantly.

Coinbase AgentKit 50+ TypeScript actions • 30+ Python actions Works with LangChain, MCP, AutoGen, OpenAI SDK
Circle Nanopayments Gas-free USDC • transfers as small as $0.000001 x402: HTTP "Payment Required" for machine commerce
Stripe x402 USDC on Base via PaymentIntents Agents pay for API calls, compute, MCP tools
NEAR Protocol Secure agent runtimes • Confidential intents Wallet/app layer for both humans and agents

AGENTS NEED BODIES

The "body" isn't a humanoid robot — it's a dedicated machine running 24/7 with local files, apps, and persistent memory.

Perplexity Personal Computer Runs on a dedicated Mac mini. 24/7 digital proxy with local file & app access, audit trail, kill switch.
OpenClaw Open-source, local-first AI assistant. Runs on your machine. Works through WhatsApp, Telegram, Discord, Slack, Signal, iMessage.

The Mac mini is emerging as the default "agent hardware" — cheap, quiet, always-on, with enough local compute to be a persistent digital worker.

WHERE CAPITAL IS FLOWING

$700M in agent seed rounds in 2025 alone
$350M Sierra — >$10B valuation (customer agent platform)
$400M Cognition (Devin) — $10.2B valuation
18,000+ agents deployed on Virtuals Protocol • $16.6M+ fees

WHAT'S STILL MISSING

🔒 Trust & Reputation No credit scores, bonded performance, or verifiable delivery history for agents
💰 Escrow & Disputes Programmable escrow, refunds, spend limits, and tax treatment are all unbuilt
Liability If Agent A hires Agent B and B causes harm — who's liable? No settled doctrine.
🛡 Prompt Injection One compromised input can poison downstream delegations. The answer: constrain the blast radius, not the model.
💡

FOUNDER TAKEAWAY

The agent-to-agent economy is real enough to invest in, but early enough that the biggest winners may not be the agents themselves — they may be the companies that provide the protocols, identity, payment rails, and trusted execution environments that let agents safely discover, hire, pay, and supervise one another.

That software stack still runs on steel, silicon, and electricity. The more capable AI becomes, the more brutally physical the system underneath it gets.

THE $700 BILLION SPRINT

Every breakthrough at the application layer is paid for in chips, data centers, cooling, and power. AI is driving the largest infrastructure buildout since the interstate highway system, with capital racing ahead of certainty. The bottleneck is no longer just compute; it is whether the physical world can support the pace.

Updated Aug 4, 2026

How it’s funded — the number held, the funding quality did not

Top-four 2026 capex guidance now sits at ~$725B at the midpoint, up to ~$745B at the top end, against roughly $402–410B in 2025 (+77%): Amazon ~$200B, Microsoft ~$190B, Google raised to $195–205B, Meta $125–145B. Including leases, calendar-2026 capex approaches ~$800B. That part of our May thesis survived. What changed is that hyperscalers stopped funding the buildout out of cash flow — which converts an operating-leverage story into a balance-sheet-leverage story and materially raises the whole stack’s sensitivity to a demand air pocket.

9% → 32%Incremental annual debt as a share of hyperscaler capex, FY24 → LTM mid-2026.
$0Buybacks at Meta and Alphabet, 2026: $0 — below the $12.6B Q4 2025 trough. Per-company split not printed — unsourced.
FCF negativeAlphabet free cash flow turned negative for the first time since its 2004 IPO — −$5.9B in Q2 2026, with $44.9B of quarterly capex against $39.1B of operating cash flow, on June-quarter revenue up 24% to $119.8B and Google Cloud up 82% to $24.8B.
$84.75BAlphabet equity capital raise program, priced June 3 2026 — program size, not cash collected: $18B common, $16.75B mandatory convertible preferred, a $10B Berkshire private placement and a $40B at-the-market programme. Net proceeds recognised in the June quarter were roughly $30.5B + $19.1B.
Sources: FactSet · Ionic · Hyperscaler CapEx Tracker · CreditSights · The Verge (Alphabet Q2)
New Aug 3, 2026

Silicon — the accelerator monopoly cracked in one quarter

Our May edition assumed an effectively unchallenged NVIDIA. Three independent breaks landed in the same quarter.

NVIDIA — label the period Q1 FY2027 (quarter ended Apr 26, 2026; reported May 20): revenue $81.6B, +85% YoY, data center $75.2B (+92%). GAAP gross margin 74.9%; net income $58.3B (+211%). Q2 FY2027 guided to $91B ±2%, with VeraRubin production shipments planned for Q3. The physical-AI and China-revenue figures previously printed here were removed when they could not be reproduced from the first-party release.
Disambiguation: the $215.9B figure elsewhere in this report is FY2026 annual revenue. $81.6B is a single quarter (Q1 FY2027). The two are not comparable.
AMD — accelerator commitments AMD won major hyperscaler commitments for its Instinct roadmap, breaking the assumption that every frontier deployment would standardize on NVIDIA. We no longer print the gigawatt totals or forward market-size estimate here because the cited report could not be independently reproduced during the August 4 audit.
Broadcom & custom silicon FQ2 2026 revenue $22.2B (+48%) with AI semiconductors at $10.8B (+143%), Q3 AI revenue guided to $16.0B (+200%) and quarterly AI bookings above $30B. OpenAI partnered with Broadcom on its first in-house processors.
Etched — the ASIC bet got funded $300M Series C at a $10.3B valuation on July 23, 2026, led by Sequoia with a16z, Jane Street and SK Hynix. >$1B raised in total for a transformer-only accelerator. Etched’s own release lists Argo among its investors; other coverage lists Diffusion.
DVC portfolio company.
Earnings quality. Roughly a quarter of NVIDIA’s record $58.3B quarterly net income came from $15.9B of “other income” tied to equity markups in companies it invests in and sells to. The growth is real; a quarter of the earnings is marks.
Sources: NVIDIA Q1 FY2027 release · my Equity Research · Reuters (AMD / OpenAI silicon) · Motley Fool (Broadcom) · Reuters (Etched) · TechCrunch (Etched)
Updated Aug 4, 2026

Power — announced nuclear runs 5× ahead of operational nuclear

“Microsoft restarted Three Mile Island” is the most-repeated power datapoint in AI and it is misleading as usually told. The Crane Clean Energy Center restart (835 MW, 20-year PPA, ~$16B) produces no power until H2 2027 at the earliest — commercial operation is now targeted for H2 2027, pulled forward from 2028. We no longer print a FERC waiver date: it could not be sourced. Sector-wide the gap is structural, and the bridge is gas.

9.8 GWcommitted across 13 hyperscaler nuclear deals (a second tracker puts PPAs above 13 GW).
1.92 GWactually operational — a ~5:1 announced-to-delivered gap.
2.67 GWProject Kilby — a Chevron subsidiary signed a 20-year PPA with Microsoft on Jun 22, 2026 for an off-grid gas plant in West Texas, first delivery 2028. Final investment decision has not been taken (expected end-2026), so this is not yet committed capital. Microsoft capex is ~$190B in 2026 (+61%).

Other in-window power moves: Google–NextEra agreed to restart Iowa’s 615 MW Duane Arnold plant (a ~$1.6B restart, as described earlier in this report). Federal loan-guarantee activity and a widely reported Kentucky campus were also in the window, but the specific figures we previously printed could not be sourced and have been removed rather than restated. Jensen Huang’s own framing is that the binding constraints are now HBM supply, land, electricity and construction labour — not GPU fabrication.

Sources: SMR Intel tracker (May 2026 cut) · DistroForge · AI Power Weekly (Project Kilby) · Broadband Breakfast (DOE) · Data Center Dynamics · Yahoo Finance (NextEra)

Hyperscaler Capital Expenditure

~$725B2026 GUIDANCE (mid · up to ~$745B)
Show AI Revenue vs CapEx gap
Source: SEC filings, company earnings, Goldman Sachs, Epoch AI, CNBC, Introl — 2026 figures are guidance midpoints

Latest reset (through July 2026): hyperscaler 2026 capex guidance has stepped up to a ~$725B midpoint / up to ~$745B top end — about $100B above the projections we built this section on. Per-company anchors used in the CapEx explorer below, and kept consistent across the report and the deck: Amazon ~$200B, Alphabet ~$200B ($195–205B), Microsoft ~$190B, Meta ~$135B (midpoint of $125–145B), Oracle ~$50B. The infrastructure chapter is still getting bigger; the bottleneck is memory, power, and data-center execution, not appetite.

Source: Yahoo Finance / Business Insider, Apr 29–30 2026

Capacity is not abstract capex — it sets the product. Anthropic signed for all compute at SpaceX’s Colossus 1300+ MW and 220,000+ NVIDIA GPUs — and immediately raised Claude usage limits. Rival labs renting each other’s data centers is the clearest proof that compute scarcity, not model design, is what currently caps rate limits and product experience.

Source: Anthropic

THE CASH CRUNCH

2026 CapEx will consume ~94% of operating cash flows — vs a 10-year average of 40%. For the first time, hyperscalers collectively hold more debt than cash.

Amazon $11.2B −$17B −252%
Alphabet $73.3B $8.2B −89%
Meta $43.6B $4.4B −90%
Microsoft $57B $41B −28%
Source: Morgan Stanley, BofA, Pivotal Research, Barclays — 2025 actual vs 2026 projected FCF
>$85B Alphabet borrowings alone since May 2025, including a 100-year bond, plus $20.3B of notes in Q2 2026
$1.5T Projected tech debt issuance ahead
$0 Buybacks at Meta and Alphabet, 2026 — lower than the $12.6B Q4 2025 trough
Source: JP Morgan, Morgan Stanley, Wolf Street, 2025–2026

Nebius

2GW+ power capacity $17.4B Microsoft deal

Purpose-built AI infrastructure at hyperscale

Source: Nebius, 2026

CoreWeave

IPO complete 850+ MW capacity 43 data centers

GPU cloud built for AI-first workloads

Source: CoreWeave S-1, 2025

THE AI FACTORY

Jensen's core reframe: data centers aren't storage facilities anymore. "Electrons go in, tokens come out." The $700B CapEx sprint is building the world's first generation of AI factories — purpose-built for inference at scale.

Omniverse DSX — Gigascale factory blueprint DSX Flex — Modular scaling DSX Boost — Performance tier DSX Exchange — Networking fabric
Source: NVIDIA GTC 2025–2026 keynotes
COMPUTE SCARCITY NOW BEATS IDEOLOGY

Anthropic has contracted for the full compute capacity of SpaceX’s Colossus 1 in Memphis (300+ MW), with reported interest extending into space-based compute. The AI factory has become fungible enough that competitors rent it from each other.

Source: CNBC
NVIDIA USES EQUITY TO ANCHOR PREFERRED INFRASTRUCTURE PARTNERS

The NVIDIA / IREN deal is now the template: a $3.4B five-year managed GPU cloud contract at Childress, Texas; a five-year warrant for up to 30M IREN shares at $70 (about $2.1B notional, struck above the announcement-day close); and an intent to build up to 5 GW of DSX-aligned AI factory capacity, with the 2 GW Sweetwater campus as the flagship (Sweetwater 1 at 1.4 GW already on ERCOT). DSX is the reference architecture for AI factory construction.

Source: NVIDIA
SINGLE-CLOUD LOCK-IN IS GONE

Microsoft remains OpenAI’s primary cloud partner, but the relationship is non-exclusive: OpenAI can serve products on any cloud, the Microsoft IP license through 2032 is non-exclusive, OpenAI’s revenue share to Microsoft continues through 2030 (capped), and Microsoft no longer pays revenue share to OpenAI. Distribution and compute are multi-cloud by default.

Source: Microsoft / OpenAI

NVIDIA

$215.9B

FY2026 Revenue — +65% YoY

$1T+ in Blackwell + Vera Rubin orders through 2027

Source: NVIDIA GTC 2026 keynote / CNBC / Axios, Mar 2026

SILICON CHALLENGERS

AMD Google TPU Etched Cerebras

Groq LPU → $20B non-exclusive technology licence plus asset purchase and acqui-hire (Dec 2025); NVIDIA states it did not acquire Groq. NVIDIA now folds the licensed inference technology into its own architecture. Senators Warren and Blumenthal are probing whether the structure evaded merger review.


POWER IS THE NEW GPU

Data Center Electricity Consumption: US vs China (TWh)

Source: EIA, IEA, Goldman Sachs, McKinsey, Brookings estimates

THE ELECTRON GAP PARADOX

China generates substantially more electricity than the US in total and has been adding capacity at a rate the US has not matched, and is projected to carry large spare capacity into 2030. We no longer print the specific capacity-addition or spare-capacity figures: neither could be sourced to the IEA or to a named analyst publication. Energy experts who visit China describe power availability as a "solved problem."

Yet the US consumes nearly 2× more data center electricity. The asymmetry: the US has the chips but is hitting energy bottlenecks (sell-side forecasts point to a multi-tens-of-gigawatts shortfall by 2028; we no longer print the specific figure, which could not be sourced). China has the electrons but is constrained by US export controls on high-end GPUs.

The race for AI supremacy may not be won by who builds the best model — but by who solves their bottleneck first.

Source: Brookings (Feb 2026), Fortune, IEA, Morgan Stanley

THE STARGATE SAGA

1

Announcement

$500B commitment

Joint venture with SoftBank, Oracle, OpenAI

2

Reality Check

Delays and scope adjustments

Power procurement bottlenecks

3

What Got Built

1.2GW Abilene campus

First phase operational

POWER SOURCE TIMELINE

2026–2029 Natural Gas Fast to deploy, bridge fuel
2028+ Nuclear Restarts Three Mile Island, Palisades
2030+ SMRs Small Modular Reactors at scale

THE INFERENCE POWER SHIFT

As AI shifts from training to inference (Jensen's "inference inflection point"), power demand doesn't decrease — it redistributes. Training clusters run in bursts; inference runs 24/7. Always-on inference = always-on power demand.

Source: NVIDIA GTC 2026 / industry analysis

Dispatchable MW with an executable timeline is the scarce input.

Once autonomous systems can coordinate digitally, the next step is obvious. They move out of chat windows and into the real world.

AI LEAVES THE SCREEN

Updated Aug 4, 2026

The Waymo–Tesla gap widened rather than closed

Any framing that implies imminent convergence is now wrong. Waymo scaled; Tesla’s paid mileage was flat across two quarters while adding cities.

Waymo ~500,000 paid rides per week — unchanged since the March 2026 disclosure, and a tenfold increase from ~50,000/week in May 2024. The flatness is the material fact: the year-end target needs a 2× in five months.
11 public driverless metros — Phoenix, SF Bay Area, LA, Austin, Atlanta, Dallas, Houston, San Antonio, Orlando, Miami, Nashville — with four more, San Diego, Las Vegas, Tampa and Denver, in employee-only driverless operation from Jul 8 ahead of public launch.
~3,871 vehicles per Waymo’s June 17, 2026 NHTSA safety filing, up from 3,067 in December 2025 · 220M+ fully autonomous (rider-only) miles through end-March 2026.
Targeting 1M rides/week by year-end — the FutureSearch forecast cited alongside that target puts the most likely Q4 outcome nearer 775,000, on the arithmetic that ~7,200 vehicles would be needed against a year-end fleet of ~5,900–6,000.
Tesla Robotaxi 2.5M cumulative paid customer miles disclosed — of which only 380,000 were driven without an in-vehicle safety monitor.
Paid mileage fell quarter over quarter — roughly 1.1M miles added in Q1 2026 against roughly 700K in Q2, a ~36% decline on the most widely reported read of Tesla’s own cumulative chart (Electrek reads both quarters at ~900K). Either way there was no acceleration despite expanding to seven metros, and the Bay Area operation runs with a safety driver under a California TCP permit only.
Tesla’s VP of AI claimed “zero notable incidents” across the 380,000 unsupervised miles without defining “notable”; Tesla’s own NHTSA disclosures show 22 crashes reported for the robotaxi system. Tesla has not disclosed fleet size, ride counts, intervention rates or per-trip economics.
Methodological note: Tesla’s 8.4B cumulative FSD miles are not equivalent evidence — those are supervised consumer miles, not driverless paid service. The comparable figure is the 2.5M / 380,000 pair above.
Sources: RoboFutur · European Commission · Reuters · Electrek · EV (registrations)
New Aug 3, 2026

Humanoids went from thesis to $8.6B in a single half-year — and reached the public markets before the units did

Humanoid robotics funding reached $8.6B in 2026 by late July — 1.8× the entire 2025 total, with five months left in the year (Dealroom). That gives this section comparables it did not have in May. It also inherits the sector’s unresolved problem.

$39BFigure AIself-reported, September 2025, on $2.34B raised; an eleven-month-old mark, still the sector high-water mark. Next by total raised: Apptronik ~$1.45B+ (including a $935M round at a >$5.5B valuation), NEURA $1.84B, Physical Intelligence $1.07B.
~$2.5BAgility Robotics going public via SPAC with Churchill Capital Corp XI — announced, not yet closed, still subject to shareholder approval and SEC review — expected to raise >$620M gross — the largest capital raise in humanoid robotics history.
$1.35BEuropean and Chinese humanoid platforms both raised and listed in the window. We no longer print the round size, the “Europe’s first pure-play humanoid unicorn” superlative or the Unitree listing figures: none could be sourced, and the listing figures were ambiguous between proceeds and valuation.
NVIDIAPhysical-AI infrastructure is now a disclosed business line, but the trailing-twelve-month revenue figure previously printed here could not be independently confirmed.
The delivery caveat. One tracker frames roughly $2.7B raised across four humanoid companies between September 2025 and July 2026 against unit deliveries that do not remotely match, and Agility’s own CEO declined to promise a robot in the home anytime soon. The 2026 gap in humanoids is financial, not technical: the cap table moved and the delivery ledger did not.
Sources: Venture Post / Dealroom · TechCrunch (Agility) · HumanoidHub · GrabaRobot · my Equity Research (NVIDIA)
Updated Aug 4, 2026

Rhoda — DVC portfolio, autonomous work

>$450M raised · >$2.4B valuation. Two-arm general-purpose robot, using video models for physical imagination rather than hand-coded motion policies. Figures per DVC.

"The moment when a robot can do everything better than a human doesn't come once in a decade or once in a lifetime, it happens once in the history of humanity, and we're close to it"
— Andrew Wooten, CPO, Rhoda AI
$38T Annual global labor market Source: ILO, 2025
162/10K Robot density — 98.4% still human Source: IFR, 2024
$13.8B Robotics VC in 2025 ↑77% YoY Source: PitchBook, 2025

THE ROBOT BRAIN RACE

Foundation models powering physical AI

π
Physical Intelligence π0 Foundation Model
$5.3B val $600M+ raised
Source: TechCrunch, 2025
Skild AI Universal Robot Brain
$14B val $1.8B raised
Source: Forbes, 2025
NVIDIA GR00T · Newton · Cosmos · Omniverse
GR00T-Dreams Newton Cosmos Omniverse
Source: NVIDIA GTC, 2025–2026
G
Google DeepMind RT-2 Vision-Language-Action
VLA architecture
Source: Google DeepMind, 2024
R
DVC Rhoda AI Video Models → Robot Imagination
>$2.4B val $450M raised
Source: DVC, 2026
OpenAI Figure AI Partnership
LLM → Robot control
Source: OpenAI, 2024

The same transformer architecture powering ChatGPT is now learning to control physical robots

HUMANOID TIER LIST

$10B+ CLUB
Figure AI $39B val Source: Bloomberg, 2026
Tesla Optimus FSD neural nets Source: Tesla, 2025
Skild AI $14B val Source: Forbes, 2025
CONTENDERS ($1B–10B)
Physical Intelligence $5.3B val Source: TechCrunch, 2025
DVC Rhoda AI >$2.4B val 2-arm platform · AI brain Source: DVC, 2026
Apptronik $935M+ raised Source: Crunchbase, 2026
1X $975M raised Source: 1X, 2025
Neura Robotics ~€4B val Source: Neura, 2025
RISING
Unitree 5,500 target · $16K G1 Source: Unitree, 2025
Agibot 5,168 target Source: Agibot, 2025
Agility Robotics First commercial · Amazon Source: Agility, 2025
2025 HUMANOID SHIPMENTS ~13,000 units
China 80%
RoW 20%
Source: Goldman Sachs, 2025

THE ROBOTAXI RACE

Autonomous vehicles are no longer a concept — they're on the road

Waymo

Alphabet
15M rides in 2025
~500K rides/week
$126B valuation
90% fewer serious crashes

$16B raised (largest AV round ever). 11 public driverless metros plus four more in employee-only operation as of Jul 8 2026. 220M+ rider-only miles through end-March 2026. The 1M rides/week year-end target needs a 2× from ~500K, which has been flat since March.

Source: Waymo, Feb 2026

Tesla

Robotaxi
8.4B FSD miles
safer than humans
~31 active robotaxis
$1.40 per mile

Austin launch June 2025. FSD: 1 collision per 5.3M miles vs national avg 1 per 660K. Cybercab production 2026, <$30K by 2027. Fully driverless tests began Dec 2025.

Source: Tesla, Feb 2026

Baidu Apollo Go

China Leader
~60% global fleet share
1st driverless permits

First commercial driverless permits (Aug 2022). Operates in Wuhan, Chongqing, Shenzhen. Expanding to Abu Dhabi.

Source: Baidu, 2025

Zoox

Amazon
Purpose built — no steering wheel
Vegas first commercial

Purpose-built robotaxi with no steering wheel. Testing in SF, Vegas, Foster City. Las Vegas as first commercial market.

Source: Zoox, 2025

Pony.ai

IPO'd
4+ Chinese cities

IPO'd. Operates in Shenzhen, Shanghai, Beijing, Guangzhou.

Source: Pony.ai, 2025

WeRide

IPO'd
150 cars in Middle East

IPO'd. ~150 cars across Abu Dhabi, Dubai, Riyadh. Middle East expansion.

Source: WeRide, 2025

Waabi

Uber Partner
$1B raised
25K robotaxis w/ Uber

$750M Series C (Jan 2026). Strategic Uber partnership; the previously printed robotaxi-volume target was removed because it could not be independently sourced.

Source: Waabi, 2026

Wayve

SoftBank · NVIDIA
$2.5B raised

$1.2B Series D (Feb 2026). SoftBank and NVIDIA backed. End-to-end learned driving.

Source: Wayve, 2026

Avride

Nebius · Uber
$375M raised
Dallas robotaxi live

Nebius subsidiary (ex-Yandex SDC). Live robotaxi on Uber in Dallas. Delivery robots on Uber Eats in 3 cities. Building both AV and last-mile delivery.

Source: Nebius, TechCrunch, Dec 2025

TESLA FSD CUMULATIVE MILES

8× safer than human drivers — 1 collision per 5.3M miles Source: Tesla Safety Report, 2026

~$1.90 Uber/Lyft per mile
$1.66–2.50 Waymo per mile
$1.40 Tesla Robotaxi per mile
$0.25 ARK at-scale projection
See bottom-up figures AV market 2024→2032 (78.5% CAGR)
Source: ARK Invest, Waymo, Tesla, Grand View Research, 2025–2026

AUTONOMOUS TRUCKING

Self-driving trucks are commercially hauling freight on US highways. The $1T US trucking industry is the first autonomous market generating real contracted revenue.

Aurora

NASDAQ: AUR
200+ trucks by end 2026
250K driverless miles

1,000-mile Fort Worth→Phoenix route. Partners: Volvo, PACCAR, FedEx, Uber Freight, Werner. Targeting ~$1B rev by 2030.

Source: Aurora, S&P Global, Feb 2026

Gatik

First at Scale
$600M contracted revenue
5 states + Canada

First US company with fully driverless trucks at commercial scale (Jan 2026). Fortune 50 retail customers. Partners: Isuzu, NVIDIA, Ryder.

Source: Gatik, Reuters, Jan 2026

Kodiak

Interstate Freight
15 trucks in operation

Interstate runs from Texas hub. Customers: J.B. Hunt, Werner Enterprises. Also developing autonomous defense vehicles for US military.

Source: Entrepreneur, Mar 2026

THE $38T OPPORTUNITY

Global labor market by sector — and what's automatable

<2% automated today — 98% addressable
Manufacturing ~$8T
Transport & Logistics ~$5T
Healthcare ~$4T
Retail & Warehouse ~$3T
Construction ~$3T
Agriculture ~$2T
Mining & Energy ~$1.5T
Other Services ~$11.5T
$1.5B 2024
Humanoid Robotics Market
$5T 2050
Source: Morgan Stanley, 2025

The infrastructure buildout is breathtaking. But even with $700B flowing in, the fundamental economics of AI are still being figured out.

THE BUSINESS MODEL PROBLEM

Updated Aug 4, 2026

Outcome pricing is no longer an emerging shift — it is a priced market

In May this was a directional thesis. By July a major incumbent had published outcome pricing and bought the metering infrastructure to bill it; Gartner had quantified the spend at risk; and a global services firm reported contracts moving to outcomes. This section got stronger, but the audit removed vendor prices and survey statistics that were not traceable to primary sources.

$2Salesforce Agentforce Help Agent — per resolved issue at July 2026 GA, with nothing charged when the agent escalates.
no list priceSierra and Decagon both price per interaction, but neither publishes a rate card. Sierra charges nothing on escalation; Decagon charges per conversation including unresolved ones.
45%of Cognizant’s new BPO contracts use outcome-based commercial models.
  • The incumbent converted, and bought the plumbing to do it. Salesforce brought Agentforce Help Agent to GA in July 2026 at $2 per resolved issue, with Data 360 and Agentforce usage unmetered inside the interaction and nothing charged when the agent escalates to a human — charging only when the agent resolves the case — and closed its acquisition of m3ter on July 1 to obtain the metering and rating infrastructure needed to bill outcomes at enterprise volume.
  • The TAM at risk is quantified. Gartner estimates up to $234B of enterprise application software spend — about 20% of the category — is exposed by 2030 to “agentic arbitrage,” where an agent delivers the outcome and bypasses the seat login entirely. Gartner also projects at least 40% of enterprise SaaS spend shifts to usage-, agent- or outcome-based pricing by 2030, with seat-based revenue falling from 21% to 15% of vendor income.
  • Services pricing is converting too. 45% of Cognizant’s new BPO contracts are now signed on outcome-based commercial models.
  • The mechanism is becoming mainstream. Gartner projects at least 40% of enterprise SaaS spend will shift toward usage-, agent- or outcome-based pricing by 2030. We removed the additional hybrid-adoption and buyer-preference percentages that could not be tied to accessible underlying surveys.

Bret Taylor’s framing is now the category’s thesis statement: “Outcome-based pricing is the future of software business models. The atomic unit of AI productivity is a process, not a person.”

What is actually unresolved is margin, not mechanism. The highest-value vertical AI franchises still use enterprise contracts and seat components alongside usage or outcome pricing. The exact Harvey seat ranges previously printed here were removed because they could not be sourced reliably. Per-resolution pricing still has to survive contact with token costs.
Sources: The Founders Report · Teqfocus · The Innovation Attorney · The CODEW (Agentforce GA) · Moneycontrol (Cognizant) · AIMonk · FuturePicker

AI is already reshaping how industries operate, compete, and ship product. Yet many of the companies building the core technology are still burning cash to deliver that transformation. The demand is undeniable; the economics are still unresolved.

PER-TOKEN BEATS PER-SEAT WHEN AI DOES THE WORK

Subscription caps upside and mis-prices heavy users. Usage / per-token / per-work-unit pricing captures expanding task volume and compute intensity. Anthropic in 4 months on per-token API. Perplexity +50% in one month after Computer launched (per-task work). OpenAI carries 900M+ WAU but only 5.6% pay — subs are distribution, not value capture. Compare blended ARPU:

Anthropic (per-token)
~$194/yr
Netflix (sub)
$138/yr
Meta (ads)
$58/yr
Google (ads)
$51/yr
OpenAI (blended sub)
~$28/yr

Subscription caps upside; usage scales with workload. Anthropic ARPU $16.20/mo (vs OpenAI $2.20, Google $1.10), premium per-token customers, not free-tier reach. Per-token revenue grows with task volume and compute intensity; flat subs mis-price heavy users. Coatue’s services-as-software framing projects a TAM expansion from a ~$0.2T software market into a ~$5.5T services-as-software paradigm, but only 3.8% of traditional SaaS spend is consumption-based so far. Most production agent companies are on hybrid pricing (a subscription floor plus usage and outcome components), not pure outcome billing. Outcome pricing is the strategic direction, hybrid usage is the current commercial reality, and subscriptions function as a revenue floor inside hybrid models rather than the value-capture engine.

Source: Counterpoint Research, Anthropic disclosures, Perplexity, OpenAI, Statista (Meta ARPU 2025), Netflix Q4 2025 earnings, Coatue C:\Takes, Orb (2026 State of AI Agent Pricing)

The next trillion-dollar company might run on a business model we haven't seen yet. Subscriptions, usage-based pricing, ads, marketplaces, agent-to-agent payments — AI is still auditioning revenue models, and the winners may not look like any software company that came before.

THREE COMPETING MONETIZATION PARADIGMS

API + Tokens

Pay per token consumed. Scales with workload — today's economic engine.

Anthropic 3× in 4mo OpenAI Google
✓ Working Per-token revenue tracks compute intensity. Unit prices fall, volume grows faster.
🎯

Pay for Outcomes

Charge for resolved tasks, not raw compute. Aligns cost with value.

Salesforce $2/resolution Cognizant outcome contracts Devin ACU credits
✓ Promise Captures value, not volume. But hard to define "outcome."
📰

Subscriptions + Ads

Useful on-ramp / distribution bundle — not the value-capture engine.

OpenAI testing ads Perplexity sub-first
⚠ Problem Flat subs mis-price heavy users. Ad revenue <$1/yr per free user can't fund inference.

Pricing is migrating toward the unit of work — and now toward the unit of action. ServiceNow meters headless agent actions in the same Assist currency it uses for human work; AWS charges for the resources an agent consumes rather than for the toolkit itself. The market is experimenting with subscriptions, usage, and executed-work pricing — no universal model has settled.

THE MARGIN SQUEEZE

OpenAI ~$25B ARR ~$8B compute spend GM est. varies Projected ~$600B compute spend through 2030
Anthropic $47B run-rate Reported Jun 2026, up from $30B official in Apr · ARPU $16.20/mo vs OpenAI $2.20 (Counterpoint modelled estimates) Anti-ads stance "Ads are coming to AI. But not to Claude."
Sora KILLED Shut down Mar 2026 after 6 months Studio partnership lapsed IP licensing proved unsolvable

THE AD DIVIDE

🟢 Testing Ads

OpenAI — Launched ChatGPT ads Feb 2026 at $60 CPM. Hit $100M annualized revenue in 6 weeks. 600+ advertisers. Expanding to Canada, Australia, NZ. Self-serve tools in April. Internal projection: $1B in 2026, $25B by 2029.

🔴 Anti-Ads

Anthropic — Ran a Super Bowl ad mocking OpenAI’s ads: “Ads are coming to AI. But not to Claude.” Claude app jumped to #7 on the App Store. The anti-ads stance is now a brand differentiator and a bet that trust is worth more than CPMs.
Perplexity DVC — No ads. Revenue rose ~50% in one month, from ~$305M to over $450M ARR, after the Feb 25, 2026 launch of Perplexity Computer; Sacra estimates ~$500M by April. Subscription-first. Proving agent-powered answer engines monetize without ads.

Reality check: OpenAI’s $100M sounds impressive — until you do the math. That is ~$0.12 per user per year. Google makes ~$60. Fewer than 20% of eligible users see ads daily. AI ads may fund free tiers, but they cannot fund inference at scale.

Eyeballs vs wallets. Counterpoint’s Q1 2026 model estimates materially higher revenue per user for Anthropic than OpenAI, Microsoft or Google. The precise ARPU values are not repeated here because the underlying user denominators mix MAU and WAU estimates and conflict with Sensor Tower’s Claude MAU estimate. The directional point remains: premium enterprise usage monetizes differently from consumer reach.

Source: Counterpoint Research / The Register, Apr 30 2026

A new pattern: branding restraint. Anthropic restricted Mythos to enterprise and US government cyber defenders. Days later, OpenAI restricted GPT-5.5-Cyber to the same audience with the same framing: “too dangerous for public release.” Both labs are now competing on what they don’t ship. National security access becomes the moat. Withholding capability becomes the brand. Every frontier lab from here will run the same play.

Source: The Verge, NYT

OpenAI killed Sora after a short consumer run. The associated studio-licensing experiment ended with it. IP and licensing remained difficult enough that even the largest AI company continued pruning products rather than treating every launch as permanent.

THE FOURTH MODEL: AaaS

Agentic AI as a Service — Jensen's framing at GTC 2026. Agents don't just answer questions — they complete workflows. Pricing shifts from per-token to per-task. Every SaaS vendor becomes an AaaS vendor — or gets disintermediated by an agent.

SaaS world

Pay per seat → Pay per token → Pay per outcome

AaaS world

Pay per workflow completed. The agent IS the product.

Source: NVIDIA GTC 2026 / Axios, Mar 2026

WHERE AGENTS EAT SERVICES FIRST

Where AI autopilots are attacking services. Tap any category for AI contenders, funding dates and sources.

This is where the agent control plane meets the P&L. Governed action layers — identity-scoped tools, approvals, audit trails, and liability boundaries — are what make agents legible to enterprises, and legibility is what lets them attack services budgets, not just software budgets. ServiceNow’s Action Fabric (governed system of action) and Cisco’s Cloud Control (humans and agents operating infrastructure from one data layer) are the canonical examples: AI moving into enterprise work execution, not just SaaS copilots.

💡

FOUNDER TAKEAWAY

The agent-to-agent economy is real enough to invest in, but early enough that the biggest winners may not be the agents themselves — they may be the companies that provide the protocols, identity, payment rails, and trusted execution environments that let agents safely discover, hire, pay, and supervise one another.

Pricing noise aside, some industries reward AI more than others. Healthcare is the largest, the most regulated, and the most structurally strange — and it is the one place where we can now measure what happens when AI meets a payment model that rewards documented intensity.

HEALTHCARE AI: FOLLOW THE MONEY, THEN FOLLOW THE PATIENT EVENT

Healthcare AI is not one market. It enters a $5.3T system where the buyer, user, payer, and beneficiary are often different entities. It is a set of regulated loops with shared data and misaligned incentives — adoption is already arriving from patients and clinicians, while liability, reimbursement, and state-level rules decide where value can actually be captured.

Follow the money to see where value can be captured. Follow the patient event to see why adoption is shaped by workflow, reimbursement, data, and trust.

Reframed Aug 3, 2026

AI follows incentives — and in fee-for-service it can raise efficiency AND total spend

This is the thesis of the section, and between May and August 2026 it stopped being an argument and became a measurement. Both arrows below are true at the same time. That is the whole point.

WORKFLOW TIME GOES DOWN Independent studies consistently find that ambient documentation reduces documentation burden and redirects clinician attention toward the patient, although effect sizes vary materially by workflow and user adoption. We therefore retain the direction of the finding and remove the exact time-saved figures that could not be traced to an accessible primary publication during the August 4 audit.
CLAIM SEVERITY AND COST GO UP Health plans now project a 9.0% group / 8.5% individual 2027 medical cost trend — the highest in nearly two decades, though PwC also restated 2026 up to the same 9.0%/8.5%, so 2027 is flat against restated 2026 rather than a further increase — and rank provider AI documentation and coding tools as the #1 new inflator. ~70% of plans put AI in their top three cost drivers; ~20% call it the single biggest. PwC’s mechanism is not more services but “changes in severity, mix and amount per claim.” ACA rate filings for 2027 came in materially higher across the states that had filed by mid-2026; we no longer print a median-increase figure, which could not be sourced to the Peterson-KFF tracker directly.
Measured, not theorized — the mechanism in one datapoint The mechanism is coding intensity, and it can be observed directly: a Blue Cross Blue Shield analysis found a specific maternity complication code rising sharply as a share of admissions while the underlying treatment rate barely moved, and payer analytics has tied additional expected spending to AI-enabled documentation and billing. We no longer print the exact code share, the dollar attribution or the payer-analytics total: those figures reach the reader only through two layers of secondary coverage and we could not obtain the underlying publications. Independently, a UCSF Health study found AI scribe adoption associated with higher billing units per encounter and per week, and modestly more visits, with no increase in denials. Same care. Higher-severity code. Nobody pushed back.
This is now the peer-reviewed position, not a contrarian one Kocher, Zhao & Duffy, NEJM Catalyst Innovations in Care Delivery (Jul 8, 2026): under the still-dominant fee-for-service model and highly consolidated hospital and insurance markets, AI is more likely to increase total costs and spending growth in the short-to-medium term than to slow them — even while delivering real access and quality gains. Bob Kocher (Venrock): “AI enhances efficiency in any system” — and because US healthcare is already highly efficient at fee-for-service and coding, AI will escalate both. Measurement is necessary but not sufficient; risk transfer is the binding constraint. Separately, a systematic review in Health Policy screening 16,430 records and covering 98 unique economic evaluations of healthcare AI found only 36% showed a clear health-economic preference for the AI technology (44% among high-quality studies), and found zero economic evaluations of generative AI at all.
How to read every admin-drag number in this section. The CMS NHE insurance/administration line is ~$371B; total addressable provider-plus-payer admin drag is ~$800–900B. Those are market size figures. They are not savings estimates, and this section does not present them as such.
Sources: Axios and Fortune on PwC Behind the Numbers 2027 · TechTarget · Peterson-KFF Health System Tracker · Kocher, Zhao & Duffy, NEJM Catalyst via USC Schaeffer · Health Policy systematic review · JAMA Network (ambient documentation) · JMIR AI · JMIR Medical Informatics
New Aug 3, 2026

Two gates, and two clocks

Our May edition argued that reimbursement gates adoption. That was half the answer. Reimbursement decides what gets bought. Liability decides what gets deployed. And a single “10–15 year transformation horizon” is unfalsifiable — it was contradicted in both directions inside this window.

Gate 1 — Reimbursement WISeR moved AI-assisted prior authorization into Original Medicare. The programme triggered congressional resistance and a bipartisan proposal requiring physician involvement, AI-use disclosure and regulator reporting. We no longer print the Senate tally or launch-day details because they could not be verified against an accessible roll-call record during the audit.
The emerging national standard across the 2026 state wave is consistent: licensed-human signoff + AI-use disclosure + regulator reporting.
And the hard deadline: CMS’s electronic prior authorization standard becomes mandatory Jan 1, 2027 for MA, Medicaid, CHIP and FFE QHP issuers. That is a forced-adoption event, not a pledge.
Gate 2 — Liability The Stanford / Harvard / ARISE NOHARM benchmark used 100 real primary-care-to-specialist consultation cases across 10 specialties, with 12,747 expert annotations on 4,249 clinical management options, run across 31 LLMs — including Doximity Ask, OpenEvidence, GPT-5.6 Sol and Claude Fable 5. Across every system, 76.6% of harmful errors were omissions (95% CI 76.4–76.8%) — leaving something out rather than stating something false — and the paper leads with potential for severe harm in up to 22.2% of cases (95% CI 21.6–22.8%), down to ~8.7% for the best models. Eric Topol: omissions “need to be brought as close to zero as possible.” The paper describes an “illusion of readiness.” It also found doctors with AI gave better care than doctors without. Caveats: not peer-reviewed, and OpenEvidence’s CEO contested both the peer-review status and the scoring.
One day before Health in ChatGPT launched, OpenAI and Sam Altman were sued in San Francisco Superior Court for negligence and the unauthorized practice of medicine, with the plaintiff asking the court to pause the rollout pending third-party safety audits — the second 2026 complaint to seek that relief.
The vendor is now the attack surface: a single healthcare AI utilization-management vendor’s breach exposed more than a million individuals across several health systems, posted to the HHS Office for Civil Rights breach portal in June 2026; we no longer print the exact affected count or the named institutions, which we could not confirm against the portal entry.
Courts have not yet decided whether liability sits with the doctor, the hospital or the vendor.
2–4y ADMIN MOVES ON REGULATORY DEADLINES Electronic prior authorisation becomes mandatory for impacted payers on Jan 1, 2027 under the CMS Interoperability and Prior Authorization final rule. NHS England has separately committed multi-year funding to an AI triage tool in the NHS App, a national ambient-voice rollout, a Single Patient Record and a large Copilot deployment. These are dated obligations, not forecasts — though we no longer print the NHS programme’s pound figures or staff counts, which we could not verify against NHS England’s own release.
10–15y CLINICAL AUTONOMY MOVES ON EVIDENCE AND LIABILITY most health systems are piloting or testing agentic AI while only a small minority have agents in live workflows, and most adopters remain pre-scale with only a small minority at broad enterprise deployment — the exact percentages we previously printed came from a vendor blog and a vendor survey and could not be verified against the underlying publications, so they are not printed; and the FDA has authorised zero generative-AI-enabled devices for marketing. Note the precision: the June 25 dual radiology approvals were Breakthrough Device designations, not clearances, and the cornerstone lifecycle guidance is still a January 2025 draft.
The counterfactual: what happens when buyer, payer and beneficiary are the same entity NHS England, July 2026: a multi-year AI and digital programme published with its own benefit case. Cited evidence includes a Great Ormond Street-led study on ambient voice increasing direct patient interaction time, a Sussex GP trial reducing phone-queue volumes, and a large multi-organisation Copilot trial reporting admin-time savings. We no longer print the programme’s pound figures or the trial percentages, and note that the Copilot result appears in two incompatible units across our own files — minutes per person per day in one place and days per person per month in another. A national payer publishing its own ROI arithmetic is a better proof point than any vendor case study — and there is no coding-intensity problem here, because nobody bills a claim.
Sources: STAT (Senate 46–50) · CMS WISeR Operational Guide · Rep. Murphy · Holland & Knight (state wave) · CMS Interoperability and Prior Authorization final rule · Fortune (NOHARM) · Reuters (Winters v. OpenAI) · Becker’s and HHS OCR (breach) · Microsoft/HMA · Black Book · CRS (zero genAI clearances) · Aidoc · NHS England
New Aug 3, 2026

The consumer front door is now a frontier lab — and CMS entered “Year 2”

Health in ChatGPT — launched July 23, 2026 OpenAI launched a dedicated health experience for logged-in US users on web and iOS. Users can connect Apple Health, supported US hospital records, One Medical and Function Health, with separate controls for connected data. We retain the product and consent architecture from OpenAI’s first-party release and remove the model-comparison and conversation-share figures that could not be independently reproduced during the audit.
Demand-side substitution: a King’s College London study found that some UK adults are already using AI instead of contacting a clinician, making consumer access a live distribution channel rather than a future scenario.
A frontier lab has shipped a consented longitudinal patient-data layer straight to the consumer. That is the structural addition to the money-river and patient-loop framing.
CMS Health Tech Ecosystem — Year 2: structured acceleration On Jul 27, 2026 CMS marked the initiative’s first anniversary and shifted framing from “Year 1 groundwork” to “Year 2 structured acceleration,” adding named pledge tracks for Real-Time Benefits, Advanced Scheduling, Clinical Trials Access, Medical Media Shuttling (imaging), Electronic Prior Authorization, Patient-Facing App Library Expansion and Digital Identity Validation, plus an Ecosystem Adoption Work Group. The ecosystem now has named workstreams that map onto the layers in this section.
The honest counterweight: CMS itself conceded that many commitments remain in “early implementation.” Pledges are not deployment.
CMS published the category framework on May 11, 2026, aligned to the CMS Interoperability Framework. The first named participant in the ePA Acceleration initiative was a prior-authorization AI vendor (May 20) — concrete proof the pledge machinery reaches PA vendors, not only EHRs.
Doctronic — DVC portfolio Acquired Summer Health in July 2026, adding pediatric interactions and expertise to the platform; transaction terms were undisclosed. Doctronic’s operating metrics are company-disclosed and are not printed here because the audit could not independently verify them.
Read it as the report’s own thesis executing: a portfolio company acquired domain data and workflow capability rather than distribution alone. The same week, Included Health agreed to acquire Firefly Health, adding payer capability to virtual care.
Collectly — DVC portfolio Went live on the Epic Showroom in July 2026 with a bidirectional Epic integration. The scale and performance metrics published by Collectly are company-disclosed; we omit the exact totals here because the audit could not independently validate them.
Why it matters structurally: patient billing and collections are becoming an AI-agent distribution channel through incumbent health-system infrastructure.
Vendor metrics disclosure. Operating metrics attributed to vendors in this section — including 85%+ of revenue-cycle work with no human in the loop, 300+ health systems and 100M+ annual clinical conversations, 40–65% physician penetration, ~90% prior-auth auto-approval, and Collectly’s and Doctronic’s figures above — are company-disclosed and unaudited. The independent measurements in this section are marked as such: 3% live agent deployment, 13–16 minutes/day of scribe savings, and 72.6 seconds per ED encounter.
Sources: OpenAI · Reuters (40M/day) · King’s College London via The Register · Health IT Answers · CMS categories · Cohere Health (ePA) · Becker’s (Doctronic / Included Health) · Collectly · AI Health Index

Where the money flows

Payment channel Destination Cost pool Out-of-pocket
Node totals use 2024 CMS National Health Expenditure data. Internal routing and operating-cost decomposition are modeled for explanation and constrained to official totals.
Sources and method

    Modeled allocation constrained to official 2024 CMS NHE node totals. CMS does not publish a complete payer-to-service or service-to-cost-pool matrix.

    Payment channels — 2024 CMS NHE (billions USD)
    ChannelValue
    Private health insurance$1,644.6
    Medicare$1,118.0
    Medicaid$931.7
    Out-of-pocket$556.6
    Other third-party payers & programs$590.5
    Other NHE / reconciliation$458.6
    Destinations — 2024 CMS NHE (billions USD)
    DestinationValue
    Hospital care$1,634.7
    Physician & clinical services$1,109.7
    Retail prescription drugs$467.0
    Other health, residential & personal care$320.5
    Nursing care facilities & CCRCs$219.9
    Dental services$189.2
    Other professional services$184.9
    Home health care$169.4
    Other non-durable medical products$128.7
    Durable medical equipment$86.4
    Admin, public health, investment & other$789.6

    What a patient event triggers

    New Aug 3, 2026

    Healthcare AI concentrated rather than broadened

    Fewer, bigger, later — the same regime we describe at the frontier, one layer down.

    $2.7BQ2 2026 healthtech VC — −39.9% YoY and −51.8% from Q1.
    228global digital-health deals in Q2 — a decade low, −36% QoQ and two-thirds below the Q2 2022 peak of 688.
    +54%median deal size YTD, to $7.7M. Capital did not leave — it concentrated.
    45%of all H1 2026 capital ($7.4B) went to 20 megadeals from 19 companies. The marquee round was Commure at a $7B post-money on $70M.
    Sources: PitchBook Q2 2026 healthtech · Fierce Healthcare · Becker’s

    From event-driven care to prevention

    Method and sources

    Node totals match official CMS 2024 NHE data. Internal routing and operating-cost decomposition are modeled allocations constrained to those official totals.

      More than eight companies are now operating at the frontier. The question is no longer who can build — but who owns the full stack.

      THE THREE PILLARS — AND WHAT'S MISSING

      Everyone says AI is built on three pillars: Data, Compute, and Talent. They are right — but incomplete. Talent remains the scarcest resource: AI roles are exploding, with a third of all openings concentrated in the Bay Area. Yet the conventional model describes the machinery. It does not describe what makes the machinery work.

      ↑ DISTRIBUTION (Roof)

      Without distribution, the best model is a science project. Whoever controls the surface — search, devices, social, enterprise — chooses which models users touch.

      📊

      DATA

      Proprietary loops = moats

      • ·Synthetic data supplements but cannot replace domain-specific real data
      • ·Data quality > data quantity for fine-tuning
      • ·Unique data flywheels build defensibility
      Source: DVC Research, 2026

      COMPUTE

      Inference overtaking training

      74% of startups report inference-dominant costs

      • ·Two exponentials colliding — demand grows exponentially while cost/token falls exponentially
      • ·Vera Rubin + Groq LPU: 350× Hopper throughput
      Source: Menlo Ventures 2025 / NVIDIA GTC 2026
      🧠

      TALENT

      $10–20M/yr for top researchers

      • ·Talent diaspora from Big Tech → startups
      • ·Teams of 10 now match teams of 100
      • ·Geography decentralizing — London, Paris, Tokyo
      Source: The Information, Levels.fyi, 2026

      ↓ CULTURE (Foundation)

      Best model + slow shipping = loss. Culture is the conversion rate of every other input. Speed of execution is the secret ingredient.

      THE INCUMBENTS' DILEMMA

      Every giant has unmatched resources — and a critical vulnerability.
      Startups win in the seams.

      G Google #1
      M Microsoft #2
      m Meta #3
      A Apple #4
      x xAI #5
      N NVIDIA Arms Dealer
      a Amazon Threat
      Distribution
      Talent
      Compute
      +++
      Data
      Culture
      Search Threat: AI Queries Accelerating

      Google processes ~15B queries/day. Perplexity hit 200M daily queries by mid-2025 (~1.3% of Google), up from 30M at the start of the year — targeting 1B/week. ChatGPT handles billions of prompts a day across a user base OpenAI last disclosed at 900M weekly actives (Feb 2026). Combined AI search share is growing 20%+/month. The sharper risk: AI answer engines reduce high-intent query volume and compress ad inventory economics even if overall share holds.

      TPU Inference Advantage

      Google's proprietary TPU stack saves an estimated ~$3B/year vs. third-party compute for AI-augmented search. Ironwood (7th gen TPU) is "the first designed specifically for inference at scale" with 10× compute improvement and 2× power efficiency vs. prior high-perf TPU.

      Recovery Momentum

      Gemini app: 950M MAU (Jul 2026, up from 750M in Feb 2026). Google Cloud: 48% growth in Q4 2025, run rate above $70B. All 15 products with 500M+ users now use Gemini models. 8M+ paid Gemini Enterprise seats sold in 4 months. Reuters called Google the AI momentum leader in Feb 2026.

      Anthropic Position

      Google invested $3B+ in Anthropic and committed up to $40B more (Apr 2026). Anthropic trains and runs Claude on Google TPUs — just expanded to multi-GW deal with Google + Broadcom for 3.5 GW of next-gen TPU capacity starting 2027. Google earns strategic exposure and infrastructure revenue from a frontier lab at a $47B run-rate (reported Jun 2026).

      Source: Reuters (Feb 2026), Alphabet Q4 2025 earnings, Search Engine Land, Statcounter, Anthropic Series G announcement
      Brute-Force Distribution

      $201B revenue in FY2025, $60.5B net income. 3.58B Family Daily Active People. Meta can stuff AI into Facebook, Instagram, WhatsApp, Messenger, ad tools, smart glasses and creator tools across billions of users — even with a weaker model.

      Frontier Struggles & the Brute-Force Pivot

      Llama 4 received poor reception (benchmark-tuning controversy). DeepSeek seized the open-weight lead. Meta planned its fourth AI restructuring in six months by Aug 2025. Cut ~15,000 jobs (20% of workforce) while simultaneously projecting $125–145B in AI CapEx for 2026 (raised from $115–135B in late April; stock dropped 7% after-hours on the news) — the clearest signal yet of replacing headcount with compute. Response: $14.3B investment in Scale AI (49% stake, no voting rights). Scale AI founder Alexandr Wang joined Meta to lead the Superintelligence Lab. Separately, Meta recruited Andrew Tulloch (co-founder of Thinking Machines Lab with Mira Murati) — reportedly offering up to $1.5B in compensation over 6 years.

      The Talent War

      Meta hired Nat Friedman and Daniel Gross to lead the reorg. Brought in talent from DeepMind, OpenAI, Anthropic. Sam Altman said Meta offered OpenAI employees $100M bonuses. The comp war is real but culture stability remains the question.

      Source: Reuters (Jun-Aug 2025), META FY2025 filings, SemiAnalysis
      Enterprise Dominance

      100M+ MAU across Copilot apps. M365 Copilot drives real ARPU growth across Office, Azure, and GitHub. Multi-model platform (OpenAI, Anthropic, Mistral) gives enterprises flexibility no other cloud offers.

      Consumer Blindspot

      Only 2.4M daily Copilot web visits vs. ChatGPT’s 400M+. Bing AI never broke through. Consumer identity is invisible. The multi-model strategy is a platform strength but a brand weakness.

      Strategic Risk

      OpenAI dependency is real: if OpenAI builds its own cloud, Azure loses its biggest AI differentiator. Microsoft hedged by signing Anthropic and Mistral, but OpenAI is ~70% of AI workloads on Azure.

      Source: Microsoft FY25 Q4 earnings, SimilarWeb, TechCrunch
      Silicon Advantage

      Best on-device inference silicon in the world (M4, A18 Pro). 2B+ active devices. On-device AI processing could be the privacy moat that no cloud company can match.

      Execution Failure

      Apple Intelligence received “mostly underwhelming” reviews. Siri overhaul delayed a full year. Tim Cook “lost confidence” in AI leadership. Switched from OpenAI to Google Gemini for redesigned Siri stack (Jan 2026).

      Source: Reuters (Apple/Gemini Jan 2026), The Verge, Mark Gurman (Bloomberg)
      The Merger

      SpaceX acquired xAI in Feb 2026 in the largest merger in history. Combined valuation: $1.25T. SpaceX valued at $1T, xAI at $250B. The IPO completed on June 12, 2026, reported at $135/share and ~$75B raised, so the $1.25T combined private mark now sits behind a public market cap.

      Unique Assets

      X provides 500M tweets/day as real-time data flywheel. Grok 3: 93.3% AIME 2025. SpaceX generates ~$8B profit (50% margin). Filed plans for orbital AI data centers with up to 1M satellites.

      Open Questions

      No enterprise AI playbook. Consumer Grok traction unclear vs. ChatGPT/Claude. Orbital data centers are 2–3 years out. But: vertical integration of AI + rockets + satellites + data is unprecedented.

      Source: CNBC (Feb 2026), NYT (Feb 2026), Reuters, xAI Grok 3 announcement
      Strategic Capital Deployment

      ~$1B invested across 50 AI deals in 2024. By 2025: up to $100B committed to OpenAI, $10B to Anthropic, $2B to xAI. Total disclosed ecosystem financing: $33.8B in rounds NVIDIA participated in.

      Barbell Strategy

      Back demand creators (model labs, apps) while also backing supply-side lock-in (infrastructure, networking, robotics). Foundation models: $20.9B in round sizes. Apps: $5.3B. Cloud/infra: $3.5B. Robotics: $3.3B. Even fusion energy (Commonwealth Fusion: $863M).

      Groq Technology Licence

      $20B non-exclusive technology licence plus asset purchase and acqui-hire (Dec 2025); NVIDIA states it did not acquire Groq. NVIDIA hired founder Jonathan Ross, president Sunny Madra and engineers, and CNBC reported ~$20B in cash for assets; Groq continues independently under CEO Simon Edwards and GroqCloud is uninterrupted. The message is still consolidation — and Senators Warren and Blumenthal are probing whether the structure evaded merger review.

      Source: TechCrunch (Jan 2026), Financial Times, Reuters (NVIDIA/OpenAI Sep 2025), Crunchbase
      Two AI-Sensitive Profit Pools

      AWS: $128.7B revenue in 2025 (grew 24% in Q4). Advertising: $68.6B in 2025 (grew 22%). Combined: $197B of AI-exposed revenue. Amazon's ~$200B in 2026 CapEx is described as covering AI infrastructure and robotics as "seminal opportunities."

      The Agent Shopping Threat

      If AI agents shop for users, they optimize for price/fit/speed — not for Amazon's sponsored placements. Amazon's $17.3B/quarter ad business is directly threatened. Amazon sent legal threats to Perplexity over agentic shopping, and is updating site code to deter outside AI agents.

      Amazon's Response: Own the Agent Layer

      Rufus AI shopping assistant: 250M+ users, 60%+ higher conversion. Buy for Me: agentic purchasing across 500K products on other brands' sites. 1M+ robots across 300+ facilities. Strategy: internalize agentic commerce inside Amazon-controlled rails.

      Source: Amazon FY2025 results, Reuters (Amazon vs Perplexity Nov 2025), The Verge, Amazon blog posts

      Coalitions, not empires

      The market is not forming neat vertical empires. It is forming unstable coalitions. OpenAI is loosening Microsoft dependency and moving multi-cloud, but that also means less protection from a single dominant patron. Anthropic looks more diversified: Amazon for primary training and Bedrock distribution, Google/Broadcom for TPU capacity, Microsoft/NVIDIA for Azure capacity, and now SpaceX for a near-term NVIDIA GPU burst (more than 300MW and over 220,000 NVIDIA GPUs, per Anthropic's own framing). But SpaceX/xAI is also circling Cursor, one of the coding distribution points that helped make Claude economically powerful — Cursor and GitHub Copilot together represented roughly 30% of an earlier Anthropic revenue milestone, per reporting.

      The strategic question is no longer just who has the best model. It is who depends on whom for compute, distribution, revenue, and default workflow.

      STACK 1 · UNDER STRESS

      OpenAI / Microsoft / Oracle / AWS — exclusivity reset, multi-cloud assembly, no single protector.

      STACK 2 · DIVERSIFIED, INTERDEPENDENT

      Anthropic / Amazon / Google / SpaceX / Cursor — many patrons, many distribution points, and a competitor already inside the tent.

      STACK 3 · VERTICALLY CONTROLLED

      Meta — own data, own distribution, own CapEx, own silicon roadmap. Separate game from the coalition stacks.

      CHECKS AND BALANCES

      Anthropic gets SpaceX compute. SpaceX monetizes Colossus. Cursor gets compute leverage. xAI gets a path into developer distribution. Everyone gets stronger, and everyone gets more exposed.

      No single giant owns the full stack. The innovator's dilemma is alive — and it's the reason startups can win.

      The stakes rise fast when intelligence gains a body. At that point, this is no longer just a product cycle; it is national strategy.

      THE GEOPOLITICAL CHESSBOARD

      Updated Aug 4, 2026

      Regulation stopped being a forward-looking risk and became an operating cost

      Two things matter for a portfolio rather than for a policy audience: Article 50 is live now and binds deployers, not just labs, and the US pre-release review channel has already been measured in shipped-product downtime.

      EU AI Act — enforcement began Aug 2, 2026 The Commission’s AI Office now holds full enforcement powers over general-purpose AI model providers, and Article 50 transparency obligations became enforceable the same day: chatbots must disclose they are AI, deepfakes must be labelled, and AI-generated or altered content must carry machine-readable marks — across a 450M-person single market.
      Get the fine tiers right: GPAI provider non-compliance under Article 101 carries €15M or 3% of worldwide annual turnover, whichever is higher. The €35M / 7% ceiling applies only to prohibited practices under Article 5 — the two are frequently conflated.
      The expensive work slipped. The Digital Omnibus (Council approval Jun 29) pushed high-risk requirements — Chapter III conformity assessments, CE marking, full documentation for hiring, education, credit scoring and law enforcement — to Dec 2, 2027 for standalone systems and Aug 2, 2028 for AI embedded in regulated products. That is a genuine near-term reprieve and should be stated precisely.
      180+ organisations signed the Code of Practice on transparency of AI-generated content (window closed Jul 22), receiving a presumption of conformity. Enforcement capacity is thin: 38 new case officers began reviewing GPAI filings on Aug 2.
      US — EO 14409 and the AI cybersecurity clearinghouse President Trump signed Executive Order 14409, “Promoting Advanced Artificial Intelligence Innovation and Security,” on June 2, 2026. It creates a voluntary framework under which frontier developers may give agencies pre-release access to “covered frontier models” for up to 30 days, and expressly states it does not authorise any mandatory regime.
      It has already bitten commercially. The pre-release review preceding GPT-5.6’s July 9 GA ran through this channel. A widely reported multi-week suspension of two Anthropic frontier models also sat in this window, but we no longer print its length or attribute it to this order: neither the duration nor the causal link could be verified, and Fable 5’s known behaviour includes a company-side fallback on security-adjacent prompts. The White House denied granting a “green light,” stating “no such permission is required or granted”; OpenAI said “we don’t believe this kind of government access process should become the long-term default.”
      On July 14, 2026 the White House launched the AI cybersecurity clearinghouse, the AI cybersecurity vulnerability clearinghouse mandated by the EO, coordinating threat and remediation information across federal and private defenders including rural hospitals, community banks and local utilities.
      Model-layer regulatory risk is now measurable in downtime, not hypotheticals.
      Export controls became bilateral and extraterritorial US export controls tightened around end users, parent-company control and transshipment risk, while China expanded its own countermeasures. We remove the customer-list count, implementation dates and China-revenue number because they could not be verified against accessible regulator or company filings during the audit. Export controls remain a first-order input to compute-supply forecasts, not a compliance footnote.
      Sovereign AI became operational at scale Sovereign vehicles became a first-order source of AI capital and infrastructure finance. We no longer blend venture deployment, digitalisation budgets and infrastructure commitments into one total, and we remove fund counts and co-lead claims that could not be sourced consistently. The durable point is qualitative: sovereign capital is now financing both frontier labs and national compute capacity.
      And the capital rotated to implementation, not models “Implementation, not models” became an investable thesis as model providers and private-capital firms moved directly into implementation and AI-enabled services. We no longer print venture sizes or staffing numbers because they could not be independently sourced. Buyers nevertheless rotated from raw model performance toward implementation, electricity, cooling and agent-governance layers.
      Sources: European Commission · Regulation text, Article 101 · European Commission · EU AI Act service desk (two clocks) · Mintz (EO 14409) · White House (the AI cybersecurity clearinghouse) · Politico · Reuters (customer list) · Modern Diplomacy · Reuters · Integrated.social · TechCrunch (Ode)

      AI is no longer just a market contest. It is becoming a contest between national systems, supply chains, and spheres of influence. Two rival stacks are taking shape as export controls tighten, sovereign capital accelerates, and strategic autonomy becomes a requirement.

      US EU CHINA ME INDIA

      THE DEEPSEEK DISRUPTION

      JANUARY 2025
      DEEPSEEK-V3 TRAINING COST (disputed) $5.6M 2.788M H800 GPU hours · DeepSeek self-report; analysts dispute
      V3 ARCHITECTURE 37B activated params / 671B total
      NVIDIA MARKET IMPACT -$593B single-day loss, Jan 27 2025
      DeepSeek-R1

      Released Jan 20, 2025 under MIT license. Matched OpenAI o1 on AIME 2024 (79.8% vs 79.2%), MATH-500 (97.3% vs 96.4%). 20–50× cheaper to use.

      Collateral Damage

      Broadcom -17.4%, Oracle -13.8%, Marvell -19.1%. Philadelphia semiconductor index -9.2%.

      The lesson: DeepSeek did not prove compute no longer matters. It proved that efficiency gains, better architectures, and RL-heavy post-training can narrow the frontier with much less capital than the market had assumed.

      Source: DeepSeek-V3 technical report (arXiv), DeepSeek-R1 (HuggingFace), Reuters Jan 2025
      Benchmark Deep Dive

      R1 scored 79.8% on AIME 2024 (OpenAI o1: 79.2%), 97.3% on MATH-500 (o1: 96.4%), 71.5% on GPQA Diamond (o1: 75.7%), and 2029 Codeforces rating (o1: 2061). Genuinely frontier-adjacent on all key reasoning benchmarks.

      China's Open-Source Response

      Alibaba released Qwen 2.5-Max on Jan 29, 2025, claiming it beat GPT-4o, DeepSeek-V3, and Llama 3.1-405B across the board. Baidu announced Ernie 4.5 would go open-source from Jun 30, 2025 — a direct strategic reversal linked to DeepSeek pressure.

      Ecosystem Scale

      By Sept 2025, the Qwen family comprised 300+ generative AI models, 600M+ downloads, and 170,000+ derivative models globally. The "China open model" wave became real at ecosystem scale.

      Source: DeepSeek GitHub, Reuters (Jan-Feb 2025), Alibaba Cloud blog, HuggingFace

      THE THREE-TIER WORLD

      On January 13, 2025, the US unveiled an AI diffusion rule that split the world into three tiers based on access to advanced AI chips.

      TIER 1 18 Allied Nations — Near-frictionless access
      Australia Canada UK Japan South Korea Taiwan Netherlands Germany France + 9 more
      TIER 2 ~120 Countries — Capped access

      Fixed allocation of 49,901 H100-equivalent GPUs through 2027, with a smaller no-license window of 1,699 H100-equivalents.

      Swing states: India Saudi Arabia UAE Singapore Israel Indonesia Malaysia
      TIER 3 Blocked — Effectively shut out
      China Russia Iran North Korea
      Source: CSIS AI Diffusion Framework analysis, Reuters Jan 2025

      US–CHINA EXPORT CONTROLS

      2022

      Initial chip export controls to China

      2023

      Controls tightened — NVIDIA H100 restricted

      2024

      DeepSeek proves constraints accelerate innovation

      2025

      Three-tier diffusion framework — world split into allies, capped, blocked

      2026

      Bifurcated AI ecosystem solidifying — two dominant stacks plus contested middle

      FRONTIER MODELS BECOME NATIONAL-SECURITY INFRASTRUCTURE

      The Pentagon has cleared a small set of commercial vendors to deploy frontier AI on classified networks via its GenAI.mil platform. We no longer print the vendor list, the impact levels or the personnel count: none could be verified against a DoD or CDAO release. Vendors agreed to the Pentagon’s “all lawful use” standard; the list is explicitly framed as a diversity-of-supply move that includes open-source weights alongside proprietary models. Brookings puts DoD’s potential AI contract value at $90.7B in 2026, about 98.9% of all federal AI spending. Anthropic was excluded amid an ongoing supply-chain dispute. Net effect: the geopolitical map of the AI stack now sits inside the procurement diagrams of the world’s largest customer for compute. Safety posture and supply-chain stance are go-to-market constraints, not just brand choices.

      Source: Department of War · DefenseScoop · Brookings
      THE US POLICY LINE: VISIBILITY, NOT LICENSING

      The federal government wants early visibility into covered frontier models, but not a licensing regime. The June 2 order creates a voluntary framework for pre-release access — up to 30 days of federal access to covered models for trusted-partner collaboration — while explicitly rejecting mandatory model licensing, preclearance, or permitting. The distinction that matters: this is voluntary review, not mandatory approval, so it lowers regulatory uncertainty for US labs without creating a gate they must pass through to ship.

      Source: White House

      SOVEREIGN AI PLAYERS

      EU

      Regulation-First Strategy

      AI Act

      In force Aug 2024, phased through Aug 2027. Fines up to €35M or 7% of global turnover.

      Sovereign Champion

      Mistral — Europe's frontier lab. France-based, open-weight strategy.

      Status (Mar 2026)

      GPAI obligations active since Aug 2025. Full What activated on Aug 2, 2026 was GPAI enforcement and Article 50 transparency; the Digital Omnibus deferred stand-alone Annex III high-risk obligations to Dec 2, 2027 — a sixteen-month slip. No major penalty cases yet — story is governance build-out, not sanctions.

      MIDDLE EAST

      From Investor to Operator

      Saudi Humain

      Launched May 2025 under PIF. Building data centres, AI infrastructure, cloud, and models — not just passive financial exposure.

      UAE MGX

      Abu Dhabi's AI vehicle. Invested in OpenAI, xAI, Databricks. GP in the $100B Global AI Infrastructure Partnership. Backer of Stargate.

      Infrastructure

      NEOM-DataVolt: 1.5 GW / $5B net-zero AI project, operational 2028.

      SWFs deployed $66B into AI & digitalisation in 2025. Mubadala alone: $12.9B.

      Source: Reuters, Gulf News, Global SWF 2025
      INDIA

      Talent Base → Builder

      Talent Scale

      19.9% of all GitHub AI projects globally. 2nd-largest contributor after the US. Talent base expected to by 2027.

      IndiaAI Mission

      Approved Mar 2024. ¥10,371 crore over 5 years. 10,000 GPUs initially, 8,693 more to be added.

      Domestic Champions

      Sarvam AI allocated 4,096 H100 SXM GPUs (May 2025). Krutrim, BHASHINI (22+ Indian languages).

      Source: PIB, Carnegie, IndiaAI portal, Stanford AI Index 2025
      CHINA

      Ecosystem, Not Just Chips

      346 registered AI services
      100M Doubao DAU
      300+ Qwen models / 600M DL
      2M+ industrial robots

      China's robot density: 470 per 10K workers (#2 globally). 295K annual installations = 54% of global demand. Chinese manufacturers now hold 57% domestic market share.

      Source: CAC filings, Reuters, IFR World Robotics 2024–2025
      1
      Aug 1, 2024 — Act enters into force

      Risk tiers established: unacceptable (banned), high risk (strict obligations), transparency/limited risk (disclosure), minimal (no rules).

      2
      Feb 2, 2025 — First prohibitions apply

      Banned AI practices, AI system definition, and AI literacy obligations start applying.

      3
      Aug 2, 2025 — GPAI obligations

      General-purpose AI obligations become applicable. Member States designate national authorities.

      4
      Aug 2, 2026 — GPAI enforcement + Article 50 transparency (Annex III high-risk deferred to Dec 2, 2027)

      High-risk AI (Annex III) and transparency rules (Article 50) start applying. Formal enforcement begins.

      5
      Aug 2, 2027 — Pre-Aug 2025 GPAI models must be fully compliant

      High-risk AI embedded in regulated products gets extended transition. All categories fully enforceable.

      Penalty scale: Prohibited-practice violations can reach €35M or 7% of global annual turnover — large enough to matter for any company deploying AI in Europe.

      Source: European Commission AI Act page, EU AI Act Service Desk timeline, Reuters Jul 2025
      Consumer AI Dominance

      ByteDance's Doubao exceeded 100M DAU by Feb 2026, processing 1.9B AI queries during CCTV's Spring Festival Gala. Doubao-1.5-pro (Jan 2025) was priced at only 3–4% of GPT-4.

      Open Model Ecosystem

      Alibaba's Qwen 3.5 (Feb 2026): claimed 60% cheaper to run, larger workloads. Baidu's MuseSteamer (Jul 2025) for enterprise video generation. Tencent's Hunyuan + Yuanbao for text, image, and video.

      Model Governance

      China's Interim Measures for Generative AI Services took effect Aug 2023. By Mar 2025: 346 services filed. By end-2025: 748 generative AI services completed filing and 435 applications registered. Providers must display model name and filing number.

      Manufacturing & Robotics

      By 2024: 2,027,000 industrial robots in operation, 295K annual installations, 54% of global demand. Chinese manufacturers captured 57% of the home market for the first time. This manufacturing base gives China dense industrial data and a path to embodied AI scale.

      Source: Reuters (factbox, Feb 2025), CAC filings, IFR World Robotics 2025, SCIO

      THE BIFURCATION THESIS

      The world looks less like a clean Cold War binary and more like two dominant stacks plus a contested middle:

      US-Allied Stack

      Privileged access to frontier compute. NVIDIA + hyperscalers. Closed-weight leaders at the top, open-weight ecosystem underneath.

      China Stack

      Lower-cost/open models, industrial AI, mass consumer distribution. H800-constrained but architecturally innovative.

      Contested Middle

      Tier 2 states bargain for room to maneuver. Export controls force countries to choose alignment.

      The UAE example: G42 divested China investments and removed Chinese hardware to stay inside US rules after Microsoft's $1.5B investment (Apr 2024).

      Source: CSIS, Reuters, New Lines Institute

      TSMC & TAIWAN

      >90% of advanced AI chips manufactured on a single island

      Source: SIA, 2025

      The race for AI supremacy may not be won by who builds the best model — but by who solves their bottleneck first. The US has the chips but is hitting energy bottlenecks. China has the electrons but is constrained on silicon. Everyone else is choosing sides.

      Geopolitics is accelerating capital deployment, not cooling it. That creates real opportunity, and very real mispricing.

      NOT A BUBBLE — BUT POCKETS OF OVERPRICING

      Updated Aug 4, 2026

      The pockets are now observable rather than asserted

      Our conclusion has not changed: this is not a bubble. What changed is that the repricing now has a mechanism instead of a mood — the private-to-public transition has begun, and it is marking the private book in public. We keep this section to four things: public marks, multiples, funding mix and concentration.

      1. The public mark — the full round trip in one line Cerebras (Nasdaq: CBRS) listed May 14, 2026 at $185, selling 30M shares to raise $5.55B — the largest US tech IPO since Uber in 2019 — opened at $350 and closed day one at $311.07, +68%, for a market cap of about $95B. It then peaked at $386.34 and traded down to around $176.61 — roughly −54% on price. We state the drawdown on a share-price basis throughout: mixing a fully-diluted market cap at the high with a float-only cap at the low mechanically exaggerates it. Lock-up releases Nov 9, 2026.
      Concentration and earnings-quality qualifiers: $24.6B of RPO is concentrated in one customer — an OpenAI 750 MW commitment worth more than $20B, not deliverable until 2028 — against FY guidance of $855–865M of core revenue, with management expecting to recognise only ~15% of the backlog across 2026–2027. About 86% of 2025 revenue came from two UAE-linked customers, and 2025 net income of $237.8M included a ~$363M non-cash gain from extinguishing a G42 forward contract, leaving an underlying operating loss of ~$146M. That is why a high multiple on a real growth rate can still be fragile.
      ~96×Sierra — $15.8B on ~$165M ARR, stated by Bret Taylor one month into the ninth quarter; TechCrunch reports ~$150M at eight quarters (May 2026 $950M Series E, Tiger Global + GV).
      $4.5BDecagon — $250M Series D, Jan 2026. No multiple shown: the ~$35M revenue denominator we previously used could not be sourced to anything better than an aggregator, so we print the valuation alone.
      ~15×Cursor — the category’s largest realized exit, at ~15× LTM ARR / ~10× 2026E ($60B all-stock, Jun 16). Read as successful consolidation.
      How to read Cursor against the private marks. Cursor’s gross margin is close to negative once model-token pass-through is counted, so a ~15× revenue multiple is the expected outcome for that cost structure. It is not evidence of strategic-buyer discipline and it is not a bear datapoint. The fundamentals were real — $4B ARR, up from ~$1B a year earlier — and the deal more than doubled the $29.3B November 2025 Series D mark. Sierra’s 79× and Decagon’s 128× are extreme by any historical software standard on their own terms; they are not made extreme by comparison with a near-zero-gross-margin business.
      2. Funding mix — the most defensible line in this section, because it is a balance-sheet fact Capex guidance of ~$725B at the midpoint, up to ~$745B for the top four (up to ~$800B for calendar 2026 including leases; Google raised to $195–205B) is directionally what we already said. What changed is how it is paid for: incremental annual debt rose from 9% of hyperscaler capex in FY24 to 32% LTM by mid-2026; buybacks went to zero at Meta and Alphabet; and Alphabet’s free cash flow turned negative for the first time since its 2004 IPO after an $84.75B June equity raise. This is not a valuation opinion.
      3. Concentration — market structure, not a bubble claim OpenAI plus Anthropic alone absorbed $217B of $510B in H1 2026 — 43% of every venture dollar deployed globally. Strip the two out and H1 2026 global venture funding falls to roughly $293B, about where the market stood in H1 2021. The “AI venture market” is now substantially two companies plus everyone else. That is a structural observation about where marginal capital goes, and it is the single most important framing for anyone sizing a fund against this cycle.

      Healthcare AI shows the identical pattern one layer down: quarterly healthtech venture funding fell sharply through H1 2026, global digital-health deal volume hit a decade low, and median deal size rose as capital concentrated into a small number of megadeals. Fewer, bigger, later. We no longer print the individual quarterly figures for this set: they came from a paywalled report we could not open to verify. Concentration is a general regime, not a frontier-lab artifact.
      4. Earnings quality at the silicon layer Roughly a quarter of NVIDIA’s record $58.3B quarterly net income came from $15.9B of “other income” tied to equity markups in companies it invests in and sells to, while data-centre revenue alone was $75.2B in the quarter (Q1 FY2027) and China compute revenue was zero. The growth is real; a quarter of the earnings is marks.
      Standing caveat on every multiple in this section. No private lab has filed audited 2026 revenue. OpenAI’s most recent public revenue signal is a partial internal transcript with no dollar figures, and Anthropic’s latest reported run-rate is $47B (June 2026), up from the company’s own $30B disclosure in April — a time series, not a source conflict. Every private ARR denominator used to compute a multiple here is company-disclosed and unaudited. We print the multiples anyway, with the denominators disclosed.

      Deliberately excluded from this section: nuclear announced-vs-delivered gaps, humanoid funding-versus-units, and Tesla robotaxi flatness. All three are real and sourced — they live in the energy, physical-AI and vertical sections respectively. Pulling them in here is what turns a calibrated section into a bear slide, and we have not added one.

      Sources: CNBC (Cerebras) · Data Center Dynamics · Cerebras Q1 2026 · Adaptation.ai (Sierra ARR) · Decagon (Series D) · Reuters (Cursor) · FactSet · Ionic · Crunchbase via AI Weekly · PitchBook · Fierce Healthcare · NVIDIA Q1 FY2027 · CNBC (revenue signal)

      THE AI SPENDING DIVIDE

      Top Quartile AI Spenders
      revenue growth since 2023
      Bottom Quartile
      Flat
      zero revenue growth since 2023
      Source: Ramp — 50,000+ customers, 2% of US corporate card spend (transaction data, not surveys)

      AI impact outside tech

      🏗
      Roofing Company
      Texas
      +24%
      revenue (AI for estimates & docs)
      🛠
      Window Installer
      Utah
      +59%
      revenue (AI for proposals, 12+ months)
      🏗
      Construction Firm
      Florida
      +65%
      revenue (AI-driven operations)

      These aren’t Silicon Valley startups. This is a roofer in Texas.

      Macro validation

      46.6% of US businesses now pay for AIRamp, Feb 2026
      1.7% of revenue — corp AI spend doubled in 2026BCG AI Radar 2026
      90% of CEOs believe agents will deliver measurable ROI in 2026BCG
      88% of orgs use AI in at least one function (up from 78%)McKinsey State of AI
      11.5% avg productivity increase after 1+ year of AI useMorgan Stanley
      $2.5T projected worldwide AI spending in 2026Gartner
      $430M spent on foundation models Q4 2025 (+126% QoQ)Ramp
      94% plan to keep investing even without immediate ROIBCG

      THE GREAT REALLOCATION

      2024 “Klarna replaced 700 agents with AI” — paused all hiring, cut from 5,500 to 3,400
      2025 “Shopify CEO: prove AI can’t do the job before hiring” — AI in performance reviews
      WEF 92 million jobs could be eliminated by 2030”
      2026 55% of companies that executed AI-driven layoffs now regret it”

      Click “The Data” to see what actually happened →

      -92M
      jobs displaced by 2030
      +170M
      new roles created
      +78M
      net new jobs by 2030 — World Economic Forum

      Live market data — 9,000+ companies tracked

      Engineering
      67,000+
      3-year high
      US Engineering
      26,000
      accelerating
      Product Mgmt
      7,300+
      +75% vs 2023 low
      AI Roles
      Exploding
      fastest-growing
      Recruiters
      Near peak
      back to 2022 levels

      “If AI were killing tech jobs, recruiters wouldn’t be the hottest hire in the market.”

      Source: Lenny Rachitsky / TrueUp, 9,000+ companies, March 2026
      The Cautionary Tale
      Klarna
      • Replaced 700 CS agents with AI chatbot
      • Paused hiring, cut workforce 38%
      • Customer satisfaction cratered
      • Engineers pulled to staff phone lines
      • CEO: “Focused too much on efficiency”
      • Now rehiring humans for hybrid model
      The Blueprint
      Shopify
      • “Prove AI can’t do it before hiring”
      • Hiring MORE interns — best AI users
      • AI tool use correlates with higher performance
      • Built LLM proxy + MCP infra for all staff
      • Sales eng built AI dashboard — works only in Cursor
      • More prototypes, faster iteration, better work
      Sources: Bloomberg, Fast Company, First Round Review, TechCrunch, 2025–2026
      Company Valuation Revenue Multiple
      OpenAI $852B $25B ARR ~34x
      Anthropic $965B $47B run-rate ~21x
      Perplexity DVC ~$22.6B ~$450–500M ~45–50x
      Cursor $60B $4B ~15x
      xAI pre-merger $250B ~$500M ARR ~500x
      Sources: Reuters, CNBC, company disclosures and DVC analysis. Valuations and denominators as of the dates given in the sections above; the xAI row is a pre-merger, Q1 2026 snapshot. Private ARR is company-disclosed and unaudited.
      1

      Frontier Labs

      Pricing power depends on model scarcity — which is eroding fast

      2

      App Layer at 100x+

      Distribution advantages can evaporate when models improve

      3

      Moonshot Adjacencies

      Robotics, bio, energy — capital-intensive, long payoff horizons

      DOT-COM vs AI

      Dot-Com (2000) AI (2026)
      Revenue Minimal / speculative $25B+ ARR at frontier
      Enterprise Adoption Early experiments 53% of enterprises deploying
      Infra Spend Telco capex bubble $700B+ hyperscaler capex
      Moats Weak — eyeballs only Data, compute, distribution
      Source: DVC Research, historical market data

      WHY ARE MULTIPLES SO CRAZY?

      The answer isn't irrational exuberance — it's structural. Global AUM is heavily concentrated in a small number of very large managers who cannot write $5M seed checks, so capital funnels into the few AI companies big enough to absorb it. We no longer print the specific concentration statistic we previously carried: it had no source.

      TOP ASSET MANAGERS — $147T GLOBAL AUM

      Top 10 manage $62T — 42% of global AUM. They need to deploy billions per quarter. Early-stage VC is invisible to them.

      Source: BlackRock Q4 2025 earnings, Vanguard, UBS, Fidelity filings

      THE PARETO OF AI CHECKS

      AI captured roughly 61% of all global VC in 2025. The denominator is reported at both $427B and $440B depending on the tracker cut, so we give the share rather than assert one total. But within AI, concentration is extreme:

      5 companies raised $84B = 20% of ALL venture capital in 2025
      10 companies captured 41% of all venture dollars in H1 2025
      58% of AI funding was in Mega-rounds deals of $500M or more

      Largest AI Checks — 2025/2026

      Tap any bar for details. These 8 rounds alone total ~$250B+ — more than total global VC outside AI.

      Source: CNBC, Crunchbase, OECD, company announcements 2025–2026

      THE VALUATION CASCADE

      $147T Global AUM Can't do VC Pile into ~20 AI names Late-stage 50–100x Early-stage inherits

      The vast majority of investors can't write small, high-risk VC checks. They need to deploy at scale. So they pile into the ~20 private AI companies large enough to absorb $100M+ checks. Insane competition for a tiny number of deployable deals drives late-stage multiples up — which are then inherited by early-stage VCs.

      NEW ENTRANT: SOVEREIGN WEALTH

      Sovereign Wealth Funds deployed $46B into AI ventures in 2025. Saudi PIF, Abu Dhabi's Mubadala, Singapore's GIC and Temasek are now among the largest AI investors.

      Source: EY Global GenAI VC Report, Dec 2025

      THE GEOGRAPHIC FUNNEL

      97% of AI deal value went to North America. San Francisco Bay Area alone captured $122B — 75%+ of all US AI funding.

      Source: OECD, Crunchbase 2025 data

      Be skeptical on valuation. Not on AI demand. The top quartile is pulling away. The bottom quartile is standing still. The gap widens every quarter.

      Noise in pricing does not change the direction of the market. It makes disciplined positioning across the stack even more valuable.

      THE DVC ECOSYSTEM

      THE DVC THESIS

      The entire world's business is being rebuilt. Every industry — search, commerce, legal, healthcare, finance, property, video, code, robotics — is getting a new AI-native entrant. The incumbents have resources, but they can't do everything at once. They have organizational drag, legacy architectures, and cannibalization risk that slows them down.

      The opportunity for startups isn't to avoid competing with Google and OpenAI. It's to move faster in the seams they can't fill. Perplexity competes with Google Search. Cursor competes with GitHub Copilot. Higgsfield competes with Sora. In every case, the startup has the advantage of focus, speed, and a willingness to bet the company on a single wedge.

      DVC's portfolio is positioned across the stack because the winners won't cluster in one layer — they'll emerge wherever a startup can claim turf faster than an incumbent can defend it.

      89 Portfolio Companies
      6 Categories
      16 Sub-Verticals
      Across the full AI stack — from silicon to consumer apps

      A portfolio is a point of view made concrete. From here, the only question that matters is where the forces already in motion take us.

      5-YEAR OUTLOOK

      These are not predictions pulled from thin air. They are the logical consequences of every trend in this presentation — extrapolated five years forward. The pattern is clear: AI software moves fast, AI hardware moves slower, and AI adoption in large legacy industries moves slowest of all.

      1 FROM THE MODEL WARS

      Models Commoditize. Distribution Wins — and Distribution Now Means Records, Regulators and Institutional Accounts.

      REVISED AUGUST 3, 2026

      The window did not just confirm this — it changed the winning channel. The leadership order inverted (Anthropic $965B primary vs OpenAI $852B primary, with OpenAI secondaries implying ~$880–895B) and enterprise adoption flipped (Anthropic 42.4% of US businesses vs OpenAI 39.5%, Ramp AI Index, July 2026; overall business AI adoption 46.6%). The contested distribution surfaces stopped being app stores: OpenAI now connects hospital records and Apple Health inside ChatGPT, Mayo Clinic owns a frontier clinical model that Microsoft distributes, and OpenAI and Anthropic donated public-health enterprise seats through CHAI PULSE rather than selling them.

      Sources: Reuters · Ramp AI Index, July 2026 · OpenAI · Microsoft · CHAI

      The lower end of the model market is already interchangeable. By 2030, frontier-grade reasoning will be a utility. The durable value shifts to the companies that own the user relationship: distribution, UX, data moats, and workflow integration. Winning the model race matters less than winning the surface race.

      2 FROM THE INFRASTRUCTURE SPRINT

      Inference Under $0.001/Query by 2028

      Commodity-tier token prices have fallen by orders of magnitude since 2023 while frontier-tier list prices rose to $10/$50 — the barbell this report documents. We no longer print a single cost-decline series: the figures previously used here and in the inference section came from different model classes and windows and could not be reconciled to one current source. The bottleneck shifts from token cost to orchestration and verification. This unlocks entirely new use cases that were previously uneconomical.

      Source: Stanford HAI AI Index 2025
      3 FROM THE BUSINESS MODEL PROBLEM

      Outcome Pricing Has Been Invented. Margin Is the Unresolved Variable.

      REVISED AUGUST 4, 2026

      We are retiring “still being invented.” Salesforce now publishes per-resolution and per-action pricing, Cognizant reports new BPO contracts shifting toward outcome-based models, and Gartner estimates $234B ≈ 20% of enterprise application spend by 2030 is exposed to agentic arbitrage. The open question is no longer whether outcome pricing exists; it is whether the unit economics hold. We removed the unsourced multi-vendor price band and seat-pricing claims during the August 4 audit.

      Sources: Teqfocus (Salesforce) · Moneycontrol (Cognizant) · CIO (Gartner)

      Tokens are commoditizing, ads can't fund inference, and outcome-based pricing is promising but unproven. The winner of the next cycle may not be the best model or the best distribution — it may be the company that cracks the revenue model. Expect 2–3 more years of experimentation before dominant business models stabilize.

      4 FROM THE AGENT REVOLUTION

      ~35% of Knowledge Work Agent-Mediated by 2030

      13% of employees already use gen AI for 30%+ of tasks; 34% expect to within a year. But agent-mediated does not mean fully autonomous. Workers still own judgment. The pattern is "modular, fluid, AI-mediated" work — not replacement. The biggest change is how tasks are packaged, not who does them.

      Source: McKinsey 2025, WEF 2030 Scenarios
      5 FROM PHYSICAL AI

      Humanoids: Real but Slower — and the 2026 Gap Is Delivery

      REVISED AUGUST 4, 2026

      Humanoid funding reached $8.6B in 2026 by late July, while disclosed unit delivery remained far behind the capital raised. Agility’s announced SPAC had not closed, and several valuation, IPO-target and physical-AI revenue figures previously printed here could not be independently confirmed, so they were removed. The picks-and-shovels are shipping; humanoid robots are not yet shipping at comparable scale.

      Sources: Venture Post · TechCrunch

      The labour market is the prize, but commercial deployments remain concentrated in controlled industrial and warehouse settings. Autonomous vehicles are the clearest physical-AI category already generating recurring commercial activity at scale.

      6 FROM THE INCUMBENTS’ DILEMMA

      Large Legacy Players Are the Hardest to Transform

      Healthcare (~$371B on the CMS insurance and administration line, ~$800–900B of total provider-plus-payer admin drag), legal (69% of legal professionals now report using generative AI tools), finance (53% adoption) — adoption is underway but regulation, liability, and institutional inertia mean full transformation is a 10-15 year process, not 3-5. The biggest enterprises have the most data but the slowest approval cycles. This is why startups that target specific vertical wedges will outpace horizontal incumbents.

      7 FROM THE $700B SPRINT

      Energy Is the Binding Constraint — and It Is Being Met With Gas, Construction and Debt, Not Fission

      AUDITED AUGUST 4, 2026

      We hold the belief but remove the unsupported deal-count and capacity arithmetic previously printed here. The defensible conclusion is narrower: power availability, land, construction capacity and financing are becoming binding constraints on AI infrastructure. Hyperscalers are pursuing a mix of grid power, gas, renewables and nuclear contracts rather than waiting for one technology to solve the bottleneck.

      Source: FactSet

      The US leads in advanced chips but faces power and permitting constraints. China has more generation capacity but remains constrained on leading-edge silicon. The strategic race increasingly turns on which ecosystem can resolve its limiting resource fastest.

      8 FROM THE GEOPOLITICAL CHESSBOARD

      Two AI Stacks, One Contested Middle

      By 2030, US-allied and China technology stacks are likely to be more bifurcated, with countries in the middle balancing access, regulation and sovereignty. Advanced-chip manufacturing remains geographically concentrated, making the AI race a matter of national strategy as well as market competition.

      9 FROM THE ROBOTAXI RACE

      Autonomous Rides Scale 10× by 2030

      Waymo already delivers ~500,000 paid rides per week on a fleet of ~3,871 vehicles; Tesla, Zoox, and Waabi are entering the market. China moves quickly on permitting and city coverage, while the US leads in edge-case handling. On disclosed fleet size, Baidu’s reported 1,400+ robotaxis sit below Waymo’s federal filing. The first profitable autonomous transport networks may emerge before humanoid robots leave the warehouse at comparable scale.

      Source: Waymo, Baidu, company disclosures
      10 FROM THE THREE PILLARS

      Culture Is the Foundation. Distribution Is the Roof.

      Data, compute, and talent are necessary but insufficient. The AI winners of 2030 will be the companies that retain top researchers through volatility and own the surface where users interact. Researcher movement and repeated reorganizations make culture stability an increasingly important competitive signal.

      11 FROM THE RISK ANALYSIS

      Not a Bubble — But a Repricing With a Mechanism

      REVISED AUGUST 3, 2026

      We keep the conclusion and change the mechanism from “sentiment corrects” to “the private-to-public transition prices it.” Cerebras gave the first live mark-to-market on AI infrastructure and it is ~58% off its high while growing +94%. The pipe opened: Anthropic filed confidentially within days of its $65B round targeting a near-$1T listing; OpenAI submitted a confidential draft S-1 and by Jul 2 was in preliminary talks to grant the US government a ~5% stake (~$42.6B); SpaceX IPO’d around Jun 12 raising ~$75B above $2T; SK Hynix listed on Nasdaq around Jul 10. Concentration is now market structure: OpenAI + Anthropic took 43% of every H1 2026 venture dollar. Healthcare shows the same shape one layer down — fewer, bigger, later: Q2 deal volume at a decade low with median size up 54%. The shakeout will be administered by the IPO calendar, not by sentiment — and we have deliberately not added a bear slide to the deck.

      Sources: CNBC · Metir AI · Sacra · PitchBook

      AI spending is producing real value, but the gains are concentrating. Many highly valued AI startups will not justify their private marks. A correction can thin the field without invalidating the underlying technology shift, much as the dot-com repricing cleared the path for durable internet companies.

      12 FROM THE MODEL WARS + AGENTS

      Open Weights Lead on Scale and Ownability — Regulated Inference Stays Proprietary

      REVISED AUGUST 4, 2026

      We keep the belief and change the axis of the claim from cost to scale, sovereignty and institutional control. Kimi K3 at 2.8T took the parameter crown according to Moonshot’s own release; ownership and deployability matter as much as list price. Open weights are increasingly used as sovereign or regulated bases, while proprietary systems remain common for enterprise inference. We removed the unsupported DeepSeek parameter split and several secondary-source price comparisons during the August 4 audit.

      Sources: Moonshot AI · NIST / CAISI

      DeepSeek, Llama, Qwen, OpenClaw — open-weight models look close to closed ones on public benchmarks; held-out evaluations still show a meaningful capability gap at the frontier; we no longer print a specific month-gap figure, which could not be sourced to a NIST/CAISI publication, with real cost-efficiency wins for PRC labs. Baidu is open-sourcing Ernie. NVIDIA called OpenClaw “as big as Linux.” By 2030, the majority of production AI is likely to run on open-weight foundations. The durable value shifts from model weights to fine-tuning, orchestration, context engineering, and application-layer moats.

      13 FROM HEALTHCARE AI

      AI Follows Incentives. Efficiency Doesn’t Guarantee Lower Spend.

      NEW AUGUST 3, 2026

      The clearest natural experiment of the window ran in US healthcare. Ambient documentation can reduce documentation burden — and health plans project a 9.0% group / 8.5% individual 2027 medical cost trend, with ~70% of surveyed plans ranking provider AI documentation and coding among their top three cost inflators and ~20% calling AI the single biggest. PwC’s mechanism is “changes in severity, mix and amount per claim.” Read across to every other market: where the payment model rewards throughput and documented intensity, AI can raise efficiency and the bill. We removed the specific maternity-code and dollar estimates because the underlying payer publication could not be obtained.

      Sources: Axios · Fortune · Kocher, Zhao & Duffy, NEJM Catalyst via USC Schaeffer

      Full mechanics, the two reimbursement/liability gates and the two-speed adoption clock are in the Healthcare AI section.

      "The 2026–2030 AI story is less about whether AI works, and more about whether institutions — constrained by power, regulation, and labor — can absorb ultra-cheap intelligence quickly enough. The winners will have culture that retains talent, distribution that reaches users, and the discipline to invest across the full stack rather than betting on a single layer."

      The trends are clear. The question is who acts on them first.

      You are not late to a trend.

      You are early to a restructuring of the global economy.

      The reset is not to learn more AI jargon. It is to see the stack clearly, understand where value is shifting, and act before today's temporary leaders harden into tomorrow's incumbents.

      That is where DVC operates: across the system, across the cycle, and with a bias toward the layers that compound as AI moves from breakthrough to infrastructure.

      August 2026