AI Intelligence · 2026 Refreshed Aug 17
Independent Research · Refreshed August 17, 2026

The AI Technology Stack
A Definitive Landscape Report

A 12-layer analysis of the AI ecosystem — silicon to vertical applications — with competitive landscapes, market sizing, funding metrics, and strategic build opportunities across funding tiers. Refreshed against publicly available information as of August 17, 2026.

Layers analyzed 12
Charts 8
Ecosystem value $2T+
Data as of August 17, 2026
Updates 8 material changes in August 17 refresh (+29 prior)
As of August 17, 2026

The report body below has been corrected inline (strikethrough = old value, green = corrected value, red superscript links back to the relevant update here). Entries 30–37 are the August 17, 2026 refresh cycle; entries 20–29 (August 10, 2026), 07–19 (July 13, 2026) and 01–06 (May 26, 2026) are retained below.

07
Anthropic valuation: $380B (Series G) $965B (Series H)
Anthropic closed a $65B Series H at a $965B post-money valuation on May 28, 2026 — the highest-valued private AI company, narrowly ahead of OpenAI. Supersedes the $380B Feb 2026 Series G mark throughout the report.
Source: Fortune, Jun 1, 2026
08
Anthropic confidentially filed for IPO (Jun 1, 2026)
Filed with the SEC four days after the Series H close. Targets an October 2026 Nasdaq listing; Goldman Sachs, JPMorgan, and Morgan Stanley leading; expected raise $60B+. Replaces "targeting an H2 2026 IPO at a potential $900B+ valuation."
Source: Fortune, Jun 1, 2026
09
Anthropic ARR: ~$43B $47B disclosed; ~$69B est. (Jul 10)
Anthropic disclosed $47B run-rate revenue in mid-May 2026 (Sacra). Yipit estimates ARR reached ~$69B by July 10, with average daily net-new ARR accelerating from ~$400M to ~$550M. The report treats $47B as the disclosed floor and $69B as a third-party estimate.
Source: Sacra, May 2026; Yipit via NextBigFuture, Jul 10, 2026
10
Anthropic flagship lineup: Opus 4.6/4.7 Claude Fable 5 / Mythos 5
June 9, 2026: Anthropic released Claude Fable 5 and Claude Mythos 5, a new Mythos-class tier above Opus; Opus 4.8 is now the next-most-capable model. Fable 5: always-on adaptive thinking, 1M-token context, $10/$50 per M tokens (2x Opus 4.8). Layer 06 product row corrected.
Source: Anthropic; TechCrunch, Jun 9, 2026
11
SpaceX went public June 12, 2026 (NASDAQ: SPCX)
Largest IPO in history: $75B raised at a $2T+ valuation; the stock rose 20% in its first full session. References to SpaceX's S-1 and "~$1T pre-IPO" status are superseded; the $1.25T combined-entity figure from the Feb xAI acquisition is retained as historical.
Source: CNN Business, Jun 12, 2026; CNBC, Jun 15, 2026
12
SpaceX exercised $60B option to acquire Anysphere (Cursor)
June 16, 2026: four days after its IPO, SpaceX exercised its April 2026 option and agreed to acquire Cursor-maker Anysphere in a $60B all-stock deal, the largest acquisition of a venture-backed startup on record. Expected close Q3 2026 pending regulatory approval. The Layer 09 "dual offer" framing is superseded.
Source: Quartz; TechCrunch, Jun 16, 2026
13
Grok 4.5 released; xAI→SpaceXAI rebrand completed
July 8, 2026: SpaceXAI released Grok 4.5, pitched as an "Opus-class" coding/agentic model at $2/$6 per M tokens, trained with Cursor data. The @xai handle became @SpaceXAI on July 6; the Grok product brand is unchanged. Not yet available in the EU.
Source: Axios; TechCrunch, Jul 8, 2026
14
NVIDIA networking run-rate: $11B/qtr $15B/qtr
Q1 FY27 (reported May 2026): revenue $81.6B (+85% YoY), data center $75.2B (+92%), data center networking $15B (~3x YoY). New segment split: Hyperscale $38B, ACIE $37B. FY26 historical figures in the report are unchanged.
Source: NVIDIA Q1 FY27 results; TIKR; Futurum, May 2026
15
Cerebras post-IPO: ~$95B Day-1 ~$60B (Jul 13)
CBRS trades at ~$215 in mid-July, a ~$60B market cap, down ~31% from the $311.07 Day-1 close. The ~$95B Day-1 figure was fully diluted (~$66–67B basic); both retained as historical. Inference-silicon thesis intact; the froth has cooled.
Source: Yahoo Finance; stockanalysis.com, Jul 2026
16
OpenAI: confidential S-1 (Jun 8); ARR plateau; ~$965B secondary
OpenAI confidentially filed a draft S-1 on June 8, 2026 but is leaning toward a 2027 listing at a targeted ~$1T valuation. ARR held near $25B from Feb through Apr, ending the $2B/month net-new pace. Secondary marks ~$965B (Jul 10). Preliminary talks reported Jul 2 on a 5% US government stake (~$42.6B); no agreement reached.
Source: Forbes; Sacra; CNBC, Jun–Jul 2026
17
CoreWeave: FY26 revenue guide $12–13B; Meta commitments $35.2B
Q1 2026 revenue $2.08B (+112% YoY); FY26 guidance $12–13B; 2027 target $30B+ run-rate (75%+ contracted). Meta expanded its commitment by $21B on Apr 9 (total $35.2B through 2032); backlog $66.8B. Debt grew from $21B (end-2025) via an $8.5B March borrowing plus new convertibles; 2026 capex $30–35B (~$2.60 per $1 of new revenue). The "$5B revenue" body figure was FY2025.
Source: CNBC, Apr 9, 2026; CoreWeave Q1 2026 results
18
Hyperscaler 2026 capex: $700B+ ~$725B
Combined big-four 2026 AI capex now tracks ~$725B, up ~77% from ~$410B in 2025: Amazon ~$200B, Google ~$175–185B, Meta ~$115–135B, Microsoft ~$110–120B. 60%+ of spend goes to power and data center shells rather than chips.
Source: CNBC, Feb 2026; Yahoo Finance / valueaddvc compilations, Jul 2026
19
SaaSpocalypse: index-level recovery by early July
The public software basket is back to green for 2026 at the index level, but gains are uneven — many names remain far below January levels (HubSpot ~46% underwater despite the bounce). The Feb crash figures (~$285B erased; Thomson Reuters -15.83%) stand as historical.
Source: SaaStr; Yahoo Finance, Jul 2026
20
Hyperscaler 2026 capex re-guided at Q2: AMZN $200B ~$220B, GOOGL $180–190B $195–205B
Q2 2026 earnings (Jul 22–30) reset every big-four capex line. Alphabet raised 2026 capex to $195–205B on Jul 22 (third raise this year: $175–185B → $180–190B in April → $195–205B); Q2 revenue $119.8B (+24%), Google Cloud $24.8B (+82%), but free cash flow went negative at −$5.9B, a first for the quarter, and the stock fell on the print. Amazon lifted 2026 capex to ~$220B on Jul 30 from $200B, explicitly citing higher memory costs, with Jassy warning capacity still will not meet demand in 2026 or 2027. Meta narrowed and lifted its floor to $130–145B on Jul 29, with Q2 capex of $31.1B, revenue $60.8B (+28%), expenses up 55%, and the stock falling on thinning free cash flow. Microsoft guided calendar-2026 capex to ~$175B (down from ~$190B in April), but this is an accounting shift, not a spending cut: extending datacenter useful life from 15 to 25 years effective Q1 FY27 moves future leases from finance to operating, and operating leases fall outside capex. Underlying investment is unchanged and FY27 capex is still guided to grow, with Q1 FY27 alone above $50B. Big-four total now ~$720–745B (midpoint ~$730B), up ~78% from ~$410B in 2025. Supersedes update 18.
Source: Bloomberg and CNBC, Jul 22, 2026 (Ashkenazi); CNBC, Jul 30, 2026 (Jassy); Microsoft FY26 Q4 call, Jul 29, 2026 (Hood); Meta Q2 2026 call, Jul 29, 2026
21
Anthropic flagship lineup extended: Claude Opus 5 (Jul 24, 2026)
Anthropic shipped Claude Sonnet 5 on Jun 30 and Claude Opus 5 on Jul 24, 2026. Opus 5 holds Opus pricing at $5/$25 per M tokens — half the cost of Fable 5 ($10/$50) — while landing within ~0.5% of Fable 5 on agentic coding benchmarks, with a 1M-token context window, 128K max output, and a May 2026 knowledge cutoff. The practical effect is a sharp drop in the price of frontier-class agentic work, which pressures every layer that resells inference. Extends update 10.
Source: Anthropic; Axios and MarkTechPost, Jul 24, 2026
22
Fireworks AI: $225M+ raised $1.5B Series D at $17.5B; >$1B annualized revenue
Fireworks closed a $1.505B Series D at a $17.5B valuation on Jul 16, 2026 (Atreides, Index, TCV co-leading; NVIDIA, Lightspeed, Bessemer, Menlo, Insight participating), and disclosed crossing $1B in annualized revenue, 5× year over year, with daily token volume up from 15T to over 40T. This re-rates the entire inference-as-a-service segment: the Layer 08 "$225M+ raised" figure understated the category by an order of magnitude.
Source: Fireworks AI; CNBC and SiliconANGLE, Jul 16, 2026
23
SambaNova: $1B Series F first close at $11B valuation (Jul 8, 2026)
General Atlantic led a $1B Series F first close at an $11B post-money valuation, with T. Rowe Price, Capital Group, Qatar Investment Authority, Battery, and Seligman participating — five months after a $350M Series E in February 2026. Alongside Cerebras and Groq, this confirms sustained investor appetite for non-NVIDIA inference silicon. The Layer 01 "~$1.1B total raised" figure is superseded.
Source: Bloomberg and TechCrunch, Jul 8, 2026; General Atlantic
24
Harvey: $11B valuation, $190M ARR $15.5B round in talks on $350M annualized revenue
Harvey is in talks to raise ~$500M at a $15.5B valuation (The Information, Aug 7, 2026), with Lightspeed reported as circling to lead. Annualized revenue passed $350M, of which roughly $300M is ARR — up more than 80% since January, against ~$190M ARR at end-2025. The valuation path is $8B (Dec 2025) → $11B (Mar 2026) → $15.5B in talks (Aug 2026). Strongest single data point behind the "AI-native vertical SaaS" thesis, but treat it as unclosed: a Harvey spokesperson declined to comment, and the figure includes the new money.
Source: The Information, Aug 7, 2026; SiliconANGLE and PYMNTS, Aug 7, 2026; CNBC and Bloomberg, Mar 25, 2026
25
Cerebras: ~$60B (Jul 13) ~$64B (Aug 9); first post-IPO earnings Aug 12
CBRS traded at ~$228 on Aug 9, 2026, a ~$64B market cap, recovering modestly off the mid-July trough but still ~27% below the $311.07 Day-1 close. Its first full post-IPO quarter reports Aug 12, 2026 — the first hard read on whether wafer-scale inference revenue is compounding behind the multiple. Third-party market-cap figures diverge ($44B–$64B) depending on basic versus fully diluted share counts.
Source: stockanalysis.com and Investing.com, Aug 9, 2026
26
Anthropic IPO: investor roadshow scheduled Aug–Sept; October Nasdaq target intact
Following the Jun 1 confidential filing, Anthropic has scheduled institutional investor meetings for August–September 2026, with the public S-1 expected in late summer or early fall and pricing ahead of an October 2026 Nasdaq listing. Goldman Sachs, JPMorgan, and Morgan Stanley lead; expected raise remains $60B+. No ticker confirmed. Disclosed run-rate revenue stands at $47B (May 2026); the ~$69B July figure remains a third-party estimate. Extends updates 08 and 09.
Source: StartupHub.ai, Jul 15, 2026; Fortune, Jun 1, 2026; Sacra
27
SpaceX–Anysphere (Cursor): expected close Q3 2026 closing by end of August 2026
The $60B all-stock acquisition is now pending only final regulatory approval, with close expected no later than end of August 2026. Anysphere common and preferred convert to SpaceX Class A stock at the $60B implied value against SpaceX's seven-day VWAP immediately before closing. Deal protection: a $10B penalty if SpaceX fails to close for its own reasons and a $4B regulatory breakup fee. Layer 09 competitive structure changes materially on close. Extends update 12.
Source: BigGo Finance, Aug 2026; Quartz and TechCrunch, Jun 16, 2026
28
NVIDIA Q2 FY27 guidance: $91B ±2%, excluding China data-center compute
Guidance issued with Q1 FY27 results is $91B ±2% — the highest quarterly revenue guide in semiconductor history — and explicitly excludes any data-center compute revenue from China. Actual Q2 FY27 results are due in late August 2026 and are not yet reported; the report's Q1 FY27 actuals ($81.6B revenue, $75.2B data center, $15B networking) stand as the latest verified quarter.
Source: NVIDIA Q1 FY27 results, May 20, 2026
29
Governance: US Executive Order on frontier AI and cybersecurity (Jun 2, 2026); EU action plan (Jul 2026)
Executive Order "Promoting Advanced Artificial Intelligence Innovation and Security" (Jun 2, 2026) does three things: accelerates AI-enabled cyber defense across federal systems and critical infrastructure, creates a voluntary pre-release engagement framework giving government early access to covered frontier models, and directs the Attorney General to prioritize criminal enforcement against AI-enabled cyberattacks. Benchmarking deliverables were due August 1, 2026. The EU followed in July 2026 with a Cybersecurity and AI action plan coordinating Member State response to advanced models. The two regimes now diverge sharply — US voluntary and partnership-framed, EU prescriptive under the AI Act — which raises the compliance surface for any vendor selling into both. First governance entry in this report; see the revised Agent Compliance and AI-Security opportunity cards.
Source: Skadden, Latham & Watkins, and Wiley client alerts, Jun 2026; European Commission, Jul 2026

Prior update cycle — July 13, 2026:

Prior update cycle — May 26, 2026:

01
xAI dissolved as a standalone company
On May 6, 2026, Musk announced xAI ceases to exist as a separate entity and was rebranded as SpaceXAI, a division of SpaceX. The "Rank 5 model lab" entry in Layer 06 no longer reflects current corporate structure.
Source: Musk public statement, May 6, 2026
02
Anthropic-SpaceX $15B/yr compute deal
May 6, 2026: Anthropic struck a deal for full capacity of SpaceX's Colossus One data center ~300 MW, 220k+ NVIDIA GPUs. Per SpaceX S-1: $15B annual run-rate. This adds SpaceX as Anthropic's fourth major compute counterparty (alongside AWS, Google Cloud, CoreWeave).
Source: SpaceX S-1, May 2026
03
Anthropic ARR: $30B ~$43B
Sacra estimate for April 2026, up from $9B at end of 2025. The $30B figure reflects the Feb 2026 Series G disclosure; growth has been steeper than originally reported.
Source: Sacra, May 2026
04
OpenAI enterprise share: 25% 27%
Per the same Menlo Ventures Dec 2025 report cited in the original, OpenAI's enterprise LLM share fell to 27% from 50% in 2023, while Anthropic rose to 40% from 24%. The 25% figure was an undercount.
Source: Menlo Ventures Enterprise LLM Report, Dec 2025
05
Cerebras Day-1 market cap: $56-66B FDV ~$95B
Cerebras priced at $185/share on May 14, raised $5.55B, and closed Day 1 at $95B market cap, well above the $56-66B FDV range reported. Strong reception validates inference-silicon thesis more aggressively than originally noted.
Source: CNBC, May 14, 2026
06
NVIDIA-Groq: licensing deal, not acquisition
Announced December 24, 2025 (not early 2026). Structured as a non-exclusive licensing agreement, not a traditional acquisition. Groq Inc. remains nominally independent. The "$20B asset deal" phrasing is technically correct but the implied acquisition framing in surrounding context is not.
Source: Techi, Dec 24, 2025
30
Data-centre energy: 1,000 TWh in 2026 ~485 TWh (2025) → ~945 TWh (2030)
The 1,000 TWh-by-2026 figure carried since the April edition was a 2024-vintage IEA high-growth projection that did not materialise. Actual global data-centre consumption was ~485 TWh in 2025, and the IEA April 2026 base case now puts ~945 TWh at 2030, roughly 3% of global electricity. The report overstated 2026 consumption by roughly 2× and attributed a 2030-scale number to 2026. Corrected in both opportunity cards. The direction of the thesis is unchanged and the constraint argument actually strengthens, because the binding limit is interconnection queues rather than aggregate consumption.
Source: IEA, Energy and AI (Apr 2026); S&P Global, Apr 2026
31
Cerebras Q2 2026 reported Aug 12: $210M core revenue, FY26 guide raised
The prior edition flagged Aug 12 as the first hard read on wafer-scale economics. Result: $210M core revenue against ~$194M expected, Q3 guide $214–216M, and FY26 core revenue guidance raised to $880–890M. The stock nonetheless fell ~14% in extended trading. This confirms the report's existing framing precisely: the technology thesis is validated and the exit multiple continues to compress.
Source: CNBC, Aug 12, 2026; Cerebras 8-K
32
EU AI Act high-risk obligations deferred to Dec 2027 / Aug 2028
On June 16, 2026 the European Parliament approved amendments deferring Annex III use-based high-risk obligations from Aug 2026 to Dec 2027 and Annex I product-regulated obligations from Aug 2027 to Aug 2028. Article 50 transparency duties and AI Office enforcement over GPAI providers did take effect on schedule Aug 2, 2026. Material for Part 3: the Agent Compliance opportunity was upgraded to HIGH partly on EU prescriptive pressure, and the near-term EU gate is now disclosure and GPAI rather than high-risk certification. The US voluntary framework is unaffected.
Source: European Parliament, Jun 16, 2026; Holland & Knight, Apr 2026
33
AI-security consolidation: Lakera, Protect AI, Astrix, Robust Intelligence all acquired
Lakera was listed in the company database as an independent Layer 10 vendor; it was acquired by Check Point (2025). Also absorbed: Protect AI by Palo Alto Networks (completed Jul 22, 2025, now Prisma AIRS), Astrix by Cisco, and Robust Intelligence by Cisco. Cyera has signed a $1B letter of intent for Oasis Security. The pattern matters more than any single deal: independents in this category are being acquired before reaching scale.
Source: Palo Alto Networks press release, Jul 22, 2025; Check Point; Cisco
34
Zenity $125M Series C (Aug 3, 2026) — missed by the Aug 10 refresh
Zenity closed a $125M Series C led by Norwest on Aug 3, 2026, bringing total funding to ~$185M, with SoftBank Vision Fund 2, Hitachi Ventures, Qumra and LG Technology Ventures participating. The round closed seven days before the last refresh and was not captured. Zenity secures AI agent actions inside enterprises, which is the Layer 09 surface the Agent Compliance opportunity in Part 3 targets. Broader context: AI-native security overtook identity as the fastest-growing VC category in Q1 2026 at $4.1B, +47% YoY.
Source: BusinessWire, Aug 3, 2026; Intel Capital
35
Demand-side evidence added: 88% adopt, ~6% capture EBIT value
Part 3 adds the demand-side lens the report previously lacked entirely. Stanford AI Index 2026 puts enterprise adoption at 88% in at least one function but under 10% fully scaled in any function; McKinsey finds only 39% attribute any EBIT impact and ~6% report impact above 5%. Gartner projects 40%+ of agentic projects cancelled by 2027. This materially qualifies the supply-side valuations throughout the report.
Source: Stanford HAI AI Index 2026; McKinsey State of AI 2026; PwC Global AI Jobs Barometer 2026
36
Layer 03 database gap closed; +55 companies indexed
Layer 03 (Networking) previously had zero companies in the search index and knowledge graph despite eight appearing in its own competitive table, so every company modal rendered an empty row for it. Arista, Cisco, Astera Labs, Credo, Cornelis, DriveNets and Celestica are now indexed, along with companies named in report tables but previously unsearchable (Hugging Face, MongoDB, Perplexity, Replit, IBM, Datadog, vLLM and others) and the security vendors introduced in Part 4.
Source: internal database audit, Aug 17, 2026
37
Layer 02 capex headline corrected: $650–700B ~$725B
The Layer 02 Market Sizing headline still read $650–700B after the Q2 2026 hyperscaler figures were corrected inline in entry 20. Its own detail text now sums to roughly $732B (AMZN ~$220B, GOOGL $195–205B, META $130–145B, MSFT ~$175B calendar 2026), and the FinOps opportunity in Part 3 already cited ~$720–745B. The headline was the last place carrying the stale range. Found by cross-checking the new layer-sizing chart against the card it summarises.
Source: Q2 2026 hyperscaler earnings, Jul 22–30, 2026; internal consistency audit

Five Key Findings

1

The AI stack remains a $2+ trillion ecosystem in 2026, spanning 12 layers from silicon to verticals. Value concentration sits at silicon (NVIDIA) and the model layer (Anthropic, OpenAI); the application layer is ascending.

2

Anthropic overtook OpenAI in annualized revenue (~$43B$47B May · ~$69B est. Jul[9] vs. ~$25B) with 40% enterprise API share (Menlo Ventures, Dec 2025) vs. 25%27%[4] for OpenAI. Targeting an H2 2026 IPO at a potential $900B+ valuation. Closed a $65B Series H at $965B post-money (May 28); confidentially filed for IPO June 1, targeting an October 2026 Nasdaq listing.[7,8]

3

SpaceX acquired xAI in an all-stock deal on Feb 2, 2026: $250B standalone within a $1.25T combined entity. xAI subsequently dissolved as a standalone company on May 6, 2026, rebranded as SpaceXAI[1]. SpaceX itself went public June 12, 2026 in the largest IPO in history ($75B raised, $2T+ valuation, NASDAQ: SPCX) and moved to acquire Cursor-maker Anysphere for $60B four days later[11,12]. A fundamental reshaping of the model layer landscape.

4

Inference workloads now exceed 55% of AI cloud infrastructure spend. Cerebras IPO'd May 14, 2026 at ~$56-66B FDV~$95B Day-1 market cap[5] (cooled to ~$60B by July 13; ~$64B Aug 9)[15][25]; OpenAI's $20B+ equity stake validates inference-specific silicon at scale. Fireworks' $1.5B Series D at $17.5B on >$1B annualized revenue re-rates the serving layer above it.[22]

5

Feb 2026 SaaSpocalypse erased ~$285B in software market cap following Anthropic's Cowork launch. Thomson Reuters dropped 15.83% in a single session (largest on record). Per-seat SaaS → outcome/fractional-FTE pricing transition is underway. By early July 2026 the public software basket was back to green for the year at the index level, though the recovery is uneven[19].

The Five Findings, in Data

Six charts carrying the quantitative spine of this report. Each has a data table beneath it, and each names its basis. Where a bar is dashed, it is a projection rather than an actual.

The adoption cliff: 88% use AI, 6% capture value from it
Share of surveyed enterprises reaching each stage. Bar length encodes the percentage; the collapse between stage 3 and stage 5 is the finding.
Use AI in ≥1 function: 88%Use AI in ≥1 function: 88%Use AI in ≥1 function88%Use generative AI in ≥1 function: 70%Use generative AI in ≥1 function: 70%Use generative AI in ≥1 function70%At least experimenting with agents: 62%At least experimenting with agents: 62%At least experimenting with agents62%Scaling an agentic system: 23%Scaling an agentic system: 23%Scaling an agentic system23%Fully scaled AI in any function: <10%Fully scaled AI in any function: <10%Fully scaled AI in any function<10%EBIT impact >5% (“high performers”): ~6%EBIT impact >5% (“high performers”): ~6%EBIT impact >5% (“high performers”)~6%
Data table
StageShareSource
Use AI in ≥1 function88%Stanford AI Index 2026
Use generative AI in ≥1 function70%Stanford AI Index 2026
At least experimenting with agents62%McKinsey State of AI
Scaling an agentic system23%McKinsey State of AI
Fully scaled AI in any function<10%Stanford AI Index 2026
EBIT impact >5%~6%McKinsey State of AI
Stanford HAI AI Index 2026; McKinsey State of AI (agentic era, 2026).
Enterprise LLM API spend, $B
The 2026 figure is a projection and is drawn dashed. Spend roughly doubled in the six months to mid-2025.
0481115Late 2024: $3.5B$3.5BLate 2024Mid 2025: $8.4B$8.4BMid 2025End 2026 (proj.): $15B$15BEnd 2026 (proj.)$B
Data table
PeriodSpendBasis
Late 2024$3.5Bactual
Mid 2025$8.4Bactual
End 2026$15Bprojected
Menlo Ventures State of Generative AI / mid-year LLM market update.
Enterprise LLM API market share
Production workloads, late 2025. Anthropic leads by a wider margin in enterprise API usage than in consumer mindshare.
Anthropic: 40%40%OpenAI: 27%27%Google: 21%21%Other: 12%12%Anthropic 40%OpenAI 27%Google 21%Other 12%
Data table
ProviderShare
Anthropic40%
OpenAI27%
Google21%
Other12%
Menlo Ventures Enterprise LLM Report.
Market size by layer, $B
Value concentrates at the ends of the stack. Figures span three orders of magnitude on a single linear scale, so the smallest layers read as slivers and their direct labels carry the value. The middle layers, where most startups get built, are the smallest markets here by a wide margin.
12 Verticals: $2.5T12 Verticals: $2.5T12 Verticals$2.5T11 App Platforms: $315B11 App Platforms: $315B11 App Platforms$315B04 Cloud: $1.1T04 Cloud: $1.1T04 Cloud$1.1T02 Compute capex: $725B02 Compute capex: $725B02 Compute capex$725B01 Silicon: $500B01 Silicon: $500B01 Silicon$500B08 Inference: $106B08 Inference: $106B08 Inference$106B06 Models (funding): $80B06 Models (funding): $80B06 Models (funding)$80B05 Data Infra: $50B05 Data Infra: $50B05 Data Infra$50B03 Networking: $20B03 Networking: $20B03 Networking$20B10 Middleware: $8–10B10 Middleware: $8–10B10 Middleware$8–10B09 Orchestration: $7.6B09 Orchestration: $7.6B09 Orchestration$7.6B07 MLOps: $5–8B07 MLOps: $5–8B07 MLOps$5–8B
Data table
LayerMarket size
12 Verticals$2.5T
11 App Platforms$315B
04 Cloud$1.1T
02 Compute capex$725B
01 Silicon$500B
08 Inference$106B
06 Models (funding)$80B
05 Data Infra$50B
03 Networking$20B
10 Middleware$8–10B
09 Orchestration$7.6B
07 MLOps$5–8B
Per-layer sizing as cited in Part 1; mixed bases (TAM, capex, funding) noted in each layer's Market Sizing card. Bars use the report's fixed layer palette.
Global data-centre electricity, TWh
The dashed 2026 column is the superseded projection this report carried until entry 30. It did not materialise: actual consumption tracked far below it.
02505007501,0002024: 41541520242025 actual: 4854852025 actual2026 (old proj.): 1,0001,0002026 (old proj.)2030 IEA base: 9459452030 IEA baseTWh
Data table
YearConsumptionBasis
2024415 TWhactual
2025485 TWhactual
20261,000 TWhsuperseded projection
2030945 TWhIEA base case
IEA Energy and AI (Apr 2026); IEA Electricity 2024 for the superseded figure.
AI security is two markets, not one
Conflating them is the standard error in vendor reports. AI applied to security is mature and large; securing AI itself is small and compounding fast.
$0B$15B$30B$45B$60BAI applied to security 2026: $30–35B$30–35BSecuring AI itself 2026: $1.65B$1.65B2026AI applied to security 2032: ~$60B est.~$60B est.Securing AI itself 2032: $13.5B$13.5B2032AI applied to securitySecuring AI itself
Data table
SegmentYearSize
AI applied to security2026$30–35B (bar uses midpoint)
AI applied to security2032~$60B (extrapolated)
Securing AI itself2026$1.65B
Securing AI itself2032$13.5B
MarketsandMarkets agentic AI security (42% CAGR); Grand View / Precedence for AI-in-cybersecurity. The 2032 AI-applied figure is extrapolated from published CAGRs and is the least firm number on this chart.

12-Layer Architecture

Layer 01

Semiconductor / Silicon

Specialized chips—GPUs, TPUs, custom ASICs, and HBM—designed to accelerate the matrix math and parallel computation underpinning ML training and inference. The foundation of all AI capability.

Market Sizing

~$500B

AI chip revenue in 2026 (Deloitte). AMD CEO Lisa Su projects a $1T TAM by 2030. Broader semiconductor market surpassed $830B in 2025 (Omdia). NVIDIA FY26 data center revenue alone reached $193.7B; Q1 FY27 (May 2026): $81.6B total revenue (+85% YoY), $75.2B data center, networking $15B/qtr[14].

Key Dynamics

NVIDIA maintains an extraordinary moat through its CUDA software ecosystem (20+ years of developer investment). Custom silicon is eroding this at the hyperscaler level. The DeepSeek efficiency shock caused a $600B single-day drop in NVIDIA's market cap. Market is supply-constrained, not demand-constrained.

Inference-Specific Silicon Photonic Computing Custom ASICs for Hyperscalers HBM Memory
Funding Metrics — Layer 01Est. from public R&D / capex disclosures & named-player funding
Cumulative to Date
~$280B
NVIDIA + AMD + Intel R&D + Cerebras / Groq totals
2025 Funding
~$95B
R&D + private rounds (Cerebras, SambaNova)
2026 YTD
~$46B
Cerebras IPO $5.55B + Groq $20B deal + SambaNova $1B @ $11B (Jul 8)[23] + capex
YoY Growth
+58%
Driven by inference-silicon mega-rounds
#CompanyProductDifferentiatorRevenue / FundingSignal
1NVIDIAH200, B200, GB300, Vera RubinFull-stack: CUDA, NVLink, GPU-aware networking, software moatFY26 total $215.9B (+65%); DC $193.7BQ1 FY27 guide $78BQ1 FY27 actual $81.6B (+85% YoY)[14]; Q2 FY27 guide $91B ±2%, ex-China[28]; Vera Rubin H2 2026
2AMDMI300X, MI400, MI450 InstinctPrice-performance; ROCm open-source stackData center growing ~60% YoY$90B OpenAI deal for 6GW MI450 (Oct 2025)
3Google (TPU)TPU v6 TrilliumGCP integration; Transformer-optimizedInternal + cloud resaleAnthropic 3.5GW + $40B Google investment (Apr 2026)
4BroadcomCustom AI ASICs (XPUs)Bespoke silicon for Google, Meta, othersNetworking + ASIC revenue surgingMajor Anthropic TPU demand
5Amazon (AWS)Trainium 2/3, InferentiaCost-optimized training/inference; Bedrock integrationInternal + cloud resaleTrainium3 3× faster than T2
6IntelGaudi 3, Falcon Shoresx86 ecosystem leverageNew AI chips Oct 2025Restructuring underway
7CerebrasWafer Scale Engine 3Entire wafer as single chip; ultra-low latency$510M 2025 rev (+76%); $238M net incomeIPO May 14, 2026: $5.55B raised @ ~$56-66B FDV~$95B Day-1 close[5] (~$60B by Jul 13; ~$64B Aug 9; Q2 reported Aug 12: $210M core rev vs ~$194M est., FY26 guide raised to $880–890M, stock −14% after hours)[15][25]; OpenAI deal $20B+
8GroqLPU / Groq 3 LPXDeterministic latency; fastest tokens-per-sec$20B asset dealnon-exclusive licensing[6] w/ NVIDIA (Dec 24, 2025)Groq Inc. nominally independent; not acquired
9SK Hynix / Samsung / MicronHBM3e, HBM4High Bandwidth Memory for all AI acceleratorsHBM $100B+ by 2030$16B (2024) → $32.6B (2026)
10MarvellCustom AI acceleratorsCustom silicon for hyperscalersStrong data center growthAI-exposed revenue accelerating
Value-Chain Dependencies — Layer 01
Upstream
TSMC and Samsung foundry capacity, ASML EUV lithography, CoWoS advanced packaging, HBM supply from SK Hynix, Micron and Samsung.
Downstream
Sets the cost and availability floor for every layer above. L02 buildout timing is gated by accelerator allocation, not by capital.
Security surface
Hardware root of trust, confidential computing and TEEs on the accelerator itself, die and package attestation, firmware integrity. Blackwell is the first TEE-I/O capable GPU, and the throughput penalty for confidential mode fell from 30–40% to near parity, which is what moved it from research to procurement.
Regulatory exposure
Export controls are the binding regulatory constraint in the whole stack. A January 2026 Commerce rule reopened H200 and MI325X sales to China; late-May 2026 rules re-tightened on Blackwell, required licences for China- and Macau-headquartered entities, and closed the offshore-subsidiary loophole. No approved H200 units have actually shipped.
Power intensity
Perf-per-watt set here compounds through every layer above. A 2× efficiency gain at silicon is the cheapest possible relief on the L02 power constraint.
Sub-Layers — Silicon 6 sub-layers
Compute acceleratorsNVIDIA, AMD, TPU, Cerebras HBM / memorySK Hynix, Micron, Samsung Interconnect siliconBroadcom, Marvell Advanced packagingTSMC CoWoS, Amkor EDA toolsSynopsys, Cadence Chip IPArm, SiFive
Layer 02

Compute Infrastructure & Data Centers

Physical infrastructure — servers, racks, cooling, power delivery — that houses AI silicon and enables large-scale training and inference. Converts electricity and hardware into usable compute capacity.

Market Sizing

$650–700B ~$725B[37]

Combined 2026 hyperscaler capex (Q1 2026 earningsQ2 2026 earnings, Jul 22–30[20]): AMZN $200B+~$220B, GOOGL $180–190B$195–205B, META $125–145B$130–145B[20], MSFT ~$120B (or ~$190B incl. off-BS leases)~$175B calendar 2026 incl. finance leases — a lease reclassification, not a spending cut[20]. Big-four total now ~$720–745B, up ~78% from ~$410B in 2025; Alphabet's Q2 free cash flow turned negative at −$5.9B, and Amazon attributed its $20B raise to higher memory costs. Global data center energy consumption projected to reach 1,000 TWh in 2026 (IEA).

Key Dynamics

The neocloud segment has exploded. CoreWeave carries $21B~$29B+[17] in long-term debt against $5B$12–13B guided 2026 (Q1: $2.08B, +112% YoY)[17] revenue — model works only while GPU demand outpaces cost of capital. Power availability is the primary bottleneck; hyperscalers pursuing nuclear, geothermal, and direct energy partnerships. UPDATE: Anthropic-SpaceX struck a deal for full capacity of Colossus One (~300 MW, 220k+ GPUs) on May 6, 2026; ~$15B annual run-rate, adding SpaceX as Anthropic's fourth major compute counterparty[2].

Neoclouds / AI-Native Clouds Liquid Cooling Infrastructure Modular/Edge Data Centers
Funding Metrics — Layer 02Hyperscaler capex + neocloud rounds (10-Qs, company disclosures)
Cumulative to Date
~$1.6T
Aggregate hyperscaler DC capex 2020–2025
2025 Funding
~$420B
Capex + neocloud rounds (CoreWeave, Lambda, RunPod)
2026 YTD
~$215B
Of $650–700B guided 2026 capex
YoY Growth
+55%
Power-constrained but demand-led
#CompanyProductDifferentiatorRevenue / FundingSignal
1CoreWeaveGPU cloud (NVIDIA-focused)AI-first neocloud; purpose-built for GPU workloads$5.13B (2025); $12–13B guide (2026)Backlog ~$88–95B post Meta+$21B & Anthropic deals; NVIDIA $2B strategic investment Jan 2026
2EquinixColocation & interconnection260+ data centers; carrier-neutral$8B+ revenueAI-driven demand accelerating
3Digital RealtyData center REITLarge-scale campus deployments$5B+ revenueMulti-hundred MW AI campuses
4Lambda LabsGPU cloudH100 at $2.99/GPU-hr; developer-friendly~$500M run rate; $1.5B Series E24MW AI factory in Kansas City
5CrusoeAI compute + energyClean energy-powered data centersVC-backedSustainability resonating with enterprises
6NebiusGPU cloud (NVIDIA partner)European neocloud; Yandex spinoutGrowing in EUQ2 2025 +625% YoY
7RunPodServerless GPU cloudDeveloper-first; low-cost inference$120M ARR; 90% YoY growthScale on just $20M seed
8VultrCloud computeGlobal edge; cost-competitiveAcquired by DigitalOcean (2025)Expanding GPU capacity
Value-Chain Dependencies — Layer 02
Upstream
L01 accelerator allocation and HBM, plus grid interconnection rights and firm generation.
Downstream
Capacity and $/GPU-hour ceiling for L04 and L08. Scarcity here passes straight through to token prices at L06 and L08.
Security surface
Tenant isolation across multi-tenant GPU fleets, GPU memory residue between tenants, BMC and firmware, physical access. Neocloud tenancy models are generally less hardened than hyperscaler equivalents, and the audit evidence gap is widening.
Regulatory exposure
Siting, permitting and water. More than 200 data-centre bills were introduced across all 50 states in 2025 and over 40 were enacted. DOE has proposed federal jurisdiction over new loads above 20MW to standardise approvals.
Power intensity
The binding constraint for the entire stack. ERCOT alone received 198GW of large-load interconnection applications in Q1 2026 against a 438GW queue. Energization runs 4–7 years in major hubs and up to a decade elsewhere. Roughly 60% of hyperscaler capex is power and shells, not chips.
Sub-Layers — Compute & Data Centers 6 sub-layers
Hyperscaler DCsAWS, Azure, GCP AI-native neocloudsCoreWeave, Lambda, Crusoe ColocationEquinix, Digital Realty Power infraConstellation, Vistra, SMR CoolingLiquid, immersion, free-air Orbital DCsStarcloud, SpaceX
Layer 03

Networking & Interconnect

High-speed networking fabrics — InfiniBand, Ethernet, NVLink, optical transceivers — that connect GPUs within and across servers. Network bandwidth and latency directly constrain AI training speed and inference throughput.

Market Sizing

~$20B

AI networking in 2025 (650 Group). NVIDIA networking Q4 FY26 reached $11.0B (+263% YoY). Spectrum-X Ethernet surpassed a $10B annualized run rate. 1.6 Tbps switches enter volume deployment in 2026.

Key Dynamics

InfiniBand held ~80% of AI back-end networks in 2023; Ethernet surpassed it in 2025. UEC 1.0 (June 2025) provides InfiniBand-like RDMA on open-standard Ethernet. NVIDIA benefits either way. Data movement costs now exceed compute costs at scale.

Ultra Ethernet (UEC 1.0) Co-Packaged Optics (CPO) Scale-Across Networking
Funding Metrics — Layer 03Public-company networking segment revenue + private rounds
Cumulative to Date
~$45B
NVIDIA Mellanox + Broadcom AI networking + Arista DC + private rounds
2025 Funding
~$22B
NVIDIA networking $25B+ FY26 annualized run-rate
2026 YTD
~$13B
NVIDIA networking $11B (Q4 FY26) → $15B (Q1 FY27)[14]
YoY Growth
+150%+
NVIDIA networking +263% YoY in Q4 FY26
#CompanyProductDifferentiatorRevenue / FundingSignal
1NVIDIA (Mellanox)InfiniBand XDR, Spectrum-X, NVLinkFull networking stack; sub-2µs latency; GPU-aware routing$11.0B/quarter networking (Q4 FY26)+263% YoY; Spectrum-X at $10B+ run rate
2BroadcomTomahawk/Jericho switch ASICsBandwidth doubling every 2 years; 51.2 Tbps switchesDominant Ethernet switch ASIC shareLeading CPO deployment
3Arista NetworksData center Ethernet switchesLeading Ethernet DC switch vendorLeading position in DC switchingAI back-end networking outperformance
4CiscoAI networking portfolioEnterprise distribution; hybrid cloudLarge installed baseExpanding AI product line
5Marvell800G/1.6T optical DSPsLeading optical interconnect; custom PHYStrong 800G leadAI networking silicon growing rapidly
6CelesticaWhite-box switches/serversODM for hyperscaler AI networkingTop 3 AI networking vendorHyperscaler custom designs
7DriveNetsAI Fabric (Ethernet)Fabric-scheduled Ethernet for AI; SDNVC-backed growth stageOpen alternative to NVIDIA lock-in
8Cornelis NetworksOmni-Path (CN5000)Cost-competitive InfiniBand alternative at 400 GbpsNiche but growingPrice-sensitive HPC/AI deployments
Value-Chain Dependencies — Layer 03
Upstream
L01 interconnect silicon from Broadcom and Marvell, plus optics and transceiver supply.
Downstream
Caps effective cluster scale for L06 training. Tail latency introduced here surfaces directly as p99 latency at L08.
Security surface
East-west encryption at RDMA line rate, fabric segmentation, and materially different trust models between InfiniBand and Ethernet deployments. Encrypting at 800G without throughput loss remains expensive.
Regulatory exposure
Light and mostly inherited. Exposure arrives via L02 siting rules and via export controls on high-end networking equipment.
Power intensity
Optics and switching are a rising share of rack power as clusters scale, and are one of the few places where efficiency gains are still relatively cheap.
Sub-Layers — Networking 5 sub-layers
Intra-rackNVLink, UALink Intra-DC fabricInfiniBand, Ultra Ethernet Long-haulDark fiber, wavelengths Optical / CPOCoherent, Lumentum DPU / smart NICBlueField, Pensando, Nitro
Layer 04

Cloud Platforms & AI Services

Managed compute, storage, and AI-specific services that abstract infrastructure complexity. The primary distribution and monetization chokepoint for enterprise AI consumption.

Market Sizing

~$1.1T

Overall cloud market in 2026 (IaaS + PaaS + SaaS). Q4 2025 cloud infrastructure spending reached $119B (+29% YoY). GenAI-specific cloud services grew 140–180% in Q2 2025.

Key Dynamics

Big Three control 63% of the market. Google Cloud fastest-growing at 50%, driven by Gemini and Anthropic partnership. Multi-cloud adoption reached 89% of enterprises (Flexera 2026). Shift from training to inference favors elastic, serverless compute models.

GPU-as-a-Service (GPUaaS) Multi-Model AI Platforms Sovereign Cloud
Funding Metrics — Layer 04Hyperscaler AI capex + cloud platform private rounds
Cumulative to Date
~$2.1T
Cloud platforms cumulative capex + neocloud equity/debt
2025 Funding
~$510B
Big Three capex + Oracle + CoreWeave $2B IPO
2026 YTD
~$260B
Of $650–700B 2026 capex guide
YoY Growth
+45%
AI services growing 140–180% subset (Synergy)
#CompanyProductDifferentiatorRevenue / FundingSignal
1AWSBedrock, SageMaker, Trainium, InferentiaBroadest catalog (200+); largest GPU fleet~$120B annualized; 30% market shareQ1 2026 grew 28%; $200B~$220B[20] capex; 5GW Anthropic expansion
2Microsoft AzureAzure AI Foundry, Copilot StudioDeepest enterprise ecosystem (M365, GitHub)~$100B+; 22% market share39% Q4 growth; Claude Opus 4.6 in Azure
3Google CloudVertex AI, Gemini API, TPU accessBest-in-class AI/ML platform; TPU infrastructure~$50B+; 14% market shareBacklog $460B+ Q1 2026; 50% Q4 growth
4Oracle CloudOCI AI, GPU SuperclusterAggressive pricing; database ecosystemGrowing from smaller basexAI Colossus on OCI; Stargate partner
5CoreWeaveAI-native GPU cloudPurpose-built for AI; best GPU performance/$$5.13B (2025); $12–13B (2026)IPO'd March 2025; ~$88–95B backlog
6IBMWatsonxEnterprise AI governance; regulated industry focusAcquired DataStaxTargeting compliance-heavy enterprises
7Alibaba CloudPAI, Tongyi QianwenDominant in China; competitive pricingLargest cloud in APACGrowing AI services in APAC
Value-Chain Dependencies — Layer 04
Upstream
L02 capacity and L03 fabric.
Downstream
The primary distribution channel for L06 models and the default buying surface for everything from L08 to L12.
Security surface
Non-human and agent identity, key management, model endpoint access control. NHI is the fastest-consolidating security category in the stack: Cisco acquired Astrix, Cyera signed a $1B letter of intent for Oasis Security, and disclosed NHI rounds passed $340M over the last year.
Regulatory exposure
Data residency and sovereignty. This is the clearest transmission path for EU obligations into US vendors, because residency is enforced at the cloud boundary rather than at the model.
Power intensity
PUE and siting economics set the floor under every price above. Efficiency gains here are largely exhausted at the hyperscalers and still available at the long tail.
Sub-Layers — Cloud Platforms 5 sub-layers
Hyperscaler AIBedrock, Foundry, Vertex AI-native GPU cloudsCoreWeave, Lambda Sovereign / regionalG42, Mistral Compute Bare-metal rentalVoltage Park, Vast.ai Managed K8s for AIRun:ai, Anyscale
Layer 05

Data Infrastructure

Systems for storing, processing, indexing, and serving data to AI models — vector databases, data lakehouses, feature stores, labeling platforms, and synthetic data generators. Data quality is the primary determinant of AI application quality.

Market Sizing

$50B+

Broader AI data infrastructure market. Vector database market alone reaches $3.2B+ in 2026 (growing to $4.3B by 2028). Databricks at $134B valuation with $4.8B ARR (+55%); Snowflake FY26 revenue ~$4.2B.

Key Dynamics

Vector database market undergoing rapid commoditization. Traditional DBs (pgvector, MongoDB, Elasticsearch) adding native vector capabilities. 61% of enterprises delay AI initiatives due to lack of trusted data. Moat shifting from vector storage to unified data platforms.

Enterprise Context Platforms Synthetic Data Generation AI-Optimized Feature Stores
Funding Metrics — Layer 05Crunchbase / PitchBook public records + company press releases
Cumulative to Date
~$32B
Databricks $14B+ raised; Scale AI $1.6B + Meta $14.3B; vector DBs
2025 Funding
~$22B
Databricks $4B + Meta-Scale $14.3B + MongoDB Voyage $220M
2026 YTD
~$1.8B
Databricks acquisitions (Tecton, Antimatter, SiftD)
YoY Growth
−30%
Vector DBs cooling (Pinecone exploring sale); consolidation phase
#CompanyProductDifferentiatorRevenue / FundingSignal
1DatabricksLakehouse, Mosaic ML, LakebaseUnified analytics + AI training on one platform$134B valuation; $4.8B ARR (+55%); $4B Series L Dec 2025S-1 ready H2 2026; acquired Tecton, Antimatter, SiftD
2SnowflakeCortex AI, Arctic embeddingsData cloud with native AI/ML; enterprise baseFY26 revenue ~$4.2BAI features driving consumption growth
3PineconeServerless vector databaseFully managed; zero-ops vector search$138M total raised at $750M valuationReportedly exploring sale; Notion churned to pgvector
4MongoDBAtlas Vector Search + Voyage AIVectors alongside operational data$1.9B+ revenue$220M Voyage AI acquisition
5WeaviateOpen-source vector DBHybrid search (BM25 + vector); modular$67M+ raisedHIPAA certified
6Scale AIData labeling + evaluationHuman-in-the-loop; government contracts$29B valuation after Meta $14.3B for 49% stake (Jun 2025)Wang departed for Meta; Droege CEO
7QdrantOpen-source vector DBRust-native; performance-optimized$38M+ raisedFastest-growing open-source vector DB
8Zilliz / MilvusEnterprise vector DBBillion-scale vector search; Kubernetes-native$113M+ raisedMost scalable open-source option
9ChromaLightweight vector DBDeveloper-friendly; embeddable; simplest API$20M+ raisedDefault for prototyping
10Fivetran / AirbyteELT for AI pipelinesData connectors to unify enterprise data for AIFivetran: $5.6B valuationEssential plumbing for RAG and fine-tuning
Value-Chain Dependencies — Layer 05
Upstream
L04 storage and compute primitives.
Downstream
Quality ceiling for L06 training and L08 retrieval. 87% of enterprises still cite data readiness as the primary impediment to AI deployment.
Security surface
Training-data poisoning, provenance and lineage, and PII or PHI leakage into corpora. This is the layer the Enterprise Context Governance opportunity in Part 3 is built on.
Regulatory exposure
Copyright and lawful basis for training data, plus residency. Regulatory risk concentrates here more than at L06, because the exposure attaches to the corpus rather than the weights.
Power intensity
Storage and ETL draw is small relative to L02 and is not a meaningful constraint.
Sub-Layers — Data Infrastructure 7 sub-layers
Vector databasesPinecone, Weaviate, Qdrant LakehousesDatabricks, Snowflake Feature storesTecton Labeling / RLHF dataScale, Surge, Mercor Synthetic dataGretel, Mostly AI ETL / streamingFivetran, Confluent Governance / lineageAtlan, Collibra
Layer 06

Model Development (Foundation Models)

Research labs and companies that train large foundation models — LLMs, multimodal models, diffusion models — which serve as the base intelligence layer for all downstream applications. Highest concentration of capital and talent in the AI stack.

Market Sizing

$80B

Foundation model funding in 2025 — 40% of all global AI investment. Anthropic crossed ~$43B (Apr)$47B (May) · Yipit est. ~$69B (Jul 10)[9] ARR; OpenAI at ~$25B ($2B/month)(growth plateaued Feb–Apr)[16]. Top two captured 14% of all global VC across all sectors in 2025.

Key Dynamics

Anthropic overtook OpenAI in revenue at 40% enterprise API share (Menlo Ventures, Dec 2025) vs. 25%27%[4] OpenAI. Claude Code holds 54% market share. SpaceX acquired xAI (Feb 2, 2026) in an all-stock deal. Open-source models (Llama, Mistral, DeepSeek) compressing pricing power. Capital intensity extreme.

Reasoning / "Thinking" Models Small Language Models (SLMs) Domain-Specific Foundation Models
Funding Metrics — Layer 06Company press releases + Sacra + media reports of named rounds
Cumulative to Date
~$280B
OpenAI ~$80B + Anthropic ~$50B + xAI ~$30B + Mistral/Cohere/etc.
2025 Funding
~$80B
40% of all global AI investment (refresh PDF)
2026 YTD
~$165B
OpenAI $122B + Anthropic Series G + xAI $20B + Mistral €1.7B + Cerebras IPO
YoY Growth
+300%+
YTD already exceeds full FY2025 (annualized ~4×)
#CompanyProductDifferentiatorRevenue / FundingSignal
1AnthropicClaude Opus 4.6/4.7, Sonnet 4.6Claude Fable 5 / Mythos 5 (Jun 9), Opus 4.8[10], Sonnet 5 (Jun 30), Opus 5 (Jul 24, $5/$25 per M)[21], Claude Code, CoworkEnterprise-first; safety; coding dominance~$43B$47B+ (est. ~$69B Jul)[9] ARR; $380B (Feb 2026 Series G)$965B (May 28 Series H, $65B raised)[7] valuation40% enterprise share; targeting IPO H2 2026IPO filed Jun 1; Oct 2026 Nasdaq target[8]; roadshow Aug–Sep 2026[26]
2OpenAIGPT-5, o3, ChatGPT, CodexConsumer brand; broadest distribution~$25B ARR ($2B/month)(plateaued Feb–Apr)[16]; $852B valuation post $122B round (Mar 31, 2026); ~$965B secondary (Jul 10)[16]Stargate $500B JV; AWS $138B; AMD 6GW MI450; confidential S-1 Jun 8, leaning 2027 listing[16]
3Google DeepMindGemini 2.5/3, Gemini FlashFull-stack Google integration; TPU advantageInternal (Google subsidiary)Powers Workspace, Android, Search
4Meta AILlama 3/4 (open-weight)Open-weight models; community ecosystemInternal (Meta subsidiary)$35.2B CoreWeave deal (Sept 25 + Apr 26); $125–145B$130–145B[20] 2026 capex
5xAISpaceXAI[1]Grok 4Grok 4.5 (Jul 8, 2026)[13]Musk-backed; X data access$230B valuation (Jan 2026 Series E, $20B raised)Acquired by SpaceX Feb 2, 2026 (all-stock): $250B within $1.25T combined entity. Dissolved as standalone May 6, 2026; now a division of SpaceX. Parent SpaceX public Jun 12, 2026 (SPCX, $2T+)[11]
6Mistral AIMistral Large 3, PixtralEuropean champion; open-source roots; cost-efficient€11.7B valuation (Sept 2025 Series C, €1.7B led by ASML for 11% stake)$830M debt March 2026 for Paris DC; $400M+ ARR
7DeepSeekDeepSeek V3, R1Efficiency breakthroughs; open-weightGovernment-backedShook industry compute assumptions
8CohereCommand R+, North, Tiny AyaEnterprise + sovereign AI focus$7B valuation; $240M ARR (Q4 2025); ~$1.7B total raisedTargeting 2026 IPO; 70% gross margins; Saab partnership
9AI21 LabsJamba modelsMoE + Mamba hybrid architecturesRumored NVIDIA acquisition targetDifferentiated architecture
10Stability AI / Black Forest LabsStable Diffusion, FLUXOpen image/video generationRestructured; FLUX gaining shareLeading open-source generative media
Value-Chain Dependencies — Layer 06
Upstream
L01 silicon and L02 capacity set the training frontier; L05 sets the data ceiling.
Downstream
Capability and price-per-token propagate to every layer above. Opus 5 shipping at half of Fable 5 pricing re-rates the unit economics of L08 through L12 without those layers changing anything.
Security surface
Weight exfiltration, dangerous-capability evaluation, and pre-release red teaming. Gray Swan raised a $40M Series A in May 2026 and sits inside the pre-release evaluation processes at OpenAI, Anthropic and Meta, cited in 11 frontier model system cards. RAND's weight-security levels are becoming the reference framework.
Regulatory exposure
The most actively regulated layer. EU GPAI obligations and Article 50 transparency took effect on schedule August 2, 2026. The June 2026 US Executive Order created a voluntary pre-release framework with government early access.
Power intensity
Training runs are the largest single power events, though inference has now overtaken training in aggregate steady-state draw.
Sub-Layers — Model Development 7 sub-layers
Frontier closedAnthropic, OpenAI, GDM Open-weightLlama, Mistral, DeepSeek, Qwen Small / on-devicePhi, Gemma Domain-specificHarvey, Bloomberg GPT MultimodalImage, video, audio, 3D Reasoning / agentic World / robotics modelsPI, Skild, 1X
Layer 07

Model Optimization & MLOps

Tools for fine-tuning, distilling, quantizing, evaluating, and managing models post-training. Bridges the gap between a raw foundation model and a production-ready model tailored to specific enterprise needs.

Market Sizing

$5–8B

MLOps market in 2026, growing at 25–30% CAGR. Hugging Face at $4.5B valuation (1M+ models hosted). Weights & Biases acquired by CoreWeave (closed 2025), now operates as CoreWeave's MLOps software stack. LLMOps segment is nascent but growing explosively.

Key Dynamics

MLOps bifurcating: traditional (tracking, registry) is mature; LLMOps (RAG evaluation, guardrails) is nascent and fragmented. Biggest gap: evaluation — enterprises cannot reliably measure AI accuracy or safety. Serving a frontier model can cost $100M+/day in compute.

LLM Evaluation & Guardrails Model Distillation at Scale AI Safety Testing
Funding Metrics — Layer 07Aggregated public funding announcements for named players
Cumulative to Date
~$2.1B
HF $395M + W&B $250M + Anyscale + Arize + Braintrust
2025 Funding
~$420M
W&B/CoreWeave deal + Arize + Braintrust extensions
2026 YTD
~$120M
LLMOps category rounds (eval-focused players)
YoY Growth
+10%
Mature category; LLMOps subset growing faster
#CompanyProductDifferentiatorRevenue / FundingSignal
1Hugging FaceModel Hub, Transformers, Inference EndpointsDe facto standard for open model distribution; 1M+ models$4.5B valuation; $395M raisedCommunity lock-in; largest model registry
2Weights & BiasesExperiment tracking, model registry, W&B ModelsBest-in-class ML experiment trackingAcquired by CoreWeave (closed 2025)CoreWeave's MLOps stack
3Databricks (Mosaic ML)Model training + fine-tuning on lakehouseUnified data + model lifecycle; governancePart of $134B DatabricksTraining integrated into data platform
4Anyscale / RayDistributed compute framework for MLOpen-source; scales fine-tuning across GPUs$1B+ valuationUsed by OpenAI, Anthropic, and others
5NVIDIA NeMoModel customization frameworkNVIDIA hardware optimized; guardrails built-inPart of NVIDIA ecosystemEnterprise focus on customization
6Predibase / LudwigFine-tuning platformLow-code fine-tuning on open-source models$20M+ raisedDeveloper-friendly abstractions
7Arize / PhoenixML observability + LLM evaluationProduction monitoring for model drift and quality$78M+ raisedLLM evaluation becoming critical
8BraintrustLLM evaluation platformPrompt engineering + evaluation in one platform$60M+ raisedGrowing developer adoption
Value-Chain Dependencies — Layer 07
Upstream
L06 base models and L05 evaluation data.
Downstream
Determines what fraction of L06 capability actually reaches L08 at acceptable cost and latency.
Security surface
Model artifact supply chain, unsafe deserialization, and malicious models on public hubs. The model registry is effectively an unguarded package manager. Palo Alto absorbed Protect AI into Prisma AIRS in July 2025; HiddenLayer remains independent on $56M raised.
Regulatory exposure
Documentation, record-keeping and third-party assessment obligations land here first whenever high-risk rules do eventually apply.
Power intensity
Quantization and distillation are the primary demand-side lever on L02 power draw, and the only one that does not require new generation.
Sub-Layers — Model Optimization & MLOps 6 sub-layers
Fine-tuningTogether, Predibase, OpenPipe RLHF / RLAIF tooling Distillation / quantizationTensorRT, llama.cpp Experiment trackingW&B, Comet Eval & observabilityBraintrust, LangSmith, Arize Model registryMLflow, Vertex Registry
Layer 08

Inference & Serving

Systems that deploy trained models to serve predictions and generate outputs at production scale. This is where AI models meet real users and generate revenue — and the fastest-growing segment of the entire stack.

Market Sizing

$106B → $255B

AI inference market 2025 → 2030 at 19.2% CAGR. Inference surpassed training in data center revenue late 2025. By early 2026, inference crossed 55% of AI cloud infrastructure spending ($37.5B).

Key Dynamics

Every AI query costs $0.0016–$0.10 in GPU compute; at billion queries/day that's $1.6M–$100M/day. This is why Meta signed a $35B CoreWeave deal focused on inference. Inference-specific silicon thesis validated by Cerebras IPO at ~$56-66B FDV~$95B Day-1 market cap[5] (May 14, 2026), ~$60B by Jul 13, ~$64B Aug 9[15][25] with OpenAI's $20B+ equity commitment, and NVIDIA's $20B Groq asset deal. On the serving side, Fireworks raised $1.5B at $17.5B on >$1B annualized revenue (Jul 16, 2026), and Anthropic's Opus 5 cut frontier-class agentic pricing in half at $5/$25 per M tokens.[22]

Speculative Decoding & Optimization Edge Inference Inference Marketplaces
Funding Metrics — Layer 08Cerebras S-1 + Together/Baseten/Fireworks/Modal disclosures
Cumulative to Date
~$26B
Cerebras ($720M raised pre-IPO + $5.55B IPO) + Groq $20B deal + Together $534M
2025 Funding
~$1.8B
Together $400M+ + Fireworks + Baseten + Modal
2026 YTD
~$7B~$8.5B[22]
Cerebras IPO $5.55B + Together $1B + Fireworks $1.5B @ $17.5B
YoY Growth
+280%
YTD already ~4× full 2025; inference-silicon mega-rounds
#CompanyProductDifferentiatorRevenue / FundingSignal
1NVIDIA (TensorRT)TensorRT-LLM, Triton Inference ServerDeepest GPU optimization; lowest latency on NVIDIA hardwarePart of NVIDIA software revenueStandard for GPU inference optimization
2vLLM (UC Berkeley)Open-source LLM serving enginePagedAttention; highest throughput open-source servingOpen-source; community-maintainedAdopted by most major AI companies
3CoreWeaveManaged inference cloudGPU-optimized serving; distributed observability$12B+ projected 20269 of 10 leading model providers on platform
4AWS (Inferentia/Bedrock)Managed inference via BedrockCost-optimized custom silicon; serverless inferencePart of AWS AI revenueTrainium/Inferentia reducing costs 40%+
5Together AIInference API + open-source servingFast, cheap inference for open-source models$1B annualized rev (Mar 2026); in talks for $1B at $7.5BDeveloper community; strong pricing
6GroqLPU / Groq 3 LPX inferenceDeterministic latency; fastest tokens-per-sec$20B asset deallicensing deal[6] w/ NVIDIA (Dec 24, 2025)Nominally independent
7Fireworks AIInference optimization platformFireAttention; compound AI system serving$225M+ raised$1.5B Series D @ $17.5B (Jul 16, 2026); >$1B annualized revenue[22]Fastest inference for many architectures40T+ tokens/day, 5× YoY revenue; NVIDIA-backed[22]
8CerebrasWafer-scale inferenceUltra-low latency; single-chip model executionIPO May 14, 2026 @ ~$56-66B FDV~$95B Day-1[5] (~$60B Jul 13; ~$64B Aug 9)[15][25]; $20B+ OpenAI deal w/ ~11% equity$510M 2025 rev (+76%)
9BasetenModel serving platformGPU-optimized serving; Truss framework$100M+ raisedStrong developer adoption
10ModalServerless inferencePay-per-second GPU; developer-friendly$75M+ raisedPopular for batch and on-demand workloads
Value-Chain Dependencies — Layer 08
Upstream
L01 inference silicon, L02 capacity, L06 weights.
Downstream
Sets unit economics for all of L09 through L12. This is where guardrails actually execute, so it is also where policy becomes enforceable rather than advisory.
Security surface
Runtime prompt injection, output filtering, redaction at the serving boundary, and token flooding. Consolidated fast: Check Point acquired Lakera and Cisco acquired Robust Intelligence, leaving Prompt Security among the few independents.
Regulatory exposure
Article 50 content marking and deepfake labelling are enforced at the serving boundary, which makes this the layer where EU transparency duties actually bind.
Power intensity
Now over 55% of AI infrastructure spend and the dominant steady-state load. Inference efficiency, not training efficiency, is what determines the 2030 power curve.
Sub-Layers — Inference 6 sub-layers
Inference siliconCerebras, Groq, SambaNova Serving frameworksvLLM, SGLang, TensorRT-LLM Inference-as-a-ServiceTogether, Fireworks, Baseten, Modal Edge / on-deviceONNX, ANE, Qualcomm Marketplaces / routersOpenRouter, Martian Speculative decoding
Layer 09

Orchestration & Developer Tooling

Frameworks and tools that help developers build AI applications — chaining model calls, managing context windows, integrating retrieval, handling agent workflows, and connecting to external tools.

Market Sizing

$12.8B

AI coding tools revenue in 2026 (more than 2× the $5.1B 2024 figure). 85% of devs use AI tools, 73% regularly (JetBrains April 2026). Average developer uses 2–4 simultaneously. AI agent market $7.6B (2025) → $50.3B (2030) at 45.8% CAGR.

Key Dynamics

Three coding tool leaders: Claude Code (complex tasks), Cursor (IDE experience; $60B SpaceX acquisition pending[12], closing by end Aug 2026[27]), GitHub Copilot (enterprise deployment). 73% of developers use AI coding tools regularly. MCP (Model Context Protocol) crossed 97M installs by March 2026, becoming the standard for AI-tool connectivity.

Agent Orchestration Platforms Model Context Protocol (MCP) AI-Native IDEs
Funding Metrics — Layer 09Press releases (Anysphere, Vercel, Replit, LangChain, CrewAI, Windsurf)
Cumulative to Date
~$6.5B
Anysphere $2B+ + Vercel $250M + Replit $200M + Windsurf $3B exit + LangChain/Llama
2025 Funding
~$1.4B
Anysphere $1B+ + Vercel + LangChain extensions
2026 YTD
~$2B
Anysphere $50B-valuation round Anysphere: $60B SpaceX acquisition (Jun 16)[12] + Windsurf/OpenAI close
YoY Growth
+200%+
YTD already exceeds 2025; fastest-growing tooling layer
#CompanyProductDifferentiatorRevenue / FundingSignal
1AnthropicClaude CodeTerminal-native; 82% SWE-bench; 54% market share$2.5B+ ARR from Claude Code46% most-loved (JetBrains Apr 2026)
2AnysphereCursorAI-native IDE; 72% autocomplete acceptance$2B ARR; negotiating $50B valuation$60B acquisition value[12]70% Fortune 1000; SpaceX dual offer Apr 2026: $60B acq or $10B JVSpaceX exercised $60B all-stock option Jun 16, 2026; close expected Q3[12]
3GitHub (Microsoft)CopilotBroadest IDE support; 4.7M paid subscribers (+75% YoY)Microsoft Copilot suiteMoving to token-based billing June 1, 2026
4LangChainLangChain, LangGraph, LangSmithLargest orchestration ecosystem; 500+ integrations$130M+ raisedLangGraph → production standard for agents
5LlamaIndexLlamaIndex, LlamaCloudBest-in-class RAG; 300+ data connectors$50M+ raised44K GitHub stars; purpose-built retrieval
6MicrosoftAgent Framework (AutoGen + Semantic Kernel)Enterprise-grade; .NET/Python/Java; Azure integrationPart of MicrosoftGA Q1 2026; Azure-heavy enterprises
7CrewAIMulti-agent orchestrationRole-based multi-agent collaboration$40M+ raisedScaling enterprise customers
8Vercelv0 (AI app generator)Frontend-focused; Next.js ecosystem$250M+ raised; $3.5B valuationPioneered "vibe coding"
9ReplitReplit AgentBrowser-based AI coding; full-stack app generation$200M+ raised; $1.16B valuationAgent 3 builds entire web apps
10Windsurf (Codeium)AI code editorFree tier; broader model supportAcquired by OpenAI (~$3B)Competition intensifies
Value-Chain Dependencies — Layer 09
Upstream
L08 serving endpoints and L10 tool connectivity.
Downstream
Agent reliability here determines whether L11 and L12 products can credibly sell outcomes rather than assistance, which is the entire premise of the pricing transition in Finding 5.
Security surface
The fastest-moving attack surface in the stack. Agent permissions, tool-call sandboxing, egress control, human-in-loop gates. Autonomy is outpacing controls: Zenity raised a $125M Series C on August 3, 2026 explicitly against this surface.
Regulatory exposure
Audit trails and action logging. Both the June 2026 US EO and EU record-keeping expectations land here, which is what turned the Agent Compliance opportunity in Part 3 into a procurement gate.
Power intensity
Negligible directly, but multiplies L08 load: a multi-step agent trace can consume 10–100× the tokens of a single completion.
Sub-Layers — Orchestration & Dev Tooling 6 sub-layers
Agent frameworksLangChain, LlamaIndex, CrewAI Durable workflowsTemporal, Inngest, Restate Coding-agent IDEsCursor, Windsurf, Claude Code Browser / computer useBrowserbase, E2B Evals as codePromptfoo, Inspect AI notebooksHex, Marimo
Layer 10

Middleware & APIs

API gateways, model routers, authentication, rate limiting, caching, and cost management systems that sit between AI models and applications. The "plumbing" that makes AI reliable, observable, and cost-effective in production.

Market Sizing

$8–10B

API management market broadly. AI-specific middleware nascent at $1–2B. AI observability market projected to reach $3B+ by 2028.

Key Dynamics

40% of companies spend $10M+/year on AI, but cost visibility tooling is primitive. AI gateways routing queries to cheapest model gaining adoption as enterprises use 2–4 models simultaneously. LLM observability being absorbed by APM vendors. Middleware layer has weak moats — with one exception now forming: the shared context layer, where governance, lineage, and accumulated company knowledge create switching costs that gateways and routers never had.

AI Cost Management / FinOps AI Gateways & Routers Prompt Caching & Optimization Middleware
Funding Metrics — Layer 10Public funding announcements for AI gateway / observability players
Cumulative to Date
~$220M
Portkey + Helicone + LiteLLM + Mem0 + Zep + OpenRouter
2025 Funding
~$85M
Multiple seed/Series A across AI gateway category
2026 YTD
~$45M
Early-stage rounds; category still fragmented
YoY Growth
+30%
Sub-$1B segment; thin moats limit large rounds
#CompanyProductDifferentiatorRevenue / FundingSignal
1OpenAIAPI platform (direct)Largest model API by volume; developer ecosystemPart of OpenAI revenueDefault API for many applications
2AnthropicClaude API (direct + reseller)Highest enterprise API share (40%); strongest reasoningPart of Anthropic revenue40% enterprise share (Menlo Dec 2025)
3OpenRouterMulti-model API gatewaySingle API for 100+ models; auto routing and fallbackVC-backedDefault aggregator for model comparison
4PortkeyAI gateway + observabilityUnified API gateway with monitoring; guardrails$20M+ raisedGrowing in enterprise AI ops
5HeliconeLLM observabilityOpen-source LLM monitoring; cost tracking$15M+ raisedDeveloper-friendly; strong community
6LiteLLMOpen-source proxyUnified API across 100+ LLM providers; load balancingOpen-sourceDe facto standard for multi-model orchestration
7DatadogAI observabilityLLM monitoring integrated into existing APMPart of $2.2B Datadog revenueLeveraging enterprise relationships
8Mem0 / ZepAI memory managementPersistent memory for AI agents across sessions$10–20M raised eachSolving context management at scale
9CoconutShared context layer for AI agentsModel-agnostic, governed, versioned company context; every fact carries its source, and gaps are flagged rather than filled with a guessPrivate (Coconut AI Inc.)Targets the context-governance gap named in Part 4. Competitive set is platform-native, not peer: Claude Projects/Cowork, ChatGPT memory, and M365 Copilot can each bundle this away, with Glean ($300M ARR) the comparison every enterprise buyer makes
Value-Chain Dependencies — Layer 10
Upstream
L08 endpoints and L09 agent runtimes.
Downstream
Gateway policy and spend controls govern what L11 and L12 can safely expose to end users.
Security surface
MCP has no standard trust model. OX Security disclosed a systemic STDIO transport flaw in April 2026 allowing unsanitized OS command execution, with an estimated 200,000 vulnerable instances across a supply chain of 150M+ package downloads. The NSA has since published MCP security guidance, and the May 2026 enterprise-ready spec shifts responsibility to operators rather than solving it in-protocol.
Regulatory exposure
The natural enforcement point for both US voluntary and EU prescriptive evidence requirements, because the gateway is the only place that sees every call.
Power intensity
Negligible.
Sub-Layers — Middleware & APIs 7 sub-layers
Model gateways / routersPortkey, Kong, Cloudflare MCP servers / connectors Shared context / groundingCoconut, Mem0, Zep Agent identity / authArcade, Anon Rate limit / costHelicone, Langfuse Safety / guardrailsLakera, Protect AI API mgmt for AIApigee, Mulesoft
Layer 11

Application Platforms

Horizontal platforms that enable businesses to build, deploy, and manage AI-powered applications — including no-code/low-code AI builders, enterprise AI platforms, and agent deployment platforms.

Market Sizing

$315B

Global SaaS market in early 2026, growing at 20% CAGR toward $1.1T by 2032. AI-native SaaS spending increased 108% YoY. Gartner: 80% of enterprises deploy GenAI-enabled apps by end of 2026.

Key Dynamics

The Feb 3–5, 2026 "SaaSpocalypse" wiped ~$285B in market cap following Anthropic's Cowork launch. Thomson Reuters dropped 15.83% in a single session (largest on record); LegalZoom -19.68%; software ETFs ~-20% YTD by March. Publicis Sapient reducing SaaS licenses by ~50%. Shift from per-seat to outcome-based pricing. Vertical SaaS has stronger defensibility than horizontal.

AI Agent Platforms AI-Native App Builders Enterprise AI Governance Platforms
Funding Metrics — Layer 11Public SaaS R&D + private app-builder rounds (Lindy, Relevance, Bolt, Lovable)
Cumulative to Date
~$1.2B
Private app-builder & agent-platform rounds (excl. public SaaS)
2025 Funding
~$650M
Lovable, Bolt, Lindy, Relevance, Notion AI rounds
2026 YTD
~$430M
Agent-platform rounds accelerating post-SaaSpocalypse
YoY Growth
+90%
SaaSpocalypse repricing pulled capital into agents
#CompanyProductDifferentiatorRevenueSignal
1SalesforceEinstein AI, AgentforceCRM-native AI; massive enterprise install base$37B+AI agents across sales, service, marketing
2ServiceNowNow Assist, AI AgentsITSM/workflow automation with embedded AI$10B+Agent-to-agent communication protocols
3MicrosoftCopilot (M365, Dynamics, Power Platform)Deepest enterprise software integrationPart of $240B+ MicrosoftCopilot across entire product suite
4PalantirAIP (Artificial Intelligence Platform)Ontology-based AI; defense/intelligence pedigree$3B+; $250B+ market capFastest-growing enterprise AI platform
5GoogleWorkspace AI (Gemini)AI across Docs, Sheets, Gmail, MeetPart of Google Cloud3B+ Workspace users
6Notion AIAI-augmented productivityConsumer-friendly; horizontal knowledge management$150M+ ARRStrong developer/startup adoption
7Zapier / MakeAI-powered workflow automationNo-code automation with AI steps; 7,000+ integrations$230M+ ARR (Zapier)AI automation becoming core use case
8Relevance AI / LindyAgent builder platformsNo-code AI agent deployment for business users$10–20M raised eachDemocratizing agent creation
Value-Chain Dependencies — Layer 11
Upstream
L09 orchestration, L10 connectivity, L05 enterprise context.
Downstream
Distribution channel into the enterprise seat base. This is the surface the February 2026 SaaSpocalypse repriced.
Security surface
DLP for copilots, oversharing through enterprise search, and shadow AI. Oversharing is consistently the most-cited blocker to copilot rollout, and it is a permissions problem inherited from L05, not a model problem.
Regulatory exposure
Workplace surveillance, accessibility and employment law. Underrated exposure, because it attaches to the deploying enterprise rather than the vendor.
Power intensity
Negligible.
Sub-Layers — Application Platforms 6 sub-layers
SaaS reimaginedAgentforce, Now Assist, Breeze Agentic work platformsCowork, Copilot Studio No-code AI buildersv0, Bolt, Lovable Comms / collab AISlack AI, Notion AI, Glean Productivity suitesM365 Copilot, Workspace AI Digital twins of peopleViven, Simile, Delphi
Layer 12

Vertical End-User Products

AI applications built for specific industries — legal AI, healthcare AI, financial AI, sales AI — that deliver complete solutions to end users. This is where AI value is ultimately captured.

Market Sizing

$2.5T

Global spending on AI-enabled applications projected in 2026, growing 44% YoY. Vertical AI software grows from $133.5B (2025) to $194B (2029) — the largest segment by revenue.

Key Dynamics

Vertical AI has the strongest moat potential in the entire stack — domain-specific data flywheels, regulatory certification, workflow lock-in. Winning formula: frontier model + proprietary domain data + purpose-built workflow automation. Harvey's leap from $5B to $11B and into talks at $15.5B on $350M annualized revenue by August 2026[24] is the clearest signal of moat strength. The pattern now extends down-market: against Ironclad at $200M ARR and a $3.2B mark, SpotDraft reached 700+ customers and ~$380M by selling contract lifecycle management to mid-market legal teams the enterprise vendors price out, and differentiating on an architecture rather than a model — the document is processed on the endpoint NPU and never reaches a server. Pricing evolving toward outcome-based and fractional-FTE models.

Agentic Vertical AI AI-Native Content & Media AI for Regulated Industries
Funding Metrics — Layer 12Harvey, Glean, Perplexity, Abridge, etc. funding announcements
Cumulative to Date
~$8.5B
Harvey $1B + Glean $765M + Perplexity $1.5B+ + Abridge + others
2025 Funding
~$3.2B
Glean Jun 2025 + Perplexity Series E + Abridge mid-2025 rounds
2026 YTD
~$1.8B
Harvey $200M Mar 2026 + Perplexity E-6 + Hexus acquisition
YoY Growth
+125%
Vertical AI valuations re-rating sharply (Harvey 2.2× in 9mo)
#CompanyProductDifferentiatorFunding / ValuationSignal
1HarveyAI for legalPurpose-built for law firms; trusted by top firms$11B valuation (Mar 2026, $200M co-led GIC+Sequoia); $190M ARR EOY 2025$15.5B in talks (Aug 2026, ~$500M raise, unclosed); $350M annualized revenue, ~$300M of it ARR[24]100K+ lawyers; 25K+ custom agents; Hexus acquired Jan 2026
2PerplexityAI search engineAnswer engine with citations; consumer + enterprise$20–22B valuation (early 2026 Series E-6); ~$200M ARRSnapchat partnership Jan 2026; $750M Azure commit
3GleanEnterprise AI searchPermissions-aware knowledge graph; connects all enterprise tools$7.2B valuation (Jun 2025); $200M ARR (Dec 2025, doubled in 9 months); $765M raised100M+ agent actions/yr; targeting 1B
4WriterEnterprise AI platformBrand-safe content generation; governed AI$1.9B valuationFortune 500 customers
5JasperAI marketing platformMarketing content generation at scale$1.7B valuationPivoting to enterprise workflows
6AbridgeAI clinical documentationReal-time clinical conversation summarization~$5.3B valuation (mid-2025 rounds)Healthcare-specific compliance
7EvenUpAI for personal injury lawAutomated demand letter generation$1B+ valuationDeep vertical specialization
8Viz.aiAI for radiology/stroke detectionFDA-cleared AI diagnostic imaging$500M+ valuationRegulatory moat
9Suki AIAI medical documentationVoice-enabled clinical documentation assistant$500M+ valuationPhysician adoption growing
10HebbiaAI for knowledge workEnterprise document analysis; financial/legal focus$700M+ valuationMatrix reasoning architecture
11IroncladEnterprise contract lifecycle managementCategory leader for enterprise legal; deepest workflow and integration surface$200M ARR (Jan 2026, +34% YoY); $3.2B valuation; $334M raisedNo new round since the Jan 2022 Series E; crossed $100M ARR in 2024 and doubled it in two years
12SpotDraftAI-native contract lifecycle management (VerifAI)Mid-market CLM sold beneath Ironclad and Icertis. Deliberate split: contract review, clause extraction, risk scoring, redlining, and playbook application run on the Snapdragon X Elite NPU inside Word, so the document never reaches a server; login, licensing, collaboration, and analytics stay cloud-side~$92M raised; ~$380M post-money (Jan 2026 Series B extension, Qualcomm Ventures)700+ customers, up from ~400 a year earlier; 1M+ contracts/yr, volume +173% YoY; ~100% revenue growth guided for 2026
Value-Chain Dependencies — Layer 12
Upstream
Everything below. Most vertical vendors are thin on infrastructure and deep on workflow, which is why they are the least power-exposed and most regulation-exposed layer.
Downstream
Terminal layer. Where realized ROI either materialises or does not, and therefore where the demand-side evidence for the whole stack has to come from.
Security surface
The spine inverts here. L01 through L11 are about securing AI; L12 is where AI is sold as security. AI-native security overtook identity as the fastest-growing VC category in Q1 2026 at $4.1B, up 47% year over year.
Regulatory exposure
Sectoral certification per vertical is the moat named in the top-ranked opportunity in Part 3: state-by-state insurance rules, HIPAA, FINRA, FDA.
Power intensity
Negligible.
Sub-Layers — Vertical End-User Products 10 sub-layers
CodingCursor, Claude Code, Copilot LegalHarvey, Ironclad, SpotDraft, EvenUp HealthcareAbridge, Ambience, OpenEvidence FinanceHebbia, Rogo, Numerai SupportSierra, Decagon, Parloa Sales / marketingClay, 11x, Regie EducationMagicSchool, Khanmigo, Speak Defense / govtechPalantir, Anduril, Scale Donovan Creative / mediaRunway, ElevenLabs, Suno, MJ Robotics / physical AIFigure, 1X, PI

What Buyers Actually Did

Part 1 is supply-side: who builds what, layer by layer. It does not answer the question an investment committee asks first, which is whether any of this is working for the people paying for it. This part is the demand side, and it is deliberately less flattering than the rest of the document.

The headline is a gap. Enterprise adoption of AI is close to universal and enterprise value capture from AI is rare. Both facts are well evidenced, they are not in tension, and the distance between them is the single most important number in this report.

Adoption
88%
of organisations use AI in at least one business function. Effectively saturated.
Scaled
<10%
have fully scaled AI in any single function. Adoption is broad and shallow.
Any EBIT impact
39%
attribute any EBIT impact at all to AI. The majority cannot point to one.
High performers
~6%
report EBIT impact above 5%. This group captures a disproportionate share of total value.
Only 39% of enterprises attribute any EBIT impact to AI
Expected versus realized return. The 171% figure is anticipated ROI from an executive survey; the 39% and 6% are self-reported outcomes.
Anticipate positive agentic ROI: 171% expectedAnticipate positive agentic ROI: 171% expectedAnticipate positive agentic ROI171% expectedAttribute any EBIT impact to AI: 39%Attribute any EBIT impact to AI: 39%Attribute any EBIT impact to AI39%EBIT impact above 5%: ~6%EBIT impact above 5%: ~6%EBIT impact above 5%~6%
Data table
MeasureValueSource
Anticipate positive agentic ROI171% (expected)PagerDuty exec survey
Attribute any EBIT impact39%McKinsey
EBIT impact >5%~6%McKinsey
McKinsey State of AI 2026; PagerDuty survey of 1,000 executives. Note the first bar is an expectation, not an outcome, and is shown in a different colour for that reason.

What separates the 6%

McKinsey's high performers are three times more likely to have fundamentally redesigned workflows around AI, and three times more likely to be scaling agents in a function. The differentiator is not model choice, spend, or technical sophistication. It is willingness to change the process rather than bolt AI onto the existing one. This is the single most actionable finding in the report for an operator, and it argues that the binding constraint on enterprise AI value is organisational, not technological.

The counter-signal

Gartner projects that more than 40% of agentic AI projects will be cancelled by 2027, citing unclear ROI and weak governance rather than capability shortfalls. Read alongside the 171% ROI executives say they anticipate, this is an expectations problem: buyers are underwriting returns the deployed systems are not yet producing. For vendors this is a near-term revenue risk that does not show up in any ARR chart, and it is the demand-side reason the governance opportunities in Part 3 are rated as highly as they are.

AI skills wage premium and job growth
Jobs requiring AI skills are growing roughly eight times faster than the overall market, and command a 62% wage premium.
Job growth: AI-skilled roles: +69%Job growth: AI-skilled roles: +69%Job growth: AI-skilled roles+69%Job growth: all roles: +9%Job growth: all roles: +9%Job growth: all roles+9%Wage premium, AI skills (2026): +62%Wage premium, AI skills (2026): +62%Wage premium, AI skills (2026)+62%Wage premium, AI skills (2025): +57%Wage premium, AI skills (2025): +57%Wage premium, AI skills (2025)+57%
Data table
MeasureValue
AI-skilled role growth+69%
Total jobs market+9%
Wage premium 2026+62%
Wage premium 2025+57%
PwC 2026 Global AI Jobs Barometer, based on over one billion job advertisements across 27 countries.
What this means for the investment case

The supply side of this stack is priced for broad enterprise value capture that has not yet occurred. That does not invalidate the infrastructure thesis, because inference demand is real and power-constrained regardless of whether buyers can attribute EBIT. But it does mean the application and vertical layers carry more execution risk than their funding multiples imply, and it makes the workflow-redesign and governance layers structurally more attractive than a pure capability play. The 6% who capture value did so by changing how work happens. Whoever sells that change profitably is positioned better than whoever sells the model.

Where to Build

Bootstrappable <$5M
● HIGH

AI Cost Management & FinOps for LLM Operations

Gap: No turnkey solution for real-time AI spend optimization across multi-model environments. Enterprises using 2–4 LLMs simultaneously cannot intelligently route queries to the cheapest model, forecast inference costs, or identify prompt-level waste.

Pain: 40% of companies spend $10M+/year on AI; Cloud Efficiency Rate dropped from 80% → 65%. 500–1,000% cost underestimation at scale (Gartner). Q2 2026 sharpened this: big-four capex reached ~$720–745B, Amazon raised its guide $20B on memory costs alone, and Alphabet posted negative Q2 free cash flow[20]. That cost base gets passed through in token prices. The countervailing force is Opus 5 at half of Fable 5 pricing[21] — per-token costs are falling while total spend rises, which is precisely the condition that makes FinOps tooling necessary rather than optional. TAM: ~$2–4B within 3 years.

Business Case

3-person team with FinOps background builds "Cloudflare Workers meets Apptio for LLMs." Monetize via % of savings identified. 5-year outcome: $500M+ ARR acquired by Datadog or ServiceNow.

● HIGH

RAG Pipeline Quality Evaluation & Testing

Gap: No turnkey RAG evaluation solution across hybrid retrieval strategies. Enterprises cannot systematically measure retrieval quality, chunk relevance, or end-to-end answer accuracy.

Pain: 87% cite data readiness as primary AI impediment. Arize, Braintrust, and Ragas offer partial solutions only. Watch the ceiling: retrieval evaluation is being absorbed into context-layer platforms that own the corpus and can score grounding natively. The standalone window is narrowing — build for the evaluation harness, not the retrieval stack. TAM: ~$1–2B within 3 years.

Business Case

Small ML team builds "Datadog for RAG" — automated quality scoring, regression testing, alerting. $500–$5,000/month per pipeline. 5-year: $200M ARR, acquired by Databricks or Datadog.

● HIGH ↑ from MEDIUM-HIGH

AI Agent Compliance & Audit Trails

Gap: No standardized solution for logging, auditing, and proving compliance of AI agent actions in regulated industries. Current frameworks (LangGraph, CrewAI) have minimal built-in governance.

Pain: Gartner: 40% of enterprise apps will embed AI agents by end of 2026. GRC software is a $15B market. Upgraded this cycle: the June 2, 2026 US Executive Order created a voluntary pre-release framework with government early access to frontier models and August 1 benchmarking deliverables, while the EU's July 2026 Cybersecurity and AI action plan pushes the opposite, prescriptive direction[29]. Any vendor selling into both now needs dual-regime evidence, and nobody produces it off the shelf. Regulation moved this from a nice-to-have to a procurement gate. TAM: ~$1–3B~$2–5B within 3 years.

Business Case

4–6 engineers with GRC or audit background build the evidence layer: immutable action logs, policy attestation, and export packs mapped to both the US voluntary framework and EU AI Act Article 12 record-keeping. Sell to the compliance buyer, not the platform team. 5-year: $150–300M ARR, acquired by a GRC incumbent (OneTrust, Drata) or ServiceNow.

● MEDIUM-HIGH

MCP Integration Platform

Gap: MCP is becoming the standard for connecting AI models to enterprise tools, but no dominant platform simplifies building, testing, and deploying MCP servers securely.

Pain: By end of 2026, successful SaaS launches will advertise "MCP-native." Building MCP servers requires custom engineering per tool. TAM: ~$500M–1.5B within 3 years.

● MEDIUM-HIGH · NEW

On-Device Inference for Privileged & Regulated Workflows

Gap: Legal, healthcare, and financial workflows handle text that is privileged, PHI, or MNPI. The default AI architecture ships that text to a third-party API. There is no standard toolkit for running extraction, classification, and risk scoring entirely on the endpoint while keeping quality near frontier-model levels.

Pain: SpotDraft proved the demand shape, and the useful detail is the split rather than the slogan. Its VerifAI runs embeddings, clause extraction, risk scoring, redlining, and playbook application on the Snapdragon X Elite NPU end-to-end inside Microsoft Word, so the document never reaches a server; login, licensing, collaboration, orchestration, and fleet analytics stay in the cloud. That is the replicable pattern — keep the privileged artifact on the endpoint, keep the control plane central — and Qualcomm Ventures funded the company to scale it. Two honest constraints: it currently lands on a narrow hardware footprint (Snapdragon X Elite laptops), and the tooling around it is hand-rolled. Quantization, offline evaluation, model update distribution, and proving to an auditor that inference stayed local are all bespoke today. The gap is that toolchain, not the silicon. TAM: ~$1–2B within 3 years.

● MEDIUM

Domain-Specific Fine-Tuning Data Marketplaces

Gap: Enterprises wanting to fine-tune models for specific domains (legal, medical, financial) lack curated, high-quality training datasets. Options are expensive custom labeling or noisy web scrapes.

Pain: Regulated industries cannot use general web data. Synthetic data tools lack domain depth. TAM: ~$500M–1B within 3 years.

Venture-Scale $5M–$50M+
● HIGH

Enterprise Context & Data Governance Layer for AI

Gap: Context engineering tools exist (LangChain, Mem0, Zep) but no unified platform governs where that context comes from. Result: 66% of enterprises get biased or misleading AI insights; 57% duplicate AI efforts across departments.

Pain: 88% have a formal context strategy, but 87% cite data readiness as a significant impediment. 61% frequently delay AI initiatives. TAM: ~$5–10B within 5 years.

Status change — the category is now contested. Coconut is building exactly this: a shared, living, model-agnostic context layer where every fact carries its source, answers are identical across Claude, ChatGPT, Copilot, and Slack, and missing context is flagged rather than guessed. Qontext raised $2.7M pre-seed in Berlin on the same thesis (HV Capital, with the founders of n8n, Neo4j, Celonis, and make.com participating). The thesis is validated; the open window is narrower than it was in July. Differentiation now has to come from governance depth and lineage, not from the idea.

Business Case

15–30 engineers build "Okta for AI context" — governance platform sitting between enterprise data systems and AI apps. $50–200/user/month. 5-year: $500M+ ARR, IPO or acquired by Databricks/Snowflake for $5–10B.

● HIGH

AI-Native Vertical SaaS for Regulated Industries

Gap: Regulated industries (healthcare, financial services, legal, government) need complete, compliant AI solutions combining frontier model capabilities with domain-specific workflow automation and regulatory certification.

Pain: Vertical SaaS growing at 23.9% CAGR — nearly double the broader SaaS market. Harvey, Abridge, EvenUp prove the pattern. Most verticals remain underserved. Strengthened this cycle: Harvey is in talks at $15.5B on $350M annualized revenue (~$300M ARR), a 41% valuation step in five months[24]. The more useful signal for a new entrant is below the top: SpotDraft reached 700+ customers and ~$380M valuation by taking contract lifecycle management to mid-market legal teams that Ironclad and Icertis price out, and won on deployment speed and on-device privacy rather than model quality. The frontier lab is not the competitor; the incumbent's implementation cycle is. TAM: ~$10–30B across verticals within 5 years.

Business Case

20–40 people with domain expertise build "Harvey for Insurance" — AI automating claims processing from intake to settlement. State-by-state regulation = enormous barrier to entry. 5-year: $1B+ ARR, 60%+ gross margins, pursuing IPO.

● HIGH

AI Agent Operations Platform (AgentOps)

Gap: No platform manages the complete agent lifecycle — deployment, monitoring, debugging, cost optimization, safety testing, rollback — as enterprises deploy autonomous agents at scale.

Pain: 66% of engineering teams report AI outputs look correct but fail during testing. Current tools (LangSmith, Arize) offer observability but not full lifecycle management. TAM: ~$3–8B within 5 years.

Business Case

15–25 engineers with APM/DevOps background build "Datadog for AI Agents." $0.01 per agent action monitored. 5-year: $300M–500M ARR; standard infrastructure for enterprise agent operations.

● MEDIUM-HIGH

Multimodal AI for Physical Industries

Gap: Physical industries (manufacturing, construction, agriculture, logistics) lack AI systems integrating vision, sensor data, robotics control, and language models into unified operational platforms.

Pain: Manufacturing AI adoption <15% despite $14T output. Edge inference hardware (Qualcomm, NVIDIA Jetson) is now capable. Capital is arriving: Atoms, Travis Kalanick's physical-AI company, raised $1.7B led by a16z in July 2026 — one of the largest private rounds of the month and a signal that the category is moving from research to buildout. TAM: ~$15–25B within 5 years.

● MEDIUM-HIGH

AI-Powered Cybersecurity for AI Systems

Gap: As AI handles sensitive data and makes autonomous decisions, new attack surfaces emerge — prompt injection, data poisoning, model extraction. No comprehensive AI-specific cybersecurity solution exists.

Pain: Critical injection defense is a top enterprise concern. Defense tooling is primitive. The June 2, 2026 Executive Order put AI-enabled cyber defense on the federal procurement path and directed criminal enforcement against AI-enabled attacks; the EU's July action plan targets the same surface from the resilience side[29]. Federal demand signal now exists where it did not in July, which shortens the enterprise sales cycle for anyone with a credible product. TAM: ~$3–5B within 5 years.

Deep-Pocketed $100M+
● HIGH

Next-Generation AI Compute Infrastructure (Post-GPU)

Gap: GPU-based compute is hitting power efficiency walls. Photonic computing, neuromorphic chips, and wafer-scale integration could deliver 10–100× efficiency gains, breaking the power constraints limiting data center scaling.

Pain: Data centre energy projected to reach 1,000 TWh in 2026reached ~485 TWh in 2025; IEA base case ~945 TWh by 2030[30]. Power availability is the #1 bottleneck, and the binding limit is interconnection queues rather than aggregate consumption. Cerebras IPO at ~$56–66B FDV (May 2026)Cerebras trades at ~$64B (Aug 9, 2026) after a ~$95B Day-1 close, with Q2 reported Aug 12: $210M core revenue against ~$194M expected and FY26 guidance raised to $880–890M, yet the stock fell ~14% after hours — technology thesis validated, multiple still compressing[31][25] and OpenAI's $20B+ equity stake validate the inference-specific silicon thesis at scale. SambaNova's $1B Series F at $11B (Jul 8) confirms private capital is still funding non-NVIDIA architectures[23]. Caution: the public comp has round-tripped ~27% off its Day-1 close, so the thesis is validated at the technology level but the exit multiple is no longer the one that priced in May. TAM: ~$50–100B within 7 years.

Business Case

Deep-tech startup develops photonic AI accelerators — 50× energy efficiency over GPUs for inference. Requires $200M+ for tape-out. 5-year: acquired by NVIDIA/AMD/Intel for $10B+, or becomes "TSMC of photonic computing."

● HIGH

Sovereign AI Infrastructure Stacks

Gap: Nations outside the US (EU, India, Saudi Arabia, Japan, South Korea) need full-stack sovereign AI: data centers, custom silicon, foundation models, and apps tailored to local languages and regulations.

Pain: EU AI Act requires data residency; Saudi Arabia invested $100B+ in AI infrastructure; Mistral raised €1.7B from ASML for European AI sovereignty. TAM: ~$100B+ within 5 years.

Business Case

Consortium partners with sovereign wealth fund to build full-stack AI platform for Middle East or Southeast Asia. Combines local data centers with Llama/Mistral models + vertical AI apps. 5-year: $5B+ revenue with government-backed demand certainty.

● HIGH

AI-First Enterprise Software Suite (Horizontal)

Gap: No company has built a comprehensive AI-first enterprise suite that replaces Salesforce + ServiceNow + SAP with an agent-native architecture from the ground up.

Pain: $285B market rout when Anthropic's Cowork plugins launched. Publicis Sapient reducing SaaS licenses by ~50%. The $315B SaaS market is being restructured. TAM: ~$50–100B within 7 years.

Business Case

Palantir-scale ambition — AI-native enterprise OS that replaces point-solution SaaS. Agents handle CRM, ITSM, HR, finance, legal through a single intelligence layer. $500M+ capital required. 5-year: "next Salesforce" at $50B+.

● MEDIUM-HIGH

Full-Stack AI for Drug Discovery & Biotech

Gap: AI has shown breakthrough potential (AlphaFold), but no company has built a full-stack platform taking a drug target from identification through lead optimization, clinical trial design, and regulatory submission.

Pain: Drug development costs $2.6B per approved drug; clinical trial failure rates exceed 90%. AI could reduce timelines 30–50% and costs 50%+. TAM: ~$30B+ within 7 years.

● HIGH ↑ from MEDIUM-HIGH

AI Energy Infrastructure (Nuclear/Geothermal for Data Centers)

Gap: AI data centre power consumption projected to reach 1,000 TWh by 2026~485 TWh in 2025 rising to ~945 TWh by 2030 (IEA base case)[30]. Grid capacity is the binding constraint: ERCOT alone received 198GW of large-load interconnection applications in Q1 2026 against a 438GW queue, with 4–7 year energization timelines in major hubs. Dedicated clean energy (SMRs, geothermal, grid-scale storage) for AI compute creates massive value.

Pain: Microsoft revived Three Mile Island; Google signed geothermal deals; power is the #1 bottleneck cited by data center operators. Upgraded this cycle on direct operator testimony: Andy Jassy said on July 30 that even at ~$220B of 2026 capex Amazon will not have enough capacity to meet demand in 2026, and expects the same in 2027[20]. When the largest buyer states publicly that money is not the constraint, the constraint is power and shells. Roughly 60% of hyperscaler capex already goes to power and data center shells rather than chips. TAM: ~$50B+ within 7 years.

Business Case

Developer secures interconnect queue positions and firm generation (SMR, geothermal, or gas-plus-storage bridge) and sells power-plus-shell as a bundled product to hyperscalers and neoclouds on 15-year contracts. Capital-intensive and permitting-bound, but demand risk is close to zero through 2030.

Top 5 Opportunities Across All Tiers

#OpportunityTierRatingKey Rationale
1AI-Native Vertical SaaS for Regulated IndustriesVenture ($5–50M)HIGHHarvey at $11BHarvey in talks at $15.5B on $350M annualized revenue, ~$300M ARR (Aug 2026, unclosed)[24]; SpotDraft proves the mid-market tier; Abridge $5.3B
2Enterprise Context & Data Governance LayerVenture ($5–50M)HIGHDatabricks' Tecton/Antimatter/SiftD acquisitions confirm consolidation; Coconut and Qontext now building it directly — validated but contested
3AI Cost Management / FinOps for LLMsBootstrappable (<$5M)HIGHCopilot on token billing; ~$720–745B big-four capex meets Opus 5 at half of Fable 5 pricing — spend up, unit cost down[20]
4AI-First Enterprise Software SuiteDeep-Pocketed ($100M+)HIGH$285B SaaSpocalypse proved replacement thesis at market scale
5Next-Gen Compute Infrastructure (Post-GPU)Deep-Pocketed ($100M+)HIGHCerebras IPO at ~$56–66B FDVCerebras ~$64B (Aug 9) + SambaNova $1B @ $11B; thesis intact, multiple compressed[25]

Recommended "Where to Play"

Limited Capital <$5M

AI Cost Management or RAG Evaluation

Both are horizontal, model-agnostic, and reach revenue quickly. AI FinOps has the largest addressable market and weakest incumbent coverage. Newly competitive alternative: Agent Compliance & Audit Trails, upgraded to HIGH this cycle — the June 2026 Executive Order and the EU July action plan turned governance evidence into a procurement gate with no off-the-shelf answer.[29]

Venture Capital $5–50M

Vertical AI for a Regulated Industry

Moat from regulatory certification, proprietary data, and workflow integration is the strongest in the stack. Go one tier below the marquee names: SpotDraft won mid-market legal on deployment speed and on-device privacy, not model quality. Second priority: Enterprise Context Governance — the "missing layer" every enterprise AI deployment needsstill the missing layer, but no longer unoccupied — Coconut and Qontext are building it now, so enter on governance depth and lineage rather than on the idea.

Deep Capital $100M+

AI-First Enterprise Software Suite

Largest market restructuring since the cloud transition. Requires world-class execution and 5–7 year horizon. Focused bet alternative: Sovereign AI infrastructure — near-certain demand backed by government budgets. Lower-variance alternative added this cycle: AI Energy Infrastructure, upgraded to HIGH after Amazon stated that ~$220B of 2026 capex still will not meet demand through 2027. Power, not capital, is the binding constraint.[20]

Avoid Across All Tiers

Pure-play model development — capital requirements are prohibitive; Anthropic and OpenAI have insurmountable leads.  ·  Pure-play vector databases — commoditizing rapidly.  ·  Undifferentiated AI middleware — thin wrappers with no moat.

Three Spines Through the Stack

Part 1 reads the stack by altitude: each layer consumes the one below and supplies the one above. Three forces do not respect that structure. Security, regulation and energy cut vertically through all twelve layers, and none of them can be located at a single altitude. Treating any of them as a thirteenth layer would be a category error, because a layer is defined by what it consumes and what it supplies, and these consume and supply at every level at once.

Each spine below is the stack transposed: one row per layer, read top to bottom. The pattern each one traces is the finding. Security runs as a defensive concern from L01 to L11 and then inverts at L12, where AI is sold as security rather than secured. Regulation concentrates at the two ends and thins in the middle. Energy binds hard at the bottom four layers and is close to irrelevant above them, which is itself worth seeing at a glance rather than asserting.

Security & Trust

Spine 01

Market Sizing

$1.65B → $13.5B

Two markets, not one, and conflating them is the standard error. AI applied to security is mature at roughly $30–35B in 2026. Securing AI itself is far smaller and far faster: MarketsandMarkets puts agentic AI security at $1.65B in 2026 reaching $13.5B by 2032 (42% CAGR), while Dell'Oro sizes AI systems security at close to $8B by 2030 from a standing start in 2024.

Key Dynamics

Consolidation is running ahead of category formation. Palo Alto absorbed Protect AI, Check Point took Lakera, Cisco took both Robust Intelligence and Astrix, and Cyera signed a $1B letter of intent for Oasis. Independents are being acquired before reaching scale, which implies the durable position in most of this spine is a feature inside a platform rather than a standalone product. The exception is the agent layer, where Zenity's $125M Series C in August 2026 shows capital still funding independents. For an investor the read is straightforward: at L07 and L08 you are underwriting an acqui-exit, at L09 and L12 you are still underwriting a company.

LayerExposure at this altitudeWho sells hereSignal
01Hardware root of trust, TEEs on-accelerator, die and package attestationNVIDIA Confidential Computing, Intel TDX, AMD SEV-SNP, Arm CCABlackwell is the first TEE-I/O capable GPU; confidential-mode throughput penalty fell from 30–40% to near parity
02Tenant isolation on shared GPU fleets, memory residue, BMC and firmwareFortanix, Edgeless Systems, operator-native controlsNeocloud tenancy is less hardened than hyperscaler; audit evidence gap widening
03East-west encryption at RDMA line rate, fabric segmentationNVIDIA BlueField DPUs, Cisco Hypershield, AristaEncryption at 800G without throughput loss remains unsolved cheaply
04Non-human and agent identity, KMS, endpoint access controlAstrix (Cisco), Oasis (Cyera $1B LOI), Entro, GitGuardian, WizFastest-consolidating category; NHI rounds passed $340M in the last year
05Training-data poisoning, provenance and lineage, PII and PHI leakageCyera, Varonis, BigID, Unity CatalogUnderpins the Enterprise Context Governance opportunity in Part 3
06Weight exfiltration, dangerous-capability evals, pre-release red teamingGray Swan, Irregular, Haize Labs, lab-internalGray Swan $40M Series A (May 2026); cited in 11 frontier model system cards
07Model artifact supply chain, unsafe deserialization, malicious hub modelsHiddenLayer, Prisma AIRS (ex-Protect AI), JFrog MLProtect AI absorbed by Palo Alto Jul 2025; HiddenLayer independent on $56M
08Runtime prompt injection, output filtering, redaction, token floodingPrompt Security, Lakera (Check Point), Robust Intelligence (Cisco)Two of the three leading independents acquired; this layer is now platform-owned
09Agent permissions, tool-call sandboxing, egress control, human-in-loop gatesZenity, Invariant Labs, E2B, BrowserbaseZenity $125M Series C (Aug 3, 2026); capital still backing independents here
10MCP server trust and authn, gateway policy, spend limits as a controlCloudflare AI Gateway, Portkey, KongOX Security disclosed ~200,000 vulnerable MCP instances (Apr 2026); NSA guidance since issued
11DLP for copilots, oversharing via enterprise search, shadow AIMicrosoft Purview, Netskope, Island, Witness AIOversharing is the most-cited blocker to copilot rollout; a permissions problem, not a model problem
12AI applied to security operations — detection, triage, responseCrowdStrike Charlotte, Microsoft Security Copilot, Dropzone, Prophet, TorqDropzone $57.4M / 100+ enterprise customers; Prophet $41M; all three majors shipped agentic SOC at RSAC 2026
Where value concentrates

L04, L09 and L12. Identity (L04) and agent control (L09) are the two surfaces where enterprises are actively writing cheques and where no incumbent yet owns the category. L12 is a genuine vertical with real revenue. Everything else is either already absorbed into platform vendors (L07, L08) or inherited from general infrastructure security (L02, L03), and should be priced as a feature rather than a market.

Regulation & Governance

Spine 02

Market Sizing

Two diverging regimes

Not a market so much as a procurement gate and a compliance cost. The adjacent GRC software market is roughly $15B. The material fact for 2026 is divergence in kind: the US June 2026 Executive Order is voluntary and access-based, while the EU is prescriptive. Any vendor selling into both now needs dual-regime evidence, and nobody produces it off the shelf.

Key Dynamics

The EU timeline just moved, and it moved in the direction that reduces near-term pressure. On June 16, 2026 the European Parliament deferred most high-risk obligations: Annex III use-based duties slip from August 2026 to December 2027, and Annex I product-regulated duties from August 2027 to August 2028. What did take effect on schedule on August 2, 2026 is Article 50 transparency and AI Office enforcement over general-purpose models. The practical consequence is that near-term obligations are about disclosure and GPAI, not high-risk certification, which pushes the compliance-tooling opportunity in Part 3 further out than the July reading implied.

LayerExposure at this altitudeWhat appliesSignal
01Export controls, entity lists, fab and tooling restrictionsJan 2026 Commerce rule reopened H200 / MI325X to China; late-May 2026 rules re-tightened on Blackwell and closed the offshore-subsidiary loopholeChinese vendors took ~41% of China's AI accelerator server market in 2025, up from NVIDIA's ~95% share in 2022 (IDC)
02Siting, permitting, water disclosure, large-load interconnection200+ state data-centre bills in 2025, 40+ enacted; DOE proposed federal jurisdiction over loads above 20MWRegulatory risk is now local and political, not federal
03Export controls on high-end networking equipmentInherited from L01 and L02Low direct exposure
04Data residency and sovereigntyEU residency requirements enforced at the cloud boundaryThe main transmission path for EU obligations into US vendors
05Copyright, lawful basis for training data, residencyActive litigation across jurisdictionsExposure attaches to the corpus, not the weights, so it concentrates here rather than at L06
06GPAI obligations, transparency, pre-release evaluationEU Article 50 and GPAI enforcement took effect Aug 2, 2026 as scheduled. High-risk obligations deferred: Annex III to Dec 2027, Annex I to Aug 2028 (Parliament vote, Jun 16, 2026)US June 2026 EO is voluntary and access-based; the two regimes are diverging in kind, not just in degree
07Documentation, record-keeping, third-party assessmentLands here first when high-risk rules applyDeferred to Dec 2027, which removes near-term urgency
08Content marking, deepfake labelling, output disclosureArticle 50 enforced at the serving boundary, live nowThe one EU obligation with immediate teeth
09Audit trails, action logging, human oversightJune 2026 US EO and EU record-keeping both land hereTurned agent governance evidence into a procurement gate
10Policy enforcement and evidence generationNatural enforcement point for both regimesThe gateway is the only component that sees every call
11Workplace surveillance, accessibility, employment lawAttaches to the deploying enterprise, not the vendorUnderrated; rarely appears in vendor risk registers
12Sectoral certification per verticalState-by-state insurance, HIPAA, FINRA, FDAThe moat named in the top-ranked opportunity in Part 3
Where value concentrates

L01 and L06, with a long tail at L12. Export controls at silicon are the highest-consequence regulatory force in the stack and the one this report has historically under-covered. L06 carries the live EU obligations. L12 carries sectoral certification, which is a moat rather than a cost. The middle of the stack (L03, L07, L10) is largely regulated by inheritance.

Energy & Physical Constraints

Spine 03

Market Sizing

~485 → ~945 TWh

Global data-centre electricity consumption was approximately 485 TWh in 2025, and the IEA base case projects roughly 945 TWh by 2030, about 3% of global electricity, up from 1.5% in 2024. Accelerated servers grow near 30% annually against 9% for conventional servers and account for close to half the net increase. This corrects the 1,000 TWh-in-2026 figure carried in prior editions, which was a 2024-vintage projection that did not materialise.

Key Dynamics

The constraint is interconnection, not generation. There is no shortage of prospective electrons; there is a shortage of approved connections and delivery timelines. ERCOT alone received 198GW of large-load applications in Q1 2026 against a 438GW queue, energization runs 4–7 years in major hubs, and DOE has moved to assert federal jurisdiction over loads above 20MW specifically to compress that. This is why Amazon can state that $220B of 2026 capex still will not meet demand: capital is not the binding input.

LayerExposure at this altitudeEvidenceSignal
01Wafer and advanced-packaging capacity; perf-per-wattCoWoS allocation, HBM supplyEvery efficiency gain here compounds through all 11 layers above
02Grid interconnection, firm generation, water, shellsERCOT took 198GW of large-load applications in Q1 2026 against a 438GW queueEnergization runs 4–7 years in major hubs, up to a decade elsewhere. ~60% of hyperscaler capex is power and shells
03Optics and switching power drawRising share of rack power as clusters scaleOne of the few remaining cheap efficiency levers
04PUE, siting economics, coolingHyperscaler PUE gains largely exhausted; long tail still has headroomSets the floor under every price above
05Storage and ETL drawSmall relative to L02Not a binding constraint
06Training-run power eventsLargest single events, but inference has overtaken training in aggregateTraining is the headline; inference is the curve
07Quantization and distillation as demand-side leversThe only relief that does not require new generationLever, not a constraint
08Steady-state token-serving loadNow over 55% of AI infrastructure spendInference efficiency, not training efficiency, determines the 2030 power curve
09Multi-step agent traces multiply L08 drawA single agent trace can consume 10–100× the tokens of one completionIndirect multiplier
10Routing and caching as efficiency leversSemantic caching and cheap-model routingLever, not a constraint
11None material—Not a binding constraint
12None material—Not a binding constraint
Where value concentrates

L01 through L04 only. The thinness of the top eight rows is the finding, not a gap in the research. Physical constraints bind at the bottom of the stack and are close to irrelevant above L04, which is precisely why application-layer companies can scale without capital intensity while infrastructure companies cannot. The one exception worth watching is L08 and L09: inference efficiency and agent-trace multiplication are demand-side levers on L02 power draw, and they are the only levers available on a shorter horizon than grid interconnection.

How This Was Built, and What to Distrust

A landscape report is only as good as its willingness to say where it is weak. This appendix states the method, the confidence tiers, and the known defects that survived this cycle.

Method

Every quantified claim is checked against public primary sources on each refresh, prioritised by volatility: valuations, ARR, IPO status and corporate structure are re-verified every cycle; silicon roadmaps and sovereign policy less often. Corrections are logged in the Updates banner with the superseded value struck through inline and a superscript link back to the entry. No figure is carried forward on the strength of having appeared in a prior edition, which is exactly how the 1,000 TWh energy error survived three cycles before entry 30 caught it.

Confidence tiers

A Filed or disclosed: SEC filings, earnings calls, official press releases. Treat as fact.
B Credible third-party estimate: Sacra, Menlo, Dell'Oro, IEA, analyst surveys. Directionally reliable, point values soft.
C Market-sizing projection or extrapolation. Useful for order of magnitude only. Every TAM in Part 3 is tier C, and the 2032 figure in the AI-security chart is the softest number in the document.

Known defects in this edition

Survey-based demand data is self-reported. The 88%, 39% and 6% figures in Part 2 come from executive surveys, which systematically over-report adoption and under-report failure. Read them as upper bounds on adoption and lower bounds on disappointment.

Market sizings use mixed bases. The layer chart places TAM, annual capex and annual funding on one axis because that is how the underlying sources report them. It is legitimate for rank order and misleading for arithmetic. Do not sum it.

Newly added companies carry structural data only. The companies added this cycle to close the Layer 03 and search-coverage gaps have verified layer placement, taglines and relationships, but financial stats only where a figure was verified this cycle. An empty stat is an absence of verified data, never an implied zero.

Private-company revenue is estimated. ARR figures for private companies derive from third-party estimators and should be treated as tier B throughout.

Stack Knowledge Graph
Drag nodes · Scroll to zoom · Esc close
Click any node Tap a company to see its stack position, key metrics, and all mapped connections.

Or type to focus a specific company.