AI Inference Infrastructure
AI inference infrastructure is the serving side of artificial intelligence: the chips, systems, clouds, networking, and power that run trained models in production. As reasoning and agentic models generate more tokens per query, this recurring workload is growing faster than training, and it spans NVIDIA, Broadcom, CoreWeave, Arista, Vertiv, and Cerebras.
| Category | Thematic / AI compute — inference & serving |
|---|---|
| Representative companies | 8 |
| Related ETF | TIGER US AI Data Center TOP4 Plus ETF (0142D0) |
| Last updated | 2026-06-15 |
- AI inference infrastructure is the serving side of AI: the chips, systems, clouds, networking, and power that run trained models in production, as opposed to the one-time cost of training them.
- The workload mix is tilting toward inference. Reasoning and agentic models emit far more output tokens per query, so the recurring cost of running models is compounding faster than the cost of training them.
- NVIDIA still anchors inference: data-center revenue was a record $75.2 billion in its February-April 2026 quarter, up 92%, and the GB200 NVL72 rack was engineered for reasoning-model inference.
- Custom silicon, neoclouds, networking, and power all scale with inference. Broadcom's AI revenue grew 143% to $10.8 billion, CoreWeave's backlog reached $99.4 billion, and Cerebras signed a 750-megawatt, $20-billion-plus inference deal with OpenAI.
- The closest Korea-listed proxies are the Akros U.S. AI Data Center and U.S. AI Semiconductor indices, tracked by the TIGER and KODEX ETFs, which hold many of these names.
What is AI inference infrastructure?
It is the part of the AI build-out that runs models after they are trained. Training teaches a model once; inference is the act of running it to answer every prompt, classify every image, or take every agent step, billions of times a day. Because that cost recurs with usage, the hardware optimized for serving has become its own investable layer. The concept holds the leaders of each part of that layer: NVIDIA for GPUs, Broadcom and Marvell for custom accelerators and networking silicon, Arista for the Ethernet fabric, CoreWeave and Nebius for neocloud capacity, Vertiv for power and cooling, and Cerebras for wafer-scale inference systems.
Why is the workload shifting from training to inference?
Because the newest models think before they answer. Reasoning and agentic systems generate long chains of intermediate tokens, and an autonomous agent can call a model dozens of times to finish one task, so the compute spent serving a model over its life can dwarf the compute spent training it. That demand is showing up in NVIDIA’s results: data-center revenue reached a record $75.2 billion in the February-April 2026 quarter, up 92% year over year, and total revenue hit $81.6 billion, up 85% (NVIDIA, May 20, 2026). NVIDIA founder and CEO Jensen Huang framed the scale of the build-out:
“The buildout of AI factories, the largest infrastructure expansion in human history, is accelerating at extraordinary speed.”
— Jensen Huang, founder and CEO, NVIDIA (NVIDIA, May 20, 2026)
Why does custom silicon rise alongside the GPU for inference?
Because inference is cost-sensitive and latency-sensitive, and hyperscalers want a cheaper lever than merchant GPUs for steady-state serving. Broadcom co-designs Google’s TPU and Meta’s MTIA accelerators and supplies the Ethernet switch silicon of AI fabrics; its AI semiconductor revenue grew 143% to $10.8 billion in fiscal Q2 2026, with $16.0 billion guided for fiscal Q3 (Broadcom 8-K, Jun 3, 2026). Marvell builds custom XPUs and the optical DSPs that move tokens inside a cluster, and data center was about 76% of its record $2.418 billion fiscal Q1 2027 revenue, up 28% (Marvell 8-K, May 27, 2026). At the system extreme, Cerebras sells a wafer-scale processor 58 times larger than NVIDIA’s B200 and signed an inference deal with OpenAI worth more than $20 billion for 750 megawatts of compute (CNBC, Jan 14, 2026).
Where do cloud, networking, and power fit in AI inference infrastructure?
They are where inference capacity is actually rented and run. Neoclouds buy GPUs at fleet scale and rent them back under multi-year contracts: CoreWeave ended Q1 2026 with a $99.4 billion revenue backlog and more than 3.5 gigawatts of contracted power (CoreWeave, May 7, 2026), while Nebius grew Q1 2026 group revenue 684% to $399 million with Microsoft and Meta as customers (Nebius, May 13, 2026). The clusters need a fabric and a power envelope: Arista’s Q1 2026 revenue rose 35% to $2.709 billion as it raised its AI-networking target (Arista, May 5, 2026), and Vertiv’s net sales rose 30% to $2.65 billion on data-center power and liquid cooling, with the Americas growing 44% organically (Vertiv 8-K, Apr 22, 2026).
Who should consider AI inference infrastructure exposure?
It suits an investor who wants the recurring, usage-driven side of AI rather than a single chip, and who can size for the fact that the whole basket leans on the same buyers. The revenue is current, not promised: CoreWeave’s backlog is $99.4 billion (CoreWeave), Arista raised guidance on AI-networking demand (Arista), and Vertiv’s order book reflects gigawatts of planned capacity (Vertiv 8-K).
The caveat is concentration. A handful of hyperscalers and AI labs fund most of the demand, so a slowdown in their capex would hit chips, clouds, networking, and power together. For an individual investor that argues for a satellite position beside a diversified core; for an institution it works as a liquid, high-beta sleeve whose factor overlap with growth and momentum books has to be budgeted, with one new-issue exception in Cerebras (Cerebras S-1).
Which companies represent AI Inference Infrastructure?
| Company | Sector | What it does |
|---|---|---|
| NVIDIA (NVDA) | Information Technology · Semiconductors | NVIDIA's GPUs run the bulk of AI inference, and its Blackwell GB200 NVL72 rack was engineered for the inference of reasoning and agentic models. Data-center revenue was a record $75.2 billion in Q1 FY2027 (February-April 2026), up 92% year over year. |
| Broadcom (AVGO) | Information Technology · Semiconductors | Broadcom co-designs the custom accelerators (Google's TPU, Meta's MTIA) that hyperscalers use to serve inference at lower cost than merchant GPUs, plus the Ethernet switch silicon that lashes inference clusters together. AI semiconductor revenue grew 143% to $10.8 billion in fiscal Q2 2026. |
| Marvell Technology (MRVL) | Information Technology · Semiconductors | Marvell builds custom XPUs and the optical DSPs and Ethernet switching that move tokens inside inference clusters. Data center was about 76% of record fiscal Q1 2027 revenue of $2.418 billion, up 28% year over year. |
| CoreWeave (CRWV) | Information Technology · AI Cloud (GPU-as-a-Service) | CoreWeave is the largest pure-play AI cloud, renting NVIDIA GPU clusters that increasingly serve inference workloads. Its revenue backlog reached $99.4 billion and contracted power exceeded 3.5 gigawatts at the end of Q1 2026. |
| Nebius Group (NBIS) | Information Technology · AI Cloud (GPU-as-a-Service) | Nebius is a fast-growing neocloud building NVIDIA GPU capacity for AI training and inference. Q1 2026 group revenue grew 684% year over year to $399 million, with Microsoft and Meta among its infrastructure customers. |
| Arista Networks (ANET) | Information Technology · Data-Center Networking | Arista's high-performance Ethernet switches are a standard fabric for AI inference and training clusters. Q1 2026 revenue rose 35% to $2.709 billion and the company raised its AI-networking target. |
| Vertiv (VRT) | Industrials · Data Center Power & Cooling | Vertiv supplies the power distribution and liquid cooling that dense inference racks require. AI demand drove Q1 2026 net sales up 30% to $2.65 billion, with organic Americas growth of 44%. |
| Cerebras Systems (CBRS) | Information Technology · Semiconductors (AI Systems) | Cerebras sells wafer-scale systems purpose-built for fast inference; its Wafer-Scale Engine is 58 times larger than NVIDIA's B200 chip. OpenAI signed a deal worth more than $20 billion for 750 megawatts of Cerebras inference compute. |
What are the risks of AI Inference Infrastructure?
The cash flows are real, but so is the dependence on a few budgets.
- Capex concentration. A small group of hyperscalers and AI labs funds most inference demand; a pause in their spending would hit every layer of the basket at once.
- Leverage in the cloud layer. Neocloud growth is debt-financed, and CoreWeave added an $8.5 billion facility in Q1 2026 alone, so higher rates raise the cost of each incremental gigawatt (CoreWeave, May 7, 2026).
- Customer concentration in new issues. Cerebras drew 86% of its 2025 revenue from two UAE-linked customers, and its valuation embeds the OpenAI ramp executing on schedule (Cerebras S-1).
- Pricing risk from efficiency. Cheaper or more efficient inference silicon, including hyperscaler ASICs, could compress pricing for any single layer even as total demand grows.
Related concepts, securities & terms
- AI Infrastructure parent
- US AI Data Center sibling
- US AI Semiconductor sibling
- Optical Networking & Photonics related
- NVIDIA (NVDA) child
- Broadcom (AVGO) child
- Marvell Technology (MRVL) child
- CoreWeave (CRWV) child
- Nebius Group (NBIS) child
- Arista Networks (ANET) child
- Vertiv (VRT) child
- Cerebras Systems (CBRS) child
- AI Inference child
- GPU-as-a-Service (Neocloud) related
- AI accelerator related
- AI On-Devices vs AI Inference Infrastructure related
Related indices & ETFs
- TIGER US AI Data Center TOP4 Plus ETF (0142D0) · Mirae Asset Global Investments Korea-listed ETF tracking the Akros U.S. AI Data Center index; holds neocloud and power names that serve inference.
- Akros U.S. AI Data Center TOP4 Plus Index (AUAIDC) · Akros Technologies, Inc. The proprietary Akros index whose constituents include the data-center and cloud layers that run AI inference.
These references describe index-tracking relationships as a matter of fact and are not a recommendation to buy any product. Akros, as the index provider, may receive licensing fees from product sponsors. Review the product's prospectus before investing.
Frequently asked questions
What is AI inference infrastructure?
It is the serving side of artificial intelligence: the silicon, systems, clouds, networking, and power that run already-trained models in production. Training builds a model once; inference runs it for every user query, every day. The basket spans GPUs (NVIDIA), custom accelerators and networking (Broadcom, Marvell, Arista), neoclouds (CoreWeave, Nebius), power and cooling (Vertiv), and wafer-scale inference systems (Cerebras).
How is inference different from AI training?
Training is the upfront, compute-heavy process of teaching a model, paid once per model version. Inference is the recurring cost of running that model to answer prompts, which scales with usage. Inference is more latency-sensitive and cost-sensitive, which is why hyperscalers deploy custom accelerators such as Google's TPU and Meta's MTIA, co-designed by Broadcom, alongside merchant GPUs (Broadcom 8-K, Jun 3, 2026).
Is inference becoming a bigger workload than training?
The mix is shifting that way. Reasoning and agentic models generate far more output tokens per query than earlier chatbots, so the compute spent serving a model can exceed the compute spent training it over the model's life. NVIDIA built its GB200 NVL72 rack specifically for reasoning-model inference, and its data-center revenue still grew 92% to a record $75.2 billion in the February-April 2026 quarter (NVIDIA, May 20, 2026).
Which companies represent AI inference infrastructure?
Eight names across the serving stack: NVIDIA (NVDA) for GPUs, Broadcom (AVGO) and Marvell (MRVL) for custom accelerators and networking silicon, Arista (ANET) for the Ethernet fabric, CoreWeave (CRWV) and Nebius (NBIS) for neocloud capacity, Vertiv (VRT) for power and cooling, and Cerebras (CBRS) for wafer-scale inference systems. CoreWeave alone ended Q1 2026 with a $99.4 billion revenue backlog (CoreWeave, May 7, 2026).
What are the main risks of AI inference infrastructure?
Concentration and capex dependence. A handful of hyperscalers and AI labs fund most of the demand, so a pause in their spending would hit chips, clouds, networking, and power at once. Neoclouds add leverage and customer-concentration risk: CoreWeave carries a debt-financed build-out, and Cerebras drew 86% of its 2025 revenue from two UAE-linked customers (Cerebras S-1). Cheaper or more efficient inference silicon could also compress pricing for any single layer.
Is AI inference infrastructure a good long-term investment for an individual investor?
It suits an investor who wants exposure to the recurring, usage-driven side of AI and can tolerate single-theme volatility, sized as a satellite rather than a core holding. The cash flows are real and dated: CoreWeave's backlog reached $99.4 billion (CoreWeave, May 7, 2026), Arista's revenue grew 35% to $2.709 billion (Arista, May 5, 2026), and Vertiv's net sales rose 30% to $2.65 billion (Vertiv 8-K, Apr 22, 2026). But the whole basket leans on the same hyperscaler budgets, so it moves together. Always check fees, holdings, and risk before investing.
How does AI inference infrastructure fit an institutional thematic mandate?
As a high-beta sleeve that captures the serving layer of AI in liquid, mostly US-listed names, with one new issue in Cerebras. The thesis maps to disclosed commitments: OpenAI's 750-megawatt, $20-billion-plus Cerebras deal (CNBC, Jan 14, 2026), Nebius signing Microsoft and Meta (Nebius, May 13, 2026), and Broadcom's $16.0 billion fiscal Q3 AI guide (Broadcom 8-K). The allocator's caveats are factor overlap with growth and momentum books and a demand base concentrated in a few capex budgets.
Is there an ETF for AI inference infrastructure?
There is no dedicated inference ETF, but the closest Korea-listed proxy is the TIGER US AI Data Center TOP4 Plus ETF (0142D0), which tracks the Akros U.S. AI Data Center index and holds neocloud and power names that serve inference (Mirae Asset). The KODEX US AI Semiconductor TOP3 Plus ETF (0151S0) covers the chip layer (Samsung Asset Management). Always check fees, holdings, and risk before investing.
Sources & references
- NVIDIA Announces Financial Results for First Quarter Fiscal 2027 · NVIDIA Corporation (Newsroom), 2026-05-20
- Broadcom announces second quarter fiscal year 2026 financial results (SEC 8-K, Exhibit 99.1) · Broadcom Inc. / SEC EDGAR, 2026-06-03
- Marvell Technology reports fiscal Q1 2027 results (Form 8-K, Exhibit 99.1) · Marvell Technology / SEC EDGAR, 2026-05-27
- CoreWeave Reports Strong First Quarter 2026 Results · CoreWeave, Inc., 2026-05-07
- Nebius reports first quarter 2026 financial results · Nebius Group N.V., 2026-05-13
- Arista Networks, Inc. Reports First Quarter 2026 Financial Results · Arista Networks, Inc., 2026-05-05
- Vertiv Reports First Quarter 2026 Results (SEC 8-K, Exhibit 99.1) · Vertiv Holdings Co / SEC EDGAR, 2026-04-22
- Cerebras scores OpenAI deal worth over $10 billion · CNBC, 2026-01-14
- Cerebras Systems Inc. Form S-1 Registration Statement · Cerebras Systems Inc. / SEC EDGAR, 2026-04-17
- TIGER US AI Data Center TOP4 Plus ETF (0142D0) — product page · Mirae Asset Global Investments, 2025-12-09
- KODEX 미국AI반도체TOP3플러스 ETF (0151S0) — product page · Samsung Asset Management, 2026-01-13