AI Inference

AI inference is the process of running a trained AI model to produce an output, such as an answer, a classification, or a prediction, for each new input. It is the recurring, usage-driven counterpart to training: training builds the model once, while inference runs it for every query, every day.

What is AI inference?

Inference is what happens every time you actually use an AI model. The model’s weights are fixed after training, and inference feeds a new input through those weights to generate an output: a chat reply, a translation, a recommendation, or a detected object. Because it runs once per request rather than once per model, inference is where AI meets real usage, and its cost scales with how many people use a product and how much each query demands. Newer reasoning and agentic models make that demand heavier, because they generate long chains of intermediate tokens and can call a model many times to finish one task.

How is AI inference used in thematic investing?

Inference is the durable demand signal behind AI infrastructure. Training spending is large but lumpy and concentrated in a few labs; inference spending grows steadily with adoption and recurs with usage, which is why it underpins the AI inference infrastructure basket of accelerators, memory, networking, and cloud capacity. The scale is visible in chipmaker results: NVIDIA’s data-center revenue, most of it tied to training and inference compute, reached a record $75.2 billion in its February-April 2026 quarter, up 92% year over year (NVIDIA, May 20, 2026). Custom inference silicon is growing even faster: Broadcom’s AI semiconductor revenue, much of it inference-oriented ASICs and networking, grew 143% to $10.8 billion in fiscal Q2 2026 (Broadcom 8-K, Jun 3, 2026). NVIDIA founder and CEO Jensen Huang framed the scale of the build-out behind it:

“The buildout of AI factories, the largest infrastructure expansion in human history, is accelerating at extraordinary speed.”

— Jensen Huang, founder and CEO, NVIDIA (NVIDIA, May 20, 2026)

For an investor, inference is the reason AI demand is more than a one-time capital cycle: as long as people use AI products, someone pays to run the models, and that spending flows to the chips, memory, networks, power, and clouds that serve it.

Frequently asked questions

What is the difference between AI training and inference?

Training is the upfront, compute-heavy process of teaching a model from data, paid once per model version. Inference is running that model to answer prompts, a recurring cost that scales with usage. Over a popular model's life, inference compute can exceed training compute.

Why is AI inference important for investors?

Because inference is the part of AI spending that recurs with usage, so it underpins the durable demand for accelerators, memory, networking, and cloud capacity, the AI inference infrastructure basket. NVIDIA's data-center revenue, most of it tied to inference and training compute, reached a record $75.2 billion in its February-April 2026 quarter.

Sources & references

  1. NVIDIA Announces Financial Results for First Quarter Fiscal 2027 · NVIDIA Corporation (Newsroom), 2026-05-20
  2. Broadcom announces second quarter fiscal year 2026 financial results (SEC 8-K, Exhibit 99.1) · Broadcom Inc. / SEC EDGAR, 2026-06-03