On-Device AI

On-device AI is artificial intelligence that runs locally on a phone, PC, car, or edge device, using a neural processing unit built into the device’s chip, instead of sending the request to a cloud data center. It lowers latency, cost, and privacy exposure, and it works offline.

What is on-device AI?

On-device AI keeps the model and the computation on the hardware in front of you. Modern phone and laptop chips include a neural processing unit (NPU), a block designed to run AI models efficiently at low power, so a summary, a translation, an image edit, or an assistant reply can be produced without a round trip to a server. The trade-off is size: a device has far less memory and power than a data center, so on-device AI runs compact models, while the largest models still run in the cloud. Hybrid designs split the work, with Apple reserving a Private Cloud Compute path for requests too large to run locally. The capability is real today: Qualcomm’s Snapdragon X Elite NPU is rated at 45 TOPS and can run models with more than 13 billion parameters locally (Qualcomm Snapdragon X Elite), and Apple runs a roughly 3-billion-parameter model on-device (Apple Machine Learning Research).

How is on-device AI used in thematic investing?

On-device AI is the demand thesis behind a different slice of the AI trade than data centers: the chips and devices in consumers’ hands. The case rests on four advantages, latency, cost, privacy, and offline use, and it favors the companies that supply the silicon and software for it, which is the basis of the AI on-devices concept. Apple is the flagship deployer, running its on-device system across more than 2.5 billion active devices, and at WWDC 2026 it pushed the software further:

“We’re delivering the next generation of Apple Intelligence across our platforms; introducing Siri AI, a profoundly more intelligent, knowledgeable, and capable Siri.”

— Craig Federighi, SVP Software Engineering, Apple (Apple, Jun 8, 2026)

For investors, on-device AI is a lower-beta way to own the AI build-out than the data-center trade: its demand tracks the device installed base and the silicon inside it rather than hyperscaler capex, and it offers optionality if inference shifts from the cloud toward the edge.

Frequently asked questions

What is an NPU?

An NPU, or neural processing unit, is a specialized block in a device's chip that runs AI models efficiently at low power. Qualcomm's Snapdragon X Elite NPU is rated at 45 TOPS, and Apple calls the equivalent block in its chips the Neural Engine.

What are the limits of on-device AI?

A phone or laptop has far less memory and power than a data center, so the largest models still run in the cloud. On-device AI handles compact models well; designs like Apple's reserve a Private Cloud Compute path for requests too big to run locally.

Sources & references

  1. Snapdragon X Elite (45 TOPS Hexagon NPU) · Qualcomm Incorporated, 2026-06-15
  2. Updates to Apple's On-Device and Server Foundation Language Models · Apple Machine Learning Research, 2025-06-09