Introduction
Every new laptop seems to carry an “AI PC” sticker these days, but that label hides a specific, easy-to-misunderstand number underneath it. The NPU TOPS rating has become the headline spec in PC marketing, yet most buyers have no idea what it actually measures or why it matters less than the packaging suggests. This article breaks down what TOPS really means, why Microsoft picked the number it did, and what genuinely determines whether a laptop can run machine learning models locally.
1. What Actually Makes a PC an AI PC
Microsoft’s specific certification, called Copilot+ PC, sets a clear bar: a Windows 11 PC needs a Neural Processing Unit capable of running at 40 or more TOPS, alongside 16GB of RAM, 256GB of storage, and Windows 11 version 24H2 or later. Meeting that bar unlocks a defined set of on-device AI features, including Recall, Live Captions, Studio Effects, and Click to Do.
That threshold is a features gate, not a performance rating. Crossing 40 TOPS switches on a specific list of Windows capabilities. It says nothing about how fast a large language model would run on that same hardware, which is where a lot of buyer confusion starts.
2. Breaking Down the 40 TOPS Number
TOPS stands for trillions of operations per second, a raw measure of how many low-precision math operations a chip’s NPU can execute every second. It is a hardware throughput number, similar in spirit to clock speed or core count, rather than a benchmark of real-world task completion time.
Current Copilot+ certified chips comfortably clear that 40 TOPS floor. Intel’s Core Ultra 200V series reaches around 48 TOPS, AMD’s Ryzen AI 300 series lands near 50 TOPS, and Qualcomm’s Snapdragon X Elite delivers about 45 TOPS. The newer Snapdragon X2 Elite pushes considerably higher, into the 80 to 85 TOPS range, currently the highest published figure among mainstream Windows NPUs.
3. Where That Number Came From
The 40 TOPS requirement was not picked out of thin air, though its early history was a little chaotic. When Microsoft first announced Copilot+ PC requirements in May 2024, the move immediately exposed a problem: Intel’s first-generation NPU inside the Core Ultra Meteor Lake processor delivered only around 11 TOPS, far short of what Microsoft was about to require. AMD’s contemporary Ryzen Hawk Point platform managed roughly 16 TOPS, also well below the coming bar.
That gap kicked off what amounted to an unplanned hardware race. Within about a year, Intel, AMD, and Qualcomm had all shipped chips clearing 40 TOPS, and by 2026 every current-generation part from the major vendors clears that floor comfortably. The number stopped being a differentiator between competing chips and became simply the minimum price of entry.
4. Why TOPS Is Not an LLM Speed Rating
Here is the part most AI PC marketing glosses over. Large language model performance is measured in tokens generated per second, not TOPS, and the two numbers do not translate directly into each other.
The scale of the mismatch is genuinely large. A discrete desktop GPU like the RTX 4090 advertises more than 1,300 TOPS, while a Copilot+ certified NPU sits somewhere between 40 and 85 TOPS, a roughly 29 times difference. If TOPS were the metric that determined LLM speed, a Copilot+ laptop would look hopelessly outmatched. But the comparison is misleading in the other direction too, because the actual bottleneck for running a language model locally is not raw compute throughput at all.
5. The Real Bottleneck for Local AI Models
What actually determines how fast a local language model runs is memory bandwidth and capacity, not the TOPS figure printed on a spec sheet. A model has to load its parameters into memory and move data through that memory continuously during inference, and a chip that cannot feed data to its processing units fast enough will sit idle no matter how high its theoretical TOPS ceiling is.
This is why architecture matters more than the headline number once a chip clears the 40 TOPS floor. The Snapdragon X2 Elite, for example, is currently the only Windows NPU able to run quantized 7 billion parameter models at genuinely usable speeds, a capability tied to its memory architecture and software optimization through Qualcomm’s Hexagon NPU stack, not simply its higher raw TOPS number. Apple’s approach makes the same point from a different angle. The company stopped publishing a specific Neural Engine TOPS figure starting with the M5 chip, instead relying on the chip’s unified memory architecture, where the CPU, NPU, and GPU all share the same high-bandwidth memory pool, to handle demanding AI workloads more efficiently than a standalone TOPS number would suggest.
6. How the Major Chips Compare in 2026
With every current chip clearing Microsoft’s 40 TOPS floor, the meaningful differences between platforms now show up elsewhere.
- Qualcomm Snapdragon X2 Elite: Leads on raw TOPS at 80 to 85, built on a third-generation 18-core Oryon CPU, and delivers 20 to 30-plus hours of battery life alongside roughly 30 percent lower power consumption than comparable x86 chips.
- Intel Core Ultra 200V: Reaches around 48 TOPS with strong ONNX Runtime software optimization supporting broad application compatibility.
- AMD Ryzen AI 300 series: Reaches around 50 TOPS, competitive across general productivity and AI-assisted workloads on Windows.
- Apple M5: Neural Engine performance lands around 38 to 40 TOPS, but benefits from a more mature Core ML software stack and unified memory that lets the M5 Pro and M5 Max handle large workloads better than the raw number implies.
Direct TOPS comparisons across Windows and macOS are not especially meaningful, since the two platforms optimize their AI software stacks differently. Comparing chips within the same platform tells a more useful story than comparing raw numbers across platforms.
7. Why an NPU Exists at All
None of this means the NPU is pointless, even with GPUs technically capable of higher raw throughput. The entire reason NPUs exist comes down to power efficiency rather than peak performance. A task that takes around 100 milliseconds on an NPU might take roughly 20 milliseconds on a discrete GPU, sounding like a clear GPU win, except that GPU draws around 150 watts doing it while the NPU draws around 2 watts.
For a battery-powered laptop running always-on AI features throughout the day, that efficiency gap is the entire point. An NPU is not trying to beat a desktop GPU on speed. It exists to make continuous, background AI processing sustainable on a device that needs to last a full day unplugged.
8. What to Actually Check Before Buying
Given that every current Copilot+ chip already clears the 40 TOPS floor, that number alone is no longer a useful way to differentiate laptops. A few more practical checks matter more:
- Confirm the specific applications you care about explicitly list your chip’s NPU as supported, rather than assuming a generic TOPS threshold guarantees compatibility.
- Prioritize memory capacity and configuration if local LLM inference matters to you, since that determines real-world model performance more than the TOPS figure does.
- Compare battery life under realistic workloads like browsing, video calls, and typing, rather than relying on a single headline battery number.
- Check whether storage is replaceable and whether memory is soldered, since many thin laptops lock in these components at the time of purchase.
- Remember that 40 TOPS is Microsoft’s floor for 2026, not a long-term target. As on-device AI models grow more capable, their compute requirements are likely to climb well past today’s baseline.
Key Takeaways
- Copilot+ requirement: A 40+ TOPS NPU, 16GB RAM, 256GB storage, and Windows 11 24H2 or later.
- What TOPS measures: Raw trillions of operations per second, a hardware throughput number, not a real-world task benchmark.
- What TOPS does not measure: Large language model speed, which is measured in tokens per second and depends heavily on memory bandwidth and capacity.
- Current chip range: Roughly 38 to 85 TOPS across Intel, AMD, Qualcomm, and Apple’s latest silicon, all clearing the Copilot+ floor.
- Why NPUs matter: Dramatically lower power draw than a GPU for the same AI task, which is essential for all-day battery life.
- What actually differs now: Memory architecture, software optimization, and battery efficiency, not the TOPS number on the box.
FAQs
What TOPS rating does a Copilot+ PC require? Microsoft requires at least 40 TOPS from a dedicated NPU, along with 16GB of RAM and 256GB of storage.
Does a higher TOPS number mean faster AI performance? Not necessarily, since large language model speed depends more on memory bandwidth and capacity than on raw TOPS.
Why did Intel’s first NPU fall short of Copilot+ requirements? Intel’s original Meteor Lake NPU delivered around 11 TOPS, far below the 40 TOPS threshold Microsoft later required.
Why do NPUs exist if GPUs are more powerful? NPUs use dramatically less power than GPUs for the same AI task, making them essential for sustained battery-powered AI features.
Is 40 TOPS enough for the future? It is the current 2026 floor set by Microsoft, but requirements are expected to rise as on-device AI models become more demanding.
Conclusion
The 40 TOPS NPU requirement behind every Copilot+ PC is a real, meaningful hardware threshold, but it answers a much narrower question than most marketing implies. It determines whether a specific set of Windows features will run locally, not how fast a language model will actually respond on that machine. With every current chip clearing that floor, the real differences between AI PCs in 2026 come down to memory architecture, software optimization, and power efficiency, the details that actually decide whether local machine learning on a laptop feels genuinely fast or just technically compliant.




