AMD XDNA
AMD's spatial dataflow NPU for on-chip AI inference.
AMD’s XDNA is a microarchitecture for deep learning processors, built to accelerate machine learning and AI workloads. It draws on technology from Xilinx, which AMD bought in 2022. Chips using XDNA appear in AMD’s Ryzen AI processors, working alongside the Zen CPU and RDNA GPU on the same die.
The architecture uses a spatial dataflow design: AI Engine (AIE) tiles process data in parallel, minimizing trips to external memory. This approach exploits parallelism and data locality for better performance and power efficiency. Each AIE tile includes a VLIW + SIMD vector processor for high-throughput compute and tensor operations, a scalar RISC-style processor for control flow, local memory blocks for weights, activations, and intermediate data, on-chip program and data memories to cut latency and power, and dedicated DMA engines with programmable interconnects for deterministic, high-bandwidth data movement between tiles. The tile arrays are modular and scalable, letting AMD configure NPUs with different tile counts to suit various power, area, and performance targets. Operating frequencies typically reach up to 1.3 GHz, adjustable based on thermal and power limits.
The first-generation XDNA NPU launched in early 2023 with the Ryzen 7040 “Phoenix” series, delivering up to 10 TOPS in mobile devices. A refresh, the Ryzen 8040 “Hawk Point” series from 2024, improved the NPU via firmware updates, higher clock speeds, and tuning, boosting performance to around 16 TOPS. The second generation, XDNA 2, debuted with the Ryzen AI 300 and PRO 300 mobile processors based on Zen 5, reaching up to 55 TOPS on flagship models.
XDNA’s core is a spatially arranged array of AI Engine tiles, enabling parallel and pipelined ML processing. Each tile has VLIW + SIMD vector cores for matrix multiplications and convolutions, a scalar control processor for instruction sequencing, on-chip SRAM for model parameters and intermediate data, and programmable DMA controllers with a low-latency interconnect for deterministic data movement. This design supports low-latency, high-bandwidth computation for real-time AI inference on edge devices.
Benefits include deterministic latency for predictable inference timing, power efficiency from reduced external DRAM access, high compute density in a small silicon area for thin-and-light devices, and scalability from mobile to enterprise-class systems.
- First generation launch
- early 2023 with Ryzen 7040 'Phoenix' series
- First generation performance
- up to 10 TOPS
- First-generation refresh (hawk point)
- 2024, Ryzen 8040 series, up to 16 TOPS
- Second generation (xdna 2)
- Ryzen AI 300 and PRO 300 mobile processors, up to 55 TOPS
- Operating frequency
- up to 1.3 GHz
- Acquisition
- Xilinx acquired by AMD in 2022
Lore & Background
XDNA employs a spatial dataflow architecture, where AI Engine (AIE) tiles process data in parallel with minimal external memory access. Each AIE tile contains a VLIW + SIMD vector processor, a scalar RISC-style processor, local memory blocks, on-chip program and data memories, and dedicated DMA engines with programmable interconnects. The tile arrays are scalable and modular, allowing AMD to configure NPUs with varying tile counts to fit different power, area, and performance targets.
The first generation XDNA NPU launched in early 2023 with the Ryzen 7040 'Phoenix' series, achieving up to 10 TOPS in mobile form factors. A first-generation refresh, the Ryzen 8040 'Hawk Point' series released in 2024, improved the NPU through firmware updates and higher clock speeds, pushing performance to around 16 TOPS. The second generation, XDNA 2, debuted with Ryzen AI 300 and PRO 300 mobile processors based on Zen 5 microarchitecture, drastically increasing AI throughput to up to 55 TOPS on flagship models.
The architecture is designed for deterministic latency, power efficiency through on-chip local memory, high compute density, and scalability from lightweight mobile devices to enterprise-class servers. XDNA is supported via AMD's ROCm and Vitis AI software stacks, enabling popular ML frameworks such as ONNX, TensorFlow, and PyTorch. Microsoft Windows ML runtime integrates AMD NPU acceleration in devices marketed as Copilot+ PCs.
Reader's Guide
XDNA represents AMD's entry into dedicated neural processing units for client devices, leveraging the spatial dataflow architecture inherited from Xilinx. Its significance lies in enabling local AI inference on thin-and-light laptops without cloud dependency, as seen in Copilot+ PCs. The architecture's modular tile design allows scaling from mobile to server-class deployments, while the use of on-chip SRAM and deterministic interconnects reduces latency and power consumption compared to traditional GPU or CPU approaches. The article notes that advertised TOPS are theoretical maximums, with actual performance varying based on thermal headroom, workload specifics, and driver optimizations. Some entry-level models disable or limit NPU functionality to save power and reduce die area. The software ecosystem and tooling are described as evolving, with continued improvements expected to fully exploit hardware capabilities. XDNA's legacy is tied to AMD's acquisition of Xilinx and its integration into Ryzen AI processors, positioning it as a competitor to other on-chip NPUs like Apple's Neural Engine and Google's Tensor Processing Unit.
Did You Know?
- XDNA is based on Xilinx technology, which AMD acquired in 2022.
- The first-generation XDNA NPU achieved up to 10 TOPS in the Ryzen 7040 'Phoenix' series.
- XDNA 2 reached up to 55 TOPS on flagship Ryzen AI 300 mobile processors.
- Each AIE tile contains a VLIW + SIMD vector processor and a scalar RISC-style processor.
Spatial Dataflow at the Core
AMD's XDNA architecture takes a fundamentally different approach to on-chip AI acceleration by organizing computation into a spatial dataflow fabric rather than relying on a single monolithic processor. At the heart of this design sit arrays of AI Engine tiles, each one a self-contained processing unit that can crunch data in parallel while keeping it close to the compute logic. Every tile pairs a VLIW-plus-SIMD vector core—tuned for the matrix multiplications and convolutions that dominate neural networks—with a smaller scalar RISC-style core that handles sequencing and control-flow tasks. To keep data from having to travel to slow external DRAM, each tile carries its own local memory blocks for weights, activations, and intermediate results, along with on-chip program and data storage. Dedicated DMA engines and a programmable interconnect fabric shuttle data between tiles with deterministic, high-bandwidth transfers. The whole tile array is modular: AMD can scale the number of tiles up or down to hit different power, area, and performance targets, with operating frequencies reaching roughly 1.3 GHz depending on thermal budgets. This design philosophy traces back to Xilinx technology, which AMD brought into its portfolio through the 2022 acquisition.
Three Generations of Growing Throughput
XDNA's journey from debut to its current form tells a story of rapid iteration. The first-generation NPU arrived in early 2023 alongside the Ryzen 7040 "Phoenix" mobile processors, delivering up to 10 TOPS of AI throughput in a laptop-class form factor. A year later, the Ryzen 8040 "Hawk Point" refresh bumped that figure to roughly 16 TOPS, not through a new silicon design but via firmware updates, higher clock speeds, and tuning refinements applied to the same underlying architecture. The real leap came with XDNA 2, which debuted in the Ryzen AI 300 and PRO 300 mobile processors built on AMD's Zen 5 CPU microarchitecture. Flagship models in that lineup push AI throughput to as high as 55 TOPS—a more-than-fivefold jump over the original Phoenix chip. Throughout all three generations, the NPU has served as a dedicated companion to the Zen CPU cores and RDNA GPU on the same die, giving AMD a three-engine approach to mixed workloads. The modular tile philosophy means each generation can retune the balance between tile count, clock speed, and power envelope to suit the target market, from thin ultrabooks to more demanding mobile workstations.
Software Stacks and the Copilot+ Push
Hardware alone cannot deliver AI acceleration, and AMD has leaned on a layered software strategy to make XDNA useful to developers. The NPU is accessible through two primary toolchains: the ROCm (Radeon Open Compute) platform and the Vitis AI stack, both of which let programmers offload inference and training workloads onto the tile array. Major machine-learning frameworks—ONNX, TensorFlow, and PyTorch—are supported through these tools, meaning researchers and engineers do not need to rewrite models from scratch to target AMD silicon. On the consumer side, Microsoft's Windows ML runtime has been integrated to expose NPU acceleration on devices marketed as Copilot+ PCs, enabling local AI inference that runs entirely on-device without round-trips to the cloud. This pairing positions AMD's mobile processors as viable alternatives in the growing category of AI-first laptops. That said, the software ecosystem is still maturing; AMD has acknowledged that tooling and driver optimizations are ongoing efforts, and full exploitation of the hardware's capabilities remains a work in progress.
Where the Promise Meets Reality
XDNA's architectural choices deliver several tangible advantages that set it apart from routing AI work through a general-purpose CPU or a discrete GPU. Because each tile keeps weights and intermediate data in on-chip SRAM, the NPU avoids the energy cost of repeatedly fetching from external DRAM, which translates into lower power draw during sustained inference. The spatial dataflow design also produces deterministic, predictable latency—valuable for real-time applications where a missed frame or delayed response is unacceptable. Compute density is another strength: packing tens of TOPS into a footprint small enough for ultrabooks and portable workstations is something traditional GPU approaches struggle to match. The modular tile philosophy extends scalability upward as well, allowing the same architectural language to scale into enterprise-class server configurations with far more tiles. Yet important caveats remain. Advertised TOPS figures represent theoretical ceilings; real-world throughput shifts with thermal headroom, specific workload characteristics, and the maturity of driver and software optimizations. In some entry-level SKUs, AMD disables or limits NPU functionality to conserve die area and power budget, meaning not every Ryzen AI chip delivers the full headline number.
Frequently Asked Questions
What is AMD XDNA?
XDNA is AMD's dedicated NPU microarchitecture built to accelerate on-chip AI inference and machine-learning workloads. It lives on the same die as the Zen CPU cores and RDNA graphics block inside Ryzen AI processors.
How does XDNA's spatial dataflow design differ from a GPU approach?
Rather than repeatedly fetching data from external memory, XDNA arranges AI Engine tiles in a grid that process data in parallel right where it resides. Each tile pairs a VLIW core with a SIMD vector unit, exploiting data locality to squeeze more throughput out of every watt.
Which Ryzen chips actually ship with XDNA?
First-gen XDNA debuted in early 2023 on the Ryzen 7040 'Phoenix' mobile series, and the 2024 Hawk Point refresh (Ryzen 8040) raised the ceiling to 16 TOPS. The second-generation XDNA 2 architecture is found in the Ryzen AI 300 and PRO 300 mobile lines.
What AI performance numbers can XDNA hit?
First generation tops out at 10 TOPS, the Hawk Point refresh reaches 16 TOPS, and XDNA 2 delivers up to 55 TOPS of AI compute. Across all generations the NPU runs at a maximum operating frequency of 1.3 GHz.
Where did the XDNA technology originally come from?
The design lineage traces back to Xilinx, the FPGA and adaptive-computing company AMD acquired in 2022. AMD adapted Xilinx's spatial-parallel processing concepts into a fixed-function NPU aimed at consumer and mobile AI workloads.
More in PC Hardware 1-24
Spotted an error? Know more?
Reader corrections go straight into our review queue. Suggest an edit · How this site is sourced
