GPU and CPU Microarchitectures Codexery

Frequently Asked Questions

The most-asked questions about gpu and cpu microarchitectures.

What exactly is a microarchitecture?

It is the internal blueprint of a processor: how transistors are grouped into execution units, caches, and control logic so that instructions actually get carried out. Two chips can implement the same ISA (x86-64, ARM, etc.) yet feel completely different in use because their microarchitectures are laid out differently.

How does a GPU microarchitecture differ from a CPU's?

GPU designs stack thousands of small shader/ALU cores behind a very wide memory interface to maximize parallel throughput, while CPU microarchitectures invest in deep out-of-order scheduling, large caches, and aggressive branch prediction to keep a handful of wide cores fed. The trade-off is uniform, massively parallel workloads versus low-latency, irregular control flow.

Who are the most-watched architects in the enthusiast community?

Jim Keller (AMD, Tesla, Apple), Mark Papermaster (AMD), and the anonymous supply-chain leakers who post on forums like TechPowerUp and the successors to AnandTech tend to dominate threads. On the GPU side, the teams behind NVIDIA's successive architectures and AMD's RDNA/CDNA lines generate the most fan speculation.

Where should a total beginner start learning about microarchitectures?

A practical path is to first understand the instruction set, then read a clear explanation of the instruction pipeline, and finally compare two named designs side by side (for example, Intel Golden Cove versus AMD Zen 4) to see the concepts in concrete form. Annotated die shots and community wikis are useful visual companions along the way.

What does 'out-of-order execution' mean and why do fans obsess over it?

It is the technique of reordering independent instructions so that execution units stay busy while waiting on memory or branch results. Fans track it because the reorder-buffer size, dispatch-port count, and scheduler design are the main levers that let a 3 GHz part and a 5.5 GHz part do the same work.

What is a chiplet and why did it change the conversation?

A chiplet is a packaging strategy that splits a die into smaller, separately fabricated tiles (compute, I/O, cache) bonded together, allowing mixed process nodes and better yield. AMD's Zen 2 was the high-profile mainstream debut, and it reshaped how both AMD and Intel structure their product roadmaps.

Which specs do fans compare the moment a new microarchitecture drops?

The usual checklist covers IPC, core count, L1/L2/L3 cache sizes and latencies, memory bandwidth, and backend execution width (ALU, FPU, load/store ports). For GPUs the focus shifts to shader-core count, rasterizer-to-RTU ratio, and memory bus width.

What is a 'leak' in this community and how reliable are they?

A leak is an unsanctioned reveal of an unshipped design, typically sourced from supply-chain partners, early silicon samples, or internal documents. Accuracy varies widely—some land within a few percent of the final spec while others misread test silicon—so fans usually wait for a second corroborating source before treating a number as real.

What's the difference between a 'wide' and a 'deep' microarchitecture?

A wide design (e.g., Intel's recent P-cores) issues many operations per cycle across a large backend, while a deep design (e.g., older ARM little-cores) relies on a longer pipeline and more in-flight instructions. Neither is inherently superior; they simply target different power and latency budgets.

Why do fans care about a 5 % IPC bump?

At the same process node and clock, IPC is the purest measure of architectural efficiency, and a 5 % gain often represents months of redesign in the scheduler or cache hierarchy. It is the metric that separates a marketing refresh from a genuine generational step, which is why the community dissects every benchmark delta.

Explore the full GPU and CPU Microarchitectures codex →