Intel and AMD Microprocessors, Part 2 Codexery

Nehalem (microarchitecture)

Intel's 45 nm microarchitecture reintroduced Hyper-threading and Turbo Boost.

Nehalem (microarchitecture)

Nehalem is the code name for a 45 nm microarchitecture from Intel that launched in November 2008. It powered the first generation of Core i5 and i7 processors and marked a significant step forward from the older Core microarchitecture found in Core 2 chips, which itself was part of the P6 microarchitecture series that began with the Pentium Pro in 1995. The name comes from the Nehalem River.

Built on a 45 nm process, Nehalem could reach higher clock speeds without losing efficiency and was more energy-efficient than the earlier Penryn microprocessors. It brought back Hyper-threading, reduced the size of the L2 cache, and introduced a larger L3 cache shared across all cores. While the architecture was a radical departure from NetBurst, it kept a few minor features from that design. A die-shrink to 32 nm later produced Westmere, and the architecture was fully replaced by the second-generation Sandy Bridge in January 2011.

Technologically, Nehalem evolved from the Core microarchitecture with several changes. Cache line blocks on L2 and L3 caches shrank from 128 bytes in NetBurst and Merom/Penryn to 64 bytes per line, matching the size used in Yonah and Pentium M. Hyper-threading returned, along with Intel Turbo Boost 1.0. Some models included 2 to 24 MiB of L3 cache using Smart Cache. The Instruction Fetch Unit gained a second-level branch predictor with a two-level Branch Target Buffer and a Return Stack Buffer, and it supported all previous predictor types like the Indirect Predictor and Loop Detector. A second-level unified translation lookaside buffer (sTLB) held 512 entries for small pages and was 4-way associative. Each core had three integer ALUs, two vector ALUs, and two AGUs. Native quad-core, hexa-core, and octa-core processors were built on a single die. High-end desktop, server, and workstation models used Intel QuickPath Interconnect, while other models used Direct Media Interface, replacing the legacy front-side bus. Each core had 64 KB of L1 cache (32 KB data and 32 KB instruction) and 256 KB of L2 cache. Mid-range models integrated PCI Express and DMI into the processor, eliminating the northbridge. An integrated memory controller supported two or three channels of DDR3 SDRAM or four FB-DIMM2 channels.

Quick Facts

Created
November 11, 2008
Cores
1-6 (1-8 Xeon) / <!-- (2-12)
Transistors
731M to 2300M 45 nm / <!-- (C0, D0)
Clock
1.06 GHz to 3.33 GHz
Dmi-Slowest
2
Qpi-Slowest
4.80
Qpi-Fastest
6.40

Facts from the source article.

Lore & Background

Nehalem evolved from the Core microarchitecture by introducing several key changes. It reduced the cache line block on L2/L3 cache from 128 bytes to 64 bytes per line, matching the size used in Yonah and Pentium M. Hyper-threading technology was reintroduced, and Intel Turbo Boost 1.0 debuted. The architecture featured a 2–24 MiB L3 cache with Smart Cache in some models, and included an Instruction Fetch Unit with a second-level branch predictor, two-level Branch Target Buffer, and Return Stack Buffer. Nehalem also supported all predictor types previously used in Intel's processors, such as the Indirect Predictor and Loop Detector.

Per core, Nehalem provided 3 integer ALUs, 2 vector ALUs, and 2 AGUs. It enabled native quad-, hex-, and octa-core processors on a single die. High-end desktop, server, and workstation models used Intel QuickPath Interconnect, while other models used Direct Media Interface, replacing the legacy front side bus. The architecture integrated PCI Express and DMI into the processor in mid-range models, replacing the northbridge, and included an integrated memory controller supporting two or three memory channels of DDR3 SDRAM or four FB-DIMM2 channels. Second-generation Intel Virtualization Technology introduced Extended Page Table support, virtual processor identifiers, and non-maskable interrupt-window exiting. SSE4.2 and POPCNT instructions were added, and macro-op fusion now worked in 64-bit mode.

Reader's Guide

Nehalem's significance lies in its focus on performance, resulting in increased core size. Compared to Penryn, it delivered 10–25% better single-threaded performance and 20–100% better multithreaded performance at the same power level, while consuming 30% less power for the same performance. On average, Nehalem provided a 15–20% clock-for-clock increase in performance per core. Overclocking was possible with Bloomfield processors and the X58 chipset; Lynnfield processors used a PCH, removing the need for a northbridge. The architecture incorporated SSE4.2 SIMD instructions, adding seven new instructions to the SSE 4.1 set from the Core 2 series, and reduced atomic operation latency by 50% to eliminate overhead on operations such as the LOCK CMPXCHG compare-and-swap instruction. Nehalem later received a die-shrink to 32 nm with Westmere, and was fully succeeded by 'second-generation' Sandy Bridge in January 2011.

Did You Know?

More in Intel and AMD Microprocessors, Part 2 1-24

Spotted an error? Know more?

Reader corrections go straight into our review queue. Suggest an edit · How this site is sourced

Comments

Loading…
Open in the interactive codex →