Intel Core (microarchitecture)
Returned to efficient pipelines, enabling multi-core with lower power.
Jacek Halicki · CC BY-SA 4.0
The Intel Core microarchitecture is a multi-core processor microarchitecture launched by Intel in mid-2006. It succeeded the Enhanced Pentium M, the previous iteration of the P6 microarchitecture series, and replaced the NetBurst microarchitecture, which suffered from high power consumption and heat intensity due to an inefficient pipeline designed for high clock rate. The architecture was developed as Merom and provisionally referred to as Next Generation Micro-architecture.
Quick Facts
- Created
- June 26, 2006 (Xeon) / July 27, 2006 (Core 2) / <!--
- Model
- Celeron Series
- Model
- P6 family (Celeron, Pentium, Pentium Dual-Core, Core 2 range, Xeon)
- Numcores
- 1–4 (2-6 Xeon)
- Transistors
- 105M to 582M (65 nm) / 228M to 1900M (45 nm) / <!-- (A1, M0)
Facts from the source article.
Lore & Background
The Core microarchitecture was developed by Intel Israel, based in part on the Pentium M processor family. It was designed to deliver superior performance despite not reaching the high clocks of NetBurst, by using a short and efficient pipeline. The first processors using this architecture were code-named Merom (mobile), Conroe (desktop), and Woodcrest (servers and workstations). While architecturally identical, the three lines differed in socket, bus speed, and power consumption. Desktop and mobile processors were branded Core 2, later expanding to Pentium Dual-Core, Pentium, and Celeron brands; server and workstation processors were branded Xeon.
Features included Macro-Ops Fusion, which combined two x86 instructions into a single micro-operation (though not in 64-bit mode), and speculative execution of loads ahead of stores with unknown addresses. All 128-bit SSE instructions achieved 1 cycle throughput (previously 2 cycles). A new power-saving design ran all components at minimum speed, raising speed dynamically as needed. The architecture provided more efficient decoding stages, execution units, caches, and buses, reducing power consumption while increasing processing capacity. Core-based processors featured multiple cores, hardware virtualization (Intel VT-x), Intel 64, and SSSE3, but lacked hyper-threading technology. The consumer version also lacked an L3 cache, though it was present in high-end versions of Core-based Xeons.
Reader's Guide
The Core microarchitecture marked a strategic shift for Intel away from the high-clock, high-power NetBurst design toward power efficiency and multi-core scalability. By returning to lower clock rates and improving the use of available clock cycles and power, it enabled the transition to dual- and multi-core CPUs across all market segments. The architecture's 14-stage pipeline—less than half of Prescott's—allowed it to sustain up to 4 instructions per cycle, compared to the 3 IPC of its predecessors. Its shared L2 cache and power-saving technologies delivered significant performance-per-watt improvements: 20% more performance for Merom at the same power level compared to Core Duo, 40% more performance for Conroe at 40% less power compared to Pentium D, and 80% more performance for Woodcrest at 35% less power compared to the original dual-core Xeon. The architecture was later shrunk to 45 nm as the Penryn/Wolfdale generation, adding SSE4.1 and a new divide/shuffle engine. The Core microarchitecture laid the foundation for the follow-on Nehalem microarchitecture, which reintroduced hyper-threading and an L3 cache on consumer parts.
Did You Know?
- The Core microarchitecture was developed under the code name Merom, with development starting in 2001.
- Its pipeline is 14 stages long, less than half the length of Prescott's pipeline.
- Macro-Ops Fusion combines two x86 instructions into a single micro-operation, but does not work in 64-bit mode.
- The architecture lacks hyper-threading, which was present in Pentium 4 processors.
From Northbridge to Die: The HD Graphics Revolution
Before 2010, Intel's integrated graphics lived outside the processor itself, embedded in the motherboard's northbridge chip as part of what the company called its Hub Architecture. These chips went by names like Intel Extreme Graphics and Intel GMA, and they carried a well-earned reputation for underwhelming performance and limited feature sets. Gamers and power users largely dismissed them, favoring discrete cards from Nvidia or ATI/AMD instead. The Platform Controller Hub redesign changed that equation by eliminating the northbridge entirely and folding graphics processing directly into the processor package. When Clarkdale and Arrandale shipped in January 2010 under the new HD Graphics banner, they brought 12 execution units capable of up to 43.2 GFLOPS at 900 MHz, along with H.264 1080p video decoding at up to 40 frames per second. That was a meaningful step beyond the GMA X4500's 10 execution units at 800 MHz, which also lacked several capabilities. For the first time, Intel's integrated solution could hold its own against rival integrated adapters, reshaping expectations for what a CPU-embedded GPU could deliver.
Generations, Tiers, and the GTx Ladder
Intel organized its pre-Xe integrated GPUs into a clear hierarchy: each generation mapped to a specific microarchitecture (Gen5 through Gen11) with a corresponding instruction set, and within each generation, tiers were labeled GT1, GT2, or GT3e to signal increasing capability. Sandy Bridge in 2011 introduced hardware video encoding and HD postprocessing effects to its HD 2000 and 3000 variants. Ivy Bridge in 2012 pushed further, offering HD 2500 and 4000, with a special HD P4000 variant on Xeon processors that supported unbuffered ECC RAM. By the time Skylake arrived in 2015, Intel had retired legacy VGA support and enabled multi-monitor configurations of up to three displays over HDMI 1.4, DisplayPort 1.2, or eDP 1.3. Kaby Lake in 2016 added full hardware acceleration for 8- and 10-bit HEVC and VP9 decoding alongside 4K UHD premium streaming. Ice Lake then brought a 10-nanometer Gen 11 design featuring two HEVC 10-bit encode pipelines, three 4K display outputs, variable rate shading, and integer scaling. Each step up the ladder represented tangible gains in media processing, display flexibility, and computational throughput.
Crystalwell: eDRAM and the Iris Pro Experiment
One of the most architecturally distinctive moves in Intel's integrated graphics history came with Haswell in June 2013. The Iris Pro GT3e tier packed 128 MB of embedded DRAM into the same physical package as the CPU, but on a separate die fabricated using a different manufacturing process. Intel branded this memory as Crystalwell and classified it as a Level 4 cache shared between both the CPU and the GPU, a design choice that gave the graphics unit a dedicated high-bandwidth memory pool without consuming the main die's real estate. The Linux community caught up quickly, with the drm/i915 driver gaining awareness of and the ability to utilize this eDRAM starting with kernel version 3.12. The concept proved popular enough that Broadwell-K desktop processors, announced in November 2013 and aimed at enthusiast users, also carried Iris Pro Graphics. The eDRAM approach persisted in select Iris Pro and Iris Plus models across subsequent generations, establishing a precedent for heterogeneous memory packaging that would influence later Intel designs.
The Xe Era and the Arc Branding Shift
The naming and architecture landscape shifted dramatically as Intel transitioned from the Gen-based microarchitectures to the Xe family. Later products adopted architecture names like Xe-LP, Xe-LPG, Xe2, and Xe3, with Xe-cores housing vector engines whose number and capabilities varied from one architecture to the next. A critical distinction emerged: Intel Xe architectures power both integrated GPUs inside processor packages and discrete graphics cards, meaning the Intel Arc brand alone does not tell you whether you are looking at an on-package GPU or a standalone card. With Meteor Lake and the subsequent Core Ultra processors, Intel began marketing its integrated GPUs under both the Intel Graphics and Intel Arc names. Integrated Arc products now include the 130V and 140V in Lunar Lake, the 130T and 140T in Arrow Lake-H, and the B370 and B390 in Panther Lake. The Xe-LP generation introduced features such as AV1 8-bit and 10-bit fixed-function hardware decoding, Sampler Feedback, Dual Queue Support, and DirectX 12 View Instancing Tier 2, while dropping FP64 support entirely. Twin Lake N-series processors followed in the first quarter of 2025, continuing this evolving lineup.
Gallery






Frequently Asked Questions
What is the Intel Core microarchitecture?
It is a multi-core CPU design that Intel shipped in mid-2006, built around a 14-stage pipeline that can dispatch up to four instructions per cycle. Each core carries 64 KB of L1 cache split evenly between data and instructions, while all cores on the die share a common L2 cache.
What did Intel Core replace, and why was that change necessary?
Core succeeded the Enhanced Pentium M (the last P6-family design) and displaced the NetBurst line, whose extremely long pipeline chased peak clock speeds at the cost of enormous power draw and heat. By returning to a shorter, more efficient pipeline, Intel could finally pack multiple cores into a package without overheating the socket.
What was the project codename before the Core launch?
During development the design was known internally as "Merom," and Intel also used the placeholder label "Next Generation Micro-architecture" before settling on the Core branding at the mid-2006 release.
How does Core's caching hierarchy work?
Every core gets its own 64 KB L1 cache, divided into 32 KB for instructions and 32 KB for data, while the second-level cache is a single shared pool accessible by all cores on the die. This shared-L2 approach trimmed total silicon area compared to giving each core a private L2.
Why does the Core microarchitecture matter in CPU history?
It marked Intel's decisive pivot away from the "higher clock is better" philosophy of NetBurst toward efficient multi-core design, a strategy that shaped virtually every x86 desktop and server chip that followed. The 14-stage, four-wide pipeline and shared-cache layout became the template Intel refined for years afterward.
More in Intel and AMD Microprocessors, Part 2 1-24
Spotted an error? Know more?
Reader corrections go straight into our review queue. Suggest an edit · How this site is sourced
