Nvidia and AMD Graphics Processors Codexery

GeForce 900 series

High-end introduction to the Maxwell microarchitecture with improved energy efficiency.

GeForce 900 series

The GeForce 900 series is a lineup of GPUs from Nvidia, following the GeForce 700 series and marking the high-end debut of the Maxwell microarchitecture, which takes its name from physicist James Clerk Maxwell. These chips were built using TSMC’s 28 nm manufacturing process. With Maxwell—the successor to Kepler—Nvidia aimed for three main improvements over the GeForce 700 and 600 series: better graphics performance, easier programming, and greater energy efficiency.

The first generation of Maxwell, known as GM10x, includes chips like the GM107 and GM108. These appeared in products such as the GeForce GTX 745, GTX 750/750 Ti, and mobile GTX 850M/860M (GM107), as well as the GT 830M/840M (GM108). Nvidia focused on power savings rather than adding many new user-facing features. They increased the L2 cache from 256 KiB on the older GK107 to 2 MiB on GM107, which reduced the need for memory bandwidth. To further cut power, they narrowed the memory bus from 192 bits on GK106 to 128 bits on GM107. The streaming multiprocessor design also changed from Kepler’s SMX to a new layout called SMM. The warp scheduler structure stayed similar to Kepler’s, letting each scheduler issue up to two independent instructions from the same warp in order. In SMM, each of the four warp schedulers controls its own set of 32 FP32 CUDA cores, 8 load/store units, and 8 special function units. This contrasts with Kepler’s SMX, where four schedulers shared a pool of six sets of 32 FP32 CUDA cores, two sets of 16 load/store units, and two sets of 16 special function units, connected by a power-hungry crossbar. Maxwell removed that crossbar. Texture units and FP64 CUDA cores remain shared. SMM allows finer allocation of resources than SMX, saving power when workloads don’t use shared resources well. Nvidia claims a 128-core SMM achieves about 86% of the performance of a 192-core SMX. Each Graphics Processing Cluster (GPC) can hold up to four SMX units in Kepler, but up to five SMM units in first-generation Maxwell.

GM107 supports CUDA Compute Capability 5.0, compared to 3.5 on GK110/GK208 and 3.0 on GK10x GPUs. Features like Dynamic Parallelism and HyperQ, first seen in GK110/GK208, are available across all Maxwell products.

Quick Facts

Codename
GM20x
Architecture
Maxwell
Created
September 18, 2014
Model
GeForce series
Transistors
2.94B (GM206)
Fab
TSMC
Process
28 nm
Midrange
GeForce GTX 950 · GeForce GTX 960
Highend
GeForce GTX 970 · GeForce GTX 980
Enthusiast
GeForce GTX 980 Ti · Nvidia Titan X

Facts from the source article.

Lore & Background

The GeForce 900 series is built on the Maxwell microarchitecture, which Nvidia introduced in two generations. First generation Maxwell (GM10x) chips, such as the GeForce GTX 750/750 Ti, focused on power efficiency by increasing L2 cache from 256 KiB on GK107 to 2 MiB on GM107 and cutting the memory bus from 192-bit to 128-bit. Nvidia redesigned the streaming multiprocessor from Kepler's SMX to SMM, removing the crossbar that connected shared resources, allowing finer-grain allocation and saving power. Each SMM contains 4 warp schedulers, each controlling 32 FP32 CUDA cores, 8 load/store units, and 8 special function units. Nvidia claimed a 128 CUDA core SMM has 86% of the performance of a 192 CUDA core SMX.

Second generation Maxwell (GM20x) introduced technologies such as Dynamic Super Resolution, Third Generation Delta Color Compression, Multi-Pixel Programming Sampling, Nvidia VXGI, VR Direct, Multi-Projection Acceleration, and Multi-Frame Sampled Anti-Aliasing (MFAA), while removing support for Coverage-Sampling Anti-Aliasing (CSAA). HDMI 2.0 support was added. The ROP to memory controller ratio changed from 8:1 to 16:1. Second generation NVENC added HEVC encoding and support for H.264 encoding at 1440p/60FPS and 4K/60FPS, whereas first generation only supported H.264 1080p/60FPS. The GM206 GPU supports full fixed function HEVC hardware decoding.

Regarding asynchronous compute, while the Maxwell series was marketed as fully DirectX 12 compliant, Oxide Games uncovered that Maxwell-based cards do not perform well when async compute is utilized. Nvidia partially implemented it through a driver-based shim, coming at a high performance cost. Asynchronous compute on Maxwell requires that both a game and the GPU driver be specifically coded for it. The driver forces a Maxwell GPU to place all tasks into one queue and execute each task in serial. Oxide claimed that Nvidia pressured them not to include the asynchronous compute feature in their benchmark so that the 900 series would not be at a disadvantage against AMD's products.

Reader's Guide

The GeForce 900 series represents Nvidia's transition to the Maxwell microarchitecture, which prioritized energy efficiency and improved graphics capabilities over its Kepler predecessors. The series introduced significant architectural changes, including the SMM streaming multiprocessor design that reduced power consumption by removing the crossbar used in Kepler's SMX units. The first generation Maxwell chips demonstrated that substantial performance could be achieved with lower memory bandwidth and a narrower memory bus, thanks to increased L2 cache. The second generation added features like Dynamic Super Resolution and HDMI 2.0, along with enhanced video encoding and decoding capabilities.

However, the series was marked by controversy. The GeForce GTX 970's specifications were found to differ from those initially announced: it had 1.75 MB of L2 cache versus 2 MB in the GTX 980, 56 ROPs versus 64, and its memory was divided into a 3.5 GB section and a 0.5 GB section, with access to the latter being 7 times slower. Nvidia issued a statement acknowledging the altered specifications and later apologized. A class-action lawsuit alleging false advertising was filed against Nvidia and Gigabyte Technology. Additionally, the series faced criticism for its implementation of asynchronous compute, which relied on a driver-based shim and performed poorly compared to AMD's hardware-based solution, leading to the feature being disabled by the driver for Maxwell.

Did You Know?

More in Nvidia and AMD Graphics Processors 1-24

Spotted an error? Know more?

Reader corrections go straight into our review queue. Suggest an edit · How this site is sourced

Comments

Loading…
Open in the interactive codex →