2  GPU Hardware and Memory

Explain how GPU execution units, the memory hierarchy, and interconnects determine the cost of moving and processing model tensors.

A GPU can execute transformer operations quickly only when enough parallel work reaches its arithmetic units and reused data stays close to them. This chapter follows tensors from registers and shared memory through L2 and HBM, then across links between GPUs and servers, so later optimizations can be tied to the hardware resource they use.

A GPU combines device-wide memory and cache with many Streaming Multiprocessors (SMs) (the GPU’s independent processing units, Section 2.2), each containing local execution and storage resources.

Two GPU devices each contain HBM, shared L2 cache, and Streaming Multiprocessors. One SM is expanded into registers, shared memory, warp scheduling, and Tensor Cores. NVLink joins device boundaries. Section labels identify chapter 2 section map: 2.1 - CPU and GPU strengths; 2.2 - GPU chip structure; 2.3 - Memory hierarchy; 2.4 - Tensor cores and shape alignment; 2.5 - Intra-node interconnects; 2.6 - Interconnect hierarchy beyond HBM; 2.7 - Hardware limits on software design; 2.8 - Work, bytes, and time.
Figure 2.1: GPU resources are nested within a device, while NVLink connects device boundaries.

High Bandwidth Memory (HBM) and the L2 cache (an on-chip cache shared by all SMs) serve the whole device. Registers, shared memory, warp scheduling (the SM’s rotation among groups of 32 threads), and Tensor Cores belong inside an SM. The second device makes the boundary explicit: NVLink (NVIDIA’s direct GPU-to-GPU link) carries peer traffic between GPUs rather than connecting an SM directly to another device’s cache.

This chapter supplies the hardware facts behind two of the levers in Section 1.6.1: move fewer bytes, because the memory hierarchy decides what reuse costs, and use faster hardware paths, because Tensor Cores reach their peak rate only for the shapes and formats they support.

2.1 CPU and GPU strengths

CPUs emphasize latency and branch-heavy control flow. They devote more chip area to caches and complex execution logic. GPUs emphasize throughput. They devote more area to arithmetic units and hardware scheduling of many simple threads (independent sequences of instructions, each working on its own data). To use a GPU effectively, a workload must expose enough parallel work to keep those arithmetic units busy and regular enough memory access to use the bandwidth efficiently. CPU-bound work, like request parsing or complex tokenization, often stays on the host.

CPU and GPU designs emphasize different kinds of work. As a broad design tendency, CPUs devote more resources to low-latency control, caches, and irregular execution, while GPUs devote more resources to throughput across many similar operations. Treat the visual as a conceptual comparison rather than a literal chip floorplan. Modern CPUs and GPUs overlap in capability, but the throughput emphasis of GPUs explains why regular tensor shapes, batching, and memory coalescing matter.
Figure 2.2: CPU and GPU design trade-off

2.2 GPU chip structure

GPU chip structure is the organization of execution units, on-chip storage (memory built into the GPU chip itself), caches, and off-chip memory that a kernel uses while it runs. Computation happens in execution units, while data is staged and reused through several storage levels with different scopes and costs.

GPU performance depends on registers, per-block shared memory, hardware-managed caches, High Bandwidth Memory (HBM), and external interconnects. Each level differs in scope, capacity, lifetime, and how software influences placement or reuse.

GPU performance depends on registers, per-block shared memory, hardware-managed caches, High Bandwidth Memory (HBM), and external interconnects. Each level differs in scope, capacity, lifetime, and how software influences placement or reuse. The capacities printed in this source slide describe a particular accelerator rather than every GPU. Do not read the tiers as one compulsory path for every access: instructions, caches, shared-memory staging, and interconnect traffic depend on the operation and kernel.
Figure 2.3: GPU memory hierarchy

The capacities printed in the figure describe a particular accelerator rather than every GPU. The tiers are not one compulsory path for every access: instructions, caches, shared-memory staging, and interconnect traffic depend on the operation and kernel.

A Streaming Multiprocessor (SM) is a GPU execution unit with a finite pool of registers, shared memory, scheduling slots, and arithmetic resources. The CPU launches a kernel, a GPU function executed by many logical threads. Its launch groups those threads into blocks, whose threads can share on-chip storage and synchronize. In the baseline CUDA execution model, each block executes on one SM, and several blocks can reside there together if resources permit. A warp is a scheduling group of 32 threads within a block on the NVIDIA GPUs described by the CUDA C++ Programming Guide 12.4. The SM schedules these warps, while L2 and HBM serve the whole device.1 Chapter 3 develops indexing and scheduling in detail.

One word, four meanings: block

Thread block: the group of threads that a kernel launch creates together, which shares on-chip memory and can synchronize. This is the meaning used here and in Section 3.2.

Block of data, also called a tile: the rectangular piece of a tensor that one Triton program instance loads and computes, sized by constants such as BLOCK_SIZE (Section 4.3). It names a piece of data, not a group of threads.

Transformer block: one layer of the model, meaning its attention and feed-forward sub-layers together (Section 1.1).

KV-cache block: the fixed number of token positions whose keys and values are stored in one allocation (Section 9.5). It names a unit of memory management.

The subject under discussion decides which one is meant.

Resident blocks and occupancy

Registers are used by individual threads, but all resident threads draw from the SM’s finite register capacity. Shared memory is allocated per block, but all resident blocks draw from the SM’s shared-memory capacity. Admission must satisfy both limits together, as well as the device’s block and thread limits.

Occupancy is the ratio of resident active warps on an SM to its maximum supported resident warps. High occupancy can hide latency, but it does not guarantee high throughput when memory access, instruction mix, or dependencies are poor.

The chip structure matters because every kernel consumes scarce per-SM resources: registers, shared memory, scheduler slots, and memory bandwidth. A kernel with too many registers per thread may reduce the number of active warps. A kernel with too much shared memory per block may reduce how many blocks can reside on an SM. A kernel with poor memory locality can keep arithmetic units idle even when many threads are present.

2.3 Memory hierarchy

GPU memory is not one uniform pool. Performance depends on which level serves the data, who can see it, and how directly the programmer can control it.

  • Registers: Fastest storage, private to one thread, and used for scalar temporaries and small per-thread fragments. Register pressure (a kernel’s demand on the finite register file) can reduce occupancy because each Streaming Multiprocessor (SM) has a finite register file.
  • Shared memory: On-chip memory shared by threads in one block, also called a Cooperative Thread Array (CTA). It is explicitly managed and useful for tiled reuse, but capacity is small and bank conflicts can waste bandwidth.
  • L2 cache: Device-wide cache shared across SMs. It can capture cross-SM or cross-kernel locality, but programmers usually control it indirectly through access patterns and kernel choices rather than direct allocation.
  • High Bandwidth Memory (HBM): Large off-chip GPU memory that stores model weights, activations, gradients, optimizer state, and KV cache. Its bandwidth is high in absolute terms, but it is still much slower than on-chip storage.
  • Host memory and interconnect: CPU memory reached over PCIe (the standard expansion bus that connects GPUs, network cards, and other devices to the CPU) or another system path. Transfers are much slower than on-device reuse, so repeated host/device movement in frequently executed code is usually wasteful.

Memory-tier caveat

The phrase ‘GPU memory’ hides large cost differences among registers, shared memory, L2, HBM, and host transfers. A kernel can fit in HBM and still be slow because it fails to reuse data closer to the SM.

Table 2.1: Memory hierarchy comparison.
Level Scope Engineering role Main risk
Registers Thread Fast scalar and vector temporaries Too many registers reduce occupancy
Shared memory Block / CTA Tile reuse and cooperative staging Bank conflicts or capacity pressure
L2 cache Device Cross-SM reuse and caching Limited control and eviction
HBM Device Large tensor storage Bandwidth and latency dominate if reuse is poor
Host / interconnect System or cluster Data loading and multi-GPU movement Transfer latency and bandwidth

GPU caches are hardware-managed storage between Streaming Multiprocessors and HBM. CUDA memory spaces such as registers, shared memory, global memory, and constant memory expose explicit placement choices. Ordinary L1 and L2 cache residency mostly follows access pattern, instruction choice, and hardware replacement policy.

The per-SM L1 or unified data cache is closest to executing warps. It can reduce latency for repeated or nearby accesses inside work scheduled on the same SM. Cross-SM reuse belongs mainly to the device-wide L2 cache, which sits between kernel load/store traffic and HBM. L2 can capture reuse across SMs or across nearby kernel launches, while remaining much smaller than model weights, activations, or long KV-cache state.

A cache line holds a fixed-size, aligned region of memory together with information identifying its address. The cache’s address-mapping rule assigns each line to a set, the group of candidate locations where it may be kept. Each location in that set is a way. Associativity is the number of ways per set. If more needed lines map to one set than it has ways, one must be evicted even when another set has free space. Returning to the evicted line produces a conflict miss. Exact NVIDIA GPU set mapping and replacement behavior are architecture-specific, not a stable CUDA programming contract.

Worked example: Free capacity with repeated cache misses

Consider a hypothetical cache with two sets and two ways per set, initially empty. Let the set be line_number mod 2, and let replacement evict the least recently used line. Accessing lines 0, 2, 4, 0 sends every access to set 0. Line 0 misses and occupies one way. Line 2 misses and fills the other. Line 4 misses and evicts line 0. The final access to line 0 misses again and evicts line 2. Set 1 remains empty throughout, despite these evictions. If a changed layout places the third item in line 5 instead, accesses 0, 2, 5, 0 put line 5 in set 1, and the final access to line 0 hits. This illustrates how addresses can change reuse without changing the number of items. The modulo mapping and replacement rule are teaching assumptions, not an NVIDIA cache specification. Actual layout changes need profiler evidence.

The practical optimization question is therefore not ‘which cache set owns this tensor?’ The useful questions are whether adjacent lanes access adjacent addresses, whether the working set fits within the level that should reuse it, whether a stride pattern repeatedly fights for the same limited cache resources, how many cache sectors (fixed-size units of cache traffic, defined below) are requested, how many requests miss to L2 or HBM, and whether reused data should be staged explicitly in shared memory or registers.

  • L1 or unified data cache: Per-SM cache path for local load/store traffic. It helps locality inside work scheduled on the same SM, but it is not a general cross-SM state store.
  • L2 cache: Device-wide cache shared by SMs. It is the main hardware cache level to discuss for cross-SM global-memory locality and L2 persistence controls.
  • Cache sector: Profiler-visible access granularity. Nsight Compute reports sectors as 32-byte chunks, with four sectors in a 128-byte L1 or L2 cache line on documented NVIDIA profiling views.
  • Associativity: The ways-per-set limit just illustrated. It matters when address layout or stride creates conflict misses even though total cache capacity seems sufficient.
  • Cache control boundaries: The hardware limits direct placement, but provides specific controls:
    • cudaFuncSetCacheConfig: Requests a preferred L1/shared-memory split on supported devices.
    • L2 persistence window: Gives selected accesses preferential retention across kernel launches within the supported cache limits.
    • PTX eviction hints: Give streaming data a lower retention priority to reduce interference with reused lines. They do not guarantee that reused data remains cached.2
    • Profiler-counter proof: Verifies hit rates to confirm cache effectiveness. None of these controls assigns an arbitrary tensor to a cache set or pins an ordinary cache line.

L2 cache hit rates (the fraction of requests found in the cache) can drop if strides repeatedly map different addresses to the same cache sets. This is an associativity or conflict miss. Reorganizing the layout or padding the dimension can reduce these misses, though the exact hardware hash function is not documented.

Cache controls help only when the access pattern creates reuse and profiler counters show that the hardware retains it. They cannot repair scattered accesses or force exact residency. Once operands reach an SM, the next question is whether their shapes and data types let Tensor Cores perform the arithmetic efficiently.

2.4 Tensor cores and shape alignment

Tensor Cores are specialized Matrix Multiply-Accumulate (MMA) units inside Streaming Multiprocessor partitions. They consume supported matrix tiles at very high throughput. One H100 SXM configuration has 132 SMs and four fourth-generation Tensor Cores per SM, but exact counts and available throughput depend on the GPU variant and enabled chip resources. This configuration is documented in NVIDIA Hopper Architecture In-Depth.3

Multiply-accumulate means multiplying corresponding operand values and adding the products into a running result. For a matrix tile, an operand of shape \(m\times k\) multiplies one of shape \(k\times n\) and adds into an \(m\times n\) accumulator. The result has \(m\) rows and \(n\) columns. Here \(k\) is the shared dimension over which products are summed, not a third output axis. A tile described as \(m\times n\times k=16\times16\times16\) therefore consumes two \(16\times16\) operand tiles and updates a \(16\times16\) result. Exact supported instruction shapes depend on architecture, dtype, and kernel path.4

Worked example: Padding a reduction dimension

Take \(A\) with rows [1,2,3] and [4,5,6], and \(B\) with rows [1,0], [0,1], and [1,1]. Their shapes are \(2\times3\) and \(3\times2\), so the output has shape \(2\times2\). With a zero starting accumulator, its upper-left value is \(1\times1+2\times0+3\times1=4\).

For an illustrative software tile with \(m=n=k=2\), the first two positions of the shared dimension contribute rows [1,2], [4,5]. The remaining position contributes [3,3], [6,6]. Padding \(A\) with a zero fourth column and \(B\) with a zero fourth row lets that second tile process two positions: its added zero products leave the result unchanged. Summing the two contributions gives output rows [4,5] and [10,11]. All four output entries are retained. The padded calculation performs 16 multiply-add pairs instead of 12, so easier tiling comes with extra work. This small software tile illustrates the arithmetic, not an available Tensor Core instruction. A kernel can also handle the remainder separately. Actual speed depends on the selected implementation.

This is why hidden sizes, head dimensions, batch shapes, and padding can affect performance. Padding may increase nominal arithmetic while improving utilization on a suitable Tensor Core path. Section 3.7 develops the full GEMM shape relation.

In modern transformer GEMM, ‘compute-bound’ (Section 2.8) often means ‘tensor-core-bound’: the useful question is whether the operation reaches the specialized matrix units efficiently. Verify this in traces rather than assuming it. Nsight Systems kernel names, Nsight Compute instruction metrics, and achieved tensor-core utilization can show whether the intended path was actually used.

Shape alignment depends on the library path and GPU architecture as well as the mathematical matrix dimensions. NVIDIA’s matrix-multiplication guidance reports that modern cuBLAS and cuDNN can often use Tensor Cores even when dimensions are imperfectly aligned, while efficiency improves when the fastest-varying matrix dimensions align to hardware-friendly byte multiples. For FP16, a common practical rule is to prefer multiples of 8 elements, with larger multiples often helping on A100-class paths. Exact requirements vary by CUDA library version, dtype, layout, and GPU generation. See Matrix Multiplication Background User’s Guide.5

A poorly aligned matrix shape can trigger internal padding, a less efficient tile path, a different kernel, or Tensor Core execution with lower utilization. A trace relates the actual kernel path, achieved utilization, and tensor dimensions.

The tensor-core figure compares two tile-utilization cases. It is a schematic utilization comparison, not a timing benchmark.

Tensor Cores execute supported matrix multiply-accumulate instructions on small hardware fragments. Kernels combine those fragments into warp-, block-, and software-level tiles that can be much larger than one instruction.

Tensor Cores execute supported matrix multiply-accumulate instructions on small hardware fragments. Kernels combine those fragments into warp-, block-, and software-level tiles that can be much larger than one instruction. The 64 by 64 regions in the visual are software or kernel tiles, not the dimensions of one Tensor Core instruction. Alignment and remainder handling can affect utilization, while exact instruction shapes and Tensor Core counts depend on GPU architecture, dtype, and SKU.
Figure 2.4: Tensor Core tiling and utilization

The 64 by 64 regions in the visual are software or kernel tiles, not the dimensions of one Tensor Core instruction. Alignment and remainder handling can affect utilization, while exact instruction shapes and Tensor Core counts depend on GPU architecture, dtype, and SKU.

Tensor Core speed requires a supported numerical format and an execution path that selects the accelerated kernel. Suitable dimension alignment can improve efficiency without being a universal requirement for Tensor Core use. When an operation spans several GPUs, the next question is whether the local interconnect can move partial results quickly enough to preserve that compute gain.

2.5 Intra-node interconnects

The earlier sections described one GPU’s execution and memory hierarchy. Scaling starts when tensors must cross a physical boundary: between GPUs in one server or between servers in a cluster. A GPU is one execution device. A board or accelerator card packages one or more GPUs and their local HBM. A node, also called a server, contains CPUs, host memory, PCIe root complexes, network interface cards, and one or more GPU boards. Intra-node traffic stays inside one server. Inter-node traffic crosses the cluster network between servers.

Intra-node topology starts before GPU-to-GPU links. Multi-socket servers often have Non-Uniform Memory Access (NUMA): CPU cores, host memory channels, PCIe root complexes, GPUs, and Network Interface Cards (NICs) (the devices that connect a server to the cluster network) are closer to some sockets than others. A process placed on the wrong CPU socket can feed a nearby-looking GPU through a slower path.

Software placement determines which of these paths carries an input batch. A process is a running operating-system program instance. A rank is its integer identity within a distributed group, not a GPU. Selecting a GPU device chooses where that process submits device work. CPU affinity restricts which CPU cores may run the process, while host-memory placement is a separate NUMA policy. Optional data-loading workers are helper processes that prepare batches for the training process. They do not automatically acquire distributed ranks or their own GPUs.

Worked example: Two ranks feeding two local GPUs

Assume one server with two CPU sockets, GPU 0 attached near socket 0 and GPU 1 near socket 1, and one training process per GPU. Process P0 has rank 0, CPU affinity on socket 0, and selects GPU 0. P1 has rank 1, runs on socket 1, and selects GPU 1. These rank-to-device matches are explicit choices in this example. Other mappings are possible.

For P1, optional loader workers prepare a batch using CPU cores and host memory placed near socket 1. P1 submits its input copy through the GPU’s local PCIe path into GPU 1’s HBM, then launches work on that data. If the batch’s host pages are instead on socket 0, the copy may also cross the inter-socket link. CPU affinity alone does not move those pages or select GPU 1. Local placement can reduce this extra traffic, but its effect needs measurement. Chapter 12 extends this setup to communication groups and multiple servers.

PCIe connects GPUs to the host and sometimes to each other, but it is much slower than on-package or board-level GPU interconnects. NVLink provides higher bandwidth GPU-to-GPU communication inside a node. NVSwitch (a switch chip that joins many NVLink GPUs) extends this idea to many GPUs. These fabrics matter for tensor parallelism and frequent collectives.

Table 2.2 provides an initial scale comparison: it shows how the available byte rate usually falls as traffic leaves one GPU. These are planning ranges, not measured ceilings for the current system.

Table 2.2: Typical movement paths and bandwidth ranges for planning. Actual values depend on the hardware generation and topology.
Path Typical scope Typical bandwidth How to use the number
HBM On one GPU Several TB/s on data-center GPUs Use as the roofline memory ceiling for local kernel reads and writes.
NVLink / NVSwitch GPU-to-GPU inside supported systems Hundreds of GB/s to around TB/s-class aggregate paths depending on generation and topology Useful for frequent tensor-parallel or collective traffic inside the fast local domain.
PCIe CPU/GPU or device path through PCIe Tens of GB/s per direction for current high-end links Good for setup and staged copies, but too slow for repeated hot-loop movement.
InfiniBand / RoCE Across nodes Tens to hundreds of GB/s effective per node only on engineered fabrics Use for scale-out communication after confirming the runtime selected the intended NIC and transport.

2.6 Interconnect hierarchy beyond HBM

An interconnect is the path used to move data between processors, accelerators, host memory, network cards, or nodes. After data leaves HBM or one GPU, the cost model changes: bandwidth usually drops, latency rises, and synchronization becomes more visible.

Moving a host batch to the GPU also depends on its memory type. Pageable memory is ordinary host memory whose physical pages the operating system may move or page out. Pinned memory is host memory held in place for transfers. Direct Memory Access (DMA) lets a device move the payload without the CPU copying each byte. On a pageable host-to-GPU route, the runtime may first copy the payload with the CPU into a pinned staging buffer. A copy engine then transfers it to HBM. Starting from a pinned buffer can remove that staging copy. The CPU still allocates buffers and submits work, and dependent computation must wait for transfer completion. Pinning consumes host resources and does not itself prove overlap. Section 3.3 gives the CUDA lifecycle.6

Remote DMA (RDMA) lets network hardware transfer data to or from registered memory on another machine, with software setting up the accessible buffers and communication. GPUDirect RDMA extends the direct path to compatible GPU memory: the NIC can access HBM through PCIe, avoiding an intermediate host payload buffer. A staged GPU-to-network path copies from HBM to host memory before the NIC sends the bytes. A supported direct path transfers between the NIC and HBM. CPU coordination and completion ordering remain necessary. Compatible devices, drivers, memory registration, and topology determine whether this path works, as described in the NVIDIA GPUDirect RDMA documentation. Section 12.4 develops those network conditions.

  • PCIe: General host/device expansion bus for CPUs, GPUs, NICs, and other devices.
  • NVLink: High-bandwidth GPU-to-GPU link available only on supported systems and topologies.
  • NVSwitch: Switch fabric that connects multiple NVLink-capable GPUs inside a server or platform.
  • NIC: Network Interface Card, the device that connects a node to the cluster network.
  • InfiniBand: High-performance cluster fabric commonly used for multi-node GPU training and serving.
  • RoCE: RDMA over Converged Ethernet. Performance depends strongly on Ethernet fabric configuration.
  • Ethernet: General-purpose network fabric. It is ubiquitous, but unoptimized paths are often weak for frequent GPU collectives.

After HBM, the next distance is usually another GPU or another node. Moving tensors across PCIe, NVLink, NVSwitch, InfiniBand, RoCE, or Ethernet is far more expensive than reading registers or shared memory. Distributed training and distributed inference therefore add explicit communication costs and synchronization points to the single-GPU execution model.

Each memory or communication path is a separate measured resource. Register and shared-memory reuse usually avoids slower device-memory traffic. An HBM load uses the GPU memory controllers. A peer transfer (a direct copy between two GPUs) uses NVLink, NVSwitch, or PCIe. A cross-node transfer adds a NIC and network fabric. Their relative cost depends on message size, topology, contention, and overlap, so the selected path, effective bandwidth, and latency require measurement.

A memory-transfer path is the hardware route used to move bytes from one memory domain to another. For GPU workloads, the important domains are CPU system memory, pinned host memory, one GPU’s HBM, another GPU’s HBM, and memory reachable through a Network Interface Card (NIC). These domains differ in latency, bandwidth, address-translation rules, and synchronization behavior.

The performance mistake is to treat every transfer as the same kind of HBM traffic. A kernel reading its own tensor, a host batch copied into GPU memory, a peer GPU copy, and a cross-node collective all move bytes, but they use different software owners and hardware paths.

Copy engines and caches solve different movement problems. A copy engine is used for explicit transfer work, such as host-device copies or some device-device copies. The cache hierarchy serves ordinary kernel load/store instructions as warps read and write tensors. A slow trace can involve both: a batch might first move through a DMA copy path into HBM, then a kernel might waste HBM bandwidth because its warp loads miss cache or access scattered sectors.

The following principles and table answer a practical design question: which software layer initiates each path, which physical route carries the bytes, and which condition should be checked before optimizing it.

Transfer design principles

Pinned-memory staging: Use page-locked host buffers when repeated host-to-device copies need efficient DMA and possible overlap.

Compute near data: Run the operation where the tensor already lives instead of moving the tensor to a more convenient but distant processor.

Direct path validation: Confirm that a peer, NVLink/NVSwitch, or RDMA path is actually available before assuming the fast route exists.

Distance by frequency: The more often data moves, the closer the communicating devices must be (Section 12.11).

After a trace identifies a transfer, Table 2.3 answers a different question: which software submitted it, which physical route likely carried it, and what evidence can distinguish a direct path from staging or fallback.

Table 2.3: Transfer paths: initiating software, physical route, and diagnostic check.
Path Typical owner Hardware route What to verify
Kernel HBM access Kernel code, vendor library, Triton, TorchInductor SM load/store through L2 and HBM controllers Uncoalesced access or poor reuse wastes HBM bandwidth
Host to GPU PyTorch tensor placement, CUDA runtime Pinned/pageable host memory through PCIe or similar host path Pageable staging, NUMA mismatch, hidden synchronization
GPU to GPU CUDA peer copy, NCCL, distributed framework NVLink, NVSwitch, or PCIe peer path Peer access unavailable, topology asymmetry, host bounce
GPU to network NCCL, UCX/MPI, NIC driver NIC/RDMA path through InfiniBand, RoCE, or Ethernet Wrong NIC, no GPUDirect RDMA, congestion, rank skew

Name the data path before choosing a fix. A profiler copy span might be a host-device DMA transfer, a peer GPU copy, a network transfer, unified-memory migration, or explicit kernel load/store traffic. Each path requires a different change.

2.7 Hardware limits on software design

Software design is constrained by the cheapest place where data can be reused. If values are reused within a thread, registers are ideal. If values are reused by a block, shared memory can help. If values are reused across kernels or requests, L2, HBM layout, and serving-engine cache policies matter. Distributed systems add another hierarchy: GPU-to-GPU, node-to-node, and storage-to-host movement.

VRAM means the GPU device-memory capacity available for tensors, model weights, activations, workspaces, and caches. On the datacenter GPUs discussed in this module, that device memory is High Bandwidth Memory (HBM).

For many current LLMs, one device cannot hold the weights at the desired precision together with KV cache, activations, temporary workspaces, and allocator headroom. The raw weight byte count and the workload-specific state together determine whether quantization, sharding, or offload is needed. Table 2.4 compares specific GPU variants because capacity and bandwidth depend on that choice. The Tesla P100 SXM2 16 GB row uses the Tesla P100 Datasheet7, the Tesla V100 SXM2 32 GB row uses the Tesla V100 Datasheet8, the A100 80 GB SXM row uses the NVIDIA A100 specifications9, the H100 SXM5 80 GB row uses the NVIDIA H100 product specifications10, and the B200 row derives per-GPU values from the NVIDIA DGX B200 system totals11.

Table 2.4: Example GPU capacity and raw-weight sizing limits.
Example GPU VRAM capacity Peak bandwidth Example model scale What else uses memory
Tesla P100 SXM2 16 GB (GP100 / Pascal) 16 GB (HBM2) 732 GB/s BERT-Large (~340M parameters) FP16 weights are about 0.68 GB; BERT-Large is a small model by current LLM-serving scale
Tesla V100 SXM2 32 GB (GV100 / Volta) 32 GB (HBM2) 900 GB/s GPT-2 (~1.5B parameters) Fits in a single GPU; requires ~3 GB in FP16, leaving room for activations
A100 80 GB SXM (GA100 / Ampere) 80 GB (HBM2e) 2.0 TB/s GPT-3 (~175B parameters) FP16 weights (~350 GB) require multi-GPU sharding. Crossing nodes depends on GPUs and usable memory per server plus workload headroom
H100 SXM5 80 GB (GH100 / Hopper) 80 GB (HBM3) 3.35 TB/s Llama 3-70B (~70B parameters) FP16 weights need two or more GPUs. Raw FP8 weights can fit one 80 GB GPU only with an explicit KV-cache and workspace budget; INT4 leaves more headroom
B200 (DGX B200 / Blackwell, HBM3e) 180 GB (HBM3e) 8.0 TB/s Llama 3-405B (~405B parameters) Requires multi-GPU sharding. An 8-GPU node is a common deployment unit; final GPU count depends on weight dtype plus KV-cache and workspace budget

Worked example: GPU count is not server count

At two bytes per parameter, 175 billion parameters occupy about 350 GB before other state. That exceeds one 80 GB GPU, so this placement needs multiple GPUs. A hypothetical server with eight such GPUs has \(8\times80=640\) GB of aggregate raw capacity. Thus 350 GB alone does not establish a need for multiple servers. The 640 GB is spread over eight devices, not one allocation pool: a sharding scheme must fit each device’s share together with activations, cache, workspaces, and other required state. The available GPUs and usable memory per server determine whether the workload must cross nodes.

2.8 Work, bytes, and time

Section 1.4 estimated that one generated token of an 8B model needs about 16 GFLOP, roughly 16 microseconds at an H100’s peak arithmetic rate. For a model whose projection weights exceed cache capacity, a small decode step typically reads most of those weights from HBM, so memory traffic can set the time. Arithmetic and data movement each impose a lower bound. Their maximum is an ideal estimate when they overlap completely. Choosing the right optimization therefore starts with knowing which of the two is slower.

To a first approximation, arithmetic takes FLOPs divided by the peak compute rate, and data movement takes bytes moved divided by the memory bandwidth. Their ratio depends on one property of the operation:

  • Arithmetic intensity: FLOPs performed per byte moved to or from memory, measured in FLOP/byte. It states how much arithmetic each fetched byte supports.
  • Ridge point: the peak compute rate divided by the memory bandwidth. At this intensity the two times are equal.
  • Memory-bound and compute-bound regions: below the ridge point, the bandwidth ceiling is lower than the compute ceiling. Above it, the compute ceiling is lower. These are the model’s limiting resources. A measured kernel can still be limited by launch overhead, latency, or poor utilization.

For an H100 SXM GPU, the dense BF16 Tensor Core peak is about 989 TFLOP/s. NVIDIA’s specification lists 1,979 TFLOP/s with structured sparsity, twice the dense rate. With 3.35 TB/s of HBM bandwidth, the ridge point is \(989\times10^{12}/3.35\times10^{12}\approx295\) FLOP/byte.12 A BF16 operation must therefore perform about 295 arithmetic operations for every byte it moves before the arithmetic units, rather than memory, become the limit.

Worked example: One projection at four batch sizes

Inputs: A BF16 weight matrix of shape \(4{,}096\times4{,}096\) multiplies a batch of \(B\) token vectors of width 4,096, the query projection from Section 1.1. Each weight is used once per row in one multiply and one add, so the operation performs \(2\times B\times4{,}096^2=33{,}554{,}432\,B\) FLOPs. It reads the weights once (\(2\times4{,}096^2=33{,}554{,}432\) bytes) and reads and writes \(B\) rows of 4,096 two-byte values (\(16{,}384\,B\) bytes). The intensity is therefore \(33{,}554{,}432\,B/(33{,}554{,}432+16{,}384\,B)=B/(1+B/2{,}048)\) FLOP/byte.

Table 2.5: Work, estimated HBM bytes, and ideal time lower bounds for one 4,096 by 4,096 projection as the row count grows.
Rows \(B\) FLOPs Bytes moved Intensity (FLOP/byte) Time at 3.35 TB/s Time at 989 TFLOP/s Limit
1 33.6 M 33.6 MB 1.0 10.0 µs 0.03 µs Memory
64 2.15 G 34.6 MB 62 10.3 µs 2.2 µs Memory
345 11.6 G 39.2 MB 295 11.7 µs 11.7 µs Balanced
2,048 68.7 G 67.1 MB 1,024 20.0 µs 69.5 µs Compute

Conclusion: With one row, which is one token of single-request decode, each weight byte supports about one FLOP, nearly 300 times below the ridge point. Under this HBM-traffic estimate, the memory-time lower bound is about 10 microseconds. Adding rows reuses each loaded weight, so this ideal bound grows slowly until about 345 rows, where the compute and memory bounds meet. This is why prefill and training, which process many tokens at once, tend to be compute-bound, while small-batch decode tends to be memory-bound. The estimate assumes that the weights come from HBM and that copying and computing overlap perfectly.

These two times give a lower bound, not a prediction. A kernel can be slower than both because of launch overhead, too little parallel work, poor memory access patterns, or synchronization. Chapter 5 turns this estimate into the roofline model, fixes the rules for counting FLOPs and bytes, and shows how profiler measurements expose the gap. Section 3.7 applies the same reasoning to matrix-vector products, and Chapter 9 to a whole decode step.

CUDA makes this hierarchy programmable. The next chapter names the host, device, kernel, grid, block, thread, warp, stream, and memory-allocation objects that software uses to express work on this hardware.


  1. NVIDIA. (2024). CUDA C++ Programming Guide 12.4. https://docs.nvidia.com/cuda/archive/12.4.0/cuda-c-programming-guide/index.html (release 2024-03-05). Support: thread hierarchy, warps of 32 threads, memory hierarchy, streams, events, and coalesced access guidance used in Sections 2.2 and 2.3. Limit: exact SM counts, cache sizes, and bandwidths are architecture specific and are cited separately where stated.↩︎

  2. NVIDIA. (2024). Parallel Thread Execution ISA 8.4, sections 9.7.8.1-9.7.8.2. Cache operators and eviction priority hints. Cache and eviction policies are performance hints, not guarantees of residency or memory ordering.↩︎

  3. Andersch, M., Palmer, G., Krashinsky, R., Stam, N., Mehta, V., Brito, G., & Ramaswamy, S. (2022). NVIDIA Hopper Architecture In-Depth (NVIDIA Technical Blog, 2022-03-22). https://developer.nvidia.com/blog/nvidia-hopper-architecture-in-depth/ . Support: H100 SXM5 has 132 SMs with 4 fourth-generation Tensor Cores per SM; full GH100 has 144 SMs; HBM3 memory and 50 MB L2 cache configuration. Limit: counts describe the stated SXM5 configuration; PCIe and NVL variants differ and exact throughput depends on clocks and enabled resources.↩︎

  4. NVIDIA. (2023). Matrix Multiplication Background User’s Guide (last updated 2023-02-01). https://docs.nvidia.com/deeplearning/performance/dl-performance-matrix-multiplication/index.html. Support: Tensor Core use with cuBLAS 11.0 and cuDNN 7.6.3 and later without strict alignment, with best efficiency for FP16 at multiples of 8 elements and larger multiples on A100 paths. Limit: only the fastest-varying dimensions must obey the rule and exact behavior varies by library version, dtype, layout, and architecture.↩︎

  5. NVIDIA. (2023). Matrix Multiplication Background User’s Guide (last updated 2023-02-01). https://docs.nvidia.com/deeplearning/performance/dl-performance-matrix-multiplication/index.html. Support: Tensor Core use with cuBLAS 11.0 and cuDNN 7.6.3 and later without strict alignment, with best efficiency for FP16 at multiples of 8 elements and larger multiples on A100 paths. Limit: only the fastest-varying dimensions must obey the rule and exact behavior varies by library version, dtype, layout, and architecture.↩︎

  6. NVIDIA. (2024). CUDA C++ Programming Guide 12.4. https://docs.nvidia.com/cuda/archive/12.4.0/cuda-c-programming-guide/index.html (release 2024-03-05). Support: thread hierarchy, warps of 32 threads, memory hierarchy, streams, events, and coalesced access guidance used in Sections 2.2 and 2.3. Limit: exact SM counts, cache sizes, and bandwidths are architecture specific and are cited separately where stated.↩︎

  7. NVIDIA. (2016). Tesla P100 Datasheet. https://images.nvidia.com/content/tesla/pdf/nvidia-tesla-p100-datasheet.pdf. Support: SPECIFICATIONS table lists Tesla P100 SXM2 with 16 GB CoWoS HBM2 and 732 GB/s memory bandwidth. Limit: the PCIe 12 GB variant provides 549 GB/s, so the table row names the SXM2 16 GB variant explicitly.↩︎

  8. NVIDIA. (2018). Tesla V100 Datasheet. https://images.nvidia.com/content/technologies/volta/pdf/tesla-volta-v100-datasheet-letter-fnl-web.pdf. Support: SPECIFICATIONS table lists Tesla V100 with 16 GB or 32 GB HBM2 and 900 GB/s memory bandwidth. Limit: the 32 GB configuration doubles the standard 16 GB offering while bandwidth stays 900 GB/s, so the table row names the 32 GB variant explicitly.↩︎

  9. NVIDIA. (n.d.). NVIDIA A100 Tensor Core GPU specifications. https://www.nvidia.com/en-us/data-center/a100/. Support: specifications table lists A100 80 GB HBM2e with 1,935 GB/s (PCIe) and 2,039 GB/s (SXM) memory bandwidth. Limit: the table rounds the SXM value to 2.0 TB/s as a planning value and names the SXM variant explicitly.↩︎

  10. NVIDIA. (n.d.). NVIDIA H100 Tensor Core GPU product specifications. https://www.nvidia.com/en-us/data-center/h100/. Support: product specifications table lists H100 SXM with 80 GB memory and 3.35 TB/s memory bandwidth. Limit: PCIe and NVL variants differ in memory type and bandwidth, so the table row names the SXM5 80 GB HBM3 variant explicitly.↩︎

  11. NVIDIA. (n.d.). NVIDIA DGX B200 specifications. https://www.nvidia.com/en-us/data-center/dgx-b200/. Support: specifications list 8 Blackwell GPUs with 1,440 GB total GPU memory and 64 TB/s total HBM3e bandwidth, which divides to 180 GB and 8 TB/s per GPU. Limit: these per-GPU values are derived from the published eight-GPU system totals, not measured transfer rates.↩︎

  12. NVIDIA. (n.d.). NVIDIA H100 Tensor Core GPU product specifications. https://www.nvidia.com/en-us/data-center/h100/. Support: the H100 SXM column lists BFLOAT16 Tensor Core performance of 1,979 teraFLOPS, footnoted “With sparsity,” and 3.35 TB/s GPU memory bandwidth. Limit: the dense value of about 989 TFLOP/s is half the listed sparse figure, and both are theoretical peaks that depend on clocks and kernel path.↩︎