diff --git a/.wordlist.txt b/.wordlist.txt index 45844d8..9e4d0f8 100644 --- a/.wordlist.txt +++ b/.wordlist.txt @@ -885,3 +885,26 @@ logfile mnt ps verboseness +AMDSMI +AST +IOD +IODs +amd +amdsmi +celerator +chiplets +vllm +misconfigurations +integrators +hyperscaler +lifecycle +prem +ASPEED +LLC +SVM +TPX +addressability +datacenter +programmability +uptime +TPX diff --git a/docs/gpu-partitioning/images/mi300a_CPX.png b/docs/gpu-partitioning/images/mi300a_CPX.png new file mode 100644 index 0000000..60e684c Binary files /dev/null and b/docs/gpu-partitioning/images/mi300a_CPX.png differ diff --git a/docs/gpu-partitioning/images/mi300a_NPS1.png b/docs/gpu-partitioning/images/mi300a_NPS1.png new file mode 100644 index 0000000..139c957 Binary files /dev/null and b/docs/gpu-partitioning/images/mi300a_NPS1.png differ diff --git a/docs/gpu-partitioning/images/mi300a_SPX.png b/docs/gpu-partitioning/images/mi300a_SPX.png new file mode 100644 index 0000000..a59139a Binary files /dev/null and b/docs/gpu-partitioning/images/mi300a_SPX.png differ diff --git a/docs/gpu-partitioning/images/mi300a_TPX.png b/docs/gpu-partitioning/images/mi300a_TPX.png new file mode 100644 index 0000000..70061e8 Binary files /dev/null and b/docs/gpu-partitioning/images/mi300a_TPX.png differ diff --git a/docs/gpu-partitioning/images/mi300x_CPX.png b/docs/gpu-partitioning/images/mi300x_CPX.png new file mode 100644 index 0000000..c868ae6 Binary files /dev/null and b/docs/gpu-partitioning/images/mi300x_CPX.png differ diff --git a/docs/gpu-partitioning/images/mi300x_NPS1.png b/docs/gpu-partitioning/images/mi300x_NPS1.png new file mode 100644 index 0000000..6432353 Binary files /dev/null and b/docs/gpu-partitioning/images/mi300x_NPS1.png differ diff --git a/docs/gpu-partitioning/images/mi300x_NPS4.png b/docs/gpu-partitioning/images/mi300x_NPS4.png new file mode 100644 index 0000000..fa30fdc Binary files /dev/null and b/docs/gpu-partitioning/images/mi300x_NPS4.png differ diff --git a/docs/gpu-partitioning/images/mi300x_SPX.png b/docs/gpu-partitioning/images/mi300x_SPX.png new file mode 100644 index 0000000..5759580 Binary files /dev/null and b/docs/gpu-partitioning/images/mi300x_SPX.png differ diff --git a/docs/gpu-partitioning/index.rst b/docs/gpu-partitioning/index.rst new file mode 100644 index 0000000..8bd77d8 --- /dev/null +++ b/docs/gpu-partitioning/index.rst @@ -0,0 +1,64 @@ +.. meta:: + :description: Learn how to partition AMD GPUs/APUs. + :keywords: AMD, GPU, APU, partitioning, ROCm, MI300X, MI300A + +************************** +AMD GPU/APU Partitioning +************************** + +Partitioning Overview +^^^^^^^^^^^^^^^^^^^^^^ + +Modern large-scale AI and HPC workloads demand fine-grained control over GPU resource allocation, memory isolation, and multi-tenant scheduling. AMD's Instinct™ MI300 series accelerators — including the MI300X GPU and MI300A APU — support flexible partitioning schemes that allow users to logically subdivide a single device into multiple independent partitions optimized for different workloads. + +This documentation portal serves as a centralized index for navigating the complete GPU partitioning workflow on AMD platforms. It links to detailed technical guides for each supported accelerator, including: + +- **Architecture deep dives** to understand partitioning capabilities. +- **Quick start instructions** to apply compute and memory partition modes using `amd-smi`. +- **Guide to run vLLM workload** for inference benchmarking. +- **Troubleshooting resources** for resolving partitioning issues in production environments. + +Compatibility Matrix +^^^^^^^^^^^^^^^^^^^^^^ + +To streamline deployment planning and reduce configuration friction, we include below a **GPU Partitioning Schemes Compatibility Matrix**. This matrix outlines which combinations of **Compute Partitioning Modes** (e.g., SPX, CPX) and **Memory Partitioning Modes** (e.g., NPS1, NPS4) are validated for each supported device. It also notes any **minimum ROCm driver version requirements** necessary to enable specific configurations. + +.. important:: + **New to partitioning modes?** Before using the compatibility matrix, it's essential to understand the core concepts of **Compute Partitioning Modes** (SPX, CPX, TPX) and **Memory Partitioning Modes** (NPS1, NPS4). These modes determine how compute and memory resources are logically divided across a single device. + + See our detailed overview here: + - :ref:`MI300X Compute Partitioning ` / :ref:`MI300A Compute Partitioning ` + - :ref:`MI300X Memory Partitioning ` / :ref:`MI300A Memory Partitioning ` + +By consolidating this matrix on the index page, users can quickly evaluate platform capabilities and navigate to device-specific documentation with full awareness of what is supported on their hardware and software stack. + +.. list-table:: GPU Partitioning Schemes Compatibility Matrix + :header-rows: 1 + :widths: 20 20 20 20 20 + + * - Instinct GPUs + - SPX + NPS1 + - TPX + NPS1 + - CPX + NPS1 + - CPX + NPS4 + * - MI300X + - ✅ + - NA + - + - ✅ (ROCm 6.4) + * - MI300A + - ✅ + - ✅ (ROCm 6.3) + - ✅ (ROCm 6.4) + - NA + +.. note:: + The compatibility matrix is a living document and will be updated as new ROCm releases and device capabilities are validated. Users are encouraged to check back frequently for the latest information. + +Device Documentation +^^^^^^^^^^^^^^^^^^^^^ + +- :doc:`AMD Instinct MI300X GPU ` — Includes guidance for MI300X GPU-specific partitioning, architecture, System compatibility, and running vLLM inference. +- :doc:`AMD Instinct MI300A APU ` — Includes guidance for APU-specific partitioning, architecture, System compatibility, and running vLLM inference. + +We recommend users start with this index page to assess compatibility, then follow device-specific documentation to implement and validate GPU partitioning configurations in their own clusters or platforms. diff --git a/docs/gpu-partitioning/mi300a/index.rst b/docs/gpu-partitioning/mi300a/index.rst new file mode 100644 index 0000000..d61dadf --- /dev/null +++ b/docs/gpu-partitioning/mi300a/index.rst @@ -0,0 +1,28 @@ +.. meta:: + :description: AMD Instinct MI300A APU + :keywords: AMD, MI300A, APU, CPU-GPU, Instinct, Overview + +******************************************* +AMD Instinct MI300A APU +******************************************* + +The AMD Instinct™ MI300A APU (Accelerated Processing Unit) is a groundbreaking compute platform that fuses AMD EPYC™ CPU cores with CDNA™ 3 GPU architecture into a single, unified package. Purpose-built for data-intensive AI, HPC, and scientific computing workloads, the MI300A offers unprecedented levels of memory bandwidth, compute density, and energy efficiency — all through a cohesive CPU-GPU heterogeneous system. + +As the world’s first data center APU based on **advanced chiplet packaging**, MI300A breaks traditional boundaries by enabling shared memory between the CPU and GPU, reducing latency and eliminating redundant data transfers. This guide provides an in-depth reference for developers, system architects, and platform integrators working with MI300A systems — from configuration and partitioning to memory access models and workload deployment. + +Key technical highlights of the MI300A platform include: + +- **24 Zen 4 CPU cores** integrated with **228 CDNA 3 CUs**, all sharing a unified addressable HBM memory pool. +- Up to **128 GB of unified HBM3 memory**, accessible by both CPU and GPU without the need for explicit memory copies. +- **Advanced GPU partitioning** support (SPX, CPX, TPX) enabling workload isolation, fine-grained scheduling, and resource optimization. +- Hardware-accelerated **coherent shared memory**, enabling low-latency CPU-GPU communication for tightly coupled compute models. +- Full compatibility with the **ROCm 6.x software stack**, including HIP, OpenMP offload, and leading AI/ML libraries. + +This guide is structured to help users get the most out of MI300A across a wide range of applications: + +- :doc:`Overview ` — Deep dive into MI300A architecture, APU topology, and compute/memory partitioning models. +- :doc:`Requirements ` — Platform setup, BIOS/kernel configuration, ROCm compatibility matrix, and supported distros. +- :doc:`Quick Start Guide ` — Walkthrough for bringing up MI300A systems and configuring partitions using `amd-smi`. +- :doc:`Troubleshooting ` — Diagnosing common errors, partition conflicts, and optimizing workload placement. + +Whether you are deploying the MI300A in an exascale supercomputer or using it to accelerate simulation, AI, or analytics workloads, this documentation serves as your go-to reference for maximizing performance, interoperability, and development agility. diff --git a/docs/gpu-partitioning/mi300a/overview.rst b/docs/gpu-partitioning/mi300a/overview.rst new file mode 100644 index 0000000..018746a --- /dev/null +++ b/docs/gpu-partitioning/mi300a/overview.rst @@ -0,0 +1,242 @@ +AMD Instinct MI300A APU Overview +================================ + +1. Introduction +---------------- + +The AMD Instinct™ MI300A Accelerated Processing Unit (APU) represents a major architectural advancement in AMD’s Compute DNA (CDNA) portfolio. As AMD’s first fully integrated CPU+GPU APU designed for high-performance computing (HPC) and artificial intelligence (AI) workloads, the MI300A combines the power of Zen 4 CPU cores with CDNA3 GPU compute units and high-bandwidth HBM3 memory into a single unified package. This integration delivers significant benefits over traditional discrete CPU-GPU systems by eliminating performance bottlenecks & memory bandwidth bottlenecks, reducing data movement overhead, addressing programmability overhead, and the need to refactor code for new hardware generations. + +The MI300A introduces a fully coherent, unified memory architecture, enabling CPU and GPU components to share data and cache efficiently, simplifying software development and improving runtime performance. By tightly integrating high-performance “Zen4” CPU cores with high-throughput GPU Compute Units (CUs) and 128GB of unified HBM3 memory into a single socket with hardware-supported cache coherence, the MI300A APU eliminates the data movement penalties commonly associated with discrete architectures. Through architectural innovations like shared last-level cache (LLC), cache-coherent memory access, direct CPU-GPU fabric interconnects, seamless task delegation across CPU and GPU, synchronization across compute domain, and support for hardware sparsity, MI300A is optimized to accelerate the convergence of HPC and AI at exascale. + +These architectural innovations are supported by AMD’s ROCm™ open software platform, which provides a consistent programming model for heterogeneous compute. The result is a single-package solution that delivers exceptional energy efficiency, programmability, and performance density — enabling next-generation exascale systems to tackle converged HPC and AI workloads with reduced complexity and improved throughput. + +This guide provides a comprehensive overview to GPU partitioning on MI300A APU platforms, focusing on supported compute partition modes, NUMA configurations or memory access models, system configuration requirements, and usage guidance. Whether for virtualized multi-tenant environments or tightly coupled HPC workloads, the MI300A’s partitioning features empower developers and administrators to tailor resource allocation for optimal system utilization. This guide also includes validation and troubleshooting guidance to help users leverage MI300A’s full potential on bare-metal deployments. + +2. GPU Architecture Summary +--------------------------- + +The MI300A APU architecture is built using AMD’s chiplet-based design principles and state-of-the-art 3D stacking technology, bringing CPU and GPU compute into a unified, high-bandwidth package. This architecture enables tight coupling between CPU and GPU resources while maximizing memory bandwidth and minimizing data latency. + +Key architectural components include: + +- **Accelerator Complex Dies (XCDs):** + - 6 XCDs per socket + - Each XCD contains 38 CDNA3-based GPU Compute Units (CUs) + - Total of 228 GPU CUs per socket + +- **CPU Chiplets:** + - 3 chiplets with 8-core AMD “Zen4” CPU dies + - Total of 24 CPU cores per socket + - CPU and GPU share a unified memory address space and a large Last Level Cache + +- **HBM3 Memory:** + - 8 x 16GB HBM3 stacks per socket + - 128GB of total unified HBM capacity + - Host memory and I/O buffers are fully interleaved across all 8 HBM stacks + +- **Last Level Cache (LLC):** + - 256 MB of shared last-level cache (LLC) per socket by both CPU and GPU clients + - Sits beyond the coherence point; access does not require cache probing or flushes + +- **Infinity Fabric:** + - High-bandwidth, coherent interconnect fabric connecting all compute elements + - Enables 2-socket (2S) and 4-socket (4S) fully connected node configurations + - Supports remote memory access and GPU virtualization + +- **Sparsity Acceleration:** + - Hardware-level 2:4 sparsity support to accelerate AI workloads + - Efficient handling of sparse matrix operations to save compute cycles and memory + +- **Partitioning Modes:** + - **SPX (Single Partition):** All 6 XCDs are grouped into one partition + - **TPX (Triple Partition):** Two XCDs per partition, yielding 3 partitions per socket + - **CPX (Core Partitioned):** Each XCD is treated as a separate partition, yielding 6 partitions per socket + +- **NUMA Mode:** Memory partitioning modes that define how HBM is allocated and accessed by logical devices + - **NPS1 (NUMA Per Socket):** Data is uniformly interleaved across all HBM stacks + - Fixed at boot time and not dynamically configurable + +Unlike MI300X, MI300A does not support discrete DDR DIMM access, and all system memory is resident within the 128GB HBM. This ensures high memory bandwidth and simplifies data placement for unified CPU-GPU workloads. Both 550W air-cooled and 760W liquid-cooled configurations are supported, making MI300A suitable for diverse datacenter environments. + +This architectural design provides the foundation for software-defined partitioning and workload orchestration — enabling MI300A to deliver balanced compute, memory, and I/O performance for the next era of converged HPC and AI workloads. + + + +3. Partitioning Concepts +------------------------- + +.. _mi300a_compute-partitioning: + +a. Compute Partitioning (SPX, TPX, CPX) +^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ + +Compute partitioning on the MI300A APU enables fine-grained resource management by dividing GPU compute resources into multiple logical devices, allowing users to optimize for workload isolation, parallel execution, and resource efficiency. Unlike discrete GPU solutions, the MI300A architecture unifies CPU, GPU, and memory into a single coherent address space backed by high-bandwidth HBM3. This allows partitioned workloads to share memory more seamlessly while benefiting from high interconnect bandwidth and cache coherence. + +MI300A supports the following compute partitioning modes: + +- **SPX (Single Partition X-celerator):** All GPU XCDs are grouped as a single monolithic device. +- **TPX (Triple Partition X-celerator):** The GPU complex is divided into three partitions, each containing two XCDs. +- **CPX (Core Partitioned X-celerator):** Each of the six XCDs is treated as a separate logical device. + +All partitioning is managed at the driver level and can be reconfigured dynamically using utilities such as ``amd-smi``. These modes enable flexible workload orchestration strategies ranging from unified execution (SPX) to strict isolation (CPX). + +**Key Benefits of Compute Partitioning:** + +- Enables multi-tenancy and workload isolation. +- Provides better scheduling granularity and performance tuning. +- Optimizes resource utilization across mixed workloads. +- Reduces contention and minimizes performance variability. + +**Partitioning Rules and Notes:** + +- MI300A includes 6 GPU XCDs per socket; partitions must use physically grouped XCDs. +- Each partition is assigned an equal number of Compute Units (CUs) and interleaved access to the unified HBM pool. +- Partitioning is spatial (hardware-aware) and operates at the kernel driver level, independent of virtualization or container runtimes. +- In MI300A only **NPS1** (NUMA Per Socket) is available, where all HBM stacks are uniformly interleaved. + +.. list-table:: MI300A Partition Modes Comparison + :header-rows: 1 + + * - Mode + - Logical Devices + - CUs per Device + - Memory per Device + - Best For + * - **SPX** + - 1 + - 228 + - 128GB + - Unified workloads, large models + * - **TPX** + - 3 + - 76 + - 32GB + - Parallel, medium-size batch jobs + * - **CPX** + - 6 + - 38 + - 16GB + - Isolation, multi-user setups, fine-grained scheduling + +i. SPX (Single Partition X-celerator) +^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ + +- **Default Mode** for MI300A platforms. +- Combines all six XCDs into a single logical GPU device. +- All compute and memory resources (228 CUs, 128GB HBM) are exposed as a unified pool. +- Ideal for applications that require high memory bandwidth, unified addressability, and hardware-level synchronization. + +**Behavior:** + +- ``amd-smi`` reports a single GPU device. +- Workloads are automatically distributed across all six XCDs. +- Unified Last Level Cache (LLC) and memory fabric ensure efficient inter-chiplet communication. +- Optimal for single-user, large-batch, or monolithic workloads such as deep learning model training or HPC simulations. + +ii. TPX (Triple Partition X-celerator) +^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ + +- Divides the GPU into **three partitions**, each comprising **two XCDs**. +- Exposes three logical GPU devices per socket. +- Each TPX partition gets access to 76 CUs and approximately 32GB of interleaved HBM memory. + +**Use Case:** + +- Balanced resource sharing across multiple jobs. +- Good for parallel model execution where each model requires moderate compute and memory. +- Enables concurrent scheduling of independent medium-sized workloads without over-provisioning. + +**Behavior:** + +- ``amd-smi`` reports three GPU devices. +- Each device can be targeted independently via HIP, OpenMP, or other ROCm-compatible programming models. +- All partitions maintain full memory coherence and uniform memory access through NPS1. + +iii. CPX (Core Partitioned X-celerator) +^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ + +- Most granular mode of operation. +- Each XCD is exposed as a distinct logical GPU device (6 devices per socket). +- Each CPX partition includes 38 CUs and 16GB of interleaved HBM memory. +- Excellent for scenarios requiring workload isolation or running multiple lightweight jobs concurrently. + +**Use Case:** + +1. **Multi-User Environments:** Allocate each CPX partition to different users or tenants to enforce hardware-level isolation. +2. **Task Parallelism:** Run multiple inference or small-batch training jobs simultaneously. +3. **Fine-Tuned Scheduling:** Better visibility and control over how jobs are assigned to physical resources. + +**Behavior:** + +- ``amd-smi`` reports six GPU devices. +- Each device operates as an independent compute target with full access to the shared memory fabric. +- Peer-to-peer (P2P) access between CPX partitions is supported and can be enabled for collective operations. +- CPX mode is particularly powerful in shared infrastructure environments or cloud-native workloads. + +.. list-table:: + :header-rows: 1 + + * - MI300A SPX + - MI300A TPX + - MI300A CPX + * - .. image:: ../images/mi300a_SPX.png + - .. image:: ../images/mi300a_TPX.png + - .. image:: ../images/mi300a_CPX.png + * - **SPX:** All 6 XCDs form a single device. + - **TPX:** Three partitions with 2 XCDs each. + - **CPX:** Six partitions, one per XCD. + +- **Diagram Note:** Dotted lines in the diagrams indicate compute partition boundaries. + +.. _mi300a_memory-partitioning: + +b. Memory Partitioning (NPS1) +^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ + +The MI300A platform operates exclusively in **NPS1** mode — or **NUMA Per Socket** — where all eight HBM stacks are uniformly interleaved and exposed as a unified memory pool. Unlike MI300X, MI300A does not support DDR memory. + +**Key Features of NPS1 Mode:** + +- The entire 128GB of HBM is accessible across all partitions, regardless of compute mode (SPX, TPX, CPX). +- Memory is interleaved across the eight HBM stacks to ensure maximum bandwidth and minimal latency. +- No memory locality enforcement across partitions — partitions can transparently access the shared memory fabric. + +.. list-table:: MI300A Memory Mode Overview + :header-rows: 1 + + * - Memory Mode + - Description + - Compatible Compute Modes + * - **NPS1** + - Interleaved HBM3 pool (128GB) accessible by all partitions + - SPX, TPX, CPX + +- The tight integration of CPU and GPU with a unified cache-coherent memory fabric eliminates the complexity of NUMA-aware memory allocation typically required in multi-socket, discrete systems. + +.. image:: ../images/mi300a_NPS1.png + :alt: MI300A NPS1 Unified Memory Layout + +- **Diagram Note:** All GPU partitions in SPX, TPX, and CPX share the same physical memory pool via the NPS1 model. + +4. Benefits of Partitioning (MI300A APU) +---------------------------------------- + +Partitioning in the MI300A APU—enabled via SPX, TPX, and CPX modes—offers a flexible architecture that balances unified memory access with compute isolation. These modes allow system architects to tune performance, resource efficiency, and workload isolation on heterogeneous CPU+GPU platforms. + +- **CPX mode in MI300A** enables fine-grained control over the GPU compute fabric by exposing each XCD as a distinct logical GPU. When paired with memory locality-aware execution, CPX mode enhances *parallelism*, *isolation*, and *throughput* for multi-user or multi-tenant systems, particularly in **high-performance computing (HPC)** and **cloud-native inference** environments. + +- **TPX mode** serves as a *balanced hybrid mode*, exposing 4 partitions with 2 XCDs each. This is especially beneficial for mid-sized models or workloads that demand more compute capacity and memory than a single XCD can provide, while still requiring workload separation. TPX enables *optimal use of shared CPU and GPU resources*, and aligns well with the shared-memory design of the MI300A APU. + +- Partitioning at the compute level enables **dynamic workload management**, where different applications or user sessions can be mapped to different GPU partitions, without interference or scheduling conflicts. This is a key enabler for *simultaneous AI, HPC, and mixed-precision scientific workloads*. + +- **Memory-coherent interconnects** in MI300A ensure that each partition can maintain high-bandwidth, low-latency communication with the CPU and system memory. Even when GPUs are logically isolated via CPX, partitions retain access to the shared HBM and DDR memory pools through the CPU’s memory controller, simplifying software complexity for multi-GPU workloads. + +- Partitioning also plays a crucial role in **fault containment and serviceability**. In the event of a GPU partition failure, workloads in other partitions can continue unaffected, enhancing system uptime and reducing recovery overheads. + +- **Driver-level flexibility** allows runtime switching between SPX, TPX, and CPX modes (subject to reboot in some configurations), enabling operators to adapt the system to workload needs without hardware reconfiguration. + +- On MI300A, GPU partitioning also interacts closely with **HMM (Heterogeneous Memory Management)** and **Shared Virtual Memory (SVM)**, enabling user applications to seamlessly share pointers and memory structures across CPU and GPU partitions. This allows for a *unified programming model* that reduces developer complexity and increases code portability. + +.. note:: + + On MI300A, while **SPX remains the default mode**, CPX and TPX offer compelling benefits for *multi-process environments*, *scientific workflows*, and *latency-sensitive inference* pipelines. Administrators should carefully benchmark their workloads across modes to identify the optimal configuration. diff --git a/docs/gpu-partitioning/mi300x/index.rst b/docs/gpu-partitioning/mi300x/index.rst new file mode 100644 index 0000000..4a401e8 --- /dev/null +++ b/docs/gpu-partitioning/mi300x/index.rst @@ -0,0 +1,28 @@ +.. meta:: + :description: AMD Instinct MI300X GPU + :keywords: AMD, MI300X, GPU, Overview + +******************************************* +AMD Instinct MI300X GPU +******************************************* + +The AMD Instinct™ MI300X GPU represents a significant leap in data center GPU design, purpose-built for large-scale AI inference, high-throughput LLM workloads, and advanced HPC deployments. Featuring cutting-edge CDNA™ 3 architecture and industry-leading memory capacity, the MI300X is designed to meet the most demanding compute and memory bandwidth requirements of today’s generative AI era. + +This documentation provides a comprehensive guide for users, system integrators, and infrastructure teams working with the MI300X, covering the complete software and runtime configuration lifecycle — from initial bring-up and partitioning to troubleshooting and running real-world workload. + +Key technical highlights of the MI300X include: + +- Up to **192 GB of HBM3 memory** with ultra-high bandwidth to accelerate large model inference. +- Support for advanced **GPU partitioning modes**, enabling logical GPU segmentation (SPX, CPX) for multi-tenant deployments and workload isolation. +- Fine-grained **memory partitioning** (NPS1, NPS4) for optimizing memory locality and performance in dense compute clusters. +- Full-stack compatibility with **ROCm 6.x**, the open software platform for AMD GPUs, enabling tight integration with PyTorch, Hugging Face, and other AI/ML frameworks. + +The sections below provide targeted guidance for each step of working with the MI300X platform: + +- :doc:`Overview ` — Architectural deep dive and partitioning model explanation. +- :doc:`Requirements ` — Platform prerequisites, supported ROCm versions, and kernel/BIOS configurations. +- :doc:`Quick Start Guide ` — Step-by-step instructions to configure GPU and memory partitions using `amd-smi`. +- :doc:`Troubleshooting ` — Common error resolutions and best practices for debugging partitioning-related issues. +- :doc:`Run a VLLM workload ` — Instructions for deploying a high-throughput LLM inference pipeline on MI300X. + +Whether you're deploying MI300X at scale in a hyperscaler data center or integrating it into an on-prem AI cluster, this guide is your central reference for maximizing performance, stability, and resource efficiency. diff --git a/docs/gpu-partitioning/mi300x/overview.rst b/docs/gpu-partitioning/mi300x/overview.rst new file mode 100644 index 0000000..a111b44 --- /dev/null +++ b/docs/gpu-partitioning/mi300x/overview.rst @@ -0,0 +1,207 @@ +AMD Instinct MI300X GPU Partitioning Overview +=============================================== + +1. Introduction +---------------- + +The AMD Instinct™ MI300X GPU represents a significant step forward in the design of modular, scalable GPU compute platforms. With its innovative architecture and ROCm software ecosystem, MI300X supports dynamic compute and memory partitioning. This capability enables developers and system administrators to treat a single GPU as multiple logical devices, allowing for efficient workload management, resource isolation, and performance tuning. + +The MI300X GPU introduces advanced support for compute and memory partitioning, enabling high-performance computing (HPC), artificial intelligence (AI), and machine learning (ML) workloads to achieve fine-grained resource allocation and isolation. + +Partitioning exposes internal GPU hardware components, specifically Compute Complexes (XCDs) and memory stacks (HBM) as discrete logical devices. This allows users to optimize system utilization, achieve better scheduling control, and tailor compute and memory resources to workload-specific requirements. + +This guide provides a detailed overview of GPU partitioning modes, primarily CPX and NPS4 on bare-metal operating systems. It includes architectural background, partitioning use cases, configuration methods, and validation steps to ensure users can fully leverage MI300X's capabilities. + +2. GPU Architecture Summary +--------------------------- + +The MI300X GPU is composed of modular chiplets, each optimized for compute or I/O tasks to achieve scalability and high-throughput performance. +Key architectural components include: + +- **XCD (Accelerator Complex Die):** Compute element of the GPU, each XCD contains 38 Compute Units (CUs), responsible for executing parallel workloads. +- **IOD (I/O Die):** Manages interconnects, memory, and data routing across the chiplets. +- **3D Stacking**: Each pair of XCDs is 3D-stacked on a single IOD allowing for tight integration and low-latency interconnects. +- **HBM (High-Bandwidth Memory):** MI300X includes 8 stacks of HBM, offering 192GB of unified memory. +- **Total GPU Configuration**: + + - 8 XCDs per GPU → 304 total CUs + - 4 IODs per GPU + - 8 HBM (High Bandwidth Memory) stacks (2 per IOD) + - 192GB of unified HBM capacity + +This layout provides the foundation for partitioning, allowing resources to be split and exposed logically to the operating system and applications. + +3. Partitioning Concepts +------------------------ + +.. _mi300x_compute-partitioning: + +a. Compute Partitioning (SPX, CPX) +^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ + +Compute partitioning (also referred to as MCP – Modular Chiplet Platform) is the division of the GPU's compute and memory resources into smaller logical units, which can then be addressed as independent devices by applications. This is implemented in the driver layer and can be dynamically adjusted at runtime using command-line tools like ``amd-smi``. + +There are two key compute partitioning modes: + +- **SPX (Single Partition X-celerator):** Treats the entire GPU as a single device. +- **CPX (Core Partitioned X-celerator):** Exposes each XCD as an individual logical GPU. + +**Key Benefits of Partitioning:** + +- Supports workload isolation and parallelism. +- Enables better scheduling granularity and performance tuning. +- Facilitates resource sharing across users or processes in multi-tenant environments. + +**Partitioning Rules:** + +- Partitions must include an even number of XCDs (e.g., 2, 4, 6, 8). +- Partitioning is spatial: each partition is composed of physically grouped XCDs. +- Compute partitioning is configured via the driver and hardware level, no VM (virtualization) or hypervisor required. + +**Partition Modes Comparison** + ++--------+------------------+----------------+-------------------+-------------------------------+ +| Mode | Logical Devices | CUs per Device | Memory per Device | Best For | ++========+==================+================+===================+===============================+ +| SPX | 1 | 304 | 192GB | Unified workloads | ++--------+------------------+----------------+-------------------+-------------------------------+ +| CPX | 8 | 38 | 24GB | Isolation, fine-grained | +| | | | | scheduling, small batch sizes | ++--------+------------------+----------------+-------------------+-------------------------------+ + +i. SPX (Single Partition X-celerator) +^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ + +- **Default Mode** for MI300X. +- Treats all 8 XCDs as a single monolithic GPU. +- All memory and compute resources are visible as one unified device. +- Implicit synchronization across XCDs is handled by the hardware. + +**Use Case**: Ideal for large-scale models or applications that require unified compute and memory access without needing explicit control over scheduling. + +**Behavior:** + +- ``amd-smi`` shows **1 GPU** with **304 CUs** and **192GB HBM**. +- Workgroups are **automatically distributed** across all XCDs (round-robin). +- The GPU will always revert back to this default SPX mode when the system is rebooted or when the amdgpu driver is unloaded and reloaded. + +ii. CPX (Core Partitioned X-celerator) +^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ + +- Each XCD is represented as a **separate logical GPU**. +- Offers granular control—each partition gets 38 CUs and 24GB of HBM. +- **Memory Allocation:** + + - In NPS1, HBM memory is interleaved across all stacks. + - In NPS4, each XCD gets a dedicated memory quadrant (2 HBM stacks) which can be used to interleave the 24GB of dedicated memory each XCD is given in this mode. + +- CPX works optimally with memory partitioning (NPS4) + +**Use Case**: + +1. **Multi-Tenant Environments:** Allocate separate GPU partitions to different users or tenants in a data center to achieve isolation. +2. **Heterogeneous Workloads:** Run AI training, inference, and HPC workloads simultaneously on different partitions where individual models/data fit within a single XCD's memory. +3. **Resource Oversubscription:** Optimize resource usage by oversubscribing partitions for workloads with varying demands. + +**Behavior:** + +- ``amd-smi`` shows **8 GPUs**, each with **38 CUs** and **24GB HBM**. +- Workgroups are explicitly launched to a specific XCD (i.e., scheduling can be + controlled at the application level). +- Peer-to-Peer (P2P) access between XCDs is available and can be enabled. + +.. - MALL (Memory Attached Last Level Cache) is shared between two CPX partitions. + +.. list-table:: + :header-rows: 1 + + * - MI300X SPX + - MI300X CPX + * - .. image:: ../images/mi300x_SPX.png + - .. image:: ../images/mi300x_CPX.png + * - **SPX:** All XCDs appear as one logical device. + - **CPX:** Each XCD appears as one logical device. + +- **Diagram Note:** Dotted lines in the diagrams indicate compute partition boundaries. + +.. _mi300x_memory-partitioning: + +b. Memory Partitioning (NPS) +^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ + +- The memory partitioning modes (known as Non-Uniform Memory Access (NUMA) Per Socket (NPS)) change the number of NUMA domains that a device exposes, which define how HBM (High Bandwidth Memory) is allocated and exposed to logical devices. +- Memory partitioning in the MI300X series involves dividing the total memory, specifically HBM stacks, which are accessible to a compute unit, into partitions. +- This is configured as application memory for XCDs, allowing for more efficient memory management and allocation. +- The memory partitioning is done at the hardware level, and the driver manages the visibility of these partitions to the operating system and applications. +- In MI300X, the number of memory partitions must be less than or equal to the number of compute partitions. +- The MI300X supports two memory partitioning modes: + + - **NPS1 (Unified Memory):** + + - All 8 HBM stacks are viewed as one unified memory pool and is accessible to all XCDs. + - `amd-smi` will show 1 device with 192GB of HBM. + - Memory is allocated interleaved across all HBM stacks. + - Best for workloads requiring unified memory. + - Compatible mode with SPX and CPX. + + - **NPS4 (Partitioned Memory):** + + - Pairs of HBM stacks forming 48GB each are viewed as separate memory partitions. Each CPX partition still only has access to 24GB of HBM memory, but the memory is interleaved across this 48GB memory partition instead of across the entire 192GB of the GPU. + - Each memory quadrant (partition) of the memory is directly visible to the logical devices in its quadrant. + - An XCD can still access all portions of memory through multi-GPU programming techniques. + - Best for workloads requiring dedicated memory resources. + - Only available with CPX mode. + - In NPS4 mode, the traffic latency to HBM (High Bandwidth Memory) is minimized because it remains on the same AID (Accelerator Interface Domain), leading to shorter latency and faster transitions from idle to full bandwidth. + - In NPS4 mode, higher bandwidth to MALL (Memory Attached Last Level Cache) can be achieved. + - In most cases, NPS4 mode is highly performant when paired with CPX mode for workloads that fit within the memory capacity of a single XCD. + +.. list-table:: Memory Partitioning Modes + :header-rows: 1 + :widths: 20 50 30 + + * - Memory Mode + - Description + - Compute Mode Compatibility + * - **NPS1** + - Unified memory pool (192GB) + - SPX, CPX + * - **NPS4** + - 4 memory partitions (48GB each). Note- Each CPX only accesses 24GB from the partition. + - CPX only + +.. list-table:: + :header-rows: 1 + + * - MI300X NPS1 + - MI300X NPS4 + * - .. image:: ../images/mi300x_NPS1.png + - .. image:: ../images/mi300x_NPS4.png + * - **NPS1:** All HBM stacks appear as a unified memory pool. + - **NPS4:** HBM stacks are segmented into memory quadrants. + +- **Diagram Note:** Dotted lines in the diagrams indicate memory partition boundaries. + +4. Benefits of Partitioning +---------------------------- + +Partitioning support in the AMD Instinct™ MI300X GPU delivers significant operational and performance advantages in large-scale AI inference and HPC environments. By logically segmenting GPU and memory resources, users can achieve fine-grained workload control, reduce overhead, and boost cluster utilization. + +Key benefits of partitioning on MI300X include: + +- **Improved performance for small to mid-sized language models:** + Partitioning the MI300X into four logical CPX GPUs (via `SPP=CPX` and `NPS=4`) allows small models (≤13B parameters) to run independently within each GPU slice. This enables higher concurrency and throughput when serving multiple models simultaneously, especially in VLLM-based inference engines. + +- **Enhanced communication efficiency for distributed workloads:** + CPX + NPS4 mode aligns well with multi-GPU collective communication patterns, delivering improved bandwidth and reduced latency for all-to-all and all-reduce operations through optimized ROCm Communication Collectives Library (RCCL) backend. + +- **Power savings and thermal optimization:** + Memory partitioning with `NPS=4` reduces the power consumed by the HBM3 memory stacks per workload, enabling energy-efficient inference and better thermal headroom under dense workloads. + +- **Dynamic resource provisioning and flexibility:** + Partitioning enables **runtime configuration of compute and memory** without requiring a full system reboot. This supports agile scheduling, workload isolation, and high system uptime in shared or multi-tenant data center environments. + +- **Improved workload packing and density:** + Logical GPU slicing allows simultaneous deployment of multiple containerized inference services per MI300X GPU. This leads to higher resource utilization and better GPU consolidation ratios when compared to monolithic deployments. + +.. note:: + Mixed memory partitioning modes (e.g., combining NPS1 and NPS4) are **not recommended** for single-node configurations due to potential performance and synchronization issues across NUMA domains. diff --git a/docs/gpu-partitioning/mi300x/quick-start-guide.rst b/docs/gpu-partitioning/mi300x/quick-start-guide.rst new file mode 100644 index 0000000..3ead07f --- /dev/null +++ b/docs/gpu-partitioning/mi300x/quick-start-guide.rst @@ -0,0 +1,798 @@ +Quick Start Guide to Partitioning MI300X GPUs +============================================== + +This guide serves as a practical and technically detailed reference for configuring compute and memory partitioning on AMD Instinct™ MI300X GPUs using the `amd-smi` utility. Partitioning is a key feature that enables system administrators, developers, and data center operators to dynamically subdivide a single MI300X GPU into multiple logical devices—each with its own dedicated compute and memory resources. This empowers users to maximize resource utilization, improve workload isolation, and optimize performance for diverse AI, HPC, and multi-tenant cloud environments. + +Partitioning on MI300X involves two dimensions: + +- **Compute Partitioning**: Divides the GPU’s compute units (XCDs) into isolated logical devices. For example, CPX mode creates eight independent compute partitions per GPU. +- **Memory Partitioning**: Splits the high-bandwidth memory (HBM) into physically separated addressable regions. NPS4 mode, for instance, creates four memory partitions of 48GB each. + +When used together, compute and memory partitioning modes such as CPX/NPS4 enable fine-grained allocation of GPU resources to individual workloads, enhancing security and manageability while maintaining high throughput. + +This Quick Start Guide walks through the necessary steps to create, verify, modify, and delete GPU partitions using standard system-level tools. Each step is accompanied by real command-line examples, expected outputs, and actionable notes to help ensure successful execution. + +Whether you are deploying MI300X GPUs for large-scale language model inference, cloud-native AI services, or isolated multi-user compute environments, understanding and leveraging GPU partitioning is essential for unlocking the full flexibility and efficiency of the Instinct platform. + + +1. Creating CPX/NPS4 Partition +------------------------------- + + - This section describes how to create a CPX/NPS4 partition on MI300X GPUs using the `amd-smi` tool. + - The partitioning process involves setting compute and memory partitioning modes to CPX and NPS4, respectively. + - The example below demonstrates how to set up a CPX/NPS4 partition on all GPUs in the system. + +To create a CPX/NPS4 partition: + +a. **Set compute partitioning mode to CPX:** + + + .. tab-set:: + + .. tab-item:: Compute Partition Command + + .. code-block:: shell-session + + # Set compute partition mode + sudo amd-smi set --gpu all --compute-partition CPX + + .. tab-item:: Shell output + + :: + + GPU: 0 + ACCELERATOR_PARTITION: Successfully set accelerator partition to CPX (profile #3) + + GPU: 1 + ACCELERATOR_PARTITION: Successfully set accelerator partition to CPX (profile #3) + + GPU: 2 + ACCELERATOR_PARTITION: Successfully set accelerator partition to CPX (profile #3) + + GPU: 3 + ACCELERATOR_PARTITION: Successfully set accelerator partition to CPX (profile #3) + + GPU: 4 + ACCELERATOR_PARTITION: Successfully set accelerator partition to CPX (profile #3) + + GPU: 5 + ACCELERATOR_PARTITION: Successfully set accelerator partition to CPX (profile #3) + + GPU: 6 + ACCELERATOR_PARTITION: Successfully set accelerator partition to CPX (profile #3) + + GPU: 7 + ACCELERATOR_PARTITION: Successfully set accelerator partition to CPX (profile #3) + +.. tip:: + + On MI300X, memory partitioning requires that the number of memory partitions not exceed the number of compute partitions. As a result, the **SPX+NPS4** configuration is invalid. + + When switching from **SPX+NPS1** to **CPX+NPS4**, setting the memory partition to **NPS4** will automatically transition the compute partition to **CPX**. You can skip the explicit CPX compute mode step, setting **NPS4** alone is sufficient to configure a valid **CPX+NPS4** partition. + + .. code-block:: shell-session + + sudo amd-smi set --memory-partition NPS4 + + This command will internally transition the compute mode from SPX to CPX, followed by the memory mode switch to NPS4 — provided all prerequisites are met and no GPU workloads are active. + + +b. **Set memory partitioning mode to NPS4:** + + .. tab-set:: + + .. tab-item:: Memory Partition Command + + .. code-block:: shell-session + + # Set memory partition mode + sudo amd-smi set --memory-partition NPS4 + + .. tab-item:: Shell output + + :: + + ****** WARNING ****** + + Setting Dynamic Memory (NPS) partition modes require users to quit all GPU workloads. + AMD SMI will then attempt to change memory (NPS) partition mode. + Upon a successful set, AMD SMI will then initiate an action to restart AMD GPU driver. + This action will change all GPU's in the hive to the requested memory (NPS) partition mode. + + Please use this utility with caution. + + Do you accept these terms? [Y/N] Y + + Trying again - Updating memory partition for gpu 0: [██████████████..........................] 50/140 secs remain + + GPU: 0 + MEMORY_PARTITION: Successfully set memory partition to NPS4 + + GPU: 1 + MEMORY_PARTITION: Successfully set memory partition to NPS4 + + GPU: 2 + MEMORY_PARTITION: Successfully set memory partition to NPS4 + + GPU: 3 + MEMORY_PARTITION: Successfully set memory partition to NPS4 + + GPU: 4 + MEMORY_PARTITION: Successfully set memory partition to NPS4 + + GPU: 5 + MEMORY_PARTITION: Successfully set memory partition to NPS4 + + GPU: 6 + MEMORY_PARTITION: Successfully set memory partition to NPS4 + + GPU: 7 + MEMORY_PARTITION: Successfully set memory partition to NPS4 + + GPU: 8 + MEMORY_PARTITION: Successfully set memory partition to NPS4 + + GPU: 9 + MEMORY_PARTITION: Successfully set memory partition to NPS4 + + GPU: 10 + MEMORY_PARTITION: Successfully set memory partition to NPS4 + + GPU: 11 + MEMORY_PARTITION: Successfully set memory partition to NPS4 + + GPU: 12 + MEMORY_PARTITION: Successfully set memory partition to NPS4 + + GPU: 13 + MEMORY_PARTITION: Successfully set memory partition to NPS4 + + GPU: 14 + MEMORY_PARTITION: Successfully set memory partition to NPS4 + + OSError: [Errno 24] Too many open files + +.. note:: + The above `amd-smi` command to set the partition mode may not show memory partition status for all GPUs. This is a known tool issue. + Despite the error, the partition mode will be set correctly across all GPUs. + +- The command will set the following: + + - **Compute Partitioning:** CPX mode (8 XCDs → 8 logical GPUs) + - **Memory Partitioning:** NPS4 mode (4 memory partitions with 2 HBM stacks each) + + +2. Verifying Partition Creation +---------------------------------- + + - After setting the partitioning modes, you can verify the partition creation using the `amd-smi` tool. + - The command will display the current partitioning status of the GPUs, including compute and memory partitioning modes. + +To confirm active partitioning state: + +Use `amd-smi` to confirm active partition states: + + .. tab-set:: + + .. tab-item:: Command + + .. code-block:: shell-session + + # Check partitioning status + amd-smi static --partition + + .. tab-item:: Shell output + + :: + + GPU: 0 + PARTITION: + COMPUTE_PARTITION: CPX + MEMORY_PARTITION: NPS4 + PARTITION_ID: 0 + + GPU: 1 + PARTITION: + COMPUTE_PARTITION: CPX + MEMORY_PARTITION: NPS4 + PARTITION_ID: 1 + + GPU: 2 + PARTITION: + COMPUTE_PARTITION: CPX + MEMORY_PARTITION: NPS4 + PARTITION_ID: 2 + + GPU: 3 + PARTITION: + COMPUTE_PARTITION: CPX + MEMORY_PARTITION: NPS4 + PARTITION_ID: 3 + + GPU: 4 + PARTITION: + COMPUTE_PARTITION: CPX + MEMORY_PARTITION: NPS4 + PARTITION_ID: 4 + + GPU: 5 + PARTITION: + COMPUTE_PARTITION: CPX + MEMORY_PARTITION: NPS4 + PARTITION_ID: 5 + + GPU: 6 + PARTITION: + COMPUTE_PARTITION: CPX + MEMORY_PARTITION: NPS4 + PARTITION_ID: 6 + + GPU: 7 + PARTITION: + COMPUTE_PARTITION: CPX + MEMORY_PARTITION: NPS4 + PARTITION_ID: 7 + + GPU: 8 + PARTITION: + COMPUTE_PARTITION: CPX + MEMORY_PARTITION: NPS4 + PARTITION_ID: 0 + + GPU: 9 + PARTITION: + COMPUTE_PARTITION: CPX + MEMORY_PARTITION: NPS4 + PARTITION_ID: 1 + + GPU: 10 + PARTITION: + COMPUTE_PARTITION: CPX + MEMORY_PARTITION: NPS4 + PARTITION_ID: 2 + + GPU: 11 + PARTITION: + COMPUTE_PARTITION: CPX + MEMORY_PARTITION: NPS4 + PARTITION_ID: 3 + + GPU: 12 + PARTITION: + COMPUTE_PARTITION: CPX + MEMORY_PARTITION: NPS4 + PARTITION_ID: 4 + + GPU: 13 + PARTITION: + COMPUTE_PARTITION: CPX + MEMORY_PARTITION: NPS4 + PARTITION_ID: 5 + + GPU: 14 + PARTITION: + COMPUTE_PARTITION: CPX + MEMORY_PARTITION: NPS4 + PARTITION_ID: 6 + + GPU: 15 + PARTITION: + COMPUTE_PARTITION: CPX + MEMORY_PARTITION: NPS4 + PARTITION_ID: 7 + + GPU: 16 + PARTITION: + COMPUTE_PARTITION: CPX + MEMORY_PARTITION: NPS4 + PARTITION_ID: 0 + + GPU: 17 + PARTITION: + COMPUTE_PARTITION: CPX + MEMORY_PARTITION: NPS4 + PARTITION_ID: 1 + + GPU: 18 + PARTITION: + COMPUTE_PARTITION: CPX + MEMORY_PARTITION: NPS4 + PARTITION_ID: 2 + + GPU: 19 + PARTITION: + COMPUTE_PARTITION: CPX + MEMORY_PARTITION: NPS4 + PARTITION_ID: 3 + + GPU: 20 + PARTITION: + COMPUTE_PARTITION: CPX + MEMORY_PARTITION: NPS4 + PARTITION_ID: 4 + + GPU: 21 + PARTITION: + COMPUTE_PARTITION: CPX + MEMORY_PARTITION: NPS4 + PARTITION_ID: 5 + + GPU: 22 + PARTITION: + COMPUTE_PARTITION: CPX + MEMORY_PARTITION: NPS4 + PARTITION_ID: 6 + + GPU: 23 + PARTITION: + COMPUTE_PARTITION: CPX + MEMORY_PARTITION: NPS4 + PARTITION_ID: 7 + + GPU: 24 + PARTITION: + COMPUTE_PARTITION: CPX + MEMORY_PARTITION: NPS4 + PARTITION_ID: 0 + + GPU: 25 + PARTITION: + COMPUTE_PARTITION: CPX + MEMORY_PARTITION: NPS4 + PARTITION_ID: 1 + + GPU: 26 + PARTITION: + COMPUTE_PARTITION: CPX + MEMORY_PARTITION: NPS4 + PARTITION_ID: 2 + + GPU: 27 + PARTITION: + COMPUTE_PARTITION: CPX + MEMORY_PARTITION: NPS4 + PARTITION_ID: 3 + + GPU: 28 + PARTITION: + COMPUTE_PARTITION: CPX + MEMORY_PARTITION: NPS4 + PARTITION_ID: 4 + + GPU: 29 + PARTITION: + COMPUTE_PARTITION: CPX + MEMORY_PARTITION: NPS4 + PARTITION_ID: 5 + + GPU: 30 + PARTITION: + COMPUTE_PARTITION: CPX + MEMORY_PARTITION: NPS4 + PARTITION_ID: 6 + + GPU: 31 + PARTITION: + COMPUTE_PARTITION: CPX + MEMORY_PARTITION: NPS4 + PARTITION_ID: 7 + + GPU: 32 + PARTITION: + COMPUTE_PARTITION: CPX + MEMORY_PARTITION: NPS4 + PARTITION_ID: 0 + + GPU: 33 + PARTITION: + COMPUTE_PARTITION: CPX + MEMORY_PARTITION: NPS4 + PARTITION_ID: 1 + + GPU: 34 + PARTITION: + COMPUTE_PARTITION: CPX + MEMORY_PARTITION: NPS4 + PARTITION_ID: 2 + + GPU: 35 + PARTITION: + COMPUTE_PARTITION: CPX + MEMORY_PARTITION: NPS4 + PARTITION_ID: 3 + + GPU: 36 + PARTITION: + COMPUTE_PARTITION: CPX + MEMORY_PARTITION: NPS4 + PARTITION_ID: 4 + + GPU: 37 + PARTITION: + COMPUTE_PARTITION: CPX + MEMORY_PARTITION: NPS4 + PARTITION_ID: 5 + + GPU: 38 + PARTITION: + COMPUTE_PARTITION: CPX + MEMORY_PARTITION: NPS4 + PARTITION_ID: 6 + + GPU: 39 + PARTITION: + COMPUTE_PARTITION: CPX + MEMORY_PARTITION: NPS4 + PARTITION_ID: 7 + + GPU: 40 + PARTITION: + COMPUTE_PARTITION: CPX + MEMORY_PARTITION: NPS4 + PARTITION_ID: 0 + + GPU: 41 + PARTITION: + COMPUTE_PARTITION: CPX + MEMORY_PARTITION: NPS4 + PARTITION_ID: 1 + + GPU: 42 + PARTITION: + COMPUTE_PARTITION: CPX + MEMORY_PARTITION: NPS4 + PARTITION_ID: 2 + + GPU: 43 + PARTITION: + COMPUTE_PARTITION: CPX + MEMORY_PARTITION: NPS4 + PARTITION_ID: 3 + + GPU: 44 + PARTITION: + COMPUTE_PARTITION: CPX + MEMORY_PARTITION: NPS4 + PARTITION_ID: 4 + + GPU: 45 + PARTITION: + COMPUTE_PARTITION: CPX + MEMORY_PARTITION: NPS4 + PARTITION_ID: 5 + + GPU: 46 + PARTITION: + COMPUTE_PARTITION: CPX + MEMORY_PARTITION: NPS4 + PARTITION_ID: 6 + + GPU: 47 + PARTITION: + COMPUTE_PARTITION: CPX + MEMORY_PARTITION: NPS4 + PARTITION_ID: 7 + + GPU: 48 + PARTITION: + COMPUTE_PARTITION: CPX + MEMORY_PARTITION: NPS4 + PARTITION_ID: 0 + + GPU: 49 + PARTITION: + COMPUTE_PARTITION: CPX + MEMORY_PARTITION: NPS4 + PARTITION_ID: 1 + + GPU: 50 + PARTITION: + COMPUTE_PARTITION: CPX + MEMORY_PARTITION: NPS4 + PARTITION_ID: 2 + + GPU: 51 + PARTITION: + COMPUTE_PARTITION: CPX + MEMORY_PARTITION: NPS4 + PARTITION_ID: 3 + + GPU: 52 + PARTITION: + COMPUTE_PARTITION: CPX + MEMORY_PARTITION: NPS4 + PARTITION_ID: 4 + + GPU: 53 + PARTITION: + COMPUTE_PARTITION: CPX + MEMORY_PARTITION: NPS4 + PARTITION_ID: 5 + + GPU: 54 + PARTITION: + COMPUTE_PARTITION: CPX + MEMORY_PARTITION: NPS4 + PARTITION_ID: 6 + + GPU: 55 + PARTITION: + COMPUTE_PARTITION: CPX + MEMORY_PARTITION: NPS4 + PARTITION_ID: 7 + + GPU: 56 + PARTITION: + COMPUTE_PARTITION: CPX + MEMORY_PARTITION: NPS4 + PARTITION_ID: 0 + + GPU: 57 + PARTITION: + COMPUTE_PARTITION: CPX + MEMORY_PARTITION: NPS4 + PARTITION_ID: 1 + + GPU: 58 + PARTITION: + COMPUTE_PARTITION: CPX + MEMORY_PARTITION: NPS4 + PARTITION_ID: 2 + + GPU: 59 + PARTITION: + COMPUTE_PARTITION: CPX + MEMORY_PARTITION: NPS4 + PARTITION_ID: 3 + + GPU: 60 + PARTITION: + COMPUTE_PARTITION: CPX + MEMORY_PARTITION: NPS4 + PARTITION_ID: 4 + + GPU: 61 + PARTITION: + COMPUTE_PARTITION: CPX + MEMORY_PARTITION: NPS4 + PARTITION_ID: 5 + + GPU: 62 + PARTITION: + COMPUTE_PARTITION: CPX + MEMORY_PARTITION: NPS4 + PARTITION_ID: 6 + + GPU: 63 + PARTITION: + COMPUTE_PARTITION: CPX + MEMORY_PARTITION: NPS4 + PARTITION_ID: 7 + +3. Modifying Partitions +------------------------ + + - This section describes how to modify the partitioning modes of MI300X GPUs using the `amd-smi` tool. + - You can switch between compute and memory partitioning modes as needed. + - The example below demonstrates how to switch between compute and memory partitioning modes. + +Use the following commands to switch compute or memory partitioning modes. + +**Compute Partition Examples:** + + .. tab-set:: + + .. tab-item:: Compute Partition Command + + .. code-block:: shell-session + + # Set compute partition mode + sudo amd-smi set --gpu all --compute-partition CPX + + .. tab-item:: Shell output + + :: + + GPU: 0 + ACCELERATOR_PARTITION: Successfully set accelerator partition to CPX (profile #3) + + GPU: 1 + ACCELERATOR_PARTITION: Successfully set accelerator partition to CPX (profile #3) + + GPU: 2 + ACCELERATOR_PARTITION: Successfully set accelerator partition to CPX (profile #3) + + GPU: 3 + ACCELERATOR_PARTITION: Successfully set accelerator partition to CPX (profile #3) + + GPU: 4 + ACCELERATOR_PARTITION: Successfully set accelerator partition to CPX (profile #3) + + GPU: 5 + ACCELERATOR_PARTITION: Successfully set accelerator partition to CPX (profile #3) + + GPU: 6 + ACCELERATOR_PARTITION: Successfully set accelerator partition to CPX (profile #3) + + GPU: 7 + ACCELERATOR_PARTITION: Successfully set accelerator partition to CPX (profile #3) + + .. tab-set:: + + .. tab-item:: Compute Partition Command + + .. code-block:: shell-session + + # Set compute partition mode + sudo amd-smi set --gpu all --compute-partition SPX + + .. tab-item:: Shell output + + :: + + GPU: 0 + ACCELERATOR_PARTITION: Successfully set accelerator partition to SPX (profile #0) + + GPU: 1 + ACCELERATOR_PARTITION: Successfully set accelerator partition to SPX (profile #0) + + GPU: 2 + ACCELERATOR_PARTITION: Successfully set accelerator partition to SPX (profile #0) + + GPU: 3 + ACCELERATOR_PARTITION: Successfully set accelerator partition to SPX (profile #0) + + GPU: 4 + ACCELERATOR_PARTITION: Successfully set accelerator partition to SPX (profile #0) + + GPU: 5 + ACCELERATOR_PARTITION: Successfully set accelerator partition to SPX (profile #0) + + GPU: 6 + ACCELERATOR_PARTITION: Successfully set accelerator partition to SPX (profile #0) + + GPU: 7 + ACCELERATOR_PARTITION: Successfully set accelerator partition to SPX (profile #0) + + .. tab-set:: + + .. tab-item:: Memory Partition Command + + .. code-block:: shell-session + + # Set memory partition mode + sudo amd-smi set --memory-partition NPS4 + + .. tab-item:: Shell output + + :: + + ****** WARNING ****** + + Setting Dynamic Memory (NPS) partition modes require users to quit all GPU workloads. + AMD SMI will then attempt to change memory (NPS) partition mode. + Upon a successful set, AMD SMI will then initiate an action to restart AMD GPU driver. + This action will change all GPU's in the hive to the requested memory (NPS) partition mode. + + Please use this utility with caution. + + Do you accept these terms? [Y/N] Y + + Trying again - Updating memory partition for gpu 0: [██████████████..........................] 50/140 secs remain + + GPU: 0 + MEMORY_PARTITION: Successfully set memory partition to NPS4 + + GPU: 1 + MEMORY_PARTITION: Successfully set memory partition to NPS4 + + GPU: 2 + MEMORY_PARTITION: Successfully set memory partition to NPS4 + + GPU: 3 + MEMORY_PARTITION: Successfully set memory partition to NPS4 + + GPU: 4 + MEMORY_PARTITION: Successfully set memory partition to NPS4 + + GPU: 5 + MEMORY_PARTITION: Successfully set memory partition to NPS4 + + GPU: 6 + MEMORY_PARTITION: Successfully set memory partition to NPS4 + + GPU: 7 + MEMORY_PARTITION: Successfully set memory partition to NPS4 + + GPU: 8 + MEMORY_PARTITION: Successfully set memory partition to NPS4 + + GPU: 9 + MEMORY_PARTITION: Successfully set memory partition to NPS4 + + GPU: 10 + MEMORY_PARTITION: Successfully set memory partition to NPS4 + + GPU: 11 + MEMORY_PARTITION: Successfully set memory partition to NPS4 + + GPU: 12 + MEMORY_PARTITION: Successfully set memory partition to NPS4 + + GPU: 13 + MEMORY_PARTITION: Successfully set memory partition to NPS4 + + GPU: 14 + MEMORY_PARTITION: Successfully set memory partition to NPS4 + + OSError: [Errno 24] Too many open files + + .. tab-set:: + + .. tab-item:: Memory Partition Command + + .. code-block:: shell-session + + # Set memory partition mode + sudo amd-smi set --memory-partition NPS1 + + .. tab-item:: Shell output + + :: + + ****** WARNING ****** + + Setting Dynamic Memory (NPS) partition modes require users to quit all GPU workloads. + AMD SMI will then attempt to change memory (NPS) partition mode. + Upon a successful set, AMD SMI will then initiate an action to restart AMD GPU driver. + This action will change all GPU's in the hive to the requested memory (NPS) partition mode. + + Please use this utility with caution. + + Do you accept these terms? [Y/N] Y + + Trying again - Updating memory partition for gpu 0: [██████████████..........................] 50/140 secs remain + + + GPU: 0 + MEMORY_PARTITION: Successfully set memory partition to NPS1 + + GPU: 1 + MEMORY_PARTITION: Successfully set memory partition to NPS1 + + GPU: 2 + MEMORY_PARTITION: Successfully set memory partition to NPS1 + + GPU: 3 + MEMORY_PARTITION: Successfully set memory partition to NPS1 + + GPU: 4 + MEMORY_PARTITION: Successfully set memory partition to NPS1 + + GPU: 5 + MEMORY_PARTITION: Successfully set memory partition to NPS1 + + GPU: 6 + MEMORY_PARTITION: Successfully set memory partition to NPS1 + + GPU: 7 + MEMORY_PARTITION: Successfully set memory partition to NPS1 + +.. note: + NPS4 is only compatible with CPX mode. Attempting to set NPS4 with SPX will result in a failure. + +4. Deleting Partitions +----------------------- + + - This section describes how to delete or reset the partitioning modes of MI300X GPUs using the `amd-smi` tool. + - You can revert the partitioning modes to their default settings. + - The example below demonstrates how to delete or reset the partitioning modes. + +To delete or reset partitions, revert both compute and memory partitioning to defaults: + +.. code-block:: shell-session + + sudo amd-smi set --gpu all --compute-partition SPX + sudo amd-smi set --memory-partition NPS1 + + diff --git a/docs/gpu-partitioning/mi300x/requirements.rst b/docs/gpu-partitioning/mi300x/requirements.rst new file mode 100644 index 0000000..80a76d6 --- /dev/null +++ b/docs/gpu-partitioning/mi300x/requirements.rst @@ -0,0 +1,162 @@ +Requirements to Partition MI300X GPUs +====================================== + +Partitioning AMD Instinct™ MI300X GPUs is a critical enabler for modern heterogeneous computing environments where isolation, resource sharing, and workload-specific optimization are paramount. By dividing a single physical GPU into multiple logical partitions, developers and system administrators can tailor computational resources to meet the unique performance, memory, and security demands of diverse applications—including large-scale AI inference, training, HPC simulations, and cloud-native deployments. + +This document provides a comprehensive overview of the system, software, and firmware requirements needed to successfully configure and operate GPU partitioning on MI300X devices. Partitioning support for the MI300X platform is tightly integrated with the ROCm software stack and relies on both hardware-level and OS-level infrastructure. As such, careful attention must be given to platform readiness, including validated driver versions, kernel support, supported memory modes, and compatibility with partitioning utilities such as `amd-smi`. + +Users should ensure their system environment meets all listed prerequisites prior to attempting partition configuration. Failure to do so may result in incomplete GPU enumeration, missing partitioning capabilities, or instability during execution. + +This guide is intended for system integrators, developers, platform architects, and IT administrators tasked with deploying MI300X-based platforms in bare-metal, production-grade environments. All configurations, tools, and commands referenced herein have been validated on supported operating systems and are based on ROCm version 6.4 or newer. + +1. Prerequisites +----------------- + +- MI300X GPUs must be installed and recognized by the system. +- ROCm stack must be correctly installed. +- Firmware and kernel must support partitioning (latest recommended). +- `amd-smi` tool is required for runtime management. +- Bare-metal OS installation—no virtualization layer. + +2. System Requirements +---------------------- + +To ensure a successful partitioning experience with MI300X GPUs, confirm the following system requirements: + +a. Hardware Requirements +~~~~~~~~~~~~~~~~~~~~~~~~ + +- **GPU**: AMD MI300X + +b. Software Requirements +~~~~~~~~~~~~~~~~~~~~~~~~~ + +- **Linux Kernel**: Version 5.15 or newer + + #. to find the kernel version, run the following command + + .. tab-set:: + + .. tab-item:: Command + + .. code-block:: shell-session + + # Check Linux kernel version + hostnamectl | grep 'Kernel' + + .. tab-item:: Shell output + + :: + + Kernel: Linux 5.15.0-134-generic + +- **AMDSMI Tool Library**: Version 25.3.0 or newer + + #. to find the amdsmi version, run the following command + + .. tab-set:: + + .. tab-item:: Command + + .. code-block:: shell-session + + # Check AMD-SMI version + amd-smi version | grep -o 'AMDSMI [^|]*' + + .. tab-item:: Shell output + + :: + + AMDSMI Tool: 25.2.0+f4ad5ee + AMDSMI Library version: 25.3.0 + +- **AMD GPU Driver**: amdgpu-build 2120656 (>= 6.12.12) + + #. to find the amdgpu version, run the following command + + .. tab-set:: + + .. tab-item:: Command + + .. code-block:: shell-session + + # Check amd gpu version + amd-smi version | grep -o 'amdgpu version: [^|]*' + + .. tab-item:: Shell output + + :: + + amdgpu version: 6.12.12 + + +c. Firmware Requirements +~~~~~~~~~~~~~~~~~~~~~~~~~ + +- **VBIOS Version**: 022.040.003.043.000001 + + #. to find the VBIOS version, run the following command + + .. tab-set:: + + .. tab-item:: Command + + .. code-block:: shell-session + + # Check VBIOS version + amd-smi static | grep -A 4 -m 1 'VBIOS' + + .. tab-item:: Shell output + + :: + + VBIOS: + NAME: AMD MI300X_HW_SRIOV_CVS_1VF + BUILD_DATE: 2024/09/25 10:52 + PART_NUMBER: 113-M3000100-102 + VERSION: 022.040.003.042.000001 + +d. Operating System Requirements +~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ + +- Ubuntu 22.04+, 24.04+ +- Oracle Linux Server 8.8+ + + #. to check the operating system version, run the following command + + .. tab-set:: + + .. tab-item:: Command + + .. code-block:: shell-session + + # Check Operating System version + hostnamectl | grep 'Operating System' + + .. tab-item:: Shell output + + :: + + Operating System: Ubuntu 22.04.5 LTS + +e. Driver Requirements +~~~~~~~~~~~~~~~~~~~~~~~ + +- **ROCm**: Version 6.4 or newer + + #. to find the ROCm version, run the following command + + .. tab-set:: + + .. tab-item:: Command + + .. code-block:: shell-session + + # Check ROCm version + amd-smi version | grep -o 'ROCm version: [^|]*' + + .. tab-item:: Shell output + + :: + + ROCm version: 6.4.0 diff --git a/docs/gpu-partitioning/mi300x/run-vllm.rst b/docs/gpu-partitioning/mi300x/run-vllm.rst new file mode 100644 index 0000000..2b24d6b --- /dev/null +++ b/docs/gpu-partitioning/mi300x/run-vllm.rst @@ -0,0 +1,178 @@ +Steps to Run a vLLM Workload +============================= + +Introduction +------------ + +This document provides a step-by-step guide to executing large language model (LLM) inference workloads using the vLLM engine on AMD Instinct™ MI300X GPUs. vLLM is a high-throughput and memory-efficient inference engine optimized for transformer-based models, capable of achieving near-zero latency and maximum throughput through continuous batching and memory planning. + +These instructions assume that your MI300X GPUs have already been correctly partitioned into CPX/NPS4 modes and are accessible to the container runtime via the ROCm software stack. The guide walks through container setup, dependency configuration, and model execution using a real-world benchmark script provided by AMD. Each section provides detailed command examples, expected behavior, and practical configuration tips to help you validate performance, resource allocation, and model compatibility with the underlying hardware. + +This document is particularly useful for AI developers, performance engineers, and infrastructure administrators seeking to validate multi-GPU inference performance, confirm partitioning behavior, and optimize GPU resource allocation for LLM workloads in ROCm-based environments. + +--- + +1. One-Time Setup +------------------ + +Pull the prebuilt AMD container for vLLM workloads. This container includes all necessary ROCm libraries and pre-installed dependencies. + +.. code-block:: bash + + # one time setup + sudo docker pull rocm/vllm:instinct_main + +--- + +2. Launching the Container +--------------------------- + +Run the container with the required privileges and device mappings to enable GPU access: + +.. code-block:: bash + + # run the docker image and open a bash terminal + sudo docker run -it --network=host --group-add=video --ipc=host --cap-add=SYS_PTRACE --security-opt seccomp=unconfined --device /dev/kfd --device /dev/dri rocm/vllm:instinct_main /bin/bash + +--- + +3. Cloning the Benchmark Suite +------------------------------- + +Inside the container, clone the ROCm/MAD repository and navigate to the vLLM benchmarking script directory: + +.. code-block:: bash + + # clone the AMD MAD repo to run vllm workload from inside the docker image + git clone https://github.com/ROCm/MAD.git + cd MAD + # install any dependencies required for the benchmark script including python packages and libraries + pip install -r requirements.txt + cd MAD/scripts/vllm + +--- + +4. Authentication - Hugging Face Token +--------------------------------------- + +To download LLM models such as `Llama-3`, you need a Hugging Face account and an access token. +Create a Hugging Face account and generate a personal access token from your Hugging Face profile. Then export it: + +.. code-block:: bash + + # create Hugging Face account and use that token below + export HF_TOKEN="your_HuggingFace_token" + +This token is used to pull LLM models (e.g., Llama-3) from Hugging Face. + +--- + +5. How to Check if an LLM Fits within a CPX GPU? +-------------------------------------------------- + +Before executing a workload, it is critical to ensure that the selected model can fit entirely within the memory available on a single CPX GPU partition, particularly when using CPX/NPS4 mode (which provides 24GB per partition). + +A model will typically fit on a single GPU partition if: + +:: + + size_of_weights + size_of_KV_cache < available_VRAM + +The size of the **Key-Value (KV) cache** is workload-dependent and driven by batch size, sequence length, and model architecture. + +**Formula for KV cache size per token (in bytes):** + +:: + + kv_cache_per_token = 2 × n_layers × hidden_dim × precision_in_bytes + +Where: + +- `2` accounts for Key and Value caches +- `n_layers` is the number of transformer layers +- `hidden_dim` = n_heads × d_head +- `precision_in_bytes` is 2 for float16 and bfloat16, 4 for float32 + +**Total KV cache size in bytes:** + +:: + + total_kv_cache = batch_size × seq_length × kv_cache_per_token + +--- + +**Example 1: Llama-2-7B (FP16)** + +:: + + Model weights ≈ 2 × 7 = 14 GB + kv_cache_per_token = 2 × 32 × 4096 × 2 = 524,288 bytes + total_kv_cache = 1 × 4096 × 524,288 ≈ 2 GB + Total memory usage = 14 GB + 2 GB = 16 GB + +✅ This model fits within a single CPX GPU partition (24 GB VRAM). + +--- + +**Example 2: Llama-2-13B (FP16)** + +:: + + Model weights ≈ 2 × 13 = 26 GB + kv_cache_per_token = 2 × 40 × 5120 × 2 = 819,200 bytes + total_kv_cache = 1 × 4096 × 819,200 ≈ 3.6 GB + Total memory usage = 26 GB + 3.6 GB ≈ 29.6 GB + +❌ This model exceeds a single 24 GB CPX partition and will require multiple partitions or tensor parallelism. + +--- + +How do above partitions affect LLM models? +~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ + + - With reduced memory in the CPX mode (24GB HBM per XCD), models may not fit within one CPX GPU. So, models have to be partitioned across multiple CPX GPUs using tensor parallelism. + - With reduced compute in the CPX mode (38CUs per XCD), models may be compute bounded if they run on lesser XCD units compared to SPX mode. As above, using tensor parallelism to split the model across multiple CPX GPUs can take advantage of more compute units. + +**Summary:** Always pre-calculate memory needs and compute needs for your selected model and batch size to determine the appropriate number of CPX partitions (i.e., GPUs) to assign. + +--- + +6. GPU Selection (Optional) +---------------------------- + +If you wish to limit the vLLM workload to a specific set of GPUs (e.g., 8 out of the total available), define the HIP_VISIBLE_DEVICES environment variable. If left unset, all GPUs are utilized. + +.. code-block:: bash + + # set the environment variables for the GPUs to be used + # leave this blank if you want to use all GPUs + # to use the first 8 GPUs, set the variable to 0,1,2,3,4,5,6,7 + export HIP_VISIBLE_DEVICES=0,1,2,3,4,5,6,7 + +--- + +7. Running the vLLM Benchmark +------------------------------ + +The benchmark script accepts several command-line options to customize the test. Here's an example that runs the meta-llama/Llama-3.1-8B-Instruct model on 8 GPUs in FP16 mode: + +.. code-block:: bash + + # from the app/MAD/scripts/vllm directory, run the following command to run the vllm workload + # -g 1 means to use 1 GPU, -g 8 means to use 8 GPUs, etc. + # please refer to the README on the ROCm/MAD GitHub repo for more details on the command line options + ./vllm_benchmark_report.sh -s all -m meta-llama/Llama-3.1-8B-Instruct -g 8 -d float16 + +--- + +**Command Breakdown:** + +- `-s all`: Run all benchmark tests (latency, throughput, etc.) +- `-m`: Hugging Face model name to use (e.g., `meta-llama/Llama-3.1-8B-Instruct`) +- `-g`: Number of GPUs to use +- `-d`: Precision mode (choose from `float16`, `bfloat16`, `float32`) + +For additional options (e.g., batch size, sequence length, tokenizer config), refer to the `MAD/scripts/vllm/README.md` file in the GitHub repository. + +.. note:: + Ensure that your container has internet access to pull models from Hugging Face during benchmarking. diff --git a/docs/gpu-partitioning/mi300x/troubleshooting.rst b/docs/gpu-partitioning/mi300x/troubleshooting.rst new file mode 100644 index 0000000..3cd95e8 --- /dev/null +++ b/docs/gpu-partitioning/mi300x/troubleshooting.rst @@ -0,0 +1,174 @@ +Troubleshooting +================== + +This section provides a comprehensive guide to diagnosing and resolving common issues that may arise during the GPU partitioning and workload execution process on AMD MI300X systems. Given the complexity of managing GPU resources across multiple partitions, users may encounter errors stemming from permission misconfigurations, resource contention, incompatible memory modes, or system-level driver conflicts. + +The troubleshooting guidance here aims to equip system administrators, ML engineers, and HPC practitioners with actionable solutions and detailed context to restore correct system behavior and accelerate productivity. Each issue is presented with relevant command examples, expected error messages, and precise resolution steps to minimize guesswork and enable swift remediation. + +Topics covered include: +- Resolving permission-denied errors when invoking `amd-smi` partitioning commands. +- Handling GPU resource conflicts during partition creation or reset. +- Diagnosing partition incompatibilities between memory and compute modes. +- Addressing missing GPU visibility in CPX mode caused by Linux kernel video driver issues. + +For optimal system operation, it is recommended to thoroughly read through each scenario and verify all preconditions (e.g., compute mode, memory mode, driver state) before applying corrective actions. In complex environments with multiple GPUs or users, coordination across software stack components — from container runtimes and kernel drivers to benchmarking frameworks — is crucial for sustained platform stability and performance. + +1. Error while running amd-smi command for partitioning +-------------------------------------------------------- + +If you receive an error when executing partitioning commands using `amd-smi`, verify that you are running the command with elevated permissions. + + .. tab-set:: + + .. tab-item:: Command + + .. code-block:: shell-session + + # Check if the GPU is in use + amd-smi set --gpu all --compute-partition SPX + + + .. tab-item:: Shell output + + :: + + amdsmi.amdsmi_exception.AmdSmiLibraryException: Error code: + 10 | AMDSMI_STATUS_NO_PERM - Permission Denied + + The above exception was the direct cause of the following exception: + + PermissionError: Command requires elevation + +**Resolution:** + +- Ensure the command is executed with `sudo`. +- Confirm that no applications or system services are currently utilizing the GPU. If so, terminate or stop them before retrying. + +2. Error while creating partitions +----------------------------------- + +Partitioning operations can fail if the GPU is actively being used by another process. + + .. tab-set:: + + .. tab-item:: Command + + .. code-block:: shell-session + + # setting memory partition mode + sudo amd-smi set --gpu all --compute-partition SPX + + .. tab-item:: Shell output + + :: + + amdsmi.amdsmi_exception.AmdSmiLibraryException: Error code: + 30 | AMDSMI_STATUS_BUSY - Device busy + + The above exception was the direct cause of the following exception: + + ValueError: Unable to set accelerator partition to SPX on GPU ID: 0 BDF:0000:11:00.0 + +**Resolution:** + +- Ensure that no compute workloads or system services are using the GPU. +- Use `amd-smi`, `ps -aux`, `top` -like tools to identify running GPU jobs. +- Terminate conflicting jobs before retrying the partition command. + +3. Error while resetting partition to SPX +------------------------------------------- + +NPS4 memory mode is only compatible with CPX compute mode. If you attempt to switch to SPX while memory mode is still set to NPS4, the operation will fail. + + .. tab-set:: + + .. tab-item:: Command + + .. code-block:: shell-session + + # Set compute partition mode + sudo amd-smi set --gpu all --compute-partition SPX + + .. tab-item:: Shell output + + :: + + Attempted to set accelerator partition to SPX (profile #0 on GPU ID: 0 BDF:0000:11:00.0 + + [AMDSMI_STATUS_SETTING_UNAVAILABLE] Please check amd-smi partition --memory --accelerator for available profiles. + Users may need to switch memory partition to another mode in order to enable the desired accelerator partition. + + amdsmi.amdsmi_exception.AmdSmiLibraryException: Error code: + 55 | AMDSMI_STATUS_SETTING_UNAVAILABLE - Setting is not available + + The above exception was the direct cause of the following exception: + + ValueError: [AMDSMI_STATUS_SETTING_UNAVAILABLE] Unable to set accelerator partition to SPX on GPU ID: 0 BDF:0000:11:00.0 + +**Resolution:** + +Before switching to SPX mode, first revert the memory partition mode to NPS1: + + .. tab-set:: + + .. tab-item:: Command + + .. code-block:: shell-session + + # Set memory partition mode + sudo amd-smi set --memory-partition NPS1 + + .. tab-item:: Shell output + + :: + + GPU: 0 + MEMORY_PARTITION: Successfully set memory partition to NPS1 + + GPU: 1 + MEMORY_PARTITION: Successfully set memory partition to NPS1 + + GPU: 2 + MEMORY_PARTITION: Successfully set memory partition to NPS1 + + GPU: 3 + MEMORY_PARTITION: Successfully set memory partition to NPS1 + + GPU: 4 + MEMORY_PARTITION: Successfully set memory partition to NPS1 + + GPU: 5 + MEMORY_PARTITION: Successfully set memory partition to NPS1 + + GPU: 6 + MEMORY_PARTITION: Successfully set memory partition to NPS1 + + GPU: 7 + MEMORY_PARTITION: Successfully set memory partition to NPS1 + +Once complete, you can safely reset compute partitioning to SPX mode. + +4. All 64 GPUs not visible in `amd-smi` output in CPX mode +----------------------------------------------------------- + +In CPX mode, the system should expose 64 logical GPUs (8 per physical MI300X device). If you observe fewer GPUs, it may be due to a known Linux kernel issue involving the BMC virtual video driver. On most systems this virtual video driver is AST (ASPEED AST media controller). + +**Resolution:** + +Unload the AST video driver and reload the AMD GPU kernel modules: + +.. code-block:: bash + + # Unload the BMC virtual video driver + sudo modprobe -r ast + + # Unload the amdgpu driver + sudo modprobe -r amdgpu + + # Load the amdgpu driver + sudo modprobe amdgpu + +After these steps, rerun `amd-smi` to verify that all 64 GPUs are now visible. + +.. note:: + If the AST driver cannot be unloaded due to it being in use, consider blacklisting the AST module in `/etc/modprobe.d/blacklist.conf` and rebooting the system. diff --git a/docs/index.md b/docs/index.md index 875eaa7..92cfbfd 100644 --- a/docs/index.md +++ b/docs/index.md @@ -27,6 +27,8 @@ The AMD Instinct documentation is organized into the following categories: :class-body: rocm-card-banner rocm-hue-12 * [System optimization](./system-optimization/index.rst) +* [GPU Partitioning](./gpu-partitioning/index.rst) + ::: :::{grid-item-card} Conceptual diff --git a/docs/sphinx/_toc.yml.in b/docs/sphinx/_toc.yml.in index 5083212..2721140 100644 --- a/docs/sphinx/_toc.yml.in +++ b/docs/sphinx/_toc.yml.in @@ -56,6 +56,30 @@ subtrees: title: AMD Instinct MI200 - file: system-optimization/mi100.md title: AMD Instinct MI100 + - file: gpu-partitioning/index.rst + title: GPU Partitioning + subtrees: + - entries: + - file: gpu-partitioning/mi300x/index.rst + title: AMD Instinct MI300X GPU + subtrees: + - entries: + - file: gpu-partitioning/mi300x/overview.rst + title: GPU Partitioning Overview + - file: gpu-partitioning/mi300x/requirements.rst + title: Requirements + - file: gpu-partitioning/mi300x/quick-start-guide.rst + title: Quick Start Guide + - file: gpu-partitioning/mi300x/troubleshooting.rst + title: Troubleshooting + - file: gpu-partitioning/mi300x/run-vllm.rst + title: Running an example VLLM workload + - file: gpu-partitioning/mi300a/index.rst + title: AMD Instinct MI300A APU + subtrees: + - entries: + - file: gpu-partitioning/mi300a/overview.rst + title: GPU Partitioning Overview - caption: Conceptual entries: - file: conceptual/iommu.rst