Hardware reference
Every hands-on page in this course is written for four kinds of machine, and every one states the smallest memory that can follow its primary path. This page is the reference for both.
The figures below are from vendor documentation, retrieved on the date shown. Bandwidth is the theoretical figure; the measured figure from the validation hardware replaces it as the validation pass reaches each machine.
| Track | Machines | Memory | Bandwidth | Compute | Checked |
|---|---|---|---|---|---|
| S NVIDIA DGX Spark | NVIDIA DGX Spark Founders Edition and the OEM systems on the same GB10 board: Acer Veriton GN100, ASUS Ascent GX10, Dell Pro Max with GB10, Gigabyte AI Top Atom, HP ZGX Nano, Lenovo ThinkStation PGX, MSI EdgeXpert | 128 GB unified | 273 GB/s | Blackwell GPU with 20-core Arm CPU (GB10), CUDA 13 on aarch64 | 2026-09-08 |
| X AMD Ryzen AI Max+ 395 | GMKtec EVO-X2 and other Ryzen AI Max+ 395 (Strix Halo) mini PCs and laptops | 64 / 128 GB unified | 256 GB/s | Radeon 8060S integrated GPU (40 RDNA 3.5 compute units, gfx1151), 16 Zen 5 cores, XDNA 2 NPU; Vulkan and ROCm | 2026-09-08 |
| M Apple silicon Mac | Mac mini, Mac Studio and MacBook Pro with M4 and later, plus the M3 Ultra Mac Studio for its 512 GB ceiling | 24 / 32 / 64 / 128 / 256 / 512 GB unified | 546 GB/s* | Apple GPU with Metal; MLX | 2026-09-08 |
| N NVIDIA desktop or laptop | Windows or Linux desktops and laptops with GeForce RTX 30, 40 and 50 series or RTX PRO GPUs | 8 / 12 / 16 / 24 / 32 / 48 / 96 GB VRAM | 1792 GB/s* | CUDA 13; WSL2 on Windows | 2026-09-08 |
* Varies by chip or card: Apple silicon Mac: By chip, from Apple's specification pages read 2026-09-09: M4 120 GB/s, M4 Pro 273, M4 Max 546, M6 153 or 170, M5 Pro 307, M5 Max 460 or 614 depending on GPU configuration, M3 Ultra 819, M5 Ultra 1200. NVIDIA desktop or laptop: By card: RTX 3090 936 GB/s, RTX 4090 1008, RTX 5090 1792, RTX PRO 6000 Blackwell 1792; laptop parts are lower.
- Track S — NVIDIA DGX Spark
Operating system: DGX OS 7.x (Ubuntu 24.04 base)
Interconnects: ConnectX-7 with two QSFP ports at up to 200 Gb/s (RoCE); 10 GbE; Wi-Fi 7
Best at: The full NVIDIA training and serving stack at 128 GB; the vendor-documented two-machine cluster over ConnectX-7.
Limits: 273 GB/s bandwidth bounds decode speed; aarch64 means some wheels and containers need NVIDIA's arm64 builds.
Sources:NVIDIA DGX Spark product page (2026-09-08) · DGX Spark release notes (2026-09-08) · DGX Spark playbooks (2026-09-08)
- Track X — AMD Ryzen AI Max+ 395
Operating system: Ubuntu 24.04 HWE or 26.04, Fedora 43, or Windows 11
Interconnects: 2.5 GbE; USB4 / Thunderbolt-class ports; Wi-Fi 7
Best at: Best price per gigabyte for large-model inference at 128 GB; the honest experience of a maturing GPU stack.
Limits: GPU-visible memory is capped below the total (about 96 GB on Windows; the GTT limit on Linux); the ROCm 10.0.0 compatibility matrix (dated 2026-08-14, read 2026-09-09) lists gfx1151 without a support-tier qualifier while AMD's install pages did not mention the chip; vLLM's GPU installation page (read 2026-09-09) lists Ryzen AI MAX / AI 300 (gfx1151/1150) among its ROCm targets with pre-built wheels for ROCm 7.0 and 7.2.1, which the validation pass has yet to exercise.
Sources:GMKtec EVO-X2 product page (2026-09-08) · ROCm on Strix Halo (2026-09-08) · ROCm compatibility matrix (2026-09-08)
- Track M — Apple silicon Mac
Operating system: macOS 26 (Tahoe); 26.2 or later for RDMA over Thunderbolt 5
Interconnects: Thunderbolt 5 (M4 Pro and above) with RDMA on macOS 26.2 or later; 10 GbE on Mac Studio
Best at: Highest memory at high bandwidth; MLX; vendor-supported clustering over Thunderbolt 5 RDMA.
Limits: No CUDA; the GPU wired-memory limit must be raised for large models; vLLM on macOS means either its CPU build or the separate vLLM-Metal plugin listed on its installation index (read 2026-09-09), not the mainline GPU path the course teaches in Part 9.
Sources:Mac Studio technical specifications (2026-09-08) · MLX (2026-09-08) · mlx-lm (2026-09-08)
- Track N — NVIDIA desktop or laptop
Operating system: Ubuntu 24.04 or later, or Windows 11 with WSL2
Interconnects: PCIe (no NVLink on the 40 and 50 series); 1 to 25 GbE
Best at: Fastest decode per pound at small and mid sizes; the deepest training-tool support.
Limits: VRAM is the ceiling (32 GB on consumer cards); system RAM offload is slow; multi-GPU is PCIe-bound.
Sources:NVIDIA GeForce RTX 50 series (2026-09-08) · CUDA on WSL user guide (2026-09-08) · CUDA Toolkit release notes (2026-09-08)
Memory tiers
| Tier | What it can follow |
|---|---|
| 8 GB | 8B-class models at Q4; every Level 1 and 2 lab on its reduced path. |
| 12–16 GB | 14B-class at Q4 to Q6; LoRA fine-tuning of 1–4B models; the Track N validation tier. |
| 24 GB | 32B-class at Q4; QLoRA of 8B; the 14B-class distillation teacher; the Track M validation tier at 24 GB unified. |
| 32 GB | 30B-class MoE at Q8; gpt-oss-20b in MXFP4 with room; RTX 5090. |
| 48–64 GB | 70B-class at Q4; 30B-class teachers at Q8; M4 Pro and mid Mac Studio configurations. |
| 96 GB | gpt-oss-120b in MXFP4; RTX PRO 6000 Blackwell. |
| 128 GB | The DGX Spark, EVO-X2 and M4 Max/M5 Max tier: 120B-class MoE, 235B-class at IQ4 on Spark and Mac, 30B-class full fine-tunes with QLoRA. |
| 256 GB | Two 128 GB machines clustered, or a 256 GB Mac Studio: 235B-class at Q8, 400B-class at FP4 on a Spark pair. |
| 512 GB | M3 Ultra and M5 Ultra Mac Studio: 671B-class MoE at Q4. |