What You Need
The short version
Section titled “The short version”Required: one machine from one of the four platform tracks below, with at least 8 GB of GPU or unified memory, a recent operating system, and enough disk for the models a tier can run.
Not required: a second machine, a GPU with more than 16 GB, a paid API, a cloud account, or any subscription. Every cluster lab has a single-machine path, and every lab states a reduced path for smaller memory.
The course can be completed end to end on a 16 GB NVIDIA laptop with smaller models, with two exceptions that need Apple silicon or a DGX Spark pair and say so.
Choose your track
Section titled “Choose your track”Your platform track
Stored in this browser only. Every lab still shows all four tracks; yours is opened first.
Every hands-on page opens your track first and keeps the other three collapsed on the same page. Change the choice here at any time.
The four tracks
Section titled “The four tracks”| Track | Machines | Memory | Bandwidth | Compute |
|---|---|---|---|---|
| S NVIDIA DGX Spark | NVIDIA DGX Spark | 128 GB unified | 273 GB/s | Blackwell GPU with 20-core Arm CPU (GB10), CUDA 13 on aarch64 |
| X AMD Ryzen AI Max+ 395 | AMD Ryzen AI Max+ 395 | 64 / 128 GB unified | 256 GB/s | Radeon 8060S integrated GPU (40 RDNA 3.5 compute units, gfx1151), 16 Zen 5 cores, XDNA 2 NPU; Vulkan and ROCm |
| M Apple silicon Mac | Apple silicon Mac | 24 / 32 / 64 / 128 / 256 / 512 GB unified | 546 GB/s* | Apple GPU with Metal; MLX |
| N NVIDIA desktop or laptop | NVIDIA desktop or laptop | 8 / 12 / 16 / 24 / 32 / 48 / 96 GB VRAM | 1792 GB/s* | CUDA 13; WSL2 on Windows |
* Varies by chip or card: Apple silicon Mac: By chip, from Apple's specification pages read 2026-09-09: M4 120 GB/s, M4 Pro 273, M4 Max 546, M6 153 or 170, M5 Pro 307, M5 Max 460 or 614 depending on GPU configuration, M3 Ultra 819, M5 Ultra 1200. NVIDIA desktop or laptop: By card: RTX 3090 936 GB/s, RTX 4090 1008, RTX 5090 1792, RTX PRO 6000 Blackwell 1792; laptop parts are lower.
Memory tiers
| Tier | What it can follow |
|---|---|
| 8 GB | 8B-class models at Q4; every Level 1 and 2 lab on its reduced path. |
| 12–16 GB | 14B-class at Q4 to Q6; LoRA fine-tuning of 1–4B models; the Track N validation tier. |
| 24 GB | 32B-class at Q4; QLoRA of 8B; the 14B-class distillation teacher; the Track M validation tier at 24 GB unified. |
| 32 GB | 30B-class MoE at Q8; gpt-oss-20b in MXFP4 with room; RTX 5090. |
| 48–64 GB | 70B-class at Q4; 30B-class teachers at Q8; M4 Pro and mid Mac Studio configurations. |
| 96 GB | gpt-oss-120b in MXFP4; RTX PRO 6000 Blackwell. |
| 128 GB | The DGX Spark, EVO-X2 and M4 Max/M5 Max tier: 120B-class MoE, 235B-class at IQ4 on Spark and Mac, 30B-class full fine-tunes with QLoRA. |
| 256 GB | Two 128 GB machines clustered, or a 256 GB Mac Studio: 235B-class at Q8, 400B-class at FP4 on a Spark pair. |
| 512 GB | M3 Ultra and M5 Ultra Mac Studio: 671B-class MoE at Q4. |
The hardware reference has the full detail for each track: operating systems, interconnects, what each is best at and where each stops.
Memory tiers
Section titled “Memory tiers”Every hands-on page states the smallest memory that can follow its primary path, from these tiers, and a reduced path below it. The tier decides which models the labs use, not whether you can do them.
| Tier | What it can follow |
|---|---|
| 8 GB | Every Level 1 and 2 lab on its reduced path with 8B-class models at Q4 |
| 12–16 GB | 14B-class models; fine-tuning of 1–4B models; the whole course on the reduced paths |
| 24 GB | 32B-class at Q4; fine-tuning of 8B models; the distillation labs with a 14B teacher |
| 32 GB | 30B-class mixture-of-experts at Q8; the RTX 5090 tier |
| 48–64 GB | 70B-class at Q4; 30B-class teachers |
| 96 GB | The 120B-class mixture-of-experts models |
| 128 GB | The DGX Spark, EVO-X2 and M4 Max tier: 235B-class at IQ4, 30B-class teachers at Q8, the cluster labs at their intended scale |
| 256–512 GB | Two 128 GB machines, or a Mac Studio: 400B-class and 671B-class models |
Operating systems
Section titled “Operating systems”- Track S: DGX OS as shipped. The labs use NVIDIA’s containers where a tool has no aarch64 wheel.
- Track X: Ubuntu 24.04 with the hardware-enablement kernel, Ubuntu 26.04 or Fedora 43 for the ROCm path; Windows 11 is shown for llama.cpp, Ollama and LM Studio only.
- Track M: macOS 26. The two-Mac cluster needs 26.2 or later on both machines.
- Track N: Ubuntu 24.04 or later, or Windows 11 with WSL2. The Linux commands are the primary path; native Windows is shown for llama.cpp, Ollama and LM Studio.
Disk and downloads
Section titled “Disk and downloads”Models are large. The reference set for the 8 GB tier is under 20 GB in total; the 128 GB tier’s set, including the 235B-class model, is several hundred gigabytes. Each lab lists the download size of the models it uses, and Part 7 teaches a model library shared between every tool so that a model is downloaded once per home. Budget 500 GB of free disk for the Level 2 and 3 labs at the 24 GB tier and 2 TB at the 128 GB tier.
Network
Section titled “Network”For the single-machine course, none beyond downloading. For the cluster parts, two machines on the same wired network; Part 18 says what each link can deliver and which parallelism it can afford, and the labs were validated over the direct ConnectX-7 link between two Sparks and over ordinary Ethernet.
Accounts and keys
Section titled “Accounts and keys”A free Hugging Face account, for the gated models the course names and for uploading your own fine-tunes to a private repository. No paid API is used anywhere in the course; the agent parts point every tool at your own server.
Versions the course was validated against
Section titled “Versions the course was validated against”Every tool the course teaches is pinned to the version its pages were written against, and the date that version was checked. The build warns when a pin is more than four months old, and the pins are reviewed every quarter. Where a page depends on a specific version, it says so inline.
| Tool | Version | Tracks | Checked | Notes |
|---|---|---|---|---|
| llama.cpp | v0.4.0 | S X M N | 2026-09-08 | Versioned release v0.4.0 (2026-09-04, ggml 0.23.0) confirmed through the GitHub API; the project also publishes rolling build tags (b10867 on 2026-09-08) whose numbers name the prebuilt archives. RPC backend with RDMA on Linux (RoCE) and macOS (Thunderbolt 5). |
| Ollama | 0.33.3 | S X M N | 2026-09-08 | MLX engine in preview on Apple silicon from 0.19. |
| LM Studio | 0.4.23 | S X M N | 2026-09-08 | Proprietary freeware; llama.cpp and MLX engines; headless daemon. Version read from the download page on 2026-09-08 (release date not shown). |
| vLLM | 0.28.0 | S N | 2026-09-08 | vLLM's GPU installation page (read 2026-09-09) lists Ryzen AI MAX / AI 300 (gfx1151/1150) among its ROCm targets with pre-built wheels for ROCm 7.0 and 7.2.1, not yet exercised by the validation pass; on macOS only the CPU build and a separate Metal plugin are listed. Disaggregated prefill is marked experimental. |
| SGLang | 0.5.19 | S N | 2026-09-08 | DGX Spark support to confirm on hardware. |
| TensorRT-LLM | 1.2.1 | S N | 2026-09-08 | Latest stable release is 1.2.1; 1.3 is a release candidate. Taught through NGC containers on Spark; local aarch64 builds are experimental. |
| NVIDIA Dynamo | 1.4.2 | N | 2026-09-08 | Surveyed, not taught hands-on: multi-node requires RDMA and GB10 is untested. |
| mlx-lm | 0.31.3 | M | 2026-09-08 | Distributed inference via mlx.launch with ring, MPI and RDMA backends. |
| exo | main | M | 2026-09-08 | README (2026-09-08): app needs macOS 26.2 or later; RDMA over Thunderbolt 5 on M4 Pro Mac mini, M4 Max Mac Studio and MacBook Pro, M3 Ultra Mac Studio; Linux runs on CPU only with GPU support under development; installed from source with uv sync --extra mlx. |
| ExLlamaV3 | 1.4.8 | N | 2026-09-08 | CUDA 12.4 or later; EXL3 format; served through TabbyAPI. |
| ktransformers | 0.7.0 | N | 2026-09-08 | CPU plus GPU hybrid inference for very large mixture-of-experts models; 0.7.0 adds AMD AVX-512 CPU support. |
| llama-swap | v255 | S X M N | 2026-09-08 | |
| LiteLLM | 1.100.0 | S X M N | 2026-09-08 | |
| Open WebUI | 0.11.3 | S X M N | 2026-09-08 | Custom Open WebUI Licence since April 2025; LibreChat and AnythingLLM (MIT) are the alternatives the course names. |
| Hugging Face CLI | 1.30.0 | S X M N | 2026-09-08 | Provides the hf command-line tool. |
| transformers | 5.16.1 | S X N | 2026-09-08 | No MLX target; the Mac training path uses mlx-lm or PyTorch MPS. |
| TRL | 1.12.0 | S X N | 2026-09-08 | Release 1.12.0 is noted on its page as an accidental duplicate of 1.11.0. Stable: SFT, DPO, KTO, GRPO, RLOO, Reward and Distillation trainers. |
| PEFT | 0.20.0 | S X N | 2026-09-08 | |
| Unsloth | 0.1.807-beta (GitHub tag) | S X N | 2026-09-08 | GitHub release tag as listed on 2026-09-08; confirm the installed pip version string on the lab machines. DGX Spark guide and NVIDIA playbook. Unsloth's AMD installation page read on 2026-09-09 names RDNA 3, 3.5 and 4 discrete cards and the MI300X and does not name the Ryzen AI Max+ (gfx1151), so Track X's primary fine-tuning path is TRL on ROCm. QAT supported. |
| LLaMA-Factory | 0.9.5 | S N | 2026-09-08 | One of NVIDIA's DGX Spark fine-tuning playbooks; 0.9.5 adds Gemma 4 and Transformers v5 support. |
| Axolotl | 0.18.0 | N | 2026-09-08 | Release 0.18.0 emphasises fine-tuning large sparse mixture-of-experts models. |
| nanochat | main | S X M N | 2026-09-08 | The pretraining lab codebase. |
| lm-evaluation-harness | 0.4.13 | S X M N | 2026-09-08 | |
| Aider | 0.86.0 | S X M N | 2026-09-08 | |
| OpenAI Codex CLI | 0.153.4 | S X M N | 2026-09-08 | Apache-2.0; release tags are rust-v<version>; --oss and model_providers in config.toml. The config reference read on 2026-09-09 documents wire_api = "responses" only, so a custom provider against a plain chat-completions gateway is not a documented path. |
| OpenCode | 1.18.29 | S X M N | 2026-09-08 | MIT; opencode.json providers. The docs read on 2026-09-09 document Ollama, llama.cpp and LM Studio; vLLM and MLX servers are configured the same way but are not named in the docs. |
| Claude Code | current | S X M N | 2026-09-08 | Proprietary; local endpoint through ANTHROPIC_BASE_URL and llama-server's Anthropic Messages endpoint. |
| Goose | 1.50.0 | S X M N | 2026-09-08 | |
| OpenHands | 1.16.0 | S X M N | 2026-09-08 | |
| iperf3 | 3.21 | S X M N | 2026-09-08 | |
| Docker Engine | current | S X N | 2026-09-08 | Podman is accepted where the labs say so. |
| uv | 0.12.11 | S X M N | 2026-09-08 | Python environment and package manager used by every lab. |
| Lemonade Server | current | S X M N | 2026-09-09 | AMD's local server (Apache-2.0) wrapping llama.cpp and the Ryzen AI NPU flow; OpenAI-compatible API on port 13305 by default. The CLI documents run, chat, status and pull (no serve subcommand) as read on 2026-09-09. Version to pin from the binary in the validation pass. |
| TabbyAPI | current | N | 2026-09-09 | OpenAI-compatible server for ExLlamaV3 (AGPL-3.0); the README states no release version, so it is pinned by commit in the validation pass. |
| mistral.rs | current | S X M N | 2026-09-09 | Rust inference engine (MIT) with CUDA, Metal and CPU backends; the README states no release version, so it is pinned by tag in the validation pass. |
| verl | main | S N | 2026-09-09 | README read 2026-09-09: Apache-2.0, ByteDance Seed. PPO, GRPO, GSPO, DAPO, Dr. GRPO, RLOO, REINFORCE++; rollouts through vLLM, SGLang or Transformers; FSDP2 or Megatron; NVIDIA, ROCm MI300X and Ascend. Surveyed in Part 14, not taught hands-on; release tag to capture in the validation pass. |
| OpenRLHF | main | N | 2026-09-09 | README read 2026-09-09: Apache-2.0; Ray, vLLM and DeepSpeed ZeRO-3; PPO with critic, GRPO, RLOO, REINFORCE++, DPO, IPO, cDPO; written around 70B models on eight 80 GB GPUs. Surveyed in Part 14, not taught hands-on; release tag to capture in the validation pass. |
| Medusa | main | S N | 2026-09-09 | Apache-2.0; draft heads trained with medusa/train/train_legacy.py on ShareGPT-style data (README read 2026-09-09). Taught in Part 17 as the exercise; not a vLLM speculative method as of that date. Pinned by commit in the validation pass. |
| EAGLE | main | S N | 2026-09-09 | Apache-2.0; EAGLE-3 training through eagle/traineagle3/main.py under DeepSpeed, with the README pointing at SpecForge (read 2026-09-09). Pinned by commit in the validation pass. |
| SpecForge | main | S N | 2026-09-09 | MIT; the SGLang project's draft-model trainer, run as specforge train with a YAML config (training guide read 2026-09-09). Pinned by commit in the validation pass. |
| LMCache | current | S N | 2026-09-09 | KV-cache layer with GPU, CPU, disk and remote tiers; vLLM connectors LMCacheConnectorV1 and LMCacheMPConnector. Docs at https://docs.lmcache.ai/ read 2026-09-09; no release version stated, to pin in the validation pass. |
| Mooncake | current | S N | 2026-09-09 | Apache-2.0. Transfer Engine plus Mooncake Store; README (read 2026-09-09) lists TCP, RDMA, NVMe-oF and NVLink transports; integrated by vLLM, SGLang and LMDeploy. Surveyed in Part 22. |
| NIXL | current | S N | 2026-09-09 | NVIDIA Inference Xfer Library, Apache-2.0. Plug-ins include UCX, GDS and POSIX; README (read 2026-09-09) states it was tested with UCX 1.23.x. Surveyed in Part 22. |
| Prometheus | current | S X M N | 2026-09-09 | Configured through prometheus.yml and rule files rather than flags; default listen port 9090. Image tag to pin from the lab machines in the validation pass. |
| Grafana | current | S X M N | 2026-09-09 | Datasources and dashboards provisioned from files; default listen port 3000; admin password from the environment. Image tag to pin in the validation pass. |
| Prometheus node exporter | current | S X M N | 2026-09-09 | Host metrics for the dashboards lab; default port 9100. Image tag to pin in the validation pass. |
| NVIDIA DCGM exporter | current | S N | 2026-09-09 | GPU telemetry for Prometheus on Tracks S and N; default port 9400. Whether the free-memory field is exported is confirmed on the reader's own /metrics. Image tag to pin in the validation pass. |
| MCP Python SDK | 2.x | S X M N | 2026-09-09 | MIT; Python 3.10 or later. Version 2 is a rework supporting the MCP specification dated 2026-07-28, and pip install mcp now installs 2.x; mcp dev and mcp run come from the cli extra (README read 2026-09-09). Exact release to pin in the validation pass. |
| Podman | current | S X M N | 2026-09-09 | Accepted where the labs say Docker or Podman; rootless by default. Version to pin on the lab machines. |
| SkyRL | v0.3.0 | S N | 2026-09-09 | README read 2026-09-09: Apache-2.0, Berkeley Sky Computing Lab with Anyscale. Full-stack RL library for multi-turn tool use and long-horizon agent tasks (skyrl-train, skyrl-agent, skyrl-gym, skyrl-tx) with vLLM integration. Surveyed in Part 27, not taught hands-on; release tag and date to confirm in the validation pass. |
| PyTorch | 2.14.0 | S X M N | 2026-09-12 | The training and inference framework behind Parts 1, 2 and 11 to 17. Wheels: cu130 and cu126 indexes for CUDA, rocm7.2 nightly index for Strix Halo, PyPI (MPS) on Mac. Version and release date read from PyPI on 2026-09-12. |
| torchvision | 0.29.0 | S X M N | 2026-09-12 | Provides the MNIST dataset loader for the Part 1 lab; installed from the same index as PyTorch. Version read from PyPI on 2026-09-12. |
| JupyterLab | 4.6.3 | S X M N | 2026-09-12 | The notebook environment set up in the Part 1 lab and used for the loss-curve plots. Version read from PyPI on 2026-09-12. |
| MLX | 0.32.2 | M | 2026-09-12 | Apple's array framework; the Part 1 lab trains the MNIST network with it side by side with PyTorch MPS, and mlx-lm (pinned separately) builds on it. Version read from PyPI on 2026-09-12. |
| Matplotlib | 3.11.2 | S X M N | 2026-09-12 | Plots the loss and validation curves in the Part 1 lab (plot-curves.py). Not in the NGC PyTorch container, so the Spark container path installs it explicitly. Version read from PyPI on 2026-09-12. |