Command reference
This page is generated from the same file the build validates commands against. Every shell block and every lab script on the site is parsed, and every option passed to a tool listed here is checked against this list. An option that does not exist fails the build.
The list is deliberately not a manual. It records which options exist, so that the course cannot invent one; what each option does is taught on the page that uses it, linked from each entry.
33 tools, 0 with option lists captured from the pinned version's help output, generated 2026-09-08.
Tools marked seed have option lists written from documentation rather than captured from --help; the build treats an unknown option on such a tool as a warning, not an error, until the capture is made on the lab machines.
aiderAider 0.86.0 · seed-h --help --model --edit-format --weak-model --api-key --set-env --message -m --test-cmd --auto-test --map-tokens --read --git --no-git --auto-commits --no-auto-commits --yes-always --cache-prompts --verify-ssl --model-settings-file --openai-api-base --openai-api-key --no-show-model-warningsTaught in:Capstone 5: The Agentic Workstation, Aider with Local Models, Challenge: The Agent That Escaped the Sandbox, Lab: One Task, Six Agents, Lab: Sandbox Your Agent: Containers, Permissions and Secrets, Project: A Local Agentic Coding Workstation, The Landscape: Terminal, Editor and Autonomous Agents
codexOpenAI Codex CLI 0.153.4 · seed-h --help --oss --local-provider --sandbox --ask-for-approval -a --dangerously-bypass-approvals-and-sandbox --yolo --full-auto --image --search --config -c --model -m --cd -CSubcommands:
exec,resume,cloud,mcp,login,logout,completion,appTaught in:Lab: Write and Connect an MCP Server, OpenAI Codex CLI and OpenCode with Local Models, Lab: One Task, Six Agents, Project: A Local Agentic Coding Workstation, The Landscape: Terminal, Editor and Autonomous Agents
exoexo main · seed-h --help -q --quiet -v --verbose -m --force-master --no-api --api-port --no-worker --no-downloads --offline --no-batch --legacy-daemon --bootstrap-peers --namespace --zenoh-port --discovery-port --fast-synch --no-fast-synchTaught in:exo: Automatic Partitioning Across Your Macs, Lab: A Two-Mac Cluster over Thunderbolt 5, Reality Check: 'Four Mac Studios Replace a GPU Server'
ggml-rpc-serverllama.cpp v0.4.0 · seed--help -h -H --host -p --port -c --cache -t --threads -d --deviceTaught in:Capstone 3: Cluster or Tiered Deployment
gooseGoose 1.50.0 · seed-h --helpSubcommands:
session,run,configureTaught in:Capstone 5: The Agentic Workstation, Goose and OpenHands: Autonomous Agents and Sandboxing, Lab: One Task, Six Agents, Project: A Local Agentic Coding Workstation, The Landscape: Terminal, Editor and Autonomous Agents
hfHugging Face CLI 1.30.0 · seed-h --help --versionSubcommands:
download,upload,authTaught in:Capstone 1: Hardware and Model Plan, Capstone 6: The Report, Lab: Look Inside a Model, Parameters, Layers and Model Size, Dense, Mixture-of-Experts and Hybrid Architectures, Lab: Build Your Model Shortlist, Lab: Prepare Your Machine, GGUF and Quantisation Types, Installing and Building llama.cpp on Your Platform, Lab: Run and Benchmark the Course Reference Models, Managing a Model Library: Storage, Naming and Versions, Lab: Same Model, Every Engine, Lab: Serve a Model to Twenty Concurrent Users, PyTorch, Transformers, Datasets and the Hugging Face Ecosystem, Merging, Exporting and Quantising a Fine-Tuned Model, Lab: Train and Deploy a Draft for Your Model, Lab: A Mixed-Platform Cluster, Lab: Run a Model Bigger Than Any One Machine, Backup, Upgrades and Reproducibility of a Model Estate, Security for Exposed Endpoints and the Model Supply Chain
iperf3iperf3 3.21 · seed-h --help -s --server -c --client -p --port -t --time -P --parallel -b --bandwidth -u --udp -R --reverse -J --json -i --interval -B --bind -A --affinity -V --verbose -4 -6 -v --version --bitrate --bidir -l --length -w --window -M --set-mss -Z --zerocopy --logfile -D --daemon -1 --one-offTaught in:Lab: Build and Measure Your Cluster Network, Networking for Home Clusters: 2.5 GbE to 200 GbE, Thunderbolt 5 and RDMA
litellmLiteLLM 1.100.0 · seed-h --help --model --port --host --config --debug --num_workers --temperature --max_tokens --drop_params --add_key --test --versionTaught in:Capstone 5: The Agentic Workstation, Capstone 3: Cluster or Tiered Deployment, Capstone 2: The Inference Service, Project: Your Local Model Gateway, Project: A Tiered Inference Architecture, Backup, Upgrades and Reproducibility of a Model Estate, Challenge: The 3 a.m. Out-of-Memory, Lab: Dashboards for Your Cluster, Observability: Metrics, Logs and Traces for LLM Serving, Routing and Model Management: llama-swap, LiteLLM, NGINX and Health Checks, Security for Exposed Endpoints and the Model Supply Chain, Lab: A Minimal Agent from Scratch, Claude Code with a Local Endpoint, Lab: One Task, Six Agents, Project: A Local Agentic Coding Workstation, Project: A Multi-Agent System on Your Cluster, Reality Check: 'Agents Are Just Loops', Lab: Fine-Tune a Small Model on Your Own Agent Trajectories and Measure the Gain
llama-benchllama.cpp v0.4.0 · seed--help -h --version -m --model -p --n-prompt -n --n-gen -pg -b --batch-size -ub --ubatch-size -ngl --n-gpu-layers -t --threads -r --repetitions -o --output --numa -sm --split-mode -ts --tensor-split -fa --flash-attn --mmap --mlock -v --verbose -ctk --cache-type-k -ctv --cache-type-v --rpc -dev --device -rpc --list-devices --no-warmupTaught in:Challenge: The Model That Runs at Two Tokens per Second, Installing and Building llama.cpp on Your Platform, Lab: Run and Benchmark the Course Reference Models, llama.cpp: The Engine That Runs Everywhere, Lab: Quantise Your Fine-Tune Five Ways and Measure Each, Challenge: The Cluster That Is Slower Than One Machine, Lab: A Mixed-Platform Cluster, Lab: Run a Model Bigger Than Any One Machine, llama.cpp RPC: Layers Across Machines, RDMA Transport, Tuning and Measuring the Split
llama-clillama.cpp v0.4.0 · seed--help -h --version -m --model -p --prompt -f --file -n --predict -c --ctx-size -b --batch-size -ngl --n-gpu-layers -t --threads --temp --top-k --top-p --min-p --repeat-penalty -i --interactive --interactive-first -cnv --conversation --color -s --seed --mlock --no-mmap --mirostat --grammar --grammar-file --chat-template -ins --instruct --multiline-input --in-prefix --in-suffix -r --reverse-prompt --log-disable --verbose-prompt --simple-io --lora --control-vector -e --escape --numa -fa --flash-attn --rpc -ts --tensor-split -dev --deviceTaught in:Challenge: The Model That Runs at Two Tokens per Second, Installing and Building llama.cpp on Your Platform, Lab: Run and Benchmark the Course Reference Models, llama-cli and llama-server, llama.cpp: The Engine That Runs Everywhere, Sampling: Temperature, Top-p, Min-p, Repetition and Determinism, Lab: Your First Training Run, Challenge: The Fine-Tune That Got Worse, Lab: Fine-Tune a 1B to 4B Model to Follow Your Format, Merging, Exporting and Quantising a Fine-Tuned Model, Project: A Specialist Assistant, Lab: Sequence-Level Distillation of a 30B-Class Teacher into a 4B Student, Challenge: The Cluster That Is Slower Than One Machine, Lab: Run a Model Bigger Than Any One Machine, llama.cpp RPC: Layers Across Machines, Lab: Fine-Tune a Small Model on Your Own Agent Trajectories and Measure the Gain
llama-imatrixllama.cpp v0.4.0 · seed--help -h -lv --verbosity -m --model -f --file -o --output-file -ofreq --output-frequency --output-format --save-frequency --process-output --in-file --parse-special --chunk --from-chunk --chunks --no-ppl --show-statistics -ngl --n-gpu-layersTaught in: no page yet.
llama-perplexityllama.cpp v0.4.0 · seed--help -h -m --model -f --file -ngl --n-gpu-layers -c --ctx-size --kl-divergence-base --kl-divergenceTaught in: no page yet.
llama-quantizellama.cpp v0.4.0 · seed--help -h --allow-requantize --leave-output-tensor --pure --imatrix --include-weights --exclude-weights --output-tensor-type --token-embedding-type --keep-split --override-kvTaught in:Capstone 4: Improve a Small Model, GGUF and Quantisation Types, llama.cpp: The Engine That Runs Everywhere, Lab: Your First Training Run, Lab: Fine-Tune a 1B to 4B Model to Follow Your Format, Merging, Exporting and Quantising a Fine-Tuned Model, Project: A Specialist Assistant, Lab: Sequence-Level Distillation of a 30B-Class Teacher into a 4B Student, Project: The Distillation Pipeline, Lab: Quantise Your Fine-Tune Five Ways and Measure Each, Post-Training Quantisation in Depth: GPTQ, AWQ, K-quants, imatrix and HQQ, Quantisation-Aware Training and the Blackwell Formats: FP8, NVFP4 and MXFP4, Lab: Fine-Tune a Small Model on Your Own Agent Trajectories and Measure the Gain
llama-serverllama.cpp v0.4.0 · seed--help -h --version -m --model -mu --model-url -hf --hf-repo --hf-file -a --alias -t --threads -tb --threads-batch -c --ctx-size -n --predict -b --batch-size -ub --ubatch-size -np --parallel --cont-batching -ngl --n-gpu-layers -sm --split-mode -ts --tensor-split -mg --main-gpu --mlock --no-mmap --numa --host --port --path --api-key --api-key-file --embedding --embeddings --reranking --metrics --slots --props -to --timeout --chat-template --chat-template-file --jinja -fa --flash-attn --rope-scaling --rope-freq-base --rope-freq-scale --yarn-orig-ctx --yarn-ext-factor --yarn-attn-factor --cache-type-k --cache-type-v --no-context-shift --grp-attn-n --grp-attn-w --log-disable -v --verbose -s --seed --temp --top-k --top-p --min-p --repeat-penalty --repeat-last-n --presence-penalty --frequency-penalty --mirostat --grammar --grammar-file --json-schema --lora --lora-scaled --control-vector -sp --special --verbose-prompt --spec-draft-model -md --spec-draft-n-max --spec-draft-n-min --spec-draft-p-min -ngld --spec-type --cache-prompt --no-cache-prompt --slot-save-path -ctk -ctv --no-slots --no-metrics --rpc -dev --device --gpu-layers --cache-reuse --reasoning-format --reasoning-budgetTaught in:Capstone 3: Cluster or Tiered Deployment, Capstone 4: Improve a Small Model, Capstone 2: The Inference Service, Choosing a Model for a Memory Budget, Challenge: The Model That Runs at Two Tokens per Second, Lab: Run and Benchmark the Course Reference Models, llama-cli and llama-server, llama.cpp: The Engine That Runs Everywhere, Sampling: Temperature, Top-p, Min-p, Repetition and Determinism, Lab: A Private Chat Service for Your Home Network, Reality Check: 'The Default Context Is Enough', Lab: Same Model, Every Engine, AMD-Native: ROCm Builds, Lemonade Server and the NPU Question, Why a Second Kind of Engine: Batching, Paged Attention and Throughput, Installing vLLM: x86 CUDA, DGX Spark, ROCm and What Does Not Work, Lab: Serve a Model to Twenty Concurrent Users, Project: Your Local Model Gateway, Tool Calling and Structured Output on the Server Side, Speculative Decoding: Draft Models, EAGLE and n-gram, Lab: Benchmark Local Models on Your Own Tasks, Local Coding Assistants: Autocomplete and Chat in Your Editor, Vision, Speech and Documents: Multimodal Locally, Privacy, Security and Serving Beyond localhost, Project: A Private Document Question-Answering Service, Prompting That Works Locally: System Prompts, Chat Templates and Thinking Modes, Retrieval-Augmented Generation: Embeddings, Chunking and Vector Stores, Structured Output and JSON Mode, Challenge: The Fine-Tune That Got Worse, Lab: Fine-Tune a 1B to 4B Model to Follow Your Format, Merging, Exporting and Quantising a Fine-Tuned Model, Project: A Specialist Assistant, Lab: DPO a Model to Prefer Your Style, Lab: GRPO on a Maths or Code Task on One Machine, Reality Check: 'RL Makes Small Models Reason', Challenge: The Student That Learned the Teacher's Mistakes, Lab: Logit Distillation with TRL's Distillation Trainers, Lab: Sequence-Level Distillation of a 30B-Class Teacher into a 4B Student, Project: The Distillation Pipeline, Generating Synthetic Data with a Local Teacher, Challenge: The Benchmark That Lied, Evaluation Harnesses: lm-evaluation-harness, lighteval, EvalPlus and Your Own, Lab: Quantise Your Fine-Tune Five Ways and Measure Each, Lab: Run a Standard Benchmark Suite on Your Model, Measuring Quantisation Damage: Perplexity, KL Divergence and Task Evaluations, Lab: Train and Deploy a Draft for Your Model, Prefix Caching and KV Reuse, Speculative Decoding Revisited: Acceptance Rates and When It Pays, Challenge: The Cluster That Is Slower Than One Machine, Lab: A Mixed-Platform Cluster, Lab: Run a Model Bigger Than Any One Machine, llama.cpp RPC: Layers Across Machines, RDMA Transport, Tuning and Measuring the Split, Multi-GPU Desktops: PCIe, Tensor Parallel Without NVLink and Expert Parallel, Lab: KV Cache Offload and Sharing, Lab: Two-Machine Prefill and Decode with vLLM, Project: A Tiered Inference Architecture, vLLM Disaggregated Prefill: Connectors and the Proxy, Challenge: The 3 a.m. Out-of-Memory, Lab: Dashboards for Your Cluster, Observability: Metrics, Logs and Traces for LLM Serving, Routing and Model Management: llama-swap, LiteLLM, NGINX and Health Checks, Security for Exposed Endpoints and the Model Supply Chain, Context Engineering: Memory, Compaction and the KV Budget, Function Calling End to End on Local Engines, Lab: A Minimal Agent from Scratch, Reasoning Models in Agent Loops, Claude Code with a Local Endpoint, Project: A Multi-Agent System on Your Cluster, Reality Check: 'Agents Are Just Loops', Lab: Fine-Tune a Small Model on Your Own Agent Trajectories and Measure the Gain
llama-swapllama-swap v255 · seed-h --help --config --listen --log-level --watch-config --versionTaught in:Capstone 2: The Inference Service, Project: Your Local Model Gateway, Lab: Fine-Tune a 1B to 4B Model to Follow Your Format, Project: A Specialist Assistant, Lab: Sequence-Level Distillation of a 30B-Class Teacher into a 4B Student, Project: The Distillation Pipeline, Backup, Upgrades and Reproducibility of a Model Estate, Challenge: The 3 a.m. Out-of-Memory, Lab: Dashboards for Your Cluster, Observability: Metrics, Logs and Traces for LLM Serving, Routing and Model Management: llama-swap, LiteLLM, NGINX and Health Checks, Project: A Local Agentic Coding Workstation, Project: A Multi-Agent System on Your Cluster, Reality Check: 'Agents Are Just Loops', Lab: Fine-Tune a Small Model on Your Own Agent Trajectories and Measure the Gain
lmsLM Studio 0.4.23 · seed-h --help --versionSubcommands:
get,ls,load,unload,server,psTaught in:LM Studio: GUI, MLX and Headless Serving, Managing a Model Library: Storage, Naming and Versions, Reality Check: 'The Default Context Is Enough', Function Calling End to End on Local Engines, Lab: Write and Connect an MCP Server, OpenAI Codex CLI and OpenCode with Local Models
mlx_lm.cache_promptunpinned · seed--model --prompt --prompt-cache-file --max-kv-size --kv-bits --quantized-kv-startTaught in: no page yet.
mlx_lm.convertmlx-lm 0.31.3 · seed--hf-path --mlx-path -q --quantize --q-bits --q-group-size --dtype --upload-repoTaught in:MLX and mlx-lm: Apple's Native Path, Lab: Quantise Your Fine-Tune Five Ways and Measure Each, Post-Training Quantisation in Depth: GPTQ, AWQ, K-quants, imatrix and HQQ, Quantisation-Aware Training and the Blackwell Formats: FP8, NVFP4 and MXFP4
mlx_lm.fuseunpinned · seed--model --save-path --adapter-path --hf-path --upload-repo --dequantize --export-gguf --gguf-pathTaught in:Lab: Fine-Tune a Small Model on Your Own Agent Trajectories and Measure the Gain
mlx_lm.generatemlx-lm 0.31.3 · seed--model --prompt --max-tokens --temp --top-p --seed --adapter-path --trust-remote-code --colorize --eos-token --draft-model --num-draft-tokens --prompt-cache-file --max-kv-size --kv-bits --kv-group-size --system-prompt --chat-template-config --ignore-chat-template --verbose --quantized-kv-startTaught in:Lab: Look Inside a Model, MLX and mlx-lm: Apple's Native Path, Lab: Your First Training Run, Challenge: The Fine-Tune That Got Worse, Lab: Fine-Tune a 1B to 4B Model to Follow Your Format, Lab: Sequence-Level Distillation of a 30B-Class Teacher into a 4B Student, Lab: Train and Deploy a Draft for Your Model, Prefix Caching and KV Reuse
mlx_lm.loramlx-lm 0.31.3 · seed--model --train --data --batch-size --iters --learning-rate --num-layers --adapter-path --test --resume-adapter-file --fine-tune-type --mask-promptTaught in:Capstone 4: Improve a Small Model, Lab: Your First Training Run, Challenge: The Fine-Tune That Got Worse, Lab: Fine-Tune a 1B to 4B Model to Follow Your Format, Project: A Specialist Assistant, Lab: Sequence-Level Distillation of a 30B-Class Teacher into a 4B Student, Project: The Distillation Pipeline, Lab: Fine-Tune a Small Model on Your Own Agent Trajectories and Measure the Gain
mlx_lm.servermlx-lm 0.31.3 · seed--model --host --port --adapter-path --max-tokens --temp --trust-remote-code --log-level --chat-templateTaught in:Lab: Same Model, Every Engine, MLX and mlx-lm: Apple's Native Path, Lab: Serve a Model to Twenty Concurrent Users, Lab: A Two-Mac Cluster over Thunderbolt 5, mlx.distributed: Ring, MPI and RDMA over Thunderbolt 5, Reality Check: 'Four Mac Studios Replace a GPU Server'
mlx.distributed_configunpinned · seed--verbose --hosts --ignore-unreachable --hostfile --over --output-hostfile --auto-setup --no-auto-setup --dot --backend --envTaught in: no page yet.
mlx.launchmlx-lm 0.31.3 · seed--hosts --backend --hostfile -n --verbose --label --env --print-python --repeat-hosts --mpi-arg --connections-per-ip --starting-port -p --cwd --nccl-port --pythonTaught in:MLX and mlx-lm: Apple's Native Path, Lab: A Two-Mac Cluster over Thunderbolt 5, mlx.distributed: Ring, MPI and RDMA over Thunderbolt 5
nvidia-smiunpinned · seed-h --help -L --list-gpus -q --query --query-gpu --format -l --loop -i --id -pm --persistence-mode --query-compute-apps -d --display --lock-gpu-clocks --lock-memory-clocks -pl --power-limit -v --versionTaught in:Lab: Measure Your Memory Bandwidth and Compute, Lab: Prepare Your Machine, Challenge: The Model That Runs at Two Tokens per Second, Lab: Run and Benchmark the Course Reference Models, Lab: Same Model, Every Engine, TensorRT-LLM, NIM and the DGX Spark Playbooks, Lab: Serve a Model to Twenty Concurrent Users, Lab: Train a 10M to 125M Parameter Model in an Afternoon, Project: A Domain Micro-Model, Lab: Serve a 400B-Class Model on Two DGX Sparks, Multi-GPU Desktops: PCIe, Tensor Parallel Without NVLink and Expert Parallel, TensorRT-LLM and Dynamo on Spark Pairs, Capacity Planning and Cost per Million Tokens at Home, Challenge: The 3 a.m. Out-of-Memory, Lab: Dashboards for Your Cluster, Observability: Metrics, Logs and Traces for LLM Serving
ollamaOllama 0.33.3 · seed-h --help -v --versionSubcommands:
run,pull,serve,create,list,show,ps,rm,cp,pushTaught in:Reality Check: 'A Small Local Model Is as Good as the Frontier', Lab: A Private Chat Service for Your Home Network, Managing a Model Library: Storage, Naming and Versions, Ollama: Models as a Service, Reality Check: 'The Default Context Is Enough', Tool Calling and Structured Output on the Server Side, Function Calling End to End on Local Engines, Reasoning Models in Agent Loops, OpenAI Codex CLI and OpenCode with Local Models
opencodeOpenCode 1.18.29 · seed-h --help -v --version --print-logs --log-level --pureSubcommands:
run,auth,models,serve,mcp,agent,session,upgradeTaught in:Lab: Write and Connect an MCP Server, OpenAI Codex CLI and OpenCode with Local Models, Lab: One Task, Six Agents, Project: A Local Agentic Coding Workstation, The Landscape: Terminal, Editor and Autonomous Agents
python -m sglang.launch_serverSGLang 0.5.19 · seed--model-path --host --port --tp-size --mem-fraction-static --context-length --quantization --trust-remote-code --chat-template --api-key --served-model-name --enable-torch-compile --grammar-backend --tool-call-parser --reasoning-parser --attention-backend --speculative-algorithm --speculative-draft-model-path --speculative-num-steps --speculative-eagle-topk --speculative-num-draft-tokens --disable-radix-cache --max-running-requests --chunked-prefill-size --enable-metrics --log-level --log-requests --kv-cache-dtype --dtype --tensor-parallel-size --model --disaggregation-mode --disaggregation-transfer-backend --disaggregation-bootstrap-port --disaggregation-ib-device --base-gpu-idTaught in:Tool Calling and Structured Output on the Server Side, SGLang: When to Choose It, Speculative Decoding: Draft Models, EAGLE and n-gram, Prefix Caching and KV Reuse, Training a Draft Model: Medusa and EAGLE at Home, SGLang PD Disaggregation and NVIDIA Dynamo
rayunpinned · seed-h --help --versionSubcommands:
start,stop,status,list,up,down,attachTaught in:Lab: Serve a 400B-Class Model on Two DGX Sparks, vLLM Multi-Node with Ray: Tensor Parallel Inside, Pipeline Parallel Across
rocm-smiunpinned · seed-h --help --showhw --showtemp --showuse --showmemuse --showpower --showclocks --setperflevel --setsclk --setmclk -d --device --json --showdriverversion -v --versionTaught in:Lab: Same Model, Every Engine, AMD-Native: ROCm Builds, Lemonade Server and the NPU Question, Lab: Train a 10M to 125M Parameter Model in an Afternoon, Project: A Domain Micro-Model, Capacity Planning and Cost per Million Tokens at Home, Challenge: The 3 a.m. Out-of-Memory, Lab: Dashboards for Your Cluster, Observability: Metrics, Logs and Traces for LLM Serving
rpc-serverllama.cpp v0.4.0 · seed--help -h -H --host -p --port -c --cache -t --threads -d --deviceTaught in: no page yet.
trtllm-serveTensorRT-LLM 1.2.1 · seed-h --help --host --port --tp_size --pp_size --max_batch_size --kv_cache_free_gpu_memory_fraction --backend --max_seq_len --max_num_tokens --kv_cache_dtype --tool_parser --reasoning_parser --served_model_name --log_level --trust_remote_code --extra_llm_api_optionsTaught in:Lab: Same Model, Every Engine, TensorRT-LLM, NIM and the DGX Spark Playbooks, Lab: Serve a 400B-Class Model on Two DGX Sparks, TensorRT-LLM and Dynamo on Spark Pairs
vllmvLLM 0.28.0 · seed-h --help --versionSubcommands:
serve,bench,chat,completeTaught in:Capstone 3: Cluster or Tiered Deployment, Lab: Same Model, Every Engine, Why a Second Kind of Engine: Batching, Paged Attention and Throughput, Installing vLLM: x86 CUDA, DGX Spark, ROCm and What Does Not Work, Lab: Serve a Model to Twenty Concurrent Users, Tool Calling and Structured Output on the Server Side, Serving with vLLM: Quantised Weights, Context, Memory and Multi-GPU, SGLang: When to Choose It, Speculative Decoding: Draft Models, EAGLE and n-gram, Merging, Exporting and Quantising a Fine-Tuned Model, Generating Synthetic Data with a Local Teacher, Evaluation Harnesses: lm-evaluation-harness, lighteval, EvalPlus and Your Own, Lab: Run a Standard Benchmark Suite on Your Model, Lab: Train and Deploy a Draft for Your Model, Prefix Caching and KV Reuse, Speculative Decoding Revisited: Acceptance Rates and When It Pays, Training a Draft Model: Medusa and EAGLE at Home, Lab: Serve a 400B-Class Model on Two DGX Sparks, Multi-GPU Desktops: PCIe, Tensor Parallel Without NVLink and Expert Parallel, vLLM Multi-Node with Ray: Tensor Parallel Inside, Pipeline Parallel Across, Lab: KV Cache Offload and Sharing, Lab: Two-Machine Prefill and Decode with vLLM, Project: A Tiered Inference Architecture, SGLang PD Disaggregation and NVIDIA Dynamo, vLLM Disaggregated Prefill: Connectors and the Proxy, Context Engineering: Memory, Compaction and the KV Budget, Function Calling End to End on Local Engines, Reasoning Models in Agent Loops