Sources and verification
The policy
Section titled “The policy”Nothing about a tool’s flags or behaviour, a machine’s capabilities, a model’s size, licence or capability, or a technique’s results appears in this course without having been checked against an official source or measured on the validation hardware. Where something could not be verified, the course says so rather than filling the gap with a confident guess.
Sources are used in this order of authority:
- The tool’s own documentation and README at the pinned version, and its
--helpoutput captured into the command reference. - Vendor hardware and platform documentation: NVIDIA for the DGX Spark, CUDA and GeForce; AMD for ROCm and the Ryzen AI Max+; Apple for MLX, Metal and macOS; Microsoft for WSL2.
- Model cards on Hugging Face and the publisher’s release notes, plus the licence text itself.
- Papers for techniques: attention, adapters, preference optimisation, reinforcement learning, distillation, speculative decoding, disaggregated serving, cited by title and identifier.
- Measurements made on the validation hardware, recorded with their full context.
Forums, blog posts, videos and social media may aid understanding and may point at a source. They never establish behaviour. Benchmark numbers from third parties are quoted only as “reported by”, with the link and the date, never as the course’s own claim.
What “validated” means here
Section titled “What “validated” means here”Every lab is written from the official documentation of the versions pinned on the what you need page, and then executed on the author’s lab machines before the course is released: two NVIDIA DGX Sparks, a GMKtec EVO-X2, an Apple silicon MacBook Pro and an NVIDIA laptop running Ubuntu. Each lab page states which tracks were run, with the versions and the date, and which were not.
That claim has limits worth stating plainly:
- The Mac used for validation has 24 GB of memory, so Track M pages with a larger floor are written from documentation and marked as such. The two-Mac cluster lab was not run.
- The NVIDIA laptop has 16 GB of VRAM, so the 24 and 32 GB tiers on Track N are documented from the 16 GB run plus arithmetic, and multi-GPU desktops were not run.
- Native Windows and WSL2 commands were checked by inspection against NVIDIA’s and Microsoft’s documentation, not executed.
If you find something that does not behave as described, the official documentation and your own measurement win, and the course is wrong. Please treat it that way and report it.
Versions age fastest
Section titled “Versions age fastest”The engines, trainers and agents this course teaches release weekly. Every page that describes a tool’s behaviour names the version and the date it was checked; the pins are reviewed every quarter, and the build warns when a pin is more than four months old. Where a vendor’s own documentation is out of date, the page says so rather than silently correcting it.
Every source cited
Section titled “Every source cited”The index below is built from the pages themselves: each page records its sources in its own metadata, and this list is assembled from those records at build time. It cannot list a citation the course does not make, and it cannot miss one that it does.
1024 distinct sources are cited across 184 of the course's 184 content pages. Pages that make no checkable claim about a tool, a machine or a model — the conceptual lessons — carry no citations, which is why the counts differ.
adk.dev
- Google ADK — Modelsretrieved 2026-09-09
Sections: Model connectors; self-hosted options
Used by:P26 Agent Frameworks Compared
- Google ADK — vLLMretrieved 2026-09-09
Sections: LiteLlm with a self-hosted endpoint
Used by:P26 Agent Frameworks Compared
ai-act-service-desk.ec.europa.eu
- Timeline for the Implementation of the EU AI Act (AI Act Service Desk)retrieved 2026-09-12
ai.google.dev
- Gemma 4 licence (Apache License 2.0), Google AI for Developersretrieved 2026-09-12
- Gemma Prohibited Use Policyretrieved 2026-09-12
Sections: Last modified 21 February 2024
- Gemma Terms of Useretrieved 2026-09-12
Sections: Scope note and Appendix; 1.1(b), (c), (e); 3.1, 3.2, 3.3; 4.5; 4.6 (last modified 1 April 2026)
aider.chat
- Aider — advanced model settingsretrieved 2026-09-09
Sections: Model settings; search order
Used by:P25 Aider with Local Models
- Aider — edit formatsretrieved 2026-09-09
Used by:P25 Aider with Local Models
- Aider — o1 tops aider's new polyglot leaderboard (how the 225 exercises were chosen)retrieved 2026-09-09, 2026-09-12
Sections: 225 exercises; selection method
Used by:P4 Reading a Model Card and a BenchmarkP25 Aider with Local Models
- Aider — Ollamaretrieved 2026-09-09
Sections: Context window; API base; ollama_chat prefix
Used by:P25 Aider with Local Models
- Aider — OpenAI compatible APIsretrieved 2026-09-09
Used by:P25 Aider with Local Models
- Aider — options referenceretrieved 2026-09-09
Sections: message; test-cmd; auto-test; yes-always
Used by:P25 Aider with Local ModelsP25 Lab: One Task, Six Agents
- Aider LLM leaderboardsretrieved 2026-09-09, 2026-09-12
Sections: Polyglot benchmark; Polyglot benchmark; last updated; Polyglot benchmark; entries; last updated
Used by:P4 Reading a Model Card and a BenchmarkP16 Evaluation Harnesses: lm-evaluation-harness, lighteval, EvalPlus and Your OwnP25 Aider with Local ModelsP25 Which Local Models Can Actually Drive an Agent
alexgarcia.xyz
- sqlite-vec — a vector search SQLite extensionretrieved 2026-09-08
Sections: Overview; vec0 virtual table
Used by:P10 Project: A Private Document Question-Answering ServiceP10 Retrieval-Augmented Generation: Embeddings, Chunking and Vector Stores
- sqlite-vec — KNN queriesretrieved 2026-09-08
Sections: MATCH and k; distance metrics; MATCH and k; distance metrics; joining to source rows
Used by:P10 Project: A Private Document Question-Answering ServiceP10 Retrieval-Augmented Generation: Embeddings, Chunking and Vector Stores
- sqlite-vec — Pythonretrieved 2026-09-08
Sections: Installation; loading the extension; Installation; loading the extension; serialize_float32
Used by:P10 Project: A Private Document Question-Answering ServiceP10 Retrieval-Augmented Generation: Embeddings, Chunking and Vector Stores
allenai.org
- Ai2 — Olmoretrieved 2026-09-08
anthropic.com
- Anthropic — Building effective agentsretrieved 2026-09-09
Sections: Agents versus workflows; agent-computer interfaces; Agents versus workflows; the augmented LLM; agent-computer interfaces; Routing; when to add complexity; Routing; parallelisation; orchestrator-workers; evaluator-optimiser; when to add complexity; When to add complexity
Used by:P24 Lab: A Minimal Agent from ScratchP24 What an Agent Is: The Loop, Tools and StateP26 Agentic Retrieval and Research AgentsP26 Multi-Agent Patterns: Router, Planner and Executor, Critic, Parallel Fan-OutP26 Reality Check: 'Agents Are Just Loops'
- Anthropic — Effective context engineering for AI agentsretrieved 2026-09-09
Sections: Context as a finite resource; compaction; note-taking; just-in-time retrieval; tool design; Context as a finite resource; the attention budget
Used by:P24 Context Engineering: Memory, Compaction and the KV BudgetP24 Reasoning Models in Agent Loops
apache.org
- Apache License, Version 2.0retrieved 2026-09-12
Sections: Section 1 (Object form, Derivative Works); sections 2, 3, 4 and 6
api.ngc.nvidia.com
- NVIDIA NGC — tensorrt-llm/release image list (compressed sizes per architecture)retrieved 2026-09-13
Sections: tag 1.3.0rc13, arm64 compressedSize
Used by:P8 Lab: Same Model, Every Engine
apple.com
- Apple Mac mini technical specificationsretrieved 2026-09-09
Sections: Chip; Memory; Connectivity; Chip; Memory; Connectivity
Used by:P5 Apple Silicon: Unified Memory, Metal and MLXP5 Why Memory, Not FLOPS, Decides What You Can RunP18 Networking for Home Clusters: 2.5 GbE to 200 GbE, Thunderbolt 5 and RDMA
- Apple Mac Studio technical specificationsretrieved 2026-09-09, 2026-09-13
Sections: Chip; Memory; Connectivity; memory bandwidth per chip; no memory speed or width published; Chip; Memory; Connectivity; Memory; Connectivity
Used by:P5 Apple Silicon: Unified Memory, Metal and MLXP5 Lab: Measure Your Memory Bandwidth and ComputeP5 Why Memory, Not FLOPS, Decides What You Can RunP18 Networking for Home Clusters: 2.5 GbE to 200 GbE, Thunderbolt 5 and RDMAP21 mlx.distributed: Ring, MPI and RDMA over Thunderbolt 5P21 Reality Check: 'Four Mac Studios Replace a GPU Server'
ar5iv.labs.arxiv.org
- Attention Is All You Need — full text (ar5iv rendering)retrieved 2026-09-12
Sections: 3.1 Encoder and Decoder Stacks; 3.2.1 Scaled Dot-Product Attention; 3.2.2 Multi-Head Attention; 3.3 Position-wise Feed-Forward Networks; 3.4 Embeddings and Softmax; 3.5 Positional Encoding
Used by:P2 Attention and the Transformer
arena.ai
- LMArenaretrieved 2026-09-12
arxiv.org
- A General Theoretical Paradigm to Understand Learning from Human Preferences (Azar et al., arXiv:2310.12036)retrieved 2026-09-09
Sections: Abstract
- A Simple and Effective Pruning Approach for Large Language Models (Sun, Liu, Bair and Kolter, arXiv:2306.11695)retrieved 2026-09-09, 2026-09-12
Sections: Abstract; 3 Wanda, Pruning by Weights and Activations (the pruning metric and its squared form; Structured N:M Sparsity); Abstract
Used by:P3 Distillation, Pruning and Quantisation: How Small Models Get GoodP15 Pruning and Compression: Prune-and-Distil
- Accelerating Large Language Model Decoding with Speculative Samplingretrieved 2026-09-09
Sections: Abstract; modified rejection sampling
Used by:P17 Speculative Decoding Revisited: Acceptance Rates and When It Pays
- Adam: A Method for Stochastic Optimizationretrieved 2026-09-09
Used by:P11 Memory Arithmetic for Training: Weights, Gradients, Optimiser States and Activations
- AgentBench: Evaluating LLMs as Agentsretrieved 2026-09-09
Sections: Abstract; eight environments; failure analysis; Open-source versus API models; failure analysis
Used by:P26 Evaluating Agents: Trajectories, Success Rates and CostP26 Reality Check: 'Agents Are Just Loops'
- An Empirical Investigation of Catastrophic Forgetting in Gradient-Based Neural Networks (Goodfellow et al., arXiv:1312.6211)retrieved 2026-09-09
Sections: Abstract
- An Empirical Study of Catastrophic Forgetting in Large Language Models During Continual Fine-tuning (Luo et al., arXiv:2308.08747)retrieved 2026-09-09
Sections: Abstract; findings
Used by:P13 Challenge: The Fine-Tune That Got WorseP13 What Fine-Tuning Changes and What It Cannot
- An Empirical Study of Qwen3 Quantization (arXiv:2505.02214)retrieved 2026-09-12
Sections: Abstract
- Attention Is All You Need (Vaswani et al., arXiv:1706.03762)retrieved 2026-09-08
Sections: Abstract
Used by:P2 Attention and the Transformer
- AWQ: Activation-aware Weight Quantization for LLM Compression and Acceleration (arXiv:2306.00978)retrieved 2026-09-09, 2026-09-12
Sections: Abstract; 3.1 Preserving 1% Salient Weights; 3.2 Activation-aware Scaling (equations 1, 2, 4 and 5); 5.1 Setups; Abstract
Used by:P3 Distillation, Pruning and Quantisation: How Small Models Get GoodP16 Measuring Quantisation Damage: Perplexity, KL Divergence and Task EvaluationsP16 Post-Training Quantisation in Depth: GPTQ, AWQ, K-quants, imatrix and HQQ
- Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preferenceretrieved 2026-09-12
- Compact Language Models via Pruning and Knowledge Distillation (Muralidharan et al., arXiv:2407.14679)retrieved 2026-09-09, 2026-09-12
Sections: Abstract; 2.2 Importance Analysis (Width); 2.3 Obtaining a Pruned Model; Abstract
Used by:P3 Distillation, Pruning and Quantisation: How Small Models Get GoodP15 Pruning and Compression: Prune-and-Distil
- DAPO: An Open-Source LLM Reinforcement Learning System at Scale (Yu et al., arXiv:2503.14476)retrieved 2026-09-09
Sections: Abstract; Clip-Higher; Dynamic Sampling; Token-Level Policy Gradient Loss; Overlong Reward Shaping; Abstract; Dynamic Sampling
Used by:P14 Reinforcement Learning with Verifiable Rewards: GRPO ExplainedP14 Reality Check: 'RL Makes Small Models Reason'
- DataComp-LM: In search of the next generation of training sets for language models (arXiv:2406.11794)retrieved 2026-09-09
Sections: Abstract
Used by:P12 Anatomy of a Training Loop: nanochat Line by Line
- Deduplicating Training Data Makes Language Models Better (Lee et al., arXiv:2107.06499)retrieved 2026-09-12
Sections: Abstract
- DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning (arXiv:2501.12948)retrieved 2026-09-09, 2026-09-12
Sections: Abstract; B.3.3 800K Supervised Data; B.4.3 Hyper-Parameters of Distillation (Table 6); Abstract
Used by:P3 Distillation, Pruning and Quantisation: How Small Models Get GoodP14 Reinforcement Learning with Verifiable Rewards: GRPO ExplainedP14 Reality Check: 'RL Makes Small Models Reason'P15 Reasoning Distillation: Teacher Traces as Training DataP15 Why Distillation Works
- DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning, first version (arXiv:2501.12948v1)retrieved 2026-09-12
Sections: 2.2.2 Reward Modeling; 2.3.1 Cold Start to 2.3.4 Reinforcement Learning for all Scenarios; 2.4 Distillation; 4.1 Distillation v.s. Reinforcement Learning
Used by:P3 Post-Training: SFT, Preference Tuning and Reinforcement Learning
- DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning, revised version (arXiv:2501.12948v2)retrieved 2026-09-12
Sections: Abstract; 1 Introduction (R1-Zero readability and language mixing)
Used by:P3 Post-Training: SFT, Preference Tuning and Reinforcement Learning
- DeepSeek-V3 Technical Reportretrieved 2026-09-12
Sections: 2.1.2 DeepSeekMoE with Auxiliary-Loss-Free Load Balancing
Used by:P4 Dense, Mixture-of-Experts and Hybrid Architectures
- DeepSeek-V3 Technical Report (arXiv 2412.19437), section 2.1.2, DeepSeekMoE with auxiliary-loss-free load balancingretrieved 2026-09-12
- DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models (Shao et al., arXiv:2402.03300)retrieved 2026-09-09
Sections: Abstract
Used by:P14 From Imitation to Preferences: Why Ranking Beats CopyingP14 Reinforcement Learning with Verifiable Rewards: GRPO Explained
- DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models (Shao et al., arXiv:2402.03300v3)retrieved 2026-09-12
Sections: 4.1 Group Relative Policy Optimization (advantage, objective, KL added to the loss)
Used by:P3 Post-Training: SFT, Preference Tuning and Reinforcement Learning
- Defining and Characterizing Reward Hacking (Skalse et al., arXiv:2209.13085)retrieved 2026-09-09
Sections: Abstract
Used by:P14 From Imitation to Preferences: Why Ranking Beats CopyingP14 Reward Functions: Maths, Code Tests, Format and Length
- Direct Preference Optimization: Your Language Model is Secretly a Reward Model (Rafailov et al., arXiv:2305.18290)retrieved 2026-09-09
Sections: Abstract
Used by:P14 DPO and Its Family: IPO, KTO, ORPO and SimPOP14 From Imitation to Preferences: Why Ranking Beats Copying
- Direct Preference Optimization: Your Language Model is Secretly a Reward Model (Rafailov et al., arXiv:2305.18290v3)retrieved 2026-09-12
Sections: Abstract; 3 Preliminaries (beta); 4 Direct Preference Optimization (Equation 7, implicit reward, gradient)
Used by:P3 Post-Training: SFT, Preference Tuning and Reinforcement Learning
- Distilling the Knowledge in a Neural Network (Hinton, Vinyals and Dean, arXiv:1503.02531)retrieved 2026-09-09, 2026-09-12
Sections: Abstract; 1 Introduction; 2 Distillation (equations 1 and 2, the T-squared scaling); 2 Distillation; Introduction; 2 Distillation
Used by:P3 Distillation, Pruning and Quantisation: How Small Models Get GoodP15 Challenge: The Student That Learned the Teacher's MistakesP15 Three Kinds of Distillation: Logit, Sequence and On-PolicyP15 Why Distillation Works
- DistServe: Disaggregating Prefill and Decoding for Goodput-optimized Large Language Model Servingretrieved 2026-09-09
Sections: Abstract
- DoRA: Weight-Decomposed Low-Rank Adaptation (Liu et al., arXiv:2402.09353)retrieved 2026-09-09
Sections: Abstract
Used by:P13 LoRA and QLoRA Explained
- EAGLE-2: Faster Inference of Language Models with Dynamic Draft Treesretrieved 2026-09-09
Sections: Abstract; context-dependent acceptance
Used by:P17 Speculative Decoding Revisited: Acceptance Rates and When It PaysP17 Training a Draft Model: Medusa and EAGLE at Home
- EAGLE-3: Scaling up Inference Acceleration of Large Language Models via Training-Time Testretrieved 2026-09-09
Sections: Abstract; throughput at batch size 64; Abstract; multi-layer feature fusion; scaling with data
Used by:P17 Speculative Decoding Revisited: Acceptance Rates and When It PaysP17 Training a Draft Model: Medusa and EAGLE at Home
- EAGLE: Speculative Sampling Requires Rethinking Feature Uncertaintyretrieved 2026-09-09
Sections: Abstract; results; Abstract; feature-level autoregression
Used by:P9 Speculative Decoding: Draft Models, EAGLE and n-gramP17 Training a Draft Model: Medusa and EAGLE at Home
- Efficient Large-Scale Language Model Training on GPU Clusters Using Megatron-LMretrieved 2026-09-09
Sections: Abstract
Used by:P18 Tensor, Pipeline, Expert, Data and Sequence ParallelismP18 Why One Box Runs Out: The Memory Wall and the Bandwidth Wall
- Efficient Memory Management for Large Language Model Serving with PagedAttention (arXiv:2309.06180)retrieved 2026-09-09, 2026-09-12
Sections: Abstract; Abstract; problem statement
Used by:P3 Inference: Prefill, Decode and Why Memory Bandwidth RulesP9 Why a Second Kind of Engine: Batching, Paged Attention and Throughput
- Evaluating Large Language Models Trained on Code (pass@k estimator, arXiv:2107.03374)retrieved 2026-09-12
- Fast Inference from Transformers via Speculative Decodingretrieved 2026-09-09
Sections: Abstract; the speculative sampling method
Used by:P17 Speculative Decoding Revisited: Acceptance Rates and When It Pays
- Fine-Tuning or Retrieval? Comparing Knowledge Injection in LLMs (Ovadia et al., arXiv:2312.05934)retrieved 2026-09-09
Sections: Abstract
Used by:P13 Building an SFT DatasetP13 Project: A Specialist AssistantP13 What Fine-Tuning Changes and What It Cannot
- FlashAttention: Fast and Memory-Efficient Exact Attention with IO-Awareness (Dao et al., arXiv:2205.14135)retrieved 2026-09-12
Sections: Abstract
Used by:P2 Attention and the Transformer
- FP8 Formats for Deep Learning (Micikevicius et al., arXiv:2209.05433)retrieved 2026-09-12
Sections: Abstract
Used by:P1 Tensors, GPUs and Precision: Why Matrix Multiplication Is Everything
- Gated Delta Networks: Improving Mamba2 with Delta Ruleretrieved 2026-09-12
Used by:P4 Dense, Mixture-of-Experts and Hybrid Architectures
- GLU Variants Improve Transformer (Shazeer, arXiv:2002.05202)retrieved 2026-09-12
Sections: Abstract
Used by:P2 Attention and the Transformer
- GPQA: A Graduate-Level Google-Proof Q&A Benchmarkretrieved 2026-09-09, 2026-09-12
Sections: Abstract; canary string; Table 5 set sizes; Abstract
Used by:P4 Reading a Model Card and a BenchmarkP16 Evaluation Harnesses: lm-evaluation-harness, lighteval, EvalPlus and Your Own
- gpt-oss-120b & gpt-oss-20b Model Card (arXiv:2508.10925)retrieved 2026-09-12
Sections: Table 1; 2.6 Evaluation; Figure 1; Table 3; 5.2.3.1 SWE-bench Verified
Used by:P2 Parameters, Layers and Model SizeP4 Reading a Model Card and a Benchmark
- GPTQ: Accurate Post-Training Quantization for Generative Pre-trained Transformers (Frantar et al., arXiv:2210.17323)retrieved 2026-09-09
Sections: Abstract
Used by:P16 Post-Training Quantisation in Depth: GPTQ, AWQ, K-quants, imatrix and HQQ
- GQA: Training Generalized Multi-Query Transformer Models from Multi-Head Checkpoints (Ainslie et al., arXiv:2305.13245)retrieved 2026-09-08, 2026-09-12
Sections: Abstract
Used by:P2 Attention and the TransformerP4 Choosing a Model for a Memory Budget
- Group Sequence Policy Optimization (Zheng et al., arXiv:2507.18071)retrieved 2026-09-09
Sections: Abstract
Used by:P14 Reinforcement Learning with Verifiable Rewards: GRPO Explained
- GShard: Scaling Giant Models with Conditional Computation and Automatic Shardingretrieved 2026-09-09
Sections: Abstract
Used by:P18 Tensor, Pipeline, Expert, Data and Sequence Parallelism
- He, Zhang, Ren and Sun — Deep Residual Learning for Image Recognition (arXiv:1512.03385)retrieved 2026-09-12
- He, Zhang, Ren and Sun — Delving Deep into Rectifiers (arXiv:1502.01852)retrieved 2026-09-12
- Instruction-Following Evaluation for Large Language Models (Zhou et al., arXiv:2311.07911)retrieved 2026-09-09
Sections: Abstract; verifiable instructions
Used by:P16 Evaluation Harnesses: lm-evaluation-harness, lighteval, EvalPlus and Your OwnP16 Lab: Run a Standard Benchmark Suite on Your Model
- Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena (Zheng et al., arXiv:2306.05685)retrieved 2026-09-08, 2026-09-09
Sections: Abstract; biases of LLM judges; Abstract; biases of LLM judges; agreement with human preferences
Used by:P10 Lab: Benchmark Local Models on Your Own TasksP16 LLM-as-Judge, Contamination and Honest Reporting
- KTO: Model Alignment as Prospect Theoretic Optimization (Ethayarajh et al., arXiv:2402.01306)retrieved 2026-09-09
Sections: Abstract
- Language Models are Few-Shot Learners (Brown et al., arXiv:2005.14165v4)retrieved 2026-09-12
Sections: Table 2.1 caption (300 billion training tokens)
Used by:P3 Post-Training: SFT, Preference Tuning and Reinforcement Learning
- LIMA: Less Is More for Alignment (Zhou et al., arXiv:2305.11206)retrieved 2026-09-09
Sections: Abstract
Used by:P13 Building an SFT DatasetP13 What Fine-Tuning Changes and What It Cannot
- LiveCodeBench: Holistic and Contamination Free Evaluation of Large Language Models for Coderetrieved 2026-09-09, 2026-09-12
Sections: 5.1 Avoiding Contamination
Used by:P4 Reading a Model Card and a BenchmarkP16 LLM-as-Judge, Contamination and Honest Reporting
- LLM Pruning and Distillation in Practice: The Minitron Approach (Sreenivas et al., arXiv:2408.11796)retrieved 2026-09-09
Sections: Abstract
- LLM-QAT: Data-Free Quantization Aware Training for Large Language Models (Liu et al., arXiv:2305.17888)retrieved 2026-09-09, 2026-09-12
Sections: Abstract
Used by:P3 Distillation, Pruning and Quantisation: How Small Models Get GoodP16 Quantisation-Aware Training and the Blackwell Formats: FP8, NVFP4 and MXFP4
- LoRA: Low-Rank Adaptation of Large Language Models (Hu et al., arXiv:2106.09685)retrieved 2026-09-09
Sections: Abstract
Used by:P13 LoRA and QLoRA Explained
- Mamba: Linear-Time Sequence Modeling with Selective State Spacesretrieved 2026-09-08
Used by:P4 Dense, Mixture-of-Experts and Hybrid Architectures
- Medusa: Simple LLM Inference Acceleration Framework with Multiple Decoding Headsretrieved 2026-09-09
Sections: Medusa-1; self-distillation; Abstract; Medusa-1 and Medusa-2; tree attention; self-distillation
Used by:P17 Lab: Train and Deploy a Draft for Your ModelP17 Training a Draft Model: Medusa and EAGLE at Home
- Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelismretrieved 2026-09-09
Sections: Abstract
Used by:P18 Tensor, Pipeline, Expert, Data and Sequence Parallelism
- Microscaling Data Formats for Deep Learning (Rouhani et al., arXiv:2310.10537)retrieved 2026-09-09
Sections: Abstract
Used by:P16 Quantisation-Aware Training and the Blackwell Formats: FP8, NVFP4 and MXFP4
- MiniLLM: Knowledge Distillation of Large Language Models (Gu, Dong, Wei and Huang, arXiv:2306.08543)retrieved 2026-09-09
Sections: Abstract
Used by:P15 Three Kinds of Distillation: Logit, Sequence and On-PolicyP15 Why Distillation Works
- MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmarkretrieved 2026-09-09, 2026-09-12
Sections: Abstract
Used by:P4 Reading a Model Card and a BenchmarkP16 Evaluation Harnesses: lm-evaluation-harness, lighteval, EvalPlus and Your Own
- Model Cards for Model Reportingretrieved 2026-09-08
- Mooncake: A KVCache-centric Disaggregated Architecture for LLM Servingretrieved 2026-09-09
Sections: Abstract
Used by:P22 The KV Cache as a Transferable Object: Mooncake, LMCache and NIXL
- On-Policy Distillation of Language Models: Learning from Self-Generated Mistakes (Agarwal et al., arXiv:2306.13649)retrieved 2026-09-09
Sections: Abstract
Used by:P15 Lab: Logit Distillation with TRL's Distillation TrainersP15 Three Kinds of Distillation: Logit, Sequence and On-PolicyP15 Why Distillation Works
- ORPO: Monolithic Preference Optimization without Reference Model (Hong et al., arXiv:2403.07691)retrieved 2026-09-09
Sections: Abstract
- PaLM: Scaling Language Modeling with Pathways (arXiv:2204.02311)retrieved 2026-09-09
Sections: Abstract
- Power Lines: Scaling Laws for Weight Decay and Batch Size in LLM Pre-training (arXiv:2505.13738)retrieved 2026-09-09
Sections: Abstract
Used by:P12 Anatomy of a Training Loop: nanochat Line by Line
- Proximal Policy Optimization Algorithms (Schulman et al., arXiv:1707.06347)retrieved 2026-09-09
Sections: Abstract
Used by:P14 From Imitation to Preferences: Why Ranking Beats CopyingP14 Reinforcement Learning with Verifiable Rewards: GRPO Explained
- QLoRA: Efficient Finetuning of Quantized LLMsretrieved 2026-09-09
Sections: Abstract
Used by:P11 Memory Arithmetic for Training: Weights, Gradients, Optimiser States and ActivationsP13 LoRA and QLoRA ExplainedP13 Project: A Specialist Assistant
- Qwen3 Technical Report (arXiv:2505.09388)retrieved 2026-09-12
Sections: 4 Post-training (Strong-to-Weak Distillation); 4.5 Strong-to-Weak Distillation; 3.1 Pre-training Data (36 trillion tokens, all Qwen3 models); 3.1 Pre-training Data (36 trillion tokens, 119 languages); 3.2 Pre-training Stage (three stages and their token counts); 4.6 Post-training Evaluation; Tables 17 and 18
Used by:P3 Distillation, Pruning and Quantisation: How Small Models Get GoodP3 Open Weights, Open Source and LicencesP3 Pretraining: Learning from Trillions of TokensP4 Reading a Model Card and a Benchmark
- Qwen3 Technical Report (arXiv:2505.09388)retrieved 2026-09-12
Sections: Pre-training data - 36 trillion tokens for all Qwen3 models
- Qwen3 Technical Report (arXiv:2505.09388v1)retrieved 2026-09-12
Sections: 4 Post-training (Long-CoT Cold Start, Reasoning RL, Thinking Mode Fusion, General RL); 4.5 Strong-to-Weak Distillation; 4.7 Discussion; 4.3 Thinking Budget; 4.5 Strong-to-Weak Distillation
Used by:P3 Post-Training: SFT, Preference Tuning and Reinforcement LearningP4 Base, Instruct, Thinking, Coder, Vision, Embedding: Reading a Model Name
- Rethinking Benchmark and Contamination for Language Models with Rephrased Samples (Yang et al., arXiv:2311.04850)retrieved 2026-09-09
Sections: Abstract; Abstract; rephrased samples; overlap findings
Used by:P16 Challenge: The Benchmark That LiedP16 LLM-as-Judge, Contamination and Honest Reporting
- Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks (Lewis et al., arXiv:2005.11401)retrieved 2026-09-08, 2026-09-13
Sections: Abstract
Used by:P10 Retrieval-Augmented Generation: Embeddings, Chunking and Vector StoresAI Problem-Solving Map
- Ring Attention with Blockwise Transformers for Near-Infinite Contextretrieved 2026-09-09
Sections: Abstract
Used by:P18 Tensor, Pipeline, Expert, Data and Sequence Parallelism
- RoFormer: Enhanced Transformer with Rotary Position Embedding (Su et al., arXiv:2104.09864)retrieved 2026-09-08
Sections: Abstract
Used by:P2 Attention and the Transformer
- Scaling Data-Constrained Language Models (Muennighoff et al., arXiv:2305.16264)retrieved 2026-09-12
Sections: Abstract
- Scaling Laws for Neural Language Models (Kaplan et al., arXiv:2001.08361)retrieved 2026-09-12
Sections: Abstract; 1.2 Summary of Scaling Laws (equations 1.1 to 1.3); notation (C ≈ 6NBS, PF-days); 2.1 Parameter and Compute Scaling (Table 1, C ≈ 6N per token)
- Scaling Laws for Precision (arXiv:2411.04330)retrieved 2026-09-12
Sections: Abstract
- Self-Instruct: Aligning Language Models with Self-Generated Instructions (Wang et al., arXiv:2212.10560)retrieved 2026-09-12
Sections: Filtering and Postprocessing; statistics of the generated data
Used by:P3 Distillation, Pruning and Quantisation: How Small Models Get Good
- Self-RAG: Learning to Retrieve, Generate, and Critique through Self-Reflectionretrieved 2026-09-09
Sections: Abstract; reflection tokens; retrieval on demand
- Sequence-Level Knowledge Distillation (Kim and Rush, arXiv:1606.07947)retrieved 2026-09-09
Sections: Abstract
Used by:P15 Three Kinds of Distillation: Logit, Sequence and On-Policy
- SGLang: Efficient Execution of Structured Language Model Programsretrieved 2026-09-09
Sections: Abstract; RadixAttention; Abstract; RadixAttention; compressed finite state machines
Used by:P9 Why a Second Kind of Engine: Batching, Paged Attention and ThroughputP9 SGLang: When to Choose It
- ShortGPT: Layers in Large Language Models are More Redundant Than You Expect (Men et al., arXiv:2403.03853)retrieved 2026-09-09, 2026-09-12
Sections: Abstract; Block Influence (definition); Layer Removal; Limitations (generative against multiple-choice tasks); Abstract
Used by:P3 Distillation, Pruning and Quantisation: How Small Models Get GoodP15 Pruning and Compression: Prune-and-Distil
- SimPO: Simple Preference Optimization with a Reference-Free Reward (Meng et al., arXiv:2405.14734)retrieved 2026-09-09
Sections: Abstract
- SmoothQuant: Accurate and Efficient Post-Training Quantization for Large Language Models (Xiao et al., arXiv:2211.10438)retrieved 2026-09-09
Sections: Abstract
Used by:P16 Post-Training Quantisation in Depth: GPTQ, AWQ, K-quants, imatrix and HQQ
- SparseGPT: Massive Language Models Can Be Accurately Pruned in One-Shot (Frantar and Alistarh, arXiv:2301.00774)retrieved 2026-09-09
Sections: Abstract
- Splitwise: Efficient Generative LLM Inference Using Phase Splittingretrieved 2026-09-09
Sections: Abstract
- SWE-Gym: Training Software Engineering Agents and Verifiersretrieved 2026-09-09
- SWE-smith: Scaling Data for Software Engineering Agentsretrieved 2026-09-09
Used by:P27 Reinforcement Learning on Agent Tasks: Tests as Rewards
- Switch Transformers: Scaling to Trillion Parameter Models with Simple and Efficient Sparsityretrieved 2026-09-09, 2026-09-12
Sections: Abstract
Used by:P4 Dense, Mixture-of-Experts and Hybrid ArchitecturesP18 Tensor, Pipeline, Expert, Data and Sequence Parallelism
- The case for 4-bit precision: k-bit Inference Scaling Laws (arXiv:2212.09720)retrieved 2026-09-12
Sections: Abstract
- The Curse of Recursion: Training on Generated Data Makes Models Forget (Shumailov et al., arXiv:2305.17493)retrieved 2026-09-12
Sections: Abstract
Used by:P3 Distillation, Pruning and Quantisation: How Small Models Get Good
- The FineWeb Datasets: Decanting the Web for the Finest Text Data at Scale (arXiv:2406.17557)retrieved 2026-09-12
Sections: Abstract; 3.2 to 3.7 (extraction, language filter, MassiveText filters, MinHash parameters, per-snapshot deduplication, C4 and custom filters, PII); 4 FineWeb-Edu
- The Leaderboard Illusion (arXiv:2504.20879)retrieved 2026-09-12
- The Llama 3 Herd of Models (arXiv 2407.21783)retrieved 2026-09-12
Sections: 3.2 Model Architecture, tokenizer and vocabulary; 1 Introduction (3.8 × 10^25 FLOPs); 3.1.1 Web Data Curation (URL, document and line-level deduplication); 3.1.2 Data Mix; 3.2 (405B on 15.6T tokens); 3.2.1 Scaling Laws (402B on 16.55T); 3.3.1 and 3.3.2 (16K H100 GPUs, BF16 MFU 38 to 43 per cent); 3.3.4 Reliability (466 interruptions in 54 days); 3.4.1 Initial Pre-Training; 3.4.3 Annealing
Used by:P2 Tokens, Tokenisers and VocabularyP3 Pretraining: Learning from Trillions of Tokens
- ToolLLM: Facilitating Large Language Models to Master 16000+ Real-world APIsretrieved 2026-09-09
- Training Compute-Optimal Large Language Models (Hoffmann et al., arXiv:2203.15556)retrieved 2026-09-09, 2026-09-12
Sections: Abstract; 1 Introduction (Chinchilla, 1.4 trillion tokens, inference cost); Table 1 (Gopher 280B on 300B tokens); 3.3 Approach 3 and equation 10 (E, A, B, alpha, beta); Table 2 (exponents a and b); Table 3 (optimal FLOPs and tokens by model size); Abstract
Used by:P3 Pretraining: Learning from Trillions of TokensP12 Data for Pretraining: FineWeb-Edu, Cosmopedia and Training a TokeniserP12 Project: A Domain Micro-ModelP12 Scaling Laws and Compute Budgets at Home
- Training Deep Nets with Sublinear Memory Costretrieved 2026-09-09
Used by:P11 Memory Arithmetic for Training: Weights, Gradients, Optimiser States and Activations
- Training language models to follow instructions with human feedback (Ouyang et al., arXiv:2203.02155)retrieved 2026-09-09
Sections: Abstract
Used by:P14 From Imitation to Preferences: Why Ranking Beats Copying
- Training language models to follow instructions with human feedback (Ouyang et al., arXiv:2203.02155v1)retrieved 2026-09-12
Sections: Abstract; 3.2 Dataset; 3.5 Models (SFT, 6B reward models, reward model loss, PPO-ptx objective); 4.2 Results on public NLP datasets (alignment tax)
Used by:P3 Post-Training: SFT, Preference Tuning and Reinforcement Learning
- Understanding deep learning requires rethinking generalization (Zhang, Bengio, Hardt, Recht and Vinyals, arXiv:1611.03530)retrieved 2026-09-12
Sections: Abstract
Used by:P1 Generalisation: Train, Validation, Test and Overfitting
- Understanding R1-Zero-Like Training: A Critical Perspective (Liu et al., arXiv:2503.20783)retrieved 2026-09-09
Sections: Abstract; optimisation bias in GRPO
Used by:P14 Reinforcement Learning with Verifiable Rewards: GRPO ExplainedP14 Lab: GRPO on a Maths or Code Task on One MachineP14 Reality Check: 'RL Makes Small Models Reason'
- τ-bench: A Benchmark for Tool-Agent-User Interaction in Real-World Domainsretrieved 2026-09-09
Sections: Abstract; database-state evaluation; pass^k; Abstract; pass^k; pass^k; consistency across trials
Used by:P26 Evaluating Agents: Trajectories, Success Rates and CostP26 Multi-Agent Patterns: Router, Planner and Executor, Critic, Parallel Fan-OutP26 Reality Check: 'Agents Are Just Loops'
brew.sh
- Homebrew — home pageretrieved 2026-09-13
Sections: install command; "The script explains what it will do and then pauses before it does it."
Used by:P5 Lab: Prepare Your Machine
build.nvidia.com
- DGX Spark playbook — Connect Multiple DGX Spark through a Switchretrieved 2026-09-09
- DGX Spark playbook — Connect Two Sparksretrieved 2026-09-09
- DGX Spark playbook — Connect Two Sparks, Run on Two Sparksretrieved 2026-09-09
Used by:P20 Connecting Two DGX Sparks over ConnectX-7P20 Lab: Serve a 400B-Class Model on Two DGX Sparks
- DGX Spark playbook — Connect Two Sparks, Troubleshootingretrieved 2026-09-09
- DGX Spark playbook — NCCL for Multiple Sparksretrieved 2026-09-09
- DGX Spark playbook — NCCL for Multiple Sparks, Run on two Sparksretrieved 2026-09-09
- DGX Spark playbook — NIM for LLMsretrieved 2026-09-09
- DGX Spark playbook — Serve LLMs with vLLM, Multi-node servingretrieved 2026-09-09
Used by:P20 Connecting Two DGX Sparks over ConnectX-7P20 Lab: Serve a 400B-Class Model on Two DGX SparksP20 vLLM Multi-Node with Ray: Tensor Parallel Inside, Pipeline Parallel Across
- DGX Spark playbook — Serve LLMs with vLLM, Troubleshootingretrieved 2026-09-09
Used by:P20 Lab: Serve a 400B-Class Model on Two DGX SparksP20 vLLM Multi-Node with Ray: Tensor Parallel Inside, Pipeline Parallel Across
- DGX Spark playbook — TRT LLM for Inferenceretrieved 2026-09-09, 2026-09-13
Sections: Model Support Matrix; Prerequisites; Model Support Matrix; Time and risk
Used by:P8 Lab: Same Model, Every EngineP8 TensorRT-LLM, NIM and the DGX Spark PlaybooksP20 Lab: Serve a 400B-Class Model on Two DGX SparksP20 TensorRT-LLM and Dynamo on Spark Pairs
- DGX Spark playbook — TRT LLM for Inference, Run on two Sparksretrieved 2026-09-09
Used by:P20 Lab: Serve a 400B-Class Model on Two DGX SparksP20 TensorRT-LLM and Dynamo on Spark Pairs
- DGX Spark playbook — vLLMretrieved 2026-09-09
- NVIDIA DGX Spark playbooksretrieved 2026-09-09
Sections: Quickstarts; vllm; Playbook list
Used by:P5 NVIDIA DGX Spark: GB10, DGX OS and the aarch64 CaveatP8 TensorRT-LLM, NIM and the DGX Spark PlaybooksP9 Installing vLLM: x86 CUDA, DGX Spark, ROCm and What Does Not WorkP9 SGLang: When to Choose It
caddyserver.com
- Caddy — Automatic HTTPSretrieved 2026-09-08, 2026-09-09
Sections: Local certificate issuance; which names qualify for public certificates; Local HTTPS; CA root; Local HTTPS; the internal certificate authority; caddy trust; The internal certificate authority; caddy trust
Used by:P28 Capstone 2: The Inference ServiceP7 Lab: A Private Chat Service for Your Home NetworkP10 Privacy, Security and Serving Beyond localhostP23 Routing and Model Management: llama-swap, LiteLLM, NGINX and Health Checks
- Caddy documentation — bind directiveretrieved 2026-09-08
Used by:P7 Lab: A Private Chat Service for Your Home Network
- Caddy documentation — Caddyfile Conceptsretrieved 2026-09-08
Sections: Environment variables; site addresses
Used by:P7 Lab: A Private Chat Service for Your Home Network
- Caddy documentation — Command Lineretrieved 2026-09-13
Sections: caddy run; caddy trust; caddy untrust (--cert); caddy version
Used by:P7 Lab: A Private Chat Service for Your Home Network
- Caddy documentation — Conventionsretrieved 2026-09-08
Sections: Data directory
Used by:P7 Lab: A Private Chat Service for Your Home Network
- Caddy documentation — reverse_proxy directiveretrieved 2026-09-13
Sections: Streaming; flush_interval (ignored for text/event-stream and unknown Content-Length)
Used by:P7 Lab: A Private Chat Service for Your Home Network
- Caddy documentation — tls directiveretrieved 2026-09-08
Sections: internal (lifetimes; installing the root from a container)
Used by:P7 Lab: A Private Chat Service for Your Home Network
catalog.ngc.nvidia.com
- NVIDIA NGC — vLLM containerretrieved 2026-09-09
Sections: Tags; running the container
Used by:P9 Installing vLLM: x86 CUDA, DGX Spark, ROCm and What Does Not Work
- NVIDIA NGC catalog — PyTorch containerretrieved 2026-09-09, 2026-09-12
Sections: Tag 26.08-py3; docker run example; Multi-Arch Support; JupyterLab included; compressed size; Pull tag; docker run example; Pull tag; docker run example; Multi-Arch Support
Used by:P1 Lab: Your Python Environment and a First Trained ModelP11 Lab: Your First Training RunP11 Setting Up a Training Environment on Each Platform
code.claude.com
- Claude Agent SDK — Overviewretrieved 2026-09-09
Sections: Comparison with other Claude tools; capabilities
Used by:P26 Agent Frameworks Compared
- Claude Code — CLI referenceretrieved 2026-09-09
Sections: print; output-format; model; permission-mode; print; permission-mode
Used by:P25 Claude Code with a Local EndpointP25 Lab: One Task, Six Agents
- Claude Code — Connect Claude Code to tools via MCPretrieved 2026-09-09
Sections: claude mcp add; scopes; .mcp.json; trust warning; .mcp.json; transports; scopes; .mcp.json; scopes
Used by:P24 Lab: Write and Connect an MCP ServerP25 Claude Code with a Local EndpointP25 Project: A Local Agentic Coding Workstation
- Claude Code — environment variablesretrieved 2026-09-09
Sections: ANTHROPIC_BASE_URL; ANTHROPIC_AUTH_TOKEN; model variables; traffic controls
- Claude Code — other LLM gatewaysretrieved 2026-09-09
Sections: What a gateway provides; ANTHROPIC_BASE_URL
Used by:P25 Claude Code with a Local EndpointP25 The Landscape: Terminal, Editor and Autonomous AgentsP26 Agent Frameworks Compared
- Claude Code — permission modesretrieved 2026-09-09
Used by:P25 Challenge: The Agent That Escaped the SandboxP25 Claude Code with a Local Endpoint
- Claude Code — settingsretrieved 2026-09-09
Sections: Settings scopes; permissions; env
code.visualstudio.com
- Visual Studio Code — language modelsretrieved 2026-09-09
Sections: Tool calling requirement; custom endpoints
Used by:P25 Editor Agents: Cline, Kilo Code, Continue and the Proprietary Editors
- Visual Studio Code documentation — Remote development over SSHretrieved 2026-09-09
Sections: System requirements; Connect to a remote host
Used by:P11 Setting Up a Training Environment on Each Platform
creativecommons.org
- Creative Commons Attribution-NonCommercial 4.0 deedretrieved 2026-09-12
- Creative Commons Attribution-ShareAlike 4.0 deedretrieved 2026-09-12
curl.se
cursor.com
deeplearningbook.org
- Deep Learning (Goodfellow, Bengio and Courville) — Chapter 5, Machine Learning Basicsretrieved 2026-09-08
Sections: Chapter 5; Chapter 6; Chapter 8
Used by:P1 Generalisation: Train, Validation, Test and OverfittingP1 Neural Networks, Activations and BackpropagationP1 What Learning Means: Data, Loss and Gradient DescentP1 What Learning Means: Data, Loss and Gradient Descent
developer.amd.com
- AMD — Lemonade getting started playbookretrieved 2026-09-09, 2026-09-13
Used by:P8 Lab: Same Model, Every EngineP8 AMD-Native: ROCm Builds, Lemonade Server and the NPU Question
developer.apple.com
- Apple — App Sandboxretrieved 2026-09-09
Used by:P25 Lab: Sandbox Your Agent: Containers, Permissions and Secrets
- Apple — TN3205: Low-latency communication with RDMA over Thunderboltretrieved 2026-09-09
Sections: Requirements; enabling RDMA; Requirements; enabling RDMA; topologies
Used by:P19 Challenge: The Cluster That Is Slower Than One MachineP19 RDMA Transport, Tuning and Measuring the Split
developer.download.nvidia.com
- NVIDIA CUDA repository package index — ubuntu2404/x86_64retrieved 2026-09-13
Sections: Packages.gz; cuda-toolkit-13 -> cuda-toolkit-13-4 13.4.1-1 and its dependencies
Used by:P5 Lab: Prepare Your Machine
developer.meta.com
- Llama 3.1 Community License Agreementretrieved 2026-09-08, 2026-09-09
Sections: Section 1, Licence rights and redistribution; Section 2, Additional commercial terms; 1.b Redistribution and Use
Used by:P3 Open Weights, Open Source and LicencesP4 Model Families and Who Makes ThemP15 Pruning and Compression: Prune-and-DistilP15 Generating Synthetic Data with a Local Teacher
developer.nvidia.com
- Introducing NVFP4 for Efficient and Accurate Low-Precision Inferenceretrieved 2026-09-09, 2026-09-12
Sections: E2M1 values; NVFP4 and MXFP4 block sizes and scale formats; NVFP4 format; comparison with MXFP4
Used by:P1 Tensors, GPUs and Precision: Why Matrix Multiplication Is EverythingP16 Quantisation-Aware Training and the Blackwell Formats: FP8, NVFP4 and MXFP4
- NVIDIA Developer — CUDA GPUs, compute capability by productretrieved 2026-09-12
Sections: GeForce RTX 50 series (12.0); NVIDIA GB10, DGX Spark (12.1)
Used by:P1 Tensors, GPUs and Precision: Why Matrix Multiplication Is Everything
- NVIDIA Technical Blog — NVIDIA Hopper Architecture In-Depthretrieved 2026-09-12
Sections: PCIe Gen 5 x16 total and per-direction bandwidth, against Gen 4
Used by:P1 Tensors, GPUs and Precision: Why Matrix Multiplication Is Everything
developers.googleblog.com
- Gemma 3 QAT Models: Bringing state-of-the-art AI to consumer GPUsretrieved 2026-09-09
Sections: Perplexity drop claim; QAT training steps; perplexity drop; VRAM figures
Used by:P16 Measuring Quantisation Damage: Perplexity, KL Divergence and Task EvaluationsP16 Quantisation-Aware Training and the Blackwell Formats: FP8, NVFP4 and MXFP4
developers.llamaindex.ai
- LlamaIndex — Framework overviewretrieved 2026-09-09
Sections: Agents; workflows
Used by:P26 Agent Frameworks Compared
- LlamaIndex — OpenAILike API referenceretrieved 2026-09-09
Sections: Class description; parameters; example
Used by:P26 Agent Frameworks Compared
- LlamaIndex — vLLM exampleretrieved 2026-09-09
Sections: Note on OpenAI-compatible servers
Used by:P26 Agent Frameworks Compared
digital-strategy.ec.europa.eu
- AI Omnibus enters into force (European Commission, 27 July 2026)retrieved 2026-09-12
- Commission Guidelines on the scope of the obligations for providers of general-purpose AI models, C(2025) 7719 final (19 November 2025)retrieved 2026-09-12
Sections: 2.1 (paragraphs 17-18); 3.1.2 (paragraph 51); 3.2 (paragraphs 60-65); 4.1 and 4.2 (paragraphs 70-89); 5.3 (paragraphs 106-112); Annex A.1 (paragraph 115), A.2.2 (paragraph 129), A.3 (paragraphs 132, 136)
- General-purpose AI models in the AI Act - questions and answers (European Commission)retrieved 2026-09-12
Sections: Definitions; obligations; open-source exemption; fine-tuning; enforcement powers (last update 9 September 2025)
distilabel.argilla.io
- Distilabel documentationretrieved 2026-09-09
Sections: Overview; components gallery
- Distilabel documentation — OpenAILLM componentretrieved 2026-09-09
Sections: Attributes; example against a local server
- Distilabel documentation — quickstartretrieved 2026-09-09
Sections: Minimal pipeline
docs.ag2.ai
- AG2 — OpenAI modelsretrieved 2026-09-09
Sections: Configuration list; package naming
Used by:P26 Agent Frameworks Compared
docs.anythingllm.com
- AnythingLLM documentationretrieved 2026-09-08
Sections: Desktop, self-hosted and cloud; LLM providers; embedder setup
docs.astral.sh
- uv documentation — Installationretrieved 2026-09-12, 2026-09-13
Sections: Standalone installer; Homebrew; Updating uv; Standalone installer with a version in the URL; Homebrew; uv self update
Used by:P1 Lab: Your Python Environment and a First Trained ModelP5 Lab: Prepare Your Machine
- uv documentation — Installing and managing Pythonretrieved 2026-09-12
Sections: uv python install; automatic downloads
Used by:P1 Lab: Your Python Environment and a First Trained Model
- uv documentation — Managing packages (uv pip install)retrieved 2026-09-12
Used by:P1 Lab: Your Python Environment and a First Trained Model
- uv documentation — Python environments (uv venv)retrieved 2026-09-12
Sections: Creating a virtual environment; Using a virtual environment; discovery order
Used by:P1 Lab: Your Python Environment and a First Trained Model
- uv documentation — Using uv with pip-compatible commandsretrieved 2026-09-09
Sections: uv venv; uv pip install
Used by:P11 Setting Up a Training Environment on Each Platform
- uv documentation — Using uv with PyTorchretrieved 2026-09-12
Sections: Automatic backend selection (--torch-backend)
Used by:P1 Lab: Your Python Environment and a First Trained Model
docs.axolotl.ai
- Axolotl documentationretrieved 2026-09-09
Sections: Overview; requirements; quickstart
Used by:P13 Unsloth, Axolotl, LLaMA-Factory and mlx-lm: Higher-Level Tools
docs.brew.sh
- Homebrew documentation — Installationretrieved 2026-09-13
Sections: default prefix /opt/homebrew; post-installation shellenv; supported macOS
Used by:P5 Lab: Prepare Your Machine
docs.cline.bot
- Cline — auto-approveretrieved 2026-09-09
Used by:P25 Challenge: The Agent That Escaped the SandboxP25 Editor Agents: Cline, Kilo Code, Continue and the Proprietary Editors
- Cline — Ollama providerretrieved 2026-09-09
Used by:P25 Editor Agents: Cline, Kilo Code, Continue and the Proprietary Editors
- Cline — OpenAI Compatible providerretrieved 2026-09-09
Used by:P25 Editor Agents: Cline, Kilo Code, Continue and the Proprietary EditorsP25 The Landscape: Terminal, Editor and Autonomous Agents
- Cline — Plan and Actretrieved 2026-09-09
Used by:P25 Editor Agents: Cline, Kilo Code, Continue and the Proprietary Editors
docs.continue.dev
- Continue — configuration referenceretrieved 2026-09-08, 2026-09-09
Sections: config.yaml keys; models; roles; models; roles; capabilities
Used by:P10 Local Coding Assistants: Autocomplete and Chat in Your EditorP25 Editor Agents: Cline, Kilo Code, Continue and the Proprietary EditorsP25 The Landscape: Terminal, Editor and Autonomous Agents
- Continue — documentationretrieved 2026-09-08
Sections: Modes; configuration
Used by:P10 Local Coding Assistants: Autocomplete and Chat in Your Editor
- Continue — llama.cpp providerretrieved 2026-09-08
Used by:P10 Local Coding Assistants: Autocomplete and Chat in Your Editor
- Continue — OpenAI providerretrieved 2026-09-09
Sections: apiBase for OpenAI-compatible servers
Used by:P25 Editor Agents: Cline, Kilo Code, Continue and the Proprietary Editors
docs.crewai.com
- CrewAI — LLMsretrieved 2026-09-09
Sections: Configuring an LLM in code; local models with Ollama
Used by:P26 Agent Frameworks Compared
docs.docker.com
- Docker — bind mountsretrieved 2026-09-09
Sections: Considerations and constraints; Syntax; considerations and constraints
Used by:P25 Challenge: The Agent That Escaped the SandboxP25 Lab: Sandbox Your Agent: Containers, Permissions and Secrets
- Docker — docker image pullretrieved 2026-09-09
Sections: Pull an image by digest
Used by:P23 Backup, Upgrades and Reproducibility of a Model Estate
- Docker — Docker securityretrieved 2026-09-09
Sections: Kernel namespaces; control groups; daemon attack surface; Docker daemon attack surface
Used by:P28 Capstone 5: The Agentic WorkstationP25 Challenge: The Agent That Escaped the SandboxP25 Goose and OpenHands: Autonomous Agents and SandboxingP25 Lab: Sandbox Your Agent: Containers, Permissions and SecretsP25 Project: A Local Agentic Coding Workstation
- Docker — isolate containers with a user namespaceretrieved 2026-09-09
Used by:P25 Lab: Sandbox Your Agent: Containers, Permissions and Secrets
- Docker — networking overviewretrieved 2026-09-09
Sections: Network drivers
Used by:P25 Lab: Sandbox Your Agent: Containers, Permissions and Secrets
- Docker — none network driverretrieved 2026-09-09
Used by:P25 Challenge: The Agent That Escaped the SandboxP25 Lab: Sandbox Your Agent: Containers, Permissions and Secrets
- Docker — runtime options with memory, CPUs and GPUsretrieved 2026-09-09
Used by:P25 Lab: Sandbox Your Agent: Containers, Permissions and Secrets
- Docker Compose — the networks top-level elementretrieved 2026-09-09
Sections: driver; internal; internal; external
Used by:P18 Lab: Build and Measure Your Cluster NetworkP25 Challenge: The Agent That Escaped the SandboxP25 Lab: Sandbox Your Agent: Containers, Permissions and SecretsP25 Project: A Local Agentic Coding Workstation
- Docker Desktop — install on Macretrieved 2026-09-09
Used by:P25 Lab: Sandbox Your Agent: Containers, Permissions and Secrets
- Docker documentation — Compose file services referenceretrieved 2026-09-08, 2026-09-09
Sections: ports; read_only; tmpfs; user; cap_drop; security_opt; pids_limit; mem_limit; cpus
Used by:P7 Lab: A Private Chat Service for Your Home NetworkP25 Lab: Sandbox Your Agent: Containers, Permissions and Secrets
- Docker documentation — docker compose cpretrieved 2026-09-13
Used by:P7 Lab: A Private Chat Service for Your Home Network
- Docker documentation — docker compose downretrieved 2026-09-13
Sections: --volumes; --rmi
Used by:P7 Lab: A Private Chat Service for Your Home Network
- Docker documentation — docker compose logsretrieved 2026-09-13
Sections: --since
Used by:P7 Lab: A Private Chat Service for Your Home Network
- Docker documentation — docker compose upretrieved 2026-09-13
Sections: --wait; --wait-timeout
Used by:P7 Lab: A Private Chat Service for Your Home Network
- Docker documentation — docker container runretrieved 2026-09-09, 2026-09-12
Sections: --gpus, --publish, --ipc, --volume, --workdir, --interactive, --tty, --rm; env; env-file; read-only; user
Used by:P1 Lab: Your Python Environment and a First Trained ModelP25 Challenge: The Agent That Escaped the Sandbox
- Docker documentation — GPU support in Docker Composeretrieved 2026-09-13
Used by:P7 Lab: A Private Chat Service for Your Home Network
- Docker documentation — GPU support in Docker Desktop for Windows; WSL 2 integrationretrieved 2026-09-09, 2026-09-13
Sections: WSL 2 backend only; validation command and its sample nbody output; Settings > Resources > WSL Integration
Used by:P5 Lab: Prepare Your MachineP11 Setting Up a Training Environment on Each Platform
- Docker documentation — Install Docker Engine on Ubuntu; Linux post-installation stepsretrieved 2026-09-13
Sections: apt repository; uninstall conflicting packages; docker-ce and plugins; docker group; log out and log back in
Used by:P5 Lab: Prepare Your Machine
- Docker documentation — Interpolationretrieved 2026-09-13
Used by:P7 Lab: A Private Chat Service for Your Home Network
- Docker documentation — Linux post-installation steps for Docker Engineretrieved 2026-09-09, 2026-09-13
Sections: Manage Docker as a non-root user; docker group privileges
Used by:P7 Lab: A Private Chat Service for Your Home NetworkP25 Challenge: The Agent That Escaped the SandboxP25 Goose and OpenHands: Autonomous Agents and SandboxingP25 Lab: Sandbox Your Agent: Containers, Permissions and Secrets
- Docker documentation — Merge Compose filesretrieved 2026-09-13
Used by:P7 Lab: A Private Chat Service for Your Home Network
- Docker documentation — Packet filtering and firewallsretrieved 2026-09-08
Sections: Docker and ufw
Used by:P7 Lab: A Private Chat Service for Your Home Network
- Docker documentation — Port publishing and mappingretrieved 2026-09-13
Sections: Publishing ports; direct routing; default bind address
Used by:P7 Lab: A Private Chat Service for Your Home Network
- Docker documentation — Pre-defined environment variables in Composeretrieved 2026-09-13
Sections: COMPOSE_FILE
Used by:P7 Lab: A Private Chat Service for Your Home Network
- Docker Engine — Bridge network driverretrieved 2026-09-09
Sections: User-defined bridges
Used by:P18 Lab: Build and Measure Your Cluster NetworkP18 The Course Reference Cluster and the Single-Machine Path
docs.github.com
- GitHub Copilot — bring your own keyretrieved 2026-09-09
Sections: Policy restrictions
Used by:P25 Editor Agents: Cline, Kilo Code, Continue and the Proprietary Editors
docs.kernel.org
- Linux kernel documentation — amdgpu miscellaneousretrieved 2026-09-13
Sections: mem_info_vram_total, mem_info_vram_used, mem_info_gtt_total, mem_info_gtt_used
Used by:P6 Challenge: The Model That Runs at Two Tokens per Second
- Linux kernel documentation — amdgpu module parametersretrieved 2026-09-09
Sections: gttsize; vramlimit
Used by:P5 AMD Ryzen AI Max+ 395: Strix Halo, ROCm and Vulkan
docs.langchain.com
- LangChain — ChatOpenAI integrationretrieved 2026-09-09
Sections: Instantiation; base_url; bind_tools; with_structured_output
Used by:P26 Agent Frameworks Compared
- LangGraph — Overviewretrieved 2026-09-09
Sections: What LangGraph is; durable execution; installation example
Used by:P26 Agent Frameworks Compared
- LangGraph — Persistenceretrieved 2026-09-09
Sections: Checkpointers; threads; checkpointer libraries; Checkpointers; threads; short-term and long-term memory
Used by:P26 Agent Frameworks ComparedP26 Multi-Agent Patterns: Router, Planner and Executor, Critic, Parallel Fan-Out
- LangGraph — Streamingretrieved 2026-09-09
Sections: Stream modes
Used by:P26 Agent Frameworks Compared
docs.litellm.ai
- LiteLLM — Anthropic /v1/messagesretrieved 2026-09-09
Sections: Usage; Usage; supported providers
Used by:P9 Project: Your Local Model GatewayP25 Claude Code with a Local Endpoint
- LiteLLM — Budgets and rate limitsretrieved 2026-09-09
Sections: rpm_limit; tpm_limit; max_budget; budget_duration; tpm_limit; rpm_limit
Used by:P23 Challenge: The 3 a.m. Out-of-MemoryP23 Routing and Model Management: llama-swap, LiteLLM, NGINX and Health ChecksP23 Security for Exposed Endpoints and the Model Supply Chain
- LiteLLM — Claude Code quickstartretrieved 2026-09-09
Sections: Unified endpoint; environment variables; troubleshooting
- LiteLLM — Deploymentretrieved 2026-09-09
Sections: Container image; required environment variables; default port; Container image; pinning a version tag
Used by:P9 Project: Your Local Model GatewayP23 Backup, Upgrades and Reproducibility of a Model Estate
- LiteLLM — Health checksretrieved 2026-09-09
Sections: Endpoints; background health checks
Used by:P9 Project: Your Local Model GatewayP23 Routing and Model Management: llama-swap, LiteLLM, NGINX and Health Checks
- LiteLLM — Loggingretrieved 2026-09-09
Sections: Callbacks; message redaction; turn_off_message_logging; callbacks; per-request redaction; turn_off_message_logging; per-request redaction
Used by:P9 Project: Your Local Model GatewayP23 Observability: Metrics, Logs and Traces for LLM ServingP23 Security for Exposed Endpoints and the Model Supply Chain
- LiteLLM — Prometheus metricsretrieved 2026-09-09
Sections: callbacks; metric names and labels; Enabling the callback; metric names and labels
Used by:P23 Lab: Dashboards for Your ClusterP23 Observability: Metrics, Logs and Traces for LLM Serving
- LiteLLM — Proxy config.yamlretrieved 2026-09-09
Sections: model_list; os.environ references; OpenAI-compatible endpoints; model_list; os.environ references; routing_strategy; model_group_alias; os.environ references; model_list; OpenAI-compatible endpoints
Used by:P9 Project: Your Local Model GatewayP23 Routing and Model Management: llama-swap, LiteLLM, NGINX and Health ChecksP23 Security for Exposed Endpoints and the Model Supply ChainP25 Project: A Local Agentic Coding WorkstationP26 Project: A Multi-Agent System on Your Cluster
- LiteLLM — Proxy overviewretrieved 2026-09-09
Sections: Starting the proxy; config.yaml
- LiteLLM — Reliability and fallbacksretrieved 2026-09-09
Sections: fallbacks; num_retries; cooldown; context_window_fallbacks; fallbacks; context_window_fallbacks; num_retries; allowed_fails; cooldown_time; fallbacks
Used by:P9 Project: Your Local Model GatewayP23 Challenge: The 3 a.m. Out-of-MemoryP23 Routing and Model Management: llama-swap, LiteLLM, NGINX and Health ChecksP25 Project: A Local Agentic Coding Workstation
- LiteLLM — Virtual Keysretrieved 2026-09-09
Sections: Budgets; rate limits; key expiry; Budgets; rate limits; key expiry and rotation; Requirements; key generation; Requirements; key generation; duration; key info; Key generation; per-key usage
Used by:P28 Capstone 5: The Agentic WorkstationP28 Capstone 2: The Inference ServiceP9 Project: Your Local Model GatewayP23 Routing and Model Management: llama-swap, LiteLLM, NGINX and Health ChecksP23 Security for Exposed Endpoints and the Model Supply ChainP25 Lab: One Task, Six AgentsP25 Project: A Local Agentic Coding Workstation
docs.lmcache.ai
- LMCache — Disaggregated prefillretrieved 2026-09-09
Sections: Two-node setup; single-node note; Single node testing note; environment; Two-node setup; single-node note; environment variables
Used by:P22 Lab: Two-Machine Prefill and Decode with vLLMP22 The KV Cache as a Transferable Object: Mooncake, LMCache and NIXLP22 vLLM Disaggregated Prefill: Connectors and the Proxy
- LMCache — documentationretrieved 2026-09-09
Sections: Overview; tiered storage; reuse across serving engines; Tiered storage; reuse across serving engines; Overview; secondary KV storage; distributed KV cache
Used by:P22 Lab: KV Cache Offload and SharingP22 Project: A Tiered Inference ArchitectureP22 The KV Cache as a Transferable Object: Mooncake, LMCache and NIXL
- LMCache — FileSystem secondary storage backendretrieved 2026-09-09
Sections: Configuration
Used by:P22 Lab: KV Cache Offload and SharingP22 The KV Cache as a Transferable Object: Mooncake, LMCache and NIXL
- LMCache — Kubernetes deploymentretrieved 2026-09-09
Sections: vLLM Production Stack
- LMCache — quickstartretrieved 2026-09-09
Sections: vLLM in-process mode; configuration keys
docs.nvidia.com
- CUDA C++ Best Practices Guide — Data Transfer Between Host and Deviceretrieved 2026-09-12
Sections: PCIe x16 Gen3 bandwidth; device memory bandwidth comparison
Used by:P1 Tensors, GPUs and Precision: Why Matrix Multiplication Is Everything
- CUDA C++ Programming Guide — Compute Capabilities, technical specifications per compute capabilityretrieved 2026-09-12
Sections: Table 31, Memory Information per Compute Capability (registers per SM; maximum shared memory per SM); Table 32, Shared Memory Capacity per Compute Capability
Used by:P1 Tensors, GPUs and Precision: Why Matrix Multiplication Is Everything
- CUDA C++ Programming Guide — Unified and System Memoryretrieved 2026-09-12
Sections: mapped memory across the CPU-GPU interconnect
Used by:P1 Tensors, GPUs and Precision: Why Matrix Multiplication Is Everything
- CUDA C++ Programming Guide — Writing SIMT Kernels (memory spaces)retrieved 2026-09-12
Sections: registers, shared memory, L1 and L2 cache, global memory
Used by:P1 Tensors, GPUs and Precision: Why Matrix Multiplication Is Everything
- NCCL documentation — Environment Variablesretrieved 2026-09-09
Sections: NCCL_SOCKET_IFNAME; NCCL_IB_HCA; NCCL_IB_GID_INDEX; NCCL_IB_DISABLE; NCCL_DEBUG; NCCL_SOCKET_IFNAME; NCCL_IB_HCA; NCCL_DEBUG; NCCL_IB_DISABLE
Used by:P20 Connecting Two DGX Sparks over ConnectX-7P20 vLLM Multi-Node with Ray: Tensor Parallel Inside, Pipeline Parallel Across
- NVIDIA Container Toolkit — Installing the NVIDIA Container Toolkitretrieved 2026-09-13
Sections: apt repository; version 1.20.0-1; nvidia-ctk runtime configure; sample workload
Used by:P5 Lab: Prepare Your Machine
- NVIDIA Container Toolkit — Running a sample workloadretrieved 2026-09-13
Used by:P7 Lab: A Private Chat Service for Your Home Network
- NVIDIA CUDA Installation Guide for Linux (CUDA 13.4)retrieved 2026-09-09, 2026-09-13
Sections: Network repository installation (cuda-keyring); meta packages; toolkit and driver independent from 13.4; post-installation PATH; /usr/local/cuda symbolic link; Pre-installation actions; Network repository installation; Post-installation actions
Used by:P5 Lab: Prepare Your MachineP5 NVIDIA Desktops and Laptops: VRAM Tiers, CUDA and WSL2
- NVIDIA CUDA on WSL User Guideretrieved 2026-09-09, 2026-09-12, 2026-09-13
Sections: Install the Windows driver only; wsl.exe --update; /usr/lib/wsl/lib/nvidia-smi; Getting started; CUDA support for WSL 2; known limitations (/usr/lib/wsl/lib); WSL kernel 5.10.16.3; Getting started with CUDA on WSL 2; CUDA support for WSL 2; Known limitations; Features not yet supported (NVML queries)
Used by:P1 Lab: Your Python Environment and a First Trained ModelP5 Lab: Prepare Your MachineP5 NVIDIA Desktops and Laptops: VRAM Tiers, CUDA and WSL2P8 Lab: Same Model, Every Engine
- NVIDIA CUDA Toolkit release notesretrieved 2026-09-12
Sections: CUDA 13.x applications run on drivers >=580; CUDA 12.8 GA needs 570.26 (Linux) / 570.65 (Windows); CUDA 12.6 GA needs 560.28.03 (Linux) / 560.76 (Windows)
Used by:P1 Lab: Your Python Environment and a First Trained Model
- NVIDIA DCGM — Install DCGM Exporterretrieved 2026-09-09
Sections: DCGM_FI_DEV_POWER_USAGE; DCGM_FI_DEV_FB_USED; Running the container; the counters CSV; field names
Used by:P23 Capacity Planning and Cost per Million Tokens at HomeP23 Challenge: The 3 a.m. Out-of-MemoryP23 Lab: Dashboards for Your ClusterP23 Observability: Metrics, Logs and Traces for LLM Serving
- NVIDIA Deep Learning Performance — Train With Mixed Precisionretrieved 2026-09-09, 2026-09-12
Sections: Half Precision Format; Loss Scaling To Preserve Small Gradient Magnitudes; Satisfying Tensor Core Shape Constraints; Loss scaling; single-precision master weights
Used by:P1 Tensors, GPUs and Precision: Why Matrix Multiplication Is EverythingP11 Memory Arithmetic for Training: Weights, Gradients, Optimiser States and Activations
- NVIDIA DGX Spark documentationretrieved 2026-09-09
- NVIDIA DGX Spark documentation — Container Runtime for Dockerretrieved 2026-09-09, 2026-09-12, 2026-09-13
Sections: NVIDIA Container Toolkit preinstalled; docker run --gpus; the docker group; Installation; optional docker group; validation; runtime not found
Used by:P1 Lab: Your Python Environment and a First Trained ModelP5 Lab: Prepare Your MachineP5 NVIDIA DGX Spark: GB10, DGX OS and the aarch64 Caveat
- NVIDIA DGX Spark documentation — Release notesretrieved 2026-09-12, 2026-09-13
Sections: DGX OS 7.5.0; CUDA Toolkit 13.0.2; GPU driver 580.159.03; Current software versions (DGX OS 7.5.0, driver 580.159.03, CUDA 13.0.2, kernel 6.17)
Used by:P1 Lab: Your Python Environment and a First Trained ModelP5 Lab: Prepare Your Machine
- NVIDIA DGX Spark User Guide — ConnectX-7 Networkingretrieved 2026-09-09
Sections: QSFP ports; interface naming; QSFP ports; QSFP ports and RoCE devices; QSFP ports and RoCE devices; supported cluster sizes; QSFP ports and RoCE devices; supported cluster sizes; approved cables
Used by:P5 NVIDIA DGX Spark: GB10, DGX OS and the aarch64 CaveatP18 Lab: Build and Measure Your Cluster NetworkP18 Networking for Home Clusters: 2.5 GbE to 200 GbE, Thunderbolt 5 and RDMAP18 The Course Reference Cluster and the Single-Machine PathP19 Challenge: The Cluster That Is Slower Than One MachineP19 Lab: Run a Model Bigger Than Any One MachineP19 RDMA Transport, Tuning and Measuring the SplitP20 Connecting Two DGX Sparks over ConnectX-7P20 Lab: Serve a 400B-Class Model on Two DGX SparksP22 The KV Cache as a Transferable Object: Mooncake, LMCache and NIXL
- NVIDIA DGX Spark User Guide — DGX Dashboardretrieved 2026-09-13
Sections: localhost:11000; SSH tunnel
Used by:P5 Lab: Prepare Your Machine
- NVIDIA DGX Spark User Guide — DGX OSretrieved 2026-09-09
Used by:P5 NVIDIA DGX Spark: GB10, DGX OS and the aarch64 Caveat
- NVIDIA DGX Spark User Guide — Hardware Overviewretrieved 2026-09-09, 2026-09-13
Sections: LPDDR5X 8533; 16 channels (256 bit); 273 GB/s; up to 1,000 TOPS; up to 1 PFLOP at FP4 precision with sparsity
Used by:P5 Lab: Measure Your Memory Bandwidth and ComputeP5 NVIDIA DGX Spark: GB10, DGX OS and the aarch64 Caveat
- NVIDIA DGX Spark User Guide — Known Issuesretrieved 2026-09-09, 2026-09-13
Sections: Memory-Usage Not Supported; cudaMemGetInfo; power adapter; nvidia-smi Memory-Usage Not Supported; memory reporting on unified memory and SWAP; nvidia-smi Memory-Usage; cudaMemGetInfo; nvidia-smi reports Memory-Usage Not Supported; reporting memory with unified memory
Used by:P5 Lab: Prepare Your MachineP5 NVIDIA DGX Spark: GB10, DGX OS and the aarch64 CaveatP6 Challenge: The Model That Runs at Two Tokens per SecondP6 Lab: Run and Benchmark the Course Reference ModelsP8 Lab: Same Model, Every Engine
- NVIDIA DGX Spark User Guide — OS and Component Update Guideretrieved 2026-09-09, 2026-09-13
Sections: Update methods; manual system updates; Founders Edition note
Used by:P5 Lab: Prepare Your MachineP5 NVIDIA DGX Spark: GB10, DGX OS and the aarch64 Caveat
- NVIDIA Dynamo — Compatibilityretrieved 2026-09-09
Sections: Platform support
- NVIDIA Dynamo — Overall Architectureretrieved 2026-09-09
Sections: Design goals; system model; request, control and storage planes; Design goals; request plane; Design goals; request plane; storage and events plane
Used by:P20 TensorRT-LLM and Dynamo on Spark PairsP22 Project: A Tiered Inference ArchitectureP22 SGLang PD Disaggregation and NVIDIA Dynamo
- NVIDIA Dynamo — RDMA Setupretrieved 2026-09-09
Sections: What RDMA is; when you need it; Why Dynamo needs RDMA
Used by:P20 TensorRT-LLM and Dynamo on Spark PairsP22 Lab: Two-Machine Prefill and Decode with vLLMP22 SGLang PD Disaggregation and NVIDIA DynamoP22 The KV Cache as a Transferable Object: Mooncake, LMCache and NIXLP22 Why Prefill and Decode Want Different Hardware
- NVIDIA Dynamo — Router designretrieved 2026-09-09
Sections: Cost function; KV overlap
Used by:P22 Project: A Tiered Inference ArchitectureP22 SGLang PD Disaggregation and NVIDIA Dynamo
- NVIDIA NeMo — Automatic Speech Recognitionretrieved 2026-09-08
Sections: Transcribing with a pretrained model
Used by:P10 Vision, Speech and Documents: Multimodal Locally
- NVIDIA NeMo AutoModel documentationretrieved 2026-09-09
Sections: Overview; PEFT; installation
Used by:P13 Unsloth, Axolotl, LLaMA-Factory and mlx-lm: Higher-Level Tools
- NVIDIA NIM documentation hubretrieved 2026-09-09
- NVIDIA PyTorch container release notes — Release 26.08retrieved 2026-09-12
Sections: Contents of the PyTorch container (Ubuntu 24.04, Python 3.12, PyTorch 2.14.0a0, CUDA 13.4.1, JupyterLab 4.6.3)
Used by:P1 Lab: Your Python Environment and a First Trained Model
- NVIDIA System Management Interface (nvidia-smi) documentationretrieved 2026-09-09, 2026-09-13
Sections: --format csv, noheader, nounits; -d/--display PERFORMANCE; Clocks Event Reasons; Selective query options; GPU Link information; clocks throttle reasons; GPU Link information; power-limit; clocks throttle reasons; topo; --query-gpu; --format=csv with nounits and noheader; --loop; --query-gpu; --format; --loop
Used by:P5 Lab: Measure Your Memory Bandwidth and ComputeP12 Lab: Train a 10M to 125M Parameter Model in an AfternoonP20 Lab: Serve a 400B-Class Model on Two DGX SparksP20 Multi-GPU Desktops: PCIe, Tensor Parallel Without NVLink and Expert ParallelP23 Capacity Planning and Cost per Million Tokens at HomeP23 Observability: Metrics, Logs and Traces for LLM Serving
docs.ollama.com
- Ollama — Tool callingretrieved 2026-09-09
Sections: Tools in the request; tool results; streaming; the agent loop; The think parameter; message.thinking; streaming
Used by:P24 Function Calling End to End on Local EnginesP24 Reasoning Models in Agent Loops
docs.openhands.dev
- OpenHands — CLI headless moderetrieved 2026-09-09
Used by:P25 Lab: One Task, Six Agents
- OpenHands — LLM configurationretrieved 2026-09-09
Sections: Model requirements
Used by:P25 Goose and OpenHands: Autonomous Agents and Sandboxing
- OpenHands — local LLMsretrieved 2026-09-09
Sections: Model prefix; context length; tool-use reliability; Context length; tool-use reliability
Used by:P25 Goose and OpenHands: Autonomous Agents and SandboxingP25 The Landscape: Terminal, Editor and Autonomous AgentsP25 Which Local Models Can Actually Drive an Agent
docs.openwebui.com
- Open WebUI — Knowledgeretrieved 2026-09-08
Sections: Collections; referencing a collection in chat; citations
Used by:P10 Project: A Private Document Question-Answering Service
- Open WebUI Docs — Backups (community tutorial)retrieved 2026-09-13
Sections: Files in persistent data store
Used by:P7 Lab: A Private Chat Service for Your Home Network
- Open WebUI Docs — Environment Variable Configurationretrieved 2026-09-13
Sections: ConfigVar environment variables; ENABLE_SIGNUP; WEBUI_ADMIN_EMAIL; DEFAULT_USER_ROLE; OLLAMA_BASE_URL; ENABLE_OPENAI_API; WEBUI_SECRET_KEY; CORS_ALLOW_ORIGIN; ENABLE_VERSION_UPDATE_CHECK
Used by:P7 Lab: A Private Chat Service for Your Home Network
- Open WebUI Docs — Featuresretrieved 2026-09-08
- Open WebUI Docs — Groupsretrieved 2026-09-13
Sections: Preview Access
Used by:P7 Lab: A Private Chat Service for Your Home Network
- Open WebUI Docs — Hardening Open WebUIretrieved 2026-09-08, 2026-09-13
Sections: Secrets; registration; network architecture; TLS; CORS; Secrets; registration; cookie settings; network architecture; TLS; CORS
Used by:P7 Front-Ends: Open WebUI and FriendsP7 Lab: A Private Chat Service for Your Home Network
- Open WebUI Docs — Licenseretrieved 2026-09-08
Sections: Branding restriction; thresholds; effective version
- Open WebUI Docs — Llama.cppretrieved 2026-09-13
Used by:P7 Lab: A Private Chat Service for Your Home Network
- Open WebUI Docs — Models (Workspace)retrieved 2026-09-13
Sections: Core configuration; Visibility
Used by:P7 Lab: A Private Chat Service for Your Home Network
- Open WebUI Docs — Quick Startretrieved 2026-09-08
Sections: Docker; Python (pip/uv); first account; connections; Docker; Docker Compose; Python (pip/uv); the first account
Used by:P7 Front-Ends: Open WebUI and FriendsP7 Lab: A Private Chat Service for Your Home Network
- Open WebUI Docs — Rolesretrieved 2026-09-13
Sections: Role details; headless admin account creation
Used by:P7 Lab: A Private Chat Service for Your Home Network
- Open WebUI Docs — Starting With OpenAI-Compatible Serversretrieved 2026-09-08
Sections: Connection steps; URL format; local server examples
Used by:P7 Front-Ends: Open WebUI and FriendsP7 Lab: A Private Chat Service for Your Home Network
- Open WebUI Docs — Task Modelsretrieved 2026-09-13
Sections: Turning individual tasks off
Used by:P7 Lab: A Private Chat Service for Your Home Network
docs.podman.io
- Podman documentation — indexretrieved 2026-09-09
Sections: Familiar CLI
Used by:P25 Lab: Sandbox Your Agent: Containers, Permissions and Secrets
docs.python.org
- Python documentation — resource, resource usage informationretrieved 2026-09-09
Sections: Availability; RLIMIT_CPU; RLIMIT_AS; RLIMIT_FSIZE; RLIMIT_NPROC; setrlimit; setrlimit; RLIMIT_CPU, RLIMIT_AS, RLIMIT_FSIZE, RLIMIT_NPROC
Used by:P14 Reward Functions: Maths, Code Tests, Format and LengthP24 Lab: A Minimal Agent from Scratch
- Python documentation — subprocess, subprocess managementretrieved 2026-09-09
Sections: subprocess.run timeout; Security considerations; subprocess.run; timeout and TimeoutExpired; shell injection
Used by:P14 Reward Functions: Maths, Code Tests, Format and LengthP24 Lab: A Minimal Agent from Scratch
docs.pytorch.org
- PyTorch 2.14 — torch.nn.functional.scaled_dot_product_attentionretrieved 2026-09-12
Sections: Signature; is_causal; scale; enable_gqa
Used by:P2 Attention and the Transformer
- PyTorch 2.14 documentation — Tensor Attributes, torch.dtyperetrieved 2026-09-12, 2026-09-13
Sections: dtype table (sign-exponent-mantissa layouts); float8 limitations; float16 S-E-M 1-5-10; bfloat16 S-E-M 1-8-7; float8_e4m3fn S-E-M 1-4-3
Used by:P1 Tensors, GPUs and Precision: Why Matrix Multiplication Is EverythingP5 Lab: Measure Your Memory Bandwidth and Compute
- PyTorch 2.14 documentation — torch.matmulretrieved 2026-09-13
Sections: 2-D and 1-D inputs return the matrix-vector product
- PyTorch 2.14 documentation — torch.mps.recommended_max_memoryretrieved 2026-09-12
- PyTorch 2.14 documentation — torch.nn.CrossEntropyLossretrieved 2026-09-12
Sections: Input expectations; loss with class indices; reduction
Used by:P1 What Learning Means: Data, Loss and Gradient Descent
- PyTorch 2.14 documentation — torch.nn.Dropoutretrieved 2026-09-12
Sections: Training and evaluation behaviour; scaling factor
Used by:P1 Generalisation: Train, Validation, Test and Overfitting
- PyTorch 2.14 documentation — torch.nn.functional.scaled_mmretrieved 2026-09-13
Sections: signature; the scaling-recipe enums are not documented on the page
- PyTorch 2.14 documentation — torch.optim.AdamWretrieved 2026-09-09, 2026-09-12
Sections: Algorithm box (decoupled weight decay); constructor defaults; Algorithm box; constructor defaults; Algorithm; parameters
Used by:P1 Generalisation: Train, Validation, Test and OverfittingP1 What Learning Means: Data, Loss and Gradient DescentP11 Memory Arithmetic for Training: Weights, Gradients, Optimiser States and Activations
- PyTorch 2.14 documentation — torch.optim.SGDretrieved 2026-09-12
Sections: Algorithm box (weight decay); constructor defaults; Algorithm box; constructor defaults
Used by:P1 Generalisation: Train, Validation, Test and OverfittingP1 What Learning Means: Data, Loss and Gradient Descent
- PyTorch 2.14 documentation — torch.Tensor.backwardretrieved 2026-09-12
Sections: Gradient accumulation note
Used by:P1 What Learning Means: Data, Loss and Gradient Descent
- PyTorch 2.14 documentation — Type Info, torch.finforetrieved 2026-09-12
Sections: torch.finfo attributes
Used by:P1 Tensors, GPUs and Precision: Why Matrix Multiplication Is Everything
- PyTorch documentation — Autograd mechanicsretrieved 2026-09-09
Sections: How autograd encodes the history; Setting requires_grad; Locally disabling gradient computation
Used by:P11 PyTorch, Transformers, Datasets and the Hugging Face Ecosystem
- PyTorch documentation — Automatic Mixed Precision package, torch.ampretrieved 2026-09-08, 2026-09-09
Sections: Autocasting; Gradient Scaling; Autocasting
Used by:P1 Tensors, GPUs and Precision: Why Matrix Multiplication Is EverythingP11 PyTorch, Transformers, Datasets and the Hugging Face EcosystemP11 Setting Up a Training Environment on Each Platform
- PyTorch documentation — torch.mpsretrieved 2026-09-09
Sections: synchronize; empty_cache; recommended_max_memory
Used by:P5 Apple Silicon: Unified Memory, Metal and MLXP5 Lab: Measure Your Memory Bandwidth and Compute
- PyTorch documentation — torch.utils.checkpointretrieved 2026-09-09
Sections: Warnings; Activation checkpointing; use_reentrant
Used by:P11 Experiment Tracking and ReproducibilityP11 Memory Arithmetic for Training: Weights, Gradients, Optimiser States and Activations
- PyTorch documentation 2.14 — MPS backendretrieved 2026-09-09, 2026-09-12
Sections: is_available and is_built; the macOS 14.0 requirement
Used by:P1 Lab: Your Python Environment and a First Trained ModelP11 Setting Up a Training Environment on Each Platform
- PyTorch documentation 2.14 — Reproducibilityretrieved 2026-09-09, 2026-09-12
Sections: results may differ across platforms and between CPU and GPU with identical seeds; Controlling sources of randomness; CUDA convolution benchmarking; DataLoader
Used by:P1 Lab: Your Python Environment and a First Trained ModelP11 Experiment Tracking and Reproducibility
- PyTorch documentation 2.14 — torch.loadretrieved 2026-09-12
Sections: weights_only default; map_location; the warning about untrusted files
Used by:P1 Lab: Your Python Environment and a First Trained Model
- PyTorch documentation 2.14 — torch.nn.initretrieved 2026-09-12
Sections: kaiming_uniform_, xavier_uniform_, constant_
- PyTorch documentation 2.14 — torch.nn.Linearretrieved 2026-09-12
Sections: Variables (weight and bias initialisation)
- PyTorch documentation 2.14 — torch.nn.Moduleretrieved 2026-09-12
Sections: state_dict; load_state_dict (strict); eval and train
Used by:P1 Lab: Your Python Environment and a First Trained Model
- PyTorch documentation 2.14 — torch.nn.SiLUretrieved 2026-09-12
- PyTorch documentation 2.14 — torch.saveretrieved 2026-09-12
Sections: zipfile-based format since 1.6; the .pt convention
Used by:P1 Lab: Your Python Environment and a First Trained Model
- PyTorch tutorial — Automatic Differentiation with torch.autogradretrieved 2026-09-12
- PyTorch tutorial — Optimizing Model Parametersretrieved 2026-09-12
Sections: Hyperparameters; Optimization Loop; Loss Function; Optimizer
Used by:P1 What Learning Means: Data, Loss and Gradient Descent
- PyTorch tutorials — Saving and Loading Modelsretrieved 2026-09-12
Sections: What is a state_dict; save/load state_dict; model.eval()
Used by:P1 Lab: Your Python Environment and a First Trained Model
- torchvision documentation — datasets.MNISTretrieved 2026-09-12
Sections: root layout MNIST/raw; download parameter
Used by:P1 Lab: Your Python Environment and a First Trained Model
docs.ray.io
- Ray documentation — Launching an On-Premise Clusterretrieved 2026-09-09
Sections: Start the Head Node; Start Worker Nodes; Troubleshooting
Used by:P20 Lab: Serve a 400B-Class Model on Two DGX SparksP20 vLLM Multi-Node with Ray: Tensor Parallel Inside, Pipeline Parallel Across
docs.searxng.org
- SearXNG — Documentationretrieved 2026-09-09
Sections: What SearXNG is; privacy; self-hosting
- SearXNG — Search APIretrieved 2026-09-09
Sections: Endpoints; parameters; output formats
docs.sglang.io
- SGLang — AMD GPU platformretrieved 2026-09-09
Sections: Supported hardware; installation
Used by:P9 SGLang: When to Choose It
- SGLang — Installationretrieved 2026-09-09
Sections: Install with pip or uv; Docker; platform pages
Used by:P9 SGLang: When to Choose It
- SGLang — PD Disaggregationretrieved 2026-09-09
Sections: Router; Motivation; Mooncake; NIXL; router; environment variables; Motivation
Used by:P22 Project: A Tiered Inference ArchitectureP22 SGLang PD Disaggregation and NVIDIA DynamoP22 Why Prefill and Decode Want Different Hardware
- SGLang — Production metricsretrieved 2026-09-09
Sections: num_running_reqs; num_queue_reqs; token_usage; Enabling metrics; Enabling metrics; metric names
Used by:P23 Challenge: The 3 a.m. Out-of-MemoryP23 Lab: Dashboards for Your ClusterP23 Observability: Metrics, Logs and Traces for LLM Serving
- SGLang — Server argumentsretrieved 2026-09-09
Sections: Model, HTTP server, memory and scheduling, API options; KV cache dtype; radix cache; metrics
Used by:P9 SGLang: When to Choose ItP17 Prefix Caching and KV Reuse
- SGLang — Speculative decodingretrieved 2026-09-09
Sections: EAGLE and EAGLE3; tuning the three parameters; EAGLE and EAGLE3 launch commands; memory notes; Supported algorithms; EAGLE3 launch command
Used by:P9 Speculative Decoding: Draft Models, EAGLE and n-gramP17 Speculative Decoding Revisited: Acceptance Rates and When It PaysP17 Training a Draft Model: Medusa and EAGLE at Home
- SGLang — Structured outputsretrieved 2026-09-09
Sections: Grammar backends; constraint parameters; Grammar backends
Used by:P9 Tool Calling and Structured Output on the Server SideP9 SGLang: When to Choose It
- SGLang — Tool parserretrieved 2026-09-09
Sections: Supported parsers
Used by:P9 Tool Calling and Structured Output on the Server SideP9 SGLang: When to Choose It
docs.vllm.ai
- llm-compressor documentationretrieved 2026-09-09
Sections: Overview; supported formats; Algorithms; supported formats
Used by:P13 Merging, Exporting and Quantising a Fine-Tuned ModelP16 Post-Training Quantisation in Depth: GPTQ, AWQ, K-quants, imatrix and HQQ
- vLLM — AutoAWQretrieved 2026-09-09
Sections: Deprecation; quant_config; serving; Deprecation notice
Used by:P9 Serving with vLLM: Quantised Weights, Context, Memory and Multi-GPUP13 Merging, Exporting and Quantising a Fine-Tuned ModelP13 Unsloth, Axolotl, LLaMA-Factory and mlx-lm: Higher-Level ToolsP20 Lab: Serve a 400B-Class Model on Two DGX Sparks
- vLLM — Automatic Prefix Caching (design)retrieved 2026-09-09
Sections: Block hashing; eviction
Used by:P9 Why a Second Kind of Engine: Batching, Paged Attention and Throughput
- vLLM — Automatic Prefix Caching (design)retrieved 2026-09-09
Sections: Block hashing; eviction
Used by:P17 Prefix Caching and KV Reuse
- vLLM — Automatic Prefix Caching (usage)retrieved 2026-09-09
Sections: Long document query; multi-round conversation; limits; Multi-round conversation; limits
Used by:P17 Prefix Caching and KV ReuseP24 Context Engineering: Memory, Compaction and the KV Budget
- vLLM — CPU installationretrieved 2026-09-09
Sections: Apple silicon requirements and limitations
Used by:P9 Installing vLLM: x86 CUDA, DGX Spark, ROCm and What Does Not Work
- vLLM — Data Parallel Deploymentretrieved 2026-09-09
Sections: Internal load balancing; multi-node
Used by:P20 vLLM Multi-Node with Ray: Tensor Parallel Inside, Pipeline Parallel Across
- vLLM — Disaggregated Prefilling (experimental)retrieved 2026-09-09
Sections: OffloadingConnector; ExampleConnector; Usage example; connectors; status; Connectors; status; Connectors; Connectors; development abstractions; Why; usage example; connectors; development; Why disaggregated prefilling; benefits
Used by:P22 Lab: KV Cache Offload and SharingP22 Lab: Two-Machine Prefill and Decode with vLLMP22 Project: A Tiered Inference ArchitectureP22 SGLang PD Disaggregation and NVIDIA DynamoP22 The KV Cache as a Transferable Object: Mooncake, LMCache and NIXLP22 vLLM Disaggregated Prefill: Connectors and the ProxyP22 Why Prefill and Decode Want Different Hardware
- vLLM — Engine argumentsretrieved 2026-09-09
Sections: Defaults; cpu-offload-gb; cpu-offload-gb; kv-cache-dtype
Used by:P9 Serving with vLLM: Quantised Weights, Context, Memory and Multi-GPUP22 Lab: KV Cache Offload and SharingP22 vLLM Disaggregated Prefill: Connectors and the Proxy
- vLLM — Expert Parallel Deploymentretrieved 2026-09-09
Sections: Single node deployment; backend selection guide; Configuration; layer behavior; backend selection
Used by:P20 Multi-GPU Desktops: PCIe, Tensor Parallel Without NVLink and Expert ParallelP20 vLLM Multi-Node with Ray: Tensor Parallel Inside, Pipeline Parallel Across
- vLLM — FP8 quantizationretrieved 2026-09-09
Sections: Hardware requirements; E4M3 and E5M2; online dynamic quantization; Hardware requirements; recipe; online dynamic quantisation
Used by:P9 Serving with vLLM: Quantised Weights, Context, Memory and Multi-GPUP13 Merging, Exporting and Quantising a Fine-Tuned Model
- vLLM — GPTQModelretrieved 2026-09-09
Used by:P9 Serving with vLLM: Quantised Weights, Context, Memory and Multi-GPU
- vLLM — GPU installationretrieved 2026-09-09
Sections: CUDA requirements; ROCm requirements and supported GPUs
Used by:P9 Installing vLLM: x86 CUDA, DGX Spark, ROCm and What Does Not Work
- vLLM — LoRA adaptersretrieved 2026-09-09
Sections: enable-lora; lora-modules; max-lora-rank; runtime updating
Used by:P13 Merging, Exporting and Quantising a Fine-Tuned Model
- vLLM — Multimodal Inputsretrieved 2026-09-08
Sections: Offline inference; online serving; --limit-mm-per-prompt
Used by:P10 Vision, Speech and Documents: Multimodal Locally
- vLLM — NVIDIA Model Optimizerretrieved 2026-09-09
Sections: NVFP4; supported checkpoint formats
Used by:P9 Serving with vLLM: Quantised Weights, Context, Memory and Multi-GPU
- vLLM — Optimization and Tuningretrieved 2026-09-09
Sections: Preemption; chunked prefill; Preemption; chunked prefill; CUDA graphs
Used by:P9 Why a Second Kind of Engine: Batching, Paged Attention and ThroughputP9 Lab: Serve a Model to Twenty Concurrent UsersP9 Serving with vLLM: Quantised Weights, Context, Memory and Multi-GPU
- vLLM — Parallelism and Scalingretrieved 2026-09-09
Sections: Multi-node with Ray; tensor and pipeline parallel sizing; Choosing a strategy; Choosing a strategy; multi-node; Choosing a parallelism strategy; Multi-node deployment; Multi-node deployment; Ray cluster setup with containers; Optimizing network communication; Distributed inference strategies; edge case, uneven GPU splits; Distributed inference strategies; Multi-node deployment; Ray cluster setup with containers; Optimizing network communication
Used by:P28 Capstone 3: Cluster or Tiered DeploymentP9 Serving with vLLM: Quantised Weights, Context, Memory and Multi-GPUP18 Tensor, Pipeline, Expert, Data and Sequence ParallelismP18 Why One Box Runs Out: The Memory Wall and the Bandwidth WallP18 The Course Reference Cluster and the Single-Machine PathP20 Lab: Serve a 400B-Class Model on Two DGX SparksP20 Multi-GPU Desktops: PCIe, Tensor Parallel Without NVLink and Expert ParallelP20 vLLM Multi-Node with Ray: Tensor Parallel Inside, Pipeline Parallel Across
- vLLM — Production metricsretrieved 2026-09-09
Sections: Metric names; endpoint; Speculative decoding metrics; Prefix cache metrics; Prefix cache metrics; KV cache usage; Metric names; kv_cache_usage_perc; num_requests_running; num_requests_waiting; Metric names and types
Used by:P9 Lab: Serve a Model to Twenty Concurrent UsersP9 Serving with vLLM: Quantised Weights, Context, Memory and Multi-GPUP9 Speculative Decoding: Draft Models, EAGLE and n-gramP17 Lab: Train and Deploy a Draft for Your ModelP17 Prefix Caching and KV ReuseP17 Speculative Decoding Revisited: Acceptance Rates and When It PaysP22 Lab: KV Cache Offload and SharingP22 Lab: Two-Machine Prefill and Decode with vLLMP22 Project: A Tiered Inference ArchitectureP22 vLLM Disaggregated Prefill: Connectors and the ProxyP22 Why Prefill and Decode Want Different HardwareP23 Challenge: The 3 a.m. Out-of-MemoryP23 Lab: Dashboards for Your ClusterP23 Observability: Metrics, Logs and Traces for LLM Serving
- vLLM — Quantizationretrieved 2026-09-09
Sections: Supported hardware matrix
Used by:P9 Serving with vLLM: Quantised Weights, Context, Memory and Multi-GPU
- vLLM — Quickstartretrieved 2026-09-09
Sections: Installation; OpenAI-compatible server
Used by:P9 Installing vLLM: x86 CUDA, DGX Spark, ROCm and What Does Not Work
- vLLM — Reasoning outputsretrieved 2026-09-09
Sections: Reasoning parsers; the reasoning field; tool calling with reasoning; Reasoning parsers; the reasoning field; disabling thinking; tool calling
Used by:P24 Function Calling End to End on Local EnginesP24 Reasoning Models in Agent Loops
- vLLM — Speculative decodingretrieved 2026-09-09
Sections: Methods; configuration; method selection; Configuration examples; method selection; limitations; Configuration keys; method selection; limitations; Common configuration keys; method selection
Used by:P9 Speculative Decoding: Draft Models, EAGLE and n-gramP17 Lab: Train and Deploy a Draft for Your ModelP17 Speculative Decoding Revisited: Acceptance Rates and When It PaysP17 Training a Draft Model: Medusa and EAGLE at Home
- vLLM — Structured outputsretrieved 2026-09-08, 2026-09-09
Sections: Request fields; backends; Parameters; backends; OpenAI-compatible response_format
Used by:P9 Tool Calling and Structured Output on the Server SideP10 Structured Output and JSON Mode
- vLLM — Tool callingretrieved 2026-09-09
Sections: Automatic function calling; parsers per model family; tool_choice; Request and response shape; tool_choice; parallel calls; parsers per family; Request and response shape; tool_choice; parallel calls; Automatic function calling; tool_choice
Used by:P9 Tool Calling and Structured Output on the Server SideP24 Function Calling End to End on Local EnginesP24 Lab: A Minimal Agent from ScratchP24 What an Agent Is: The Loop, Tools and State
- vLLM — Using Dockerretrieved 2026-09-09
Sections: Official image; shared memory
Used by:P9 Installing vLLM: x86 CUDA, DGX Spark, ROCm and What Does Not Work
- vLLM — vllm bench serveretrieved 2026-09-09
Sections: Options; Options; percentile reporting; goodput
Used by:P9 Why a Second Kind of Engine: Batching, Paged Attention and ThroughputP9 Lab: Serve a Model to Twenty Concurrent Users
- vLLM — vllm serve CLI referenceretrieved 2026-09-09
Sections: Options; kv-offloading-backend; kv-offloading-size; cpu-offload-gb
Used by:P9 Lab: Serve a Model to Twenty Concurrent UsersP9 Serving with vLLM: Quantised Weights, Context, Memory and Multi-GPUP20 vLLM Multi-Node with Ray: Tensor Parallel Inside, Pipeline Parallel AcrossP22 Lab: KV Cache Offload and SharingP22 Lab: Two-Machine Prefill and Decode with vLLMP22 vLLM Disaggregated Prefill: Connectors and the Proxy
- vLLM documentation — Generative modelsretrieved 2026-09-09
Sections: LLM.generate; LLM.chat
- vLLM documentation — Installationretrieved 2026-09-09
Sections: Supported platforms; hardware plugins
Used by:P5 AMD Ryzen AI Max+ 395: Strix Halo, ROCm and VulkanP5 Apple Silicon: Unified Memory, Metal and MLXP9 Installing vLLM: x86 CUDA, DGX Spark, ROCm and What Does Not Work
- vLLM documentation — Offline inferenceretrieved 2026-09-09
Sections: Overview
docs.wandb.ai
- Weights & Biases documentation — Experimentsretrieved 2026-09-09
documentation.ubuntu.com
- Ubuntu Server documentation — Install a root CA certificate in the trust storeretrieved 2026-09-13
Sections: Install a PEM-format certificate; Uninstall a PEM-format certificate (update-ca-certificates --fresh)
Used by:P7 Lab: A Private Chat Service for Your Home Network
- Ubuntu Server documentation — Install and configure an NFS serverretrieved 2026-09-09
Sections: Installation; configuration; client
Used by:P18 Lab: Build and Measure Your Cluster NetworkP18 The Course Reference Cluster and the Single-Machine Path
- Ubuntu Server documentation — Install NVIDIA driversretrieved 2026-09-13
Sections: ubuntu-drivers list; ubuntu-drivers install
Used by:P5 Lab: Prepare Your Machine
download.pytorch.org
- PyTorch wheel index — cu130, rocm7.2 and cpu indexesretrieved 2026-09-09, 2026-09-12
Sections: torch 2.14.0 wheels listed per index, including linux_aarch64 under cu130; cu126 carries 2.14.0+cu126 and torchvision 0.29.0+cu126; the cu128 index stops at torch 2.11.0; cu128; cu130; rocm6.4; rocm7.0; cpu
Used by:P1 Lab: Your Python Environment and a First Trained ModelP12 Lab: Train a 10M to 125M Parameter Model in an Afternoon
genai.owasp.org
- OWASP LLM01:2025 Prompt Injectionretrieved 2026-09-08, 2026-09-09
Sections: Definition; direct and indirect; prevention and mitigation; Indirect prompt injection; segregating external content; Direct and indirect injection; prevention; Prevention and mitigation; Definition; direct and indirect injection; prevention and mitigation
Used by:P10 Privacy, Security and Serving Beyond localhostP10 Project: A Private Document Question-Answering ServiceP23 Security for Exposed Endpoints and the Model Supply ChainP24 Model Context Protocol: Servers, Clients and TransportsP24 What an Agent Is: The Loop, Tools and StateP26 Agentic Retrieval and Research AgentsP26 Safety: Prompt Injection, Tool Permissions and Human-in-the-Loop
- OWASP Top 10 for LLM Applicationsretrieved 2026-09-08
Sections: The 2025 list
github.com
- Aider benchmark harness — README, benchmark.py and prompts.pyretrieved 2026-09-09, 2026-09-12
Sections: Running the benchmark in Docker
Used by:P4 Reading a Model Card and a BenchmarkP25 Aider with Local Models
- Aider leaderboard page source (columns pass_rate_2 and percent_cases_well_formed)retrieved 2026-09-12
- Aider repository — licenceretrieved 2026-09-09
Sections: LICENSE.txt
Used by:P25 The Landscape: Terminal, Editor and Autonomous Agents
- AutoAWQ repositoryretrieved 2026-09-09
Sections: README; deprecation notice; README; deprecation notice; quant_config
Used by:P16 Lab: Quantise Your Fine-Tune Five Ways and Measure EachP16 Post-Training Quantisation in Depth: GPTQ, AWQ, K-quants, imatrix and HQQ
- Caddy v2.11.4 source — logging.go and modules/caddypki/ca.goretrieved 2026-09-13
Sections: console encoder in an interactive terminal, JSON otherwise; root installation log messages (ca.go, pki.go)
Used by:P7 Lab: A Private Chat Service for Your Home Network
- Crush repositoryretrieved 2026-09-09
Sections: README; Local Models; licence
Used by:P25 The Landscape: Terminal, Editor and Autonomous Agents
- EleutherAI — lm-evaluation-harnessretrieved 2026-09-09
Sections: README; reproducibility and task implementation guidelines; README; reproducibility of publicly available prompts; Install; model backends; example commands
Used by:P28 Capstone 4: Improve a Small ModelP28 Capstone 6: The ReportP16 Evaluation Harnesses: lm-evaluation-harness, lighteval, EvalPlus and Your OwnP16 Lab: Run a Standard Benchmark Suite on Your Model
- EvalPlus repositoryretrieved 2026-09-09
Sections: README; HumanEval+ and MBPP+
Used by:P16 Evaluation Harnesses: lm-evaluation-harness, lighteval, EvalPlus and Your Own
- ExLlamaV3 — READMEretrieved 2026-09-09
Sections: Installation; quantization; what's missing
Used by:P8 Lab: Same Model, Every EngineP8 Specialist Engines: ExLlamaV3, ktransformers and mistral.rs
- ExLlamaV3 — Releasesretrieved 2026-09-09
Used by:P8 Specialist Engines: ExLlamaV3, ktransformers and mistral.rs
- ExLlamaV3 v1.4.8 — conversion/convert_model.py argument definitionsretrieved 2026-09-13
Used by:P8 Lab: Same Model, Every Engine
- ExLlamaV3 v1.4.8 — modules/embedding.py (prefer_cpu)retrieved 2026-09-13
Used by:P8 Lab: Same Model, Every Engine
- exo - API technical referenceretrieved 2026-09-09
Sections: Instance Management; Inference; Complete Endpoint Summary; Instance Management; Benchmarked Chat Completions
Used by:P21 exo: Automatic Partitioning Across Your MacsP21 Lab: A Two-Mac Cluster over Thunderbolt 5
- exo — READMEretrieved 2026-09-09
Sections: RDMA over Thunderbolt 5; Features; Quick Start; Enabling RDMA on macOS; Environment Variables; Benchmarking; Hardware Accelerator Support; Quick Start; Enabling RDMA on macOS; Environment Variables; Benchmarking; Features; Enabling RDMA on macOS; Features; Benchmarks
Used by:P18 Networking for Home Clusters: 2.5 GbE to 200 GbE, Thunderbolt 5 and RDMAP21 exo: Automatic Partitioning Across Your MacsP21 Lab: A Two-Mac Cluster over Thunderbolt 5P21 mlx.distributed: Ring, MPI and RDMA over Thunderbolt 5P21 Reality Check: 'Four Mac Studios Replace a GPU Server'
- exo source - src/exo/main.pyretrieved 2026-09-09
Sections: argument parser
- FastChat rating_systems.py (Bradley-Terry fit, 400-point scale, style features)retrieved 2026-09-12
- FasterDecoding/Medusa repositoryretrieved 2026-09-09
Sections: Training command; data preparation; Training; data preparation; self-distillation
Used by:P17 Lab: Train and Deploy a Draft for Your ModelP17 Training a Draft Model: Medusa and EAGLE at Home
- Gemini CLI repositoryretrieved 2026-09-09
Sections: README; authentication
Used by:P25 The Landscape: Terminal, Editor and Autonomous Agents
- GGUF file format specificationretrieved 2026-09-09, 2026-09-12
Sections: Design goals; File structure; Standardized key-value pairs; naming convention; Design goals; File structure; Metadata; Naming convention; Design goals; File structure
Used by:P2 Parameters, Layers and Model SizeP6 GGUF and Quantisation TypesP6 llama.cpp: The Engine That Runs Everywhere
- Gin v1.10.0 source — logger.goretrieved 2026-09-13
Sections: defaultLogFormatter
Used by:P7 Lab: A Private Chat Service for Your Home Network
- GPTQModel repositoryretrieved 2026-09-09
Sections: README; supported methods; quantisation API
Used by:P16 Post-Training Quantisation in Depth: GPTQ, AWQ, K-quants, imatrix and HQQ
- Hermes-Function-Callingretrieved 2026-09-09
Sections: Prompt format
- HQQ — Half-Quadratic Quantization repositoryretrieved 2026-09-09
Sections: README; supported bits and group sizes
Used by:P16 Post-Training Quantisation in Depth: GPTQ, AWQ, K-quants, imatrix and HQQ
- Hugging Face Transformers v5.16.1 - cache_utils.pyretrieved 2026-09-12
Sections: DynamicLayer.update
Used by:P3 Inference: Prefill, Decode and Why Memory Bandwidth Rules
- Hugging Face Transformers v5.16.1 — attention_interface.md (documentation source at the pinned tag)retrieved 2026-09-12
Sections: Attention backends; Create a new attention function; Pass a custom 4D attention mask; Bidirectional attention
Used by:P2 Attention and the Transformer
- Hugging Face Transformers v5.16.1 — modeling_outputs.py (CausalLMOutputWithPast docstring)retrieved 2026-09-12
Sections: CausalLMOutputWithPast
Used by:P2 Attention and the TransformerP2 Embeddings: Meaning as GeometryP2 From Autocomplete to Assistant: Next-Token Prediction
- Hugging Face Transformers v5.16.1 — modeling_rope_utils.pyretrieved 2026-09-12
Sections: inv_freq computation; _compute_yarn_parameters
Used by:P2 Attention and the Transformer
- Hugging Face Transformers v5.16.1 — Tool use (docs/source/en/chat_extras.md)retrieved 2026-09-12
Sections: passing tools to apply_chat_template; the JSON schema format of a tool definition
- huggingface_hub v1.30.0 — cli/_output.py and utils/_detect_agent.py (human and agent output formats)retrieved 2026-09-12
Sections: OutputFormat; result(); is_agent()
Used by:P2 Lab: Look Inside a Model
- huggingface_hub v1.30.0 source — _local_folder.py, read_download_metadataretrieved 2026-09-12
- huggingface_hub v1.30.0 source — constants.py, utils/_auth.py, file_download.pyretrieved 2026-09-13
Sections: HF_TOKEN_PATH and HF_STORED_TOKENS_PATH under HF_HOME; token written with mode 600; relative cache symlinks survive a move
Used by:P5 Lab: Prepare Your Machine
- huggingface_hub v1.30.0 source — file_download.py, local-directory up-to-date check and _download_to_tmp_and_moveretrieved 2026-09-12
Used by:P4 Lab: Build Your Model ShortlistP4 Model Families and Who Makes Them
- huggingface_hub v1.30.0 source — hf_api.py (get_dataset_leaderboard, DatasetLeaderboardEntry) and _eval_results.py (EvalResultEntry)retrieved 2026-09-12
Used by:P4 Reading a Model Card and a BenchmarkP4 Model Families and Who Makes Them
- Kilo Code repositoryretrieved 2026-09-09
Sections: README; licence; Kilo CLI lineage; README; licence; Kilo CLI
Used by:P25 Editor Agents: Cline, Kilo Code, Continue and the Proprietary EditorsP25 The Landscape: Terminal, Editor and Autonomous Agents
- ktransformers — READMEretrieved 2026-09-09
Sections: Installation; supported models; CPU kernels
Used by:P8 Specialist Engines: ExLlamaV3, ktransformers and mistral.rs
- Lemonade — built-in model registry (server_models.json)retrieved 2026-09-13
Used by:P8 Lab: Same Model, Every Engine
- Lemonade — repository READMEretrieved 2026-09-09
Sections: Supported backends; hardware acceleration; licence
Used by:P8 AMD-Native: ROCm Builds, Lemonade Server and the NPU Question
- Lemonade repository — data/lemond.service.in and data/lemond-user.service.in (systemd units)retrieved 2026-09-13
Used by:P8 Lab: Same Model, Every Engine
- Linux kernel v6.17 — drivers/gpu/drm/amd/amdgpu/amdgpu_ttm.cretrieved 2026-09-13
Sections: gtt_size from ttm_tt_pages_limit; "M of GTT memory ready" (same in v6.14)
Used by:P5 Lab: Prepare Your Machine
- Linux kernel v6.17 — drivers/gpu/drm/ttm/ttm_device.cretrieved 2026-09-13
Sections: ttm_global_init, num_pages /= 2
Used by:P5 Lab: Prepare Your Machine
- linux-rdma/perftest — READMEretrieved 2026-09-09
Sections: Tests; running
Used by:P18 Lab: Build and Measure Your Cluster NetworkP18 Networking for Home Clusters: 2.5 GbE to 200 GbE, Thunderbolt 5 and RDMA
- LiveCodeBench repository README (dataset versions, n=10, pass@1)retrieved 2026-09-12
- Llama 3.1 Acceptable Use Policy, text in meta-llama/llama-modelsretrieved 2026-09-12
Sections: Prohibited Uses
- Llama 3.1 Community License Agreement, text in meta-llama/llama-modelsretrieved 2026-09-12, 2026-09-13
Sections: Acceptance paragraph; 1.a, 1.b.i to 1.b.iv; 2; 5.b, 5.c; 6; 7; 2, Additional Commercial Terms
Used by:P3 Open Weights, Open Source and LicencesP4 Model Families and Who Makes Them
- Llama 3.1 model card (meta-llama/llama-models, MODEL_CARD.md)retrieved 2026-09-12, 2026-09-13
Sections: Training data (~15 trillion tokens); Training Time (GPU hours) table; hardware (H100-80GB, TDP of 700W); licence
Used by:P3 Pretraining: Learning from Trillions of TokensP4 Model Families and Who Makes Them
- LLaMA-Factory — READMEretrieved 2026-09-09
Sections: Features; hardware; CLI; licence
Used by:P13 Unsloth, Axolotl, LLaMA-Factory and mlx-lm: Higher-Level Tools
- llama-swap — example configurationretrieved 2026-09-09
Sections: Macros; models; groups; ttl; unloadTimeout; groups; healthCheckTimeout; checkEndpoint; ttl; unloadTimeout; groups
Used by:P9 Project: Your Local Model GatewayP23 Challenge: The 3 a.m. Out-of-MemoryP23 Routing and Model Management: llama-swap, LiteLLM, NGINX and Health Checks
- llama-swap — READMEretrieved 2026-09-09
Sections: Configuration; command-line flags; endpoints; container images; Container images and tags; endpoints; /running; /api/models/unload; /logs; /metrics; /running; Endpoints; /metrics; /logs; Endpoints; configuration keys; groups
Used by:P9 Project: Your Local Model GatewayP23 Backup, Upgrades and Reproducibility of a Model EstateP23 Challenge: The 3 a.m. Out-of-MemoryP23 Lab: Dashboards for Your ClusterP23 Observability: Metrics, Logs and Traces for LLM ServingP23 Routing and Model Management: llama-swap, LiteLLM, NGINX and Health Checks
- llama.cpp - llama-server README at build b10867retrieved 2026-09-12
Sections: --batch-size, --ubatch-size, --parallel
Used by:P3 Inference: Prefill, Decode and Why Memory Bandwidth Rules
- llama.cpp — Build guideretrieved 2026-09-09
Sections: Vulkan; HIP; CUDA; Vulkan; HIP; Metal; Notes about GPU-accelerated backends; CPU build; CUDA; Metal; Vulkan; HIP; Notes about GPU-accelerated backends; Notes about GPU-accelerated backends; CPU build; CUDA; Metal; Vulkan; HIP; HIP; Vulkan; CUDA; Metal; Vulkan; HIP
Used by:P5 AMD Ryzen AI Max+ 395: Strix Halo, ROCm and VulkanP6 Challenge: The Model That Runs at Two Tokens per SecondP6 Installing and Building llama.cpp on Your PlatformP6 Lab: Run and Benchmark the Course Reference ModelsP6 llama-cli and llama-serverP6 llama.cpp: The Engine That Runs EverywhereP8 AMD-Native: ROCm Builds, Lemonade Server and the NPU QuestionP19 Lab: A Mixed-Platform ClusterP19 llama.cpp RPC: Layers Across Machines
- llama.cpp — common/arg.cppretrieved 2026-09-09
Sections: --rpc and --tensor-split definitions
- llama.cpp — convert_hf_to_gguf.pyretrieved 2026-09-09
Sections: parse_args
Used by:P6 GGUF and Quantisation Types
- llama.cpp — Docker documentationretrieved 2026-09-13
Sections: Images
Used by:P7 Lab: A Private Chat Service for Your Home Network
- llama.cpp — Function callingretrieved 2026-09-09
Sections: Native formats; generic fallback; template overrides; Native formats; generic fallback; parallel_tool_calls; Native tool-call model families; generic fallback
Used by:P9 Tool Calling and Structured Output on the Server SideP24 Function Calling End to End on Local EnginesP25 Which Local Models Can Actually Drive an Agent
- llama.cpp — GBNF grammar guideretrieved 2026-09-08
Sections: Background; syntax; JSON schema conversion; limitations
- llama.cpp — llama-bench READMEretrieved 2026-09-09
Sections: Usage; output columns; Usage; output columns; output formats; Usage; output formats; Usage and options; Usage and options; list-devices
Used by:P6 Challenge: The Model That Runs at Two Tokens per SecondP6 Installing and Building llama.cpp on Your PlatformP6 Lab: Run and Benchmark the Course Reference ModelsP6 llama.cpp: The Engine That Runs EverywhereP16 Lab: Quantise Your Fine-Tune Five Ways and Measure EachP19 Challenge: The Cluster That Is Slower Than One MachineP19 Lab: A Mixed-Platform ClusterP19 Lab: Run a Model Bigger Than Any One MachineP19 RDMA Transport, Tuning and Measuring the Split
- llama.cpp — llama-imatrix READMEretrieved 2026-09-09
Sections: Usage; examples; Options; examples
Used by:P6 GGUF and Quantisation TypesP16 Lab: Quantise Your Fine-Tune Five Ways and Measure EachP16 Post-Training Quantisation in Depth: GPTQ, AWQ, K-quants, imatrix and HQQ
- llama.cpp — llama-perplexity READMEretrieved 2026-09-09
Sections: KL divergence mode; What perplexity measures; KL divergence mode; output fields
Used by:P16 Lab: Quantise Your Fine-Tune Five Ways and Measure EachP16 Measuring Quantisation Damage: Perplexity, KL Divergence and Task Evaluations
- llama.cpp — llama-quantize READMEretrieved 2026-09-09
Sections: Options; quantisation types and bits per weight; Usage; quantisation types; Usage; options; quantisation types; Quantisation types; bits per weight
Used by:P6 GGUF and Quantisation TypesP11 Lab: Your First Training RunP15 Project: The Distillation PipelineP16 Lab: Quantise Your Fine-Tune Five Ways and Measure EachP16 Measuring Quantisation Damage: Perplexity, KL Divergence and Task EvaluationsP16 Post-Training Quantisation in Depth: GPTQ, AWQ, K-quants, imatrix and HQQ
- llama.cpp — llama-server READMEretrieved 2026-09-08, 2026-09-09
Sections: Command-line options; Usage; command-line options; API endpoints; Web UI; Command-line options; sampling parameters; Parallel slots; continuous batching; metrics endpoint; Command-line options; chat templates; grammars; Speculative decoding options; Sampling options; response_format; --alias; /infill endpoint; prompt caching; --api-key and --host; --host; --api-key and --api-key-file; --slots; --props; --embedding and /v1/embeddings; --reranking and /v1/rerank; response_format; Chat template options; prompt caching; /completion parameters; --embedding and /v1/embeddings; --reranking and /v1/rerank; response_format; json_schema and grammar parameters on /completion; --jinja; --chat-template; --chat-template-file; LoRA options; LoRA options; --jinja; --alias; LoRA options; chat template options; OpenAI-compatible endpoints; /completion n_probs; OpenAI-compatible endpoints; OpenAI-compatible endpoints; /completion n_probs; Command-line options; parallel decoding; Command-line options; --alias; Completion endpoint; logprobs; Speculative decoding options; metrics; slots; Prompt caching; KV cache types; slots endpoints; metrics; Speculative decoding options; metrics endpoint; Command-line options; the timings object; Command-line options; GET /metrics; the timings object; GET /metrics; GET /slots; the timings object; split-mode; tensor-split; main-gpu; n-gpu-layers; Prompt caching; KV cache types; slots save and restore; Prompt caching; slots; parallel; --ctx-size; --parallel; --cache-type-k; --cache-type-v; --metrics; /metrics; --metrics; GET /metrics; --metrics; GET /metrics; GET /health; GET /slots; --api-key; --props; --slots; /health; --api-key-file; --props; --slots; --ctx-size; --cache-type-k and --cache-type-v; --cache-prompt; --jinja; --reasoning-format; --reasoning-budget; --reasoning-format; --reasoning-budget; Anthropic-compatible API endpoints; --jinja; --cache-reuse
Used by:P6 Challenge: The Model That Runs at Two Tokens per SecondP6 Lab: Run and Benchmark the Course Reference ModelsP6 llama-cli and llama-serverP6 llama.cpp: The Engine That Runs EverywhereP6 Sampling: Temperature, Top-p, Min-p, Repetition and DeterminismP8 Lab: Same Model, Every EngineP9 Why a Second Kind of Engine: Batching, Paged Attention and ThroughputP9 Lab: Serve a Model to Twenty Concurrent UsersP9 Tool Calling and Structured Output on the Server SideP9 Speculative Decoding: Draft Models, EAGLE and n-gramP10 Lab: Benchmark Local Models on Your Own TasksP10 Local Coding Assistants: Autocomplete and Chat in Your EditorP10 Privacy, Security and Serving Beyond localhostP10 Project: A Private Document Question-Answering ServiceP10 Prompting That Works Locally: System Prompts, Chat Templates and Thinking ModesP10 Retrieval-Augmented Generation: Embeddings, Chunking and Vector StoresP10 Structured Output and JSON ModeP13 Challenge: The Fine-Tune That Got WorseP13 Lab: Fine-Tune a 1B to 4B Model to Follow Your FormatP13 Merging, Exporting and Quantising a Fine-Tuned ModelP14 Lab: DPO a Model to Prefer Your StyleP14 Lab: GRPO on a Maths or Code Task on One MachineP14 The RL Toolchain: TRL, Unsloth, verl, OpenRLHF and vLLM RolloutsP15 Lab: Sequence-Level Distillation of a 30B-Class Teacher into a 4B StudentP15 Project: The Distillation PipelineP15 Generating Synthetic Data with a Local TeacherP16 Challenge: The Benchmark That LiedP16 Measuring Quantisation Damage: Perplexity, KL Divergence and Task EvaluationsP17 Lab: Train and Deploy a Draft for Your ModelP17 Prefix Caching and KV ReuseP17 Speculative Decoding Revisited: Acceptance Rates and When It PaysP19 Challenge: The Cluster That Is Slower Than One MachineP19 Lab: A Mixed-Platform ClusterP19 Lab: Run a Model Bigger Than Any One MachineP19 llama.cpp RPC: Layers Across MachinesP19 RDMA Transport, Tuning and Measuring the SplitP20 Multi-GPU Desktops: PCIe, Tensor Parallel Without NVLink and Expert ParallelP22 Lab: KV Cache Offload and SharingP22 Lab: Two-Machine Prefill and Decode with vLLMP22 vLLM Disaggregated Prefill: Connectors and the ProxyP23 Challenge: The 3 a.m. Out-of-MemoryP23 Lab: Dashboards for Your ClusterP23 Observability: Metrics, Logs and Traces for LLM ServingP23 Routing and Model Management: llama-swap, LiteLLM, NGINX and Health ChecksP23 Security for Exposed Endpoints and the Model Supply ChainP24 Context Engineering: Memory, Compaction and the KV BudgetP24 Function Calling End to End on Local EnginesP24 Reasoning Models in Agent LoopsP25 Claude Code with a Local Endpoint
- llama.cpp — Multimodal supportretrieved 2026-09-08
Sections: libmtmd; --mmproj and -hf; supported model families
Used by:P10 Vision, Speech and Documents: Multimodal Locally
- llama.cpp — READMEretrieved 2026-09-09
Sections: Quick start; Quick start; Supported backends; Description
Used by:P6 Installing and Building llama.cpp on Your PlatformP6 llama-cli and llama-serverP6 llama.cpp: The Engine That Runs Everywhere
- llama.cpp — requirements directoryretrieved 2026-09-09
Used by:P11 Lab: Your First Training Run
- llama.cpp — RPC backend READMEretrieved 2026-09-09
Sections: Status and security warning; --rpc and tensor split; RDMA; Overview; tensor split; Overview; cache; tensor split; Overview; building; RDMA; Overview; Usage; RDMA transport; Troubleshooting; Overview; Usage; Local cache; Troubleshooting; Usage; Local cache; RDMA transport; Troubleshooting; Overview; Usage; Local cache; RDMA transport; Troubleshooting; RDMA transport; Local cache; Troubleshooting; RDMA; usage
Used by:P28 Capstone 3: Cluster or Tiered DeploymentP18 Lab: Build and Measure Your Cluster NetworkP18 Networking for Home Clusters: 2.5 GbE to 200 GbE, Thunderbolt 5 and RDMAP18 Tensor, Pipeline, Expert, Data and Sequence ParallelismP18 Why One Box Runs Out: The Memory Wall and the Bandwidth WallP18 The Course Reference Cluster and the Single-Machine PathP19 Challenge: The Cluster That Is Slower Than One MachineP19 Lab: A Mixed-Platform ClusterP19 Lab: Run a Model Bigger Than Any One MachineP19 llama.cpp RPC: Layers Across MachinesP19 RDMA Transport, Tuning and Measuring the SplitP21 Lab: A Two-Mac Cluster over Thunderbolt 5P21 mlx.distributed: Ring, MPI and RDMA over Thunderbolt 5P21 Reality Check: 'Four Mac Studios Replace a GPU Server'
- llama.cpp — tools/rpc/rpc-server.cppretrieved 2026-09-09
Sections: print_usage and the argument parser
Used by:P19 Challenge: The Cluster That Is Slower Than One MachineP19 Lab: A Mixed-Platform ClusterP19 Lab: Run a Model Bigger Than Any One MachineP19 llama.cpp RPC: Layers Across Machines
- llama.cpp b10868 — llama-context.cppretrieved 2026-09-13
Sections: n_ctx_seq
Used by:P7 Lab: A Private Chat Service for Your Home Network
- llama.cpp b10868 — llama-server READMEretrieved 2026-09-13
Sections: --ctx-size; --parallel; --kv-unified; --alias; GET /health; GET /v1/models
Used by:P7 Lab: A Private Chat Service for Your Home Network
- llama.cpp b10868 — tools/server/server-context.cpp and server-http.cppretrieved 2026-09-13
Sections: launch_slot_with_task "processing task" log line; per-request logger disabled in server-http.cpp
Used by:P7 Lab: A Private Chat Service for Your Home Network
- llama.cpp pull request 15293 — server, add SWA checkpointsretrieved 2026-09-12
Used by:P4 Dense, Mixture-of-Experts and Hybrid Architectures
- llama.cpp releases — b10936 assetsretrieved 2026-09-09, 2026-09-13
Sections: ubuntu-arm64, ubuntu-vulkan-arm64 and ubuntu-x64 archives; no Linux CUDA archive; Assets of the current build tag
Used by:P6 Challenge: The Model That Runs at Two Tokens per SecondP6 Installing and Building llama.cpp on Your PlatformP6 llama.cpp: The Engine That Runs Everywhere
- llama.cpp v0.4.0 - ggml-quants.c (quantize_row_q4_0_ref, quantize_row_q4_0_impl, make_qx_quants)retrieved 2026-09-12
Used by:P3 Distillation, Pruning and Quantisation: How Small Models Get Good
- llama.cpp v0.4.0 - imatrix.cpp (squared activations accumulated per input column)retrieved 2026-09-12
Used by:P3 Distillation, Pruning and Quantisation: How Small Models Get GoodP4 Dense, Mixture-of-Experts and Hybrid Architectures
- llama.cpp v0.4.0 - llama-imatrix READMEretrieved 2026-09-12
Sections: Usage; --show-statistics
Used by:P3 Distillation, Pruning and Quantisation: How Small Models Get Good
- llama.cpp v0.4.0 - llama-quant.cpp (remap_layer and the block count written under --prune-layers; the output tensor's default type)retrieved 2026-09-12
Sections: tensor_allows_quantization; use_more_bits; Q4_K_M rules for attn_v, ffn_down and output; MXFP4_MOE
Used by:P3 Distillation, Pruning and Quantisation: How Small Models Get GoodP4 Choosing a Model for a Memory BudgetP4 Dense, Mixture-of-Experts and Hybrid Architectures
- llama.cpp v0.4.0 - llama-quantize READMEretrieved 2026-09-12
Sections: Options; Advanced options; bits per weight table; examples (the naive Q4_K_M quantisation with default settings)
Used by:P3 Distillation, Pruning and Quantisation: How Small Models Get Good
- llama.cpp v0.4.0 - src/llama-vocab.cppretrieved 2026-09-12
Sections: FIM token detection by text; end-of-generation token detection by text; print_info
Used by:P4 Base, Instruct, Thinking, Coder, Vision, Embedding: Reading a Model Name
- llama.cpp v0.4.0 - tools/server/server-common.cppretrieved 2026-09-12
Sections: oaicompat_chat_params_parse, image input check
Used by:P4 Base, Instruct, Thinking, Coder, Vision, Embedding: Reading a Model Name
- llama.cpp v0.4.0 - tools/server/server-context.cppretrieved 2026-09-12, 2026-09-13
Sections: /infill token check; exceed_context_size_error; slot context capping; context size reduced by --fit; YaRN n_ctx_train adjustment
Used by:P4 Base, Instruct, Thinking, Coder, Vision, Embedding: Reading a Model NameP7 Reality Check: 'The Default Context Is Enough'P8 Lab: Same Model, Every Engine
- llama.cpp v0.4.0 — Build guide at the release tagretrieved 2026-09-13
Sections: CUDA Unified Memory and System Memory Fallback; HIP Unified Memory; Metal; Notes about GPU-accelerated backends; CUDA unified memory and System Memory Fallback; HIP unified memory; Metal and --n-gpu-layers 0; Notes about GPU-accelerated backends (-ngl 0 and --device none)
Used by:P6 Challenge: The Model That Runs at Two Tokens per SecondP6 Lab: Run and Benchmark the Course Reference Models
- llama.cpp v0.4.0 — common/arg.cpp (--ctx-size, --rope-scaling, --rope-scale, --yarn-orig-ctx)retrieved 2026-09-12, 2026-09-13
Sections: --ctx-size; --rope-scaling; --rope-scale; --yarn-orig-ctx
Used by:P2 Attention and the TransformerP8 Lab: Same Model, Every Engine
- llama.cpp v0.4.0 — common/log.h (log levels)retrieved 2026-09-13
Sections: LOG_TRC at LOG_LEVEL_TRACE (4); default threshold LOG_LEVEL_INFO (3)
- llama.cpp v0.4.0 — fitting parameters to device memoryretrieved 2026-09-13
Sections: common_params_fit_impl log lines; common_fit_params; common_memory_breakdown_print; tools/fit-params/README.md; common_memory_breakdown_print; src/llama-kv-cache.cpp KV size line; ggml/src/ggml-common.h block sizes
Used by:P6 Challenge: The Model That Runs at Two Tokens per SecondP6 Lab: Run and Benchmark the Course Reference Models
- llama.cpp v0.4.0 — ggml-cuda.cu (free memory on unified-memory systems) and ggml-metal-device.m (recommendedMaxWorkingSetSize)retrieved 2026-09-12, 2026-09-13
Sections: ggml_backend_cuda_device_get_memory; ggml_backend_cuda_get_available_uma_memory; ggml_metal_device_get_memory; ggml_cuda_device_malloc and GGML_CUDA_ENABLE_UNIFIED_MEMORY; cudaMalloc failed; UMA free memory excluded for HIP; ggml-vulkan.cpp ggml_backend_vk_get_device_memory; ggml-metal-device.m working-set warning; ggml_backend_cuda_device_get_memory (MemAvailable for integrated devices on Linux); ggml/src/ggml-metal/ggml-metal-device.m ggml_metal_device_get_memory; ggml/src/ggml-blas/ggml-blas.cpp device memory; common/arg.cpp common_print_available_devices; src/llama-context.cpp n_ctx padding
Used by:P4 Choosing a Model for a Memory BudgetP6 Challenge: The Model That Runs at Two Tokens per SecondP6 Lab: Run and Benchmark the Course Reference Models
- llama.cpp v0.4.0 — ggml/src/ggml-common.h (block structs and their static_assert sizes)retrieved 2026-09-12
Sections: block_q8_0, block_q4_K, block_q5_K, block_q6_K, block_q4_0, block_iq4_xs, block_mxfp4, block_q3_K, block_q2_K
Used by:P4 Choosing a Model for a Memory BudgetP4 Dense, Mixture-of-Experts and Hybrid Architectures
- llama.cpp v0.4.0 — llama-bench README and sourceretrieved 2026-09-13
Sections: defaults (n_gpu_layers -1, fit target off); llama_null_log_callback without --verbose; get_backend; markdown columns; test_prompt, test_gen and the timed loop; get_ts; get_backend; markdown column selection
Used by:P6 Challenge: The Model That Runs at Two Tokens per SecondP6 Lab: Run and Benchmark the Course Reference Models
- llama.cpp v0.4.0 — llama-bench README at the release tagretrieved 2026-09-13
Sections: Syntax (options and defaults); JSON output example; prefilled context
Used by:P6 Lab: Run and Benchmark the Course Reference Models
- llama.cpp v0.4.0 — llama-server READMEretrieved 2026-09-12, 2026-09-13
Sections: --ctx-size, --cache-type-k, --cache-type-v, --flash-attn, --swa-full, --fit, --fit-target, --fit-ctx, --gpu-layers (default auto), --parallel, --kv-unified, --cache-ram, --ubatch-size, --log-verbosity (default 3); --swa-full, --cpu-moe, --n-cpu-moe, --ctx-checkpoints; --hf-repo; --reasoning-budget; --reasoning-budget-message; --embedding; --pooling; POST /reranking; --n-gpu-layers auto; --fit, --fit-target, --fit-ctx; --ctx-size; --no-kv-offload; --cache-ram; --parallel; tools/cli/README.md --single-turn; --ctx-size, --fit, --parallel, --verbose, --host, --port; GET /health; POST /v1/chat/completions (reasoning_effort); GET /props; --ctx-size, --parallel, --fit, --context-shift, --rope-scaling, --rope-scale, --yarn-orig-ctx; OpenAI-compatible Chat Completions API; Tool call support; Timings and context usage; --cache-prompt; --reasoning-format; --parallel; --kv-unified; --cache-prompt; --cache-ram; --slot-prompt-similarity; GET /slots; GET /metrics (available metrics); timings and cache_n
Used by:P4 Choosing a Model for a Memory BudgetP4 Dense, Mixture-of-Experts and Hybrid ArchitecturesP4 Base, Instruct, Thinking, Coder, Vision, Embedding: Reading a Model NameP6 Challenge: The Model That Runs at Two Tokens per SecondP6 Lab: Run and Benchmark the Course Reference ModelsP7 Reality Check: 'The Default Context Is Enough'P8 Lab: Same Model, Every EngineP9 Lab: Serve a Model to Twenty Concurrent Users
- llama.cpp v0.4.0 — log verbosity and argument handlingretrieved 2026-09-13
Sections: common_log_get_verbosity (library INFO logged at trace level 4, default threshold 3); common/arg.cpp no usable GPU warning, --ctx-size 0 and fit_params_min_ctx
Used by:P6 Challenge: The Model That Runs at Two Tokens per Second
- llama.cpp v0.4.0 — src/llama-context.cpp (n_ctx_train warning)retrieved 2026-09-12
Sections: llama_context constructor; n_ctx_seq > n_ctx_train; n_ctx padding to 256; n_ctx_seq for unified and split caches; quantized V cache requires Flash Attention
Used by:P2 Attention and the TransformerP4 Choosing a Model for a Memory Budget
- llama.cpp v0.4.0 — src/llama-kv-cache.cpp and src/llama-kv-cache-iswa.cppretrieved 2026-09-12
Sections: SWA cache size; llama_kv_cache size log line; failed to allocate buffer for kv cache
- llama.cpp v0.4.0 — src/llama-kv-cache.cpp and src/llama-model.cpp (load log formats)retrieved 2026-09-13
Used by:P8 Lab: Same Model, Every Engine
- llama.cpp v0.4.0 — src/llama-model.cpp (layers placed on the GPU)retrieved 2026-09-13
Sections: n_gpu_layers below 0 means every layer plus the output layer; "offloaded %d/%d layers to GPU"; load_tensors (i_gpu_start, input layer on the CPU, offloaded N/M layers, model buffer size); src/llama.cpp llama_supports_gpu_offload; src/models/qwen3.cpp tied output
Used by:P4 Choosing a Model for a Memory BudgetP6 Challenge: The Model That Runs at Two Tokens per Second
- llama.cpp v0.4.0 — tools/fit-params/README.md and common/fit.cpp (memory breakdown table; layer overflow)retrieved 2026-09-13
Sections: context reduction no lower than the minimum context; layer overflow to system memory, MoE tensors first; "n_gpu_layers already set by user"; breakdown rows printed with LOG_TRC
- llama.cpp v0.4.0 — tools/server/server.cpp (automatic slot count; memory breakdown on shutdown)retrieved 2026-09-13
Sections: common_memory_breakdown_print called after the main loop ends
- llama.vim — local LLM-assisted text completionretrieved 2026-09-08
Sections: README; recommended models by memory; context reuse
Used by:P10 Local Coding Assistants: Autocomplete and Chat in Your Editor
- llama.vscode — local LLM-assisted text completion for VS Coderetrieved 2026-09-08
Sections: README
Used by:P10 Local Coding Assistants: Autocomplete and Chat in Your Editor
- LLM Compressor — AWQ exampleretrieved 2026-09-09
Sections: Recipe; oneshot
Used by:P16 Lab: Quantise Your Fine-Tune Five Ways and Measure Each
- LLM-Pruner repositoryretrieved 2026-09-09
Sections: README; supported models; update history
- llm.c — READMEretrieved 2026-09-09
Sections: Overview; quick start (CPU)
Used by:P12 Anatomy of a Training Loop: nanochat Line by Line
- MCP Python SDK — READMEretrieved 2026-09-09
Sections: Installation; the server in 15 lines; the client in 10 lines
- meta-llama/llama-models — models/llama3/tokenizer.pyretrieved 2026-09-12
Sections: pat_str; num_reserved_special_tokens; special token names
- Mistral Vibe repositoryretrieved 2026-09-09
Sections: README; custom domains; licence
Used by:P25 The Landscape: Terminal, Editor and Autonomous Agents
- mistral.rs — READMEretrieved 2026-09-09
Sections: Installation; server; quantisation; tool calling
Used by:P8 Specialist Engines: ExLlamaV3, ktransformers and mistral.rs
- MLX LM — generate.py at v0.31.3retrieved 2026-09-12
Sections: setup_arg_parser; chat template application
Used by:P2 Lab: Look Inside a Model
- MLX LM — models/qwen3.py and models/base.py at v0.31.3retrieved 2026-09-12
Sections: Attention.__call__; scaled_dot_product_attention
Used by:P2 Lab: Look Inside a Model
- MLX source - python/mlx/_distributed_utils/config.pyretrieved 2026-09-09
Sections: IPConfigurator.setup; extract_connectivity; argument parser
- MLX source - python/mlx/_distributed_utils/launch.pyretrieved 2026-09-09
Sections: argument parser
Used by:P21 mlx.distributed: Ring, MPI and RDMA over Thunderbolt 5
- mlx-examples — transformer_lmretrieved 2026-09-09
Sections: Transformer language model training
Used by:P12 Lab: Train a 10M to 125M Parameter Model in an Afternoon
- mlx-lm - distributed inference exampleretrieved 2026-09-09
Sections: docstring; arguments
Used by:P21 Lab: A Two-Mac Cluster over Thunderbolt 5P21 mlx.distributed: Ring, MPI and RDMA over Thunderbolt 5
- mlx-lm — generate.py argument parserretrieved 2026-09-09
Sections: setup_arg_parser
Used by:P17 Lab: Train and Deploy a Draft for Your ModelP17 Prefix Caching and KV Reuse
- mlx-lm — LoRA and QLoRA fine-tuningretrieved 2026-09-09
Sections: Run; Data; Evaluate; Generate; Fuse; Run; Data; Run; fuse; data format; Run; fine-tune type; data format; fuse; Fuse; Run; Data; Fuse; Data; tools format; mask-prompt; num-layers; Data; tools format; mask-prompt; num-layers; fuse
Used by:P11 Lab: Your First Training RunP11 Setting Up a Training Environment on Each PlatformP13 Challenge: The Fine-Tune That Got WorseP13 Lab: Fine-Tune a 1B to 4B Model to Follow Your FormatP13 Merging, Exporting and Quantising a Fine-Tuned ModelP13 Project: A Specialist AssistantP13 Unsloth, Axolotl, LLaMA-Factory and mlx-lm: Higher-Level ToolsP15 Lab: Sequence-Level Distillation of a 30B-Class Teacher into a 4B StudentP15 Project: The Distillation PipelineP27 Fine-Tuning for Tool Use and Your CodebaseP27 Lab: Fine-Tune a Small Model on Your Own Agent Trajectories and Measure the Gain
- mlx-lm — READMEretrieved 2026-09-09, 2026-09-12
Sections: Conversion and quantisation; LoRA fine-tuning; conversion and quantisation; Quick Start; Python API; Command Line; Installation; generate; convert; Python API; MLX Community; Feature list; command line tools; Prompt caching
Used by:P28 Capstone 1: Hardware and Model PlanP28 Capstone 4: Improve a Small ModelP2 Lab: Look Inside a ModelP8 MLX and mlx-lm: Apple's Native PathP14 Lab: DPO a Model to Prefer Your StyleP14 The RL Toolchain: TRL, Unsloth, verl, OpenRLHF and vLLM RolloutsP17 Prefix Caching and KV Reuse
- mlx-lm — server documentationretrieved 2026-09-09
Sections: Starting the server; endpoints; request fields
Used by:P8 Lab: Same Model, Every EngineP8 MLX and mlx-lm: Apple's Native Path
- mlx-lm v0.31.3 — models/cache.py (KVCache step)retrieved 2026-09-13
Used by:P8 Lab: Same Model, Every Engine
- mlx-lm v0.31.3 — server.py (options, model resolution, prompt cache, reasoning and usage fields)retrieved 2026-09-13
Sections: --decode-concurrency; --prompt-concurrency; --prompt-cache-size; batchable requests; usage and cached_tokens
Used by:P8 Lab: Same Model, Every EngineP9 Lab: Serve a Model to Twenty Concurrent Users
- mlx-vlmretrieved 2026-09-08
Sections: README; installation; CLI and server
Used by:P10 Vision, Speech and Documents: Multimodal Locally
- Moby (Docker Engine) source — daemon/runtime_unix.go and daemon/errors.goretrieved 2026-09-13
Sections: error strings "unknown or invalid runtime name: %s" and "could not select device driver %q with capabilities: %v"
Used by:P5 Lab: Prepare Your Machine
- modded-nanogpt — READMEretrieved 2026-09-09
Sections: Overview; Overview; techniques list
Used by:P12 Anatomy of a Training Loop: nanochat Line by LineP12 Scaling Laws and Compute Budgets at Home
- Mooncake — repository READMEretrieved 2026-09-09
Sections: Integrations; Transfer Engine; Mooncake Store; supported transports; integrations
Used by:P22 SGLang PD Disaggregation and NVIDIA DynamoP22 The KV Cache as a Transferable Object: Mooncake, LMCache and NIXL
- nanochat — READMEretrieved 2026-09-09
Sections: Overview; Getting started; Research; Running on CPU / MPS; Precision / dtype; File structure; Acknowledgements; Time-to-GPT-2 Leaderboard; Getting started; Research; Running on CPU / MPS; Precision / dtype; Time-to-GPT-2 Leaderboard; Research; Overview; Time-to-GPT-2 Leaderboard; Getting started; Running on CPU / MPS
Used by:P12 Anatomy of a Training Loop: nanochat Line by LineP12 Data for Pretraining: FineWeb-Edu, Cosmopedia and Training a TokeniserP12 Lab: Train a 10M to 125M Parameter Model in an AfternoonP12 Scaling Laws and Compute Budgets at HomeP12 What Pretraining Teaches You That Fine-Tuning Cannot
- nanoGPT — READMEretrieved 2026-09-09
Sections: Quick start; Reproducing GPT-2
Used by:P12 Anatomy of a Training Loop: nanochat Line by Line
- NIXL — NVIDIA Inference Xfer Libraryretrieved 2026-09-09
Sections: Overview; plugins; prerequisites
Used by:P22 The KV Cache as a Transferable Object: Mooncake, LMCache and NIXL
- NVIDIA DGX Spark playbooks — Connect Two Sparksretrieved 2026-09-09
Sections: Network interface configuration; troubleshooting; Prerequisites; physical hardware connection; Overview; physical hardware connection; network interface configuration
Used by:P19 Challenge: The Cluster That Is Slower Than One MachineP19 Lab: Run a Model Bigger Than Any One MachineP19 RDMA Transport, Tuning and Measuring the Split
- NVIDIA Model Optimizer (TensorRT Model Optimizer) repositoryretrieved 2026-09-09
Sections: README; quantization formats; QAT
Used by:P16 Quantisation-Aware Training and the Blackwell Formats: FP8, NVFP4 and MXFP4
- NVlabs/Minitron repositoryretrieved 2026-09-09
Sections: README; released models
- Ollama - API documentationretrieved 2026-09-09, 2026-09-12
Sections: Generate a chat completion (parameters, think, response fields, tokens-per-second formula); List Local Models; List Running Models; Version; Chat request with tools; format parameter
Used by:P3 Reality Check: 'A Small Local Model Is as Good as the Frontier'P9 Tool Calling and Structured Output on the Server Side
- Ollama - README (install commands and quickstart)retrieved 2026-09-12
Used by:P3 Reality Check: 'A Small Local Model Is as Good as the Frontier'
- Ollama documentation - FAQretrieved 2026-09-12
Sections: Exposing Ollama on the network (default bind address); How do I know if my model was loaded onto the GPU; context window size; where models are stored; keep-alive
Used by:P3 Reality Check: 'A Small Local Model Is as Good as the Frontier'
- Ollama documentation - GPUretrieved 2026-09-12
Sections: AMD ROCm supported GPUs (gfx1151); Vulkan; Metal
Used by:P3 Reality Check: 'A Small Local Model Is as Good as the Frontier'
- Ollama documentation - Linuxretrieved 2026-09-12
Sections: Install; manual install (amd64, arm64, ROCm); systemd service; logs; uninstall
Used by:P3 Reality Check: 'A Small Local Model Is as Good as the Frontier'
- Ollama documentation - macOSretrieved 2026-09-12
Sections: Install; CLI link; file locations and logs
Used by:P3 Reality Check: 'A Small Local Model Is as Good as the Frontier'
- Ollama documentation - Modelfile referenceretrieved 2026-09-12
Sections: Valid parameters and values
Used by:P3 Reality Check: 'A Small Local Model Is as Good as the Frontier'
- Ollama documentation - Thinkingretrieved 2026-09-12
Sections: Supported models; the think field; message.thinking; CLI quick reference
Used by:P3 Reality Check: 'A Small Local Model Is as Good as the Frontier'
- Ollama documentation - Troubleshootingretrieved 2026-09-12
Sections: Log locations per platform
Used by:P3 Reality Check: 'A Small Local Model Is as Good as the Frontier'
- Ollama documentation - Windowsretrieved 2026-09-12
Sections: Install; environment variables; file locations and logs
Used by:P3 Reality Check: 'A Small Local Model Is as Good as the Frontier'
- Ollama documentation v0.33.3 — Context lengthretrieved 2026-09-13
Sections: defaults by memory; App slider; ollama ps
Used by:P7 Lab: A Private Chat Service for Your Home Network
- Ollama documentation v0.33.3 — GPUretrieved 2026-09-13
Sections: AMD, SELinux (container_use_devices)
Used by:P7 Lab: A Private Chat Service for Your Home Network
- Ollama documentation v0.33.3 — Troubleshootingretrieved 2026-09-13
Sections: server logs on Mac (~/.ollama/logs/server.log)
Used by:P7 Lab: A Private Chat Service for Your Home Network
- Ollama releases - v0.33.3retrieved 2026-09-12
Used by:P3 Reality Check: 'A Small Local Model Is as Good as the Frontier'
- Ollama source - envconfig/config.go at v0.33.3retrieved 2026-09-12
Sections: OLLAMA_CONTEXT_LENGTH default
Used by:P3 Reality Check: 'A Small Local Model Is as Good as the Frontier'
- Ollama source - llm/server.go at v0.33.3retrieved 2026-09-12
Sections: DoneReason values
Used by:P3 Reality Check: 'A Small Local Model Is as Good as the Frontier'
- Ollama source - scripts/install.sh at v0.33.3retrieved 2026-09-12
Sections: BINDIR selection and the ollama symlink
Used by:P3 Reality Check: 'A Small Local Model Is as Good as the Frontier'
- Ollama source at v0.33.3 — llm/server.go, llm/llama_server.go and server/sched.go (trained-context cap, llama-server arguments, automatic context reduction)retrieved 2026-09-13
- Ollama source at v0.33.3 — server/routes.go (VRAM-based default context) and server/prompt.go (chat history truncation)retrieved 2026-09-13
- Ollama v0.33.3 source — server/routes.go and go.modretrieved 2026-09-13
Sections: gin.Default() request logger; gin v1.10.0
Used by:P7 Lab: A Private Chat Service for Your Home Network
- Open WebUI v0.11.3 source — AdvancedParams.svelte (num_ctx for Ollama)retrieved 2026-09-13
- Open WebUI v0.11.3 source — authentication routerretrieved 2026-09-13
Sections: signup; add_user
Used by:P7 Lab: A Private Chat Service for Your Home Network
- Open WebUI v0.11.3 source — command-line entry pointretrieved 2026-09-13
Sections: serve (host and port defaults)
Used by:P7 Lab: A Private Chat Service for Your Home Network
- Open WebUI v0.11.3 source — configurationretrieved 2026-09-13
Sections: run_migrations (Alembic at start-up)
Used by:P7 Lab: A Private Chat Service for Your Home Network
- Open WebUI v0.11.3 source — database connectionretrieved 2026-09-13
Sections: DATABASE_ENABLE_SQLITE_WAL, journal_mode
Used by:P7 Lab: A Private Chat Service for Your Home Network
- Open WebUI v0.11.3 source — model filteringretrieved 2026-09-13
Sections: get_filtered_models
Used by:P7 Lab: A Private Chat Service for Your Home Network
- Open WebUI v0.11.3 source — pyproject.tomlretrieved 2026-09-13
Sections: dependencies (onnxruntime==1.26.0); requires-python
Used by:P7 Lab: A Private Chat Service for Your Home Network
- OpenAI — Harmony response formatretrieved 2026-09-09
Sections: Channels; tool namespaces; Channels; instruction hierarchy
Used by:P24 Function Calling End to End on Local EnginesP24 Reasoning Models in Agent Loops
- OpenAI Codex CLI repositoryretrieved 2026-09-09
Sections: README; licence
Used by:P25 OpenAI Codex CLI and OpenCode with Local ModelsP25 The Landscape: Terminal, Editor and Autonomous Agents
- openai/gpt-2 — src/encoder.pyretrieved 2026-09-12
Sections: bytes_to_unicode; the pre-tokenisation pattern; bpe merge by lowest rank
- OpenHands repositoryretrieved 2026-09-09
Sections: Licence; quickstart container command
Used by:P25 Goose and OpenHands: Autonomous Agents and Sandboxing
- OpenRLHF — READMEretrieved 2026-09-09
Sections: Overview; supported algorithms; Ray, vLLM and DeepSpeed; licence
Used by:P14 The RL Toolchain: TRL, Unsloth, verl, OpenRLHF and vLLM Rollouts
- OpenTelemetry — GenAI semantic conventions repositoryretrieved 2026-09-09
Sections: Repository description; spans, metrics and events for GenAI clients and MCP
Used by:P26 Evaluating Agents: Trajectories, Success Rates and Cost
- Podman — rootless tutorialretrieved 2026-09-09
Used by:P25 Lab: Sandbox Your Agent: Containers, Permissions and Secrets
- Prometheus node exporter — READMEretrieved 2026-09-09
Sections: Default port; running in Docker; collectors
Used by:P23 Lab: Dashboards for Your ClusterP23 Observability: Metrics, Logs and Traces for LLM Serving
- Qwen Code — model providersretrieved 2026-09-09
Sections: Local self-hosted models
Used by:P25 The Landscape: Terminal, Editor and Autonomous Agents
- Qwen3-Coder repository READMEretrieved 2026-09-12
Sections: Fill in the middle with Qwen3-Coder
Used by:P4 Base, Instruct, Thinking, Coder, Vision, Embedding: Reading a Model Name
- ROCm SMI — rocm_smi.py (--showmeminfo)retrieved 2026-09-13
Used by:P8 Lab: Same Model, Every Engine
- Roo Code repositoryretrieved 2026-09-09
Sections: Archive notice; README; README; archive notice
Used by:P25 Editor Agents: Cline, Kilo Code, Continue and the Proprietary EditorsP25 The Landscape: Terminal, Editor and Autonomous Agents
- SafeAILab/EAGLE repositoryretrieved 2026-09-09
Sections: Training; hardware; Training; hardware; SpecForge recommendation
Used by:P17 Lab: Train and Deploy a Draft for Your ModelP17 Training a Draft Model: Medusa and EAGLE at Home
- safetensors — format specificationretrieved 2026-09-08, 2026-09-12
Sections: Format; Format; Notes
Used by:P2 Lab: Look Inside a ModelP2 Parameters, Layers and Model Size
- SentencePiece — READMEretrieved 2026-09-12
Sections: Description; whitespace escaped as ▁ (U+2581); lossless detokenisation
- SkyRLretrieved 2026-09-09
Sections: README; components; releases
Used by:P27 Reinforcement Learning on Agent Tasks: Tests as Rewards
- SpecForge — training guideretrieved 2026-09-09
Sections: Training entry point; configuration; data sources
Used by:P17 Lab: Train and Deploy a Draft for Your ModelP17 Training a Draft Model: Medusa and EAGLE at Home
- SWE-bench repository READMEretrieved 2026-09-12
- TabbyAPI — common/auth.py (api_tokens.yml, bearer keys)retrieved 2026-09-13
Used by:P8 Lab: Same Model, Every Engine
- TabbyAPI — config_sample.ymlretrieved 2026-09-13
Used by:P8 Lab: Same Model, Every Engine
- TabbyAPI — READMEretrieved 2026-09-09
Sections: Features; Docker; licence
Used by:P8 Lab: Same Model, Every EngineP8 Specialist Engines: ExLlamaV3, ktransformers and mistral.rs
- TabbyAPI — Tool calling documentationretrieved 2026-09-13
Used by:P8 Lab: Same Model, Every Engine
- TensorRT-LLM v1.2.1 — serve/openai_protocol.py (response_format handling)retrieved 2026-09-13
Used by:P8 Lab: Same Model, Every Engine
- TensorRT-LLM v1.2.1 — serve/openai_server.py (model name reported for a local directory)retrieved 2026-09-13
Sections: lines 128-132
Used by:P8 Lab: Same Model, Every Engine
- Terminal-Bench repositoryretrieved 2026-09-12
- torchtune repositoryretrieved 2026-09-09
Sections: Maintenance notice
Used by:P13 Unsloth, Axolotl, LLaMA-Factory and mlx-lm: Higher-Level Tools
- transformers v5.16.1 - Qwen2-VL image processor (smart_resize)retrieved 2026-09-12
Used by:P4 Base, Instruct, Thinking, Coder, Vision, Embedding: Reading a Model Name
- transformers v5.16.1 — configuration_qwen3.py (initializer_range)retrieved 2026-09-12
- Transformers v5.16.1 — docs/source/en/chat_templating.md at the tagretrieved 2026-09-12
Sections: Using apply_chat_template; add_generation_prompt; Model training
Used by:P2 From Autocomplete to Assistant: Next-Token PredictionP3 Post-Training: SFT, Preference Tuning and Reinforcement Learning
- Transformers v5.16.1 — docs/source/en/generation_strategies.md at the tagretrieved 2026-09-12
Used by:P2 From Autocomplete to Assistant: Next-Token Prediction
- Transformers v5.16.1 — modeling_gpt_oss.py (GptOssTopKRouter)retrieved 2026-09-12
Used by:P4 Dense, Mixture-of-Experts and Hybrid Architectures
- Transformers v5.16.1 — modeling_nemotron_h.py (NemotronHTopkRouter, Mamba-2 mixer and state shapes)retrieved 2026-09-12
Used by:P4 Dense, Mixture-of-Experts and Hybrid Architectures
- Transformers v5.16.1 — modeling_qwen3_moe.py (Qwen3MoeTopKRouter, load_balancing_loss_func)retrieved 2026-09-12
Used by:P4 Dense, Mixture-of-Experts and Hybrid Architectures
- Transformers v5.16.1 — modeling_qwen3_moe.py, modeling_gpt_oss.py and cache_utils.py (tensor layout and the DynamicCache API)retrieved 2026-09-12
- Transformers v5.16.1 — modeling_qwen3_next.py (Qwen3NextSparseMoeBlock, Qwen3NextGatedDeltaNet, torch_recurrent_gated_delta_rule)retrieved 2026-09-12
Used by:P4 Dense, Mixture-of-Experts and Hybrid Architectures
- transformers v5.16.1 — modeling_qwen3.py (Qwen3MLP, Qwen3RMSNorm, the norms in a block)retrieved 2026-09-12
Sections: Qwen3RMSNorm; Qwen3MLP; rotate_half; apply_rotary_pos_emb; repeat_kv; eager_attention_forward; Qwen3Attention; Qwen3DecoderLayer; Qwen3ForCausalLM.forward - logits_to_keep, past_key_values
Used by:P1 Neural Networks, Activations and BackpropagationP2 Attention and the TransformerP2 Embeddings: Meaning as GeometryP2 From Autocomplete to Assistant: Next-Token PredictionP3 Inference: Prefill, Decode and Why Memory Bandwidth Rules
- transformers v5.16.1 — modeling_utils.py (PreTrainedModel._init_weights: std = config.initializer_range, init.normal_)retrieved 2026-09-12
- Transformers v5.16.1 — src/transformers/loss/loss_utils.py (ForCausalLMLoss, the label shift)retrieved 2026-09-12
Used by:P2 From Autocomplete to Assistant: Next-Token Prediction
- transformers v5.16.1 source — NemotronH configuration (hybrid_override_pattern letters)retrieved 2026-09-12
- TRL — GRPOTrainer and GRPOConfig source at v1.12.0retrieved 2026-09-12
Sections: advantage computation (nanstd, scale_rewards, + 1e-4); grpo_config.py scale_rewards default
Used by:P3 Post-Training: SFT, Preference Tuning and Reinforcement Learning
- TRL — SFTConfig source at v1.12.0retrieved 2026-09-12
Sections: completion_only_loss; assistant_only_loss; loss_type
Used by:P3 Post-Training: SFT, Preference Tuning and Reinforcement Learning
- TRL v1.12.0 — trl/trainer/sft_trainer.py (completion_only_loss docstring, prompt labels set to -100)retrieved 2026-09-12
Used by:P2 From Autocomplete to Assistant: Next-Token Prediction
- verl — READMEretrieved 2026-09-09
Sections: Overview; supported algorithms; backends; hardware; licence
Used by:P14 The RL Toolchain: TRL, Unsloth, verl, OpenRLHF and vLLM Rollouts
- vLLM — disaggregated serving examplesretrieved 2026-09-09
Sections: README; disagg_proxy_demo.py; README
Used by:P22 Lab: Two-Machine Prefill and Decode with vLLMP22 Project: A Tiered Inference ArchitectureP22 vLLM Disaggregated Prefill: Connectors and the Proxy
- vLLM — example connector, prefill exampleretrieved 2026-09-09
Sections: KVTransferConfig; shared_storage_path
Used by:P22 Lab: Two-Machine Prefill and Decode with vLLMP22 vLLM Disaggregated Prefill: Connectors and the Proxy
- vLLM v0.28.0 — docs/configuration/optimization.md and docs/features/automatic_prefix_caching.mdretrieved 2026-09-13
Sections: Limits
- vLLM v0.28.0 — Reasoning outputsretrieved 2026-09-13
Used by:P8 Lab: Same Model, Every Engine
- vLLM v0.28.0 — Tool callingretrieved 2026-09-13
Sections: Qwen models
Used by:P8 Lab: Same Model, Every Engine
- vLLM v0.28.0 — vllm/benchmarks/serve.py and vllm/benchmarks/datasets/datasets.pyretrieved 2026-09-13
Sections: add_cli_args; TPOT definition; result printout; random dataset options
- vLLM v0.28.0 — vllm/config/cache.pyretrieved 2026-09-12, 2026-09-13
Sections: gpu_memory_utilization; enable_prefix_caching default; kv_cache_memory_bytes; cache_dtype; DEFAULT_BLOCK_SIZE
Used by:P4 Choosing a Model for a Memory BudgetP8 Lab: Same Model, Every EngineP9 Lab: Serve a Model to Twenty Concurrent Users
- vLLM v0.28.0 — vllm/config/model.py (_get_and_verify_max_len)retrieved 2026-09-12
Sections: _get_and_verify_max_len; VLLM_ALLOW_LONG_MAX_MODEL_LEN
Used by:P2 Attention and the Transformer
- vLLM v0.28.0 — vllm/entrypoints/openai/cli_args.pyretrieved 2026-09-13
Sections: enable_prompt_tokens_details
- vLLM v0.28.0 — vllm/transformers_utils/model_arch_config_convertor.py (derived max length and its key)retrieved 2026-09-12
Sections: possible_keys; max_len_key
Used by:P2 Attention and the Transformer
- vLLM v0.28.0 — vllm/v1/core/kv_cache_utils.pyretrieved 2026-09-13
Sections: get_num_blocks; get_max_concurrency_for_kv_cache_config; "GPU KV cache size" log line; insufficient-memory errors
- vLLM v0.28.0 — vllm/v1/metrics/loggers.pyretrieved 2026-09-13
Sections: periodic stats line; metric names and labels
- vLLM v0.28.0 — vllm/v1/worker/gpu_worker.py and vllm/v1/worker/utils.pyretrieved 2026-09-13
Sections: kv_cache_memory_bytes path; "Available KV cache memory"; request_memory error
- whisper.cppretrieved 2026-09-08
Sections: README; licence; supported platforms
Used by:P10 Vision, Speech and Documents: Multimodal Locally
gmktec.com
- GMKtec EVO-X2 AI Mini PC product pageretrieved 2026-09-09, 2026-09-13
Sections: Specifications; Onboard LPDDR5X (non-upgradeable), 8000MHz
Used by:P5 AMD Ryzen AI Max+ 395: Strix Halo, ROCm and VulkanP5 Lab: Measure Your Memory Bandwidth and Compute
goose-docs.ai
- Goose — CLI commandsretrieved 2026-09-09
Sections: goose run; goose session; flags; goose run; recipe; max-turns; no-session
Used by:P25 Goose and OpenHands: Autonomous Agents and SandboxingP25 Lab: One Task, Six Agents
- Goose — environment variablesretrieved 2026-09-09
Sections: GOOSE_PROVIDER; GOOSE_MODE; GOOSE_MAX_TURNS
Used by:P25 Goose and OpenHands: Autonomous Agents and Sandboxing
- Goose — providersretrieved 2026-09-09
Sections: OpenAI-compatible; Ollama; tool-calling requirement
Used by:P25 Goose and OpenHands: Autonomous Agents and SandboxingP25 The Landscape: Terminal, Editor and Autonomous Agents
- Goose — recipe referenceretrieved 2026-09-09
Sections: Top-level keys; extensions block
Used by:P25 Goose and OpenHands: Autonomous Agents and Sandboxing
- Goose — using extensionsretrieved 2026-09-09
Sections: Built-in extensions; adding an MCP server; Adding an MCP server
Used by:P25 Goose and OpenHands: Autonomous Agents and SandboxingP25 Project: A Local Agentic Coding Workstation
grafana.com
- Grafana — Configure Grafana with Dockerretrieved 2026-09-09
Sections: Environment variables; GF_SECURITY_ADMIN_PASSWORD
- Grafana — Provisioningretrieved 2026-09-09
Sections: Dashboards from files; Data sources; dashboards; environment variables
Used by:P23 Backup, Upgrades and Reproducibility of a Model EstateP23 Lab: Dashboards for Your ClusterP23 Observability: Metrics, Logs and Traces for LLM Serving
gutenberg.org
- Project Gutenberg — Permission, How Toretrieved 2026-09-09
Sections: Public domain in the US; trademark; other countries
hub.docker.com
- Docker Hub — caddy official imageretrieved 2026-09-13
Sections: /data and /config; do not mount the Caddyfile directly
Used by:P7 Lab: A Private Chat Service for Your Home Network
- Docker Hub — nvidia/cuda tag 13.0.1-devel-ubuntu24.04retrieved 2026-09-13
Sections: compressed size per architecture (arm64 3,904,944,943 bytes)
Used by:P5 Lab: Prepare Your Machine
huggingface.co
- Accelerate documentation — indexretrieved 2026-09-09
Used by:P11 PyTorch, Transformers, Datasets and the Hugging Face Ecosystem
- bartowski/Meta-Llama-3.1-8B-Instruct-GGUF file listingretrieved 2026-09-12
- bartowski/Meta-Llama-3.1-8B-Instruct-GGUF model cardretrieved 2026-09-12
Used by:P4 Lab: Build Your Model ShortlistP4 Base, Instruct, Thinking, Coder, Vision, Embedding: Reading a Model Name
- bitsandbytes — Installationretrieved 2026-09-09
Sections: AMD ROCm; CPU; NVIDIA CUDA; AMD ROCm; CPU
Used by:P13 Lab: Fine-Tune a 1B to 4B Model to Follow Your FormatP13 LoRA and QLoRA ExplainedP13 Project: A Specialist Assistant
- bitsandbytes documentation — 8-bit optimizersretrieved 2026-09-09
Used by:P11 Memory Arithmetic for Training: Weights, Gradients, Optimiser States and Activations
- Cosmopedia dataset cardretrieved 2026-09-09
Sections: Dataset description; Dataset splits; Dataset creation
Used by:P12 Data for Pretraining: FineWeb-Edu, Cosmopedia and Training a Tokeniser
- Datasets documentation — Loadretrieved 2026-09-09
Sections: JSON; Local and remote files; Hugging Face Hub; Local and remote files; JSON
Used by:P11 Datasets: Formats, Chat Templates, Tokenisation and PackingP11 PyTorch, Transformers, Datasets and the Hugging Face Ecosystem
- Datasets documentation — Processretrieved 2026-09-09
Sections: Split; Shuffle; Map; Map; Split; Shuffle
Used by:P11 Datasets: Formats, Chat Templates, Tokenisation and PackingP11 PyTorch, Transformers, Datasets and the Hugging Face Ecosystem
- DeepSeek on Hugging Face — organisation pageretrieved 2026-09-08
- DeepSeek-R1 model cardretrieved 2026-09-09
Sections: License
- DeepSeek-R1-Distill-Qwen-7B and Qwen2.5-Math-7B tokenizer and config filesretrieved 2026-09-12
Sections: tokenizer.json added_tokens ids 151643 to 151659; tokenizer_config.json chat_template; config.json vocab_size
Used by:P4 Base, Instruct, Thinking, Coder, Vision, Embedding: Reading a Model Name
- DeepSeek-R1-Distill-Qwen-7B model cardretrieved 2026-09-09, 2026-09-12
Sections: Model Downloads, DeepSeek-R1-Distill Models (table and the note on configs and tokenizers); Usage recommendations; License; Model summary; base model; licence; Licence; Model summary; licence; Distillation procedure; licence; usage recommendations
Used by:P4 Base, Instruct, Thinking, Coder, Vision, Embedding: Reading a Model NameP15 Reasoning Distillation: Teacher Traces as Training DataP15 Generating Synthetic Data with a Local TeacherP15 Why Distillation WorksP27 Distilling a Big Agent Model into a Small Local One
- DeepSeek-V3.2 model cardretrieved 2026-09-12
- DeepSeek-V4-Flash model cardretrieved 2026-09-12
- Devstral-Small-2-24B-Instruct-2512 model cardretrieved 2026-09-12
- EvalEval EEE datastore — gpt-oss-120b GPQA Diamond record (wasp 0.3.0)retrieved 2026-09-12
- Gemma 3 27B instruction-tuned model cardretrieved 2026-09-08
- Gemma 3 27B instruction-tuned, QAT q4_0 GGUF model cardretrieved 2026-09-08, 2026-09-09
Sections: Model description; licence
Used by:P4 Base, Instruct, Thinking, Coder, Vision, Embedding: Reading a Model NameP16 Quantisation-Aware Training and the Blackwell Formats: FP8, NVFP4 and MXFP4
- Gemma 4 E4B instruction-tuned model cardretrieved 2026-09-12
- ggml-org/gpt-oss-120b-GGUF - file listing; tensor table of gpt-oss-120b-MXFP4.ggufretrieved 2026-09-12
Used by:P3 Inference: Prefill, Decode and Why Memory Bandwidth Rules
- ggml-org/gpt-oss-120b-GGUF — gpt-oss-120b-MXFP4.gguf header, read by HTTP range requestretrieved 2026-09-13
Sections: tensor table - expert weights MXFP4; attention, token_embd and output Q8_0; router and biases F32
- ggml-org/gpt-oss-120b-GGUF file listingretrieved 2026-09-12
Used by:P4 Choosing a Model for a Memory BudgetP4 Dense, Mixture-of-Experts and Hybrid ArchitecturesP4 Lab: Build Your Model Shortlist
- ggml-org/gpt-oss-120b-GGUF model repositoryretrieved 2026-09-13
Used by:P6 Lab: Run and Benchmark the Course Reference Models
- ggml-org/gpt-oss-20b-GGUF file listingretrieved 2026-09-12
Used by:P4 Choosing a Model for a Memory BudgetP4 Lab: Build Your Model Shortlist
- ggml-org/gpt-oss-20b-GGUF model repositoryretrieved 2026-09-09
Used by:P6 GGUF and Quantisation TypesP6 Lab: Run and Benchmark the Course Reference Models
- ggml-org/Qwen2.5-VL-7B-Instruct-GGUF on Hugging Faceretrieved 2026-09-08
Sections: Quantised files; usage with the llama.cpp server
Used by:P10 Vision, Speech and Documents: Multimodal Locally
- GLM-4.6 model cardretrieved 2026-09-12
- Google on Hugging Face — organisation pageretrieved 2026-09-08
- gpt-oss-120b - config.jsonretrieved 2026-09-12
Sections: layer_types, sliding_window, num_local_experts, num_experts_per_tok, quantization_config
Used by:P3 Inference: Prefill, Decode and Why Memory Bandwidth Rules
- gpt-oss-120b config.jsonretrieved 2026-09-12
Used by:P4 Choosing a Model for a Memory BudgetP4 Dense, Mixture-of-Experts and Hybrid Architectures
- gpt-oss-120b model cardretrieved 2026-09-09, 2026-09-12
Sections: Highlights - parameters, active parameters, MXFP4 quantization; Licence; evaluation references
Used by:P3 Inference: Prefill, Decode and Why Memory Bandwidth RulesP4 Dense, Mixture-of-Experts and Hybrid ArchitecturesP4 Reading a Model Card and a BenchmarkP4 Model Families and Who Makes ThemP25 Which Local Models Can Actually Drive an Agent
- gpt-oss-20b — config.jsonretrieved 2026-09-12
- gpt-oss-20b config.jsonretrieved 2026-09-12
- Granite 4.0 H Small model cardretrieved 2026-09-12
- GSM8K dataset card (openai/gsm8k)retrieved 2026-09-09
Sections: Dataset summary; data fields; licence; Dataset summary; licence
Used by:P14 Lab: GRPO on a Maths or Code Task on One MachineP14 Reality Check: 'RL Makes Small Models Reason'P14 Reward Functions: Maths, Code Tests, Format and Length
- Hugging Face — hugging-quants/Meta-Llama-3.1-405B-Instruct-AWQ-INT4retrieved 2026-09-09
- Hugging Face — hugging-quants/Meta-Llama-3.1-70B-Instruct-AWQ-INT4retrieved 2026-09-09
- Hugging Face — Qwen/Qwen3-32B-AWQretrieved 2026-09-09
- Hugging Face — unsloth/Qwen3-4B-GGUF file listingretrieved 2026-09-13
Used by:P7 Lab: A Private Chat Service for Your Home Network
- Hugging Face Datasets documentation — Streamretrieved 2026-09-09
Sections: Split dataset; Shuffle; Map; Filter
Used by:P12 Data for Pretraining: FineWeb-Edu, Cosmopedia and Training a Tokeniser
- Hugging Face Hub — CLI guideretrieved 2026-09-09
Sections: Download files; cache management; Download; cache management and verification
Used by:P28 Capstone 1: Hardware and Model PlanP28 Capstone 6: The Report
- Hugging Face Hub — Command Line Interface (hf), v1.30.0retrieved 2026-09-12, 2026-09-13
Sections: hf auth login; hf auth whoami; hf download; Dry-run mode; Download to a local folder; Quiet mode; Download timeout; hf auth login; hf auth whoami; hf download; hf cache ls; hf env
Used by:P2 Lab: Look Inside a ModelP5 Lab: Prepare Your Machine
- Hugging Face Hub — file listings for unsloth/Qwen3-8B-GGUF, Qwen/Qwen3-8B-AWQ and Qwen/Qwen3-8Bretrieved 2026-09-13
Sections: file sizes; Qwen3-8B-Q4_K_M.gguf size and sha256; AWQ config.json
- Hugging Face Hub — Repository settingsretrieved 2026-09-09
Sections: Repository visibility
Used by:P13 Merging, Exporting and Quantising a Fine-Tuned ModelP13 Project: A Specialist Assistant
- Hugging Face Hub API - model info, tags and file trees for the repositories named on this pageretrieved 2026-09-12
Sections: read with huggingface_hub 1.30.0 (model_info, list_repo_tree, get_safetensors_metadata)
Used by:P4 Base, Instruct, Thinking, Coder, Vision, Embedding: Reading a Model Name
- Hugging Face Hub API — file listings of Qwen/Qwen3-0.6B, Qwen3-1.7B and Qwen3-8B (identical tokenizer.json, tokenizer_config.json and generation_config.json object ids)retrieved 2026-09-12
Used by:P2 From Autocomplete to Assistant: Next-Token PredictionP3 Post-Training: SFT, Preference Tuning and Reinforcement Learning
- Hugging Face Hub API — model listings (author, createdAt, gated, license tags, childrenModelCount) for the fourteen namespaces on this pageretrieved 2026-09-12
- Hugging Face Hub API — unsloth/Qwen3-4B-GGUF and unsloth/Qwen3-8B-GGUF (GGUF context_length, file sizes)retrieved 2026-09-13
- Hugging Face Hub documentation - Gated modelsretrieved 2026-09-08, 2026-09-09, 2026-09-12
Sections: Access gated models as a user; Download files; Access gated models as a user; Manage gated models as a model author
Used by:P3 Open Weights, Open Source and LicencesP4 Lab: Build Your Model ShortlistP4 Reading a Model Card and a BenchmarkP4 Model Families and Who Makes ThemP11 PyTorch, Transformers, Datasets and the Hugging Face Ecosystem
- Hugging Face Hub documentation - Licenses (identifiers)retrieved 2026-09-12
- Hugging Face Hub documentation - Model Cards (specifying a license, a base model and datasets)retrieved 2026-09-12
Sections: Specifying a base model; Specifying a task (pipeline_tag)
Used by:P3 Open Weights, Open Source and LicencesP4 Reading a Model Card and a BenchmarkP4 Base, Instruct, Thinking, Coder, Vision, Embedding: Reading a Model Name
- Hugging Face Hub documentation — Command Line Interface (CLI)retrieved 2026-09-09
Sections: hf auth login; hf download; Download a specific revision
Used by:P11 PyTorch, Transformers, Datasets and the Hugging Face Ecosystem
- Hugging Face Hub documentation — Command Line Interface (hf)retrieved 2026-09-08, 2026-09-09, 2026-09-12
Sections: Getting started; Output formatting; hf auth login; hf auth whoami; hf download; hf cache verify; hf cache prune; hf env; hf download; hf cache; hf auth login; hf upload; hf cache verify
Used by:P4 Lab: Build Your Model ShortlistP7 Managing a Model Library: Storage, Naming and VersionsP13 Merging, Exporting and Quantising a Fine-Tuned ModelP23 Backup, Upgrades and Reproducibility of a Model EstateP23 Security for Exposed Endpoints and the Model Supply Chain
- Hugging Face Hub documentation — Dataset cardsretrieved 2026-09-09
Sections: Dataset card metadata
Used by:P11 Datasets: Formats, Chat Templates, Tokenisation and Packing
- Hugging Face Hub documentation — Download files from the Hubretrieved 2026-09-09, 2026-09-12
Sections: Download files to a local folder; Dry-run mode; Faster downloads; Downloading a specific revision; the LFS SHA-256
Used by:P4 Lab: Build Your Model ShortlistP23 Backup, Upgrades and Reproducibility of a Model EstateP23 Security for Exposed Endpoints and the Model Supply Chain
- Hugging Face Hub documentation — Environment variablesretrieved 2026-09-12
Sections: HF_HOME; HF_HUB_CACHE; HF_TOKEN; HF_HUB_DOWNLOAD_TIMEOUT; HF_HUB_DISABLE_XET; HF_XET_HIGH_PERFORMANCE; HF_XET_RECONSTRUCT_WRITE_SEQUENTIALLY; HF_HUB_ENABLE_HF_TRANSFER (deprecated)
- Hugging Face Hub documentation — Evaluation Results (.eval_results, badges, benchmark datasets)retrieved 2026-09-12
- Hugging Face Hub documentation — GGUF, quantization typesretrieved 2026-09-08, 2026-09-12
Sections: Quantization Types table
Used by:P1 Tensors, GPUs and Precision: Why Matrix Multiplication Is EverythingP4 Choosing a Model for a Memory BudgetP4 Lab: Build Your Model ShortlistP4 Base, Instruct, Thinking, Coder, Vision, Embedding: Reading a Model Name
- Hugging Face Hub documentation — Understand cachingretrieved 2026-09-08
Sections: File-based caching; refs, blobs, snapshots, trees; limitations; CACHEDIR.TAG; inspect, verify and clean your cache
Used by:P7 Managing a Model Library: Storage, Naming and Versions
- Hugging Face Hub documentation — User access tokensretrieved 2026-09-09
Sections: What are User Access Tokens; Best practices
Used by:P11 PyTorch, Transformers, Datasets and the Hugging Face Ecosystem
- Hugging Face Hub model API records — safetensors parameter counts for Qwen3-4B, Qwen3-8B, Qwen3-14B, Qwen3-32B, Qwen3-30B-A3B, Qwen3-235B-A22B and gpt-oss-120bretrieved 2026-09-12
- Hugging Face Hub OpenAPI specification (Markdown rendering)retrieved 2026-09-12
Sections: GET /api/models/{namespace}/{repo}/tree/{rev}/{path}
- Hugging Face Hub v1.30.0 — Environment variablesretrieved 2026-09-13
Sections: HF_HOME; HF_HUB_CACHE; HF_TOKEN_PATH
Used by:P5 Lab: Prepare Your Machine
- Hugging Face Tokenizers — documentation indexretrieved 2026-09-08
Used by:P2 Lab: Look Inside a Model
- Hugging Face Tokenizers — Quicktourretrieved 2026-09-09, 2026-09-12
Sections: Build a tokenizer from scratch; Using the tokenizer; alignment tracking and offsets; Build a tokenizer from scratch; Training the tokenizer
Used by:P2 Tokens, Tokenisers and VocabularyP12 Data for Pretraining: FineWeb-Edu, Cosmopedia and Training a Tokeniser
- Hugging Face Tokenizers 0.23.2 — API reference, Tokenizerretrieved 2026-09-12
Sections: The pipeline; from_file; encode; decode; train_from_iterator; get_vocab_size; token_to_id; id_to_token
- Hugging Face Tokenizers 0.23.2 — API reference, Trainersretrieved 2026-09-12
Sections: BpeTrainer parameters
- Hugging Face Transformers — Attention backendsretrieved 2026-09-08, 2026-09-12
Sections: Attention backends; Set an attention backend; Set an attention backend
Used by:P2 Attention and the TransformerP2 Lab: Look Inside a Model
- Hugging Face Transformers — Cache strategies (KV cache)retrieved 2026-09-12
Sections: Default cache; Cache offloading; Quantized cache; Prefill a cache
- Hugging Face Transformers — Chat templatesretrieved 2026-09-08, 2026-09-09, 2026-09-12
Sections: Using apply_chat_template; add_generation_prompt; Using apply_chat_template; the special-tokens warning; add_generation_prompt; Model training; Using apply_chat_template; add_generation_prompt; Model training; Model training; apply_chat_template; apply_chat_template; add_generation_prompt; Model training; Model training
Used by:P2 Lab: Look Inside a ModelP2 From Autocomplete to Assistant: Next-Token PredictionP2 Tokens, Tokenisers and VocabularyP11 Datasets: Formats, Chat Templates, Tokenisation and PackingP13 Building an SFT DatasetP13 Challenge: The Fine-Tune That Got WorseP13 Fine-Tuning with TRL and PEFT
- Hugging Face Transformers — Chat templatesretrieved 2026-09-08, 2026-09-09
Sections: Using apply_chat_template; add_generation_prompt; continue_final_message; Model training; add_generation_prompt
Used by:P10 Prompting That Works Locally: System Prompts, Chat Templates and Thinking ModesP27 Fine-Tuning for Tool Use and Your Codebase
- Hugging Face Transformers — Generation strategiesretrieved 2026-09-08, 2026-09-12
Sections: Greedy search; Sampling; Greedy search; Sampling; Beam search
Used by:P2 Lab: Look Inside a ModelP2 From Autocomplete to Assistant: Next-Token Prediction
- Hugging Face Transformers — Loading modelsretrieved 2026-09-08, 2026-09-09
Sections: Custom models; trust_remote_code; revision pinning; Custom models; trust_remote_code; loading from a specific revision
Used by:P10 Privacy, Security and Serving Beyond localhostP23 Security for Exposed Endpoints and the Model Supply Chain
- Hugging Face Transformers — Model outputsretrieved 2026-09-08, 2026-09-12
Sections: BaseModelOutput; attentions; BaseModelOutput; hidden_states; CausalLMOutput; attentions; logits
Used by:P2 Attention and the TransformerP2 Embeddings: Meaning as GeometryP2 Lab: Look Inside a Model
- Hugging Face Transformers — Tokenization algorithms (tokenizer summary)retrieved 2026-09-12
Sections: Byte pair encoding; Byte-level BPE; Unigram; SentencePiece; WordPiece; Word-level; Character-level
- huggingface_hub 1.30.0 — Command Line Interfaceretrieved 2026-09-12
Sections: hf download; Download a single file; Download to a local folder
- huggingface_hub v1.30.0 - HfApi reference (model_info, dataset_info, ModelInfo, DatasetInfo, card data)retrieved 2026-09-12
- HuggingFaceFW/fineweb - dataset cardretrieved 2026-09-12
Sections: What is being released; data processing pipeline; deduplication; licence (ODC-By 1.0 and CommonCrawl Terms of Use); token counts with the gpt2 tokenizer
- HuggingFaceFW/fineweb-edu - dataset cardretrieved 2026-09-09, 2026-09-12
Sections: Educational classifier (Llama3-70B-Instruct annotations, score threshold 3, about 92 per cent removed); deduplication ablation; licence; What is it; Dataset curation; Annotation; Licensing Information; Licensing Information
Used by:P3 Pretraining: Learning from Trillions of TokensP12 Data for Pretraining: FineWeb-Edu, Cosmopedia and Training a TokeniserP12 Project: A Domain Micro-Model
- HuggingFaceTB on Hugging Face — organisation pageretrieved 2026-09-08
- IBM Granite on Hugging Face — organisation pageretrieved 2026-09-08
- karpathy/climbmix-400b-shuffle dataset cardretrieved 2026-09-09
Used by:P12 Data for Pretraining: FineWeb-Edu, Cosmopedia and Training a Tokeniser
- Kimi-K2-Instruct-0905 model cardretrieved 2026-09-12
- Kimi-K2-Instruct-0905 tokenizer_config.json and file list (tokenization_kimi.py, tiktoken.model)retrieved 2026-09-12
- lighteval documentation — quick tourretrieved 2026-09-09
Sections: Available commands; task specification
Used by:P16 Evaluation Harnesses: lm-evaluation-harness, lighteval, EvalPlus and Your Own
- livecodebench/code_generation_lite loading script (release files)retrieved 2026-09-12
- Llama 3.1 8B Instruct model cardretrieved 2026-09-09
Sections: Model card; tokenizer
Used by:P15 Challenge: The Student That Learned the Teacher's Mistakes
- Llama-3.1-Minitron-4B-Width-Base model cardretrieved 2026-09-09
Sections: Model architecture; training; licence; Licence; model architecture
Used by:P15 Pruning and Compression: Prune-and-DistilP15 Generating Synthetic Data with a Local Teacher
- Llama-4-Scout-17B-16E-Instruct model cardretrieved 2026-09-12
- MathArena AIME 2025 dataset cardretrieved 2026-09-12
- Meta Llama on Hugging Face — organisation pageretrieved 2026-09-08
- meta-llama/Llama-3.1-8B-Instruct model API record (gated field)retrieved 2026-09-12
- Microsoft on Hugging Face — organisation pageretrieved 2026-09-08
- MiniMax on Hugging Face — organisation pageretrieved 2026-09-08
- MiniMax-M1-80k model cardretrieved 2026-09-12
- MiniMax-M2 model cardretrieved 2026-09-12
- Mistral AI on Hugging Face — organisation pageretrieved 2026-09-08
- Mixture of Experts Explained — Hugging Face blogretrieved 2026-09-12
Used by:P4 Dense, Mixture-of-Experts and Hybrid Architectures
- mlx-community/Qwen3-8B-4bit model repositoryretrieved 2026-09-09, 2026-09-13
Used by:P8 Lab: Same Model, Every EngineP8 MLX and mlx-lm: Apple's Native Path
- Moonshot AI on Hugging Face — organisation pageretrieved 2026-09-08
- mradermacher/Qwen3-0.6B-Base-GGUF and unsloth/Qwen3-0.6B-GGUF - GGUF header metadataretrieved 2026-09-13
Sections: tokenizer.ggml.eos_token_id read from the smallest file of each repository; Hub base_model tags
Used by:P4 Base, Instruct, Thinking, Coder, Vision, Embedding: Reading a Model Name
- NVIDIA on Hugging Face — organisation pageretrieved 2026-09-08
- NVIDIA-Nemotron-3-Nano-30B-A3B-BF16 config.jsonretrieved 2026-09-12
Used by:P4 Dense, Mixture-of-Experts and Hybrid Architectures
- NVIDIA-Nemotron-3-Nano-30B-A3B-BF16 model cardretrieved 2026-09-12
Used by:P4 Dense, Mixture-of-Experts and Hybrid ArchitecturesP4 Model Families and Who Makes Them
- nvidia/ClimbMix dataset cardretrieved 2026-09-09
Sections: Dataset description; Licence
Used by:P12 Data for Pretraining: FineWeb-Edu, Cosmopedia and Training a Tokeniser
- Olmo-3-7B-Instruct model cardretrieved 2026-09-12
- openai/gpt-oss-20b model cardretrieved 2026-09-08, 2026-09-09, 2026-09-12
Sections: MXFP4 quantisation of the MoE weights; licence; parameter counts; Model card; memory footprint; Highlights (MXFP4 quantization); licence; Highlights; Reasoning levels; chat_template.jinja; Model description; quantisation; Reasoning levels; harmony response format; Licence; MXFP4 quantisation; evaluation; Harmony format; reasoning effort; agentic capabilities
Used by:P1 Tensors, GPUs and Precision: Why Matrix Multiplication Is EverythingP2 Parameters, Layers and Model SizeP3 Distillation, Pruning and Quantisation: How Small Models Get GoodP4 Model Families and Who Makes ThemP4 Base, Instruct, Thinking, Coder, Vision, Embedding: Reading a Model NameP6 GGUF and Quantisation TypesP6 Sampling: Temperature, Top-p, Min-p, Repetition and DeterminismP10 Prompting That Works Locally: System Prompts, Chat Templates and Thinking ModesP15 Generating Synthetic Data with a Local TeacherP16 Quantisation-Aware Training and the Blackwell Formats: FP8, NVFP4 and MXFP4P24 Reasoning Models in Agent Loops
- parakeet-tdt-0.6b-v3 model cardretrieved 2026-09-08, 2026-09-12
Sections: Model overview; licence; input format; long-form audio
Used by:P4 Model Families and Who Makes ThemP10 Vision, Speech and Documents: Multimodal Locally
- PEFT — LoRA developer guideretrieved 2026-09-09
Sections: Merging adapters; merge_and_unload; Rank and alpha; Target modules; Merging adapters; Initialization; Rank and alpha; Target modules; rsLoRA; DoRA; Merging adapters; Merging adapters; merge_adapter; merge_and_unload
Used by:P13 Challenge: The Fine-Tune That Got WorseP13 Fine-Tuning with TRL and PEFTP13 Lab: Fine-Tune a 1B to 4B Model to Follow Your FormatP13 LoRA and QLoRA ExplainedP13 Merging, Exporting and Quantising a Fine-Tuned Model
- PEFT — Quantizationretrieved 2026-09-09
Sections: torchao caveats; Quantize a model; LoraConfig; QLoRA-style training; torchao caveats; LoftQ; Quantize a model; QLoRA-style training
Used by:P13 Challenge: The Fine-Tune That Got WorseP13 LoRA and QLoRA ExplainedP13 Merging, Exporting and Quantising a Fine-Tuned ModelP13 Project: A Specialist Assistant
- PEFT documentation — LoRA (conceptual guide)retrieved 2026-09-09
Sections: Merging; Low-rank decomposition; rank; merging
Used by:P11 PyTorch, Transformers, Datasets and the Hugging Face EcosystemP13 Fine-Tuning with TRL and PEFTP13 LoRA and QLoRA Explained
- PEFT documentation — LoRA developer guideretrieved 2026-09-09
Sections: LoraConfig; merge_and_unload; LoraConfig; print_trainable_parameters; LoraConfig; merge_and_unload
Used by:P11 Lab: Your First Training RunP11 Memory Arithmetic for Training: Weights, Gradients, Optimiser States and ActivationsP11 PyTorch, Transformers, Datasets and the Hugging Face EcosystemP14 Lab: DPO a Model to Prefer Your StyleP15 Lab: Logit Distillation with TRL's Distillation TrainersP15 Lab: Sequence-Level Distillation of a 30B-Class Teacher into a 4B StudentP15 Project: The Distillation Pipeline
- Phi-4 model cardretrieved 2026-09-12
- Qwen on Hugging Face — organisation pageretrieved 2026-09-08
- Qwen/Qwen3-235B-A22B-GGUF file listingretrieved 2026-09-12
- Qwen/Qwen3-8B — config.jsonretrieved 2026-09-12
Sections: hidden_size, intermediate_size, num_hidden_layers, num_attention_heads, num_key_value_heads, head_dim, vocab_size, tie_word_embeddings, initializer_range; vocab_size
Used by:P1 Neural Networks, Activations and BackpropagationP1 Tensors, GPUs and Precision: Why Matrix Multiplication Is EverythingP1 What Learning Means: Data, Loss and Gradient DescentP4 Choosing a Model for a Memory Budget
- Qwen/Qwen3-8B-AWQ model repositoryretrieved 2026-09-13
Used by:P8 Lab: Same Model, Every Engine
- Qwen3-0.6B — config.jsonretrieved 2026-09-12
Sections: num_hidden_layers, num_key_value_heads, head_dim, tie_word_embeddings
Used by:P2 Attention and the TransformerP3 Inference: Prefill, Decode and Why Memory Bandwidth Rules
- Qwen3-0.6B config.jsonretrieved 2026-09-09
Used by:P11 Lab: Your First Training RunP11 Memory Arithmetic for Training: Weights, Gradients, Optimiser States and ActivationsP15 Lab: Logit Distillation with TRL's Distillation Trainers
- Qwen3-0.6B model card and config.jsonretrieved 2026-09-09, 2026-09-12
Sections: Model overview; config.json; Model overview; licence
Used by:P2 Embeddings: Meaning as GeometryP11 Lab: Your First Training Run
- Qwen3-0.6B tokenizer_config.jsonretrieved 2026-09-09
Used by:P15 Lab: Logit Distillation with TRL's Distillation TrainersP15 Why Distillation Works
- Qwen3-0.6B-Base and Qwen3-0.6B repository filesretrieved 2026-09-12
Sections: generation_config.json and tokenizer_config.json of both repositories, and of the 1.7B pair
Used by:P4 Base, Instruct, Thinking, Coder, Vision, Embedding: Reading a Model Name
- Qwen3-1.7B — config.jsonretrieved 2026-09-12
Used by:P2 Attention and the TransformerP2 Lab: Look Inside a Model
- Qwen3-1.7B — files and versionsretrieved 2026-09-12
Sections: pre_tokenizer regex; added_tokens; chat_template; file checksums, compared with Qwen3-0.6B and Qwen3-8B
Used by:P2 Lab: Look Inside a ModelP2 Tokens, Tokenisers and Vocabulary
- Qwen3-1.7B config.jsonretrieved 2026-09-09, 2026-09-12
Used by:P4 Lab: Build Your Model ShortlistP11 Memory Arithmetic for Training: Weights, Gradients, Optimiser States and ActivationsP15 Lab: Logit Distillation with TRL's Distillation TrainersP15 Lab: Sequence-Level Distillation of a 30B-Class Teacher into a 4B Student
- Qwen3-1.7B model card and config.jsonretrieved 2026-09-09, 2026-09-12
Sections: Model overview; config.json; Model Overview; Quickstart; Switching Between Thinking and Non-Thinking Mode; Best Practices; Model Overview; Switching Between Thinking and Non-Thinking Mode; Best Practices; Model overview; licence
Used by:P2 Embeddings: Meaning as GeometryP2 Lab: Look Inside a ModelP2 From Autocomplete to Assistant: Next-Token PredictionP3 Post-Training: SFT, Preference Tuning and Reinforcement LearningP14 Lab: DPO a Model to Prefer Your StyleP14 Lab: GRPO on a Maths or Code Task on One Machine
- Qwen3-1.7B-Base — files and versionsretrieved 2026-09-12
Used by:P2 Lab: Look Inside a Model
- Qwen3-1.7B-Base model cardretrieved 2026-09-08, 2026-09-09, 2026-09-12
Sections: Model Overview; Pretraining data (36 trillion tokens, 119 languages); parameter counts (1.7B, 1.4B non-embedding); licence; Model overview; Training stage
Used by:P2 Lab: Look Inside a ModelP2 From Autocomplete to Assistant: Next-Token PredictionP3 Post-Training: SFT, Preference Tuning and Reinforcement LearningP3 Pretraining: Learning from Trillions of TokensP12 Data for Pretraining: FineWeb-Edu, Cosmopedia and Training a TokeniserP12 Project: A Domain Micro-ModelP12 What Pretraining Teaches You That Fine-Tuning Cannot
- Qwen3-14B config.jsonretrieved 2026-09-12
- Qwen3-14B model cardretrieved 2026-09-09
Sections: Model overview; licence
Used by:P15 Lab: Sequence-Level Distillation of a 30B-Class Teacher into a 4B Student
- Qwen3-235B-A22B config.jsonretrieved 2026-09-12
Used by:P4 Choosing a Model for a Memory BudgetP4 Dense, Mixture-of-Experts and Hybrid Architectures
- Qwen3-30B-A3B — config.jsonretrieved 2026-09-12
Sections: num_experts, num_experts_per_tok, num_key_value_heads, head_dim
Used by:P2 Parameters, Layers and Model SizeP3 Inference: Prefill, Decode and Why Memory Bandwidth Rules
- Qwen3-30B-A3B config.jsonretrieved 2026-09-09, 2026-09-12
Used by:P4 Choosing a Model for a Memory BudgetP4 Dense, Mixture-of-Experts and Hybrid ArchitecturesP11 Memory Arithmetic for Training: Weights, Gradients, Optimiser States and Activations
- Qwen3-30B-A3B model cardretrieved 2026-09-08, 2026-09-09, 2026-09-12
Sections: Model overview; Model overview - parameters and experts; Model overview and licence; Model overview; licence; best practices; Licence; best practices; Switching between thinking and non-thinking mode; best practices; config.json
Used by:P2 Parameters, Layers and Model SizeP3 Inference: Prefill, Decode and Why Memory Bandwidth RulesP3 Reality Check: 'A Small Local Model Is as Good as the Frontier'P4 Dense, Mixture-of-Experts and Hybrid ArchitecturesP15 Lab: Sequence-Level Distillation of a 30B-Class Teacher into a 4B StudentP15 Project: The Distillation PipelineP15 Reasoning Distillation: Teacher Traces as Training DataP15 Generating Synthetic Data with a Local TeacherP19 llama.cpp RPC: Layers Across Machines
- Qwen3-30B-A3B-Instruct-2507 model cardretrieved 2026-09-12
Sections: Model overview note on non-thinking mode; Best practices
Used by:P4 Model Families and Who Makes ThemP4 Base, Instruct, Thinking, Coder, Vision, Embedding: Reading a Model Name
- Qwen3-30B-A3B-Thinking-2507 model cardretrieved 2026-09-12
Sections: Model overview note on thinking mode; long-context evaluation notes; Best practices
Used by:P4 Base, Instruct, Thinking, Coder, Vision, Embedding: Reading a Model Name
- Qwen3-32B config.jsonretrieved 2026-09-12
Used by:P4 Choosing a Model for a Memory BudgetP4 Dense, Mixture-of-Experts and Hybrid Architectures
- Qwen3-32B model cardretrieved 2026-09-12
Used by:P4 Dense, Mixture-of-Experts and Hybrid Architectures
- Qwen3-4B — config.jsonretrieved 2026-09-12
Used by:P2 Attention and the Transformer
- Qwen3-4B config.jsonretrieved 2026-09-09, 2026-09-12
Used by:P4 Choosing a Model for a Memory BudgetP11 Memory Arithmetic for Training: Weights, Gradients, Optimiser States and ActivationsP15 Lab: Sequence-Level Distillation of a 30B-Class Teacher into a 4B Student
- Qwen3-4B model cardretrieved 2026-09-09, 2026-09-12, 2026-09-13
Sections: Model overview; switching between thinking and non-thinking mode; best practices; licence; Model Overview; Processing Long Texts; Licence; Best Practices; Model overview; licence; best practices; Licence; model overview; Best practices; sampling settings; Agentic use; licence; best practices; Licence; best practices
Used by:P3 Reality Check: 'A Small Local Model Is as Good as the Frontier'P7 Reality Check: 'The Default Context Is Enough'P13 Lab: Fine-Tune a 1B to 4B Model to Follow Your FormatP13 Project: A Specialist AssistantP13 What Fine-Tuning Changes and What It CannotP14 Lab: GRPO on a Maths or Code Task on One MachineP15 Project: The Distillation PipelineP27 Distilling a Big Agent Model into a Small Local OneP27 Fine-Tuning for Tool Use and Your CodebaseP27 Lab: Fine-Tune a Small Model on Your Own Agent Trajectories and Measure the Gain
- Qwen3-4B tokenizer_config.jsonretrieved 2026-09-09
Used by:P15 Challenge: The Student That Learned the Teacher's MistakesP15 Lab: Logit Distillation with TRL's Distillation TrainersP15 Reasoning Distillation: Teacher Traces as Training DataP15 Why Distillation Works
- Qwen3-8B — config.jsonretrieved 2026-09-12
Sections: num_hidden_layers, num_key_value_heads, head_dim, tie_word_embeddings
Used by:P2 Parameters, Layers and Model SizeP3 Inference: Prefill, Decode and Why Memory Bandwidth RulesP3 Distillation, Pruning and Quantisation: How Small Models Get GoodP4 Reading a Model Card and a BenchmarkP4 Model Families and Who Makes Them
- Qwen3-8B — model-00001-of-00005.safetensors header (tensor names and shapes)retrieved 2026-09-12
Sections: JSON header
Used by:P2 Attention and the Transformer
- Qwen3-8B model card and config.jsonretrieved 2026-09-08, 2026-09-09, 2026-09-12, 2026-09-13
Sections: Model Overview; Processing Long Texts; config.json; Model overview; config.json; Model overview; Processing long texts; Context length; enable_thinking; apply_chat_template example; vocab_size, bos_token_id, eos_token_id; Metadata; Model Overview; Processing Long Texts; Best Practices; Switching between thinking and non-thinking mode; Best practices; Model overview; context length; Best Practices; thinking and non-thinking modes; Switching between thinking and non-thinking mode; Best Practices; recommended sampling settings; Switching Between Thinking and Non-Thinking Mode; Best Practices; Model overview; licence; Best practices; Model overview; Model overview; licence; best practices; Model overview; best practices; Best Practices; benchmark evaluation; Context length; YaRN; Thinking and non-thinking modes; agentic use; enable_thinking; sampling settings per mode; Thinking and non-thinking modes; enable_thinking; sampling settings; output length
Used by:P2 Attention and the TransformerP2 Embeddings: Meaning as GeometryP2 From Autocomplete to Assistant: Next-Token PredictionP2 Parameters, Layers and Model SizeP2 Tokens, Tokenisers and VocabularyP4 Choosing a Model for a Memory BudgetP4 Reading a Model Card and a BenchmarkP4 Model Families and Who Makes ThemP4 Base, Instruct, Thinking, Coder, Vision, Embedding: Reading a Model NameP6 llama-cli and llama-serverP6 Sampling: Temperature, Top-p, Min-p, Repetition and DeterminismP7 Reality Check: 'The Default Context Is Enough'P8 Lab: Same Model, Every EngineP10 Lab: Benchmark Local Models on Your Own TasksP10 Prompting That Works Locally: System Prompts, Chat Templates and Thinking ModesP14 Reality Check: 'RL Makes Small Models Reason'P15 Challenge: The Student That Learned the Teacher's MistakesP15 Lab: Logit Distillation with TRL's Distillation TrainersP15 Lab: Sequence-Level Distillation of a 30B-Class Teacher into a 4B StudentP15 Reasoning Distillation: Teacher Traces as Training DataP16 Challenge: The Benchmark That LiedP16 Lab: Run a Standard Benchmark Suite on Your ModelP16 LLM-as-Judge, Contamination and Honest ReportingP24 Context Engineering: Memory, Compaction and the KV BudgetP24 Function Calling End to End on Local EnginesP24 Lab: A Minimal Agent from ScratchP24 Reasoning Models in Agent Loops
- Qwen3-8B-FP8 model cardretrieved 2026-09-12
Sections: Note on FP8
Used by:P4 Base, Instruct, Thinking, Coder, Vision, Embedding: Reading a Model Name
- Qwen3-8B-GGUF file listing (Qwen's own conversion)retrieved 2026-09-12, 2026-09-13
Used by:P4 Choosing a Model for a Memory BudgetP4 Lab: Build Your Model ShortlistP4 Model Families and Who Makes Them
- Qwen3-Coder-30B-A3B-Instruct model card (including its benchmark image)retrieved 2026-09-08, 2026-09-09, 2026-09-12
Sections: Model overview; Best practices; Model overview; Best Practices; Context length; Tool calling; non-thinking mode; context length; Non-thinking mode; context length; sampling settings; Model overview; licence; Agentic coding; context length; licence; Agentic coding; function calling; context length
Used by:P4 Reading a Model Card and a BenchmarkP4 Base, Instruct, Thinking, Coder, Vision, Embedding: Reading a Model NameP10 Local Coding Assistants: Autocomplete and Chat in Your EditorP10 Prompting That Works Locally: System Prompts, Chat Templates and Thinking ModesP24 Context Engineering: Memory, Compaction and the KV BudgetP24 Function Calling End to End on Local EnginesP24 Reasoning Models in Agent LoopsP25 Which Local Models Can Actually Drive an AgentP27 Distilling a Big Agent Model into a Small Local OneP27 Fine-Tuning for Tool Use and Your Codebase
- Qwen3-Embedding-0.6B model cardretrieved 2026-09-08, 2026-09-12
Sections: Model overview; usage (last_token_pool, F.normalize, padding_side); instruction format; MRL dimensions; Instruction format; dimensions; licence; Model overview; instruction format; dimensions
Used by:P2 Embeddings: Meaning as GeometryP4 Base, Instruct, Thinking, Coder, Vision, Embedding: Reading a Model NameP10 Project: A Private Document Question-Answering ServiceP10 Retrieval-Augmented Generation: Embeddings, Chunking and Vector Stores
- Qwen3-Embedding-0.6B repository files — config.json, 1_Pooling/config.json, modules.json, config_sentence_transformers.json, model.safetensors headerretrieved 2026-09-12
- Qwen3-Next-80B-A3B-Instruct config.jsonretrieved 2026-09-12
Used by:P4 Dense, Mixture-of-Experts and Hybrid Architectures
- Qwen3-Next-80B-A3B-Instruct model cardretrieved 2026-09-12
Used by:P4 Dense, Mixture-of-Experts and Hybrid Architectures
- Qwen3-Reranker-0.6B model card and repository filesretrieved 2026-09-08, 2026-09-12
Sections: Transformers usage; modules.json; 1_LogitScore/config.json; Prompt format; evaluation setup; licence; Model overview; prompt format; evaluation setup
Used by:P4 Base, Instruct, Thinking, Coder, Vision, Embedding: Reading a Model NameP10 Project: A Private Document Question-Answering ServiceP10 Retrieval-Augmented Generation: Embeddings, Chunking and Vector Stores
- Qwen3-VL-8B-Instruct model cardretrieved 2026-09-08, 2026-09-12
Sections: README; config.json; preprocessor_config.json; Capabilities; context length; deployment
Used by:P4 Model Families and Who Makes ThemP4 Base, Instruct, Thinking, Coder, Vision, Embedding: Reading a Model NameP10 Vision, Speech and Documents: Multimodal Locally
- Qwen3.8-2.4T-A95B model card (licence metadata)retrieved 2026-09-12
- Qwen3.8-Flash-Next model card (licence metadata)retrieved 2026-09-12
- safetensors — documentation indexretrieved 2026-09-08, 2026-09-09, 2026-09-12
Sections: Load tensors; Overview
Used by:P2 Lab: Look Inside a ModelP2 Parameters, Layers and Model SizeP4 Base, Instruct, Thinking, Coder, Vision, Embedding: Reading a Model NameP10 Privacy, Security and Serving Beyond localhostP23 Security for Exposed Endpoints and the Model Supply Chain
- smolagents — Guided tourretrieved 2026-09-09
Sections: CodeAgent versus ToolCallingAgent; multi-agents; Multi-agents; managed_agents; name and description
Used by:P26 Agent Frameworks ComparedP26 Multi-Agent Patterns: Router, Planner and Executor, Critic, Parallel Fan-Out
- smolagents — Introductionretrieved 2026-09-09
Sections: Key features; code agents and tool-calling agents; The agent loop in a thousand lines
Used by:P26 Agent Frameworks ComparedP26 Reality Check: 'Agents Are Just Loops'
- smolagents — Models referenceretrieved 2026-09-09
Sections: OpenAIModel; LiteLLMModel; VLLMModel; MLXModel
Used by:P26 Agent Frameworks Compared
- SmolLM2-1.7B model cardretrieved 2026-09-09
Sections: Training; License
Used by:P12 Data for Pretraining: FineWeb-Edu, Cosmopedia and Training a Tokeniser
- SmolLM3-3B model cardretrieved 2026-09-12
Sections: Enabling and disabling extended thinking mode
Used by:P4 Model Families and Who Makes ThemP4 Base, Instruct, Thinking, Coder, Vision, Embedding: Reading a Model Name
- SWE-bench Verified dataset card (FAIL_TO_PASS, PASS_TO_PASS)retrieved 2026-09-12
- Transformers documentation — Customizing modelsretrieved 2026-09-09
Sections: Upload
Used by:P11 PyTorch, Transformers, Datasets and the Hugging Face Ecosystem
- Transformers documentation — Tool useretrieved 2026-09-09
Sections: Passing tools; JSON schemas; get_json_schema; Passing tools; JSON schemas
Used by:P27 Collecting Trajectories from Your AgentsP27 Lab: Fine-Tune a Small Model on Your Own Agent Trajectories and Measure the Gain
- Transformers documentation — Trainerretrieved 2026-09-09
Sections: train; save_model; save_state; Trainer; train; evaluate; save_model
Used by:P11 Experiment Tracking and ReproducibilityP11 PyTorch, Transformers, Datasets and the Hugging Face Ecosystem
- transformers documentation — TrainingArgumentsretrieved 2026-09-12
Sections: weight_decay (the page served v5.17.0 on the retrieval date)
Used by:P1 Generalisation: Train, Validation, Test and Overfitting
- TRL — Dataset formats and typesretrieved 2026-09-09
Sections: Standard and conversational formats; language modeling; prompt-completion; Prompt-completion; conversational; Conversational prompt-completion; Preference; Unpaired preference; Which dataset type to use; Preference; Unpaired preference; Prompt-only; Which dataset type to use; Preference; Conversational; Tool Calling; the tools column; Tool Calling; Tool Calling; the tools column
Used by:P13 Building an SFT DatasetP13 Fine-Tuning with TRL and PEFTP13 Lab: Fine-Tune a 1B to 4B Model to Follow Your FormatP14 DPO and Its Family: IPO, KTO, ORPO and SimPOP14 From Imitation to Preferences: Why Ranking Beats CopyingP14 Lab: DPO a Model to Prefer Your StyleP27 Collecting Trajectories from Your AgentsP27 Fine-Tuning for Tool Use and Your CodebaseP27 Lab: Fine-Tune a Small Model on Your Own Agent Trajectories and Measure the Gain
- TRL — SFT Trainerretrieved 2026-09-09
Sections: Expected dataset type and format; Train on completion only; Train on assistant messages only; SFTConfig parameters; Logged metrics; Quick start; Customization; Logged metrics; SFTConfig parameters; Quick start; Train adapters with PEFT; SFTConfig parameters; SFTTrainer parameters; quantization_config; Train adapters with PEFT; Train with Unsloth; Looking deeper into the SFT method; Computing the loss
Used by:P13 Building an SFT DatasetP13 Challenge: The Fine-Tune That Got WorseP13 Fine-Tuning with TRL and PEFTP13 Lab: Fine-Tune a 1B to 4B Model to Follow Your FormatP13 Project: A Specialist AssistantP13 Unsloth, Axolotl, LLaMA-Factory and mlx-lm: Higher-Level ToolsP13 What Fine-Tuning Changes and What It Cannot
- TRL — Transformers Reinforcement Learningretrieved 2026-09-09
Sections: Trainer taxonomy; offline, online and distillation methods; What's New; Taxonomy; What's New; Taxonomy; Knowledge distillation
Used by:P28 Capstone 4: Improve a Small ModelP15 Lab: Logit Distillation with TRL's Distillation TrainersP15 Three Kinds of Distillation: Logit, Sequence and On-Policy
- TRL 1.12.0 documentation — SFTConfigretrieved 2026-09-12
Sections: SFTConfig signature (weight_decay default) and the list of defaults that differ from TrainingArguments
Used by:P1 Generalisation: Train, Validation, Test and Overfitting
- TRL documentation — Async Distillation Trainerretrieved 2026-09-09
Sections: teacher_server_urls; tokenizer requirement; Overview; How it differs; AsyncDistillationConfig
Used by:P15 Challenge: The Student That Learned the Teacher's MistakesP15 Three Kinds of Distillation: Logit, Sequence and On-Policy
- TRL documentation — Dataset formats and typesretrieved 2026-09-09
Sections: Overview; Standard; Conversational; Prompt-completion; Which dataset type to use; Prompt-completion
Used by:P11 Datasets: Formats, Chat Templates, Tokenisation and PackingP11 Lab: Your First Training Run
- TRL documentation — Distillation Trainerretrieved 2026-09-09
Sections: DistillationTrainer parameters; DistillationConfig; Overview; Quick start; Computing the loss; Expected dataset type; Logged metrics; Train adapters with PEFT; DistillationConfig; Overview; Computing the loss; Expected dataset type; DistillationConfig; Overview; DistillationTrainer parameters
Used by:P15 Challenge: The Student That Learned the Teacher's MistakesP15 Lab: Logit Distillation with TRL's Distillation TrainersP15 Three Kinds of Distillation: Logit, Sequence and On-PolicyP15 Why Distillation Works
- TRL documentation — DPO Trainerretrieved 2026-09-09
Sections: Expected dataset type and format; Looking deeper into the DPO method; Loss Types; Logged metrics; DPOConfig; Looking deeper into the DPO method; Logged metrics; Quick start; Expected dataset type and format; Train adapters with PEFT; DPOConfig; Logged metrics
Used by:P14 DPO and Its Family: IPO, KTO, ORPO and SimPOP14 From Imitation to Preferences: Why Ranking Beats CopyingP14 Lab: DPO a Model to Prefer Your Style
- TRL documentation — Generalized Knowledge Distillation Trainerretrieved 2026-09-09
Sections: Usage tips; GKDConfig
Used by:P15 Lab: Logit Distillation with TRL's Distillation TrainersP15 Three Kinds of Distillation: Logit, Sequence and On-Policy
- TRL documentation — GRPO Trainerretrieved 2026-09-09
Sections: Generating completions; Computing the advantage; Estimating the KL divergence; Computing the loss; GRPOConfig; Logged metrics; Quick start; Using custom reward functions; GRPOConfig; Speeding up training with vLLM; Logged metrics; GRPOConfig; Logged metrics; Using custom reward functions; GRPOConfig reward_weights; Logged metrics; Quick start; GRPOConfig; Speeding up training with vLLM; rollout_func; Tools; Environments; custom reward function; num_generations; max_tool_calling_iterations
Used by:P14 Reinforcement Learning with Verifiable Rewards: GRPO ExplainedP14 Lab: GRPO on a Maths or Code Task on One MachineP14 Reality Check: 'RL Makes Small Models Reason'P14 Reward Functions: Maths, Code Tests, Format and LengthP14 The RL Toolchain: TRL, Unsloth, verl, OpenRLHF and vLLM RolloutsP27 Reinforcement Learning on Agent Tasks: Tests as Rewards
- TRL documentation — KTO Trainerretrieved 2026-09-09
Sections: Expected dataset type and format; Batch size recommendations; Learning rate recommendations; Imbalanced data; KTOConfig
- TRL documentation — MiniLLM Trainerretrieved 2026-09-09
Sections: Overview; MiniLLMConfig
Used by:P15 Three Kinds of Distillation: Logit, Sequence and On-Policy
- TRL documentation — ORPO Trainerretrieved 2026-09-09
Sections: Overview; Expected dataset type; Logged metrics; ORPOConfig
- TRL documentation — Reducing memory usageretrieved 2026-09-09
Sections: Truncation; Packing; Padding-free; Chunked cross-entropy; Gradient checkpointing; Truncation
Used by:P11 Datasets: Formats, Chat Templates, Tokenisation and PackingP11 Memory Arithmetic for Training: Weights, Gradients, Optimiser States and Activations
- TRL documentation — SFT Trainerretrieved 2026-09-09
Sections: Expected dataset type and format; Train on completion only; Train on assistant messages only; Packing; SFTConfig; Logged metrics; Quick start; Expected dataset type and format; Train adapters with PEFT; Train on completion only; Computing the loss; Packing; Quick start; Expected dataset type and format; Train adapters with PEFT
Used by:P11 Datasets: Formats, Chat Templates, Tokenisation and PackingP11 Experiment Tracking and ReproducibilityP11 Lab: Your First Training RunP11 Memory Arithmetic for Training: Weights, Gradients, Optimiser States and ActivationsP11 PyTorch, Transformers, Datasets and the Hugging Face Ecosystem
- TRL documentation — SFT Trainerretrieved 2026-09-09
Sections: Expected dataset type and format; Train on completion only; Train adapters with PEFT; Expected dataset type and format; Train adapters with PEFT; Expected dataset type and format; Train on completion only; Tool Calling with SFT; Train on assistant messages only; Train on assistant messages only; Tool Calling with SFT; Train adapters with PEFT
Used by:P15 Lab: Sequence-Level Distillation of a 30B-Class Teacher into a 4B StudentP15 Project: The Distillation PipelineP15 Reasoning Distillation: Teacher Traces as Training DataP27 Distilling a Big Agent Model into a Small Local OneP27 Fine-Tuning for Tool Use and Your CodebaseP27 Lab: Fine-Tune a Small Model on Your Own Agent Trajectories and Measure the Gain
- TRL documentation — vLLM integrationretrieved 2026-09-09
Sections: How TRL uses the server; Modes of using vLLM during training; Supported versions; Modes of using vLLM during training; Supported versions; How TRL uses the server; Modes of using vLLM during training; Advanced usage
Used by:P14 Reinforcement Learning with Verifiable Rewards: GRPO ExplainedP14 Lab: GRPO on a Maths or Code Task on One MachineP14 The RL Toolchain: TRL, Unsloth, verl, OpenRLHF and vLLM Rollouts
- TRL v1.12.0 documentation - General Online Logit Distillation (GOLD) Trainerretrieved 2026-09-12
Sections: Overview; How Token Merging Works
Used by:P3 Distillation, Pruning and Quantisation: How Small Models Get Good
- turboderp/Qwen3-8B-exl3 model repository (branch 4.0bpw)retrieved 2026-09-13
Used by:P8 Lab: Same Model, Every Engine
- unsloth/Qwen3-1.7B-GGUF file listingretrieved 2026-09-13
- unsloth/Qwen3-14B-GGUF file listingretrieved 2026-09-12, 2026-09-13
Used by:P4 Choosing a Model for a Memory BudgetP4 Lab: Build Your Model Shortlist
- unsloth/Qwen3-235B-A22B-GGUF file listingretrieved 2026-09-12
- unsloth/Qwen3-235B-A22B-GGUF file listing (sharded; sizes summed per quantisation)retrieved 2026-09-12
- unsloth/Qwen3-235B-A22B-GGUF model repositoryretrieved 2026-09-09, 2026-09-13
Sections: IQ4_XS directory, three shards; Files and versions; the IQ4_XS directory; Files and versions; IQ4_XS; config.json
Used by:P6 Lab: Run and Benchmark the Course Reference ModelsP19 Lab: Run a Model Bigger Than Any One MachineP19 llama.cpp RPC: Layers Across Machines
- unsloth/Qwen3-30B-A3B-GGUF - file listing; tensor table of Qwen3-30B-A3B-Q4_K_M.ggufretrieved 2026-09-12
Used by:P3 Inference: Prefill, Decode and Why Memory Bandwidth Rules
- unsloth/Qwen3-30B-A3B-GGUF file listingretrieved 2026-09-12, 2026-09-13
Used by:P4 Choosing a Model for a Memory BudgetP4 Dense, Mixture-of-Experts and Hybrid ArchitecturesP4 Lab: Build Your Model Shortlist
- unsloth/Qwen3-30B-A3B-GGUF model repositoryretrieved 2026-09-09
Sections: Files and versions
Used by:P19 Lab: A Mixed-Platform ClusterP19 Lab: Run a Model Bigger Than Any One Machine
- unsloth/Qwen3-32B-GGUF — Qwen3-32B-UD-IQ2_XXS.gguf header, read by HTTP range requestretrieved 2026-09-13
Sections: tensor table - token_embd.weight Q3_K, output.weight Q5_K, 9,271,120,896 tensor bytes
- unsloth/Qwen3-32B-GGUF file listingretrieved 2026-09-12, 2026-09-13
Used by:P4 Choosing a Model for a Memory BudgetP4 Dense, Mixture-of-Experts and Hybrid ArchitecturesP4 Lab: Build Your Model Shortlist
- unsloth/Qwen3-4B-GGUF file listingretrieved 2026-09-12, 2026-09-13
Used by:P4 Choosing a Model for a Memory BudgetP4 Lab: Build Your Model Shortlist
- unsloth/Qwen3-8B-GGUF - file listing; tensor table of Qwen3-8B-Q4_K_M.ggufretrieved 2026-09-12
Used by:P3 Inference: Prefill, Decode and Why Memory Bandwidth Rules
- unsloth/Qwen3-8B-GGUF — file listing (Hub tree API)retrieved 2026-09-12, 2026-09-13
Sections: Qwen3-8B-Q4_K_M.gguf, 5,027,784,512 bytes; tensor list read from the file's first 16 MiB; file names, byte sizes and LFS SHA-256 for unsloth/Qwen3-4B, 8B, 14B, 30B-A3B, 32B and 235B-A22B GGUF and ggml-org/gpt-oss-20b and 120b GGUF
Used by:P2 Parameters, Layers and Model SizeP4 Choosing a Model for a Memory BudgetP4 Lab: Build Your Model ShortlistP4 Model Families and Who Makes ThemP6 Challenge: The Model That Runs at Two Tokens per SecondP6 Lab: Run and Benchmark the Course Reference Models
- unsloth/Qwen3-8B-GGUF model cardretrieved 2026-09-09, 2026-09-12, 2026-09-13
Sections: Files and quantisations; Recommended settings; Qwen3-8B-Q4_K_M.gguf file listing and GGUF tensor table
Used by:P4 Lab: Build Your Model ShortlistP4 Base, Instruct, Thinking, Coder, Vision, Embedding: Reading a Model NameP6 GGUF and Quantisation TypesP6 Installing and Building llama.cpp on Your PlatformP6 Lab: Run and Benchmark the Course Reference ModelsP6 Sampling: Temperature, Top-p, Min-p, Repetition and DeterminismP8 Lab: Same Model, Every Engine
- Z.ai on Hugging Face — organisation pageretrieved 2026-09-08
invariantlabs.ai
- Invariant Labs — MCP security notification, tool poisoning attacksretrieved 2026-09-09
Sections: Tool poisoning; rug pulls; Tool poisoning; rug pulls; tool shadowing; Definition; the addition-tool example; rug pull; shadowing; mitigations
Used by:P24 Lab: Write and Connect an MCP ServerP24 Model Context Protocol: Servers, Clients and TransportsP26 Safety: Prompt Injection, Tool Permissions and Human-in-the-Loop
jeffgeerling.com
- 1.5 TB of VRAM on Mac Studio - RDMA over Thunderbolt 5retrieved 2026-09-09
Sections: Enabling RDMA; Stability Issues; Baseline; HPL and Llama.cpp; Enabling RDMA; Stability Issues
Used by:P21 exo: Automatic Partitioning Across Your MacsP21 Reality Check: 'Four Mac Studios Replace a GPU Server'
jetbrains.com
jupyterlab.readthedocs.io
- JupyterLab documentation — Installationretrieved 2026-09-09, 2026-09-12
Used by:P1 Lab: Your Python Environment and a First Trained ModelP11 Setting Up a Training Environment on Each Platform
- JupyterLab documentation — Starting JupyterLabretrieved 2026-09-12
Sections: jupyter lab; the working directory
Used by:P1 Lab: Your Python Environment and a First Trained Model
keith.github.io
- sandbox-exec manual pageretrieved 2026-09-09
Sections: Deprecation notice
Used by:P25 Lab: Sandbox Your Agent: Containers, Permissions and Secrets
kilo.ai
- Kilo — Ollama providerretrieved 2026-09-09
Used by:P25 Editor Agents: Cline, Kilo Code, Continue and the Proprietary Editors
langfuse.com
- Langfuse — Documentationretrieved 2026-09-09
Sections: What Langfuse is; self-hosting; traces, sessions and scores; OpenTelemetry
Used by:P26 Evaluating Agents: Trajectories, Success Rates and Cost
learn.chatgpt.com
- Codex — advanced configurationretrieved 2026-09-09
Sections: OSS mode; provider examples
- Codex — agent approvals and securityretrieved 2026-09-09
Sections: Approval policies; bypass flag; Approval policies; CLI flags; codex exec; sandbox and approval flags
Used by:P25 Challenge: The Agent That Escaped the SandboxP25 OpenAI Codex CLI and OpenCode with Local ModelsP25 Lab: One Task, Six Agents
- Codex — AGENTS.mdretrieved 2026-09-09
Sections: Discovery precedence; size limit
- Codex — configuration referenceretrieved 2026-09-09
Sections: model_providers; oss_provider; wire_api; mcp_servers; mcp_servers; sandbox_mode; approval_policy; model_providers; oss_provider; wire_api
Used by:P25 OpenAI Codex CLI and OpenCode with Local ModelsP25 Project: A Local Agentic Coding WorkstationP25 The Landscape: Terminal, Editor and Autonomous Agents
- Codex — sandboxingretrieved 2026-09-09
Sections: Sandbox modes
- OpenAI Codex — MCPretrieved 2026-09-09
Sections: codex mcp add; mcp_servers in config.toml
learn.microsoft.com
- Advanced settings configuration in WSLretrieved 2026-09-13
Sections: .wslconfig; memory and swap defaults; the 8 second rule; wsl --shutdown
Used by:P6 Challenge: The Model That Runs at Two Tokens per Second
- Microsoft Learn — Accessing network applications with WSLretrieved 2026-09-13
Sections: Mirrored mode networking; Hyper-V firewall (New-NetFirewallHyperVRule)
Used by:P7 Lab: A Private Chat Service for Your Home Network
- Microsoft Learn — Basic commands for WSLretrieved 2026-09-12, 2026-09-13
Sections: wsl --update; wsl --status; running a Linux command as wsl <command>; wsl --install; wsl --update; wsl --status
Used by:P1 Lab: Your Python Environment and a First Trained ModelP5 Lab: Prepare Your Machine
lemonade-server.ai
- Lemonade — CLI referenceretrieved 2026-09-09, 2026-09-13
Sections: pull; load; unload; status; backends install; global options
Used by:P8 Lab: Same Model, Every EngineP8 AMD-Native: ROCm Builds, Lemonade Server and the NPU Question
- Lemonade — llama.cpp backend optionsretrieved 2026-09-13
Used by:P8 Lab: Same Model, Every Engine
- Lemonade — OpenAI-compatible APIretrieved 2026-09-09, 2026-09-13
Used by:P8 Lab: Same Model, Every EngineP8 AMD-Native: ROCm Builds, Lemonade Server and the NPU Question
- Lemonade — project siteretrieved 2026-09-09
Used by:P8 AMD-Native: ROCm Builds, Lemonade Server and the NPU Question
librechat.ai
- LibreChat documentationretrieved 2026-09-08
Sections: Installation; custom endpoints
livecodebench.github.io
- LiveCodeBenchretrieved 2026-09-08, 2026-09-09
Sections: Scenarios; contamination; Date-stamped problems; contamination analysis
Used by:P4 Reading a Model Card and a BenchmarkP16 Evaluation Harnesses: lm-evaluation-harness, lighteval, EvalPlus and Your OwnP16 LLM-as-Judge, Contamination and Honest Reporting
lmarena.ai
- LMArena text leaderboard (overall, style control)retrieved 2026-09-12
lmstudio.ai
- LM Studio — Downloadretrieved 2026-09-08
- LM Studio — Tool useretrieved 2026-09-09
Sections: Native and default tool use; supported models; streaming
- LM Studio Docs — Download an LLMretrieved 2026-09-08
- LM Studio Docs — full documentation text (SDK prediction config, contextOverflowPolicy)retrieved 2026-09-13
- LM Studio Docs — Headless / service moderetrieved 2026-09-08
- LM Studio Docs — Headless Moderetrieved 2026-09-08
Sections: llmster daemon; run on login; just-in-time model loading
- LM Studio Docs — homeretrieved 2026-09-08, 2026-09-09
Sections: Overview; llama.cpp and MLX engines; Developer; OpenAI compatibility API
Used by:P7 LM Studio: GUI, MLX and Headless ServingP9 Tool Calling and Structured Output on the Server Side
- LM Studio Docs — Import Modelsretrieved 2026-09-08
Sections: Expected directory structure
Used by:P7 LM Studio: GUI, MLX and Headless ServingP7 Managing a Model Library: Storage, Naming and Versions
- LM Studio Docs — lms CLIretrieved 2026-09-08
Sections: Subcommands and flags
Used by:P7 LM Studio: GUI, MLX and Headless ServingP7 Managing a Model Library: Storage, Naming and Versions
- LM Studio Docs — lms loadretrieved 2026-09-13
- LM Studio Docs — MCP Serversretrieved 2026-09-08, 2026-09-09
Sections: mcp.json; cautions; Configuration; cautions
Used by:P7 LM Studio: GUI, MLX and Headless ServingP24 Lab: Write and Connect an MCP ServerP24 Model Context Protocol: Servers, Clients and Transports
- LM Studio Docs — OpenAI Compatibility APIretrieved 2026-09-08
Sections: Supported endpoints; base URL and port
- LM Studio Docs — REST API v1, list your modelsretrieved 2026-09-13
- LM Studio Docs — System Requirementsretrieved 2026-09-08
lmsys.org
- Fast and Expressive LLM Inference with RadixAttention and SGLangretrieved 2026-09-09
Sections: RadixAttention; cache reuse patterns
Used by:P17 Prefix Caching and KV Reuse
maa.org
- Mathematical Association of America — AIME descriptionretrieved 2026-09-12
man.freebsd.org
- FreeBSD manual pages - netstat (1)retrieved 2026-09-09
Sections: Options -i, -b, -I
- FreeBSD manual pages — ping (8)retrieved 2026-09-09
Sections: Options -c, -s, -D, -q
man.openbsd.org
- OpenBSD manual pages — ssh-keygen (1)retrieved 2026-09-09
Sections: Options -t, -f, -C
man7.org
- setsebool(8) manual pageretrieved 2026-09-13
Sections: -P
Used by:P7 Lab: A Private Chat Service for Your Home Network
manpages.ubuntu.com
- Ubuntu manual pages — ip-link (8)retrieved 2026-09-09
Sections: ip link set; ip link show
- Ubuntu manual pages — ping (8)retrieved 2026-09-09
Sections: Options -c, -s, -M, -q
- Ubuntu manual pages — ssh-copy-id (1)retrieved 2026-09-09
Sections: Synopsis; -i
microsoft.github.io
- AutoGen — Modelsretrieved 2026-09-09
Sections: OpenAI-compatible endpoints note; model_info
Used by:P26 Agent Frameworks Compared
ml-explore.github.io
- MLX 0.32.2 documentation — mlx.core.device_inforetrieved 2026-09-13
Used by:P5 Lab: Measure Your Memory Bandwidth and ComputeP5 Lab: Prepare Your Machine
- MLX 0.32.2 documentation — mlx.core.quantizeretrieved 2026-09-13
Sections: modes affine, mxfp4, mxfp8, nvfp4; group sizes and bits per mode; biases only for affine
- MLX 0.32.2 documentation — mlx.core.quantized_matmulretrieved 2026-09-13
- MLX documentation — Build and Installretrieved 2026-09-09, 2026-09-12
Sections: Python Installation; Requirements; the Rosetta check; the mlx[cpu] Linux build; Python Installation; Requirements
Used by:P1 Lab: Your Python Environment and a First Trained ModelP11 Setting Up a Training Environment on Each Platform
- MLX documentation — Devices and Streamsretrieved 2026-09-12
Sections: default_device; set_default_device
Used by:P1 Lab: Your Python Environment and a First Trained Model
- MLX documentation — Distributed Communicationretrieved 2026-09-09
Sections: Thunderbolt ring; JACCL backend; Ring backend; JACCL backend; Backends; hostfile; RDMA over Thunderbolt; topologies; Backends; Getting Started with JACCL; Getting Started with Ring; Thunderbolt Ring; Getting Started with JACCL; Enabling RDMA; Defining a Mesh; Backends; Selecting Backend; Getting Started with Ring; Getting Started with JACCL; Getting Started with MPI; Distributed Without mlx.launch
Used by:P18 Lab: Build and Measure Your Cluster NetworkP18 Networking for Home Clusters: 2.5 GbE to 200 GbE, Thunderbolt 5 and RDMAP18 The Course Reference Cluster and the Single-Machine PathP19 RDMA Transport, Tuning and Measuring the SplitP21 exo: Automatic Partitioning Across Your MacsP21 Lab: A Two-Mac Cluster over Thunderbolt 5P21 mlx.distributed: Ring, MPI and RDMA over Thunderbolt 5P21 Reality Check: 'Four Mac Studios Replace a GPU Server'
- MLX documentation — Function Transformsretrieved 2026-09-12
- MLX documentation — index (unified memory model)retrieved 2026-09-09, 2026-09-12
Sections: Key features
Used by:P1 Tensors, GPUs and Precision: Why Matrix Multiplication Is EverythingP8 MLX and mlx-lm: Apple's Native Path
- MLX documentation — Launching Distributed Programsretrieved 2026-09-09
Sections: mlx.distributed_config; mlx.launch; Ring Specifics; JACCL Specifics; mlx.distributed_config; mlx.launch; Providing Hosts; Ring Specifics; JACCL Specifics; MPI Specifics
Used by:P8 MLX and mlx-lm: Apple's Native PathP21 Lab: A Two-Mac Cluster over Thunderbolt 5P21 mlx.distributed: Ring, MPI and RDMA over Thunderbolt 5
- MLX documentation — Lazy Evaluationretrieved 2026-09-09
- MLX documentation — Loss functionsretrieved 2026-09-12
Sections: cross_entropy; reduction defaults to none
Used by:P1 Lab: Your Python Environment and a First Trained Model
- MLX documentation — Memory managementretrieved 2026-09-09
Used by:P5 Apple Silicon: Unified Memory, Metal and MLXP5 Lab: Measure Your Memory Bandwidth and Compute
- MLX documentation — mlx.core.set_wired_limitretrieved 2026-09-09, 2026-09-13
Used by:P5 Apple Silicon: Unified Memory, Metal and MLXP5 Lab: Measure Your Memory Bandwidth and ComputeP5 Lab: Prepare Your Machine
- MLX documentation — Neural Networks (mlx.nn)retrieved 2026-09-12
Sections: Module; mx.eval of parameters; value_and_grad; save_weights and load_weights
Used by:P1 Lab: Your Python Environment and a First Trained Model
- MLX documentation — Optimizersretrieved 2026-09-12
Sections: the update / mx.eval loop; SGD
Used by:P1 Lab: Your Python Environment and a First Trained Model
mlflow.org
- MLflow documentation — Trackingretrieved 2026-09-09
Sections: Runs and experiments; logging functions; mlflow server
modelcontextprotocol.io
- Model Context Protocol — Governance and stewardshipretrieved 2026-09-09
Sections: Project policies; technical governance; SEPs
Used by:P24 Model Context Protocol: Servers, Clients and Transports
- Model Context Protocol — Key changes in 2026-07-28retrieved 2026-09-09
Sections: Major changes; deprecated features
Used by:P24 Model Context Protocol: Servers, Clients and Transports
- Model Context Protocol — Security best practicesretrieved 2026-09-09
Sections: Token passthrough; local server compromise; scope minimisation
Used by:P26 Safety: Prompt Injection, Tool Permissions and Human-in-the-Loop
- Model Context Protocol — Specificationretrieved 2026-09-09
Sections: Overview; Security and Trust & Safety; Overview; features; Security and Trust & Safety
Used by:P24 Lab: Write and Connect an MCP ServerP24 Model Context Protocol: Servers, Clients and TransportsP24 What an Agent Is: The Loop, Tools and State
- Model Context Protocol — Specification (2025-06-18)retrieved 2026-09-09
Sections: Security and Trust and Safety; key principles
- Model Context Protocol — stdio transportretrieved 2026-09-09
Sections: Framing; stdout and stderr rules; shutdown
Used by:P24 Lab: Write and Connect an MCP ServerP24 Model Context Protocol: Servers, Clients and Transports
- Model Context Protocol — Streamable HTTP transportretrieved 2026-09-09
Sections: Security and endpoint; request metadata headers; Security and endpoint; sending and receiving messages
Used by:P24 Lab: Write and Connect an MCP ServerP24 Model Context Protocol: Servers, Clients and Transports
- Model Context Protocol — Toolsretrieved 2026-09-09
Sections: Deterministic ordering of tools/list; Tool definitions; error handling; security considerations; User interaction model; tool definitions; error handling; security considerations; User interaction model; tool definitions; error handling; User interaction model; untrusted annotations
Used by:P24 Context Engineering: Memory, Compaction and the KV BudgetP24 Lab: Write and Connect an MCP ServerP24 Model Context Protocol: Servers, Clients and TransportsP24 What an Agent Is: The Loop, Tools and StateP26 Safety: Prompt Injection, Tool Permissions and Human-in-the-Loop
- Model Context Protocol — Transports overviewretrieved 2026-09-09
Sections: Standard bindings; messages; custom transports
Used by:P24 Model Context Protocol: Servers, Clients and Transports
- Model Context Protocol — Versioningretrieved 2026-09-09
Sections: Revision states; the current version; negotiation
Used by:P24 Model Context Protocol: Servers, Clients and Transports
networking-docs.nvidia.com
- NVIDIA Networking — RDMA over Converged Ethernet (RoCE)retrieved 2026-09-09
Sections: RoCEv1; RoCEv2; flow control
Used by:P18 Networking for Home Clusters: 2.5 GbE to 200 GbE, Thunderbolt 5 and RDMA
news.lmarena.ai
- Does Style Matter? Disentangling style and substance in Chatbot Arenaretrieved 2026-09-12
nginx.org
- nginx — Configuring HTTPS serversretrieved 2026-09-09
Sections: The minimal server block; session cache; ssl_certificate; ssl_certificate_key; default protocols
Used by:P23 Routing and Model Management: llama-swap, LiteLLM, NGINX and Health ChecksP23 Security for Exposed Endpoints and the Model Supply Chain
- nginx — ngx_http_limit_req_moduleretrieved 2026-09-09
Sections: limit_req_zone; limit_req; limit_req_status; limit_req; burst; nodelay; limit_req_status
Used by:P23 Routing and Model Management: llama-swap, LiteLLM, NGINX and Health ChecksP23 Security for Exposed Endpoints and the Model Supply Chain
- nginx — ngx_http_proxy_moduleretrieved 2026-09-09
Sections: proxy_pass; proxy_set_header; proxy_buffering; proxy_read_timeout
Used by:P23 Routing and Model Management: llama-swap, LiteLLM, NGINX and Health Checks
nvidia.com
- NVIDIA DGX Spark product pageretrieved 2026-09-09
Sections: Specifications
Used by:P5 NVIDIA DGX Spark: GB10, DGX OS and the aarch64 CaveatP5 Why Memory, Not FLOPS, Decides What You Can RunP20 Connecting Two DGX Sparks over ConnectX-7
- NVIDIA GeForce graphics card comparisonretrieved 2026-09-09
Sections: Specifications, RTX 5090, RTX 5080, RTX 4090, RTX 4080
Used by:P5 NVIDIA Desktops and Laptops: VRAM Tiers, CUDA and WSL2P5 Why Memory, Not FLOPS, Decides What You Can RunP20 Multi-GPU Desktops: PCIe, Tensor Parallel Without NVLink and Expert Parallel
- NVIDIA GeForce RTX 5090retrieved 2026-09-13
Sections: Memory interface width 512-bit; 3352 AI TOPS
- NVIDIA H100 Tensor Core GPU — specificationsretrieved 2026-09-12
Sections: memory bandwidth; interconnect (PCIe Gen5); H100 SXM, BFLOAT16 Tensor Core 1,979 teraFLOPS with sparsity; max TDP
Used by:P1 Tensors, GPUs and Precision: Why Matrix Multiplication Is EverythingP3 Pretraining: Learning from Trillions of Tokens
- NVIDIA Nemotron Open Model Licenseretrieved 2026-09-12
Sections: Preamble; sections 2, 3, 7, 9 and 10 (v. December 15, 2025)
- NVIDIA RTX PRO 6000 Blackwellretrieved 2026-09-09, 2026-09-13
Sections: 1792 GB/sec memory bandwidth; 4000 TOPS, theoretical FP4 TOPS using sparsity; Specifications
Used by:P5 Lab: Measure Your Memory Bandwidth and ComputeP5 NVIDIA Desktops and Laptops: VRAM Tiers, CUDA and WSL2P5 Why Memory, Not FLOPS, Decides What You Can RunP20 Multi-GPU Desktops: PCIe, Tensor Parallel Without NVLink and Expert Parallel
nvidia.github.io
- TensorRT-LLM — Container imagesretrieved 2026-09-09, 2026-09-13
Sections: docker run flags; release image tag; Pre-built images on NGC
Used by:P8 Lab: Same Model, Every EngineP8 TensorRT-LLM, NIM and the DGX Spark PlaybooksP20 TensorRT-LLM and Dynamo on Spark Pairs
- TensorRT-LLM — Parallelism in TensorRT LLMretrieved 2026-09-09
Sections: Overview of parallelism strategies; attention module; FFN module; Wide-EP
- TensorRT-LLM — Quantizationretrieved 2026-09-09
Sections: Supported formats; hardware support matrix; ModelOpt
- TensorRT-LLM — Quick Start Guideretrieved 2026-09-09
- TensorRT-LLM — trtllm-serve CLI referenceretrieved 2026-09-09
Sections: serve; disaggregated; embeddings; serve
Used by:P8 Lab: Same Model, Every EngineP8 TensorRT-LLM, NIM and the DGX Spark PlaybooksP20 TensorRT-LLM and Dynamo on Spark Pairs
- TensorRT-LLM 1.2.1 — Guided decodingretrieved 2026-09-13
Sections: Online API, trtllm-serve
Used by:P8 Lab: Same Model, Every Engine
- TensorRT-LLM 1.2.1 — KV cache systemretrieved 2026-09-13
Sections: How much memory is allocated to KV cache; cross-request reuse
Used by:P8 Lab: Same Model, Every Engine
- TensorRT-LLM 1.2.1 — trtllm-serve CLI referenceretrieved 2026-09-13
Sections: serve options; --tool_parser; --reasoning_parser; --extra_llm_api_options
Used by:P8 Lab: Same Model, Every Engine
- TensorRT-LLM documentation — homeretrieved 2026-09-09
Sections: Getting started; deployment guide; features
ollama.com
openai.com
- Introducing SWE-bench Verifiedretrieved 2026-09-09
Sections: 500 samples; annotation; limitations
openai.github.io
- OpenAI Agents SDK — Modelsretrieved 2026-09-09
Sections: Using other LLM providers; common issues
Used by:P26 Agent Frameworks Compared
opencode.ai
- OpenCode — CLIretrieved 2026-09-09
Sections: run; models; serve; global flags; opencode run flags
Used by:P25 OpenAI Codex CLI and OpenCode with Local ModelsP25 Lab: One Task, Six Agents
- OpenCode — configurationretrieved 2026-09-09
Sections: model; small_model; instructions
Used by:P25 OpenAI Codex CLI and OpenCode with Local ModelsP25 Project: A Local Agentic Coding Workstation
- OpenCode — MCP serversretrieved 2026-09-09
Sections: Local and remote server configuration
- OpenCode — permissionsretrieved 2026-09-09
Used by:P25 Challenge: The Agent That Escaped the SandboxP25 OpenAI Codex CLI and OpenCode with Local Models
- OpenCode — providersretrieved 2026-09-09
Sections: Custom provider; Ollama; llama.cpp; LM Studio; Ollama num_ctx guidance
Used by:P25 OpenAI Codex CLI and OpenCode with Local ModelsP25 The Landscape: Terminal, Editor and Autonomous AgentsP25 Which Local Models Can Actually Drive an Agent
- OpenCode — rulesretrieved 2026-09-09
Sections: AGENTS.md precedence
opendatacommons.org
- Open Data Commons Attribution License summaryretrieved 2026-09-12
opensource.org
- Open Source AI Definition - FAQretrieved 2026-09-12
Sections: Kinds of training data; validation phase; legal nature of parameters
- The MIT Licenseretrieved 2026-09-12
- The Open Source AI Definition 1.0retrieved 2026-09-12
Sections: What is Open Source AI; Preferred form to make modifications
opentelemetry.io
- OpenTelemetry — Semantic conventions for generative AIretrieved 2026-09-09
Sections: Notice that the conventions have moved
Used by:P26 Evaluating Agents: Trajectories, Success Rates and Cost
packages.ubuntu.com
- Ubuntu Packages — linux-oem-24.04c in noble-updatesretrieved 2026-09-13
Sections: version 6.17.0-1032.32
Used by:P5 Lab: Prepare Your Machine
pipx.pypa.io
- pipx — CLI reference and examplesretrieved 2026-09-13
Sections: pipx install PACKAGE_SPEC with a version specifier
Used by:P5 Lab: Prepare Your Machine
proceedings.mlr.press
prometheus.io
- Prometheus — Alerting overviewretrieved 2026-09-09
Sections: Prometheus and Alertmanager
Used by:P23 Lab: Dashboards for Your ClusterP23 Observability: Metrics, Logs and Traces for LLM Serving
- Prometheus — Alerting practicesretrieved 2026-09-09
Sections: What to alert on; symptom-based alerting
- Prometheus — Alerting rulesretrieved 2026-09-09
Sections: Rule syntax; for; Rule syntax; for; labels; annotations
Used by:P23 Challenge: The 3 a.m. Out-of-MemoryP23 Lab: Dashboards for Your ClusterP23 Observability: Metrics, Logs and Traces for LLM Serving
- Prometheus — Configurationretrieved 2026-09-09
Sections: global; scrape_configs; static_configs; rule_files; global; scrape_configs; rule_files
Used by:P23 Lab: Dashboards for Your ClusterP23 Observability: Metrics, Logs and Traces for LLM Serving
- Prometheus — Getting startedretrieved 2026-09-09
Sections: Minimal configuration; default port; --config.file; Minimal configuration; default port; starting it
Used by:P23 Lab: Dashboards for Your ClusterP23 Observability: Metrics, Logs and Traces for LLM Serving
- Prometheus — Querying the HTTP APIretrieved 2026-09-09
Sections: Range queries; Instant and range queries; the JSON response
Used by:P23 Capacity Planning and Cost per Million Tokens at HomeP23 Challenge: The 3 a.m. Out-of-Memory
py.sdk.modelcontextprotocol.io
- MCP Python SDK — Client transportsretrieved 2026-09-09
Sections: Stdio subprocess; Streamable HTTP
- MCP Python SDK — Handling errorsretrieved 2026-09-09
Sections: ToolError; unhandled exceptions
- MCP Python SDK — Running your serverretrieved 2026-09-09
Sections: stdio and streamable-http; mcp dev; mcp run
pydantic.dev
- Pydantic — Getting startedretrieved 2026-09-08
Sections: BaseModel; validation errors; JSON Schema
Used by:P10 Project: A Private Document Question-Answering ServiceP10 Structured Output and JSON Mode
- Pydantic AI — Agentsretrieved 2026-09-09
Sections: Agent construction; running agents; usage limits; UsageLimits; request_limit; Agent construction; output_type; UsageLimits; run_sync
Used by:P26 Agent Frameworks ComparedP26 Multi-Agent Patterns: Router, Planner and Executor, Critic, Parallel Fan-OutP26 Project: A Multi-Agent System on Your Cluster
- Pydantic AI — OpenAI modelsretrieved 2026-09-09
Sections: OpenAI-compatible providers; OpenAIChatModel; Responses versus Chat Completions; OpenAIChatModel; OpenAIProvider; base_url
Used by:P26 Agent Frameworks ComparedP26 Project: A Multi-Agent System on Your Cluster
- Pydantic AI — Overviewretrieved 2026-09-09
Sections: Feature list; durable execution; model-agnostic
Used by:P26 Agent Frameworks Compared
- Pydantic AI — Toolsretrieved 2026-09-09
Sections: Registering tools; RunContext; docstring extraction; agent.tool; RunContext; docstring extraction
Used by:P26 Agent Frameworks ComparedP26 Project: A Multi-Agent System on Your Cluster
pymupdf.readthedocs.io
- PyMuPDF — Aboutretrieved 2026-09-08
Sections: Licensing; capabilities
Used by:P10 Vision, Speech and Documents: Multimodal Locally
pypi.org
- amd-debug-tools 0.2.21 — amd_debug/ttm.py and common.py (PyPI wheel)retrieved 2026-09-13
Sections: TTM_PARAM_PATH; MODPROBE_CONF_PATH; MAX_MEMORY_PERCENTAGE; gb_to_pages; set() regenerates the initramfs, clear() does not; --version
Used by:P5 Lab: Prepare Your Machine
- Distilabel on PyPIretrieved 2026-09-09
Sections: Release history
- PyPI — mlx-metal 0.32.2 release filesretrieved 2026-09-12
Sections: required by mlx 0.32.2 on Darwin; wheel sizes 42.5 MB (macOS 14 and 15) and 64.4 MB (macOS 26)
Used by:P1 Lab: Your Python Environment and a First Trained Model
- PyPI — onnxruntime 1.26.0 filesretrieved 2026-09-13
Sections: macOS wheels (macosx_14_0_arm64 only)
Used by:P7 Lab: A Private Chat Service for Your Home Network
- PyPI — torch 2.14.0 release files and dependenciesretrieved 2026-09-12
Sections: Wheel sizes per platform (x86_64 554.6 MB, aarch64 454.0 MB, macOS arm64 127.3 MB); requires_dist (cuda-toolkit, cudnn, nccl, cusparselt, nvshmem on Linux)
Used by:P1 Lab: Your Python Environment and a First Trained Model
- rustbpe on PyPIretrieved 2026-09-09
Sections: Published wheels
Used by:P12 Lab: Train a 10M to 125M Parameter Model in an Afternoon
raw.githubusercontent.com
- AnythingLLM — LICENSE fileretrieved 2026-09-08
- ggml - ggml.h at llama.cpp build b10867 (enum ggml_type)retrieved 2026-09-12
Used by:P3 Inference: Prefill, Decode and Why Memory Bandwidth Rules
- ggml — ggml-common.h at llama.cpp build b10867 (block_q4_0, block_q8_0, block_q4_K)retrieved 2026-09-12
Used by:P1 Tensors, GPUs and Precision: Why Matrix Multiplication Is Everything
- LibreChat — LICENSE fileretrieved 2026-09-08
- Llama 3.1 model card (meta-llama/llama-models)retrieved 2026-09-09
Sections: Instruction-tuned model evaluation results
Used by:P16 Lab: Run a Standard Benchmark Suite on Your Model
- llama.cpp — convert_hf_to_gguf.pyretrieved 2026-09-09
Sections: parse_args; Command-line arguments
Used by:P11 Lab: Your First Training RunP13 Lab: Fine-Tune a 1B to 4B Model to Follow Your FormatP13 Merging, Exporting and Quantising a Fine-Tuned Model
- llama.cpp — convert_lora_to_gguf.pyretrieved 2026-09-09
Sections: Command-line arguments
Used by:P13 Merging, Exporting and Quantising a Fine-Tuned Model
- lm-evaluation-harness — command-line interface documentationretrieved 2026-09-09
Sections: --apply_chat_template; --fewshot_as_multiturn; --gen_kwargs; --log_samples; Command-line flags; --log_samples; --seed; --limit
Used by:P16 Challenge: The Benchmark That LiedP16 Evaluation Harnesses: lm-evaluation-harness, lighteval, EvalPlus and Your OwnP16 Lab: Run a Standard Benchmark Suite on Your ModelP16 LLM-as-Judge, Contamination and Honest Reporting
- lm-evaluation-harness — GPQA task variantsretrieved 2026-09-09
Used by:P16 Evaluation Harnesses: lm-evaluation-harness, lighteval, EvalPlus and Your Own
- lm-evaluation-harness — gsm8k task configurationretrieved 2026-09-09
Sections: num_fewshot; generation_kwargs; filter_list
Used by:P16 Challenge: The Benchmark That LiedP16 Evaluation Harnesses: lm-evaluation-harness, lighteval, EvalPlus and Your OwnP16 Lab: Run a Standard Benchmark Suite on Your Model
- lm-evaluation-harness — ifeval task configurationretrieved 2026-09-09
Sections: metric_list
Used by:P16 Challenge: The Benchmark That LiedP16 Evaluation Harnesses: lm-evaluation-harness, lighteval, EvalPlus and Your OwnP16 Lab: Run a Standard Benchmark Suite on Your Model
- lm-evaluation-harness — new task guideretrieved 2026-09-09
Sections: Task YAML fields; output_type
Used by:P16 Evaluation Harnesses: lm-evaluation-harness, lighteval, EvalPlus and Your Own
- mlx-lm — convert.py argument definitionsretrieved 2026-09-09
Sections: setup_arg_parser; setup_arg_parser; --q-mode
Used by:P8 Lab: Same Model, Every EngineP8 MLX and mlx-lm: Apple's Native PathP16 Lab: Quantise Your Fine-Tune Five Ways and Measure EachP16 Post-Training Quantisation in Depth: GPTQ, AWQ, K-quants, imatrix and HQQP16 Quantisation-Aware Training and the Blackwell Formats: FP8, NVFP4 and MXFP4
- mlx-lm — fuseretrieved 2026-09-09
Sections: Command-line arguments
Used by:P13 Merging, Exporting and Quantising a Fine-Tuned Model
- mlx-lm — generate.py argument definitionsretrieved 2026-09-09
- mlx-lm — server.py argument definitions and routesretrieved 2026-09-09
- nanochat — nanochat/dataloader.pyretrieved 2026-09-09
Used by:P12 Anatomy of a Training Loop: nanochat Line by Line
- nanochat — nanochat/dataset.pyretrieved 2026-09-09
Sections: parquets_iter_batched; list_parquet_files
Used by:P12 Anatomy of a Training Loop: nanochat Line by LineP12 Data for Pretraining: FineWeb-Edu, Cosmopedia and Training a TokeniserP12 Project: A Domain Micro-Model
- nanochat — nanochat/gpt.pyretrieved 2026-09-09, 2026-09-12
Sections: estimate_flops docstring; forward; num_scaling_params; estimate_flops; num_scaling_params
Used by:P1 Tensors, GPUs and Precision: Why Matrix Multiplication Is EverythingP12 Anatomy of a Training Loop: nanochat Line by LineP12 Lab: Train a 10M to 125M Parameter Model in an AfternoonP12 Project: A Domain Micro-ModelP12 Scaling Laws and Compute Budgets at Home
- nanochat — nanochat/loss_eval.pyretrieved 2026-09-09
Sections: evaluate_bpb
Used by:P12 Anatomy of a Training Loop: nanochat Line by LineP12 Project: A Domain Micro-Model
- nanochat — nanochat/optim.pyretrieved 2026-09-09
Used by:P12 Anatomy of a Training Loop: nanochat Line by Line
- nanochat — nanochat/tokenizer.pyretrieved 2026-09-09
Used by:P12 Anatomy of a Training Loop: nanochat Line by Line
- nanochat — pyproject.tomlretrieved 2026-09-09
Used by:P12 Lab: Train a 10M to 125M Parameter Model in an Afternoon
- nanochat — runs/runcpu.shretrieved 2026-09-09
Used by:P12 Lab: Train a 10M to 125M Parameter Model in an Afternoon
- nanochat — runs/speedrun.shretrieved 2026-09-09
Used by:P12 Anatomy of a Training Loop: nanochat Line by LineP12 Data for Pretraining: FineWeb-Edu, Cosmopedia and Training a TokeniserP12 Lab: Train a 10M to 125M Parameter Model in an Afternoon
- nanochat — scripts/base_eval.pyretrieved 2026-09-09
Used by:P12 Anatomy of a Training Loop: nanochat Line by LineP12 Lab: Train a 10M to 125M Parameter Model in an Afternoon
- nanochat — scripts/base_train.pyretrieved 2026-09-09
Sections: Scaling laws and muP extrapolations; learning-rate schedule; training loop logging
Used by:P12 Anatomy of a Training Loop: nanochat Line by LineP12 Lab: Train a 10M to 125M Parameter Model in an AfternoonP12 Scaling Laws and Compute Budgets at Home
- nanochat — scripts/chat_cli.pyretrieved 2026-09-09
Used by:P12 Anatomy of a Training Loop: nanochat Line by Line
- nanochat — scripts/tok_eval.pyretrieved 2026-09-09
Used by:P12 Data for Pretraining: FineWeb-Edu, Cosmopedia and Training a Tokeniser
- nanochat — scripts/tok_train.pyretrieved 2026-09-09
Used by:P12 Anatomy of a Training Loop: nanochat Line by LineP12 Data for Pretraining: FineWeb-Edu, Cosmopedia and Training a TokeniserP12 Project: A Domain Micro-Model
- Ollama API documentationretrieved 2026-09-08
Sections: Generate a chat completion; Show model information; List running models
Used by:P7 Ollama: Models as a Service
- Ollama API documentation — OpenAI compatibilityretrieved 2026-09-08
Sections: Supported request fields; Setting the context size
Used by:P7 Ollama: Models as a Service
- Ollama documentation — CLI referenceretrieved 2026-09-08
Used by:P7 Ollama: Models as a Service
- Ollama documentation — Context lengthretrieved 2026-09-08
Used by:P7 Ollama: Models as a Service
- Ollama documentation — Dockerretrieved 2026-09-08
Sections: Nvidia GPU; AMD GPU; Vulkan support
Used by:P7 Lab: A Private Chat Service for Your Home NetworkP7 Ollama: Models as a Service
- Ollama documentation — FAQretrieved 2026-09-08
Sections: How can I expose Ollama on my network; setting environment variables on Mac; keep alive; context window size; concurrency; Where are models stored; how do I set them to a different location; Context window size; where models are stored; keep alive; concurrency; multiple GPUs; K/V cache quantisation
Used by:P7 Lab: A Private Chat Service for Your Home NetworkP7 Managing a Model Library: Storage, Naming and VersionsP7 Ollama: Models as a Service
- Ollama documentation — Hardware supportretrieved 2026-09-08
Sections: Nvidia; AMD Radeon; Metal; Vulkan GPU support; GPU selection
Used by:P7 Ollama: Models as a Service
- Ollama documentation — Linuxretrieved 2026-09-08
Sections: Install; ARM64 install; AMD GPU install; startup service; customising
Used by:P7 Ollama: Models as a Service
- Ollama documentation — macOSretrieved 2026-09-08
Sections: System requirements; filesystem requirements; troubleshooting
Used by:P7 Ollama: Models as a Service
- Ollama documentation — Modelfile referenceretrieved 2026-09-08
Sections: FROM — build from a GGUF file; Instructions; PARAMETER; TEMPLATE; SYSTEM; ADAPTER; LICENSE; REQUIRES
Used by:P7 Managing a Model Library: Storage, Naming and VersionsP7 Ollama: Models as a Service
- Ollama documentation at v0.33.3 — API (load a model, list running models, version)retrieved 2026-09-13
- Ollama documentation at v0.33.3 — Context lengthretrieved 2026-09-13
- Ollama documentation at v0.33.3 — FAQretrieved 2026-09-13
Sections: How can I specify the context window size; How do I configure Ollama server; concurrency; K/V cache quantization
- Ollama documentation at v0.33.3 — Modelfile referenceretrieved 2026-09-13
Sections: PARAMETER — valid parameters and values
- Ollama documentation at v0.33.3 — OpenAI compatibilityretrieved 2026-09-13
Sections: Supported request fields (reasoning_effort); Setting the context size
- Ollama documentation at v0.33.3 — Troubleshooting (where the logs are)retrieved 2026-09-13
- Open WebUI — LICENSE fileretrieved 2026-09-08
Sections: Clause 4
- PEFT — target-module validation implementationretrieved 2026-09-13
Sections: inject_adapter target-module validation
Used by:P13 Lab: Fine-Tune a 1B to 4B Model to Follow Your FormatP13 LoRA and QLoRA Explained
- Qwen2.5-Coder repositoryretrieved 2026-09-08
Sections: File-level code completion (fill in the middle)
Used by:P10 Local Coding Assistants: Autocomplete and Chat in Your Editor
rfc-editor.org
- RFC 8375: Special-Use Domain 'home.arpa.'retrieved 2026-09-08, 2026-09-09
Sections: Abstract; Sections 1 and 3; Sections 1 and 3
Used by:P7 Lab: A Private Chat Service for Your Home NetworkP18 Lab: Build and Measure Your Cluster NetworkP21 Lab: A Two-Mac Cluster over Thunderbolt 5
rocm.docs.amd.com
- AMD — Install Ryzen Software for Linux with ROCm (ROCm on Radeon and Ryzen, 7.2.1)retrieved 2026-09-13
Sections: Prepare the system; amdgpu-install; --no-dkms; groups; rocminfo; Configure shared memory; amd-ttm
Used by:P5 Lab: Prepare Your Machine
- AMD ROCm 10.0.0 — Install AMD ROCmretrieved 2026-09-13
Sections: OEM kernel for Ryzen APUs; uninstall ROCm 7.2.4 or older first
Used by:P5 Lab: Prepare Your Machine
- AMD SMI — Using the AMD SMI CLI toolretrieved 2026-09-09
Sections: metric -p for power; JSON output; metric -p; metric -v; JSON output; metric; power and memory; JSON output
Used by:P23 Capacity Planning and Cost per Million Tokens at HomeP23 Lab: Dashboards for Your ClusterP23 Observability: Metrics, Logs and Traces for LLM Serving
- AMD SMI documentationretrieved 2026-09-09
Used by:P12 Lab: Train a 10M to 125M Parameter Model in an Afternoon
- AMD SMI documentationretrieved 2026-09-09
Sections: Relationship to ROCm SMI
Used by:P23 Observability: Metrics, Logs and Traces for LLM Serving
- ROCm — System requirements (Linux)retrieved 2026-09-09
Sections: Supported GPUs; supported operating systems
Used by:P5 AMD Ryzen AI Max+ 395: Strix Halo, ROCm and VulkanP8 AMD-Native: ROCm Builds, Lemonade Server and the NPU Question
- ROCm 10.0.0 compatibility matrixretrieved 2026-09-09, 2026-09-13
Sections: System requirements and information; AMD APU series; Ryzen APU; AMD Ryzen AI Max+ 395 (Radeon 8060S) (gfx1151); supported Ubuntu versions 26.04 and 24.04.4; inbox kernel driver; Supported GPUs; supported operating systems; AMD APU series; supported operating systems
Used by:P5 AMD Ryzen AI Max+ 395: Strix Halo, ROCm and VulkanP5 Lab: Prepare Your MachineP6 Challenge: The Model That Runs at Two Tokens per SecondP6 Installing and Building llama.cpp on Your PlatformP8 AMD-Native: ROCm Builds, Lemonade Server and the NPU QuestionP11 Setting Up a Training Environment on Each PlatformP12 Lab: Train a 10M to 125M Parameter Model in an Afternoon
- ROCm documentation — Installing PyTorch for ROCmretrieved 2026-09-09, 2026-09-12
Sections: Using a wheels package (nightly rocm7.2 index); Testing the PyTorch installation; docker, video and render groups; Using a wheels package; Using a wheels package; Using a Docker image; Testing the PyTorch installation
Used by:P1 Lab: Your Python Environment and a First Trained ModelP8 AMD-Native: ROCm Builds, Lemonade Server and the NPU QuestionP11 Lab: Your First Training RunP11 Setting Up a Training Environment on Each Platform
- ROCm documentation — Post-installation instructionsretrieved 2026-09-12
Sections: rocminfo; amd-smi version
Used by:P1 Lab: Your Python Environment and a First Trained Model
- ROCm documentation — Prerequisitesretrieved 2026-09-12
Sections: Configuring permissions for GPU access; the statement on integrated graphics
Used by:P1 Lab: Your Python Environment and a First Trained Model
- ROCm documentation — Quick start installation guideretrieved 2026-09-09, 2026-09-12
Sections: amdgpu driver and ROCm packages per distribution; usermod for the render and video groups
Used by:P1 Lab: Your Python Environment and a First Trained ModelP5 AMD Ryzen AI Max+ 395: Strix Halo, ROCm and Vulkan
ryzenai.docs.amd.com
- AMD Ryzen AI Software — Installation instructionsretrieved 2026-09-09
Used by:P8 AMD-Native: ROCm Builds, Lemonade Server and the NPU Question
- AMD Ryzen AI Software — LLM flow overviewretrieved 2026-09-09
Used by:P8 AMD-Native: ROCm Builds, Lemonade Server and the NPU Question
- AMD Ryzen AI Software documentationretrieved 2026-09-09
Used by:P5 AMD Ryzen AI Max+ 395: Strix Halo, ROCm and Vulkan
- AMD Ryzen AI Software documentationretrieved 2026-09-09
Used by:P8 AMD-Native: ROCm Builds, Lemonade Server and the NPU Question
scikit-learn.org
- scikit-learn — Choosing the right estimatorretrieved 2026-09-13
Used by:AI Problem-Solving Map
- scikit-learn — Metrics and scoringretrieved 2026-09-13
Used by:AI Problem-Solving Map
simonwillison.net
- Simon Willison — The lethal trifecta for AI agentsretrieved 2026-09-09
Sections: The three capabilities; The three capabilities; the advice
Used by:P24 Model Context Protocol: Servers, Clients and TransportsP26 Project: A Multi-Agent System on Your ClusterP26 Safety: Prompt Injection, Tool Permissions and Human-in-the-Loop
software.es.net
- iperf3 — Invoking iperf3retrieved 2026-09-09
Sections: Options; defaults; Options
Used by:P18 Lab: Build and Measure Your Cluster NetworkP18 Networking for Home Clusters: 2.5 GbE to 200 GbE, Thunderbolt 5 and RDMA
sqlite.org
- SQLite — Write-Ahead Loggingretrieved 2026-09-13
Sections: Checkpointing; the last connection closing
Used by:P7 Lab: A Private Chat Service for Your Home Network
support.apple.com
- Apple — set up other users on your Macretrieved 2026-09-09
Used by:P25 Lab: Sandbox Your Agent: Containers, Permissions and Secrets
support.google.com
- Android Help — Add & remove certificatesretrieved 2026-09-13
Used by:P7 Lab: A Private Chat Service for Your Home Network
swebench.com
- SWE-benchretrieved 2026-09-12
Sections: Leaderboard results data, bash-only entry 20250807_mini-v1.7.0_gpt-oss-120b
- SWE-bench Verified leaderboard descriptionretrieved 2026-09-09, 2026-09-12
Sections: Description; leaderboard entries
Used by:P4 Reading a Model Card and a BenchmarkP25 Which Local Models Can Actually Drive an Agent
tbench.ai
- Terminal-Benchretrieved 2026-09-09, 2026-09-12
Sections: Leaderboard; benchmark description
Used by:P4 Reading a Model Card and a BenchmarkP25 Which Local Models Can Actually Drive an Agent
tensorflow.org
- TensorBoard — Get startedretrieved 2026-09-09
unsloth.ai
- Unsloth documentation — AMD installationretrieved 2026-09-09
Sections: Supported GPUs; ROCm versions; Supported GPUs; ROCm versions; installation
Used by:P13 Lab: Fine-Tune a 1B to 4B Model to Follow Your FormatP13 Unsloth, Axolotl, LLaMA-Factory and mlx-lm: Higher-Level Tools
- Unsloth documentation — Fine-tuning LLMs with NVIDIA DGX Spark and Unslothretrieved 2026-09-09
Sections: Docker image; models trained
Used by:P13 Unsloth, Axolotl, LLaMA-Factory and mlx-lm: Higher-Level Tools
- Unsloth documentation — homeretrieved 2026-09-09
Sections: Overview; supported platforms
Used by:P13 Unsloth, Axolotl, LLaMA-Factory and mlx-lm: Higher-Level Tools
- Unsloth documentation — macOS installationretrieved 2026-09-09
Sections: Unsloth Desktop; Unsloth Studio
Used by:P13 Unsloth, Axolotl, LLaMA-Factory and mlx-lm: Higher-Level Tools
- Unsloth documentation — notebooksretrieved 2026-09-09
Sections: Qwen3 notebooks
Used by:P13 Unsloth, Axolotl, LLaMA-Factory and mlx-lm: Higher-Level Tools
- Unsloth documentation — pip installretrieved 2026-09-09
Sections: Installation commands; requirements
Used by:P13 Unsloth, Axolotl, LLaMA-Factory and mlx-lm: Higher-Level Tools
- Unsloth documentation — Reinforcement learning and GRPO guideretrieved 2026-09-09
Sections: Dataset and step recommendations; Memory requirements; model size guidance; dataset and step recommendations
Used by:P14 Lab: GRPO on a Maths or Code Task on One MachineP14 Reality Check: 'RL Makes Small Models Reason'P14 The RL Toolchain: TRL, Unsloth, verl, OpenRLHF and vLLM Rollouts
verl.readthedocs.io
- verl documentation — Agentic RL Trainingretrieved 2026-09-09
Sections: Agent loop; rollout backends; configuration
Used by:P27 Reinforcement Learning on Agent Tasks: Tests as Rewards
vulkan.lunarg.com
- Vulkan SDK — Getting started on Linuxretrieved 2026-09-13
Sections: mesa-vulkan-drivers vulkan-tools; Verify the SDK installation
Used by:P5 Lab: Prepare Your Machine
- Vulkan SDK — vulkaninforetrieved 2026-09-13
Sections: --summary
Used by:P5 Lab: Prepare Your Machine
zed.dev
- Zed — agent panelretrieved 2026-09-09
Sections: Tool calling
Used by:P25 Editor Agents: Cline, Kilo Code, Continue and the Proprietary Editors