THE DEV BENCH

AI & GPU Infrastructure

Run the infrastructure behind AI — NVIDIA cert prep, MLOps, GPU job scheduling with Slurm, and the datacenter operations behind model training and inference.

Start here

AI/ML Vocabulary51 cards — step 1 of 16

Then get hands-on: Slurm lab 8 scenarios.

AI Infrastructure & Platform Engineer (LLMOps + GPU/Cloud Ops) — learning pathThe combined role in one place — a hands-on GPU/LLMOps lab track (ten runnable runbooks), the concept decks behind the vocabulary, and the cert paths that feed it. Free-tier-first: most labs run on a 16GB card at $0.NVIDIA AI Infrastructure & Operations — learning pathCert roadmap, decks, labs, a skill map, and verified reading — one workspace for the GPU-infra ops skill set.AI Hardware — Rack, Interconnect and Identification — the courseThe physical layer rather than the operations — safety, rack, node, the PCIe/SXM/OAM split, then cabling, testing and bring-up, which are the bulk of the real job. Anchored on NVIDIA's NCP-ARI technician blueprint. Syllabus scaffolded; units not built yet.

Flashcard decks

16 decks · 704 cards · in study order

Vocabulary — start here

1 deck · 51 cards

LLM foundations

1 deck · 16 cards

The physical stack & what it costs

3 decks · 97 cards

How the models work

1 deck · 30 cards

Model serving

1 deck · 35 cards

LLM app orchestration

1 deck · 34 cards

GPU operations

1 deck · 39 cards

Cloud AI platforms compared

1 deck · 31 cards

Legacy cert decks — being merged

4 decks · 313 cards

Interview prep

2 decks · 58 cards

Slurm lab

live
Job scheduling on a real single-node cluster — the scheduler that runs GPU clusters. Learn mode covers sbatch, resources and the queue; Break-fix mode fixes drained nodes, impossible resource requests and dead daemons.

Curated resources

verified July 2026

The core subject for the AI Infrastructure & Platform Engineer role. Vendor and project docs lead — this stack moves fast enough that third-party tutorials go stale within months. The two papers are worth reading in full; everything else is reference you will return to.

Model serving — the centre of the role

GPU operations

How the models work

LLM application layer

Certification