Build an AI serving platform, run it on two clouds, and move it into an environment that cannot reach the internet. Track A is what the job assumes you already know. Track B is the job itself: nineteen units, thirteen of them ending in something you built, measured or broke on purpose. This is a job simulator rather than a certification path, which is why almost every unit in Track B hands over a deliverable instead of a score.
1 of 19 units built
Track B unit 1 is ready: the rented-GPU harness is built and has been proven end to end, with a pod created, connected to, and destroyed with the meter verified off. The remaining eighteen units are written as specifications and their rigs are not built yet, so they are marked planned and say so on the page. Track A deliberately carries no reading lists yet: a link that has not been fetched does not go in a curriculum here, and enumerating them is its own piece of work.
Track A — what the job assumes you know
The knowledge the work leans on, stated once so the labs can assume it. Serving as a system rather than a model, two clouds as one subject with two vocabularies, the delivery spine, what crossing a boundary really costs, designing against a framework, and the half of a source-control platform that an administrator never meets.
A1
The serving stack: what actually runs a model in production
KRplanned
Almost everything written about running models is about training them. The job is serving: one process, holding a very large amount of GPU memory, answering requests with wildly variable cost. Until the shape of that is clear, every decision downstream is guesswork.
A2
Two clouds, one subject: Azure and Google Cloud as dialects
Kplanned
Learning two clouds as two subjects doubles the work and teaches neither well. Learned as one subject with two vocabularies, the second one costs a fraction of the first, and the comparison itself is what teaches the design tradeoff.
A3
The delivery spine: Kubernetes, Terraform, CI/CD, GitOps
Kplanned
This is the half most people bring to the job. It is here to be stated once and then assumed, not taught from scratch, so the units that follow can lean on it without pretending it is new.
A4
Build low, move high: what crossing a boundary actually costs
KRplanned
Moving software into a restricted environment is not a deployment problem, it is a provenance problem. Every image, chart, module, package and set of model weights has to arrive with a story about where it came from that a security officer will accept.
A5
Designing for the exam and the job: well-architected, and securing AI
KRplanned
Architecture certifications reward judgement rather than recall, and so does the job. The framework vocabulary is worth learning because it gives a shared language for arguing about a tradeoff, which is most of what a technical lead does.
A6
GitLab as a practitioner, not an administrator
Kplanned
Running the server and using the product are different skills, and years of the first teach almost none of the second. Where all the code, issues and work live in one platform, the user surface is the daily job.
Track B — the job itself
Four labs, thirteen units, every one ending in a deliverable. Rent a GPU and serve a model on Kubernetes, tune it one knob at a time, break it on purpose and diagnose it cold. Build the same landing zone in two clouds from code, with a pipeline that holds no keys. Move software across a boundary and refuse what is not signed. Then design an architecture from constraints and build it.
B1
Lab 1.1 — Stand up the rig, and prove you can destroy it
S ~45m
Every hour of GPU work after this one depends on being able to create a box and destroy it without thinking. Teardown discipline is not housekeeping here: a forgotten instance bills overnight, and the habit is the difference between a lab that costs pennies and one that costs a weekend.
After this you can
Create a GPU pod from a script, connect to it, and destroy it
Prove afterwards that nothing survived, rather than assuming the API call worked
What to watch for
The pod reports RUNNING before it has an IP address or port mappings. Anything scripted has to poll for the address rather than trust the status, which is the same class of mistake as trusting an HTTP 200 to mean the page changed.
Brief · B1-a
A verified create-and-destroy cycle
Stand up one GPU pod using the smallest card that fits the work, connect to it over SSH, confirm the GPU is visible to the operating system, then destroy the pod and prove the meter is off.
What you hand over
A terminal capture showing the pod created with its hourly cost, nvidia-smi output from inside it, and a teardown that reports no pods remaining.
Done when — read this before you start
The GPU name, memory and driver version were read from inside the pod, not from the provider console
Teardown was verified by re-listing pods, not by the delete call returning success
The provider console independently shows no running instances
B2
Lab 1.2 — Serve a model on Kubernetes with a real GPU
SRplanned
Running an inference engine in a container on one box is a solved problem and teaches little. Scheduling it on Kubernetes, where the GPU is a countable resource that has to be advertised, requested, tolerated and made visible to the container, is the thing the job actually asks for and the thing almost nobody has done.
B3
Lab 1.3 — The tuning ladder: one knob at a time
Splanned
Tuning is where opinion is thickest and evidence is thinnest. Changing one setting at a time and recording the number is unglamorous and is the only thing that separates a claim from a measurement.
B4
Lab 1.4 — Break it on purpose: the inference fault ladder
SRplanned
Building the stack is the smaller half of the job. Supporting it is the larger half, and it is the half nothing else on this platform teaches. A fault ladder is the only way to practise diagnosis without waiting for real outages.
B5
Lab 1.5 — Model weights are artifacts, and artifacts need provenance
Splanned
A set of weights is a multi-gigabyte binary from somewhere else, which is exactly the thing a restricted environment cares most about. Treating it as a supply chain problem rather than a download is what connects this lab to the boundary work later.
B6
Lab 1.6 — Multi-GPU and the interconnect
Splanned
This is the one unit that cannot be done on owned hardware. Splitting a model across cards, and measuring what the link between them actually contributes, requires a machine with more than one GPU and a real interconnect.
B7
Lab 2.1 — The same landing zone, twice, in two clouds
Splanned
Building the same thing in both clouds from code is the single most efficient exercise available: it teaches each cloud, teaches the difference between them, and proves the infrastructure is genuinely in code rather than merely described by it.
B8
Lab 2.2 — A pipeline that authenticates to both clouds with no keys
SRplanned
This single exercise covers continuous delivery, infrastructure as code, both clouds and the GitLab practitioner gap at once, and it removes the most common serious finding in any cloud review: a long-lived credential in a variable.
B9
Lab 2.3 — Managed Kubernetes with a GPU node pool
Splanned
This is where the two halves meet, and it is the closest thing in the whole path to the real job: the serving stack from Lab 1, running on a managed cluster built by the code from Lab 2.
B10
Lab 2.4 — Cloud break-fix, injected and reverted in seconds
SRplanned
Cloud faults are where most real incident time goes, and almost all of the cheap ones can be injected and reverted in seconds. This is the same drill as the inference ladder, pointed at the layer underneath it.
B11
Lab 3.1 — Sign it, verify it, and refuse it
SRplanned
Moving something across a boundary is only half the exercise. The half that matters to a security officer is proving that what arrived is what left, and that anything unsigned is rejected rather than merely noticed.
B12
Lab 3.2 — Size and updates: the two things that break real transfers
Splanned
A first install across a boundary is the easy case and the one every demonstration shows. The real programme is large artifacts and the second version, and both behave differently enough to invalidate the assumptions made during the first.
B13
Lab 4 — Build the case study instead of reading it
SRplanned
Published architecture case studies are business scenarios with constraints, which makes them buildable. Designing an architecture from constraints and then building it teaches the tradeoff in a way that reading the reference answer never does, and it doubles as landing zone practice.
This is a job simulator rather than a certification path. Track B does not resemble the work, it is the work: you rent a GPU, serve a model on Kubernetes, tune it one setting at a time, then break it on purpose and diagnose it cold. Almost every unit hands over a deliverable rather than a score, because nobody is paid to know that a device plugin exists.
Track B costs real money, unlike every other path here. GPU boxes and cloud accounts bill by the hour, so each lab ends by destroying what it built and verifying the meter is off. The first unit exists mainly to make that habit automatic before anything expensive depends on it.
Most units are still specifications rather than built rigs, and the page marks each one honestly. See the full list of learning paths for the ones with practice exams and browser labs already finished.