Inference Engineering
Inference Engineering on AI-ML Companion: Serve trained models in production - latency, throughput, and cost per token. 11 interactive modules with live visualizations, quizzes, and hands-on Python coding.
Start free: What Inference Engineering Is is fully open to everyone, no account required. The other 10 modules are part of AI-ML Companion Premium; every title and summary is listed below so you can see exactly what the track covers before deciding.
Modules in this track
- What Inference Engineering Is (free) - Prefill vs decode, the bandwidth ceiling, the KV cache, and the three numbers
- Inference Engine Internals (premium) - PagedAttention, continuous batching, FlashAttention v3, prefix caching, chunked prefill
- Quantization & Compression (premium) - FP16 to FP8 to INT8 to INT4, AWQ vs GPTQ vs SmoothQuant, FP8 hardware on H100/B200
- Collective Comms & Networking (premium) - NCCL allreduce, NVLink, NVSwitch, InfiniBand, RDMA - what makes multi-GPU serving work
- Hardware Fluency (premium) - H100/H200/B200 vs MI300X vs Trainium2 - HBM bandwidth, FP8 support, MIG economics
- GPU Scheduling on Kubernetes (premium) - KubeRay, gang scheduling, MIG management, Karpenter on GPU spot, topology-aware placement
- Multi-Tenant Routing & SLOs (premium) - Prefix-cache-aware routing, priority queues, SLO admission control, fairness vs head-of-line blocking
- LLM Observability + Cost (premium) - TTFT, ITL/TPOT, KV-cache hit rate, $/1M tokens - the metrics every inference platform needs
- LoRA Serving at Scale (premium) - S-LoRA, Punica - serving thousands of fine-tuned adapters on one base model
- Advanced Decoding (premium) - Speculative decoding, Medusa, EAGLE, constrained output - the latency wins beyond standard primitives
- Cold-Start & Model Loading (premium) - Safetensors, streaming load, Alluxio/Fluid, pre-warm pools - get 90s cold-start down to 5-10s