Mock Teaching Demonstration

Training at Scale: From Notebook to Production

Building Rigorous, Reproducible, and Cost-Efficient Training Systems

Audience

ECE Graduate Students (ML Systems Track)

Duration

10 Minutes

Level

Intermediate ML Engineering

Context for Hiring Committee

This demonstration addresses the gap between academic ML (training in notebooks with model.fit()) and production ML systems. Students learn to train models but lack the engineering rigor for production: experiment tracking, distributed training, cost optimization, and reproducibility. This lesson bridges that gap by teaching systematic evaluation, scalable infrastructure, and cloud-portable training pipelines—skills critical for modern ML engineering roles.

Learning Objectives (Bloom's Taxonomy)
Understand: Explain why model.fit() in a notebook is insufficient for production training systems
Apply: Implement experiment tracking with MLflow to ensure reproducibility across teams
Analyze: Calculate cost-performance tradeoffs when choosing distributed training strategies
Evaluate: Design evaluation frameworks with baselines, slicing, and cost-sensitive metrics for production models

10-Minute Demo Structure: Foundations-First Approach

1

The Problem: Why Notebook Training Fails at Scale

1 min

📂 Visual: Show File Explorer with 47 Jupyter Notebooks

"You just joined a startup. Your predecessor left 47 notebooks: 'model_final.ipynb', 'model_final_v2.ipynb', 'model_ACTUALLY_final.ipynb'. Which hyperparameters are in production? No one knows. Training takes 8 hours on a single GPU—$200 per run. One failed experiment costs a day's work. This is the reality of unscalable training. Let's fix it."

🎯 PEDAGOGICAL STRATEGY:

  • Relatable Frustration: Students have lived this—messy notebooks, lost experiments
  • Real Cost: $200/run makes it tangible—wasting time = wasting money
  • Visual Chaos: File tree with duplicate notebooks creates visceral reaction
  • Problem Framing: "Let's fix it" signals this is solvable, not inevitable

💬 Key Question to Class:

"How many of you have lost track of which model version was 'the good one'? [Show of hands] Exactly. model.fit() doesn't give you experiment tracking, distributed training, or cost optimization. Production training needs infrastructure."

2

Theoretical Foundation 1: SGD Convergence & Batch Size Selection

2.5 min

📐 Optimization Theory

"Training is optimization: find θ* = argmin L(θ) where L is loss over dataset. SGD doesn't converge instantly—there's a mathematical rate. Understanding this helps you choose batch sizes, learning rates, and when to stop training."

🎯 PEDAGOGICAL STRATEGY:

  • Theory Before Practice: Math explains why distributed training doesn't scale linearly
  • Concrete Implications: Convergence rate determines training time—practical consequence
  • Visual Formulas: Show equations, then explain in plain English

🔬 SGD Convergence Rate

E[f(θ_T) - f(θ*)] ≤ O(1/√T)

where T = number of iterations

• Batch size B: Reduces gradient noise by factor √B

• Learning rate α: Must decay as α_t = α₀/√t for convergence

• Convergence: Need O(1/ε²) steps to reach ε-suboptimal solution

Translation: To halve your error, you need 4× more training steps. Doubling batch size gives you √2 better gradient estimates, not 2×. This is why large batches (B=32K) underperform—you need MORE steps to converge, not fewer.

⚖️ Batch Size Tradeoff Analysis

Training Time = (N_samples / B) × t_step(B)

where t_step(B) = t_compute(B) + t_comm

• Small batch (B=256): Many steps, low communication overhead

• Large batch (B=8192): Fewer steps, high communication overhead

• Optimal: B ∈ [512, 2048] balances noise reduction vs steps

B = 256 (Small)

Steps: 10,000

Gradient noise: High

Communication: 5% time

Better convergence

B = 8192 (Large)

Steps: 312

Gradient noise: Low

Communication: 40% time

Worse final accuracy

3

Theoretical Foundation 2: Amdahl's Law & Parallel Scaling Limits

2.5 min

⚡ The Scaling Question

"You have 8 GPUs. Naively, you expect 8× speedup. Reality: 5× if you're lucky. Why? Amdahl's Law—one of the most important theorems in parallel computing. It tells you the MAXIMUM possible speedup given serial bottlenecks."

🎯 PEDAGOGICAL STRATEGY:

  • Classic CS Theory: Amdahl's Law from architecture courses applies to ML
  • Realistic Expectations: Students learn why distributed training isn't magic
  • Formula to Intuition: Show math, then graph, then real numbers

📊 Amdahl's Law: Theoretical Speedup Limit

Speedup(N) = 1 / (S + P/N)

where:

• S = fraction of serial (non-parallelizable) work

• P = fraction of parallel work (P = 1 - S)

• N = number of workers (GPUs)

Even with N → ∞, Speedup ≤ 1/S

If S=0.25, max speedup is 4×, no matter how many GPUs!

For ML Training: Serial work includes data loading, gradient synchronization, checkpointing, logging. Typically S ≈ 0.20-0.30.

N = 2

1.7×

N = 4

3.2×

N = 8

4.6×

N = 64

5.8×

Diminishing returns: 64 GPUs gives 5.8×, not 64×! This is why Facebook used 256 GPUs but only got 30× speedup.

🔬 What Causes Serial Bottlenecks?

Gradient Synchronization (All-Reduce)

Every step, all GPUs must sync gradients. Can't parallelize this—inherently serial.

Data Loading & Preprocessing

CPUs bottleneck. Even with parallel dataloaders, I/O is slower than compute.

Model Initialization & Checkpointing

Single GPU writes checkpoint to disk. Others wait. Serialized by filesystem.

4

Theoretical Foundation 3: Communication Complexity Analysis

2 min

🌐 The Hidden Cost: Network Communication

"Adding GPUs means gradients must be synchronized across machines. For a 1 billion parameter model, that's 4GB of data (at FP32) transmitted EVERY STEP. If your network bandwidth is 10Gbps, that's 3.2 seconds of communication overhead per step. This is often LONGER than the forward+backward pass!"

🎯 PEDAGOGICAL STRATEGY:

  • Concrete Numbers: 4GB per step makes communication cost tangible
  • Complexity Analysis: O(P) communication ties to CS theory (algorithms course)
  • Practical Solutions: FP16, gradient compression, FSDP are engineering responses

📐 All-Reduce Communication Cost

Time_comm = (P × M × 2) / Bandwidth

where:

• P = number of parameters

• M = bytes per parameter (4 for FP32, 2 for FP16)

• 2 = ring all-reduce needs 2 passes

• Bandwidth = network speed (bytes/sec)

Example: GPT-2 (1.5B params), FP32, 100Gbps network

Time_comm = (1.5B × 4 × 2) / (100Gbps/8) = 960ms

If compute = 500ms, communication overhead is 192%!

⚡ Reducing Communication Overhead

Strategy 1: FP16 Mixed Precision

Halves communication (M = 2 instead of 4)

Speedup: 2× comm reduction, 1.7× overall speedup

Strategy 2: Gradient Compression

Send top-k largest gradients only (compress to 1% of size)

Speedup: 100× comm reduction, ~5× overall (with accuracy tradeoff)

Strategy 3: FSDP (Fully Sharded Data Parallel)

Shard model across GPUs—each GPU only stores 1/N parameters

Speedup: Reduces memory AND communication (allows larger batch sizes)

🔢 Bandwidth-Limited Training Time

Total Training Time Formula:

T_total = N_steps × (T_compute + T_comm)

As N_gpus increases, T_compute decreases (parallelism) but T_comm increases (more sync).

Optimal N_gpus occurs when dT_total/dN = 0 (calculus optimization)

For most networks: Optimal N ∈ [8, 32] GPUs. Beyond this, communication dominates.

5

From Theory to Practice: Building Production Pipelines

1.5 min

⚡ The 8-Hour Problem

"Single GPU: 8 hours, $200. But GPUs are parallel processors—why use just one? With 8 GPUs using data parallelism, training finishes in 1.5 hours (~5x speedup). Cost: $300 total, but you saved 6.5 hours. Time is money. But why NOT 8x speedup? Amdahl's Law explains."

📊 Amdahl's Law: Theoretical Speedup Limit

Speedup(N) = 1 / (S + P/N)

where:

• S = fraction of serial (non-parallelizable) work

• P = fraction of parallel work (P = 1 - S)

• N = number of workers (GPUs)

Even with N → ∞, Speedup ≤ 1/S

For ML Training: S = 0.25 (data loading, gradient sync, checkpointing takes 25%)

N = 4 GPUs

3.2x

N = 8 GPUs

4.6x

N = 16 GPUs

5.3x

Notice: Diminishing returns! 16 GPUs only gives 5.3x speedup (not 16x) because serial portion limits scaling.

🌐 Communication Complexity Analysis

Communication Cost per Step:

• All-Reduce gradient sync: O(P × M / bandwidth)

• P = model parameters (e.g., 175B for GPT-3)

• M = bytes per parameter (4 for FP32, 2 for FP16)

Example: 1B param model, FP32, 100Gbps network

Comm time = (1B × 4 bytes) / (100Gbps / 8) ≈ 320ms

If forward+backward = 500ms, comm is 39% overhead!

Solutions: (1) Use FP16 mixed precision (halves comm), (2) Gradient compression (trades accuracy for speed), (3) FSDP—shard parameters across GPUs to reduce memory and comm.

🎯 PEDAGOGICAL STRATEGY:

  • Simple Math: 8 hours → 1.5 hours is visceral—students get it immediately
  • Cost Analysis: $200 → $300 BUT 6.5 hours saved—teaches ROI thinking
  • Why Not 8x?: Communication overhead—sets up realistic expectations
  • Visual Aid: Timeline comparing single vs distributed training

🚀 Data Parallelism Strategy

❌ Single GPU

Batch size: 256

Steps: 10,000

Time per step: 3s

Total: 8.3 hours

✅ 8 GPUs (Data Parallel)

Batch size: 2048 (256×8)

Steps: 1,250

Time per step: 4s (+comm)

Total: 1.4 hours (5.9x speedup)

Why not 8x speedup? Communication overhead for gradient synchronization takes ~25% of time.

6

Experiment Tracking & Reproducibility

0.5 min

💰 The $10K Mistake

"On-demand GPUs: $3/hour. Run 100 experiments at 8 hours each = $2,400. Spot instances: Same GPU for $0.90/hour (70% cheaper). Same 100 experiments = $720. Saved $1,680. Tradeoff: Spot instances can be interrupted. Solution: Checkpointing. Let's analyze the math behind this decision."

📐 Cost-Performance Optimization

Total Cost = (Cost_compute + Cost_checkpointing) × (1 + P_interrupt × Overhead)

where:

• Cost_compute = price/hr × training_hours

• Cost_checkpointing = storage + I/O overhead

• P_interrupt = probability of spot interruption (~5%)

• Overhead = wasted work before checkpoint

On-Demand: $3/hr × 8hr = $24 (no interruption risk)

Spot: $0.90/hr × 8hr × 1.05 = $7.56 (5% interrupt tax)

Savings: $24 - $7.56 = $16.44 per run (68% reduction)

⚖️ Break-Even Analysis

When does checkpointing overhead exceed savings?

Let t_ckpt = checkpoint overhead (time to save state)

Break-even when: t_ckpt × spot_price > savings_per_checkpoint

For spot price = $0.90/hr, savings = $2.10/hr vs on-demand

Result: Checkpointing pays off if t_ckpt < 2.3× training time

(Typical t_ckpt = 1-2 min per epoch, always worth it!)

🎯 PEDAGOGICAL STRATEGY:

  • Concrete Savings: $1,680 saved is real money—students care about cost
  • Tradeoff Thinking: Cheaper but interruptible—teaches engineering constraints
  • Practical Solution: Checkpointing mitigates risk—actionable technique

📉 Cost Comparison Table

Instance TypeCost/Hour100 Experiments (8h)Risk
On-Demand GPU$3.00$2,400None
Spot Instance$0.90$720Interruptible (~5%)
Spot + Checkpointing$0.90$750 (+$30 overhead)Minimal (resume on interrupt)

Strategy: Use spot instances with checkpointing for 70% cost savings with negligible risk.

7

Synthesis & Production Workflow

2 min

🔄 End-to-End Training Pipeline

1.

Evaluation Design: Define baselines (majority, simple model) + metrics (cost-sensitive)

2.

Experiment Tracking: Use MLflow to log params, metrics, artifacts automatically

3.

Distributed Training: Scale to 8 GPUs with data parallelism (5-6x speedup)

4.

Cost Optimization: Use spot instances + checkpointing (70% savings)

5.

Deploy Best Model: Compare experiments in MLflow, deploy winner to production

🎯 PEDAGOGICAL CLOSURE:

  • Five-Step Framework: Repeatable workflow students can apply to any project
  • Actionable Takeaway: Students leave knowing how to build production training systems
  • Realistic Expectations: 5-6x speedup (not 8x) + 70% cost savings are achievable

🎓 Key Mathematical Insights for Students

  • →SGD Convergence: O(1/√T) rate means halving error requires 4× more steps—batch size affects noise, not convergence speed.
  • →Amdahl's Law: 25% serial work limits max speedup to 4×, no matter how many GPUs—communication is the bottleneck.
  • →Communication Complexity: O(P×M/bandwidth) per step—FP16 halves comm time, FSDP shards parameters.
  • →Cost Optimization: Break-even analysis: checkpointing overhead < savings → spot instances always win.
  • →Baselines Save Time: Always compare complex models to simple ones—avoid 3-month complexity traps.
Pedagogical Innovations
  • Cost-Benefit Framing: Every concept tied to time/money savings—students see practical value
  • Realistic Expectations: 5-6x speedup (not 8x), 70% savings—teaches engineering constraints
  • Chaos → Order Arc: Start with messy notebooks, end with production pipeline—satisfying narrative
  • Live Demos: MLflow UI, code execution, cost tables—multi-modal learning
ECE Program Alignment
  • ECE 1505 (Convex Optimization): SGD convergence theory connects to optimization foundations
  • ECE 1508 (ML Systems): Distributed training, cloud infrastructure, production pipelines
  • Industry Relevance: MLflow, distributed training, spot instances are standard at Google/Meta/Uber
Research-Informed Design

This demo reflects my research on adaptive AI systems and agent-centric pedagogy. The lesson structure embodies cognitive load theory (progressive complexity), constructivism (students discover cost tradeoffs), and experiential learning (live coding demos). By framing training as a systems engineering problem (not just algorithm tuning), students develop the production mindset critical for industry ML roles.

📚 Key References

  • • Sculley et al. (2015) - ML Systems Technical Debt
  • • Dean et al. (2012) - Large Scale Distributed Deep Networks
  • • Bloom's Taxonomy - Anderson & Krathwohl (2001)

🔬 My Related Research

  • • Adaptive AI systems with scalable training pipelines
  • • Agent-centric pedagogy for ML engineering education
  • • Cost-aware optimization for production ML systems

Why This Demonstrates Teaching Excellence

Real-World Impact

Students learn to save $1,680 in cloud costs and 6.5 hours per training run—skills directly applicable to industry ML engineering roles.

Systems Thinking

Bridges academic ML (model.fit()) and production engineering (distributed training, cost optimization, reproducibility)—critical gap in ML education.

Actionable Framework

Five-step production training pipeline students can immediately apply to research projects and internships—not just theory, but practical workflow.

Download Lecture Slides
Complete slide deck for "Training at Scale: Mathematical Foundations" workshop

Download the full 21-slide presentation covering:

  • SGD convergence theory and batch size selection
  • Amdahl's Law and parallel scaling limits
  • Communication complexity analysis and optimization strategies
  • Cost-performance tradeoffs and economic analysis
  • End-to-end production system design from first principles

📄 PDF format • 21 slides • Landscape orientation • Optimized for screen sharing

Additional Materials

💻 Jupyter Notebook

PyTorch code for distributed training with FSDP, checkpointing

📊 Benchmark Dataset

Performance metrics for different configurations

📋 Facilitator Guide

Complete script with timing and expected responses

References & Further Reading

Mathematical Foundations

  • Bottou, L. (2012). "Stochastic Gradient Descent Tricks." Neural Networks: Tricks of the Trade. Springer.[PDF]
  • Amdahl, G. M. (1967). "Validity of the single processor approach to achieving large scale computing capabilities." AFIPS Conference Proceedings, 30, 483-485.
  • Goyal, P., et al. (2017). "Accurate, Large Minibatch SGD: Training ImageNet in 1 Hour." arXiv:1706.02677.[arXiv]

Distributed Training Systems

  • Dean, J., et al. (2012). "Large Scale Distributed Deep Networks." NIPS 2012.[Link]
  • Rasley, J., et al. (2020). "DeepSpeed: System Optimizations Enable Training Deep Learning Models with Over 100 Billion Parameters." KDD 2020.[arXiv]
  • Zhao, Y., et al. (2023). "PyTorch FSDP: Experiences on Scaling Fully Sharded Data Parallel." VLDB 2023.[arXiv]

Production ML Systems

  • Sculley, D., et al. (2015). "Hidden Technical Debt in Machine Learning Systems." NIPS 2015.[Link]
  • Zaharia, M., et al. (2018). "Accelerating the Machine Learning Lifecycle with MLflow." IEEE Data Engineering Bulletin, 41(4).[Docs]
  • Baylor, D., et al. (2017). "TFX: A TensorFlow-Based Production-Scale Machine Learning Platform." KDD 2017.[Link]

Communication Optimization

  • Lin, Y., et al. (2018). "Deep Gradient Compression: Reducing the Communication Bandwidth for Distributed Training." ICLR 2018.[arXiv]
  • Micikevicius, P., et al. (2018). "Mixed Precision Training." ICLR 2018.[arXiv]

Educational Foundations

  • Anderson, L. W., & Krathwohl, D. R. (2001). "A Taxonomy for Learning, Teaching, and Assessing: A Revision of Bloom's Taxonomy of Educational Objectives." Longman.
  • Sweller, J. (1988). "Cognitive Load During Problem Solving: Effects on Learning." Cognitive Science, 12(2), 257-285.

Prepared for University of Toronto ECE Department

Dr. Priyamvada Tripathi • Application for Assistant Teaching Professor

© 2026 Dr. Priyamvada Tripathi. All rights reserved.

You are free to share and adapt this content with attribution for non-commercial purposes under the same license.