ML Monitoring: Detecting Silent Failures in Production
When Your Model Fails Without Errors
ECE Graduate Students (ML/AI Track)
10 Minutes
Intermediate ML / Production Systems
Context for Hiring Committee
This demonstration addresses a critical gap in ML engineering education: students learn to train models with high validation accuracy but lack the systems thinking to maintain them in production. This lesson connects classical ML theory (statistical testing, distribution analysis) with modern MLOps practice (monitoring, drift detection, automated recovery), preparing students for industry roles where model reliability is paramount.
10-Minute Demo Structure
The Silent Failure (Hook)
2 min
📊 Visual: Project Monitoring Dashboard
"Look at this production model. January: 95% accuracy. February: 94%. March: 91%... June: 60%. But here's what's terrifying—" [switch to error logs] "—zero errors. Zero exceptions. The code runs perfectly. This is the nightmare of production ML: silent degradation."
🎯 PEDAGOGICAL STRATEGY:
- Cognitive Hook: Create dissonance—system is 'working' but failing
- Real-World Relevance: Students have used buggy software; this is different and scarier
- Visual Impact: Graph of declining accuracy with empty error logs makes abstract concrete
- Socratic Question: "Why is the model failing if there are no bugs?" forces causal reasoning
💬 Key Question to Class:
"If there are no code errors, what's failing? [Pause for responses] Exactly—the model's assumptions about the world. Traditional monitoring tracks uptime and errors. ML monitoring must track something else: statistical assumptions."
The Core Problem: Data Drift
3 min
📱 Concrete Example: COVID Impact on Fraud Detection
"Your fraud model was trained in 2019. Features: transaction_amount, time_of_day, merchant_category. COVID hits. Everyone shops online. High-value grocery orders. Midnight Amazon purchases. All your features SHIFTED. The model sees patterns it thinks are 'fraud' but they're actually normal now."
🎯 PEDAGOGICAL STRATEGY:
- Relatable Example: COVID changed behavior—students lived through this shift
- Multiple Modalities: Verbal explanation + histogram visualization + numerical data
- Conceptual Before Technical: Understand WHAT drift is before HOW to detect it
📊 Live Demo: Show Distribution Shift
Training Data (2019)
Mean transaction: $45
Std dev: $30
Midnight purchases: 5%
Production Data (2020)
Mean transaction: $180
Std dev: $120
Midnight purchases: 28%
Detection Method: The Kolmogorov-Smirnov Test
3 min
💻 Live Coding: Implement Drift Detection
from scipy.stats import ks_2samp
# Compare training vs production distributions
statistic, p_value = ks_2samp(
training_data['transaction_amount'],
production_data['transaction_amount']
)
if p_value < 0.05:
print(f"🚨 DRIFT DETECTED! p={p_value:.4f}")
print("Distributions are statistically different")
else:
print("✅ No significant drift")🎯 PEDAGOGICAL STRATEGY:
- Code Before Theory: Students see HOW to detect before diving into statistical mechanics
- Minimal Viable Implementation: 5 lines of code demonstrates core concept
- Immediate Feedback: p-value gives binary answer (drift or no drift)
- Active Learning: Students code along in their notebooks during demo
🧠Conceptual Explanation
"The KS test asks: 'Could these two datasets come from the same distribution?' It measures maximum distance between cumulative distributions. If p-value < 0.05, we reject the hypothesis that they're the same—drift detected."
[Show visualization: overlaid histograms with KS statistic highlighted]
Interactive Challenge: Debug the Mystery Failure
2 min
🎮 Student Activity: Pair Debugging
"Here's production data from last week. Your fraud model's false positive rate jumped from 2% to 18%. I've logged features for 1000 transactions. On your laptops: run KS test on each feature. Which feature drifted? You have 90 seconds. GO!"
# Students will discover: 'time_of_day' has p-value = 0.001 (massive drift)
# All other features: p-values > 0.4 (no drift)
# Conclusion: Time pattern changed, but model still uses 2019 patterns
🎯 PEDAGOGICAL STRATEGY:
- Active Application: Students immediately apply KS test to real scenario
- Time Pressure: 90 seconds creates urgency, simulates production debugging
- Discovery Learning: Students find the drifted feature themselves (not told)
- Peer Collaboration: Pairs discuss findings, compare p-values
- Bloom's Analysis Level: Debugging requires comparing multiple distributions, isolating root cause
Beyond Detection: Recovery Strategies
2 min
🔄 When Drift is Detected: Four Options
1. Automated Retraining
Use recent data, validate offline, deploy if better
When: PSI < 0.25, labels available
2. Model Rollback
Revert to previous version immediately
When: PSI > 0.25, severe accuracy drop
3. Human-in-the-Loop
Model predicts, human reviews before action
When: High-stakes (medical, financial)
4. Graceful Degradation
Fallback to rule-based system when uncertain
When: Model confidence < threshold
🎯 PEDAGOGICAL STRATEGY:
- Decision Framework: Not just technical solutions, but WHEN to use each
- Bloom's Evaluation: Students must judge appropriate response based on context
- Industry Connection: These are strategies used by Google, Netflix, Uber
- Quick Poll: "PSI=0.3 detected—what would you do?" → students vote → discuss
Synthesis & Takeaways
1 min
🎓 Key Insights for Students
- →ML ≠Traditional Software: No errors doesn't mean no failures. Statistical assumptions can break silently.
- →Monitor Distributions, Not Just Accuracy: Drift detection (KS test, PSI) catches problems before users complain.
- →Have a Response Plan: Detection without recovery strategy is incomplete. Define thresholds and actions.
🎯 PEDAGOGICAL CLOSURE:
- Explicit Synthesis: Recap three core concepts explicitly for retention
- Actionable Takeaway: Students leave knowing how to add drift detection to their projects
- Bridge to Next Topic: "This is just one failure mode—in a full course, we'd cover 4 more: training-serving skew, feedback loops, data leakage, and concept drift."
- Bloom's Scaffolding: Progresses from Understand → Apply → Analyze → Evaluate in 10 minutes
- Immediate Application: Students run real code during demo, not passive observation
- Real-World Context: COVID case study makes drift tangible and memorable
- Socratic Questioning: "Why?" before "How?"—builds conceptual understanding
- ECE 1505 (Convex Optimization): Statistical testing connects to hypothesis testing frameworks
- ECE 1508 (ML Systems): Production deployment challenges and monitoring infrastructure
- Industry Relevance: Prepares students for MLE roles at tech companies where model reliability is critical
During Demo (Continuous Check-ins):
- Minute 2: "Raise hand if this scenario has happened to you—model worked in dev, failed in prod"
- Minute 5: "Turn to neighbor: explain data drift in your own words"
- Minute 8: Quick poll—"Which recovery strategy for severe drift? Vote now."
Post-Demo (Exit Ticket - 2 minutes):
- "On a scale of 1-5, how well do you understand data drift detection now?"
- "Name one production ML failure mode you learned about today"
- "Would you feel confident adding drift monitoring to your next ML project?"
Week 1-2: Foundations
This 10-minute demo as lecture hook. Extended lab: implement complete monitoring system with dashboard.
Week 3-4: Advanced Topics
Cover remaining failure modes: training-serving skew, feedback loops, data leakage. Case studies from industry.
Final Project:
Deploy ML model with complete monitoring, trigger artificial drift, demonstrate automated recovery. Graded on system design and reliability.
Download the full 14-slide presentation covering:
- 4 types of silent ML failures (data drift, concept drift, training-serving skew, feedback loops)
- Hands-on KS test implementation code
- Monitoring dashboard design
- Recovery strategies and production ML tools
- Discussion questions for graduate seminars
📄 PDF format • 14 slides • Landscape orientation • Optimized for screen sharing
💻 Jupyter Notebook
Complete code for KS test, PSI calculation, dashboard creation
📊 Sample Dataset
Production data with induced drift for student debugging activity
📋 Facilitator Guide
Complete script with timing, questions, expected student responses
This demo is grounded in my research on human-AI collaboration and adaptive systems. The pedagogy reflects cognitive load theory (progressive disclosure), constructivism (students discover drift through debugging), and learning sciences (immediate feedback via code execution).
📚 Key References
- • Sculley et al. (2015) - Hidden Technical Debt in ML Systems
- • Breck et al. (2017) - ML Monitoring Best Practices
- • Bloom's Taxonomy (Revised) - Anderson & Krathwohl
🔬 My Related Research
- • Adaptive AI systems that detect and respond to distribution shifts
- • Human-agent collaboration in uncertain environments
- • Agent-centric pedagogy for production ML education
Why This Demonstrates Teaching Excellence
Deep Technical Content
Covers advanced ML engineering concepts (statistical testing, production systems) accessible to graduate students in 10 minutes through clear examples.
Active Learning Design
Students code, debug, and make decisions—not passive. 90-second pair activity ensures engagement and immediate application.
Industry-Relevant Skills
Addresses real pain point in industry: ML engineers who can train but not maintain models. Prepares students for production ML roles.
Prepared for University of Toronto ECE Department
Dr. Priyamvada Tripathi • Application for Assistant Teaching Professor