ML Model Health Dashboard
Production Fraud Detection Model - Real-time Monitoring
CRITICAL: Immediate Action Required
Last updated: 9:10:21 PM
Model Performance Over Time
- Accuracy (%)
- Precision (%)
- Recall (%)
64.4%
Current Accuracy
Below Threshold
67.5%
Precision
59.7%
Recall
Alert Triggered
Accuracy dropped below 80% threshold. Model performance has declined30.6% since deployment.
Implementation: Drift Detection Code
import numpy as np
from scipy.stats import ks_2samp
class DriftDetector:
"""Production ML drift monitoring system"""
def __init__(self, alert_threshold_psi=0.25):
self.threshold = alert_threshold_psi
self.training_stats = {}
def fit(self, training_data):
"""Store training data statistics"""
for feature in training_data.columns:
self.training_stats[feature] = training_data[feature].values
def detect_drift(self, production_data):
"""Check for drift in production data"""
alerts = []
for feature, train_values in self.training_stats.items():
prod_values = production_data[feature].values
# Kolmogorov-Smirnov Test
statistic, p_value = ks_2samp(train_values, prod_values)
# Population Stability Index
psi = self.calculate_psi(train_values, prod_values)
if psi > self.threshold:
alerts.append({
'feature': feature,
'psi': psi,
'ks_statistic': statistic,
'p_value': p_value,
'status': 'severe' if psi > 0.25 else 'moderate'
})
return alerts
def calculate_psi(self, expected, actual, bins=10):
"""Calculate Population Stability Index"""
expected_percents = np.histogram(expected, bins=bins)[0] / len(expected)
actual_percents = np.histogram(actual, bins=bins)[0] / len(actual)
psi = np.sum((actual_percents - expected_percents) *
np.log((actual_percents + 1e-10) / (expected_percents + 1e-10)))
return psi
# Usage in production
detector = DriftDetector(alert_threshold_psi=0.25)
detector.fit(training_data)
# Run daily batch job
alerts = detector.detect_drift(production_data_last_week)
if alerts:
for alert in alerts:
send_slack_alert(
f"🚨 Drift detected in {alert['feature']}: PSI={alert['psi']:.3f}"
)
log_to_monitoring_db(alert)📚 How to Use This Dashboard:
- Deploy this monitoring system alongside your ML model in production
- Log all predictions and input features to a database (e.g., PostgreSQL, BigQuery)
- Run drift detection daily/weekly as batch job comparing recent data to training baseline
- Set up alerts (Slack, email, PagerDuty) when PSI exceeds thresholds
- Create response runbook: PSI > 0.25 → investigate + retrain, Accuracy < 80% → rollback
Technology Stack for Production Monitoring
Data Logging & Storage:
- • PostgreSQL / BigQuery: Store predictions, features, ground truth
- • Log Structure: timestamp, model_version, features (JSON), prediction, actual (when available)
- • Retention: Keep 6-12 months for drift analysis
Visualization & Alerting:
- • Streamlit / Gradio: Quick dashboard prototypes (like this one!)
- • Grafana / Datadog: Production-grade monitoring dashboards
- • Evidently AI: Open-source ML monitoring with drift detection built-in
- • Slack / PagerDuty: Alerting when thresholds exceeded