Gradient-Aware Multi-Task Learning for Predictive Maintenance
AMNL, GABA, and GRACE for RUL Prediction Under Multi-Condition Degradation
A research-grade walkthrough of three multi-task learning strategies for Remaining Useful Life prediction — AMNL (accuracy-first), GABA (safety-first), and GRACE (balanced) — all built on a shared CNN-BiLSTM-Attention backbone. Validated across 335 experiments on NASA C-MAPSS and N-CMAPSS DS02, beating the published SOTA (DKAMFormer) on multi-condition data.
24 chapters— in publication order.
Part I · Chapter 01 · 4 sections · 44 min
Why Remaining Useful Life prediction matters, the cost of being late, and the three deployment regimes that motivate this book.
- 1.1What is Predictive Maintenance?10m
- 1.2The RUL Prediction Problem12m
- 1.3Why Safety Matters: NASA Score vs. RMSE12m
- 1.4Three Deployment Regimes (Accuracy / Safety / Balanced)10m
Part I · Chapter 02 · 4 sections · 54 min
The two benchmarks that define progress in turbofan RUL prediction, and why multi-condition data is the real challenge.
- 2.1NASA C-MAPSS Overview12m
- 2.2FD001 to FD004: Single vs. Multi-Condition15m
- 2.3N-CMAPSS DS02: Realistic Flight Envelopes15m
- 2.4Why Multi-Condition Datasets Are Hard12m
Part I · Chapter 03 · 5 sections · 72 min
The minimal math you need: time series tensors, 1D convolution, recurrent networks, attention, and softmax cross-entropy.
- 3.1Time Series & Tensors12m
- 3.21D Convolution for Sensor Streams15m
- 3.3Recurrent Networks & LSTM Cells18m
- 3.4Self-Attention15m
- 3.5Softmax & Cross-Entropy12m
Part I · Chapter 04 · 4 sections · 54 min
Shared backbones, task-specific heads, and the loss-combination problem that the rest of the book is dedicated to solving.
- 4.1Why Multi-Task Learning?12m
- 4.2Shared vs. Task-Specific Parameters12m
- 4.3The Loss-Combination Problem15m
- 4.4A Gradient-Level View of MTL15m
Part II · Chapter 05 · 5 sections · 66 min
Sensor catalog, operating conditions, fault modes, and the file formats you will actually load into PyTorch.
- 5.1C-MAPSS File Structure12m
- 5.2The 21 Sensors and 3 Operational Settings15m
- 5.3Selecting 14 Informative Sensors12m
- 5.4Operating-Condition Discovery15m
- 5.5Fault Modes (HPC, Fan, and Combinations)12m
Part II · Chapter 06 · 4 sections · 49 min
The silent hero of the framework: why global Z-score fails on multi-condition data and how per-condition normalization fixes it.
- 6.1Why Global Z-Score Fails12m
- 6.2Discovering Operating Conditions (k-Means)12m
- 6.3Per-Condition Z-Score Implementation15m
- 6.4Preventing Test-Set Leakage10m
Part II · Chapter 07 · 4 sections · 55 min
Building the (B, 30, 17) input tensor: sliding windows, the piecewise-linear RUL cap, three-class health labels, and a reusable PyTorch Dataset.
- 7.1Sliding-Window Sequences (length = 30)15m
- 7.2Piecewise-Linear RUL Cap (R_max = 125)12m
- 7.3Health-State Discretization (3 Classes)10m
- 7.4PyTorch Dataset & DataLoader18m
Part III · Chapter 08 · 4 sections · 54 min
Three 1D conv layers (17 → 64 → 128 → 64) extract local degradation patterns from sensor streams.
- 8.11D Convolution for Sensor Series12m
- 8.2Three-Layer Conv Stack15m
- 8.3BatchNorm and Dropout for Stability12m
- 8.4PyTorch Implementation15m
Part III · Chapter 09 · 4 sections · 60 min
Two-layer BiLSTM (h=256) captures long-range temporal dependencies in degradation signatures.
- 9.1Why Bidirectional Beats Unidirectional12m
- 9.2LSTM Cell Mathematics18m
- 9.3Two-Layer BiLSTM Design (h = 256)15m
- 9.4PyTorch Implementation15m
Part III · Chapter 10 · 4 sections · 57 min
Eight-head self-attention with residual connection lets the model focus on degradation-relevant timesteps.
- 10.1Scaled Dot-Product Attention15m
- 10.2Multi-Head Attention with 8 Heads15m
- 10.3Residual Connection and LayerNorm12m
- 10.4PyTorch Implementation15m
Part III · Chapter 11 · 4 sections · 52 min
Two task-specific heads (RUL regression + 3-class health classification) on top of a shared 32-d feature, totaling ~3.5 M parameters.
- 11.1Shared 32-Dimensional Feature Vector10m
- 11.2RUL Regression Head12m
- 11.3Health Classification Head (3 Classes)12m
- 11.4Complete DualTaskModel and Parameter Count18m
Part IV · Chapter 12 · 4 sections · 60 min
The empirical discovery that motivates the rest of the book: regression gradients exceed classification gradients by 500x on shared parameters.
- 12.1Computing Per-Task Gradient Norms15m
- 12.2Why MSE Gradients Dominate Cross-Entropy18m
- 12.3Empirical Measurement (n = 4,120 samples)15m
- 12.4Consequences for Shared Feature Learning12m
Part IV · Chapter 13 · 4 sections · 54 min
Why low RMSE coincides with high NASA score, and why the tradeoff cannot be hidden behind a single metric.
- 13.1NASA Score: The Asymmetric Cost of Lateness15m
- 13.2Visualizing the RMSE-NASA Pareto Frontier15m
- 13.3Three Deployment Regimes Revisited12m
- 13.4Mapping Regimes to AMNL, GABA, and GRACE12m
Part V · Chapter 14 · 4 sections · 51 min
Up-weighting near-failure samples so the regressor pays attention where errors hurt the most.
- 14.1Why Equal-Weight MSE Underweights Failure12m
- 14.2The Linear-Decay Sample Weight w(RUL)15m
- 14.3Choosing w_max = 2.0 (Not 5.0 or 10.0)12m
- 14.4PyTorch Implementation12m
Part V · Chapter 15 · 5 sections · 74 min
Fixed 0.5/0.5 task weighting + failure-biased MSE, with the optimizer, scheduler, and EMA tricks that hold it together.
- 15.1The Fixed 0.5/0.5 Combined Loss12m
- 15.2Optimizer & Scheduler (AdamW + Warmup + Plateau)15m
- 15.3Gradient Clipping and Weight EMA12m
- 15.4Per-Dataset Dropout Tuning (Legacy-Pipeline Note)10m
- 15.5Full Training Script Walkthrough25m
Part V · Chapter 16 · 4 sections · 49 min
Best-in-literature RMSE on FD002/FD003, the FD001 NASA penalty, and the cross-pipeline caveat you must report.
- 16.1Best RMSE in the Literature (FD002, FD003)15m
- 16.2The FD001 NASA Penalty12m
- 16.3Cross-Pipeline Caveats12m
- 16.4When to Choose AMNL10m
Part VI · Chapter 17 · 4 sections · 54 min
Equalize each task's contribution to the shared backbone by giving lower weight to whichever task has bigger gradients.
- 17.1Equalizing Task Contributions12m
- 17.2Why Inverse-Proportional Weights Work15m
- 17.3The Two-Task Closed Form12m
- 17.4How GABA Differs from GradNorm15m
Part VI · Chapter 18 · 5 sections · 74 min
Per-step gradient norms, EMA smoothing (β = 0.99), minimum floor (λ_min = 0.05), and a 100-step warmup — the full pseudocode walked end to end.
- 18.1Per-Step Gradient Norm Computation15m
- 18.2Exponential Moving Average (β = 0.99)15m
- 18.3Minimum Floor and Renormalization12m
- 18.4Warmup (First 100 Steps)10m
- 18.5Full PyTorch Implementation22m
Part VI · Chapter 19 · 3 sections · 39 min
GABA viewed as a proportional feedback controller with an IIR filter and anti-windup floor — the property that gives it stronger stability guarantees than GradNorm.
- 19.1GABA as a Proportional Feedback Controller15m
- 19.2EMA as a First-Order IIR Low-Pass Filter12m
- 19.3Floor as Anti-Windup; Bounded-Weight Guarantee12m
Part VI · Chapter 20 · 4 sections · 52 min
GABA + standard MSE: best NASA among adaptive methods, no auxiliary loss, no learned parameters, and a single λ that converges within 10 epochs.
- 20.1GABA + Standard MSE Training Pipeline15m
- 20.2Watching the Weights Converge15m
- 20.3Best NASA Among Adaptive Methods12m
- 20.4When to Choose GABA10m
Part VII · Chapter 21 · 3 sections · 39 min
Adaptive weighting and loss-shape are orthogonal — GRACE composes them and resolves the accuracy-safety tradeoff.
- 21.1Separation of Concerns: Adaptation vs. Loss Shape12m
- 21.2The GRACE Loss Equation12m
- 21.3Why It Works (and When It Does Not — FD003)15m
Part VII · Chapter 22 · 4 sections · 64 min
The full reproducible pipeline: 5 seeds, AdamW, ReduceLROnPlateau, EMA, gradient clipping, and exact hyperparameters.
- 22.1Putting It All Together15m
- 22.2Hyperparameters and Defaults12m
- 22.3Reproducibility (Seeds, Environment, Determinism)12m
- 22.4Full Training Script Walkthrough25m
Part VII · Chapter 23 · 4 sections · 55 min
Best NASA on multi-condition C-MAPSS, the RMSE-NASA Pareto picture, and the only method to win on N-CMAPSS DS02.
- 23.1Best NASA on Multi-Condition C-MAPSS15m
- 23.2The RMSE-NASA Pareto Picture15m
- 23.3Best Overall on N-CMAPSS DS0215m
- 23.4When to Choose GRACE10m
Part VIII · Chapter 24 · 1 section · 10 min
Fixed weighting, Homoscedastic Uncertainty, GradNorm, and DWA — the published baselines we ran inside the same framework.
Start chapter 24Where the book lands in practice.
Failure-Biased Weighted MSE
Up-weighting near-failure samples so the regressor pays attention where errors hurt the most.
Open chapterAMNL Training Pipeline
Fixed 0.5/0.5 task weighting + failure-biased MSE, with the optimizer, scheduler, and EMA tricks that hold it together.
Open chapterAMNL Results & When to Use It
Best-in-literature RMSE on FD002/FD003, the FD001 NASA penalty, and the cross-pipeline caveat you must report.
Open chapterInverse-Gradient Balancing: The Idea
Equalize each task's contribution to the shared backbone by giving lower weight to whichever task has bigger gradients.
Open chapterThe GABA Algorithm
Per-step gradient norms, EMA smoothing (β = 0.99), minimum floor (λ_min = 0.05), and a 100-step warmup — the full pseudocode walked end to end.
Open chapterControl-Theoretic Interpretation
GABA viewed as a proportional feedback controller with an IIR filter and anti-windup floor — the property that gives it stronger stability guarantees than GradNorm.
Open chapterTraining GABA & Results
GABA + standard MSE: best NASA among adaptive methods, no auxiliary loss, no learned parameters, and a single λ that converges within 10 epochs.
Open chapterCombining GABA + Weighted MSE
Adaptive weighting and loss-shape are orthogonal — GRACE composes them and resolves the accuracy-safety tradeoff.
Open chapterGRACE Training Pipeline
The full reproducible pipeline: 5 seeds, AdamW, ReduceLROnPlateau, EMA, gradient clipping, and exact hyperparameters.
Open chapterGRACE Results & the Pareto Frontier
Best NASA on multi-condition C-MAPSS, the RMSE-NASA Pareto picture, and the only method to win on N-CMAPSS DS02.
Open chapter95 sections. Begin with one.
Chapter 1 — Predictive Maintenance & RUL — is where every reader starts.
In progress — 95 of 121 lessons published
26 more sections are still being written and are not part of this count.