Module 06 Summary
What this module established
A convolution is a question asked of every patch; a layer is a bank of such questions and depth is questions about answers. Early layers transfer almost always, late ones only between similar domains - and the distance determines how many layers you must retrain.
Carry forward
- Train the head frozen, then unfreeze at a much lower rate. Unfreezing early or at the original rate destroys what you were transferring.
- Preprocessing belongs to the checkpoint. Scale, channel order and normalisation constants must match exactly, and a mismatch degrades silently.
- Compute the majority-class baseline before training. Three architectures at 91, 92 and 94% look like progress until the constant predictor scores 94.
Before moving on
Move on when your preprocessing is asserted against the checkpoint and your model is reported against both baselines on the metric that matters.
