Why does He initialization use instead of Xavier's ?
mediumAnswer
- ReLU zeros out half of its inputs on average, halving the effective variance of activations.
- He init doubles the variance to compensate: .
- This keeps activation variance stable through ReLU layers and prevents signals from collapsing to zero.
- Default for any modern ReLU / Leaky-ReLU / GELU network.
Check yourself — multiple choice
- He init is for tanh activations
- ReLU halves activation variance; He compensates with
- He init sets all weights to 1
- He init requires no math
ReLU zeros half of activations ⇒ He doubles the init variance.
#initialization
Practise Deep Learning
214 interview questions in this topic.