Study the paper
Deep Residual Learning for Image Recognition
Lessons, visuals, quizzes, flashcards, and resources—organized in teaching order.
Empirical Analysis of Layer Responses
Standard deviations of layer responses on CIFAR-10
Standard deviations of layer responses on CIFAR-10
An empirical analysis of the standard deviations (std) of layer responses on CIFAR-10 was conducted to investigate the preconditioning hypothesis. The responses are measured at the outputs of each layer, after Batch Normalization (BN) and before nonlinearity.
Sources
S4.F7
Figure 7: Standard deviations (std) of layer responses on CIFAR-10. The responses are the outputs of each 3×\times3 layer, after BN and before nonlinearity. Top: the layers are shown in their original order. Bottom: the responses are ranked in descending order.
The results show that residual networks (ResNets) generally have smaller response standard deviations compared to their plain counterparts. This supports the hypothesis that residual learning acts as a good preconditioner, easing the optimization of extremely deep networks.
Sources
S4.F7
Figure 7: Standard deviations (std) of layer responses on CIFAR-10. The responses are the outputs of each 3×\times3 layer, after BN and before nonlinearity. Top: the layers are shown in their original order. Bottom: the responses are ranked in descending order.
| Measure | Value |
|---|
Sources
S4.F7
Figure 7: Standard deviations (std) of layer responses on CIFAR-10. The responses are the outputs of each 3×\times3 layer, after BN and before nonlinearity. Top: the layers are shown in their original order. Bottom: the responses are ranked in descending order.