Study the paper
Deep Residual Learning for Image Recognition
Lessons, visuals, quizzes, flashcards, and resources—organized in teaching order.
Bottleneck Architectures and Layer Responses
Standard Deviations of Layer Responses on CIFAR-10
Standard Deviations of Layer Responses on CIFAR-10
An analysis of the standard deviations (std) of layer responses on CIFAR-10 reveals key insights into the behavior of residual learning. The responses are measured at the outputs of each layer, after Batch Normalization (BN) and before the nonlinearity.
Sources
S4.F7
Figure 7: Standard deviations (std) of layer responses on CIFAR-10. The responses are the outputs of each 3×\times3 layer, after BN and before nonlinearity. Top: the layers are shown in their original order. Bottom: the responses are ranked in descending order.
The empirical results show that residual networks generally have responses with smaller standard deviations compared to their plain counterparts. This supports the understanding that residual functions are closer to zero than non-residual functions, easing optimization.
Sources
S4.F7
Figure 7: Standard deviations (std) of layer responses on CIFAR-10. The responses are the outputs of each 3×\times3 layer, after BN and before nonlinearity. Top: the layers are shown in their original order. Bottom: the responses are ranked in descending order.
| Measure | Value |
|---|
Sources
S4.F7
Figure 7: Standard deviations (std) of layer responses on CIFAR-10. The responses are the outputs of each 3×\times3 layer, after BN and before nonlinearity. Top: the layers are shown in their original order. Bottom: the responses are ranked in descending order.