You are reading immutable version 1. The current guide may be newer.

Study the paper

Deep Residual Learning for Image Recognition

Lessons, visuals, quizzes, flashcards, and resources—organized in teaching order.

All activities

Empirical Analysis of Layer Responses

Standard deviations of layer responses on CIFAR-10

Standard deviations of layer responses on CIFAR-10

An empirical analysis of the standard deviations (std) of layer responses on CIFAR-10 was conducted to investigate the preconditioning hypothesis. The responses are measured at the outputs of each 3×33\times3 layer, after Batch Normalization (BN) and before nonlinearity.

Sources

S4.F7

Figure 7: Standard deviations (std) of layer responses on CIFAR-10. The responses are the outputs of each 3×\times3 layer, after BN and before nonlinearity. Top: the layers are shown in their original order. Bottom: the responses are ranked in descending order.

The results show that residual networks (ResNets) generally have smaller response standard deviations compared to their plain counterparts. This supports the hypothesis that residual learning acts as a good preconditioner, easing the optimization of extremely deep networks.

Sources

S4.F7

Figure 7: Standard deviations (std) of layer responses on CIFAR-10. The responses are the outputs of each 3×\times3 layer, after BN and before nonlinearity. Top: the layers are shown in their original order. Bottom: the responses are ranked in descending order.
Reported values
MeasureValue
Sources

S4.F7

Figure 7: Standard deviations (std) of layer responses on CIFAR-10. The responses are the outputs of each 3×\times3 layer, after BN and before nonlinearity. Top: the layers are shown in their original order. Bottom: the responses are ranked in descending order.