Deep Learning2015intermediate12 min read
Highway Networks
شبكات الطريق السريع
Srivastava, R. K. · Greff, K. · Schmidhuber, J. — ICML Deep Learning Workshop
The problem
Deeper neural networks can represent more complex functions, but by 2015 networks deeper than ~20 layers was extremely difficult. Gradients vanished or exploded across many layers, got stuck, and simply adding layers made performance worse instead of better — even with careful initialization schemes like Xavier and He init.
The contribution
Highway Networks: an architecture inspired by gating that adds a learned T and a C = 1−T to each . The output becomes y = H(x)·T(x) + x·(1−T(x)), allowing each layer to smoothly interpolate between transforming its input and passing it through unchanged. With a simple negative initialization for the gates, networks with hundreds of layers can be trained directly with SGD — no staged training, no teacher networks, no special initialization for H.
The impact
Highway Networks were the first architecture to demonstrate that networks with 100+ layers could be trained with plain SGD. Published months before ResNet, they introduced the principle of learned shortcut paths that inspired residual connections. ResNet simplified the gating to a fixed identity shortcut and became the dominant architecture, but Highway Networks proved the fundamental insight: give gradients an unobstructed path through the network, and depth becomes a strength rather than a barrier.
Imagine a multi-lane highway. On a regular road (a plain network), every car must exit at every toll booth, get processed, then re-enter — by the 50th booth, traffic has slowed to a crawl and most cars have been lost.
A highway network adds an express lane at each booth. A smart traffic gate looks at each car and decides: "This one needs processing — send it through the booth" or "This one should pass straight through." The gate's decision is learned, not fixed.
The result: even with 100 toll booths, traffic flows smoothly. Some cars zip through dozens of express lanes unchanged, while others get transformed exactly where they need to be. This is the magic of gated information flow.
The problem: depth should help, but it doesn't
By 2015, theory was clear: deeper networks can represent exponentially more complex functions than shallow ones. A 20-layer network can express function classes that would need an exponentially wider 2-layer network. ImageNet accuracy had climbed from 84% to 95% by going deeper.
But there was a wall. Beyond roughly 20 layers, networks stopped learning:
-
Vanishing gradients. Each layer multiplies the by its . After 50 such multiplications, the gradient reaching the early layers is astronomically small — those layers effectively stop updating. You saw this decay in the LSTM chapter.
-
Optimization difficulty. Even with -preserving initialization (Xavier, He init), deeper plain networks converge to worse solutions. The surface becomes riddled with saddle points and poor local minima.
-
Degradation paradox. A deeper network should be at least as good as a shallower one — it could always learn identity mappings for the extra layers. But in practice, adding layers made training error worse, not just test error. The problem was optimization, not capacity.
The insight: borrow gating from LSTM
The key insight came from an unlikely source: recurrent networks. LSTM had already solved the problem — but across time steps, not layers. An LSTM cell uses gates to decide what to remember, what to forget, and what to output. The critical mechanism is the : an information highway that runs through time, protected by gates that control what enters and what leaves.
Srivastava, Greff, and Schmidhuber asked: what if we apply the same principle to depth? Instead of gating information across time steps, gate it across layers. Give each layer the ability to say "I'll transform this input" or "I'll let it pass through unchanged" — or any smooth blend between the two.
This is a fundamentally different philosophy from simply making layers wider or adding more parameters. Instead of forcing every layer to transform its input, let the network learn which layers should transform and which should carry information through. The network designs its own depth.
The mechanism: transform, gate, carry
A plain feedforward layer takes an input x and produces output y = H(x), where H is typically a linear transform followed by a nonlinear activation. The input is always fully transformed — there is no option to skip.
A highway layer adds two more components:
-
A transform gate T(x) — a function that outputs a value between 0 and 1 for each , deciding how much of the transformed signal to let through.
-
A carry gate C(x) — set to 1 − T(x) for simplicity — deciding how much of the original input to carry forward unchanged.
The output becomes a blend of the transformation and the original input, controlled by the gate. Think of it as a mixing board: the gate is a slider for each dimension. Pushed all the way to "transform", you get a plain layer. Pushed all the way to "carry", the layer becomes an identity function. The network learns where to set each slider.
Why it works: unobstructed gradient highways
The magic is in the gradient. For a plain layer, the Jacobian is dy/dx = H'(x) — the derivative of the transformation. Stack 50 of these and the gradient is the product of 50 Jacobians. If each is slightly less than 1, the product vanishes exponentially.
For a highway layer, the Jacobian becomes a mixture:
When T → 0 (carry mode), dy/dx = I — the identity matrix. The gradient passes through untouched, exactly as if the layer wasn't there. This is the "highway" for gradients.
When T → 1 (transform mode), dy/dx = H'(x) — same as a plain layer. The gradient flows through the transformation.
In between, the gradient is a smooth blend. The critical insight is that even if some layers transform heavily, the carry paths in other layers ensure that gradients can always reach the early layers. One open highway lane is enough to keep the whole network trainable.
What the gates learn: selective transformation
After training, the gate activity reveals a striking pattern. In experiments on MNIST (50 layers) and CIFAR-100 (50 layers), the authors found:
-
Most layers operate in carry mode. The majority of transform gates stay close to 0, meaning the input passes through with minimal change. The network discovers that it doesn't need all 50 layers to transform.
-
Transformation concentrates in early layers. Most of the actual computation happens in the first ~10 layers for MNIST and ~30 for CIFAR-100. After that, the information highways take over and the stabilizes.
-
Gates are highly selective. For a single input, the gate activity across blocks is very sparse — only a few blocks activate per layer. Different inputs activate different blocks. The network has learned input-dependent routing.
-
The "stripe" pattern. Visualizing block outputs across layers shows horizontal stripes — values that remain constant through many consecutive layers. These are the "information highways" in action: specific features that the network has decided to preserve unchanged.
Highway Networks vs ResNet: learned vs fixed shortcuts
ResNet, published months after Highway Networks, took a simpler approach. Instead of a learned gate, ResNet uses a fixed identity shortcut:
y = H(x) + x
No gate, no sigmoid, no extra parameters. The layer always adds the transformation on top of the identity. This simplicity turned out to be remarkably effective — ResNet won ImageNet 2015 and became the backbone of modern vision.
The key differences:
-
Highway: y = H(x)·T(x) + x·(1−T(x)) — the gate can suppress the transformation or the identity. More flexible but more parameters.
-
ResNet: y = H(x) + x — always adds both. Simpler, fewer parameters, easier to scale.
Highway Networks are the general case: set T = 0.5 everywhere and you get something close to ResNet. But ResNet's constraint that every layer contributes additively turned out to be a powerful inductive bias. It forces every layer to learn a residual — the difference between the desired output and the input — which is often easier to learn than the full mapping.
Experiments: depth without degradation
The authors ran a clean comparison on MNIST: both plain and highway networks with identical architectures at depths of 10, 20, 50, and 100 layers. Each highway layer had 50 units; each plain layer had 71 units (to match count). Hyperparameters were optimized via random search for both architectures.
Results were dramatic:
-
At 10 layers, both architectures performed well.
-
At 20 layers, plain networks began to degrade. Highway networks converged faster.
-
At 50 layers, plain networks barely learned. Highway networks matched 10-layer performance.
-
At 100 layers, plain networks essentially failed. Highway networks achieved their best results — about one order of magnitude better training error than the 10-layer plain network.
On CIFAR-10, highway networks with 19 layers matched or exceeded Fitnets (which required a pre-trained teacher network for training), using only plain . A 900-layer highway network on CIFAR-100 showed no signs of optimization difficulty.
Building highway networks: practical details
A few practical details make highway networks work smoothly:
-
Dimensionality matching. The highway equation y = H(x)·T + x·(1−T) requires x and y to have the same size. When a dimension change is needed, use a plain (non-highway) layer first to resize, then continue with highway layers.
-
Convolutional highways. The same principle applies to CNNs: H and T share local receptive fields and use sharing. Zero-padding keeps feature maps the same size.
-
Gate initialization. Set b_T to −1 or −3. This biases the network toward carry at the start, allowing gradients to flow freely while the transformation weights are still random. Training naturally opens the gates where needed.
-
Any . Because carry mode provides gradient flow regardless of H's properties, highway networks work with , , or any other activation — no special initialization scheme for H is needed.
Simplified to show the idea — not the real implementation.
import torch import torch.nn as nn
class HighwayLayer(nn.Module):
"""A single highway layer with learned gating."""
def __init__(self, size, bias=-1.0):
super().__init__()
self.H = nn.Linear(size, size) # Transform
self.T = nn.Linear(size, size) # Transform gate
self.T.bias.data.fill_(bias) # Negative bias → start in carry mode
def forward(self, x):
h = torch.relu(self.H(x)) # Normal transformation
t = torch.sigmoid(self.T(x)) # Gate: 0=carry, 1=transform
return h * t + x * (1 - t) # Blend transform + carryLegacy: the road to residual learning
Highway Networks proved a principle that changed forever: if you give gradients an unobstructed path through the network, you can train arbitrary depth. Before this paper, 20 layers was deep. After it, 100+ layers became routine.
ResNet took the simpler path — a fixed identity addition instead of a learned gate — and became the standard architecture. But the intellectual lineage is clear: Highway Networks → ResNet → DenseNet → Transformers with residual streams. Every modern deep network carries this DNA.
The paper also introduced a powerful analytical lens: examining what gates learn tells you how the network uses its depth. The discovery that most layers operate in carry mode foreshadowed the "lottery ticket" hypothesis and research into network efficiency.
1997
LSTM — gating across time
Hochreiter & Schmidhuber introduce LSTM with forget, input, and output gates. The cell state provides an unobstructed path for gradients across time steps — the intellectual ancestor of highway connections.
2015
Highway Networks — gating across layers
Srivastava, Greff, and Schmidhuber apply LSTM-style gating to feedforward depth. Networks with 100+ layers train successfully with SGD for the first time. Published at ICML 2015 Deep Learning Workshop.
2015
ResNet — fixed identity shortcuts
He et al. simplify the highway idea to y = H(x) + x — no gate, just add. 152 layers, ImageNet champion. The simpler design scales better and becomes the dominant architecture.
2017
Recurrent Highway Networks
Zilly et al. extend highway connections to recurrent networks, bridging the gap between the original depth-gating idea and sequence modeling.
2017
Transformer — residual streams everywhere
Vaswani et al. build the Transformer with residual connections around every sub-layer. The residual stream is a descendant of both Highway and ResNet, enabling training of arbitrarily deep attention stacks.
CitationSrivastava, Greff, Schmidhuber. Highway Networks. ICML Deep Learning Workshop, 2015.
Terms in this paper
- Residual Connectionالوصلة التجاوزية
- LSTMشبكة الذاكرة الطويلة قصيرة المدى
- Gradientالتدرج التفاضلي
- Vanishing Gradientاضمحلال متجهات الميل
- Transform Gateبوابة التحويل
- Carry Gateبوابة الحمل
- Skip Connectionsالوصلات المتخطّية