—Image input
[B, 3, 64, 64]no weights · 48 KB
Skip connections
Your plain convolutional stack got worse when you made it deeper — not overfit, just worse on the training set too. The 2015 answer was to change the shape of the block itself.
Build a convolutional block whose output is added back to an earlier tensor, then classify into 10 classes. The residual only connects when the shapes across the span match exactly.
The meter
[B, 64, 64, 64]Parameters
1.8K
FLOPs / example
7.1M
Activations
1.0 MB
per example
The rack
—[B, 3, 64, 64]no weights · 48 KB
[B, 3, 64, 64][B, 64, 64, 64]1.8K params · 7.1M FLOPs · 1.0 MB
Export
Every block knows its shapes and its weights, so it knows its own constructor. The result is a plain nn.Module with no dependency on Azimuth — paste it into your notebook.
The checker
Shapes chain cleanly
Every block accepts what the one before it produces.
Produces the required output
Needs [B, 10] · rack currently ends at [B, 64, 64, 64]
Uses the blocks this idea needs
Still missing: Residual add, Classifier head
Blocks are in a workable order
Conv2D must come before Residual add.
Within the size budget
1.8K of 3.0M budget
Shelf
Sources
Core layers
Reshaping
Normalization & regularization
Activations
Heads
Dials
Pick a block on the rack to tune its dials and read what it does.