—Tabular input
[B, 64]no weights · 256 B
Vanishing gradients
A colleague hands you this model. It compiles, the shapes all line up, and it trains for hours without the loss moving. Nothing about the diagram is wrong — and that is the point.
The symptomLoss sits at its initial value. Gradients measured at the first layer are ~1e-9; at the last layer they are normal.
Repair it. Keep the depth — the fix is not to make the network shallower — and give the gradient a path back to the first layer.
The meter
[B, 2]Parameters
58K
FLOPs / example
58K
Activations
4 KB
per example
The rack
—[B, 64]no weights · 256 B
[B, 64][B, 128]8.3K params · 8.2K FLOPs · 512 B
[B, 128][B, 128]no weights · 512 B
[B, 128][B, 128]17K params · 16K FLOPs · 512 B
[B, 128][B, 128]no weights · 512 B
[B, 128][B, 128]17K params · 16K FLOPs · 512 B
[B, 128][B, 128]no weights · 512 B
[B, 128][B, 128]17K params · 16K FLOPs · 512 B
[B, 128][B, 128]no weights · 512 B
[B, 128][B, 2]258 params · 256 FLOPs · 8 B
Export
Every block knows its shapes and its weights, so it knows its own constructor. The result is a plain nn.Module with no dependency on Azimuth — paste it into your notebook.
The checker
Shapes chain cleanly
Every block accepts what the one before it produces.
Produces the required output
Needs [B, 2] · rack currently ends at [B, 2]
Uses the blocks this idea needs
Still missing: Residual add, Layer norm
Depth is within range
10 blocks · allowed 10+
No saturating activation left in the stack
0 of 4 set to anything but Sigmoid
Shelf
Sources
Core layers
Reshaping
Normalization & regularization
Activations
Heads
Dials
Pick a block on the rack to tune its dials and read what it does.