—Text input
[B, 128]no weights
Pre-norm transformer blocks
Every modern language model is one block repeated dozens of times. Build that block once, correctly, and you understand the shape of all of them.
Assemble a decoder block that scores the whole vocabulary at every position. Attention must be causal — a generative model may never read the future.
The meter
[B, 128, 512]Parameters
16M
FLOPs / example
0
Activations
256 KB
per example
The rack
—[B, 128]no weights
[B, 128][B, 128, 512]16M params · 256 KB
Export
Every block knows its shapes and its weights, so it knows its own constructor. The result is a plain nn.Module with no dependency on Azimuth — paste it into your notebook.
The checker
Shapes chain cleanly
Every block accepts what the one before it produces.
Produces the required output
Needs [B, ·, ·] · rack currently ends at [B, 128, 512]
Uses the blocks this idea needs
Still missing: Positional encoding, Layer norm, Multi-head attention, Residual add, Language-model head
Blocks are in a workable order
Layer norm must come before Multi-head attention.
Depth is within range
2 blocks · allowed 8+
Attention is causal
No Multi-head attention on the rack yet.
Shelf
Sources
Embedding
Core layers
Reshaping
Normalization & regularization
Activations
Heads
Dials
Pick a block on the rack to tune its dials and read what it does.