Pre-norm transformer blocks

The generative decoder block

Every modern language model is one block repeated dozens of times. Build that block once, correctly, and you understand the shape of all of them.

Assemble a decoder block that scores the whole vocabulary at every position. Attention must be causal — a generative model may never read the future.

The meter

What this architecture costs

[B, 128, 512]

Parameters

16M

FLOPs / example

0

Activations

256 KB

per example

The rack

What you have built

  1. —

    Text input

    [B, 128]

    no weights

  2. [B, 128]

    Token embedding

    [B, 128, 512]

    16M params · 256 KB

Export

This rack, as PyTorch

Every block knows its shapes and its weights, so it knows its own constructor. The result is a plain nn.Module with no dependency on Azimuth — paste it into your notebook.

The checker

What holds so far

17%
  • Shapes chain cleanly

    Every block accepts what the one before it produces.

  • Produces the required output

    Needs [B, ·, ·] · rack currently ends at [B, 128, 512]

  • Uses the blocks this idea needs

    Still missing: Positional encoding, Layer norm, Multi-head attention, Residual add, Language-model head

  • Blocks are in a workable order

    Layer norm must come before Multi-head attention.

  • Depth is within range

    2 blocks · allowed 8+

  • Attention is causal

    No Multi-head attention on the rack yet.

Shelf

Blocks available

Sources

Embedding

Core layers

Reshaping

Normalization & regularization

Activations

Heads

Dials

No block selected

Pick a block on the rack to tune its dials and read what it does.