Vanishing gradients

The stack that will not learn

A colleague hands you this model. It compiles, the shapes all line up, and it trains for hours without the loss moving. Nothing about the diagram is wrong — and that is the point.

The symptomLoss sits at its initial value. Gradients measured at the first layer are ~1e-9; at the last layer they are normal.

Repair it. Keep the depth — the fix is not to make the network shallower — and give the gradient a path back to the first layer.

The meter

What this architecture costs

[B, 2]

Parameters

58K

FLOPs / example

58K

Activations

4 KB

per example

The rack

What you have built

  1. —

    Tabular input

    [B, 64]

    no weights · 256 B

  2. [B, 64]

    Dense (linear)

    [B, 128]

    8.3K params · 8.2K FLOPs · 512 B

  3. [B, 128]

    Activation

    [B, 128]

    no weights · 512 B

  4. [B, 128]

    Dense (linear)

    [B, 128]

    17K params · 16K FLOPs · 512 B

  5. [B, 128]

    Activation

    [B, 128]

    no weights · 512 B

  6. [B, 128]

    Dense (linear)

    [B, 128]

    17K params · 16K FLOPs · 512 B

  7. [B, 128]

    Activation

    [B, 128]

    no weights · 512 B

  8. [B, 128]

    Dense (linear)

    [B, 128]

    17K params · 16K FLOPs · 512 B

  9. [B, 128]

    Activation

    [B, 128]

    no weights · 512 B

  10. [B, 128]

    Classifier head

    [B, 2]

    258 params · 256 FLOPs · 8 B

Export

This rack, as PyTorch

Every block knows its shapes and its weights, so it knows its own constructor. The result is a plain nn.Module with no dependency on Azimuth — paste it into your notebook.

The checker

What holds so far

60%
  • Shapes chain cleanly

    Every block accepts what the one before it produces.

  • Produces the required output

    Needs [B, 2] · rack currently ends at [B, 2]

  • Uses the blocks this idea needs

    Still missing: Residual add, Layer norm

  • Depth is within range

    10 blocks · allowed 10+

  • No saturating activation left in the stack

    0 of 4 set to anything but Sigmoid

Shelf

Blocks available

Sources

Core layers

Reshaping

Normalization & regularization

Activations

Heads

Dials

No block selected

Pick a block on the rack to tune its dials and read what it does.