—Text input
[B, 128]no weights
Attention & pooling
A customer-support team receives thousands of Arabic messages a day across several dialects. They want each one sorted into positive or negative before a human sees it.
Build a rack that turns a batch of tokenized sentences into two class scores. Attention must do the contextual work, and the model has to fit in a 30M-parameter budget.
The meter
[B, 128]Parameters
0
FLOPs / example
0
Activations
0 B
per example
The rack
—[B, 128]no weights
Export
Every block knows its shapes and its weights, so it knows its own constructor. The result is a plain nn.Module with no dependency on Azimuth — paste it into your notebook.
The checker
Shapes chain cleanly
Every block accepts what the one before it produces.
Produces the required output
Needs [B, 2] · rack currently ends at [B, 128]
Uses the blocks this idea needs
Still missing: Token embedding, Multi-head attention, Dropout, Classifier head
Blocks are in a workable order
Token embedding must come before Multi-head attention.
Within the size budget
0 of 30M budget
Shelf
Sources
Embedding
Core layers
Reshaping
Normalization & regularization
Activations
Heads
Dials
Pick a block on the rack to tune its dials and read what it does.