Attention & pooling

A sentiment classifier for Arabic

A customer-support team receives thousands of Arabic messages a day across several dialects. They want each one sorted into positive or negative before a human sees it.

Build a rack that turns a batch of tokenized sentences into two class scores. Attention must do the contextual work, and the model has to fit in a 30M-parameter budget.

The meter

What this architecture costs

[B, 128]

Parameters

0

FLOPs / example

0

Activations

0 B

per example

The rack

What you have built

  1. —

    Text input

    [B, 128]

    no weights

Export

This rack, as PyTorch

Every block knows its shapes and its weights, so it knows its own constructor. The result is a plain nn.Module with no dependency on Azimuth — paste it into your notebook.

The checker

What holds so far

40%
  • Shapes chain cleanly

    Every block accepts what the one before it produces.

  • Produces the required output

    Needs [B, 2] · rack currently ends at [B, 128]

  • Uses the blocks this idea needs

    Still missing: Token embedding, Multi-head attention, Dropout, Classifier head

  • Blocks are in a workable order

    Token embedding must come before Multi-head attention.

  • Within the size budget

    0 of 30M budget

Shelf

Blocks available

Sources

Embedding

Core layers

Reshaping

Normalization & regularization

Activations

Heads

Dials

No block selected

Pick a block on the rack to tune its dials and read what it does.