Parameter-efficient fine-tuning

Adapting a large model to clinical notes

A hospital wants a triage model over its own clinical notes. The data cannot leave the building, and the building has one GPU. Fine-tuning every weight of a large pretrained model is not on the table.

Freeze the pretrained base and train adapters beside it. Sort notes into three urgency levels while keeping trainable weights under 400K — even though the model itself is far larger.

The meter

What this architecture costs

[B, 128, 768]

Parameters

25M

Trainable

0

25M frozen

FLOPs / example

0

Activations

384 KB

per example

The rack

What you have built

  1. —

    Text input

    [B, 128]

    no weights

  2. [B, 128]

    Token embedding

    [B, 128, 768]

    25M params (25M frozen) · 384 KB

Export

This rack, as PyTorch

Every block knows its shapes and its weights, so it knows its own constructor. The result is a plain nn.Module with no dependency on Azimuth — paste it into your notebook.

The checker

What holds so far

40%
  • Shapes chain cleanly

    Every block accepts what the one before it produces.

  • Produces the required output

    Needs [B, 3] · rack currently ends at [B, 128, 768]

  • Uses the blocks this idea needs

    Still missing: LoRA adapter, Multi-head attention, Classifier head

  • The pretrained attention stays frozen

    No Multi-head attention on the rack yet.

  • Within the trainable budget

    0 of 400K budget

Shelf

Blocks available

Sources

Embedding

Core layers

Reshaping

Normalization & regularization

Adapters

Heads

Dials

No block selected

Pick a block on the rack to tune its dials and read what it does.