Metric Learning1993foundational10 min read

Signature Verification Using a "Siamese" Time Delay Neural Network

التحقّق من التوقيعات بشبكة عصبية سيامية ذات تأخير زمني

Bromley, J. · Bentz, J. W. · Bottou, L. · Guyon, I. · LeCun, Y. · Moore, C. · Säckinger, E. · Shah, R. — NIPS

The problem

In the early 1990s, verifying handwritten signatures on pen-input tablets was done either by hand-crafted — counting strokes, measuring angles — or by memorizing entire signatures, which fails when people sign slightly differently each time. A credit card's magnetic strip can store only 80 bytes, so any system must compress a signature into a tiny, fixed-size while remaining robust to natural variation and resistant to forgery.

The contribution

The "Siamese" : two identical Time Delay Neural Networks (TDNNs) that share all weights. Each sub-network converts a raw signature (position, speed, acceleration over time) into a compact . The cosine of the angle between the two vectors measures similarity. During , genuine pairs are pushed toward cosine 1 and forgery pairs toward cosine −1. At verification time, only one sub-network runs — the output is a vector that can be compared to a stored template using just 38 bytes.

The impact

This paper introduced the term "" and the core idea of learning an space where distance encodes similarity — the foundation of modern . The architecture directly inspired FaceNet's triplet loss, Sentence-BERT's sentence embeddings, SimCLR's contrastive framework, and SimSiam's simplified self-supervised learning. Every modern face verification, image retrieval, and sentence similarity system traces back to this 1993 paper.

Imagine two identical twins working as security guards at a bank. Each twin examines one document — a stored signature and a new signature on a check. Because they're identical, they notice the same things: pen pressure, stroke rhythm, the little flourish at the end.

After examining their documents, each twin writes down a short summary on a card. A manager then holds the two cards side by side: if the summaries match, the check is approved. If they differ, it's flagged as a possible forgery.

The Siamese network is those twins — two identical neural networks that produce comparable summaries, trained so that genuine signatures produce similar summaries and forgeries produce different ones.

The problem: how do you teach a machine to compare?

Traditional classifiers learn to say "this is class A" or "this is class B." But signature verification isn't — it's comparison. The system must answer: "does this new signature match the one on file?" without having seen this particular person's signature during training.

This is the core challenge of verification (also called metric learning): learn a representation where similar things are close and different things are far apart, even for categories never seen during training. The same challenge appears in face verification ("is this the same person?"), speaker verification ("is this the same voice?"), and document matching ("are these about the same topic?").

The architecture: twin networks, shared weights

The Siamese network consists of two identical sub-networks. "Identical" means every , every , every connection is exactly the same — if you change a weight in one network, it changes in the other. This is called (or weight tying).

Why shared weights? Because comparison must be symmetric: if signature A matches signature B, then signature B must match signature A. Shared weights guarantee that both inputs are processed the same way — the same features are extracted regardless of which input slot a signature enters.

Each sub-network is a (TDNN) — a 1D convolutional network that slides learned filters across the temporal of the signature. The input is a time series: 200 time steps × 8 features (pen position, speed, acceleration, direction, curvature, and pen up/down). The TDNN extracts local temporal patterns (like loops or sudden stops) and gradually compresses them into a fixed-length feature vector.

Open in Lab
The two sub-networks share all weights. Each converts a raw signature into a feature vector. The cosine similarity between the vectors determines whether the signatures match.
The demo wakes as you arrive…

Think of each sub-network as a compression pipeline: raw temporal signal → local pattern detection → progressive summarization → compact feature vector. The network learns what matters — pen pressure patterns, speed profiles, curvature sequences — and throws away like exact position or slight size differences.

Preprocessing: making signatures comparable

Before reaching the network, signatures undergo several steps to remove irrelevant variations while preserving discriminative information:

  • Position normalization: subtract the linear trend from x(t) and y(t), making the representation invariant to where on the pad the person signed.
  • Size normalization: divide by the standard deviation of y, so a small signature and a large one by the same person look similar.
  • Time resampling: resample all signatures to exactly 200 time steps using linear interpolation, since the network requires fixed-length input. This preserves temporal spacing — a forger who writes too slowly or too quickly will be penalized.
  • Feature extraction: compute 8 features at each time step — pen up/down, x and y positions, speed, centripetal and tangential acceleration, direction cosine and sine. These carry both shape information (to prevent forgers from imitating rhythm alone) and dynamic information (to penalize forgers who copy only the shape).
Open in Lab
Toggle between raw and preprocessed views to see how normalization removes irrelevant variation.
The demo wakes as you arrive…

Measuring similarity: the cosine angle

After both sub-networks produce their feature vectors, the network computes the cosine of the angle between them. This is the final output of the Siamese network — a single number between −1 and +1.

Why cosine rather than ? Cosine measures direction, not magnitude. Two feature vectors can differ in length (one signature pressed harder overall) but still point in the same direction if they capture the same writing pattern. This makes the comparison robust to global scaling differences.

cos⁡(θ)=a⋅b∥a∥ ∥b∥=∑iaibi∑iai2  ∑ibi2\cos(\theta) = \frac{\mathbf{a} \cdot \mathbf{b}}{\|\mathbf{a}\| \, \|\mathbf{b}\|} = \frac{\sum_i a_i b_i}{\sqrt{\sum_i a_i^2}\;\sqrt{\sum_i b_i^2}}
Cosine similarity — the direction-based distance metric — a and b are the two feature vectors · cos(θ) = 1 means identical direction (genuine match) · cos(θ) = −1 means opposite directions (forgery) · the magnitude ‖a‖ cancels out, so only direction matters
Open in Lab
Drag the two vectors to see how cosine similarity changes with direction. Notice that length doesn't matter.
The demo wakes as you arrive…

Training: learning to measure similarity

Training a Siamese network requires pairs of inputs with a label: same (genuine pair) or different (forgery pair). The network sees both inputs simultaneously and adjusts its shared weights to make genuine pairs produce feature vectors with cosine ≈ 1, and forgery pairs produce vectors with cosine ≈ −1.

The training set included 982 genuine signatures from 108 signers and 402 forgeries. From these, up to 7,701 signature pairs were constructed: 50% genuine–genuine, 40% genuine– forgery, and 10% genuine–random (where "random forgeries" are simply other people's genuine signatures, simulating zero-effort attacks).

Training used a modified version of . The key constraint: both sub-networks must update identically, because they share the same weights. In practice, this means computing gradients from both branches and averaging them before updating.

Open in Lab
Watch how the network pushes genuine pairs together and pulls forgery pairs apart in embedding space.
The demo wakes as you arrive…

Verification: one network, one comparison

At verification time, the Siamese structure simplifies beautifully. Only one sub-network is needed. During enrollment, a person signs six times; each signature passes through the sub-network to produce a feature vector. These six vectors define a statistical of the person's signature — a multivariate normal density.

When a new signature arrives, the sub-network converts it to a feature vector. This vector is compared against the stored model: if the that it belongs to the genuine exceeds a threshold, the signature is accepted. Otherwise it's rejected.

The entire stored model fits in 38 bytes — one byte per dimension of the feature vector. This is well within the 80-byte limit of a credit card magnetic strip, leaving room for the model to be updated with each successful verification.

Open in Lab
Follow a signature through the full verification pipeline: preprocessing → feature extraction → comparison with stored template → accept/reject decision.
The demo wakes as you arrive…

The same idea in code

Siamese network with cosine similarity, minimal implementationpython

Simplified to show the idea — not the real implementation.

import numpy as np

def cosine_similarity(a, b):
    """Cosine of the angle between two vectors."""
    return np.dot(a, b) / (np.linalg.norm(a) * np.linalg.norm(b))

class SiameseNetwork:
    """Two identical sub-networks with shared weights."""

    def __init__(self, sub_network):
        self.network = sub_network  # one set of weights, used twice

    def forward(self, sig_a, sig_b):
        """Process both signatures with the SAME network."""
        feat_a = self.network(sig_a)  # feature vector for signature A
        feat_b = self.network(sig_b)  # feature vector for signature B
        return cosine_similarity(feat_a, feat_b)

    def verify(self, new_sig, stored_template):
        """At test time: run ONE sub-network, compare to stored model."""
        feat = self.network(new_sig)
        return cosine_similarity(feat, stored_template)

# Training: genuine pairs → push cosine toward +1
#           forgery pairs → push cosine toward -1
# Backpropagation updates ONE set of shared weights.
# At deployment, only the forward pass of ONE sub-network runs.

Results and practical performance

The best network (Architecture 1 with cleaned data) achieved 95.5% genuine acceptance while detecting 80% of forgeries. When first-time signatures were excluded (people needed a few tries to get accustomed to the tablet), performance rose to 97.0%.

The 38-dimensional feature vector could be stored in just 38 bytes (one byte per value) with no loss of performance — meeting the 80-byte credit card constraint with room to spare. The model could even be updated with each successful use, becoming more accurate over time.

Most errors came from inconsistent signers who omitted letters or added strokes, and from pen-up trajectories (pen movement above the pad) which proved both hard to imitate and hard to reproduce consistently.

Why it mattered: the birth of metric learning

Three core ideas from this paper became permanent fixtures of :

  • Weight sharing between twin networks — guaranteeing symmetric, consistent feature extraction. This principle reappears in every framework.
  • Learning embeddings, not labels — training a network to produce representations where distance is meaningful, rather than class probabilities. This is the basis of all modern embedding models.
  • as a trainable metric — using angular distance as a differentiable objective that backpropagation can optimize. Later work generalized this to Euclidean distance () and triplet loss.
  1. 1993

    Siamese Networks (this paper)

    Introduced the Siamese architecture with shared weights and cosine similarity for signature verification. First use of "learning to compare" in neural networks.

  2. 2005

    Contrastive Loss (Chopra, Hadsell, LeCun)

    Formalized the contrastive loss function with a margin parameter, extending Siamese networks to Euclidean distance and enabling dimensionality reduction by learning invariant mappings.

  3. 2015

    FaceNet — Triplet Loss

    Extended the pairwise idea to triplets (anchor, positive, negative), achieving human-level face verification. Direct descendant of the Siamese paradigm.

  4. 2019

    Sentence-BERT

    Applied Siamese BERT encoders to produce sentence embeddings optimized for semantic similarity — Siamese networks conquering language.

  5. 2020

    SimCLR — Contrastive Visual Learning

    Simplified contrastive learning for vision: augment an image twice, push the two views together, push different images apart. The Siamese idea at scale.

  6. 2021

    SimSiam — No Negatives Needed

    Showed that Siamese networks can learn without negative pairs at all, using a Stop Gradient trick. Simplified self-supervised learning to its essence.

From pen strokes on a tablet in 1993 to billion- models verifying faces, matching sentences, and learning visual representations — the core idea remains the same: pass two inputs through identical networks, measure the distance, and let backpropagation learn what "similar" means.

CitationBromley, Bentz, Bottou, Guyon, LeCun, Moore, Säckinger, Shah. Signature Verification Using a "Siamese" Time Delay Neural Network. NIPS, 1993.

Terms in this paper