Language Models2019intermediate11 min read

Sentence-BERT: Sentence Embeddings Using Siamese BERT-Networks

Sentence-BERT: تضمينات الجمل باستخدام شبكات BERT السيامية

Reimers, N. · Gurevych, I. — EMNLP-IJCNLP

The problem

BERT achieves state-of-the-art on sentence-pair tasks like — but it uses a architecture: both sentences must be fed together. This means comparing 10,000 sentences requires ~50 million forward passes (~65 hours on a V100 GPU). Naively extracting sentence embeddings from BERT (averaging outputs or using the [CLS] token) produces representations worse than simple GloVe averages when compared with . BERT was therefore unusable for semantic search, clustering, or any task requiring fast pairwise comparison at scale.

The contribution

Sentence-BERT (SBERT): a modification of BERT using siamese and triplet network structures to produce fixed-size sentence embeddings that are semantically meaningful and comparable via cosine similarity. Fine-tuned on NLI data (SNLI + MultiNLI) with a classification objective, SBERT reduces the search for the most similar pair among 10,000 sentences from 65 hours to ~5 seconds while maintaining BERT-level accuracy. On seven STS benchmarks, SBERT outperforms InferSent by 11.7 points and Universal Sentence Encoder by 5.5 points ().

The impact

SBERT made BERT embeddings practical for real-world applications: semantic search, duplicate detection, clustering, and retrieval-augmented generation. It spawned the sentence-transformers library — now one of the most popular embedding frameworks — and established the paradigm that powers modern systems like DPR, Contriever, and every major vector search engine. SBERT proved that task-specific of pre-trained encoders with siamese structures is the recipe for production-grade embeddings.

Imagine a courtroom where a judge must decide if two witnesses tell the same story. With BERT, the judge insists on hearing both witnesses together in every session — and with 10,000 witnesses, that's 50 million sessions. Sentence-BERT works differently: it gives each witness a passport photo — a compact summary of their testimony. Now the judge can simply compare photos, finishing in seconds what used to take days. The trick is the camera (a ) so that witnesses with similar stories get similar-looking photos.

The problem: BERT cannot compare sentences at scale

BERT achieves excellent results on sentence-pair tasks by acting as a cross-encoder: both sentences are concatenated with a [SEP] token, fed together through all layers, and a prediction is made from the joint representation. This works brilliantly for accuracy — but it is computationally catastrophic for comparison at scale.

Finding the most similar pair in a collection of n sentences requires n(n−1)/2 forward passes. For n = 10,000, that is ~50 million comparisons — roughly 65 hours on a V100 GPU. For the 40 million questions on Quora, a single similarity query would take over 50 hours.

The natural workaround — extracting a single vector per sentence from BERT — fails badly. Averaging BERT's output tokens gives embeddings that score only 54.81 on STS benchmarks (Spearman correlation), and using the [CLS] token scores a dismal 29.19 — both worse than simple GloVe averages (61.32). BERT's internal representations were never trained to be useful as standalone sentence vectors.

Open in Lab
Compare the cross-encoder (BERT) approach vs the bi-encoder (SBERT) approach. Watch how the number of comparisons explodes with scale.
The demo wakes as you arrive…

The solution: siamese BERT with pooling

The core idea of Sentence-BERT is elegant: pass each sentence independently through a BERT encoder that shares weights (a siamese network), add a layer on top to compress the variable-length output into a fixed-size vector, then fine-tune the whole system so that similar sentences land nearby in vector space.

Think of it as training twins. Two identical copies of BERT (sharing all parameters) each process one sentence. A pooling operation — by default, the mean of all output token vectors — squeezes each sentence into a single embedding. These embeddings are then compared using a task-appropriate objective function during training. At time, you simply encode each sentence once and compare with cosine similarity.

This architecture is directly inspired by FaceNet's siamese and triplet networks, which solved the same scaling problem for face recognition: instead of comparing face images in pairs, FaceNet learns a compact embedding per face, enabling instant lookup.

Open in Lab
Explore the siamese architecture: two BERT copies share weights, each produces a pooled embedding, and the objective function brings similar sentences closer.
The demo wakes as you arrive…

Pooling strategies: compressing BERT output into one vector

BERT outputs one vector per token. To get a single , SBERT adds a pooling layer on top. Three strategies were tested:

  • (default): average all output token vectors. This treats every token's representation equally, producing a balanced summary of the full sentence.
  • : take the element-wise maximum across all token vectors. This captures the most activated feature in each dimension, emphasizing the strongest signals.
  • CLS pooling: use the output of the [CLS] token directly, as BERT was originally designed for classification.

The showed that MEAN pooling works best with the classification objective (80.78 Spearman on STS) and is on par with CLS for regression. MAX pooling performed significantly worse for the regression objective (69.92 vs 87.44 for MEAN). The default choice is MEAN for its consistent performance across objectives.

Open in Lab
See how MEAN, MAX, and CLS pooling compress token vectors into a single sentence embedding.
The demo wakes as you arrive…

Three training objectives

The siamese architecture can be trained with different objective functions depending on the available labeled data:

1. Classification objective (NLI). For data with labels (entailment, contradiction, neutral), the two sentence embeddings u and v are combined as [u; v; |u−v|] — the concatenation of both embeddings plus their element-wise absolute difference — and fed to a softmax classifier. The element-wise difference |u−v| is the most important signal: it directly measures how the two embeddings diverge in each dimension.

2. Regression objective (STS). For sentence similarity scores (continuous labels), cosine similarity between u and v is computed and trained with mean squared error loss.

3. Triplet objective. Given an anchor sentence a, a positive p, and a negative n, the model is trained so that the distance between a and p is smaller than the distance between a and n by at least a margin ε. This follows the from FaceNet.

o=softmax ⁣(Wt⋅[u;  v;  ∣u−v∣])o = \text{softmax}\!\bigl(W_t \cdot [u;\; v;\; |u - v|]\bigr)
Classification objective — concatenation with element-wise difference — u and v are the sentence embeddings (pooled BERT outputs). The concatenation [u; v; |u−v|] is multiplied by a trainable weight matrix Wₜ of shape (3d × k), where d is the embedding dimension and k is the number of labels (3 for NLI). The element-wise difference |u−v| acts like a distance detector for each feature dimension.
Ltriplet=max⁡(∥a−p∥−∥a−n∥+ε,  0)\mathcal{L}_{\text{triplet}} = \max\bigl(\|a - p\| - \|a - n\| + \varepsilon,\; 0\bigr)
Triplet loss — anchor, positive, negative — The loss pushes the anchor a closer to the positive p and farther from the negative n. The margin ε (set to 1 in the paper) ensures a minimum gap. When the constraint is already satisfied, the loss is zero and no gradient flows. Euclidean distance is used as the metric.
Open in Lab
Click an objective function to see how it trains the siamese network.
The demo wakes as you arrive…

Training details

SBERT is fine-tuned on the combination of SNLI (570,000 sentence pairs) and MultiNLI (430,000 pairs), totaling about 1 million training examples with three labels: entailment, contradiction, and neutral. The classification objective with the softmax head is used.

Training takes less than 20 minutes — a dramatic contrast with methods like InferSent that train from scratch. This is because SBERT starts from pre-trained BERT weights and only needs one epoch of fine-tuning. Key hyperparameters: 16, Adam optimizer with learning rate 2×10⁻⁵, and linear over 10% of the training data.

The insight behind using NLI data for sentence embeddings is that the entailment/contradiction/neutral task forces the model to learn deep semantic comparison — sentences that entail each other must map nearby, contradictions must map far apart. This is exactly the geometric structure needed for good embeddings.

The idea in code

Sentence-BERT — siamese architecture with mean poolingpython

Simplified to show the idea — not the real implementation.

import torch
import torch.nn as nn

class SBERT(nn.Module):
    """Siamese BERT for sentence embeddings."""

    def __init__(self, bert_model, hidden_dim=768, num_labels=3):
        super().__init__()
        self.bert = bert_model       # shared encoder (the "twins")
        # Classification head: 3d → num_labels
        self.classifier = nn.Linear(hidden_dim * 3, num_labels)

    def mean_pool(self, token_embeds, attention_mask):
        """Average token embeddings, ignoring padding tokens."""
        mask = attention_mask.unsqueeze(-1).float()       # (B, T, 1)
        summed = (token_embeds * mask).sum(dim=1)         # (B, d)
        counts = mask.sum(dim=1).clamp(min=1e-9)          # (B, 1)
        return summed / counts                            # (B, d)

    def encode(self, input_ids, attention_mask):
        """Encode one sentence → fixed-size embedding."""
        output = self.bert(input_ids, attention_mask=attention_mask)
        return self.mean_pool(output.last_hidden_state, attention_mask)

    def forward(self, ids_a, mask_a, ids_b, mask_b):
        """Siamese forward: encode both sentences, classify."""
        u = self.encode(ids_a, mask_a)   # sentence A embedding
        v = self.encode(ids_b, mask_b)   # sentence B embedding

        # Combine: [u; v; |u - v|]  — the |u-v| is the key signal
        combined = torch.cat([u, v, (u - v).abs()], dim=-1)
        logits = self.classifier(combined)
        return logits

# At inference: encode once, compare with cosine similarity
# emb_a = model.encode(sentence_a)
# emb_b = model.encode(sentence_b)
# similarity = cosine_similarity(emb_a, emb_b)

Results: faster and better embeddings

SBERT-NLI-large achieved an average Spearman correlation of 76.55 across seven STS benchmarks — outperforming InferSent (65.01) by 11.5 points and Universal Sentence Encoder (71.22) by 5.3 points. On SentEval transfer tasks, SBERT improved 2+ points over both baselines, achieving the best results in 5 out of 7 tasks.

When trained additionally on the STS benchmark dataset (two-step training: NLI then STS), SBERT-NLI-STSb-large reached a Spearman correlation of 86.10 on the STSb test set. The cross-encoder BERT scored 88.77 — only 2.7 points higher, but at the cost of O(n²) comparisons.

On the Argument Facet Similarity corpus, SBERT performed nearly on par with the BERT cross-encoder in within-topic evaluation. On the Wikipedia Sections task, SBERT achieved 80.42% accuracy versus 74% for the previous best BiLSTM approach.

Open in Lab
Compare SBERT against baseline embedding methods across STS benchmarks. Click a method to highlight its performance.
The demo wakes as you arrive…

Computational efficiency: from 65 hours to 5 seconds

The efficiency gain is the paper's headline result. Finding the most similar pair in 10,000 sentences:

  • BERT cross-encoder: ~50 million pair comparisons → 65 hours
  • SBERT bi-encoder: 10,000 embeddings (~5 seconds) + cosine similarity matrix (~0.01 seconds) → ~5 seconds total

That is a 47,000× speedup. With optimized index structures like FAISS, searching 40 million Quora questions drops from 50 hours to milliseconds.

On GPU throughput, SBERT with smart batching (grouping sentences of similar length to reduce padding) processes 2,042 sentences per second — 9% faster than InferSent and 55% faster than Universal Sentence Encoder. On CPU, the transformer architecture is slower than InferSent's single BiLSTM layer, but the quality-speed tradeoff overwhelmingly favors SBERT.

Open in Lab
Drag the slider to change the number of sentences and watch the time comparison between BERT cross-encoder and SBERT.
The demo wakes as you arrive…

Ablation study highlights

The ablation study reveals which design choices matter most:

Concatenation strategy matters more than pooling. For the classification objective on NLI data, switching pooling from MEAN (80.78) to CLS (79.80) loses only 1 point. But removing |u−v| from the concatenation drops performance by over 10 points. The element-wise difference is the training signal that teaches embeddings to be geometrically meaningful.

MEAN pooling is the safe default. MEAN wins with the classification objective and performs competitively with regression. MAX pooling fails badly for regression (69.92 vs 87.44 MEAN) — likely because max-over-time collapses too much information for fine-grained similarity scoring.

Two-step training helps. Training first on NLI, then on STS data, improves performance by 1–2 points for SBERT and 3–4 points for the BERT cross-encoder. NLI pre-training provides a strong geometric prior that STS fine-tuning refines.

What SBERT unlocked

  1. 2019

    Sentence-BERT

    Siamese BERT with mean pooling and NLI training. 65 hours → 5 seconds for pairwise search. The sentence-transformers library is born.

  2. 2020

    DPR — Dense Passage Retrieval

    Applied the bi-encoder paradigm from SBERT to open-domain question answering. Two separate BERT encoders for questions and passages, trained with in-batch negatives.

  3. 2020

    Multilingual SBERT

    Extended SBERT to 50+ languages using knowledge distillation from English SBERT to multilingual student models.

  4. 2022

    Contriever

    Unsupervised contrastive pre-training for dense retrieval. Extended the bi-encoder idea to work without labeled data, using cropping-based positive pairs.

  5. 2022

    Sentence Transformers ecosystem

    The library now hosts 5,000+ models on Hugging Face, powers vector databases, RAG pipelines, and semantic search at every major tech company.

SBERT's deepest contribution is architectural, not just empirical. It established the bi-encoder as the standard architecture for production embedding systems. Before SBERT, the dominant approaches were either slow cross-encoders or weak averaging methods. SBERT proved that you can have both: BERT-quality understanding with instant comparison. This bi-encoder paradigm now powers DPR, Contriever, ColBERT, E5, and virtually every modern embedding model.

CitationReimers, Gurevych. Sentence-BERT: Sentence Embeddings using Siamese BERT-Networks. EMNLP-IJCNLP, 2019.

Terms in this paper