Language Models2019intermediate11 min read
Sentence-BERT: Sentence Embeddings Using Siamese BERT-Networks
Sentence-BERT: تضمينات الجمل باستخدام شبكات BERT السيامية
Reimers, N. · Gurevych, I. — EMNLP-IJCNLP
The problem
BERT achieves state-of-the-art on sentence-pair tasks like — but it uses a architecture: both sentences must be fed together. This means comparing 10,000 sentences requires ~50 million forward passes (~65 hours on a V100 GPU). Naively extracting sentence embeddings from BERT (averaging outputs or using the [CLS] token) produces representations worse than simple GloVe averages when compared with . BERT was therefore unusable for semantic search, clustering, or any task requiring fast pairwise comparison at scale.
The contribution
Sentence-BERT (SBERT): a modification of BERT using siamese and triplet network structures to produce fixed-size sentence embeddings that are semantically meaningful and comparable via cosine similarity. Fine-tuned on NLI data (SNLI + MultiNLI) with a classification objective, SBERT reduces the search for the most similar pair among 10,000 sentences from 65 hours to ~5 seconds while maintaining BERT-level accuracy. On seven STS benchmarks, SBERT outperforms InferSent by 11.7 points and Universal Sentence Encoder by 5.5 points ().
The impact
SBERT made BERT embeddings practical for real-world applications: semantic search, duplicate detection, clustering, and retrieval-augmented generation. It spawned the sentence-transformers library — now one of the most popular embedding frameworks — and established the paradigm that powers modern systems like DPR, Contriever, and every major vector search engine. SBERT proved that task-specific of pre-trained encoders with siamese structures is the recipe for production-grade embeddings.
Imagine a courtroom where a judge must decide if two witnesses tell the same story. With BERT, the judge insists on hearing both witnesses together in every session — and with 10,000 witnesses, that's 50 million sessions. Sentence-BERT works differently: it gives each witness a passport photo — a compact summary of their testimony. Now the judge can simply compare photos, finishing in seconds what used to take days. The trick is the camera (a ) so that witnesses with similar stories get similar-looking photos.
The problem: BERT cannot compare sentences at scale
BERT achieves excellent results on sentence-pair tasks by acting as a cross-encoder: both sentences are concatenated with a [SEP] token, fed together through all layers, and a prediction is made from the joint representation. This works brilliantly for accuracy — but it is computationally catastrophic for comparison at scale.
Finding the most similar pair in a collection of n sentences requires n(n−1)/2 forward passes. For n = 10,000, that is ~50 million comparisons — roughly 65 hours on a V100 GPU. For the 40 million questions on Quora, a single similarity query would take over 50 hours.
The natural workaround — extracting a single vector per sentence from BERT — fails badly. Averaging BERT's output tokens gives embeddings that score only 54.81 on STS benchmarks (Spearman correlation), and using the [CLS] token scores a dismal 29.19 — both worse than simple GloVe averages (61.32). BERT's internal representations were never trained to be useful as standalone sentence vectors.
The solution: siamese BERT with pooling
The core idea of Sentence-BERT is elegant: pass each sentence independently through a BERT encoder that shares weights (a siamese network), add a layer on top to compress the variable-length output into a fixed-size vector, then fine-tune the whole system so that similar sentences land nearby in vector space.
Think of it as training twins. Two identical copies of BERT (sharing all parameters) each process one sentence. A pooling operation — by default, the mean of all output token vectors — squeezes each sentence into a single embedding. These embeddings are then compared using a task-appropriate objective function during training. At time, you simply encode each sentence once and compare with cosine similarity.
This architecture is directly inspired by FaceNet's siamese and triplet networks, which solved the same scaling problem for face recognition: instead of comparing face images in pairs, FaceNet learns a compact embedding per face, enabling instant lookup.
Pooling strategies: compressing BERT output into one vector
BERT outputs one vector per token. To get a single , SBERT adds a pooling layer on top. Three strategies were tested:
- (default): average all output token vectors. This treats every token's representation equally, producing a balanced summary of the full sentence.
- : take the element-wise maximum across all token vectors. This captures the most activated feature in each dimension, emphasizing the strongest signals.
- CLS pooling: use the output of the
[CLS]token directly, as BERT was originally designed for classification.
The showed that MEAN pooling works best with the classification objective (80.78 Spearman on STS) and is on par with CLS for regression. MAX pooling performed significantly worse for the regression objective (69.92 vs 87.44 for MEAN). The default choice is MEAN for its consistent performance across objectives.
Three training objectives
The siamese architecture can be trained with different objective functions depending on the available labeled data:
1. Classification objective (NLI). For data with labels (entailment, contradiction, neutral), the two sentence embeddings u and v are combined as [u; v; |u−v|] — the concatenation of both embeddings plus their element-wise absolute difference — and fed to a softmax classifier. The element-wise difference |u−v| is the most important signal: it directly measures how the two embeddings diverge in each dimension.
2. Regression objective (STS). For sentence similarity scores (continuous labels), cosine similarity between u and v is computed and trained with mean squared error loss.
3. Triplet objective. Given an anchor sentence a, a positive p, and a negative n, the model is trained so that the distance between a and p is smaller than the distance between a and n by at least a margin ε. This follows the from FaceNet.
Training details
SBERT is fine-tuned on the combination of SNLI (570,000 sentence pairs) and MultiNLI (430,000 pairs), totaling about 1 million training examples with three labels: entailment, contradiction, and neutral. The classification objective with the softmax head is used.
Training takes less than 20 minutes — a dramatic contrast with methods like InferSent that train from scratch. This is because SBERT starts from pre-trained BERT weights and only needs one epoch of fine-tuning. Key hyperparameters: 16, Adam optimizer with learning rate 2×10⁻⁵, and linear over 10% of the training data.
The insight behind using NLI data for sentence embeddings is that the entailment/contradiction/neutral task forces the model to learn deep semantic comparison — sentences that entail each other must map nearby, contradictions must map far apart. This is exactly the geometric structure needed for good embeddings.
The idea in code
Simplified to show the idea — not the real implementation.
import torch
import torch.nn as nn
class SBERT(nn.Module):
"""Siamese BERT for sentence embeddings."""
def __init__(self, bert_model, hidden_dim=768, num_labels=3):
super().__init__()
self.bert = bert_model # shared encoder (the "twins")
# Classification head: 3d → num_labels
self.classifier = nn.Linear(hidden_dim * 3, num_labels)
def mean_pool(self, token_embeds, attention_mask):
"""Average token embeddings, ignoring padding tokens."""
mask = attention_mask.unsqueeze(-1).float() # (B, T, 1)
summed = (token_embeds * mask).sum(dim=1) # (B, d)
counts = mask.sum(dim=1).clamp(min=1e-9) # (B, 1)
return summed / counts # (B, d)
def encode(self, input_ids, attention_mask):
"""Encode one sentence → fixed-size embedding."""
output = self.bert(input_ids, attention_mask=attention_mask)
return self.mean_pool(output.last_hidden_state, attention_mask)
def forward(self, ids_a, mask_a, ids_b, mask_b):
"""Siamese forward: encode both sentences, classify."""
u = self.encode(ids_a, mask_a) # sentence A embedding
v = self.encode(ids_b, mask_b) # sentence B embedding
# Combine: [u; v; |u - v|] — the |u-v| is the key signal
combined = torch.cat([u, v, (u - v).abs()], dim=-1)
logits = self.classifier(combined)
return logits
# At inference: encode once, compare with cosine similarity
# emb_a = model.encode(sentence_a)
# emb_b = model.encode(sentence_b)
# similarity = cosine_similarity(emb_a, emb_b)Results: faster and better embeddings
SBERT-NLI-large achieved an average Spearman correlation of 76.55 across seven STS benchmarks — outperforming InferSent (65.01) by 11.5 points and Universal Sentence Encoder (71.22) by 5.3 points. On SentEval transfer tasks, SBERT improved 2+ points over both baselines, achieving the best results in 5 out of 7 tasks.
When trained additionally on the STS benchmark dataset (two-step training: NLI then STS), SBERT-NLI-STSb-large reached a Spearman correlation of 86.10 on the STSb test set. The cross-encoder BERT scored 88.77 — only 2.7 points higher, but at the cost of O(n²) comparisons.
On the Argument Facet Similarity corpus, SBERT performed nearly on par with the BERT cross-encoder in within-topic evaluation. On the Wikipedia Sections task, SBERT achieved 80.42% accuracy versus 74% for the previous best BiLSTM approach.
Computational efficiency: from 65 hours to 5 seconds
The efficiency gain is the paper's headline result. Finding the most similar pair in 10,000 sentences:
- BERT cross-encoder: ~50 million pair comparisons → 65 hours
- SBERT bi-encoder: 10,000 embeddings (~5 seconds) + cosine similarity matrix (~0.01 seconds) → ~5 seconds total
That is a 47,000× speedup. With optimized index structures like FAISS, searching 40 million Quora questions drops from 50 hours to milliseconds.
On GPU throughput, SBERT with smart batching (grouping sentences of similar length to reduce padding) processes 2,042 sentences per second — 9% faster than InferSent and 55% faster than Universal Sentence Encoder. On CPU, the transformer architecture is slower than InferSent's single BiLSTM layer, but the quality-speed tradeoff overwhelmingly favors SBERT.
Ablation study highlights
The ablation study reveals which design choices matter most:
Concatenation strategy matters more than pooling. For the classification objective on NLI data, switching pooling from MEAN (80.78) to CLS (79.80) loses only 1 point. But removing |u−v| from the concatenation drops performance by over 10 points. The element-wise difference is the training signal that teaches embeddings to be geometrically meaningful.
MEAN pooling is the safe default. MEAN wins with the classification objective and performs competitively with regression. MAX pooling fails badly for regression (69.92 vs 87.44 MEAN) — likely because max-over-time collapses too much information for fine-grained similarity scoring.
Two-step training helps. Training first on NLI, then on STS data, improves performance by 1–2 points for SBERT and 3–4 points for the BERT cross-encoder. NLI pre-training provides a strong geometric prior that STS fine-tuning refines.
What SBERT unlocked
2019
Sentence-BERT
Siamese BERT with mean pooling and NLI training. 65 hours → 5 seconds for pairwise search. The sentence-transformers library is born.
2020
DPR — Dense Passage Retrieval
Applied the bi-encoder paradigm from SBERT to open-domain question answering. Two separate BERT encoders for questions and passages, trained with in-batch negatives.
2020
Multilingual SBERT
Extended SBERT to 50+ languages using knowledge distillation from English SBERT to multilingual student models.
2022
Contriever
Unsupervised contrastive pre-training for dense retrieval. Extended the bi-encoder idea to work without labeled data, using cropping-based positive pairs.
2022
Sentence Transformers ecosystem
The library now hosts 5,000+ models on Hugging Face, powers vector databases, RAG pipelines, and semantic search at every major tech company.
SBERT's deepest contribution is architectural, not just empirical. It established the bi-encoder as the standard architecture for production embedding systems. Before SBERT, the dominant approaches were either slow cross-encoders or weak averaging methods. SBERT proved that you can have both: BERT-quality understanding with instant comparison. This bi-encoder paradigm now powers DPR, Contriever, ColBERT, E5, and virtually every modern embedding model.
CitationReimers, Gurevych. Sentence-BERT: Sentence Embeddings using Siamese BERT-Networks. EMNLP-IJCNLP, 2019.
Terms in this paper
- Siamese Networkالشبكة السيامية
- Triplet Lossخسارة الثلاثيّات
- Cosine Similarityتشابه جيب التمام
- Sentence Embeddingتضمين الجملة
- Bi-Encoderالمُرمِّز الثنائي
- Cross-Encoderالمُرمِّز التقاطعي
- Poolingالتجميع المكاني
- Natural Language Inferenceالاستدلال اللغوي الطبيعي
- Semantic Textual Similarityالتشابه الدلالي النصي
- Contrastive Learningالتعلم التبايُني
- Mean Poolingالتجميع بالمتوسط
- Fine-Tuningالضبط الدقيق
- Information Retrievalاسترجاع المعلومات
- Dense Retrievalالاسترجاع الكثيف
- Embedding Spaceفضاء التضمين