Recommender Systems2019intermediate11 min read
BERT4Rec: Sequential Recommendation with Bidirectional Encoder Representations from Transformer
BERT4Rec: التوصية التسلسلية بتمثيلات المرمِّز ثنائي الاتجاه من المحوِّل
Sun, F. · Liu, J. · Wu, J. · Pei, C. · Lin, X. · Ou, W. · Jiang, P. — CIKM
The problem
models before 2019 processed user interaction histories in a strict left-to-right order. Markov-chain methods like FPMC only looked at the last few items. RNN-based models like GRU4Rec processed the sequence step by step but could only see past items at each step. Even SASRec, which introduced to sequential recommendation, used a that blocked each position from seeing future items. All these approaches suffer from the same limitation: they build each item's representation using only leftward context, ignoring potentially valuable signals from future interactions in the sequence.
The contribution
BERT4Rec applies BERT's objective to sequential recommendation. Instead of predicting the next item from left context only, it randomly masks items in the user's interaction sequence and trains a deep bidirectional to predict the masked items using context from both directions. At time, it appends a special [mask] token at the end of the sequence to predict the next item. On four benchmark datasets (Beauty, Steam, MovieLens-1m, MovieLens-10m), BERT4Rec outperforms SASRec, GRU4Rec, and Caser with improvements of 7–12% in and .
The impact
BERT4Rec demonstrated that the masked pre- paradigm from NLP transfers powerfully to recommendation. It became one of the most cited sequential recommendation papers and inspired a wave of Transformer-based recommenders. The insight that user behavior is not strictly left-to-right — a purchase today can retrospectively explain a browse from yesterday — opened the door to bidirectional modeling across the recommendation community. Follow-up work includes S3-Rec, ALBERT4Rec, and various contrastive learning extensions.
Imagine you're a detective examining a suspect's shopping trail: shoes → jacket → ??? → sunglasses → watch. Previous recommender systems work like a detective who can only read the trail left to right — at the mystery item, all they know is "shoes, then jacket."
BERT4Rec is a detective who sees the full trail at once: shoes and jacket came before, sunglasses and watch came after. The blank could be a scarf (complements the jacket and matches the sunglasses), which the left-to-right detective would never guess because they can't see the sunglasses yet.
The problem: one-way reading loses context
Sequential recommendation tries to predict what a user will interact with next, given their chronological history of clicks, purchases, or ratings. The key insight is that order matters: watching Inception then Interstellar signals something different from watching Interstellar then Inception.
Before BERT4Rec, models attacked this problem in three waves. Markov-chain methods (FPMC) assumed the next item depends only on the last one or two — like predicting weather from today alone. RNN-based models (GRU4Rec) processed the sequence step by step, carrying a hidden state forward, but suffered from the problem on long sequences. Self- models (SASRec) let every item attend to all previous items at once — solving the long-range problem — but still used a causal mask that blocked each position from looking ahead.
The causal mask is borrowed from language modeling, where you predict the next word. It makes sense there: you cannot peek at future words during generation. But in recommendation, the entire user history is available at inference time. Blocking rightward context is an artificial constraint that throws away useful information.
The idea: mask items, predict from all directions
BERT4Rec borrows the task (also called masked language modeling) from BERT. In NLP, BERT masks random words and predicts them from surrounding context. BERT4Rec does the same with items: given a user's interaction sequence, randomly mask some items and train the model to predict the original items using the full bidirectional context.
Concretely, for a user sequence , the model might mask positions 2 and 4 to produce . The model must reconstruct and using all remaining items — including items that come after the masked positions. This is impossible with a causal mask.
Why does this matter for recommendation? Because user behavior is not strictly sequential. A user who browses a camera, then a tripod, then a memory card, then a camera bag — the camera bag purchase retroactively explains why they looked at the tripod. Bidirectional context captures these retrospective dependencies that left-to-right models miss.
The architecture: stacked Transformer layers for items
The BERT4Rec architecture is a stack of Transformer encoder layers — the same building block as BERT, without the causal mask. Each layer contains two sub-layers: multi-head self-attention followed by a position-wise feed-forward network. Both sub-layers use residual connections and .
The input representation is the sum of two embeddings. The item maps each item ID to a -dimensional vector. The encodes where each item sits in the sequence (position 1, 2, …, ). Unlike the original Transformer which used sinusoidal functions, BERT4Rec uses learnable position embeddings — the same choice BERT made.
The model processes a fixed-length sequence of the most recent items. If the user has fewer than interactions, the sequence is left-padded with a special padding token. If the user has more, only the last items are kept. This is a practical constraint that keeps computation bounded.
Training: the masked item prediction objective
Training BERT4Rec follows the Cloze procedure. For each training sequence, a proportion of items are randomly selected and replaced with the special [mask] token. The model then predicts the original item at each masked position.
The loss function is the negative log-likelihood (cross-entropy) of the correct item at each masked position. Only masked positions contribute to the loss — unmasked items provide context but receive no gradient signal. This is exactly how BERT's MLM works, transferred to the item domain.
One important design choice: the masking ratio . Too low and training is slow because each sequence produces few prediction signals. Too high and there isn't enough context left to make accurate predictions. The paper finds (masking 20% of items) works best for short sequences, while is better for longer ones — more context is available, so more items can be masked.
Inference: predicting the next item
At inference time, the goal shifts from filling in blanks to predicting the next item. BERT4Rec handles this elegantly: append a [mask] token at the end of the user's sequence and take the model's prediction at that position.
The model outputs a hidden vector at the mask position. This vector is projected onto the item embedding matrix to produce a score for every item in the catalog. The top- items by score become the recommendations.
This approach means the training objective (predict masked items anywhere) and the inference objective (predict the next item at the end) are slightly mismatched. However, the rich bidirectional representations learned during training more than compensate — and this is exactly what the experiments demonstrate.
The idea in code
Simplified to show the idea — not the real implementation.
import numpy as np
def mask_sequence(items, num_items, mask_id, mask_prob=0.2):
"""Apply Cloze-style masking to a user interaction sequence."""
masked = items.copy()
labels = np.full_like(items, -1) # -1 = don't predict here
for i in range(len(items)):
if items[i] == 0: # skip padding
continue
if np.random.random() < mask_prob:
labels[i] = items[i] # remember the real item
masked[i] = mask_id # replace with [mask]
return masked, labels
def bert4rec_loss(model, items, mask_id, num_items):
"""One training step: predict masked items from bidirectional context."""
masked_input, labels = mask_sequence(items, num_items, mask_id)
# Forward pass: FULL attention — no causal mask!
# Every item sees every other item (past AND future).
hidden = model.encode(masked_input) # (seq_len, d_model)
loss = 0
count = 0
for i in range(len(items)):
if labels[i] != -1:
# Project hidden state onto item embeddings → score per item
logits = hidden[i] @ model.item_embeddings.T
loss += cross_entropy(logits, labels[i])
count += 1
return loss / count
def predict_next(model, user_history, mask_id, top_k=10):
"""Inference: append [mask] and predict the next item."""
sequence = user_history + [mask_id] # append [mask] at the end
hidden = model.encode(sequence)
mask_hidden = hidden[-1] # representation at [mask]
scores = mask_hidden @ model.item_embeddings.T
return top_k_items(scores, top_k)Results: bidirectional beats unidirectional across the board
BERT4Rec was evaluated on four datasets spanning different domains and scales: Amazon Beauty (small, sparse), Steam (medium, gaming), MovieLens-1m, and MovieLens-10m (large, dense). The evaluation protocol uses leave-one-out: the last item in each user's sequence is held out for testing, the second-to-last for validation.
Key results across all datasets:
-
BERT4Rec outperforms SASRec (the strongest unidirectional baseline) by 7.2% on average in Hit Rate@10 and 11.7% in NDCG@10. This is a direct apples-to-apples comparison — same Transformer architecture, only the masking direction differs.
-
BERT4Rec beats GRU4Rec and its improved variant GRU4Rec+ by even larger margins, confirming that self-attention outperforms recurrence for sequential recommendation.
-
On the largest dataset (MovieLens-10m), BERT4Rec with 2 layers already beats SASRec with the best-tuned depth, showing that bidirectional context is more valuable than adding layers.
-
Deeper models (more Transformer layers) help, but the biggest single gain comes from switching unidirectional to bidirectional — the direction matters more than the depth.
Key design decisions
The paper carefully ablates several design choices:
Masking ratio : On short sequences (Beauty), is optimal. On long sequences (MovieLens-1m), works better because more context items remain even with more masks. This mirrors BERT's finding that 15% masking worked well for language — the exact ratio depends on sequence length and density.
Number of layers : Performance improves from 1 to 2 layers on all datasets, and 2 to 3 layers on larger datasets. Beyond that, gains plateau or even decline due to . The sweet spot is 2 layers for most settings.
Maximum sequence length : Longer histories help up to a point ( for small datasets, for MovieLens-10m). Beyond the optimal length, padding dominates and performance degrades.
Number of attention heads : The default of works well. More heads don't consistently help because each head gets fewer dimensions, and item relationships are simpler than linguistic structures.
What BERT4Rec unlocked
2016
GRU4Rec — RNNs for session-based recommendation
First successful application of recurrent neural networks to sequential recommendation. Processed item sequences step by step with a GRU, but suffered from vanishing gradients on long histories.
2018
SASRec — self-attention for sequential recommendation
Replaced RNNs with a Transformer-style self-attention mechanism. Each item attends to all previous items at once, solving the long-range dependency problem. But still uses a causal mask — left-to-right only.
2019
BERT4Rec — bidirectional masking for recommendation
Dropped the causal mask and adopted BERT's Cloze task. Every item sees every other item, enabling richer representations. 7–12% improvement over SASRec across four datasets.
2020
S3-Rec — self-supervised pre-training for recommendation
Extended BERT4Rec with additional self-supervised tasks (item attribute prediction, segment prediction) for richer pre-training signals in cold-start scenarios.
2021
Contrastive learning meets sequential recommendation
CL4SRec and CoSeRec combined BERT4Rec's masked prediction with contrastive objectives, learning more robust representations by contrasting augmented views of the same sequence.
BERT4Rec's lasting contribution is a simple but powerful transfer of ideas: what if we treat user interaction sequences the way BERT treats sentences? Mask some items, predict them from full context, and let the model learn which items explain which other items — regardless of their order in time. The result is a recommendation model that understands user intent more deeply because it reads the story from both ends, not just the beginning.
CitationSun, Liu, Wu, Pei, Lin, Ou, Jiang. BERT4Rec: Sequential Recommendation with Bidirectional Encoder Representations from Transformer. CIKM, 2019.
Terms in this paper
- Sequential Recommendationالتوصية التسلسلية
- Masked Language Modeling (MLM)نمذجة اللغة المُقنَّعة (MLM)
- Self-Attentionالانتباه الذاتي
- Clozeملء الفراغ
- Collaborative Filteringالتصفية التعاونية
- Implicit Feedbackالتغذية الراجعة الضمنية
- Multi-Head Attentionالانتباه المتعدد المسارات
- Positional Embeddingالتضمين الموضعي
- Autoregressive Modelالنموذج التوليدي التراجعي
- Causal Maskقناع سببي
- Dropoutالإسقاط العشوائي للعصبونات
- Embeddingالتضمين
- Softmaxسوفت ماكس
- Cross-Entropy Lossفقد الإنتروبيا التقاطعية
- Cold Startالبداية الباردة