Recommender Systems2018intermediate10 min read
Deep Interest Network for Click-Through Rate Prediction
شبكة الاهتمام العميق للتنبؤ بمعدّل النقر
Zhou, G. · Song, C. · Zhu, X. · Fan, Y. · Zhu, H. · Ma, X. · Yan, Y. · Jin, J. · Li, H. · Gai, K. — KDD
The problem
Traditional deep CTR models (&) compress all of a user's historical behaviors into a single fixed-length — the same vector regardless of which ad is being scored. This bottleneck cannot capture diverse interests: a young mother who browses coats, earrings, and children's toys is represented identically whether the candidate ad is a handbag or a phone. Expanding the vector dimension causes on sparse industrial data with hundreds of millions of features.
The contribution
DIN introduces a local activation unit — an mechanism that soft-searches a user's behavior history with respect to each candidate ad. The result is an adaptive user that varies per ad, spotlighting relevant past behaviors. The paper also contributes two techniques: mini-batch aware that makes L2 feasible at billion- scale, and Dice — a data-adaptive that generalizes PReLU by shifting the rectification point to the input distribution's mean. Deployed in Alibaba's display advertising, DIN achieved 10% CTR lift in online A/B tests.
The impact
DIN pioneered the idea that user representations should be ad-aware, not static. This local activation paradigm spawned a family of models — DIEN (interest evolution), DSIN (session interest), and BERT4Rec (sequential recommendation) — and influenced industrial recommender systems at Alibaba, Google, and Meta. The concept of adaptive, query-dependent user modeling is now standard in large-scale CTR prediction.
Imagine you walk into a giant library carrying a slip of paper describing a handbag you want to buy. The old system hands you a pre-packed box of "everything you've ever read" — novels, cookbooks, travel guides — the same box regardless of your request.
DIN works like a smart librarian: she reads your slip, scans the shelves, and pulls only the fashion magazines and accessory catalogs. For a different slip — say, running shoes — she would pull the sports section instead. The box changes every time because which books matter depends on what you're looking for right now.
The problem: one box for all interests
In e-commerce advertising, every time a user visits a page the system must decide which ad to show. The decision hinges on predicting (CTR) — the probability that this user will click this particular ad.
By 2017 the dominant approach was Embedding&MLP: encode each user's behavioral history (browsed goods, clicked categories) as sparse multi-hot vectors, embed them into dense vectors, pool those embeddings into a single fixed-length vector via sum or average , concatenate it with ad features, and pass everything through a multi- perceptron.
The fatal flaw: the user vector is the same no matter which ad is being scored. A young mother who browsed coats, earrings, tote bags, and children's toys is compressed into one point in embedding space. When the candidate ad is a handbag, the system can't zoom into the bag-related part of her history — it has to use the whole compressed lump. Expanding the vector dimension is the obvious fix, but with goods_id features of dimension 600 million, it leads to catastrophic overfitting.
Feature representation: sparse to dense
CTR data is naturally multi-group categorical. An instance looks like [weekday=Friday, gender=Female, visited_cate_ids={Bag,Book}, ad_cate_id=Book]. Each group is encoded as a one-hot or multi-hot binary vector. For example, visited_cate_ids with two active categories becomes a multi-hot vector with two 1s.
These binary vectors are astronomically large — goods_id alone has ~600 million dimensions in Alibaba's system. The embedding layer maps each active ID to a learned dense vector of dimension (typically 12 in production). One-hot features yield a single embedding vector; multi-hot features yield a list of embedding vectors whose length varies per user.
To feed this variable-length list into a fixed-size MLP, a pooling layer — sum or average — collapses the list into one vector. This is the bottleneck DIN aims to break.
The DIN idea: let the ad choose which memories matter
DIN's insight is beautifully simple: not all past behaviors are equally relevant to every ad. When you see a handbag ad, your past browsing of tote bags and leather bags matters far more than the running shoes you looked at last week. DIN introduces a local activation unit — a small feed-forward network that scores each historical behavior against the candidate ad.
Think of it as a spotlight on a stage: the candidate ad is the director calling out which actors to illuminate. Bags light up for a handbag ad; shoes light up for a sneaker ad. The illuminated actors (high- behaviors) dominate the user's representation, while the dimmed ones (irrelevant behaviors) fade to the background.
Crucially, unlike standard attention in NMT, DIN does not normalize the weights to sum to 1. This preserves the intensity of interest: a user with 90% clothing history will naturally produce a higher total activation for a T-shirt ad than a phone ad. Normalizing would erase this difference.
The activation function takes three inputs: the behavior embedding , the ad embedding , and their outer product . The outer product is explicit knowledge that helps the small network learn relevance patterns faster. The output is a scalar weight — how much this behavior matters for this ad.
DIN architecture — end to end
DIN keeps the same Embedding&MLP skeleton as the base model, but replaces the plain pooling layer for features with the local activation unit. Here is the full data flow:
1. Embedding layer — every ID (goods_id, shop_id, cate_id) is mapped to a dense vector of dimension .
2. Local activation unit — for each behavior in the user's history, the activation network scores it against the candidate ad. The output weights are used in a weighted sum (not average) of behavior embeddings to produce .
3. Concatenation — is concatenated with the ad embedding, user profile features, and context features.
4. MLP — fully connected layers (192 → 200 → 80 → 2 in production) learn nonlinear feature interactions.
5. — standard binary cross-entropy (negative log-likelihood) for click/no-click.
Training at industrial scale: two key techniques
Training deep networks on Alibaba's data — 2 billion samples, 600 million goods_id features — brings two challenges that the paper solves with novel techniques.
The same idea in code
Simplified to show the idea — not the real implementation.
import numpy as np
def activation_unit(behavior_emb, ad_emb):
"""Score one behavior against the candidate ad.
behavior_emb: (D,) ad_emb: (D,) -> scalar weight"""
outer = behavior_emb * ad_emb # element-wise (Hadamard) product
concat = np.concatenate([behavior_emb, ad_emb, outer]) # (3D,)
# Small MLP: 3D -> 36 -> 1
h = np.maximum(0, concat @ W1 + b1) # hidden layer
return (h @ W2 + b2).item() # scalar activation weight
def din_user_vector(behavior_list, ad_emb):
"""Weighted-sum pooling: each behavior weighted by relevance to ad."""
weights = [activation_unit(b, ad_emb) for b in behavior_list]
# NO softmax — preserve intensity of interest
weighted = sum(w * b for w, b in zip(weights, behavior_list))
return weighted # shape: (D,) — changes per adExperimental results
DIN was evaluated on three datasets of increasing scale: Amazon Electronics (1.7M samples), MovieLens (20M samples), and Alibaba production data (2.14 billion samples). On all three, DIN outperformed BaseModel, Wide&Deep, PNN, and DeepFM.
The most striking result was on Amazon, where user behaviors are richest: DIN achieved 5.35% relative improvement (RelaImpr) over BaseModel, jumping to 6.82% with Dice. On the massive Alibaba dataset, DIN with both MBA regularization and Dice achieved 11.65% RelaImpr — a 0.0113 absolute AUC gain, which is commercially significant at Alibaba's scale.
In online A/B testing over nearly a month, DIN delivered a 10% CTR lift and 3.8% RPM (Revenue Per Mille) increase compared to the previous production model. DIN has since been deployed as Alibaba's main display advertising model.
Design insight: why DIN skips softmax
In standard attention (as in the ), weights are normalized via to sum to 1. DIN deliberately removes this constraint. Why?
Consider a user with 90% clothing and 10% electronics in their history. For a T-shirt ad, almost all behaviors are relevant — the activation weights sum to a large number, signaling high interest intensity. For a phone ad, only 10% of behaviors fire — the sum is small, signaling low interest intensity. If we applied softmax, both sums would equal 1, and this intensity signal would be lost. The downstream MLP would have to rediscover intensity from scratch.
This is a conscious trade-off: DIN gives up the probabilistic interpretation of weights in exchange for richer information about how strongly the user's history connects to the ad.
Why it mattered
2016
Wide&Deep (Google)
Combined hand-crafted feature crosses ("wide") with learned embeddings ("deep"). Became the industry standard, but user vectors remained static.
2017
NCF — Neural Collaborative Filtering
Replaced matrix factorization's dot product with a neural network to model user-item interactions. User vectors still fixed per user.
2018
DIN — Deep Interest Network
First to make the user vector ad-dependent via local activation. Deployed at Alibaba scale with MBA regularization and Dice activation.
2019
DIEN — Deep Interest Evolution Network
Extended DIN with GRU-based interest evolution — modeling how interests change over time, not just which ones are relevant now.
2019
BERT4Rec
Applied BERT's masked prediction to sequential recommendation. Bidirectional context let the model predict user interests from both past and future behaviors.
DIN's core idea — that the representation of a user should depend on what you're asking about — is a specific instance of a universal principle: context-dependent representation. The Transformer's attention makes word representations sentence-dependent. BERT makes them bidirectionally dependent. DIN showed this principle works just as powerfully in recommender systems.
CitationZhou, Song, Zhu, Fan, Zhu, Ma, Yan, Jin, Li, Gai. Deep Interest Network for Click-Through Rate Prediction. KDD, 2018.
Terms in this paper
- Attentionآلية الانتباه
- Click-Through Rateمعدّل النقر
- Embeddingالتضمين
- User Behaviorسلوك المستخدم
- Activation Functionدالة التنشيط
- Regularizationالضبط الهيكلي
- Recommender Systemنظام التوصية
- Softmaxسوفت ماكس
- Poolingالتجميع المكاني
- Multi-Layer Perceptron (MLP)البيرسبترون متعدد الطبقات
- Collaborative Filteringالتصفية التعاونية
- Featureميزة / سمة
- Overfittingفرط التخصيص