Recommender Systems2018intermediate10 min read

Deep Interest Network for Click-Through Rate Prediction

شبكة الاهتمام العميق للتنبؤ بمعدّل النقر

Zhou, G. · Song, C. · Zhu, X. · Fan, Y. · Zhu, H. · Ma, X. · Yan, Y. · Jin, J. · Li, H. · Gai, K. — KDD

The problem

Traditional deep CTR models (&) compress all of a user's historical behaviors into a single fixed-length — the same vector regardless of which ad is being scored. This bottleneck cannot capture diverse interests: a young mother who browses coats, earrings, and children's toys is represented identically whether the candidate ad is a handbag or a phone. Expanding the vector dimension causes on sparse industrial data with hundreds of millions of features.

The contribution

DIN introduces a local activation unit — an mechanism that soft-searches a user's behavior history with respect to each candidate ad. The result is an adaptive user that varies per ad, spotlighting relevant past behaviors. The paper also contributes two techniques: mini-batch aware that makes L2 feasible at billion- scale, and Dice — a data-adaptive that generalizes PReLU by shifting the rectification point to the input distribution's mean. Deployed in Alibaba's display advertising, DIN achieved 10% CTR lift in online A/B tests.

The impact

DIN pioneered the idea that user representations should be ad-aware, not static. This local activation paradigm spawned a family of models — DIEN (interest evolution), DSIN (session interest), and BERT4Rec (sequential recommendation) — and influenced industrial recommender systems at Alibaba, Google, and Meta. The concept of adaptive, query-dependent user modeling is now standard in large-scale CTR prediction.

Imagine you walk into a giant library carrying a slip of paper describing a handbag you want to buy. The old system hands you a pre-packed box of "everything you've ever read" — novels, cookbooks, travel guides — the same box regardless of your request.

DIN works like a smart librarian: she reads your slip, scans the shelves, and pulls only the fashion magazines and accessory catalogs. For a different slip — say, running shoes — she would pull the sports section instead. The box changes every time because which books matter depends on what you're looking for right now.

The problem: one box for all interests

In e-commerce advertising, every time a user visits a page the system must decide which ad to show. The decision hinges on predicting (CTR) — the probability that this user will click this particular ad.

By 2017 the dominant approach was Embedding&MLP: encode each user's behavioral history (browsed goods, clicked categories) as sparse multi-hot vectors, embed them into dense vectors, pool those embeddings into a single fixed-length vector via sum or average , concatenate it with ad features, and pass everything through a multi- perceptron.

The fatal flaw: the user vector is the same no matter which ad is being scored. A young mother who browsed coats, earrings, tote bags, and children's toys is compressed into one point in embedding space. When the candidate ad is a handbag, the system can't zoom into the bag-related part of her history — it has to use the whole compressed lump. Expanding the vector dimension is the obvious fix, but with goods_id features of dimension 600 million, it leads to catastrophic overfitting.

Open in Lab
Left: traditional pooling gives the same user vector for every ad. Right: DIN adapts the vector to spotlight behaviors relevant to each candidate ad.
The demo wakes as you arrive…

Feature representation: sparse to dense

CTR data is naturally multi-group categorical. An instance looks like [weekday=Friday, gender=Female, visited_cate_ids={Bag,Book}, ad_cate_id=Book]. Each group is encoded as a one-hot or multi-hot binary vector. For example, visited_cate_ids with two active categories becomes a multi-hot vector with two 1s.

These binary vectors are astronomically large — goods_id alone has ~600 million dimensions in Alibaba's system. The embedding layer maps each active ID to a learned dense vector of dimension DD (typically 12 in production). One-hot features yield a single embedding vector; multi-hot features yield a list of embedding vectors whose length varies per user.

To feed this variable-length list into a fixed-size MLP, a pooling layer — sum or average — collapses the list into one vector. This is the bottleneck DIN aims to break.

Open in Lab
Click on a user behavior to see how it maps from sparse one-hot/multi-hot to a dense embedding, then gets pooled.
The demo wakes as you arrive…

The DIN idea: let the ad choose which memories matter

DIN's insight is beautifully simple: not all past behaviors are equally relevant to every ad. When you see a handbag ad, your past browsing of tote bags and leather bags matters far more than the running shoes you looked at last week. DIN introduces a local activation unit — a small feed-forward network that scores each historical behavior against the candidate ad.

Think of it as a spotlight on a stage: the candidate ad is the director calling out which actors to illuminate. Bags light up for a handbag ad; shoes light up for a sneaker ad. The illuminated actors (high- behaviors) dominate the user's representation, while the dimmed ones (irrelevant behaviors) fade to the background.

Crucially, unlike standard attention in NMT, DIN does not normalize the weights to sum to 1. This preserves the intensity of interest: a user with 90% clothing history will naturally produce a higher total activation for a T-shirt ad than a phone ad. Normalizing would erase this difference.

vU(A)=f(vA,e1,e2,…,eH)=∑j=1Ha(ej,vA)⋅ej=∑j=1Hwj⋅ej\mathbf{v}_U(A) = f(\mathbf{v}_A, e_1, e_2, \ldots, e_H) = \sum_{j=1}^{H} a(e_j, \mathbf{v}_A) \cdot e_j = \sum_{j=1}^{H} w_j \cdot e_j
Local activation unit — the core of DIN — {e₁, …, eₕ} = behavior embeddings · vₐ = candidate ad embedding · a(·) = activation network that outputs weight wⱼ · the user vector vᵤ(A) changes per ad

The activation function a(ej,vA)a(e_j, \mathbf{v}_A) takes three inputs: the behavior embedding eje_j, the ad embedding vA\mathbf{v}_A, and their outer product ej⊗vAe_j \otimes \mathbf{v}_A. The outer product is explicit knowledge that helps the small network learn relevance patterns faster. The output is a scalar weight — how much this behavior matters for this ad.

Open in Lab
Choose a candidate ad and watch how the activation weights change across the user's behavior history. Higher bars = more relevant behaviors.
The demo wakes as you arrive…

DIN architecture — end to end

DIN keeps the same Embedding&MLP skeleton as the base model, but replaces the plain pooling layer for features with the local activation unit. Here is the full data flow:

1. Embedding layer — every ID (goods_id, shop_id, cate_id) is mapped to a dense vector of dimension DD.

2. Local activation unit — for each behavior in the user's history, the activation network a(⋅)a(\cdot) scores it against the candidate ad. The output weights are used in a weighted sum (not average) of behavior embeddings to produce vU(A)\mathbf{v}_U(A).

3. Concatenation — vU(A)\mathbf{v}_U(A) is concatenated with the ad embedding, user profile features, and context features.

4. MLP — fully connected layers (192 → 200 → 80 → 2 in production) learn nonlinear feature interactions.

5. — standard binary cross-entropy (negative log-likelihood) for click/no-click.

Open in Lab
Click on any layer to see its role in the DIN pipeline.
The demo wakes as you arrive…

Training at industrial scale: two key techniques

Training deep networks on Alibaba's data — 2 billion samples, 600 million goods_id features — brings two challenges that the paper solves with novel techniques.

wj←wj−η(1∣Bm∣∑(x,y)∈Bm∂L∂wj+λαmjnjwj)w_j \leftarrow w_j - \eta \left( \frac{1}{|B_m|} \sum_{(x,y) \in B_m} \frac{\partial L}{\partial w_j} + \lambda \frac{\alpha_{mj}}{n_j} w_j \right)
Mini-batch aware L2 update rule — αₘⱼ = 1 if feature j appears in batch Bₘ, else 0 · nⱼ = global frequency of feature j · only present features pay the regularization cost
p(s)=11+e−s−E[s]Var[s]+ϵp(s) = \frac{1}{1 + e^{-\frac{s - E[s]}{\sqrt{Var[s] + \epsilon}}}}
Dice control function — s = one dimension of the layer input · E[s], Var[s] = mini-batch mean and variance · the sigmoid centers on the data mean instead of 0
Open in Lab
Drag the mean slider to see how Dice shifts the activation curve while PReLU stays fixed at 0.
The demo wakes as you arrive…

The same idea in code

DIN local activation unit — completepython

Simplified to show the idea — not the real implementation.

import numpy as np

def activation_unit(behavior_emb, ad_emb):
    """Score one behavior against the candidate ad.
    behavior_emb: (D,)  ad_emb: (D,)  ->  scalar weight"""
    outer = behavior_emb * ad_emb           # element-wise (Hadamard) product
    concat = np.concatenate([behavior_emb, ad_emb, outer])  # (3D,)
    # Small MLP: 3D -> 36 -> 1
    h = np.maximum(0, concat @ W1 + b1)     # hidden layer
    return (h @ W2 + b2).item()              # scalar activation weight

def din_user_vector(behavior_list, ad_emb):
    """Weighted-sum pooling: each behavior weighted by relevance to ad."""
    weights = [activation_unit(b, ad_emb) for b in behavior_list]
    # NO softmax — preserve intensity of interest
    weighted = sum(w * b for w, b in zip(weights, behavior_list))
    return weighted  # shape: (D,) — changes per ad

Experimental results

DIN was evaluated on three datasets of increasing scale: Amazon Electronics (1.7M samples), MovieLens (20M samples), and Alibaba production data (2.14 billion samples). On all three, DIN outperformed BaseModel, Wide&Deep, PNN, and DeepFM.

The most striking result was on Amazon, where user behaviors are richest: DIN achieved 5.35% relative improvement (RelaImpr) over BaseModel, jumping to 6.82% with Dice. On the massive Alibaba dataset, DIN with both MBA regularization and Dice achieved 11.65% RelaImpr — a 0.0113 absolute AUC gain, which is commercially significant at Alibaba's scale.

In online A/B testing over nearly a month, DIN delivered a 10% CTR lift and 3.8% RPM (Revenue Per Mille) increase compared to the previous production model. DIN has since been deployed as Alibaba's main display advertising model.

Open in Lab
AUC comparison across models on all three datasets.
The demo wakes as you arrive…

Design insight: why DIN skips softmax

In standard attention (as in the ), weights are normalized via to sum to 1. DIN deliberately removes this constraint. Why?

Consider a user with 90% clothing and 10% electronics in their history. For a T-shirt ad, almost all behaviors are relevant — the activation weights sum to a large number, signaling high interest intensity. For a phone ad, only 10% of behaviors fire — the sum is small, signaling low interest intensity. If we applied softmax, both sums would equal 1, and this intensity signal would be lost. The downstream MLP would have to rediscover intensity from scratch.

This is a conscious trade-off: DIN gives up the probabilistic interpretation of weights in exchange for richer information about how strongly the user's history connects to the ad.

Why it mattered

  1. 2016

    Wide&Deep (Google)

    Combined hand-crafted feature crosses ("wide") with learned embeddings ("deep"). Became the industry standard, but user vectors remained static.

  2. 2017

    NCF — Neural Collaborative Filtering

    Replaced matrix factorization's dot product with a neural network to model user-item interactions. User vectors still fixed per user.

  3. 2018

    DIN — Deep Interest Network

    First to make the user vector ad-dependent via local activation. Deployed at Alibaba scale with MBA regularization and Dice activation.

  4. 2019

    DIEN — Deep Interest Evolution Network

    Extended DIN with GRU-based interest evolution — modeling how interests change over time, not just which ones are relevant now.

  5. 2019

    BERT4Rec

    Applied BERT's masked prediction to sequential recommendation. Bidirectional context let the model predict user interests from both past and future behaviors.

DIN's core idea — that the representation of a user should depend on what you're asking about — is a specific instance of a universal principle: context-dependent representation. The Transformer's attention makes word representations sentence-dependent. BERT makes them bidirectionally dependent. DIN showed this principle works just as powerfully in recommender systems.

CitationZhou, Song, Zhu, Fan, Zhu, Ma, Yan, Jin, Li, Gai. Deep Interest Network for Click-Through Rate Prediction. KDD, 2018.

Terms in this paper