Recommender Systems2017intermediate11 min read

Neural Collaborative Filtering

التصفية التعاونية العصبية

He, X. · Liao, L. · Zhang, H. · Nie, L. · Hu, X. · Chua, T.-S. — WWW

The problem

(MF) dominated for a decade after the Netflix Prize, but its — a simple of user and item latent vectors — is fundamentally linear. It cannot model the complex, nonlinear patterns in how users interact with items, especially with (clicks, views, purchases) where the signal is noisy and there is no explicit negative feedback.

The contribution

NCF: a general neural framework that replaces MF's inner product with a learnable neural architecture. Three instantiations — GMF (Generalized Matrix Factorization), MLP (), and NeuMF (Neural Matrix Factorization) that fuses both. GMF recovers classical MF as a special case. MLP stacks dense layers on concatenated embeddings to learn nonlinear interactions. NeuMF combines both paths, trained with and , outperforming BPR and eALS on MovieLens and Pinterest.

The impact

NCF established that neural networks can and should replace handcrafted interaction functions in recommender systems. It opened the door for deep recommendation models — from DeepFM and Wide & Deep to BERT4Rec and DIN — and shifted the field from linear latent factor models to end-to-end learned architectures. Today, virtually every production recommender uses neural interaction functions descended from this insight.

Imagine you run a restaurant and want to predict which dishes a customer will order. The old approach (matrix factorization) summarizes each customer as a short list of taste numbers — spicy: 0.8, sweet: 0.3 — and each dish with matching numbers. To predict whether someone will like a dish, you simply multiply and sum these numbers. It works, but it misses nuance: "likes spicy food except desserts" can't be captured by a simple multiplication.

NCF replaces this multiplication with a smart tasting panel: a that takes the customer's profile and the dish's profile, passes them through layers of learned pattern recognition, and outputs a nuanced compatibility score. The panel can learn any complex taste rule — even ones that would be impossible for multiplication to express.

The challenge: learning from silence

Most recommendation research before 2017 focused on explicit feedback — star ratings, thumbs up/down. But in practice, implicit feedback dominates: clicks, views, purchases, time spent. The problem is that implicit feedback only tells you what a user did interact with. A zero doesn't mean dislike — it could mean the user never saw the item. There are no real negatives, only observed positives and a sea of unknowns.

The standard approach was matrix factorization: represent each user as a vector pu\mathbf{p}_u and each item as a vector qi\mathbf{q}_i, then predict interaction as their inner product y^ui=puTqi\hat{y}_{ui} = \mathbf{p}_u^T \mathbf{q}_i. This was popularized by the Netflix Prize and refined by BPR (Bayesian Personalized Ranking) for implicit feedback.

Open in Lab
Explicit feedback has clear signals (ratings). Implicit feedback only records interactions — zeros are ambiguous.
The demo wakes as you arrive…

Why the inner product hits a ceiling

Matrix factorization's inner product is a linear combination of latent dimensions — each dimension is independent and contributes equally. This limits its expressiveness. Consider the paper's famous example: four users with interaction patterns that cannot be faithfully represented in a low-dimensional space using inner products alone.

User u4u_4 is most similar to u1u_1, then u3u_3, then u2u_2. But once you have placed u1u_1, u2u_2, and u3u_3 in the , there is no position for u4u_4 that preserves all three rankings simultaneously. The inner product creates a geometry too rigid for real interaction patterns — it's like trying to draw a map of friendships on a flat piece of paper when the real relationships are tangled in many dimensions.

Open in Lab
Try placing user 4 in the latent space. No matter where you put it, at least one similarity ranking is violated — the inner product is too rigid.
The demo wakes as you arrive…

The NCF framework: learn the interaction function

The key insight of NCF is deceptively simple: replace the fixed inner product with a neural network that learns the interaction function from data. Instead of assuming that user-item compatibility is a , let the network discover any function — linear or nonlinear — that best predicts interactions.

The framework works in four stages. First, the user ID and item ID are converted to one-hot vectors. Second, an layer projects these sparse vectors into dense latent vectors. Third, the embeddings are fed into neural collaborative filtering layers — this is where the magic happens and where different model variants differ. Fourth, the output layer produces a predicted score between 0 and 1.

y^ui=f(pu,qi∣Θ)\hat{y}_{ui} = f(\mathbf{p}_u, \mathbf{q}_i \mid \Theta)
The NCF prediction — a learned function replaces the inner product — Instead of the fixed p·q, f can be any neural architecture parameterized by Θ. This is the paper's central idea: let the data determine the interaction function.
Open in Lab
Click each layer to see its role. The embedding layer converts IDs to vectors; the neural CF layers learn the interaction; the output predicts interaction probability.
The demo wakes as you arrive…

Three instantiations: GMF, MLP, and NeuMF

NCF is a framework, not a single model. The paper proposes three instantiations that differ in how they combine the user and item embeddings:

GMF (Generalized Matrix Factorization) takes the of the user and item embedding vectors, then passes the result through a weighted output layer with a activation. When the output weights are all ones and there is no activation, GMF reduces exactly to standard MF. But the learnable weights allow each latent dimension to contribute differently — some dimensions matter more than others.

MLP (Multi-Layer Perceptron) concatenates the user and item embeddings into a single vector and passes it through multiple dense layers with activations. Each layer progressively distills higher-order interaction patterns. The tower architecture (e.g., 64→32→16→8) gradually narrows, compressing complex interactions into a compact prediction signal.

NeuMF (Neural Matrix Factorization) is the fusion: run GMF and MLP in parallel with separate embedding sets, concatenate their last hidden layers, and feed the combined vector through a final prediction layer. This lets the model capture both the linearity of MF and the nonlinearity of deep networks simultaneously.

Open in Lab
Toggle between GMF, MLP, and NeuMF to see how each model processes user and item embeddings differently.
The demo wakes as you arrive…

GMF: matrix factorization goes neural

The first NCF instantiation recovers classical MF and then generalizes it. The idea is elegant: compute the element-wise product of user and item embeddings (which is what MF does inside the inner product), but instead of summing all dimensions equally, pass the product through a learnable output layer.

y^ui=σ ⁣(hT(pu⊙qi))\hat{y}_{ui} = \sigma\!\left(\mathbf{h}^T (\mathbf{p}_u \odot \mathbf{q}_i)\right)
GMF — element-wise product with learned weights — ⊙ is element-wise multiplication. h is a learned weight vector. σ is sigmoid. When h = 1 (all ones) and σ is identity, this reduces to standard MF.

MLP: learning nonlinear interactions

The MLP path takes a fundamentally different approach. Instead of multiplying embeddings element-wise, it concatenates them into a single vector and feeds it through a tower of dense layers. Each layer applies a linear transformation followed by a ReLU activation. Think of it as a conversation between the user representation and the item representation — each layer allows them to interact in increasingly abstract ways.

The tower architecture shrinks progressively. If the predictive factors are 8, the MLP layers might be 32→16→8 with an embedding size of 16 on each side. This bottleneck forces the network to distill the most predictive interaction patterns.

y^ui=σ ⁣(hT ϕL(…ϕ2(ϕ1([pu; qi]))))\hat{y}_{ui} = \sigma\!\left(\mathbf{h}^T \, \phi_L(\ldots \phi_2(\phi_1([\mathbf{p}_u;\,\mathbf{q}_i])))\right)
MLP prediction — stacked dense layers on concatenated embeddings — [p;q] is concatenation. Each φ is a dense layer with ReLU. L layers of progressively abstracted features compress to a final prediction through h and sigmoid.

NeuMF: the best of both worlds

NeuMF is the paper's flagship model. It runs GMF and MLP in parallel, each with its own set of embeddings. This is crucial — sharing embeddings would force both paths to use the same representation, limiting the model. Separate embeddings let GMF learn the linear patterns it's good at while MLP focuses on nonlinear patterns.

The last hidden layers of both paths are concatenated and fed into a single output neuron with sigmoid activation. A α controls the trade-off between the two paths during initialization.

y^ui=σ ⁣(hT[puG⊙qiG⏟GMF;  ϕL(…(puM,qiM))⏟MLP])\hat{y}_{ui} = \sigma\!\left(\mathbf{h}^T \left[\underbrace{\mathbf{p}_u^G \odot \mathbf{q}_i^G}_{\text{GMF}} ;\; \underbrace{\phi_L(\ldots(\mathbf{p}_u^M, \mathbf{q}_i^M))}_{\text{MLP}}\right]\right)
NeuMF — fusing GMF and MLP — Two parallel paths with separate embeddings (G for GMF, M for MLP). Their outputs are concatenated and projected to a final score. The model combines linear and nonlinear interaction modeling.
Open in Lab
Watch how a user-item pair flows through both the GMF and MLP paths before merging into a final prediction.
The demo wakes as you arrive…

Training: binary cross-entropy and negative sampling

A key design choice in NCF is treating the recommendation task as binary classification rather than regression. For each observed interaction (u,i)(u, i), the label is 1. For unobserved pairs, items are randomly sampled as negative examples with label 0. The number of negative samples per positive instance is a tunable ratio — the paper found that 3–6 negatives per positive works well.

The is binary cross-entropy (log loss), which naturally fits the probabilistic interpretation: the model outputs the probability that user uu will interact with item ii. This is more principled than squared error for binary targets, because it penalizes confident wrong predictions much more heavily.

L=−∑(u,i)∈Y+log⁡y^ui−∑(u,j)∈Y−log⁡(1−y^uj)L = -\sum_{(u,i) \in \mathcal{Y}^+} \log \hat{y}_{ui} - \sum_{(u,j) \in \mathcal{Y}^-} \log(1 - \hat{y}_{uj})
Binary cross-entropy with negative sampling — Y⁺ = observed interactions (positives), Y⁻ = sampled unobserved pairs (negatives). The first term pushes predictions toward 1 for positives; the second toward 0 for negatives.
Open in Lab
Click on cells to observe how negative sampling creates training pairs from a sparse interaction matrix.
The demo wakes as you arrive…

Pre-training: bootstrapping NeuMF

Training NeuMF from random initialization can be tricky because the model is non-convex and has many local minima. The paper's solution: pre-train GMF and MLP separately using Adam optimizer, then use their learned parameters to initialize NeuMF. Only the output weights of the fused model are initialized from scratch — specifically as a weighted combination controlled by α (set to 0.5 for equal contribution).

After initialization, NeuMF is fine-tuned with vanilla rather than Adam. The reasoning: Adam maintains momentum information that would be mismatched with the pre-trained parameters from a different optimization trajectory. This two-stage training — pre-train components, then fine-tune the whole — became a common pattern in later recommendation models.

The same idea in code

NeuMF — the complete forward passpython

Simplified to show the idea — not the real implementation.

import numpy as np

def sigmoid(x):
    return 1 / (1 + np.exp(-np.clip(x, -500, 500)))

def relu(x):
    return np.maximum(0, x)

def neumf_predict(user_id, item_id, params):
    """Forward pass of NeuMF: GMF + MLP in parallel."""
    # --- GMF path ---
    p_gmf = params['user_emb_gmf'][user_id]   # user embedding for GMF
    q_gmf = params['item_emb_gmf'][item_id]   # item embedding for GMF
    gmf_out = p_gmf * q_gmf                    # element-wise product

    # --- MLP path ---
    p_mlp = params['user_emb_mlp'][user_id]
    q_mlp = params['item_emb_mlp'][item_id]
    x = np.concatenate([p_mlp, q_mlp])         # concatenate embeddings
    for W, b in params['mlp_layers']:
        x = relu(W @ x + b)                    # dense + ReLU

    # --- Fusion ---
    concat = np.concatenate([gmf_out, x])      # merge both paths
    score = sigmoid(params['h'] @ concat)       # final prediction
    return score  # probability of interaction

What the experiments showed

The paper evaluated on two datasets: MovieLens 1M (1 million ratings from 6,000 users on 4,000 movies, binarized to implicit feedback) and Pinterest (1.5 million interactions from 55,000 users on 9,900 images, naturally implicit). The evaluation metric was Hit Rate and NDCG at top-K, using leave-one-out evaluation.

The results were clear and consistent: NeuMF outperformed all baselines on both datasets. The performance ranking was always NeuMF > MLP > GMF > BPR > eALS > ItemKNN. Deeper MLP architectures improved performance, confirming that adding nonlinear layers helps. The negative sampling ratio mattered — too few negatives (1 per positive) degraded performance, while 3–6 negatives struck the right balance.

Perhaps most importantly, the pre-training strategy consistently improved NeuMF over random initialization, validating the two-stage training approach.

Why NCF changed recommendation research

  1. 2009

    BPR — Bayesian Personalized Ranking

    Rendle et al. formalized pairwise learning for implicit feedback, optimizing the ranking of observed over unobserved items. The dominant baseline before NCF.

  2. 2017

    NCF — Neural Collaborative Filtering

    He et al. replaced the inner product with neural networks, proving MF is a special case and that deeper models give better recommendations.

  3. 2018

    Deep Interest Network (DIN)

    Alibaba's DIN added attention mechanisms to capture which historical behaviors are relevant to a candidate item, extending neural recommendation to sequences.

  4. 2019

    BERT4Rec

    Applied BERT's masked language modeling idea to sequential recommendation — masking items in a user's history and predicting them from bidirectional context.

  5. 2020

    Neural Graph Collaborative Filtering

    Extended NCF's ideas by propagating embeddings on the user-item interaction graph, capturing higher-order connectivity patterns.

NCF's insight — that the interaction function should be learned — was simple, but it redirected an entire field. Every major recommendation model since 2017, from DIN to BERT4Rec, builds on this foundation.

CitationHe, Liao, Zhang, Nie, Hu, Chua. Neural Collaborative Filtering. WWW, 2017.

Terms in this paper