AI Summary

Netflix's LLM ranker beat a model tuned for years, using 40x less training data

GenRec, Netflix's LLM-backed ranker, beat a production model built on thousands of engineered features in a 4-week A/B test on about 10% of traffic, while using roughly 40x fewer post-training examples and a prompt cut from 5,000 to 1,700 tokens.

Netflix has put an LLM in the ranking seat of its recommender and it beat the incumbent. GenRec, a decoder-only model fine-tuned from an in-house foundation LLM, outperformed a production ranker “tuned over many years with thousands of engineered features” in a 4-week A/B test on about 10% of Netflix traffic. It did that with roughly 40x fewer labeled post-training examples and fewer input signals, and it reads the user’s history as a plain-text prompt instead of a feature vector.

  • Offline: about +1.6% relative MRR over the production ranker, and still climbing as data and signals scale.
  • Online: statistically significant wins on both short-term and long-term metrics. The core metric moved +0.006% relative, which is small in isolation but significant at Netflix scale.
  • 40x less Phase-2 training data than the production model needed to reach and pass parity.
  • Prompt cut from about 5,000 to 1,700 tokens with negligible quality loss, cutting serving cost to roughly one-third.
  • Freshness matters more than the base model: task-specific post-training added +35 to 50% MRR over the foundation model alone, growing to about +80% two weeks later.

From feature engineering to context engineering

The classic Netflix ranker is a discriminative model fed hand-built features across users, titles, and their interactions, with custom architectures per surface. Every new business need (games, live events, podcasts) means more features and more re-engineering. Off-the-shelf LLMs don’t fix this on their own: the paper notes they over-recommend popular titles, hallucinate things not in the catalog, and barely personalize.

GenRec’s answer is to turn the user into text. A verbalizer writes context (surface, device, time, locale), profile (country, tenure, plan), watch history, and item metadata into a prompt, and the model is left to infer the feature interactions itself. The authors frame the shift directly: the job becomes “deciding what information to present to the model, how to express it within a limited token budget, and what objective and rewards to optimize for.”

The verbalizer’s rules are the new feature engineering:

  1. Keep in full: high-signal events like long plays and thumbs-up, with rich metadata.
  2. Drop: recent low-signal noise such as very short plays and stray clicks.
  3. Compress: repetitive behavior, like a binge session, into a summary.
  4. Elaborate: new releases and cold-start titles, where the model has little else to go on.

Short- to medium-term history gets the most detail; older history is dropped or folded into a brief interest summary.

How the model ranks

GenRec never generates a title name. It runs prefill only: the prompt goes through the LLM once, a pooled hidden state is compared against a learned embedding for every item in the catalog, and a softmax over the catalog produces the ranking. One forward pass scores every candidate, and because outputs are constrained to catalog items, hallucinated titles are impossible by construction. It’s served on vLLM, and the authors note the stack now looks like LLM infrastructure (GPUs, batching, caching) rather than classic RecSys MLPs and factorization models.

Training happens in two phases on different clocks:

  • Phase 1 adapts an open-source LLM to Netflix data for user and content understanding. It updates rarely and isn’t constrained by serving cost.
  • Phase 2 post-trains for ranking. It refreshes often to track new releases and shifting tastes, and has to be cheap to serve. This is the paper’s focus.

The Phase 2 loss mixes a catalog-wide ranking objective with a language-modeling objective (to keep prompt-based steering possible). Raw engagement alone would drift toward what gets clicked, so the ranking loss is reward-weighted: reward models estimate which short-term engagements predict long-term retention and catalog exploration, and others rebalance across movies, games, live, and podcasts. Without this, the paper says, the model “might over-recommend binge-watching over discovery, favor videos over games.” GRPO-style RL showed extra gains in early tests but was too expensive to train, so it’s deferred.

Where the offline gains come from

Relative MRR improvement from each stage, as reported (upper bound of range shown). Bars scaled to 80%.

Phase 1 backbone vs. off-the-shelf LLM+10 to 20%
Phase 2 on a fresh Phase 1 model+35 to 50%
Phase 2, two weeks after Phase 1 cutoffabout +80%

The Phase 2 gap widens over time because the foundation model goes stale on popularity and interests. That’s the argument for splitting the phases: the expensive model can update slowly, as long as a cheap post-training step keeps it current. And since Phase 2 needs 40x fewer labels than the old ranker, refreshing it often is affordable.

Cost is a prompt-length problem

The paper’s rule of thumb: inference cost scales roughly with model size times context length. So Netflix attacked both. On the model side, it explored distillation into smaller backbones. On the context side, it ran a three-step compaction: pick and compress events, sweep history length to find the elbow where extra tokens stop helping, then trim wording and drop few-shot examples.

Prompt budget per ranking request

Offline ranking quality was essentially unchanged. Serving cost fell to about one-third.

Original verbalizationabout 5,000 tokens
Compacted verbalizationabout 1,700 tokens

Scaling behaves like it does for LLMs generally. MRR rose monotonically as Phase 2 data went from 1x to 20x, for both a roughly 1B and a roughly 10B model, and at a fixed GPU budget the 10B model consistently scored higher. Netflix then picked a point on that curve on purpose: big enough to capture most of the gain, small enough to serve.

Caveats

  • The online win is tiny in absolute terms. +0.006% relative on the core metric is statistically significant at Netflix’s scale, but the paper doesn’t name the metric, and most of the evidence is directional.
  • Few absolute numbers. Model sizes are “about 1B” and “about 10B,” scaling results are relative, and there’s no public dataset or code. This is an industry report, not a reproducible benchmark.
  • Batch surfaces only. The A/B test ran on batch-compute surfaces, where latency is forgiving. Real-time ranking with an LLM in the loop is a harder serving problem the paper doesn’t claim to have solved.
  • It’s not really generative. Despite the name, production GenRec is an LLM encoder with a catalog classification head. Explanations and free-form steering are listed as future work.
  • The comparison is deliberately lopsided. GenRec ran in a “low-data, low-signals configuration” against the full production system. That makes the data-efficiency claim strong, but offline gains were still growing, so the final margin is unknown.

The takeaway for anyone running a recommender: the valuable work moved from designing features and bespoke architectures to deciding what goes in the prompt and what the reward should value. Netflix’s biggest serving win came from cutting two-thirds of the prompt, not from a new model.

#research #llms #recommendation-systems #system-design

Liked this? Engineer's Codex sends one deep dive and a link roundup every week.

No spam. Unsubscribe anytime.