Most recommender designs answer the question "which items, in what order?" A streaming service's homepage asks a harder one. The screen is a grid: a vertical list of rows, each a horizontal list of titles, each shown with a particular image. The system has to decide which rows to show, in what order, which titles go in each, and how each title is presented, and it has to do so for every profile on every device, in a fraction of a second, with a fallback when anything fails.

Netflix has described much of this publicly: a 2013 Tech Blog post by Xavier Amatriain and Justin Basilico laid out an offline, nearline and online split for its recommendation computation, and a 2017 post described choosing artwork per member with contextual bandits. This page builds a Netflix-style design from those ideas and first principles. It does not reproduce Netflix's internals or numbers. The generic retrieval and ranking funnel lives in recommender architecture; a single-feed design that learns from every swipe is in TikTok's architecture. Here the unit of output is the page.

Advertisement

The problem is a page, not a list

The homepage is a two-dimensional ranking under constraints. Members scan the top rows and the left of each row most, so position is precious in both directions. Rows carry meaning: "Because you watched X" explains a recommendation, a genre row lets someone browse a mood, and "Continue watching" serves a different intent from discovery. A good page therefore needs more than relevance per title. It needs coverage of the member's different interests, little repetition across rows, rows whose labels make sense, and room for new titles the system knows less about.

Breaking the problem into stages keeps each one tractable:

  1. Row generation: produce many candidate rows, each from its own source: continue watching, top picks, because-you-watched for recent titles, genre and theme rows, trending in your region, new releases.
  2. Within-row ranking: order the titles in each candidate row for this profile.
  3. Page construction: choose and order the rows, removing duplicates and balancing interests.
  4. Presentation: pick the image and supporting text for each title shown.

Three clocks: offline, nearline and online

Different parts of the work need different freshness, and the cost of computing them varies by orders of magnitude. The 2013 Netflix post split the work by when it runs, and the split still holds as a design tool.

Offline, nearline and online: three clocks for one homepageOffline (hours)train models, candidate rowsNearline (seconds-minutes)react to plays and list addsOnline (milliseconds)context, final ranking, pageData lakeimpressions, plays, catalogEvent streamplay, pause, add to listPrecomputed storecandidate rows and page per profile and devicewriteupdatereadDevice requestprofile, device, timeFallbacksstale page, then unpersonalized page by region
Offline jobs train models and precompute candidate rows; nearline consumers update them within seconds or minutes of member actions; the online tier reads the precomputed state, adds request context and builds the page. Fallbacks keep a page on screen when any tier fails.
  • Offline jobs run on the data lake over hours: train the ranking models, compute embeddings and similarity, generate candidate rows for every profile, and precompute a default page. They can use any amount of history and any model size, but the result is hours old.
  • Nearline consumers read the event stream and react to actions within seconds or minutes: a member finishes a series, so remove it from continue watching and refresh the because-you-watched row it seeds; a member adds a title to their list, so update that row. They run models that are cheap enough to call per event, and write results into the same store the online tier reads.
  • Online code runs inside the request: it knows the device, the time of day, the session so far, and what the member just did. It reranks precomputed candidates with that context and builds the page. It must finish in tens of milliseconds, so it does little retrieval and no heavy model training.

The design rule is to push work to the slowest tier that is fresh enough. Candidate generation over the whole catalog belongs offline; reacting to a play belongs nearline; adapting to a phone versus a television belongs online.

Advertisement

Row generation and within-row ranking

Each row source is a small recommender of its own. Continue watching is a lookup ordered by recency and the chance of resuming. Because-you-watched rows start from a recently watched title and retrieve similar ones, often with embeddings and an approximate nearest-neighbour index of the kind described in vector search at scale. Genre and theme rows filter the catalog by tags and rank by predicted interest. Trending rows rank by recent popularity in the member's region and filter by the profile's maturity setting.

A shared ranking model then orders titles inside each row for the profile. Using one model across rows keeps scores comparable, which page construction needs. Its features combine the member (history, taste embeddings, recency of activity), the title (genre, age, popularity, availability in the region) and their interaction, read from a feature store so that training and serving compute them the same way; see feature store architecture.

Generate several dozen rows, each longer than what is visible, so page construction has choices.

Page construction: choosing rows together

Ranking rows independently and showing the top ones produces a bad page: the same strong titles appear in top picks, trending and two because-you-watched rows. Page construction treats the page as a whole. A greedy builder is the simplest version that works: repeatedly add the row with the best score after penalizing titles already shown, reorder each row so unseen titles come first, and skip rows that would add too little that is new.

from collections import Counter

# Candidate rows: (row_id, row_score, titles in ranked order)
ROWS = [
    ("continue_watching", 0.95, ["t1", "t2"]),
    ("top_picks",        0.90, ["t3", "t4", "t5", "t6", "t7"]),
    ("because_t9",       0.80, ["t4", "t5", "t8", "t10", "t11"]),
    ("trending",         0.70, ["t3", "t12", "t4", "t13"]),
    ("crime_tv",         0.65, ["t8", "t15", "t16", "t17", "t18"]),
    ("new_releases",     0.60, ["t12", "t19", "t20", "t21", "t22"]),
]
VISIBLE = 3          # titles visible per row without scrolling
MIN_FRESH = 2        # a row needs this many not-yet-shown titles up front
DUP_PENALTY = 0.08   # per duplicated visible title

def build_page(rows, max_rows=4):
    shown, page = Counter(), []
    pool = list(rows)
    while pool and len(page) < max_rows:
        best, best_val, best_titles = None, -1.0, None
        for row_id, score, titles in pool:
            fresh = [t for t in titles if not shown[t]]
            reused = [t for t in titles if shown[t]]
            ordered = (fresh + reused)[:VISIBLE]   # push repeats to the right
            dups = sum(1 for t in ordered if shown[t])
            if len(fresh) < MIN_FRESH:
                continue
            val = score - DUP_PENALTY * dups
            if val > best_val:
                best, best_val, best_titles = (row_id, score, titles), val, ordered
        if best is None:
            break
        pool.remove(best)
        page.append((best[0], round(best_val, 2), best_titles))
        shown.update(best_titles)
    return page

for row_id, val, titles in build_page(ROWS):
    print(f"{row_id:18} {val:.2f} {titles}")

Running it prints:

continue_watching  0.95 ['t1', 't2']
top_picks          0.90 ['t3', 't4', 't5']
because_t9         0.80 ['t8', 't10', 't11']
crime_tv           0.65 ['t15', 't16', 't17']

Trending has a higher base score than the crime row (0.70 against 0.65), but only two of its titles are new, so its third visible slot would repeat a title already on screen. The penalty drops it to 0.62 and the crime row takes the fourth slot, which also adds a different interest to the page. Because-you-watched keeps its place, but with t4 and t5 moved right, since top picks already shows them.

Production builders add constraints: pinned rows, limits per row type, a minimum number of genres on the first screen, slots for new titles and device render limits. Some replace the greedy loop with a model that scores a row given the rows above it. Either way, evaluate the page as a unit.

Artwork personalization with contextual bandits

The same title can be shown with different images: one emphasizing a lead actor, another the genre, another a particular scene. Netflix's 2017 artwork post framed choosing among them as a contextual bandit problem. For each impression the system picks one image given the member's context, observes whether the member played the title, and learns which images work for whom.

The bandit framing matters because the system only learns about images it shows. It must sometimes show an image it does not think is best, which is exploration, and it must log how likely its policy was to show each image, so that later analysis can correct for the choices it made. Multi-armed bandits covers the algorithms; the serving pattern looks like this:

def choose_artwork(title, context, model, epsilon=0.05):
    arms = artwork_candidates(title)                  # e.g. 3-8 images per title
    scores = {a: model.predict(context, title, a) for a in arms}
    best = max(scores, key=scores.get)
    if random() < epsilon:
        chosen = choice(arms)                         # explore
    else:
        chosen = best                                 # exploit
    # Probability that THIS policy shows `chosen` in THIS context.
    propensity = epsilon / len(arms) + (1 - epsilon) * (chosen == best)
    log_impression(title, chosen, context, propensity, model.version)
    return chosen

The logged propensity is the key field. With it, an offline replay can estimate how a new policy would have performed on logged traffic by weighting each logged impression by the ratio of the new policy's probability to the old one's. Without it, the logs only describe what the current policy liked to show, and any new policy trained on them inherits its blind spots.

The serving path and its fallbacks

A homepage request flows through a short path. Authenticate and load the profile; read the precomputed page and candidate rows for that profile and device class from a low-latency store; rerank with request context; apply business and legal filters, such as regional availability and maturity rating; choose artwork; render the response for the device. Each step has a budget, and the whole request has to fit inside the time the app waits before showing something.

Personalization is valuable but not essential for every request; a page on screen is essential. Design a fallback chain and test it:

  1. If online reranking times out, serve the precomputed page as it is.
  2. If the precomputed page is missing or the store is unavailable, serve the last page cached for that profile, even if it is a day old.
  3. If there is nothing for the profile, serve an unpersonalized page for the region and maturity level, built offline from popularity.

Every level must still apply availability and maturity filters, and the rate of each fallback must be measured, because a silent rise in fallback traffic looks like a quiet drop in engagement.

Logging impressions, not just plays

Every model above trains on the join of what was shown with what happened. Log each impression with profile, device, row and position, title position, artwork id, model versions and propensities; log plays, their duration and list adds; join them within a window.

Without impressions, a model cannot tell "not interested" from "never saw it". And a title in the first slot gets attention for being there, so record position and correct for it in training.

Evaluation: offline, interleaving, A/B

Offline metrics, such as recall of later plays in the top titles of a row or the page, are cheap and good for discarding bad ideas, but they are measured on logs produced by the old system. Netflix has also written about interleaving: mixing two rankers' results in one list for the same member and seeing whose titles get played, which detects ranking differences with far fewer members than a full A/B test. The final decision belongs to an A/B test on the outcomes the business cares about, such as retention and viewing over weeks rather than clicks in a session; A/B testing for ML covers the mechanics.

Worked example: sizing the precomputed store

Take a hypothetical service with 50 million profiles that use an average of two device classes. Precompute 40 candidate rows of 40 titles each per profile and device, stored as 4-byte title ids with a 2-byte score: about 9.6 KB per page before overhead. That is 50 million times 2 times 9.6 KB, about 960 GB, or around 2 TB with replication and overhead, which a distributed key-value store or a sharded in-memory cache handles comfortably.

Rebuilding every page nightly writes 100 million entries, about 10,000 writes per second over three hours; nearline updates add a few thousand per second at peak. At a peak of 100,000 homepage loads per second, the store must serve reads in low single-digit milliseconds so reranking keeps most of the budget. The numbers are invented; the point is that precomputation turns a model problem into a storage and freshness problem that is far easier to scale.

Failure modes

  • Duplicate-heavy pages: rows ranked independently show the same few titles four times.
  • Stale state after actions: a finished series stays in continue watching because the nearline path is lagging.
  • Feedback loops: titles shown more get played more and are shown even more; without exploration and position correction, the catalog narrows.
  • Missing propensities: offline replay of a new artwork policy becomes impossible or biased.
  • Fallbacks that skip filters: a cached page shows a title no longer available, or one above the profile's maturity level.
  • Optimizing the wrong metric: a policy raises plays per session and lowers long-term retention, which only a long A/B test reveals.

Trade-offs

Precomputation is fast and cheap but only as fresh as the nearline path; online reranking adds context but spends latency. Exploration costs a little engagement now for better models later, and diversity penalties broaden pages but can bury a member's strongest interest. Decide each against A/B outcomes, not an offline score.

What to do next

  1. Write down your page's structure: row types, rows per screen and titles visible per row on each device.
  2. List every row source, and for each decide whether it is computed offline, nearline or online.
  3. Add page-level deduplication and diversity to page construction, and evaluate whole pages, not single rows.
  4. Log impressions with row, position, artwork, model version and propensity, and join them to outcomes.
  5. Build and test the fallback chain, applying availability and maturity filters at every level, and alert on fallback rate.
  6. Add a small, logged exploration budget to artwork or row choice, and use replay with propensities to screen new policies.
  7. Decide launches with A/B tests on retention and viewing over weeks, using interleaving to screen rankers quickly.
Key takeaway: A Netflix-style homepage is a page-construction problem: generate many candidate rows from different sources, rank titles within each row with one shared model, then choose and order rows together with deduplication and diversity, and choose artwork per member. Split computation into offline, nearline and online tiers so heavy work runs ahead of the request and actions still show up within seconds, and keep stale and unpersonalized fallbacks. Log impressions, positions and propensities, and decide with long A/B tests.