Instagram's main feed looks like the classic news-feed problem: show each user recent posts from the accounts they follow. In practice it is a mixing problem. A page interleaves posts from followed accounts, suggested posts from accounts the user does not follow, and ads, each with its own retrieval and scoring, and every item is a photo, carousel or video that must reach the phone fast enough that scrolling never stalls.
Our Twitter timeline architecture and news feed fan-out guide already cover timeline caches, fan-out on write versus read, and multi-stage ranking. This article assumes those and focuses on what an Instagram-shaped feed adds: per-surface rankers, blending three candidate streams under diversity and ad-load rules, seen-state and stable pagination, and media delivery. The internals described are a reference architecture consistent with Instagram's public explanations, not a description of Meta's private systems.
One app, several rankers
The one structural fact Instagram has stated plainly, in its 2023 explainer on ranking, is that there is no single algorithm. Feed, Stories, Explore, Reels and Search each have their own ranking because people use them differently. Stories rank only accounts you follow, weighted by how often you view them. Explore and Reels are mostly recommendations from accounts you do not follow. Feed is the hybrid: followed posts first, suggested posts after or between them. In 2022 Instagram also added Following and Favorites views that show followed accounts in chronological order, so the ranked feed is a default, not the only option.
Architecturally, that means you build a shared platform (feature store, model serving, impression logging, media pipeline) and several thin surface pipelines on top, each owning its candidate sources, objective and blending rules. One mega-ranker for every surface would force Stories to trade off against Reels in a single loss, and every change would need sign-off from every surface owner.
Candidate inventory: followed, suggested and ads
For followed content the useful unit is the unseen inventory: posts from accounts you follow, created since a lookback horizon, that you have not yet been shown. Because a typical user follows a few hundred accounts that mostly post rarely, inventory is cheapest to build on read: query each followed author's recent-posts index, a short list keyed by author and sorted by time, and merge. Accounts with huge audiences make fan-out on write expensive, while accounts that post often and have few followers make it cheap; the hybrid described in the fan-out guide applies unchanged.
What changes for Instagram is the horizon. If a user opens the app twice a week, a strict 24-hour window shows them almost nothing they missed. Grow the lookback until the inventory reaches a target size, say 500 candidates, or a cap of several days. Then subtract seen posts. This is why a returning user sees posts from two days ago at the top: they are the best unseen candidates, not the newest.
Suggested candidates come from interest retrieval, typically embedding nearest-neighbour search over recent posts seeded by the user's engagement history, which the short-video recommendation article covers in depth. Ads are not candidates in the same sense: an ads system runs its own auction and returns a few sponsored items with bids and predicted value, and the feed only decides where they may go.
Ranking with a value function
Instagram's explainer lists the signals feed ranking uses: your activity (what you liked, shared, saved, commented on), information about the post (popularity, time, length, location), information about the author (how people have engaged with them recently), and your history with that author. Models turn these into predicted probabilities of specific actions, and a value function combines them into one score. A minimal version:
# predicted probabilities come from the model server, one vector per candidate
WEIGHTS = {"like": 1.0, "comment": 4.0, "share": 6.0, "save": 5.0,
"dwell_10s": 2.0, "profile_tap": 1.5, "hide": -40.0, "report": -200.0}
def value(pred, cand, user):
v = sum(WEIGHTS[k] * pred[k] for k in WEIGHTS)
if cand.source == "followed":
v *= 1.0 + 0.3 * user.closeness(cand.author_id) # relationship strength
if cand.is_reshare_of_original_elsewhere:
v *= 0.7 # prefer original content
return v
def rank(cands, user, model, budget_ms=60):
light = model.light_score(cands, user) # cheap model, all ~1,500
top = sorted(cands, key=light.get, reverse=True)[:300]
preds = model.heavy_predict(top, user, timeout_ms=budget_ms)
if preds is None: # timeout: degrade, never fail the page
return [(c, light[c]) for c in top]
return sorted(((c, value(preds[c], c, user)) for c in top),
key=lambda t: t[1], reverse=True)The weights are product decisions expressed as numbers. Negative weights on hide and report are what make the feed avoid content people dislike, and they must be large because those events are rare. The two-stage shape, light model on everything and heavy model on a few hundred, is what keeps the scoring budget bounded as inventory grows.
Blending: why sorted order is not a feed
Sorting by score produces a bad feed. Five posts in a row from the friend who posted a holiday album, or a run of suggestions with no followed posts, both make users leave. Instagram has said it avoids showing too many posts from the same account in a row and too many suggested posts back to back. A blender enforces such constraints greedily over the ranked list:
def blend(ranked, ads, page_size=12, max_same_author_run=1,
max_suggested_run=2, ad_every=5, min_ad_gap=4):
page, author_last, sugg_run, since_ad = [], None, 0, 0
pool, ads = list(ranked), list(ads)
while len(page) < page_size and pool: # ads only go between organic items
if ads and since_ad >= ad_every - 1 and since_ad >= min_ad_gap:
page.append(ads.pop(0)); since_ad = 0
continue
for i, (cand, _) in enumerate(pool): # best item that obeys the rules
if cand.author_id == author_last and max_same_author_run <= 1:
continue
if cand.source == "suggested" and sugg_run >= max_suggested_run:
continue
break
else:
i = 0 # nothing legal: relax rather than stall
cand, _ = pool.pop(i)
page.append(cand)
author_last = cand.author_id
sugg_run = sugg_run + 1 if cand.source == "suggested" else 0
since_ad += 1
return pageIntegrity filters run before blending: remove posts the user cannot see (blocked, private, deleted since retrieval) and down-rank content flagged by classifiers. Ads also obey frequency caps per advertiser across sessions, which is a read from the same seen-state store.
Seen-state and stable pagination
Two user-visible bugs dominate feed complaints: seeing the same post twice, and posts jumping around while scrolling. Both are state problems, not ranking problems.
- Impression logging. The client reports a post as seen when it was on screen long enough, not when it was sent, and the server writes it to a per-user seen set with a TTL of a week or two. A Bloom filter per user keeps the membership check cheap; false positives hide a post the user never saw, which is acceptable, while false negatives would show repeats and do not occur.
- Ranked pages are snapshots. When the first page is requested, the server ranks a larger window, say 100 items, stores the ordered ids under a session key with a short TTL, and returns page one with an opaque cursor naming the session and offset. Later pages read from the snapshot, so new scores cannot reorder what the user is scrolling through.
- Refresh is explicit. Pull-to-refresh, or returning after a gap, starts a new session. New items appear at the top; the old snapshot is discarded.
def get_page(user, cursor=None, size=12):
if cursor is None or not snapshots.exists(cursor.session):
ids = [c.id for c in blend(rank(gather(user), user, model), ads_for(user), page_size=100)]
session = snapshots.put(user.id, ids, ttl_s=1800)
offset = 0
else:
session, offset = cursor.session, cursor.offset
ids = snapshots.slice(session, offset, size)
ids = [i for i in ids if visible(user, i)] # re-check deletes and blocks
return hydrate(ids), Cursor(session, offset + size)
Media delivery on the read path
Feed ranking can be perfect and the product still feel slow if images arrive late. Each upload is transcoded into several renditions (widths for different screens, modern and fallback codecs, video at multiple bitrates) and stored once; the feed response carries URLs for the renditions suited to the device and network, not the bytes. Those URLs are served from a CDN whose cache key includes the rendition, so the popular files stay at the edge.
The client prefetches the media for the next few items while the user is still looking at the current one, and the server sizes the first page to cover what fits on screen plus that prefetch window. Prefetching too far wastes cellular data on posts the user scrolls past; too little produces grey placeholders. Make the prefetch depth a client setting tied to network type rather than a server constant.
Worked example: one feed request
A user who follows 400 accounts opens the app after 36 hours away. The numbers below are illustrative, chosen to show where the budget goes.
- The Feed API checks for a live snapshot, finds none, and fans out in parallel: recent-posts lookups for 400 authors return 620 posts within a 72-hour horizon; the seen filter removes 190, leaving 430 followed candidates; interest retrieval returns 1,000 suggested candidates; the ads system returns 6 eligible ads. Elapsed: about 40 ms, bounded by the slowest source.
- The light model scores 1,430 candidates and keeps 300. The heavy model predicts action probabilities for those 300 in about 60 ms.
- The value function orders them; integrity filtering removes 4 posts from an account the user blocked yesterday.
- The blender builds a 100-item snapshot: followed posts dominate the top, no author appears twice in a row, suggestions never run more than two, and ads land roughly every fifth slot.
- The first 12 items are hydrated with captions, counts and rendition URLs and returned with a cursor. Total server time is around 150 ms; the phone begins fetching images for items 1 to 5 immediately.
Failure modes and degradation
- Heavy model timeout. Serve light-model order rather than an error; log the degraded rate, because a silent fallback hides a capacity problem.
- Suggested source down. Serve followed content only. A shorter, all-followed feed is better than an empty one.
- Seen-store unavailable. Fall back to a short client-side seen list sent with the request, accepting a few repeats.
- Stale snapshot after a delete. Re-check visibility at hydration time, as in the code above.
- Feedback loops. Ranking trains on what it showed, so it learns its own biases. Keep a small exploration share and evaluate on held-out traffic.
- Media origin overload. A viral post can miss the CDN in many regions at once; use request coalescing at the edge and an origin shield.
Trade-offs
Ranking versus chronology. Ranking raises engagement but makes the feed feel unpredictable; offering a chronological Following view is a cheap pressure valve. Snapshot size. Larger snapshots make pagination stable for longer but spend ranking work on items many users never reach. Suggestions. More suggested content grows discovery for creators at the cost of users seeing less from friends; make the cap a tuned parameter, not a constant. Per-surface pipelines. They let teams move independently, but duplicate infrastructure unless the shared layers are real platforms.
What to do next
- Write down each surface you have and its objective; do not let one ranker serve them all.
- Build followed inventory on read with an adaptive lookback, and subtract a seen set fed by client impressions.
- Use a two-stage ranker with a fixed candidate count for the heavy model and a light-model fallback on timeout.
- Implement the blender with explicit author-run, suggestion-run and ad-gap rules, and unit-test them.
- Serve pages from a stored ranked snapshot with an opaque cursor, and re-check visibility at hydration.
- Return rendition URLs, not bytes, and tune client prefetch depth per network type.
- Instrument degraded-mode rates and repeat-impression rates as first-class feed health metrics.