Method

Exactly how the recommendations are produced

No black box and no machine learning. Two vector computations and a weighted average — all of it readable below, and all of it unit-tested.

Step 1

Your ratings become a vector

Every film you rate contributes one component: +1 for a like, 0 for neutral, −1 for a dislike. Films you skip are absent, not zero — a blank must never be read as indifference, and it must never dilute the denominator.

Step 2

Each critic is a vector too

Every critic's published verdicts are normalised to the same −1…+1 scale. Reviews scored 3/4, B+, 8/10 or 4/5 all map onto that axis, and the published Fresh/Rotten label is treated as authoritative for the sign.

Step 3

Cosine similarity over films you both covered

sim(u,c) = Σm ∈ M(u,c) ru,m · rc,m ── / ── ( √(Σ ru,m²) · √(Σ rc,m²) )

M(u,c) is the co-rated film set. A critic needs at least 3 co-rated films to be eligible at all, and only critics with a positive similarity can be selected — a critic pointing the other way is not a top critic, and a negative weight would punish films they loved.

One deliberate refinement. Raw cosine is unstable when the overlap is tiny: three coincidental agreements already produce 1.000. So ranking uses adjusted = similarity × n / (n + 5) where n is the co-rated count. A 99% match on 4 films scores 0.44; a 76% match on 16 films scores 0.58 and therefore ranks higher. The raw cosine is still shown, unchanged.
Step 4

Your top 10 critics, weighted

The surviving critics are sorted by adjusted similarity, truncated to 10, and their weights normalised to sum to 1. Those ten weights are your profile.

Step 5 · the joint score

Films neither of you has rated, scored by both sets of critics

Score(m) = ( Σi ∈ Top_A wA,i · ri,m + Σj ∈ Top_B wB,j · rj,m ) ── / ── ( Σ wA,i + Σ wB,j )

The denominator is the full weight of both profiles — which is the strict consensus reading. A film only ranks highly when it is broadly liked across both sets of critics, not merely adored by one enthusiastic specialist.

That strictness has a blind spot: a film reviewed by two of the twenty critics is capped low no matter how much those two loved it. So the deck exposes a second ordering — Strongest advocates — which uses the same numerator but divides only by the weights of the critics who actually filed. Both numbers are on every card, so the choice is yours rather than ours.

Broad consensus

Sorted by Score(m). Favours films with wide agreement and solid coverage. The default.

Strongest advocates

Sorted by the coverage-adjusted score. Favours films a few of your critics were passionate about.

Ties break toward films both profiles rate positively on average, then toward broader evidence, then editorial popularity. A final pass caps the deck at two consecutive films sharing a primary genre, so a single taste cannot swallow the queue.

Live layer

In theatres

The release list is fetched in this order, each layer falling back to the next: TMDB now_playing when an API key is configured; otherwise a scrape of the Rotten Tomatoes in-theatres browse page, which is server-rendered and yields real titles, posters, opening dates and critic scores; otherwise the bundled snapshot set. Results are cached in SQLite for 24 hours.

Per-title critic quotes are attempted from TMDB's reviews endpoint, then from RT film pages, then matched against the bundled archive restricted to your top critics. RT stopped server-rendering critic review text, so that middle layer usually yields nothing — which is precisely why the archive layer exists and why every result carries its provenance in the UI.

RT browse scraping: disabledTMDB key: not configured
Provenance

Where the data comes from

600films
130critics
15,087real reviews
64%of them Fresh
  • Metacritic export matched to 503/600 films: audience score attached to all of them, and 158 synopses upgraded to the longer description. Audience data is display-only and is never an input to critic matching or the joint score.
  • 456 reviews were dropped because the export had no review text for them.
  • Bundled in-theatres fallback covers 16 films released 2018+ in the snapshot era (Joker, Once Upon a Time In Hollywood, Us, Avengers: Endgame…). Live sources take priority.
  • Source: Kaggle Rotten Tomatoes exports (17,712 films / 1,130,017 critic reviews, snapshot Nov 2020). Critic names, publications, review text and Fresh/Rotten verdicts are real; per-critic taste parameters (era bias, mainstream bias, contrarianism, harshness) are computed from that data.
Honest limitations

What this does not do

  • The snapshot is frozen. The corpus is a Rotten Tomatoes export from late 2020. New releases only enter through the live theatrical layer, and only with whatever metadata that layer exposes.
  • Similarity is not causation. Agreeing with a critic on the films you have both covered does not mean you will agree on the next one. The match percentage is a correlation over a handful of films, nothing more.
  • Thin overlaps stay thin. A 20-film quiz produces 3–18 co-rated films per critic. The evidence prior makes that honest rather than pretending otherwise, but more ratings genuinely mean a better profile.
  • Audience scores are decoration. The Metacritic audience figure on each card is displayed for context and is never an input to the matching or the joint score.
  • Not affiliated with Rotten Tomatoes. Critic names, publications and review text are reproduced from a public dataset for demonstration. This is a portfolio project, not a commercial product.