Exactly how the recommendations are produced
No black box and no machine learning. Two vector computations and a weighted average — all of it readable below, and all of it unit-tested.
Your ratings become a vector
Every film you rate contributes one component: +1 for a like, 0 for neutral, −1 for a dislike. Films you skip are absent, not zero — a blank must never be read as indifference, and it must never dilute the denominator.
Step 2Each critic is a vector too
Every critic's published verdicts are normalised to the same −1…+1 scale. Reviews scored 3/4, B+, 8/10 or 4/5 all map onto that axis, and the published Fresh/Rotten label is treated as authoritative for the sign.
Step 3Cosine similarity over films you both covered
M(u,c) is the co-rated film set. A critic needs at least 3 co-rated films to be eligible at all, and only critics with a positive similarity can be selected — a critic pointing the other way is not a top critic, and a negative weight would punish films they loved.
Your top 10 critics, weighted
The surviving critics are sorted by adjusted similarity, truncated to 10, and their weights normalised to sum to 1. Those ten weights are your profile.
Films neither of you has rated, scored by both sets of critics
The denominator is the full weight of both profiles — which is the strict consensus reading. A film only ranks highly when it is broadly liked across both sets of critics, not merely adored by one enthusiastic specialist.
That strictness has a blind spot: a film reviewed by two of the twenty critics is capped low no matter how much those two loved it. So the deck exposes a second ordering — Strongest advocates — which uses the same numerator but divides only by the weights of the critics who actually filed. Both numbers are on every card, so the choice is yours rather than ours.
Sorted by Score(m). Favours films with wide agreement and solid coverage. The default.
Sorted by the coverage-adjusted score. Favours films a few of your critics were passionate about.
Ties break toward films both profiles rate positively on average, then toward broader evidence, then editorial popularity. A final pass caps the deck at two consecutive films sharing a primary genre, so a single taste cannot swallow the queue.
In theatres
The release list is fetched in this order, each layer falling back to the next: TMDB now_playing when an API key is configured; otherwise a scrape of the Rotten Tomatoes in-theatres browse page, which is server-rendered and yields real titles, posters, opening dates and critic scores; otherwise the bundled snapshot set. Results are cached in SQLite for 24 hours.
Per-title critic quotes are attempted from TMDB's reviews endpoint, then from RT film pages, then matched against the bundled archive restricted to your top critics. RT stopped server-rendering critic review text, so that middle layer usually yields nothing — which is precisely why the archive layer exists and why every result carries its provenance in the UI.
Where the data comes from
- Metacritic export matched to 503/600 films: audience score attached to all of them, and 158 synopses upgraded to the longer description. Audience data is display-only and is never an input to critic matching or the joint score.
- 456 reviews were dropped because the export had no review text for them.
- Bundled in-theatres fallback covers 16 films released 2018+ in the snapshot era (Joker, Once Upon a Time In Hollywood, Us, Avengers: Endgame…). Live sources take priority.
- Source: Kaggle Rotten Tomatoes exports (17,712 films / 1,130,017 critic reviews, snapshot Nov 2020). Critic names, publications, review text and Fresh/Rotten verdicts are real; per-critic taste parameters (era bias, mainstream bias, contrarianism, harshness) are computed from that data.
What this does not do
- The snapshot is frozen. The corpus is a Rotten Tomatoes export from late 2020. New releases only enter through the live theatrical layer, and only with whatever metadata that layer exposes.
- Similarity is not causation. Agreeing with a critic on the films you have both covered does not mean you will agree on the next one. The match percentage is a correlation over a handful of films, nothing more.
- Thin overlaps stay thin. A 20-film quiz produces 3–18 co-rated films per critic. The evidence prior makes that honest rather than pretending otherwise, but more ratings genuinely mean a better profile.
- Audience scores are decoration. The Metacritic audience figure on each card is displayed for context and is never an input to the matching or the joint score.
- Not affiliated with Rotten Tomatoes. Critic names, publications and review text are reproduced from a public dataset for demonstration. This is a portfolio project, not a commercial product.