Closet Picks

Method

How this is counted

Everything here is scraped from criterion.com, checked daily and rebuilt whenever Criterion publishes something new. Nothing is hand-entered, so what is on this site is whatever Criterion published.

The source

Three pages carry the whole dataset. The Closet Picks search page lists every visit. Each visit's own page lists what that guest chose, in the order the page shows them. And the collection list gives every film Criterion sells, with its spine number, director, country and year, which is the denominator for anything phrased as a share.

Box sets are one choice

Guests do not only pick single films. 102 different box sets have been chosen, and a set is one decision however many discs it holds. Counting Ingmar Bergman’s Cinema as its 40 films would put Ingmar Bergman far above everyone else on packaging alone, so rankings on this site count a set once.

Where the question is about films rather than choices, sets are opened up: that is the difference between 851 films picked directly (49% of the collection) and 1273 reached in total (73%).

What this misses

Where the cover art comes from

The cover art is Criterion's, from the same collection list every other fact on this site is read off. We do not link to their images: each one is fetched once, re-encoded to 300 pixels wide as WebP, and served from this site's own origin, so browsing here never spends Criterion's bandwidth and never breaks when they move a bucket. 1,736 of 1,736 films carry one (100%), and 101 of the 102 box sets anyone has taken. Anything without one renders an empty box at the same size rather than a broken image, so no list ever shifts as the pictures arrive. The one exception is the poster frame on a visit with no video, which shows Criterion's own image from their server: two visits are in that position and neither is worth a second image pipeline.

Every name is a link

Every film in the collection has a page, including the 463 counted as never reached: a dead end is the worst place to leave a reader who has just been told something was never picked, and that page is where the count says what it cannot see. Every one of the 707 credited directors has a page too, not only the 162 the ranking publishes, and so does every box set anyone has taken. A name we have no page for stays plain text rather than becoming a guess at a search URL.

What comes from TMDB

Criterion publishes director, country, year and spine, and nothing else. Genre and runtime come from TMDB, and so does the rating for any film IMDb does not itself cover. A candidate is accepted only when its title agrees with Criterion's, and then only when the director's surname agrees too, or failing that a release year within a year either way. A shared surname on its own is not a match: it would attach another film's rating and nothing downstream could tell. 1727 of 1736 films matched (99%).

The 9 that did not are mostly television, which a movie search cannot see, and compilations with no single entry. Matching is not the same as coverage: TMDB records no runtime for some of the films it does know, so every page built on these fields counts the field it actually reads rather than the match. 1720 films carry a genre and 1715 carry a runtime.

A rating from fewer than ten TMDB voters is dropped rather than averaged, and a runtime TMDB records as zero is treated as unknown rather than as a film of no length.

What comes from IMDb

TMDB's own vote counts are thin for this collection: a median of 201 votes per film, against a median of 9,150 on IMDb, forty-five times deeper. IMDb is therefore the primary rating wherever it has one, with TMDB standing in only for the films IMDb does not cover. The two are never averaged together for the same film: different scales, different populations, and blending them would describe neither. Exactly one source wins per film, and which one is recorded so a chart can name it. 1705 of 1736 films carry a rating from either source: 1692 from IMDb and 13 falling back to TMDB.

The join adds no matching of its own: the id IMDb is looked up by comes straight off the TMDB record that already matched, so it never widens or narrows which films are covered beyond what TMDB found. A rating from fewer than a hundred IMDb voters is dropped rather than published, a higher floor than TMDB's ten because IMDb's own audience per film here runs so much larger.

The line below, and the same line in the footer of every page, is IMDb's own wording rather than ours. IMDb publishes these datasets for personal and non-commercial use and sets out the conditions itself: take the data from the files they publish, agree to their conditions of use, keep the use non-commercial, and acknowledge the source in exactly the sentence they specify. So the permission is real, but it is a standing grant to anyone who meets those conditions rather than an arrangement this site negotiated, and IMDb can withdraw it whenever they like. The sentence stays word for word because changing it would break one of the conditions.

The same page permits keeping a local copy, which is what happens here: the ratings file is downloaded, cached and read from disk rather than queried per film. Nothing on this site is sold and it carries no advertising; the tip jar in the footer is a donation towards the hosting, not a sale. If that ever changed the licence would stop applying and the IMDb ratings would have to come out.

Information courtesy of IMDb (https://www.imdb.com). Used with permission.

Where the video comes from

Criterion posts every visit to Vimeo, and their Vimeo uploads are restricted to criterion.com, so an embed of one here renders a privacy notice rather than the film. The same visits are on their YouTube channel, gathered in one playlist, and those embed normally. Matching the two lists is done once by hand and the result is committed, so a rebuild reads a file rather than calling YouTube. 401 of 414 visits have a video on record.

Most of them match on the guest's name alone. The ones that do not are pinned by hand rather than guessed: nine guests have been in the closet twice, so their name matches two videos, and putting the wrong decade's visit on a page is the kind of mistake that looks completely fine. The 13 with no match at all are not on the channel under any title we could find, and they show a poster frame linking to Criterion's own copy instead of a player that would not play.

Three questions, and what the answers would not carry

Three things somebody would reasonably want a section built on. Each was tested; none of them supports the section it would have been, and one of them came back significant and still does not. They are here because a chart that was not drawn is part of the method, and because the next person to have the idea should be able to see what happened to it.

The order of a visit's list carries almost nothing

Criterion's page lists a visit's choices in some order, and it is tempting to read that order: the first pick as the headline, the last as the deep cut. Permuting position within each visit 2,000 times, the strongest thing the order carries is a rank correlation of 0.053 with film year. That is two to three standard errors from zero over 2,965 picks, so it is real and it is also a fifth of one percent of the variance. A list drifts very slightly from older and more canonical towards newer and more obscure, and nothing in it supports reading one visit as a path through film history. The one difference big enough to see is that box sets sit later in a list than single films (0.539 against 0.494 on a nought-to-one scale).

Position in the list Picks Mean film year Mean rating Subtitled
first fifth 662 1973.5 7.596 39%
second fifth 534 1976.1 7.52 38%
middle fifth 521 1976.6 7.512 39%
fourth fifth 519 1978.8 7.499 36%
last fifth 729 1976.3 7.544 35%

Relative position is 0 for the first item on Criterion's page and 1 for the last, over visits that chose more than one thing. Each test permutes position within each visit 2000 times, which holds every visit's contents fixed and asks only whether the order inside it carries the measure. mean_picks_of_film is size-biased by construction and is a reading aid rather than an average over films.

Picks of the same film cluster in time, and that is not influence

Do repeat picks of a film cluster, so that one guest taking it makes the next guest likelier to? The spread of pick dates is tighter than chance at a p-value of 0.0005, which is as small as 2,000 draws allow. That is not evidence that guests follow each other, and this section exists to stop that sentence being written.

The null holds every visit's list and every film's pick count fixed and only shuffles which visit sat at which date. So it rejects whenever a visit's contents and its date are related at all, and the collection arriving in instalments does that on its own: a film that reached the shelf in 2024 could only ever be picked late, and this statistic calls that clustering. The other candidate mechanism is measured and set aside, since visits did not get systematically longer or shorter as the series went on (rank correlation -0.026, p 0.6052).

Films Observed spread Expected by chance Standard deviations p
Every film with 3 or more picks 409 80.6 86.28 -4.1 0.0005
Only spine 600 and below 174 88.03 86.47 0.72 0.7581

Restricted to the older half of the shelf the effect disappears (0.72 standard deviations against -4.1). That is suggestive and it is not a control: a spine number is release order, not a release date, low-spine films go out of print and come back, and the subset is smaller. It shows that the departure from chance shrinks among films that are on average older. It cannot show that availability is why.

A guest's own director is no more likely to be reached

205 films in the collection were directed by somebody who has stood in the closet. They are reached 75% of the time against 73% for everything else, a difference of +2.0 pt with an interval from -4.7 pt to +7.9 pt. So the data would have hidden a difference of up to about six points either way, and within that, being a guest's own director does nothing. Read it beside the zero on the lineage page: nobody has picked back, and nobody was reaching more than usual in the first place.

Reach rate, films by a guest against everything else
Directed by a guest +2.0 pt
Table view
Reach rate, films by a guest against everything else
By a guest Everything else Difference 95% interval Note
Directed by a guest 75% (205) 73% (1,531) +2.0 pt -4.7 pt to +7.9 pt A film counts as a guest's when one of its credited directors is also a guest.

Every floor this site applies

Sixteen analyses drop rows below a threshold, and each one is defensible on its own: an index over a five-film decade is noise, a shared-choice score over four picks is a coin toss. What none of them was, until this table, is visible. If something you expected is missing from a ranking, it is usually one of these rather than an omission. Every value here is read out of the module that applies it, so the table cannot drift from the code.

Analysis Floor What it means
directors 2 visits A director needs this many reaching visits to be in the ranking.
directors 3 films A director represented by fewer films than this is a shelf fact rather than a taste signal, so they are left out of the ranking but keep a page.
directors 0.3333 share A box set counts towards a director only if they made at least this share of it, which keeps the maker of one supplement out of the ranking.
women 0.9 share Below this share of credited directors carrying a sourced gender, the women's page publishes no share at all. An empty table would otherwise read as a true zero rather than as an absent answer.
women 1 visits A woman director needs this many reaching visits to be in the women's ranking. Deliberately lower than the directors ranking's shelf floor, which would drop a director Criterion stocks one film by.
lenses 3 films Same floor as the directors ranking, so the three counts are drawn over the same directors.
eras 15 films A decade needs this many films before its index is printed: one pick in a five-film decade reads as ten times over.
countries 10 films The same floor for a country, lower because the collection's countries are thinner than its decades.
genres 15 films The same floor for a genre.
languages 15 films The same floor for a language.
citations 3 visits A guest-director needs this many visits reaching their work to appear on the citation leaderboard.
ratings_gap 5 picks A guest needs this many rated picks before their mean rating is ranked.
series_time 20 visits A year needs this many visits before its month shares are published as a seasonal pattern rather than as bare counts.
series_time 25 picks A visit year needs this many single-film picks before it carries quartiles and before it counts towards the age trend.
rarity 5 choices A visit needs this many choices before its mean rarity is ranked.
similarity 5 choices A visit needs this many choices to be compared with another visit.
similarity 2 choices Two visits need this many choices in common to be called close.
similarity 30 pairs A year needs this many visit pairs on each side of the trade comparison to get its own row, though the test itself uses every same-year pair.
cooccurrence 4 visits A pair of items needs this many visits taking both before it is listed.
cooccurrence 4 visits The same floor for a pair of directors, counted after the films they made together are removed from both sides.
cooccurrence 20 visits For the pairs that are rarely together, each director must be reached this often before a low lift is a fact about them and not the sample.
trendsetters 3 visits A first pick needs this many later visits taking the same thing before it counts as followed.
trades 12 visits A trade needs this many solo visits before it gets a row.
trades 25 picks A year needs this many picks from a group before that year can carry the comparison.
spine_gravity 5 picks A guest needs this many single-film picks before their mean spine percentile is ranked.
runtimes 4 films A visit needs this many reached films before its runtime total is ranked.
runtimes 0.8 share A visit's runtime total is published only when this share of its reached films carry a runtime, and it is still a floor rather than a total.
visit_shape 4 films Same floor as the runtime ranking, for the per-choice figure.
spread 8 picks A director needs this many single-film picks before their spread is measured.
spread 3 films And this many films in the collection, so a ceiling exists to normalise against.
reach_screen 20 films A stratified cell needs this many films on each side before its difference is printed.
repeat_visitors 2 visits A person needs this many visits to appear as a returning guest.
checks 3 picks A film needs this many picks before the spread of their dates is measured.
checks 2000 draws How many permutations every test in the checks payload draws.
checks 20260819 seed The seed those permutations use. It is published because the outputs are committed, so a reseeded run would rewrite them for no reason.

Refresh

Criterion's two listings are checked every day, and everything here is rebuilt on the days they have gained or lost something. This build ran 2026-09-25 and read 414 of 414 visits and 1736 films.