Documentary Style DNAA Filmmaker Preset Explorer
ObservationalParticipatoryPoeticExpositoryReflexivePerformative

Reference

Data & Confidence

Where the corpus's numbers come from: the cinemetrics tradition this project extends, the four shot-length figures that carry an actual measurement, and the per-filmmaker confidence score that says how much documented material backs each profile.

Part of the Methodology. Read that page first for why the corpus exists and what the six Nichols modes mean.

Barry Salt, cinemetrics, and what this project extends

This system does explicitly what cinemetrics does (measuring style through hard numbers) but widens the lens well past the single variable that field was built around.

Cinemetrics, the tradition most associated with film historian Barry Salt, treats measurable variables (chiefly average shot length) as evidence of a filmmaker's or a period's style, built from frame-by-frame counts across a body of work. This corpus keeps that evidentiary discipline but extends it in two directions: each preset declares a shot-length figure alongside a standard deviation and a confidence score wherever a real measurement exists, and it adds six further layers cinemetrics never attempted to quantify: epistemological stance, sound-design philosophy, ethical boundaries, and narrative-angle distribution.

FilmmakerAvg. shot durationStatus
Frederick Wiseman252sreported (source unconfirmed)
Agnès Varda18sreported (source unconfirmed)
Ken Burns8.5sreported (source unconfirmed)
Michael Moore4.2sreported (source unconfirmed)
12 other presetsCALIBRATION_PENDING

A note on those four figures. The same large-language-model pass reported them as measured, and they are plausible, but the corpus owner has not located the original source (a cinemetrics dataset such as Barry Salt's is one candidate). Until the source is confirmed, the site labels them reported, source unconfirmed rather than measured.

Only four numeric values in the whole corpus are reported with any specificity beyond a qualitative note; the model behind them has not been independently verified, so the table above labels them accordingly rather than as measured fact. The other twelve are honestly marked CALIBRATION_PENDING rather than assigned an invented number, because documentary cinema (especially outside a handful of canonical American and European names) remains underrepresented in the cinemetrics database itself. The extension this corpus attempts is real, but it inherits cinemetrics' central discipline: declare only what was actually measured, and say so plainly when it wasn't. Nothing here should be read as academic rigor equivalent to a real frame-by-frame cinemetrics study. It's a deliberately visual way to place, say, Varda's radar shape next to Moore's and see structurally why an 18-second average and a 4.2-second average come from two different relationships to time, not just two different numbers.

Confidence scores

Each preset carries a single confidence score, generated during the heuristic build and adjusted where the corpus owner flagged a mismatch. It sits alongside the score, never inside the eight axes themselves.

The score is one number per filmmaker, not one per axis. It answers a narrower question than the radar does: how much documented material backs this profile up, not how accurate any single axis reading is. A profile can score high on confidence and still carry a caution field warning against a specific misreading. The two are tracking different things.

Every score in this corpus sits at 0.6 or above. There is no low-confidence tier here. That is a property of the corpus, not a claim that documentary style is easy to model: the sixteen filmmakers chosen all left enough public record, whether through decades of critical writing, extensive interviews, or a small but closely studied filmography, to support a working profile. A filmmaker without that record was left out rather than scored low.

FilmmakerCodeConfidence
Ken BurnsBU0.90
Frederick WisemanWI0.85
Michael MooreMO0.80
Albert & David MayslesMA0.75
Agnès VardaVA0.75
Werner HerzogHE0.75
Errol MorrisMR0.75
Wang BingWB0.70
Anand PatwardhanAP0.70
Joshua OppenheimerOP0.70
Rithy PanhRP0.70
Nanfu WangNW0.65
Chris MarkerMK0.65
Trinh T. Minh-haTR0.65
Byun Young-jooBY0.60
Grace LeeGL0.60

The pattern in that table is not random. The filmmakers with a broadcast footprint, measured shot durations, and a large body of critical writing (Burns, Wiseman, Moore) land at the top. The ones whose style depends on a smaller filmography, calibration-pending shot data, or a practice that resists conventional measurement (Byun, Grace Lee, Wang Bing, Nanfu Wang) sit lower, without dropping into territory that would call the profile unusable. That correlation is what a working heuristic should produce: a model that tracks how much evidence it had, rather than a flat score handed out the same way to every entry.

Each score carries a status field, ADOPTED_HEURISTIC across the whole corpus. That label means the implementer proposed the value from the documented record and the corpus owner signed off on it. It is not a statistical confidence interval and it was not derived from inter-rater agreement. Read it the same way you would read the shot-duration line above: an honest marker of how the number was arrived at, not a claim of precision it cannot support.

Confidence scores describe how much the corpus had to work with for a given filmmaker, not how correct any single axis value is. A high score and a caution field can sit on the same profile. Read the full scope & limitations →