Why this corpus
The corpus was assembled in two passes, with an explicit rule at each step: cover every mode, and inside each mode, choose filmmakers who genuinely diverge rather than restate the same approach.
The first pass (12 presets) established at least two exemplars per Nichols mode and deliberately extended the corpus past its most obvious American/European anchors. It added Wang Bing, Trinh T. Minh-ha, and Rithy Panh alongside the more familiar Wiseman, Varda, Moore, Herzog, Burns, Morris, Marker, Maysles, and Oppenheimer.
The second pass (4 presets) rebalanced the two thinnest modes. Participatory and Expository each had only one profile. Moore alone standing in for an entire mode reads as "confrontational polemic = participatory," which flattens the mode to one register. Nanfu Wang and Byun Young-joo were added as two structurally distinct participatory registers (embedded investigation under real risk; trust built over more than a year of cohabitation), and Grace Lee and Anand Patwardhan were added to prove expository documentary is not inherently a PBS/BBC-broadcast form. Patwardhan alone represents fifty years of independent, partisan Indian documentary practice.
Throughout, one safeguard was applied in both directions: a filmmaker's stylistic affinity to non-Western narrative frameworks (Layer 7 of each preset) is treated as an empirical property of the actual work, never assumed from a filmmaker's background. Two of the strongest such affinities in the corpus belong to a French essayist (Marker) and a Cambodian testimonial filmmaker (Panh); two American filmmakers (Moore, Burns) carry honest, low, empty values, and so do Nanfu Wang and Grace Lee, both of Asian heritage, where the films themselves don't support the affinity. The corpus tries to demonstrate this rule holds both ways, not just one.
The six Nichols modes
Documentary theorist Bill Nichols proposed six broad "modes": recurring ways a documentary can organize its relationship to its subject and its viewer. Every preset in this corpus is classified into one of them, then the corpus deliberately fills each mode with filmmakers who diverge structurally.
Observational
The camera watches life unfold with minimal intervention: no narration, no to-camera interviews, no staged setups. Often called "fly on the wall" or Direct Cinema. Duration itself becomes an argument.
Wiseman's institutional long take, Maysles's intimate proximity, Wang Bing's durational witnessing of people history has left behind: three different relationships to "not intervening."
"The camera is a fly on the wall. No narration, no music, just observe."
"We don't tell stories. We watch them unfold."
"To film the time of those whom history has left behind, at the speed of their own lives."
Participatory
The filmmaker steps into the frame as an active presence in the events being filmed (asking, provoking, negotiating) rather than staying outside them.
Moore's confrontational provocateur, Nanfu Wang's embodied investigator working inside real danger, Byun Young-joo's presence built through more than a year of trust before the camera appears: three distinct registers, not three variations on one.
"Comfort the afflicted, afflict the comfortable."
"To investigate the systems that shaped me, with my own life offered as part of the evidence."
"Live with them first; the camera is only allowed in once it is no longer a stranger."
Poetic
Mood, association, and visual or aural rhythm take priority over argument or straightforward narrative. Meaning accumulates through form, not exposition.
Varda's participatory-poetic gleaning versus Herzog's ecstatic, stylized truth-through-fabrication.
"Je ne filme pas pour expliquer, je filme pour partager."
"There are deeper strata of truth in cinema, and there is such a thing as poetic, ecstatic truth. It is mysterious and elusive, and can be reached only through fabrication and imagination and stylization."
Expository
The film addresses the viewer directly and builds a case, typically through authoritative narration or a clear rhetorical structure, by marshalling evidence toward a thesis.
Burns's invisible orchestrator, Grace Lee's visible and contested interlocutor, Patwardhan's partisan first-person witness working fifty years outside the PBS/BBC tradition: proof the mode isn't inherently a broadcast-anchored one.
"History is not a fixed thing. It's a conversation between the past and the present."
"Synthesis made in conversation: biography assembled with its subject, not pronounced over them."
"The documentary as an instrument of the secular conscience: evidence gathered in the street, argued in the edit, defended in court."
Reflexive
The film draws attention to its own construction: how documentary meaning gets made, and what authority the filmmaker claims to have in making it.
Morris's investigative first-person cinema, Marker's essayistic memory-work, Trinh's deconstruction of the very authority to "speak about" a subject.
"The pursuit of truth is more interesting than truth itself."
"I will have spent my life trying to understand the function of remembering, which is not the opposite of forgetting, but rather its lining."
"I do not intend to speak about; just speak nearby."
Performative
Subjective, embodied experience (the filmmaker's or the subjects') takes priority over a claim to detached, objective truth.
Oppenheimer's perpetrators reenacting their own crimes, Panh's absence made material through hand-carved figures where the image of atrocity does not exist.
"How do we live with ourselves and the roles we've played? Can cinema be a space for perpetrators to see themselves?"
"There are many things that man should not see or know. Should he see them, he would be better off dying. But if any of us sees or knows these things, then we must live to tell of them."
Barry Salt, cinemetrics, and what this project extends
This system does explicitly what cinemetrics does (measuring style through hard numbers) but widens the lens well past the single variable that field was built around.
Cinemetrics, the tradition most associated with film historian Barry Salt, treats measurable variables (chiefly average shot length) as evidence of a filmmaker's or a period's style, built from frame-by-frame counts across a body of work. This corpus keeps that evidentiary discipline but extends it in two directions: each preset declares a shot-length figure alongside a standard deviation and a confidence score wherever a real measurement exists, and it adds six further layers cinemetrics never attempted to quantify: epistemological stance, sound-design philosophy, ethical boundaries, and narrative-angle distribution.
| Filmmaker | Avg. shot duration | Status |
|---|---|---|
| Frederick Wiseman | 252s | reported (source unconfirmed) |
| Agnès Varda | 18s | reported (source unconfirmed) |
| Ken Burns | 8.5s | reported (source unconfirmed) |
| Michael Moore | 4.2s | reported (source unconfirmed) |
| 12 other presets | — | CALIBRATION_PENDING |
A note on those four figures. The same large-language-model pass reported them as measured, and they are plausible, but the corpus owner has not located the original source (a cinemetrics dataset such as Barry Salt's is one candidate). Until the source is confirmed, the site labels them reported, source unconfirmed rather than measured.
Only four numeric values in the whole corpus are reported with any specificity beyond a qualitative note; the model behind them has not been independently verified, so the table above labels them accordingly rather than as measured fact. The other twelve are honestly marked CALIBRATION_PENDING rather than assigned an invented number, because documentary cinema (especially outside a handful of canonical American and European names) remains underrepresented in the cinemetrics database itself. The extension this corpus attempts is real, but it inherits cinemetrics' central discipline: declare only what was actually measured, and say so plainly when it wasn't. Nothing here should be read as academic rigor equivalent to a real frame-by-frame cinemetrics study. It's a deliberately visual way to place, say, Varda's radar shape next to Moore's and see structurally why an 18-second average and a 4.2-second average come from two different relationships to time, not just two different numbers.
Confidence scores
Each preset carries a single confidence score, generated during the heuristic build and adjusted where the corpus owner flagged a mismatch. It sits alongside the score, never inside the eight axes themselves.
The score is one number per filmmaker, not one per axis. It answers a narrower question than the radar does: how much documented material backs this profile up, not how accurate any single axis reading is. A profile can score high on confidence and still carry a caution field warning against a specific misreading. The two are tracking different things.
Every score in this corpus sits at 0.6 or above. There is no low-confidence tier here. That is a property of the corpus, not a claim that documentary style is easy to model: the sixteen filmmakers chosen all left enough public record, whether through decades of critical writing, extensive interviews, or a small but closely studied filmography, to support a working profile. A filmmaker without that record was left out rather than scored low.
| Filmmaker | Code | Confidence |
|---|---|---|
| Ken Burns | BU | 0.90 |
| Frederick Wiseman | WI | 0.85 |
| Michael Moore | MO | 0.80 |
| Albert & David Maysles | MA | 0.75 |
| Agnès Varda | VA | 0.75 |
| Werner Herzog | HE | 0.75 |
| Errol Morris | MR | 0.75 |
| Wang Bing | WB | 0.70 |
| Anand Patwardhan | AP | 0.70 |
| Joshua Oppenheimer | OP | 0.70 |
| Rithy Panh | RP | 0.70 |
| Nanfu Wang | NW | 0.65 |
| Chris Marker | MK | 0.65 |
| Trinh T. Minh-ha | TR | 0.65 |
| Byun Young-joo | BY | 0.60 |
| Grace Lee | GL | 0.60 |
The pattern in that table is not random. The filmmakers with a broadcast footprint, measured shot durations, and a large body of critical writing (Burns, Wiseman, Moore) land at the top. The ones whose style depends on a smaller filmography, calibration-pending shot data, or a practice that resists conventional measurement (Byun, Grace Lee, Wang Bing, Nanfu Wang) sit lower, without dropping into territory that would call the profile unusable. That correlation is what a working heuristic should produce: a model that tracks how much evidence it had, rather than a flat score handed out the same way to every entry.
Each score carries a status field, ADOPTED_HEURISTIC across the whole corpus. That label means the implementer proposed the value from the documented record and the corpus owner signed off on it. It is not a statistical confidence interval and it was not derived from inter-rater agreement. Read it the same way you would read the shot-duration line: an honest marker of how the number was arrived at, not a claim of precision it cannot support.
Filmmaker portraits
Grouped by Nichols mode. Manifesto lines are editorial paraphrases of each filmmaker's documented stance, not verbatim quotations, except where drawn from a filmmaker's own published writing.
Scope and limitations
Read this before reading anything else on this site as a claim about real people.
This project offers a directional preset system for documentary narration and editing, not an exact scientific model of auteur identity. Each preset translates familiar cinematic shorthand ("Varda-like," "Wiseman-like," "Moore-like") into a structured set of narrative tendencies, with a confidence score attached to show how well-supported each part of the profile is.
Think of it as a structural sandbox: a way of testing how differently organized minds might handle the same material. Each profile encodes recurring structural preferences, epistemological stances, temporal signatures, sound-design philosophies, narrative-angle distributions, and ethical boundaries: an interpretive layer for creative planning, not a claim to have reconstructed a filmmaker's full authorship.
What it's for
The scope is deliberately limited to supporting narrative planning, segmentation, and stylistic direction, including hybrid sequences, like an intimate Varda-style opening followed by a systemic Moore-style middle section. It doesn't attempt to capture everything about how a filmmaker works, and makes no claim to explain genius, historical context, or the full complexity of an artistic career.
Why confidence scores matter
The confidence framework is central to that honesty. High-confidence attributes reflect patterns that are relatively easy to observe across a body of work; lower-confidence attributes mark interpretive or underdetermined judgment calls, and should be read as provisional. That makes the system more transparent and better suited to assisted creative exploration than to strict scholarly classification.
The per-filmmaker breakdown, and what the distribution across the corpus suggests, sits in its own section above.
A note on corpus versions 2.1.0 and 2.2.0
Readers who bookmarked specific Proximity Matrix percentages before July 2026 may notice shifts of a few points for pairs involving Errol Morris, Agnès Varda, and Werner Herzog: three axis scores were corrected at corpus v2.1.0 after a review of the underlying evidence (Morris's ambiguity tolerance, Varda's poetic angle, and Herzog's reflexivity level). At v2.2.0, every preset's confidence score moved from a provisional review status to an adopted working figure, still a modeled estimate, not an externally audited one. The similarity formula, the eight axes themselves, and the matrix's colour-band thresholds did not change at either revision.
Explicit limitations
A preset can compress a living, contradictory, evolving practice into a stable label, which flattens nuance by construction. The system is not neutral: it inherits real theoretical choices about how documentary form gets divided into parameters, modes, and ethical stances. And it is likely stronger as operational guidance for creative direction than as an explanation of why any filmmaker works the way they do.
Where the numbers come from. The eight axis values for all 16 presets were generated by an advanced large language model, which read each filmmaker's body of work and critical reception and estimated the values. They were not measured, coded by a human panel, or taken from an existing dataset. The corpus owner reviewed and adopted them (see the v2.1.0 / v2.2.0 note above), but "adopted" here means accepted as a reasonable working estimate, without independent verification. Average shot duration is a partial exception: four values are reported with a specific figure (see the table above), though their original source has not been confirmed, and the rest carry CALIBRATION_PENDING rather than a filled-in guess. The eight axes are also a deliberate flattening of a larger, private KAITSA schema, which uses more layers than this site shows. The reduction is for teaching clarity; the framework was not invented for this site.
In short
Read this as a practical, modular reference for assisted documentary design, not a definitive theory of documentary authorship. Its value is in making creative direction more legible and more honest about its own uncertainty, not in settling questions about authorship this project was never built to settle.
From structural claim to filmic evidence
The model becomes more credible when structural observations can be checked against concrete passages from films.