Filmmaking is a sequence of judgments.

SF Locale studies how human decisions can guide video systems across shots, scenes and complete films.

Trajectory
5
feature (70+ min) films completed
1M+
generations recorded
4K
director annotation hours

The production process creates a dataset of intent, alternatives, choices, revisions and consequences.

Scene 4 production map transitions to a complete keyframe trajectory and then to a candidate reward and prompt inspector.
Trajectory walkthrough. A scene map opens into a single keyframe route, then reveals the prompt and multi-head audit for one candidate.

The research begins inside production.

SF Locale began using generative models in January 2026. The studio has completed five feature-length films since then.

5feature-length films have been completed.
1M+image and video generations have been recorded.
50K+revision chains have been preserved.
4,000hours of director annotation have been collected.
100Kshots are represented in the archive.

Each record can connect a shot's intent, prompt, references, candidate family, director choice, revision reason and final position in sequence.

Scene 4 production map with 35 keyframes, candidate counts and generated-reference lineage.
Production map. The Scene 4 archive exposes 35 keyframe families, their candidate counts and cross-keyframe reference reuse.

The trajectory is the learning object.

A finished shot contains only the result. The production trajectory preserves the alternatives that were considered, the actions that produced them and the judgment that moved the film forward.

The archive is being reconstructed as a branching decision graph. A route can be locked, revised, retained for later use or terminated. A selected image can also become a reference for a later shot.

Figure 1. The graph is event-sourced. New decisions extend the record without erasing earlier states.
τₜ = (Xₜ, Aₜ, Oₜ, Hₜ, Fₜ, Xₜ₊₁, Gₜ, Cₜ, Yₜ)
Trajectory record. X is the production state, A the generation action, O the candidate family, H the human decision, F the feedback, G the lineage, C the active canon and Y the later outcome.
A complete keyframe trajectory with three action families and fifteen candidate images.
One route in full. Each action preserves its references, candidate family and the decision state it left unresolved.

The model needs the state of the film.

A prompt does not describe everything that makes a candidate useful. The same image can be correct in one sequence and unusable in another. A useful representation needs the context that shaped the decision.

Active canon
The committed facts of the film, including character identity, location, wardrobe, props, time and visual rules.
Lineage
The parent candidates, references, branches and workflow operations that produced the current state.
Human history
The locks, shortlists, rejections, revisions, retained branches and prior reasons around the shot.
Shot objective
The narrative purpose, desired performance, framing, movement and relation to surrounding shots.
Production pressure
The time, compute, budget and delivery constraints active when the decision was made.

The first experiment tests whether context improves prediction.

The first benchmark can compare a simple critic that sees a candidate and prompt with a contextual critic that also sees canon, references, lineage and earlier decisions. Both models predict the filmmaker's recorded choice.

Figure 2. Chronological and cross-film splits reduce leakage from sibling candidates, repeated assets and future decisions.

Production context and trajectory history should improve the prediction of filmmaker choices, repairs and stopping decisions.

The evaluation can measure pairwise accuracy, top-k lock recall, ranking quality, calibration, abstention, next-action accuracy and repair-cost prediction. Performance should be reported over time and across held-out films.

Human judgment has several dimensions.

A lock means that a candidate was selected for a specific role in a specific production state. It does not mean that the candidate is universally good. An unselected candidate may become useful later, and an unseen candidate provides no preference evidence.

A critic can keep separate estimates for intent, reference fidelity, canon consistency, performance, cinematography, technical quality, narrative function and production utility. The weight of each estimate can change with the production state.

Q(Sₜ, s) = immediate fit + future utility − expected repair cost
Value hypothesis. The value of a candidate depends on its immediate role, its later usefulness and the repairs it may create.

Research note. The archive does not yet establish a calibrated value function or a causal measure of future repair cost.

A selected counterfactual candidate alongside its reward vector, prompt and action references.
Candidate audit. The inspector makes the current heuristic transparent: its prompt, references, individual reward heads and evidence limitations remain visible.

The archive supports several learning problems.

Selection

Rank candidates for the current shot and explain which constraints drive the ranking.

Orchestration

Predict the next production action, including revision, model choice, workflow operation and branch change.

Reference policy

Choose which prior images, characters, locations and style references should condition the next action.

Stopping

Estimate when a route is ready to lock and when further generation has low expected value.

Continuity

Predict when a local choice will create later repairs or reduce creative options elsewhere in the film.

Scene policy

Coordinate shot scale, composition, movement, duration and order across a complete scene.

The research proceeds from measurement to assistance.

The current program treats existing generators as components of the production environment. The immediate work is to make the trajectory auditable and establish reliable baselines.

ReconstructRecover candidate sets, lineage, exposure, choices, feedback and delayed outcomes.
MeasureAudit missingness, censoring, branch termination, reference reuse and disagreement.
LearnTrain state-conditioned preference models on chronological held-out data.
PredictTest next actions, stopping decisions, likely failures and future repairs.
AssistEvaluate decision support with visible uncertainty and human override.
PilotRun a limited human-supervised production study after offline evaluation.

Current status. The archive and system definitions exist. Model training and validation remain open research work. The page makes no claim that a production reward model, orchestration policy or debt model has been trained.

The studio funds research through film production.

SF Locale is a studio of filmmakers and engineers from IIT Bombay and IIT BHU. The studio makes films for clients and for its own projects. Production funds the research and continues to grow the trajectory archive.