Inside Runs · exploratory research · closed
What we learned building a first-cut engine.
We wanted to know whether finished edits could help turn new footage into a stronger, more familiar starting point.
So we reconstructed real editorial decisions, compared creator profiles and ran a blind first-cut experiment. The results changed the question Runs is built to answer.
Ayu Koene ·
Where it started
The first version was built while editing a film about my mum.
Before moving abroad, I filmed conversations with my mum, Luz, about the paintings I had grown up with. They became a reel series and a documentary: a daughter discovering the artist she thought she already knew.
I cut the first episode by hand. After that, AY-CUT v1 created the series intro and outro, cut every remaining episode and assembled the YouTube documentary—coming soon.
That handoff became the starting point for Runs: carry the repeatable first-cut work forward while leaving the story and final judgment with the editor.
What actually happened
Finished work became context for the next cut.
This is the real lineage behind the research—not a synthetic benchmark assembled for a demo.
- My earlier manual edits, in CapCut and Edits
- AY-CUT measures the patterns in them
- New interview footage
- AY-CUT proposes the edits
- I review and correct them
- Published videos
- Later, the research reconstructs the choices
The evidence base we were actually working with
15
finished edits examined
426
editing decisions reconstructed
1758
real alternatives compared
1
creator in the core corpus
15 fold units from one creator across four bodies of work: two Luz interview sessions, a DJI Mic reel, and two hand-cut personal reels. Three edits are pure hand cuts; twelve are agent/engine proposals with light corrections. Fifteen folds are not fifteen independent demonstrations of taste. The figures below describe this corpus; they are not product performance claims.
The first hypothesis
Editing style leaves a trace.
If earlier edits are useful context, the same footage should not produce the same rhythm for everyone. We ran identical footage and the same voiceover through two real creator profiles. Only the profile changed.
Ayu @ayukoene
5.4cuts / 10s
21 shots · median 1.92s
A real second creator @arctic__fever
3.8cuts / 10s
15 shots · median 2.36s
The second profile was learned from 12 of that creator's published reels. Two profiles over the same footage produce measurably different cuts, and these figures come from the generated cut lists, which is the sound way to measure pacing. What it does not show is that either output lands inside its creator's own distribution — that is the replication test, and it ships today only on loudness and caption geometry.
The finding that changed the question
Taste is not an answer key.
Then we tested the assumption underneath the whole programme. We showed the creator 12 of her own hardest past decisions again, blind—her final choice mixed with the options she had passed over, unlabelled and shuffled.
She chose the exact same clip again 4 of 12 times. Yet the historical choice was still acceptable in 11 of 12, and she accepted 71.7% of all the options overall.
The important finding was not inconsistency. It was plurality: several choices can genuinely work, and the one you reach for can change with the day. Exact historical agreement was the wrong definition of a good draft.
Blind re-judgment · 2026-07-27
4 / 12
of her own final-cut choices the creator re-picked exactly, re-judging the 12 hardest decisions blind from shuffled raw previews.
11 / 12 of those historical choices she still rated acceptable.
Solid: acceptable on re-judgment · outlined: rejected. On average 2.75 of 3.83 options per case were acceptable.
46 options across 12 cases, mean difficulty 0.923: 12 rated best, 21 fine, 13 wrong. Exact self-agreement 33.3% (CI 13.8%–60.9%) against 27.2% chance — on these cases a single right answer barely exists, which makes exact-match accuracy a floor. What does exist is the 28.3% of options she rejects. One judge, one sitting — snapshot at experiments/2026-07-27-recut-probe.
One creator, one sitting, twelve decisions—a probe, not a study.
MethodologyThe real first-cut experiment
More context did not magically make the cut better.
We indexed 11 raw clips into 264 addressable moments and asked a model to cut a real reel 6 times. It saw the brief alone, then the creator’s finished work, then her past corrections.
All 6 plans were valid on the first attempt: real footage, in-bounds timings, nothing invented. That proved the production plumbing. It did not prove that adding references made the drafts better.
- Brief only
- References cited
- 0
- “I’d post this”
- 0 of 2
- Plus her finished work
- References cited
- 5
- “I’d post this”
- 1 of 2
- Plus her past corrections
- References cited
- 19
- “I’d post this”
- 0 of 2
What actually decided the verdict
Five of six change-notes named the same thing — and it was in none of the payloads. Asked to scope it, she said it was not an editing rule at all: it was that footage, that morning, how she felt she looked in it.
The engine produced six valid, editable cuts and never invented footage. It could not show whether her finished work improves a draft: every cut broke the same unstated, project-specific preference, so five of six verdicts landed on the same value. That question is open, not answered.
Two runs of the same condition shared only 13.0%–23.0% of their chosen moments. A first cut is one plausible answer, not the answer.
The strategic turn
The research changed the product, not just the score.
The old question
Can Runs predict every invisible choice the creator would make?
The better question
Is starting from a Runs draft substantially better than starting from an empty timeline?
One strong draft
Runs gives you a credible direction to continue, not a finished answer you must accept or reject.
A real timeline
Editing is half the product, not the fallback after the AI. Final judgment stays visible and reversible.
Learning from the finish
What survives from draft to export is a better signal than trying to make every preference explicit first.
What carries forward
Keep the useful machinery. Retire the mythology.
What survived
- Ground every decision in footage that really exists.
- Use finished work as context, not as an answer key.
- Keep every first cut editable.
- Let corrections become context for the next project.
What was retired
- The promise of reproducing every invisible choice.
- The early hand-cut advantage, which vanished with more examples.
- The dramatic final-hold claim, which was really a title card.
- The idea that a theory of taste must come before a useful product.
Known corrections to the figures on this page
- The scope-ablation flattened figure (0.3142) was measured on the 401-case Luz-only corpus and has not been re-run on the 426-case corpus; scoped rates track the current D_dossier by-kind scores.
- An earlier draft labelled the DJI Mic reel human_hand_cut. AYCUT_LESSONS.md records it as an early agent/engine cut with light creator adjustment, so the corpus label stays machine_proposed_human_corrected — but HARDENING_AND_ARCHIVE_RESULTS.md §2.3 still calls it a creator-endorsed hand cut. One of the two is wrong; the split here follows the corpus.
- The hand-cut advantage (0.610 vs 0.412) was retired on 2026-07-27 after two more hand-cut edits took independent authorship to 3 of 15 and the gap collapsed to +0.020.
- The re-decision probe is twelve cases and one sitting. It is the newest and least-replicated measurement here, and also the one that most changes how the older numbers should be read.
- Pacing in the divergence exhibit is read off generated cut lists, which is sound. Pacing measured from finished renders appears nowhere on this page — the scene detector miscounts editorial cuts. See the craft signal audit.
- Programme closed 2026-07-28. Experiment 001 (references vs brief-only, 6 cuts) validated the harness — all plans structurally valid, no hallucinated footage — but could not separate its arms: every cut broke the same unstated, project-specific preference. It did NOT establish that references improve drafts at scale, nor that taste is mostly negative constraints.
The honest close
What this story does not prove.
The target moves
Asked to re-decide twelve of her own hard choices blind, the creator picked the clip she originally kept 4 times out of 12. Every accuracy figure on this page is agreement with one recording of a choice the creator does not herself reproduce.
One creator
Every selection number here comes from one person's body of work. 15 held-out tests are not 15 independent demonstrations of taste, and the craft side has two creators only on the signals that survived the audit.
The corpus is not reproducible
Rebuilding the same ledgers with the same toolchain changed 131 of 417 cases through transcription drift alone. The chosen answers held; the language features every condition reads did not — and the drift is the size of the effect.
No real generic-model baseline yet
The “no editing profile” approach is a simple rule, not a capable language model. Given no context it falls back to the same longest-first rule as the approach above it, which is why the two score almost identically. Runs has not yet been compared against a strong generic AI.
Whether references improve drafts at scale
The engine produced six valid, editable cuts and never invented footage. It could not show whether her finished work improves a draft: every cut broke the same unstated, project-specific preference, so five of six verdicts landed on the same value. That question is open, not answered. The effect remains open.
The research is closed. The product now earns its evidence in use.
Runs makes the first cut, you finish it in the timeline, and the difference between the two becomes the next useful piece of context. The measure is no longer whether software can read a creator’s mind. It is whether the draft saves enough time and preserves enough intent to be worth continuing.
The editor remains available through the closed test.