Inside Runs · exploratory research · closed

What we learned building a first-cut engine.

We wanted to know whether finished edits could help turn new footage into a stronger, more familiar starting point.

So we reconstructed real editorial decisions, compared creator profiles and ran a blind first-cut experiment. The results changed the question Runs is built to answer.

Ayu Koene ·

Where it started

The first version was built while editing a film about my mum.

Before moving abroad, I filmed conversations with my mum, Luz, about the paintings I had grown up with. They became a reel series and a documentary: a daughter discovering the artist she thought she already knew.

I cut the first episode by hand. After that, AY-CUT v1 created the series intro and outro, cut every remaining episode and assembled the YouTube documentary—coming soon.

That handoff became the starting point for Runs: carry the repeatable first-cut work forward while leaving the story and final judgment with the editor.

What actually happened

Finished work became context for the next cut.

This is the real lineage behind the research—not a synthetic benchmark assembled for a demo.

  1. My earlier manual edits, in CapCut and Edits
  2. AY-CUT measures the patterns in them
  3. New interview footage
  4. AY-CUT proposes the edits
  5. I review and correct them
  6. Published videos
  7. Later, the research reconstructs the choices

The evidence base we were actually working with

15

finished edits examined

426

editing decisions reconstructed

1758

real alternatives compared

1

creator in the core corpus

15 fold units from one creator across four bodies of work: two Luz interview sessions, a DJI Mic reel, and two hand-cut personal reels. Three edits are pure hand cuts; twelve are agent/engine proposals with light corrections. Fifteen folds are not fifteen independent demonstrations of taste. The figures below describe this corpus; they are not product performance claims.

01

The first hypothesis

Editing style leaves a trace.

If earlier edits are useful context, the same footage should not produce the same rhythm for everyone. We ran identical footage and the same voiceover through two real creator profiles. Only the profile changed.

Ayu @ayukoene

5.4cuts / 10s

21 shots · median 1.92s

A real second creator @arctic__fever

3.8cuts / 10s

15 shots · median 2.36s

The second profile was learned from 12 of that creator's published reels. Two profiles over the same footage produce measurably different cuts, and these figures come from the generated cut lists, which is the sound way to measure pacing. What it does not show is that either output lands inside its creator's own distribution — that is the replication test, and it ships today only on loudness and caption geometry.

02

The finding that changed the question

Taste is not an answer key.

Then we tested the assumption underneath the whole programme. We showed the creator 12 of her own hardest past decisions again, blind—her final choice mixed with the options she had passed over, unlabelled and shuffled.

She chose the exact same clip again 4 of 12 times. Yet the historical choice was still acceptable in 11 of 12, and she accepted 71.7% of all the options overall.

The important finding was not inconsistency. It was plurality: several choices can genuinely work, and the one you reach for can change with the day. Exact historical agreement was the wrong definition of a good draft.

Fig. 1Blind re-judgment of 12 hard decisions: the same choice 4 times, the original still acceptable in 11.

Blind re-judgment · 2026-07-27

On the hardest calls, most options are acceptable — even to the creator

4 / 12

of her own final-cut choices the creator re-picked exactly, re-judging the 12 hardest decisions blind from shuffled raw previews.

11 / 12 of those historical choices she still rated acceptable.

look what the cat brought home today3/4
look what the cat brought home today1/4
the teacher who believed in me3/3
the teacher who believed in me3/4
do you discover while painting3/5
do you discover while painting4/5
a flood made me see my old paintings differently2/3
paint angry paint happy4/4
when is it easy to let a painting go1/3
the beauty of something frail3/3
my moms hidden art collection v102/3
not for sale rough v14/5

Solid: acceptable on re-judgment · outlined: rejected. On average 2.75 of 3.83 options per case were acceptable.

46 options across 12 cases, mean difficulty 0.923: 12 rated best, 21 fine, 13 wrong. Exact self-agreement 33.3% (CI 13.8%60.9%) against 27.2% chance — on these cases a single right answer barely exists, which makes exact-match accuracy a floor. What does exist is the 28.3% of options she rejects. One judge, one sitting — snapshot at experiments/2026-07-27-recut-probe.

One creator, one sitting, twelve decisions—a probe, not a study.

Methodology
03

The real first-cut experiment

More context did not magically make the cut better.

We indexed 11 raw clips into 264 addressable moments and asked a model to cut a real reel 6 times. It saw the brief alone, then the creator’s finished work, then her past corrections.

All 6 plans were valid on the first attempt: real footage, in-bounds timings, nothing invented. That proved the production plumbing. It did not prove that adding references made the drafts better.

Brief only
References cited
0
“I’d post this”
0 of 2
Plus her finished work
References cited
5
“I’d post this”
1 of 2
Plus her past corrections
References cited
19
“I’d post this”
0 of 2

What actually decided the verdict

Five of six change-notes named the same thing — and it was in none of the payloads. Asked to scope it, she said it was not an editing rule at all: it was that footage, that morning, how she felt she looked in it.

The engine produced six valid, editable cuts and never invented footage. It could not show whether her finished work improves a draft: every cut broke the same unstated, project-specific preference, so five of six verdicts landed on the same value. That question is open, not answered.

Two runs of the same condition shared only 13.0%23.0% of their chosen moments. A first cut is one plausible answer, not the answer.

The strategic turn

The research changed the product, not just the score.

The old question

Can Runs predict every invisible choice the creator would make?

The better question

Is starting from a Runs draft substantially better than starting from an empty timeline?

One strong draft

Runs gives you a credible direction to continue, not a finished answer you must accept or reject.

A real timeline

Editing is half the product, not the fallback after the AI. Final judgment stays visible and reversible.

Learning from the finish

What survives from draft to export is a better signal than trying to make every preference explicit first.

What carries forward

Keep the useful machinery. Retire the mythology.

What survived

  • Ground every decision in footage that really exists.
  • Use finished work as context, not as an answer key.
  • Keep every first cut editable.
  • Let corrections become context for the next project.

What was retired

  • The promise of reproducing every invisible choice.
  • The early hand-cut advantage, which vanished with more examples.
  • The dramatic final-hold claim, which was really a title card.
  • The idea that a theory of taste must come before a useful product.
Known corrections to the figures on this page
  • The scope-ablation flattened figure (0.3142) was measured on the 401-case Luz-only corpus and has not been re-run on the 426-case corpus; scoped rates track the current D_dossier by-kind scores.
  • An earlier draft labelled the DJI Mic reel human_hand_cut. AYCUT_LESSONS.md records it as an early agent/engine cut with light creator adjustment, so the corpus label stays machine_proposed_human_corrected — but HARDENING_AND_ARCHIVE_RESULTS.md §2.3 still calls it a creator-endorsed hand cut. One of the two is wrong; the split here follows the corpus.
  • The hand-cut advantage (0.610 vs 0.412) was retired on 2026-07-27 after two more hand-cut edits took independent authorship to 3 of 15 and the gap collapsed to +0.020.
  • The re-decision probe is twelve cases and one sitting. It is the newest and least-replicated measurement here, and also the one that most changes how the older numbers should be read.
  • Pacing in the divergence exhibit is read off generated cut lists, which is sound. Pacing measured from finished renders appears nowhere on this page — the scene detector miscounts editorial cuts. See the craft signal audit.
  • Programme closed 2026-07-28. Experiment 001 (references vs brief-only, 6 cuts) validated the harness — all plans structurally valid, no hallucinated footage — but could not separate its arms: every cut broke the same unstated, project-specific preference. It did NOT establish that references improve drafts at scale, nor that taste is mostly negative constraints.

The honest close

What this story does not prove.

The target moves

Asked to re-decide twelve of her own hard choices blind, the creator picked the clip she originally kept 4 times out of 12. Every accuracy figure on this page is agreement with one recording of a choice the creator does not herself reproduce.

One creator

Every selection number here comes from one person's body of work. 15 held-out tests are not 15 independent demonstrations of taste, and the craft side has two creators only on the signals that survived the audit.

The corpus is not reproducible

Rebuilding the same ledgers with the same toolchain changed 131 of 417 cases through transcription drift alone. The chosen answers held; the language features every condition reads did not — and the drift is the size of the effect.

No real generic-model baseline yet

The “no editing profile” approach is a simple rule, not a capable language model. Given no context it falls back to the same longest-first rule as the approach above it, which is why the two score almost identically. Runs has not yet been compared against a strong generic AI.

Whether references improve drafts at scale

The engine produced six valid, editable cuts and never invented footage. It could not show whether her finished work improves a draft: every cut broke the same unstated, project-specific preference, so five of six verdicts landed on the same value. That question is open, not answered. The effect remains open.

The research is closed. The product now earns its evidence in use.

Runs makes the first cut, you finish it in the timeline, and the difference between the two becomes the next useful piece of context. The measure is no longer whether software can read a creator’s mind. It is whether the draft saves enough time and preserves enough intent to be worth continuing.

The editor remains available through the closed test.