Evidence
Measurements taken from saved projects, generated by tooling in this repository so the numbers quoted elsewhere can be traced and regenerated. Each report says what it can and cannot support. None of this is a held-out evaluation.
Regenerating a report
swift run EasycutTextCleanup report ~/Documents/EasycutData/projects/Test23.easycut docs/evidence/test23-rough-cut-2026-09.md
The command copies the project to a temporary directory, opens the copy (which applies any pending schema migration to the copy only), reads the current AI Rough Cut, the Creator Cut derived from it, the transcript rows, the review notes and the edit history, and writes Markdown. It makes no cloud request and does not modify the project. The private project data stays outside the repository.
Test23, September 2026 — report
The one real session with the shipped Rough Cut: 39 clips, 1 h 06 min of source, 607 speech lines. What the numbers show, and what they do not:
- The AI removed 353 of 607 speech lines and kept 227 whole plus 27 in part, bringing 66 minutes to 34 minutes 38 seconds. 106 lines carry a review note; all are validator findings from the request format of that day (the run predates the coded reasons introduced later).
- The creator reversed none of the 353 speech removals in 32 minutes of editing sessions and removed one more kept line. That is consistent with the cut being acceptable to its author, but it is not a precision figure: the project records no full preview and no approval, so it cannot show that every removed line was reviewed, and the creator is also the developer.
- The creator's time went into pacing, not speech. 211 no-speech rows were removed by hand or through Remove enclosed gaps (the AI had removed 105 at analysis time), taking the cut from 34:38 to 29:08 and from 303 to 97 media segments. Manual gap work is the largest visible cost of the session and the clearest lead for a later improvement.
- Session spans, 4½ and 27½ minutes, run from opening the review window to the last saved edit and include playback and reading; the app records no active-review clock, so the brief's 5–15 minute target cannot be measured from this data.
The product brief states the precision and review-time figures as targets for a future held-out evaluation and cites this report as the only development evidence so far.
This page is generated from docs/evidence/README.md in the Telltake source; the document is the text of record.