An agent-driven video editor. Describe the cut in plain English, an AI writes the JSON shot list, FFmpeg renders it in both 16:9 and 9:16 — no timeline, no mouse, no editor.
Every shot is an in/out timecode in a JSON file. Change one number, re-render. Audio levels and scene-change detection tell you where the dead air is.
sheet turns a recording into a grid of timecoded frames an AI can actually read — so the model picks the cuts, then sets pan and zoom per shot in the shot list.
Whisper transcribes your recording to timecoded SRT. Narration is synthesised locally from 25+ natural voices — nothing uploaded, nothing billed per word.
Burns titles and captions straight into the render, and outputs 16:9 and 9:16 from the same shot list — one source, every platform.
Record once. Hand the footage to an agent. Get a finished cut back.