Video editing
I asked an AI agent to edit my video. Here are all three versions.
The first cut had low captions, repetitive graphics, and a long pause at the end. These are the actual videos and the instructions I gave Codex to improve them with Tesseract.
Aaron Makelky · · 10 min read + 3 short videos
Mirage sponsored the video shown here and provided early access to Tesseract.
1. Give the agent the job, the footage, and the brand
I recorded a short video asking whether AI could edit as well as I can. I wanted the raw take and the edited version beside each other so people could judge the result. I also gave some instructions out loud while recording, including a diagram, a zoom on my coffee cup, and quiet retro music.
I was working in Codex with Tesseract by Mirage, which lets an agent build and render an editable video project locally. Linear supplied the task brief, and GitHub supplied my video brand files.
There were three different video ideas on the card. I had to tell the agent which one we were editing. Here is my first request, with the original spelling and app mentions preserved.
Eight clips were copied from the card and checked against the originals. The selected complete take ran 57.375 seconds. The other concepts stayed separate. I then pointed the agent at my existing video assets instead of describing a new look from scratch.
The videos use my navy, blue, light blue, white, and teal palette, with Syne and DM Sans. The font files mattered as much as the color names. The first font import needed a compatibility fix before the rendered text actually used the intended letterforms.
Watch the original take and read what I said on camera
So, can AI really edit a video as well as me? I'm going to show you side by side the raw version uncut, and then the version I'm going to edit with this new tool called Tesseract. This video is sponsored by Mirage. Thank you for giving me early access. I'm going to have you put on the screen right now a little diagram that explains how I go from camera footage to ChatGPT Work to edit with Tesseract, and it doesn't have to open a standalone video editing tool. A couple other rules: when I point at something like my name on my coffee cup right there, I need you to zoom in and put a gentle sound effect. The captions and on-screen graphics don't cover my face. And then make sure the background music is lo-fi and kind of retro, but it's not loud enough to distract from the narrative. So, how did Tesseract do? Can this edit a video as well as I can?
2. Watch the first cut before fixing it
The first version gave me a playable comparison with captions and graphics. It also left the things I would immediately change in an edit: a caption rectangle sitting too low, explainers that felt static, and too much footage after I finished talking.
Watch the opening captions, then jump to the coffee and the ending. Each player shows the original on the left and that version's edit on the right. You hear one edited mix.
The file passed its technical checks, but I still didn't like the treatment. I described the changes in terms of things I could see and hear.
That gave the agent specific work: move the captions onto my torso, remove their background, animate the explainers, make the coffee cue noticeable, and end the clip after the final thought. Asking for a more engaging edit had left too many of those decisions unspecified.
3. Move the captions and give the graphics a purpose
Version 2 moved the words up onto my torso and removed the caption boxes. Words appeared at their spoken timing, while short phrases stayed together on screen. The graphic elements gained entrances and exits, with the camera, ChatGPT Work, and Tesseract explanation appearing in sequence.
The coffee zoom got a stronger warm sound cue, and the music dipped around it. The ending lost 2.625 seconds of recording-stop footage. A closing push-in carried the final question, and the clip stopped at 54.75 seconds.
I said, "that version works for now." Seeing the result at the final comparison size helped. A caption can look fine in a full-height portrait preview and become too small when the video shares the frame with the original.
I kept the complete take in the editable project. Both sides of the revised comparison use the same source timing, and both end together. Trimming the end of the delivery did not delete the original recording.
4. Give the captions more range and stop repeating the ding
The next request was smaller. I wanted stronger captions and more variety in the sound effects.
Version 3 kept the same 167 verified words and their timing. It regrouped them into 42 natural phrases, down from 70 shorter groups in v2. Supporting words made a small entrance. Selected words appeared larger, rose a little, and settled once. Five blue underlines drew attention to particular phrases.
The captions also stayed white while the explainer used teal. I could read the words without having two competing teal elements asking for attention.
The nine sound cues now have different jobs. The camera graphic gets a double-click, the chat action gets short typing and send ticks, and the Tesseract handoff gets a snap. Air sounds handle transitions. The coffee moment keeps its warm accent, and the ending gets a gentle impact. The words themselves do not each make a noise.
My response was "banger, well done." I asked for v3 to be uploaded and its location recorded on the task.
Where the audio actually came from
I didn't need to supply a sound folder for these effects. Tesseract's local helper can make simple procedural taps, clicks, pops, pings, whooshes, risers, downlifters, and impacts. Combining those sparingly produced the different accents.
The music was a separate job. This edit used an instrumental bed synthesized locally with Python and imported into Tesseract. Calling that Tesseract music generation would be inaccurate. If I wanted a particular licensed track or recorded sound, I would supply it or ask the agent to use another authorized source.
The captions also depended on local transcription outside Tesseract and a verified word map. Tesseract handled the editable text and animation; it did not transcribe the recording for us.
5. Keep my eyes level when the camera moves
After approving v3, I added another instruction. Long stretches of narration need subtle camera movement, especially when there is no graphic on screen or when a framing change can help cover a cut. During a zoom, my eyes should stay at the same height in the frame.
That became a rule in /video-edit-tesseract. A typical narration move starts with a slow 2 to 5 percent push or release over 3 to 6 seconds. The agent pairs the zoom with a position adjustment so the camera movement does not pull the eyes toward the top of the frame. It checks the position after cropping, including inside the smaller comparison panel.
This broader narration movement has not been rendered into v3. The cup close-up and closing push-in were already there. The new rule is saved for the next edit, with a calculator and instructions to inspect the start, middle, and end of each move. Ordinary slow camera movement stays silent.
The saved editing instructions
What /video-edit-tesseract can do
The skill points to my video assets and records the treatment we arrived at through these revisions. It also tells the agent how to check the export and keep the project editable.
Choose a stage to see what my saved skill directs the agent to do. The labels distinguish native editing features from my instructions and the tools used alongside Tesseract.
- Engine
- Tesseract's editable document and renderer
- Workflow
- My custom skill and editorial decisions
- Helper
- Local scripts, media, and other tools
Source & brand
Start with the intended footage and the actual brand assets.
- WorkflowRead the brief and select the complete intended take. Keep separate video concepts in separate projects.
- HelperCopy camera-card files and compare hashes before editing. Preserve the originals.
- WorkflowRead the video asset repository, inspect a real reference, and use its fonts, colors, and graphics.
- EngineImport local footage, audio, images, and fonts into an editable project.
My video palette uses navy #000957, blue #344CB7, light blue #577BC1, white #FFFFFF, and teal #2DD4BF. Syne 700 and DM Sans 600 supply the video type. The website has its own type system.
Read all nine stages without switching views
Source & brand
Start with the intended footage and the actual brand assets.
- Workflow: Read the brief and select the complete intended take. Keep separate video concepts in separate projects.
- Helper: Copy camera-card files and compare hashes before editing. Preserve the originals.
- Workflow: Read the video asset repository, inspect a real reference, and use its fonts, colors, and graphics.
- Engine: Import local footage, audio, images, and fonts into an editable project.
My video palette uses navy #000957, blue #344CB7, light blue #577BC1, white #FFFFFF, and teal #2DD4BF. Syne 700 and DM Sans 600 supply the video type. The website has its own type system.
Words & timing
Know what was actually said before drawing a caption.
- Helper: Start with a supplied transcript or local speech recognition. Tesseract does not bundle transcription.
- Workflow: Check product names, disclosures, quiet word beginnings, and the final line against the audio.
- Workflow: Keep source time and edit time mapped when footage is cut or moved.
- Workflow: Plan each caption, graphic, camera move, and sound against a spoken cue.
Speech recognition can suggest words. It cannot certify that those words are correct.
Cut & frame
Keep the complete thought and remove the recording-stop pause.
- Engine: Use video layers with source ranges, active ranges, fit, crop, position, and scale.
- Workflow: Tighten dead handles while preserving complete words and natural breaths.
- Workflow: Match framing across cuts and keep important hands, products, and UI in view.
- Workflow: End on the final thought with a motivated move and a clean cut.
An exported clip can end earlier while the original source remains available inside the project.
Captions
Readable words on the torso, with emphasis chosen by meaning.
- Engine: Editable text layers use imported fonts, outlines, colors, transforms, and timed keyframes.
- Workflow: Use one or two natural lines with no caption rectangle. Reveal words at their verified onsets while the phrase block stays in place.
- Workflow: Use white support words, selective teal emphasis, navy outlines, and occasional blue underlines.
- Workflow: At a 1080×1920 portrait canvas, begin near 90 to 100 px for support and 110 to 130 px for emphasis, then check actual fit.
- Workflow: Give selected words one quick settle. Keep the final phrase readable and move the captions clear of a product close-up.
These are custom style instructions. A generic Tesseract install does not know my approved caption treatment.
Camera movement
Add gentle movement without pulling the speaker's eyes up the frame.
- Workflow: Look for static narration, especially passages with no graphic. Consider a 2 to 5% push or release over 3 to 6 seconds.
- Helper: Calculate paired scale and position values from the eye midpoint after crop and fit.
- Engine: Animate the footage group with matching timing and easing for scale and position.
- Workflow: Keep captions and diagrams in screen space. Preserve natural head motion and match eye height across cuts.
- Workflow: An intentional cup close-up may change the focal point. Restore the established speaker framing on return.
This broader narration rule was added after v3. The illustration below explains the next-edit rule; it is not footage of a fourth version.
Graphics & motion
Show a relationship or process the spoken sentence cannot show by itself.
- Engine: Build editable text, shapes, groups, imported media, masks, adjustments, and keyframed effects.
- Workflow: Stage nodes and connectors in speech order, with an entrance, readable hold, and quiet exit.
- Workflow: Adapt existing assets and real supplied screenshots. Avoid invented product UI and decorative boxes repeating the dialogue.
- Engine: Use supported procedural scripts and custom WGSL effects when the installed schema and local resources support them.
Advanced compositing is version-dependent. Do not assume person segmentation, mesh import, rigging, simulation, or every 3D effect is available.
Voice, music & effects
Keep the voice dominant and give different visual actions different sounds.
- Engine: Place separate audio layers with gain, fades, and timed music ducking.
- Helper: Create local procedural tap, click, pop, ping, whoosh, riser, downlifter, and impact accents. Adjust duration, level, and seed.
- Workflow: Match effects to meaningful arrivals, exits, or focal moves. Leave ordinary camera drift silent.
- Helper: Import supplied or cleared music. The Python-made bed in this edit was imported, not generated by Tesseract.
- Workflow: Preserve voice character, room tone, and breaths. Compare any restrained cleanup at matched loudness.
No bundled full music-generation service, voice generator, or dedicated dereverberation model. Timed gain ducking is not automatic signal-driven compression.
Review the export
Check the actual video people will watch.
- Engine: Render native previews and the final composition.
- Workflow: Inspect caption entrances and holds, graphic overlaps, zoom start/middle/end, cut boundaries, and the final second.
- Helper: Verify dimensions, duration, frame count, codec, full decode, audio timing, loudness, and true peak.
- Workflow: Listen to the final mix where playback is available. A meter or waveform match is not a listening review.
- Workflow: Fix creative problems even when a technical render check passes.
Preserve an approved mix. For a new mix, use delivery requirements and keep true peak at or below -1 dBTP.
Files & versions
Deliver a playable video and preserve the editable project.
- Engine: Save a portable .tsrct document with native footage, text, vector, audio, and animation layers.
- Engine: Export MP4, or supported transparent ProRes 4444 overlays. Inspect the actual output dimensions and frame rate.
- Helper: Make a smaller review encode, posters, or filmstrips without changing the edit.
- Workflow: Make a square original/edited comparison when requested, with a shared source clock and one audio mix.
- Workflow: Keep the prior version and record the actual version and location. Update a task or upload only through an authorized connector workflow.
The package runs locally on supported macOS/Windows hosts. It does not cloud-render, publish, sync with Premiere/After Effects, or bundle generative video. Upload, creator approval, sponsor approval, and publication are separate events.
6. Start the next edit with the instructions already saved
The skill now tells the agent to verify the words before captions, plan graphics around spoken cues, keep captions on the torso, vary sound by purpose, and trim the recording-stop pause. It points to the approved edit for reference. I should not have to discover the same low caption boxes again on every first cut.
It also keeps the work editable. I get an MP4 to watch and a native .tsrct project that retains the footage, text, graphics, and sound layers. The agent checks the encoded video as well as its previews, including caption overlap, cue timing, the ending, and whether the file plays all the way through.
This is a new prompt for the next job, built from the saved instructions. It is not the prompt that produced v3 in one pass.
To try the underlying tool, start with Mirage's Tesseract repository and installation instructions. My style settings sit on top of its video-editing and motion-graphics skills. The separate video about tracking coding agents in Linear and the coffee edit are still waiting. I can use this skill when I pick those up.