# I asked an AI agent to edit my video. Here are all three versions.

Source: https://aaronmakelky.com/blog/tesseract-video-edit-process
Last updated: 2026-09-21

My exact prompts, three before-and-after edits, and the caption, sound, and camera rules I saved for the next Tesseract video.

[Back to writing](/blog)

Video editing

The first cut had low captions, repetitive graphics, and a long pause at the end. These are the actual videos and the instructions I gave Codex to improve them with Tesseract.

Aaron Makelky · September 21, 2026 · 10 min read + 3 short videos

Mirage sponsored the video shown here and provided early access to Tesseract.

[The request](#start) [Watch v1](#v1) [Watch v2](#v2) [Watch v3](#v3) [The camera rule](#camera) [Explore the skill](#skill)

## 1. Give the agent the job, the footage, and the brand

I recorded a short video asking whether AI could edit as well as I can. I wanted the raw take and the edited version beside each other so people could judge the result. I also gave some instructions out loud while recording, including a diagram, a zoom on my coffee cup, and quiet retro music.

I was working in Codex with [Tesseract by Mirage](https://github.com/mirage-hq/Tesseract), which lets an agent build and render an editable video project locally. Linear supplied the task brief, and GitHub supplied my video brand files.

There were three different video ideas on the card. I had to tell the agent which one we were editing. Here is my first request, with the original spelling and app mentions preserved.

My exact starting prompt Copy exact prompt

> pull the videos from today on my camera sd card, then edit the first video needed in the [@Linear](plugin://linear@openai-curated-remote) task about tesseract by mirage. there is a separate video about linear to keep track of coding agent work, and finally a third distinct video abot tesseract editing a cool edit of me drinking my coffee. we will work on those two later, for now, its just he sponsored tesseract video we have an opportunity folder and sponsored post for. read the docs and lets get this task going

Eight clips were copied from the card and checked against the originals. The selected complete take ran 57.375 seconds. The other concepts stayed separate. I then pointed the agent at my existing video assets instead of describing a new look from scratch.

My exact brand instruction Copy exact prompt

> make sure you use my personal brand visuals from the video assets folder in [@GitHub](plugin://github@openai-curated-remote) for this video. you can improve and modify them, but follow the brand guidelines on colors for the graphics and captions

The videos use my navy, blue, light blue, white, and teal palette, with Syne and DM Sans. The font files mattered as much as the color names. The first font import needed a compatibility fix before the rendered text actually used the intended letterforms.

### Watch the original take and read what I said on camera

Your browser cannot play this video. [Open the Original take MP4](/media/tesseract-edit-process/original.mp4).

The complete 57.375-second camera take, resized for the web. The original audio is preserved.

So, can AI really edit a video as well as me? I'm going to show you side by side the raw version uncut, and then the version I'm going to edit with this new tool called Tesseract. This video is sponsored by Mirage. Thank you for giving me early access. I'm going to have you put on the screen right now a little diagram that explains how I go from camera footage to ChatGPT Work to edit with Tesseract, and it doesn't have to open a standalone video editing tool. A couple other rules: when I point at something like my name on my coffee cup right there, I need you to zoom in and put a gentle sound effect. The captions and on-screen graphics don't cover my face. And then make sure the background music is lo-fi and kind of retro, but it's not loud enough to distract from the narrative. So, how did Tesseract do? Can this edit a video as well as I can?

## 2. Watch the first cut before fixing it

The first version gave me a playable comparison with captions and graphics. It also left the things I would immediately change in an edit: a caption rectangle sitting too low, explainers that felt static, and too much footage after I finished talking.

Watch the opening captions, then jump to the coffee and the ending. Each player shows the original on the left and that version's edit on the right. You hear one edited mix.

Your browser cannot play this video. [Open the Version 1 MP4](/media/tesseract-edit-process/v1.mp4).

Start 0:18 · Diagram 0:32 · Coffee 0:50 · Ending

Version 1 · 57.375 seconds. The full take includes the pause before I stop recording.

The file passed its technical checks, but I still didn't like the treatment. I described the changes in terms of things I could see and hear.

My exact feedback on v1 Copy exact prompt

> update [@Linear](plugin://linear@openai-curated-remote) with your progress, yes. edits: There is no sound effect when the camera zooms in on my coffee cup. - The ending leaves way too much of a handle on the clip after I'm done talking. We need to use some kind of zoom or engaging motion to end the clip instead of just dead air where I look down and pause the recording. - The captions are too bland and sit too low in the frame. They need to be on the torso of the speaker and use some more engaging elements, like no background or having the words pop in one at a time for more of a TikTok social media reel style instead of just a bland rectangle with the captions. - The little explainer graphics that pop up need to be animated. - There should be elements of motion and gentle sound effects as they come in and out of frame instead of just static rectangles that are conveying the same thing already said in the dialogue.

That gave the agent specific work: move the captions onto my torso, remove their background, animate the explainers, make the coffee cue noticeable, and end the clip after the final thought. Asking for a more engaging edit had left too many of those decisions unspecified.

## 3. Move the captions and give the graphics a purpose

Version 2 moved the words up onto my torso and removed the caption boxes. Words appeared at their spoken timing, while short phrases stayed together on screen. The graphic elements gained entrances and exits, with the camera, ChatGPT Work, and Tesseract explanation appearing in sequence.

The coffee zoom got a stronger warm sound cue, and the music dipped around it. The ending lost 2.625 seconds of recording-stop footage. A closing push-in carried the final question, and the clip stopped at 54.75 seconds.

Your browser cannot play this video. [Open the Version 2 MP4](/media/tesseract-edit-process/v2.mp4).

Start 0:18 · Diagram 0:32 · Coffee 0:50 · Ending

Version 2 · 54.75 seconds. Torso captions, animated explainers, a clearer coffee cue, and a tighter ending.

My exact comparison request Copy exact prompt

> great. now show me the side by side versions that will be the final format

I said, "that version works for now." Seeing the result at the final comparison size helped. A caption can look fine in a full-height portrait preview and become too small when the video shares the frame with the original.

I kept the complete take in the editable project. Both sides of the revised comparison use the same source timing, and both end together. Trimming the end of the delivery did not delete the original recording.

## 4. Give the captions more range and stop repeating the ding

The next request was smaller. I wanted stronger captions and more variety in the sound effects.

My exact v3 request Copy exact prompt

> lets take the captions up a notch, and diversify the "ding" sound effects. does tesseract offer you options for music and audio effect generation, or do you need a folder of those from me?

Version 3 kept the same 167 verified words and their timing. It regrouped them into 42 natural phrases, down from 70 shorter groups in v2. Supporting words made a small entrance. Selected words appeared larger, rose a little, and settled once. Five blue underlines drew attention to particular phrases.

The captions also stayed white while the explainer used teal. I could read the words without having two competing teal elements asking for attention.

Your browser cannot play this video. [Open the Version 3 MP4](/media/tesseract-edit-process/v3.mp4).

Start 0:18 · Diagram 0:32 · Coffee 0:50 · Ending

Version 3 · 54.75 seconds. This is the edit I approved. The timing, music bed, diagram motion, coffee zoom, and closing move carry over from v2.

The nine sound cues now have different jobs. The camera graphic gets a double-click, the chat action gets short typing and send ticks, and the Tesseract handoff gets a snap. Air sounds handle transitions. The coffee moment keeps its warm accent, and the ending gets a gentle impact. The words themselves do not each make a noise.

My response was "banger, well done." I asked for v3 to be uploaded and its location recorded on the task.

### Where the audio actually came from

I didn't need to supply a sound folder for these effects. Tesseract's local helper can make simple procedural taps, clicks, pops, pings, whooshes, risers, downlifters, and impacts. Combining those sparingly produced the different accents.

The music was a separate job. This edit used an instrumental bed synthesized locally with Python and imported into Tesseract. Calling that Tesseract music generation would be inaccurate. If I wanted a particular licensed track or recorded sound, I would supply it or ask the agent to use another authorized source.

The captions also depended on local transcription outside Tesseract and a verified word map. Tesseract handled the editable text and animation; it did not transcribe the recording for us.

## 5. Keep my eyes level when the camera moves

After approving v3, I added another instruction. Long stretches of narration need subtle camera movement, especially when there is no graphic on screen or when a framing change can help cover a cut. During a zoom, my eyes should stay at the same height in the frame.

My exact camera and skill request Copy exact prompt

> one edition to this learning, we need more subtle camera movement to make the video more engaging. especially with a long period of narration with no on-screen graphic, or covering a cut. zooms should keep the speakers eyes at the same level in the frame. take what youve learned from this edit task with tesseract, and create a new skill.md that points at my video assets folder and has clear instructions to get the similar style output, captions, graphics, sound effects, etc what we achieved through iterations as a first time output when the skill is invoked. name it / video-edit-tesseract.

That became a rule in `/video-edit-tesseract`. A typical narration move starts with a slow 2 to 5 percent push or release over 3 to 6 seconds. The agent pairs the zoom with a position adjustment so the camera movement does not pull the eyes toward the top of the frame. It checks the position after cropping, including inside the smaller comparison panel.

This broader narration movement has not been rendered into v3. The cup close-up and closing push-in were already there. The new rule is saved for the next edit, with a calculator and instructions to inspect the start, middle, and end of each move. Ordinary slow camera movement stays silent.

The saved editing instructions

## What /video-edit-tesseract can do

The skill points to my video assets and records the treatment we arrived at through these revisions. It also tells the agent how to check the export and keep the project editable.

Choose a stage to see what my saved skill directs the agent to do. The labels distinguish native editing features from my instructions and the tools used alongside Tesseract.

Engine Tesseract's editable document and renderer

Workflow My custom skill and editorial decisions

Helper Local scripts, media, and other tools

Source & brand Words & timing Cut & frame Captions Camera movement Graphics & motion Voice, music & effects Review the export Files & versions

### Source & brand

Start with the intended footage and the actual brand assets.

- Workflow Read the brief and select the complete intended take. Keep separate video concepts in separate projects.
- Helper Copy camera-card files and compare hashes before editing. Preserve the originals.
- Workflow Read the video asset repository, inspect a real reference, and use its fonts, colors, and graphics.
- Engine Import local footage, audio, images, and fonts into an editable project.

My video palette uses navy #000957, blue #344CB7, light blue #577BC1, white #FFFFFF, and teal #2DD4BF. Syne 700 and DM Sans 600 supply the video type. The website has its own type system.

### Read all nine stages without switching views

### Source & brand

Start with the intended footage and the actual brand assets.

- Workflow: Read the brief and select the complete intended take. Keep separate video concepts in separate projects.
- Helper: Copy camera-card files and compare hashes before editing. Preserve the originals.
- Workflow: Read the video asset repository, inspect a real reference, and use its fonts, colors, and graphics.
- Engine: Import local footage, audio, images, and fonts into an editable project.

My video palette uses navy #000957, blue #344CB7, light blue #577BC1, white #FFFFFF, and teal #2DD4BF. Syne 700 and DM Sans 600 supply the video type. The website has its own type system.

### Words & timing

Know what was actually said before drawing a caption.

- Helper: Start with a supplied transcript or local speech recognition. Tesseract does not bundle transcription.
- Workflow: Check product names, disclosures, quiet word beginnings, and the final line against the audio.
- Workflow: Keep source time and edit time mapped when footage is cut or moved.
- Workflow: Plan each caption, graphic, camera move, and sound against a spoken cue.

Speech recognition can suggest words. It cannot certify that those words are correct.

### Cut & frame

Keep the complete thought and remove the recording-stop pause.

- Engine: Use video layers with source ranges, active ranges, fit, crop, position, and scale.
- Workflow: Tighten dead handles while preserving complete words and natural breaths.
- Workflow: Match framing across cuts and keep important hands, products, and UI in view.
- Workflow: End on the final thought with a motivated move and a clean cut.

An exported clip can end earlier while the original source remains available inside the project.

### Captions

Readable words on the torso, with emphasis chosen by meaning.

- Engine: Editable text layers use imported fonts, outlines, colors, transforms, and timed keyframes.
- Workflow: Use one or two natural lines with no caption rectangle. Reveal words at their verified onsets while the phrase block stays in place.
- Workflow: Use white support words, selective teal emphasis, navy outlines, and occasional blue underlines.
- Workflow: At a 1080×1920 portrait canvas, begin near 90 to 100 px for support and 110 to 130 px for emphasis, then check actual fit.
- Workflow: Give selected words one quick settle. Keep the final phrase readable and move the captions clear of a product close-up.

These are custom style instructions. A generic Tesseract install does not know my approved caption treatment.

### Camera movement

Add gentle movement without pulling the speaker's eyes up the frame.

- Workflow: Look for static narration, especially passages with no graphic. Consider a 2 to 5% push or release over 3 to 6 seconds.
- Helper: Calculate paired scale and position values from the eye midpoint after crop and fit.
- Engine: Animate the footage group with matching timing and easing for scale and position.
- Workflow: Keep captions and diagrams in screen space. Preserve natural head motion and match eye height across cuts.
- Workflow: An intentional cup close-up may change the focal point. Restore the established speaker framing on return.

This broader narration rule was added after v3. The illustration below explains the next-edit rule; it is not footage of a fourth version.

### Graphics & motion

Show a relationship or process the spoken sentence cannot show by itself.

- Engine: Build editable text, shapes, groups, imported media, masks, adjustments, and keyframed effects.
- Workflow: Stage nodes and connectors in speech order, with an entrance, readable hold, and quiet exit.
- Workflow: Adapt existing assets and real supplied screenshots. Avoid invented product UI and decorative boxes repeating the dialogue.
- Engine: Use supported procedural scripts and custom WGSL effects when the installed schema and local resources support them.

Advanced compositing is version-dependent. Do not assume person segmentation, mesh import, rigging, simulation, or every 3D effect is available.

### Voice, music & effects

Keep the voice dominant and give different visual actions different sounds.

- Engine: Place separate audio layers with gain, fades, and timed music ducking.
- Helper: Create local procedural tap, click, pop, ping, whoosh, riser, downlifter, and impact accents. Adjust duration, level, and seed.
- Workflow: Match effects to meaningful arrivals, exits, or focal moves. Leave ordinary camera drift silent.
- Helper: Import supplied or cleared music. The Python-made bed in this edit was imported, not generated by Tesseract.
- Workflow: Preserve voice character, room tone, and breaths. Compare any restrained cleanup at matched loudness.

No bundled full music-generation service, voice generator, or dedicated dereverberation model. Timed gain ducking is not automatic signal-driven compression.

### Review the export

Check the actual video people will watch.

- Engine: Render native previews and the final composition.
- Workflow: Inspect caption entrances and holds, graphic overlaps, zoom start/middle/end, cut boundaries, and the final second.
- Helper: Verify dimensions, duration, frame count, codec, full decode, audio timing, loudness, and true peak.
- Workflow: Listen to the final mix where playback is available. A meter or waveform match is not a listening review.
- Workflow: Fix creative problems even when a technical render check passes.

Preserve an approved mix. For a new mix, use delivery requirements and keep true peak at or below -1 dBTP.

### Files & versions

Deliver a playable video and preserve the editable project.

- Engine: Save a portable .tsrct document with native footage, text, vector, audio, and animation layers.
- Engine: Export MP4, or supported transparent ProRes 4444 overlays. Inspect the actual output dimensions and frame rate.
- Helper: Make a smaller review encode, posters, or filmstrips without changing the edit.
- Workflow: Make a square original/edited comparison when requested, with a shared source clock and one audio mix.
- Workflow: Keep the prior version and record the actual version and location. Update a task or upload only through an authorized connector workflow.

The package runs locally on supported macOS/Windows hosts. It does not cloud-render, publish, sync with Premiere/After Effects, or bundle generative video. Upload, creator approval, sponsor approval, and publication are separate events.

[Read Mirage's Tesseract documentation](https://github.com/mirage-hq/Tesseract)

## 6. Start the next edit with the instructions already saved

The skill now tells the agent to verify the words before captions, plan graphics around spoken cues, keep captions on the torso, vary sound by purpose, and trim the recording-stop pause. It points to the approved edit for reference. I should not have to discover the same low caption boxes again on every first cut.

It also keeps the work editable. I get an MP4 to watch and a native `.tsrct` project that retains the footage, text, graphics, and sound layers. The agent checks the encoded video as well as its previews, including caption overlap, cue timing, the ending, and whether the file plays all the way through.

This is a new prompt for the next job, built from the saved instructions. It is not the prompt that produced v3 in one pass.

A prompt for the next edit Copy template

> /video-edit-tesseract Edit the complete take for this brief using my video assets and brand guidelines. Use the saved torso-caption style, animated explainers, and varied gentle sound effects. Add subtle camera movement during long narration and at cuts, keeping my eyes at the same height. Preserve the spoken meaning and finish cleanly after the final thought. Show me the review video and keep the Tesseract project editable.

To try the underlying tool, start with [Mirage's Tesseract repository and installation instructions](https://github.com/mirage-hq/Tesseract). My style settings sit on top of its video-editing and motion-graphics skills. The separate video about tracking coding agents in Linear and the coffee edit are still waiting. I can use this skill when I pick those up.

## Related pages

- [More writing](/blog)

## Sources and references

- [Tesseract on GitHub](https://github.com/mirage-hq/Tesseract)
