# I one-shot a CarPlay video edit with Claude and Tesseract

Source: https://aaronmakelky.com/blog/one-shot-video-edit-claude-tesseract
Last updated: 2026-10-04

I sent one prompt and gave zero revision notes. Here are the real prompts, Linear ticket, brand rules, and outputs behind a 2:11 ChatGPT dot vs Grok Bot CarPlay video, and how to set up your own version.

[Back to blog](/blog)

Field notes

October 4, 2026 · 9 min read

Claude Code · Tesseract · AI Video Editing · Linear · Creator Workflows

My followers asked how ChatGPT dot and Grok Bot voice calls actually work over Apple CarPlay. So I sat in my car, called both, and recorded the whole thing in one take.

I sent Claude one prompt about the edit. The first version it rendered is the one I posted. No notes.

[Download the edited MP4.](/videos/carplay-dot-vs-grok-bot.mp4)

The finished edit: 2:11 cut from a 2:52 take. Captions on this player come from the VTT file the edit produced.

[Watch the post on X](https://x.com/theaaron/status/2106817466349519090) if you want the replies too.

## What one-shot means here

The last time I wrote about [editing with Tesseract](/blog/tesseract-video-edit-process), it took three rounds of corrections to get it right. This one took zero rounds of feedback.

The first message described the video I wanted. The second one only told Claude where the file was, because that session ran in the cloud and couldn't see my Mac. The third was four words on the Mac that pointed a local session at the ticket. Nothing about the edit changed after the first message, and I never asked for a revision.

## The gear

- Camera: Insta360 Luna Ultra, 4K at 24 fps in a wide 16:9 frame. The source file was 3840 by 2160 HEVC, 2 minutes 52 seconds, 1.6 GB.
- Mic: Insta360 Mic Pro, the one with the e-ink display.
- Model: Claude Opus 5.5 on medium effort in Claude Code.
- Editor: Tesseract, driven by my `video-edit-tesseract` skill.
- Tracking: Linear, plus Claude Remote Control to follow the local session from the Claude app.

I sat on one side of the frame and the car's screen sat on the other. That made a 16:9 edit easy to read and a vertical crop basically impossible, which the agent figured out on its own.

## The prompt

The prompt, typos included:

```
take the video i just shot here /Users/aaronmakelky/Library/Mobile\ Documents/com\~apple\~CloudDocs/Personal\ Brand\ Content/gpt\ dot\ vs\ grok\ bot\ apple\ car\ play

and use tesseract plus my tesseract video editing skill to edit it into an engaging, informative, clear video for my followers who asked how chat gpt dot vs grok bot voice calls work on apple car play. use my personal brand style and colors for captions, don't obscure any of the screens while im referencing them or cover the speakers face.

track this project in linear with clear documentation on progress and assets so other coding agents can follow and possibly help
```

1. An audience and a question. "My followers who asked how..." gave the edit a job.
2. Two places graphics can't go. Not over the screens while I'm pointing at them, and not over my face.
3. Pointers instead of descriptions. "My personal brand style" and "my tesseract video editing skill" point at files that already exist. I didn't describe a single color or font.

## The cloud session wrote the ticket first

The session I typed that into was running in the cloud. It couldn't reach iCloud, Tesseract, or my local skill. So it opened Linear ticket AAR-270 and wrote the brief down as a spec any agent could pick up.

![Linear ticket AAR-270 showing the goal, source, tooling, edit requirements, deliverables, and done-when sections.](/images/blog/one-shot-video-edit-claude-tesseract/linear-aar-270.png)

AAR-270, the ticket both sessions worked from.

It found the brand rules from my older video tickets and wrote them out as numbered, testable requirements. This is the exact block from the ticket:

> "Captions in Aaron's video brand style: navy #000957, blue #344CB7, light blue #577BC1, white #FFFFFF / soft white #F8FAFC, teal #2DD4BF as the single accent (one teal group per frame). Syne 700 for hook/titles, DM Sans 600 for captions."
>
>
>
> "Never obscure a screen while Aaron is referencing it (phone, CarPlay display, ChatGPT/Grok UI). When a screen is on camera, move captions to a clear zone or zoom/punch-in on the screen; don't lay graphics over it."
>
>
>
> "Accuracy: only claims Aaron makes on tape. No invented features, UI, or benchmarks."

It also wrote a deliverables list I never asked for: an editable Tesseract project, a review MP4, SRT and VTT caption files, an edit decision list with source timecodes, a suggested post caption, and SHA-256 hashes for every file.

## The second message, and Remote Control

I replied with where the file was:

```
run this now, video is in google drive and here locally: /Users/aaronmakelky/Downloads/VID_20261004_095823_261.mp4

this is running on mac studio in claude desktop app now, so you should have access to tesseract and needed local tooling.
```

It wasn't. The session was still in the cloud, and a 1.6 GB file is too big to pull through the Google Drive connector. So it added both file locations to the ticket, logged the blocker, and told me to start a local session on the Mac Studio and point it at AAR-270.

So I opened a local Claude Code session on the Mac Studio, on Opus 5.5 at medium effort, and typed four words:

```
Work aar-270 from linear
```

Claude Remote Control let me follow along from the Claude app on my phone without sitting at the desk.

![The Claude iPhone app in a Remote Control session titled aar-270 Linear, showing the prompt Work aar-270 from linear and the agent loading the skill and transcribing.](/images/blog/one-shot-video-edit-claude-tesseract/remote-control.png)

Following the Mac Studio session from the Claude app on my phone.

A little over two hours after my first prompt, the ticket moved to In Review with v1 attached.

## What the agent actually did

The local session posted its progress as Linear comments, in this order:

1. Checked the source. It confirmed the iCloud and Downloads copies were byte-identical with a SHA-256 hash, then worked in a separate `edit/` folder so the original was never touched.
2. Transcribed and timed every word. Whisper (medium) for the transcript, then forced alignment for exact word onsets. Captions only use words I actually said.
3. Cut 2:52 down to 2:11 across 30 kept segments. A false start got cut, and every cut is listed with source in and out points in the edit decision list.
4. Counted the rings from the audio. One call rang five times over about 17 seconds. It turned that wait into five one-second jump cuts with a "RING 1-5" counter.
5. Punched in on the car screen every time I pointed at it, and moved captions to the band above the display during those shots.
6. Lifted the call audio. The assistants' voices came through about 15 dB quieter than me. It brought them up about 11 dB and tagged their lines "DOT" or "GROK BOT" so you always know who's talking.
7. Built a closing comparison card that only uses claims I made on camera.

![Punch-in on the car display with the brand caption band placed above the screen instead of over it.](/images/blog/one-shot-video-edit-claude-tesseract/carplay-punch-in.jpg)

A punch-in on the car display. The caption moves up so the screen stays clear.

![Navy ring counter card reading RING with the count, shown during a jump cut of the call ringing.](/images/blog/one-shot-video-edit-claude-tesseract/ring-counter.jpg)

The ring counter turns 17 seconds of ringing into five quick cuts.

![Closing comparison card in navy, white, and teal comparing ChatGPT dot and Grok Bot calls over CarPlay.](/images/blog/one-shot-video-edit-claude-tesseract/comparison-card.jpg)

The closing card. Every line on it is something I said on tape.

## The phone chat the agent blurred

In three close-ups, the Grok Bot chat on my phone was readable. It showed personal schedule and account details I didn't want on the internet.

The agent caught it, added a blur that tracks the phone text, and then ran OCR over every frame of the final export to prove none of those strings were readable. It also deleted the earlier exports that came before the blur, including iCloud's conflict copies.

![Close-up of the phone during the call with the chat text blurred out.](/images/blog/one-shot-video-edit-claude-tesseract/phone-blur.jpg)

The tracked blur on the phone. OCR on the final export found none of the private text.

At 1:55.8 the word "It's" sits over the blurred phone for about a third of a second. The agent flagged it in the ticket. I didn't change it.

## What came back

Everything landed in one `edit/` folder next to the source, with a README on top:

- `GPT-dot-vs-Grok-Bot-CarPlay-v1-review.mp4`, an 88 MB review copy
- `GPT-dot-vs-Grok-Bot-CarPlay-v1.mp4`, the 1 GB native Tesseract export
- `GPT-dot-vs-Grok-Bot-CarPlay-v1.tsrct`, the editable project with fonts packaged
- `GPT-dot-vs-Grok-Bot-CarPlay-v1.srt` and `.vtt` caption files
- `EDL-v1.md`, source in and out points for all 30 cuts
- `POST-CAPTION-v1.md`, post copy and hashtag options
- A filmstrip preview

![Filmstrip of frames from the finished edit showing captions on the torso, punch-ins on the car display, and the closing card.](/images/blog/one-shot-video-edit-claude-tesseract/filmstrip.jpg)

The filmstrip preview the agent made so I could check it without opening the video.

It also posted its checks to the ticket. The full FFmpeg decode passed on both MP4s (3,150 frames). Loudness came in at -14.0 LUFS integrated with a true peak of -1.6 dBTP, and the call voices sat about 2 dB under mine. The output was 1920 by 1080 at 24 fps.

It didn't listen to the mix and said so in the ticket. The audio numbers come from meters.

It skipped the vertical cut. The car screen and I sit on opposite sides of the frame, so a 9:16 crop would lose the demo, and the agent said so in the ticket.

## Why this was one-shottable

### 1. A brand system written as rules

My video brand lives in a GitHub repo with a `DESIGN.md` and the font files. Every color has a job and every motion has a limit:

```
# Video brand rules

## Color (each one has a job)
- Panel background: navy #000957
- Secondary panel: blue #344CB7
- Text: white #FFFFFF
- Accent: teal #2DD4BF, one accent group per frame, never a full background

## Type (ship the font files with the project)
- Titles and hooks: Syne 700
- Captions and labels: DM Sans 600

## Motion
- Enter with 12-40px slide plus fade, 0.35-0.65s
- No bounce, glow, gradients, emoji, or platform brand colors

## Placement
- Never over the speaker's face
- Never over a screen while I'm pointing at it
- Captions sit on the torso band or in empty frame
```

### 2. A skill that holds your editing workflow

My `video-edit-tesseract` skill is a saved set of instructions Claude loads when I name it. You can see its habits in the steps above: work on a copy, time every word, cut on what was said, place captions by rule, render, then run real checks before calling anything done.

If you haven't built one yet, the [Tesseract walkthrough](/blog/tesseract-video-edit-process) shows the three rounds of corrections from my first Tesseract edit.

### 3. A ticket that defines done

- It let a second session on a different machine start exactly where the first one stopped.
- It listed deliverables up front: the SRT, VTT, and EDL I never asked for.
- It held the progress log, so I could check status from my phone instead of reading a terminal.

### 4. Footage shot for the edit

I used a dedicated mic for clean audio and shot in 4K so punch-ins stay sharp. I also left room in the frame for captions that won't cover anything. The agent can't add pixels or remove wind noise.

## Copy this

Here's a template built from my prompt:

```
Edit the video at [path] for [audience] who asked [question].

Use [your editing skill] and the brand rules in [path or repo]/DESIGN.md
for captions and graphics. Never cover my face, and never cover a screen
while I'm pointing at it.

Track it in [Linear/your tracker] so another agent could pick it up.
Done means: editable project, review MP4, SRT and VTT captions, an edit
decision list with source timecodes, a suggested post caption, and a
note of every check you ran and every check you couldn't run.
```

## Related pages

- [All field notes](/blog)
- [My first Tesseract edit, all three versions](/blog/tesseract-video-edit-process)
- [Editing a video from my phone with Codex and HyperFrames](/blog/codex-hyperframes-video-editing-workflow)
- [How to make your own ChatGPT dots in Grok Bot](/blog/chatgpt-dots-in-grok-bot)
- [About Aaron Makelky](/about)

## Sources and references

- [Published video on X](https://x.com/theaaron/status/2106817466349519090)
- [Tesseract on GitHub](https://github.com/mirage-hq/Tesseract)
- [Claude Code Remote Control](https://code.claude.com/docs/en/remote-control)
- [Linear](https://linear.app)

## Common questions

### What does one-shot mean for this video edit?

Aaron sent one creative prompt and approved the first rendered version with no revision notes. Two later messages only told Claude where the source file lived and handed the Linear ticket to a local session.

### What tools were used to edit the video?

Claude Code running Claude Opus 5.5 at medium effort, Tesseract with Aaron's video-edit-tesseract skill, Linear for tracking, Claude Remote Control to follow the local session, Whisper and forced alignment for word timing, and FFmpeg for checks.

### What camera and microphone recorded the source?

An Insta360 Luna Ultra recording 4K at 24 fps in a wide 16:9 frame, with audio from an Insta360 Mic Pro, the version with the e-ink display.

### Do I need Aaron's brand system to get the same result?

No. You need your own written brand rules: a palette with a job for each color, two font files, motion limits, and places graphics are never allowed to go.

### Why did the edit include SRT and VTT caption files?

The Linear ticket listed them as deliverables before any editing started.
