Captions for Final Cut Pro

Drop your clip in here and download an iTT, the caption format Apple built for Final Cut Pro. The transcription runs on your own Mac, so the footage never leaves it.

Drop your video here

MP4, MOV, WebM or any audio file. It is read by this page and never uploaded.

Beast

A looping preview of the caption styles this tool produces, currently showing the Beast style, and cycling through: No watermark, ever, Nothing is uploaded, It all runs on your device, No sign up, no account, Caption as much as you like, Fix any word in one click, A file for your editor, Or the finished MP4, Then post it with PublishQ.

Final Cut Pro takes iTT natively, and iTT is the better file

Apple built iTT for exactly this workflow, and Final Cut Pro reads it without a plugin, a conversion or a paid step: File, Import, Captions, and the words are on your timeline positioned and styled. This page writes it to Apple's own published profile, so it behaves like a caption file Final Cut made itself. Worth knowing while you are here: Final Cut accepts three formats in total, CEA-608, iTT and SRT, and Advanced SubStation Alpha is not one of them, so iTT is both the best file for this and the one to take.

iTT is worth having rather than tolerating. It is a restricted profile of TTML, so it is XML rather than plain text, and it carries a region and a style alongside each line: where the caption sits, how it is aligned, what weight it is. An SRT carries the words and the timecodes and nothing else, which means Final Cut has to guess the rest. Apple pins the profile down tightly, and our export follows it: exactly one div per document, timecodes on an SMPTE timebase with a frame rate, non-drop, and sansSerif as the typeface, encoded UTF-8 with no byte-order mark.

One thing to know before you start: recent Final Cut Pro can transcribe spoken English on the Mac itself. If your clip is English and you are happy with a plain caption track, that may be all you need. Reach for this page when you want another language, when you want the word-by-word animation rather than a caption track, or when you want the captions burned into the picture so nothing downstream can drop them.

Getting captions into Final Cut Pro

Take the iTT. Apple built it for this workflow, and unlike SRT it carries positioning and styling as well as the words.

  1. Drop your clip above and let the words be timed on your Mac.
  2. Correct anything the model misheard in the transcript.
  3. Download the iTT file, or burn the captions into an MP4.
  4. In Final Cut Pro, choose File then Import then Captions and pick the .itt.

Captioning for somewhere else

Same transcription, and the file to take differs by destination. These pages say which one and why.

This tool is free, forever. No accounts, no paywalls, no catch.

I build these solo. If you want to support my work, donating a coffee would be highly appreciated! ☕

Frequently Asked Questions

Everything you need to know about Captions for Final Cut Pro from PublishQ

CEA-608 (.scc), iTT (.itt) and SRT. That is the whole list, and it is why this page leads with iTT: of the three, it is the one that carries the text colour and the word-by-word highlight rather than the words alone. Final Cut has no iTT control for a background box, a font or a size, so those stay in the burned-in MP4. ASS is not supported at all, so do not download it for this workflow.
iTunes Timed Text, Apple's own caption format and a restricted profile of the W3C TTML standard. Being XML rather than plain text, it can say where a caption sits on screen and how it is styled, not just what it says and when. Apple uses it throughout its own delivery pipeline, which is why Final Cut Pro and Compressor both read it natively.
Drop the video onto this page. A speech recognition model runs inside your browser, works out what is said and when each individual word is spoken, and the captions appear over your clip straight away. Pick a style, fix any word with a click, then either burn the captions into a new MP4 or download a subtitle file. There is no account and nothing to install.
No. The file is opened by the page, its audio is decoded by your own browser, and the transcription runs in a background thread on your own hardware. Open your network tab while it works and you will see the speech model coming down to you once, and nothing going back up. After that first download the tool keeps working with your connection switched off.
Every style, the MP4 export and all five subtitle formats cost nothing, and nothing is written onto your video. Because the work happens on your device there is no per-video server cost to recover, so there is nothing to meter and no better version of the output held back. PublishQ makes its money from the social media scheduler.
Instead of a whole sentence sitting on screen, each word lights up at the exact moment it is spoken. The movement keeps the eye on the text, and the text keeps the viewer through the first few seconds where most short-form video is lost. It needs a timing for every single word rather than for each line, which is what the model here produces.
Ten of them, and all ten are free: Normal, Beast, Karaoke, Pill, Subtitle, Neon, Lava, Boxed, Typewriter and Bounce. They differ in what happens on the spoken word, which is the part that does the work: a colour change, a pop, a pill sliding behind it, a glow, a gradient, or a word arriving one at a time with a caret waiting for the next. You can also set the vertical position, the type size and how many words sit on screen at once.
Chrome or Edge on a computer, because those include the WebCodecs hardware video encoding this uses. That is also why a burn-in finishes in well under the time the clip takes to play. On Safari and Firefox the preview and every subtitle file still work, so you can generate an SRT there and burn it in your editor.
SRT, VTT, ASS, iTT and plain text. SRT is the one that works everywhere. VTT is the web standard for an HTML5 video track or a YouTube upload. ASS is the only one that keeps the per-word timing, through karaoke tags, so an editor can rebuild the animation. iTT is Apple's own format, which is what Final Cut Pro takes natively. Plain text is just the transcript, for a description or a blog post.
Yes, and the format to pick differs. CapCut names SRT and TXT as the subtitle files it imports, and that import lives in CapCut Desktop and CapCut Web rather than the mobile app. Premiere Pro lists SRT among its supported sidecar files, and DaVinci Resolve imports SRT on the free version as well as Studio. Final Cut Pro accepts SRT, iTT and CEA-608 and does not read ASS at all, so use the iTT file there. Each of those has a page of its own with the steps on it.
Clear speech in English comes back close to word perfect. Accents, background music, several people talking over each other and technical vocabulary are where any speech model gets weaker, and this one is no exception. The transcript below the video is editable for exactly that reason: clicking a word and correcting it changes the burned-in caption and every file you download, and it keeps the timing the model worked out.
Whisper handles around a hundred, with auto-detect on by default, and eighteen of the common ones are in the picker so you can name the language and skip detection. Accuracy is best in English. If another language comes back rough, switch to the maximum accuracy model, which is better outside English and costs a 759 MB download and several times the running time.
The first run downloads the speech model: 290 MB on a machine with a GPU, 206 MB on the slower fallback path. Your browser caches it, so every clip after that skips the download entirely and spends its time on the listening alone. Come back tomorrow and the model is still there.
Up to 500 MB per file, and no limit on how many you do. Length is bounded by your own device rather than by us: everything is held in memory while it works, so a long clip on a phone is the case that struggles. If you are captioning something long, do it on a computer.
Yes. The video is yours, it never leaves your device, and the open source model behind the tool puts no conditions on what you do with the result. Post it, sell it, run it as an ad.
Those upload your video to a GPU server, which is why they cost what they cost, cap your minutes and put a watermark on the free tier. Here the same work happens on hardware you already own. What you give up is honest: a large hosted model in a datacentre still wins on difficult audio, and a phone will be slower than their servers. What you keep is your footage, your minutes and a clean export.
Start for Free

No credit card required • Set up in under 3 minutes

Alexandro - Founder
PublishQ

— me 👋

Hi, I'm Alexandro 👋

I left my Software Engineer role at Amazon to build tools that solve real problems — the kind big companies ignore because they read spreadsheets instead of using their own products.

I was spending over an hour daily just scheduling 2 shorts across 3 platforms — logging in, reformatting, uploading one by one. That felt broken. So I built PublishQPublishQ . Now I create 4 shorts in 3 minutes and schedule them to 4 platforms in under 30 seconds.

PublishQ is bootstrapped. No investors, no vanity metrics. I build what actually helps you — because I use it every day myself.

Thank you,

Alexandro

Captions burned in. Now get it posted.

Use PublishQPublishQ to schedule that clip to every account you manage from one calendar. No more uploading the same video five times.

Connect Your Accounts

Add once, use forever

Schedule Posts

Plan your content ahead

Multi-Platform

Post everywhere at once

Start for Free

No credit card required • Set up in under 3 minutes

Plans from $14/mo

2 months free on yearly billing. Upgrade or cancel anytime.

Free
$0/mo
Solo
$14/mo
Starter
$29/mo
Professional
$49/mo
Business
$99/mo
Compare all plans and features