Captions for CapCut

Drop your clip in here, get the words timed, and download an SRT that CapCut imports directly. The transcription runs on your device, so nothing is uploaded and there is no queue.

Drop your video here

MP4, MOV, WebM or any audio file. It is read by this page and never uploaded.

Beast

A looping preview of the caption styles this tool produces, currently showing the Beast style, and cycling through: No watermark, ever, Nothing is uploaded, It all runs on your device, No sign up, no account, Caption as much as you like, Fix any word in one click, A file for your editor, Or the finished MP4, Then post it with PublishQ.

Two ways in, and one of them needs no import at all

On CapCut Desktop or CapCut Web, an SRT from this page drops straight in and every timing comes with it. CapCut states the detail in its own help pages: it imports external subtitle files as SRT or TXT, and that import lives in Desktop and Web rather than the phone app. So if you edit on your phone there is a second path, and it is arguably the better one anyway.

There are two ways round it. Either move the project to CapCut on a computer or in the browser, where File import handles the SRT normally, or skip the subtitle file altogether and burn the captions into the video here. A burned-in MP4 needs no import at all: CapCut treats it as ordinary footage, because that is what it is, and the captions are already in the picture.

Burning them in also survives everything that happens afterwards. A subtitle track is a separate object that an app can drop when you export, re-upload or hand the file to someone else. Pixels cannot be dropped.

CapCut has auto-captions of its own, and for a clip you are finishing and posting inside CapCut they are the shortest path. Where this page earns its place is the work around the edges: audio you would rather not send to a server, a language its captioning does not cover, a transcript you want to correct once and reuse in several edits, and the word-by-word highlight animation with the font, outline and highlight colour set by you rather than chosen from a preset list that changes.

Getting captions into CapCut

Take the SRT. CapCut lists SRT and TXT as the subtitle files it imports, so SRT is the one that carries your timings.

  1. Drop your clip above and let the words be timed on your device.
  2. Fix any name or brand the model misheard by clicking it in the transcript.
  3. Download the SRT, or burn the captions straight into an MP4.
  4. In CapCut on desktop or web, import the SRT alongside your clip.

Captioning for somewhere else

Same transcription, and the file to take differs by destination. These pages say which one and why.

This tool is free, forever. No accounts, no paywalls, no catch.

I build these solo. If you want to support my work, donating a coffee would be highly appreciated! ☕

Frequently Asked Questions

Everything you need to know about Captions for CapCut from PublishQ

No. CapCut documents subtitle import as a CapCut Desktop and CapCut Web feature, and the mobile app has no equivalent. If you are editing on a phone, burn the captions into the MP4 here instead and bring that in as normal footage.
CapCut names SRT and TXT as the subtitle formats it imports, so treat SRT as the file to bring. Our ASS export exists for players and tools built on libass, such as VLC, mpv and ffmpeg, where the per-word karaoke timing survives.
Drop the video onto this page. A speech recognition model runs inside your browser, works out what is said and when each individual word is spoken, and the captions appear over your clip straight away. Pick a style, fix any word with a click, then either burn the captions into a new MP4 or download a subtitle file. There is no account and nothing to install.
No. The file is opened by the page, its audio is decoded by your own browser, and the transcription runs in a background thread on your own hardware. Open your network tab while it works and you will see the speech model coming down to you once, and nothing going back up. After that first download the tool keeps working with your connection switched off.
Every style, the MP4 export and all five subtitle formats cost nothing, and nothing is written onto your video. Because the work happens on your device there is no per-video server cost to recover, so there is nothing to meter and no better version of the output held back. PublishQ makes its money from the social media scheduler.
Instead of a whole sentence sitting on screen, each word lights up at the exact moment it is spoken. The movement keeps the eye on the text, and the text keeps the viewer through the first few seconds where most short-form video is lost. It needs a timing for every single word rather than for each line, which is what the model here produces.
Ten of them, and all ten are free: Normal, Beast, Karaoke, Pill, Subtitle, Neon, Lava, Boxed, Typewriter and Bounce. They differ in what happens on the spoken word, which is the part that does the work: a colour change, a pop, a pill sliding behind it, a glow, a gradient, or a word arriving one at a time with a caret waiting for the next. You can also set the vertical position, the type size and how many words sit on screen at once.
Chrome or Edge on a computer, because those include the WebCodecs hardware video encoding this uses. That is also why a burn-in finishes in well under the time the clip takes to play. On Safari and Firefox the preview and every subtitle file still work, so you can generate an SRT there and burn it in your editor.
SRT, VTT, ASS, iTT and plain text. SRT is the one that works everywhere. VTT is the web standard for an HTML5 video track or a YouTube upload. ASS is the only one that keeps the per-word timing, through karaoke tags, so an editor can rebuild the animation. iTT is Apple's own format, which is what Final Cut Pro takes natively. Plain text is just the transcript, for a description or a blog post.
Yes, and the format to pick differs. CapCut names SRT and TXT as the subtitle files it imports, and that import lives in CapCut Desktop and CapCut Web rather than the mobile app. Premiere Pro lists SRT among its supported sidecar files, and DaVinci Resolve imports SRT on the free version as well as Studio. Final Cut Pro accepts SRT, iTT and CEA-608 and does not read ASS at all, so use the iTT file there. Each of those has a page of its own with the steps on it.
Clear speech in English comes back close to word perfect. Accents, background music, several people talking over each other and technical vocabulary are where any speech model gets weaker, and this one is no exception. The transcript below the video is editable for exactly that reason: clicking a word and correcting it changes the burned-in caption and every file you download, and it keeps the timing the model worked out.
Whisper handles around a hundred, with auto-detect on by default, and eighteen of the common ones are in the picker so you can name the language and skip detection. Accuracy is best in English. If another language comes back rough, switch to the maximum accuracy model, which is better outside English and costs a 759 MB download and several times the running time.
The first run downloads the speech model: 290 MB on a machine with a GPU, 206 MB on the slower fallback path. Your browser caches it, so every clip after that skips the download entirely and spends its time on the listening alone. Come back tomorrow and the model is still there.
Up to 500 MB per file, and no limit on how many you do. Length is bounded by your own device rather than by us: everything is held in memory while it works, so a long clip on a phone is the case that struggles. If you are captioning something long, do it on a computer.
Yes. The video is yours, it never leaves your device, and the open source model behind the tool puts no conditions on what you do with the result. Post it, sell it, run it as an ad.
Those upload your video to a GPU server, which is why they cost what they cost, cap your minutes and put a watermark on the free tier. Here the same work happens on hardware you already own. What you give up is honest: a large hosted model in a datacentre still wins on difficult audio, and a phone will be slower than their servers. What you keep is your footage, your minutes and a clean export.
Start for Free

No credit card required • Set up in under 3 minutes

Alexandro - Founder
PublishQ

— me 👋

Hi, I'm Alexandro 👋

I left my Software Engineer role at Amazon to build tools that solve real problems — the kind big companies ignore because they read spreadsheets instead of using their own products.

I was spending over an hour daily just scheduling 2 shorts across 3 platforms — logging in, reformatting, uploading one by one. That felt broken. So I built PublishQPublishQ . Now I create 4 shorts in 3 minutes and schedule them to 4 platforms in under 30 seconds.

PublishQ is bootstrapped. No investors, no vanity metrics. I build what actually helps you — because I use it every day myself.

Thank you,

Alexandro

Captions burned in. Now get it posted.

Use PublishQPublishQ to schedule that clip to every account you manage from one calendar. No more uploading the same video five times.

Connect Your Accounts

Add once, use forever

Schedule Posts

Plan your content ahead

Multi-Platform

Post everywhere at once

Start for Free

No credit card required • Set up in under 3 minutes

Plans from $14/mo

2 months free on yearly billing. Upgrade or cancel anytime.

Free
$0/mo
Solo
$14/mo
Starter
$29/mo
Professional
$49/mo
Business
$99/mo
Compare all plans and features