Captions for Premiere Pro

Drop your clip in here, get every word timed, and download an SRT that drops straight into Premiere Pro as a caption track. The work happens on your device.

Drop your video here

MP4, MOV, WebM or any audio file. It is read by this page and never uploaded.

Beast

A looping preview of the caption styles this tool produces, currently showing the Beast style, and cycling through: No watermark, ever, Nothing is uploaded, It all runs on your device, No sign up, no account, Caption as much as you like, Fix any word in one click, A file for your editor, Or the finished MP4, Then post it with PublishQ.

Premiere reads SRT and TTML, and it makes captions a real track

Adobe publishes the list, and it is longer than most: as sidecar files Premiere takes SCC, MCC, XML, STL, SRT and DFXMP, and on the XML side it reads W3C TTML (also called DFXP), SMPTE-TT and EBU-TT. Advanced SubStation Alpha is on none of those lists, so bring the SRT. What arrives is not a graphic pinned to a clip but a caption track in the Captions panel, which you can restyle, split, retime and export again in whichever delivery format the job wants.

That is the real reason to bring a file rather than a burned-in video into Premiere. The moment the words are a track, the whole caption toolset applies to them, including per-word captions and the styling controls, and you keep the ability to hand a broadcaster an SCC or a streamer a TTML from the same timeline.

Premiere also has Speech to Text of its own, and it is good. Where this page earns its place is when you would rather not send the audio anywhere, when the clip is not in a language your Premiere install has a language pack for, or when what you actually want is the animated word-by-word look burned into the picture, which a caption track cannot do.

Getting captions into Premiere Pro

Take the SRT. It is on Adobe's own list of supported caption sidecar files and it lands as an editable caption track.

  1. Drop your clip above and let the words be timed on your device.
  2. Fix any names or jargon in the transcript before you export.
  3. Download the SRT, or burn the captions into an MP4.
  4. In Premiere, open the Captions panel and import the SRT onto your sequence.

Captioning for somewhere else

Same transcription, and the file to take differs by destination. These pages say which one and why.

This tool is free, forever. No accounts, no paywalls, no catch.

I build these solo. If you want to support my work, donating a coffee would be highly appreciated! ☕

Frequently Asked Questions

Everything you need to know about Captions for Premiere Pro from PublishQ

Adobe lists SCC, MCC, XML, STL, SRT and DFXMP as sidecar files, plus W3C TTML (DFXP), SMPTE-TT and EBU-TT on the XML side. SRT is the one to take from here: it is the most portable of them and it arrives as an editable caption track rather than a graphic.
Yes, if you bring the SRT. It lands in the Captions panel as a proper caption track, so you can restyle it, resplit lines, nudge timings and export it again in another format. If you bring the burned-in MP4 instead, the captions are pixels and are no longer text, which is the trade you make for them being impossible to drop.
Drop the video onto this page. A speech recognition model runs inside your browser, works out what is said and when each individual word is spoken, and the captions appear over your clip straight away. Pick a style, fix any word with a click, then either burn the captions into a new MP4 or download a subtitle file. There is no account and nothing to install.
No. The file is opened by the page, its audio is decoded by your own browser, and the transcription runs in a background thread on your own hardware. Open your network tab while it works and you will see the speech model coming down to you once, and nothing going back up. After that first download the tool keeps working with your connection switched off.
Every style, the MP4 export and all five subtitle formats cost nothing, and nothing is written onto your video. Because the work happens on your device there is no per-video server cost to recover, so there is nothing to meter and no better version of the output held back. PublishQ makes its money from the social media scheduler.
Instead of a whole sentence sitting on screen, each word lights up at the exact moment it is spoken. The movement keeps the eye on the text, and the text keeps the viewer through the first few seconds where most short-form video is lost. It needs a timing for every single word rather than for each line, which is what the model here produces.
Ten of them, and all ten are free: Normal, Beast, Karaoke, Pill, Subtitle, Neon, Lava, Boxed, Typewriter and Bounce. They differ in what happens on the spoken word, which is the part that does the work: a colour change, a pop, a pill sliding behind it, a glow, a gradient, or a word arriving one at a time with a caret waiting for the next. You can also set the vertical position, the type size and how many words sit on screen at once.
Chrome or Edge on a computer, because those include the WebCodecs hardware video encoding this uses. That is also why a burn-in finishes in well under the time the clip takes to play. On Safari and Firefox the preview and every subtitle file still work, so you can generate an SRT there and burn it in your editor.
SRT, VTT, ASS, iTT and plain text. SRT is the one that works everywhere. VTT is the web standard for an HTML5 video track or a YouTube upload. ASS is the only one that keeps the per-word timing, through karaoke tags, so an editor can rebuild the animation. iTT is Apple's own format, which is what Final Cut Pro takes natively. Plain text is just the transcript, for a description or a blog post.
Yes, and the format to pick differs. CapCut names SRT and TXT as the subtitle files it imports, and that import lives in CapCut Desktop and CapCut Web rather than the mobile app. Premiere Pro lists SRT among its supported sidecar files, and DaVinci Resolve imports SRT on the free version as well as Studio. Final Cut Pro accepts SRT, iTT and CEA-608 and does not read ASS at all, so use the iTT file there. Each of those has a page of its own with the steps on it.
Clear speech in English comes back close to word perfect. Accents, background music, several people talking over each other and technical vocabulary are where any speech model gets weaker, and this one is no exception. The transcript below the video is editable for exactly that reason: clicking a word and correcting it changes the burned-in caption and every file you download, and it keeps the timing the model worked out.
Whisper handles around a hundred, with auto-detect on by default, and eighteen of the common ones are in the picker so you can name the language and skip detection. Accuracy is best in English. If another language comes back rough, switch to the maximum accuracy model, which is better outside English and costs a 759 MB download and several times the running time.
The first run downloads the speech model: 290 MB on a machine with a GPU, 206 MB on the slower fallback path. Your browser caches it, so every clip after that skips the download entirely and spends its time on the listening alone. Come back tomorrow and the model is still there.
Up to 500 MB per file, and no limit on how many you do. Length is bounded by your own device rather than by us: everything is held in memory while it works, so a long clip on a phone is the case that struggles. If you are captioning something long, do it on a computer.
Yes. The video is yours, it never leaves your device, and the open source model behind the tool puts no conditions on what you do with the result. Post it, sell it, run it as an ad.
Those upload your video to a GPU server, which is why they cost what they cost, cap your minutes and put a watermark on the free tier. Here the same work happens on hardware you already own. What you give up is honest: a large hosted model in a datacentre still wins on difficult audio, and a phone will be slower than their servers. What you keep is your footage, your minutes and a clean export.
Start for Free

No credit card required • Set up in under 3 minutes

Alexandro - Founder
PublishQ

— me 👋

Hi, I'm Alexandro 👋

I left my Software Engineer role at Amazon to build tools that solve real problems — the kind big companies ignore because they read spreadsheets instead of using their own products.

I was spending over an hour daily just scheduling 2 shorts across 3 platforms — logging in, reformatting, uploading one by one. That felt broken. So I built PublishQPublishQ . Now I create 4 shorts in 3 minutes and schedule them to 4 platforms in under 30 seconds.

PublishQ is bootstrapped. No investors, no vanity metrics. I build what actually helps you — because I use it every day myself.

Thank you,

Alexandro

Captions burned in. Now get it posted.

Use PublishQPublishQ to schedule that clip to every account you manage from one calendar. No more uploading the same video five times.

Connect Your Accounts

Add once, use forever

Schedule Posts

Plan your content ahead

Multi-Platform

Post everywhere at once

Start for Free

No credit card required • Set up in under 3 minutes

Plans from $14/mo

2 months free on yearly billing. Upgrade or cancel anytime.

Free
$0/mo
Solo
$14/mo
Starter
$29/mo
Professional
$49/mo
Business
$99/mo
Compare all plans and features