Captions Generator
Drop in a clip. Every word gets timed, styled and burned into a finished MP4, right here in this tab.
Drop your video here
MP4, MOV, WebM or any audio file. It is read by this page and never uploaded.
Beast
A looping preview of the caption styles this tool produces, currently showing the Beast style, and cycling through: No watermark, ever, Nothing is uploaded, It all runs on your device, No sign up, no account, Caption as much as you like, Fix any word in one click, A file for your editor, Or the finished MP4, Then post it with PublishQ.
Nothing to pay, now or later
Every style, the MP4 export and all five subtitle formats. No account, no credits, no minutes to run out of, and no watermark on anything you make.
Your video never leaves your device
It is read by this page and nothing is uploaded. Open your network tab and watch: the speech model comes down to you once, and nothing goes back.
No limit on how much you caption
No file size cap, no queue, no daily allowance. The work happens on your machine, so there is nothing on our side to meter.
Ten animated styles
Word-by-word highlight, karaoke, a sliding pill, neon, lava, typewriter, bounce and more, with the position, the size and the words per line yours to set.
Fix a word in one click
The transcript is editable, so a name or a brand the model mishears is corrected once and lands in the burned video and in every file you download.
A file for whichever editor you use
SRT for CapCut and Premiere, iTT for Final Cut Pro, ASS with per-word karaoke timing, VTT for the web, and the plain transcript for show notes.
The finished video, captions burned in
Hardware-encoded in your browser and downloaded straight to you. No render queue, no email when it is done, no watermark in the corner.
Then post it everywhere, if you want
Optional, and free to ignore. PublishQ can schedule the finished video to every account you post from, out of one calendar.
This tool is free, forever. No accounts, no paywalls, no catch.
I build these solo. If you want to support my work, donating a coffee would be highly appreciated! ☕
Your video never leaves this tab
This is the part worth understanding, because it decides everything else about the tool. There is no server in the path. The file is opened by the page, your browser decodes its audio, and a speech recognition model runs in a background thread on your own hardware. The only thing that crosses the network is the model itself, coming down to you the first time you use the page.
You can check that rather than take our word for it. Open your browser network tab, drop a clip in, and watch: the model arrives, and nothing goes back the other way. Once it has arrived, the tool keeps working with your connection switched off.
Two things follow from it. Nothing is metered, because serving you costs almost nothing, so there is no reason to cap your minutes or hold the clean export behind a payment. And footage you would rather not hand to a company you have never met stays on the machine it is already on: an unreleased edit, a client project under embargo, a medical or legal recording, or simply your own face.
How a word gets its own moment on screen
A speech model returns text, not timings. Getting the animation to land on the right syllable takes a per-word start and end, and that is a separate thing the model has to be asked for. Whisper produces it by aligning the attention it paid to the audio against the audio itself, which places each word to within a fraction of a second. That list of words and times is the whole input to everything you see afterwards.
Those words are then grouped into the pages you actually read. A page ends when it is full, when the speaker stops, or when the sentence does, and it stays on screen until the next one begins so captions never blink off mid-thought. Set how many words share a page and the grouping is recomputed instantly, because the timings never change.
The preview and the export are the same renderer, called twice. What you approve on screen is what gets encoded, frame by frame, into the MP4. That is why there is no render queue to wait in and no surprise when the file lands.
Which file to take into which editor
If you are posting straight from here, take the MP4: the captions are already in the picture and no platform can strip them. If you are cutting the video somewhere else first, take a subtitle file, and which one depends on where it is going.
- SRT is the one that works everywhere. CapCut, Premiere Pro, DaVinci Resolve, YouTube, Final Cut Pro, VLC. Text and timings, nothing else.
- ASSis the only one that keeps the per-word timing, in karaoke tags, so the word-by-word animation can be rebuilt rather than flattened into lines. Anything built on libass reads it: VLC, mpv, and ffmpeg's subtitles filter. The editors are stricter, and none of the four below lists it.
- iTTis Apple's own format and the right answer for Final Cut Pro, which accepts SRT, iTT and CEA-608 and nothing else. It carries positioning and styling as well as the words.
- VTT is the web standard, for a subtitle track on your own site or a YouTube upload.
- Text is the transcript on its own, which is what you want for a description, a blog post or show notes.
Getting better captions on the first pass
- Name the language instead of leaving detection to guess. It is one dropdown and it removes a whole class of error, especially on a clip that opens with music before anyone speaks.
- Read the transcript before you export. Names, brands and technical words are what a speech model gets wrong, and fixing one is a click. The timing it worked out is kept.
- Fewer words per page for fast delivery, more for a calm voice. Three or four suits most short-form talking; one word at a time is the aggressive look, and it needs speech to match.
- Keep captions clear of the interface. Every app puts its own buttons and text over the bottom of a vertical video, so drag the position up until the preview looks safe. You can check the exact covered area for each platform in the post previewer.
- Caption on a computer if the clip is long. Everything is held in memory while it works, so length is bounded by your device rather than by any limit we set.
Where the captions are going
Which file to take depends on what is at the other end, and the editors disagree with each other about what they will accept. These pages say what each one takes and why.
Recording the clip in the first place? The online teleprompter scrolls your script while you look at the lens, and your phone can drive it. Once the captioned MP4 is ready, PublishQ schedules it to every account you manage from one calendar.
Frequently Asked Questions
Everything you need to know about Captions Generator from PublishQ
No credit card required • Set up in under 3 minutes
Similar Tools
Explore more tools to enhance your social media content creation

Carousel generator
Make a carousel post in your browser. Pick a template, write the slides, download them at full size. No signup, no watermark, no design skills.

AI Background Remover
Cut the background out of a photo without uploading it anywhere, then download the full-quality transparent PNG. No downscaled preview, no signup.

Solid Background Remover
Clear a solid background of any colour in about a second, with no AI model to download. Nothing is uploaded and the PNG keeps its full quality.

Post Previewer
See exactly how your posts will look on Twitter, Instagram, LinkedIn, TikTok, and more before you publish. No signup required.

Online Teleprompter
Read your script straight down the lens while the words scroll themselves. Open it on a second device, and your phone becomes the remote.

Media Downloader
Paste a link to a post and save the photo or video behind it. One click, free, with no ads and nothing to install.
YouTube Thumbnail Downloader
Paste a YouTube link and every size of its thumbnail appears at once, up to the full 1280 by 720. Save any of them, with no watermark.

— me 👋
Hi, I'm Alexandro 👋
I left my Software Engineer role at Amazon to build tools that solve real problems — the kind big companies ignore because they read spreadsheets instead of using their own products.
I was spending over an hour daily just scheduling 2 shorts across 3 platforms — logging in, reformatting, uploading one by one. That felt broken. So I built PublishQ . Now I create 4 shorts in 3 minutes and schedule them to 4 platforms in under 30 seconds.
PublishQ is bootstrapped. No investors, no vanity metrics. I build what actually helps you — because I use it every day myself.
Thank you,
Alexandro
Captions burned in. Now get it posted.
Use PublishQ to schedule that clip to every account you manage from one calendar. No more uploading the same video five times.
Connect Your Accounts
Add once, use forever
Schedule Posts
Plan your content ahead
Multi-Platform
Post everywhere at once
No credit card required • Set up in under 3 minutes
Plans from $14/mo
2 months free on yearly billing. Upgrade or cancel anytime.
- Free
- $0/mo
- Solo
- $14/mo
- Starter
- $29/mo
- Professional
- $49/mo
- Business
- $99/mo
Explore related
AI-Generated Shorts in 30 seconds
Turn an idea into a ready-to-post short video. Pick a template, let AI handle the rest, and schedule it — all without leaving PublishQ.
Write Once, Post Everywhere
Preview how your post looks on each platform, customize the text per account if needed, then publish or schedule to all of them at once.
How Much Does YouTube Pay for 1,000 Views in 2026?
YouTube pays $0.50 to $29 per 1,000 views depending on niche, country, and content type. Full RPM breakdown with real data by niche and country.
