Skip to main content
To all resources

Subtitles

September 4, 2026

How to Auto-Generate Subtitles: Tools, Accuracy and the Review Step Everyone Skips

Automatic subtitle generation: a waveform turning into timed subtitle lines on a video player

Auto-generating subtitles means letting speech recognition turn a video's audio into a timed text file, usually SRT. On clean audio the first pass is genuinely good. The work that remains is not transcription — it is the four things every engine gets wrong in the same way, every time.

This is the practical version: which tool to use, what to expect from it, and what to fix before the file goes anywhere.

Key Takeaways

  • Automatic subtitles are a first draft, not a finished file. Budget review time, not correction time after someone complains.
  • Accuracy depends on audio quality far more than on the tool. Clean single-speaker audio transcribes well everywhere; noise, accents and overlapping speech degrade every engine.
  • The four predictable errors: product names, speaker changes, numbers with units, and lines that run too long to read.
  • Set the source language manually. Auto-detection is where multilingual videos go wrong.
  • If the video will be translated afterwards, generate subtitles in a tool that also translates — a handover between tools is where timing breaks.

Step 1 — Pick the tool

The choice is simpler than the tool lists suggest, because it comes down to one question: does the video need to exist in another language afterwards?

If it stays in one language, a standalone transcriber is the cheapest route. Editors like Descript and Premiere Pro generate subtitles directly, and web tools do it without an install. You get an SRT and you are done.

If it needs translating, generate the subtitles where the translation happens. Exporting an SRT from one tool and importing it into another is where timing drifts, because the second tool re-flows the text without the audio. Full video platforms — including Dubly — transcribe, translate and export in one pass, which removes that handover. The category comparison is on the alternatives page.

If nobody will watch with sound, you also need burned-in subtitles, not just a file. Social feeds autoplay muted and will not load a separate track.

Step 2 — Generate the first pass

Two settings decide most of the outcome, and both are usually one click away.

Set the source language explicitly. Auto-detection is reliable on monolingual audio and unreliable the moment a second language appears — a German presentation with English product terms is the classic failure case, and the engine will sometimes switch mid-sentence.

Use the best audio you have. If the video was edited from a separate audio track, feed the tool that track rather than the exported mix. Music beds and room tone cost accuracy that no amount of correcting gets back cleanly.

Expect the first pass to take roughly as long as the video, sometimes less.

Step 3 — Correct the four predictable errors

Every engine fails in the same four places. Reviewing for these four specifically is faster than reading the file top to bottom.

  • Product and brand names. Speech recognition maps unfamiliar words onto familiar ones. Your product becomes a common noun that sounds similar. This is the single most visible error, because it is the word your audience knows best.
  • Speaker changes. Automatic output rarely marks who is talking. In interviews and panels, two speakers merge into one paragraph and the viewer cannot follow.
  • Numbers, units and dates. "Fifteen hundred" comes out as words, "1.5" comes out as "one point five", and currency symbols get dropped. Anything a viewer might act on needs checking.
  • Lines that are too long. The engine breaks on pauses, not on reading speed. Any line beyond about two lines of text has to be split, or the viewer is still reading when it disappears.

A useful discipline: read the file at playback speed once, with the video muted. Errors that survive a silent read are the ones that matter.

Step 4 — Export the right format

SRT is the default and works nearly everywhere — social platforms, video players, editing software.

VTT is the browser format. Use it when the video plays in an HTML5 player and you need styling or positioning.

Burned-in subtitles are pixels, not text: they cannot be switched off, and they cannot be translated later without regenerating the video. Use them where a separate file will not be loaded, which in practice means muted autoplay feeds.

Keep the SRT even when you burn in. It is the source for every later translation, and regenerating it from a finished video means starting over.

What accuracy to actually expect

Vendors publish accuracy percentages and they are all measured on clean audio, so they tell you little about your material. The honest framing is comparative: on studio-quality single-speaker audio, current engines are close enough that the tool choice barely matters. On a recorded meeting with three people and a laptop microphone, all of them need heavy correction, and the difference between the best and the worst is a matter of degree.

What consistently helps, in order: better source audio, an explicitly set language, and a glossary of your own terms where the tool supports one. Dubly's glossary is the mechanism for the third — fixed terms stay fixed across every language rather than being re-guessed per line.

When automatic subtitles are not the goal

Generating subtitles is often a proxy for a different problem: the video needs to work for people who do not speak the language it was recorded in. Subtitles solve that partially — the viewer reads instead of listening, and attention splits.

If that is the actual job, the comparison worth running is subtitles against dubbing, not one subtitle tool against another. We wrote it up in AI dubbing vs subtitles, and the cost side is in what subtitle translation costs.

Conclusion

Auto-generating subtitles is a solved problem for the first 90% and an unsolved one for the last 10% — and the last 10% is the part your audience notices, because it contains your product names and your numbers.

Generate the draft in whatever tool fits the destination, review for the four predictable errors, export SRT even if you burn in, and keep the file. If the video is heading for another language, do the generating where the translating happens.

Related: How to translate subtitles · What subtitle translation costs

On clean, single-speaker audio, modern speech recognition is accurate enough that most corrections are cosmetic. On noisy recordings, strong accents or overlapping speakers, every engine degrades and the file needs real editing. Audio quality predicts the result better than the choice of tool.
Most tools transcribe in the source language first and translate afterwards, because translating unreliable transcription multiplies the errors. Generating and translating in the same tool avoids a handover where the timing usually breaks.
An SRT is a separate text file the player displays and the viewer can switch off or translate later. Burned-in subtitles are part of the picture — always visible, and not translatable without regenerating the video. Social feeds that autoplay muted need burned-in subtitles.
Yes. Every project transcribes the source audio and exports an SRT file plus burned-in subtitles in two styles with a custom accent colour — included in every plan. The same pass also produces the translated, lip-synced version if you need one.
Generation itself usually takes about as long as the video, often less. The review is the real time cost: plan roughly a quarter of the video's runtime for correcting names, speaker changes, numbers and over-long lines.

About the author

Leon Bach

Leon Bach

Growth Marketing Manager