A training video and seven languages
A client sends a thirty-minute training video and asks for subtitles in seven languages. Your agency does mostly documents. The PM finds a freelancer to transcribe the English, who returns a Word document with no timings. Someone else times it in Subtitle Edit. The SRT file is then sent to translators, some of whom handle subtitles well and some of whom translate the text without regard to line length. The Polish comes back with lines far too long to read in the time on screen.
Then the client asks for WebVTT for their learning platform as well as SRT for YouTube, and an extra burned-in version for social media.
Why subtitling is awkward
- Transcription and timing are separate skills and often separate people.
- Translators need to respect reading speed and characters per line, which normal CAT workflows do not check.
- Clients have different specs: line length, reading speed, number of lines, positioning.
- Output formats vary: SRT, WebVTT, platform-specific formats, burned-in video.
- Checking subtitles means watching the video, which is slow for seven languages.
What it costs
Rework when subtitles are too long to read or out of sync. PM time chasing transcribers, timers and translators. Clients who notice poor subtitles immediately, because everybody watches video. And a service you may be turning away, because it is too much trouble to run, while clients increasingly have video to translate.
| Stage | Manual | Pipeline |
|---|---|---|
| Transcript | Typed from scratch | Speech-to-text draft for a transcriber to correct |
| Timing | Set by hand | Draft timings from the audio, adjusted by a person |
| Translation | Text translated without limits | Limits shown and checked per subtitle |
| Checks | Watch every video | Reading speed, length and overlap checked automatically |
| Delivery | Converted by hand | All formats generated from one master |
How we build the subtitling pipeline
- The video is uploaded through your portal or intake, and the audio is run through a speech-to-text service such as OpenAI Whisper, Azure Speech or AWS Transcribe, producing a draft transcript with word timings.
- A transcriber corrects the draft, and the text is segmented into subtitles using the client's spec for line length, number of lines and minimum and maximum duration.
- A person adjusts timings in a subtitle editor, such as Subtitle Edit, working from the draft rather than from nothing.
- The master subtitle file is sent for translation in your CAT tool, where it supports subtitle formats, or in a subtitle editor, with character limits shown for each subtitle.
- Translated files are checked automatically for reading speed, line length, overlaps and gaps, against the client's spec, and issues go back to the translator.
- A reviewer watches the video with subtitles for a final check, with flagged subtitles highlighted so they know where to look first.
- Output files are generated in every format the client needs, and burned-in versions are rendered where requested.
Speech-to-text drafts save typing but are not accurate enough to ship, especially with accents, jargon or several speakers. A person always corrects them.
What changes
Subtitling becomes a repeatable process rather than a special project. Transcribers and timers correct drafts instead of starting from nothing. Translators see the limits as they work. Mechanical problems are caught before a reviewer spends time watching. And delivering several formats is a click rather than a conversion job. Your agency can say yes to video work with confidence.
Does subtitling feel like this?
- Transcripts are typed from scratch.
- Translated subtitles come back too long to read.
- Converting between SRT, WebVTT and other formats is manual.
- Checking means watching every video in every language.
- You turn down video work because it is hard to run.