Subtitles & Captions

Subtitles & Captions gigs from Buxonline freelancers, starting at $1.

No gigs in this category yet.

About subtitles & captions

Subtitles and captions are text overlays that display spoken dialogue, sound effects, and other audio information during video playback. Subtitles typically translate foreign-language speech for viewers who speak a different language, while captions transcribe all relevant audio—including speaker identification, music cues, and ambient sounds—for viewers who are deaf or hard of hearing. Both appear synchronised to the video, timed to match when words are spoken or sounds occur.

Creating them involves transcribing audio, segmenting text into readable chunks, timing each segment to appear and disappear at precise moments, and formatting the result according to technical specifications. The work requires balancing readability—keeping line length, display duration, and reading speed comfortable—with accuracy to the source material. Subtitle files are delivered in specific formats like SRT, VTT, or SBV, each with slightly different syntax for encoding timing and styling information.

Poor work shows up immediately: text that flashes by too quickly, lines broken mid-phrase forcing awkward reading patterns, mistimed cues that lag behind the speaker, or translations that miss cultural context. Well-executed subtitles feel invisible, allowing viewers to follow dialogue naturally without conscious effort, while maintaining enough screen time that even slower readers can keep pace without pausing.

Guides related to subtitles & captions

Subtitles & Captions — questions and answers

What reading speed should subtitles target, and how does that affect timing?
Most subtitles aim for 17 to 21 characters per second, which allows comfortable reading without rushing. Slower-paced content can drop to 15 characters per second, while fast dialogue may push to 23, though anything beyond that risks losing viewers. This reading speed directly determines how long each subtitle remains on screen and often requires condensing verbose speech.
How are line breaks decided when a sentence spans multiple subtitle frames?
Line breaks should respect linguistic units—keeping subjects with verbs, prepositions with their objects, and avoiding splits between adjectives and nouns. When a sentence continues across frames, the break typically happens at a natural pause like a comma or conjunction. Poor breaks force readers to mentally hold incomplete thoughts, disrupting comprehension and making the experience feel clumsy.
What's the difference between open captions and closed captions?
Open captions are burned directly into the video image and cannot be turned off—every viewer sees them. Closed captions exist as a separate text track that viewers can toggle on or off. Open captions guarantee visibility regardless of platform or player support, while closed captions offer flexibility and are required for accessibility compliance in many jurisdictions.
Do subtitle files work across different video platforms, or do they need reformatting?
Most platforms accept common formats like SRT or VTT, but each may interpret styling differently or impose character limits and timing restrictions. YouTube, Vimeo, Facebook, and broadcast systems often require adjustments to line length, positioning, or metadata. A file that works perfectly on one platform may need tweaking for another, particularly for font rendering or special characters.
How is speaker identification handled in captions when multiple people talk?
Captions typically use conventions like em dashes, chevrons, or colour coding to distinguish speakers. When two people speak in quick succession, each line might start with a dash. For overlapping dialogue, both lines appear simultaneously, positioned differently or colour-coded. The method depends on caption format capabilities and platform support—SRT handles basic text, while advanced formats allow styling.
Should sound effects and music be described in captions, and how detailed should that be?
Captions for deaf and hard-of-hearing audiences include relevant non-speech audio using bracketed descriptions like [door slams] or [tense music playing]. The detail depends on narrative importance—a doorbell that triggers action gets captioned, while ambient background noise might not. Music is described by mood or genre when it affects tone, and lyrics are transcribed when audible and relevant.