Closed captions are the toggleable text track a viewer can switch on or off — dialogue plus the non-speech audio (music, sound effects, speaker labels) that a deaf or hard-of-hearing viewer would otherwise miss. They are one of three things creators sloppily call "captions," and confusing them costs reach and, increasingly, legal exposure. This guide draws the lines that matter: closed captions (a separate track the platform serves behind a CC button) versus open captions (burned into the picture, always on) versus subtitles (dialogue-only, for viewers who can hear but not follow the language) versus SDH (subtitle-delivered captions for the deaf and hard of hearing). It covers the delivery layer most creators never see — the CEA-608 and CEA-708 broadcast standards, the WebVTT and SRT sidecar files the web uses, and where auto-generated ASR tracks fit — and the accessibility law that turned captions from a courtesy into a requirement: the CVAA and FCC rules in the US, the ADA and Section 508, and the European Accessibility Act that became binding in June 2025. Then it gets practical. It explains why the answer for social video is almost always open captions even though the accessibility term is closed captions, gives a platform-by-platform decision (YouTube serves and indexes a real CC track; TikTok, Reels, and Shorts reward burned-in open captions in the safe zone), covers the accuracy problem that makes raw auto-captions non-compliant, and ends on the part that actually breaks at volume: captioning every clip, on-brand and correctly placed, across nine platforms, week after week, without a person retyping subtitles into a desktop editor one file at a time.
A closed caption is a text version of everything a viewer would hear in a video — the spoken dialogue plus the non-speech audio a deaf or hard-of-hearing person would otherwise miss: music cues, sound effects, laughter, a phone ringing, and who is speaking. The word "closed" is the important part. It means the text lives on a separate track that the viewer can switch on or off, almost always through a button labeled "CC." The captions are not part of the picture; they are data the player overlays on demand, which is why you can turn them off, restyle them, or have them translated.
That single property — toggleable, because the text is a separate track rather than baked-in pixels — is what distinguishes closed captions from everything else creators loosely call "captions," and it is the source of most of the confusion. It is worth separating three meanings up front. On Instagram or LinkedIn, "caption" usually means the post-body text under the media, covered in the caption glossary entry. On a TikTok, "captions" usually means the animated on-screen words synced to speech. And in the accessibility and broadcast world, "closed captions" means specifically the toggleable audio-transcript track this guide is about. Same word, three different things — and the distinction has real consequences for reach and for legal exposure.
The cleanest way to hold these apart is to ask two questions: can the viewer turn the text off, and does the text assume the viewer can hear? Closed captions are toggleable and assume the viewer cannot hear (so they include non-speech audio). Open captions are not toggleable — they are burned directly into the video frame as pixels, always visible, and cannot be removed. Subtitles assume the viewer can hear but not understand the language, so they carry dialogue only, usually translated. SDH — subtitles for the deaf and hard of hearing — is the hybrid: caption content (non-speech cues, speaker labels) delivered through the subtitle mechanism, which is what you get on most streaming services and Blu-rays because their delivery format is subtitle-based rather than the old broadcast caption stream.
The open-versus-closed split is the one that matters most for creators, because it maps directly onto a production decision. Closed captions can be turned off, restyled by the player, translated on the fly, and — critically — read by search engines, because the text exists as machine-readable data. Open captions have none of that flexibility, but they have one enormous advantage: they are indestructible. Because they are part of the picture, they survive a re-upload to a different platform, they render identically everywhere regardless of whether that platform even supports a caption track, and they play by default with no viewer action required. On a muted, autoplaying social feed, "plays by default" is the whole game — which is why the accessibility-standard term is closed captions but the thing most short-form creators actually ship is open captions.
Closed captions have to be encoded and carried somehow, and the format depends on where the video runs. On North American broadcast and cable, two standards dominate: CEA-608 (also called line 21, the older analog-era standard, limited to a monospaced style and a handful of positions) and CEA-708 (the digital-TV standard, which adds fonts, colors, sizing, and flexible placement). You will see these referenced whenever captions touch traditional TV or professional video-on-demand pipelines. On the web, captions travel as sidecar files — small text files that sit alongside the video and sync by timestamp. WebVTT (.vtt) is the native HTML5 format; SRT (.srt) is the ubiquitous simple format most tools import and export; TTML/IMSC shows up in broadcast-derived streaming. A sidecar file is what you upload to YouTube when you "add captions," and it is what a platform generates for you when it auto-captions with speech recognition.
Auto-generated captions sit inside this same layer. When a platform or tool produces captions automatically, it runs the audio through an automatic speech recognition (ASR) model — Whisper-class systems are the common engine — and emits a timed track, usually as VTT or SRT. That track is a real closed-caption file: toggleable, translatable, indexable. It is also, out of the box, frequently wrong in ways that matter, which is the accuracy problem covered further down.
Captions crossed from courtesy to requirement more than a decade ago, and the obligations have only widened. In the United States, the Twenty-First Century Communications and Video Accessibility Act (CVAA), enforced through FCC rules, requires captions on internet video that previously aired on US television — so a broadcaster cannot strip captions when a clip moves online. The Americans with Disabilities Act (ADA) and Section 508 create separate captioning obligations for many private businesses, places of public accommodation, government bodies, and federally funded organizations. The FCC also sets quality rules for the captions it governs: they must be accurate, synchronized with the audio, complete from start to finish, and properly placed on screen.
The obligation is not US-only. The European Accessibility Act became binding on 28 June 2025 and requires accessible audiovisual content — captions among them — across EU member states, pushing accessibility from something large broadcasters did to something a far wider set of publishers must now build in. The exact scope of who is covered varies by jurisdiction and by the nature of the publisher, so treat "does this specific video legally require captions" as a question for the relevant rules rather than something to infer from a general summary. The safe operating posture for anyone publishing at scale is to caption everything to a real quality bar, because the cost of doing so is low and the exposure of not doing so is rising.
Even where captions are not legally mandated, the reach case is decisive. The large majority of people watch social video with the sound off — feeds autoplay muted, and much of the viewing happens in places where sound is socially impossible. A video without visible text in that environment is a video most people scroll past before they ever consider unmuting. This is the argument for designing video for silence from the first frame, laid out in the captions-first video strategy guide and the retention-focused how-to. Beyond raw watchability, platform research has consistently found that captions lift completion, recall, and how much viewers like a clip, and completion is increasingly the exact signal the algorithm optimizes for.
There is a discovery angle specific to the closed (toggleable, machine-readable) form. Because a real caption track is text data, search engines and platform search can read it. On YouTube in particular, an accurate closed-caption file is indexable content that helps the video surface for spoken phrases that appear nowhere in the title or description. Burned-in open captions give you none of that — they are pixels, invisible to a crawler. That is the strongest single reason to supply a closed-caption track in addition to burning captions into the frame on long-form and on-demand video: the open captions win the muted feed, and the closed track wins the search index.
The practical resolution is not "pick one." It is knowing which job each form does and matching it to the platform. Open (burned-in) captions win wherever the video autoplays muted, gets cross-posted, and needs to look identical everywhere with full styling control — that is short-form social. Closed captions win wherever the viewer benefits from being able to toggle, translate, or the platform indexes the track — that is YouTube, other long-form, and any surface with a real accessibility requirement. On the highest-stakes content, do both: burn open captions for the feed and upload a closed-caption file for search and accessibility.
YouTube is the clearest case for a real closed-caption track: it serves captions behind the CC button, lets viewers auto-translate them, and indexes the text for search, so uploading an accurate .srt or .vtt is pure upside — and you can still burn open captions into a Short on top. TikTok, Instagram Reels, and YouTube Shorts are open-caption territory in practice: viewers watch muted, and burned-in captions in the 9:16 safe zone (the middle band, clear of the UI chrome at top and bottom) are what hold attention; each platform also has its own auto-caption feature, but styling and placement control is why creators burn their own. LinkedIn and X video autoplay muted in-feed and reward burned-in captions the same way. The per-platform mechanics of sizing, placement, and the description-versus-on-screen split are covered in the TikTok captions guide. For reaching non-native audiences, a translated closed-caption track (or a burned-in localized version) is the lever, covered in multilingual auto-translated captions and how AI video translation works.
Auto-captioning is now a baseline feature — every major platform and editor ships it, a shift examined in green screen and auto-captions are baseline features now. But "the platform can generate captions" is not the same as "the captions are correct." ASR reliably mishears jargon, brand names, acronyms, proper nouns, and fast or accented speech, and it does not punctuate or add non-speech cues on its own. Raw auto-captions therefore fail the very quality bar the accessibility rules demand — accurate, synchronized, complete — which means shipping them unedited is both a reach risk (a wrong word on screen reads as sloppy) and, for regulated content, a compliance risk. The rule of thumb: auto-generate to save the typing, then always proof, fixing the names and terms specific to your niche. A practical end-to-end workflow for this is in how to add captions to a video.
Captioning one video is a solved problem — open a desktop editor, auto-generate, fix the errors, burn it in, export. The problem is that a real content operation does not make one video. It makes a week of them across formats and platforms, and the naive workflow does not survive that multiplication: each clip gets captioned by hand, in a different tool, with inconsistent fonts and placement, and nobody keeps the closed-caption sidecar files organized alongside the burned-in versions. Captioning quietly becomes the bottleneck that caps how much video a small team can ship.
This is the gap Kompozy is built to close, and it is a production problem, not an editing one. Kompozy is a full AI content generation and multi-platform publishing engine, and its talking-video formats — Persona Shorts and Clipped Shorts — come captioned by default: each transcribes the spoken audio and burns the words into the 9:16 safe zone as open captions in your brand font and styling, as part of generating the clip rather than as a separate manual pass. The look stays identical across every clip because the caption styling is part of the render, not retyped per file. So the muted-feed reach case is handled automatically on the formats that carry spoken dialogue.
The reason the styling stays on-brand across hundreds of clips is the same reason the copy does: one Persona Brief governs voice and a banned-word filter governs terminology, so the words on screen match the words in the post match the words in the newsletter. From there, a single queue fans each captioned clip across the eight social platforms plus blog and email, behind a per-post review gate where a human approves what ships by default — which is exactly where you proof the auto-generated text before it goes out, turning the accuracy problem into a fast yes/no instead of a retype. Autopilot can take that approval step off your plate per source, gated instead by its own automated quality checks, once you trust it on that source. The boundary, stated honestly: if you need broadcast-grade CEA-708 compliance captions for a regulated TV pipeline, that is a specialist captioning workflow, not this. Kompozy owns the creator case — captioning every social and long-form clip, on-brand and correctly placed, at the volume a real publishing cadence demands.
"Captions" is three different things, and the money is in not confusing them. Closed captions are the toggleable, machine-readable track that serves accessibility, legal compliance, and search indexing. Open captions are the always-on burned-in pixels that win the muted social feed. Subtitles and SDH sit alongside for translation and for the deaf-and-hard-of-hearing case on subtitle-delivered platforms. The correct posture for most creators is not to choose but to match: burn open captions for the feed, upload a closed-caption track where the platform serves and indexes it, proof anything a machine transcribed, and — the part that actually decides whether this survives contact with a real posting schedule — caption every clip on-brand and at volume rather than one file at a time in a desktop editor.
Closed captions are a text version of a video's audio that the viewer can turn on or off, usually via a "CC" button. Unlike a plain transcript or foreign-language subtitle, they include non-speech audio — music cues, sound effects, and speaker identification — because they are written for someone who cannot hear the soundtrack at all. Technically they travel as a separate track (a sidecar file or an embedded data stream), not as pixels burned into the picture, which is what makes them toggleable.
The difference is who controls them. Closed captions are a separate track the viewer switches on or off, so they can be turned off, restyled, translated, or read by search engines. Open captions are burned directly into the video frame as pixels, so they are always visible and cannot be removed — which is exactly why most short-form social creators use them: they survive re-uploads, autoplay muted, and render identically on every platform regardless of whether that platform exposes a caption track.
No, though the words are used loosely. Subtitles assume the viewer can hear the audio but not understand the language, so they transcribe or translate dialogue only. Captions assume the viewer cannot hear at all, so they add non-speech information — sound effects, music, speaker labels. SDH (subtitles for the deaf and hard of hearing) is the hybrid: caption-style content delivered through the subtitle mechanism, common on streaming and discs where the delivery format is subtitle-based rather than the older broadcast caption stream.
Often, yes, depending on where the content runs and who publishes it. In the US, the CVAA and FCC rules require captions on internet video that previously aired on television, and the ADA and Section 508 create captioning obligations for many businesses, public bodies, and federally funded entities. The EU's European Accessibility Act, binding since 28 June 2025, requires accessible audiovisual content across member states. The rules also demand quality — captions must be accurate, synchronized, and complete, which raw auto-generated tracks frequently are not.
For short-form social video — TikTok, Reels, Shorts — burned-in open captions almost always win, because most viewers watch muted, the caption must survive cross-posting, and you want full control of styling and placement in the 9:16 safe zone. For YouTube and other long-form or on-demand surfaces, also supply a real closed-caption track: YouTube serves it behind the CC button, indexes it for search, and it lets viewers translate or turn it off. The strongest setup is often both — burned-in open captions plus an uploaded closed-caption file.
Closed captions are a text version of a video's audio — dialogue plus non-speech sound like music, effects, and speaker labels — delivered as a separate track the viewer can toggle on or off with a CC button. They differ from open captions, which are burned permanently into the picture, and from subtitles, which translate dialogue for viewers who can hear. For creators, closed captions serve accessibility, legal compliance, and search indexing, while burned-in open captions serve sound-off reach on short-form social feeds.
Get started → · ← All guides · Compare Kompozy vs other tools