Free tool · Translation

Free YouTube transcript translator

Paste a link, pick a language, and read the video's transcript in it. Timestamps survive the trip, so you can jump back into the video from any line — or switch to the SRT view and copy a subtitle file that is already stamped.

Translate into

The chips fill the box — type over them for anything else. The second field is the caption language to read before translating, so a Spanish video going to German needs es here.

Translation runs on the tightest allowance on the site because it spends model tokens on top of the caption fetch. A free account gives you your own allowance, downloads and an API key.

How it works

This is a two-step tool, and it is worth knowing which two steps. First it asks YouTube for the caption track attached to the video — the same track the CC button shows, whether a human wrote it or YouTube's speech recognition did. Then it sends that text through a language model and gives you the result, line by line, with each line still carrying the start time of the cue it came from.

What it does not do is listen. There is no speech-to-text step here. If the uploader disabled captions and automatic captions never ran, there is no text to translate and the tool will say so plainly rather than inventing something. That is also why it is fast for its category: reading an existing track and translating it is much less work than transcribing an hour of audio from scratch.

The source-language field matters more than it looks. It tells the tool which caption track to read before anything is translated. Left at en it reads the English track — so if you want to translate a YouTube video to German from a Spanish original, set the source to es first. Getting this wrong is the single most common reason a translation comes back empty.

Why this tool has a tighter daily limit

Every other free tool here gives you more runs per day than this one, and the reason is simple enough to just say: translation costs money to run and plain extraction does not.

Pulling a transcript is a fetch. We ask YouTube for the caption track, hand you the text, and the only thing spent is bandwidth on our proxy. Translating that transcript does the same fetch and then pushes every line through a language model, and model tokens are billed per request. A long video is a lot of tokens. So a translation is counted against the AI actions allowance — deliberately the tightest one on the site — while a plain extraction is not.

The tool shows you the allowance the server reports, along with how much of it is left, so you are never guessing. A free account moves you off the shared anonymous pool onto your own, and comes with an API key if you would rather do this from a script.

Downloads need an account

Reading and copying the translation is open to anyone — including the finished SRT text. Saving it as a real file is the one thing behind the free signup, along with a bigger allowance and bulk access through the API.

What the output is good for — and what limits it

The quality ceiling here is the source captions, not the translation. If a human wrote and punctuated the original track, you will get clean, readable sentences in the target language. YouTube's automatic captions are a different animal: they arrive as an unpunctuated stream with no capitals, no sentence boundaries and no speaker labels, and proper nouns are a coin flip. Translating that gives you something readable and useful, but it carries every one of those gaps across the language barrier and sometimes adds a new one, because a model guessing at sentence boundaries will occasionally guess wrong.

So: for following along with a talk, skimming a lecture in a language you do not read, checking whether a video is worth watching, or building a searchable archive of YouTube subtitles in another language, this is genuinely good. For subtitles that go out under a brand, it is a first draft that saves you the boring 80% — not a replacement for someone who speaks the language.

The timestamps are what make it more than a wall of text. Each translated line keeps the timing of its original cue, so the timestamped view links back into the video at the exact second, and the SRT view is a working subtitle file — the server renders those cues, not the browser, so what you copy is exactly what the file contains. Paste it into a subtitle editor, a video editor, or a player and it behaves like any other SRT. Bear in mind that translated text changes length: German runs long, Japanese usually short, and a cue timed for the original can feel cramped. The timings are left exactly where the source put them rather than being quietly stretched.

Three ways to read the same result

One request produces all three views, and switching between them costs nothing — no second run, no second charge against your allowance.

ViewWhat you getGood for
TimestampedOne line per caption cue, each stamped and linked to that second of the videoReading along, jumping to a moment, quoting with a source
Plain textThe whole translation as continuous prose, no stampsPasting into a doc, summarising, translating further
SRTA finished .srt file, rendered server-side with real cue numbering and timingsSubtitle editors, video editors, players

Copy works on whichever view is showing. To download the translated SRT from YouTube as an actual file rather than copied text, you need a free account — that button is the one thing signup unlocks here.

Questions

Does this translate the audio, or the captions?

The captions. The tool asks YouTube for the caption track that already exists on the video, then translates that text. It never listens to the audio, so a video with captions switched off — and no automatic ones — has nothing here to work with, and you will get a clear message rather than a silent empty result. This also means the translation inherits whatever the caption track got right or wrong in the first place.

Why is the daily allowance smaller than for a plain transcript?

Because translation costs more to run. A plain extraction is a fetch: we ask YouTube for the caption track and hand it back, and the only resource it burns is bandwidth. A translation does that fetch and then sends every line through a language model, which spends tokens we pay for per request. That is the whole reason this tool sits on the tightest allowance on the site — two runs a day for anonymous visitors, counted against the AI actions bucket rather than the general one. A free account gives you your own allowance instead of sharing the anonymous pool.

Which languages can I translate into?

Anything you can name. The chips cover the thirteen requested most often — German, Spanish, French, Portuguese, Italian, Dutch, Polish, Turkish, Arabic, Hindi, Japanese, Korean and Chinese — but the field underneath is free text up to 40 characters, so a language name or an ISO code both work. Type 'de' or 'German' and you get the same thing. Less common languages will work as well as the model behind them does, which is honestly variable, so check a few lines before you trust a whole file.

Is the translation good enough to publish?

Treat it as a solid draft, not a finished localisation. It is machine translation of machine-or-human captions, and two things limit it: the quality of the source track, and the fact that the model translates line by line without seeing your video. Names, jargon and jokes are where it slips. For understanding a video, following along, or building a searchable archive, it is more than good enough. For subtitles that go out under your name, have someone who speaks the language read them first.

Do the timestamps still line up after translating?

Yes. Every translated line keeps the start time and duration of the caption cue it came from, which is why the SRT view drops straight into a subtitle workflow. Translated text does change length — German runs longer than English, Japanese usually shorter — so a line can feel tight or roomy against its cue. The timings are not moved to compensate; a subtitle editor is the right place to nudge them if reading speed matters.