How to get YouTube transcripts with an API (Python, Node.js and curl)
YouTube transcripts are one of the most useful text sources on the internet. Conference talks, tutorials, earnings calls, podcasts and product demos all end up on YouTube, and most of them have captions. If you can get that text into your application, you can search it, summarise it, translate it, or feed it to a language model.
This guide shows how to do that reliably from code. We'll cover where transcripts come from, why the popular do-it-yourself approach breaks as soon as you deploy it, and then walk through working examples in curl, Python and Node.js, including languages, subtitle files, playlists and chunking for retrieval.
Where YouTube transcripts actually come from
A YouTube "transcript" is a caption track. Every video can have zero or more of them:
- Creator-uploaded captions: written or corrected by the channel. Usually the most accurate.
- Auto-generated captions: produced by YouTube's speech recognition. Available for many videos in major languages, with no punctuation guarantees and the occasional misheard word.
- Auto-translated captions: YouTube can machine-translate an existing track into other languages on the fly.
Each caption track is a list of short segments with a start time and a duration. The "transcript" most people want is those segments joined into one block of text, but the timestamps are valuable too: they let you link an answer back to the exact second in a video.
Two consequences follow. First, if a video has no caption track, there is no transcript to fetch. Getting text from those videos needs speech-to-text on the audio, which is a different (and much more expensive) job. Second, the language you get depends on which tracks exist, so your code should state a preference and handle fallbacks.
The official YouTube Data API does have a captions endpoint, but downloading caption tracks through it generally requires OAuth authorisation tied to the video's owner, so it doesn't help when you want transcripts of other people's public videos.
Why the DIY approach breaks on servers
The usual first attempt is an open-source library that fetches the caption data the YouTube web player uses. In Python that's typically youtube-transcript-api. It works beautifully on a laptop:
from youtube_transcript_api import YouTubeTranscriptApi
transcript = YouTubeTranscriptApi().fetch("jNQXAC9IVRw")
Then you deploy it to AWS, GCP, a VPS or a serverless function, and requests start failing with errors like RequestBlocked or IpBlocked. YouTube is much stricter with traffic from cloud provider IP ranges than with home connections. The library's own documentation has a section about working around IP bans, and the usual fix is routing requests through rotating residential proxies, which you then have to buy, configure and monitor.
That's the real cost of the DIY route: not the ten lines of code, but keeping them working in production.
Getting transcripts through an API
With the Siftwright API you send the video URL and get the transcript back as JSON. The service connects directly and falls back to a proxy automatically when it needs to, so the same request works from your laptop, a container or a cloud function. Videos without captions come back with an error status and aren't billed.
All examples below assume your key is in an environment variable:
export SIFTWRIGHT_API_KEY="sw_live_..."
curl
curl https://siftwright.com/v1/youtube-transcript \
-H "Authorization: Bearer $SIFTWRIGHT_API_KEY" \
-H "Content-Type: application/json" \
-d '{"url": "https://www.youtube.com/watch?v=jNQXAC9IVRw", "outputFormats": ["text", "segments"]}'
The response contains a results array with one item per video, plus how many results were billed and how much of your plan is left:
{
"endpoint": "youtube-transcript",
"count": 1,
"billed": 1,
"results": [
{
"videoId": "jNQXAC9IVRw",
"status": "ok",
"title": "Me at the zoo",
"channelName": "jawed",
"durationSeconds": 19,
"language": "en",
"isAutoGenerated": false,
"wordCount": 39,
"text": "All right, so here we are, in front of the elephants ...",
"segments": [
{ "start": 1.2, "duration": 2.16, "text": "All right, so here we are, in front of the elephants" }
]
}
],
"usage": { "plan": "starter", "used": 1, "quota": 10000, "remaining": 9999, "period_end": "..." }
}
Python
import os
import requests
API = "https://siftwright.com/v1/youtube-transcript"
HEADERS = {"Authorization": f"Bearer {os.environ['SIFTWRIGHT_API_KEY']}"}
def get_transcripts(urls, formats=("text", "segments"), language="en"):
r = requests.post(
API,
headers=HEADERS,
json={"urls": list(urls), "outputFormats": list(formats), "language": language},
timeout=300,
)
r.raise_for_status()
body = r.json()
ok = [v for v in body["results"] if v["status"] == "ok"]
failed = [v for v in body["results"] if v["status"] != "ok"]
return ok, failed
ok, failed = get_transcripts([
"https://www.youtube.com/watch?v=jNQXAC9IVRw",
"https://youtu.be/dQw4w9WgXcQ",
])
for video in ok:
print(f"{video['title']}: {video['wordCount']} words ({video['language']})")
for video in failed:
print("No transcript:", video.get("url"), video["status"])
Two details matter here. The timeout is generous because the request is synchronous: batches can take a while. And the code separates successful items from failures instead of assuming every video has captions.
Node.js (18 or newer)
const API = "https://siftwright.com/v1/youtube-transcript";
async function getTranscript(url, outputFormats = ["text"]) {
const res = await fetch(API, {
method: "POST",
headers: {
Authorization: `Bearer ${process.env.SIFTWRIGHT_API_KEY}`,
"Content-Type": "application/json",
},
body: JSON.stringify({ url, outputFormats }),
signal: AbortSignal.timeout(300_000),
});
const body = await res.json();
if (!res.ok) throw new Error(`${body.error.code}: ${body.error.message}`);
const [video] = body.results;
if (video.status !== "ok") throw new Error(`No transcript: ${video.status}`);
return video;
}
const video = await getTranscript("https://youtu.be/jNQXAC9IVRw");
console.log(video.title, "-", video.text.slice(0, 200));
Errors come back as JSON with a stable code (for example invalid_api_key, quota_exceeded or rate_limited), so you can branch on the code rather than parsing messages. The full list is in the docs.
Choosing a language, or translating
language takes a comma-separated preference list. The first language that has a caption track wins; if none match, you get the first available track, and the response tells you which one was used:
{"url": "https://www.youtube.com/watch?v=VIDEO_ID", "language": "es,en"}
Check language, isAutoGenerated and availableLanguages on the result to decide whether it's good enough for your use case. By default, creator-uploaded captions are preferred over auto-generated ones (preferManual: true).
To machine-translate the transcript into a single target language regardless of the source, add translateTo:
{"url": "https://www.youtube.com/watch?v=VIDEO_ID", "translateTo": "de"}
Translation uses YouTube's built-in caption translation, so quality is similar to what you'd see in the player.
Subtitle files: SRT and VTT
If you're building subtitles rather than feeding text to a model, ask for srt or vtt. You get the complete file contents as a string that you can write straight to disk:
ok, _ = get_transcripts(["https://youtu.be/jNQXAC9IVRw"], formats=["srt", "vtt"])
video = ok[0]
with open(f"{video['videoId']}.srt", "w", encoding="utf-8") as f:
f.write(video["srt"])
with open(f"{video['videoId']}.vtt", "w", encoding="utf-8") as f:
f.write(video["vtt"])
Playlists and channels
Pass a playlist URL or a channel URL (/@handle or /channel/UC...) and it expands to its videos. Use maxVideosPerSource to cap how many (1 to 25 per source through the API; channels return the newest uploads first):
ok, failed = get_transcripts(
["https://www.youtube.com/@SomeChannel"],
formats=["text"],
)
Only videos that actually return a transcript are billed, so a channel where half the uploads lack captions costs half as much. One request accepts up to 10 sources. For bigger jobs, loop over batches and keep the rate limit (60 requests per minute per key) in mind.
Preparing transcripts for an LLM
Raw transcripts are long and unpunctuated in places. Two habits make them much more useful for retrieval-augmented generation (RAG) and summarisation.
Chunk by time, not by characters. Segments already carry timestamps, so group them into windows of roughly 60 to 90 seconds. Every chunk then maps to a real position in the video, and your application can cite it as a deep link:
def chunk_segments(video, window=75.0):
chunks, current, start = [], [], None
for seg in video["segments"]:
if start is None:
start = seg["start"]
current.append(seg["text"])
if seg["start"] + seg["duration"] - start >= window:
chunks.append({"start": start, "text": " ".join(current)})
current, start = [], None
if current:
chunks.append({"start": start, "text": " ".join(current)})
for c in chunks:
c["url"] = f"https://www.youtube.com/watch?v={video['videoId']}&t={int(c['start'])}s"
return chunks
Keep the metadata. Title, channel, publish date and duration are returned for free and make great filters in a vector store ("only talks from this channel", "only videos from this year").
For summarisation, the plain text field is usually all you need. For question answering with citations, embed the chunks above and store url alongside each one.
Handling failures gracefully
Transcripts fail for predictable reasons: the video is private or removed, it has no captions, or it's a live stream that hasn't finished. Every item carries a status, and failed items are never billed. A robust pipeline:
- Logs failed items with their status so you can see patterns.
- Doesn't retry "no captions" failures: they won't succeed later unless the creator adds captions.
- Retries transport errors (
upstream_error,upstream_timeout) with backoff; nothing is charged for those either.
Cost
A transcript counts as one result. On the Starter plan ($29/month for 10,000 results) that's $2.90 per 1,000 transcripts if you use the whole plan, falling to $0.67 per 1,000 on Enterprise. If you'd rather not subscribe, the same tool runs pay-per-event on the Apify Store at $3 per 1,000 transcripts. Our pricing page has a calculator, and we wrote up when pay-per-result beats a subscription with worked examples.
Summary
- Transcripts are caption tracks. No captions, no transcript.
- DIY scraping works locally but tends to get blocked from cloud servers.
- An API call with a URL gets you text, timestamped segments, SRT or VTT, plus metadata.
- Chunk by timestamps for RAG, and handle per-item failures instead of assuming success.
You can try a transcript request right now, without signing up, in the live demo.
Related guides
Website screenshot API in Python and Node.js: full-page, mobile, retina and PDF
How to capture website screenshots from code without running your own headless browser fleet. Working Python and Node.js examples for full-page, mobile, retina and PDF captures, late-loading pages, and a batch script with retries.
HTML to PDF at scale: print CSS, fonts, signed URLs and batch rendering
A practical guide to generating PDFs from HTML for invoices, reports and certificates. Print CSS that works, page sizes and breaks, waiting for fonts and charts, rendering private documents via signed URLs, and a batch pipeline for thousands of PDFs.
Get the next guide by email
Practical tutorials on transcripts, screenshots, news data and agent tooling. A couple of emails a month at most.