How to get YouTube transcripts with an API (Python, Node.js and curl)

By Siftwright team7 min read

YouTube transcripts are one of the most useful text sources on the internet. Conference talks, tutorials, earnings calls, podcasts and product demos all end up on YouTube, and most of them have captions. If you can get that text into your application, you can search it, summarise it, translate it, or feed it to a language model.

This guide shows how to do that reliably from code. We'll cover where transcripts come from, why the popular do-it-yourself approach breaks as soon as you deploy it, and then walk through working examples in curl, Python and Node.js, including languages, subtitle files, playlists and chunking for retrieval.

Where YouTube transcripts actually come from

A YouTube "transcript" is a caption track. Every video can have zero or more of them:

  • Creator-uploaded captions: written or corrected by the channel. Usually the most accurate.
  • Auto-generated captions: produced by YouTube's speech recognition. Available for many videos in major languages, with no punctuation guarantees and the occasional misheard word.
  • Auto-translated captions: YouTube can machine-translate an existing track into other languages on the fly.

Each caption track is a list of short segments with a start time and a duration. The "transcript" most people want is those segments joined into one block of text, but the timestamps are valuable too: they let you link an answer back to the exact second in a video.

Two consequences follow. First, if a video has no caption track, there is no transcript to fetch. Getting text from those videos needs speech-to-text on the audio, which is a different (and much more expensive) job. Second, the language you get depends on which tracks exist, so your code should state a preference and handle fallbacks.

The official YouTube Data API does have a captions endpoint, but downloading caption tracks through it generally requires OAuth authorisation tied to the video's owner, so it doesn't help when you want transcripts of other people's public videos.

Why the DIY approach breaks on servers

The usual first attempt is an open-source library that fetches the caption data the YouTube web player uses. In Python that's typically youtube-transcript-api. It works beautifully on a laptop:

from youtube_transcript_api import YouTubeTranscriptApi

transcript = YouTubeTranscriptApi().fetch("jNQXAC9IVRw")

Then you deploy it to AWS, GCP, a VPS or a serverless function, and requests start failing with errors like RequestBlocked or IpBlocked. YouTube is much stricter with traffic from cloud provider IP ranges than with home connections. The library's own documentation has a section about working around IP bans, and the usual fix is routing requests through rotating residential proxies, which you then have to buy, configure and monitor.

That's the real cost of the DIY route: not the ten lines of code, but keeping them working in production.

Getting transcripts through an API

With the Siftwright API you send the video URL and get the transcript back as JSON. The service connects directly and falls back to a proxy automatically when it needs to, so the same request works from your laptop, a container or a cloud function. Videos without captions come back with an error status and aren't billed.

All examples below assume your key is in an environment variable:

export SIFTWRIGHT_API_KEY="sw_live_..."

curl

curl https://siftwright.com/v1/youtube-transcript \
  -H "Authorization: Bearer $SIFTWRIGHT_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{"url": "https://www.youtube.com/watch?v=jNQXAC9IVRw", "outputFormats": ["text", "segments"]}'

The response contains a results array with one item per video, plus how many results were billed and how much of your plan is left:

{
 "endpoint": "youtube-transcript",
 "count": 1,
 "billed": 1,
 "results": [
  {
   "videoId": "jNQXAC9IVRw",
   "status": "ok",
   "title": "Me at the zoo",
   "channelName": "jawed",
   "durationSeconds": 19,
   "language": "en",
   "isAutoGenerated": false,
   "wordCount": 39,
   "text": "All right, so here we are, in front of the elephants ...",
   "segments": [
    { "start": 1.2, "duration": 2.16, "text": "All right, so here we are, in front of the elephants" }
   ]
  }
 ],
 "usage": { "plan": "starter", "used": 1, "quota": 10000, "remaining": 9999, "period_end": "..." }
}

Python

import os
import requests

API = "https://siftwright.com/v1/youtube-transcript"
HEADERS = {"Authorization": f"Bearer {os.environ['SIFTWRIGHT_API_KEY']}"}


def get_transcripts(urls, formats=("text", "segments"), language="en"):
    r = requests.post(
        API,
        headers=HEADERS,
        json={"urls": list(urls), "outputFormats": list(formats), "language": language},
        timeout=300,
    )
    r.raise_for_status()
    body = r.json()
    ok = [v for v in body["results"] if v["status"] == "ok"]
    failed = [v for v in body["results"] if v["status"] != "ok"]
    return ok, failed


ok, failed = get_transcripts([
    "https://www.youtube.com/watch?v=jNQXAC9IVRw",
    "https://youtu.be/dQw4w9WgXcQ",
])
for video in ok:
    print(f"{video['title']}: {video['wordCount']} words ({video['language']})")
for video in failed:
    print("No transcript:", video.get("url"), video["status"])

Two details matter here. The timeout is generous because the request is synchronous: batches can take a while. And the code separates successful items from failures instead of assuming every video has captions.

Node.js (18 or newer)

const API = "https://siftwright.com/v1/youtube-transcript";

async function getTranscript(url, outputFormats = ["text"]) {
  const res = await fetch(API, {
    method: "POST",
    headers: {
      Authorization: `Bearer ${process.env.SIFTWRIGHT_API_KEY}`,
      "Content-Type": "application/json",
    },
    body: JSON.stringify({ url, outputFormats }),
    signal: AbortSignal.timeout(300_000),
  });
  const body = await res.json();
  if (!res.ok) throw new Error(`${body.error.code}: ${body.error.message}`);
  const [video] = body.results;
  if (video.status !== "ok") throw new Error(`No transcript: ${video.status}`);
  return video;
}

const video = await getTranscript("https://youtu.be/jNQXAC9IVRw");
console.log(video.title, "-", video.text.slice(0, 200));

Errors come back as JSON with a stable code (for example invalid_api_key, quota_exceeded or rate_limited), so you can branch on the code rather than parsing messages. The full list is in the docs.

Choosing a language, or translating

language takes a comma-separated preference list. The first language that has a caption track wins; if none match, you get the first available track, and the response tells you which one was used:

{"url": "https://www.youtube.com/watch?v=VIDEO_ID", "language": "es,en"}

Check language, isAutoGenerated and availableLanguages on the result to decide whether it's good enough for your use case. By default, creator-uploaded captions are preferred over auto-generated ones (preferManual: true).

To machine-translate the transcript into a single target language regardless of the source, add translateTo:

{"url": "https://www.youtube.com/watch?v=VIDEO_ID", "translateTo": "de"}

Translation uses YouTube's built-in caption translation, so quality is similar to what you'd see in the player.

Subtitle files: SRT and VTT

If you're building subtitles rather than feeding text to a model, ask for srt or vtt. You get the complete file contents as a string that you can write straight to disk:

ok, _ = get_transcripts(["https://youtu.be/jNQXAC9IVRw"], formats=["srt", "vtt"])
video = ok[0]
with open(f"{video['videoId']}.srt", "w", encoding="utf-8") as f:
    f.write(video["srt"])
with open(f"{video['videoId']}.vtt", "w", encoding="utf-8") as f:
    f.write(video["vtt"])

Playlists and channels

Pass a playlist URL or a channel URL (/@handle or /channel/UC...) and it expands to its videos. Use maxVideosPerSource to cap how many (1 to 25 per source through the API; channels return the newest uploads first):

ok, failed = get_transcripts(
    ["https://www.youtube.com/@SomeChannel"],
    formats=["text"],
)

Only videos that actually return a transcript are billed, so a channel where half the uploads lack captions costs half as much. One request accepts up to 10 sources. For bigger jobs, loop over batches and keep the rate limit (60 requests per minute per key) in mind.

Preparing transcripts for an LLM

Raw transcripts are long and unpunctuated in places. Two habits make them much more useful for retrieval-augmented generation (RAG) and summarisation.

Chunk by time, not by characters. Segments already carry timestamps, so group them into windows of roughly 60 to 90 seconds. Every chunk then maps to a real position in the video, and your application can cite it as a deep link:

def chunk_segments(video, window=75.0):
    chunks, current, start = [], [], None
    for seg in video["segments"]:
        if start is None:
            start = seg["start"]
        current.append(seg["text"])
        if seg["start"] + seg["duration"] - start >= window:
            chunks.append({"start": start, "text": " ".join(current)})
            current, start = [], None
    if current:
        chunks.append({"start": start, "text": " ".join(current)})
    for c in chunks:
        c["url"] = f"https://www.youtube.com/watch?v={video['videoId']}&t={int(c['start'])}s"
    return chunks

Keep the metadata. Title, channel, publish date and duration are returned for free and make great filters in a vector store ("only talks from this channel", "only videos from this year").

For summarisation, the plain text field is usually all you need. For question answering with citations, embed the chunks above and store url alongside each one.

Handling failures gracefully

Transcripts fail for predictable reasons: the video is private or removed, it has no captions, or it's a live stream that hasn't finished. Every item carries a status, and failed items are never billed. A robust pipeline:

  1. Logs failed items with their status so you can see patterns.
  2. Doesn't retry "no captions" failures: they won't succeed later unless the creator adds captions.
  3. Retries transport errors (upstream_error, upstream_timeout) with backoff; nothing is charged for those either.

Cost

A transcript counts as one result. On the Starter plan ($29/month for 10,000 results) that's $2.90 per 1,000 transcripts if you use the whole plan, falling to $0.67 per 1,000 on Enterprise. If you'd rather not subscribe, the same tool runs pay-per-event on the Apify Store at $3 per 1,000 transcripts. Our pricing page has a calculator, and we wrote up when pay-per-result beats a subscription with worked examples.

Summary

  • Transcripts are caption tracks. No captions, no transcript.
  • DIY scraping works locally but tends to get blocked from cloud servers.
  • An API call with a URL gets you text, timestamped segments, SRT or VTT, plus metadata.
  • Chunk by timestamps for RAG, and handle per-item failures instead of assuming success.

You can try a transcript request right now, without signing up, in the live demo.

Related guides

Get the next guide by email

Practical tutorials on transcripts, screenshots, news data and agent tooling. A couple of emails a month at most.