Website screenshot API in Python and Node.js: full-page, mobile, retina and PDF

By Siftwright team6 min read

Taking a screenshot of a web page sounds like a one-liner. Spin up a headless browser, open the URL, save a PNG. And for one page on your laptop, it is.

In production it turns into a small infrastructure project: browsers that leak memory, pages that never fire their load event, cookie banners covering the content, fonts that haven't loaded yet, lazy images that only appear when you scroll, and servers that need far more RAM than the rest of your app. That's why screenshot APIs exist.

This guide shows how to capture screenshots with an API from Python and Node.js, covering the options that matter in practice: full-page versus viewport, mobile sizes, retina scaling, waiting for late content, PDFs, and batch jobs with retries.

When to use an API instead of running Playwright yourself

Running Playwright or Puppeteer yourself is a good choice when you capture a handful of pages, need to log in to a site, or must interact with the page (click, type, scroll to an element) before capturing.

An API is the better choice when:

  • You capture public pages at volume, or at unpredictable times.
  • Your app runs on serverless or small containers where a 300 MB browser doesn't fit.
  • You don't want to maintain browser versions, fonts, and timeouts.
  • You want failed captures to be clearly reported and not billed.

The rest of this post uses the Siftwright screenshot endpoint, POST /v1/screenshot, which renders pages in a real Chromium browser. Examples assume your key is in SIFTWRIGHT_API_KEY.

Your first screenshot

curl

curl https://siftwright.com/v1/screenshot \
  -H "Authorization: Bearer $SIFTWRIGHT_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{"url": "https://example.com", "format": "png", "fullPage": false}'

The response lists one result per URL:

{
 "endpoint": "screenshot",
 "count": 1,
 "billed": 1,
 "results": [
  {
   "url": "https://example.com",
   "status": "ok",
   "format": "png",
   "fileUrl": "https://api.apify.com/v2/key-value-stores/.../records/capture-....png?signature=...",
   "sizeBytes": 19887,
   "durationMs": 327
  }
 ],
 "usage": { "...": "..." }
}

fileUrl is a signed download link. It works without your API key, so you can hand it to another service, and it stays valid while the file is retained (7 days). If you need the image longer, download it and store it yourself.

Python

import os
import requests

API = "https://siftwright.com/v1/screenshot"
HEADERS = {"Authorization": f"Bearer {os.environ['SIFTWRIGHT_API_KEY']}"}


def capture(url, path, **options):
    r = requests.post(API, headers=HEADERS, json={"url": url, **options}, timeout=300)
    r.raise_for_status()
    shot = r.json()["results"][0]
    if shot["status"] != "ok":
        raise RuntimeError(f"{url}: {shot['status']}")
    img = requests.get(shot["fileUrl"], timeout=60)
    img.raise_for_status()
    with open(path, "wb") as f:
        f.write(img.content)
    return shot


capture("https://example.com", "example.png", format="png", fullPage=True)

Node.js (18 or newer)

import { writeFile } from "node:fs/promises";

const API = "https://siftwright.com/v1/screenshot";

async function capture(url, path, options = {}) {
  const res = await fetch(API, {
    method: "POST",
    headers: {
      Authorization: `Bearer ${process.env.SIFTWRIGHT_API_KEY}`,
      "Content-Type": "application/json",
    },
    body: JSON.stringify({ url, ...options }),
    signal: AbortSignal.timeout(300_000),
  });
  const body = await res.json();
  if (!res.ok) throw new Error(`${body.error.code}: ${body.error.message}`);
  const [shot] = body.results;
  if (shot.status !== "ok") throw new Error(`${url}: ${shot.status}`);
  const file = await fetch(shot.fileUrl);
  await writeFile(path, Buffer.from(await file.arrayBuffer()));
  return shot;
}

await capture("https://example.com", "example.png", { fullPage: true });

Full page or viewport?

fullPage: true (the default) scrolls the whole document into one tall image. It's what you want for archiving, visual regression and "show me the whole landing page".

fullPage: false captures only the visible viewport, which is what a visitor sees before scrolling. Use it for link previews, thumbnails and social cards, where a 10,000-pixel-tall image is useless.

The viewport defaults to 1280 × 800. You can set viewportWidth from 320 to 3840 and viewportHeight from 240 to 2160.

Mobile and retina captures

Responsive sites render a completely different layout at phone widths. To see it, set a phone-sized viewport:

capture(
    "https://example.com",
    "example-mobile.png",
    viewportWidth=390,
    viewportHeight=844,
    deviceScaleFactor=3,
    fullPage=False,
)

deviceScaleFactor multiplies the pixel density: 2 or 3 gives sharp images on high-density screens and in print. The trade-off is file size, which grows with the square of the factor. A 1280 × 800 capture at factor 2 is really 2560 × 1600 pixels.

Keep in mind that a narrow viewport isn't a full mobile emulation. Sites that sniff the user agent rather than using responsive CSS may still serve the desktop version.

JPEG or PNG?

  • PNG is lossless. Best for pages with text, UI and sharp edges, and for pixel-diffing.
  • JPEG is much smaller for photo-heavy pages and is fine for thumbnails.

Set "format": "jpeg" when size matters more than perfect text rendering.

Pages that load late

Modern pages often render content after the load event: client-side apps fetch data, charts animate in, images lazy-load. Two options help:

  • waitForSelector: wait until an element exists. Pick something that only appears once the content you care about has rendered, like #pricing-table or [data-loaded="true"].
  • delaySec: wait a fixed number of seconds (up to 10) after the page loads.
capture(
    "https://example.com/dashboard/public",
    "dashboard.png",
    waitForSelector="#chart svg",
    delaySec=1,
)

Prefer waitForSelector where you can, since it's precise. Use delaySec for animations and cases where there's no reliable element to wait for.

Cookie-consent banners are dismissed on a best-effort basis by default (blockCookieBanners: true). If you need to see the banner, set it to false.

PDFs from the same endpoint

Set "format": "pdf" to get a print-ready PDF instead of an image, and choose the paper size with pdfPageFormat (A4, A3, Letter or Legal):

capture("https://example.com", "example.pdf", format="pdf", pdfPageFormat="Letter")

PDFs render with the page's print styles, so CSS like @media print and @page applies. We wrote a separate guide on HTML to PDF at scale covering templates, fonts and page breaks.

Batch captures with retries

Each request accepts up to 5 URLs. For a longer list, send batches and retry only the pages that failed for transient reasons. Here's a complete Python script:

import os
import time
import pathlib
import requests

API = "https://siftwright.com/v1/screenshot"
HEADERS = {"Authorization": f"Bearer {os.environ['SIFTWRIGHT_API_KEY']}"}
OUT = pathlib.Path("shots")
OUT.mkdir(exist_ok=True)


def batches(items, size=5):
    for i in range(0, len(items), size):
        yield items[i:i + size]


def run(urls, attempts=3):
    pending = list(urls)
    for attempt in range(1, attempts + 1):
        failed = []
        for batch in batches(pending):
            r = requests.post(
                API,
                headers=HEADERS,
                json={"urls": batch, "format": "png", "fullPage": False},
                timeout=300,
            )
            if r.status_code == 429 and r.json()["error"]["code"] == "rate_limited":
                time.sleep(int(r.headers.get("Retry-After", "60")))
                failed.extend(batch)
                continue
            if r.status_code >= 500:
                failed.extend(batch)  # upstream error: nothing was billed
                continue
            r.raise_for_status()
            for shot in r.json()["results"]:
                if shot["status"] == "ok":
                    name = shot["url"].split("//", 1)[-1].replace("/", "_")[:120] + ".png"
                    (OUT / name).write_bytes(requests.get(shot["fileUrl"], timeout=60).content)
                else:
                    failed.append(shot["url"])
        if not failed:
            return []
        pending = failed
        time.sleep(5 * attempt)
    return pending


leftover = run([
    "https://example.com",
    "https://www.wikipedia.org",
    "https://news.ycombinator.com",
])
print("Could not capture:", leftover)

A few design choices worth copying:

  • Retry at the URL level. A batch can partly succeed; only the failures go back into the queue.
  • Respect Retry-After. The rate limit is 60 requests per minute per key.
  • Don't retry forever. Some pages will never load (dead domains, pages that block automated browsers). Report them.
  • Failed captures cost nothing, so retries don't double-bill you.

Common pitfalls

Blank or half-rendered images. The page probably renders client-side. Add waitForSelector for an element in the main content.

Huge files. Full-page captures of long pages at deviceScaleFactor: 3 can be many megabytes. Use viewport captures, JPEG, or factor 1 for thumbnails.

Login-only content. The API captures what a logged-out visitor sees. For authenticated pages, run a browser yourself or expose a public, tokenised preview URL.

Consent walls. Cookie banners are dismissed best-effort. Full-screen consent walls that require a choice to continue may still appear in some regions.

Cost

Every successful capture (PNG, JPEG or PDF) counts as one result. Plans start at $29/month for 10,000 results, which is $2.90 per 1,000 captures, and drop to $0.67 per 1,000 on the Enterprise plan. Pay-per-event on the Apify Store costs $1.50 per 1,000 captures with no subscription. Timeouts, DNS failures and error pages are never billed.

Summary

  • Use an API when you capture public pages at volume or from small servers.
  • Pick fullPage for archives and diffs, viewport captures for previews.
  • Combine viewportWidth, viewportHeight and deviceScaleFactor for mobile and retina.
  • Use waitForSelector for late-rendering pages.
  • Batch 5 URLs per request and retry only what failed.

Want to see it first? Capture any public page in the live demo, or read the screenshot API overview.

Related guides

Get the next guide by email

Practical tutorials on transcripts, screenshots, news data and agent tooling. A couple of emails a month at most.