HTML to PDF at scale: print CSS, fonts, signed URLs and batch rendering
Most applications eventually need to produce PDFs: invoices, receipts, reports, certificates, tickets, contracts. The most maintainable way to do that today is to design the document as an HTML page and let a real browser print it. You get the full power of CSS, you can preview in any browser, and your designers already know the tools.
The hard parts are not the conversion itself. They're print layout (page sizes, margins, page breaks), making sure everything has loaded before printing (fonts, images, charts), keeping private documents private, and running it reliably for thousands of documents. This guide covers each of those, with examples that use the Siftwright screenshot endpoint in PDF mode.
The approach: template → URL → PDF
The Siftwright API renders URLs, not raw HTML strings, in a Chromium browser. So the pipeline has three steps:
- Your application renders a template (Jinja, Handlebars, React on the server, anything) to HTML and serves it at a URL.
- You call
POST /v1/screenshotwith"format": "pdf"and that URL. - You download the PDF from the returned
fileUrland store it.
Serving the document at a URL has real advantages beyond "that's what the API accepts". You can open the exact same URL in your own browser to debug layout, press Ctrl+P to preview the print version, and reuse your existing app's routing, auth and styles.
A minimal request:
curl https://siftwright.com/v1/screenshot \
-H "Authorization: Bearer $SIFTWRIGHT_API_KEY" \
-H "Content-Type: application/json" \
-d '{"url": "https://example.com", "format": "pdf", "pdfPageFormat": "A4"}'
Print CSS that works
Browsers switch to print styles when producing a PDF, so everything you know about @media print applies. These are the rules that make the biggest difference.
Set the page size and margins in CSS
@page {
size: A4;
margin: 18mm 16mm 20mm;
}
The API's pdfPageFormat (A4, A3, Letter, Legal) sets the paper size. Keep your @page size consistent with it, and control margins in CSS so they live with the template.
Use print-friendly units and colours
Millimetres and points behave predictably on paper; viewport units don't. Background colours and images are often dropped when printing unless you ask for them:
html {
-webkit-print-color-adjust: exact;
print-color-adjust: exact;
}
Control page breaks
Nothing looks less professional than a table row split across two pages, or a heading stranded at the bottom of one.
h2, h3 { break-after: avoid; }
tr, .line-item, figure { break-inside: avoid; }
.page-break { break-before: page; }
thead { display: table-header-group; } /* repeat table headers on each page */
Hide what doesn't belong on paper
@media print {
nav, .cookie-banner, .no-print, button { display: none !important; }
a { color: inherit; text-decoration: none; }
}
Headers and footers
The API doesn't expose separate header and footer templates. Put them in the page instead. For a footer that appears on every page, a position: fixed element at the bottom of the page works in Chromium's print mode. For page numbers, most teams either render totals server-side ("Invoice 1042 · Page 1 of 1" for short documents) or accept the limitation for long reports.
Make sure everything has loaded
A PDF rendered too early has fallback fonts, missing images or empty charts. Two API options solve this:
waitForSelector: wait until a specific element exists.delaySec: wait a fixed number of seconds (0 to 10) after load.
The most robust pattern is to add a "ready" marker to the page once everything is done, and wait for it:
<script>
Promise.all([
document.fonts.ready,
...Array.from(document.images).map(img =>
img.complete ? Promise.resolve() : new Promise(r => { img.onload = img.onerror = r; })
),
window.renderCharts ? window.renderCharts() : Promise.resolve(),
]).then(() => {
const done = document.createElement("div");
done.id = "pdf-ready";
done.hidden = true;
document.body.appendChild(done);
});
</script>
Then request the PDF with "waitForSelector": "#pdf-ready". The hidden element won't show up in the output, but it exists in the DOM as soon as the page is truly ready.
Self-host your fonts (or use a fast font CDN) and keep images reasonably sized. Every extra second of loading is a second of rendering time.
Keeping private documents private
Invoices and reports usually shouldn't be public. The renderer must be able to load the page, but nobody else should. The standard solution is a signed, short-lived URL:
- Generate a random token or an HMAC signature over the document ID and an expiry time.
- Serve the document only when the signature is valid and not expired.
- Give the renderer that URL.
Here's a minimal version in Python with Flask:
import hashlib
import hmac
import os
import time
from flask import Flask, abort, render_template, request
app = Flask(__name__)
SECRET = os.environ["PDF_URL_SECRET"].encode()
def sign(doc_id, expires):
msg = f"{doc_id}:{expires}".encode()
return hmac.new(SECRET, msg, hashlib.sha256).hexdigest()
def signed_url(doc_id, ttl=300):
expires = int(time.time()) + ttl
return f"https://app.example.com/print/invoice/{doc_id}?exp={expires}&sig={sign(doc_id, expires)}"
@app.get("/print/invoice/<doc_id>")
def print_invoice(doc_id):
exp = int(request.args.get("exp", "0"))
sig = request.args.get("sig", "")
if exp < time.time() or not hmac.compare_digest(sig, sign(doc_id, exp)):
abort(404)
invoice = load_invoice(doc_id) # your data access
return render_template("invoice_print.html", invoice=invoice)
A five-minute expiry is plenty, since rendering takes seconds. Return 404 rather than 403 for invalid signatures so the route doesn't confirm that documents exist.
If your HTML lives in object storage instead, a pre-signed GET URL from S3, Google Cloud Storage or Cloudflare R2 does the same job with no code on your side.
Rendering and storing the PDF
With a signed URL in hand, request the PDF and copy it to your own storage. The fileUrl in the response is itself a signed link that works without your API key for 7 days, but your storage should be the long-term home.
import requests
API = "https://siftwright.com/v1/screenshot"
HEADERS = {"Authorization": f"Bearer {os.environ['SIFTWRIGHT_API_KEY']}"}
def render_pdfs(doc_ids, page_format="A4"):
urls = {signed_url(d): d for d in doc_ids}
r = requests.post(
API,
headers=HEADERS,
json={
"urls": list(urls),
"format": "pdf",
"pdfPageFormat": page_format,
"waitForSelector": "#pdf-ready",
},
timeout=300,
)
r.raise_for_status()
done, failed = {}, []
for item in r.json()["results"]:
doc_id = urls.get(item["url"])
if item["status"] == "ok":
done[doc_id] = requests.get(item["fileUrl"], timeout=60).content
else:
failed.append(doc_id)
return done, failed
Match results back to documents by URL rather than by position, so a partial failure can't attach the wrong PDF to the wrong customer.
Batch rendering thousands of documents
Each request accepts up to 5 URLs, and a key allows 60 requests per minute. For month-end invoice runs or report generation, put document IDs on a queue and have a worker drain it:
import time
def render_all(doc_ids, store):
queue = list(doc_ids)
attempts = {d: 0 for d in queue}
while queue:
batch, queue = queue[:5], queue[5:]
try:
done, failed = render_pdfs(batch)
except requests.HTTPError as err:
if err.response is not None and err.response.status_code == 429:
code = err.response.json()["error"]["code"]
if code == "quota_exceeded":
raise # stop and alert: the plan is used up
time.sleep(int(err.response.headers.get("Retry-After", "60")))
done, failed = {}, batch
for doc_id, pdf in done.items():
store(doc_id, pdf)
for doc_id in failed:
attempts[doc_id] += 1
if attempts[doc_id] < 3:
queue.append(doc_id)
else:
print("giving up on", doc_id)
Notes on this design:
- Idempotent storage.
store()should overwrite by document ID so a retry never creates duplicates. - Bounded retries. A template bug fails every time; three attempts is enough to ride out transient errors.
- Stop on quota.
quota_exceededmeans the plan's results are used up for the period; retrying won't help. - Failures are free. Only PDFs that render successfully are counted, so retries don't cost extra.
Throughput is simple arithmetic: at 5 PDFs per request, a sequential worker whose requests take 5 to 10 seconds each renders roughly 1,800 to 3,600 documents an hour. Run a few workers in parallel for more, staying under 60 requests per minute. Rendering time varies with page complexity, so measure with your own templates.
Debugging checklist
- Wrong fonts: the font file didn't load before rendering. Wait on
document.fonts.readyvia the ready marker. - Missing background colours: add
print-color-adjust: exact. - Content cut off at the right edge: fixed-width layouts wider than the paper. Use
max-width: 100%and relative widths in print CSS. - Blank pages: an element with
break-before: pageat the very start, or aheight: 100vhcontainer. - Every PDF fails: open the signed URL in a private browser window. If it doesn't load for you, it won't load for the renderer.
Cost
Each successfully rendered PDF counts as one result, the same as a screenshot. Plans run from $29/month for 10,000 results ($2.90 per 1,000) down to $0.67 per 1,000 on Enterprise; the same tool is $1.50 per 1,000 pay-per-event on the Apify Store. Failed renders are never billed.
Summary
- Design documents as HTML pages and render them in a real browser.
- Put page size, margins and page breaks in print CSS.
- Wait for a "ready" marker so fonts, images and charts are in the PDF.
- Serve private documents at signed, short-lived URLs.
- Batch 5 URLs per request, retry failures a bounded number of times, and store PDFs in your own storage.
Read the HTML to PDF API overview for every option, or capture a page in the live demo to see the rendering quality.
Related guides
How to get YouTube transcripts with an API (Python, Node.js and curl)
A practical guide to fetching YouTube transcripts from code: why DIY scrapers get IP-blocked on servers, working Python, Node.js and curl examples, languages and translation, SRT/VTT, playlists, and chunking transcripts for LLMs.
Website screenshot API in Python and Node.js: full-page, mobile, retina and PDF
How to capture website screenshots from code without running your own headless browser fleet. Working Python and Node.js examples for full-page, mobile, retina and PDF captures, late-loading pages, and a batch script with retries.
Get the next guide by email
Practical tutorials on transcripts, screenshots, news data and agent tooling. A couple of emails a month at most.