Build a news monitoring pipeline with a Google News API (Python)
Knowing when your company, your competitors or your market show up in the news is valuable, and the commercial tools that do it can cost hundreds of dollars a month per seat. The core of a news monitor is simple enough to build yourself: run a set of searches on a schedule, keep what's new, and send a digest.
This guide builds that pipeline in Python. By the end you'll have a script that:
- Runs a list of Google News queries.
- Normalises and deduplicates the results.
- Stores every article in SQLite so it only reports each one once.
- Prints (or emails) a daily digest grouped by query.
It's about 120 lines, has one dependency (requests), and costs cents per day to run.
Why Google News, and why an API
Google News aggregates thousands of publishers, ranks them reasonably well, and supports search operators. The catch is that Google doesn't offer an official News API; the old one was retired long ago. Scraping the pages yourself means dealing with markup changes, consent pages and blocking.
We'll use the Siftwright Google News endpoint, which returns the results for any query as JSON: title, link, source, publish time and snippet, for a chosen language and country edition. You only pay for articles actually returned, and an empty search is free. Examples assume your key is in SIFTWRIGHT_API_KEY.
A quick test from the command line:
curl https://siftwright.com/v1/google-news \
-H "Authorization: Bearer $SIFTWRIGHT_API_KEY" \
-H "Content-Type: application/json" \
-d '{"query": "web scraping", "maxItems": 3}'
Each result looks like this:
{
"query": "web scraping",
"title": "Web scraping startup Firecrawl closes $75M investment",
"link": "https://news.google.com/rss/articles/CBMi...?oc=5",
"source": "SiliconANGLE",
"publishedAt": "2026-09-22T22:46:00.000Z",
"snippet": "Web scraping startup Firecrawl closes $75M investment SiliconANGLE",
"status": "ok"
}
Step 1: design your queries
Good queries matter more than good code. A monitor usually tracks three kinds of topics:
- Your brand: the company name, product names, and your domain. Quote multi-word names so they match as a phrase:
"Acme Analytics". - Competitors: one query per competitor, same pattern.
- Your market: a few phrases that describe the space, like
"usage-based pricing" SaaS.
Google News understands quotes and OR, which lets you fold variants into one query:
QUERIES = {
"brand": '"Acme Analytics" OR acmeanalytics.com',
"competitor-globex": '"Globex" analytics',
"market": '"product analytics" OR "usage-based pricing"',
}
Keep each query tight. A vague query like analytics returns thousands of irrelevant articles, and you pay for every one you ask for. That's also why maxItems matters: a daily run rarely needs more than 20 to 30 articles per query.
Pick the edition with language and country. en/US is the default; en/GB, de/DE or hi/IN return different publishers and rankings. If you care about several markets, run the same query per edition.
Step 2: fetch results
A small function that runs one query and returns only successful items:
import os
import requests
API = "https://siftwright.com/v1/google-news"
HEADERS = {"Authorization": f"Bearer {os.environ['SIFTWRIGHT_API_KEY']}"}
def search(query, max_items=25, language="en", country="US"):
r = requests.post(
API,
headers=HEADERS,
json={"query": query, "maxItems": max_items, "language": language, "country": country},
timeout=300,
)
if r.status_code == 429:
raise RuntimeError(r.json()["error"]["message"])
r.raise_for_status()
return [a for a in r.json()["results"] if a.get("status") == "ok"]
Step 3: normalise and deduplicate
The same story often appears several times: syndicated to multiple publishers, or matched by two of your queries. Two cheap rules remove most duplicates:
- Same link is the same article.
- Same normalised title is almost always the same story, even from a different source.
Google News titles often end with " - Publisher Name". Strip that, lowercase, and collapse punctuation:
import hashlib
import re
def title_key(title, source):
t = title.strip()
if source and t.endswith(" - " + source):
t = t[: -len(" - " + source)]
t = re.sub(r"[^a-z0-9 ]+", " ", t.lower())
t = re.sub(r"\s+", " ", t).strip()
return hashlib.sha1(t.encode()).hexdigest()
You could go further with fuzzy matching, but in practice exact matches on normalised titles catch the bulk of repeats without false positives.
Step 4: store what you've seen
SQLite is perfect for this: a single file, no server, and a UNIQUE constraint does the deduplication for us. INSERT OR IGNORE means an article we've already stored is silently skipped, and cursor.rowcount tells us whether it was new.
import sqlite3
SCHEMA = """
CREATE TABLE IF NOT EXISTS articles (
id INTEGER PRIMARY KEY,
topic TEXT NOT NULL,
title TEXT NOT NULL,
source TEXT,
link TEXT NOT NULL UNIQUE,
title_key TEXT NOT NULL UNIQUE,
published_at TEXT,
first_seen TEXT NOT NULL DEFAULT (datetime('now'))
);
"""
def open_db(path="news.db"):
db = sqlite3.connect(path)
db.executescript(SCHEMA)
return db
def save(db, topic, article):
cur = db.execute(
"INSERT OR IGNORE INTO articles (topic, title, source, link, title_key, published_at) "
"VALUES (?, ?, ?, ?, ?, ?)",
(
topic,
article["title"],
article.get("source"),
article["link"],
title_key(article["title"], article.get("source")),
article.get("publishedAt"),
),
)
return cur.rowcount == 1
Because both link and title_key are unique, an article is new only if we've never seen its link or its story.
Step 5: build the digest
Now tie it together: run every query, save the results, and collect the new ones for the digest.
from collections import defaultdict
from datetime import datetime, timezone
def run(queries):
db = open_db()
new = defaultdict(list)
for topic, query in queries.items():
for article in search(query):
if save(db, topic, article):
new[topic].append(article)
db.commit()
return new
def digest(new):
today = datetime.now(timezone.utc).strftime("%d %B %Y")
lines = [f"News digest for {today}", ""]
if not new:
lines.append("Nothing new today.")
for topic, articles in new.items():
lines.append(f"## {topic} ({len(articles)} new)")
for a in sorted(articles, key=lambda a: a.get("publishedAt") or "", reverse=True):
lines.append(f"- {a['title']} ({a.get('source', 'unknown')})")
lines.append(f" {a['link']}")
lines.append("")
return "\n".join(lines)
if __name__ == "__main__":
print(digest(run(QUERIES)))
The first run reports everything it finds. From the second run on, you'll only see articles that appeared since the last run.
Step 6: deliver it
Printing is enough to start. To email it, use Python's standard library with any SMTP provider:
import smtplib
from email.message import EmailMessage
def send(text, to_addr):
msg = EmailMessage()
msg["Subject"] = "Daily news digest"
msg["From"] = os.environ["DIGEST_FROM"]
msg["To"] = to_addr
msg.set_content(text)
with smtplib.SMTP_SSL(os.environ["SMTP_HOST"], 465) as s:
s.login(os.environ["SMTP_USER"], os.environ["SMTP_PASSWORD"])
s.send_message(msg)
Slack and Microsoft Teams both accept incoming webhooks: POST {"text": digest_text} to the webhook URL and you're done.
Step 7: schedule it
Any scheduler works. On a Linux server, a cron entry that runs every morning at 07:00:
0 7 * * * cd /opt/newsmonitor && /usr/bin/python3 monitor.py >> monitor.log 2>&1
Every few hours is reasonable for brand monitoring; once a day is plenty for market news.
Following links to the publisher
Google News links point to news.google.com and redirect to the publisher's article. For the digest that's fine: clicking works. If you want the final URL (for example, to group by domain), keep in mind that the redirect may happen in the browser rather than as a plain HTTP redirect, so resolving it reliably can require a headless browser. For most monitoring, the source field already tells you who published it.
What it costs
Each returned article counts as one result. Three queries with maxItems=25, run once a day, is at most 75 articles a day, or about 2,250 a month. That fits comfortably in the Starter plan ($29/month for 10,000 results) alongside other usage. Pay-per-event on the Apify Store would cost about $3.40 a month for the same volume ($1.50 per 1,000 articles), so for a single small monitor that's the cheaper option. See our write-up on pay-per-result vs subscription pricing for the break-even maths.
Ways to spend less:
- Lower
maxItemsfor broad queries. New articles show up at the top. - Run market queries daily and brand queries more often, not everything every hour.
- Tighten queries that keep returning noise.
Where to take it next
- Sentiment and relevance: pass each new title and snippet to a language model and ask for a relevance score and a one-line summary.
- Alerts: send an instant message when a brand query returns something from a list of important publishers.
- Trends: count new articles per topic per day in SQLite and chart it.
- Agents: expose
search()as a tool so an AI assistant can check today's coverage before answering. Our guide on web data for AI agents shows how.
Summary
A news monitor is a loop: query, dedupe, store, report. Tight queries and a unique constraint in SQLite do most of the work. Try a query in the live demo to see the data, then check the Google News API page for every option.
Related guides
How to get YouTube transcripts with an API (Python, Node.js and curl)
A practical guide to fetching YouTube transcripts from code: why DIY scrapers get IP-blocked on servers, working Python, Node.js and curl examples, languages and translation, SRT/VTT, playlists, and chunking transcripts for LLMs.
Website screenshot API in Python and Node.js: full-page, mobile, retina and PDF
How to capture website screenshots from code without running your own headless browser fleet. Working Python and Node.js examples for full-page, mobile, retina and PDF captures, late-loading pages, and a batch script with retries.
Get the next guide by email
Practical tutorials on transcripts, screenshots, news data and agent tooling. A couple of emails a month at most.