Find the bugs your users are complaining about: mining App Store reviews with Python
Your one-star reviews are the cheapest bug reports you will ever get. People take the time to write down what broke, on which device, after which update, and they do it in public. The trouble is reading them: the App Store shows one country at a time, mixes praise with complaints, and has no export button.
This guide pulls the low-star reviews for an app into Python, then answers three questions a product or QA team actually asks:
- Where are the complaints coming from?
- Which app versions collect the most of them?
- What are people complaining about?
Everything below was run on 30 September 2026 against a real app (Notion's iOS app, picked because it is popular and has plenty of reviews). The output you see is the real output, not a mock-up.
What you need
- Python 3.9 or newer.
- An Apify account and API token. The free plan includes monthly credits, which is far more than this tutorial uses.
- The official client:
pip install apify-client. The code uses version 3.x, where runs come back as objects (run.default_dataset_id). On the older 2.x client, userun["defaultDatasetId"]instead.
We use the App Store Reviews Scraper, which we built and publish on the Apify Store. It reads Apple's public customer-review feed, so you get the same reviews anyone can read on the App Store, as JSON. You can search by app name, pass several countries in one run, and filter by stars before anything is charged. The price is $0.0001 per review returned ($0.10 per 1,000), so the 300 reviews in this tutorial cost three cents.
Set your token once:
export APIFY_TOKEN="apify_api_..."
pip install apify-client
Step 1: fetch only the reviews you care about
The key idea is to filter at the source. You don't want 5-star "love it!" reviews for a bug hunt, so ask for 1 and 2 stars only with maxRating: 2. Filtered-out reviews are never returned or billed.
run_input = {
"apps": ["Notion"], # a name, an App Store link or a numeric id
"countries": ["us", "gb", "ca", "au"],
"maxReviewsPerApp": 300, # across all countries
"maxRating": 2, # only 1- and 2-star reviews are returned (and billed)
}
run = client.actor("siftwright/app-store-reviews-scraper").call(run_input=run_input)
reviews = list(client.dataset(run.default_dataset_id).iterate_items())
A few details worth knowing:
- App names work, but links are exact. The name is resolved with Apple's own App Store search and the top match is used. "Notion" matched Notion: Notes, Tasks, AI by Notion Labs. If a name is ambiguous, paste the App Store link or the numeric id instead.
- Every country is a separate review pool. A review written in the UK store does not show up in the US store. Listing several countries is how you get a fuller picture; the scraper de-duplicates across them.
maxReviewsPerAppis a total, not per country. Countries are read in order until the cap is reached. In our run the cap of 300 was hit after the US and UK, so Canada and Australia were never fetched and never charged. Raise the cap if you want every listed country.- Apple's feed has a ceiling. It returns up to 500 reviews per app, per country, per sort order (most recent or most helpful). For a big app that is the most recent few weeks or months, not its whole history.
Each item looks like this (a real row, shortened):
{
"appName": "Notion: Notes, Tasks, AI",
"country": "us",
"rating": 1,
"title": "Keep flickering",
"text": "It just keep pulling everything up. Unusable on Mac.",
"appVersionReviewed": "1.7.340",
"date": "2026-09-23T...Z"
}
The appVersionReviewed field is the one that makes this useful for engineering teams: it tells you which build the reviewer was on.
Step 2: check what you paid for
Pay-per-result pricing is only fair if you can check it. Charges are recorded on the run itself, and the counters settle a moment after the run finishes, so re-read the run before printing them:
run = client.run(run.id).get()
print(f"{len(reviews)} low-star reviews, charged: {run.charged_event_counts}")
Our run printed 300 low-star reviews, charged: {'apify-actor-start': 1, 'review': 300}: one charge per review in the dataset, plus Apify's standard start fee of $0.00005.
Step 3: where are complaints coming from?
print(Counter(r["country"] for r in reviews).most_common())
Our output was [('us', 207), ('gb', 93)]. With a higher cap you would see all four countries. For apps with localised content, pricing or payment options, a single country standing out is often the first clue that a problem is regional.
Step 4: which versions collect the most low-star reviews?
App versions are strings like 1.7.341, so sort them numerically rather than alphabetically (otherwise 1.7.99 sorts after 1.7.341):
by_version = Counter(r["appVersionReviewed"] for r in reviews)
for version, n in sorted(by_version.items(), key=lambda kv: [int(x) for x in re.findall(r"\d+", kv[0])], reverse=True)[:6]:
print(f"{version:>10} {n:3d} {'#' * n}")
For a fast-shipping app like Notion, which releases every few days, the counts per version are small (2 to 5 for the most recent builds in our run). The signal gets much stronger when you run this on a schedule and compare one release with the previous one: a version that suddenly collects three times the usual low-star reviews is worth a look before the next release goes out.
Step 5: tag reviews with explainable themes
You could send every review to an LLM, and for deep analysis you probably should. But a first pass with plain keyword themes is instant, free, and, most importantly, explainable: anyone on the team can see why a review was tagged.
THEMES = {
"crash / bug": r"crash\w*|bugs?|buggy|glitch\w*|broken|flicker\w*|not working|doesn.t work",
"slow / lag": r"slow|lag\w*|loading|takes forever|freez\w*",
"sync / offline": r"sync\w*|offline|no internet|connection",
"login / account": r"log ?in|sign ?in|account|password|locked out",
"ai features": r"ai|artificial intelligence",
"pricing": r"price|pricing|subscription|paywall|expensive",
"mobile / ipad": r"mobile|iphone|ipad|phone app",
}
PATTERNS = {theme: re.compile(rf"\b(?:{p})\b") for theme, p in THEMES.items()}
Two lessons from getting this right on real data:
- Wrap every alternative in word boundaries. Our first draft had
\bonly in front of the first word, solagalso matched "flag" andaicould match inside other words.\b(?:...)\baround the whole group fixes that. - Then allow the endings people actually type. With strict boundaries,
bugno longer matched "bugs" or "buggy", and the crash theme dropped from 110 reviews to 45.bugs?|buggy|crash\w*brought it back. Always spot-check a few matches per theme before trusting the numbers.
Counting is then a simple loop, keeping one example title per theme so the output is self-explaining (see the full script below). Our run printed:
crash / bug 110 (37%) e.g. [us 1.7.341] what's happening?
ai features 90 (30%) e.g. [us 1.7.340] AI is terrible
mobile / ipad 89 (30%) e.g. [us 1.7.341] Can’t make checklists in Mobil app
slow / lag 40 (13%) e.g. [us 1.7.339] Too slow and takes up too much data storage
login / account 15 (5%) e.g. [us 1.7.328] Great before the most recent update (2026)
pricing 8 (3%) e.g. [us 1.7.333] Why so bloated?
sync / offline 6 (2%) e.g. [us 1.7.327] Ready to switch to an offline note app
A review can carry several themes, so the percentages add up to more than 100%. Read the result as "what share of unhappy reviewers mention this", not as a pie chart. For this app, stability and the mobile experience dominate, and AI features come up in almost a third of low-star reviews, which is exactly the kind of thing a product team wants to know about a recent feature push.
Step 6: save it for a spreadsheet or an LLM
with open("low_star_reviews.csv", "w", newline="", encoding="utf-8") as f:
w = csv.DictWriter(f, fieldnames=["date", "country", "rating", "appVersionReviewed", "title", "text"], extrasaction="ignore")
w.writeheader()
w.writerows(reviews)
The CSV opens in any spreadsheet. It is also a good input for an LLM prompt such as "group these reviews into the five most common problems, quote one review for each, and name the app version where each problem first appears". Keeping the version and country columns in the file lets the model answer that second part.
We left author names out of the CSV on purpose. You rarely need them for product analysis, and not storing personal data you don't need keeps you on the right side of GDPR and similar laws.
The full script
This is the exact script that produced the output above:
import csv
import os
import re
from collections import Counter, defaultdict
from apify_client import ApifyClient
client = ApifyClient(os.environ["APIFY_TOKEN"])
run_input = {
"apps": ["Notion"], # a name, an App Store link or a numeric id
"countries": ["us", "gb", "ca", "au"],
"maxReviewsPerApp": 300, # across all countries
"maxRating": 2, # only 1- and 2-star reviews are returned (and billed)
}
run = client.actor("siftwright/app-store-reviews-scraper").call(run_input=run_input)
reviews = list(client.dataset(run.default_dataset_id).iterate_items())
run = client.run(run.id).get() # re-read the run: charge counters settle a moment after it finishes
print(f"{len(reviews)} low-star reviews, charged: {run.charged_event_counts}")
# 1. Where do the complaints come from?
print(Counter(r["country"] for r in reviews).most_common())
# 2. Which versions collect the most low-star reviews?
by_version = Counter(r["appVersionReviewed"] for r in reviews)
for version, n in sorted(by_version.items(), key=lambda kv: [int(x) for x in re.findall(r"\d+", kv[0])], reverse=True)[:6]:
print(f"{version:>10} {n:3d} {'#' * n}")
# 3. Tag each review with simple, explainable themes.
THEMES = {
"crash / bug": r"crash\w*|bugs?|buggy|glitch\w*|broken|flicker\w*|not working|doesn.t work",
"slow / lag": r"slow|lag\w*|loading|takes forever|freez\w*",
"sync / offline": r"sync\w*|offline|no internet|connection",
"login / account": r"log ?in|sign ?in|account|password|locked out",
"ai features": r"ai|artificial intelligence",
"pricing": r"price|pricing|subscription|paywall|expensive",
"mobile / ipad": r"mobile|iphone|ipad|phone app",
}
PATTERNS = {theme: re.compile(rf"\b(?:{p})\b") for theme, p in THEMES.items()}
theme_counts = Counter()
examples = defaultdict(list)
for r in reviews:
text = f"{r['title']} {r['text']}".lower()
for theme, pattern in PATTERNS.items():
if pattern.search(text):
theme_counts[theme] += 1
if len(examples[theme]) < 2:
examples[theme].append(f"[{r['country']} {r['appVersionReviewed']}] {r['title']}")
for theme, n in theme_counts.most_common():
print(f"{theme:<16} {n:3d} ({n / len(reviews):.0%}) e.g. {examples[theme][0]}")
# 4. Save everything for a spreadsheet or an LLM pass later.
with open("low_star_reviews.csv", "w", newline="", encoding="utf-8") as f:
w = csv.DictWriter(f, fieldnames=["date", "country", "rating", "appVersionReviewed", "title", "text"], extrasaction="ignore")
w.writeheader()
w.writerows(reviews)
print("saved low_star_reviews.csv")
And its complete output:
300 low-star reviews, charged: {'apify-actor-start': 1, 'review': 300}
[('us', 207), ('gb', 93)]
1.7.341 2 ##
1.7.340 5 #####
1.7.339 5 #####
1.7.338 2 ##
1.7.337 2 ##
1.7.336 3 ###
crash / bug 110 (37%) e.g. [us 1.7.341] what's happening?
ai features 90 (30%) e.g. [us 1.7.340] AI is terrible
mobile / ipad 89 (30%) e.g. [us 1.7.341] Can’t make checklists in Mobil app
slow / lag 40 (13%) e.g. [us 1.7.339] Too slow and takes up too much data storage
login / account 15 (5%) e.g. [us 1.7.328] Great before the most recent update (2026)
pricing 8 (3%) e.g. [us 1.7.333] Why so bloated?
sync / offline 6 (2%) e.g. [us 1.7.327] Ready to switch to an offline note app
saved low_star_reviews.csv
Making it a weekly habit
A one-off report is interesting; a weekly one is useful. A few ways to run it on a schedule:
- Apify schedules: save the input as a task and schedule it; the dataset of every run stays available through the API.
- Filter by date: pass
since(for example the date of your last run) so each run only returns, and only charges for, new reviews. - Watch competitors too: add competing apps to
apps. What users hate about a rival is a feature list for your roadmap. - Alert on spikes: compare this week's crash-theme count with last week's and post to Slack when it doubles.
Limits to keep in mind
- Apple's feed gives the most recent (or most helpful) 500 reviews per app, country and sort order. It is not a full archive.
- Developer replies are not in the feed.
- Star ratings without text (ratings-only) are not reviews and are not included.
- Keyword themes are a first pass. They miss sarcasm, other languages and unusual wording. Use them to decide where to look, then read the reviews.
Wrapping up
In about 60 lines you went from "we should read our reviews" to a per-version, per-theme picture of what unhappy users are saying, for three cents. The same script works for any app on the App Store and, with a list of competitors, for your whole market.
If you want to try it, the App Store Reviews Scraper runs on Apify's free plan credits. Questions or ideas are welcome at support@siftwright.com.
Related guides
Build a news monitoring pipeline with a Google News API (Python)
A step-by-step guide to monitoring brand, competitor and industry news in Python. Design queries, fetch Google News results as JSON, deduplicate, store in SQLite, and send a daily digest, in about 120 lines of code.
How to get YouTube transcripts with an API (Python, Node.js and curl)
A practical guide to fetching YouTube transcripts from code: why DIY scrapers get IP-blocked on servers, working Python, Node.js and curl examples, languages and translation, SRT/VTT, playlists, and chunking transcripts for LLMs.
Get the next guide by email
Practical tutorials on transcripts, screenshots, news data and agent tooling. A couple of emails a month at most.