Find every Shopify, WordPress and HubSpot site in a lead list: tech stack lookup with evidence in Python
If you sell a Shopify app, a WordPress plugin or a migration service, the first question about any lead is the same: what is their website actually built on? A lead list with that one extra column is worth far more than one without it, because you can stop pitching a Shopify app to companies that run Magento.
Tools like Wappalyzer and BuiltWith answer the question, but their APIs are priced for teams that look up thousands of sites a month. If you have a list of a few hundred companies from a conference, a CRM export or a directory, you want something closer to "pay a fraction of a cent per website and get the answer back as JSON".
This guide does exactly that. We take a small CSV of company websites, detect the technologies on each one, and write an enriched CSV with four new columns: platform, the evidence for that platform, email provider and hosting. Then we print the list grouped by platform, so you can see your segments at a glance.
Everything here was run on 1 October 2026 against real, public websites. The output shown is the real output.
What you need
- Python 3.9 or newer.
- An Apify account and API token. The free plan includes monthly credits, which covers this tutorial many times over.
- The official client:
pip install apify-client. The code uses version 3.x, where runs come back as objects (run.default_dataset_id). On the older 2.x client, userun["defaultDatasetId"]instead.
We use the Tech Stack Detector, which we built and publish on the Apify Store. It loads each website once over plain HTTPS, reads the response headers, cookies, meta tags, script and stylesheet URLs, and checks the domain's MX, SPF and NS records. It costs $0.002 per website scanned ($2 per 1,000), and sites that fail to load are free. The ten sites in this tutorial cost under two cents.
Set your token once:
export APIFY_TOKEN="apify_api_..."
The lead list
Here is the input, leads.csv. It is deliberately mixed: two direct-to-consumer brands, a news site, a magazine, a few software companies, a non-profit and one domain that does not exist, so we can see how failures are handled.
company,website
Allbirds,allbirds.com
Gymshark,gymshark.com
TechCrunch,techcrunch.com
The New Yorker,newyorker.com
Basecamp,basecamp.com
Notion,notion.so
HubSpot,hubspot.com
Mozilla,mozilla.org
Ghost,ghost.org
Example (dead),this-domain-does-not-exist-sw123.com
In real life this is your CRM export. The only column the code relies on is website, and it accepts bare domains or full URLs.
One call for the whole list
The detector takes a list of websites and scans them in parallel, so there is no need to loop and call it once per row:
import csv
import os
from apify_client import ApifyClient
client = ApifyClient(os.environ["APIFY_TOKEN"])
with open("leads.csv", newline="") as f:
leads = list(csv.DictReader(f))
run = client.actor("siftwright/tech-stack-detector").call(
run_input={"urls": [row["website"] for row in leads], "includeDns": True}
)
# apify-client 3.x returns a Run object; on 2.x use run["defaultDatasetId"]
results = {item["input"]: item for item in client.dataset(run.default_dataset_id).iterate_items()}
We key the results by input, the exact string we sent, because the site you ask for is often not the site you land on. notion.so redirects to www.notion.com, and gymshark.com ended up on a regional checkout domain. Matching on the input keeps every result attached to the right row of your CSV.
includeDns turns on the MX, SPF and NS checks. They are what tell you a company uses Google Workspace or Microsoft 365 for email, which is often as useful for targeting as the website platform itself. It costs nothing extra.
What comes back
Each website produces one item. Here is the one for ghost.org, shortened:
{
"input": "ghost.org",
"finalUrl": "https://ghost.org/",
"status": "ok",
"httpStatus": 200,
"generator": "Hugo 0.119.0",
"technologyCount": 11,
"byCategory": {
"DNS provider": ["Cloudflare DNS"],
"Email provider": ["Google Workspace"],
"Hosting": ["Netlify"],
"Static site generator": ["Hugo"],
"Transactional email": ["Mandrill"],
"UI framework": ["Tailwind CSS"]
},
"technologies": [
{ "name": "Netlify", "categories": ["Hosting"],
"evidence": ["header server: Netlify", "header x-nf-request-id: 01M3TJ4..."] },
{ "name": "Google Workspace", "categories": ["Email provider"],
"evidence": ["dns mx: alt1.aspmx.l.google.com"] }
]
}
Two fields do most of the work:
byCategorygroups the technology names by what they are. It is the quickest way to ask "which CMS?" or "which email provider?".technologiescarries theevidencefor every detection: the header, meta tag, asset URL or DNS record that matched. This matters more than it looks. When a salesperson asks "are we sure they're on Shopify?", you can answer withheader powered-by: Shopifyinstead of "the tool said so".
A failed site comes back as a row too, with status: "error" and a reason, and it is not charged:
{
"input": "this-domain-does-not-exist-sw123.com",
"status": "error",
"error": "Domain not found (DNS lookup failed)"
}
Turning detections into segments
A lead list needs one platform per company, not fifteen technology names. So we pick the first match from a short list of categories, in order of how specific they are: an ecommerce platform beats a CMS, and a CMS beats a static site generator.
from collections import defaultdict
PLATFORM_CATEGORIES = ["Ecommerce", "CMS", "Headless CMS", "Static site generator"]
def first_in(item, categories):
for category in categories:
names = item.get("byCategory", {}).get(category)
if names:
return names[0]
return ""
def evidence_for(item, name):
for tech in item.get("technologies", []):
if tech["name"] == name:
return tech["evidence"][0]
return ""
Then we walk the original leads, in their original order, and build one output row per company:
segments = defaultdict(list)
rows = []
for lead in leads:
item = results.get(lead["website"], {})
if item.get("status") != "ok":
rows.append({**lead, "status": item.get("error", "not scanned")})
continue
platform = first_in(item, PLATFORM_CATEGORIES)
email = first_in(item, ["Email provider"])
segments[platform or "Custom / unknown"].append(lead["company"])
rows.append({
**lead,
"status": "ok",
"platform": platform,
"platform_evidence": evidence_for(item, platform) if platform else "",
"email_provider": email,
"hosting": first_in(item, ["Hosting", "CDN"]),
"analytics": ", ".join(item.get("byCategory", {}).get("Analytics", [])
+ item.get("byCategory", {}).get("Tag manager", [])),
"tech_count": item["technologyCount"],
})
Failed sites stay in the file with their error message instead of silently disappearing. That way nobody wonders why a company is missing, and you can fix typos in the domain column and run just those rows again.
Writing the enriched CSV
with open("leads_enriched.csv", "w", newline="") as f:
fields = ["company", "website", "status", "platform", "platform_evidence",
"email_provider", "hosting", "analytics", "tech_count"]
writer = csv.DictWriter(f, fieldnames=fields)
writer.writeheader()
writer.writerows(rows)
for platform, companies in sorted(segments.items(), key=lambda kv: -len(kv[1])):
print(f"{platform:<18} {len(companies)} {', '.join(companies)}")
Running the whole script printed this:
Custom / unknown 3 The New Yorker, Basecamp, Mozilla
Shopify 2 Allbirds, Gymshark
WordPress 1 TechCrunch
Contentful 1 Notion
HubSpot CMS 1 HubSpot
Hugo 1 Ghost
And leads_enriched.csv came out like this:
| company | platform | platform_evidence | email_provider | hosting |
|---|---|---|---|---|
| Allbirds | Shopify | header powered-by: Shopify | Microsoft 365 | Cloudflare |
| Gymshark | Shopify | header powered-by: Shopify | Cloudflare | |
| TechCrunch | WordPress | meta generator: WordPress 6.9.9 | ||
| The New Yorker | Google Workspace | Amazon CloudFront | ||
| Basecamp | Cloudflare | |||
| Notion | Contentful | asset https://images.ctfassets.net | Vercel | |
| HubSpot | HubSpot CMS | header x-hs-hub-id: 53 | Google Workspace | Cloudflare |
| Mozilla | Google Workspace | Google Cloud | ||
| Ghost | Hugo | meta generator: Hugo 0.119.0 | Google Workspace | Netlify |
| Example (dead) |
The run reported nine websites scanned and one failed, and Apify's billing showed nine charged events: nine websites at $0.002 is $0.018. The dead domain cost nothing.
Reading the results honestly
A few things in that table are worth understanding before you hand it to a sales team.
"Custom / unknown" usually means custom. Basecamp and The New Yorker run their own application stacks, so no off-the-shelf platform shows up. That is a real answer, and for a Shopify app vendor it is the right one: not a fit.
Blank email provider does not mean "no email". TechCrunch and Gymshark route mail through security gateways (Mimecast and Proofpoint showed up under "Email security"), which hide the mailbox provider behind them. If email matters to your pitch, add the Email security category to that column.
Redirects are followed. Notion's notion.so is now notion.com, and the result describes the site people actually see. Check finalUrl if a result looks surprising.
One page, no browser. The detector reads the HTML, headers and DNS of the page you give it and does not execute JavaScript. Tools that only appear after client-side scripts run can be missed, and the catalogue (around 200 technologies) is smaller than Wappalyzer's or BuiltWith's. For "which CMS, which shop platform, which email provider, which host" on a list of companies, that has been plenty in our tests. For a full inventory of every marketing pixel on a single site, use a browser-based tool.
Making it repeatable
Once the script works, a few small changes make it useful week after week:
- Run only new leads. Keep the enriched CSV, and before calling the Actor, drop rows whose
websitealready has astatusofok. - Cap the spend. Pass
maxTotalChargeUsdin the run options, for exampleclient.actor(...).call(run_input=..., max_total_charge_usd=1), and the run stops at that budget. At $0.002 per website, a dollar covers 500 sites. - Schedule it. Apify can run the Actor on a schedule with a saved input, and you can pull the latest dataset from a cron job or an n8n, Make or Zapier flow.
- Use it from an agent. The same Actor is available through Apify's MCP server, so an AI agent can look up a prospect's stack mid-conversation:
https://mcp.apify.com?tools=siftwright/tech-stack-detector.
The full script
Here is everything in one file, exactly as we ran it. Save it as segment_leads.py next to your leads.csv and run python segment_leads.py.
import csv
import os
from collections import defaultdict
from apify_client import ApifyClient
client = ApifyClient(os.environ["APIFY_TOKEN"])
with open("leads.csv", newline="") as f:
leads = list(csv.DictReader(f))
run = client.actor("siftwright/tech-stack-detector").call(
run_input={"urls": [row["website"] for row in leads], "includeDns": True}
)
# apify-client 3.x returns a Run object; on 2.x use run["defaultDatasetId"]
results = {item["input"]: item for item in client.dataset(run.default_dataset_id).iterate_items()}
PLATFORM_CATEGORIES = ["Ecommerce", "CMS", "Headless CMS", "Static site generator"]
def first_in(item, categories):
for category in categories:
names = item.get("byCategory", {}).get(category)
if names:
return names[0]
return ""
def evidence_for(item, name):
for tech in item.get("technologies", []):
if tech["name"] == name:
return tech["evidence"][0]
return ""
segments = defaultdict(list)
rows = []
for lead in leads:
item = results.get(lead["website"], {})
if item.get("status") != "ok":
rows.append({**lead, "status": item.get("error", "not scanned")})
continue
platform = first_in(item, PLATFORM_CATEGORIES)
email = first_in(item, ["Email provider"])
segments[platform or "Custom / unknown"].append(lead["company"])
rows.append({
**lead,
"status": "ok",
"platform": platform,
"platform_evidence": evidence_for(item, platform) if platform else "",
"email_provider": email,
"hosting": first_in(item, ["Hosting", "CDN"]),
"analytics": ", ".join(item.get("byCategory", {}).get("Analytics", []) + item.get("byCategory", {}).get("Tag manager", [])),
"tech_count": item["technologyCount"],
})
with open("leads_enriched.csv", "w", newline="") as f:
fields = ["company", "website", "status", "platform", "platform_evidence", "email_provider", "hosting", "analytics", "tech_count"]
writer = csv.DictWriter(f, fieldnames=fields)
writer.writeheader()
writer.writerows(rows)
for platform, companies in sorted(segments.items(), key=lambda kv: -len(kv[1])):
print(f"{platform:<18} {len(companies)} {', '.join(companies)}")
If your list is large or you need the data refreshed on a schedule without running anything yourself, the Tech Stack Detector page has the pricing and limits, and our custom data feeds cover scheduled deliveries.
Related guides
Find the bugs your users are complaining about: mining App Store reviews with Python
Pull 1- and 2-star Apple App Store reviews for any app by name across several countries, group them by app version and tag them with explainable themes, in about 60 lines of Python. Tested on real data.
How to get YouTube transcripts with an API (Python, Node.js and curl)
A practical guide to fetching YouTube transcripts from code: why DIY scrapers get IP-blocked on servers, working Python, Node.js and curl examples, languages and translation, SRT/VTT, playlists, and chunking transcripts for LLMs.
Get the next guide by email
Practical tutorials on transcripts, screenshots, news data and agent tooling. A couple of emails a month at most.