Turn a list of company websites into contact details: emails, phones and socials in Python
You have a list of company websites and nothing else: no email, no phone, no LinkedIn page. Maybe it came from a conference exhibitor page, a directory export or a CRM with half the columns empty. Every prospecting workflow starts the same way: open the site, hunt for the contact page, copy an address into a spreadsheet.
That is mechanical work, so it can be scripted. But there is a trap on the way: scraping every address a website mentions gives you a column full of legal@, privacy@ and copyright@ mailboxes, plus the occasional named employee. Extraction is the easy half. The useful half is deciding which of the addresses you found is one you should actually write to.
This guide does both. We take six websites, extract public contact details from each, and produce an enriched CSV with one best email per company. Everything below was run on 2 October 2026 against real sites, and the output shown is the real output, including the cases where the honest answer is "nothing usable".
What you need
- Python 3.9 or newer (the code uses
str.removeprefix, which arrived in 3.9). - An Apify account and API token. The free plan includes monthly credits, far more than this tutorial uses.
- The official client:
pip install apify-client. The code uses version 3.x, where runs come back as objects (run.default_dataset_id). On the 2.x client, userun["defaultDatasetId"].
We use the Contact Details Extractor, which we publish on the Apify Store. Give it homepage URLs and it checks the homepage plus pages that look like contact, about or support pages, up to a page limit you set. It reads emails from visible text and mailto: links, phone numbers from text and tel: links, and social profiles from structured data (sameAs) and the site's header, footer and navigation. It costs $0.002 per website ($2 per 1,000), and sites that fail to load are free. This whole tutorial cost about one cent.
export APIFY_TOKEN="apify_api_..."
The lead list
leads.csv is deliberately small and mixed: a few software companies and one domain that does not exist, so we can see how failures look.
company,website
Basecamp,basecamp.com
Mozilla,mozilla.org
Ghost,ghost.org
Buffer,buffer.com
Plausible,plausible.io
Example (dead),this-domain-does-not-exist-sw123.com
Step 1: one run for the whole list
The Actor takes all the websites in one run and scans them in parallel, so you do not loop over rows:
run = client.actor("siftwright/contact-details-extractor").call(
run_input={"startUrls": [row["website"] for row in leads], "maxPagesPerDomain": 6}
)
results = {
item["domain"].removeprefix("www."): item
for item in client.dataset(run.default_dataset_id).iterate_items()
}
maxPagesPerDomain is the crawl depth per website (default 6, maximum 15). It does not change the price, because you are charged once per website, so raising it costs only time.
Notice that we key the results by domain, not by the string we sent. The Actor normalises each input to a full URL (basecamp.com comes back as https://basecamp.com), so a lookup on the original CSV value would miss. We normalise our own rows the same way with a small helper:
def domain_of(site):
host = urlparse(site if "//" in site else "https://" + site).hostname or ""
return host.removeprefix("www.")
What comes back
One item per website. Here is the real item for plausible.io:
{
"domain": "plausible.io",
"url": "https://plausible.io",
"status": "ok",
"emails": ["hello@plausible.io"],
"phones": [],
"facebook": null,
"twitter": "https://twitter.com/PlausibleHQ",
"linkedin": "https://www.linkedin.com/company/plausible-analytics/",
"instagram": null,
"youtube": null,
"pagesScanned": 6
}
The dead domain comes back as a row too, with a reason and no charge:
{
"domain": "this-domain-does-not-exist-sw123.com",
"status": "error",
"error": "fetch failed",
"pagesScanned": 0
}
Across the six sites, five loaded and the Actor's log reported "5 website(s) scanned, 1 failed to load (not charged)".
What the raw data looks like
This is where honest tooling matters, so here is every result for the five sites that loaded:
| Site | Emails found | Phones | Socials |
|---|---|---|---|
| basecamp.com | one address belonging to a named person | none | none |
| ghost.org | support@ghost.org |
none | none |
| plausible.io | hello@plausible.io |
none | X, LinkedIn |
| mozilla.org | trademark-permissions@mozilla.com |
none | LinkedIn, Instagram |
| buffer.com | hello@, legal@, copyright@, privacy@, security@ |
two | Facebook, X, LinkedIn, Instagram |
Three things stand out, and none of them is a bug:
- Most sites publish very little. Companies deliberately keep their public contact surface small. Two of five sites gave us no phone number at all, because they do not publish one. An extractor can only report what a site shows.
- The addresses are not all equal. Buffer publishes five. Only one is a way to start a business conversation. Mozilla's only public address on the pages checked is a trademark-permissions mailbox, which is a department for one specific purpose.
- The data is as the website wrote it. Buffer's second phone number came back as
+1-800)-952-5210, with a stray bracket exactly as the page formatted it. We clean that on our side.
Raw extraction output is a list of candidates, not a finished lead list. Hence step 2.
Step 2: rank the emails
We sort addresses into three buckets by their local part, the bit before the @:
# Role mailboxes you can usually write to about business, best first.
PREFERRED = ["sales", "hello", "contact", "info", "team", "partnerships", "support"]
# Mailboxes that exist for something else: never use them for outreach.
AVOID = ("legal", "privacy", "security", "copyright", "abuse", "dmca", "press",
"trademark", "noreply", "no-reply", "careers", "jobs", "billing")
def rank_email(email):
local = email.split("@")[0].lower()
if any(word in local for word in AVOID):
return 99
for i, word in enumerate(PREFERRED):
if local == word:
return i
return 50 # looks like a named person: keep it out of automated outreach
def best_email(emails):
ranked = sorted(emails, key=rank_email)
if ranked and rank_email(ranked[0]) < 50:
return ranked[0]
return ""
The rule is conservative on purpose. A role address like hello@ is a mailbox the company set up to receive messages from strangers. A legal@ or privacy@ mailbox is for formal requests, and a cold pitch there is both ineffective and annoying. An address that looks like one person's name, such as the one we found on Basecamp's site, is a personal address. Even when it is published, it is better left out of automated mailings, and it is personal data in the legal sense, which matters if you or your recipients are in the EU or UK. best_email returns an empty string in that case, and we keep the original in an all_emails column so a human can decide.
Phones need tidying because they arrive as they appeared on the page:
def clean_phone(phone):
# The Actor returns phones as found on the page, so tidy them here.
phone = re.sub(r"[^\d+]", "", phone)
return phone if 8 <= len(phone) <= 15 else ""
Stripping everything except digits and + turns +1-800)-952-5210 into +18009525210, and the length check (E.164 numbers have at most 15 digits) drops fragments that are not phone numbers.
Step 3: write the enriched CSV
rows = []
for lead in leads:
item = results.get(domain_of(lead["website"]), {})
if item.get("status") != "ok":
rows.append({**lead, "status": item.get("error", "not scanned")})
continue
emails = item["emails"]
rows.append({
**lead,
"status": "ok",
"best_email": best_email(emails),
"all_emails": " ".join(emails),
"phone": next((p for p in map(clean_phone, item["phones"]) if p), ""),
"linkedin": item["linkedin"] or "",
"twitter": item["twitter"] or "",
})
Failed sites stay in the file with their error instead of vanishing, so nobody wonders why a company is missing and you can fix the typo and re-run only those rows. The full script, which also writes leads_enriched.csv, is at the end. Its real output:
Basecamp ok - -
Mozilla ok - -
Ghost ok support@ghost.org -
Buffer ok hello@buffer.com +18007787879
Plausible ok hello@plausible.io -
Example (dead) fetch failed - -
Three of the five loaded sites yield a usable address (two hello@ and one support@), and two yield none, which is a realistic result for modern software companies that hide behind contact forms. Lists of local businesses, agencies and shops usually do better, since they tend to publish a plain info@ address and a phone number. Test on a sample of your own list first; at $2 per 1,000 websites that costs next to nothing.
What this does not do
- It does not guess or verify addresses. It reports what the site publishes. It does not generate
firstname@company.compatterns or check whether a mailbox exists. - It does not read contact forms, JavaScript-only pages or anything behind a login. If a site shows its address only after a script runs, it will be missed.
- It checks a limited number of pages. The default is 6 pages per site; an address buried on page 40 is out of reach.
- Social links come from structured data and header, footer and navigation only. That is deliberate: a page-wide scan picks up links to other people's profiles in blog posts and comments.
Use the data responsibly
Public does not mean unrestricted. Before you email anyone, check the rules that apply to you: CAN-SPAM in the US, and GDPR plus the ePrivacy rules in the EU and UK, which treat a named person's work address as personal data and set conditions for marketing to it. Role mailboxes at companies are the easiest case. Always identify yourself, say why you are writing, and honour opt-outs. This guide is not legal advice.
The full script
import csv
import os
import re
from urllib.parse import urlparse
from apify_client import ApifyClient
def domain_of(site):
host = urlparse(site if "//" in site else "https://" + site).hostname or ""
return host.removeprefix("www.")
client = ApifyClient(os.environ["APIFY_TOKEN"])
with open("leads.csv", newline="") as f:
leads = list(csv.DictReader(f))
run = client.actor("siftwright/contact-details-extractor").call(
run_input={"startUrls": [row["website"] for row in leads], "maxPagesPerDomain": 6}
)
results = {
item["domain"].removeprefix("www."): item
for item in client.dataset(run.default_dataset_id).iterate_items()
}
# Role mailboxes you can usually write to about business, best first.
PREFERRED = ["sales", "hello", "contact", "info", "team", "partnerships", "support"]
# Mailboxes that exist for something else: never use them for outreach.
AVOID = ("legal", "privacy", "security", "copyright", "abuse", "dmca", "press",
"trademark", "noreply", "no-reply", "careers", "jobs", "billing")
def rank_email(email):
local = email.split("@")[0].lower()
if any(word in local for word in AVOID):
return 99
for i, word in enumerate(PREFERRED):
if local == word:
return i
return 50 # looks like a named person: keep it out of automated outreach
def best_email(emails):
ranked = sorted(emails, key=rank_email)
if ranked and rank_email(ranked[0]) < 50:
return ranked[0]
return ""
def clean_phone(phone):
# The Actor returns phones as found on the page, so tidy them here.
phone = re.sub(r"[^\d+]", "", phone)
return phone if 8 <= len(phone) <= 15 else ""
rows = []
for lead in leads:
item = results.get(domain_of(lead["website"]), {})
if item.get("status") != "ok":
rows.append({**lead, "status": item.get("error", "not scanned")})
continue
emails = item["emails"]
rows.append({
**lead,
"status": "ok",
"best_email": best_email(emails),
"all_emails": " ".join(emails),
"phone": next((p for p in map(clean_phone, item["phones"]) if p), ""),
"linkedin": item["linkedin"] or "",
"twitter": item["twitter"] or "",
})
fields = ["company", "website", "status", "best_email", "all_emails", "phone", "linkedin", "twitter"]
with open("leads_enriched.csv", "w", newline="") as f:
writer = csv.DictWriter(f, fieldnames=fields)
writer.writeheader()
writer.writerows(rows)
for row in rows:
print(f"{row['company']:<15} {row['status']:<12} {row.get('best_email') or '-':<20} {row.get('phone') or '-'}")
Next steps
- Combine this with the Tech Stack Detector to segment the list first, as in find every Shopify, WordPress and HubSpot site in a lead list, and only extract contacts for the segment you actually sell to.
- If you prefer a plain HTTP API with an API key over the Apify client, the same extraction is available as
POST /v1/contact-detailsin the Siftwright API. For the cost comparison, see pay-per-result vs subscription pricing.
Questions or a site that should work and does not? Write to support@siftwright.com.
Related guides
Find every Shopify, WordPress and HubSpot site in a lead list: tech stack lookup with evidence in Python
Take a CSV of company websites, detect each site's CMS, ecommerce platform, email provider and hosting, and split the list into segments you can act on. About 70 lines of Python, tested on real sites, with the evidence behind every detection.
Find the bugs your users are complaining about: mining App Store reviews with Python
Pull 1- and 2-star Apple App Store reviews for any app by name across several countries, group them by app version and tag them with explainable themes, in about 60 lines of Python. Tested on real data.
Get the next guide by email
Practical tutorials on transcripts, screenshots, news data and agent tooling. A couple of emails a month at most.