Turn a list of company websites into contact details: emails, phones and socials in Python

By Siftwright team8 min read

You have a list of company websites and nothing else: no email, no phone, no LinkedIn page. Maybe it came from a conference exhibitor page, a directory export or a CRM with half the columns empty. Every prospecting workflow starts the same way: open the site, hunt for the contact page, copy an address into a spreadsheet.

That is mechanical work, so it can be scripted. But there is a trap on the way: scraping every address a website mentions gives you a column full of legal@, privacy@ and copyright@ mailboxes, plus the occasional named employee. Extraction is the easy half. The useful half is deciding which of the addresses you found is one you should actually write to.

This guide does both. We take six websites, extract public contact details from each, and produce an enriched CSV with one best email per company. Everything below was run on 2 October 2026 against real sites, and the output shown is the real output, including the cases where the honest answer is "nothing usable".

What you need

  • Python 3.9 or newer (the code uses str.removeprefix, which arrived in 3.9).
  • An Apify account and API token. The free plan includes monthly credits, far more than this tutorial uses.
  • The official client: pip install apify-client. The code uses version 3.x, where runs come back as objects (run.default_dataset_id). On the 2.x client, use run["defaultDatasetId"].

We use the Contact Details Extractor, which we publish on the Apify Store. Give it homepage URLs and it checks the homepage plus pages that look like contact, about or support pages, up to a page limit you set. It reads emails from visible text and mailto: links, phone numbers from text and tel: links, and social profiles from structured data (sameAs) and the site's header, footer and navigation. It costs $0.002 per website ($2 per 1,000), and sites that fail to load are free. This whole tutorial cost about one cent.

export APIFY_TOKEN="apify_api_..."

The lead list

leads.csv is deliberately small and mixed: a few software companies and one domain that does not exist, so we can see how failures look.

company,website
Basecamp,basecamp.com
Mozilla,mozilla.org
Ghost,ghost.org
Buffer,buffer.com
Plausible,plausible.io
Example (dead),this-domain-does-not-exist-sw123.com

Step 1: one run for the whole list

The Actor takes all the websites in one run and scans them in parallel, so you do not loop over rows:

run = client.actor("siftwright/contact-details-extractor").call(
    run_input={"startUrls": [row["website"] for row in leads], "maxPagesPerDomain": 6}
)
results = {
    item["domain"].removeprefix("www."): item
    for item in client.dataset(run.default_dataset_id).iterate_items()
}

maxPagesPerDomain is the crawl depth per website (default 6, maximum 15). It does not change the price, because you are charged once per website, so raising it costs only time.

Notice that we key the results by domain, not by the string we sent. The Actor normalises each input to a full URL (basecamp.com comes back as https://basecamp.com), so a lookup on the original CSV value would miss. We normalise our own rows the same way with a small helper:

def domain_of(site):
    host = urlparse(site if "//" in site else "https://" + site).hostname or ""
    return host.removeprefix("www.")

What comes back

One item per website. Here is the real item for plausible.io:

{
  "domain": "plausible.io",
  "url": "https://plausible.io",
  "status": "ok",
  "emails": ["hello@plausible.io"],
  "phones": [],
  "facebook": null,
  "twitter": "https://twitter.com/PlausibleHQ",
  "linkedin": "https://www.linkedin.com/company/plausible-analytics/",
  "instagram": null,
  "youtube": null,
  "pagesScanned": 6
}

The dead domain comes back as a row too, with a reason and no charge:

{
  "domain": "this-domain-does-not-exist-sw123.com",
  "status": "error",
  "error": "fetch failed",
  "pagesScanned": 0
}

Across the six sites, five loaded and the Actor's log reported "5 website(s) scanned, 1 failed to load (not charged)".

What the raw data looks like

This is where honest tooling matters, so here is every result for the five sites that loaded:

Site Emails found Phones Socials
basecamp.com one address belonging to a named person none none
ghost.org support@ghost.org none none
plausible.io hello@plausible.io none X, LinkedIn
mozilla.org trademark-permissions@mozilla.com none LinkedIn, Instagram
buffer.com hello@, legal@, copyright@, privacy@, security@ two Facebook, X, LinkedIn, Instagram

Three things stand out, and none of them is a bug:

  1. Most sites publish very little. Companies deliberately keep their public contact surface small. Two of five sites gave us no phone number at all, because they do not publish one. An extractor can only report what a site shows.
  2. The addresses are not all equal. Buffer publishes five. Only one is a way to start a business conversation. Mozilla's only public address on the pages checked is a trademark-permissions mailbox, which is a department for one specific purpose.
  3. The data is as the website wrote it. Buffer's second phone number came back as +1-800)-952-5210, with a stray bracket exactly as the page formatted it. We clean that on our side.

Raw extraction output is a list of candidates, not a finished lead list. Hence step 2.

Step 2: rank the emails

We sort addresses into three buckets by their local part, the bit before the @:

# Role mailboxes you can usually write to about business, best first.
PREFERRED = ["sales", "hello", "contact", "info", "team", "partnerships", "support"]
# Mailboxes that exist for something else: never use them for outreach.
AVOID = ("legal", "privacy", "security", "copyright", "abuse", "dmca", "press",
         "trademark", "noreply", "no-reply", "careers", "jobs", "billing")


def rank_email(email):
    local = email.split("@")[0].lower()
    if any(word in local for word in AVOID):
        return 99
    for i, word in enumerate(PREFERRED):
        if local == word:
            return i
    return 50  # looks like a named person: keep it out of automated outreach


def best_email(emails):
    ranked = sorted(emails, key=rank_email)
    if ranked and rank_email(ranked[0]) < 50:
        return ranked[0]
    return ""

The rule is conservative on purpose. A role address like hello@ is a mailbox the company set up to receive messages from strangers. A legal@ or privacy@ mailbox is for formal requests, and a cold pitch there is both ineffective and annoying. An address that looks like one person's name, such as the one we found on Basecamp's site, is a personal address. Even when it is published, it is better left out of automated mailings, and it is personal data in the legal sense, which matters if you or your recipients are in the EU or UK. best_email returns an empty string in that case, and we keep the original in an all_emails column so a human can decide.

Phones need tidying because they arrive as they appeared on the page:

def clean_phone(phone):
    # The Actor returns phones as found on the page, so tidy them here.
    phone = re.sub(r"[^\d+]", "", phone)
    return phone if 8 <= len(phone) <= 15 else ""

Stripping everything except digits and + turns +1-800)-952-5210 into +18009525210, and the length check (E.164 numbers have at most 15 digits) drops fragments that are not phone numbers.

Step 3: write the enriched CSV

rows = []
for lead in leads:
    item = results.get(domain_of(lead["website"]), {})
    if item.get("status") != "ok":
        rows.append({**lead, "status": item.get("error", "not scanned")})
        continue
    emails = item["emails"]
    rows.append({
        **lead,
        "status": "ok",
        "best_email": best_email(emails),
        "all_emails": " ".join(emails),
        "phone": next((p for p in map(clean_phone, item["phones"]) if p), ""),
        "linkedin": item["linkedin"] or "",
        "twitter": item["twitter"] or "",
    })

Failed sites stay in the file with their error instead of vanishing, so nobody wonders why a company is missing and you can fix the typo and re-run only those rows. The full script, which also writes leads_enriched.csv, is at the end. Its real output:

Basecamp        ok           -                    -
Mozilla         ok           -                    -
Ghost           ok           support@ghost.org    -
Buffer          ok           hello@buffer.com     +18007787879
Plausible       ok           hello@plausible.io   -
Example (dead)  fetch failed -                    -

Three of the five loaded sites yield a usable address (two hello@ and one support@), and two yield none, which is a realistic result for modern software companies that hide behind contact forms. Lists of local businesses, agencies and shops usually do better, since they tend to publish a plain info@ address and a phone number. Test on a sample of your own list first; at $2 per 1,000 websites that costs next to nothing.

What this does not do

  • It does not guess or verify addresses. It reports what the site publishes. It does not generate firstname@company.com patterns or check whether a mailbox exists.
  • It does not read contact forms, JavaScript-only pages or anything behind a login. If a site shows its address only after a script runs, it will be missed.
  • It checks a limited number of pages. The default is 6 pages per site; an address buried on page 40 is out of reach.
  • Social links come from structured data and header, footer and navigation only. That is deliberate: a page-wide scan picks up links to other people's profiles in blog posts and comments.

Use the data responsibly

Public does not mean unrestricted. Before you email anyone, check the rules that apply to you: CAN-SPAM in the US, and GDPR plus the ePrivacy rules in the EU and UK, which treat a named person's work address as personal data and set conditions for marketing to it. Role mailboxes at companies are the easiest case. Always identify yourself, say why you are writing, and honour opt-outs. This guide is not legal advice.

The full script

import csv
import os
import re
from urllib.parse import urlparse

from apify_client import ApifyClient


def domain_of(site):
    host = urlparse(site if "//" in site else "https://" + site).hostname or ""
    return host.removeprefix("www.")


client = ApifyClient(os.environ["APIFY_TOKEN"])

with open("leads.csv", newline="") as f:
    leads = list(csv.DictReader(f))

run = client.actor("siftwright/contact-details-extractor").call(
    run_input={"startUrls": [row["website"] for row in leads], "maxPagesPerDomain": 6}
)
results = {
    item["domain"].removeprefix("www."): item
    for item in client.dataset(run.default_dataset_id).iterate_items()
}

# Role mailboxes you can usually write to about business, best first.
PREFERRED = ["sales", "hello", "contact", "info", "team", "partnerships", "support"]
# Mailboxes that exist for something else: never use them for outreach.
AVOID = ("legal", "privacy", "security", "copyright", "abuse", "dmca", "press",
         "trademark", "noreply", "no-reply", "careers", "jobs", "billing")


def rank_email(email):
    local = email.split("@")[0].lower()
    if any(word in local for word in AVOID):
        return 99
    for i, word in enumerate(PREFERRED):
        if local == word:
            return i
    return 50  # looks like a named person: keep it out of automated outreach


def best_email(emails):
    ranked = sorted(emails, key=rank_email)
    if ranked and rank_email(ranked[0]) < 50:
        return ranked[0]
    return ""


def clean_phone(phone):
    # The Actor returns phones as found on the page, so tidy them here.
    phone = re.sub(r"[^\d+]", "", phone)
    return phone if 8 <= len(phone) <= 15 else ""


rows = []
for lead in leads:
    item = results.get(domain_of(lead["website"]), {})
    if item.get("status") != "ok":
        rows.append({**lead, "status": item.get("error", "not scanned")})
        continue
    emails = item["emails"]
    rows.append({
        **lead,
        "status": "ok",
        "best_email": best_email(emails),
        "all_emails": " ".join(emails),
        "phone": next((p for p in map(clean_phone, item["phones"]) if p), ""),
        "linkedin": item["linkedin"] or "",
        "twitter": item["twitter"] or "",
    })

fields = ["company", "website", "status", "best_email", "all_emails", "phone", "linkedin", "twitter"]
with open("leads_enriched.csv", "w", newline="") as f:
    writer = csv.DictWriter(f, fieldnames=fields)
    writer.writeheader()
    writer.writerows(rows)

for row in rows:
    print(f"{row['company']:<15} {row['status']:<12} {row.get('best_email') or '-':<20} {row.get('phone') or '-'}")

Next steps

Questions or a site that should work and does not? Write to support@siftwright.com.

Related guides

Get the next guide by email

Practical tutorials on transcripts, screenshots, news data and agent tooling. A couple of emails a month at most.