All articles
Lead GenerationBy Efe Berke Çolaker 10 min read

Auditing a Lead List You Did Not Build

A one hour audit for an inherited or purchased list: seven checks, the numbers to compare against, and the point where rebuilding beats cleaning.

ON THIS PAGE
  1. 01The seven checks, in order
  2. 02What to compare the numbers against
  3. 03The decision the audit produces
  4. 04The provenance question, which outranks
  5. 05Writing the audit down
  6. 06Sources and method
  7. 07FAQ
Auditing a Lead List You Did Not Build

By Efe Berke Colaker, Founder at GetleadReviewed by the Getlead editorial team for accuracy. Last updated August 2026.

Someone hands you a file. It came from a vendor, a previous employee or an old campaign, and the only thing anyone can say about it is that it has 40,000 rows.

The temptation is to start cleaning. The better move is an hour of measurement first, because the audit frequently shows that cleaning is not the right investment.

KEY TAKEAWAYS
Audit before you clean. Cleaning a list that should be discarded is the most common way teams spend a week to reach the same conclusion.
Seven checks, one hour, no sending. Nothing in this audit requires a campaign or risks a domain.
The baseline to compare against: 43.4% confirmed valid, 23.9% invalid and 16.7% catch-all on raw B2B data in our verification.
A list with no provenance is a list you cannot defend, whatever its accuracy turns out to be.

The seven checks, in order

A lead list audit is a measurement pass that establishes what a file contains before any decision is made about using it.

  1. Provenance. Can anyone say where each record came from and when? Not the file, the records.
  2. Duplicates. Rate on the address, then on the canonical company domain.
  3. Role addresses. Share of shared inboxes such as info@, sales@ and support@.
  4. ICP match. Sample 50 rows by hand and check title, size and geography against your profile.
  5. Verification sample. Run 500 addresses through SMTP verification and record the four states.
  6. Age indicators. Titles and company names that no longer exist, which date the file.
  7. Suppression collisions. How many rows are already customers, open opportunities or prior opt-outs.

For example, an audit that finds 12% duplicates, 9% role addresses and a 31% confirmed valid rate has told you the file is a scrape, without needing to ask the person who supplied it.

Methodology: we analyzed 383,368 email addresses through live SMTP verification and measured 34,973 tracked outbound sends inside Getlead, aggregated and anonymized at campaign level. Every platform number here is what the mail servers and the campaigns returned, not a vendor claim. Sample and limitations are in the benchmark study.

What to compare the numbers against

An audit without a baseline produces numbers nobody can interpret. These are the reference points worth using.

CHECKHEALTHYINVESTIGATE
Confirmed validNear or above 43%Below 35%
InvalidNear 24% on raw dataAbove 35%
Catch-allAround 17%Above 30%
Duplicates on domainUnder 5%Above 10%
Role addressesUnder 5%Above 10%
ICP match on sampleAbove 70%Below 50%

The first three come from our verification of 383,368 raw B2B addresses, which returned 43.4% confirmed valid, 23.9% invalid and 16.7% catch-all. That is what ordinary raw data looks like, so a file well below it is not ordinary.

An unusually high catch-all share is the clearest tell of speculative data. Catch-all domains accept everything, so pattern guessed addresses accumulate there and never resolve to a confirmation.

The decision the audit produces

Three outcomes, and the middle one is where most inherited files land.

  • Usable. Verify, deduplicate, suppress and send. The audit found ordinary data with traceable origins.
  • Salvageable in part. One segment holds up and the rest does not. Keep the segment, retire the remainder.
  • Discard. No provenance, poor verification results, heavy role and catch-all share. Cleaning this costs more than sourcing fresh records.
43.4%confirmed valid baseline
16.7%catch-all baseline
0.51%our bounce after verification

For example, a 40,000 row file where only the UK subset scores well is not a 40,000 row asset. It is a 6,000 row asset plus 34,000 rows of storage cost and risk, and treating it as the former is how a domain gets burned.

Discarding feels wasteful because someone paid for the file. The sunk cost is already spent, and the choice in front of you is only whether to add sending reputation to the loss.

Skip the inherited file
Getlead includes a 420M+ verified B2B database, SMTP verification at export, warm-up and cold email sending. From $19.90 a month.
See pricing

The provenance question, which outranks accuracy

A file can verify well and still be unusable, because accuracy and defensibility are separate properties.

Provenance is the per record answer to where a contact came from and when. Without it you cannot tell a sourced record from a guessed one, cannot prioritise by quality, and cannot answer a recipient who asks.

Where recipients are in the EU or UK, that answer is not optional: obtaining personal data from a third party triggers a notice duty under Article 14 of the GDPR, and a file nobody can trace makes compliance a matter of hope.

If the audit cannot establish provenance, record that explicitly rather than leaving the field empty, so the limitation travels with the data instead of being rediscovered a year later.

Writing the audit down

One page, six numbers and a recommendation. The point is that the next person does not repeat the work.

  1. Row count, and how many survive deduplication.
  2. Verification split across confirmed valid, invalid, catch-all and unknown on the sample.
  3. Role address share and ICP match rate from the manual sample.
  4. Suppression collisions found against customers and prior opt-outs.
  5. Provenance status: traceable, partially traceable, or unknown.
  6. The recommendation, with the reason stated in one sentence.

Attach the audit to the records themselves where you can, as a source note. A file that arrives with its own measurement history is worth materially more than the same rows without it.

Sources and method

First-party data (Getlead, 2026): the verification split of 43.4% confirmed valid, 23.9% invalid, 16.7% catch-all and 16.0% unknown comes from 383,368 addresses analyzed through live SMTP verification, and the 0.51% bounce rate comes from 34,973 tracked sends, aggregated and anonymized at campaign level. Full method in our cold email benchmark study.

External sources: notice duties when personal data is obtained from a third party are set out in Article 14 of the GDPR; US commercial email obligations including a working opt-out come from the FTC CAN-SPAM compliance guide.

Baselines describe raw B2B data in our own sample and will differ by market and sourcing method. Checked in August 2026.

Frequently asked questions

How do I audit a lead list?

Seven checks in one hour: provenance per record, duplicate rate on address and canonical domain, role address share, ICP match on a manual sample of 50, a verification sample of 500 addresses, age indicators such as defunct titles, and collisions with customers and prior opt-outs.

What verification results should I expect from a normal list?

Our verification of 383,368 raw B2B addresses returned 43.4% confirmed valid, 23.9% invalid and 16.7% catch-all. A file well below 35% confirmed valid, or well above 30% catch-all, is not ordinary raw data.

What does a high catch-all share indicate?

Usually speculative or pattern guessed data. Catch-all domains accept mail for every address, so guessed addresses accumulate there and never resolve to a confirmation, which makes an unusually high share the clearest tell that a file was assembled rather than sourced.

When should I discard a list instead of cleaning it?

When there is no provenance, verification results are far below baseline, and role plus catch-all share is high. The money spent on the file is already gone, and the only remaining decision is whether to add sender reputation damage to that loss.

Why does provenance matter more than accuracy?

Because a file can verify well and still be indefensible. Without a per record source you cannot distinguish sourced from guessed data, cannot prioritise by quality, and cannot meet the Article 14 notice duty when recipients are in the EU or UK.

Does auditing require sending any email?

No, and it should not. Every check runs against the file itself or through SMTP verification, which asks the receiving server whether a mailbox exists without delivering anything, so no sending reputation is at risk during the audit.

Popular resources

15 best lead generation tools12 best sales prospecting toolsLead scrapers for 10+ sourcesLead scraping tool (50K leads/mo)B2B email lists by industryB2B lead generation guideBest lead gen tools for agenciesUplead vs ApolloZoominfo vs LushaWiza vs Lusha

More in B2B Data Ops

What Is a Lead Database? The Difference From a ListHow to Build a B2B Prospect List That Is Ready to ContactThe B2B Marketing Database: Build It, Segment It, Keep It AliveLead Deduplication: Merging Records Without Losing the Good OneLead Scoring for Outbound, Where Nobody Raised Their Hand YetCleaning a Lead CSV: The Ten Passes That Matter
Open the full b2b data ops guide

Customer reviews

2,400+ users. Real results.

Don't take our word for it

Replace your whole lead gen stack

Lead scraping, a 420M+ B2B database, email verification and cold email sending in one subscription. No credits, no seat pricing, cancel anytime.

Start from $19.90/mo
14-day money-back guarantee Instant access 12,400+ teams