Why every contact in your CRM needs a source
A contact without a source is a liability: you can't judge if it's current, explain where it came from, or defend using it. Here's what good provenance looks like and how to add it to the CRM you already have.
- CRM
- data quality
- provenance
- GDPR
Open any sales or fundraising CRM and pick a contact at random. Now answer three questions about it:
- Where did this email address come from?
- When was it last known to be correct?
- Why are we allowed to contact this person?
For most teams, the honest answer to all three is "no idea". The contact was in a spreadsheet someone imported, or a list bought two years ago, or an address a tool guessed. It looks exactly like a contact you found yesterday on the person's own website, and that's the problem.
This article makes the case that every contact, and ideally every field, should carry a source, and shows how to get there without rebuilding your stack.
What "source" means
Provenance is a small set of facts attached to a piece of data:
- Where it came from: a URL, a filing, a file you imported, a conversation.
- When it was found or last confirmed.
- How it was obtained: published on a page, inferred from a pattern, provided by the person, bought from a vendor.
- What has been checked since: for an email, the verification result and date.
Good provenance is field-level, not record-level. A person's name and title might come from an SEC filing, their email from their company's team page, and their phone number from a conference speaker list. Each has a different age and a different reliability. Collapsing them into one "source: web" label throws that away.
Reason 1: data decays, and you need to know how old it is
People change jobs, companies rebrand, domains lapse and roles get reorganised. A contact that was accurate when you found it drifts out of date quietly, and nothing in a typical CRM tells you.
With a found date on every field, you can make sensible decisions:
- Re-verify email addresses that haven't been checked in the last 90 days before a campaign.
- Prefer a title found last month over one found two years ago.
- Spot entire imports that are old, and treat them with suspicion.
Without dates, every contact looks equally fresh, so you either trust stale data or distrust everything.
Reason 2: you can't judge quality without knowing the method
"Found on the person's own company website" and "generated by combining a first name with a domain" are very different claims about an email address. The first is strong evidence the address exists. The second is a guess that needs verification before you send anything, and even then may be a catch-all domain accepting mail for any address.
Recording the method lets you treat each contact appropriately. It also keeps you honest: if a large share of your pipeline is guessed addresses, you'll see it, and you'll understand your bounce rate.
Reason 3: the law increasingly expects it
Privacy rules are converging on the idea that you should be able to explain where personal data came from.
- Under the GDPR, when you collect personal data from somewhere other than the person, Article 14 requires you to tell them the categories of data and the source it came from, including whether it came from publicly accessible sources. You can't do that if you don't know.
- Data subjects can ask for access to their data under Article 15, which also includes any available information about its source.
- In the United States, several states require companies that sell data about people they don't have a direct relationship with to register as data brokers. California's Delete Act goes further, creating a state platform through which people can ask registered data brokers to delete their data, with brokers required to begin processing those requests from August 2026.
Even where the law is lighter, a documented source is your best answer when someone asks, "How did you get my email?" A specific, truthful reply ("from the team page on your company's website, on September 12") ends most of those conversations well. "We bought a list" rarely does.
Reason 4: it makes outreach better
Provenance isn't only defensive. The source of a contact is often the best reason to contact them.
- An investor listed on a recent fund filing has just raised money. Say that.
- A researcher who published on your topic last month is thinking about it right now. Mention the paper.
- A head of operations quoted in a trade article about the problem you solve has told you what they care about.
"I found you here, and here's why that made me think of you" is honest personalization. It's more relevant than a merge tag and it reads like a human wrote it, because one did.
Reason 5: removals only work if you can find the data
When someone asks to be removed or never contacted again, you have to find every copy of their data: in the CRM, in old imports, in sequences, in exports teammates made. Source records make that possible. They also let you honour a removal permanently, by recognising the same address if it shows up in a future import and suppressing it.
How to add provenance to the CRM you already have
You don't need to migrate to get most of the benefit. Start with four fields on your contact object:
| Field | Example | Notes |
|---|---|---|
| Source URL | https://example.com/team | Or a file name and row for imports |
| Source type | Published, inferred, provided, purchased | A short controlled list |
| Found at | 2026-09-12 | The date the value was found or confirmed |
| Lawful basis | Legitimate interest | Required thinking for EU and UK contacts |
Then:
- Backfill imports honestly. For existing records, set the source to the import file and date if you know them. If you don't, mark them "unknown" rather than inventing something.
- Make source mandatory for new contacts. A required field at creation is the only way this sticks.
- Re-verify on a schedule. Anything older than a quarter gets checked before it's used in a campaign.
- Decide what to do with the unknowns. Contacts with no source, no recent verification and no engagement are costing you deliverability and risk. Consider archiving them.
- Keep a suppression list separate from contacts. Opt-outs should outlive the contact record, so a re-import can't bring someone back.
How OmniLead does it
We built OmniLead around this idea. Every field on every contact, whether an investor from an SEC Form D filing, a CTO from a company's team page, a researcher from OpenAlex or ORCID, or a grant from Grants.gov, stores the URL it came from, the provider and the date we found it. Click the link icon on any result and you see the receipt.
Nothing is generated by a language model. When an email isn't published, OmniLead ranks candidates from the pattern of real addresses found on the same domain, verifies them, and labels the result as pattern-based so you know what you're looking at. Imports keep their file and date as the source. Every lead stores a lawful basis, and removals from our opt-out portal apply across every workspace.
It's a small discipline that pays off every time someone asks where a contact came from, including you.
Start with contacts that come with receipts
Search investors, companies, researchers and grants with a source on every field. Free to start, no card needed.