Years of exports, banker books and badge scans have filled your pipeline with records — and quietly fixed its boundaries. We screen the entire active web against your thesis and show you exactly what your list is missing, with evidence.
None of these sources is wrong. Each one is a filter that was applied before you ever saw the company — and the filters compound. A gap analysis measures what they filtered out.
Profile databases index companies they previously found and tagged. Your export inherits their coverage, their taxonomy and their refresh lag — invisibly.
Processes you were shown, teasers you passed on, lists assembled for someone else's mandate. Curated for a fee, not for your thesis.
Exhibitors and attendees from the shows your team happened to attend. A marketing-budget sample of the market, three years stale.
Companies that found you, or that a contact remembered. Valuable — and structurally biased toward the visible and the already-networked.
Leading company databases start from the companies they found. We start from the entire active web — 100M+ classified domains, 24.7M of them business and finance sites — and run LLM analysis against your written thesis.
"Your target list, reconciled against a full-web screen of the same thesis: every fitting company you hold confirmed and enriched, every fitting company you lack surfaced, and every non-fitting record you hold flagged — each with quoted evidence and a source URL."
The output is three clean buckets, not a pile of lookalikes. Your team decides what to do with each bucket; nothing is overwritten, and nothing arrives without its reason attached.
The same two-pass architecture behind our full-web method — a broad triage, then deep extraction — bracketed by a normalization step in and a reconciliation step out.
We ingest your CRM export and database lists, resolve domains, and deduplicate. No credentials needed — a CSV of names and websites is enough.
The relevant slice of the 100M+ universe is cut by category and geography, then triaged for live, operating, in-scope companies.
Survivors get full-site reading against your thesis: 15 signals, three scores weighted 70/20/10, and verbatim evidence per claim.
The scored universe is matched back to your records. Every company lands in one of three buckets, every bucket ships with its proof.
This is the honest comparison. Your CRM plus profile databases are genuinely good at some of these rows. The rows they can't answer are where mandates quietly leak.
| Coverage question | CRM + database exports | Full-web gap run |
|---|---|---|
| Starting universe | The companies someone previously found — thousands of records | 100M+ classified domains; 24.7M business & finance sites; 700+ categories |
| How a company gets in | Registered, exhibited, got profiled, or met your team | Operates a live website. That is the entire entry requirement. |
| Fits without the obvious keywords | Invisible — search runs on labels the company never used | Read, not searched: a fifth or more of confirmed fits lack their category's obvious homepage keywords |
| Group-ownership screening | Rarely recorded; usually discovered mid-outreach | Explicitly screened — roughly 1 in 10 keyword-perfect candidates is already group-owned |
| Evidence per record | A profile compiled somewhere, sometime, by someone | Verbatim snippet + source URL for every inclusion and every exclusion |
| Refresh | Whenever the vendor or an analyst touches the record | Re-runs on any subset with a custom ICP; monthly deltas on the annual tier |
Records in your CRM that survive the full screen — returned with 15 extracted signals, three scores, certifications as exact claim text, and quoted evidence, so your team can re-rank the pipeline before spending another outreach cycle on it.
Companies that match your thesis and appear nowhere in your systems — consistently a larger bucket than teams expect, because many fits never used the keywords your existing sources search on.
Records you're carrying that fail the screen — group-owned subsidiaries, wrong service model, dormant sites — each with the documented reason, because deleting dead records is worth almost as much as finding live ones.
From a real run: a single industrial-services category held 367,478 domains globally. A 25,000-domain US-focused triage confirmed ~17,300 live operating companies — far more than any target list we have audited actually held.
All 15 signals ship with every record. These four do the heaviest lifting when the question is "what is wrong with the list we already have."
The single most common defect in aged CRM data — a company absorbed into a group after its record was created, invisible until outreach. We read for subsidiary language, "part of the X family" phrasing, and acquisition announcements, zeroing the transition score on group ownership; roughly one in ten keyword-perfect candidates fails here.
Database exports rarely carry ownership context; websites state it constantly — "second generation", "founder-led", "family-owned since 1981" — and we capture that language verbatim. In industrial categories, roughly half of confirmed fits carry explicit founder or family evidence on their own pages, making this the column your CRM is missing.
Founding year, decades-in-operation claims, and independence language, taken from what the company says about itself — separating a durable target from a two-year-old site wearing the same keywords, and correcting the quiet CRM failure where directory-sourced records carry no age context at all.
The most recent dated content on the site — posts, project pages, press items — tells you whether a company is visibly active or coasting on a 2019 rebuild, because aged pipelines accumulate dormant records and flagging them saves the outreach cycles that would have discovered the silence one email at a time.
Our public specimen follows the same 8 / 5 / 5 / 2 discipline: eight top fits, five keyword-missed fits, five documented exclusions, two insufficient-evidence cases. Names are anonymized on the public site; clients receive source URLs.
"a fourth-generation, family-owned CNC company"
Held in the client's CRM as a bare name. Returned with AS9100D and ITAR captured as exact claim text, founder context, and a 15-signal profile — the record went from a row to a case.
"family-built, American-owned since 1965"
Absent from every list the client held. The site never uses the category's standard keywords — a keyword search cannot find it, and none had. Full-site reading scored it as a direct thesis fit.
"a wholly owned subsidiary"
Sat in the pipeline for two years as an active target. One quoted line from its own About page ends the debate — already consolidated, transition score zeroed, record closed with the reason attached.
The uncomfortable arithmetic: a CRM audit almost never shrinks the pipeline. It swaps records you would have wasted cycles on for companies you did not know existed — and attaches proof to both moves.
A gap analysis reads what companies publish. That boundary is what makes every claim checkable — and it excludes some things other vendors are happy to guess at.
The deliverable is a flat, import-ready file plus a written read of the results. Most teams load the three buckets as list views and work them on different cadences.
A proof project from €4,900 audits one thesis-shaped slice of your CRM. The full universe with deep shortlist runs from €9,900; annual monitoring from €18,000 per thesis keeps it reconciled.
bucket: missing_from_crmmandate_fit: 92.7 · outreach: 88 · transition: 71founder_status: founder led — "the founder established the company in 1981"certifications: UL508 — captured as exact claim textevidence: 12/12 snippets verified against site textA CSV export with company names and websites, plus your thesis in plain language — no CRM access or credentials needed. Domains are the matching key; where your export lacks them, we resolve names to domains during normalization and flag ambiguous cases rather than guessing.
An export adds rows from the same kind of source you already have — companies someone previously found and tagged — and cannot tell you what all such sources missed. A gap run starts from the entire active web, reads sites against your thesis, and reconciles to your records: an audit with evidence, not a bigger pile.
It depends on your niche's keyword density, and we won't pretend to know before running it — but the structural finding is consistent: a fifth or more of confirmed fits lack their category's obvious homepage keywords, so keyword-fed lists systematically under-cover. The proof project exists precisely to answer this on one slice before you commit to the full universe.
No, and we'd encourage suspicion of anyone who says yes — no website signal reveals intent, and we don't manufacture one. What we do document: founder-associated, long-established, independently positioned businesses with identifiable decision-makers, stated on their own pages and quoted verbatim, so your team draws conclusions from evidence.
Nothing is deleted — every disqualified record ships with its documented reason and quoted source. Most teams archive them with the reason attached, which also stops the same companies re-entering the pipeline through next year's conference scan.
Every deep-extraction read happens at run time against the live site — no aged profile layer between you and the evidence. Annual monitoring (from €18,000 per thesis) re-screens the universe and delivers monthly deltas: new fits, ownership changes, and disqualifications reconciled to your list each cycle.
Send us one thesis and one export. We'll return the three buckets — confirmed, missing, disqualified — with quoted evidence for every call, starting from €4,900.