Home Guides How to Screen Out False Positives in Target Lists
Guide · Screening discipline

How to Screen Out False Positives in Target Lists

Five species of false positive survive keyword screening, and each is caught by a different test. A 14-minute guide to the taxonomy, the two-pass architecture that runs the tests in the right order, and the logging that turns exclusions into infrastructure.

~1/3
of keyword-plausible domains fail basic triage
1 in 10
keyword-perfect candidates already group-owned
5
documented exclusions in every specimen report

The arithmetic of a polluted list

Every target list carries a contamination rate, and most teams never measure theirs. In one specimen triage, nearly a third of 25,000 keyword-plausible domains failed basic liveness tests before any thesis logic ran.

Hidden waste

BD hours on trade associations, letters to branch managers without authority to sell, outreach cadence burned on businesses that ceased operating years ago but kept paying for hosting.

Corrupted feedback

When a third of your list cannot respond meaningfully, response-rate data stops telling you anything about your thesis. You cannot tune a process whose denominator is fiction.

Root cause

Contamination is the natural output of keyword-driven list building. Keywords select for vocabulary, and vocabulary is shared by operators, directories, publications, resellers, and the acquired alike.

What this guide covers

A working taxonomy of false positives, the evidence tests that catch each species, and the logging discipline that turns exclusions into infrastructure. Every test works manually or at census scale.

A taxonomy: the five species of false positive

Naming the species matters because each is caught by a different test, and a screen that only runs one test passes the other four.

  • Species one — non-companies. Directories, trade associations, industry publications, job boards, and lead-generation sites wearing the vertical's keywords more fluently than the operators do (fluent vocabulary is their business model). Loud in every keyword export; trivially caught by anyone, or anything, that actually reads the page and asks: does this entity perform the work, or aggregate those who do?
  • Species two — wrong-side businesses. The residential contractor in a commercial-services thesis, the consumer brand in a B2B screen, the hobbyist supplier in an industrial one. The keywords match perfectly; the customers do not. The test is evidence of who is served: named commercial clients, industrial case studies, B2B-shaped service pages — versus booking widgets and homeowner testimonials.
  • Species three — wrong business models. Distributors in a manufacturing thesis, brokers in an asset-based thesis, resellers in a services thesis, franchisors in an independent-operator thesis. The vertical's vocabulary is fully shared; the economics are not. The test is how revenue visibly happens: shop-floor and field-service language, fleet and facility evidence, versus catalog structures, dealer locators, and “authorized reseller” badges.
  • Species four — already-owned companies. Subsidiaries, platform branches, and franchise locations running legacy websites that still read independent. The costliest species per instance — courtship of someone else's portfolio — and reliably one in ten keyword-perfect candidates in our industrial runs. The test is the ownership read: footers, about-page acquisition narratives, news announcements, parent-branded careers portals.
  • Species five — ghosts. Companies that stopped operating while their websites did not. Hosting is cheap, forgetting is easy, and a five-year-dead site can look pristine. The test is activity evidence: most recent dated content, current-year copyright, open roles, recent case studies or news — any artifact that requires a living organization to produce.

Why keyword screens breed all five

The root cause is a category error: treating vocabulary as evidence of a business model. A keyword filter selects for a market's conversation, not its participants — and the same error runs in reverse, producing false negatives.

False positives

Directories, magazines, distributors, and acquirers all emit the vertical's keywords more fluently than operators — because publishing vocabulary is their business model.

False negatives

Operators who describe their work idiosyncratically vanish from keyword universes. In our runs, a fifth or more of confirmed fits lacked the category's obvious homepage vocabulary.

The fix: two passes

Pass one (triage) runs cheap species tests at population scale. Pass two (deep reading) runs expensive judgments on survivors only — business model, independence, thesis fit, with evidence quoted.

Architecture, not increments: never spend a deep read on a domain that a thirty-second test would have removed. This two-pass structure handles six-figure domain populations and works equally well for a manual screen of four hundred.

The independence read, in depth

Species four is the most expensive to miss and the most subtle to catch. Group ownership hides in plain sight, and an aggressive keyword rule creates its own false positives.

Where ownership hides

Footer (“a division of…”), about page (“proudly joined…”), news page (dated acquisition announcements), careers page (applications routing to a parent's portal).

Real specimen exclusions

“wholly owned subsidiary of a global industrial group” · “owned by a publicly listed industrial group” · “acquired by a global industrial group”

Family purchase

“Acquired by the Harrison family in 2001” is not consolidation — it is an independent business most theses want.

ESOPs

Neither family business nor subsidiary — its own structure with its own transaction dynamics, deserving its own accurate label.

“Partnered with”

May be a marketing flourish or a PE announcement. Only the sentence around the keyword decides.

Our rule: no company is excluded for group ownership without a quotable disclosure, captured verbatim with its source URL. Related signals — footprint branches, shared identity, centralized careers portals — sharpen the read. Independence is a conclusion assembled from evidence, not a checkbox.

Exclusion is information — log it like inventory

The instinct on finding a false positive is deletion. The better practice is documentation: species, disqualifying evidence quoted verbatim, source URL, date.

Memory

Undocumented deletions reappear in next quarter's rebuilt list, get re-researched, and re-deleted. We have watched teams pay for the same discovery three times.

Calibration

Exclusion patterns are rubric feedback. A cluster of wrong-model exclusions tells you your inclusion criteria are underspecified upstream.

Credibility

“Here is what we excluded and the sentence that disqualified each one” is the most trust-building page in any deliverable — it proves inclusions were tested against something.

Our specimen format: every vertical publishes five documented exclusions alongside fits and flags. A list that shows only winners is unauditable. Exclusion reasons must not overreach their evidence — the moment reasons outrun quotes, the log becomes a record of moods.

Insufficient evidence is a category, not a failure

Some sites do not support a confident call in either direction — six-page brochure sites with no team page, no dated content, and services described in generalities. A screen has three options.

Option 1: Guess in

Pollutes the pipeline with unverifiable rows.

Option 2: Silently drop

Quietly lies about coverage. A vendor with no unknowns has laundered uncertainty into confident rows.

Option 3: Flag honestly

Label it insufficient evidence with thin signals found, route to the instrument that resolves it: a phone call.

Workflow value

The insufficient-evidence tranche becomes a calling list for a junior team member. Ten-minute calls resolve most cheaply, and logged resolutions improve the rubric's sense of which thin-site patterns hide real companies. Our specimen format carries two such flags per twenty companies, roughly matching field rates.

A QC regime you can actually run

Screening quality is a process property, not a heroic one-time effort. The process is small enough to write on one page.

1

Sample-audit every batch

Pull a random 5% of inclusions — verify evidence quotes against live pages. Pull 5% of exclusions — check the disqualifying sentence still exists. Track both failure rates over time; treat movement as a rubric bug.

2

False-negative check quarterly

Take confirmed-good companies learned from other channels — conferences, banker teasers, portfolio CEO mentions — and test whether your screen would have found them. Every miss is a vocabulary or rubric gap with a specific fix.

3

Close the loop from outreach

Every “we were acquired last year” reply and every bounced ghost is a free audit result. Log it against the screen that passed it. Our pipeline publishes per-file verification counts because invisible QC is a QC anecdote.

Teams that run this regime converge on contamination rates low enough that response data becomes signal again. The productized version ships in every engagement from buy-side long lists to full thesis universe mapping; the manual version costs a rubric, a log, and the willingness to quote your own evidence.

A contamination ledger, end to end

Walk one specimen vertical through the full screen to fix magnitudes. The screen did not shrink the opportunity — it revealed which fraction was ever real.

25,000
raw keyword domains
17,300
after triage (S1+S5)
~90%
after model + side reads
-10%
already group-owned
702
eligible independents

Machining subvertical shown. Other universes: 545 repair houses, 254 calibration labs, 93 surface-finishing operators.

The ledger as market intelligence
Species 4 heavy

Late consolidation cycle. The ledger names the acquirers and dates the acquisitions.

Species 5 heavy

Aging or migrating vertical. Operators leaving faster than new ones forming.

Species 3 heavy

Vocabulary contested between business models — predicts where database classifications fail.

Frequently asked questions

Layered, and larger than intuition suggests. In our specimen industrial triage, nearly a third of keyword-plausible domains failed basic liveness-and-identity tests (species one and five). Of the survivors that looked thesis-perfect, roughly one in ten carried a group-ownership disclosure on inspection (species four), before counting wrong-side and wrong-model businesses, which vary most by vertical. Compounded, it is common for well under half of a raw keyword export to survive a full evidence screen — which is also why response-rate math on unscreened lists misleads.

Especially there, because small screens get rebuilt most often. The log costs one spreadsheet row per exclusion — name, species, quoted sentence, URL, date — perhaps ninety seconds each. It repays itself the first time the list is refreshed (no re-research of known rejects), the first time a stakeholder asks how the list was built, and the first time a pattern in the log reveals a rubric gap. Deletion feels efficient in the moment and costs triple later; the log is the cheapest infrastructure in the entire sourcing stack.

The company tells you, in predictable places: footer, about page, news page, careers portal. Read those four locations for group language — “division of”, “family of companies”, acquisition narratives, parent-routed applications — and record what you find verbatim. Registries and databases help as cross-checks where they exist, but the operative disclosures at small companies usually live only on their own sites, and often only as a single sentence. The rule that keeps the screen honest: no exclusion without a quotable disclosure; no independence label when the site is silent — that is “unverified”, a third state worth tracking.

It can, if the architecture forces evidence. The failure mode of casual LLM screening is impression-based classification — plausible summaries unmoored from the page. The safeguard is structural: require a verbatim quote and source URL for every claim, verify the quotes against the site text mechanically, and fail any classification that cannot cite its sentence. Our pipeline publishes those verification counts per file. Under that discipline, the model's advantage — reading every page of every site, at census scale, without fatigue — dominates its risks, and the audit trail makes the output checkable by anyone.

Borderlines are rubric information, not annoyances. The distributor that also runs a service division, the commercial contractor with a residential arm, the integrator that resells hardware — each forces a threshold question your thesis has not answered yet: how much of the wrong model disqualifies? Decide once, write it into the rubric with a proxy (“service content on two or more of the top-level pages”), and log borderline calls with their evidence either way. Over a quarter, the borderline log becomes the sharpest available description of what your thesis actually means at its edges.

Keep reading

Screening false positives at census scaleBuy-side long listsFounder-led website signalsDatabases vs full-web mappingSearch fund list constructionOur standards
What we refuse to sell: no “ready to sell” flags, no revenue or EBITDA guesses, no owner-age profiling, no distress detection — and no engagements in consumer-captive verticals. Read our standards; serious buyers tell us this page is why they trusted the rest.

See what the full universe looks like for your thesis

One email. We send the specimen report the same day — every company scored and ranked with full signal transcripts, plus the exclusions we documented and why.

Request the specimen report