Every candidate website is read against your written thesis — first in a fast triage pass, then in a deep extraction pass. Every claim in the deliverable traces back to a sentence a company published about itself.
We operate several AI platforms and maintain large-scale specialized datasets. More than 300 enterprise organisations run on our data — among them one of Europe's largest telecom operators, adtech and cybersecurity platforms, and a leading airline metasearch.
That classification of 100M+ active domains is where every engagement starts. Your thesis maps to categories; we pull every domain in scope. Nobody pre-filtered the list, so nothing was silently missed before you arrived.
Starting from every classified domain in a category — not from a database of companies somebody previously found, tagged, and decided was worth keeping.
Every drop between rows is logged with a reason — nothing exits silently.
A fast pass over everything, a deep pass over survivors. That split is what makes reading tens of thousands of websites economically sane — without ever sampling or skipping.
Your thesis maps to categories in the 100M+ domain classification. Every domain in scope is pulled — typically tens of thousands.
An LLM reads every homepage: operating company or directory? Independent or branch? Which subvertical, really? Dead domains fall away, with reasons.
Survivors get a full crawl — about, team, history, services, certifications, careers — and structured extraction of all 15 signals.
Three weighted composite scores rank the universe against your thesis. Group-owned companies zero out of outreach, whatever their fit.
Evidence is re-verified against site text, exclusions are audited, and the package ships: ranked CSV, report, exclusion log.
These are the real numbers from the published industrial-services specimen — the same run the sample report is drawn from.
Every signal is extracted only from what a company publishes about itself. Where the site is silent, we record "not visible" — never an estimate, never an inference from a name or a photo.
Reads: explicit founder or family language — "second generation", "founder-led", "family-owned" — captured as the exact phrase on the page.
Why it matters: roughly half of confirmed industrial fits carry this evidence, and outreach to an identifiable owner is a different conversation entirely.
Reads: how many principals are actually named on the site, and whether the visible bench extends beyond one or two people.
Why it matters: a shallow published bench shapes transition context; a deep one signals professionalized management worth meeting.
Reads: stated founding year, years in operation, and independence language, taken from the company's own telling of its history.
Why it matters: long-established independents are the population most theses actually target — and the hardest to find in profile databases.
Reads: subsector, services, and customer types, matched line by line against the buyer's written thesis.
Why it matters: this is the principal signal — it drives the Mandate Fit score that carries 70% of the ranking weight.
Reads: headquarters, branch locations, and the stated service area — with owned locations distinguished from partner mentions.
Why it matters: footprint decides whether a company is a platform, an add-on, or out of territory before anyone books a call.
Reads: whether revenue visibly comes from field service, manufacturing, distribution, or a hybrid of them.
Why it matters: two companies with identical keywords can have opposite business models — this signal is how a thesis avoids buying the wrong one.
Reads: service contracts, maintenance agreements, scheduled programs, consumables — the published language of repeat revenue.
Why it matters: recurring language on the site is the closest website-visible proxy for revenue quality a screen can honestly offer.
Reads: customer industries evidenced by case studies and named end markets — not guessed from homepage keywords.
Why it matters: documented aerospace, food, or utility exposure is what separates a niche specialist from a generalist with a good copywriter.
Reads: signs the company already belongs to a group or runs its own acquisition program — press pages, "part of the X family" lines, investor tabs.
Why it matters: this is usually a disqualifier. About one in ten keyword-perfect candidates fails here and becomes a documented exclusion.
Reads: visible non-founder functions — finance, operations, HR, marketing roles named on the site.
Why it matters: named second-layer management tells you how much of the company walks out the door with the founder.
Reads: open roles and their types, from careers pages and posted listings on the site itself.
Why it matters: what a company hires for is a candid window into growth, capacity constraints, and operational maturity.
Reads: the most recent dated content anywhere on the site — news posts, project updates, certifications with issue years.
Why it matters: visibly active and visibly dormant companies deserve different outreach priority, and the difference is checkable.
Reads: named OEM partnerships, distributorships, buying groups, and association memberships published on the site.
Why it matters: channel position often is the moat in industrial services — and it transfers, or doesn't, in an acquisition.
Reads: explicit certifications only — ISO, AS9100, ITAR, ISO/IEC 17025, ASME — captured as the exact claim text.
Why it matters: certifications gate entire end markets. A claim recorded verbatim can be checked; a checkbox in a database cannot.
Reads: online quoting, customer portals, e-commerce, pricing transparency — how the company sells through its own site.
Why it matters: digital maturity is a cheap, honest indicator of how investable the operation is beyond its machines and vans.
The fifteen common signals are the floor, not the ceiling. Each engagement adds signals that only make sense in your niche — the accreditations, stamps, and authorizations that decide who is real in that trade.
Each added signal follows the same discipline as the core fifteen: exact claim text, source URL, and a "not visible" record when the site doesn't say. The signal list is agreed with you before extraction starts — it is your thesis, operationalized.
Fifteen signals collapse into three composite scores. The weighting is intentionally unbalanced: fit does the work, suitability gates the list, and context is capped so it can never flatter a weak match.
How closely the company matches the written thesis: subsector, service model, recurring offering, footprint, compliance posture. The score that does the heavy lifting, and the one your ICP re-runs re-weight.
Is there an identifiable decision-maker and an independent owner to talk to? Founder or family association, visible principals, and no existing group ownership feed this score.
Long operating history, limited visible bench, self-described generational ownership — website-visible context only, and deliberately capped at a tenth of the total.
The zero-out rule: a company already owned by a group or consolidator scores zero on outreach, whatever its fit. Roughly one in ten otherwise-perfect candidates fails exactly here — and each becomes a documented exclusion, not a silent omission.
Deliverables don't say "founder-led: yes". They quote the line, name the page, and link the URL — so your associate can check our work in one click. Below, real entries from the published specimen, anonymized for this page.
If a signal is not on the website, the field says "not visible". It is never estimated, never borrowed from a third-party profile, and never inferred from a name, a photo, or a hunch.
LLM extraction at scale is powerful and fallible in equal measure. The QC pass exists because we assume errors, then go looking for them.
Every quoted snippet is matched programmatically against the crawled page text. The verification count ships in the report — "16/16 verified", or honestly, "11/13". A snippet that can't be found is removed, and the signal reverts to "not visible".
A human analyst reads the top of the ranked list and the complete exclusion log. Misclassified subverticals, mistaken independence calls, and over-generous fits get corrected before anything ships.
Every company flagged as group-owned is spot-checked against its own site, because a false ownership call costs you a real candidate. The one-in-ten exclusion rate is measured, not assumed.
Ranked CSV, specimen-style report — 8 top fits, 5 keyword-missed fits, 5 documented exclusions, 2 insufficient-evidence cases — plus the full evidence and exclusion logs, delivered CRM-ready.
A method you can trust has edges you can see. These are ours, stated plainly rather than discovered mid-engagement.
No financials, no revenue estimates, no valuation guesses. A website shows what an owner chose to say — the method refuses to pretend otherwise, which is why the standards page exists.
We do not claim to know whether any owner would entertain a conversation. No website signal supports that claim, and vendors who sell it are selling a guess.
A two-page website yields two pages of evidence. Those companies are flagged "insufficient evidence" — two per specimen report — rather than scored on imagination.
Sites change after we read them. A screen is a snapshot; the monitoring tier exists precisely because universes go stale. Re-screens catch what single passes cannot.
The specimen report answers most of these with worked examples — these are the short versions.
One email. We send the specimen report the same day — every company scored and ranked with full signal transcripts, including the ones we refused to include and why.
Request the specimen report