Home Use cases Custom Taxonomy Classification
Use case

Your categories. The whole web. Classified.

Standard industry codes answer standard questions.

When a thesis lives or dies on a distinction no code set contains — field-service-led versus depot-led, OEM-authorized versus independent aftermarket — we take your scheme, sharpen it into screenable definitions, and classify the universe into it, with quoted evidence behind every assignment.

700+
categories in production
100M+
classified domains
100%
assignments evidenced

The distinctions your thesis needs don't exist in any code set

Standard codes flatten the splits that define a thesis — field-service versus depot, OEM-authorized versus independent, certified versus uncertified. Every one of those distinctions is visible on company websites; none exists in any standard schema.

Invisible to codes

Industry codes cannot separate a distributor with a service arm from a service firm that resells parts, or a UL 508A panel builder from an electrician with a website. Your thesis lives on those splits.

Visible on websites

Route density, authorization status, certification tiers, revenue texture — companies publish these distinctions to win customers. We read them at census scale and classify every company into your scheme.

Three workarounds that fail

Force-fitting codes blurs the thesis. Hand-coding drifts between analysts. DIY with an LLM discovers the real work: definitions, edge cases, quality control, and re-applying refinements consistently.

Our production system

700+ categories maintained across 100M+ domains — the same machinery, loaded with your scheme instead of ours. Your categories, your definitions, applied uniformly with evidence on every assignment.

From category idea to screenable definition

The difference between a taxonomy that works and one that embarrasses everyone is entirely in the definitions. We spend the front of every engagement making yours precise enough to screen with — and honest enough to admit what a website cannot show.

1

Definition workshop

Each category gets a written definition with inclusion rules, disqualifiers, and the visible markers that evidence membership — the language, pages, and claims a site would actually carry. Criteria that are not website-visible get flagged now, not discovered after the run.

2

Stress test on live sites

Before anything runs at scale, we classify a few hundred real candidate sites against the draft scheme and review the results together.

Ambiguous cases become explicit edge rules; categories that will not separate get merged or redefined. The scheme that goes to production has already survived contact with reality.

3

Full-scale run

The finished scheme applies across the whole target universe through the same two-pass pipeline as our standard categories — every assignment carrying its quote, page reference, and confidence grade.

How classification runs at web scale

Two passes turn a raw domain pool into evidenced category assignments. Your taxonomy's decision rules run against full site extractions, not homepage snippets.

Pass 1 — Triage

Separates live operating companies from directories, parked domains, and dead sites. In one industrial category, 367K domains narrowed to ~17,300 genuine candidates.

Pass 2 — Deep extraction

Reads services pages, about pages, case studies, certifications, team pages. Your taxonomy's rules run against the full extraction — every assignment records verbatim evidence and source page.

Scoring & ranking

Category membership and acquisition scores come from one reading: Mandate Fit 70%, Outreach Suitability 20%, Transition Context 10%. Group ownership evidence zeroes the total.

Evidence discipline

Sites without enough evidence are graded insufficient-evidence — reported honestly alongside fits and exclusions. An auditable "cannot tell" beats a coin-flip assignment.

Living scheme

Because the scheme is code plus definitions, it stays alive. Re-runs with refined ICPs are included; definition changes re-apply across everything already classified.

Dimensions clients actually commission

Most custom schemes are not exotic — they are ordinary strategic distinctions that standard data simply does not carry.

Service delivery model

Field-service-led versus depot-led versus resident-technician programs. The economics differ completely — route density, response-time commitments, installed-base intimacy — and websites state the model plainly to their customers even though no database field captures it.

Channel authorization status

OEM-authorized versus independent aftermarket, quoted from authorization claims and partner pages. For consolidation theses in compressed air, material handling, or power transmission, this single split often defines the buy-box — and the pricing logic.

Certification-defined tiers

Categories bounded by compliance claims: ISO/IEC 17025-accredited labs versus uncertified calibration shops, ASME-stamped fabricators, NADCAP finishing houses, UL 508A panel builders. Exact claim text, captured per site, becomes a hard category edge.

Revenue texture

Contract-and-program operators versus project-and-job operators, read from the language of maintenance agreements, chemical service programs, rental fleets, and scheduled offerings. The distinction most industrial buyers ask about first, applied as a classification rather than a hunch.

End-market exposure

Who the customer base actually is, evidenced from case studies and named clients rather than inferred from keywords — municipal versus industrial water treatment being the classic example where the same equipment serves two entirely different theses.

Ecosystem role

Manufacturer, distributor, integrator, service provider, or hybrid — assigned from what the site describes doing, not from a self-selected directory label. The most common source of misfiled companies in standard data, and usually the first dimension a scheme needs.

The signal framework underneath the categories

Custom categories are not extracted from thin air — they compose from the same 15-signal extraction that runs on every deep-analyzed site. Five signals carry most custom schemes.

Strategic fit to thesis

In a custom engagement this signal is literally your taxonomy — the classifier is built from your words, stress-tested with you, and every assignment traces back to a definition you signed off. Ambiguities surface as documented edge cases, not silent inconsistency.

Service-led vs product-led model

The workhorse of industrial taxonomies: whether revenue arrives through field service, manufacturing, distribution, or a hybrid is stated on nearly every industrial site. The same "pump company" keyword covers a manufacturer, a distributor with a repair bench, and a field-service operation — no thesis treats those three alike.

Partner & channel ecosystem position

Authorization-based categories classify cleanly because companies advertise their channel positions — "authorized distributor," "certified service center," named OEM partnerships. We quote claims verbatim; for aftermarket and channel theses, this signal usually is the taxonomy.

Compliance & regulated-market readiness

Certification claims make the sharpest category edges because they are binary, dated, and checkable: a lab either publishes ISO/IEC 17025 accreditation or it does not. Our capture of exact claim text keeps the edge auditable years later.

Vertical specialization & documented end-market exposure

End-market categories fail when guessed from keywords, so we build them from documented exposure — case studies, project galleries, named customer industries. The confidence grade is part of the category assignment, not decoration.

Worked example: ten subverticals no code set contains

The specimen work we publish is itself a custom taxonomy.

One buyer-shaped scheme split US industrial services into ten subverticals — precision machining, equipment repair, automation integration, material handling, compressed air, calibration and testing, water treatment, boiler and steam, filtration, surface finishing — none of which maps one-to-one onto any standard code.

Applied through the two-pass pipeline, the scheme produced eligible independent US counts per category: 702 in precision machining, 545 in equipment repair, 534 in automation integration, down to 93 in surface finishing.

Real denominators, per custom category, each company carrying its evidence.

The interesting rows are the ones a keyword approach gets wrong. A fifth or more of confirmed fits in these runs lacked their category's obvious homepage keywords — the specimen file's "hidden fit" class.

One equipment-repair fit's homepage never says the category's name; its About page evidence reads "Founded in 1991 by the founder and the founder", and its services live three clicks deep.

A calibration-testing hidden fit is a materials lab whose Careers page. not its homepage.

carries the line "we are a family-owned business." Keyword tagging files these companies under the wrong label or misses them entirely;. definition-driven classification, reading full sites, puts them where they belong.

The exclusions are equally instructive: about one in ten keyword-perfect candidates classified out as group-owned, evidenced in the companies' own words — "privately held by a national distribution group", as one excluded calibration business states.

A taxonomy that cannot document its exclusions is a list with opinions; the specimen format exists to show both directions of the discipline.

What we refuse to sell: no “ready to sell” flags, no revenue or EBITDA guesses, no owner-age profiling, no distress detection — and no engagements in consumer-captive verticals. Read our standards; serious buyers tell us this page is why they trusted the rest.

Four ways to get categories, compared

The fair comparison is not "custom taxonomy versus nothing" — it is against the three approaches teams actually use today.

ApproachDistinction fidelityScaleEvidenceMaintainability
Standard industry codes Coarse; thesis-critical splits usually absent Universal None — assignments are self-reported or inferred Static by design
Keyword tagging Brittle; misses the fifth of fits without obvious keywords, misfiles hybrids High The keyword itself, which proves little Every refinement is a new regex debate
Manual analyst coding High on a good day; drifts between analysts and across months Hundreds of companies, then fatigue Whatever the analyst noted Re-coding after a definition change rarely happens
Full-web LLM taxonomy Your definitions, applied uniformly; ambiguity surfaced as edge cases Thousands to full-web Quote + source page + confidence, per assignment Definition changes re-apply across the classified universe

Where custom taxonomies fail — and where ours stop

Some category ideas cannot be built honestly, and the useful moment to hear that is before the run.

Categories defined by non-public facts — revenue bands, precise headcount, margin profile, backlog — are not classifiable from websites, and we will not pretend otherwise by proxy-guessing; that refusal is the same one our standards apply to every deliverable.

Categories that depend on intent or interior state ("companies open to partnership") are stories, not classifications.

And every scheme has an ambiguity floor: some real companies genuinely straddle two categories, and a hybrid distributor-integrator does not stop being hybrid because a schema would prefer it picked a side.

We handle those with explicit multi-membership or documented tie-break rules — but a scheme whose categories overlap heavily will produce arguments, not insight, and we say so in the stress-test phase.

Two more boundaries.

Coverage is web-visible coverage: a company with no meaningful web presence is outside the census, which in B2B is a thin sliver but not zero, and we state the boundary rather than extrapolating across it.

And confidence grades are real: an insufficient-evidence assignment means the site did not publish enough to classify, and buying a taxonomy from us includes accepting that class in the output.

typically a low single-digit percentage of deep-analyzed sites, honestly labeled instead of forced into a bucket.

What you own, and how it plugs in

The scheme and its classified output are engagement deliverables, confidential to you — including the definitions themselves, which by the end of the stress test usually encode real strategic thinking.

Output arrives as structured files keyed by domain: category assignments, confidence grades, evidence quotes, source pages, plus the standard signal extractions underneath.

Teams load it into CRMs as custom fields, into BI tools as the segmentation layer for market sizing, or into sourcing workflows as the buy-box filter that finally matches the thesis language in the IC memo.

A taxonomy also compounds.

The same scheme that sizes a market this quarter can drive an ICP discovery run for the commercial team next quarter and become the category backbone of ongoing monitoring after that — new companies classified on arrival, category migrations flagged as deltas.

Engagements start at proof-project scale (from €4,900, one scheme on one subvertical) so the definitions can prove themselves on a bounded universe before you commit them to a full category;.

details on pricing, and a same-day specimen report shows the evidence format before anything is signed.

Common questions

The pipeline runs schemes from three categories to several hundred — our own production taxonomy holds 700+. The practical constraint is never count; it is definition quality. Ten sharply defined, website-evidenceable categories outperform fifty aspirational ones every time, and the stress-test phase exists to find out which of your fifty are actually ten. We would rather merge categories before the run than deliver confident-looking noise after it.

Yes — the scheme applies to any universe: the full classified web, one of our category censuses, or a list you export from your CRM or data provider. Classifying an existing list is often the cheapest first step, and it doubles as a data-quality audit: expect some rows to come back as group-owned, misfiled, or dead, the same way our gap-analysis runs do.

A grade reflecting how directly the site's language supports the category decision. A quoted authorization claim on a partner page is high confidence; membership inferred from a case-study pattern is medium and labeled as such; and where evidence is genuinely thin the assignment is insufficient-evidence rather than a guess. Confidence travels with every row, so your team can set its own bar per use — strict for outreach, looser for market sizing.

You do. Definitions, edge-case rules, and classified output are deliverables under the engagement, confidential to you, and we do not reuse a client's scheme for anyone else. If you later want the scheme maintained — new companies classified as they appear, definition updates re-applied — that runs as a monitoring arrangement, but nothing obliges you to it; the files are yours either way.

Definitions, edge cases, quality control, and scale — the four places ad-hoc tagging quietly fails. A one-shot prompt applies whatever the model decides each category means that day, to whatever text it happens to fetch, with no evidence trail and no way to re-apply a refinement consistently. Our runs apply written, stress-tested definitions through a production pipeline with per-assignment quotes and confidence grades, across full sites rather than homepages. The difference shows up precisely on the ambiguous rows — which are the rows that matter.

Bring us the categories only your team can define

One email starts the definition workshop. The specimen report shows the evidence format the same day — including how we classify the companies that refuse to fit neatly.

Request the specimen report