Blog / Method / No.010

Reading a company like a salesperson, not a keyword

The word scraping appears on the site of our best customer, our nearest competitor and an agency that has never written a line of code. So we read something else.

Mohamad · KnockwireJune 30, 202610 min readMethod

Keyword matching is the first thing anyone builds and the first thing that has to go. It cannot tell a buyer from a reseller, and it certainly cannot tell either of them from a competitor, because the word on the page is identical in all three cases. Prospect enrichment only becomes useful at the point where it stops counting words and starts reading the company the way a decent salesperson would: what do you sell, who buys it, what does it cost you to make, and only then, what are you built on.

01One word, three companies that share nothing

Take a page that says the company works with web data at scale. Three very different businesses publish that sentence, and a keyword filter scores all three the same.

  • A vendor. They sell collection itself. They are a competitor, and a message from us is an embarrassment rather than an opportunity.
  • A reseller or agency. They pass someone else's data through to their own clients and add a report on top. The collection bill is not theirs, so a better collection price changes nothing they can feel.
  • A buyer. Data is the product they sell, collection is their cost of goods, and the bill grows every time a target site tightens up. This is the only one of the three worth a message.

Nothing about the phrase separates them. Neither does the technology stack, because all three might run the same tooling. The separation only appears once you read what the company sells and to whom, which is what a person does in the first thirty seconds on a website and what a keyword filter never does at all.

The word is not a signal. It is a coincidence that happens to rank well.

02Prospect enrichment starts with what the company sells

The first read is the product itself, taken from the company's own site: the homepage claim, the product pages, the pricing page if there is one, and the phrasing of what a customer actually receives. We are looking for one distinction. Is data the thing being sold, or is it something the business happens to touch on the way to selling something else.

This distinction is what everything downstream rests on. A retail intelligence company sells a data product and prices it by coverage, so more coverage costs them more money and they know it to the decimal. A logistics company that scrapes two carrier sites to enrich its own tracking page has a data activity, not a data product, and no amount of collection improvement moves a number anyone in that building watches.

This is also the reason enrichment reads the actual website rather than trusting a directory listing. A directory tells you what the company said about itself when it registered, which may be three years old and was written to be found rather than to be accurate. Directories are excellent at telling you a company exists and belongs to a category. They are poor at telling you what it sells this quarter, and that is the only version we care about.

One refusal belongs here. We do not read the personal profiles of the people who work at the company. Enrichment reads the business, its public site, its public repositories and its public postings, and it builds no picture of any individual. A company is a legitimate research subject. The people inside it are not our business until one of them chooses to reply.

03Then who buys it, and what it costs them to make

The second read is the customer, because it settles the reseller question that the first read leaves open. A company selling a data feed to hedge funds owns its collection. A company selling dashboards to brand managers may be buying that same feed wholesale and never touching a collector.

The tells are ordinary once you look for them. Case studies name the sort of customer. Pricing pages priced by seat point at software, priced by record or by request point at data. Documentation aimed at engineers means the buyer builds; documentation aimed at analysts means the buyer consumes. A partner or integration page listing a data provider is often the clearest signal of all, and it is usually the moment a promising candidate turns out to be a reseller.

The third read is cost of goods, which is the one that actually predicts a reply. We are asking whether collection sits on the expensive side of their books. Coverage counts stated in the millions, freshness promised in hours, a status page that reports collection health, a changelog that announces new sources by name: all of these mean somebody is paying for volume and watching it.

Where that answer is no, the candidate is turned down even when the first two reads were perfect. A company that sells data but sources it from an API that costs the same every month has no problem we can improve. They are a lovely business and there is nothing to say to them.

04Then whether their targets block casual traffic

The fourth read is the one most systems skip entirely, and it is the one that best predicts whether the message lands. We look at what the company collects from, and then at whether those sources are hostile to casual traffic.

Our own reference set makes the case. The top accounts by lifetime revenue at Crawlbase cluster on three target families: marketplaces, travel, and people or company records. Every one of those families blocks casual traffic hard and gets harder every year. That is not a coincidence about those customers, it is the reason they are those customers.

So a company collecting from sources that serve plain pages to anyone who asks is turned down here, and it is a genuine loss when the rest of the profile is strong. They are a real data business with a real cost of goods, and the cost simply is not the part we could change. Writing to them would mean opening with a problem they do not have, which is the fastest way to be filed as noise.

This read is also the one that most often rescues a candidate. A modest company collecting from three brutal marketplaces is a better prospect than a large company collecting from a public register, and no measure of company size will ever tell you that.

05Only now, the technology signals

Technology signals go last, deliberately, because they are confirmatory rather than diagnostic. A competitor SDK in a public repository, a job advert naming a proxy vendor, a headless browser dependency in a public manifest, a status page apologising about collection: each of these sharpens a picture. None of them draws one.

Read first, a competitor SDK looks like the strongest signal available. Read last, it is simply a fact about a company we have already classified. If the four earlier reads say reseller, a competitor SDK means their supplier picked it. If they say vendor, it means a competitor is using another competitor. The same evidence points in three directions depending on what you already know, which is exactly why it cannot come first.

The cost of doing it this way is that enrichment is slow and expensive compared with a keyword pass. It reads pages, it follows a handful of links, and it does that politely and single threaded rather than hammering anyone. We have accepted that cost because the alternative was measurably wrong.

We kept a cheap keyword filter running alongside enrichment for a quarter to grade one against the other. The keyword filter passed 1,204 candidates. Enrichment agreed with 411 of them and overturned 793, which is a shade under two thirds of its verdicts.

  • 341 sold collection themselves. Competitors and near competitors, every one of them a message we are glad we did not send.
  • 224 mentioned collection in a job advert or a blog post, but the product they sell is something else and the collection is incidental.
  • 152 resold data they did not gather, so the collection bill belongs to their supplier and not to them.
  • 76 were consultancies or agencies billing time rather than volume, where a lower unit cost is not a number anyone in the business tracks.

The errors ran the other way too, and that half is easier to miss. We took a random 600 of the candidates the keyword filter had failed and enriched them anyway. Sixty-three of them should have passed, mostly companies whose site described a product whose cost of goods was obviously collection without the vocabulary ever appearing. Extrapolated across the quarter that is roughly 1,180 companies, and it is most of the reason Gate 1 passed 1,592 rather than 411.

Two thirds wrong in one direction and about a tenth wrong in the other is not a filter that needs tuning. It is a filter reading the wrong thing. Reading order is not a detail of the implementation, it is the method.

ScopeEnrichment reads a company's public website, its public repositories and its public job postings, at a polite single-threaded pace. It reads no personal profiles, buys no contact database, and never treats a directory listing as a substitute for the site itself.

Knockwire reads the internet, throws out the companies that will never buy from you, and knocks on the doors of the ones that will. Run it on your own domain and read your own refusals.

Run it on my siteAll posts