Crawlbase is client one, so the first serious run of Knockwire was aimed at our own business. We wrote an ideal customer profile from what we believed about our best accounts, pointed the system at it, and then sat down and read the first hundred refusals in order. The profile was wrong in two places. Both errors were visible in the companies we turned down and invisible in the companies that qualified.
01Read the refusals, not the wins
The refusal ledger is the part of this product that publishes what was declined and why. Every rejection carries a written reason rather than a code, which means the log can be read as prose instead of counted, and reading it end to end is a different exercise from scrolling the qualified list.
The qualified list agrees with you. It has to. A company qualified because it matched the profile you wrote, so a page of qualified companies is a page of your own assumptions handed back with logos attached. The refusals are the only place where the profile meets something it did not predict. A company that clears every gate but one, with a reason written on the gate that stopped it, is a small argument against the rule that stopped it.
Two of those groups say nothing about us. Dead companies are hygiene, and no usable door is a fact about a website rather than about our profile. The other three are the profile being tested, and two of them were arguing with us in writing.
02First error: we had written enterprises that need data
Our stated target was enterprises that need data. It is the kind of sentence that survives a workshop because nobody can disagree with it, and it was wrong in the way broad sentences usually are. It admitted far too much, and the gates behind it spent the entire run throwing the surplus back out.
Thirty-four of the first hundred refusals carried a version of the same written reason: data is an input here, not the product. Every one of those companies had been waved through the profile check. A logistics firm scraping carrier rates has a real need for data. So does a retailer watching competitor prices, and a bank screening sanctions lists. In all three cases collection is a line item that somebody in finance is trying to reduce, and that fact sets the ceiling on what they will ever spend.
One refusal like that reads as a near miss. Thirty-four of them read as a definition. The evidence gate had worked out the attribute our profile sentence was missing, and had been writing it down in plain English, once per company, for a fortnight before anyone read the log from the top.
Our top twelve accounts by lifetime revenue look nothing like those thirty-four. Every one of the twelve sells a data product. Collection is their cost of goods rather than a side activity, which is why their spend rises when the business goes well instead of falling when the budget is reviewed. They cluster on three target families, marketplace, travel, and people or company data, and every one of those targets blocks casual traffic.
So the target is not an enterprise that needs data. It is a company whose product is data, which is already scraping, and is struggling with it. That version is narrower, sharper, and answerable at read time, because you can look at a company and see what it sells. You cannot look at a company and see how badly it needs something.
03Second error: we counted heads instead of volume
The second error was quieter and cost us more, because it rejected companies rather than merely wasting reads on them. We had put a headcount floor in the profile on the assumption that company size predicts budget.
Size does predict budget, weakly, and in the wrong direction for this market. Several of the strongest accounts on our book are small teams with very heavy volume: single-product companies where a handful of engineers run continuous collection against targets that fight back. Headcount is a proxy for spend, and it is a poor proxy once the product is data, because the volume comes from the crawl target rather than from the org chart. Eight people indexing a marketplace with millions of listings spend more than four hundred people who pull a competitor price list every Monday.
In the log this showed up as a contradiction rather than a pattern. Fourteen companies were turned down on size while carrying strong evidence on everything else, so their written reasons read as arguments against themselves: qualified on target family, qualified on collection evidence, below headcount floor. Once you have a dozen refusals shaped like that, you do not have a headcount floor. You have a bug you have not admitted to yet.
We removed the floor and put two observable questions in its place: does the target family block casual traffic, and does the public surface imply continuous collection rather than a project with an end date. Both are answerable from evidence a source actually returned. Neither asks us to guess at a payroll.
04An ideal customer profile is a hypothesis
The general lesson is about what kind of object a profile is. We had been treating ours as a description of customers we already had, when it was really a hypothesis: written before the evidence, carried forward unexamined, and never once tested against the thing it excluded. A description is checked by looking. A hypothesis is checked by watching what it throws away, and a refusal log is a record of everything a profile threw away, with the reason attached.
That changes what you do with the first hundred of anything. The instinct is to read the qualified companies and decide whether the system works. The useful pass is to read the refusals and decide whether the profile works. In our case the system did exactly what it was told, twice, and what it was told was wrong.
Which brings us to the thing we refuse to do here. We do not silently adjust the profile behind the operator. Nothing in this pipeline auto-tunes the ideal customer profile from approval and rejection behaviour, although the data to do it is sitting right there and the demo would look excellent.
The reason is that a self-adjusting profile makes the refusal ledger unreadable. The ledger is only worth publishing if a reason means the same thing in March as it meant in January. If the rules drift on their own, nobody can tell whether a company was declined because its circumstances changed or because the rule moved underneath it, and every historical refusal turns into a claim about a version of the profile that was never written down.
So the system proposes and the operator decides. When refusals cluster, we surface the cluster: here are thirty-four companies declined for the same written reason, here is the gate that declined them, here is the change to the profile that would have let them through, and here is what that change would have done to the rest of the run. The operator reads it and makes the call. A change to the ideal customer profile is a decision with a date and a name on it, and both of the errors above were fixed that way, in the open, by a person.