In June we ran the pipeline at its usual setting. In July we dropped the fit threshold by fifteen points and left everything else alone, because a filter you never loosen is a filter you cannot price. The question was whether our outbound reply rate is good because the targeting works, or good because we have never once tested the boundary. Four weeks later we had 164 percent more messages delivered and two fewer replies.
01What we changed, and what we deliberately did not
A Read is a candidate that a discovery source hands us. Fit scoring runs after enrichment and before anything is written, on a scale of zero to a hundred. Our production threshold is 72. For four weeks we set it to 57 and told nobody on the review side what the number was.
We publish what we turn down and why. A refusal ledger is easy to be proud of and hard to audit, because a company we turned down never received a message and so never had the chance to prove the reason wrong. Loosening the threshold was the cheapest honest way to check whether those reasons were correct or merely well formatted.
Nothing else moved. Same four discovery sources, same nightly read budget, same enrichment, same drafting, same named person approving every message before it went anywhere. One variable, because at this sample size two variables would have told us nothing at all.
- The other gates stayed exactly where they were. A company with no usable public contact route was still turned down for want of a Door, and a jurisdiction that requires prior consent still turned its candidates down.
- Every message was a Draft until a person read it. We do not deliver anything a human has not approved, and we did not relax that in order to get the volume up.
- Delivery stayed on the target company's own public contact form, under a real named sender at a real address.
- No email. Email sending is scheduled for Q4 2026 and was not part of this test in any form.
- Any form guarded by a CAPTCHA was parked for a human rather than solved. That is the designed behaviour and it did not bend for an experiment.
02The outbound reply rate did not survive the extra volume
June, threshold 72: 4,182 Read, 118 Qualified, 96 of those with a Door. Ninety-six Drafts went to review, the reviewer turned down 15, and 81 were delivered as Knocks. Eleven companies replied.
July, threshold 57: 4,306 Read, 341 Qualified, 268 with a Door. Two hundred and sixty-eight Drafts went to review, the reviewer turned down 54, and 214 were delivered as Knocks. Nine companies replied.
So 13.6 replies per hundred Knocks became 4.2 replies per hundred. That is the whole result and it is worth sitting with for a moment before anyone reaches for an explanation.
Splitting July by score makes the shape plain. Of the 214 Knocks, 78 scored 72 or above and would have qualified anyway, and those returned 8 replies. The other 136 scored between 57 and 71 and returned exactly one.
One reply from 136 messages is not a thin seam worth mining. At that rate we cannot tell the band apart from zero, and we would need several more months of it to try.
The 223 extra companies that qualified in July were not obviously junk, which is the uncomfortable part. Most sold something with a real data component, most were still trading, most had a plausible reason to care about the cost of collection. They read well in a spreadsheet. What separated them from the original 118 was that we could not point at a specific current reason to write to them this week rather than any other week.
03Where the extra cost actually landed
The compute cost went up and it is not worth a sentence. The real cost was the reviewer, and it was large enough that we noticed it in the second week rather than in the numbers afterwards.
Approval runs at roughly four minutes a Draft: read the research, read the message, check the Door resolves, then approve or turn it down. June was 96 Drafts, about six and a half hours. July was 268, close to eighteen. Eleven and a half extra hours of one person's attention bought two fewer replies.
The reviewer also turned down a bigger share as the month went on: 15 of 96 in June, 54 of 268 in July. Sixteen percent became twenty. That is a person doing by eye what the threshold had been doing by arithmetic, slower and with worse consistency, and it is the least defensible line in this whole experiment.
Then there is the cost we cannot put a number against. One hundred and thirty-six companies received a message from us that, on this evidence, was not worth their time. Each of those landed with a person whose job is to triage inbound. We spent their attention, not only our own.
A slower cost sits underneath that and four weeks is too short to see it. Our per-domain rate limit and the tenant suppression list mean we get very few attempts at any one company, ever. A wasted attempt is therefore close to permanent, and 136 of them is a real number of doors we have made harder for ourselves.
04What one month proves, and what it plainly does not
One month, one tenant, 214 delivered messages. That is a single data point on a single business and we are not going to dress it up as a law. If we ran July again we would not expect exactly nine replies, and anyone who would is reading the number too hard.
There is also a result inside the result that we cannot explain. The above-threshold cohort fell too, from 13.6 replies per hundred in June to 10.3 in July. Either July was a worse month for reasons entirely outside the experiment, or 78 Knocks is simply too few to read a rate from. Both are plausible, this data cannot separate them, and pretending otherwise is exactly the sort of claim the refusal ledger exists to prevent.
It does not prove that 72 is the right number either. All we learned is that 57 is worse than 72 on one business over four weeks. The correct threshold might well be 78. Testing upward is the obvious next experiment and it is the harder one, because raising the bar shrinks an already small sample until there is nothing left to measure.
What we do take from it is narrower than a rule about volume. In a channel where delivery is free and the recipient is a machine, more is usually more. This is not that channel. Here a person opens the message, decides within a few seconds whether we understood their business, and that judgement carries over to every message we might ever send them again. Precision compounds in a channel like that. Volume does not.
The threshold went back to 72 on the first of August. The 136 low-scoring companies are in the refusal ledger with their reason recorded, which is where they should have been the whole time.