Every link building agency publishes its wins. The placements page, the DR screenshots, the logos. What nobody publishes is the other side of the funnel: the sites they looked at and refused, and why. Which is strange, because the refusals are where the quality actually happens.
So here is ours. Over the last five months our prospecting pipeline evaluated 6,563 candidate sites across 5,490 domains. It rejected 2,307 of them, 35%, before a single email was written. Only 19% of everything discovered ever earned a pitch. This piece is the ledger: what got killed, for what reason, from which sources, with the caveats stated plainly.
The funnel, in plain numbers
From February 12 to July 16, 2026, the pipeline looked like this:
- 6,563 prospect records evaluated across 5,490 distinct domains
- 2,307 rejected outright (35.2%) before any outreach
- 1,261 (19.2%) earned a pitch: those sites received at least one of the 5,174 emails we sent in the window
- The remainder are still in evaluation or died of logistics rather than judgment: contact discovery failed, or the site could not be fetched reliably
Read that middle number again. For every five sites that look like prospects, one gets pitched. That ratio is not an inefficiency we are trying to fix. It is the product. The whole promise of our process is that the sites that survive it are worth an editor’s time and a client’s money.
What actually gets a prospect killed
Of the 2,307 rejections, 2,075 have a recorded reason (232 came before we logged reasons properly; more on that below). Here is the full breakdown:
| Reason | Count | Share of reasoned rejections |
|---|---|---|
| Organic traffic below our floor | 1,151 | 55% |
| Batch cap reached (we found enough better sites) | 229 | 11% |
| No reachable contact found | 100 | 5% |
| Wrong niche entirely | 89 | 4% |
| Site unfetchable / bot wall | 83 | 4% |
| Niche fit scored below floor | 66 | 3% |
| Stale or recycled list source | 44 | 2% |
| Does not accept guest posts | 16 | 1% |
| Client already listed there | 2 | <1% |
| Pipeline and manual rejections (operator blacklist, client vetoes, judgment calls) | 295 | 14% |
One filter dominates everything: more than half of all reasoned rejections are sites with no real organic traffic. Not sites with low DR. Sites where third-party traffic data says nobody is actually reading. This is why we qualify on verified traffic and never DR alone: domain metrics are cheap to inflate, sustained readership is not. An independent audit of guest-post marketplaces found 85.3% of marketplace inventory rated low quality, and our own ledger says the underlying problem is upstream of marketplaces: most of the web that accepts placed content has no audience to place in front of.
The smaller buckets are the texture of real operations. A hundred sites had no findable human to write to. Eighty-three hid behind bot walls we chose not to fight. Sixteen said plainly that they do not take guest posts, and we filed that as a rejection rather than a challenge. Two were killed because the client was already on them, which our pipeline checks automatically, because a duplicate placement helps nobody.
Where the junk comes from: rejection rate by discovery channel
The ledger gets most interesting when you sort it by where prospects came from:
| Discovery channel | Rejected | Evaluated | Rejection rate |
|---|---|---|---|
| Custom per-client list | 33 | 260 | 13% |
| Fresh SERP prospecting (v1) | 458 | 2,775 | 17% |
| Fresh SERP prospecting (v2, stricter scoring) | 1,461 | 3,083 | 47% |
| Scraped lists | 254 | 329 | 77% |
| Link-gap mining | 84 | 86 | 98% (small n) |
| Past-publisher re-checks | 17 | 21 | 81% (small n) |
Two honest notes before the conclusion. The v2 SERP number is higher than v1 because v2 scores harder, not because the web got worse: it front-loads rejection so humans review less junk. And the last two rows are small samples; we report them because we have them, not because 86 records prove anything.
But the scraped-lists row is the story. Sites that entered from scraped or purchased lists failed our vetting 77% of the time. Fresh prospecting, whether a hand-built client list or live SERP discovery, fails at 13 to 47%. A circulating list is a pre-pitched, pre-placed, pre-squeezed list: by the time it reaches you, the sites with real readers have been harvested and what remains is the residue. This is the data behind a promise we make on every engagement: fresh prospecting for your niche, no inventory list, no recycled catalog.
What the qualifier actually reads
The traffic floor is arithmetic, but plenty of rejections are editorial. Our qualifier reads pages, not just metrics, and its rejection notes read like an editor’s margin comments. Three real examples from the ledger, lightly trimmed:
- “This is a video/webinar landing page, not a listicle. It describes a panel discussion about CRM trends.”
- “A chaotic multi-niche content-marketing funnel with 2,000+ unstructured tags.”
- “Not a listicle. This is a comprehensive how-to guide about customer feedback analysis.”
None of those sites are scams. They are simply the wrong page for the placement being considered, and a link there would read as what it is: paid decoration. The cheapest moment to catch that is before the pitch, which is the entire argument for spending 35% of prospecting output on the reject pile.
The caveats, stated before the conclusions
Three of them, so the numbers above read at their true weight.
First, 232 of the 2,307 rejections carry no recorded reason, because reason logging was added after the pipeline went live. We report them as unreasoned instead of guessing. Second, ten rejections read “rejected by client”: sites we proposed and the client vetoed. We count those as the gate working, but they are the client’s judgment, not ours. Third, five months is a young dataset, and we label it as such: these are our first 6,563 records, the same way our reply rates are our first 3,945 sends. The percentages will move. The shape of them, we suspect, will not.
What this means if you are buying links
The rejection ledger is a vendor-vetting tool. Anyone selling you links is running some version of this funnel, and the two questions that expose it are:
- What share of prospects do you reject before outreach, and for what reasons? A real operation can answer from records, the way this piece does. A catalog reseller cannot, because a catalog rejects nothing.
- Where did the sites you are proposing come from? Fresh prospecting for your niche, or a standing list? Our data says that single variable moves failure rates from the teens to 77%.
The industry average price of a quality editorial link is about $509 per Reporter Outreach’s survey of 500 practitioners. At that price, the vetting is what you are paying for. If nobody can show you the reject pile, assume there is not one.
The takeaways, if you only skim
- 35% of everything our prospecting discovered in five months was rejected before a single pitch. Only 19% earned an email.
- The traffic floor is the whole ballgame: 55% of reasoned rejections are sites with no real organic readership. Verified traffic, never DR alone.
- Scraped lists failed vetting 77% of the time; fresh prospecting fails at 13 to 47%. Recycled inventory is rejected inventory.
- Stricter automated scoring (our v2) rejects more so humans review less junk. Front-loading rejection is a feature.
- Editorial rejections matter as much as numeric ones: the wrong page type gets killed even when the metrics pass.
- Young data, labeled as such: 232 early rejections carry no reason, and we say so instead of backfilling.
- Ask any vendor for their rejection rate and their reasons. The answer, or the silence, is the audit.
If you want to see what survives this gate, the fastest way is the free baseline: where AI engines cite you today, who they cite instead, and every unlinked mention of your brand we can find. Free, takes us about a day, yours either way.
Questions this piece answers
Why does a link building agency reject so many prospects?
Because most of what looks like a prospect is not one. In our last five months of production prospecting, 35% of discovered sites were rejected before any outreach, and the single biggest reason, roughly half of all recorded rejections, was failing our organic-traffic floor: the site simply has no real readers. Pitching those sites would produce links on pages nobody visits, which is exactly the inventory problem the industry has.
What should the quality bar for a guest post site be?
Verified organic traffic first, never domain metrics alone. DR and DA can be inflated cheaply; sustained organic traffic is much harder to fake. Our floor is traffic-based, checked against third-party data at prospecting time, and it kills more prospects than every other filter we run combined.
Are bought or scraped link lists worth using?
In our data, barely. Sites that entered our pipeline from scraped lists failed vetting 77% of the time, against 13 to 17% for sites found by fresh, per-client prospecting. A list that has been circulating is a list that has been pitched, placed on, and squeezed already. That is why we prospect fresh for every client instead of reselling an inventory.
How can I tell if an agency recycles inventory?
Ask two questions. What percentage of prospects do you reject before outreach, and can you show the reasons? And where did the sites you are proposing come from: fresh prospecting for me, or a standing list? An agency running a real quality gate can answer both from records. If the answer is a blank look, you are buying from a catalog.