HomeAllAboutServicesContactArticles

About 1 result (0.09 seconds)

2026-07-24 · 7 min read

Your SEO Platform Can't See Half Your Content — Here's Why That Matters

I ran an audit recently on a mid-market SaaS company's website. Good team, real budget, an SEO platform subscription they actually use. They were doing the work.

Their two best-performing blog posts — the ones pulling the most organic traffic on the entire site — appeared in no XML sitemap. Not the post sitemap, not the page sitemap, nowhere. Meanwhile the sitemap they did publish was full of event landing pages from 2021 that no longer mattered to anyone.

Their SEO platform never flagged it. It wasn't a bug. It's how these tools are built.

What do crawl-based SEO tools miss?

Crawl-based SEO tools miss anything they cannot reach: drafts, orphaned URLs, custom post types, and pages left out of XML sitemaps.

Semrush's Site Audit gives you four options for what to crawl: your website (following links from the homepage), the sitemap listed in robots.txt, a sitemap you specify by URL, or a file of URLs you upload. Semrush's own Site Audit configuration documentation recommends using your sitemap as the crawl source, because that's what search engines use. Ahrefs' Site Audit works the same way.

That's a completely reasonable design. These tools exist to tell you how you're performing in search, and search engines only care about pages they can reach. Measuring what Google sees is exactly the right job.

But it means the audit inherits your site's blind spots. If you crawl by sitemap, anything missing from the sitemap is invisible to the audit. If you crawl by following links, anything unlinked is invisible. And in both cases, a whole category of content never shows up at all:

  • Drafts and scheduled posts — the half-finished piece someone started eight months ago covering the exact topic you're about to assign
  • Unlinked or orphaned pages — live, indexable, and reachable by nobody
  • Non-public custom post types — case study libraries, resource collections, and internal structures that never render as standard pages
  • Anything your sitemap forgot — which, based on how often I find this, is more sites than you'd expect

None of that is a knock on Semrush or Ahrefs. You cannot crawl your way to a list of things that aren't crawlable. It's a boundary of the approach, not a flaw in the product.

Why are performance data and content inventory different questions?

Performance data measures search visibility. Content inventory maps the assets you own, including content that search tools may never see.

Here's the distinction that matters when you're deciding what tools to buy:

"How are we performing?" is a question about the outside world. What do we rank for, who's beating us, where are we losing ground, what's the traffic trend. This requires an index of the web, competitor data, and rank tracking infrastructure. Semrush and Ahrefs are excellent at this, and if you're serious about SEO you should be paying one of them.

"What do we actually have?" is a question about your own house. How many pages exist. Which ones say the same thing. Whether the article being commissioned this week was already written in 2023. Which content is decaying and which is fine. What's sitting in drafts.

The second question requires reading your content library directly — from the database, not from a crawl. Different data source, different job, different tool.

Most teams have bought a very good answer to the first question and no answer at all to the second one.

What does the inventory gap cost?

The inventory gap costs money through duplicate assignments, unmanaged content growth, late cannibalization diagnosis, and stale pages that nobody notices.

This isn't theoretical. The gap shows up as specific, recurring, expensive problems:

How do teams commission content they already own?

A content manager assigns an article. Nobody remembers that a similar piece was published two years ago under a different title, or that a draft covering it is sitting unfinished. You pay a writer to produce something you already had, and then the two pieces compete with each other in search.

Why does the library outgrow the team?

Every site past a few hundred pages hits a point where nobody can hold the whole library in their head. New people inherit it and have no map. The instinct becomes "when in doubt, publish something new," which accelerates the problem.

How does cannibalization get diagnosed backwards?

SEO platforms catch keyword cannibalization by noticing two of your URLs competing in the SERPs. That works — but only after both pages are live, indexed, and fighting. Reading the library directly catches the overlap before the second page is ever written.

Why is content decay invisible until traffic drops?

Content that no longer reflects your product, pricing, or positioning doesn't announce itself. It just quietly underperforms. Without an inventory that tracks freshness against your own library, you find out from a traffic graph months later.

What should the two-tool setup look like?

Use the SEO platform for market performance and the content inventory for owned-library visibility. They solve different problems.

For any site past roughly 100 pages, the sane configuration is both:

The questionThe toolWhat it reads
How are we performing in search?Semrush, AhrefsThe web — SERPs, competitors, backlinks, crawled pages
What do we own, and what's redundant?A content inventoryYour content library, directly

They're complementary, not competing. The platform tells you where you stand. The inventory tells you what you're standing on.

Where does Content Miner fit?

Content Miner fits on the inventory side: it reads WordPress directly and turns the library into evidence for editorial decisions.

Content Miner is the tool I built for the second question. It's a WordPress plugin that reads your library from the database — posts, pages, drafts, scheduled content, and public custom post types — and builds a persistent inventory you can actually act on.

What that makes possible:

  • Exact-copy and near-duplicate detection across the full library, including content no crawler would reach
  • Pre-commission validation — check a proposed title, question, or brief against everything you already have, and get a Create, Update, Differentiate, or Localize recommendation before the work is assigned
  • Readiness scoring on four separate tracks — SEO, AEO, GEO, and AIO — at both site and page level
  • Thin and stale content flagging against your own library, with GA4 evidence attached
  • Content families and coverage mapping, so you can see the shape of what you own

It reads canonical metadata from Yoast, Rank Math, and AIOSEO, so it respects the SEO configuration you already have rather than duplicating it.

And it is deliberately advisory. It never deletes, redirects, rewrites, or publishes anything. Every recommendation comes with evidence, and a person decides. For content operations at any real scale, that constraint is the point — automated content changes are how sites get quietly wrecked.

When don't you need a content inventory tool?

If your site is under 50 pages and one person wrote all of it, you already have the inventory. It's in your head. Buy the SEO platform, skip the inventory tool, revisit when the library outgrows your memory or the person who wrote it leaves.

The tool earns its place when the library gets bigger than any one person's recall — which, in my experience, happens somewhere between 100 and 200 pages, or the moment a second content hire starts.

How do you find the content gap?

Ask how many pages your site has before you open a tool. If the real number surprises you, you have an inventory problem.

Try this: without opening a tool, how many pages does your site have? Now check.

If those two numbers don't match, that's the gap. Your SEO platform isn't going to close it — that's not what it's for. But you can't run a content strategy on a library you can't see.

See what Content Miner does →

Gary Corriston runs Corriston Consulting, working with agencies and in-house marketing teams on paid media, SEO, marketing operations, and demand gen infrastructure. He's also building Campaign Budget Optimizer, an AI-native cross-platform budget allocation tool launching May 2026.

Contact →