A reproducible protocol with XML checks, three separate clocks and observations from our own publication.

Data analysis on a screen, illustrating marketing measurement
Illustrative photograph. Credits
In this article
  1. Collect evidence that supports a decision
  2. Inspect the deployed sitemap
  3. Separate the three clocks
  4. Read the individual inspection correctly
  5. Use segmentation to investigate populations

An accepted sitemap means Google could process that file. It does not certify that every listed page was crawled, indexed or shown for a search. Diagnose a missing URL by comparing four separate observations: the public file, its recorded read date, the page's index status and a current fetch of that page.

This protocol is designed for a blog or store with a recent publishing batch. It complements our robots.txt versus noindex guide. The question here is whether a Google observation describes the current deployment or an earlier version.

Collect evidence that supports a decision

Choose six representative URLs: the home page, a category, a recent article, an older article, a product and an alternate-language page. For a publication without products, use a resource page. This is a diagnostic sample, not an estimate of the whole site's indexed percentage.

For each URL, record its exact address, publication date, final HTTP response, declared canonical, noindex status, latest recorded crawl and live-test outcome. Keep raw account exports private; identifiers and unrelated interface details add no value to a public case study.

Observation What it establishes Next decision
Sitemap returns 200 The file is reachable from this test location Check XML, destinations and dates
Search Console accepts the sitemap Google processed the submission Compare its read date with deployment
Live URL test succeeds Google's test tool fetches the page now Read the separate index result
URL is on Google Inspection knows an indexed version Check impressions and reporting dates
Zero clicks No organic visits measured in this view Do not automatically diagnose removal

Inspect the deployed sitemap

Download the file from the production domain. A successful local build does not establish which deployment is publicly served.

curl --fail --silent --show-error https://example.com/sitemap.xml -o sitemap.xml

The following standard-library Python snippet counts URL entries and duplicates in a urlset. A sitemap index requires checking its child files instead.

import collections
import xml.etree.ElementTree as ET

root = ET.parse("sitemap.xml").getroot()
ns = {"s": "http://www.sitemaps.org/schemas/sitemap/0.9"}
if root.tag != "{http://www.sitemaps.org/schemas/sitemap/0.9}urlset":
    raise SystemExit("Inspect the child sitemaps of this index.")
urls = [node.text for node in root.findall("s:url/s:loc", ns)]
counts = collections.Counter(urls)
print("URLs:", len(urls))
print("Duplicates:", [url for url, count in counts.items() if count > 1])

Compare sample entries with HTML canonical tags and navigation links. A redirect to another language or product needs a destination decision, not repeated submission of the original URL. Test GET responses as well; HEAD behavior can differ.

Google's sitemap documentation recommends absolute URLs intended for search results and accurate significant-modification dates. Replacing every lastmod with today's date on every build discards useful information about actual changes.

Separate the three clocks

Deployment time, sitemap read time and page crawl time can differ. If new articles were published after the displayed sitemap read, the discovered count describes an earlier inventory. It does not establish rejection of the newer pages.

An actual Search Engine Trends check on October 3, 2026 found 547 URLs in the public sitemap. Search Console showed an October 2 read and 383 discovered pages, while its aggregate indexing report was still processing. Those observations cannot produce a reliable rejection rate for the current publishing batch. Individual inspections showed the home page, English resource page and an English Q4 article as indexed.

This is a dated observation of our own publication. It is not a promised processing time for another merchant. Its practical value is the comparison method: every inventory count needs a corresponding observation date.

Read the individual inspection correctly

Open the Google Index result first and let it finish loading. Then run a live test for a representative URL. Record crawl permission, successful retrieval, indexing permission and declared canonical.

The official inspection guide distinguishes the two views. A current test does not predict the eventual Google-selected canonical or establish every quality and security condition. A technical pass establishes present accessibility; it does not produce a ranking.

If retrieval fails, fix the specific DNS, certificate, server or robots problem indicated by the evidence. If retrieval succeeds but index information still says unknown, verify internal discovery and sitemap inclusion before reviewing the page's substance. Changing its slug without a migration reason creates another URL to discover.

For a catalog, check the visible product information as well: the retrieved page should explain the actual offer, delivery context and variant. An accessible empty template is a different problem from an unreachable useful product page. For an article, confirm that the important explanation exists in fetched HTML rather than only in a client-side interaction.

Use segmentation to investigate populations

Separate French articles, English articles and utility pages into diagnostic sitemaps when that helps interpretation. Keep full coverage and consistent language counterparts. The purpose is to compare groups of useful URLs; additional files are not additional ranking votes.

After a meaningful fix, inspect priority pages once and record the date. Subsequent checks should establish whether the sitemap read advances, the sample receives crawls and impressions appear within an available reporting period. An old performance window cannot evaluate a new article.

The next action depends on the observation: unavailable index data calls for better evidence; indexed pages without impressions call for relevance and distinctive value; impressions without clicks call for query, title and promise analysis. Our free tools and resource worksheets can support planning, while your actual Search Console property remains the source for search performance.