Crawl budget only matters for very large or rapidly changing sites, so the first step is to establish that a problem exists. Google's own guidance frames it as an advanced topic for sites with roughly a million or more unique pages, or 10,000-plus pages changing daily, or a large share of URLs stuck as Discovered - currently not indexed. For most stores, a current sitemap and the Page Indexing report are enough.
What crawl budget actually means
Google describes crawl budget as having two elements: crawl capacity limit and crawl demand (Official source). Crawl capacity is how much Google is willing to spend on your host; crawl demand is how much it wants your URLs based on popularity and freshness. Google treats a site as a unique hostname, so www.example.com and shop.example.com are separate budgets.
That definition matters because it rules out a common assumption. A crawl budget problem is not "Google is ignoring me." It is a mismatch between the URLs Google spends time on and the URLs you need indexed. If Google crawls thousands of filter combinations and never reaches your new season products, that is a real problem. If it crawls your catalog steadily and indexes the pages that matter, you do not have one.
Do you qualify as a large site?
Google's guide is explicitly for advanced cases, and it gives rough, non-exact thresholds: large sites of 1 million+ unique pages with content changing about weekly, medium or larger sites of 10,000+ unique pages with daily-changing content, or sites with a large portion of URLs classified as Discovered - currently not indexed. These are estimates for classification, not pass/fail rules.
A mid-size UK or US retailer with 5,000 stable product pages and weekly stock updates usually falls outside this. A marketplace with millions of variant URLs, or a classifieds-style catalog where inventory turns over daily, falls inside it.
How do you prove a problem with evidence?
Before changing anything, gather three independent signals and check whether they agree.
- Search Console Page Indexing report. Look at the proportion of URLs in Discovered - currently not indexed versus Crawled - currently not indexed. The first suggests Google knows the URL but has not fetched it; the second suggests it fetched and chose not to index. These point to different causes.
- Server logs. Filter requests by verified Googlebot user agent and reverse DNS, then count hits per URL pattern (product, category, filter, search result, pagination). Compare the share of crawl hits going to URLs you consider valuable against the share going to parameter or session URLs.
- Sitemap coverage. Compare the URLs in your sitemap with the URLs actually crawled. A large gap between submitted and crawled URLs, sustained over weeks, is a signal worth investigating.
If logs show Googlebot spending most of its time on low-value parameter URLs while new products wait days, you have a case. If logs show steady coverage of new products and the index report is clean, you do not.
A worked example (hypothetical)
Imagine a hypothetical US outdoor gear retailer with 40,000 product pages and a faceted navigation generating roughly 600,000 crawlable filter URLs. Server logs over 30 days show 78% of Googlebot hits going to filter combinations, 12% to category pages, and 10% to products. New products appear in the index after 9 to 14 days. These numbers are illustrative only.
That pattern is consistent with crawl waste: Google is spending capacity on URLs that should not be indexed, and product discovery is slow. The fix is not "increase crawl budget." It is to stop generating crawlable filter URLs, consolidate duplicates, and make the product set easier to reach. The same retailer with 40,000 products but clean canonical signals and 3-day product indexing would not need this work.
What actually reduces crawl waste
Google's guidance is specific about a few mechanisms. Consolidate duplicate content so Google does not spend time on near-identical URLs. Avoid letting faceted or session parameters create infinite crawlable spaces. And do not use robots.txt to temporarily redirect crawl budget toward other pages; use it only to block resources you never want crawled (Official source).
A practical sequence for a large catalog:
- Fix internal linking so important products are reachable within a few clicks.
- Canonicalise or block parameter URLs that duplicate category content.
- Keep the sitemap accurate and limited to canonical, indexable URLs.
- Re-measure logs and the Page Indexing report after 4 to 6 weeks.
What the evidence does not tell you
Google does not publish a numeric crawl quota, and the thresholds in its guide are explicitly rough. You cannot calculate "your crawl budget" from a formula. You can only observe behaviour: what Googlebot fetches, how often, and what ends up indexed. Treat any tool that promises a precise crawl budget number with caution.
It is also worth separating crawl problems from indexing problems. A page can be crawled daily and still excluded because of quality, duplication, or thin content. In that case, crawl optimisation will not help; the content decision will.
When to leave it alone
If your pages are crawled the same day they are published, Google's guide says you do not need it. For most ecommerce sites, keeping the sitemap current and reviewing the Page Indexing report regularly is adequate. Spending engineering time on crawl budget when there is no measurable problem is a common and avoidable cost.
If you are diagnosing a broader store issue, the same discipline applies elsewhere: diagnose the ecommerce journey before testing and prepare the offer before the traffic.
Follow-up questions
Does a large number of Discovered - currently not indexed URLs always mean a crawl budget problem?
No. It is one of the signals Google lists for classifying a site as advanced, but it can also come from weak internal linking, orphan pages, or low-value content. Check server logs and sitemap coverage before concluding it is a crawl capacity issue.
Can I increase my crawl budget directly?
Not through a setting. Crawl demand responds to signals like popularity and freshness, while crawl capacity is Google's decision. The practical lever is reducing waste so Google spends its existing capacity on URLs you care about.