Robots.txt controls crawling, not indexing. For product images, the decisive facts are which host serves the file, which user-agent group applies, and whether the image URL is disallowed. Google can still index a disallowed image URL if other pages link to it, but it will not fetch the file itself. Check the image host's own robots.txt, not just the page domain's.

Why product images fail differently from product pages

Google's image guidance states that the technical requirements for appearing in Search apply to images too, and that images have additional requirements because they are a different format. It also notes that Google can find images in the src attribute of an <img> element, including when nested inside elements such as <picture>, and that Google does not index CSS images. That matters for ecommerce because many storefronts render the main product shot as a background image, a lazy-loaded element, or a JavaScript-injected node. If the image is only in CSS, it is not a candidate for image indexing regardless of robots.txt. Official source

Robots.txt is a separate layer. Google's introduction to robots.txt describes it as a way to tell crawlers which URLs they can access, used mainly to manage crawler traffic. For media files, the same documentation says robots.txt can be used to prevent image, video and audio files from appearing in Google Search results, while noting that this will not prevent other pages or users from linking to the file. It also warns that a disallowed URL can still be found and indexed if it is linked from elsewhere on the web. Official source

That combination produces the confusing cases store owners report: an image missing from Google Images even though the product page ranks, or an image URL appearing in results with no thumbnail. The first is usually a discovery or eligibility problem. The second is usually a crawl-blocked file that Google knows about from links.

The host is the first thing to verify

A product image is often served from a different hostname than the page. Common patterns include a CDN subdomain, a separate image domain, or a marketplace asset host. Robots.txt is host-specific. A rule on www.example.com/robots.txt does not govern cdn.example.net. You must fetch the robots.txt of the host that appears in the image URL.

Google's image documentation addresses this directly: image sitemaps can include URLs from other domains in the <image:loc> elements, which lets sites use CDNs to host images. Official source

A worked example (hypothetical)

Imagine a UK retailer whose product page lives at www.example.co.uk and whose images are served from images.example-cdn.net. The page-level robots.txt allows everything. The CDN robots.txt contains a broad rule disallowing a path segment used for resized derivatives. The main product image is served from that path. In this scenario, the page is crawlable, the image file is not, and the image may still surface as a URL if other pages link to it. Nothing about the page's own robots.txt would reveal the problem. This is illustrative only, not a measured case.

A diagnostic you can run in order

  1. Extract the exact image URL from the rendered page, not from the source template. Use the browser's element inspector on the visible product image and copy the resolved src or srcset value.
  2. Identify the host in that URL. Write it down separately from the page host.
  3. Fetch that host's robots.txt at the root, for example https://images.example-cdn.net/robots.txt. If it returns 404, there is no host-level block from that file.
  4. Match the applicable user-agent group. Check which user-agent group applies to the crawler you care about before concluding anything from the wildcard section. Read the group that applies to the crawler you care about before concluding anything from the wildcard section.
  5. Test the specific path, not the whole file. A disallow for /tmp/ or /search/ is irrelevant to /products/.
  6. Check whether the image is in HTML at all. If the visible image is a CSS background, a canvas, or injected only after interaction, robots.txt is not the limiting factor. Official source
  7. Confirm the page-level markup uses a standard <img> element with a descriptive alt attribute, since alt text is one of the few signals available when the file itself is not fetched.
  8. Use Search Console URL Inspection on the image URL where you have verified ownership of the host.

What robots.txt does and does not decide

Layer What it controls What it does not control
robots.txt on the image host Whether the crawler may fetch the image file Whether the URL can appear in results if linked elsewhere
robots.txt on the page host Whether the crawler may fetch the HTML page Whether the image file is crawlable
<img> markup and alt text Whether the image is discoverable and describable Whether the file is allowed to be fetched
Image sitemap Discovery of image URLs, including cross-domain CDN URLs Crawl permission for those URLs

A useful mental model: robots.txt is a door, markup is a signpost, and the sitemap is a directory. A locked door does not remove the signpost. A clear door does not help if there is no signpost.

Where the evidence is thin

Google's documentation does not promise that a crawlable, well-marked-up image will appear in Google Images, nor does it specify a ranking formula for images. It describes eligibility and discovery, not outcomes. It also does not enumerate every CDN configuration or provide a universal checklist for third-party image hosts. Where a specific host's behaviour is unclear, the reliable move is to fetch that host's robots.txt and inspect the actual URL, rather than inferring from the page domain. Claims that a particular robots.txt change will increase image traffic should be treated as hypotheses to test, not guarantees.

For the surrounding page context, see our guide to the product page details buyers actually need and the ecommerce product photography shot list. Both assume the images can be crawled in the first place.

Follow-up questions

Does blocking an image in robots.txt remove it from Google Images?

Not necessarily. Google's documentation states that a disallowed URL can still be found and indexed if it is linked from other places on the web, and that robots.txt for media files prevents the file appearing in results but does not stop other pages linking to it. If removal is the goal, robots.txt is the wrong tool; use removal methods appropriate to the asset. Official source

Do I need a separate robots.txt for my CDN?

Only if the CDN host serves a robots.txt file that applies to crawlers. If the CDN returns 404 for /robots.txt, there is no host-level rule from that file. If it returns rules, those rules govern crawling of image URLs on that host, independent of your main site's robots.txt. Google's image guidance recommends verifying ownership of the CDN's domain name in Search Console so that Google can inform you of any crawl errors it may find. Official source

SEARCH ENGINE TRENDS

Put the idea into practice.

All articles