Magento 2 XML Sitemap and Robots.txt Best Practices

Magento 2 XML Sitemap and Robots.txt Best Practices

December 18, 2025 · By Magento Company
Magento 2 XML Sitemap and Robots.txt Best Practices

The XML sitemap and robots.txt are the two files that mediate between your store and every crawler that visits it. Magento generates both natively, but the defaults are generic - and on a store with layered navigation, internal search and multiple store views, generic defaults waste crawl budget or block the wrong things. Here is the configuration we deploy as standard.

XML Sitemap: Generation Settings

Marketing > SEO & Search > Site Map, plus defaults under Stores > Configuration > Catalog > XML Sitemap:

  • Products: include; priority and change frequency are largely ignored by Google, so do not agonise over them
  • Categories: include
  • CMS pages: include
  • Images: enable “Add Images into Sitemap” - image search is real traffic for retail
  • Generation: schedule nightly via cron (Generation Settings > Enabled, pick a low-traffic hour)

For stores over ~50k URLs, Magento splits sitemaps automatically (50k URLs / 50MB limits). Multi-store setups: generate one sitemap per store view, each in its own directory (/pub/sitemap/uk/, /pub/sitemap/de/), and reference each from that store view’s robots.txt.

Submitting and Verifying

Submit each sitemap in Google Search Console per property, and check the coverage report after a week: a large gap between “discovered” and “indexed” URLs is diagnostic gold - usually parameter URLs or thin pages you did not mean to expose.

robots.txt: What to Block

Magento’s default robots.txt (Stores > Configuration > General > Design > Search Engine Robots) is minimal. Our standard additions:

User-agent: *
Disallow: /catalogsearch/
Disallow: /checkout/
Disallow: /customer/
Disallow: /review/
Disallow: /*?*product_list_order=
Disallow: /*?*product_list_dir=
Disallow: /*?*product_list_limit=
Sitemap: https://www.example.com/sitemap/uk/sitemap.xml

The reasoning:

  • /catalogsearch/ - internal search results are infinite, thin and duplicate; classic crawl-budget sink
  • /checkout/, /customer/ - no value in crawling, and they generate session-parameter noise
  • Sort/direction/limit parameters - pure duplicates of the clean category URL

What NOT to Block

  • Layered navigation parameters wholesale unless your canonical/noindex policy is fully settled - blocking in robots.txt means Google can never see the noindex tag (it cannot crawl the page to read it). Pick one strategy: robots block (saves crawl budget, loses signal) or crawlable with noindex (costs budget, communicates clearly). For most mid-size catalogs we noindex facets and keep robots focused on search/checkout/account.
  • CSS and JS - blocking assets breaks rendering in Google’s crawler
  • Anything you actually want indexed, obviously - we have inherited stores whose staging robots.txt (Disallow: /) shipped to production. Add a deploy check.

Maintenance Rhythm

  • robots.txt reviewed at every theme change and platform upgrade
  • Sitemap submission re-verified after domain or HTTPS changes
  • Quarterly Search Console coverage review feeding back into noindex decisions

Two small files, disproportionate leverage. Configure them once with intent, review them on a schedule, and crawlers will spend their budget on the pages that make you money.

SEO Operations Magento 2