The XML sitemap and robots.txt are the two files that mediate between your store and every crawler that visits it. Magento generates both natively, but the defaults are generic - and on a store with layered navigation, internal search and multiple store views, generic defaults waste crawl budget or block the wrong things. Here is the configuration we deploy as standard.
XML Sitemap: Generation Settings
Marketing > SEO & Search > Site Map, plus defaults under Stores > Configuration > Catalog > XML Sitemap:
- Products: include; priority and change frequency are largely ignored by Google, so do not agonise over them
- Categories: include
- CMS pages: include
- Images: enable “Add Images into Sitemap” - image search is real traffic for retail
- Generation: schedule nightly via cron (Generation Settings > Enabled, pick a low-traffic hour)
For stores over ~50k URLs, Magento splits sitemaps automatically (50k URLs / 50MB limits). Multi-store setups: generate one sitemap per store view, each in its own directory (/pub/sitemap/uk/, /pub/sitemap/de/), and reference each from that store view’s robots.txt.
Submitting and Verifying
Submit each sitemap in Google Search Console per property, and check the coverage report after a week: a large gap between “discovered” and “indexed” URLs is diagnostic gold - usually parameter URLs or thin pages you did not mean to expose.
robots.txt: What to Block
Magento’s default robots.txt (Stores > Configuration > General > Design > Search Engine Robots) is minimal. Our standard additions:
User-agent: *
Disallow: /catalogsearch/
Disallow: /checkout/
Disallow: /customer/
Disallow: /review/
Disallow: /*?*product_list_order=
Disallow: /*?*product_list_dir=
Disallow: /*?*product_list_limit=
Sitemap: https://www.example.com/sitemap/uk/sitemap.xml
The reasoning:
/catalogsearch/- internal search results are infinite, thin and duplicate; classic crawl-budget sink/checkout/,/customer/- no value in crawling, and they generate session-parameter noise- Sort/direction/limit parameters - pure duplicates of the clean category URL
What NOT to Block
- Layered navigation parameters wholesale unless your canonical/noindex policy is fully settled - blocking in robots.txt means Google can never see the noindex tag (it cannot crawl the page to read it). Pick one strategy: robots block (saves crawl budget, loses signal) or crawlable with noindex (costs budget, communicates clearly). For most mid-size catalogs we noindex facets and keep robots focused on search/checkout/account.
- CSS and JS - blocking assets breaks rendering in Google’s crawler
- Anything you actually want indexed, obviously - we have inherited stores whose staging robots.txt (
Disallow: /) shipped to production. Add a deploy check.
Maintenance Rhythm
- robots.txt reviewed at every theme change and platform upgrade
- Sitemap submission re-verified after domain or HTTPS changes
- Quarterly Search Console coverage review feeding back into noindex decisions
Two small files, disproportionate leverage. Configure them once with intent, review them on a schedule, and crawlers will spend their budget on the pages that make you money.