Indexing and technical SEO guide
Fix a WordPress Sitemap That Google Cannot Fetch or Parse
Fetch the exact sitemap address Search Console holds, separate access failures from XML errors, repair the cause at its source, and check every child file before resubmitting.
On this page
Start with the exact address in Search Console’s Sitemaps report and fetch it yourself, signed out. If the response is not a 200 with XML in the body, you have an access problem: a wrong address, a redirect, a block, or an HTML page served in place of the file. If it is a 200 with XML and the report says Has errors, you have a parsing problem inside the document or one of its child files. The two need different fixes, and changing plugin settings solves neither when the cause is a firewall, a host rule, or stray output from a theme.
Google’s Sitemaps report help, checked September 27, 2026, uses three statuses. Success means the file was fetched and read without errors. Has errors means it was fetched but contains one or more errors, and the URLs Google could parse are still queued. Couldn’t fetch means Google could not retrieve the file at all. The same page lists fetch causes such as a robots.txt block, an incorrect URL or 404, server unavailability, and low crawl demand for the sitemap.
Fetch the exact submitted sitemap URL
Copy the address from the report rather than typing it. Search Console stores what was submitted, which may be an old plugin’s file name, http instead of https, or a hostname without www.
SITEMAP="https://garden.example/sitemap.xml"
curl -sI "$SITEMAP"
curl -sIL "$SITEMAP" | grep -iE '^HTTP|^location|content-type|x-robots-tag'
The first command shows the first response. The second follows redirects and prints each hop. The address is illustrative; use the one from your report.
Read the result against this table:
| Observation | Kind of problem | Where to look |
|---|---|---|
404 or 410 |
Wrong address | The generator’s current file name; remove the stale submission |
301 or 302 to another host or path |
Address mismatch | Submit the final address; check the site’s home URL setting |
401 or 403 |
Access control | Password protection, a firewall, or a security plugin |
5xx or a timeout |
Server | Host logs, PHP errors, memory limits on large files |
200 with Content-Type: text/html |
HTML served instead of XML | A maintenance page, login page, or cache of an error page |
200 with an XML content type |
Access works from your machine | Go to the XML checks below |
An X-Robots-Tag: noindex header on the sitemap applies to the sitemap’s own address, not to the pages it lists. Google’s robots meta tag documentation, checked September 27, 2026, describes the header as part of the HTTP response for a given URL. Google’s Sitemaps report help, checked September 27, 2026, also lists a noindex response header, in its section on deleting a sitemap, among the ways to stop Google from continuing to visit the sitemap. WP Visibility’s sitemap sends X-Robots-Tag: noindex, follow by design.
Separate access failures from XML errors
For an access failure, find the component that answered. Look at the response headers for a server or CDN name, a cache status, or a challenge page. Then check robots.txt, since a Disallow rule covering the sitemap path stops Google from fetching it:
curl -s "https://garden.example/robots.txt"
A request from your own machine is not the same as a request from Googlebot. Firewalls and bot-protection rules can treat crawlers differently from browsers, and sending Googlebot’s user agent string with curl does not make your request come from Google’s network. If your fetch works but Search Console reports Couldn’t fetch, ask whoever runs the firewall or CDN for the log lines for your sitemap path around the time of Google’s attempt. Google’s crawler verification instructions, checked September 27, 2026, describe checking an IP with a reverse DNS lookup followed by a forward lookup, or matching it against Google’s published IP ranges. That tells you whether a blocked request really was Googlebot.
WordPress causes of access failures to check:
- A staging or maintenance mode left on. The file returns a holding page. Turn the mode off or exclude the sitemap path.
- Security plugin or firewall rules that challenge unknown clients. Allow verified crawlers by the method the firewall documents.
- An address change. After moving to
httpsor a new domain, the old submission redirects or fails. Submit the address on the current property. - Search engine visibility discouraged. WordPress core disables its own sitemap when Settings, Reading is set to discourage search engines, according to the WordPress 5.5 sitemaps announcement, checked September 27, 2026. A site launched from staging with that box still ticked may have no core sitemap to fetch.
Find the XML error
When the file downloads but the report says Has errors, open it and check what comes first. The XML 1.0 specification, checked September 27, 2026, places the <?xml declaration at the start of the document’s prolog and reserves the xml name, so a blank line or space printed before it makes the file malformed. In WordPress that output can come from a theme or plugin PHP file with whitespace after a closing ?> tag.
# Show the first bytes, including invisible ones
curl -s "$SITEMAP" | head -c 64 | od -c | head -n 4
# Validate the whole document if xmllint is installed
curl -s "$SITEMAP" -o sitemap.xml && xmllint --noout sitemap.xml
After the offset, the first line of od output should begin with < ? x m l. A \n or a space before it is the fault. The specification permits a UTF-8 byte order mark, which od shows as 357 273 277, so that alone is not the error. xmllint reports each error it finds with its line number, or prints nothing if the document is well formed.
Map the report’s error to its likely source. The error names come from Google’s Sitemaps report help:
| Report error | What to check in WordPress |
|---|---|
| Leading whitespace | Blank lines or spaces before <?xml; Google says this alone won’t stop it processing the file |
| Parsing error | A PHP warning printed into the file, or an unescaped & in a URL |
| Invalid date | A lastmod that is not a W3C date, or gives a time without a time zone; check custom filters |
| URL not allowed | URLs on a different host or protocol from the sitemap, for example after a domain change |
| URLs not followed | Listed URLs with too many redirects for Google to follow, or relative URLs |
| Too many URLs | A child file above the 50,000 URL limit |
| Nested sitemap indexes | An index that lists another index, for example when one generator’s index links another’s |
A PHP notice or warning printed into the XML means PHP is displaying errors on the live site. WordPress’s debugging documentation, checked September 27, 2026, says setting WP_DEBUG_DISPLAY to false hides errors and should be paired with WP_DEBUG_LOG so they can be reviewed later. The wp_debug_mode() source, checked September 27, 2026, applies that setting only while WP_DEBUG is on; with WP_DEBUG off, turn off PHP’s display_errors in the server configuration instead. Then find and fix the warning itself.
Check for two generators
If more than one sitemap generator is active, you can submit one while Google is also reading another, or one index can link another. List every sitemap address the site answers:
for p in sitemap.xml sitemap_index.xml wp-sitemap.xml; do
printf '%s ' "$p"; curl -s -o /dev/null -w '%{http_code}\n' "https://garden.example/$p"
done
curl -s "https://garden.example/robots.txt" | grep -i '^sitemap:'
Keep one generator, submit its index, and remove other submissions from Search Console. WP Visibility’s Sitemaps module turns off WordPress’s /wp-sitemap.xml while it runs and serves its own index at /sitemap.xml; the file layout is in Submit Your Sitemap. Another SEO plugin with its own sitemap enabled at the same time is a second generator, and that one needs switching off in its own settings.
Validate the repaired file and its children
A clean index can still point to broken child files. Fetch every child and check its status and first bytes:
curl -s "$SITEMAP" | grep -o '<loc>[^<]*</loc>' | sed 's/<[^>]*>//g' | while read -r child; do
code=$(curl -s -o child.xml -w '%{http_code}' "$child")
first=$(head -c 5 child.xml)
echo "$code $first $child"
done
Each line should read 200 <?xml followed by the child address. Then run xmllint --noout on any file that looks wrong.
If a fix does not appear, the old response may be cached. WP Visibility caches sitemap files for up to a week and marks them stale when a post is saved or deleted or a term is edited or deleted. Other changes, such as sitemap settings, do not clear that cache; wp visibility flush marks every file stale so the next request rebuilds it. A page cache or CDN keeps its own copy, so purge the sitemap paths there too. The cache tracing guide shows how to tell which layer answered.
Resubmit and read the result
Once every file returns 200 and parses cleanly:
- In Search Console’s Sitemaps report, enter the index address and click Submit. Google’s help says a submitted sitemap should be fetched immediately.
- Wait for the status and Last read date to update.
- If the status is Success, move on to individual URLs. Google’s sitemap guidelines, checked September 27, 2026, describe a submitted sitemap as a hint that does not guarantee Google will download it or use it for crawling.
- If one page is missing from a working sitemap, that is a different problem: see why a WordPress page is missing from your XML sitemap.
Keep a note of the cause and the fix with the date. Sitemap failures can return after a theme update, a new security rule, or a domain change, and the note tells the next person where to look first.
