Skip to content
WP Visibility

Indexing and technical SEO guide

Fix a robots.txt Block Without Exposing Private Pages

Find the robots.txt rule that matches the blocked URL, narrow it to what should stay uncrawled, and protect private pages with a login, not robots.txt.

Published

On this page

Fetch the robots.txt file Google is actually reading, find the rule that matches the blocked URL, and narrow that rule so it covers only what should stay uncrawled. First check whether a physical robots.txt file on the server is overriding the one WordPress generates. Anything genuinely private needs a password or login, because robots.txt asks crawlers to stay away; it does not stop anyone.

This guide covers Google’s crawling rules only. Rules for AI crawlers are covered in Set the AI Crawler Policy, and a firewall or host blocking Googlebot outright is covered in Googlebot access errors.

Two statuses, two different problems

The Page Indexing report can show a robots.txt block in two ways, and they call for opposite fixes. Both descriptions are from Google’s Page Indexing report help, checked September 27, 2026.

Status Google’s description If you want the page in search If you want it out of search
URL blocked by robots.txt “This page was blocked by your site’s robots.txt file.” Remove or narrow the rule Leave it, or switch to noindex for a cleaner removal
Indexed, though blocked by robots.txt “The page was indexed despite being blocked by your website’s robots.txt file.” Remove or narrow the rule Remove the block and add noindex

The second row surprises people. Google’s robots.txt introduction, checked September 27, 2026, says robots.txt “is not a mechanism for keeping a web page out of Google,” and that a disallowed page “can still be indexed if linked to from other sites,” in which case the result appears without a description. Google’s noindex documentation, checked September 27, 2026, says that when robots.txt blocks a page, “the crawler will never see the noindex rule,” which is why the fix for unwanted indexing is to unblock and noindex. Find an accidental noindex covers the opposite case.

Find the robots.txt Google is reading

Google’s robots.txt specification, checked September 27, 2026, says the file must sit in the top-level directory of a host and that its rules apply only to that host, protocol, and port. https://www.example.com/robots.txt and https://example.com/robots.txt are different files, and so is a staging subdomain.

curl -s https://example.com/robots.txt
curl -s https://www.example.com/robots.txt

Then work out where the file comes from:

  • WordPress’s virtual file. When no physical file exists, WordPress answers /robots.txt itself. The do_robots() reference, checked September 27, 2026, shows the default output: User-agent: *, a Disallow for the admin path, and an Allow for admin-ajax.php. Plugins add lines through the robots_txt filter.
  • A physical file. If a real robots.txt exists in the site’s root folder, the web server serves it and WordPress never runs: WordPress’s standard Apache rewrite rules, checked September 27, 2026, hand a request to WordPress only when no matching file exists. Check with your host’s file manager or over SFTP. A physical file left over from an old site, a developer, or a staging setup can hold stale rules, and no WordPress setting changes it.
  • An edge rule. Some hosts and CDNs can serve or rewrite robots.txt. Cloudflare’s managed robots.txt, checked September 27, 2026, prepends its own AI crawler rules to your file when enabled, or serves a new file when none exists. If the file you fetch matches neither WordPress’s output nor a file on disk, ask your host.

One outdated belief is worth clearing up. The same do_robots() reference notes that since WordPress 5.3 the Discourage search engines option in Settings > Reading no longer writes Disallow: / into the virtual robots.txt; it uses a robots meta tag instead. A Disallow: / in a current WordPress robots.txt therefore came from a physical file, a plugin or theme, or the host.

In Search Console, the robots.txt report, checked September 27, 2026, shows which robots.txt files Google found for your top hosts, when each was last checked, and the fetch status. Compare that with what you fetched.

Find the rule that matches the URL

Google applies robots.txt in two steps, per its specification. It picks the single most specific group whose User-agent line matches the crawler (groups naming the same crawler are merged into one) and ignores the rest, so a User-agent: Googlebot group replaces the User-agent: * group for Googlebot rather than adding to it. Within that group, the most specific matching path wins, and when rules conflict Google “uses the least restrictive rule.” Paths are prefixes: * matches any run of characters and $ marks the end of the URL.

Here is an illustrative file for the fictional garden.example, with what each line does:

User-agent: *
Disallow: /wp-admin/
Allow: /wp-admin/admin-ajax.php
Disallow: /guides
Disallow: /*?
Disallow: /wp-content/

User-agent: Googlebot
Disallow: /checkout/
Line What it blocks Problem
Disallow: /guides /guides/, and also /guides-for-beginners/ Prefix match catches more than intended; /guides/ with a trailing slash is narrower
Disallow: /*? Every URL with a query string Also blocks URLs you may want crawled
Disallow: /wp-content/ Theme and plugin CSS, JavaScript, and uploaded images Google’s robots.txt introduction warns it won’t do a good job of analyzing pages that depend on blocked resources
User-agent: Googlebot group Only /checkout/ for Googlebot Googlebot ignores the * group entirely, so none of the rules above apply to it

The last row shows why reading the whole file matters. Adding a Googlebot group to block one path silently removes every other rule for Googlebot.

Three tools, three jobs. robots.txt: Asks crawlers not to fetch a path. noindex: Keeps a crawlable page out of results. Login or password: Stops anyone without access.
Each control does a different job. Only authentication restricts access; robots.txt and noindex are instructions that crawlers choose to follow.

Correct access without exposing private pages

Before you delete a rule, ask why it was added. Rules that “protect” staging copies, member areas, or draft files are the ones to replace, not remove.

Consider a private staging site at staging.garden.example that relies on Disallow: /. That rule does not stop a person or a crawler that ignores robots.txt from loading the pages, and Google’s robots.txt introduction notes that rules “may not be supported by all search engines.” It also does not reliably keep a linked URL out of Google’s index. For a staging site, put the whole host behind HTTP authentication or a login. A crawler that gets a 401 cannot read the content, and nothing needs to be listed in a public file. The staging hostname here is illustrative; do not publish a real one to test this.

For each rule in your file, decide:

  1. Public and should be crawled: remove the rule, or narrow its path so it no longer matches.
  2. Public but not useful in search: allow crawling and use noindex, so Google can read the instruction.
  3. Private: require a login or password, then the robots.txt rule is optional.
  4. Crawl waste with no search value, such as endless filter combinations: a narrow Disallow is reasonable.

Where the file comes from decides where you edit. For WordPress’s virtual file, change the setting or code in the plugin or theme that adds the rule. WP Visibility adds a Sitemap: line to the virtual file while its Sitemaps module is on. It adds AI crawler rules only when you block a crawler, and a Content-Signal: line only when you set a Content-Signals policy. It adds nothing while search engines are discouraged, and it does not detect or edit a physical file. For a physical file, edit or delete it on the server.

Verify the change

  1. Fetch robots.txt again for every host you use, with and without www.
  2. In the Search Console robots.txt report, open the more settings menu next to the file and choose Request a recrawl. Google’s specification says it generally caches robots.txt for up to 24 hours.
  3. Inspect a previously blocked URL and run Test live URL. The URL Inspection tool help, checked September 27, 2026, says Crawl allowed? shows whether a robots.txt rule blocked Google; it should now report that crawling is allowed.
  4. Open the reason in the Page Indexing report and click Validate fix. Google’s help says validation typically takes up to about two weeks.

Google’s specification also covers what happens when robots.txt itself fails to load: on a server error, Google stops crawling the site for the first 12 hours, then uses the last good copy it cached for up to 30 days. If the robots.txt report says Google could not fetch the file for a reason other than a 404, rather than showing a rule problem, treat it as a server problem and see Googlebot access errors.

Before you finish, confirm:

  • You edited the source that actually serves the file: physical file, plugin, or host.
  • Every Googlebot-specific group repeats the general rules it still needs.
  • Private areas are protected by authentication, not by a Disallow line.
  • Pages you want kept out of search are crawlable and carry noindex.

Read next

WordPress SEO with your own assistant.

WP Visibility is $99 a year for unlimited sites, client sites included, with a 30-day refund. Use its SEO tools in WordPress or connect a supported assistant. Read how proposal review and permissions work.