Indexing and technical SEO guide
Keep a WordPress Staging Site Out of Search and Verify Launch Settings
Put staging behind a login, keep the discourage setting as a backup, and run a launch check so production does not inherit staging's noindex, hostnames, or robots rules.
On this page
Put the staging site behind real access control: an HTTP password, a host-level login, or an IP allowlist, so an anonymous request gets 401, 403, or a login redirect instead of the page. Keep WordPress’s Discourage search engines from indexing this site setting on as a second layer, but do not rely on it alone. Then, at launch, check production for anything staging passed along: a site-wide noindex, canonicals or links naming the staging hostname, and robots or header rules meant for the test copy.
The reason for the order is that search directives are requests, not locks. WordPress’s Reading settings documentation, checked September 27, 2026, says the discourage option does not block access to the site and that it is up to search engines to honor the request. Google’s guidance on controlling what you share, checked September 27, 2026, says confidential or private content needs password protection so only authorized users can reach it.
Protect staging with real access controls
Use the strongest control your host offers. In rough order of preference:
- The host’s staging protection. If your host offers a password or login in front of staging environments, turn it on and confirm it covers every path, including
/wp-json/,/sitemap.xml, and uploads. - HTTP basic authentication at the web server or CDN. Every request without the credentials, crawlers included, gets a
401challenge, as RFC 7617, checked September 27, 2026, describes. - An IP allowlist for the team’s addresses, if everyone works from known networks.
A maintenance-mode or coming-soon plugin is weaker. It shows visitors a holding page, but depending on the plugin, some paths, such as REST routes, feeds, or media files, can stay reachable. If you use one, test those paths too.
Check the protection from outside, signed out, with no stored cookies:
STAGE="https://staging.garden.example"
for p in "/" "/sample-page/" "/wp-json/" "/sitemap.xml" "/robots.txt" "/feed/"; do
printf '%s ' "$p"; curl -s -o /dev/null -w '%{http_code}\n' "$STAGE$p"
done
The hostname is illustrative. Every line should show 401 or 403, or a redirect to the host’s login page. A 200 on any path means that path is public. Use dummy content on staging where you can. Do not copy real customer data into an environment whose protection you have not tested.
Keep the search settings as a second layer
With access control in place, also set staging to discourage search engines. Under Settings, Reading, tick Discourage search engines from indexing this site, or from the command line:
wp option update blog_public 0 --url="$STAGE"
WordPress then prints a noindex robots meta tag on every page that uses wp_head, as the Reading settings documentation cited above describes. The WordPress 5.5 sitemaps announcement, checked September 27, 2026, says core’s sitemap is disabled when this option is set. WP Visibility replaces core’s sitemap while its Sitemaps module is on, and in 2.10.3 that sitemap still answers at /sitemap.xml on a discouraged site, which is one reason the access check above includes that path. The noindex tag matters if the password ever lapses: a crawler that reaches the page reads noindex. It does nothing while the password is working, because a crawler that gets 401 never sees the page.
Do not add a Disallow: / rule in robots.txt as your staging protection. A crawler that obeys it never fetches the pages, so it never reads their noindex, and robots.txt does not stop anyone from opening a URL directly; Google’s robots.txt introduction, checked September 27, 2026, says its rules cannot enforce crawler behavior. Google’s noindex documentation, checked September 27, 2026, says a page must not be blocked by robots.txt for its noindex to work.
Check the outbound side too. Anything on staging that notifies search engines should be off. In WP Visibility 2.10.3, a site set to discourage search engines still sends IndexNow submissions if the IndexNow module is on, so keep that module off on staging and development copies, as Turn On IndexNow notes. On the pages WordPress prints, WP Visibility’s per-post Index setting does not override the site-wide discourage switch, so those pages stay noindexed even if a post says otherwise.
Inspect what production inherits at launch
Launches go wrong in a few predictable ways when a staging database or configuration becomes production:
| Inherited from staging | Symptom on production | Check |
|---|---|---|
| Discourage search engines left on | noindex on every page |
Robots meta tag on the home page and one post |
| Staging hostname in the database | Canonicals, links, images, or sitemap entries naming staging | Search the HTML and sitemap for the staging hostname |
X-Robots-Tag set in staging server config |
noindex header, invisible in the HTML |
curl -I on production pages |
| Staging robots.txt file | Disallow: / on production |
Fetch /robots.txt |
| Basic auth copied to production config | 401 on the live site |
Signed-out fetch |
| Staging-only plugins | Maintenance mode or debugging output | Plugin list comparison |
Run the checks against production right after the switch:
LIVE="https://garden.example"
STAGE_HOST="staging.garden.example"
wp option get blog_public --url="$LIVE" # expect 1
wp option get home --url="$LIVE"
wp option get siteurl --url="$LIVE"
curl -sI "$LIVE/" | grep -iE '^HTTP|x-robots-tag'
curl -s "$LIVE/" | grep -ioE '<meta[^>]*name=.robots.[^>]*>|<link[^>]*rel=.canonical.[^>]*>'
curl -s "$LIVE/robots.txt"
curl -s "$LIVE/" | grep -c "$STAGE_HOST" # expect 0
If the staging hostname appears in the HTML, preview a database replacement before running it:
wp search-replace "https://$STAGE_HOST" "$LIVE" --all-tables-with-prefix --dry-run
Read the dry run’s table counts, take a backup, then run it without --dry-run. The WP-CLI search-replace documentation, checked September 27, 2026, says the command handles PHP serialized data, so use it rather than a raw SQL replace.
WP Visibility’s audit includes a blog_public check that fails when the site discourages search engines, which makes it a quick launch test: wp visibility audit. On staging that failure is expected. The checks are described in Run the Site Audit.
Verify production after the handoff
A short sequence, run by someone other than the person who did the migration if possible:
- Robots. Home page, one post, one page, and one archive: no
noindexin the HTML or headers, unless intended. - Canonicals. Each names the production URL of that page.
- robots.txt. No leftover
Disallow: /, and aSitemap:line naming the production sitemap if your generator adds one. WP Visibility adds that line only when the site is not discouraging search engines. - Sitemap. Fetch it, confirm it returns XML, and check a sample of
<loc>entries use the production hostname. Then submit it in Search Console as described in Submit Your Sitemap. - Staging still locked. Rerun the staging access check. Launch work can loosen staging protection.
- Search Console. Use URL Inspection’s live test, checked September 27, 2026, on the home page to confirm Google can fetch it and sees no
noindex.
Write the results down with the date. If a later deployment pushes staging settings again, that record tells you what changed.
If staging pages are already in search
Close the access first: add authentication so the pages return 401 or 403. Search Console’s Removals tool, checked September 27, 2026, can hide URLs on a Search Console property you own, but Google says a successful request lasts only about six months and that the tool alone will not remove a URL permanently; the page also has to be removed, password protected, or noindexed. For a staging copy, the password is the permanent step.
Then check production is not the one pointing at staging. Search production’s HTML and sitemap for the staging hostname as above. A single canonical or sitemap entry naming staging can keep sending crawlers there. The indexability chapter of the field guide explains how robots.txt, noindex, and access control differ.
