Internal links and site structure guide
Find Orphan Pages in WordPress and Give Useful Pages a Home
Compare the URLs you have published with the pages a crawl from your homepage reaches, check each gap, then link the pages that deserve a route in.
On this page
To find orphan pages, make two lists and compare them: every URL WordPress has published, and every URL a visitor can reach by following links from the homepage. A published page missing from the second list is a candidate orphan. Check each candidate before acting, because some are gaps in the crawl rather than missing links, and some pages are meant to have no links at all. Then add links to the pages that deserve readers.
Google’s link guidance, checked September 27, 2026, puts the goal plainly: “Every page you care about should have a link from at least one other page on your site.” A sitemap does not replace that link. Google says submitting a sitemap is “merely a hint” in its guide to building sitemaps, checked the same day, and it gives visitors no route to the page.
List every URL you have published
With WP-CLI, one command gives you the published posts and pages with their permalinks:
wp post list --post_type=post,page --post_status=publish --field=url > published.txt
Add any public custom post types to --post_type, for example --post_type=post,page,service. The url field is one of the optional fields listed in the wp post list documentation, checked September 27, 2026.
Without shell access, use the sitemap instead. Open /sitemap.xml, or /wp-sitemap.xml if no plugin has replaced the sitemap index WordPress core added in 5.5 (checked September 27, 2026). Open each sitemap the index lists and copy the page URLs into published.txt, one per line. Know the gap: a sitemap can leave out pages set to noindex. WP Visibility’s sitemap, for example, omits noindexed and password-protected posts, as Submit your sitemap explains. Noindexed pages can still be orphans that readers need, so the WP-CLI list is the more complete one.
Crawl the site from the homepage
Run the crawl.py script from the small-site link audit against your homepage, signed out:
python crawl.py https://garden.example/ 500
Set the limit well above the number of lines in published.txt. Archives, paginated lists and other linked URLs count toward it too, and a crawl cut off early never reaches the deep pages. The script writes pages.txt (pages that answered 200) and links.csv (every internal link on the pages it crawled). Then save this as compare.py in the same folder and run it:
# compare.py: list published URLs that the crawl never reached.
import csv
published = {line.strip() for line in open("published.txt", encoding="utf-8") if line.strip()}
reached = {line.strip() for line in open("pages.txt", encoding="utf-8") if line.strip()}
with open("links.csv", encoding="utf-8") as f:
for row in csv.DictReader(f):
reached.update((row["destination"], row["final_url"]))
for url in sorted(published - reached):
print(url)
Each line it prints is a published URL that no link found in the crawl led to. That is your candidate list, not yet your orphan list.
Separate orphans from crawl failures
Before adding a single link, check why each candidate was missed. For a site called garden.example, an illustrative candidate list might contain a planting calendar, a thank-you page, an old pricing page and a guide linked only from a script-driven menu. Only one of those is an orphan worth linking.
| What you see | Likely reason | Check |
|---|---|---|
| The crawl hit its limit | The page is deeper than the crawl went | Rerun with a higher limit |
| The page is linked from a menu built with JavaScript | The script reads server HTML only | View the page source and search for the URL |
The URL differs by http, www or a trailing slash |
The same page under two addresses | Compare with the final_url column |
| The page is only in an archive past page one | Pagination links were not reached | Open the archive and follow its next links |
| A thank-you, ad landing or confirmation page | It is meant to be reached by a form or campaign | Leave it unlinked on purpose |
| No link anywhere | A true orphan | Decide what it needs, below |
If you use WP Visibility, its Link Graph module (off by default) gives you a second list. Turn it on and build its index once with wp visibility links rebuild; the orphans report then lists published, indexable posts that no other published post links to in its content (the first 100 unless you pass a higher --limit, up to 500). It does not read menus, widgets or templates, so a page linked only from the menu appears there even though a visitor can reach it. That difference is useful. A page that is in the Link Graph list but reached by the crawl is probably linked only from a menu, an archive, a template or another place outside post content, with no contextual link, which is often worth fixing too. (Links in content written as relative paths without a leading slash are not counted either, so check the page before assuming.) Read the links report covers the commands.
Decide what each orphan needs
Open each true orphan and choose one of four outcomes:
- Link it. The page answers a real question and belongs on the site. Find where readers need it.
- Merge it. It overlaps with a better page. Move anything unique into that page, then redirect the old URL.
- Leave it unlinked. It exists for a form, a campaign or a private process. Note the reason so the next audit skips it.
- Retire it. It is out of date and has no replacement. Remove it and decide on the response for the old URL; Redirect a changed URL covers the options.
Do not keep a weak page merely because it is published, and do not delete a useful page merely because nothing links to it yet. The question is whether a reader would be glad to land on it.
Add links where readers need them
For each page you keep, find two or three existing passages that raise the question it answers. Search the site for the topic, read the paragraph, and add the link in the sentence where the reader would want the next step. For the planting calendar in the example, that might be the sentence in the pruning guide that mentions timing, and the step in the service page that asks when work should start.
Adding a link from an unrelated popular page creates a route few readers will take. A link in the menu or footer makes a page reachable but tells the reader little about why it matters, so use navigation for the site’s main sections and contextual links for the rest.
When the edits are saved, run the crawl and compare.py again. The pages you linked should disappear from the output. Make the same check part of publishing: every new page gets at least one link from an existing, related page on the day it goes live.
Orphan check summary
published.txtfrom WP-CLI, or the sitemap with its noindex gap noted.- A signed-out crawl with a limit well above the number of published URLs.
- Every candidate checked for crawl limits, script menus, URL variants and pagination.
- Each true orphan linked, merged, deliberately left unlinked, or retired.
- A second comparison run after the edits.
