Skip to content
WP Visibility

Internal links and site structure guide

Run an Internal Link Audit on a Small WordPress Site

Crawl your own site, record every internal link with its anchor and response, then sort the defects by what readers need so you fix the important paths first.

Published

On this page

An internal link audit should end with a short, ordered repair list, not a spreadsheet of every link on the site. Crawl the site signed out, record each link’s source, anchor text, destination and response, then fix the defects on the pages people depend on before you touch anything else. On a site with a few hundred URLs you can do this in an afternoon with the script below.

The WordPress SEO field guide explains why internal links matter and how to think about reading paths. This guide is the procedure for checking the links you already have.

Set the scope before you crawl

Write down three things before collecting any data:

  1. The site and environment. Audit the live site as a visitor sees it, or a staging copy with the same content. A logged-in browser can show admin links and private posts that visitors never see; WordPress documents private content as only visible to site admins and editors (checked September 27, 2026).
  2. The important pages. List the five to fifteen pages that matter most: the main services or products, contact and booking pages, and the guides people arrive on. Every finding is ranked against this list later.
  3. The limit. Decide the maximum number of pages to crawl. Set it well above the number of posts, pages and archive pages the site has, because paginated archives and linked files such as PDFs count toward it too. A limit that is too low leaves part of the site unread.

Only crawl a site you own or have permission to test. The script below does not read robots.txt and requests every internal URL it finds, following redirects, pausing half a second after each page it crawls.

A small link audit. Crawl signed out: Record source, anchor, destination, status. Sort and rank: Defect type first, then page importance. Fix and recrawl: Edit the source link, then compare crawls.
The audit moves from a bounded crawl to a ranked repair list and a second crawl. It covers only links in the HTML the server sends, up to the limit you set.

Collect URLs, anchors, and destinations

Google’s link guidance, checked September 27, 2026, says Google can generally crawl a link only when it is an <a> element with an href attribute, and that good anchor text is descriptive and relevant to the page it links to. A useful audit records exactly those two things for every link.

Save this as crawl.py. It needs Python 3 and nothing else.

# crawl.py: a bounded crawl of your own site. Python 3, standard library only.
# Usage: python crawl.py https://garden.example/ 200
import csv, sys, time, urllib.error, urllib.request
from html.parser import HTMLParser
from urllib.parse import urldefrag, urljoin, urlparse

START, LIMIT = sys.argv[1], int(sys.argv[2]) if len(sys.argv) > 2 else 200
HOST = urlparse(START).netloc

class Links(HTMLParser):
    def __init__(self):
        super().__init__()
        self.links, self.href, self.text = [], None, []
    def handle_starttag(self, tag, attrs):
        if tag == "a":
            self.href, self.text = dict(attrs).get("href"), []
        elif tag == "img" and self.href is not None:
            self.text.append(dict(attrs).get("alt") or "")
    def handle_data(self, data):
        if self.href is not None:
            self.text.append(data)
    def handle_endtag(self, tag):
        if tag == "a" and self.href is not None:
            self.links.append((self.href, " ".join(" ".join(self.text).split())))
            self.href = None

def fetch(url):
    request = urllib.request.Request(url, headers={"User-Agent": "small-link-audit"})
    try:
        with urllib.request.urlopen(request, timeout=15) as r:
            html = "html" in r.headers.get("Content-Type", "")
            return r.status, r.geturl(), r.read().decode("utf-8", "replace") if html else ""
    except urllib.error.HTTPError as e:
        return e.code, e.geturl(), ""
    except Exception as e:
        return "error: " + type(e).__name__, url, ""

queue, seen, results, links = [START], {START}, {}, []
while queue and len(results) < LIMIT:
    page = queue.pop(0)
    status, final, body = results[page] = fetch(page)
    if urlparse(final).netloc != HOST or (final != page and final in seen):
        continue
    seen.add(final)
    parser = Links()
    parser.feed(body)
    for href, anchor in parser.links:
        target = urldefrag(urljoin(final, (href or "").strip()))[0]
        if urlparse(target).scheme not in ("http", "https") or urlparse(target).netloc != HOST:
            continue
        links.append((final, href, target, anchor))
        if target not in seen:
            seen.add(target)
            queue.append(target)
    time.sleep(0.5)

with open("links.csv", "w", newline="", encoding="utf-8") as f:
    out = csv.writer(f)
    out.writerow(["source", "href", "destination", "anchor", "status", "final_url", "redirected"])
    for source, href, target, anchor in links:
        if target not in results:
            results[target] = fetch(target)
        status, final, _ = results[target]
        out.writerow([source, href, target, anchor, status, final, "yes" if final != target else ""])

with open("pages.txt", "w", encoding="utf-8") as f:
    f.write("\n".join(sorted({r[1] for r in results.values() if r[0] == 200})) + "\n")
print(len(results), "URLs checked,", len(links), "internal links written to links.csv")

Run it with your homepage and your limit. Use the address the site actually serves, with or without www: the script treats any other host as external, so a start URL that redirects to another host writes no links.

python crawl.py https://garden.example/ 300

It writes two files. links.csv has one row per internal link: the page it sits on, the href exactly as written, the resolved destination, the anchor text (an image’s alt text counts as its anchor), the final status, the final URL after any redirects, and a redirected flag. pages.txt lists the final URL of everything that answered 200.

Know what it cannot see. It reads the HTML the server sends and does not run JavaScript, so a menu built by a script is invisible to it. It records the final status after redirects, not the number of hops. It stops crawling at your limit. It still checks the status of every link on the pages it crawled, but it does not read the links on pages past the limit, so a page missing from the results may simply be past it. If you prefer a desktop crawler, export the same columns and the rest of this guide still applies.

Sort the findings into defect types

Open links.csv in a spreadsheet and add a problem column. Most findings fall into one of these types:

Defect How to spot it in links.csv Where to fix it
Broken destination status is 404, 410, 500 or an error Repair broken internal links
Link through a redirect redirected is yes Update links that point through redirects
Vague or misleading anchor Anchors such as “click here”, “read more”, or blank Rewrite the anchor in its sentence
Important page with no contextual links Only menu or footer rows point to it Add links from related passages
Wrong destination The anchor promises one thing and the page answers another Edit the link or the sentence

Menus and footers repeat on every page, so they inflate link counts. To see how many different pages point to each destination, run this in the same folder:

import csv, collections
pairs = {(r["source"], r["final_url"]) for r in csv.DictReader(open("links.csv", encoding="utf-8"))}
for url, n in collections.Counter(dest for _, dest in pairs).most_common():
    print(n, url)

A destination that every page points to is probably in the menu. Your important pages should also appear as destinations from a few pages in the body of related content. Published pages you know about that never appear at all are candidates for the orphan page check.

Prioritize by what readers need

Rank each finding with two questions: how important is the page the link sits on or points to, and how badly does the defect interrupt the reader? A workable order:

  1. Broken links in the main menu, footer, and on the important pages.
  2. Important pages that are reachable only from the menu, with no link from the content that discusses them.
  3. Wrong destinations and misleading anchors on important pages.
  4. Broken links on low-traffic pages.
  5. Links through redirects, which still work for readers, starting with those in templates.
  6. Vague anchors elsewhere.

This list describes defects you observed. Google’s link guidance says there is no ideal number of links a page should contain, and its technical requirements, checked September 27, 2026, state that meeting them does not guarantee indexing. Record a fixed broken link as a fixed broken link, not as a predicted ranking change.

Repair a sample, then crawl again

Fix five to ten findings from the top of the list first. Edit the source, whether that is a post, a menu, a synced pattern or a template, then rename the first links.csv (the script overwrites it), run the crawl again and compare the two files. A fix that worked disappears from the defect list. A fix that changed a template can also change links on pages you did not open, so check that the number of rows is roughly what you expect.

Keep a worksheet with these columns so the next audit starts from a record rather than from memory:

source destination anchor problem priority fix fixed on rechecked

If you use WP Visibility, the Link Graph module gives you a second input: an index of which published posts link to which, taken from post content and reported as orphans and most-linked posts. It does not read menus, footers or templates, and it is available from WP-CLI and REST rather than an admin screen. Read the links report explains the commands. It complements a crawl; it does not replace one, and neither does the public scanner, which does not crawl every page.

Audit checklist

  • Scope, environment, important pages and crawl limit written down.
  • links.csv and pages.txt saved with the date.
  • Every row with a problem has a type and a priority.
  • Top findings fixed at the source, not only redirected.
  • Second crawl run and compared with the first.

Read next

WordPress SEO with your own assistant.

WP Visibility is $99 a year for unlimited sites, client sites included, with a 30-day refund. Use its SEO tools in WordPress or connect a supported assistant. Read how proposal review and permissions work.