Skip to content
WP Visibility

Schema and search appearance guide

Audit the Structured Data Your WordPress Site Actually Publishes

Collect the structured data from one URL per template, record each object's type, identifier and source, then fix contradictions and misleading markup first.

Published

On this page

A structured data audit starts from the public pages, not from plugin settings. Pick one URL for each template, collect every block of structured data those pages publish, and write down each object’s type, its @id, and the component that printed it. Then sort what you found into three groups: objects that complement each other, objects that contradict each other or the page, and markup that describes content the page does not show. Fix the last two groups. Repetition on its own is a lower priority.

WordPress sites commonly have several sources of structured data at once: an SEO plugin, the theme, a page builder, plugins for events or recipes, and code someone added years ago. The WordPress SEO field guide covers why the delivered page and its description should agree. This guide is the inventory method.

Pick one URL per template

Structured data is usually generated per template, so one example of each is enough to start. For garden.example, a fictional garden design business, an illustrative list:

Template Example URL
Homepage https://garden.example/
Blog post https://garden.example/guides/pruning-tomatoes/
Page https://garden.example/about/
Service page https://garden.example/services/garden-design/
Category archive https://garden.example/category/vegetables/
Author archive https://garden.example/author/sam/
Any special template Events, products, recipes, landing pages

Add a second post if some posts use a different layout or a page builder. Save the list as urls.txt, one URL per line.

Collect the markup each page publishes

Google’s introduction to structured data, checked September 27, 2026, lists three supported formats: JSON-LD, which it recommends, Microdata and RDFa. JSON-LD sits in <script type="application/ld+json"> blocks. Microdata and RDFa are attributes inside the page’s HTML, such as itemscope, itemtype and typeof.

This script lists every JSON-LD object on each URL in urls.txt, including objects inside an @graph, and reports blocks that are not valid JSON:

# schema_inventory.py: list the JSON-LD each URL publishes. Python 3, standard library only.
# Usage: python schema_inventory.py urls.txt
import json, re, sys, urllib.request

SCRIPT = re.compile(r'<script[^>]*application/ld\+json[^>]*>(.*?)</script>', re.S | re.I)

def nodes(data):
    if isinstance(data, list):
        for item in data:
            yield from nodes(item)
    elif isinstance(data, dict):
        yield from nodes(data["@graph"]) if "@graph" in data else [data]

for url in (line.strip() for line in open(sys.argv[1], encoding="utf-8") if line.strip()):
    request = urllib.request.Request(url, headers={"User-Agent": "schema-inventory"})
    html = urllib.request.urlopen(request, timeout=15).read().decode("utf-8", "replace")
    blocks = SCRIPT.findall(html)
    print(f"\n{url}: {len(blocks)} JSON-LD script(s)")
    for n, block in enumerate(blocks, 1):
        try:
            data = json.loads(block)
        except json.JSONDecodeError as error:
            print(f"  script {n}: invalid JSON, line {error.lineno} column {error.colno}: {error.msg}")
            continue
        for node in nodes(data):
            print(f"  script {n}: {node.get('@type', '(no @type)')}  {node.get('@id', '(no @id)')}")

It reads the HTML the server sends. It does not run JavaScript, and it lists top-level objects only, not ones nested inside another object’s properties. To cover the rest:

  1. Microdata and RDFa. View the page source and search for itemtype and typeof. The Schema Markup Validator fetches a URL and extracts JSON-LD, RDFa and Microdata, according to schema.org’s description of the tool, checked September 27, 2026.
  2. Markup added by JavaScript, for example through a tag manager. Google’s page on generating structured data with JavaScript, checked September 27, 2026, says Google can process structured data present in the rendered DOM, and recommends testing such pages with the URL input of the Rich Results Test rather than the code input. Compare what that test finds with the script’s output; an object only the test finds was probably added by JavaScript, is Microdata or RDFa (the test reads all three formats, checked September 27, 2026), or sits nested inside another object. The test covers only the rich result types listed on its help page, checked September 27, 2026, so it may not list every object the script finds.

Find which component printed each object

For every object, record its source. The quickest clues:

  • Markup around the script. Some plugins add an HTML comment or a class attribute near their output. Search the source a few lines above each script.
  • The identifier pattern. WP Visibility, for example, prints one JSON-LD script per page (none on search results or the 404 page) with an @graph whose @id values end in #organization or #person, #website, #webpage, #breadcrumb, and on blog posts #article plus #author for the post’s author. WooCommerce product pages also get #product. Set your site identity describes the graph. Custom JSON-LD entered in a post’s Structured data panel is added to that same graph when it is an object, or a list of objects, with its own @type; a pasted @graph wrapper is not printed.
  • The code. On a copy of the site, search the theme and plugin folders for ld+json and for itemtype:
grep -rl -e "ld+json" -e "itemtype" wp-content/themes wp-content/plugins wp-content/mu-plugins
  • Elimination. On a staging copy, deactivate one plugin at a time, or switch to a default theme, and rerun the inventory. The object that disappears belongs to the component you turned off.

Do the elimination on staging, not the live site. Record the source even when it seems obvious; the next audit will need it.

Sorting what the audit finds. Complementary: Different facts that agree with the page. Conflicting: Two versions of one fact, or markup against the page. Not representative: Describes content readers cannot see.
Fix conflicting and unrepresentative markup first. Repeated objects that agree are untidy but lower priority.

Separate repetition from conflict

Now read the inventory template by template. Google’s structured data general guidelines, checked September 27, 2026, say markup “must be a true representation of the page content” and tell you not to mark up content that is not visible to readers. Those two rules decide most cases.

What you see Group Why
A plugin’s Organization in its graph, plus a theme’s breadcrumb Microdata with the same path Complementary Different facts, no disagreement
The same Organization printed twice with the same name and URL Repetition Untidy, but the facts agree
Two Organization objects with different names or logos Conflict Two versions of one fact
Two breadcrumb trails with different paths Conflict Two answers to where the page sits
An Article with an author who is not in the byline Conflict with the page Markup disagrees with what readers see
Review or rating markup with no reviews on the page Not representative Describes content the page does not show
Markup for a feature Google has retired Harmless, no rich result Of its retired sitelinks search box, Google said unsupported structured data like it won’t cause issues in Search, checked September 27, 2026. See Google’s retired rich results

A larger graph is not a more accurate one. An object that repeats agreed facts can wait. An object that contradicts the page cannot.

Prioritize and assign repairs

Turn the inventory into a worksheet with one row per object:

URL Script Type @id Source Issue Owner Fix

Then work in this order:

  1. Invalid JSON. A block that does not parse describes nothing. See Fix JSON-LD errors in WordPress.
  2. Facts that contradict the page, such as the wrong author, business name or breadcrumb path.
  3. Markup for content that is hidden or absent.
  4. Conflicting duplicates. Decide which component owns each type, then turn off the other one in its settings, rather than deleting output with code.
  5. Missing recommended properties, only for features you care about.

Assign each row an owner: the SEO plugin settings, the theme, a named plugin, or a developer for custom code. After each change, rerun the inventory on every URL in urls.txt, not only the one you were fixing. A setting that removes a duplicate on posts can also remove something useful from another template.

Schema audit checklist

  • One URL per template in urls.txt.
  • JSON-LD listed by script; Microdata, RDFa and JavaScript output checked separately.
  • Every object has a recorded source.
  • Each object marked complementary, repeated, conflicting or not representative.
  • Repairs ordered, assigned, and verified by rerunning the inventory on all templates.

Read next

WordPress SEO with your own assistant.

WP Visibility is $99 a year for unlimited sites, client sites included, with a 30-day refund. Use its SEO tools in WordPress or connect a supported assistant. Read how proposal review and permissions work.