Schema and search appearance guide
Audit the Structured Data Your WordPress Site Actually Publishes
Collect the structured data from one URL per template, record each object's type, identifier and source, then fix contradictions and misleading markup first.
On this page
A structured data audit starts from the public pages, not from plugin settings. Pick one URL for each template, collect every block of structured data those pages publish, and write down each object’s type, its @id, and the component that printed it. Then sort what you found into three groups: objects that complement each other, objects that contradict each other or the page, and markup that describes content the page does not show. Fix the last two groups. Repetition on its own is a lower priority.
WordPress sites commonly have several sources of structured data at once: an SEO plugin, the theme, a page builder, plugins for events or recipes, and code someone added years ago. The WordPress SEO field guide covers why the delivered page and its description should agree. This guide is the inventory method.
Pick one URL per template
Structured data is usually generated per template, so one example of each is enough to start. For garden.example, a fictional garden design business, an illustrative list:
| Template | Example URL |
|---|---|
| Homepage | https://garden.example/ |
| Blog post | https://garden.example/guides/pruning-tomatoes/ |
| Page | https://garden.example/about/ |
| Service page | https://garden.example/services/garden-design/ |
| Category archive | https://garden.example/category/vegetables/ |
| Author archive | https://garden.example/author/sam/ |
| Any special template | Events, products, recipes, landing pages |
Add a second post if some posts use a different layout or a page builder. Save the list as urls.txt, one URL per line.
Collect the markup each page publishes
Google’s introduction to structured data, checked September 27, 2026, lists three supported formats: JSON-LD, which it recommends, Microdata and RDFa. JSON-LD sits in <script type="application/ld+json"> blocks. Microdata and RDFa are attributes inside the page’s HTML, such as itemscope, itemtype and typeof.
This script lists every JSON-LD object on each URL in urls.txt, including objects inside an @graph, and reports blocks that are not valid JSON:
# schema_inventory.py: list the JSON-LD each URL publishes. Python 3, standard library only.
# Usage: python schema_inventory.py urls.txt
import json, re, sys, urllib.request
SCRIPT = re.compile(r'<script[^>]*application/ld\+json[^>]*>(.*?)</script>', re.S | re.I)
def nodes(data):
if isinstance(data, list):
for item in data:
yield from nodes(item)
elif isinstance(data, dict):
yield from nodes(data["@graph"]) if "@graph" in data else [data]
for url in (line.strip() for line in open(sys.argv[1], encoding="utf-8") if line.strip()):
request = urllib.request.Request(url, headers={"User-Agent": "schema-inventory"})
html = urllib.request.urlopen(request, timeout=15).read().decode("utf-8", "replace")
blocks = SCRIPT.findall(html)
print(f"\n{url}: {len(blocks)} JSON-LD script(s)")
for n, block in enumerate(blocks, 1):
try:
data = json.loads(block)
except json.JSONDecodeError as error:
print(f" script {n}: invalid JSON, line {error.lineno} column {error.colno}: {error.msg}")
continue
for node in nodes(data):
print(f" script {n}: {node.get('@type', '(no @type)')} {node.get('@id', '(no @id)')}")
It reads the HTML the server sends. It does not run JavaScript, and it lists top-level objects only, not ones nested inside another object’s properties. To cover the rest:
- Microdata and RDFa. View the page source and search for
itemtypeandtypeof. The Schema Markup Validator fetches a URL and extracts JSON-LD, RDFa and Microdata, according to schema.org’s description of the tool, checked September 27, 2026. - Markup added by JavaScript, for example through a tag manager. Google’s page on generating structured data with JavaScript, checked September 27, 2026, says Google can process structured data present in the rendered DOM, and recommends testing such pages with the URL input of the Rich Results Test rather than the code input. Compare what that test finds with the script’s output; an object only the test finds was probably added by JavaScript, is Microdata or RDFa (the test reads all three formats, checked September 27, 2026), or sits nested inside another object. The test covers only the rich result types listed on its help page, checked September 27, 2026, so it may not list every object the script finds.
Find which component printed each object
For every object, record its source. The quickest clues:
- Markup around the script. Some plugins add an HTML comment or a class attribute near their output. Search the source a few lines above each script.
- The identifier pattern. WP Visibility, for example, prints one JSON-LD script per page (none on search results or the 404 page) with an
@graphwhose@idvalues end in#organizationor#person,#website,#webpage,#breadcrumb, and on blog posts#articleplus#authorfor the post’s author. WooCommerce product pages also get#product. Set your site identity describes the graph. Custom JSON-LD entered in a post’s Structured data panel is added to that same graph when it is an object, or a list of objects, with its own@type; a pasted@graphwrapper is not printed. - The code. On a copy of the site, search the theme and plugin folders for
ld+jsonand foritemtype:
grep -rl -e "ld+json" -e "itemtype" wp-content/themes wp-content/plugins wp-content/mu-plugins
- Elimination. On a staging copy, deactivate one plugin at a time, or switch to a default theme, and rerun the inventory. The object that disappears belongs to the component you turned off.
Do the elimination on staging, not the live site. Record the source even when it seems obvious; the next audit will need it.
Separate repetition from conflict
Now read the inventory template by template. Google’s structured data general guidelines, checked September 27, 2026, say markup “must be a true representation of the page content” and tell you not to mark up content that is not visible to readers. Those two rules decide most cases.
| What you see | Group | Why |
|---|---|---|
| A plugin’s Organization in its graph, plus a theme’s breadcrumb Microdata with the same path | Complementary | Different facts, no disagreement |
| The same Organization printed twice with the same name and URL | Repetition | Untidy, but the facts agree |
| Two Organization objects with different names or logos | Conflict | Two versions of one fact |
| Two breadcrumb trails with different paths | Conflict | Two answers to where the page sits |
| An Article with an author who is not in the byline | Conflict with the page | Markup disagrees with what readers see |
| Review or rating markup with no reviews on the page | Not representative | Describes content the page does not show |
| Markup for a feature Google has retired | Harmless, no rich result | Of its retired sitelinks search box, Google said unsupported structured data like it won’t cause issues in Search, checked September 27, 2026. See Google’s retired rich results |
A larger graph is not a more accurate one. An object that repeats agreed facts can wait. An object that contradicts the page cannot.
Prioritize and assign repairs
Turn the inventory into a worksheet with one row per object:
| URL | Script | Type | @id | Source | Issue | Owner | Fix |
|---|
Then work in this order:
- Invalid JSON. A block that does not parse describes nothing. See Fix JSON-LD errors in WordPress.
- Facts that contradict the page, such as the wrong author, business name or breadcrumb path.
- Markup for content that is hidden or absent.
- Conflicting duplicates. Decide which component owns each type, then turn off the other one in its settings, rather than deleting output with code.
- Missing recommended properties, only for features you care about.
Assign each row an owner: the SEO plugin settings, the theme, a named plugin, or a developer for custom code. After each change, rerun the inventory on every URL in urls.txt, not only the one you were fixing. A setting that removes a duplicate on posts can also remove something useful from another template.
Schema audit checklist
- One URL per template in
urls.txt. - JSON-LD listed by script; Microdata, RDFa and JavaScript output checked separately.
- Every object has a recorded source.
- Each object marked complementary, repeated, conflicting or not representative.
- Repairs ordered, assigned, and verified by rerunning the inventory on all templates.
