
Only Google Reads Your Schema Correctly
Quick answer: Google changed its JSON-LD extraction in August 2026 to apply a single pass of HTML unescaping instead of two. A census of the Tranco top 10,000 found that change altered how 10 domains are read — 0.34% of those publishing JSON-LD. The 199 domains (6.70%) whose markup Googlebot and every standards-conformant parser already read differently is the twenty-times-larger problem, and nothing in Search Console reports it.
The structured data panic of August 2026 was aimed at the wrong number. Google narrowed its JSON-LD parser, a dozen agency posts told everyone to audit their schema immediately, and the actual blast radius of the change was ten domains in a ten-thousand-site sample. Meanwhile the defect that change exposed — markup that only renders correctly inside Google and garbles everywhere else — sits on twenty times as many sites and is getting more expensive every quarter, because Googlebot is no longer the only machine reading your page.
Here is what changed, why one unescaping pass is still one too many, and how to tell in five minutes which side of the line your markup is on.
What Google actually changed
Google Search Central announced the change on LinkedIn in August 2026, in wording relayed by Search Engine Roundtable: “To bring our parser up to JSON and other standards, we changed our JSON-LD extraction and are now only applying a single pass of HTML unescaping.” Gary Illyes pointed to RFC 8259 section 7 in the same thread, which is the part of the JSON spec that defines how strings escape characters.
Read that carefully. Google did not stop unescaping HTML entities in JSON-LD. It went from two passes to one. That distinction is the whole story.
The change is not in Google's documentation. It is not in the structured data general guidelines, and it is not in the Search Central documentation updates log, whose August and September 2026 entries cover favicons, the site reputation policy, preferred sources, VideoObject, aggregator units and the Search profile badge — and nothing about JSON-LD parsing. The only written artefact is a social post. That is worth noting before you build process around it: vendor reference docs lag vendor announcements, and this one has not caught up at all.
One pass is still one too many
The HTML specification treats script as a raw text element. Character references inside it are not decoded by the HTML parser. So when a conformant reader pulls the contents of a <script type="application/ld+json"> block and hands it to a JSON parser, the correct number of HTML unescaping passes is zero. The bytes go straight into the JSON parser as written.
Google is now at one. Which means there is still a class of value where Googlebot and everything else disagree, and it is exactly the class where an HTML entity appears in a JSON string.
| Bytes in your script block | Conformant parser (0 passes) | Googlebot (1 pass) |
|---|---|---|
| Smith & Sons | Smith & Sons | Smith & Sons |
| Smith & Sons | Smith & Sons | Smith & Sons |
| Smith &amp; Sons | Smith &amp; Sons | Smith & Sons |
Row two is the common case and the one that matters. A templating layer HTML-escaped your brand name on the way into the script block. Google silently repairs it. Every other consumer stores the literal entity as part of the string, which is how a company ends up filed in a knowledge graph as “Smith & Sons”.
Row three is what broke in August. Sites that had double-escaped were relying on Google's second pass, and when the second pass went away the extra entity started surviving into the parsed value. Very few sites were in row three. A lot are in row two.
The numbers: 10 versus 199
A parser conformance census sampled the Tranco top 10,000 domains on 20 August 2026, fetched each homepage over HTTPS and extracted the JSON-LD blocks. 6,509 domains returned content. 2,969 of those published JSON-LD, across 5,020 blocks. Each string value was then read three ways — at zero unescaping passes, at one, and at two — with Google's own validator used as the oracle for what Googlebot does, rather than taking the announcement on faith.
The split, out of the 2,969 domains publishing JSON-LD:
| Finding | Domains | Share |
|---|---|---|
| Read differently by Googlebot than by conformant parsers | 199 | 6.70% |
| Invalid JSON for every parser | 36 | 1.21% |
| Actually affected by Google's single-pass change | 10 | 0.34% |
Twenty to one. The census also names the visible casualties — Dell and Investopedia showing corruption inside Google Search results — and, more instructive, the sites in the other bucket: Samsung, Meta, NASA and the White House, all with markup that reads correctly in Google and incorrectly anywhere a standards-conformant parser is doing the reading.
Treat the sample honestly: homepages only, one snapshot, large domains. A smaller site on a hand-rolled template is likelier to be in the 6.70% than Samsung is. But the ratio is the finding, and the ratio is not close.
Where the extra escaping comes from
Nobody double-escapes on purpose. It arrives through one of four routes, and all four are a layer doing its job in the wrong order.
Template auto-escaping
Jinja, Blade, ERB, Twig and Handlebars escape interpolated values for HTML by default. That is correct for body copy and wrong inside a script block. Serialise the object to JSON first, then emit it through the template's raw or unescaped filter — not the other way round.
Rich-text CMS fields
A WYSIWYG field stores entities because it stores HTML. Pipe that field into a description property untouched and you have shipped entities into a JSON string. Strip tags and decode entities on the way out, before serialisation.
Double serialisation in React and friends
Putting your JSON-LD in JSX as a child of a script tag gets it escaped as text. The pattern that works is a single JSON.stringify on a plain object, injected raw. Two stringify calls, or a stringify plus an HTML-escape helper, produce row three.
Copy-paste from a generator into a template
Blocks are correct when they leave a generator and get re-escaped when they land in a templating system that treats them as content. If you build blocks with the FAQ schema generator, the HowTo schema generator or the breadcrumb schema generator, check the rendered page, not the generator output. The bug is introduced downstream of the tool every time.
How to check yours in five minutes
View source on a rendered page — the bytes the server sent, not your framework source — and search the script block for an ampersand followed by letters and a semicolon. Any hit is a finding. Then paste the block into the JSON formatter: if it fails to parse you are in the 1.21% and no search engine is reading it at all, which is a worse and simpler problem.
Do this on your highest-value templates rather than one page. Escaping bugs are template-level, so one bad partial is thousands of bad pages.
One trap: a rich results test showing the corrected value tells you Googlebot repaired it, not that the markup is right. Google's validator is the oracle for Google's behaviour and nothing else. To see what a conformant reader gets, parse the raw block yourself.
Why the 6.70% is getting more expensive
For twenty years, markup that only Googlebot read correctly was markup that worked. That assumption expired. Your JSON-LD is now consumed by AI answer engines, retrieval pipelines, social unfurlers, price aggregators and whatever the marketing team wired into a spreadsheet — and none of them inherited Google's forgiveness. A brand name filed with a stray entity is a wrong entity everywhere except the one place you were testing.
This is the same pattern as every other structured data story this year. Google narrowed what it does with markup when it retired FAQ rich results, tightened what earns review stars on software listings, and added properties to VideoObject without making them required. In each case the value shifted away from a Google SERP feature and toward being a clean, machine-readable description of the thing. Correctness is now the feature — not a rich result, but a record that survives whatever reads it. If you are still weighing how much of the AI-readability toolkit is real, our llms.txt reality check makes the same argument from the other direction.
Fixing an escaping bug is a one-line change in a serialiser. It is the cheapest structured data work available, and unlike most of it, the outcome does not depend on what Google decides next quarter.
FAQ
Did Google's change break my structured data?
Almost certainly not. In the Tranco top 10,000 census, 10 of 2,969 domains publishing JSON-LD (0.34%) had values that Google's move from two unescaping passes to one changed. You only fall in that group if your markup was double-escaped, meaning an HTML entity's ampersand was itself escaped.
How many HTML unescaping passes should a JSON-LD parser apply?
Zero. The HTML specification treats script as a raw text element, so character references inside it are not decoded, and the contents go to the JSON parser as written. Googlebot applies one pass, which is closer to the standard than the two it used before but still not the same as zero.
How do I tell if my JSON-LD is double-escaped?
Look at the raw HTML the server returns, not your template source, and search the ld+json script block for an ampersand followed by letters and a semicolon. Any HTML entity inside a JSON string value is the bug. Paste the block into a JSON formatter to confirm the block itself is valid.
Is this change documented by Google?
No. It was announced in a Google Search Central LinkedIn post in August 2026. Neither the structured data general guidelines nor the Search Central documentation updates log mentions single-pass unescaping, so there is no reference page to cite internally.
Does this affect Microdata or RDFa too?
No. Microdata and RDFa live in ordinary HTML attributes and text nodes, where the HTML parser decodes character references for you, so a single escape is the correct representation. The double-decoding problem is specific to JSON embedded in a raw text element.
Google's forgiveness was always a loan, and it is being called in one parser at a time.


