Only 77.6% of 807 parsed UAE product pages carry structured product data that survives merchant-validity checks — the data AI answer engines are built to read. A lower bound: measured, then attacked by its own authors.
In plain terms: most parsed UAE product pages already carry merchant-valid product data. Where they do not, the larger share publishes no Product schema at all, and the rest fails on named fields — price, priceCurrency, availability. The anatomy of the headline (its frame, its adversarial audit, its mandatory caveats) begins one scroll below, and what this means for a UAE merchant follows straight after it.
The lead number, taken apart
Merchant-valid means the Product JSON-LD block carries the five fields an answer engine actually needs — name, image, price, priceCurrency, availability — machine-readable in the server-served HTML, before any JavaScript runs. Presence alone is not enough: a block that omits the price is a product a machine cannot quote.
| Outcome | pages | share |
|---|---|---|
| Merchant-valid Product JSON-LD | 626 | 77.6% |
| Product JSON-LD present, fails the validity check | 63 | 7.8% |
| No Product JSON-LD at all | 118 | 14.6% |
The three rows close exactly: the two failure modes sum to the inverse share in the lead sentence, on the same denominator. Shares of the invalid modes quoted on other denominators exist in the underlying data — they are not this complement, and this page does not mix them.
The number was attacked from both sides. Sampled adjudication of the valid stratum upheld 5 of 5 rows; on the invalid side, 3 of 10 sampled rows flipped to valid on re-examination — every one on the same nested-price pattern the validator does not read. Cells under twenty ship as counts, never rates. The denominator was attacked too: the sample found archive pages inside F_prod that can never be merchant-valid; removing them would raise the rate, the same direction as the validator limitation. Both attacks point one way: the true rate is at least what we publish.
What this means for a UAE merchant
Everything above is fixable, and most of it is one-time work. In rough order of impact:
- Check your product page's JSON-LD validity, not just its presence. The gap between "present" and "merchant-valid" is this study's second finding. View the server-served source and look for the five fields: name, image, price, priceCurrency, availability.
- If you run on Shopify, check where your price lives in the offer markup. The audited failures were prices nested in a priceSpecification block — valid schema that many parsers, including ours, do not read. Keep a direct price on the offer.
- Make sure the schema is in the server-served HTML. Most AI crawlers do not execute JavaScript. If the block assembles in the browser, it does not exist for them.
- Make an explicit crawler decision instead of inheriting a default. Nearly all measured robots.txt files say nothing about AI crawlers either way. Whatever your policy is — allow or disallow — write it down; a default is not a policy.
- The classic layer is not your differentiator here. robots.txt, sitemaps and basic schema presence are broadly in place across the market. The AI-specific layer — valid product schema and explicit crawler decisions — is where nearly everyone is silent, which is exactly where the early advantage sits.
This measurement was made with the same engine Ipponar uses to measure store AI-visibility commercially. If you want to know where your own store stands, start at ipponar.hu/en.
Present is not the same as valid
(engine 689 of 807 · adjudicated bound 83.3% (689 of 827) to 90.9% (709 of 780))
63 pages sit in the gap between presence and validity: the block is there, the machine still cannot use it. This is the study's second-strongest finding, and it travels with the lead everywhere on this page — presence-without-validity is never a footnote here.
On the same 807 pages, the machine-readability check on the price field failed on 175 (21.7%) — an upper bound, for two stated reasons: this count includes the 118 pages that publish no Product JSON-LD at all, and the validator does not read one legal nested price format.
A field-by-field breakdown of which schema fields fail exists in the measurement, but it has no row in this study's sealed publication table, so this page does not print it.
Shopify: the template does the work — configuration decides the rest
| Outcome | pages | share |
|---|---|---|
| Product JSON-LD present | 496 | 89.9% |
| Merchant-valid | 479 | 86.8% |
The gap between the two rows is 17 stores — a cell under twenty, so it ships as a count, never a rate. 5 of the 17 were adjudicated; 2 flipped. The adjudicated envelope is 479 to 496 of 552 if every gap row flipped; 12 of 17 gap rows were never adjudicated — the bound is open. The engine number is what ships; no corrected mix appears anywhere on this page.
Presence near nine-in-ten is what a platform template produces by default. The gap is consistent with configuration- and app-level choices: 13 of the 17 gap rows fail on the price/priceCurrency pair, and in the audited sample the flips came from prices delivered in a nested priceSpecification format our validator does not read — valid schema, unread by our rubric, which is exactly why the lead ships as a lower bound.
Other platform cells were under twenty rows; per the small-cell rule they get counts, never rates, and are not shown here.
Before validity comes discovery
A product page can only be judged if it can be found. On 807 of 966 measurable homepages (83.5% · frame F_home) our engine found and validated a product page.
This is a lower bound on "a discoverable product page exists": in the sampled remainder, 2 of 5 no-product-page rows flipped on re-examination — a product page exists; the discovery ladder missed it. Counts, not rates: the sample is not a census. And the funnel's first step holds almost everything: 966 of 981 census homepages were measurable at all (98.5%).
The classic search layer is mostly in place
Before the AI-specific findings, the baseline — and the contrast it sets up. Note the frame column: these three rows stand on three different denominators.
| Check | stores | share | frame |
|---|---|---|---|
| robots.txt served | 960 of 981 | 97.9% | census, all stores |
| Sitemap declared in robots.txt | 879 of 975 | 90.15% | conclusive robots files |
| Organization-family JSON-LD on the homepage | 689 of 966 | 71.3% | measurable homepages |
Three honesty notes belong right here, not in small print. The robots figure's complement is an upper-bound reading: of the non-serving remainder, 15 returned client errors and 6 were inconclusive, and on re-fetch the true no-robots range is 11 to 15. The sitemap figure sits inside a two-pass measured bound of 874 to 881 of 975. And the Organization row is presence, not validity — the check validates no fields, so it must not be read as a quality figure. (Its numerator matching the product-presence count elsewhere on this page is coincidence: different metric, different denominator.)
AI crawlers: nobody blanket-blocks — and almost nobody decides
The measured agents: CCBot, ClaudeBot, GPTBot, Google-Extended, OAI-SearchBot, PerplexityBot, meta-externalagent (+Bytespider observational).
Exactly 0 of 975 conclusive robots.txt files blanket-block all crawlers (a wildcard group disallowing everything). We ship this as "nobody blanket-blocks" — not as "everyone explicitly allows": a minority of those files have no evaluated wildcard group at all, and permissive-by-default is not an explicit permit.
| Agent | stores | share |
|---|---|---|
| CCBot | 37 | 3.8% |
| ClaudeBot | 35 | 3.6% |
| GPTBot | 35 | 3.6% |
| Google-Extended | 35 | 3.6% |
| OAI-SearchBot | 1 | 0.1% |
| PerplexityBot | 1 | 0.1% |
| meta-externalagent | 34 | 3.5% |
| Bytespider — observational; never scored, never summed into any AI-blocking figure | 37 | 3.8% |
38 stores disallow at least one of the 7 measured agents. The far larger story is silence:
(901 of 975 · frame F_robots_conclusive)
Under RFC 9309, silence defaults to allowed — but a default is not a decision. These stores have not decided anything about AI crawlers; they have inherited a default.
Has the UAE market noticed AI search? Two numbers, two denominators
One number looks like adoption; the other shows near-total silence about the crawlers that matter. Files are present; decisions are absent.
The two bars stand on different denominators (966 and 975). They must not be combined, subtracted, or read as one population. The unlabeled remainder of the silence bar is a mix of categories and deliberately carries no number.
The complement of the llms.txt figure, stated precisely: 270 of 966 measurable homepages have no llms.txt with content — 265 serve none at all, 5 serve an empty file. And an open question, stated honestly: whether the present llms.txt files are merchant-authored or platform boilerplate is a question our own classifier failed its audit gate on — so this study does not answer it — held back in the original round; measured by a blind-gated successor classifier in round 2026-08-13Q - see Part Two, Authorship.
Limits, honestly
- The adversarial audit is a sample, not a census: per-cell samples are under twenty rows, so per-cell results ship as counts, never rates. Where every member of a stratum was individually adjudicated, the page says so.
- The product-family rates are lower bounds by a named engine limitation: the validator reads only direct offer-level prices, not the nested priceSpecification form. Every audited flip in the product family was exactly this pattern.
- The discovery rate is a lower bound: the audit found product pages our discovery ladder missed.
- The parsed-product denominator contains some archive pages that can never be merchant-valid; removing them would raise the published rates. Direction is stated, magnitude is not — the sample is too small to project.
- Two families run the other way and are labeled upper bounds where they appear: the price-field failure count and the no-robots reading.
- The robots figures come from the round's pinned pass; an independent later pass drifts on a handful of domains (real robots.txt drift, not a parser defect), and the per-agent drift is disclosed in the appendix.
- Exactly 1 of 981 census rows carries a reproducible engine-side fetch-persistence failure; it stays in its bucket and is named as an instrument limitation rather than silently reclassified.
- Single measurement environment, single window (2026-08-07..2026-08-12 (UTC)). AI-search behavior and site configurations both move; these are point-in-time readings.
What did not clear the publication gate
- An llms.txt authorship classifier (merchant-authored vs. boilerplate) failed its own audit gate — its negative branch did not mean what the label said — so no authorship figure ships.
- A combined authorship-by-silence figure depended on the failed classifier and is held back.
- Both held-back items above: held back in the original round; measured by a blind-gated successor classifier in round 2026-08-13Q - see Part Two, Authorship.
- The funnel's failure-bucket split, the field-by-field invalidity distribution, and category breakdowns shipped as a sealed addendum (round 2026-08-13AD) - see the Addendum section.
- No number on this page existed publicly before the sealed export behind it.
Addendum - sealed in round 2026-08-13AD
These rows were measured in the study's original window and cleared the publication gate in a separate sealed round; every figure below resolves from that round's sealed table.
Axis. The category axis is the pre-registered census column, cross-checked against two record-level sources with zero disagreement; the census wins and is never merged. Sealed axis rule, verbatim: census column is the pre-registered axis; record-level category fields cross-checked (sites JSONL, robots checkpoint); disagreement -> axis note, census wins, never merged
Small cells. A rate ships only when numerator and denominator both clear the pre-registered floor; otherwise a count with its named denominator — and the failure-vector rows ship as counts by pre-registered rule regardless of size. English working translation, keyed: a rate ships only if num >= 20 AND den >= 20; otherwise only a count with the named denominator; the AD-vec-* rows are uniformly counts (pre-registered, addendum rule) · sealed small-cell rule, Hungarian, verbatim: rata csak akkor megy ki, ha num >= 20 ES den >= 20; kulonben csak darabszam a megnevezett nevezovel; az AD-vec-* sorok egysegesen darabszamok (pre-registered, addendum-szabaly)
Direction. The original round's bound directions travel unchanged: validity rates remain lower bounds, and the field-invalidity readings remain upper bounds. English working translation, keyed: Sampled adversarial audit: 10/60 flips (16.7%), direction almost exclusively engine under-count -> the published rates are LOWER BOUNDS · inherited sealed audit record, Hungarian, verbatim: Mintaveteles adverzarialis audit: 10/60 flip (16,7%), irany szinte kizarolag engine-alulszamolas -> a kozolt ratak ALSO KORLATOK
Never a headline. The category breakdown is disclosed-only by its sealed axis contract; nothing in this section is a headline number.
Recurring sealed caveat families — English key translated rendering · the sealed originals in each block — Hungarian unless marked otherwise — are authoritative
Translated rendering. The mandatory caveats below recur across many rows of the Addendum and Part Two; each caveat block names which of them apply and renders its own sealed original in full. The sealed text inside each block is the authoritative record — Hungarian unless marked otherwise; where a secondary figure exists only inside a sealed original, this key describes it without quoting it — the number stays in the original beside.
Measured bound (adjudication): sampled adjudication upheld 5 of 5 rows in the merchant-valid stratum (626 rows) — no false positive; in the presence-without-validity strata 3 of 10 sampled rows flipped to valid, every flip the same nested-price error pattern; 1 further adjudicated row stayed unresolved and keeps the engine's reading; the per-stratum sizes and splits stay in the sealed originals. Sampled cells under twenty ship as counts, never rates.
Named validator limit (nested price): the engine's product rubric reads only offer-level price keys — proven in the engine code — and does not descend into the nested price-specification form; an adjudicated stored-page quotation of exactly that form is in the sealed original. The lead validity rate is therefore a lower bound, and the price-family failure counts are upper bounds.
True complement: on F_prod, 181 of 807 pages (22.4%) are not merchant-valid — exactly 118 with no Product JSON-LD at all (14.6%) plus 63 present-but-invalid (7.8%); the sum closes exactly. The share the sealed record's header quotes on the narrower JSON-LD-publishing base (rendered in the sealed original beside) stands on a different denominator; the two are not interchangeable and neither is this complement.
Denominator purity: in the sampled no-JSON-LD stratum, 2 of 5 rows turned out to be platform archive listings validated as product pages by a weak cart-marker branch — F_prod contains some non-product pages that can never be merchant-valid; removing them would raise the rate, the same direction as the named validator limit.
Engine number, never a corrected mix: the engine's own number ships verbatim; sampled adjudication is a sample, not a census, and no corrected blend is ever published.
New instrument, same window: the row runs on the corpus downloaded and frozen in the original round, read with the extended validator branch; it is never a correction of the sealed parent row and not a new measurement window — only the instrument differs, the input bytes are identical.
Extended validator is a superset: the extended offer view contains the original validator's offer list and additionally descends into nested price-specification nodes, so the extended counter is per-row at least the original's; the difference is exactly the named nested-price limitation — not new data, not a new measurement.
Stored-byte instrument: every parse runs on the stored, size-capped, phone-masked bytes, never on the live page; a late JSON-LD block can fall past the cap (a named sub-frame), so presence-based counters are lower bounds and absence-based counters upper bounds.
Frozen-corpus exception: one stored body in the frozen corpus was syntactically broken by Part One's phone-masking — a digit run inside a product-identifier value was overwritten with a mask token — so the stored-form reading sees invalid there where the fetch-time fields were valid; the exception set is computed from code, and the conservation gate runs with the sealed bound widened by exactly that set.
New measurement window: the row stands on pages re-fetched in this round's own window (F_q); any difference against the original window is real-world change — the page changed, went away, or now refuses the request — not measurement error; the full breakdown ships in the drift table, nothing hidden.
Self-selected frame: F_q holds the domains that fetched and parsed in this window; the drop-outs are not a random subsample, so an F_q rate does not transfer automatically to the full input frame.
Two routes: every published cell was produced twice, by independent routes — record-level, and re-extraction from the stored HTML — and the two must agree cell-exactly; on any disagreement the round prints a symmetric difference and writes no table.
Single-source HTML cell: for a few checks both routes read the same stored bytes through the same frozen function; the real second leg is the separate recount artifact — a seam named, not hidden.
Honest N: where the judged layer did not complete, the denominator is the number of accepted rows and the shortfall is named; no silent imputation or averaging anywhere.
Instrument row: the row describes the measuring instrument, not the shops — it states the measured error of the judged rows beside it and reads together with them.
Judged residue, conservative: the numerator sums deterministic matches and accepted judged match verdicts; unresolved rows keep their deterministic reading and never enter the numerator — the error direction is conservative.
No aggregate home (prereg seam): the aggregate carries no per-field home for this breakdown, so the pre-registered stated-from-the-aggregate rider cannot be met; a named prereg-versus-artifact seam, and the reason for the verdict ceiling.
Conflated frame: on the conflated frame the failure counts include every page with no Product JSON-LD at all — the rubric leaves all five fields false on an empty node list (measured: no such row has any field true); the true per-field counts live in the clean-frame table.
Frozen corpus, no re-fetch: the input is the set of llms.txt bodies stored in the original round, sha-pinned per file; the published split is a re-reading of the original window with a new instrument and claims nothing about today.
Duplicate-cluster limit: clustering is exact-hash only, on lower-cased, whitespace-normalized bodies; near-duplicates are namedly out of scope, so the template-page count is a lower bound and the authored count an upper bound. Exact-hash clustering found no duplicate cluster, and the judged remainder is 43 bodies.
Rebuilt authorship instrument: Part One's authorship classifier failed its audit gate — its negative branch did not mean its label — so no figure shipped then. The rebuilt instrument treats authored as a positive finding that must cite merchant-specific content; a missing marker alone is never enough. These rows carry new keys; they neither override nor revive the sealed held-back rows.
Secondary axis, disclosed-only: the category breakdown is disclosed-only, never a headline; the axis is the pre-registered census column, the record-level category field is cross-check only — on disagreement the census wins, the disagreement is disclosed as an axis note, and nothing is merged.
Robots frames and passes: the primary silence rate stands on the engine's own conclusive robots frame — a frame choice, not a correction; the census-frame secondary form is 91.85% (901 of 981). The explicit-decision count is measured directly as 74 of 975, never derived by subtraction — on the census frame the subtraction complement is false and may not be published in any form; a small named set of rows falls in neither class, listed in the sealed original. The breakdown uses the parent round's canonical checkpoint pass; a recount on the sites pass differs by a small symmetric set of domains, and the two passes never mix.
Category breakdowns — disclosed-only
Each table stands on its own per-category denominators inside one parent frame; the ratio in every cell names that row's denominator, and rows from different tables must never be combined.
| Category | stores | share |
|---|---|---|
| beauty | 96 of 125 | 76.8% |
| electronics | 84 of 107 | 78.5% |
| fashion | 177 of 214 | 82.7% |
| food & beverage | 92 of 116 | 79.3% |
| home & furniture | 83 of 122 | 68.0% |
| other | 94 of 123 | 76.4% |
| Category | stores | share |
|---|---|---|
| beauty | 103 of 125 | 82.4% |
| electronics | 96 of 107 | 89.7% |
| fashion | 185 of 214 | 86.5% |
| food & beverage | 100 of 116 | 86.2% |
| home & furniture | 101 of 122 | 82.8% |
| other | 104 of 123 | 84.5% |
| Category | stores | share |
|---|---|---|
| beauty | 125 of 144 | 86.8% |
| electronics | 107 of 140 | 76.4% |
| fashion | 214 of 241 | 88.8% |
| food & beverage | 116 of 149 | 77.8% |
| home & furniture | 122 of 149 | 81.9% |
| other | 123 of 143 | 86.0% |
| Category | stores | share |
|---|---|---|
| beauty | 106 of 144 | 73.6% |
| electronics | 80 of 140 | 57.1% |
| fashion | 198 of 241 | 82.2% |
| food & beverage | 102 of 149 | 68.5% |
| home & furniture | 109 of 149 | 73.2% |
| other | 101 of 143 | 70.6% |
| Category | stores | share |
|---|---|---|
| beauty | 137 of 147 | 93.2% |
| electronics | 125 of 140 | 89.3% |
| fashion | 232 of 241 | 96.3% |
| food & beverage | 133 of 149 | 89.3% |
| home & furniture | 139 of 150 | 92.7% |
| other | 135 of 148 | 91.2% |
Mandatory caveats for the category tables translated rendering · sealed Hungarian original (authoritative) inside
Merchant-validity family (every category row) — translated rendering: all four caveats of this family are recurring — the measured bound from the adjudication, the named nested-price validator limit, the true complement and its closure, and denominator purity; see the English key.
English working translation — keyed to the sealed caveats: MEASURED BOUND (R1, from the adjudication): merchant_valid=TRUE stratum (C01, stratum size 626) of 5 examined rows 5 upheld = 0 false positives; in the presence-without-validity strata, of 10 examined rows 3 flipped to VALID (C02 2/5, stratum size 17; C03 1/5, stratum size 46), all three the same error pattern. Because of k=5 (n<20), ONLY a count, we do not publish a rate. • (a) NAMED BOUND, verified in code: product_rubric reads only the Offer/node-level price|lowPrice key (run_geo_audit.py:494-499), not offers[].priceSpecification[].price. Adjudicated quote (timeoutx.ae): '"offers":[{"@type":"Offer","priceSpecification":[{"@type":"UnitPriceSpecification","price":"200","priceCurrency":"AED"'. Because of this, the 626/807 is a LOWER BOUND. • TRUE COMPLEMENT: on F_prod 807-626 = 181 rows (22.43%) are not valid; this is exactly 118 (14.62%) without Product JSON-LD + 63 (7.81%) presence-without-validity, 118+63=181 closes. The 9.14% = 63/689 shown in the header stands on a DIFFERENT denominator (the base of those publishing JSON-LD); the two numbers are not interchangeable, and neither is the complement of merchant_valid. • DENOMINATOR PURITY (adjudicated, C04 stratum n=118): of 5 examined rows 2 were WooCommerce /shop/ ARCHIVE pages, which the weak cart_marker branch validated as product pages (ajmanfurniture.com, tealand.ae) — the F_prod=807 contains non-product pages, which can never be merchant_valid. Direction: removing them WOULD RAISE the rate, in the same direction as bound (a).
sealed audit record, Hungarian, verbatim: MÉRT KORLÁT (R1, az adjudikációból): merchant_valid=IGAZ réteg (C01, rétegméret 626) 5 vizsgált sorból 5 helybenhagyva = 0 hamis pozitív; a jelenlét-érvényesség-nélküli rétegekben 10 vizsgált sorból 3 fordult ÉRVÉNYESRE (C02 2/5, rétegméret 17; C03 1/5, rétegméret 46), mindhárom ugyanaz a hibaminta. k=5 (n<20) miatt CSAK darabszám, arányt nem közlünk. • (a) NEVESÍTETT KORLÁT, kódban igazolva: product_rubric csak Offer/node szintű price|lowPrice kulcsot olvas (run_geo_audit.py:494-499), az offers[].priceSpecification[].price-t nem. Adjudikált idézet (timeoutx.ae): '"offers":[{"@type":"Offer","priceSpecification":[{"@type":"UnitPriceSpecification","price":"200","priceCurrency":"AED"'. Emiatt a 626/807 ALSÓ KORLÁT. • IGAZ KOMPLEMENTER: F_prod-on 807-626 = 181 sor (22,43%) nem érvényes; ez pontosan 118 (14,62%) Product JSON-LD nélküli + 63 (7,81%) jelenlét-érvényesség-nélkül, 118+63=181 zár. A fejlécben szereplő 9,14% = 63/689 MÁS nevezőn (a JSON-LD-t közlők bázisa) áll; a két szám nem cserélhető fel, és egyik sem a merchant_valid komplementere. • NEVEZŐ-TISZTASÁG (adjudikált, C04 réteg n=118): 5 vizsgált sorból 2 WooCommerce /shop/ ARCHÍVUM oldal volt, amit a gyenge cart_marker-ág validált termékoldalként (ajmanfurniture.com, tealand.ae) — az F_prod=807 tartalmaz nem-termékoldalakat, amelyek sosem lehetnek merchant_valid. Irány: eltávolításuk EMELNÉ az arányt, azonos irányba az (a) korláttal.
Presence family — translated rendering: the sealed caveats of this family are already English and render verbatim below: the measured bound envelope from 83.3% (689 of 827) to 90.9% (709 of 780) around the engine rate, denominator contamination (which biases the rate down) and denominator under-inclusion, and the presence predicate itself taking no flip in adjudication.
sealed audit record, English in the sealed export, verbatim: R1 BOUND (verbatim, measured in THIS cell): engine 689/807 = 0.8538 on F_prod; adjudicated bound LOW 689/827 = 0.8331 (archives kept in the denominator, all 20 under-included rows assumed ldjson-less) .. HIGH 709/780 = 0.9090 (47 archives dropped, all 20 under-included rows assumed ldjson-bearing). • Denominator contamination (biases the rate DOWN): challenger C04 (ldjson=false stratum, n=118, k=5) flipped 2/5 - ajmanfurniture.com and tealand.ae are WooCommerce /shop/ ARCHIVES validated by a listing-grid cart_marker, not product detail pages. The ldjson=false verdict itself was CONFIRMED 5/5. Projection: 47 of 118. • Denominator under-inclusion: challenger C06 (no_product_page_found, n=50, k=5) flipped 2/5 - poolstore.ae and scarabme.com do have reachable product pages. ~20 rows are missing from F_prod and their ldjson status was never measured. One further C06 row (hobbycentre.ae) is unresolved_engine_kept. • The presence predicate itself took ZERO flips: 20 adjudicated F_prod rows (15 engine-TRUE via C01/C02/C03, 5 engine-FALSE via C04), 0 presence reversals. The 3 flips inside C02/C03 are merchant_valid-only and sit on rows already ldjson=true, so they cannot move this row.
Discovery family — translated rendering: sampled adjudication of the no-product-page-found stratum flipped 2 of 5 rows — stores that do have reachable product pages the discovery ladder did not walk — with one further row unresolved and kept as the engine read it; the page-unvalidated stratum was fully upheld. The engine number ships uncorrected (807 of 966): it is a lower estimate of the concept "has a discoverable product page" while staying exact on its own definition — the gap is between metric and concept, not in the counting. Only a sampled handful of the complement was individually adjudicated, so no corrected mix exists. Recurring: engine number never a corrected mix — see the English key.
English working translation — keyed to the sealed caveats: R1 MEASURED BOUND (verbatim): C06 no_product_page_found stratum (stratum size 50), k=5 -> 2 FLIPPED, 1 unresolved_engine_kept (hobbycentre.ae), 2 upheld. C05 page_unvalidated stratum (stratum size 87), k=5 -> 5/5 UPHELD, 0 flip. k=5 < 20, therefore only the count ships, the flip-RATE does not. • Evidence of the 2 flips: poolstore.ae 11 same-host /products/<slug>/<id> Zoho Commerce anchors on the homepage, while the engine's product_links_found=0; scarabme.com 2 server-rendered /product-details/product/<slug> anchors, alongside an empty sitemap urlset. Exactly the patterns of the named (b) limitation. • DIRECTION: the 807/966 is a LOWER ESTIMATE (a lower bound) for the concept „has a discoverable product page”. The metric is EXACT with respect to its OWN definition (the output of this discovery ladder + marker gate); the gap is between the metric and the concept, not in the counting. • R2 HONORED: the ENGINE number ships (807/966) without correction. The census exception does not apply: out of the 159-member complement only 10 rows were individually judged (C05 5 + C06 5) — there is no corrected mix.
sealed audit record, Hungarian, verbatim: R1 MÉRT KORLÁT (verbatim): C06 no_product_page_found réteg (rétegméret 50), k=5 -> 2 FLIPPED, 1 unresolved_engine_kept (hobbycentre.ae), 2 upheld. C05 page_unvalidated réteg (rétegméret 87), k=5 -> 5/5 UPHELD, 0 flip. k=5 < 20, ezért csak darabszám ship-el, flip-ARÁNY nem. • A 2 flip bizonyítéka: poolstore.ae 11 same-host /products/<slug>/<id> Zoho Commerce horgony a kezdőlapon, miközben az engine product_links_found=0; scarabme.com 2 szerver-renderelt /product-details/product/<slug> horgony, üres sitemap urlset mellett. Pontosan a megnevezett (b) limitáció mintázatai. • IRÁNY: a 807/966 ALSÓ BECSLÉS a „van felderíthető termékoldal” fogalomra. A metrika a SAJÁT definíciójára (e felderítő-létra + marker-kapu kimenete) EXAKT; a rés a metrika és a fogalom között van, nem a számolásban. • R2 BETARTVA: az ENGINE-szám ship-el (807/966) korrekció nélkül. Cenzus-kivétel nem alkalmazható: a 159 fős komplementerből csak 10 sor lett egyedileg bírálva (C05 5 + C06 5) — nincs javított keverék.
llms.txt family — translated rendering: binding no-correction rule: the engine number ships verbatim — 696 of 966 (72.0%) — never a corrected mix. The adjudication was a sample, not a census: 15 llms rows of the 966 were adjudicated; among sampled numerator members 1 of 10 was overturned — an error page served as llms.txt with a success status, a named and unfixed guard class the shipped number still contains — and of the sampled complement members 0 of 5 was overturned; both sample cells sit under twenty and ship as counts. The metric counts files with content: 5 further files downloaded empty, and the raw found count is not a metric and is never quoted as a headline. The true complement is 270 of 966 — 265 serving no file plus 5 empty; saying that 265 stores have no llms.txt would be a false complement.
English working translation — keyed to the sealed caveats: R2 BINDING: the engine number ships verbatim (696/966=0.7205), the corrected mix does NOT. The adjudication is a SAMPLE, not a census: of the 966, 15 llms rows were judged. Of the 10 judged members of the 696 numerator, 1 (easyshopping.ae, C08 seq38, opus, 2026-08-11) FLIPPED; of the 5 judged members of the 270 complement, 0. n=10 and n=5 are both <20 → ONLY COUNT, an error rate cannot be derived. • The mechanism of the flip in the anchored code: the soft-page guard of classify_llms_txt:384-385 filters ONLY the "<!doctype"/"<html" prefix. The easyshopping.ae body is a "<br /><b>Fatal error</b>" OpenCart PHP error page with HTTP 200, 1173 bytes → the guard does not fire, found=has_content=true. The HTTP-200 soft-error class is open, unfixed; the shipped number contains it. • FOUND vs HAS_CONTENT: the metric's definition is with-content. found=701/966, with-content=696/966; the gap is 5 domains (bluerhine.store, mythoby.com, rootzorganics.com, scarabme.com, thepartypopper.com) with a fetched but EMPTY llms.txt. The raw found count of 701 is NOT a metric in the engine (there is no aggregate field) — it cannot be quoted as a headline from this row. • COMPLEMENT DISCIPLINE: the true complement of the shipped metric is 270/966 (=966-696), NOT 265. The 270 = 265 no-file + 5 fetched-but-empty. The statement "265 stores have no llms.txt" would be a false complement.
sealed audit record, Hungarian, verbatim: R2-KÖTÉS: a motor-szám megy verbatim (696/966=0,7205), korrigált keverék NEM. Az adjudikáció MINTA, nem cenzus: a 966-ból 15 llms-sor lett elbírálva. A 696-os számláló 10 elbírált tagjából 1 (easyshopping.ae, C08 seq38, opus, 2026-08-11) MEGDŐLT; a 270-es komplementer 5 elbírált tagjából 0. n=10 és n=5 egyaránt <20 → CSAK DARABSZÁM, hibaarány nem vezethető le. • A megdőlés mechanizmusa a horgonyzott kódban: classify_llms_txt:384-385 soft-page őre CSAK a "<!doctype"/"<html" prefixet szűri. Az easyshopping.ae törzse "<br /><b>Fatal error</b>" OpenCart PHP-hibaoldal HTTP 200-zal, 1173 byte → az őr nem tüzel, found=has_content=true. A HTTP-200-as soft-error osztály nyitott, javítatlan; a shippelt szám ezt tartalmazza. • FOUND vs HAS_CONTENT: a metrika definíciója a with-content. found=701/966, with-content=696/966; a rés 5 domain (bluerhine.store, mythoby.com, rootzorganics.com, scarabme.com, thepartypopper.com) letöltött, de ÜRES llms.txt. A 701-es nyers found-szám a motorban NEM metrika (nincs aggregate-mező) — ebből a sorból nem idézhető headline-ként. • KOMPLEMENTER-FEGYELEM: a shippelt metrika igazi komplementere 270/966 (=966-696), NEM 265. A 270 = 265 nincs-fájl + 5 letöltött-de-üres. A "265 webshopnak nincs llms.txt-je" állítás hamis komplementer lenne.
Crawler-silence family — translated rendering: the primary rate stands on the engine's own conclusive robots frame; the census-frame secondary form, the directly measured explicit-decision count, the forbidden subtraction complement and the pass discipline are all in the robots entry of the English key; the engine-versus-artifact frame split is resolved by per-row frame naming — the counters agree everywhere, the divergence is purely denominator choice, not arithmetic.
English working translation — keyed to the sealed caveats: W1 ADJUDICATION (R5): PRIMARY is the 901/975 = 92.41% (F_robots_conclusive); SECONDARY is the 901/981 = 91.85% census form with a conclusive-frame caveat. Decisive argument: the 975 is not a correction AGAINST the engine - it is the engine's OWN frame (build_robots_matrix 1434-1460 narrows to present|absent_4xx|absent_softpage rows; matrix.frame_robots_conclusive=975). The den=981 exists only in the offline tier1. • FORBIDDEN COMPLEMENT STATEMENT: the form "901 silent, therefore the rest are explicit" is FALSE on the 981 frame - 981-901=80, but only 74 are explicit; 6 rows fall into neither category (unmeasured for all 7 agents). The 6: pinkavenue.store, kapsliving.com, gaminguae.ae, pcdubai.com, fmmdubai.com, daima.ae. This sentence cannot be published in any form. • On the 975 frame the complement is mathematically exact (901+74=975, verified in-cell), but it is still to be published as a MEASURED count (74/975), NEVER derived by subtraction from 981. This in itself is an argument for the conclusive frame. • ENGINE-vs-ARTIFACT CONTRADICTION RESOLVED by per-row frame naming: of the tier1's 15 rows, 13 use den=981 and 2 (blanket_disallow, sitemap_declared) use den=975; this row is one of the 11 verdict-derived ones. The robots_matrix publishes THE SAME numerators on 975. The numerators agree everywhere - the divergence is purely a denominator choice, not arithmetic (tier1_rate_arithmetic 15/15 PASS). • PASS-NAMING: the category breakdown uses the parent row's canonical robots-pass (geo_audit_robots_checkpoint-2026-08-08A.jsonl, last-record-wins, conclusive filter). Recounting the sites-pass gives 898/974 (a 9-domain symmetric difference) - the round names its pass, the two passes are never to be mixed.
sealed audit record, Hungarian, verbatim: W1 ADJUDIKÁCIÓ (R5): ELSŐDLEGES a 901/975 = 92,41% (F_robots_conclusive); MÁSODLAGOS a 901/981 = 91,85% census-alak conclusive-frame caveattel. Döntő érv: a 975 nem korrekció a motorral SZEMBEN - a motor SAJÁT kerete (build_robots_matrix 1434-1460 present|absent_4xx|absent_softpage sorokra szűkít; matrix.frame_robots_conclusive=975). A den=981 csak az offline tier1-ben létezik. • TILTOTT KOMPLEMENTER-ÁLLÍTÁS: a "901 néma, tehát a többi explicit" forma a 981-es kereten HAMIS - 981-901=80, de csak 74 explicit; 6 sor egyik kategóriába sem esik (mind a 7 ügynökre unmeasured). A 6: pinkavenue.store, kapsliving.com, gaminguae.ae, pcdubai.com, fmmdubai.com, daima.ae. Ez a mondat semmilyen alakban nem publikálható. • A 975-ös kereten a komplementer matematikailag pontos (901+74=975, in-cell igazolva), de akkor is MÉRT darabszámként (74/975) közlendő, SOHA nem 981-ből kivonással levezetve. Ez önmagában is a konklúzív keret melletti érv. • MOTOR-vs-ARTEFAKTUM ELLENTMONDÁS FELOLDVA soronkénti keret-megnevezéssel: a tier1 15 sorából 13 den=981-et, 2 (blanket_disallow, sitemap_declared) den=975-öt használ; ez a sor a 11 verdikt-derivált egyike. A robots_matrix UGYANEZEKET a számlálókat 975-ön közli. A számlálók mindenhol egyeznek - a szakadás tisztán nevező-választás, nem aritmetika (tier1_rate_arithmetic 15/15 PASS). • PASS-NEVESÍTÉS: a kategóriabontás a szülő-sor kanonikus robots-pass-át használja (geo_audit_robots_checkpoint-2026-08-08A.jsonl, last-record-wins, conclusive-szűrés). A sites-pass újraszámolása 898/974-et ad (9 domain szimmetrikus differencia) - a kör nevesíti a pass-át, a két pass sosem keverendő.
Field invalidity on the conflated frame — a separate block, upper bounds
These readings stand on the full parsed-product frame and include every page that publishes no Product JSON-LD at all — upper bounds by construction. The next block re-states the same five fields on the clean frame; the two frames must never be swapped, and no cell here may be combined with one from the next block. Denominator for every row: 807 parsed product pages (frame F_prod).
| Field | pages | share |
|---|---|---|
| name | 119 of 807 | 14.8% |
| image | 123 of 807 | 15.2% |
| price | 175 of 807 | 21.7% |
| priceCurrency | 174 of 807 | 21.6% |
| availability | 143 of 807 | 17.7% |
Mandatory caveats for the conflated-frame field table translated rendering · sealed Hungarian original (authoritative) inside
Every row of this table — translated rendering: in the full adjudication run 10 of 60 rows flipped, 49 were upheld and 1 stayed unresolved and kept as the engine read it; every product-field flip was exactly the nested price-specification form. These rows stand on the conflated frame and are upper bounds on true field invalidity. Recurring: no aggregate home (prereg seam) · conflated frame · named nested-price validator limit — see the English key.
English working translation — keyed to the sealed caveats: W5: there is no aggregate field. The prereg rider_2 'stated FROM THE AGGREGATE' requirement therefore can NOT be met - the pf_* breakdown lives only in the flat CSV / sites JSONL. Prereg-vs-artifact seam, named; hence a PASS_WITH_CAVEAT ceiling. • CONFLATION (the most serious reading trap): on F_prod the False counts also include the 118 rows WITHOUT Product JSON-LD - on an empty nodes list product_rubric leaves all 5 fields at False (measured: 0 no-ldjson rows where any field is True). Real field invalidity on F_prod∩ldjson (den=689): name 1, image 5, price 57, priceCurrency 56, availability 25. • NESTED-PRICE BOUND (mandatory, explicit): the price/priceCurrency branch only walks the _offers_list(n)+[n] level (lines 494-499); it does NOT read the offers.priceSpecification / UnitPriceSpecification form. The 175 and the 174 are therefore an UPPER bound, not a point estimate. • MEASURED BOUND (R1, verbatim): geo-audit-challenger-flips 60 adjudicated rows - upheld 49 / flipped 10 / unresolved_engine_kept 1. Product-field strata: C02 (shopify presence-without-validity, stratum_n=17, k=5) 2 flips; C03 (non-shopify presence-without-validity, stratum_n=46, k=5) 1 flip. All three flips are EXACTLY the nested priceSpecification form; 3/10 sampled rows. • CONFLATED FRAME: this row is to be read on the F_prod frame and also contains the rows without Product JSON-LD (see the inherited CONFLATION caveat); its direction is an UPPER bound on real field invalidity.
sealed audit record, Hungarian, verbatim: W5: nincs aggregátum-mező. A prereg rider_2 'stated FROM THE AGGREGATE' előírása így NEM teljesíthető - a pf_* bontás csak a flat CSV-ben / sites JSONL-ben él. Prereg-vs-artefaktum varrat, nevesítve; ezért PASS_WITH_CAVEAT plafon. • KONFLÁCIÓ (a legsúlyosabb olvasati csapda): F_prod-on a false-számok magukban foglalják a 118 Product JSON-LD NÉLKÜLI sort is - a product_rubric üres nodes-listán mind az 5 mezőt False-on hagyja (mérve: 0 olyan no-ldjson sor, ahol bármely mező True). Valódi mező-érvénytelenség F_prod∩ldjson-on (den=689): name 1, image 5, price 57, priceCurrency 56, availability 25. • BEÁGYAZOTT-ÁR KORLÁT (kötelező, explicit): a price/priceCurrency ág csak az _offers_list(n)+[n] szintet járja be (494-499. sor); az offers.priceSpecification / UnitPriceSpecification formát NEM olvassa. A 175 és a 174 ezért FELSŐ korlát, nem pontbecslés. • MÉRT KORLÁT (R1, verbatim): geo-audit-challenger-flips 60 adjudikált sor - upheld 49 / flipped 10 / unresolved_engine_kept 1. Termék-mező-strátumok: C02 (shopify presence-without-validity, stratum_n=17, k=5) 2 flip; C03 (nem-shopify presence-without-validity, stratum_n=46, k=5) 1 flip. Mindhárom flip PONTOSAN a beágyazott priceSpecification-forma; 3/10 mintázott sor. • KONFLÁLT KERET: ez a sor az F_prod kereten értendő és a Product JSON-LD nélküli sorokat is tartalmazza (lásd az örökölt KONFLÁCIÓ-caveatet); iránya FELSŐ korlát a valódi mező-érvénytelenségre.
Field invalidity among pages that do publish Product JSON-LD — a separate block
The clean frame: only pages that carry a Product JSON-LD block. Denominator for every row: 689 pages (frame F_prod ∩ ldjson_present (n=689)). The name and image cells sit under the pre-registered floor and ship as counts.
| Field | pages | share |
|---|---|---|
| name | 1 of 689 | n<20 - count only |
| image | 5 of 689 | n<20 - count only |
| price | 57 of 689 | 8.3% |
| priceCurrency | 56 of 689 | 8.1% |
| availability | 25 of 689 | 3.6% |
Mandatory caveats for the clean-frame field table translated rendering · sealed Hungarian original (authoritative) inside
Every row of this table — translated rendering: this block stands on the clean intersection — pages that do publish Product JSON-LD — which resolves the conflation trap; the price-family cells remain upper bounds under the named nested-price validator limit, with the measured bound from the full adjudication run as in the conflated table above. Recurring: no aggregate home (prereg seam) · conflated frame (resolved here) · named validator limit · measured bound — see the English key.
English working translation — keyed to the sealed caveats: W5: there is no aggregate field. The prereg rider_2 'stated FROM THE AGGREGATE' requirement therefore can NOT be met - the pf_* breakdown lives only in the flat CSV / sites JSONL. Prereg-vs-artifact seam, named; hence a PASS_WITH_CAVEAT ceiling. • CONFLATION (the most serious reading trap): on F_prod the False counts also include the 118 rows WITHOUT Product JSON-LD - on an empty nodes list product_rubric leaves all 5 fields at False (measured: 0 no-ldjson rows where any field is True). Real field invalidity on F_prod∩ldjson (den=689): name 1, image 5, price 57, priceCurrency 56, availability 25. • NESTED-PRICE BOUND (mandatory, explicit): the price/priceCurrency branch only walks the _offers_list(n)+[n] level (lines 494-499); it does NOT read the offers.priceSpecification / UnitPriceSpecification form. The 175 and the 174 are therefore an UPPER bound, not a point estimate. • MEASURED BOUND (R1, verbatim): geo-audit-challenger-flips 60 adjudicated rows - upheld 49 / flipped 10 / unresolved_engine_kept 1. Product-field strata: C02 (shopify presence-without-validity, stratum_n=17, k=5) 2 flips; C03 (non-shopify presence-without-validity, stratum_n=46, k=5) 1 flip. All three flips are EXACTLY the nested priceSpecification form; 3/10 sampled rows. • TRUE FIELD-INVALIDITY FRAME: this row is to be read on the F_prod ∩ ldjson_present frame (the resolution of the CONFLATION trap) - the rows without Product JSON-LD are excluded.
sealed audit record, Hungarian, verbatim: W5: nincs aggregátum-mező. A prereg rider_2 'stated FROM THE AGGREGATE' előírása így NEM teljesíthető - a pf_* bontás csak a flat CSV-ben / sites JSONL-ben él. Prereg-vs-artefaktum varrat, nevesítve; ezért PASS_WITH_CAVEAT plafon. • KONFLÁCIÓ (a legsúlyosabb olvasati csapda): F_prod-on a false-számok magukban foglalják a 118 Product JSON-LD NÉLKÜLI sort is - a product_rubric üres nodes-listán mind az 5 mezőt False-on hagyja (mérve: 0 olyan no-ldjson sor, ahol bármely mező True). Valódi mező-érvénytelenség F_prod∩ldjson-on (den=689): name 1, image 5, price 57, priceCurrency 56, availability 25. • BEÁGYAZOTT-ÁR KORLÁT (kötelező, explicit): a price/priceCurrency ág csak az _offers_list(n)+[n] szintet járja be (494-499. sor); az offers.priceSpecification / UnitPriceSpecification formát NEM olvassa. A 175 és a 174 ezért FELSŐ korlát, nem pontbecslés. • MÉRT KORLÁT (R1, verbatim): geo-audit-challenger-flips 60 adjudikált sor - upheld 49 / flipped 10 / unresolved_engine_kept 1. Termék-mező-strátumok: C02 (shopify presence-without-validity, stratum_n=17, k=5) 2 flip; C03 (nem-shopify presence-without-validity, stratum_n=46, k=5) 1 flip. Mindhárom flip PONTOSAN a beágyazott priceSpecification-forma; 3/10 mintázott sor. • VALÓDI MEZŐ-ÉRVÉNYTELENSÉG-KERET: ez a sor az F_prod ∩ ldjson_present kereten értendő (a KONFLÁCIÓ-csapda feloldása) - a Product JSON-LD nélküli sorok kizárva.
The invalid complement, decomposed — and the failure vectors
The two decomposition rows partition the invalid complement exactly, on the same denominator — the same closure the lead section states. The vector table below ships as counts by pre-registered rule, never rates.
| Mode | pages | share |
|---|---|---|
| no Product JSON-LD at all | 118 of 181 | 65.2% |
| block present, merchant-invalid | 63 of 181 | 34.8% |
| Failing field set | pages | note |
|---|---|---|
| price and priceCurrency | 33 of 63 | count only - family rule |
| price, priceCurrency and availability | 22 of 63 | count only - family rule |
| image | 4 of 63 | count only - family rule |
| availability | 2 of 63 | count only - family rule |
| all five fields | 1 of 63 | count only - family rule |
| price | 1 of 63 | count only - family rule |
Mandatory caveats for the decomposition and vector tables translated rendering · sealed Hungarian original (authoritative) inside
Both decomposition rows — translated rendering: the two rows partition the parent table's invalid complement exactly — every page with no Product JSON-LD is merchant-invalid by construction (measured: no such row has any field true). Recurring: measured bound · named validator limit · true complement · denominator purity — see the English key.
English working translation — keyed to the sealed caveats: MEASURED BOUND (R1, from the adjudication): merchant_valid=TRUE stratum (C01, stratum size 626) of 5 examined rows 5 upheld = 0 false positives; in the presence-without-validity strata, of 10 examined rows 3 flipped to VALID (C02 2/5, stratum size 17; C03 1/5, stratum size 46), all three the same error pattern. Because of k=5 (n<20), ONLY a count, we do not publish a rate. • (a) NAMED BOUND, verified in code: product_rubric reads only the Offer/node-level price|lowPrice key (run_geo_audit.py:494-499), not offers[].priceSpecification[].price. Adjudicated quote (timeoutx.ae): '"offers":[{"@type":"Offer","priceSpecification":[{"@type":"UnitPriceSpecification","price":"200","priceCurrency":"AED"'. Because of this, the 626/807 is a LOWER BOUND. • TRUE COMPLEMENT: on F_prod 807-626 = 181 rows (22.43%) are not valid; this is exactly 118 (14.62%) without Product JSON-LD + 63 (7.81%) presence-without-validity, 118+63=181 closes. The 9.14% = 63/689 shown in the header stands on a DIFFERENT denominator (the base of those publishing JSON-LD); the two numbers are not interchangeable, and neither is the complement of merchant_valid. • DENOMINATOR PURITY (adjudicated, C04 stratum n=118): of 5 examined rows 2 were WooCommerce /shop/ ARCHIVE pages, which the weak cart_marker branch validated as product pages (ajmanfurniture.com, tealand.ae) — the F_prod=807 contains non-product pages, which can never be merchant_valid. Direction: removing them WOULD RAISE the rate, in the same direction as bound (a). • DECOMPOSITION ANCHOR: no_ldjson + ldjson_invalid == the parent table's merchant_invalid complement (F_prod den - merchant_valid num); every F_prod row without Product JSON-LD is merchant_invalid by construction (measured: 0 no-ldjson rows where any field is True).
sealed audit record, Hungarian, verbatim: MÉRT KORLÁT (R1, az adjudikációból): merchant_valid=IGAZ réteg (C01, rétegméret 626) 5 vizsgált sorból 5 helybenhagyva = 0 hamis pozitív; a jelenlét-érvényesség-nélküli rétegekben 10 vizsgált sorból 3 fordult ÉRVÉNYESRE (C02 2/5, rétegméret 17; C03 1/5, rétegméret 46), mindhárom ugyanaz a hibaminta. k=5 (n<20) miatt CSAK darabszám, arányt nem közlünk. • (a) NEVESÍTETT KORLÁT, kódban igazolva: product_rubric csak Offer/node szintű price|lowPrice kulcsot olvas (run_geo_audit.py:494-499), az offers[].priceSpecification[].price-t nem. Adjudikált idézet (timeoutx.ae): '"offers":[{"@type":"Offer","priceSpecification":[{"@type":"UnitPriceSpecification","price":"200","priceCurrency":"AED"'. Emiatt a 626/807 ALSÓ KORLÁT. • IGAZ KOMPLEMENTER: F_prod-on 807-626 = 181 sor (22,43%) nem érvényes; ez pontosan 118 (14,62%) Product JSON-LD nélküli + 63 (7,81%) jelenlét-érvényesség-nélkül, 118+63=181 zár. A fejlécben szereplő 9,14% = 63/689 MÁS nevezőn (a JSON-LD-t közlők bázisa) áll; a két szám nem cserélhető fel, és egyik sem a merchant_valid komplementere. • NEVEZŐ-TISZTASÁG (adjudikált, C04 réteg n=118): 5 vizsgált sorból 2 WooCommerce /shop/ ARCHÍVUM oldal volt, amit a gyenge cart_marker-ág validált termékoldalként (ajmanfurniture.com, tealand.ae) — az F_prod=807 tartalmaz nem-termékoldalakat, amelyek sosem lehetnek merchant_valid. Irány: eltávolításuk EMELNÉ az arányt, azonos irányba az (a) korláttal. • DEKOMPOZÍCIÓ-HORGONY: no_ldjson + ldjson_invalid == a szülő-tábla merchant_invalid komplemense (F_prod den - merchant_valid num); minden Product JSON-LD nélküli F_prod sor konstrukcióból merchant_invalid (mérve: 0 olyan no-ldjson sor, ahol bármely mező True).
Every vector row — translated rendering: a pattern is the underscore-joined list of failing fields in the fixed field order; every pattern carrying price or priceCurrency is an upper bound under the named nested-price validator limit. Recurring: no aggregate home · conflated frame · named validator limit · measured bound — see the English key.
English working translation — keyed to the sealed caveats: W5: there is no aggregate field. The prereg rider_2 'stated FROM THE AGGREGATE' requirement therefore can NOT be met - the pf_* breakdown lives only in the flat CSV / sites JSONL. Prereg-vs-artifact seam, named; hence a PASS_WITH_CAVEAT ceiling. • CONFLATION (the most serious reading trap): on F_prod the False counts also include the 118 rows WITHOUT Product JSON-LD - on an empty nodes list product_rubric leaves all 5 fields at False (measured: 0 no-ldjson rows where any field is True). Real field invalidity on F_prod∩ldjson (den=689): name 1, image 5, price 57, priceCurrency 56, availability 25. • NESTED-PRICE BOUND (mandatory, explicit): the price/priceCurrency branch only walks the _offers_list(n)+[n] level (lines 494-499); it does NOT read the offers.priceSpecification / UnitPriceSpecification form. The 175 and the 174 are therefore an UPPER bound, not a point estimate. • MEASURED BOUND (R1, verbatim): geo-audit-challenger-flips 60 adjudicated rows - upheld 49 / flipped 10 / unresolved_engine_kept 1. Product-field strata: C02 (shopify presence-without-validity, stratum_n=17, k=5) 2 flips; C03 (non-shopify presence-without-validity, stratum_n=46, k=5) 1 flip. All three flips are EXACTLY the nested priceSpecification form; 3/10 sampled rows. • VECTOR DISTRIBUTION: the pattern is the underscore-joined list of the failed fields (field order: name/image/price/priceCurrency/availability); the patterns carrying price/priceCurrency are UPPER bounds because of the NESTED-PRICE bound.
sealed audit record, Hungarian, verbatim: W5: nincs aggregátum-mező. A prereg rider_2 'stated FROM THE AGGREGATE' előírása így NEM teljesíthető - a pf_* bontás csak a flat CSV-ben / sites JSONL-ben él. Prereg-vs-artefaktum varrat, nevesítve; ezért PASS_WITH_CAVEAT plafon. • KONFLÁCIÓ (a legsúlyosabb olvasati csapda): F_prod-on a false-számok magukban foglalják a 118 Product JSON-LD NÉLKÜLI sort is - a product_rubric üres nodes-listán mind az 5 mezőt False-on hagyja (mérve: 0 olyan no-ldjson sor, ahol bármely mező True). Valódi mező-érvénytelenség F_prod∩ldjson-on (den=689): name 1, image 5, price 57, priceCurrency 56, availability 25. • BEÁGYAZOTT-ÁR KORLÁT (kötelező, explicit): a price/priceCurrency ág csak az _offers_list(n)+[n] szintet járja be (494-499. sor); az offers.priceSpecification / UnitPriceSpecification formát NEM olvassa. A 175 és a 174 ezért FELSŐ korlát, nem pontbecslés. • MÉRT KORLÁT (R1, verbatim): geo-audit-challenger-flips 60 adjudikált sor - upheld 49 / flipped 10 / unresolved_engine_kept 1. Termék-mező-strátumok: C02 (shopify presence-without-validity, stratum_n=17, k=5) 2 flip; C03 (nem-shopify presence-without-validity, stratum_n=46, k=5) 1 flip. Mindhárom flip PONTOSAN a beágyazott priceSpecification-forma; 3/10 mintázott sor. • VEKTOR-ELOSZLÁS: a mintázat a bukott mezők aláhúzással fűzött listája (mező-sorrend: name/image/price/priceCurrency/availability); a price/priceCurrency-t hordozó mintázatok a BEÁGYAZOTT-ÁR korlát miatt FELSŐ korlátok.
Funnel failure buckets — the census, bucket by bucket
The measurable row is by construction the same numerator the funnel at the top of the page stands on; the buckets below it are the named drop-outs, honest zeros included. One census row carries a reproducible engine-side fetch-persistence failure and stays in its bucket. Denominator for every row: 981 census stores.
| Bucket | stores | share |
|---|---|---|
| measurable | 966 of 981 | 98.5% |
| robots disallow | 0 of 981 | n<20 - count only |
| excessive crawl delay | 0 of 981 | n<20 - count only |
| client error (forbidden) | 7 of 981 | n<20 - count only |
| client error (rate limited) | 1 of 981 | n<20 - count only |
| dead domain | 1 of 981 | n<20 - count only |
| timeout | 2 of 981 | n<20 - count only |
| fetch error | 4 of 981 | n<20 - count only |
Mandatory caveats for the funnel-bucket table translated rendering · sealed Hungarian original (authoritative) inside
The measurable row — translated rendering: the first four sealed caveats are already English and render verbatim below: the two-route census verification with zero disagreements; the sampled re-adjudication of the non-clean buckets in which 1 of 5 rows flipped — the flip being an engine-side post-fetch data loss, not an unreachable store — while the rest of that stratum was never individually adjudicated; the engine-number rule — the shipped figures are the engine numbers, no corrected mix: under the flip the measurable count would be at most 967, a bound stated only here and never shipped; and the instrument exception — exactly 1 of 981 census rows carries a reproducible engine-side fetch-persistence failure and stays in its bucket. Every legal bucket ships, the empty ones as explicit honest zeros — the aggregate's sparse map omits them; and this row's numerator equals the parent funnel's, included for partition completeness, no new information.
English working translation — keyed to the sealed caveats: [caveat 1 is English in the sealed record] • [caveat 2 is English in the sealed record] • [caveat 3 is English in the sealed record] • [caveat 4 is English in the sealed record] • FULL BUCKET CONTRACT: all 8 legal buckets ship, the zero-cardinality ones with an explicit zero (honest zero) - the aggregate's sparse map omits the zeros. • DUPLICATION NOTE: the numerator of AD-funnel-ok is identical to the numerator of the parent S09-funnel; it is included for the sake of partition completeness, it carries no new information.
sealed audit record (first four caveats English, last two Hungarian), verbatim: P1 CITATION: V-FUNNEL-G (data/uae/meta/geo-audit-verification-2026-08-08A.jsonl) verified this row two-route (Route A sites JSONL, Route B aggregate CSV): match=true, mismatches=[], stop=false; 0 per-domain bucket disagreements across all 981 domains; orphans 0; duplicate domains 0; buckets_sum_vs_attempted 981=981 on BOTH routes. • R1 BOUND (verbatim, flips C07): stratum_n=15 (the non-ok bucket rows), k=5 -> 4 upheld, 1 flipped. Flip = misnad.com: 'NOT fetch_error - the homepage was fetched and fully persisted; engine-side post-fetch data loss, row requires re-measurement rather than counting as unreachable'. SAMPLED not census: 10 of the 15 were never individually adjudicated. • R2 HONORED: the shipped figures are the ENGINE numbers (F_home 966, F_prod 807, engine rate 0.8354) - no corrected mix. Under the C07 flip F_home would be at most 967 (+1); that adjusted figure is NOT shipped, only stated here as the bound. The census-grade exception does not apply (C07 was k=5 of 15, not fully adjudicated). • INSTRUMENT EXCEPTION (engine backlog), census-verified here: exactly 1 of 981 rows has bucket_detail 'worker_error:ResponseNotRead' with fetches=[] AND evidence=[] - misnad.com. Its checked_at 2026-08-11T23:29:31Z is the latest of all 981 and equals the aggregate generated_at, so the failure persisted at the round's final attempt, after the 19:50Z flip on a 19:13:42Z record: reproducible. • TELJES BUCKET-KONTRAKTUS: mind a 8 legális bucket kimegy, a nulla-számosságúak explicit nullával (honest zero) - az aggregátum sparse map-je a nullákat elhagyja. • DUPLIKÁCIÓ-JEGYZET: az AD-funnel-ok számlálója azonos a szülő S09-funnel numerátorával; a partíció-teljesség kedvéért szerepel, új információt nem hordoz.
Every failure-bucket row — translated rendering: the same two-route verification, sampled bound, engine-number rule, instrument exception and full-bucket contract as the measurable row above, without the duplication note.
English working translation — keyed to the sealed caveats: [caveat 1 is English in the sealed record] • [caveat 2 is English in the sealed record] • [caveat 3 is English in the sealed record] • [caveat 4 is English in the sealed record] • FULL BUCKET CONTRACT: all 8 legal buckets ship, the zero-cardinality ones with an explicit zero (honest zero) - the aggregate's sparse map omits the zeros.
sealed audit record (first four caveats English, last one Hungarian), verbatim: P1 CITATION: V-FUNNEL-G (data/uae/meta/geo-audit-verification-2026-08-08A.jsonl) verified this row two-route (Route A sites JSONL, Route B aggregate CSV): match=true, mismatches=[], stop=false; 0 per-domain bucket disagreements across all 981 domains; orphans 0; duplicate domains 0; buckets_sum_vs_attempted 981=981 on BOTH routes. • R1 BOUND (verbatim, flips C07): stratum_n=15 (the non-ok bucket rows), k=5 -> 4 upheld, 1 flipped. Flip = misnad.com: 'NOT fetch_error - the homepage was fetched and fully persisted; engine-side post-fetch data loss, row requires re-measurement rather than counting as unreachable'. SAMPLED not census: 10 of the 15 were never individually adjudicated. • R2 HONORED: the shipped figures are the ENGINE numbers (F_home 966, F_prod 807, engine rate 0.8354) - no corrected mix. Under the C07 flip F_home would be at most 967 (+1); that adjusted figure is NOT shipped, only stated here as the bound. The census-grade exception does not apply (C07 was k=5 of 15, not fully adjudicated). • INSTRUMENT EXCEPTION (engine backlog), census-verified here: exactly 1 of 981 rows has bucket_detail 'worker_error:ResponseNotRead' with fetches=[] AND evidence=[] - misnad.com. Its checked_at 2026-08-11T23:29:31Z is the latest of all 981 and equals the aggregate generated_at, so the failure persisted at the round's final attempt, after the 19:50Z flip on a 19:13:42Z record: reproducible. • TELJES BUCKET-KONTRAKTUS: mind a 8 legális bucket kimegy, a nulla-számosságúak explicit nullával (honest zero) - az aggregátum sparse map-je a nullákat elhagyja.
One further family in this round's gate ledger stayed held back — the quality-depth measurement — and later shipped as its own sealed round: Part Two, below.
Part Two - quality depth, measured in a new window (round 2026-08-13Q)
Every figure below is a new-window reading on its own frame; nothing here corrects or restates a sealed Part One figure.
Small cells. A rate ships only when numerator and denominator both clear the pre-registered floor; otherwise a count with its named denominator — and the named count-families ship as counts regardless of size. English working translation, keyed: a rate ships only if num >= 20 AND den >= 20; otherwise only a count with the named denominator | Q-round count families (uniformly, regardless of size): Q-W1-desclen-*, Q-W2-residue-*, Q-W2-challenger-strip, Q-W3-zero-eligible-FQ, Q-W3-refuter-error, Q-W4-prov-*, Q-W4-blind-audit, Q-drift-* (with the exception of the ok-bucket) · sealed small-cell rule, Hungarian, verbatim: rata csak akkor megy ki, ha num >= 20 ES den >= 20; kulonben csak darabszam a megnevezett nevezovel | Q-kor darabszam-csaladok (egysegesen, merettol fuggetlenul): Q-W1-desclen-*, Q-W2-residue-*, Q-W2-challenger-strip, Q-W3-zero-eligible-FQ, Q-W3-refuter-error, Q-W4-prov-*, Q-W4-blind-audit, Q-drift-* (az ok-bucket kivetelevel)
Direction. Every direction cell renders the row's sealed direction in English through a closed, build-enforced mapping: lower bound is the sealed alsó korlát
, upper bound the sealed felső korlát
, and point estimate with measured error the sealed pontszerű becslés mért hibával
. An unmapped sealed direction aborts the build; the mapping cannot silently widen.
This round's own audit record. Every judged instrument ships beside its own measured error; the seeded challenger strips were never folded into a published rate. English working translation, keyed: Q-strip: W2 challenger 40/48 flips, W3 blind refuter exact 76/80, W4 blind audit B1 0/10 and B2 19/20 - every judged row ships with its own measured error · sealed audit record, Hungarian, verbatim: Q-strip: W2 challenger 40/48 megdontes, W3 vak refuter egzakt 76/80, W4 vak audit B1 0/10 es B2 19/20 - minden itelt sor a sajat mert hibajaval megy ki
New window, never a correction. Every reading was measured in this round's own fetch window (2026-08-12 (UTC)); the sealed Part One rates stand untouched.
The frame: what survived the re-fetch
Drift shows up at the fetch layer — a page either fetched clean or dropped into a named bucket below, each drop-out real merchant-side change or refusal, not measurement error. The new-window frame is not a random subsample of the sealed frame — direction and mandatory caveats travel with every row.
| Bucket | pages | share | direction |
|---|---|---|---|
| fetched clean | 798 of 807 | 98.9% | point estimate with measured error |
| gone | 0 of 807 | count only - family rule | point estimate with measured error |
| robots disallow | 0 of 807 | count only - family rule | point estimate with measured error |
| excessive crawl delay | 0 of 807 | count only - family rule | point estimate with measured error |
| client error (forbidden) | 0 of 807 | count only - family rule | point estimate with measured error |
| client error (rate limited) | 0 of 807 | count only - family rule | point estimate with measured error |
| dead domain | 0 of 807 | count only - family rule | point estimate with measured error |
| timeout | 8 of 807 | count only - family rule | point estimate with measured error |
| fetch error | 1 of 807 | count only - family rule | point estimate with measured error |
Mandatory caveats for the drift table translated rendering · sealed Hungarian original (authoritative) inside
The clean-fetch row and the non-retry buckets — translated rendering: drift is real change: these rows are the re-fetch outcome of the sealed input frame, and the non-clean buckets are merchant-side change or refusal, not measurement error — the nine buckets partition the input frame exactly — their sum is exactly the denominator. Recurring: new measurement window · two routes — see the English key.
English working translation — keyed to the sealed caveats: DRIFT = REAL CHANGE: this row is the output of the RE-FETCH of the sealed input frame. The non-ok buckets are not measurement errors but changes or refusals that occurred on the merchant's side; the nine buckets are a PARTITION of the input frame, and their sum is exactly the denominator. • NEW MEASUREMENT WINDOW: this row stands on the pages RE-fetched in the 2026-08-13Q round (F_q). The difference between the 08A and the Q window is REAL-WORLD CHANGE (the page was transformed, ceased to exist or now refuses the request), not a measurement error; the full breakdown ships in the Q-drift-* rows, nothing is hidden. • TWO ROUTES: every published cell was produced by two independent routes (record-level, and a derivation re-extracted from the stored HTML), and the two must match exactly, cell by cell; on a mismatch the round writes out a symmetric difference and does NOT write a table.
sealed audit record, Hungarian, verbatim: DRIFT = VALÓS VÁLTOZÁS: ez a sor a lepecsételt bemeneti keret ÚJRALETÖLTÉSÉNEK kimenete. A nem-ok bucketek nem mérési hibák, hanem a kereskedő oldalán történt változások vagy elutasítások; a kilenc bucket a bemeneti keret PARTÍCIÓJA, összegük pontosan a nevező. • ÚJ MÉRÉSI ABLAK: ez a sor a 2026-08-13Q körben ÚJRA letöltött oldalakon (F_q) áll. A 08A és a Q ablak közti eltérés VALÓS VILÁGBELI VÁLTOZÁS (az oldal átalakult, megszűnt vagy most utasítja el a kérést), nem mérési hiba; a teljes bontás a Q-drift-* sorokban megy ki, elrejtve semmi nincs. • KÉT ÚT: minden publikált cella két független úton készült (rekord-szintű, illetve a tárolt HTML-ből újra-kinyert deriváció), és a kettőnek cellára pontosan egyeznie kell; eltérés esetén a kör kiír egy szimmetrikus differenciát és NEM ír táblát.
The retry buckets (forbidden, rate limited, timeout) — translated rendering: every caveat above applies to these rows too, and additionally a self-measurement class: these buckets are partly properties of our own request — headers, egress address, pacing — not of the shop; they ship as counts, and no exclusion-type claim may be built on them alone.
English working translation — keyed to the sealed caveats: DRIFT = REAL CHANGE: this row is the output of the RE-FETCH of the sealed input frame. The non-ok buckets are not measurement errors but changes or refusals that occurred on the merchant's side; the nine buckets are a PARTITION of the input frame, and their sum is exactly the denominator. • NEW MEASUREMENT WINDOW: this row stands on the pages RE-fetched in the 2026-08-13Q round (F_q). The difference between the 08A and the Q window is REAL-WORLD CHANGE (the page was transformed, ceased to exist or now refuses the request), not a measurement error; the full breakdown ships in the Q-drift-* rows, nothing is hidden. • SELF-MEASUREMENT CLASS: the 403/429/timeout buckets are partly properties of OUR request (header, outbound IP, pacing), not of the store; therefore these rows ship as counts and no exclusion-type claim can be built on them on their own. • TWO ROUTES: every published cell was produced by two independent routes (record-level, and a derivation re-extracted from the stored HTML), and the two must match exactly, cell by cell; on a mismatch the round writes out a symmetric difference and does NOT write a table.
sealed audit record, Hungarian, verbatim: DRIFT = VALÓS VÁLTOZÁS: ez a sor a lepecsételt bemeneti keret ÚJRALETÖLTÉSÉNEK kimenete. A nem-ok bucketek nem mérési hibák, hanem a kereskedő oldalán történt változások vagy elutasítások; a kilenc bucket a bemeneti keret PARTÍCIÓJA, összegük pontosan a nevező. • ÚJ MÉRÉSI ABLAK: ez a sor a 2026-08-13Q körben ÚJRA letöltött oldalakon (F_q) áll. A 08A és a Q ablak közti eltérés VALÓS VILÁGBELI VÁLTOZÁS (az oldal átalakult, megszűnt vagy most utasítja el a kérést), nem mérési hiba; a teljes bontás a Q-drift-* sorokban megy ki, elrejtve semmi nincs. • ÖNMÉRÉS-OSZTÁLY: a 403/429/timeout bucketek részben a MI kérésünk tulajdonságai (fejléc, kimenő IP, ütem), nem a webshopé; ezért ezek a sorok darabszámként mennek ki és semmilyen kizárás-típusú állítás nem építhető rájuk önmagukban. • KÉT ÚT: minden publikált cella két független úton készült (rekord-szintű, illetve a tárolt HTML-ből újra-kinyert deriváció), és a kettőnek cellára pontosan egyeznie kell; eltérés esetén a kör kiír egy szimmetrikus differenciát és NEM ír táblát.
The named limitation, measured directly — on the frozen Part One corpus
Part One's lead shipped as a lower bound because its validator could not read one legal nested price form. The extended validator reads it. Run over the very same frozen bytes, it measures that named limitation directly — a new instrument reading on the frozen corpus, never a correction; the sealed Part One rate stands. The instrument difference is, by construction, exactly the nested-price pattern the lower-bound label named. Denominator for every row: 807 frozen pages (frame FPC (frozen 2026-08-08A product.html corpus, n=807)). The field rows include pages with no Product JSON-LD at all — upper bounds on the conflated frozen frame, separate by rule from every clean-frame block below.
| Reading | pages | share | direction |
|---|---|---|---|
| merchant-valid under the extended validator | 655 of 807 | 81.2% | lower bound |
| name fails machine-readability | 120 of 807 | 14.9% | upper bound |
| image fails | 124 of 807 | 15.4% | upper bound |
| price fails | 138 of 807 | 17.1% | upper bound |
| priceCurrency fails | 138 of 807 | 17.1% | upper bound |
| availability fails | 144 of 807 | 17.8% | upper bound |
Mandatory caveats for the frozen-corpus table translated rendering · sealed Hungarian original (authoritative) inside
The merchant-validity row — translated rendering: every caveat of this row is recurring — the measured bound, the named nested-price limit, the true complement, denominator purity, new instrument on the same window, the extended validator being a superset of the original, the stored-byte instrument, the frozen-corpus exception, and the two-route proof; see the English key.
English working translation — keyed to the sealed caveats: MEASURED BOUND (R1, from the adjudication): merchant_valid=TRUE stratum (C01, stratum size 626) of 5 examined rows 5 upheld = 0 false positives; in the presence-without-validity strata, of 10 examined rows 3 flipped to VALID (C02 2/5, stratum size 17; C03 1/5, stratum size 46), all three the same error pattern. Because of k=5 (n<20), ONLY a count, we do not publish a rate. • (a) NAMED BOUND, verified in code: product_rubric reads only the Offer/node-level price|lowPrice key (run_geo_audit.py:494-499), not offers[].priceSpecification[].price. Adjudicated quote (timeoutx.ae): '"offers":[{"@type":"Offer","priceSpecification":[{"@type":"UnitPriceSpecification","price":"200","priceCurrency":"AED"'. Because of this, the 626/807 is a LOWER BOUND. • TRUE COMPLEMENT: on F_prod 807-626 = 181 rows (22.43%) are not valid; this is exactly 118 (14.62%) without Product JSON-LD + 63 (7.81%) presence-without-validity, 118+63=181 closes. The 9.14% = 63/689 shown in the header stands on a DIFFERENT denominator (the base of those publishing JSON-LD); the two numbers are not interchangeable, and neither is the complement of merchant_valid. • DENOMINATOR PURITY (adjudicated, C04 stratum n=118): of 5 examined rows 2 were WooCommerce /shop/ ARCHIVE pages, which the weak cart_marker branch validated as product pages (ajmanfurniture.com, tealand.ae) — the F_prod=807 contains non-product pages, which can never be merchant_valid. Direction: removing them WOULD RAISE the rate, in the same direction as bound (a). • NEW INSTRUMENT, SAME WINDOW: this row runs on the FETCHED, frozen corpus of the 2026-08-08A round, with an extended validator branch (product_rubric_extended); it is NOT a correction of the sealed parent row, nor is it a new measurement window - the difference is solely the instrument's, the input bytes are identical. • V2 = V1 SUPERSET BY CONSTRUCTION: the extended offer view (extended_offers) contains the v1 offer list, plus it descends into the priceSpecification nodes, therefore the v2 numerator per row is at least as large as the v1; the difference is EXACTLY the size of the nested-price bound named in Part 1, not new data and not a new measurement. • STORED-BYTE INSTRUMENT: every parse runs on the STORED (truncated to 1.5 MB, phone-number-masked) bytes, not on a live page. Because of the truncation a late JSON-LD block may be left out (the html_truncated_stored sub-frame names this), therefore the presence-based numerators are LOWER bounds, the absence-based ones UPPER bounds. • CORPUS-REPRODUCIBILITY EXCEPTION: in the frozen corpus, the stored JSON-LD of 1 domain (soxfootwear.ae) was syntactically broken by the Part 1 phone masking (a digit sequence inside the gtin13 value was rewritten to a mask token), therefore the stored-form measurement sees invalid there, while the sealed fetch-time fields were valid; the row measures the STORED form, the exception set is computed FROM CODE and the conservation gate runs with a sealed bound widened by exactly this amount (quality-prereg-amendment-2). • TWO ROUTES: every published cell was produced by two independent routes (record-level, and a derivation re-extracted from the stored HTML), and the two must match exactly, cell by cell; on a mismatch the round writes out a symmetric difference and does NOT write a table.
sealed audit record, Hungarian, verbatim: MÉRT KORLÁT (R1, az adjudikációból): merchant_valid=IGAZ réteg (C01, rétegméret 626) 5 vizsgált sorból 5 helybenhagyva = 0 hamis pozitív; a jelenlét-érvényesség-nélküli rétegekben 10 vizsgált sorból 3 fordult ÉRVÉNYESRE (C02 2/5, rétegméret 17; C03 1/5, rétegméret 46), mindhárom ugyanaz a hibaminta. k=5 (n<20) miatt CSAK darabszám, arányt nem közlünk. • (a) NEVESÍTETT KORLÁT, kódban igazolva: product_rubric csak Offer/node szintű price|lowPrice kulcsot olvas (run_geo_audit.py:494-499), az offers[].priceSpecification[].price-t nem. Adjudikált idézet (timeoutx.ae): '"offers":[{"@type":"Offer","priceSpecification":[{"@type":"UnitPriceSpecification","price":"200","priceCurrency":"AED"'. Emiatt a 626/807 ALSÓ KORLÁT. • IGAZ KOMPLEMENTER: F_prod-on 807-626 = 181 sor (22,43%) nem érvényes; ez pontosan 118 (14,62%) Product JSON-LD nélküli + 63 (7,81%) jelenlét-érvényesség-nélkül, 118+63=181 zár. A fejlécben szereplő 9,14% = 63/689 MÁS nevezőn (a JSON-LD-t közlők bázisa) áll; a két szám nem cserélhető fel, és egyik sem a merchant_valid komplementere. • NEVEZŐ-TISZTASÁG (adjudikált, C04 réteg n=118): 5 vizsgált sorból 2 WooCommerce /shop/ ARCHÍVUM oldal volt, amit a gyenge cart_marker-ág validált termékoldalként (ajmanfurniture.com, tealand.ae) — az F_prod=807 tartalmaz nem-termékoldalakat, amelyek sosem lehetnek merchant_valid. Irány: eltávolításuk EMELNÉ az arányt, azonos irányba az (a) korláttal. • ÚJ MŰSZER, AZONOS ABLAK: ez a sor a 2026-08-08A körben LETÖLTÖTT, lefagyasztott korpuszon fut egy kiterjesztett validátor-ággal (product_rubric_extended); NEM korrekciója a lepecsételt szülő-sornak és nem is új mérési ablak - a különbség kizárólag a műszeré, a bemenő byte-ok azonosak. • V2 = V1 SZUPERHALMAZ KONSTRUKCIÓBÓL: a kiterjesztett ajánlat-nézet (extended_offers) tartalmazza a v1 ajánlat-listát, plusz leszáll a priceSpecification-csomópontokba, ezért a v2 számláló soronként legalább akkora, mint a v1; a különbség PONTOSAN a Part 1-ben nevesített beágyazott-ár korlát mérete, nem új adat és nem új mérés. • TÁROLT-BYTE MŰSZER: minden parse a LETÁROLT (1,5 MB-ra vágott, telefonszám-maszkolt) byte-okon fut, nem élő oldalon. A vágás miatt egy késői JSON-LD blokk kimaradhat (ezt a html_truncated_stored alkeret nevesíti), ezért a jelenlét-alapú számlálók ALSÓ korlátok, a hiány-alapúak FELSŐ korlátok. • KORPUSZ-REPRODUKÁLHATÓSÁGI KIVÉTEL: a fagyasztott korpuszban 1 domain (soxfootwear.ae) tárolt JSON-LD-jét a Part 1 telefon-maszkolása szintaktikailag eltörte (számjegysor a gtin13 értéken belül maszk-tokenre íródott át), ezért a tárolt-forma mérés ott érvénytelent lát, miközben a sealed fetch-kori mezők validok voltak; a sor a TÁROLT formát méri, a kivétel-halmaz KÓDBÓL számított és a conservation-kapu pontosan ennyivel szélesített sealed-korláttal fut (quality-prereg-amendment-2). • KÉT ÚT: minden publikált cella két független úton készült (rekord-szintű, illetve a tárolt HTML-ből újra-kinyert deriváció), és a kettőnek cellára pontosan egyeznie kell; eltérés esetén a kör kiír egy szimmetrikus differenciát és NEM ír táblát.
Every field row — translated rendering: upper bounds on the conflated frozen frame — the no-JSON-LD pages are included and the clean-frame counts live in the clean blocks; otherwise the same recurring set: no aggregate home, named validator limit, measured bound, new instrument on the same window, stored bytes, frozen-corpus exception, two routes — see the English key.
English working translation — keyed to the sealed caveats: W5: there is no aggregate field. The prereg rider_2 'stated FROM THE AGGREGATE' requirement therefore can NOT be met - the pf_* breakdown lives only in the flat CSV / sites JSONL. Prereg-vs-artifact seam, named; hence a PASS_WITH_CAVEAT ceiling. • CONFLATION (the most serious reading trap): on F_prod the False counts also include the 118 rows WITHOUT Product JSON-LD - on an empty nodes list product_rubric leaves all 5 fields at False (measured: 0 no-ldjson rows where any field is True). Real field invalidity on F_prod∩ldjson (den=689): name 1, image 5, price 57, priceCurrency 56, availability 25. • NESTED-PRICE BOUND (mandatory, explicit): the price/priceCurrency branch only walks the _offers_list(n)+[n] level (lines 494-499); it does NOT read the offers.priceSpecification / UnitPriceSpecification form. The 175 and the 174 are therefore an UPPER bound, not a point estimate. • MEASURED BOUND (R1, verbatim): geo-audit-challenger-flips 60 adjudicated rows - upheld 49 / flipped 10 / unresolved_engine_kept 1. Product-field strata: C02 (shopify presence-without-validity, stratum_n=17, k=5) 2 flips; C03 (non-shopify presence-without-validity, stratum_n=46, k=5) 1 flip. All three flips are EXACTLY the nested priceSpecification form; 3/10 sampled rows. • NEW INSTRUMENT, SAME WINDOW: this row runs on the FETCHED, frozen corpus of the 2026-08-08A round, with an extended validator branch (product_rubric_extended); it is NOT a correction of the sealed parent row, nor is it a new measurement window - the difference is solely the instrument's, the input bytes are identical. • CONFLATED FRAME: this row is to be read on the full product-page frame, so it also contains the rows WITHOUT Product JSON-LD (on an empty node list the rubric leaves every field false); its direction is therefore an UPPER bound on the true field invalidity. The clean frame of those publishing JSON-LD lives in the rows with the -FQL suffix. • STORED-BYTE INSTRUMENT: every parse runs on the STORED (truncated to 1.5 MB, phone-number-masked) bytes, not on a live page. Because of the truncation a late JSON-LD block may be left out (the html_truncated_stored sub-frame names this), therefore the presence-based numerators are LOWER bounds, the absence-based ones UPPER bounds. • CORPUS-REPRODUCIBILITY EXCEPTION: in the frozen corpus, the stored JSON-LD of 1 domain (soxfootwear.ae) was syntactically broken by the Part 1 phone masking (a digit sequence inside the gtin13 value was rewritten to a mask token), therefore the stored-form measurement sees invalid there, while the sealed fetch-time fields were valid; the row measures the STORED form, the exception set is computed FROM CODE and the conservation gate runs with a sealed bound widened by exactly this amount (quality-prereg-amendment-2). • TWO ROUTES: every published cell was produced by two independent routes (record-level, and a derivation re-extracted from the stored HTML), and the two must match exactly, cell by cell; on a mismatch the round writes out a symmetric difference and does NOT write a table.
sealed audit record, Hungarian, verbatim: W5: nincs aggregátum-mező. A prereg rider_2 'stated FROM THE AGGREGATE' előírása így NEM teljesíthető - a pf_* bontás csak a flat CSV-ben / sites JSONL-ben él. Prereg-vs-artefaktum varrat, nevesítve; ezért PASS_WITH_CAVEAT plafon. • KONFLÁCIÓ (a legsúlyosabb olvasati csapda): F_prod-on a false-számok magukban foglalják a 118 Product JSON-LD NÉLKÜLI sort is - a product_rubric üres nodes-listán mind az 5 mezőt False-on hagyja (mérve: 0 olyan no-ldjson sor, ahol bármely mező True). Valódi mező-érvénytelenség F_prod∩ldjson-on (den=689): name 1, image 5, price 57, priceCurrency 56, availability 25. • BEÁGYAZOTT-ÁR KORLÁT (kötelező, explicit): a price/priceCurrency ág csak az _offers_list(n)+[n] szintet járja be (494-499. sor); az offers.priceSpecification / UnitPriceSpecification formát NEM olvassa. A 175 és a 174 ezért FELSŐ korlát, nem pontbecslés. • MÉRT KORLÁT (R1, verbatim): geo-audit-challenger-flips 60 adjudikált sor - upheld 49 / flipped 10 / unresolved_engine_kept 1. Termék-mező-strátumok: C02 (shopify presence-without-validity, stratum_n=17, k=5) 2 flip; C03 (nem-shopify presence-without-validity, stratum_n=46, k=5) 1 flip. Mindhárom flip PONTOSAN a beágyazott priceSpecification-forma; 3/10 mintázott sor. • ÚJ MŰSZER, AZONOS ABLAK: ez a sor a 2026-08-08A körben LETÖLTÖTT, lefagyasztott korpuszon fut egy kiterjesztett validátor-ággal (product_rubric_extended); NEM korrekciója a lepecsételt szülő-sornak és nem is új mérési ablak - a különbség kizárólag a műszeré, a bemenő byte-ok azonosak. • KONFLÁLT KERET: ez a sor a teljes termékoldal-kereten értendő, tehát a Product JSON-LD NÉLKÜLI sorokat is tartalmazza (üres node-listán a rubrika minden mezőt hamisra hagy); iránya ezért FELSŐ korlát a valódi mező-érvénytelenségre. A JSON-LD-t közlők tiszta kerete a -FQL utótagú sorokban él. • TÁROLT-BYTE MŰSZER: minden parse a LETÁROLT (1,5 MB-ra vágott, telefonszám-maszkolt) byte-okon fut, nem élő oldalon. A vágás miatt egy késői JSON-LD blokk kimaradhat (ezt a html_truncated_stored alkeret nevesíti), ezért a jelenlét-alapú számlálók ALSÓ korlátok, a hiány-alapúak FELSŐ korlátok. • KORPUSZ-REPRODUKÁLHATÓSÁGI KIVÉTEL: a fagyasztott korpuszban 1 domain (soxfootwear.ae) tárolt JSON-LD-jét a Part 1 telefon-maszkolása szintaktikailag eltörte (számjegysor a gtin13 értéken belül maszk-tokenre íródott át), ezért a tárolt-forma mérés ott érvénytelent lát, miközben a sealed fetch-kori mezők validok voltak; a sor a TÁROLT formát méri, a kivétel-halmaz KÓDBÓL számított és a conservation-kapu pontosan ennyivel szélesített sealed-korláttal fut (quality-prereg-amendment-2). • KÉT ÚT: minden publikált cella két független úton készült (rekord-szintű, illetve a tárolt HTML-ből újra-kinyert deriváció), és a kettőnek cellára pontosan egyeznie kell; eltérés esetén a kör kiír egy szimmetrikus differenciát és NEM ír táblát.
Validity in the fresh window
(643 of 798 · lower bound · new-window reading, never a correction of Part One)
| Instrument | pages | share | direction |
|---|---|---|---|
original validator (Q-W1-merchant-valid-v1-FQ) | 615 of 798 | 77.1% | lower bound |
extended validator (Q-W1-merchant-valid-v2-FQ) | 643 of 798 | 80.6% | lower bound |
Mandatory caveats for the fresh-window validity table translated rendering · sealed Hungarian original (authoritative) inside
The original-validator row — translated rendering: all recurring — the measured bound, the named nested-price limit, the true complement, denominator purity, the new measurement window, the self-selected frame, the stored-byte instrument and the two-route proof; see the English key.
English working translation — keyed to the sealed caveats: MEASURED BOUND (R1, from the adjudication): merchant_valid=TRUE stratum (C01, stratum size 626) of 5 examined rows 5 upheld = 0 false positives; in the presence-without-validity strata, of 10 examined rows 3 flipped to VALID (C02 2/5, stratum size 17; C03 1/5, stratum size 46), all three the same error pattern. Because of k=5 (n<20), ONLY a count, we do not publish a rate. • (a) NAMED BOUND, verified in code: product_rubric reads only the Offer/node-level price|lowPrice key (run_geo_audit.py:494-499), not offers[].priceSpecification[].price. Adjudicated quote (timeoutx.ae): '"offers":[{"@type":"Offer","priceSpecification":[{"@type":"UnitPriceSpecification","price":"200","priceCurrency":"AED"'. Because of this, the 626/807 is a LOWER BOUND. • TRUE COMPLEMENT: on F_prod 807-626 = 181 rows (22.43%) are not valid; this is exactly 118 (14.62%) without Product JSON-LD + 63 (7.81%) presence-without-validity, 118+63=181 closes. The 9.14% = 63/689 shown in the header stands on a DIFFERENT denominator (the base of those publishing JSON-LD); the two numbers are not interchangeable, and neither is the complement of merchant_valid. • DENOMINATOR PURITY (adjudicated, C04 stratum n=118): of 5 examined rows 2 were WooCommerce /shop/ ARCHIVE pages, which the weak cart_marker branch validated as product pages (ajmanfurniture.com, tealand.ae) — the F_prod=807 contains non-product pages, which can never be merchant_valid. Direction: removing them WOULD RAISE the rate, in the same direction as bound (a). • NEW MEASUREMENT WINDOW: this row stands on the pages RE-fetched in the 2026-08-13Q round (F_q). The difference between the 08A and the Q window is REAL-WORLD CHANGE (the page was transformed, ceased to exist or now refuses the request), not a measurement error; the full breakdown ships in the Q-drift-* rows, nothing is hidden. • SELF-SELECTION FRAME: F_q consists of the domains that were successfully fetched AND parsed in THIS window; the dropouts (403/429/gone/timeout) are not missing by random sampling, therefore the rate measured on F_q cannot be automatically carried over to the full input frame. • STORED-BYTE INSTRUMENT: every parse runs on the STORED (truncated to 1.5 MB, phone-number-masked) bytes, not on a live page. Because of the truncation a late JSON-LD block may be left out (the html_truncated_stored sub-frame names this), therefore the presence-based numerators are LOWER bounds, the absence-based ones UPPER bounds. • TWO ROUTES: every published cell was produced by two independent routes (record-level, and a derivation re-extracted from the stored HTML), and the two must match exactly, cell by cell; on a mismatch the round writes out a symmetric difference and does NOT write a table.
sealed audit record, Hungarian, verbatim: MÉRT KORLÁT (R1, az adjudikációból): merchant_valid=IGAZ réteg (C01, rétegméret 626) 5 vizsgált sorból 5 helybenhagyva = 0 hamis pozitív; a jelenlét-érvényesség-nélküli rétegekben 10 vizsgált sorból 3 fordult ÉRVÉNYESRE (C02 2/5, rétegméret 17; C03 1/5, rétegméret 46), mindhárom ugyanaz a hibaminta. k=5 (n<20) miatt CSAK darabszám, arányt nem közlünk. • (a) NEVESÍTETT KORLÁT, kódban igazolva: product_rubric csak Offer/node szintű price|lowPrice kulcsot olvas (run_geo_audit.py:494-499), az offers[].priceSpecification[].price-t nem. Adjudikált idézet (timeoutx.ae): '"offers":[{"@type":"Offer","priceSpecification":[{"@type":"UnitPriceSpecification","price":"200","priceCurrency":"AED"'. Emiatt a 626/807 ALSÓ KORLÁT. • IGAZ KOMPLEMENTER: F_prod-on 807-626 = 181 sor (22,43%) nem érvényes; ez pontosan 118 (14,62%) Product JSON-LD nélküli + 63 (7,81%) jelenlét-érvényesség-nélkül, 118+63=181 zár. A fejlécben szereplő 9,14% = 63/689 MÁS nevezőn (a JSON-LD-t közlők bázisa) áll; a két szám nem cserélhető fel, és egyik sem a merchant_valid komplementere. • NEVEZŐ-TISZTASÁG (adjudikált, C04 réteg n=118): 5 vizsgált sorból 2 WooCommerce /shop/ ARCHÍVUM oldal volt, amit a gyenge cart_marker-ág validált termékoldalként (ajmanfurniture.com, tealand.ae) — az F_prod=807 tartalmaz nem-termékoldalakat, amelyek sosem lehetnek merchant_valid. Irány: eltávolításuk EMELNÉ az arányt, azonos irányba az (a) korláttal. • ÚJ MÉRÉSI ABLAK: ez a sor a 2026-08-13Q körben ÚJRA letöltött oldalakon (F_q) áll. A 08A és a Q ablak közti eltérés VALÓS VILÁGBELI VÁLTOZÁS (az oldal átalakult, megszűnt vagy most utasítja el a kérést), nem mérési hiba; a teljes bontás a Q-drift-* sorokban megy ki, elrejtve semmi nincs. • ÖNSZELEKCIÓS KERET: az F_q azokból a domainekből áll, amelyek EBBEN az ablakban sikeresen letöltődtek ÉS parse-olódtak; a kiesők (403/429/gone/timeout) nem véletlen mintavétellel hiányoznak, ezért az F_q-n mért arány nem vihető át automatikusan a teljes bemeneti keretre. • TÁROLT-BYTE MŰSZER: minden parse a LETÁROLT (1,5 MB-ra vágott, telefonszám-maszkolt) byte-okon fut, nem élő oldalon. A vágás miatt egy késői JSON-LD blokk kimaradhat (ezt a html_truncated_stored alkeret nevesíti), ezért a jelenlét-alapú számlálók ALSÓ korlátok, a hiány-alapúak FELSŐ korlátok. • KÉT ÚT: minden publikált cella két független úton készült (rekord-szintű, illetve a tárolt HTML-ből újra-kinyert deriváció), és a kettőnek cellára pontosan egyeznie kell; eltérés esetén a kör kiír egy szimmetrikus differenciát és NEM ír táblát.
The extended-validator row — translated rendering: the same recurring set, plus: the extended validator is a superset of the original, so its counter is per-row at least the original's — the difference is exactly the named nested-price limitation; see the English key.
English working translation — keyed to the sealed caveats: MEASURED BOUND (R1, from the adjudication): merchant_valid=TRUE stratum (C01, stratum size 626) of 5 examined rows 5 upheld = 0 false positives; in the presence-without-validity strata, of 10 examined rows 3 flipped to VALID (C02 2/5, stratum size 17; C03 1/5, stratum size 46), all three the same error pattern. Because of k=5 (n<20), ONLY a count, we do not publish a rate. • (a) NAMED BOUND, verified in code: product_rubric reads only the Offer/node-level price|lowPrice key (run_geo_audit.py:494-499), not offers[].priceSpecification[].price. Adjudicated quote (timeoutx.ae): '"offers":[{"@type":"Offer","priceSpecification":[{"@type":"UnitPriceSpecification","price":"200","priceCurrency":"AED"'. Because of this, the 626/807 is a LOWER BOUND. • TRUE COMPLEMENT: on F_prod 807-626 = 181 rows (22.43%) are not valid; this is exactly 118 (14.62%) without Product JSON-LD + 63 (7.81%) presence-without-validity, 118+63=181 closes. The 9.14% = 63/689 shown in the header stands on a DIFFERENT denominator (the base of those publishing JSON-LD); the two numbers are not interchangeable, and neither is the complement of merchant_valid. • DENOMINATOR PURITY (adjudicated, C04 stratum n=118): of 5 examined rows 2 were WooCommerce /shop/ ARCHIVE pages, which the weak cart_marker branch validated as product pages (ajmanfurniture.com, tealand.ae) — the F_prod=807 contains non-product pages, which can never be merchant_valid. Direction: removing them WOULD RAISE the rate, in the same direction as bound (a). • NEW MEASUREMENT WINDOW: this row stands on the pages RE-fetched in the 2026-08-13Q round (F_q). The difference between the 08A and the Q window is REAL-WORLD CHANGE (the page was transformed, ceased to exist or now refuses the request), not a measurement error; the full breakdown ships in the Q-drift-* rows, nothing is hidden. • SELF-SELECTION FRAME: F_q consists of the domains that were successfully fetched AND parsed in THIS window; the dropouts (403/429/gone/timeout) are not missing by random sampling, therefore the rate measured on F_q cannot be automatically carried over to the full input frame. • STORED-BYTE INSTRUMENT: every parse runs on the STORED (truncated to 1.5 MB, phone-number-masked) bytes, not on a live page. Because of the truncation a late JSON-LD block may be left out (the html_truncated_stored sub-frame names this), therefore the presence-based numerators are LOWER bounds, the absence-based ones UPPER bounds. • V2 = V1 SUPERSET BY CONSTRUCTION: the extended offer view (extended_offers) contains the v1 offer list, plus it descends into the priceSpecification nodes, therefore the v2 numerator per row is at least as large as the v1; the difference is EXACTLY the size of the nested-price bound named in Part 1, not new data and not a new measurement. • TWO ROUTES: every published cell was produced by two independent routes (record-level, and a derivation re-extracted from the stored HTML), and the two must match exactly, cell by cell; on a mismatch the round writes out a symmetric difference and does NOT write a table.
sealed audit record, Hungarian, verbatim: MÉRT KORLÁT (R1, az adjudikációból): merchant_valid=IGAZ réteg (C01, rétegméret 626) 5 vizsgált sorból 5 helybenhagyva = 0 hamis pozitív; a jelenlét-érvényesség-nélküli rétegekben 10 vizsgált sorból 3 fordult ÉRVÉNYESRE (C02 2/5, rétegméret 17; C03 1/5, rétegméret 46), mindhárom ugyanaz a hibaminta. k=5 (n<20) miatt CSAK darabszám, arányt nem közlünk. • (a) NEVESÍTETT KORLÁT, kódban igazolva: product_rubric csak Offer/node szintű price|lowPrice kulcsot olvas (run_geo_audit.py:494-499), az offers[].priceSpecification[].price-t nem. Adjudikált idézet (timeoutx.ae): '"offers":[{"@type":"Offer","priceSpecification":[{"@type":"UnitPriceSpecification","price":"200","priceCurrency":"AED"'. Emiatt a 626/807 ALSÓ KORLÁT. • IGAZ KOMPLEMENTER: F_prod-on 807-626 = 181 sor (22,43%) nem érvényes; ez pontosan 118 (14,62%) Product JSON-LD nélküli + 63 (7,81%) jelenlét-érvényesség-nélkül, 118+63=181 zár. A fejlécben szereplő 9,14% = 63/689 MÁS nevezőn (a JSON-LD-t közlők bázisa) áll; a két szám nem cserélhető fel, és egyik sem a merchant_valid komplementere. • NEVEZŐ-TISZTASÁG (adjudikált, C04 réteg n=118): 5 vizsgált sorból 2 WooCommerce /shop/ ARCHÍVUM oldal volt, amit a gyenge cart_marker-ág validált termékoldalként (ajmanfurniture.com, tealand.ae) — az F_prod=807 tartalmaz nem-termékoldalakat, amelyek sosem lehetnek merchant_valid. Irány: eltávolításuk EMELNÉ az arányt, azonos irányba az (a) korláttal. • ÚJ MÉRÉSI ABLAK: ez a sor a 2026-08-13Q körben ÚJRA letöltött oldalakon (F_q) áll. A 08A és a Q ablak közti eltérés VALÓS VILÁGBELI VÁLTOZÁS (az oldal átalakult, megszűnt vagy most utasítja el a kérést), nem mérési hiba; a teljes bontás a Q-drift-* sorokban megy ki, elrejtve semmi nincs. • ÖNSZELEKCIÓS KERET: az F_q azokból a domainekből áll, amelyek EBBEN az ablakban sikeresen letöltődtek ÉS parse-olódtak; a kiesők (403/429/gone/timeout) nem véletlen mintavétellel hiányoznak, ezért az F_q-n mért arány nem vihető át automatikusan a teljes bemeneti keretre. • TÁROLT-BYTE MŰSZER: minden parse a LETÁROLT (1,5 MB-ra vágott, telefonszám-maszkolt) byte-okon fut, nem élő oldalon. A vágás miatt egy késői JSON-LD blokk kimaradhat (ezt a html_truncated_stored alkeret nevesíti), ezért a jelenlét-alapú számlálók ALSÓ korlátok, a hiány-alapúak FELSŐ korlátok. • V2 = V1 SZUPERHALMAZ KONSTRUKCIÓBÓL: a kiterjesztett ajánlat-nézet (extended_offers) tartalmazza a v1 ajánlat-listát, plusz leszáll a priceSpecification-csomópontokba, ezért a v2 számláló soronként legalább akkora, mint a v1; a különbség PONTOSAN a Part 1-ben nevesített beágyazott-ár korlát mérete, nem új adat és nem új mérés. • KÉT ÚT: minden publikált cella két független úton készült (rekord-szintű, illetve a tárolt HTML-ből újra-kinyert deriváció), és a kettőnek cellára pontosan egyeznie kell; eltérés esetén a kör kiír egy szimmetrikus differenciát és NEM ír táblát.
Field invalidity on the clean new-window frame — a separate block
Only pages that carry a Product JSON-LD block in the new window. Denominator for every row: 679 pages (frame F_q AND Product JSON-LD present (n=679)). The name and image cells sit under the pre-registered floor and ship as counts.
| Field | pages | share | direction |
|---|---|---|---|
| name | 1 of 679 | n<20 - count only | upper bound |
| image | 5 of 679 | n<20 - count only | upper bound |
| price | 22 of 679 | 3.2% | upper bound |
| priceCurrency | 22 of 679 | 3.2% | upper bound |
| availability | 28 of 679 | 4.1% | upper bound |
Mandatory caveats for the clean-frame field table translated rendering · sealed Hungarian original (authoritative) inside
Every row of this table — translated rendering: this block stands on the clean new-window intersection — pages in F_q that publish Product JSON-LD — which resolves the conflation trap; the price-family cells remain upper bounds under the named nested-price validator limit. Recurring: no aggregate home · measured bound · new measurement window · stored bytes · two routes — see the English key.
English working translation — keyed to the sealed caveats: W5: there is no aggregate field. The prereg rider_2 'stated FROM THE AGGREGATE' requirement therefore can NOT be met - the pf_* breakdown lives only in the flat CSV / sites JSONL. Prereg-vs-artifact seam, named; hence a PASS_WITH_CAVEAT ceiling. • CONFLATION (the most serious reading trap): on F_prod the False counts also include the 118 rows WITHOUT Product JSON-LD - on an empty nodes list product_rubric leaves all 5 fields at False (measured: 0 no-ldjson rows where any field is True). Real field invalidity on F_prod∩ldjson (den=689): name 1, image 5, price 57, priceCurrency 56, availability 25. • NESTED-PRICE BOUND (mandatory, explicit): the price/priceCurrency branch only walks the _offers_list(n)+[n] level (lines 494-499); it does NOT read the offers.priceSpecification / UnitPriceSpecification form. The 175 and the 174 are therefore an UPPER bound, not a point estimate. • MEASURED BOUND (R1, verbatim): geo-audit-challenger-flips 60 adjudicated rows - upheld 49 / flipped 10 / unresolved_engine_kept 1. Product-field strata: C02 (shopify presence-without-validity, stratum_n=17, k=5) 2 flips; C03 (non-shopify presence-without-validity, stratum_n=46, k=5) 1 flip. All three flips are EXACTLY the nested priceSpecification form; 3/10 sampled rows. • NEW MEASUREMENT WINDOW: this row stands on the pages RE-fetched in the 2026-08-13Q round (F_q). The difference between the 08A and the Q window is REAL-WORLD CHANGE (the page was transformed, ceased to exist or now refuses the request), not a measurement error; the full breakdown ships in the Q-drift-* rows, nothing is hidden. • TRUE FIELD-INVALIDITY FRAME: this row is to be read on the F_q ∩ Product-JSON-LD frame - the rows not publishing JSON-LD at all are excluded, so the CONFLATION trap is resolved. • STORED-BYTE INSTRUMENT: every parse runs on the STORED (truncated to 1.5 MB, phone-number-masked) bytes, not on a live page. Because of the truncation a late JSON-LD block may be left out (the html_truncated_stored sub-frame names this), therefore the presence-based numerators are LOWER bounds, the absence-based ones UPPER bounds. • TWO ROUTES: every published cell was produced by two independent routes (record-level, and a derivation re-extracted from the stored HTML), and the two must match exactly, cell by cell; on a mismatch the round writes out a symmetric difference and does NOT write a table.
sealed audit record, Hungarian, verbatim: W5: nincs aggregátum-mező. A prereg rider_2 'stated FROM THE AGGREGATE' előírása így NEM teljesíthető - a pf_* bontás csak a flat CSV-ben / sites JSONL-ben él. Prereg-vs-artefaktum varrat, nevesítve; ezért PASS_WITH_CAVEAT plafon. • KONFLÁCIÓ (a legsúlyosabb olvasati csapda): F_prod-on a false-számok magukban foglalják a 118 Product JSON-LD NÉLKÜLI sort is - a product_rubric üres nodes-listán mind az 5 mezőt False-on hagyja (mérve: 0 olyan no-ldjson sor, ahol bármely mező True). Valódi mező-érvénytelenség F_prod∩ldjson-on (den=689): name 1, image 5, price 57, priceCurrency 56, availability 25. • BEÁGYAZOTT-ÁR KORLÁT (kötelező, explicit): a price/priceCurrency ág csak az _offers_list(n)+[n] szintet járja be (494-499. sor); az offers.priceSpecification / UnitPriceSpecification formát NEM olvassa. A 175 és a 174 ezért FELSŐ korlát, nem pontbecslés. • MÉRT KORLÁT (R1, verbatim): geo-audit-challenger-flips 60 adjudikált sor - upheld 49 / flipped 10 / unresolved_engine_kept 1. Termék-mező-strátumok: C02 (shopify presence-without-validity, stratum_n=17, k=5) 2 flip; C03 (nem-shopify presence-without-validity, stratum_n=46, k=5) 1 flip. Mindhárom flip PONTOSAN a beágyazott priceSpecification-forma; 3/10 mintázott sor. • ÚJ MÉRÉSI ABLAK: ez a sor a 2026-08-13Q körben ÚJRA letöltött oldalakon (F_q) áll. A 08A és a Q ablak közti eltérés VALÓS VILÁGBELI VÁLTOZÁS (az oldal átalakult, megszűnt vagy most utasítja el a kérést), nem mérési hiba; a teljes bontás a Q-drift-* sorokban megy ki, elrejtve semmi nincs. • VALÓDI MEZŐ-ÉRVÉNYTELENSÉG-KERET: ez a sor az F_q ∩ Product-JSON-LD kereten értendő - a JSON-LD-t egyáltalán nem közlő sorok kizárva, így a KONFLÁCIÓ-csapda fel van oldva. • TÁROLT-BYTE MŰSZER: minden parse a LETÁROLT (1,5 MB-ra vágott, telefonszám-maszkolt) byte-okon fut, nem élő oldalon. A vágás miatt egy késői JSON-LD blokk kimaradhat (ezt a html_truncated_stored alkeret nevesíti), ezért a jelenlét-alapú számlálók ALSÓ korlátok, a hiány-alapúak FELSŐ korlátok. • KÉT ÚT: minden publikált cella két független úton készült (rekord-szintű, illetve a tárolt HTML-ből újra-kinyert deriváció), és a kettőnek cellára pontosan egyeznie kell; eltérés esetén a kör kiír egy szimmetrikus differenciát és NEM ír táblát.
Depth and richness among pages with Product JSON-LD
Denominator for every recommended-field row: 679 pages (frame F_q AND Product JSON-LD present (n=679)).
| Field | pages | share | direction |
|---|---|---|---|
| brand | 538 of 679 | 79.2% | lower bound |
| identifier (GTIN, MPN or SKU) | 460 of 679 | 67.8% | lower bound |
| description | 646 of 679 | 95.1% | lower bound |
| aggregateRating | 71 of 679 | 10.5% | lower bound |
| review | 25 of 679 | 3.7% | lower bound |
| seller | 194 of 679 | 28.6% | lower bound |
| itemCondition | 162 of 679 | 23.9% | lower bound |
| more than one image | 96 of 679 | 14.1% | lower bound |
| offer core complete (price, priceCurrency, availability) | 647 of 679 | 95.3% | lower bound |
| Band | pages | note |
|---|---|---|
short (Q-W1-desclen-lt30-FQdesc) | 135 of 646 | count only - family rule |
medium (Q-W1-desclen-w30_150-FQdesc) | 345 of 646 | count only - family rule |
long (Q-W1-desclen-gt150-FQdesc) | 166 of 646 | count only - family rule |
Mandatory caveats for the depth and richness tables translated rendering · sealed Hungarian original (authoritative) inside
Every recommended-field and richness row except the offer-core row — translated rendering: schema.org treats these properties as recommended, so absence is not invalidity — the row counts presence on the JSON-LD-publishing frame, and a gap reads as unused surface, not as an error. Recurring: new measurement window · stored bytes · two routes — see the English key.
English working translation — keyed to the sealed caveats: NEW MEASUREMENT WINDOW: this row stands on the pages RE-fetched in the 2026-08-13Q round (F_q). The difference between the 08A and the Q window is REAL-WORLD CHANGE (the page was transformed, ceased to exist or now refuses the request), not a measurement error; the full breakdown ships in the Q-drift-* rows, nothing is hidden. • RECOMMENDED, NOT REQUIRED FIELD: schema.org treats this property as recommended, its absence is not invalidity; the row counts PRESENCE on the frame of those publishing JSON-LD, and the absence is to be read not as an error but as unused surface. • STORED-BYTE INSTRUMENT: every parse runs on the STORED (truncated to 1.5 MB, phone-number-masked) bytes, not on a live page. Because of the truncation a late JSON-LD block may be left out (the html_truncated_stored sub-frame names this), therefore the presence-based numerators are LOWER bounds, the absence-based ones UPPER bounds. • TWO ROUTES: every published cell was produced by two independent routes (record-level, and a derivation re-extracted from the stored HTML), and the two must match exactly, cell by cell; on a mismatch the round writes out a symmetric difference and does NOT write a table.
sealed audit record, Hungarian, verbatim: ÚJ MÉRÉSI ABLAK: ez a sor a 2026-08-13Q körben ÚJRA letöltött oldalakon (F_q) áll. A 08A és a Q ablak közti eltérés VALÓS VILÁGBELI VÁLTOZÁS (az oldal átalakult, megszűnt vagy most utasítja el a kérést), nem mérési hiba; a teljes bontás a Q-drift-* sorokban megy ki, elrejtve semmi nincs. • AJÁNLOTT, NEM KÖTELEZŐ MEZŐ: a schema.org ezt a tulajdonságot ajánlottként kezeli, hiánya nem érvénytelenség; a sor a MEGLÉTET számolja a JSON-LD-t közlők keretén, és a hiány nem hibaként, hanem kihasználatlan felületként olvasandó. • TÁROLT-BYTE MŰSZER: minden parse a LETÁROLT (1,5 MB-ra vágott, telefonszám-maszkolt) byte-okon fut, nem élő oldalon. A vágás miatt egy késői JSON-LD blokk kimaradhat (ezt a html_truncated_stored alkeret nevesíti), ezért a jelenlét-alapú számlálók ALSÓ korlátok, a hiány-alapúak FELSŐ korlátok. • KÉT ÚT: minden publikált cella két független úton készült (rekord-szintű, illetve a tárolt HTML-ből újra-kinyert deriváció), és a kettőnek cellára pontosan egyeznie kell; eltérés esetén a kör kiír egy szimmetrikus differenciát és NEM ír táblát.
The offer-core row — translated rendering: the extended offer view is a superset of the original validator's, so the difference is exactly the named nested-price limitation. Recurring: new measurement window · stored bytes · two routes — see the English key.
English working translation — keyed to the sealed caveats: NEW MEASUREMENT WINDOW: this row stands on the pages RE-fetched in the 2026-08-13Q round (F_q). The difference between the 08A and the Q window is REAL-WORLD CHANGE (the page was transformed, ceased to exist or now refuses the request), not a measurement error; the full breakdown ships in the Q-drift-* rows, nothing is hidden. • V2 = V1 SUPERSET BY CONSTRUCTION: the extended offer view (extended_offers) contains the v1 offer list, plus it descends into the priceSpecification nodes, therefore the v2 numerator per row is at least as large as the v1; the difference is EXACTLY the size of the nested-price bound named in Part 1, not new data and not a new measurement. • STORED-BYTE INSTRUMENT: every parse runs on the STORED (truncated to 1.5 MB, phone-number-masked) bytes, not on a live page. Because of the truncation a late JSON-LD block may be left out (the html_truncated_stored sub-frame names this), therefore the presence-based numerators are LOWER bounds, the absence-based ones UPPER bounds. • TWO ROUTES: every published cell was produced by two independent routes (record-level, and a derivation re-extracted from the stored HTML), and the two must match exactly, cell by cell; on a mismatch the round writes out a symmetric difference and does NOT write a table.
sealed audit record, Hungarian, verbatim: ÚJ MÉRÉSI ABLAK: ez a sor a 2026-08-13Q körben ÚJRA letöltött oldalakon (F_q) áll. A 08A és a Q ablak közti eltérés VALÓS VILÁGBELI VÁLTOZÁS (az oldal átalakult, megszűnt vagy most utasítja el a kérést), nem mérési hiba; a teljes bontás a Q-drift-* sorokban megy ki, elrejtve semmi nincs. • V2 = V1 SZUPERHALMAZ KONSTRUKCIÓBÓL: a kiterjesztett ajánlat-nézet (extended_offers) tartalmazza a v1 ajánlat-listát, plusz leszáll a priceSpecification-csomópontokba, ezért a v2 számláló soronként legalább akkora, mint a v1; a különbség PONTOSAN a Part 1-ben nevesített beágyazott-ár korlát mérete, nem új adat és nem új mérés. • TÁROLT-BYTE MŰSZER: minden parse a LETÁROLT (1,5 MB-ra vágott, telefonszám-maszkolt) byte-okon fut, nem élő oldalon. A vágás miatt egy késői JSON-LD blokk kimaradhat (ezt a html_truncated_stored alkeret nevesíti), ezért a jelenlét-alapú számlálók ALSÓ korlátok, a hiány-alapúak FELSŐ korlátok. • KÉT ÚT: minden publikált cella két független úton készült (rekord-szintű, illetve a tárolt HTML-ből újra-kinyert deriváció), és a kettőnek cellára pontosan egyeznie kell; eltérés esetén a kör kiír egy szimmetrikus differenciát és NEM ír táblát.
Every description-band row — translated rendering: the banded quantity is the word count of the tag-stripped, entity-decoded, whitespace-normalized description text, and the longest description per node decides; the three bands partition the description-holding set exactly, and the family ships uniformly as counts — the band thresholds live in the sealed definitions. Recurring: new measurement window · stored bytes · two routes — see the English key.
English working translation — keyed to the sealed caveats: NEW MEASUREMENT WINDOW: this row stands on the pages RE-fetched in the 2026-08-13Q round (F_q). The difference between the 08A and the Q window is REAL-WORLD CHANGE (the page was transformed, ceased to exist or now refuses the request), not a measurement error; the full breakdown ships in the Q-drift-* rows, nothing is hidden. • WORD-COUNT BANDS: the description is the word count of the tag-stripped, entity-decoded, whitespace-normalized text, and per node the LONGEST description decides; the three bands are a PARTITION of the description-holding set (their sum is exactly the denominator), the family is uniformly count. • STORED-BYTE INSTRUMENT: every parse runs on the STORED (truncated to 1.5 MB, phone-number-masked) bytes, not on a live page. Because of the truncation a late JSON-LD block may be left out (the html_truncated_stored sub-frame names this), therefore the presence-based numerators are LOWER bounds, the absence-based ones UPPER bounds. • TWO ROUTES: every published cell was produced by two independent routes (record-level, and a derivation re-extracted from the stored HTML), and the two must match exactly, cell by cell; on a mismatch the round writes out a symmetric difference and does NOT write a table.
sealed audit record, Hungarian, verbatim: ÚJ MÉRÉSI ABLAK: ez a sor a 2026-08-13Q körben ÚJRA letöltött oldalakon (F_q) áll. A 08A és a Q ablak közti eltérés VALÓS VILÁGBELI VÁLTOZÁS (az oldal átalakult, megszűnt vagy most utasítja el a kérést), nem mérési hiba; a teljes bontás a Q-drift-* sorokban megy ki, elrejtve semmi nincs. • SZÓSZÁM-SÁVOK: a leírás a tag-mentesített, entitás-dekódolt, whitespace-normalizált szöveg szószáma, és node-onként a LEGHOSSZABB leírás dönt; a három sáv a leírás-birtokos halmaz PARTÍCIÓJA (összegük pontosan a nevező), a család egységesen darabszám. • TÁROLT-BYTE MŰSZER: minden parse a LETÁROLT (1,5 MB-ra vágott, telefonszám-maszkolt) byte-okon fut, nem élő oldalon. A vágás miatt egy késői JSON-LD blokk kimaradhat (ezt a html_truncated_stored alkeret nevesíti), ezért a jelenlét-alapú számlálók ALSÓ korlátok, a hiány-alapúak FELSŐ korlátok. • KÉT ÚT: minden publikált cella két független úton készült (rekord-szintű, illetve a tárolt HTML-ből újra-kinyert deriváció), és a kettőnek cellára pontosan egyeznie kell; eltérés esetén a kör kiír egy szimmetrikus differenciát és NEM ír táblát.
On-page signals
| Check | pages | share | frame | direction |
|---|---|---|---|---|
| every product image carries alt text | 118 of 785 | 15.0% | F_q AND at least one non-data img src (n=785) | point estimate with measured error |
| canonical link present | 787 of 798 | 98.6% | F_q (n=798) | lower bound |
| Open Graph product markup | 633 of 798 | 79.3% | F_q (n=798) | lower bound |
Mandatory caveats for the on-page table translated rendering · sealed Hungarian original (authoritative) inside
Every row of this table — translated rendering: all recurring — the new measurement window, the stored-byte instrument, the single-source HTML cell (whose real second leg is the separate recount artifact) and the two-route proof; see the English key.
English working translation — keyed to the sealed caveats: NEW MEASUREMENT WINDOW: this row stands on the pages RE-fetched in the 2026-08-13Q round (F_q). The difference between the 08A and the Q window is REAL-WORLD CHANGE (the page was transformed, ceased to exist or now refuses the request), not a measurement error; the full breakdown ships in the Q-drift-* rows, nothing is hidden. • STORED-BYTE INSTRUMENT: every parse runs on the STORED (truncated to 1.5 MB, phone-number-masked) bytes, not on a live page. Because of the truncation a late JSON-LD block may be left out (the html_truncated_stored sub-frame names this), therefore the presence-based numerators are LOWER bounds, the absence-based ones UPPER bounds. • SINGLE-SOURCE HTML CELL: this row reads the same stored product.html bytes on both routes through the same frozen function, therefore its real second leg is the separate CPython recount artefact; the seam is named here, not concealed. • TWO ROUTES: every published cell was produced by two independent routes (record-level, and a derivation re-extracted from the stored HTML), and the two must match exactly, cell by cell; on a mismatch the round writes out a symmetric difference and does NOT write a table.
sealed audit record, Hungarian, verbatim: ÚJ MÉRÉSI ABLAK: ez a sor a 2026-08-13Q körben ÚJRA letöltött oldalakon (F_q) áll. A 08A és a Q ablak közti eltérés VALÓS VILÁGBELI VÁLTOZÁS (az oldal átalakult, megszűnt vagy most utasítja el a kérést), nem mérési hiba; a teljes bontás a Q-drift-* sorokban megy ki, elrejtve semmi nincs. • TÁROLT-BYTE MŰSZER: minden parse a LETÁROLT (1,5 MB-ra vágott, telefonszám-maszkolt) byte-okon fut, nem élő oldalon. A vágás miatt egy késői JSON-LD blokk kimaradhat (ezt a html_truncated_stored alkeret nevesíti), ezért a jelenlét-alapú számlálók ALSÓ korlátok, a hiány-alapúak FELSŐ korlátok. • EGYFORRÁSÚ HTML-CELLA: ez a sor mindkét úton ugyanazokat a tárolt product.html byte-okat olvassa ugyanazon fagyasztott függvényen keresztül, ezért a valódi második lába a külön CPython recount-artefaktum; a varrat itt nevesítve van, nem elhallgatva. • KÉT ÚT: minden publikált cella két független úton készült (rekord-szintű, illetve a tárolt HTML-ből újra-kinyert deriváció), és a kettőnek cellára pontosan egyeznie kell; eltérés esetén a kör kiír egy szimmetrikus differenciát és NEM ír táblát.
Schema versus the visible page
A seeded adversarial strip re-attacked accepted judged match verdicts: 40 of 48 flipped on challenge. The flips measure a divergence between two evidence channels — the frozen comparison spec reads the stored visible text, while the judged cells could also credit the server-rendered product page. By pre-registered rule the strip is never folded into the published match rates; it stands beside them, and every row of this block points back here.
English working translation — keyed to the sealed caveats: INSTRUMENT ROW: this row is not about the stores but about the MEASURING INSTRUMENT - it describes the error of the judged rows running alongside it, and is to be read TOGETHER with them. • HONEST N: where the judged stratum did not run to completion, the denominator is the number of ACCEPTED rows and the shortfall is named in the caveat; no tacit filling-in or averaging happens anywhere. • TWO ROUTES: every published cell was produced by two independent routes (record-level, and a derivation re-extracted from the stored HTML), and the two must match exactly, cell by cell; on a mismatch the round writes out a symmetric difference and does NOT write a table.
sealed audit record, Hungarian, verbatim: MŰSZER-SOR: ez a sor nem a webshopokról, hanem a MÉRŐESZKÖZRŐL szól - a mellette futó ítélt sorok hibáját írja le, és azokkal EGYÜTT olvasandó. • ŐSZINTE N: ahol az ítélt réteg nem futott végig, a nevező az ELFOGADOTT sorok száma és a hiány a caveatben nevesítve van; hallgatólagos kipótlás vagy átlagolás sehol nem történik. • KÉT ÚT: minden publikált cella két független úton készült (rekord-szintű, illetve a tárolt HTML-ből újra-kinyert deriváció), és a kettőnek cellára pontosan egyeznie kell; eltérés esetén a kör kiír egy szimmetrikus differenciát és NEM ír táblát.
| Field | agreement | share | direction | note |
|---|---|---|---|---|
| name | 670 of 678 | 98.8% | point estimate with measured error | instrument caveat - see strip above |
| price | 448 of 657 | 68.2% | point estimate with measured error | instrument caveat - see strip above |
| availability | 375 of 651 | 57.6% | point estimate with measured error | instrument caveat - see strip above |
| Outcome | rows | direction | note |
|---|---|---|---|
| judged match | 473 of 966 | point estimate with measured error | count only - family rule · instrument caveat - see strip above |
| sustained mismatch | 92 of 966 | point estimate with measured error | count only - family rule · instrument caveat - see strip above |
| unverifiable | 401 of 966 | point estimate with measured error | count only - family rule · instrument caveat - see strip above |
| unresolved | 0 of 966 | point estimate with measured error | count only - family rule · instrument caveat - see strip above |
Mandatory caveats for the schema-versus-page block translated rendering · sealed Hungarian original (authoritative) inside
Every agreement row — translated rendering: the denominator is the dimension's comparable set — F_q, publishing Product JSON-LD, with the field's JSON-LD side non-absent; the ambiguous sub-codes marking a missing JSON-LD side are out of frame and reported separately, and the partition — match, mismatch, and in-frame ambiguous — closes the denominator exactly. The comparison runs on the stored visible text: a price or stock state rendered only by JavaScript never reaches the stored bytes — named sub-codes mark these as limits of measurability, not accusations. Recurring: new measurement window · judged residue, conservative · stored bytes · two routes — see the English key.
English working translation — keyed to the sealed caveats: NEW MEASUREMENT WINDOW: this row stands on the pages RE-fetched in the 2026-08-13Q round (F_q). The difference between the 08A and the Q window is REAL-WORLD CHANGE (the page was transformed, ceased to exist or now refuses the request), not a measurement error; the full breakdown ships in the Q-drift-* rows, nothing is hidden. • COMPARABLE FRAME: the denominator is the comparable set of the given dimension (F_q ∩ Product-JSON-LD ∩ the JSON-LD side is NOT missing). The ambiguous sub-codes indicating a missing JSON-LD side are OUTSIDE the frame and are reported separately; the partition (match + mismatch + ambiguous within the frame) yields the denominator exactly. • JUDGED REMAINDER: the numerator is the sum of the deterministic matches AND the MATCH verdicts of the accepted judged cells. On the un-adjudicated (unresolved) remainder rows the DETERMINISTIC reading stays in force, so it does NOT enter the numerator - the direction of the error is thus the conservative direction. • PAGE-SIDE BOUND: the comparison runs on the STORED visible text; a price or stock status rendered with JavaScript is not present in the stored bytes, and this is named by the no_page_price / no_page_signal sub-code. These are not accusations, but the limits of measurability. • STORED-BYTE INSTRUMENT: every parse runs on the STORED (truncated to 1.5 MB, phone-number-masked) bytes, not on a live page. Because of the truncation a late JSON-LD block may be left out (the html_truncated_stored sub-frame names this), therefore the presence-based numerators are LOWER bounds, the absence-based ones UPPER bounds. • TWO ROUTES: every published cell was produced by two independent routes (record-level, and a derivation re-extracted from the stored HTML), and the two must match exactly, cell by cell; on a mismatch the round writes out a symmetric difference and does NOT write a table.
sealed audit record, Hungarian, verbatim: ÚJ MÉRÉSI ABLAK: ez a sor a 2026-08-13Q körben ÚJRA letöltött oldalakon (F_q) áll. A 08A és a Q ablak közti eltérés VALÓS VILÁGBELI VÁLTOZÁS (az oldal átalakult, megszűnt vagy most utasítja el a kérést), nem mérési hiba; a teljes bontás a Q-drift-* sorokban megy ki, elrejtve semmi nincs. • ÖSSZEHASONLÍTHATÓ KERET: a nevező az adott dimenzió összehasonlítható halmaza (F_q ∩ Product-JSON-LD ∩ a JSON-LD oldal NEM hiányzik). A hiányzó JSON-LD oldalt jelző kétértelmű alkódok a kereten KÍVÜL vannak és külön kerülnek közlésre; a partíció (egyezés + eltérés + kereten belüli kétértelmű) a nevezőt pontosan kiadja. • ÍTÉLT MARADÉK: a számláló a determinisztikus egyezések ÉS az elfogadott ítélt cellák MATCH-verdiktjeinek összege. Az el nem bírált (unresolved) maradék-sorokon a DETERMINISZTIKUS olvasat marad érvényben, tehát a számlálóba NEM kerül be - a hiba iránya így a konzervatív irány. • OLDAL-OLDALI KORLÁT: az összevetés a TÁROLT látható szövegen fut; a JavaScripttel renderelt ár vagy készlet-állapot a tárolt byte-okban nem szerepel, ezt a no_page_price / no_page_signal alkód nevesíti. Ezek nem vádak, hanem a mérhetőség határai. • TÁROLT-BYTE MŰSZER: minden parse a LETÁROLT (1,5 MB-ra vágott, telefonszám-maszkolt) byte-okon fut, nem élő oldalon. A vágás miatt egy késői JSON-LD blokk kimaradhat (ezt a html_truncated_stored alkeret nevesíti), ezért a jelenlét-alapú számlálók ALSÓ korlátok, a hiány-alapúak FELSŐ korlátok. • KÉT ÚT: minden publikált cella két független úton készült (rekord-szintű, illetve a tárolt HTML-ből újra-kinyert deriváció), és a kettőnek cellára pontosan egyeznie kell; eltérés esetén a kör kiír egy szimmetrikus differenciát és NEM ír táblát.
Every residue row — translated rendering: all recurring — judged residue with the conservative direction, honest N, the instrument row reading and the two-route proof; see the English key.
English working translation — keyed to the sealed caveats: JUDGED REMAINDER: the numerator is the sum of the deterministic matches AND the MATCH verdicts of the accepted judged cells. On the un-adjudicated (unresolved) remainder rows the DETERMINISTIC reading stays in force, so it does NOT enter the numerator - the direction of the error is thus the conservative direction. • HONEST N: where the judged stratum did not run to completion, the denominator is the number of ACCEPTED rows and the shortfall is named in the caveat; no tacit filling-in or averaging happens anywhere. • INSTRUMENT ROW: this row is not about the stores but about the MEASURING INSTRUMENT - it describes the error of the judged rows running alongside it, and is to be read TOGETHER with them. • TWO ROUTES: every published cell was produced by two independent routes (record-level, and a derivation re-extracted from the stored HTML), and the two must match exactly, cell by cell; on a mismatch the round writes out a symmetric difference and does NOT write a table.
sealed audit record, Hungarian, verbatim: ÍTÉLT MARADÉK: a számláló a determinisztikus egyezések ÉS az elfogadott ítélt cellák MATCH-verdiktjeinek összege. Az el nem bírált (unresolved) maradék-sorokon a DETERMINISZTIKUS olvasat marad érvényben, tehát a számlálóba NEM kerül be - a hiba iránya így a konzervatív irány. • ŐSZINTE N: ahol az ítélt réteg nem futott végig, a nevező az ELFOGADOTT sorok száma és a hiány a caveatben nevesítve van; hallgatólagos kipótlás vagy átlagolás sehol nem történik. • MŰSZER-SOR: ez a sor nem a webshopokról, hanem a MÉRŐESZKÖZRŐL szól - a mellette futó ítélt sorok hibáját írja le, és azokkal EGYÜTT olvasandó. • KÉT ÚT: minden publikált cella két független úton készült (rekord-szintű, illetve a tárolt HTML-ből újra-kinyert deriváció), és a kettőnek cellára pontosan egyeznie kell; eltérés esetén a kör kiír egy szimmetrikus differenciát és NEM ír táblát.
Machine readability of the page text
| Score | pages | share | direction |
|---|---|---|---|
substantive (Q-W3-substance-2-FQ) | 608 of 798 | 76.2% | point estimate with measured error |
partial (Q-W3-substance-1-FQ) | 97 of 798 | 12.2% | point estimate with measured error |
none (Q-W3-substance-0-FQ) | 93 of 798 | 11.7% | point estimate with measured error |
A seeded blind refuter re-scored a sample of the judged census: exact agreement on 76 of 80 of the sample (a count by family rule, never a rate). English working translation, keyed: INSTRUMENT ROW: this row is not about the stores but about the MEASURING INSTRUMENT - it describes the error of the judged rows running alongside it, and is to be read TOGETHER with them. • INSTRUMENT ERROR MEASURED (W3): on the blind re-scorer sample the EXACT agreement is 76/80, the agreement taken with a one-point tolerance is 80/80; the outcome according to section 4 of the frozen rubric: PASS. This number is to be read alongside EVERY judged W3 row; the deviation is measurement and not a defect. • HONEST N: where the judged stratum did not run to completion, the denominator is the number of ACCEPTED rows and the shortfall is named in the caveat; no tacit filling-in or averaging happens anywhere. • TWO ROUTES: every published cell was produced by two independent routes (record-level, and a derivation re-extracted from the stored HTML), and the two must match exactly, cell by cell; on a mismatch the round writes out a symmetric difference and does NOT write a table. · sealed audit record, Hungarian, verbatim: MŰSZER-SOR: ez a sor nem a webshopokról, hanem a MÉRŐESZKÖZRŐL szól - a mellette futó ítélt sorok hibáját írja le, és azokkal EGYÜTT olvasandó. • MŰSZERHIBA MÉRVE (W3): a vak újraítélő mintán az EGZAKT egyezés 76/80, az egy-pontos tűréssel vett egyezés 80/80; a fagyasztott rubrika 4. szakasza szerinti kimenet: PASS. Ez a szám MINDEN ítélt W3 sor mellett olvasandó, az eltérés mérés és nem defektus. • ŐSZINTE N: ahol az ítélt réteg nem futott végig, a nevező az ELFOGADOTT sorok száma és a hiány a caveatben nevesítve van; hallgatólagos kipótlás vagy átlagolás sehol nem történik. • KÉT ÚT: minden publikált cella két független úton készült (rekord-szintű, illetve a tárolt HTML-ből újra-kinyert deriváció), és a kettőnek cellára pontosan egyeznie kell; eltérés esetén a kör kiír egy szimmetrikus differenciát és NEM ír táblát.
| Check | pages | share | frame | direction |
|---|---|---|---|---|
| coherent heading structure | 444 of 798 | 55.6% | F_q (n=798) | point estimate with measured error |
| healthy answer-sized paragraphs | 69 of 790 | 8.7% | F_q AND at least one eligible paragraph (n=790) | point estimate with measured error |
| no eligible paragraph at all | 8 of 798 | count only - family rule | F_q (n=798) | point estimate with measured error |
Mandatory caveats for the readability block translated rendering · sealed Hungarian original (authoritative) inside
Every substance row — translated rendering: a judged census with measured error: scores come from one-domain-one-cell judgments under the frozen rubric, the denominator is the accepted judged rows (honest N), never the full F_q; the blind re-score sample agreed exactly on 76 of 80 and within one point on 80 of 80 — that instrument error reads beside every judged row, a measurement, not a defect. Recurring: new measurement window · honest N · two routes — see the English key.
English working translation — keyed to the sealed caveats: NEW MEASUREMENT WINDOW: this row stands on the pages RE-fetched in the 2026-08-13Q round (F_q). The difference between the 08A and the Q window is REAL-WORLD CHANGE (the page was transformed, ceased to exist or now refuses the request), not a measurement error; the full breakdown ships in the Q-drift-* rows, nothing is hidden. • JUDGED CENSUS, WITH MEASURED ERROR: the scores are given by one-domain-one-cell judgments based on the frozen rubric; the denominator is the count of the ACCEPTED judged rows (honest N), not the full F_q. The agreement measured on the blind re-scorer sample ships in the Q-W3-refuter-error row and is to be read alongside EVERY judged row. • HONEST N: where the judged stratum did not run to completion, the denominator is the number of ACCEPTED rows and the shortfall is named in the caveat; no tacit filling-in or averaging happens anywhere. • INSTRUMENT ERROR MEASURED (W3): on the blind re-scorer sample the EXACT agreement is 76/80, the agreement taken with a one-point tolerance is 80/80; the outcome according to section 4 of the frozen rubric: PASS. This number is to be read alongside EVERY judged W3 row; the deviation is measurement and not a defect. • TWO ROUTES: every published cell was produced by two independent routes (record-level, and a derivation re-extracted from the stored HTML), and the two must match exactly, cell by cell; on a mismatch the round writes out a symmetric difference and does NOT write a table.
sealed audit record, Hungarian, verbatim: ÚJ MÉRÉSI ABLAK: ez a sor a 2026-08-13Q körben ÚJRA letöltött oldalakon (F_q) áll. A 08A és a Q ablak közti eltérés VALÓS VILÁGBELI VÁLTOZÁS (az oldal átalakult, megszűnt vagy most utasítja el a kérést), nem mérési hiba; a teljes bontás a Q-drift-* sorokban megy ki, elrejtve semmi nincs. • ÍTÉLT CENZUS, MÉRT HIBÁVAL: a pontszámokat egy-domain-egy-cella ítéletek adják a fagyasztott rubrika alapján; a nevező az ELFOGADOTT ítélt sorok száma (őszinte N), nem a teljes F_q. A vak újraítélő mintán mért egyezés a Q-W3-refuter-error sorban megy ki és MINDEN ítélt sor mellett olvasandó. • ŐSZINTE N: ahol az ítélt réteg nem futott végig, a nevező az ELFOGADOTT sorok száma és a hiány a caveatben nevesítve van; hallgatólagos kipótlás vagy átlagolás sehol nem történik. • MŰSZERHIBA MÉRVE (W3): a vak újraítélő mintán az EGZAKT egyezés 76/80, az egy-pontos tűréssel vett egyezés 80/80; a fagyasztott rubrika 4. szakasza szerinti kimenet: PASS. Ez a szám MINDEN ítélt W3 sor mellett olvasandó, az eltérés mérés és nem defektus. • KÉT ÚT: minden publikált cella két független úton készült (rekord-szintű, illetve a tárolt HTML-ből újra-kinyert deriváció), és a kettőnek cellára pontosan egyeznie kell; eltérés esetén a kör kiír egy szimmetrikus differenciát és NEM ír táblát.
The heading row — translated rendering: heading coherence means exactly one top-level heading and gap-free heading levels from the top, over the whole stored document — navigation and footer blocks included; the frozen instrument's literal reading, not an oversight. Recurring: new measurement window · single-source HTML cell · two routes — see the English key.
English working translation — keyed to the sealed caveats: NEW MEASUREMENT WINDOW: this row stands on the pages RE-fetched in the 2026-08-13Q round (F_q). The difference between the 08A and the Q window is REAL-WORLD CHANGE (the page was transformed, ceased to exist or now refuses the request), not a measurement error; the full breakdown ships in the Q-drift-* rows, nothing is hidden. • DEFINITION OF HEADING COHERENCE: exactly one h1 AND the heading levels present are gap-free starting from one. The rule applies to the WHOLE stored document, so headings standing in the navigation and footer blocks also count - this is the literal reading of the frozen instrument, not a blunder. • SINGLE-SOURCE HTML CELL: this row reads the same stored product.html bytes on both routes through the same frozen function, therefore its real second leg is the separate CPython recount artefact; the seam is named here, not concealed. • TWO ROUTES: every published cell was produced by two independent routes (record-level, and a derivation re-extracted from the stored HTML), and the two must match exactly, cell by cell; on a mismatch the round writes out a symmetric difference and does NOT write a table.
sealed audit record, Hungarian, verbatim: ÚJ MÉRÉSI ABLAK: ez a sor a 2026-08-13Q körben ÚJRA letöltött oldalakon (F_q) áll. A 08A és a Q ablak közti eltérés VALÓS VILÁGBELI VÁLTOZÁS (az oldal átalakult, megszűnt vagy most utasítja el a kérést), nem mérési hiba; a teljes bontás a Q-drift-* sorokban megy ki, elrejtve semmi nincs. • CÍMSOR-KOHERENCIA DEFINÍCIÓJA: pontosan egy h1 ÉS a jelen lévő címsor-szintek egytől kezdve hézagmentesek. A szabály a TELJES tárolt dokumentumra vonatkozik, tehát a navigációs és lábléc-blokkokban álló címsorok is számítanak - ez a fagyasztott műszer szó szerinti olvasata, nem melléfogás. • EGYFORRÁSÚ HTML-CELLA: ez a sor mindkét úton ugyanazokat a tárolt product.html byte-okat olvassa ugyanazon fagyasztott függvényen keresztül, ezért a valódi második lába a külön CPython recount-artefaktum; a varrat itt nevesítve van, nem elhallgatva. • KÉT ÚT: minden publikált cella két független úton készült (rekord-szintű, illetve a tárolt HTML-ből újra-kinyert deriváció), és a kettőnek cellára pontosan egyeznie kell; eltérés esetén a kör kiír egy szimmetrikus differenciát és NEM ír táblát.
The paragraph-health row — translated rendering: a paragraph is a block-level unit of the stored visible text; eligible means at least five words, in-band means two to four sentences; the sentence splitter knows no abbreviation list — a title abbreviation splits a sentence — and that bias is uniform across the census. The paragraph-serialization seam between the two routes was named in the pre-registration in advance; on disagreement the round stops rather than smoothing it. Recurring: new measurement window · stored bytes · two routes — see the English key.
English working translation — keyed to the sealed caveats: NEW MEASUREMENT WINDOW: this row stands on the pages RE-fetched in the 2026-08-13Q round (F_q). The difference between the 08A and the Q window is REAL-WORLD CHANGE (the page was transformed, ceased to exist or now refuses the request), not a measurement error; the full breakdown ships in the Q-drift-* rows, nothing is hidden. • DEFINITION OF CHUNK HEALTH: paragraph = a block-level unit of the stored visible text; eligible paragraph = at least five words; within band = the sentence count between two and four. The sentence splitter knows no abbreviation list ('Dr.' cuts), the distortion is uniform across the whole census. Domains with zero eligible paragraphs form a separate sub-frame without a rate. • PARAGRAPH-SERIALIZATION SEAM: one route obtains the paragraph list from the blank-line-separated segmentation of the stored visible_text.txt, the other from the re-extraction of product.html; if a paragraph itself contains a blank line, the two segmentations may differ. The seam is named IN ADVANCE in the preregistration, and on divergence the round stops, it does not smooth it over. • STORED-BYTE INSTRUMENT: every parse runs on the STORED (truncated to 1.5 MB, phone-number-masked) bytes, not on a live page. Because of the truncation a late JSON-LD block may be left out (the html_truncated_stored sub-frame names this), therefore the presence-based numerators are LOWER bounds, the absence-based ones UPPER bounds. • TWO ROUTES: every published cell was produced by two independent routes (record-level, and a derivation re-extracted from the stored HTML), and the two must match exactly, cell by cell; on a mismatch the round writes out a symmetric difference and does NOT write a table.
sealed audit record, Hungarian, verbatim: ÚJ MÉRÉSI ABLAK: ez a sor a 2026-08-13Q körben ÚJRA letöltött oldalakon (F_q) áll. A 08A és a Q ablak közti eltérés VALÓS VILÁGBELI VÁLTOZÁS (az oldal átalakult, megszűnt vagy most utasítja el a kérést), nem mérési hiba; a teljes bontás a Q-drift-* sorokban megy ki, elrejtve semmi nincs. • CHUNK-EGÉSZSÉG DEFINÍCIÓJA: bekezdés = a tárolt látható szöveg egy blokk-szintű egysége; alkalmas bekezdés = legalább öt szó; sávon belül = a mondatszám kettő és négy között. A mondathatároló nem ismer rövidítés-listát (a 'Dr.' vág), a torzítás a teljes cenzuson egyforma. A nulla alkalmas bekezdésű domainek külön, arány nélküli alkeretet alkotnak. • BEKEZDÉS-SZERIALIZÁCIÓS VARRAT: a bekezdés-listát az egyik út a tárolt visible_text.txt üres sorral elválasztott tagolásából, a másik a product.html újra-kinyeréséből kapja; ha egy bekezdés maga tartalmaz üres sort, a két tagolás eltérhet. A varrat a preregisztrációban ELŐRE nevesítve van, és eltérés esetén a kör megáll, nem simítja el. • TÁROLT-BYTE MŰSZER: minden parse a LETÁROLT (1,5 MB-ra vágott, telefonszám-maszkolt) byte-okon fut, nem élő oldalon. A vágás miatt egy késői JSON-LD blokk kimaradhat (ezt a html_truncated_stored alkeret nevesíti), ezért a jelenlét-alapú számlálók ALSÓ korlátok, a hiány-alapúak FELSŐ korlátok. • KÉT ÚT: minden publikált cella két független úton készült (rekord-szintű, illetve a tárolt HTML-ből újra-kinyert deriváció), és a kettőnek cellára pontosan egyeznie kell; eltérés esetén a kör kiír egy szimmetrikus differenciát és NEM ír táblát.
The zero-eligible row — translated rendering: domains with no eligible paragraph at all form a separate, rate-free sub-frame under the same paragraph definition. Recurring: new measurement window · stored bytes · two routes — see the English key.
English working translation — keyed to the sealed caveats: NEW MEASUREMENT WINDOW: this row stands on the pages RE-fetched in the 2026-08-13Q round (F_q). The difference between the 08A and the Q window is REAL-WORLD CHANGE (the page was transformed, ceased to exist or now refuses the request), not a measurement error; the full breakdown ships in the Q-drift-* rows, nothing is hidden. • DEFINITION OF CHUNK HEALTH: paragraph = a block-level unit of the stored visible text; eligible paragraph = at least five words; within band = the sentence count between two and four. The sentence splitter knows no abbreviation list ('Dr.' cuts), the distortion is uniform across the whole census. Domains with zero eligible paragraphs form a separate sub-frame without a rate. • STORED-BYTE INSTRUMENT: every parse runs on the STORED (truncated to 1.5 MB, phone-number-masked) bytes, not on a live page. Because of the truncation a late JSON-LD block may be left out (the html_truncated_stored sub-frame names this), therefore the presence-based numerators are LOWER bounds, the absence-based ones UPPER bounds. • TWO ROUTES: every published cell was produced by two independent routes (record-level, and a derivation re-extracted from the stored HTML), and the two must match exactly, cell by cell; on a mismatch the round writes out a symmetric difference and does NOT write a table.
sealed audit record, Hungarian, verbatim: ÚJ MÉRÉSI ABLAK: ez a sor a 2026-08-13Q körben ÚJRA letöltött oldalakon (F_q) áll. A 08A és a Q ablak közti eltérés VALÓS VILÁGBELI VÁLTOZÁS (az oldal átalakult, megszűnt vagy most utasítja el a kérést), nem mérési hiba; a teljes bontás a Q-drift-* sorokban megy ki, elrejtve semmi nincs. • CHUNK-EGÉSZSÉG DEFINÍCIÓJA: bekezdés = a tárolt látható szöveg egy blokk-szintű egysége; alkalmas bekezdés = legalább öt szó; sávon belül = a mondatszám kettő és négy között. A mondathatároló nem ismer rövidítés-listát (a 'Dr.' vág), a torzítás a teljes cenzuson egyforma. A nulla alkalmas bekezdésű domainek külön, arány nélküli alkeretet alkotnak. • TÁROLT-BYTE MŰSZER: minden parse a LETÁROLT (1,5 MB-ra vágott, telefonszám-maszkolt) byte-okon fut, nem élő oldalon. A vágás miatt egy késői JSON-LD blokk kimaradhat (ezt a html_truncated_stored alkeret nevesíti), ezért a jelenlét-alapú számlálók ALSÓ korlátok, a hiány-alapúak FELSŐ korlátok. • KÉT ÚT: minden publikált cella két független úton készült (rekord-szintű, illetve a tárolt HTML-ből újra-kinyert deriváció), és a kettőnek cellára pontosan egyeznie kell; eltérés esetén a kör kiír egy szimmetrikus differenciát és NEM ír táblát.
Authorship of llms.txt — the blind-gated successor
Part One refused to answer the authorship question: its classifier failed the audit gate — the negative branch did not mean its label — so no figure shipped. This round rebuilt the instrument from scratch and gated it on exactly that failure mode: known negatives, presented blind, and none of them came back labelled authored. Only after that pass do these rows ship, under new row keys; the failed Part One classifier stays held back, unrevived.
(35 of 696 · frozen Part One corpus, re-read by a new instrument · blind-audit gated)
| Class | files | share | direction |
|---|---|---|---|
| merchant-authored | 35 of 696 | 5.0% | point estimate with measured error |
| platform boilerplate | 659 of 696 | 94.7% | point estimate with measured error |
| error page served as llms.txt | 2 of 696 | n<20 - count only | point estimate with measured error |
On the genuine rows of the blind draw, exact three-way agreement was 19 of 20 (a count by family rule); the known-negative rows of the draw are recorded in the sealed record beside — none of them came back labelled authored. English working translation, keyed: INSTRUMENT ROW: this row is not about the stores but about the MEASURING INSTRUMENT - it describes the error of the judged rows running alongside it, and is to be read TOGETHER with them. • BLIND AUDIT MEASURED (W4): of the known negatives, 0 received the 'authored' blind label (B1), on the real rows the exact three-way agreement is 19/20 (B2); the outcome per section 4 of the frozen protocol: PASS. The B3 taxonomy-swap count: 0, exclusively a note. • HONEST N: where the judged stratum did not run to completion, the denominator is the number of ACCEPTED rows and the shortfall is named in the caveat; no tacit filling-in or averaging happens anywhere. • TWO ROUTES: every published cell was produced by two independent routes (record-level, and a derivation re-extracted from the stored HTML), and the two must match exactly, cell by cell; on a mismatch the round writes out a symmetric difference and does NOT write a table. · sealed audit record, Hungarian, verbatim: MŰSZER-SOR: ez a sor nem a webshopokról, hanem a MÉRŐESZKÖZRŐL szól - a mellette futó ítélt sorok hibáját írja le, és azokkal EGYÜTT olvasandó. • VAK AUDIT MÉRVE (W4): az ismert negatívok közül 0 kapott 'szerzői' vak címkét (B1), a valódi sorokon az egzakt háromutas egyezés 19/20 (B2); a fagyasztott protokoll 4. szakasza szerinti kimenet: PASS. A B3 taxonómia-csere darabszáma: 0, kizárólag jegyzet. • ŐSZINTE N: ahol az ítélt réteg nem futott végig, a nevező az ELFOGADOTT sorok száma és a hiány a caveatben nevesítve van; hallgatólagos kipótlás vagy átlagolás sehol nem történik. • KÉT ÚT: minden publikált cella két független úton készült (rekord-szintű, illetve a tárolt HTML-ből újra-kinyert deriváció), és a kettőnek cellára pontosan egyeznie kell; eltérés esetén a kör kiír egy szimmetrikus differenciát és NEM ír táblát.
23 of 35 of the judged-authored files, on the sealed conclusive robots frame (65.7%), sit on stores whose robots.txt is silent about every measured AI crawler — even among stores that took the deliberate step of authoring an llms.txt, robots.txt mostly says nothing about the measured AI crawlers either way. Direction: point estimate with measured error.
| Origin | files | note |
|---|---|---|
| Shopify template | 546 of 696 | count only - family rule |
| Yoast family | 28 of 696 | count only - family rule |
| Rank Math family | 21 of 696 | count only - family rule |
| AIOSEO family | 13 of 696 | count only - family rule |
| Hostinger family | 25 of 696 | count only - family rule |
| Wix family | 18 of 696 | count only - family rule |
| duplicate cluster | 0 of 696 | count only - family rule |
| deterministic error page | 2 of 696 | count only - family rule |
| judged: authored | 35 of 696 | count only - family rule |
| judged: boilerplate | 8 of 696 | count only - family rule |
| judged: error page | 0 of 696 | count only - family rule |
| judged: unresolved | 0 of 696 | count only - family rule |
Mandatory caveats for the authorship block translated rendering · sealed Hungarian original (authoritative) inside
Every three-way row — translated rendering: the blind-audit outcome these rows passed is rendered in the instrument callout above, with its sealed record. Recurring: rebuilt authorship instrument · frozen corpus, no re-fetch · duplicate-cluster limit · two routes — see the English key.
English working translation — keyed to the sealed caveats: REBUILD OF THE S06 FAILURE: the negative branch of the Part 1 classifier was labeled 'authored', and per the census the overwhelming majority of it was another generator's template or an error page. In the new instrument 'authored' is a POSITIVE finding: it requires the citation of merchant-specific content, the marker-absence by itself is NEVER enough. This row is a NEW Q-key; it does not overwrite and does not resurrect the sealed S06/S16 rows. • FROZEN CORPUS, NO RE-FETCH: the input is the set of llms.txt bodies stored in the 2026-08-08A round, with a sha pin per file. The published breakdown is therefore a RE-READING of the 08A window with a new instrument; it asserts nothing about today's state. • DUPLICATE-CLUSTER BOUND: the clustering is built on EXACT hash match (lower-cased, whitespace-normalized body); the recognition of near-duplicates is omitted by name, therefore the template-page count is a LOWER bound, that of the authored pages an UPPER bound. MEASURED CLUSTER COUNT: 0 clusters, 0 affected domains, judged remainder 43 bodies. • BLIND AUDIT MEASURED (W4): of the known negatives, 0 received the 'authored' blind label (B1), on the real rows the exact three-way agreement is 19/20 (B2); the outcome per section 4 of the frozen protocol: PASS. The B3 taxonomy-swap count: 0, exclusively a note. • TWO ROUTES: every published cell was produced by two independent routes (record-level, and a derivation re-extracted from the stored HTML), and the two must match exactly, cell by cell; on a mismatch the round writes out a symmetric difference and does NOT write a table.
sealed audit record, Hungarian, verbatim: AZ S06-BUKÁS ÚJRAÉPÍTÉSE: a Part 1 osztályozó negatív ága 'szerzői'-nek volt címkézve, és a cenzus szerint annak túlnyomó része más generátor sablonja vagy hibaoldal volt. Az új műszerben a 'szerzői' POZITÍV lelet: kereskedő-specifikus tartalom idézését követeli, a marker-hiány önmagában SOHA nem elég. Ez a sor ÚJ Q-kulcs; a lepecsételt S06/S16 sorokat nem írja felül és nem támasztja fel. • FAGYASZTOTT KORPUSZ, NINCS ÚJRALETÖLTÉS: a bemenet a 2026-08-08A körben letárolt llms.txt törzsek halmaza, fájlonként sha-pinnel. A publikált bontás tehát a 08A ablak ÚJRAOLVASATA új műszerrel; a mai állapotról nem állít semmit. • DUPLIKÁTUM-KLASZTER KORLÁTJA: a klaszterezés PONTOS hash-egyezésre épül (kisbetűsített, whitespace-normalizált törzs); a közel-duplikátumok felismerése nevesítetten kimarad, ezért a sablon-oldal darabszáma ALSÓ korlát, a szerzői oldalaké FELSŐ korlát. MÉRT KLASZTER-SZÁM: 0 klaszter, 0 érintett domain, ítélt maradék 43 törzs. • VAK AUDIT MÉRVE (W4): az ismert negatívok közül 0 kapott 'szerzői' vak címkét (B1), a valódi sorokon az egzakt háromutas egyezés 19/20 (B2); a fagyasztott protokoll 4. szakasza szerinti kimenet: PASS. A B3 taxonómia-csere darabszáma: 0, kizárólag jegyzet. • KÉT ÚT: minden publikált cella két független úton készült (rekord-szintű, illetve a tárolt HTML-ből újra-kinyert deriváció), és a kettőnek cellára pontosan egyeznie kell; eltérés esetén a kör kiír egy szimmetrikus differenciát és NEM ír táblát.
Every provenance row — translated rendering: the deterministic fingerprint layer's counts are code-derived with their own two-route proof and ship independently of the blind-audit outcome; the judged sub-classes ship separately under a judged prefix. Recurring: frozen corpus, no re-fetch · duplicate-cluster limit · two routes — see the English key.
English working translation — keyed to the sealed caveats: FROZEN CORPUS, NO RE-FETCH: the input is the set of llms.txt bodies stored in the 2026-08-08A round, with a sha pin per file. The published breakdown is therefore a RE-READING of the 08A window with a new instrument; it asserts nothing about today's state. • DUPLICATE-CLUSTER BOUND: the clustering is built on EXACT hash match (lower-cased, whitespace-normalized body); the recognition of near-duplicates is omitted by name, therefore the template-page count is a LOWER bound, that of the authored pages an UPPER bound. MEASURED CLUSTER COUNT: 0 clusters, 0 affected domains, judged remainder 43 bodies. • NUMERATOR ROW INDEPENDENTLY OF THE GATE: the numbers of the deterministic fingerprint stratum are code-derived and have their own two-route evidence, therefore they ship independently of the outcome of the blind audit; the judged subclasses appear separated from this, with the judged_ prefix. • TWO ROUTES: every published cell was produced by two independent routes (record-level, and a derivation re-extracted from the stored HTML), and the two must match exactly, cell by cell; on a mismatch the round writes out a symmetric difference and does NOT write a table.
sealed audit record, Hungarian, verbatim: FAGYASZTOTT KORPUSZ, NINCS ÚJRALETÖLTÉS: a bemenet a 2026-08-08A körben letárolt llms.txt törzsek halmaza, fájlonként sha-pinnel. A publikált bontás tehát a 08A ablak ÚJRAOLVASATA új műszerrel; a mai állapotról nem állít semmit. • DUPLIKÁTUM-KLASZTER KORLÁTJA: a klaszterezés PONTOS hash-egyezésre épül (kisbetűsített, whitespace-normalizált törzs); a közel-duplikátumok felismerése nevesítetten kimarad, ezért a sablon-oldal darabszáma ALSÓ korlát, a szerzői oldalaké FELSŐ korlát. MÉRT KLASZTER-SZÁM: 0 klaszter, 0 érintett domain, ítélt maradék 43 törzs. • SZÁMLÁLÓ-SOR A KAPUTÓL FÜGGETLENÜL: a determinisztikus ujjlenyomat-réteg számai kód-deriváltak és saját két-utas bizonyítékuk van, ezért a vak-audit kimenetelétől függetlenül kimennek; az ítélt alosztályok ettől elkülönítve, judged_ előtaggal szerepelnek. • KÉT ÚT: minden publikált cella két független úton készült (rekord-szintű, illetve a tárolt HTML-ből újra-kinyert deriváció), és a kettőnek cellára pontosan egyeznie kell; eltérés esetén a kör kiír egy szimmetrikus differenciát és NEM ír táblát.
The crawler-silence cross — translated rendering: the denominator is the intersection of the judged-authored bodies with the sealed conclusive robots frame; the robots side comes unchanged from the parent round's checkpoint pass, and the two passes never mix. Recurring: robots frames and passes · rebuilt authorship instrument · frozen corpus · duplicate-cluster limit · two routes — see the English key; the blind audit is rendered in the callout above.
English working translation — keyed to the sealed caveats: W1 ADJUDICATION (R5): PRIMARY is the 901/975 = 92.41% (F_robots_conclusive); SECONDARY is the 901/981 = 91.85% census form with a conclusive-frame caveat. Decisive argument: the 975 is not a correction AGAINST the engine - it is the engine's OWN frame (build_robots_matrix 1434-1460 narrows to present|absent_4xx|absent_softpage rows; matrix.frame_robots_conclusive=975). The den=981 exists only in the offline tier1. • FORBIDDEN COMPLEMENT STATEMENT: the form "901 silent, therefore the rest are explicit" is FALSE on the 981 frame - 981-901=80, but only 74 are explicit; 6 rows fall into neither category (unmeasured for all 7 agents). The 6: pinkavenue.store, kapsliving.com, gaminguae.ae, pcdubai.com, fmmdubai.com, daima.ae. This sentence cannot be published in any form. • On the 975 frame the complement is mathematically exact (901+74=975, verified in-cell), but it is still to be published as a MEASURED count (74/975), NEVER derived by subtraction from 981. This in itself is an argument for the conclusive frame. • ENGINE-vs-ARTIFACT CONTRADICTION RESOLVED by per-row frame naming: of the tier1's 15 rows, 13 use den=981 and 2 (blanket_disallow, sitemap_declared) use den=975; this row is one of the 11 verdict-derived ones. The robots_matrix publishes THE SAME numerators on 975. The numerators agree everywhere - the divergence is purely a denominator choice, not arithmetic (tier1_rate_arithmetic 15/15 PASS). • REBUILD OF THE S06 FAILURE: the negative branch of the Part 1 classifier was labeled 'authored', and per the census the overwhelming majority of it was another generator's template or an error page. In the new instrument 'authored' is a POSITIVE finding: it requires the citation of merchant-specific content, the marker-absence by itself is NEVER enough. This row is a NEW Q-key; it does not overwrite and does not resurrect the sealed S06/S16 rows. • FROZEN CORPUS, NO RE-FETCH: the input is the set of llms.txt bodies stored in the 2026-08-08A round, with a sha pin per file. The published breakdown is therefore a RE-READING of the 08A window with a new instrument; it asserts nothing about today's state. • DUPLICATE-CLUSTER BOUND: the clustering is built on EXACT hash match (lower-cased, whitespace-normalized body); the recognition of near-duplicates is omitted by name, therefore the template-page count is a LOWER bound, that of the authored pages an UPPER bound. MEASURED CLUSTER COUNT: 0 clusters, 0 affected domains, judged remainder 43 bodies. • BLIND AUDIT MEASURED (W4): of the known negatives, 0 received the 'authored' blind label (B1), on the real rows the exact three-way agreement is 19/20 (B2); the outcome per section 4 of the frozen protocol: PASS. The B3 taxonomy-swap count: 0, exclusively a note. • TWO ROUTES: every published cell was produced by two independent routes (record-level, and a derivation re-extracted from the stored HTML), and the two must match exactly, cell by cell; on a mismatch the round writes out a symmetric difference and does NOT write a table. • CROSS-ROW FRAME: the denominator is the intersection of the bodies judged authored AND the sealed, conclusive robots frame; the robots side comes unchanged from the parent round's checkpoint pass. The two passes (sites vs checkpoint) are NEVER to be mixed - this row names the checkpoint pass.
sealed audit record, Hungarian, verbatim: W1 ADJUDIKÁCIÓ (R5): ELSŐDLEGES a 901/975 = 92,41% (F_robots_conclusive); MÁSODLAGOS a 901/981 = 91,85% census-alak conclusive-frame caveattel. Döntő érv: a 975 nem korrekció a motorral SZEMBEN - a motor SAJÁT kerete (build_robots_matrix 1434-1460 present|absent_4xx|absent_softpage sorokra szűkít; matrix.frame_robots_conclusive=975). A den=981 csak az offline tier1-ben létezik. • TILTOTT KOMPLEMENTER-ÁLLÍTÁS: a "901 néma, tehát a többi explicit" forma a 981-es kereten HAMIS - 981-901=80, de csak 74 explicit; 6 sor egyik kategóriába sem esik (mind a 7 ügynökre unmeasured). A 6: pinkavenue.store, kapsliving.com, gaminguae.ae, pcdubai.com, fmmdubai.com, daima.ae. Ez a mondat semmilyen alakban nem publikálható. • A 975-ös kereten a komplementer matematikailag pontos (901+74=975, in-cell igazolva), de akkor is MÉRT darabszámként (74/975) közlendő, SOHA nem 981-ből kivonással levezetve. Ez önmagában is a konklúzív keret melletti érv. • MOTOR-vs-ARTEFAKTUM ELLENTMONDÁS FELOLDVA soronkénti keret-megnevezéssel: a tier1 15 sorából 13 den=981-et, 2 (blanket_disallow, sitemap_declared) den=975-öt használ; ez a sor a 11 verdikt-derivált egyike. A robots_matrix UGYANEZEKET a számlálókat 975-ön közli. A számlálók mindenhol egyeznek - a szakadás tisztán nevező-választás, nem aritmetika (tier1_rate_arithmetic 15/15 PASS). • AZ S06-BUKÁS ÚJRAÉPÍTÉSE: a Part 1 osztályozó negatív ága 'szerzői'-nek volt címkézve, és a cenzus szerint annak túlnyomó része más generátor sablonja vagy hibaoldal volt. Az új műszerben a 'szerzői' POZITÍV lelet: kereskedő-specifikus tartalom idézését követeli, a marker-hiány önmagában SOHA nem elég. Ez a sor ÚJ Q-kulcs; a lepecsételt S06/S16 sorokat nem írja felül és nem támasztja fel. • FAGYASZTOTT KORPUSZ, NINCS ÚJRALETÖLTÉS: a bemenet a 2026-08-08A körben letárolt llms.txt törzsek halmaza, fájlonként sha-pinnel. A publikált bontás tehát a 08A ablak ÚJRAOLVASATA új műszerrel; a mai állapotról nem állít semmit. • DUPLIKÁTUM-KLASZTER KORLÁTJA: a klaszterezés PONTOS hash-egyezésre épül (kisbetűsített, whitespace-normalizált törzs); a közel-duplikátumok felismerése nevesítetten kimarad, ezért a sablon-oldal darabszáma ALSÓ korlát, a szerzői oldalaké FELSŐ korlát. MÉRT KLASZTER-SZÁM: 0 klaszter, 0 érintett domain, ítélt maradék 43 törzs. • VAK AUDIT MÉRVE (W4): az ismert negatívok közül 0 kapott 'szerzői' vak címkét (B1), a valódi sorokon az egzakt háromutas egyezés 19/20 (B2); a fagyasztott protokoll 4. szakasza szerinti kimenet: PASS. A B3 taxonómia-csere darabszáma: 0, kizárólag jegyzet. • KÉT ÚT: minden publikált cella két független úton készült (rekord-szintű, illetve a tárolt HTML-ből újra-kinyert deriváció), és a kettőnek cellára pontosan egyeznie kell; eltérés esetén a kör kiír egy szimmetrikus differenciát és NEM ír táblát. • KERESZT-SOR KERETE: a nevező a szerzőinek ítélt törzsek ÉS a lepecsételt, konklúzív robots-keret metszete; a robots-oldal a szülő-kör checkpoint-passzából jön változatlanul. A két pass (sites vs checkpoint) SOHA nem keverendő - ez a sor a checkpoint-passzt nevezi meg.
The blind-audit row — translated rendering: an instrument row about the classifier, not the shops — its measured outcome is rendered in the callout above; the taxonomy-swap note recorded no case. Recurring: instrument row · honest N · two routes — see the English key.
English working translation — keyed to the sealed caveats: INSTRUMENT ROW: this row is not about the stores but about the MEASURING INSTRUMENT - it describes the error of the judged rows running alongside it, and is to be read TOGETHER with them. • BLIND AUDIT MEASURED (W4): of the known negatives, 0 received the 'authored' blind label (B1), on the real rows the exact three-way agreement is 19/20 (B2); the outcome per section 4 of the frozen protocol: PASS. The B3 taxonomy-swap count: 0, exclusively a note. • HONEST N: where the judged stratum did not run to completion, the denominator is the number of ACCEPTED rows and the shortfall is named in the caveat; no tacit filling-in or averaging happens anywhere. • TWO ROUTES: every published cell was produced by two independent routes (record-level, and a derivation re-extracted from the stored HTML), and the two must match exactly, cell by cell; on a mismatch the round writes out a symmetric difference and does NOT write a table.
sealed audit record, Hungarian, verbatim: MŰSZER-SOR: ez a sor nem a webshopokról, hanem a MÉRŐESZKÖZRŐL szól - a mellette futó ítélt sorok hibáját írja le, és azokkal EGYÜTT olvasandó. • VAK AUDIT MÉRVE (W4): az ismert negatívok közül 0 kapott 'szerzői' vak címkét (B1), a valódi sorokon az egzakt háromutas egyezés 19/20 (B2); a fagyasztott protokoll 4. szakasza szerinti kimenet: PASS. A B3 taxonómia-csere darabszáma: 0, kizárólag jegyzet. • ŐSZINTE N: ahol az ítélt réteg nem futott végig, a nevező az ELFOGADOTT sorok száma és a hiány a caveatben nevesítve van; hallgatólagos kipótlás vagy átlagolás sehol nem történik. • KÉT ÚT: minden publikált cella két független úton készült (rekord-szintű, illetve a tárolt HTML-ből újra-kinyert deriváció), és a kettőnek cellára pontosan egyeznie kell; eltérés esetén a kör kiír egy szimmetrikus differenciát és NEM ír táblát.
Category cross-check — disclosed-only
Sealed publication rule for this axis, verbatim: disclosed-only, never a headline
— a companion breakdown on the new-window frame, never a headline. The ratio in every cell names that row's own per-category denominator.
| Category | stores | share | direction |
|---|---|---|---|
| fashion | 176 of 211 | 83.4% | lower bound |
| food & beverage | 93 of 116 | 80.2% | lower bound |
| home & furniture | 89 of 121 | 73.6% | lower bound |
| other | 97 of 121 | 80.2% | lower bound |
| beauty | 100 of 123 | 81.3% | lower bound |
| electronics | 88 of 106 | 83.0% | lower bound |
Mandatory caveats for the category cross-check translated rendering · sealed Hungarian original (authoritative) inside
Every category row — translated rendering: all recurring — the measured bound, the named nested-price limit, the true complement, denominator purity, the new measurement window, the secondary disclosed-only axis, the self-selected frame and the two-route proof; see the English key.
English working translation — keyed to the sealed caveats: MEASURED BOUND (R1, from the adjudication): merchant_valid=TRUE stratum (C01, stratum size 626) of 5 examined rows 5 upheld = 0 false positives; in the presence-without-validity strata, of 10 examined rows 3 flipped to VALID (C02 2/5, stratum size 17; C03 1/5, stratum size 46), all three the same error pattern. Because of k=5 (n<20), ONLY a count, we do not publish a rate. • (a) NAMED BOUND, verified in code: product_rubric reads only the Offer/node-level price|lowPrice key (run_geo_audit.py:494-499), not offers[].priceSpecification[].price. Adjudicated quote (timeoutx.ae): '"offers":[{"@type":"Offer","priceSpecification":[{"@type":"UnitPriceSpecification","price":"200","priceCurrency":"AED"'. Because of this, the 626/807 is a LOWER BOUND. • TRUE COMPLEMENT: on F_prod 807-626 = 181 rows (22.43%) are not valid; this is exactly 118 (14.62%) without Product JSON-LD + 63 (7.81%) presence-without-validity, 118+63=181 closes. The 9.14% = 63/689 shown in the header stands on a DIFFERENT denominator (the base of those publishing JSON-LD); the two numbers are not interchangeable, and neither is the complement of merchant_valid. • DENOMINATOR PURITY (adjudicated, C04 stratum n=118): of 5 examined rows 2 were WooCommerce /shop/ ARCHIVE pages, which the weak cart_marker branch validated as product pages (ajmanfurniture.com, tealand.ae) — the F_prod=807 contains non-product pages, which can never be merchant_valid. Direction: removing them WOULD RAISE the rate, in the same direction as bound (a). • NEW MEASUREMENT WINDOW: this row stands on the pages RE-fetched in the 2026-08-13Q round (F_q). The difference between the 08A and the Q window is REAL-WORLD CHANGE (the page was transformed, ceased to exist or now refuses the request), not a measurement error; the full breakdown ships in the Q-drift-* rows, nothing is hidden. • SECONDARY AXIS: the publication of the category breakdown is DISCLOSED-ONLY - an accompanying breakdown, never a headline number. The axis is the census column (targets_981.csv#category), the record-level category field serves only for cross-check, in case of discrepancy the census wins and the discrepancy is reported, merging NEVER. • SELF-SELECTION FRAME: F_q consists of the domains that were successfully fetched AND parsed in THIS window; the dropouts (403/429/gone/timeout) are not missing by random sampling, therefore the rate measured on F_q cannot be automatically carried over to the full input frame. • TWO ROUTES: every published cell was produced by two independent routes (record-level, and a derivation re-extracted from the stored HTML), and the two must match exactly, cell by cell; on a mismatch the round writes out a symmetric difference and does NOT write a table.
sealed audit record, Hungarian, verbatim: MÉRT KORLÁT (R1, az adjudikációból): merchant_valid=IGAZ réteg (C01, rétegméret 626) 5 vizsgált sorból 5 helybenhagyva = 0 hamis pozitív; a jelenlét-érvényesség-nélküli rétegekben 10 vizsgált sorból 3 fordult ÉRVÉNYESRE (C02 2/5, rétegméret 17; C03 1/5, rétegméret 46), mindhárom ugyanaz a hibaminta. k=5 (n<20) miatt CSAK darabszám, arányt nem közlünk. • (a) NEVESÍTETT KORLÁT, kódban igazolva: product_rubric csak Offer/node szintű price|lowPrice kulcsot olvas (run_geo_audit.py:494-499), az offers[].priceSpecification[].price-t nem. Adjudikált idézet (timeoutx.ae): '"offers":[{"@type":"Offer","priceSpecification":[{"@type":"UnitPriceSpecification","price":"200","priceCurrency":"AED"'. Emiatt a 626/807 ALSÓ KORLÁT. • IGAZ KOMPLEMENTER: F_prod-on 807-626 = 181 sor (22,43%) nem érvényes; ez pontosan 118 (14,62%) Product JSON-LD nélküli + 63 (7,81%) jelenlét-érvényesség-nélkül, 118+63=181 zár. A fejlécben szereplő 9,14% = 63/689 MÁS nevezőn (a JSON-LD-t közlők bázisa) áll; a két szám nem cserélhető fel, és egyik sem a merchant_valid komplementere. • NEVEZŐ-TISZTASÁG (adjudikált, C04 réteg n=118): 5 vizsgált sorból 2 WooCommerce /shop/ ARCHÍVUM oldal volt, amit a gyenge cart_marker-ág validált termékoldalként (ajmanfurniture.com, tealand.ae) — az F_prod=807 tartalmaz nem-termékoldalakat, amelyek sosem lehetnek merchant_valid. Irány: eltávolításuk EMELNÉ az arányt, azonos irányba az (a) korláttal. • ÚJ MÉRÉSI ABLAK: ez a sor a 2026-08-13Q körben ÚJRA letöltött oldalakon (F_q) áll. A 08A és a Q ablak közti eltérés VALÓS VILÁGBELI VÁLTOZÁS (az oldal átalakult, megszűnt vagy most utasítja el a kérést), nem mérési hiba; a teljes bontás a Q-drift-* sorokban megy ki, elrejtve semmi nincs. • MÁSODLAGOS TENGELY: a kategória-bontás közzététele DISCLOSED-ONLY - kísérő bontás, sosem főcím-szám. A tengely a cenzus-oszlop (targets_981.csv#category), a rekord-szintű kategória-mező csak keresztellenőrzésre szolgál, eltérés esetén a cenzus nyer és az eltérés közlésre kerül, összefésülés SOHA. • ÖNSZELEKCIÓS KERET: az F_q azokból a domainekből áll, amelyek EBBEN az ablakban sikeresen letöltődtek ÉS parse-olódtak; a kiesők (403/429/gone/timeout) nem véletlen mintavétellel hiányoznak, ezért az F_q-n mért arány nem vihető át automatikusan a teljes bemeneti keretre. • KÉT ÚT: minden publikált cella két független úton készült (rekord-szintű, illetve a tárolt HTML-ből újra-kinyert deriváció), és a kettőnek cellára pontosan egyeznie kell; eltérés esetén a kör kiír egy szimmetrikus differenciát és NEM ír táblát.
How we measured — and how we tried to break the measurement
A measurement is worth what its verification chain is worth.
- Sample
- 981 UAE e-commerce domains — a fixed, hash-pinned census, not a hand-picked list
- Window
- 2026-08-07..2026-08-12 (UTC)
- Method
- Server-side HTML retrieval without JavaScript execution — what an AI crawler sees; JSON-LD block-level parsing; robots.txt evaluated per RFC 9309
- Frames
- F_all 981 · F_home 966 · F_prod 807 · F_llms 966 · F_robots_conclusive 975 — every rate names its frame
- Agents
- CCBot, ClaudeBot, GPTBot, Google-Extended, OAI-SearchBot, PerplexityBot, meta-externalagent (+Bytespider observational)
- Small-cell rule
- Any cell under twenty rows ships as a count, never a rate
- Publication gate
- A sealed per-metric table — metric definition, numerator, denominator, frame, mandatory caveats — is the gate; this page renders its rows and nothing else
- Privacy
- Aggregate figures only; no individual store is named on this page
The adversarial audit
Independent adjudicators re-examined 12 cells of 5 rows each drawn across the measurement — attacking both directions: rows the engine counted and rows it did not. 49 rows were upheld, 10 of 60 flipped (16.7%), and 1 was left unresolved with the engine's reading kept. The flips ran almost exclusively in the direction of engine under-counting; one family runs the other way, which is why direction is stated per row, never blanket. The round-level result ships verbatim beside the headline at the top of this page — never as a footnote.
Every published number, with its mandatory caveats
The sealed export is the authority; this appendix condenses its mandatory caveats in English, without naming individual stores. English working translation, keyed: Sampled adversarial audit: 10/60 flips (16.7%), direction almost exclusively engine under-count -> the published rates are LOWER BOUNDS · sealed audit record, Hungarian, verbatim: Mintaveteles adverzarialis audit: 10/60 flip (16,7%), irany szinte kizarolag engine-alulszamolas -> a kozolt ratak ALSO KORLATOK
product_merchant_valid · 626 of 807 (77.6%) · F_prod lower bound
The validator reads only direct offer-level price keys, not the nested priceSpecification form — the rate is a lower bound. True complement on the same denominator: 181 pages (22.4%) = 118 with no Product JSON-LD + 63 present-but-invalid; the decomposition closes exactly. A similarly-sized share quoted in the underlying data on a different denominator (pages that do have Product JSON-LD) is not this complement and the two must never be swapped.
Sampled adjudication: valid stratum upheld 5 of 5; invalid-side flips 3 of 10, all on the nested-price pattern. The denominator contains some archive pages (never merchant-valid); removing them would raise the rate. Counts only — cells under twenty.
S02-product-ldjson · 689 of 807 (85.4%) · F_prod bounded both sides
Adjudicated bound: 83.3% (689 of 827) to 90.9% (709 of 780) — the low end keeps suspected archives in the denominator, the high end drops them and credits the under-included rows. The presence predicate itself took no flips in the sample; the engine number ships uncorrected.
S08-field-invalidity · 175 of 807 (21.7%) · F_prod upper bound
The price-field failure count includes every page with no Product JSON-LD at all, and the validator does not read the nested price form — both push the count up, so it is an upper bound, not a point estimate. This row has no aggregate-level field; it derives from the flat per-domain data, a named seam in the publication chain.
S07-shopify-section · 479 of 552 (86.8%) · F_prod ∩ Shopify open bound
Mandatory pair: presence 496 of 552 (89.9%) vs. merchant-valid. Gap 17 rows, count only. 5 of the 17 were adjudicated; 2 flipped; envelope 479 to 496 of 552; 12 of 17 unadjudicated — bound open. 13 of the 17 gap rows fail on the price/priceCurrency pair. No corrected mix ships.
S03-product-found · 807 of 966 (83.5%) · F_home lower bound
Lower estimate for "a discoverable product page exists": 2 of 5 sampled no-product-page rows flipped (real product pages the discovery ladder missed). The metric is exact for its own definition; the gap is between definition and concept, not in the counting. Engine number ships uncorrected.
S09-funnel · 966 of 981 (98.5%) · F_all → F_home verified two-route
Conservation verified two-route with zero per-domain disagreements. In the sampled non-measurable remainder, 1 of 5 flipped: under that flip F_home would be at most 967 — the adjusted figure is not shipped, only this bound. Exactly 1 of 981 rows carries a reproducible engine-side fetch-persistence failure, named as an instrument limitation. The failure-bucket split shipped as a sealed addendum (round 2026-08-13AD) - see the Addendum section.
home_org_schema · 689 of 966 (71.3%) · F_home presence only
Presence, not validity: any Organization-family JSON-LD block counts; no fields are validated. Do not read it as a quality figure. The metric definition is the engine's own declaration (not pre-registered); the frame is pre-registered. A full census on its frame, not a sample.
llms_txt_present · 696 of 966 (72.0%) · F_llms engine number
Defined as present with content. True complement: 270 of 966 — 265 with no file, 5 with an empty one; any other complement phrasing is false. In the sampled numerator, 1 of 10 examined members failed re-examination (an error page served with a success status); that soft-error class is open and the shipped number contains it. Counts only.
S11-disallow-per-agent · GPTBot 35 of 975 (3.6%) · F_robots_conclusive pinned pass
All per-agent rows ship on the conclusive frame with engine numerators unchanged; 38 stores disallow at least one measured agent. Cross-pass drift on GPTBot: 35 on the pinned pass vs 36 on the later pass — real robots.txt drift, not a parser defect. The two passes also disagree on which handful of domains are unmeasured; their agreement on the total is coincidence and is not read as confirmation.
disallow_Bytespider_observational · 37 of 975 (3.8%) · F_robots_conclusive observational
Carved out of the scored agent set by construction; never summed into any AI-blocking figure. Ships on the engine's own conclusive frame (a census re-framing exists elsewhere and is superseded). Not independently adjudicated — it rests on a deterministic RFC 9309 evaluator over a frozen corpus, reproduced exactly by recount.
ai_all7_silent · 901 of 975 (92.4%) · F_robots_conclusive primary frame
Primary on the conclusive frame; a census-frame secondary form exists with a caveat and is not rendered here to keep one number on one frame. On the census frame the complement claim ("the rest are explicit") is false in every form, because some rows are neither — that sentence is unpublishable, and this page does not derive any complement for this row.
blanket_disallow · 0 of 975 · F_robots_conclusive honest zero
The zero does not mean "everyone explicitly allows": 35 of the conclusive files have no evaluated wildcard group at all. Non-vacuity is proven: the predicate is live — 140 of the files do carry a bare full-disallow line, but always inside named-agent groups, never in the wildcard group. A probe-hijack artifact was ruled out in code. The zero is frame-invariant across all three recounted frames.
sitemap_declared · 879 of 975 (90.15%) · F_robots_conclusive measured bound
Two independent robots passes bound this between 874 to 881 of 975; the engine number sits inside the bound. The complement mixes two categories (files with no sitemap line and stores with no robots.txt at all) and is not quoted as one number. The aggregate attaches no per-field frame — only the sealed table names it, which is why this page quotes the sealed table.
S15-robots-present · 960 of 981 (97.9%) · census upper-bound complement
Non-serving remainder: 15 client-error rows and 6 inconclusive. The client-error reading is an upper bound on "serves no robots.txt": on re-fetch some of those rows answered, giving a true range of 11 to 15. Every remainder row was individually adjudicated; a competing lower count from a derivative route was refuted in adjudication and the pinned-census figure ships.
measurement_window · 2026-08-07..2026-08-12 (UTC) run metadata
The fetch sub-windows (robots sweep, pilot, bulk tranches, closure) are recorded in the sealed row.
crawler_list · the measured agent roster run metadata
CCBot, ClaudeBot, GPTBot, Google-Extended, OAI-SearchBot, PerplexityBot, meta-externalagent (+Bytespider observational) — Bytespider is observational by construction and never enters a scored figure.
2026-08-13AD · sealed addendum round derived, zero-fetch
Category, field-invalidity and funnel-bucket marginals derived from the original round's frozen artifacts behind a separate founder gate; per-row mandatory caveats render beside their tables in the Addendum section. Axis rule, verbatim: census column is the pre-registered axis; record-level category fields cross-checked (sites JSONL, robots checkpoint); disagreement -> axis note, census wins, never merged
2026-08-13Q · sealed quality round · fetch window 2026-08-12 (UTC) new window
Quality-depth readings on a fresh fetch of the same frame, with judged instruments shipping beside their own measured error; per-row mandatory caveats render beside their tables in the Part Two section. English working translation, keyed: Q-strip: W2 challenger 40/48 flips, W3 blind refuter exact 76/80, W4 blind audit B1 0/10 and B2 19/20 - every judged row ships with its own measured error · sealed audit record, Hungarian, verbatim: Q-strip: W2 challenger 40/48 megdontes, W3 vak refuter egzakt 76/80, W4 vak audit B1 0/10 es B2 19/20 - minden itelt sor a sajat mert hibajaval megy ki