Caravanserai Research · Report 01 · Annex
Registered 31 August 2026 · Published unchanged
What we expected to find, dated before we found it.
This document was written and dated while the harness was built and validated and the full run had not started. It is published unchanged whatever the results say. If the results contradict it, the contradiction is the finding and this file is the evidence that we did not go looking for the answer we wanted.
- Registered
- 31 August 2026
- Status at registration
- Harness built and validated; the full run had not started.
- Frame
- 6,989 merchants.
- Sample
- 1,165 merchants, drawn and frozen before this document was written. Seed 20260831.
Disclosure, first
Two things a reader is entitled to know before the hypotheses.
01
A 20-merchant dry run was executed before this registration, to validate the harness. It was all florists, because the sample was category-ordered at the time. It found Shopify stores serving llms.txt and products.json at 6/6 while Square served 0/4 — which is consistent with H1 below. H1 was written before that dry run, dated and committed. But we cannot claim H1 is untouched by data, and we are not going to pretend otherwise. The dry run rows are in the results file and are not excluded from analysis.
02
The sample was drawn before this document. Its composition is fixed and stated above. We cannot enlarge, shrink or re-stratify it after seeing results without saying so here.
Hypotheses
H1
Outcome, added 1 September 2026 · Confirmed
The Shopify cliff
Agent-readability is close to binary on platform rather than continuous. Shopify merchants are legible without having done anything; merchants on other platforms are substantially less legible.
- Predicted
- Shopify stores show materially higher rates of machine-readable price, structured product data and a published discovery surface than Square, Toast, Clover, Wix, Squarespace, WooCommerce and custom sites. We expect the gap to look like a step, not a gradient.
- Falsified if
- The difference between Shopify and the next-best platform is small, or readability varies smoothly across platforms with no discontinuity.
H2
Outcome, added 1 September 2026 · Directional
The pretty-site paradox
Sites built on design-led, client-rendered stacks are less readable than plainer ones. The more a merchant spent on the website, the less an agent sees.
- Predicted
- Custom/React/Webflow/Squarespace sites show lower rates of server-rendered price and structured data than WooCommerce or plain template sites.
- Falsified if
- Client-rendered stacks match or beat plainer stacks on server-rendered price and structured data.
- Known limitation, stated in advance
- A plain HTTP fetch proves a price is absent from the server response. It cannot prove JavaScript would have supplied one. Rows where no price is found are flagged and a headless render is run against that subset only; rows resolved that way are reported separately and never merged into the fetch-only figures.
H3
Outcome, added 1 September 2026 · Confirmed at the limit
Nobody publishes a promise
Price is common, live stock is rare, and a machine-readable delivery commitment to a specific address is close to nonexistent.
- Predicted
- hasAddressableCommitment at or near 0%.
- Falsified if
- A non-trivial share of merchants publish structured delivery commitments.
Measures
Fixed before the run, defined in the harness and tested against fixtures.
- robots.txt posture toward GPTBot, ClaudeBot, PerplexityBot, Google-Extended, OAI-SearchBot
- JSON-LD Product / Offer: price, priceCurrency, availability
- Price present in server-rendered HTML, recorded with which page it came from
- Platform fingerprint
- Discovery surface: sitemap.xml, products.json, llms.txt, .well-known/ai-plugin.json
- OfferShippingDetails with a deliveryTime — the H3 measure
Analysis plan
Committed in advance, so that a choice made after seeing the numbers would be visible as a change.
- Primary breakdowns: by platform (H1), by site-technology class (H2), by measure (H3).
- Wilson binomial confidence intervals at 95% on every proportion. No point estimate reported without one.
- Cross-category aggregates use the per-category weights in the sample. The draw is capped, not proportional, and unweighted aggregates would misstate the population.
- Non-response is reported in the open: robots-blocked, access-denied, unreachable, timed out and failed-to-parse, each as its own row. A study that hides its denominator is the thing this study is accusing the market of.
- Cells with n < 30 are reported with their n beside them and are not used to headline a claim.
What we are not doing.
- No merchant is named publicly. Per-store reports stay private.
- No robots.txt is worked around. A block is a data point.
- No hypothesis is added after seeing results. Anything unexpected is reported as exploratory and labelled as such.
