Skip to content

Caravanserai Research · Report 01 · Annex

Registered 31 August 2026 · Published unchanged

What we expected to find, dated before we found it.

This document was written and dated while the harness was built and validated and the full run had not started. It is published unchanged whatever the results say. If the results contradict it, the contradiction is the finding and this file is the evidence that we did not go looking for the answer we wanted.

Registered
31 August 2026
Status at registration
Harness built and validated; the full run had not started.
Frame
6,989 merchants.
Sample
1,165 merchants, drawn and frozen before this document was written. Seed 20260831.

Disclosure, first

Two things a reader is entitled to know before the hypotheses.

  1. 01

    A 20-merchant dry run was executed before this registration, to validate the harness. It was all florists, because the sample was category-ordered at the time. It found Shopify stores serving llms.txt and products.json at 6/6 while Square served 0/4 — which is consistent with H1 below. H1 was written before that dry run, dated and committed. But we cannot claim H1 is untouched by data, and we are not going to pretend otherwise. The dry run rows are in the results file and are not excluded from analysis.

  2. 02

    The sample was drawn before this document. Its composition is fixed and stated above. We cannot enlarge, shrink or re-stratify it after seeing results without saying so here.

Hypotheses

H1

Outcome, added 1 September 2026 · Confirmed

The Shopify cliff

Agent-readability is close to binary on platform rather than continuous. Shopify merchants are legible without having done anything; merchants on other platforms are substantially less legible.

Predicted
Shopify stores show materially higher rates of machine-readable price, structured product data and a published discovery surface than Square, Toast, Clover, Wix, Squarespace, WooCommerce and custom sites. We expect the gap to look like a step, not a gradient.
Falsified if
The difference between Shopify and the next-best platform is small, or readability varies smoothly across platforms with no discontinuity.

H2

Outcome, added 1 September 2026 · Directional

The pretty-site paradox

Sites built on design-led, client-rendered stacks are less readable than plainer ones. The more a merchant spent on the website, the less an agent sees.

Predicted
Custom/React/Webflow/Squarespace sites show lower rates of server-rendered price and structured data than WooCommerce or plain template sites.
Falsified if
Client-rendered stacks match or beat plainer stacks on server-rendered price and structured data.
Known limitation, stated in advance
A plain HTTP fetch proves a price is absent from the server response. It cannot prove JavaScript would have supplied one. Rows where no price is found are flagged and a headless render is run against that subset only; rows resolved that way are reported separately and never merged into the fetch-only figures.

H3

Outcome, added 1 September 2026 · Confirmed at the limit

Nobody publishes a promise

Price is common, live stock is rare, and a machine-readable delivery commitment to a specific address is close to nonexistent.

Predicted
hasAddressableCommitment at or near 0%.
Falsified if
A non-trivial share of merchants publish structured delivery commitments.

Measures

Fixed before the run, defined in the harness and tested against fixtures.

  • robots.txt posture toward GPTBot, ClaudeBot, PerplexityBot, Google-Extended, OAI-SearchBot
  • JSON-LD Product / Offer: price, priceCurrency, availability
  • Price present in server-rendered HTML, recorded with which page it came from
  • Platform fingerprint
  • Discovery surface: sitemap.xml, products.json, llms.txt, .well-known/ai-plugin.json
  • OfferShippingDetails with a deliveryTime — the H3 measure

Analysis plan

Committed in advance, so that a choice made after seeing the numbers would be visible as a change.

  • Primary breakdowns: by platform (H1), by site-technology class (H2), by measure (H3).
  • Wilson binomial confidence intervals at 95% on every proportion. No point estimate reported without one.
  • Cross-category aggregates use the per-category weights in the sample. The draw is capped, not proportional, and unweighted aggregates would misstate the population.
  • Non-response is reported in the open: robots-blocked, access-denied, unreachable, timed out and failed-to-parse, each as its own row. A study that hides its denominator is the thing this study is accusing the market of.
  • Cells with n < 30 are reported with their n beside them and are not used to headline a claim.

What we are not doing.

  • No merchant is named publicly. Per-store reports stay private.
  • No robots.txt is worked around. A block is a data point.
  • No hypothesis is added after seeing results. Anything unexpected is reported as exploratory and labelled as such.