Why we measured this
Plenty has been written about generative optimisation, but almost always on American data. We wanted to know how the websites we actually know are doing — Slovak and Czech e-shops, and the agencies that build them.
So we looked at 822 e-shops and at the portfolios of 454 agencies — that is, the sites those agencies build and maintain for their clients. We measured three technical signals that decide whether a site makes it into a language model's answer at all.
What we measured
We picked signals that can be verified from the outside and that are entirely under the site operator's control:
llms.txt— the file that tells language models what matters on a site. The equivalent ofrobots.txt, but for AI.- Structured data (schema.org) — whether a site describes its content in a machine-readable way, and whether it does so for products too.
- AI crawler access — whether
robots.txtblocks GPTBot, ClaudeBot, PerplexityBot and the rest.
We split the sample into two groups, because their problems differ: e-shops, and the agencies that build and run sites for them.
What came out of it
The first finding is uncomfortably simple: the overwhelming majority of sites have no llms.txt at all. It is not misconfigured — it does not exist. At that rate this is not a competitive edge for a handful of pioneers; it is a standard nobody has adopted yet.
The second finding is more serious. A share of sites actively block AI crawlers in robots.txt — usually not by the owner's decision, but through an inherited configuration or a template default. Such a site does not have reduced visibility. It has none. It cannot enter an answer, however good its content is.
The third finding concerns structured data. A substantial share of sites has basic schema markup, but product schema is missing far more often — and product schema is what decides whether a model can say something concrete about an item or simply passes over it in silence.
For agencies this plays out differently than for e-shops. An agency runs dozens of client sites, so one bad default in its standard configuration is multiplied across the entire portfolio at once.
What to do about it
The order follows what hurts most and what can be fixed fastest:
- Check your
robots.txt. If it blocks AI crawlers, everything else is pointless. The fix takes minutes. - Add product schema. Without it, a model has nothing to say about your products.
- Add
llms.txt. Today it buys you a head start; in a year it will be table stakes.
Methodology, and what this study does not claim
We measured externally available technical signals, at a single point in time. We do not claim that a site with llms.txt gets cited more often — that conclusion would require measuring the models' answers themselves over time, which is a different study.
We claim something more modest, but verifiable: these sites are closing doors on themselves that they mostly do not know exist.
The sample is not random — it is drawn from public directories of e-shops and partner agencies, so it represents the active part of the market rather than the internet as a whole.
Want to know where your site stands? Enter your domain on the home page and you will get the breakdown in a few minutes.