Publishers and media
For you this is about access, not commerce
The decisions that matter are who may read you, what they may do with it, and whether the ones you want are getting through.
What agents are doing here
Two very different things arrive at a publisher, and they deserve different answers. Crawlers collect at scale for training or indexing, with nobody waiting. Assistants fetch a specific piece because a reader asked about it, with someone waiting on the answer and a link back to you at the end of it.
Which of these actually apply to you
There are two standards and they do not both matter to every kind of site.
UCP
Does not applyUCP describes selling goods. Unless you run a shop alongside the publication it has nothing to say to you, and you can safely ignore it. If you sell subscriptions, that is a payment relationship rather than the catalogue and checkout model UCP is built around.
What UCP isWebMCP
Partly appliesUseful but not urgent. Declaring search over your archive, or retrieval of a specific piece, makes an assistant far more likely to quote you correctly rather than approximate you from a page it half read.
What WebMCP isWhere it usually goes wrong
None of these produce an error. Each is a working site doing something an agent cannot follow.
One rule for both kinds of visitor
Blocking training crawlers and blocking assistants that cite you are different decisions with different consequences. A single rule for all non-human traffic answers both at once, usually by accident.
robots.txt doing less than you think
Google, OpenAI and Perplexity all document that user-triggered fetches may ignore robots.txt. Anthropic alone applies it uniformly. Whatever you decide has to be enforced at the edge to actually apply.
Defaults changing under you
From September 2026, new domains onboarding to Cloudflare get defaults that block agent-classified traffic on pages showing ads. Worth knowing before it applies to you rather than after.
Paywalls that read as a block
An assistant refused at a paywall and an assistant refused by bot rules look identical from outside. Only one of those was a decision you made.
What to do about it
- 1 Separate the two decisions: content used for training, and assistants fetching a piece for a waiting reader.
- 2 Enforce whichever you choose at the edge, because robots.txt will not do it for you.
- 3 Watch what you are refusing. Turning away the assistants that would have cited you is the expensive mistake.
- 4 Consider declaring archive search, so you are quoted accurately rather than approximately.
See where your site stands
One address, no account. You get a grade, what is worth fixing, and how to fix it.
Run a different kind of site? See the other sectors.