
Web Bot Auth: Crawlers That Sign Their Requests
Quick answer: Web Bot Auth replaces "trust this user-agent string" with an Ed25519 signature you can actually verify. Cloudflare, Google and AWS have all shipped support, and the IETF working group draft went active on 1 September 2026. But Google explicitly warns that it does not sign every request from a given agent, so this is an additive signal for identifying good bots — not a gate you can close against bad ones.
Bot identification has run on an honour system for twenty years. A crawler announces itself in the User-Agent header, and you decide whether to believe it. The only real check is a reverse DNS lookup or matching against a published IP list, which means maintaining allowlists that change without notice and trusting that the publisher keeps them current.
Web Bot Auth does the obvious thing instead: the client signs the request with a private key, publishes the public key at a well-known URL, and you verify the signature. The mechanism is not novel — it is RFC 9421 HTTP Message Signatures, a Proposed Standard since February 2024. What is new is that three of the largest edge providers now implement the same profile of it, and that the spec finally has a working group behind it.
What a signed request looks like on the wire
Three headers. Signature and Signature-Input come straight from RFC 9421. The third, Signature-Agent, is what Web Bot Auth adds: a dictionary structured header whose values are HTTPS URLs pointing at the key material, with a type parameter naming the discovery mechanism. The draft defines three: directory (the default), jwks_uri and cimd.
Google's implementation sends Signature-Agent: g="https://agent.bot.goog", and publishes its keys at https://agent.bot.goog/.well-known/http-message-signatures-directory. The signature parameters are constrained rather than free-form. The working group draft requires created, expires (with a recommended maximum window of 24 hours), keyid as a base64url JWK SHA-256 thumbprint, and tag set to the literal string web-bot-auth. The covered components must include either @authority or @target-uri, which is what stops a captured signature being replayed against a different host.
Shared secrets are ruled out. The earlier architecture draft put it in capitals: implementations must not use shared HMAC. Test vectors in the working group draft cover RSASSA-PSS with SHA-512 and EdDSA over Curve25519, and Ed25519 is what both Cloudflare and AWS document using.
Who has shipped it
| Party | Role | Status |
|---|---|---|
| Cloudflare | Verifier, and co-author of the spec | Signatures accepted as a Verified Bots method since July 2025 |
| Signer | Experimental; a subset of Google-Agent requests, documented 4 May 2026 | |
| AWS | Verifier | WAF Bot Control labels, announced 14 July 2026 |
| IETF | Standards body | Working group draft active, updated 1 September 2026 |
AWS is the most operationally concrete of the three. WAF Bot Control appends labels you can write rules against: web_bot_auth:verified, invalid, expired and unknown_bot, plus vendor and bot-specific labels. Support arrived in Bot Control rule group version 4.0 for CloudFront in November 2025 and version 6.0 extended it across standard WAF resource types. Note the unknown_bot label carefully: a valid signature from a key you have never seen is not the same thing as a trusted crawler, and conflating the two is the first mistake available here.
The spec has been renamed three times — check which draft you read
This matters if you are implementing from a blog post. The work started as draft-meunier-web-bot-auth-architecture, with a companion directory draft. Both expired and were archived. They were consolidated into draft-meunier-webbotauth-httpsig-protocol, which also expired, and the live document is now draft-ietf-webbotauth-httpsig-protocol — the ietf prefix marking working group adoption.
One concrete consequence: the well-known path in the current working group draft is /.well-known/http-message-signatures-directory, served as application/http-message-signatures-directory+json, and that plural form is what Google's live directory uses. Cloudflare's 2025 announcement post documented the singular http-message-signature-directory. Read the current vendor documentation rather than an announcement from last year, and if you are publishing a directory, test that verifiers can actually fetch it before assuming the path is right.
What this does not solve
The single most important sentence in Google's documentation is the caveat: "We don't sign every request of a particular agent. Be sure that you fall back to the established methods of bot verification." Partial signing is the default state, not a transitional glitch. Any logic of the form "unsigned means not Google" will produce false negatives against real Google traffic today.
So a signature is a positive signal and its absence is not a negative one. That asymmetry is the whole design constraint. You can use verified to skip rate limits, serve a cheaper cached path, or exempt a crawler from a challenge. You cannot use missing signatures to deny, unless you are prepared to lose indexing.
It also says nothing about permission. Verifying who a crawler is does not tell you whether it may train on your content — that is what the licensing and preference signals in robots.txt and adjacent proposals try to express, and it remains a separate, much less settled layer. Identity and authorisation are different problems, and Web Bot Auth only fixes the first. Nor does it change what a crawler can see once admitted: many AI crawlers still do not execute JavaScript, so a signed request against a client-rendered page gets the same empty shell an unsigned one would — which is a content structure problem, not an identity one.
What to do about it now
If you sit behind Cloudflare or AWS WAF, you are already a verifier and the work is configuration: find the labels, decide what verified should earn, and leave the unsigned path alone Both vendors document this under their bot-management products rather than their CDN settings, so it may sit with a different team than you expect; the security and privacy category lists the main players. If you terminate your own TLS, verification is a modest amount of code — fetch the JWKS, cache per Cache-Control, delete expired keys, verify per RFC 9421 — but it is not worth writing before you have measured how much signed traffic you actually get. Log the Signature-Agent header for a fortnight first. On most sites the honest answer will be "almost none yet".
If you are building an agent rather than defending against one, the incentive is clearer. Signing gets you onto verified-bot paths at three major edges, and AWS notes that registration APIs are still on its roadmap — so manual registration is the current cost of entry. Amazon Bedrock AgentCore Browser signs automatically; everything else implements RFC 9421 itself.
The broader lesson for anyone optimising for AI referrals is that the retrieval layer is getting formal machinery — signatures, registries, IANA-registered well-known paths — while the ranking layer stays opaque. That is the reverse of the usual order, and it is worth keeping in mind when a vendor sells you a file that no consumer has confirmed reading. We went through that with llms.txt, and the measurement gap is the same one that makes Search Console's AI report so hard to act on. For tooling in this space, TechLogHub's AI tools and platforms category and the free ai.txt generator are both a reasonable starting point.
FAQ
Can I block crawlers that do not send a Web Bot Auth signature?
Not safely. Google's documentation states that it does not sign every request from a given agent and instructs site owners to fall back to established verification methods. Treating the absence of a signature as evidence of a fake crawler will block legitimate traffic, including Google's own.
Is Googlebot signing its requests?
Google's documentation names the Google-Agent user agent and says a subset of its requests are signed. It does not claim that all Google user agents participate, and the feature is labelled experimental because the underlying specification is still a draft.
Which well-known path should a directory be served from?
The current IETF working group draft registers /.well-known/http-message-signatures-directory with the media type application/http-message-signatures-directory+json, and Google's live directory uses that plural form. Older material, including Cloudflare's 2025 announcement, documented a singular variant, so verify against current vendor documentation.
Does a signature prove a crawler is allowed to use my content?
No. Web Bot Auth answers identity only. What a verified crawler may do with what it fetches is governed by separate licensing and preference mechanisms, which are considerably less settled than the signature machinery.
How long is a signature valid for?
The working group draft requires both created and expires parameters and recommends a maximum validity window of 24 hours. Because the covered components must include the target authority or URI, a signature captured from one host cannot be replayed against another.
Is this a finished standard?
Not yet. RFC 9421, the HTTP Message Signatures standard it builds on, is published. The Web Bot Auth profile on top of it is an active working group draft as of 1 September 2026, and three earlier individual drafts have already expired and been replaced, so header and path details have moved more than once.
Cryptographic identity for crawlers is a genuine improvement on a twenty-year honour system. It is also, for now, a signal that tells you who to be nicer to — not who to shut out.


