terms.txt is a proposed replacement for robots.txt that states, per path and per purpose, whether an AI client may read your pages, must pay, or is refused. It is a single-author preprint, not a standard, and nothing in the paper says any site or vendor has adopted it. It is still worth understanding, because it describes access as a priced, receipted exchange rather than a yes/no gate.

What the paper proposes

The preprint is arXiv:2609.11152, submitted 10 September 2026 by independent researcher Rajarshi Chowdhury to an IEEE Internet Computing special issue. Its full text argues that robots.txt cannot express identity, purpose, terms or price, and that newer alternatives are largely proprietary CDN features.

Its answer has two halves:

  • A file. An origin publishes /.well-known/terms.txt, in robots.txt-like syntax. Each Path: block sets a rule per purpose: allow, charge (with a price per request) or deny. The paper’s example vocabulary is search, agent, train-ai, archive and research.
  • An exchange. The file enforces nothing by itself. Enforcement comes from a request flow at the origin: a Web Bot Auth signature, a signed Access-Intent header, an optional delegation token, HTTP 402 when payment is required, and a signed Access-Receipt on delivery.

The paper separates four client classes: training crawlers, search crawlers, service-operated agents, and user-delegated agents. The last inherits the entitlements of the person it acts for, without revealing that person’s identity to the origin.

Why an agent-commerce reader should care

The paper’s objections section makes the point directly: an agent can compare twenty retailers, buy from one and deliver a conversion without a referral trail. If value no longer flows through clicks, it argues, accounting has to move to the request itself.

That is the same shift as treating access as a transaction. A signed request, a declared purpose and a receipt are what let a store tell an agent that is browsing from one that is buying on someone’s behalf.

What it costs and what it cannot do

The author built a reference implementation of about 600 lines of dependency-free JavaScript. On one shared 2.1 GHz vCPU over loopback without TLS, it reports 0.20 to 0.65 ms added per request. Identity verification is the largest single cost. A forged signature is rejected in about 0.21 ms, cheaper than serving a paid request. These are loopback figures; the paper says real-traffic deployment remains to be studied.

The limits are stated plainly in the paper:

  • A signed purpose proves who declared it, not that the declaration is true. False claims are handled by audit and revoking standing.
  • Once content leaves the origin, HTTP cannot govern what happens to it. What a model does with lawfully received content stays contractual.
  • It does not stop unauthenticated scraping, because it cannot force a client to sign.

Where it sits among existing mechanisms

The paper’s comparison table lists what exists today. robots.txt (RFC 9309) is advisory. The IETF AI Preferences drafts can express preferred use but, by charter, do not enforce it. Cloudflare’s Pay Per Crawl enforces and prices, but only through Cloudflare. Web Bot Auth verifies identity but not terms; the paper reports that the IETF working group adopted it as a Standards Track document on 1 September 2026 (see our write-up).

terms.txt is one rung earlier than all of these: it has no grammar or registered well-known location yet, and the paper lists both as still needing standardization. It also says it would be falsified if sites publishing terms and receipts see no change in crawler behavior over a year.

FAQ

What is terms.txt?

terms.txt is a proposed robots.txt-style file, published at /.well-known/terms.txt, that states per path and per purpose whether machine clients are allowed, charged or denied. It comes from a September 2026 arXiv preprint by a single author. The file itself enforces nothing; a paired signed-request exchange does.

Is terms.txt an adopted standard?

No. It is an unreviewed preprint submitted to an IEEE special issue, and the paper itself says the grammar and well-known location still need standardization. Nothing we fetched shows a CDN, agent operator or IETF working group implementing it.

Does terms.txt replace robots.txt today?

No. robots.txt remains the file crawlers actually read. The paper positions terms.txt as a way to carry enforceable, priced terms alongside it, using the same purpose vocabulary so the two can state consistent policy.

Sources

AgentReady’s thesis in one line: the next access question is not “is the bot allowed in” but “can an agent identify itself, state its purpose and complete a transaction on your terms.”