CONTENTS07 +
How the Web Gallery of Art gives language models a plain Markdown version of every record: llms.txt, Accept negotiation, strict limits and edge caching.
A fair share of the traffic on the Web Gallery of Art (WGA) isn’t people. Some of it is crawlers, and a growing part is language models fetching pages on someone’s behalf. They get the same HTML a visitor gets, navigation, scripts and a session cookie they’ll never use included, and every one of those requests costs the server the same work as a real visit. So I gave them a door of their own: a plain Markdown version of every artist and artwork, generated once a night and cheap enough to serve from a cache.
Why not let them read the page
For an agent, most of a page is packaging. What it wants from an artwork is a title, a few facts and a paragraph of commentary, wrapped in everything a browser needs to show them. The packaging also costs the most to hand out. Every artist and artwork page is built on request, and because public pages take part in a visitor’s itinerary session (the route a visitor can plan through the collection), the response can carry a cookie, so it can’t safely be served to someone else from a cache.
WGA sits behind Cloudflare, and Cloudflare offers to solve exactly this: its Markdown for Agents feature converts HTML to Markdown at the edge. It’s only on the paid plans, though, and I wasn’t going to upgrade for one feature. Converting at request time would still start from the page it converts. Publishing the Markdown myself, from the catalogue data instead of the HTML, means the free plan can cache it like any other file.
What an agent gets
Here’s the document for Jean-Joseph-Xavier Bidauld’s View of the Waterfalls at Tivoli, as the beta site serves it, with the list of related works shortened:
1# View of the Waterfalls at Tivoli23[View the canonical artwork record](<https://beta.wga.hu/artists/bidauld-jean-joseph-xavier-r46d21a0de50786/view-of-the-waterfalls-at-tivoli-r1a269997dadeec>)45## Artwork67- Artist: [BIDAULD, Jean-Joseph-Xavier](<https://beta.wga.hu/artists/bidauld-jean-joseph-xavier-r46d21a0de50786>)8- Date: 1751-18009- Technique: Oil on paper on canvas, 51 x 38 cm1011## Commentary1213This painting is a typical example of the many studies, painted in oils on paper, that Bidauld executed in the open air on his extensive trips through the countryside during his Italian sojourn from 1785 through 1790. They are not sketches but pictures finished in the studio.1415## Related artworks1617- [Italianate Landscape](<https://beta.wga.hu/artists/bidauld-jean-joseph-xavier-r46d21a0de50786/italianate-landscape-rff080d33475a9a>)18- [View of the Ponte Rocco, Tivoli \(detail\)](<https://beta.wga.hu/artists/bidauld-jean-joseph-xavier-r46d21a0de50786/view-of-the-ponte-rocco-tivoli-detail-r466bcf04593f13>)19- …2021## Attribution2223[Web Gallery of Art](<https://beta.wga.hu/pages/about>)
It has the record and nothing else: a link back to the canonical page, the facts as a list, the commentary as prose, up to twelve related works by the same artist, and an attribution line. Every link is an absolute canonical URL. The angle brackets around each URL keep parentheses in it from ending the link early, and the backslashes in \(detail\) do the same job for titles.
How agents find it
Agents that read instructions get /llms.txt, a short Markdown file at the root of the site. It says what the collection is, that the canonical site is authoritative, and how to turn a page URL into a Markdown one: take the ID after the last hyphen and put it into /agents/artworks/{id}.md or /agents/artists/{id}.md. It links the sitemap for discovering records and the about page for terms of use. Every HTML page points to it, with rel="describedby" both in the response headers and in the <head>, and the footer links to it as well.
Each artist and artwork page also names its own Markdown version, in a Link header and in the <head>:
1Link: </llms.txt>; rel="describedby"; type="text/markdown"2Link: <https://beta.wga.hu/agents/artworks/r1a269997dadeec.md>; rel="alternate"; type="text/markdown"
Asking for Markdown
An agent that reads nothing can still ask. HTTP has let a client say which formats it prefers for much longer than anyone has wanted a page for a language model: the Accept header, with optional quality values. When a request for an artist or artwork page prefers text/markdown, WGA answers with a temporary redirect to the generated file:
1GET /artists/bidauld-jean-joseph-xavier-r46d21a0de50786/view-of-the-waterfalls-at-tivoli-r1a269997dadeec2Accept: text/markdown34HTTP/2 3075Location: /agents/artworks/r1a269997dadeec.md6Vary: Accept
“Prefers” matters there. Browsers send something like text/html,…,*/*;q=0.8, and that wildcard matches Markdown as well, at a lower quality, so a naive check would send every visitor to a text file. WGA takes the best match for each type and lets Markdown win only when its quality is higher, or equal and Markdown was named explicitly:
Accept | Result |
|---|---|
text/html,application/xhtml+xml,…,*/*;q=0.8 (a browser) | the page |
*/* | the page |
text/markdown | 307 to the Markdown |
text/markdown, text/html | 307 to the Markdown |
Vary: Accept tells Cloudflare, and any other cache, that the same URL has two answers.
The check runs before the session and page middleware. It reads the generated files, not the database, and redirects only when there’s a file for the record and the requested URL is one of that record’s published addresses. Anything else, such as an old alias or an unpublished record, falls through to the normal handler and gets its usual answer. A successful redirect costs a file lookup; building the page never starts.
What it deliberately leaves out
Making a catalogue easy for a model to read also makes it easy to copy, so the Markdown is limited on purpose. There’s no llms-full.txt with the whole collection in one file. It would be the most convenient thing to give an agent, an even more convenient thing to give a scraper, and a large file to regenerate every night. llms.txt says as much, and adds that the reproductions may carry rights belonging to the institutions holding the works, so their availability shouldn’t be read as a licence to reuse them.
Each document is built from a projection of the public record, not from the record itself. The Go types the renderer accepts have fields for a title, the facts, the commentary, the links and the attribution, and nothing else: no session or itinerary state, no administrative fields, nothing from the database’s own bookkeeping. Private data can’t reach a file even by mistake, because there’s nowhere to put it.
Everything in a document is bounded. Titles are cut at 300 characters, each fact at 500, the commentary at 4,000 and URLs at 2,048, and the related works stop at twelve. Every link has to be an absolute URL with no query string, fragment or credentials. Links to other records have to stay on the site’s own host, and anything pointing at the admin interface or the API is refused. Markdown’s own punctuation in the catalogue text is escaped, so a title with brackets in it stays a title. Records without a name or title, and artworks without a published artist, are skipped; anything else that breaks a rule stops the whole run, and the previous set of files stays live.
How it stays cheap
The files come from the same job that builds the sitemap, every night and once when the application starts. The job writes a complete set into a staging directory and checks it: exactly the files it expected, none of them empty, no symlinks. Only then does it switch the single marker that says which set is current, and remove the old ones. Readers always go through that marker, so an agent sees last night’s files or tonight’s, never a mixture, and a run that fails part-way leaves the previous set in place.
Serving a file means reading it from disk, with no database query on the way, so a request for a record that doesn’t exist gets a 404 without the database noticing. A successful response looks like this:
1HTTP/2 2002Content-Type: text/markdown; charset=utf-83Cache-Control: public, no-cache, must-revalidate4Link: <https://beta.wga.hu/artists/bidauld-jean-joseph-xavier-r46d21a0de50786/view-of-the-waterfalls-at-tivoli-r1a269997dadeec>; rel="canonical"5ETag: "e05a91ebfa132ab782fd5f2848654271667c1d1f1186718d15ca86155e128fe7"
There’s no cookie, and the canonical Link points back to the HTML page. no-cache lets Cloudflare keep a copy, as long as it checks with WGA before using it. The ETag is a SHA-256 of the file, so that check gets an empty 304 Not Modified until a night’s run actually changes the content.
1sequenceDiagram2 participant A as Agent3 participant C as Cloudflare4 participant W as WGA5 A->>C: GET /artists/…/view-of-the-waterfalls-at-tivoli-…<br/>Accept: text/markdown6 C->>W: GET (Vary: Accept keeps the two answers apart)7 W-->>C: 307 Location: /agents/artworks/r1a269997dadeec.md8 C-->>A: 3079 A->>C: GET /agents/artworks/r1a269997dadeec.md10 alt nothing cached yet11 C->>W: GET12 W-->>C: 200, the file read from disk, ETag13 else a copy is cached14 C->>W: If-None-Match: the ETag15 W-->>C: 304 Not Modified until a nightly run changes the file16 end17 C-->>A: 200 text/markdown
Cache misses still count. Requests under /agents/ go through the same per-client rate limits as the artist and artwork pages, at Cloudflare and again in the application, so a cheap file isn’t a cheap way around the limits.
Wrapping up
Nothing changes for people. A browser never prefers Markdown, so it gets the page as before, plus two headers it ignores. An agent that reads llms.txt or sends the right Accept header gets the same record in about 2 KB instead of 72 KB of HTML (before scripts and images), usually from Cloudflare’s copy. Agents that do neither still get the page, through the same rate limits as everyone else.
CITE THIS POST
Referencing this in a paper or thesis? Here’s a BibTeX entry.
1@misc{galicz2026how,2 author = {Galicz, Mikl{\'o}s},3 title = {{How the Web Gallery of Art Serves Markdown to LLMs}},4 year = {2026},5 month = oct,6 howpublished = {\url{https://blackfyre.ninja/blog/how-the-web-gallery-of-art-serves-markdown-to-llms/}},7 url = {https://blackfyre.ninja/blog/how-the-web-gallery-of-art-serves-markdown-to-llms/},8 urldate = {2026-10-10},9 note = {Blog post, blackfyre.ninja}10}