BLACKFYRE
MG—01
LANG
←
ALL POSTS
05 · Writing / 2026
09 OCT 2026 · 7 MIN READ

How the Web Gallery of Art Serves Markdown to LLMs

How the Web Gallery of Art gives language models a plain Markdown version of every record: llms.txt, Accept negotiation, strict limits and edge caching.

CONTENTS
07 +

A fair share of the traffic on the Web Gallery of Art (WGA) isn’t people. Some of it is crawlers, and a growing part is language models fetching pages on someone’s behalf. They get the same HTML a visitor gets, navigation, scripts and a session cookie they’ll never use included, and every one of those requests costs the server the same work as a real visit. So I gave them a door of their own: a plain Markdown version of every artist and artwork, generated once a night and cheap enough to serve from a cache.

Why not let them read the page

For an agent, most of a page is packaging. What it wants from an artwork is a title, a few facts and a paragraph of commentary, wrapped in everything a browser needs to show them. The packaging also costs the most to hand out. Every artist and artwork page is built on request, and because public pages take part in a visitor’s itinerary session (the route a visitor can plan through the collection), the response can carry a cookie, so it can’t safely be served to someone else from a cache.

WGA sits behind Cloudflare, and Cloudflare offers to solve exactly this: its Markdown for Agents feature converts HTML to Markdown at the edge. It’s only on the paid plans, though, and I wasn’t going to upgrade for one feature. Converting at request time would still start from the page it converts. Publishing the Markdown myself, from the catalogue data instead of the HTML, means the free plan can cache it like any other file.

What an agent gets

Here’s the document for Jean-Joseph-Xavier Bidauld’s View of the Waterfalls at Tivoli, as the beta site serves it, with the list of related works shortened:

/agents/artworks/r1a269997dadeec.md
MARKDOWN
1# View of the Waterfalls at Tivoli
2
3[View the canonical artwork record](<https://beta.wga.hu/artists/bidauld-jean-joseph-xavier-r46d21a0de50786/view-of-the-waterfalls-at-tivoli-r1a269997dadeec>)
4
5## Artwork
6
7- Artist: [BIDAULD, Jean-Joseph-Xavier](<https://beta.wga.hu/artists/bidauld-jean-joseph-xavier-r46d21a0de50786>)
8- Date: 1751-1800
9- Technique: Oil on paper on canvas, 51 x 38 cm
10
11## Commentary
12
13This painting is a typical example of the many studies, painted in oils on paper, that Bidauld executed in the open air on his extensive trips through the countryside during his Italian sojourn from 1785 through 1790. They are not sketches but pictures finished in the studio.
14
15## Related artworks
16
17- [Italianate Landscape](<https://beta.wga.hu/artists/bidauld-jean-joseph-xavier-r46d21a0de50786/italianate-landscape-rff080d33475a9a>)
18- [View of the Ponte Rocco, Tivoli \(detail\)](<https://beta.wga.hu/artists/bidauld-jean-joseph-xavier-r46d21a0de50786/view-of-the-ponte-rocco-tivoli-detail-r466bcf04593f13>)
19- …
20
21## Attribution
22
23[Web Gallery of Art](<https://beta.wga.hu/pages/about>)

It has the record and nothing else: a link back to the canonical page, the facts as a list, the commentary as prose, up to twelve related works by the same artist, and an attribution line. Every link is an absolute canonical URL. The angle brackets around each URL keep parentheses in it from ending the link early, and the backslashes in \(detail\) do the same job for titles.

How agents find it

Agents that read instructions get /llms.txt, a short Markdown file at the root of the site. It says what the collection is, that the canonical site is authoritative, and how to turn a page URL into a Markdown one: take the ID after the last hyphen and put it into /agents/artworks/{id}.md or /agents/artists/{id}.md. It links the sitemap for discovering records and the about page for terms of use. Every HTML page points to it, with rel="describedby" both in the response headers and in the <head>, and the footer links to it as well.

Each artist and artwork page also names its own Markdown version, in a Link header and in the <head>:

HTTP
1Link: </llms.txt>; rel="describedby"; type="text/markdown"
2Link: <https://beta.wga.hu/agents/artworks/r1a269997dadeec.md>; rel="alternate"; type="text/markdown"

Asking for Markdown

An agent that reads nothing can still ask. HTTP has let a client say which formats it prefers for much longer than anyone has wanted a page for a language model: the Accept header, with optional quality values. When a request for an artist or artwork page prefers text/markdown, WGA answers with a temporary redirect to the generated file:

HTTP
1GET /artists/bidauld-jean-joseph-xavier-r46d21a0de50786/view-of-the-waterfalls-at-tivoli-r1a269997dadeec
2Accept: text/markdown
3
4HTTP/2 307
5Location: /agents/artworks/r1a269997dadeec.md
6Vary: Accept

“Prefers” matters there. Browsers send something like text/html,…,*/*;q=0.8, and that wildcard matches Markdown as well, at a lower quality, so a naive check would send every visitor to a text file. WGA takes the best match for each type and lets Markdown win only when its quality is higher, or equal and Markdown was named explicitly:

AcceptResult
text/html,application/xhtml+xml,…,*/*;q=0.8 (a browser)the page
*/*the page
text/markdown307 to the Markdown
text/markdown, text/html307 to the Markdown

Vary: Accept tells Cloudflare, and any other cache, that the same URL has two answers.

The check runs before the session and page middleware. It reads the generated files, not the database, and redirects only when there’s a file for the record and the requested URL is one of that record’s published addresses. Anything else, such as an old alias or an unpublished record, falls through to the normal handler and gets its usual answer. A successful redirect costs a file lookup; building the page never starts.

What it deliberately leaves out

Making a catalogue easy for a model to read also makes it easy to copy, so the Markdown is limited on purpose. There’s no llms-full.txt with the whole collection in one file. It would be the most convenient thing to give an agent, an even more convenient thing to give a scraper, and a large file to regenerate every night. llms.txt says as much, and adds that the reproductions may carry rights belonging to the institutions holding the works, so their availability shouldn’t be read as a licence to reuse them.

Each document is built from a projection of the public record, not from the record itself. The Go types the renderer accepts have fields for a title, the facts, the commentary, the links and the attribution, and nothing else: no session or itinerary state, no administrative fields, nothing from the database’s own bookkeeping. Private data can’t reach a file even by mistake, because there’s nowhere to put it.

Everything in a document is bounded. Titles are cut at 300 characters, each fact at 500, the commentary at 4,000 and URLs at 2,048, and the related works stop at twelve. Every link has to be an absolute URL with no query string, fragment or credentials. Links to other records have to stay on the site’s own host, and anything pointing at the admin interface or the API is refused. Markdown’s own punctuation in the catalogue text is escaped, so a title with brackets in it stays a title. Records without a name or title, and artworks without a published artist, are skipped; anything else that breaks a rule stops the whole run, and the previous set of files stays live.

How it stays cheap

The files come from the same job that builds the sitemap, every night and once when the application starts. The job writes a complete set into a staging directory and checks it: exactly the files it expected, none of them empty, no symlinks. Only then does it switch the single marker that says which set is current, and remove the old ones. Readers always go through that marker, so an agent sees last night’s files or tonight’s, never a mixture, and a run that fails part-way leaves the previous set in place.

Serving a file means reading it from disk, with no database query on the way, so a request for a record that doesn’t exist gets a 404 without the database noticing. A successful response looks like this:

HTTP
1HTTP/2 200
2Content-Type: text/markdown; charset=utf-8
3Cache-Control: public, no-cache, must-revalidate
4Link: <https://beta.wga.hu/artists/bidauld-jean-joseph-xavier-r46d21a0de50786/view-of-the-waterfalls-at-tivoli-r1a269997dadeec>; rel="canonical"
5ETag: "e05a91ebfa132ab782fd5f2848654271667c1d1f1186718d15ca86155e128fe7"

There’s no cookie, and the canonical Link points back to the HTML page. no-cache lets Cloudflare keep a copy, as long as it checks with WGA before using it. The ETag is a SHA-256 of the file, so that check gets an empty 304 Not Modified until a night’s run actually changes the content.

request path
MERMAID
RENDERING…
1sequenceDiagram
2 participant A as Agent
3 participant C as Cloudflare
4 participant W as WGA
5 A->>C: GET /artists/…/view-of-the-waterfalls-at-tivoli-…<br/>Accept: text/markdown
6 C->>W: GET (Vary: Accept keeps the two answers apart)
7 W-->>C: 307 Location: /agents/artworks/r1a269997dadeec.md
8 C-->>A: 307
9 A->>C: GET /agents/artworks/r1a269997dadeec.md
10 alt nothing cached yet
11 C->>W: GET
12 W-->>C: 200, the file read from disk, ETag
13 else a copy is cached
14 C->>W: If-None-Match: the ETag
15 W-->>C: 304 Not Modified until a nightly run changes the file
16 end
17 C-->>A: 200 text/markdown

Cache misses still count. Requests under /agents/ go through the same per-client rate limits as the artist and artwork pages, at Cloudflare and again in the application, so a cheap file isn’t a cheap way around the limits.

Wrapping up

Nothing changes for people. A browser never prefers Markdown, so it gets the page as before, plus two headers it ignores. An agent that reads llms.txt or sends the right Accept header gets the same record in about 2 KB instead of 72 KB of HTML (before scripts and images), usually from Cloudflare’s copy. Agents that do neither still get the page, through the same rate limits as everyone else.

CITE THIS POST

Referencing this in a paper or thesis? Here’s a BibTeX entry.

galicz2026how.bib
BIBTEX
1@misc{galicz2026how,
2 author = {Galicz, Mikl{\'o}s},
3 title = {{How the Web Gallery of Art Serves Markdown to LLMs}},
4 year = {2026},
5 month = oct,
6 howpublished = {\url{https://blackfyre.ninja/blog/how-the-web-gallery-of-art-serves-markdown-to-llms/}},
7 url = {https://blackfyre.ninja/blog/how-the-web-gallery-of-art-serves-markdown-to-llms/},
8 urldate = {2026-10-10},
9 note = {Blog post, blackfyre.ninja}
10}
TAGS
06 · CONTACT

Have a system that needs building?

GET IN TOUCH → PROJECTS →
PRODUCT DATA SHEET
MG—01
BLACKFYRE
S/N MG-1985-1027
Miklós Galicz — Golang Advocate · Solution Architect
MODEL
MG—01 "Miklós Galicz"
SERIES
1985
ORIGIN
Nagykovácsi, Hungary
FUNCTION
Senior Full Stack Engineer · Solution Architect
CORE LANGUAGES
Go · PHP · JavaScript
SPOKEN
Hungarian · English · German
SERVICE LIFE
~20 years in software, ongoing
POWER SUPPLY
Coffee, 2–4 cups / day
DIMENSIONS
1 × human, standard size
OPERATING TEMP.
Calm under production incidents
CONNECTIVITY
[email protected] · github.com/blackfyre · linkedin.com/in/galiczmiklos
Less, but better. Specifications subject to continuous improvement.
● ● ●
MG—01 · SERIES 1985
№ MG-1985-1027
CERTIFICATE OF OPERATION
Certified Operator
Has located every documented feature of the MG—01 without reading the manual. Probably.
TIME
—
FEATURES
—
DATE
—
SIGNED
Miklós Galicz