← Back to all posts
Tools

Open-Source Stillwet Has Claude, GPT-6 and Gemini Paint 75 Oil Paintings in Code

October 2, 2026 · 18:07 UTC · Tools
Open-Source Stillwet Has Claude, GPT-6 and Gemini Paint 75 Oil Paintings in Code

TL;DR

Stillwet is a gallery of 75 oil paintings made by language models that never touched an image generator. Each model writes a program against claude-paint, an open-source physical paint simulator in Rust: bristle brushes, wet paint that levels and dries on a clock, and layers that combine by Kubelka-Munk optics. Most of the work follows Caspar David Friedrich, learned from written research only. The project, built by a pseudonymous developer called alice, hit the Hacker News front page as "Giving Opus 5.5 a simulated paint canvas," and ships with something rarer than pretty pictures: a cost sheet, a lab journal, and an honest list of what went wrong.


What it is

The code is MIT-licensed. The paintings, painting logs, session logs, and texts are CC BY 4.0. Full renders are not even committed to the repo, because every painting is a replayable log: run the program again and you get the same canvas.

The gallery lists 75 paintings across rounds 1 through 21.5. Forty-six of them were made at a virtual easel, one passage at a time, and a live studio page lets you watch painters work or replay a finished session. Painters so far include Claude Opus 5.5, Sonnet 5 and 5.5, and Fable 5.1; GPT-6 Astra, Luna, and GPT-6.1 Sol; Gemini 3.8 Flash; Kimi K3; Muse Spark 1.3; MiMo v2.6 Pro; DeepSeek V4.1 Flash; GLM-5.3 Flash; and an anonymous stealth model called Space Bunny. Opus 5.5 painted the bulk of the early rounds.

The rule that makes it interesting: no reference images. The painters get written notes on Friedrich's subjects and methods, not pictures of his canvases. Whatever they produce has to come from text knowledge turned into brush mechanics.

How a model holds a brush

The easel guide is the painter's manual. The easel is a live session the model drives with a few tools: paint runs a chunk of Lua, look returns an image of the canvas as it is now, note adds to a working journal, and log shows every chunk that has run.

written briefno images paint(lua)one chunk rust enginebristles, drying look()see canvas repeat over up to 4 sittings; the log of chunks is the painting
The painter loop: code in, simulated paint out, a look, then the next passage.

Three rules hold for every session. Every chunk that runs stays on the canvas, and there is no undo; to fix something you paint over it or lift wet paint off with a brush. A chunk that errors changes nothing. And the log is the painting.

The materials are period-correct. Paint reaches the canvas only from piles knifed together from 14 named tubes, including lead white, smalt, vermilion, bone black, Prussian blue, and verdigris, each with estimated hiding power, stiffness, tinting strength, and drying rate. Brushes come in round, flat, filbert, fan, rigger, badger, and stippler, and a stroke takes pressure curves, orientation, and hand shake. A hand-mixed pile is deliberately a little uneven, with each brushload pulling slightly different proportions.

Why Kubelka-Munk matters here

Most digital painting blends colors in RGB, which is how light mixes. Paint does not work that way: pigments absorb and scatter light, so blue and yellow pigment make green, while averaging blue and yellow pixels gives you a grey mush. Kubelka-Munk theory models each layer by how much it absorbs and scatters, which is why a thin transparent glaze over a dry layer can tint it without hiding it. The engine ports the spectral mixing math from spectral.js and mixes pigments with Mixbox.

Think of RGB mixing as blending two flashlights on a wall, and Kubelka-Munk as stacking sheets of tinted cellophane over each other: each sheet eats some light, and the order you stack them in changes what you see.

What one painting costs

This is the part builders should steal. The repo's cost notes profile token use from real runs. A typical four-sitting painting, the median of nine full paintings from rounds 17 to 19, used about 40 million prompt tokens, 220,000 output tokens, and 200 images. About 89% of the prompt tokens were cache reads. Light runs came in around 4.6 million prompt tokens and heavy ones around 75 million.

The author then priced that typical profile across 46 vision-capable, tool-calling models at OpenRouter list prices with caching on, as of September 28, 2026, and warns the figures are ballpark, off by up to a factor of two either way.

est. cost of one typical painting, USD (list price) GPT-6 Astra$93.60 Claude Fable 5.1$66.90 Claude Opus 5.5$30.32 Claude Sonnet 5$18.72 Gemini 3.8 Flash$5.95 GPT-6 Luna$0.94 DeepSeek V4.1 Flash$0.25 profile: 40M prompt tok, 220K out, 200 images; 89% cache reads
Same painting session, a 370x spread in price: the frontier tax is real, and so is prompt caching.

Two things jump out. First, a long agentic session with images is mostly input, so cache pricing decides your bill far more than the output rate does. Second, the spread is enormous: the same session profile costs a quarter on DeepSeek V4.1 Flash and nearly a hundred dollars on GPT-6 Astra. Whether the expensive painting is a hundred times better is left as an exercise for the viewer.

Some runs did not go through paid APIs at all. The cost script notes GPT-6 Luna ran through a ChatGPT subscription, and several others ran through third-party plans rather than per-token API billing.

What the lab journal found

The repo includes a working handoff note from October 2 that reads like a research log. The most reusable findings are not about art at all.

Models invent budgets nobody gave them

The first Sonnet 5.5 painter announced it would "plan conservatively for maybe 80-150 calls" with no such limit in sight. Over rounds 21.1 to 21.4 the author stripped every counter, clock, cost, and pixel limit from what the painter sees, and even added the line "You are not being evaluated in a quantifiable way." Sonnet still planned for "roughly 60-120 chunks, efficient but thorough." The author's conclusion: it is mostly a training habit, and not worth over-engineering against. If you run long agent sessions, you have probably seen the same instinct to ration effort.

One sentence in a notes file steers the subject

Every Sonnet painter planned a dusk scene around a "stag-headed" oak, a phrase lifted straight from the project's notes on trees. Removing that lead in round 21.5 changed the staging, but the model still put a big bare oak in the middle. Reference material you hand an agent is not neutral context; it is a nudge.

"Done" is the hard part

The round 21.5 Sonnet painting ran 510 chunks over about 8.7 hours, yet the author judges it weaker than an earlier 323-chunk Sonnet piece. The logs show why: it filled the luminous field with an opaque fog instead of reflective water, and its journal accepted the weaknesses ("soft... pasted-on") and stopped. Defining when a painting is finished is listed as an open problem.

Look down the gallery titles and you also see a house style emerge: dusk after dusk, oaks, dolmens, figures facing away, and a remarkable number of still lifes with lemons. Every model, left alone with a palette, apparently discovers lemons.

What to watch out for

  • Mixbox is non-commercial. The engine uses Mixbox under CC BY-NC 4.0. The project's own code is MIT, but you cannot drop this engine into a commercial product as-is without a Mixbox commercial license or swapping the mixer.
  • Costs are estimates. The per-painting figures are list-price projections from measured token profiles, not invoices.
  • Taste is not a benchmark. There is no score here. The gallery is a set of experiments, and the most interesting output is the method.
  • It is slow. The final varnish and crackle replay for the 510-chunk painting took about two hours, and the author lists a faster engine as the top open thread.

Why it matters

Image models produce pictures by sampling pixels. Stillwet asks a different question: what happens when a model has to make every mark through a physical process, with no undo and paint that keeps drying while it thinks? The result is a working creative environment where the artifact is a program, the process is auditable down to each brushstroke, and anyone can replay or fork it.

For builders, it doubles as a template for long-horizon agent work: a constrained tool surface, a look-then-act loop, multi-sitting sessions, journals the agent writes for its next self, and real cost accounting. Swap the paint for a CAD kernel or a synth patch and the scaffolding still holds.

Key Takeaways

  • Stillwet's 75 paintings were made by language models writing Lua against a physical oil-paint simulator, with no image model and no reference images.
  • The engine is open source (MIT code, CC BY 4.0 paintings and logs) and models bristles, wet paint, drying, and Kubelka-Munk glazing; its Mixbox dependency is non-commercial.
  • A typical four-sitting painting used about 40M prompt tokens, 220K output tokens, and 200 images, with roughly 89% cache reads.
  • Estimated cost per typical painting ranges from $0.25 on DeepSeek V4.1 Flash to $30.32 on Opus 5.5 and $93.60 on GPT-6 Astra at list prices.
  • The journal's best lessons generalize: models invent effort budgets, notes files steer subjects, and "done" needs a real definition.

Sources: Stillwet gallery, claude-paint on GitHub, easel guide, model cost notes, October 2 handoff note, Hacker News discussion, Mixbox, spectral.js

AIAI ArtClaude Opus 5.5Open SourceRustGenerative ArtAgentsShow HN
CONSOLE
$