Historian Benjamin Breen Uses Opus 5.5 Agents to Find an Unnoticed 1615 Dodo Record
TL;DR
Historian Benjamin Breen published a step-by-step account on October 1 of using Claude Opus 5.5 agents to search the GLOBALISE corpus of Dutch East India Company records. The run surfaced a 1615 ship's log in which a Dutch crew on Mauritius "caught many tortoises, dodos [dodeersen], and some geese and parrots," a reference Breen could not find anywhere in the specialist literature. It lands one day after security researcher Carter Church's writeup of how GPT-6 Astra broke a Napoleonic cipher letter in about six hours of model time. Two independent cases, one pattern: point an agent swarm at a big, well-edited source base, keep a human expert on the question and the verdict, and old backlogs start moving.
The find: a dodo in a ship's log
The document is a log kept aboard the VOC merchant ship Wapen van Amsterdam, probably written by its captain, Isbrant Cornelisz van Petten. The ship made landfall on Mauritius in April 1615 to take on water and food before sailing on to the East Indies. The dodo line sits on folio 141 of VOC 1.04.02, inv. 1059 at the Nationaal Archief in The Hague, and the scan is public.
Breen, who has researched the early modern exotic animal trade and is writing a book on early modern extinctions, checked the result against the secondary literature, including Parrish's 2013 book The Dodo and the Solitaire. His conclusion: it "genuinely looks like this is a new addition to the timeline of dodo, which previously had a gap in the 1611-16 period." He is also candid that it is "not exactly earth-shattering," but argues it is publishable alongside a second find from the same search.
That second find is a 1638 account describing "field-hens" (velthoenderen), a Dutch word experts already associate with the extinct red rail. The reference was missed because a French scholar in 1890 translated it as perdrix, partridges. Opus 5.5 went back to the original manuscript and caught the error. A 136-year-old mistranslation is the kind of bug no linter will ever flag.
The workflow, which is the part you can steal
Nothing here needed a custom model. Breen's recipe is five steps, and the interesting engineering is in the second and third:
- Start from expertise. A specific research question from someone who already knows the field.
- Pick a large, well-edited corpus. GLOBALISE has spent 2022 to 2026 using AI to transcribe over 5 million handwritten pages of the VOC's Overgekomen Brieven en Papieren series, and its transcription viewer makes them searchable.
- Embed it. Download the sources and run them through an embedding model so you can search by meaning instead of spelling.
- Fan out. Use semantic search to pull candidate passages, then have frontier models read them and return a ranked shortlist for human review.
- Iterate or move on. If nothing answers the question, change the terms or the archive.
The embedding step does the heavy lifting with 17th-century Dutch, where one word can turn up in several spellings. Keyword search is a librarian who only finds the book if you spell the title exactly as the cataloguer did. Embedding search is the librarian who hears "big flightless bird, Mauritius, 1615" and walks you to the right shelf even when the page says dodeersen.
The agent layer is what changed this year. Breen describes Opus 5.5 spawning dozens of copies of itself to read sources in multiple languages, and when he handed the agents an API key, they ran their own embedding searches over new sources they found mid-run. They also wrote up their own process: Opus 5.5 produced a research dossier explaining how it located prior archival dodo finds before hunting for new ones.
The second data point: a Napoleonic cipher in six hours
Breen opens his post with Church's result, and it is the cleaner of the two as a reproducible build. A letter to General Auguste de Marmont, known only from a single plate in J. Vilcoq's 1969 article in the Revue historique des Armées, sat on the unsolved list at Satoshi Tomokiyo's historical-cipher site Cryptiana, dated 1807.
Church gave GPT-6 Astra one 1,202 by 1,836 pixel image. Astra cut the plate into rows, read the signs, and decided which of 175 hand-drawn marks were really the same sign written twice, collapsing them to 155 distinct signs across 1,300 cipher units. Only then did cryptanalysis start. It is a homophonic substitution cipher, the kind where one plaintext letter can be written as several different symbols, so simple letter-frequency counting fails by design.
Daniel Tant's previously published table gave values for 33 signs, covering 435 of the 1,300 units. The other two-thirds came from the letter itself, using simulated annealing scored against French n-gram statistics drawn from Hugo, Dumas, and Marmont's own memoirs. Five signs remain only partly resolved.
Two details make this more than a party trick. First, Church reran the solver from scratch with Marmont's memoirs and every Napoleonic text removed from its language model, and it recovered the same reading, which rules out the model simply pattern-matching to a text it had seen. Second, he shipped the receipts: a downloadable package with the 155-sign key, the revised transcription, and verify.py and reproduce.py scripts. Tomokiyo reviewed the solution, and Cryptiana now lists the letter as solved, noting that Church contacted him on 20 September 2026 and corrected the date to 1809. The plaintext is a military briefing on French and allied troop positions in the weeks before Austria invaded Bavaria.
Church's framing, quoted by Breen, is the one worth pinning above your desk: "I spent six hours of model time and a few evenings of my own. Tomokiyo says he's receiving solutions faster than he can record them."
Where it breaks
Breen is unusually honest about the failure modes, and they matter more than the wins if you are building research agents:
- No sense of significance. "What they can't currently do is ask the right questions or determine the significance of finds." Opus 5.5 and GPT-6 runs repeatedly spawned up to a dozen agents that "drilled down into minutiae and got utterly lost in the weeds."
- Unverifiable output. One multi-hour run produced an analysis of Inca khipu that Breen says he has no way to judge. A finding nobody qualified can check is not yet a finding.
- Happy accidents are still accidents. One wandering agent turned up an English captain, Jonathan Hide, who left the Dutch colony on Mauritius with ebony, "two sea cows," and about 20 giant tortoises he claimed he would release on St Helena. Great story. Not the question.
- Speculation stays speculation. The most tantalizing lead, that a dodo described by a Portuguese Jesuit in 1616 as an "ostrich" may be the one painted for the Mughal emperor Jahangir by Ustad Mansur, has no established chain of transmission. Breen calls it a hunch to dig into, and so should you.
His summary of the bottleneck is the line builders should take away: "The bottleneck will soon become not research findings themselves, but the attention of experts in niche topics." Generation got cheap. Verification did not.
Why builders should care
Both cases succeeded on the same three conditions: a corpus already digitized and cleaned by someone else, a narrow question set by a person who could recognize the answer, and an artifact that a third party could check (an archive folio, a reproducible script, a site maintainer's sign-off). That is a product spec. If you are building agentic search over legal discovery, patents, support tickets, or your company's own document graveyard, the lesson transfers directly: the corpus and the verifier are the moat, not the model.
It also hints at a market. Somebody has to sit between curious people running agent swarms and the handful of experts who can confirm what they found. Right now that job is Satoshi Tomokiyo's inbox.
Key Takeaways
- A new dodo record: Opus 5.5 agents surfaced a 1615 VOC ship's log entry recording dodo hunting on Mauritius, filling a 1611-16 gap that Breen could not find covered in the literature.
- A fixed mistranslation: the same run caught a 1638 red rail reference that an 1890 French translation had rendered as partridges.
- The recipe is ordinary: expert question, a big edited corpus (GLOBALISE's 5 million AI-transcribed pages), embeddings, agent readers, human verdict.
- A second, reproducible case: GPT-6 Astra solved the 1,300-unit Marmont cipher in about six hours, Church shipped verify and reproduce scripts, and Cryptiana marked it solved.
- The limit is judgment: agents wander into minutiae and cannot rank significance, so expert attention is the new bottleneck.
Sources: Benjamin Breen, Res Obscura, Nationaal Archief, VOC 1.04.02 inv. 1059, GLOBALISE (Huygens Institute), Carter Church, Breaking the Marmont Cipher, Cryptiana unsolved ciphers list, Vilcoq 1969, Revue historique des Armées (Persée), Runtime Wire