← Back to all posts
Tools

OpenTPU Runs Qwen3 and LFM2 on a Decommissioned Kintex-7 FPGA With AI-Written Verilog

October 7, 2026 · 00:21 UTC · Tools
OpenTPU Runs Qwen3 and LFM2 on a Decommissioned Kintex-7 FPGA With AI-Written Verilog

TL;DR

openTPU is an Apache 2.0 LLM accelerator whose SystemVerilog, instruction set, simulator, compiler and profiler were written by coding agents, and it runs on real silicon. Felipe Sens Bonetto, a developer in Florianópolis, Brazil, posted it to Hacker News on October 6, where it drew 210 points and 275 comments. On a decommissioned Inspur datacenter card built around an AMD Kintex-7 FPGA with 4 GiB of DDR3, it decodes LFM2.5-230M at 82 tokens per second in 4-bit, runs ten models with their real weights, and produces the same tokens as its own simulator bit for bit. The hardware was evolved by a tournament loop in which Claude Opus 5.5 proposes and implements RTL changes and a verifier throws most of them out.


What is in the box

The whole accelerator lives in one repo of 537 files, 26 of them RTL, pitched as something you can read end to end: the hardware in SystemVerilog, an ISA of eight 32-bit words per instruction, a Python ISA simulator that doubles as the spec, a kernel language called ol with its compiler, and the host tools. A sequencer issues one instruction per cycle to a DMA unit, a four-column systolic matrix unit that multiplies int8 weights streamed from DRAM, an fp32 vector unit, and a quantizer. There is no cache and no hidden scheduling, so a trace shows exactly where the cycles went.

The card is an Inspur YPCB-00338, which Bonetto calls "a datacenter decommissioned board, really popular among hobbyists." It carries a Xilinx Kintex-7 xc7k480t, two DDR3-1066 channels with a 17.1 GB/s combined peak, and a PCIe link. The production bitstream runs at 133.33 MHz on LiteDRAM controllers that a small CPU inside the memory core calibrates in 12 seconds. Everything except the card runs on a laptop with Python and Verilator 5.

The numbers

The README's benchmark table covers 18 configurations measured on the card between September 29 and October 1. Decode is 64 greedy tokens after a 512-token prompt, with the host's argmax in the loop.

decode, 4-bit weights, int8 head, tokens/s (wall, higher is better) LFM2.5-230M82.1 Qwen3-0.6B30.7 Qwen3.5-0.8B23.3 Gemma 4 E2B12.1 LFM2-2.6B10.9 SmolLM3-3B8.7 Phi-4-mini 3.8B6.6 source: openTPU README, measured on the card 2026-09-29 to 2026-10-01
Decode speed tracks bytes per token: 82 tok/s for a 230M model, 6.6 tok/s at 3.8B, all at 82 to 94% of the DRAM peak.

The column the chart leaves out is the one that matters: DRAM bandwidth while decoding sits at 82 to 94% of the 17.1 GB/s peak in every row. Single-token decode on a card like this is memory bound. Think of the chip as a fast reader with a slow librarian: for every token, the entire weight set has to be carried from DRAM to the matrix unit, so the only lever is keeping the librarian walking at full speed. Bonetto says the same on HN: "for a TPU focused on inference the name of the game is memory bandwidth."

The 4-bit path, documented in docs/quant.md, uses FP4 values with two-level block scales at 4.25 bits per weight and keeps the LM head in int8. That cuts bytes per token by about a third and lifts decode by 40% on Qwen3.5-0.8B and 45% on Qwen3-0.6B and LFM2.5, at a perplexity cost the doc reports per model. Prefill is capped by the matrix unit's multiply rate instead, at 335 tok/s for LFM2.5-230M. The host is nearly out of the loop: the card runs one decode program compiled once and streams logits back while still running, adding 0.17 to 0.30 ms of host time per token.

Models bigger than the card

Mixture-of-experts models that do not fit in 4 GiB run with their experts streamed from host storage, per docs/offload.md. The host copies only the experts missing from per-layer slots in the card's DRAM. Measured on October 1 with 4-bit experts, LFM2.5-8B-A1B (8.5B parameters, 1.7B active) decodes at 10.6 tok/s with 98.5% of expert uses hitting the slots. Qwen3.5-35B-A3B (34.7B parameters, 3.0B active) manages 3.95 tok/s with a 62% hit rate and 153 MB streamed per token at 1.41 GB/s over PCIe, matching Hugging Face's 16 greedy tokens. A 35B model on a card with 4 GiB of DDR3 is slow, but it is running.

How the agents built it

Bonetto's top HN comment gives the origin: "After using AI to develop risc-v CPU cores, the same technique was used for developing openTPU." The accelerator "started able to produce only a few tokens per second" and reached 80-plus through what he calls a recursive self-improvement loop. The loop is documented in docs/tourney.md, and it is more verifier than agent.

hypothesisOpus 5.5, high implementOpus 5.5, high 6 gates: sandbox,lint, bit-exact,board, perf, synth accept rulearea or fmax championbranch a human merges a champion into main; agents never touch main or the gates smoke slot: $1.12 and 4.7 min with Opus 5.5 at high effort
One tournament slot. K slots run in parallel git worktrees; the first failed gate ends a slot as broken.

Each RTL component has its own champion branch, and every round K agents run in parallel git worktrees off it. A hypothesis agent sees the component source, the champion's numbers, a lessons file and a rotating focus, and may write only HYPOTHESIS.md. An implementation agent edits only the component's allowed files. Then the gates run in order: a sandbox check, Verilator lint, bit-exact RTL-versus-simulator tests, the same tests at the board's micro-architecture, a Qwen3-0.6B decode proxy that may not get more than 0.2% slower, and synthesis for area and an estimated fmax. A change is accepted only if it beats the champion on area at or above the target clock, or on clock speed without growing. Merging a champion into main is, in the doc's words, "a human decision."

The agents are Claude Code by default or Codex CLI, run unattended with least privilege: edits only inside the slot worktree, shell limited to lint, test and read-only commands, no web tools. The default model is claude-opus-5-5 at high effort for hypothesis and implementation, and at low effort for a scribe that logs one lesson per slot. The doc's measured smoke round, one slot on the collective unit, cost $1.12 and took 4.7 minutes; its budget guidance is $1 to $5 and 5 to 45 minutes per slot. That round shows the kind of thing the loop finds: the agent replaced three address multiplies in the GATHER instruction, nine chained DSP48 blocks, with running-sum registers. DSP count went from 9 to 0, estimated fmax from 75 to 128 MHz, and the area score fell 28%, with the perf proxy unchanged.

one night of component tournaments, 32 accepted changes (2026-09-24) LUT before131,915 LUT after82,224 est. fmax before41 MHz est. fmax after106 MHz yosys estimates of accelerator logic, excluding XDMA and MIG IP; not signoff
Per the September 24 status doc: 32 accepted winners overnight cut LUTs by 38% and lifted the estimated clock from 41 to 106 MHz.

The repo's status doc from September 24, the day before the first Vivado build, records what one night of tournaments did: 32 accepted winners across the fp operators, quantizer, sequencer, matrix unit, vector unit and AXI adapter, and DSPs down from 416 to 267 on top of the LUT and clock gains in the chart. The same doc then hands the human a to-do list that begins "Install Vivado into the vivado-docker volume. This needs your AMD login." The agents did the night shift and left Bonetto the chores that need a password.

The RISC-V rehearsal

The technique was proven first on auto-arch-tournament, an autonomous loop pointed at an RV32IM CPU. Each round an agent proposes a hypothesis, implements it in an isolated worktree, and runs it through riscv-formal, Verilator co-simulation and three-seed FPGA place-and-route on a Tang Nano 20K. Only hypotheses that beat the champion on CoreMark per MHz merge. The headline run took 73 hypotheses and under 10 hours to move the core from 301 to 577 iterations per second, 26% above VexRiscv's published 2.30 CoreMark/MHz, with 40% fewer LUTs. The verifier rejected 63 of the 73. The README's conclusion: "the loop isn't the moat - the loop is commodity. The artifact that survived 10 wins past 63 rejections wasn't the agent; it was the verifier."

HWE Bench best fitness (CoreMark iter/s, higher is better) Opus 5.5 xhigh983 GPT-5.5 xhigh525 VexRiscv (human)370 baseline core283 hwebench.com, best repetition per model, 15 rounds x 3 slots, 38 runs total
On the author's HWE Bench leaderboard, Opus 5.5 at xhigh effort is the only model so far to clear the human-designed VexRiscv reference.

That loop became a benchmark. HWE Bench scores models on how far they can push the same core in 15 rounds of 3 slots. The leaderboard lists Opus 5.5 at xhigh effort at 983, GPT-5.5 at xhigh at 525, and the human-designed VexRiscv reference at 370, from a baseline of 283 across 38 runs. Bonetto's one-liner on HN: "Opus 5.5 was the first to beat the human baseline." Several configurations have a single run, and the README calls the ordering "a record of runs, not a tested ranking." The repo credits Karpathy's autoresearch as the inspiration.

What the skeptics said

  • Who built it, really. From user mbgerring: "AI is now capable of developing its own inference hardware. No, it isn't. A human prompted an LLM to build a software simulation environment for hardware design, enabling an LLM, when prompted by a human, to optimize hardware designs against constraints in the simulation." That is a fair description of the tournament. The gates, accept rules and sandbox are hand-written and off limits to the agents.
  • FPGAs do not scale to frontier models. One commenter noted that even the largest FPGA on the market cannot hold a small Whisper model's weights on chip. Bonetto's reply is that FPGAs are for proving an architecture "before committing 100's of millions into a custom ASIC," and that GPUs are faster but "you can't make your own arch on GPUs."
  • A unit is not a chip. User chris_money202 called the design "the simplest unit of an entire AI chip," missing the PCIe fabric, Ethernet and multi-chip links that make a production TPU a system.
  • Numerics. A reader found non-standard floating-point behavior. Bonetto acknowledged "bugs/non standard behavior regarding very small or very big floating point values," which he says were "proven harmless for LLM inference." The status doc also notes that int8 greedy decoding can pick a different token than Hugging Face when the top two are nearly tied.
  • Margins and replication. The design closes 133.33 MHz "only just," with +0.032 ns of slack, and the tournament's fmax figures are yosys estimates meant to rank candidates, not signoff. The repo was created on September 24, had 204 stars when we looked, and every performance figure comes from the author's own card and scripts.

What to steal

Even if you never buy the card, the loop is the reusable part, and it is cheap: a slot costs about the price of a coffee and runs in minutes, because the agent only ever sees one component and the verifier does the heavy lifting. Three design choices make it work. The agent cannot edit its own grader. Every candidate must stay bit-exact with a reference simulator before anyone looks at speed. And a regression gate on a real workload stops "wins" that quietly slow decode, added after a vector-unit winner cut batched-attention speed. That is a well-run CI with the humans moved out of the proposing seat and kept in the merging seat.

The board bring-up is walked through in docs/board.md. Bonetto's next step is bigger hardware: "There are a few beasty FPGAs used in crypto mining coming my way." The stated goal is a chip that runs the inference of the models that improve it. Today those models run on someone else's GPUs, so for now the snake is only sniffing its tail.

Key Takeaways

  • openTPU is an Apache 2.0 LLM accelerator, from RTL to compiler to profiler, produced by Claude and Codex agents in a verifier-gated tournament and running on a decommissioned Kintex-7 PCIe card.
  • Ten models run with real weights. In 4-bit, LFM2.5-230M decodes at 82 tok/s and Qwen3-0.6B at 31, with DRAM at 82 to 94% of peak, and every configuration matches the simulator token for token.
  • Expert streaming lets a 35B mixture-of-experts model run from a 4 GiB card at 3.95 tok/s, moving 153 MB per token over PCIe.
  • The tournament uses Opus 5.5 at high effort to propose and implement, six gates to reject, and an accept rule on area or clock. A measured slot cost $1.12 and 4.7 minutes; one night produced 32 accepted changes and cut LUTs by 38%.
  • The same loop on a RISC-V core beat VexRiscv by 26% on CoreMark/MHz, and on the author's HWE Bench only Opus 5.5 has cleared the human reference so far.
  • Caveats: a human built the simulator and the gates, the numbers are the author's own, timing margin is thin, and FPGAs prove architectures rather than serve frontier models.

Sources: openTPU on GitHub, openTPU docs/tourney.md, openTPU docs/status.md, openTPU docs/board.md, auto-arch-tournament on GitHub, HWE Bench leaderboard, Hacker News discussion

AIOpen SourceFPGAHardwareInferenceClaude CodeAgentsHomelab
CONSOLE
$