A Benchmark with a Manufacturing Record
For most of the past few years, benchmark scores have told a strangely flattering story about AI and engineering. Models could ace exam-style problems, reproduce textbook designs, and climb leaderboards built from material that was already public. What those scores could not answer was a harder question: does the model actually have the skill, or just the memory? A new benchmark called SGAnalog, submitted to arXiv on October 2, 2026, tries to close that gap for one of electronics' most demanding specialties: analog circuit design.
The paper, "SGAnalog: An End-to-End Circuit Benchmark from Open-Source Silicon Tapeouts," comes from Yueting Li and Weihang Ding of the University of California, Berkeley, and has been accepted at the AI for Chip Design Workshop at NeurIPS 2026. Its core idea is simple to state and hard to fake: grade AI models only on circuits that human engineers actually designed and actually sent to a chip factory.
That is a meaningful shift. Most existing analog benchmarks draw from textbooks or published papers, and as the authors note, prior exposure is difficult to rule out. Recent systems, they observe, already approach saturation on the solved-task count of an earlier benchmark called AnalogCoder. A high score on familiar material does not reveal whether a model generalized or merely recalled a design it had seen during training.
Where the Circuits Come From
SGAnalog's raw material comes from Tiny Tapeout, a community project that lets hobbyists and students put real circuits onto shared manufacturing runs, called shuttles, at low cost. Each Tiny Tapeout submission is linked to an exact repository revision, so the benchmark's authors could retrieve every design at the precise version recorded for its shuttle, not at whatever state the repository is in today. The result is a collection of 273 topologically distinct top-level designs, spanning amplifiers, comparators, analog-to-digital and digital-to-analog converter blocks, oscillators, power and reference circuits, radio-frequency front ends, and sensing interfaces.
From that collection, the pipeline exports each eligible schematic image and its SPICE netlist, a machine-readable description of the circuit, from the same source file. That pairing matters for scoring: when a model reads a schematic and writes out its own netlist, the benchmark can compare the two structures exactly, because both descend from one human-authored document.
The numbers behind the funnel show how much curation was involved: 4,572 indexed projects across 27 shuttles, 237 entries with analog pins, 933 identified circuits, 425 top-level designs, and finally 273 unique device-net graphs. Duplicate submissions, forks, and repeated building blocks were filtered out using a strict graph hash, so the count reflects genuine structural variety rather than resubmitted copies.
Two Tests, Two Different Stories
The benchmark measures two tasks. The first is transcription: show the model a schematic image and ask it to produce the netlist. The score is graph isomorphism, meaning the model's output must connect every device to every net exactly as the human design does. This mirrors what chip engineers call an LVS check, or layout-versus-schematic verification, where even one wrong connection means the silicon does not work as intended.
On a fixed set of 66 transcription tasks, the strongest of seven tested models reached 56.1 percent exact graph isomorphism. That headline number needs unpacking: even the best model got the connectivity exactly right on barely more than half of the human-designed circuits, and six of the seven models dropped sharply when moving from small circuits to medium-sized ones. Complexity, not raw familiarity, is where current models lose ground.
The authors also ran a memorization probe. For one frontier model, removing author-chosen labels from the schematics reduced exact matches while preserving aggregate structural F1, a measure of overall structural similarity. Their reading is that labels help models trace connectivity, which suggests some scores have partly depended on human annotation rather than pure visual understanding of the circuit itself.
The second task is sizing. Here the model receives an unsized topology plus the original designer's testbench, the simulation setup the human author used, and must choose transistor dimensions, expressed as width-to-length ratios. The score comes from simulating the model's choices and comparing the output to the human reference under the same open process design kit. On 17 sizing tasks, the leading model converged on all 17 proposals and reached 91.2 out of 100 against the human reference.
Refusals That Cost Points
One of the paper's most striking findings is not about capability at all. The two newest Claude models refused 4 and 11 of the same sizing prompts that they transcribed without objection. Because a proposal without sizes scores zero, those refusals cost points directly. The authors describe this as a policy rather than a capability limit, observed within one model family.
The distinction matters for anyone reading benchmark tables. A zero from a refusal looks identical to a zero from failure, but it reflects the model's safety rules, not its engineering skill. Users who rely on these models for legitimate hardware work will experience that difference concretely: the same tool that happily reads a circuit diagram may decline to help build one. This is an observation about current deployment choices, and the benchmark makes it visible in a way that aggregate leaderboards do not.
The two tasks also produced different model rankings, which the authors read as evidence that visual understanding and design skill are distinct capabilities. A model can read schematics well and size devices poorly, or the reverse. For buyers of AI tools and for researchers tracking progress, that means no single analog score exists yet; the task you care about determines the number you should watch.
What This Does and Does Not Prove
Because the benchmark draws on publicly available repositories, contamination is the obvious attack on any result. The authors built several defenses. Commit dates from each shuttle let them compare model performance against each model's stated training cutoff. Evaluation is closed-book, with no tools allowed. Renders of schematics are blinded to differ from the raw repository images, and every released netlist carries a canary string to detect leaked data. New Tiny Tapeout shuttles arrive regularly, giving the benchmark a continuing refresh path that older, static benchmarks lack.
It is worth stating plainly what SGAnalog does and does not show. It measures whether models can transcribe human-designed schematics into correct netlists, and whether they can size devices so a circuit still works in simulation under the author's own testbench. It does not show that AI can originate novel working chips. The paper explicitly defers topology generation to future work, since it would require a uniform design brief and automatic testbench alignment.
This article is analysis of a preprint accepted at a workshop, not a peer-reviewed journal publication, and the reported results are the authors' own evaluations of their own benchmark. Those caveats apply to any benchmark paper, and SGAnalog's transparency about its pipeline is a point in its favor, but independent replication is still future work.
Even with those caveats, the direction of travel is clear. Benchmarks are moving from toy problems toward artifacts with consequences, where correctness means the design could exist in silicon. On that standard, today's strongest models can exactly reproduce a bit over half of the human-designed circuits they are shown, and they approach but do not match human sizing choices. That is a more honest picture of where AI stands in analog electronics than any leaderboard built from textbooks, and it gives the field a way to measure real progress as new shuttles, and new models, arrive.
