Contents

Generating a Standard-Cell Library with an AI Agent

How Claude Code and the open-source LibreCell generator got within 4% of a hand-drawn 180 nm library

The flip-flop in Figure 1 was generated by LibreCell, an open-source layout generator, on a commercial 180 nm CMOS process. The dashed outline beside it is the exact footprint of its hand-drawn counterpart: a cell drawn by a professional layout engineer for a production standard-cell library used as the internal reference. I cannot show you that cell (it is production IP, and the generated layouts are pixelated for the same reason), but the outline carries the message. After eighteen logged phases of work, the generated library comes out at 1.04× the area of the hand-drawn one, at exactly the same row height, LVS-clean everywhere.

And here is the part that makes this write-up different from other layout-automation stories: most of the engineering was done by Claude Code, Anthropic’s CLI coding agent. The technology files, the DRC deck, the measurement scripts, the router patches, the geometric impossibility proofs, and the engineering logbook this post is sourced from: the agent wrote them, in sessions where I set goals and reviewed results, and it worked autonomously between review checkpoints.

Pixelated layout of the generated DFFPOSX1 flip-flop beside a dashed outline showing the hand-drawn cell’s smaller footprint
Figure 1: The generated DFFPOSX1 (pixelated for confidentiality) beside the footprint of its hand-drawn counterpart, extended with the same substrate-bias rail bands (shaded) so the two footprints are directly comparable. Same row height; the generated cell is one routing column wider.
Results at a glance
MetricResult
Library core area vs hand-drawn1.0415× (from ~2× at first bring-up)
Row height10.0 µm rail to rail, exact parity
LVSclean, 32/32 cells
DRC (in-house KLayout deck)clean on 29/32; the 3 residuals proven geometrically infeasible
Commercial sign-off (Cadence Assura)LVS clean on every cell; DRC residuals confined to the predicted class
EngineeringClaude Code, under goal-setting and review supervision

I want to convince you of two things at once: that, in this case, open-source silicon tooling could close to within 4% of hand-crafted quality, and that an AI agent can do the kind of deep, unglamorous, verification-driven engineering this requires. Neither claim survives hype, so the failures are left in. There are several.

Why this experiment matters

Standard cells are the bricks of digital design: inverters, NANDs, flip-flops, drawn once and instantiated millions of times. At work we design ultra-low-power circuits that operate near threshold, which means our library is custom, hand-drawn, and expensive to extend. Every new drive strength, every new cell variant, is layout-engineer time.

A generated library changes those economics. The layout stops being a frozen artifact and becomes a build product of two inputs: the netlists and the technology file. Need a different drive strength? Resize the transistors in the netlist, resimulate, regenerate; the layout follows in minutes. Need the library on a different process? Write a new technology file and rebuild: the router and placer work described below carries over, because it fixed the generator, not one library. Iteration and portability are the real prize this experiment was probing. The area comparison is just the test of whether the prize costs too much silicon.

The AI-in-EDA literature has an odd gap here. LLMs have been domain-adapted for chip-design assistance (ChipNeMo) and LLM agents can orchestrate complete digital RTL-to-GDS flows (ChatEDA). Reinforcement learning has generated standard cells at advanced nodes (NVCell). But there is, to my knowledge, no published work on an LLM agent doing transistor-level cell synthesis against a real process: adapting the tools, writing the technology description, debugging the router, and being graded by DRC and LVS. That intersection (AI agents and analog-flavored, transistor-level design) is exactly where I work, so I ran the experiment.

The test is honest because the reference is honest: a production library, drawn by a professional, on a real 180 nm process. Not a teaching PDK, not a synthetic benchmark. The question was never “can a tool draw an inverter”; it was “how close to a good human can an open-source generator get, when an AI agent does the adaptation work”.

The setup: a 2021 tool, a locked-down VM, and a real process

LibreCell (by Thomas Kramer) generates standard cells from SPICE netlists: it places transistors, routes them with a negotiated-congestion router, and emits GDS and LEF. It is research-grade software, last touched around 2021. Under the hood it is a showcase of classical combinatorial optimization: SMT solvers for placement legality, graph algorithms for transistor pairing, and a PathFinder-style router that resolves congestion by iterative negotiation.

Figure 2: The per-cell generation pipeline, and the loop the agent lived in: generate, verify, measure, patch, regenerate (each patch reruns placement and routing from scratch). Once the library froze, the commercial sign-off (Cadence Assura DRC and LVS) closed the loop.

The working environment was deliberately unglamorous: an ARM Linux VM with no sudo and no git. First task for Claude: get the tool running at all. That meant downloading tarballs instead of cloning, bootstrapping pip into a venv that shipped without it, relaxing a page of 2021-era dependency pins (one of them an outright invalid version specifier that modern pip refuses to parse), and patching an API ambiguity introduced by newer KLayout releases. Within the first session, a dummy-technology smoke test produced a GDS with “LVS result: SUCCESS” on it. None of this is hard engineering. All of it is the kind of friction that usually eats a weekend. It took the agent about an hour, narrated.

Then the actual work started: teaching LibreCell a real process. A LibreCell technology file is plain Python that defines everything: layer mapping, design-rule values, routing grid, cell geometry. Claude wrote one for our target, a near-threshold library with a channel length deliberately longer than the process minimum, two-metal routing, and the production library’s conventions.

Two constraints shaped everything that followed.

We could not run the foundry’s sign-off DRC locally. It only runs in commercial tools the VM did not have. So Claude built a partial DRC deck in KLayout’s rule language covering the core logic rules, plus a batch runner that summarizes violations per rule. This deck became the referee for the entire project. It is worth pausing on this: the agent wrote its own examiner, then spent seventeen phases being graded by it.

The process uses a single active layer with selective implants. LibreCell wants to draw n-active and p-active as separate layers, so Claude wrote a post-processing step that converts the generator’s output into proper full-width implant bands, positioned so that any two cells placed side by side have seamlessly merging implants. Cell abutment (the property that a row of cells is DRC-clean as a whole) became a recurring character in this story, and it is harder than it sounds.

First silicon-shaped objects

The first three cells (inverter, NAND, NOR) came out LVS- and DRC-clean after the usual geometry tuning, done by the agent itself: iterating transistor offsets, routing-grid parameters, and implant-band geometry against its own DRC deck until the violation count reached zero. Some early lessons set the tone:

  • Row height is bounded by well rules, not transistors. Cells that pass LVS can still be unmanufacturable because the well clearances don’t fit. The agent learned to check this class of failure before celebrating.
  • The generator cannot make a transistor-less cell, so the well-tap cell was built by a standalone script, parameterized to match whatever row geometry the logic cells use and invoked by the same build driver as the rest of the library.
  • “LVS passed” does not mean “electrically connected”. When we moved the power rails onto metal-1 and later narrowed them to the production library’s width, we found configurations where LVS passed on net labels while the physical supply strap was never drawn. Claude’s response was characteristic: it wrote a verification script that merges the actual conductor geometry (metal, contacts, diffusion, vias) into connected clusters and checks that the cluster containing each power rail also contains a transistor source or drain. From then on, “rails physically connected” was a checked property, not an assumption.

The deepest bug in this period was a one-character-class problem in the tool itself: the routing-grid generator used an exclusive range, so the track that was supposed to land on the top power rail was silently never created. PMOS sources had no path to VDD, while LVS kept passing on labels. The fix was small. Finding it required the agent to stop trusting the abstraction and dump the actual routing graph.

We also added a 4-rail row architecture (separate substrate-bias rails above and below the power rails) for body biasing, a requirement of our low-leakage design style that the implant-band post-processor and tap cell had to support adaptively.

The 2× wake-up call

Then I handed over the real test: the production library, three dozen cells from inverters to flip-flops with set and reset, as production-format netlists, plus the hand-drawn GDS as the golden reference. Claude wrote the netlist converter (including inferring port directions from flat netlists, which it got right on the first pass), brought up the library at the production row height, and we measured.

The generated cells were roughly 2× the area of the hand-drawn ones.

This kicked off the longest arc of the project. The first round of surgery targeted width:

  • A via-pad-aware router model. The router treated every grid node as if it might carry a full contact pad, which forced a coarse routing pitch everywhere. Claude rewrote the conflict model so that only actual via locations project pad-sized keep-outs while plain wires use wire-sized ones; vias get staggered dynamically instead of the whole grid paying for them. The flagship flip-flop went from 18.7 µm to 12.5 µm wide.
  • Variable gate spacing. The hand cells place gates at minimum poly pitch wherever diffusion is shared, and at contacted pitch only where a contact must sit between gates. Claude added this to the placer. The flip-flop reached width parity with the hand cell, a 44% total reduction.

This phase also introduced the project’s longest-running honest residual: with everything compacted, a handful of dense cells carried sub-minimum slivers between neighboring poly-contact heads, a violation class that would take another five phases to truly understand.

Quality gates a real library needs

A standard-cell library is more than small cells. Over two phases I raised the requirements a physical-design team would raise: power rails at the production library’s width, clock nets routed first and kept short in sequential cells, never stringing more than two gates on one poly net.

Each looks like a config flag, and none of them was. The narrow rails exposed the missing-track bug above. Clock-first routing turned out to work by ordering alone: Claude initially also discounted the clock net’s edge costs, measured the result, and found the discount made clock wirelength worse (the net started detouring around node costs). Pure priority ordering won, measured at −12% clock wirelength on the flagship flip-flop, with combinational cells proven geometrically identical via layer XOR. The poly rule was first enforced by banning horizontal poly routing entirely, which later turned out to be too blunt. More on that below.

The density campaign: learning from the human

With widths fixed, all remaining excess was vertical. I gave the instruction that defined the rest of the project: study the hand-made cells and learn how they do it.

Claude wrote an analysis script and dissected the hand flip-flop programmatically. The findings read like a layout designer’s bag of tricks: most poly contact heads jogged off the gate column with L-shaped poly; heads placed at free-form vertical positions in the channel between the transistor rows; mixed gate pitch, tight where diffusion is shared and relaxed where contacts need room; actives hugging the power rails at the minimum gap; second-layer metal barely used.

Each observation became a change:

  • Poly jogs were re-allowed. The earlier two-gates-per-poly rule is now enforced by a post-generation checker instead of by banning horizontal poly. Measure the property you care about; don’t outlaw the mechanism.
  • The routing grid got a half-pitch vertical grid, giving the router hand-like freedom to stagger contact heads.
  • And then a sequence of router model fixes, each found the same way: a cell fails to route, the agent reads the diagnostics, finds the model is wrong about physical reality, fixes the model. The diagnostics themselves were the first fix. Claude extended the router’s failure path from a bare “Failed to route” exception into a report naming the contested grid nodes, the nets fighting over them, and the mechanism by which each net claims the node (own wire, spacing footprint, or via-pad keep-out). Every subsequent fix traces back to one of these reports: phantom via pads projected by virtual terminal nodes, diagonal slip-throughs caused by using the wrong distance norm for pad keep-outs, mid-segment collisions missed by endpoint-only legality checks, gate terminals funneled through a single grid node when several rows were physically valid.
Terminal output of the router’s congestion diagnostics, listing contested grid nodes, the nets fighting over them, and the mechanism each net uses to claim the node
Figure 3: What the agent sees when a cell fails to route: the diagnostics it added to the router name the contested grid nodes, the nets fighting over each one, and how each net claims it. Real output from an XOR cell that was genuinely unroutable at this configuration.

One moment from this phase taught me more about agent engineering than any benchmark. An earlier build had “successfully” routed at a lower height, and Claude discovered the success was fake: the routes only fit through holes in the legality model. Fixing the model made the result worse on paper (the floor rose by two grid tracks), and the agent reported exactly that, unprompted: the honest model shrinks the geometry less than the broken one. If you want to trust an agent’s wins, watch how it reports its losses.

Core height came down from 1.21× of hand to 1.107×, with most of the library at the exact hand-parity rail gap.

Height parity, and a proof that three cells can’t be perfect

The next goal I set was concrete: library-wide core area within 1.10× of hand, DRC and LVS clean. Claude’s first move was a per-cell vertical audit, hand versus generated, which found the remaining waste immediately: hand actives are exactly as tall as the transistor width; ours were one contact-pad taller, because the generator parked the bottom source/drain contact on the first routing track with its diffusion pad sticking out below the channel. The hand cells keep that pad inside the channel.

The fix sounds trivial and isn’t: decouple the transistor’s vertical position from the routing grid. With the active edge placed independently and the contact landing inside the channel span, the cells dropped to the hand library’s exact row height, 10.0 µm rail to rail, for most of the library.

The head-dense cells (flip-flops, latch, mux, XOR class) needed a second idea, and this one the agent derived geometrically rather than empirically: raise the actives just enough that the static gate-poly bottoms clear a near-rail contact row by exactly the poly spacing rule. At one specific combination of row height, routing pitch, and offset, a fully legal contact-head row appears on every gate column next to each rail, and the flip-flops route cleanly there. The previous “raised offset” tiers from earlier phases had been physically illegal all along (head pads a few tens of nanometers from neighboring gates: the source of that long-running sliver class). The new model simply refused to construct them, and the legal tier replaced them.

Then the wall: XOR, XNOR, and the mux would not route cleanly at any offset. After roughly 25 failed configurations (seed sweeps, two finer grids, the hierarchical placer, which routed them but two columns wider, cheap second-metal, banned vertical second-metal, cheaper turns, asymmetric offsets, a router cost-model experiment), Claude stopped and did something better: it proved the wall is geometric. A legal near-rail head row requires the actives above a certain height; legal mid-channel head positions require them below a lower one. The intervals don’t intersect. The hand cells escape because a human pre-places L-shaped heads at transistor level, free of any routing grid, a capability the generator simply doesn’t have. Three cells therefore keep a small, documented residual violation class behind a per-cell model escape hatch, and the build system says so out loud.

I want to underline that sequence: an agent that stops burning compute and proves infeasibility instead is doing engineering, not search.

Result of the phase: row height equal to the hand library, area at 1.096×, 29 of 32 cells fully clean. Standard-cell libraries are a conspiracy of millimeter-scale constraints agreeing with each other at nanometer scale.

The width endgame

The final push targeted the last 9.6%: width. Two findings closed most of it.

Every cell carried identical dead margin at its edges. The outermost contact pads sat a full spacing rule away from the cell boundary, but abutment correctness only requires half the rule per side; the neighbor brings the other half. That margin, trimmed on both sides of every cell, recovered 7.7 µm across the library, found by a measurement script and harvested by a post-processing trim pass.

Pixelated layout of the generated INVX1 inverter beside a dashed outline showing the hand-drawn cell’s footprint
Figure 4: The smallest cell in the library after edge trimming (pixelated for confidentiality) beside the footprint of its hand-drawn counterpart, bias-rail bands included. The edge contact pads sit at the half-rule distance from the cell boundary, exactly like the hand cell’s.

The trim produced this project’s best catch-by-verification. The first implementation classified “rail-like” metal shapes by bounding box and replaced them with boxes, not knowing that the generator writes each net’s metal-1 as one merged polygon: rails, straps, and pads together. The “rail” replacement turned an entire power net into a solid metal sheet: shorted nets that DRC alone wouldn’t flag. It was caught because the agent diffed metal area between builds as a sanity check, saw +145%, and went looking. The fix (clip, never box) is two lines. The habit that found it is the whole methodology.

Placements are seed-stable except where they aren’t. Sweeping placement seeds across the fat cells showed nearly every cell converges to one width, except the largest flip-flop, which had a placement 5% narrower and DRC-clean at one particular seed. And then the punchline: regenerating it inside the build system with the same seed produced the wide placement again. The tool’s placement order is sensitive to process context beyond the random seed. Rather than chase non-determinism, the verified narrow cell is pinned as an artifact in the repository, with the story documented.

Final accounting:

StageLibrary area vs hand
First bring-up of the production library~2×
Via-pad router model + variable gate spacingwidth parity on the flagship
Densification campaign1.21× → 1.107× (height)
Row-height parity at 10.0 µm1.096×
Edge trim + placement hunt1.0415×
Table 1: The area campaign, stage by stage.
Bar chart showing generated library area falling from about 2 times the hand-drawn reference to 1.042 times across four project stages
Figure 5: The whole journey in one chart: generated library area relative to the hand-drawn reference, per project stage.

And the remaining 4%? It is not mysterious. It is the contacted-pitch tax: our source/drain contacts are routing-grid nodes, which forces the contacted gate pitch up to the grid pitch, while the human places contacts off-grid at the physical minimum. The per-cell excess list reads exactly like a ranking of contacted-gap counts: the big flip-flop, the clock buffers, the high-drive gates. We know precisely where every nanometer of the gap lives.

Six pixelated generated standard-cell layouts side by side at a common scale, from the small INVX1 inverter to the large DFFPOSX1 flip-flop
Figure 6: Six of the generated cells at a common scale, from the smallest inverter to the flagship flip-flop (pixelated for confidentiality).

Sign-off in a commercial tool

Everything above was graded by the in-house KLayout deck, which is partial by construction. After the library was frozen, we ran the real thing: the foundry’s sign-off DRC and LVS decks in a commercial tool, on the full library. LVS matched on every cell. DRC confirmed the in-house verdict: the violations are confined to the predicted three-cell residual class, plus a single stray violation on the half-adder. The referee the agent built for itself, from a subset of the rules, predicted the sign-off outcome.

What an AI agent is actually like as a layout engineer

Everything above is also a field report on working with Claude Code on a months-shaped problem: eighteen logged phases of supervised sessions. The workflow that emerged: I set a goal (“library-wide area within 1.10×, DRC and LVS clean”, later “reach parity”), and the agent planned, probed, built, and verified autonomously, running generation sweeps and DRC batches in the background, then wrote up the session in the project logbook it maintained. This post is sourced from that logbook. The agent documented its own failures well enough for me to quote them against it.

What it was genuinely good at:

  • Relentless systematic probing. Dozens of router configurations, seed sweeps, and height probes, executed and recorded without fatigue or attachment to any of them.
  • Reading and repairing research-grade code. The router fixes above are real algorithmic surgery in an unfamiliar codebase, driven by diagnostics the agent added itself.
  • Measuring before opining. Almost every claim about the hand cells came from a script interrogating the GDS, not from squinting at pictures.
  • Honest negative results. “The honest model makes the result worse”, “the success was fake”, “this is geometrically impossible and here is the proof”: all verbatim moves from the log.

Where it needed me:

  • Goals and acceptance criteria. “DRC clean” is a judgment call when your deck is partial and a violation class is provably unfixable without new tool capabilities. Deciding that three cells may carry a documented residual was my call, not its.
  • Direction changes. “The cell is 2× too wide”, “stop, go study the hand cells”: the highest-leverage sentences in the project were all human.
  • Skepticism as a service. The agent verifies what it thinks to verify. The habit of demanding physical-connectivity checks, area diffs, and abutment rows came from each bug that slipped, and the agent then automated the skepticism permanently.

And the failure modes, plainly: ~25 dead-end configurations on one cell class before stepping back to prove infeasibility; a post-processing bug that would have shipped shorted power nets had a sanity diff not existed; and a non-reproducible placement that had to be pinned like a museum specimen. None of these are disqualifying. All of them are familiar. They are the failure modes of an engineer: a tireless, fast, occasionally over-persistent engineer who documents everything.

What’s left, and advice if you try this

True parity needs two structural features in the generator, both now precisely specified by this project’s dead ends: off-grid source/drain contacts (to buy back the contacted-pitch tax) and transistor-level pre-placement of L-shaped poly heads (to make the last three cells clean). Both are router-and-placer projects, not configuration. After that: parasitic extraction and timing characterization, which we haven’t started. The payoff for finishing the flow end to end is the one hand layout can never offer: a library you can resize, resimulate, and reoptimize in a loop, and a port to the next process that starts with one Python file instead of months of polygon pushing.

If you attempt something like this, on any process, with any generator and any AI agent:

  1. Build the referee first. A runnable DRC deck and an LVS check, however partial, turn every subsequent claim into a pass/fail fact. Agents thrive on referees.
  2. Make the agent measure the golden reference, not describe it. Scripts that interrogate layout geometry beat any amount of looking at pictures, including the agent’s own.
  3. Set goals as verifiable predicates (“area ≤ 1.10×, DRC clean, LVS clean”), then let it run. The autonomy is real, but it is only as good as the predicate.
  4. Keep a logbook from day one, and have the agent write it. Mine doubled as the agent’s own long-term memory between sessions and as the source for this post.
  5. Demand negative results. The most valuable outputs of this project were a raised floor, a fake success exposed, and an impossibility proof. An agent that only reports progress is an agent you can’t trust.

The hand-drawn library took a professional layout designer a long time to craft, and it shows. I have now read those cells more closely than any layout I have ever drawn myself, and the respect is enormous. Getting an open-source generator and an AI agent to within 4% of that, at the same row height, with the gap itemized to the nanometer, took eighteen phases of logged work. The last 4% is not magic either. It is two well-specified features away. And there is no physical law saying the generated library has to stop at human parity. This project made no attempt to outperform the hand-drawn cells; it only tried to match them with a generator that was still missing two structural layout capabilities.

The LibreCell patches described here live in a private fork for now; I am happy to discuss the non-confidential details. The methodology needs no license at all.

References

  1. T. Kramer, LibreCell: layout generator for CMOS standard cells, codeberg.org/tok/librecell
  2. M. Liu et al., “ChipNeMo: Domain-Adapted LLMs for Chip Design,” arXiv:2311.00176, 2023
  3. Z. He et al., “ChatEDA: A Large Language Model Powered Autonomous Agent for EDA,” arXiv:2308.10204, 2023
  4. H. Ren and M. Fojtik, “NVCell: Standard Cell Layout in Advanced Technology Nodes with Reinforcement Learning,” Proc. DAC, 2021, research.nvidia.com
  5. KLayout: high performance layout viewer and editor, klayout.de
  6. Claude Code, Anthropic, claude.com/claude-code