Written in the open, and in progress. Live, evolving work that will keep changing. How this book is written →

Appendix B — A Map of AI-for-Architecture Work

Author
Affiliation

Harvard John A. Paulson School of Engineering and Applied Sciences

Published

August 11, 2026

In this appendix, we consolidate the prior body of AI-for-architecture work that Architecture 2.0 builds on. We group established results into families, map each family to the principal lifecycle stages where cited work demonstrates activity, and show how our discipline in this book differs from any single family. We intend this map as a positioning statement rather than a comprehensive survey.

We have seen machine learning appear in computer architecture for decades. Early work proposed learned in-chip mechanisms and models of architectural design spaces. More recent systems couple learned methods directly to simulators, compilers, and implementation tools. The families below form the foundation our work draws on; because existing literature already provides comprehensive surveys (L. Chen et al. 2020), we offer a positioning aid instead (Table B.1).

When we examine these method families, most fall into three regimes, and each regime supports a distinct claim. Learned in-chip components occupy the mechanism regime, where we evaluate them against the mechanism’s own budget, such as mispredict rate at fixed storage and latency. Learned cost models occupy the design-space-model regime, where we evaluate them based on their predictive fidelity against the tool they approximate. Families that participate in the design process itself occupy the third regime. Benchmarks and datasets sit alongside these regimes as supporting evaluation infrastructure. We take the normative position that we must judge a complete third-regime result at the level of the architecture study rather than by the generated artifact alone. Confusing these regimes allows a bounded branch-prediction result to become evidence for trusting an agent with a floorplan, which it is not.

Table B.1: The unit of work in Architecture 2.0 is the architecture study and its reviewable claim, not any single family’s artifact. Each row names prior results, the principal lifecycle stages demonstrated by the cited work, and the chapters that develop it. The stage column uses only the six labels defined in Chapter 3.
Family of work Representative results Principal lifecycle stages demonstrated Developed in Key sources
Learned in-chip components Perceptron branch predictors and learned prefetchers, where learning is embedded in the mechanism itself and does a bounded tactical job Implement, Evaluate Chapter 1, Chapter 5 (Jiménez and Lin 2001; L. Chen et al. 2020)
Learned cost models and predictive DSE Surrogates and regression models that predict performance and power across a Design Space Exploration (DSE) environment from cheap features Explore, Evaluate Chapter 4, Chapter 5 (Ipek et al. 2006; Lee and Brooks 2006; Bai et al. 2021)
Autotuning and autoschedulers Search and cost models for loop transformations, array layouts, and accelerator lowerings Implement, Evaluate Chapter 5, Chapter 9 (T. Chen et al. 2018; Zheng et al. 2020)
Heuristic and BO design-space search Bayesian optimization, genetic algorithms, and evolutionary strategies for parameter selection Explore Chapter 5 (Bai et al. 2021; Krishnan et al. 2023)
RL for physical design, datapath, and flow tuning Reinforcement Learning (RL) for macro placement and datapath synthesis, and commercial design-space tools Explore, Implement Chapter 3, Chapter 9, Chapter 11 (Mirhoseini et al. 2021; C.-K. Cheng et al. 2023; Roy et al. 2021; Synopsys 2023)
LLMs for RTL and HLS generation Register-Transfer-Level (RTL) generated from natural-language specifications, with High-Level Synthesis (HLS) as the older lowering lineage Implement Chapter 6, Chapter 5 (Thakur, Ahmad, et al. 2023; Thakur, Blocklove, et al. 2023; Coussy and Morawiec 2008; Fu et al. 2023)
Agentic, tool-using exploration AgentDSE, an agentic framework for design-space exploration, demonstrates simulator-in-the-loop exploration in which an agent proposes configurations, invokes the simulator, and uses returned results to choose another candidate Explore, Implement, Evaluate Chapter 3, Chapter 6, Chapter 8 (Wang et al. 2026)
Benchmarks and datasets QuArch, a benchmark dataset for architectural reasoning, tests architecture knowledge through question answering Evaluate Chapter 4, Chapter 10 (Prakash et al. 2025)
ML for verification and validation LLM-drafted assertions and testbenches, treated as candidate verification collateral that requires independent execution and review Implement, Evaluate Chapter 7 (Yan et al. 2025; Qiu et al. 2024; Kang et al. 2024)
Chen, Lizhong, Drew Penney, and Daniel A. Jiménez. 2020. AI for Computer Architecture: Principles, Practice, and Prospects. Synthesis Lectures on Computer Architecture. Morgan & Claypool Publishers. https://doi.org/10.2200/S01052ED1V01Y202009CAC055.
Ipek, Engin, Sally A. McKee, Rich Caruana, Bronis R. de Supinski, and Martin Schulz. 2006. “Efficiently Exploring Architectural Design Spaces via Predictive Modeling.” Proceedings of the 12th International Conference on Architectural Support for Programming Languages and Operating Systems, ASPLOS XII, 195–206. https://doi.org/10.1145/1168857.1168882.
Lee, Benjamin C., and David M. Brooks. 2006. “Accurate and Efficient Regression Modeling for Microarchitectural Performance and Power Prediction.” Proceedings of the 12th International Conference on Architectural Support for Programming Languages and Operating Systems, ASPLOS XII, 185–94. https://doi.org/10.1145/1168857.1168881.
Zheng, Lianmin, Chengfan Jia, Minmin Sun, et al. 2020. Ansor: Generating High-Performance Tensor Programs for Deep Learning.” 14th USENIX Symposium on Operating Systems Design and Implementation (OSDI 20), November, 863–79. https://www.usenix.org/conference/osdi20/presentation/zheng.
Cheng, Chung-Kuan, Andrew B. Kahng, Sayak Kundu, Yucheng Wang, and Zhiang Wang. 2023. “Assessment of Reinforcement Learning for Macro Placement.” Proceedings of the 2023 International Symposium on Physical Design (ISPD). https://doi.org/10.1145/3569052.3578926.
Roy, Rajarshi, Jonathan Raiman, Neel Kant, et al. 2021. PrefixRL: Optimization of Parallel Prefix Circuits Using Deep Reinforcement Learning.” Proceedings of the 58th ACM/IEEE Design Automation Conference (DAC), DAC ’21, 853–58. https://doi.org/10.1109/DAC18074.2021.9586094.
Thakur, Shailja, Baleegh Ahmad, Hammond Pearce, et al. 2023. VeriGen: A Large Language Model for Verilog Code Generation.” arXiv Preprint arXiv:2308.00708, ahead of print. https://doi.org/10.48550/arXiv.2308.00708.
Thakur, Shailja, Jason Blocklove, Hammond Pearce, Benjamin Tan, Siddharth Garg, and Ramesh Karri. 2023. “AutoChip: Automating HDL Generation Using LLM Feedback.” arXiv Preprint arXiv:2311.04887, ahead of print. https://doi.org/10.48550/arXiv.2311.04887.
Coussy, Philippe, and Adam Morawiec, eds. 2008. High-Level Synthesis: From Algorithm to Digital Circuit. Springer. https://doi.org/10.1007/978-1-4020-8588-8.
Fu, Yonggan, Yongan Zhang, Zhongzhi Yu, et al. 2023. “GPT4AIGChip: Towards Next-Generation AI Accelerator Design Automation via Large Language Models.” Proceedings of the 2023 IEEE/ACM International Conference on Computer-Aided Design (ICCAD). https://arxiv.org/abs/2309.10730.
Wang, Chenyu, Jiahe Caroline Shi, David Kong, et al. 2026. AgentDSE: Reasoning-Augmented Architectural Design Space Exploration. https://doi.org/10.48550/arXiv.2606.21836.
Prakash, Shvetank, Andrew Cheng, Jason Yik, et al. 2025. QuArch: A Question-Answering Dataset for AI Agents in Computer Architecture.” IEEE Computer Architecture Letters 24 (1): 105–8. https://doi.org/10.1109/LCA.2025.3541961.

These families are not mutually exclusive. When we run an agentic design-space search, our workflow might call a learned cost model, invoke an autoscheduler, and emit RTL. A single architecture study in our work can touch several rows at once.

The learned in-chip components family carries a constraint that the others do not. A proposed component is only useful when its inference fits the mechanism’s storage and latency budget. The perceptron predictor provides a bounded example, keeping within a 4 KB hardware budget and delivering a prediction in one cycle (Jiménez and Lin 2001). This result ties predictive accuracy directly to the hardware resources we consume inside the mechanism.

Jiménez, Daniel A., and Calvin Lin. 2001. “Dynamic Branch Prediction with Perceptrons.” Proceedings of the Seventh International Symposium on High-Performance Computer Architecture (HPCA), 197–206. https://doi.org/10.1109/HPCA.2001.903263.
Kahng, Andrew B. 2018. “Machine Learning Applications in Physical Design: Recent Results and Directions.” Proceedings of the 2018 International Symposium on Physical Design, ISPD ’18, 68–73. https://doi.org/10.1145/3177540.3177554.
Yan, Zhiyuan, Wenji Fang, Mengming Li, et al. 2025. AssertLLM: Generating Hardware Verification Assertions from Design Specifications via Multi-LLMs.” Proceedings of the 30th Asia and South Pacific Design Automation Conference (ASP-DAC), ASP-DAC ’25, 614–21. https://doi.org/10.1145/3658617.3697756.
Qiu, Ruidi, Grace Li Zhang, Rolf Drechsler, Ulf Schlichtmann, and Bing Li. 2024. “AutoBench: Automatic Testbench Generation and Evaluation Using LLMs for HDL Design.” Proceedings of the 2024 ACM/IEEE International Symposium on Machine Learning for CAD (MLCAD). https://doi.org/10.1145/3670474.3685956.
Kang, Minwoo, Mingjie Liu, Ghaith Bany Hamad, Syed Suhaib, and Haoxing Ren. 2024. FVEval: Understanding Language Model Capabilities in Formal Verification of Digital Hardware. https://doi.org/10.48550/arXiv.2410.23299.

These families also occupy distinct regions of the territory we map in Figure 1. The learned-components, cost-model, and autotuning families live where computer architecture meets machine learning. Reinforcement learning for physical design and LLM-driven RTL generation live where machine learning meets electronic design automation (EDA). Learned models can use earlier-stage features to predict routing hotspots and downstream-flow outcomes (Kahng 2018). We view these surrogates as the cost-model family applied to physical design, where a cheap prediction guides our iteration before we run a more expensive downstream flow. Generated assertions and testbenches span both regions, but they remain candidate verification collateral that we must independently execute and review (Yan et al. 2025; Qiu et al. 2024; Kang et al. 2024). AgentDSE places tool feedback inside exploration, while QuArch contributes an Evaluate-stage test of architecture knowledge and question answering.

The stage column carries a telling asymmetry: Explore appears in four rows, Implement in six, and Evaluate in six, yet no row demonstrates Formulate, Interpret and Explain, or Review and Decide. We add those responsibilities directly to our normative standard for a complete architecture study.

Because we often combine these methods, what separates Architecture 2.0 from any single family is our unit of work. A learned branch predictor changes a single mechanism. A macro-placement agent proposes a candidate for one implementation step. An RTL-generating model produces a single artifact. Prior work judges each result on that bounded output alone.

Our work instead takes the complete architecture study and its reviewable claim as our fundamental unit. For us as architects, the question is not just whether a model produced a plausible artifact, but whether the whole loop supports an architectural decision that another architect can inspect and challenge. The families mapped here supply our methods. Our discipline supplies the standard of judgment that decides what their outputs are actually worth.

B.1 The Latent Architectural Space Map

A central theme across Architecture 2.0 is that machine learning models do not operate directly on raw physical silicon. Instead, they transform complex, discrete, high-dimensional hardware artifacts (such as Abstract Syntax Trees, Control-Data Flow Graphs, place-and-route netlists, and memory trace logs) into continuous or structured latent space representations. Evaluating how different AI paradigms construct, navigate, and decode these latent spaces determines their resolving power and failure modes (Table B.2).

Four paradigms organize this territory: Token & Code Embeddings, Graph & Topological Embeddings, Continuous Surrogate & DSE Spaces, and Policy & Reinforcement State Spaces. Whatever the paradigm, a candidate latent point earns trust only through an independent physical grounding path, and each paradigm fails in its own way when its predictions diverge from physical hardware behavior.

Table B.2: Mapping latent architectural spaces across representation paradigms. Each paradigm embeds discrete hardware artifacts into continuous or graph-based latent vectors. Physical grounding through independent signoff engines remains mandatory to prevent latent hallucinations from reaching silicon.
Latent Space Paradigm Underlying Representation Search & Inference Mechanism Physical Grounding & Decoupling Primary Latent Failure Mode
Token & Code Embeddings High-dimensional dense vectors derived from Large Language Model (LLM) Transformer activations over Verilog, SystemC, or Triton DSL code Autoregressive token sampling, temperature decoding, and beam search Executable compilation via Verilator or Yosys and functional testbenches Hallucinated Syntax & Vacuous Syntax: Generated code looks syntactically plausible but violates ABI, bit-width, or structural invariants.
Graph & Topological Embeddings Node and edge embeddings (Graph Neural Networks / GNNs, Hypergraphs) capturing spatial connectivity, CDFGs, or netlist floorplans Node classification, edge prediction, and spatial clustering Physical place-and-route tools via OpenROAD or Synopsys Innovus and static timing analyzers via OpenSTA Non-Routable Clustering: Surrogates cluster macros optimally in graph space, but extracted physical layouts collapse under severe metal layer congestion.
Continuous Surrogate & DSE Spaces Low-dimensional continuous manifolds mapping discrete architecture configurations (cache size, pipeline depth, bus width) to scalar PPA estimates Bayesian Optimization (BO), Gaussian Processes, and active learning acquisition functions Cycle-accurate simulators like SCALE-Sim or gem5 and gate-level signoff via Ansys RedHawk-SC Out-of-Distribution Exploitation (Optimizer’s Curse): Optimizers select extreme latent points whose predicted gains are inflated by estimation error rather than by real PPA improvements (Smith and Winkler 2006).
Policy & Reinforcement State Spaces Markov Decision Process (MDP) state-action vectors representing placement steps, tile schedules, or circuit transformations Deep Q-Networks (DQN), Proximal Policy Optimization (PPO), and Monte Carlo Tree Search (MCTS) Downstream physical DRC/LVS signoff and formal property checkers via Cadence JasperGold Reward Hacking: The policy maximizes proxy reward functions (e.g., estimated wirelength) by injecting false timing paths or unfeasible constraints.
Smith, James E., and Robert L. Winkler. 2006. “The Optimizer’s Curse: Skepticism and Postdecision Surprise in Decision Analysis.” Management Science 52 (3): 311–22. https://doi.org/10.1287/mnsc.1050.0451.

Mapping AI methods to their underlying latent space paradigms clarifies the evidence required to validate a candidate design. A model operating in token space establishes code syntax; validating its hardware claim requires decoding that token sequence into RTL and evaluating it through physical synthesis and static timing signoff.

These four latent paradigms (token vectors, topological graphs, continuous manifolds, and policy MDP state-action pairs) meet physical reality through one structural decoupling, latent candidate generation above and physical silicon signoff below (Figure B.1). To prevent latent failure modes (vacuous syntax, non-routable metal congestion, optimizer exploitation, and reward hacking) from corrupting the design loop, candidates must pass through an independent, evaluator-owned physical grounding funnel. This funnel forces candidate proposals through deterministic simulation, formal property verification, place-and-route netlists, and static timing signoff, ensuring that only physically grounded candidates receive human commitment authority approval before tapeout freeze.

Architecture diagram showing 4 Latent Space Paradigms at top, an Independent Evaluator-Owned Physical Signoff Barrier in the middle, and Verified Physical Silicon Signoff with Human Architect Commitment Authority at bottom.
Figure B.1: The Latent Architectural Space Map decouples latent candidate generation from physical silicon signoff. Machine learning models operate over four core representation paradigms (Token & Code Embeddings, Graph & Topological Embeddings, Continuous Surrogates, and Policy State Spaces). Each paradigm relies on an independent, evaluator-owned physical grounding path (Verilator, OpenROAD, SCALE-Sim, JasperGold) to prevent latent failure modes (Vacuous Syntax, Non-Routable Congestion, Optimizer’s Curse, Reward Hacking) from reaching silicon. Verified candidates require human architect and commitment authority signoff across five clean physical signoff checks before tapeout freeze.

B.2 Candidates, Not Decisions

In our framework, these model outputs are always candidates, never final decisions. An autoscheduler’s proposed schedule supports a performance claim only after we measure its execution on the target hardware. A proposed placement requires downstream place-and-route, timing, power, and congestion checks. Generated RTL requires functional and implementation checks. We in our accountable organization retain acceptance responsibility, while our verification and implementation owners retain signoff for their respective domains.

Candidates can mislead us through false confidence or hidden total costs. Generated assertions can pass vacuously while checking nothing that actually matters, so their apparent success is not independent evidence (Beer et al. 2001; Clarke et al. 2018) (Chapter 7). When we assess macro-placement methods such as AlphaChip, a reinforcement learning macro placement tool, a fair total-cost comparison must include the pretraining and fine-tuning runs actually used for the reported design, together with their compute (C.-K. Cheng et al. 2026). Neither apparent verification success nor a favorable cost comparison replaces our required acceptance and signoff.

Beer, Ilan, Shoham Ben-David, Cindy Eisner, and Yoav Rodeh. 2001. “Efficient Detection of Vacuity in Temporal Model Checking.” Formal Methods in System Design 18 (2): 141–63. https://doi.org/10.1023/A:1008779610539.
Clarke, Edmund M., Thomas A. Henzinger, Helmut Veith, and Roderick Bloem, eds. 2018. Handbook of Model Checking. Springer. https://doi.org/10.1007/978-3-319-10575-8.

B.3 Survey of Existing AI-Assisted Systems

When we survey the public examples available through July 2026, we see individual capabilities rather than a complete end-to-end answer. We can better understand their strongest claims by making the checks behind each result explicit. For each system, we must ask what artifact it produced, what test could reject it, how much of the broader architecture problem that test actually observes, and what we can inspect from the outside. Following those checks reveals both our progress and the remaining gaps, without mistakenly treating fundamentally different systems as steps in a single ranking.

Exact or inexpensive checks let us support strong claims over a narrow task. For example, AlphaTensor found a 4 by 4 matrix-multiplication algorithm for a two-element domain that uses 47 scalar multiplications rather than Strassen’s 49 (Fawzi et al. 2022). Because every candidate admitted an exact algebraic check, the system could confidently accept or reject solutions. Similarly, AutoTVM evaluates many tensor-program schedules on real hardware; we can measure the result directly, and most failures stay contained inside the tuning run (T. Chen et al. 2018). While these reliable checks make search practical, neither result gives us a complete system architecture.

When we rely on proxy and simulator feedback, we can support broader exploration, though we adopt the limits of our checking models. We see AlphaChip use wirelength and congestion proxies to guide macro placement (Mirhoseini et al. 2021). Meanwhile, ArchGym demonstrated that no single search algorithm dominates across our design spaces when we hold simulators and sample budgets fixed (Krishnan et al. 2023). To reduce costly target evaluations, Apollo transfers knowledge across related accelerator design spaces, whereas BOOM-Explorer pays hours per sample to execute a documented 7 nm flow (Yazdanbakhsh et al. 2021; Bai et al. 2021). We find each result meaningful within its chosen feedback source, yet we require stronger evidence before any of them can support a wider physical or system claim.

Functional tests, post-synthesis metrics, and fabricated silicon establish vastly different levels of hardware progress. Since functional success alone cannot guarantee hardware design quality, we look to systems like VerilogEval to check functional behavior (Liu, Pinckney, et al. 2023; Pinckney et al. 2024), while RTLLM layers on post-synthesis design-quality measures, though still stopping short of physical signoff (Lu et al. 2024). Pushing further, Chip-Chat reached tapeout for a small processor through extensive human-guided revision, and Enlightenment-1 produced a fabricated RISC-V core by expanding Boolean-function representations from input-output examples, an approach its authors frame as eliminating the manual verify-and-debug loop (Blocklove et al. 2023; S. Cheng et al. 2024). We view these as implementation achievements, yet neither system manages to deliver the full multi-tool system-architecture capability our moonshot demands.

Pinckney, Nathaniel, Christopher Batten, Mingjie Liu, Haoxing Ren, and Brucek Khailany. 2024. Revisiting VerilogEval: A Year of Improvements in Large-Language Models for Hardware Code Generation. https://doi.org/10.48550/arXiv.2408.11053.

We find human ratings and proprietary deployment reports much harder to evaluate from the outside. For instance, ChipNeMo reports task-specific internal gains, ChatEDA checks tool execution alongside human-rated request fulfillment, and DSO.ai announces 100 commercial tapeouts without a matched methodology (Liu, Ene, et al. 2023; He et al. 2024; Synopsys 2023). While these reports show our community moving learned methods into real workflows, they cannot single-handedly establish independent functional correctness, guarantee complete physical signoff, or prove a net benefit over a strong alternative.

Together, these examples demonstrate useful pieces of a broader design capability. We can use learned search and measured autotuning to explore our choices; we can leverage placement and RTL generation to produce physical artifacts; we can build tool-using systems to invoke real flows; we can rely on domain adaptation to support selected engineering tasks; and we can push generated designs all the way to silicon under tightly bounded conditions. Our moonshot aims to connect these demonstrated capabilities through current project states and appropriate design tools, all while retaining independent checks robust enough to reject an unsound result. We recognize that no single example has established that complete system yet.

Across these systems, we see that any given result only supports the specific claim its check can inspect. When we step back, we realize that check cost and fidelity, task scope, baseline quality, workflow coverage, and access to supporting material all limit what we as architects can conclude (Table B.3).

Table B.3: The strength of each result depends on its scope and how we evaluate it. Exact, inexpensive checks support strong scoped claims, whereas narrow measurements, proxies, weak baselines, and hidden supporting data, methods, or artifacts all limit what we can ultimately conclude.
System What it establishes The telltale limit
AlphaTensor (Fawzi et al. 2022) Learned search improved a fifty-year-old matrix-multiplication count for 4 by 4 matrices modulo two The algebraic construction admits an exact mathematical check, and a single scalar scores it
AutoTVM (T. Chen et al. 2018) Measured search can improve schedules when many candidates run quickly on real hardware The result covers tensor-program schedules, where failures usually remain safely contained within the tuning run
AlphaChip (Mirhoseini et al. 2021; C.-K. Cheng et al. 2026) A learned policy produces macro placements a production flow accepts The reward relies on proxy wirelength and congestion, and the matched baseline remains disputed
ArchGym (Krishnan et al. 2023) At equal sample budgets no search algorithm dominates, so our chosen harness and budget ultimately shape the result Because every candidate is graded by a simulator, the result inherently adopts its fidelity limits
Apollo and BOOM-Explorer (Yazdanbakhsh et al. 2021; Bai et al. 2021) Learned design-space exploration can successfully exploit domain-specific evidence Apollo transfers knowledge across related accelerator spaces to reduce costly target evaluations; BOOM-Explorer uses an open-source core but depends on commercial tools and pays hours per sample
VerilogEval and RTLLM (Liu, Pinckney, et al. 2023; Lu et al. 2024) We can measure RTL generation on shared tasks VerilogEval emphasizes functional simulation; RTLLM adds post-synthesis quality metrics; neither actually reaches physical signoff
Chip-Chat and Enlightenment-1 (Blocklove et al. 2023; S. Cheng et al. 2024) Generated logic has reached silicon, including a core that successfully boots Linux One relies on extensive human-guided revision; the other synthesizes a processor from input-output examples with a provable accuracy bound, rather than the full multi-tool system-architecture capability our moonshot poses
ChatEDA (He et al. 2024) An agent can plan and invoke a real tool flow directly from stated intent The benchmark checks tool execution and human-rated request fulfillment on 50 ChatEDA-Bench tasks using a simplified OpenROAD wrapper, lacking independent functional correctness or complete signoff
ChipNeMo (Liu, Ene, et al. 2023) Domain adaptation can improve selected internal engineering tasks The gains remain task-specific, general models stay stronger on certain subtasks, and outsiders cannot reproduce the proprietary evaluation
DSO.ai (Synopsys 2023) Learned flow-parameter search has reached 100 vendor-reported commercial tapeouts The evidence standard remains a press release, lacking any public methodology or matched baseline
Fawzi, Alhussein, Matej Balog, Aja Huang, et al. 2022. “Discovering Faster Matrix Multiplication Algorithms with Reinforcement Learning.” Nature 610 (7930): 47–53.
Chen, Tianqi, Lianmin Zheng, Eddie Q. Yan, et al. 2018. “Learning to Optimize Tensor Programs.” Advances in Neural Information Processing Systems 31: 3393–404. https://proceedings.neurips.cc/paper/2018/hash/8b5700012be65c9da25f49408d959ca0-Abstract.html.
Mirhoseini, Azalia, Anna Goldie, Mustafa Yazgan, et al. 2021. “A Graph Placement Methodology for Fast Chip Design.” Nature 594 (7862): 207–12. https://doi.org/10.1038/s41586-021-03544-w.
Cheng, Chung-Kuan, Andrew B. Kahng, Sayak Kundu, Yucheng Wang, and Zhiang Wang. 2026. “An Updated Assessment of Reinforcement Learning for Macro Placement.” IEEE Transactions on Computer-Aided Design of Integrated Circuits and Systems 45 (8): 3654–68. https://doi.org/10.1109/TCAD.2025.3644293.
Krishnan, Srivatsan, Amir Yazdanbakhsh, Shvetank Prakash, et al. 2023. ArchGym: An Open-Source Gymnasium for Machine Learning Assisted Architecture Design.” Proceedings of the 50th Annual International Symposium on Computer Architecture, ISCA ’23, 14:1–16. https://doi.org/10.1145/3579371.3589049.
Yazdanbakhsh, Amir, Christof Angermueller, Berkin Akin, et al. 2021. Apollo: Transferable Architecture Exploration. https://doi.org/10.48550/arXiv.2102.01723.
Bai, Chen, Qi Sun, Jianwang Zhai, Yuzhe Ma, Bei Yu, and Martin D. F. Wong. 2021. BOOM-Explorer: RISC-V BOOM Microarchitecture Design Space Exploration Framework.” 2021 IEEE/ACM International Conference on Computer-Aided Design (ICCAD), 1–9. https://doi.org/10.1109/ICCAD51958.2021.9643455.
Liu, Mingjie, Nathaniel Pinckney, Brucek Khailany, and Haoxing Ren. 2023. “VerilogEval: Evaluating Large Language Models for Verilog Code Generation.” IEEE/ACM International Conference on Computer-Aided Design (ICCAD). https://arxiv.org/abs/2309.07544.
Lu, Yao, Shang Liu, Qijun Zhang, and Zhiyao Xie. 2024. “RTLLM: An Open-Source Benchmark for Design RTL Generation with Large Language Model.” Proceedings of the 29th Asia and South Pacific Design Automation Conference (ASP-DAC). https://arxiv.org/abs/2308.05345.
Blocklove, Jason, Siddharth Garg, Ramesh Karri, and Hammond Pearce. 2023. Chip-Chat: Challenges and Opportunities in Conversational Hardware Design.” 2023 ACM/IEEE 5th Workshop on Machine Learning for CAD (MLCAD), 1–6. https://doi.org/10.1109/MLCAD58807.2023.10299874.
Cheng, Shuyao, Pengwei Jin, Qi Guo, et al. 2024. “Automated CPU Design by Learning from Input-Output Examples.” Proceedings of the Thirty-Third International Joint Conference on Artificial Intelligence (IJCAI). https://doi.org/10.24963/ijcai.2024/425.
He, Zhuolun, Haoyuan Wu, Xinyun Zhang, et al. 2024. “ChatEDA: A Large Language Model Powered Autonomous Agent for EDA.” IEEE Transactions on Computer-Aided Design of Integrated Circuits and Systems, ahead of print. https://doi.org/10.1109/TCAD.2024.3383347.
Liu, Mingjie, Teodor-Dumitru Ene, Robert Kirby, et al. 2023. ChipNeMo: Domain-Adapted LLMs for Chip Design. arXiv preprint arXiv:2311.00176. https://arxiv.org/abs/2311.00176.
Synopsys. 2023. AI-Designed Chips Reach Scale with First 100 Commercial Tape-Outs Using Synopsys Technology. Synopsys press release. https://news.synopsys.com/2023-02-07-AI-designed-Chips-Reach-Scale-with-First-100-Commercial-Tape-outs-Using-Synopsys-Technology.

We present this table as a set of scoped claims, not as a definitive ranking of systems. When we embrace shared tasks, we make it easier to inspect exactly where each claim stops.