Can ChatGPT Do Construction Takeoffs? New Data Says Not Yet
New benchmark tests how accurately AI can read blueprints, measure spaces and produce the numbers contractors use to bid.
Leading AI models, including ChatGPT and Claude, scored just 2% to 34% accuracy on a new benchmark that tests whether AI can read real construction blueprints.
Called Construction’s Last Exam (CLE), the benchmark from Miami-based Togal.AI puts AI models through 33 real-world construction tasks. The models have to read plans, measure spaces and materials, and produce the numbers contractors use to build bids and takeoffs.For example, a model is given a set of plans and asked to measure every balcony in a building. Its answers are then checked against the correct measurements to produce an accuracy score.
AI Challenges in Construction
Most AI progress over the past few years has happened in language, chatbots, writing, and code. Vision is a harder problem, and blueprints are some of the hardest images to read: dense, technical documents packed with symbols, scales and notations that even advanced AI models still struggle to interpret with the precision it takes to actually build something. That gap is a quiet bottleneck behind cost overruns, slower timelines and, at scale, the U.S. housing shortage. It also raises a less obvious risk: those same plans carry sensitive project and client information, and most AI tools aren’t built to safeguard it.
“When you hand your blueprints to a piece of software, you’re handing over everything about a project. Most AI tools can’t tell you where that data goes or who sees it, and in this industry, that’s not a small thing,” said Patrick E. Murphy, founder and CEO of Togal.AI. “Frontier models still struggle to read a blueprint accurately, let alone protect it.”
Since 2019, Togal.AI has built proprietary machine-vision models specifically to understand construction drawings, training them on its own proprietary dataset of tens of millions of construction plans. That approach is fundamentally different from general-purpose large language models like ChatGPT and Claude.
A Growing Trend
As AI moves into data-heavy industries, a pattern is emerging: build a benchmark to test how an industry-built AI stacks up against general-purpose tools like ChatGPT and Claude.
Other high-profile examples include Humanity’s Last Exam, which tests AI against expert-level academic questions; SWE-bench, which measures whether AI can solve real-world software engineering problems; and ARC-AGI, which tests advanced reasoning and problem-solving.
CLE is the first benchmark built specifically to measure AI performance in construction.
About Togal.AI: Based in Miami, Togal.AI builds proprietary AI technology specifically for the construction industry. Trained on a proprietary dataset of tens of millions of construction plans, the company has built one of the industry’s largest labeled proprietary construction datasets. Its patented AI helps contractors estimate, review and understand construction documents faster and more accurately. Togal.AI is SOC 2 certified, reflecting its commitment to rigorous data security and privacy standards for construction information. Following three consecutive years of roughly 300% growth, Togal.AI now serves more than 10,000 users across 30 countries, including 100 of the largest in the USA. The company’s mission is to make construction faster, more accurate and more efficient, starting with the very first estimate.






