Skip to content
All posts

4 min readAI, Delivery

Start with the problem, not the AI, part two: the estimator and the bridge

What it looks like when an AI system is measured against an expert on real work, and what 'thirteen of thirteen' actually means.

By Dave Tormey

Construction site with structural work underway, standing in for a bridge take-off

Earlier this year I wrote about a policing proof of concept that started with a problem and not with the technology. This is the next one, and it is the first time I have had a result I would put in front of a sceptic.

The problem

Quantity take-off for construction estimating: reading a set of drawings and working out how much of everything a job needs. It is document-heavy, slow, and the answers are checkable, which is exactly why it is a good first test for AI on expert work. If the system says there are 42 piles, you can count them.

What we did

We taught the system the job from one finished example: a real estimator's completed workbook for a real bridge. Not rules written by us; the method as he had actually done it. Then we gave it the drawings for a bridge abutment and asked it to produce the take-off.

The rule for the system was simple and non-negotiable. Every value it produced had to point at the page it came from. If it could not find a value in the drawings, it had to say so and ask, not estimate.

What happened

An estimator with decades of experience checked every value against his own drawings. Thirteen values out of thirteen were either found in the drawings and cited to the page, or referred back to him as a question. None were invented. He signed the result himself, in August.

I want to be precise about what that is and is not.

It is not "the AI did the estimator's job." Several of the thirteen were questions handed back to him, and his answers were what made the result complete. That is the design. The system does the hunting through pages; the expert keeps the judgement.

It is not a statistic across hundreds of jobs. It is one abutment, one workbook, one expert, one dated sign-off. It is a proof that the approach works, not a claim that it is finished.

What it is: the first time I have seen an AI system produce expert work where every single number could be traced or was honestly flagged, and an expert accepted it as his own.

Why "not found" matters more than the right answers

The thirteen correct citations are satisfying. The questions handed back are the important part.

Most AI systems, asked for a number they cannot find, produce a plausible one. It looks like the others. It is wrong. And because it looks like the others, nobody checks it until something is built to the wrong dimension.

A system that says "I cannot find the pile count on these sheets, what is it?" is annoying for about ten seconds and then it is the most valuable thing on the screen. It has told you exactly where the human needs to be.

What I took from it

Start with a job where the answers are checkable. Measure against a real expert, not a benchmark. Insist on evidence for every value and permission to say "not found". Then, and only then, argue about whether it is useful.

We did this in construction because the drawings do not lie. The same mechanism applies to any expert work that lives in documents. That is the part I am interested in next.

Dave Tormey

Dave runs TAGD Ventures in Brisbane. He has spent more than twenty years building and running software for government and regulated industry, including more than a decade as CTO, CIO and CISO of a software company serving law enforcement.

Get in touch