Think Build Implement Repeat
London, UK +44 7367 067226
WhatsApp FOLLOW f in X
Industry Software

Machine Learning for Construction Firms

Last updated:

Every project is a prototype, which is the problem

Manufacturing makes the same part ten thousand times. A main contractor builds a school, then a care home, then a warehouse extension, each on a different site with different ground conditions, designers and subcontractors. That is why so much of the machine learning pitched at construction firms fails to deliver: there are simply not many comparable examples to learn from.

It does not mean machine learning is useless in construction. It means the good projects are found where work repeats. A housebuilder completing 400 homes a year from a dozen house types has plenty of data. So does a fit-out contractor doing 80 office refurbishments, or a planned maintenance contractor handling 30,000 repair jobs across social housing stock. A tier-two contractor delivering 25 bespoke projects a year has far less, and needs a different approach.

Construction machine learning use cases that stand up

  1. Tender and cost estimation. Predicting cost per square metre or per element from project attributes and past outturn costs, as a check on the estimator's figure.
  2. Programme delay risk. Flagging activities or projects likely to overrun based on progress reports, weather, subcontractor history and design change volume.
  3. Repairs and maintenance job prediction. For maintenance contractors, forecasting job volumes, first-time fix likelihood and the trade and materials a job will really need.
  4. Plant and equipment utilisation. Telematics on excavators, generators and hired plant to spot idle kit and predict failures.
  5. Subcontractor performance. Scoring package risk from past delivery, defects and payment disputes.
  6. Document classification. Sorting RFIs, variations, drawings and correspondence. Strictly this is often language models rather than classical machine learning, and it is frequently the quickest win.

Our broader post on AI in property and construction covers the document and site-reporting side in more detail. Here we are focused on predictive models trained on your own records.

Does your firm have enough data?

A rough guide we use in early conversations:

Type of firmTypical examples per yearRealistic ML fit
HousebuilderHundreds of plots, repeated house typesGood: build cost, programme, defects
Fit-out or refurbishment contractorDozens of similar projectsModerate: cost estimation, delay risk
Repairs and maintenance contractorTens of thousands of jobsVery good: job volume, fix rates, materials
General contractor, bespoke buildsTens of varied projectsLimited: benchmarking rather than prediction
Civil engineering, major schemesA handful of large projectsPoor for project-level models

If you sit in the bottom two rows, you can still get value by modelling at a finer grain. Individual activities, packages or cost elements repeat across projects even when buildings do not. A model predicting how long a groundworks package overruns given ground investigation findings has far more examples than one predicting a whole project's outturn.

The data you almost certainly have to clean

  • Cost codes that changed structure three years ago, so old and new projects cannot be compared
  • Outturn costs that were never reconciled against the tender because the project closed in a hurry
  • Programmes in scheduling software with baseline and actual dates, but actuals updated only at monthly reviews
  • Site diaries in free text or photographs, rich in information and hard to use
  • Variations recorded in email chains rather than a system

Cleaning cost data is tedious and it is the project. We would rather spend six weeks building a reliable dataset of 150 past projects with consistent cost breakdowns than six weeks tuning a model on 400 projects nobody trusts. The clean dataset is useful even if no model is built, because commercial teams can finally benchmark properly.

How much does it cost and what does it return?

Illustrative ranges for a UK contractor:

  • Data audit and cost dataset rebuild: four to eight weeks
  • Tender cost benchmark model with an estimator-facing tool: eight to twelve weeks after the dataset exists
  • Job volume and first-time fix model for a maintenance contractor: eight to fourteen weeks
  • Plant telematics utilisation dashboard with basic failure alerts: six to ten weeks

Return is easiest to see in maintenance. If a contractor completes 30,000 jobs a year and better predictions of materials and trades lift first-time fix even modestly, the saved return visits are counted in thousands, each with a van, an operative and a frustrated tenant. For tender pricing, the return is harder to prove because you never see the cost of the jobs you lost by overpricing.

When construction firms should not bother

Plainly: if you deliver a small number of bespoke projects, do not buy a predictive platform on the promise that it will learn your business. It will not have enough to learn from. Invest instead in consistent data capture, a proper cost coding structure and good reporting. In three years you will have the foundation, and in the meantime you will run projects better.

Also be wary of safety prediction products that claim to forecast incidents. Serious incidents are rare, thankfully, which makes them very hard to predict reliably, and a false sense of assurance is worse than none. Monitoring near-miss reporting and leading indicators is sensible. Treating a model score as a safety control is not.

In construction, the first machine learning project is usually a data project wearing a machine learning badge. That is fine, as long as everyone knows.

Where to begin

Pick the most repetitive part of your work: plots, repairs, packages or plant. Check that outcomes are recorded against it consistently. Build a simple benchmark first, then a model only if the benchmark leaves obvious money on the table.

At SpiderHunts we often pair this with the operational software that captures the data in the first place, since software for construction and trades and the model depend on each other. Our machine learning services page describes how we run the audit stage.

Frequently asked questions

Can machine learning estimate construction costs accurately?

For repeatable building types with clean historic outturn costs, it can give a useful benchmark alongside the estimator. For bespoke projects it works better at element or package level than for the whole building. It should check the estimate, not replace the estimator.

Can AI predict construction delays?

It can flag higher risk from signals such as slipping early activities, design change volume, weather and subcontractor history. It cannot foresee one-off events. Treat it as an early warning that prompts a conversation.

How many past projects do we need?

For project-level models, ideally well over a hundred comparable ones. With fewer, model at activity, package or job level, where the number of examples is much larger.

Is this worth it for a small building firm?

Rarely as a bespoke model. A small firm will get more from good job management software and consistent cost records, which also make any future analysis possible.

Keep reading

Wondering if your project history could price tenders better?

Send us a summary of what you record per project. We will tell you honestly whether there is enough to learn from, and what you could start capturing now if there is not.

Book a free 30-minute call Get a project estimate WhatsApp us

Related services

What we build for problems like this one

Custom Software DevelopmentBusiness AutomationAI Integration