Published: 
Aug. 26, 2026

Key takeaways

Nine in ten large infrastructure projects overrun on cost, schedule, or both, because variables like ground conditions, cost-schedule dependencies, and demand forecasts are genuinely hard to pin down upfront. @RISK Agent solves the speed problem, turning a project's alignment, ground conditions, and contract structure into a full working risk model in hours instead of days—it just still needs a human to catch which risks move together, like correlated ground conditions in adjacent tunnel sections. Working together, AI and analyst turn a fast first draft into a defensible, board-ready number that better prepares you for the tougher conversations around contingency and risk.

Why the toughest risks in infrastructure are still the ones buried in the ground

Nine out of ten large infrastructure projects go over budget, over schedule, or both. That's the finding from Oxford research by Flyvbjerg* spanning seventy years of megaproject data across every region with reliable records. Rail projects overrun on cost by an average of 44.7%, while overstating demand by 51.4%. Roads fare a bit better—a 20.4% average cost overrun, with roughly a coin-flip's chance that demand is also off by a fifth or more.

Better software, better contracts, more careful planning: none of it has moved these numbers. They've held steady for seven decades.

The reason isn't that people are bad at their jobs. It's that the variables driving infrastructure risk are genuinely hard to pin down before a project starts—what's actually in the ground, how a delay in one trade ripples into the next, whether a traffic forecast made today still holds up in year twenty of a concession.

Monte Carlo simulation has always been the right tool for that kind of uncertainty, and AI is now very good at building the scaffolding of a probabilistic model around all of that. What it isn't good at, on its own, is knowing which of those variables move together. That's the gap this article is about.

In our AI prompt engineering for risk managers field guide, we covered how to prompt @RISK Agent well. In this article, we put it to work inside a major infrastructure decision.

 

The case: A metro tunnel extension

A transit authority is putting together the business case for a 14-kilometre twin-bore metro extension. Budget: €1.3 billion. The alignment crosses three river sections and a stretch of mixed ground under a protected heritage district. The board doesn't want a single cost figure—it wants a distribution, and it wants to know what's driving the range.

Build with AI.

The analyst describes the alignment, the ground investigation summary, and the contract structure to @RISK Agent through Lumivero's MCP layer. The agent comes back with a proposed model: distributions for ground class in each tunnel section, productivity ranges for the TBM drives, contingency for utility diversions and the archaeological watching brief in the heritage zone, and a schedule network tying construction sequence to cost.

What might have taken a few days to build from scratch is ready to run in hours. The analyst checks it against the geotechnical baseline report, adjusts a few ranges, and runs it.

Verify with @RISK.

Ten thousand iterations later, the tornado chart shows ground conditions in two adjacent tunnel sections as the biggest cost drivers. Not a shock, given the geology in that stretch.

The correlation matrix is where the real issue shows up: the AI had modeled ground class in each section as independent of the others. Ground conditions in this kind of setting are spatially correlated—bad ground in one section makes bad ground in the next section more likely, because they share the same formation. This is established in tunneling work; it isn't a quirk of this project.

Fixing the correlation doesn't move the P50 cost estimate much. It widens the P90 quite a bit, because now the model is capturing a real tail risk instead of two independent long shots that mostly canceled each other out. That's a different contingency conversation with the board than the one the uncorrected model would have supported.

Defend with both.

With the correlation fixed, the analyst uses @RISK Agent's natural language interface to ask the model direct questions.

  • What happens to the schedule if TBM productivity in the two hardest sections comes in 20% below plan?
  • Which single mitigation—better ground treatment, resequencing the drives, more contingency—does the most to pull in the P90 tail?

The answers come from the simulation itself, not from whatever the model happens to know about tunnelling in general, and they're specific enough to drop straight into the business case.

 

Where AI fits

Away from this particular case, AI is already earning its keep across project infrastructure risk work more broadly: reading tender specifications and flagging clauses that won't hold up before they turn into disputes, working through geotechnical and inspection reports and documentation to surface the risky sections faster than a manual read would, proposing a first cut at distribution shapes from historical cost and schedule data.

None of it replaces the engineering judgement about ground conditions, contract exposure, or constructability. What AI does, when used inside @RISK, is get a defensible first draft in front of the person who has that judgement, faster.

 

4 things AI alone will not tell you

A few caveats are specific enough to infrastructure and civil engineering to be worth stating plainly:

  1. Ground risk is correlated, not itemized.
    Treating each borehole, each section, each structure as its own independent line ignores how geology actually behaves—adjacent ground conditions share the same formation and modeling them as independent understates the real tail risk.
  2. Cost and schedule aren't sequential, they're tangled together.
    A delay in one workstream rarely stays put. Models built section by section, rather than as a connected network, tend to miss this.
  3. The tail is where the rare events live, not the average case.
    A utility strike here, an archaeological find there, a consent that takes three months longer than planned. Any one of them is unlikely. Across an alignment this long, something in that category is close to certain. Base case estimates rarely capture it. The P90 case must.
  4. Demand forecasts don't hold up well over a long asset horizon.
    Flyvbjerg's data* shows rail projects overrunning cost by an average of 44.7% while overstating demand by 51.4%—meaning the two errors compound rather than cancel out. A model that stops at commissioning only tells half the story on anything financed against future revenue.

 

The modeler stays in control

AI accelerates. It does not decide. It speeds up the scaffolding of the model given relevant data, but it doesn't decide what's correlated, what's rare-but-catastrophic, or which demand assumption deserves a second look. That's still down to the engineer and the analyst, who know the specific geology, contract, and politics of the project sitting in front of them—none of which a general-purpose model knows unless someone tells it, and even then, unless someone checks.

Worth mentioning: a corrected model like this one shouldn't just live in a project folder and then disappear. A better model doesn't automatically mean better alignment—even a well-verified risk analysis can lose its value if it stays locked in one analyst's spreadsheet, disconnected from the rest of the program.

This is where Predict! and SharpCloud come in. Predict! centralizes risk data across the portfolio—turning individual models like this tunnel analysis into part of a single, audit-ready source of truth. SharpCloud then connects that data into a shared, visual decision layer, so the correlation the analyst just caught isn't just fixed in one model—it's visible to everyone tracking cost, schedule, and resourcing across the wider program.

Fed into that view, the model becomes part of the golden thread connecting cost, schedule, and resource decisions across the whole portfolio—not a spreadsheet that gets built once and never looked at again outside the project team.

This is one installment in a series on AI and Monte Carlo simulation in risk management. Next month we move to manufacturing, where the correlation problem shows up again in a different shape—shared supplier dependencies and yield variability, rather than ground conditions, are the thing AI is most likely to get wrong by assuming independence.

Catch up on the rest of the series:

Build with AI. Verify with @RISK. Defend with both. Buy @RISK today to get started.

Buy now

*Flyvbjerg, Bent. "What You Should Know About Megaprojects and Why: An Overview." Project Management Journal 45, no. 2 (2014): 6–19.

Manuel Carmona

Manuel Carmona 

Director, EdytrAIng, Ltd. 

Manuel Carmona is the author of Artificial Intelligence and Project Risk Analysis (Taylor & Francis, 2026) and founder of EdytrAIning. He’s a PhD researcher at the University of Westminster in London and consults and trains organizations in quantitative risk analysis, Monte Carlo simulation, and AI-integrated decision-making across the energy, infrastructure, and commercial sectors.