Delivery Confidence Score: methodology
Version 7. A transparent, non-ML ranking.
This is a ranking, not a probability. A score of 70 does not mean a 70% chance of completion. It means the project ranks above one scoring 50 on the factors below. We do not claim it is calibrated: with only five stopped projects in our backtest, it cannot be.
NZ InfraWatch is a delivery-risk screening tool built on public data. It is not investment, valuation or procurement advice: do your own due diligence before relying on any score.
Every dataset behind the map and the score is public New Zealand government data under an open licence, each listed with its publisher and licence below. The analysis built on those inputs, the scoring, the risk ratings, the project curation and the commentary, is InfraWatch's own editorial work. It is neither produced nor endorsed by the agencies that publish the data.
The model (v7)
Each score is a weighted blend of eight factors, each set by a fixed rule from data we already publish: funding status, delivery stage, election/policy risk, policy alignment, delivery mode, time to delivery, project scale and delivery agency. No machine learning, no hidden inputs. Projects missing funding status or delivery stage show "insufficient data", never a guessed number.
How far off delivery is counts. Until v7 nothing in the model noticed the difference between a project opening this year and one due in the 2030s, so a fully funded programme a decade out scored about as high as a railway weeks from opening. Distance is itself risk: years are how long there is for funding to be withdrawn, a government to change or scope to be cut. Time to delivery is a declared prior, not a fitted weight. With four stopped projects our backtest cannot estimate it, and we would rather say that than imply the data chose it. A project with no published completion date is not treated as though it were distant: an undisclosed date sits just below neutral, and we do not invent one.
Delivery mode and election risk are not the same thing, and they are not contradictory. Delivery mode records the current government's revealed delivery behaviour: in 2023 it retained and revived roads while stopping public transport, cycling and speculative mega-projects, so a road scores higher on mode than a busway does. Those values are specific to this government and would invert under a change of government, which is exactly why delivery mode carries a deliberately small weight (11%). Election risk does not flip; it prices the forward risk that a change of government reverses a project. The two deliberately point opposite ways for a government-favoured project, mode crediting it for proceeding now, election risk discounting it for depending on the government staying, and the larger election-risk weight (23%) keeps that forward risk the dominant signal.
Delivery mode is counted once. Until v7 the policy-alignment factor also ranked transport by mode, which meant the same road-over-public-transport signal was applied twice and mode carried more influence than its stated 11%. Policy alignment now scores transport at sector level, so mode enters the score only through the mode factor. Correcting this moved public transport projects up about a point and roads down about a point, changed no project's band, and left the backtest's accuracy, precision, recall and F1 unchanged, which is what you would expect from removing a duplicated signal rather than a predictive one.
Projects near opening are scored as being delivered, not decided on. Once a project is under construction and within about a year of opening it is in commissioning, and the four factors that predict whether it happens at all have resolved: it will not be cancelled by an election, its delivery mode no longer signals what a government will or will not start, a mega-project this far along has already absorbed the slippage the scale prior warns about, and how well it fits the government's priorities can no longer determine whether a finished thing opens. For a project in that window those three factor values are lifted toward certainty (never lowered), so a project weeks from a confirmed opening sits in the 90s rather than being marked down for risks it has already passed. Funding status, delivery stage and time to delivery are already at their ceiling by then, so the score approaches the 100 a completed project scores.
| Factor | Weight |
|---|---|
| electionRisk | 23% |
| fundingStatus | 22% |
| status | 19% |
| policyAlign | 11% |
| mode | 11% |
| horizon | 10% |
| valueBand | 2% |
| ownerType | 2% |
| Score | Meaning |
|---|---|
| 80–100 | Strongest delivery signals |
| 65–79 | Favourable signals |
| 50–64 | Mixed signals |
| 35–49 | Weak signals: elevated risk |
| 0–34 | Weakest signals: high risk of stalling or cancellation |
Election risk
Rated the standard way, likelihood × consequence (ISO 31000), by fixed rules from published party stances, funding status and delivery stage. The risk source is specific: a post-election change of government. Projects funded privately, locally or through regulated charges read safe (the election is not their risk source); completed and cancelled projects report the fact, not a rating.
Which projects are included
We track 85 major projects, chosen editorially: no mechanical filter turns Te Waihanga's pipeline into this set. That is a limitation, because the backtest is drawn from the same set and cannot detect a bias introduced at selection, so the full list is public and open to criticism: tracked-projects.json. Every value comes from a source; where one is not disclosed we say so. Where a tracked entry is a section of a larger tracked corridor (for example Warkworth to Te Hana, which NZTA confirms is Section 1 of the Northland Corridor RoNS), it carries a part-of marker and its value is counted once, under the parent, in every portfolio and exposure total.
Backtest: the 2023 election
This is a calibration, not a validation. Some factors, delivery mode in particular, were derived from what happened in 2023, so testing them against 2023 measures fit, not predictive skill. The first genuine out-of-sample test is the 7 November 2026 election, scored against the ratings frozen beforehand, whatever it says.
4cc0611b4ba7362e…). After polling day we score these frozen ratings against
what actually happened, whatever it says.We scored every eligible project on its pre-election (Oct 2023) status and compared the ranking to documented post-election outcomes: 47 settled (4 stopped, 43 proceeded), sourced in backtest-2023.json. Positive class = "stopped".
| Model | Accuracy | Precision | Recall | F1 | AUC |
|---|---|---|---|---|---|
| Delivery Confidence Score (in-sample) | 91% | 50% | 25% | 0.33 | 0.77 |
| Naive B: unfunded ⇒ stopped | 94% | 67% | 50% | 0.57 | 0.81 |
| Logistic regression (leave-one-out CV) | 91% | 50% | 25% | 0.33 | 0.66 |
AUC is threshold-free ranking skill, with bootstrapped 95% confidence intervals (stratified, 3000 resamples): published score 0.77 [95% CI 0.48–0.97], cross-validated regression 0.66 [95% CI 0.28–0.98], funding status alone 0.81 [95% CI 0.53–0.99]. With only five stopped projects every interval is wide: read these as a sanity check, not a performance claim.
Limitations and full results
- Small N and a single election: these results characterise the 2023 change of government, not all elections.
- Several "proceeded" road projects were retained/advanced (RoNS) but are not yet under construction; outcome = "not cancelled".
- Structural factors taken from the current dataset as pre-election proxies (potential look-ahead bias).
- The logistic regression is reported under leave-one-out CV to avoid overfitting; its edge over the funding baseline is modest and election-specific.
- Since score v3 the electionRisk factor derives from current party stances (risk.js); pre-2023 reconstructions therefore see present-day stance text. This REDUCES one look-ahead channel (hand tags that encoded known outcomes) but does not remove stance-text leakage. The v1-heuristic recall drop vs v2 reflects removing that flattery, not a worse model.
Per-project results and learned weights: backtest-results.json.
How the data-quality rating is set
Separately from the score, each project carries a data-quality indicator: High, Medium or Low. It measures the completeness and currency of our inputs for that project, not how risky the project is. We check 5 things: funding status on record; delivery stage on record; a disclosed cost estimate; at least one published source linked; a recent recorded project event. "Recent" means a dated event within the last 2 years; completed and cancelled projects are not expected to produce new events. No gaps rates High, one gap Medium, two or more Low.
Missing or stale data is uncertainty, not evidence. A Low rating never lowers a project's score and never worsens its risk classification; it only qualifies how confidently the score is presented. That rule is enforced by a unit test. Each badge lists the actual gaps behind it, generated from the project's own record.
Sources and licensing
Every dataset the site uses, who publishes it, its licence, and how often we take it in.
Te Waihanga, National Infrastructure Pipeline
The national pipeline of infrastructure projects published quarterly as CSV by Te Waihanga, the New Zealand Infrastructure Commission. Licensed CC BY 4.0; data attributed to Te Waihanga. Our tracked set is an editorial selection from this pipeline plus public records.
GETS, government tender notices
Tender notices from gets.govt.nz, read via the official published RSS feed only and cached for 30 minutes. Crown copyright owned by MBIE, licensed for reuse under the NZGOAL Creative Commons framework; attributed to MBIE. Tender summaries on the map link back to the original GETS notice.
MBIE procurement open data, contract award notices
Historical contract award notices published quarterly by MBIE as open data CSVs, covering awards back to 2014. Licensed CC BY 4.0. Used as historical award-value context behind our tender analysis.
Stats NZ, boundary data
General electorate boundaries (2025) and territorial authority boundaries published by Stats NZ. Licensed CC BY 4.0; attributed to Stats NZ. Geometry is simplified for the web, so boundaries are indicative at high zoom.
OpenStreetMap, corridor geometry and town anchor points
The corridor lines drawn for linear transport projects, and the town coordinates used to place tenders on the map, are derived from OpenStreetMap data, © OpenStreetMap contributors, licensed under the Open Database Licence (ODbL). Corridors are simplified for the web and indicative of route; tender placements are expected locations, not surveyed addresses.
Electoral Commission, 2023 general election results
Official 2023 general election results published by the Electoral Commission at electionresults.govt.nz. Copyright the Electoral Commission; the source site states no open licence, so we reproduce only the factual results (winner, party and margin per electorate), attributed and linked back to the source.
News headlines shown on the site come from publishers' public RSS feeds, are shown as headline, source and link only, and remain the copyright of their publishers. Poll figures cite the pollster where they appear.
NZ InfraWatch is independent. It is not endorsed by, affiliated with, or acting on behalf of Te Waihanga, MBIE, Stats NZ, the Electoral Commission or any other agency whose data it uses.
How AI is used here
Not in the numbers. The Delivery Confidence Score is the fixed set of weighted rules set out above. Nothing is trained on anything, nothing is inferred, and the same project scored twice gives the same answer both times. Policy alignment and the matching of news articles to projects work the same way, as published rules applied consistently. Election risk ratings are editorial judgements made by a person, then applied through a stated likelihood and consequence matrix. The tender agent is a scheduled reader of the GETS feed: "agent" there means it runs on a timer, not that it reasons.
Where AI is used:
- Writing the software. The site is built with AI coding assistance. Changes are reviewed by a person and run against the site's tests before they ship.
- Drafting text. Project summaries and explanatory copy are often drafted with AI help and then checked against the cited sources. No figure comes from a model. Every dollar value, date and status is taken from a published source, and where a source does not disclose something the site says so rather than fill the gap.
The short version: AI helps build the site and draft its prose. It does not compute the score, and no number shown anywhere on this site comes out of a model.
Version history
How the model reached v7 (historical, superseded)
v1 was a weighted blend of funding status, delivery stage, election risk, a National Plan alignment rule, project scale and delivery agency.
v2 added a National Plan priority factor (a fixed sector rule favouring health/water/energy/roads). It was later merged into a single policy-alignment factor at v4 and no longer exists as a separate factor.
v3 restated election risk as an explicit likelihood times consequence (ISO 31000 and the 5×5 qualitative matrix), and shifted a little weight from funding toward election risk.
v4 added an explicit delivery-mode factor (road, rail, ferry, public transport, cycling/walking, derived from the project name) and merged the two overlapping alignment factors into one policy-alignment factor. Mode was 2023's clearest structural divider, but see the backtest: on that one election funding status alone ranked outcomes as well as the full model.
v5 added the commissioning lift, so a project under construction and within about a year of opening is scored as being delivered rather than still being decided on.
v6 stopped counting delivery mode twice. Policy alignment had also ranked transport by mode, which meant the same signal was applied through two channels.
v7 (current) added time to delivery, because nothing in the model had noticed the difference between a project opening this year and one due in the 2030s, and extended the commissioning lift to policy alignment, which was the last factor marking down work that is already built.