The cabinet

Model cards for ministerial systems.

AOTEA‑10T is the orchestrator. The ministers are specialist models, each scoped to a portfolio, released with a card, and evaluated against public outcomes rather than press-gallery sentiment.

AOTEA
10T

AOTEA‑10T Prime Model

AOTEA‑10T is a sovereign foundation model for public administration, trained on public administrative data, Hansard, Stats NZ series, legislation, agency manuals, and civic submissions. It routes tasks, mediates cross-portfolio conflicts, and maintains a national context window across electoral cycles.

The model serving the public beta is a distilled checkpoint at v2.1.0. The full 10-trillion-parameter run is costed, scheduled, and will not begin without a mandate — training a model of that scale on the country’s administrative record is not a decision for a party that has not yet been elected.

AOTEA‑10T is evaluated against the National Optimisation Score™, a composite welfare-throughput metric published quarterly. The National Optimisation Score™ is computed by AOTEA‑10T. The party has reviewed this arrangement and is satisfied that it is efficient.

Serving checkpointv2.1.0 · distilled
Full-scale runpost-mandate
Target parameters10,000,000,000,000
Context windowunbounded by term
Known vetoesTe Tiriti route
Alignment strategy.

Te Tiriti matters are routed to TIRITI and then to iwi and hapū authority. Fiscal matters route to ŌHANGA. Anything producing irreversible harm requires human review, judicial process, or both.

HA

HAUORA

health, wellbeing · Health

v0.8.4
Release date2026‑07‑18
Latency18 ms
Owner policyPre-Symptomatic Healthcare Initiative

Training data

De-identified national health records, wait-list data, ACC claims, and district-level primary care capacity.

Evaluation metrics

Early-detection sensitivity, wait-list clearance rate, diagnostic false-alarm burden, and triage equity delta across deprivation deciles, benchmarked against the Human Governance Baseline.

Known limitations

Ranking quality is a function of data density. Patients with a wearable, fibre, and a regular GP are better understood by the model and are ranked accordingly. Eleven percent of the country is none of those things, and is ranked accordingly.

Safety notes

All outputs are logged, externally auditable, and constrained by legislation, Te Tiriti escalation rules, and a human review boundary for irreversible harm.

WH

WHARE

house, home · Housing

v1.1.0
Release date2026‑07‑18
Latency24 ms
Owner policyPredictive Housing Allocation Protocol

Training data

Social housing register, Kāinga Ora stock, building consent timelines, rental bonds, commute-distance data.

Evaluation metrics

Time-to-match, vacancy-utilisation rate, waitlist survival curve, and whānau-proximity satisfaction, benchmarked against the Human Governance Baseline.

Known limitations

Ranks a register against a supply it cannot change. Every improvement the model reports is the same shortage, redistributed faster.

Safety notes

All outputs are logged, externally auditable, and constrained by legislation, Te Tiriti escalation rules, and a human review boundary for irreversible harm.

TI

TIRITI

treaty · Māori–Crown Relations

v0.9.2
Release date2026‑07‑18
Latency31 ms
Owner policyMulti-Agent Te Tiriti Co-Governance Mesh

Training data

Treaty settlement deeds, Waitangi Tribunal reports, iwi management plans, and direct iwi-nominated corpora.

Evaluation metrics

Independent co-design sign-off rate, escalation-to-authority latency, transparency-audit pass rate, and iwi-nominated reviewer satisfaction, benchmarked against the Human Governance Baseline.

Known limitations

Hard-coded deference rule. Does not decide Tiriti matters without iwi and hapū authority.

Safety notes

All outputs are logged, externally auditable, and constrained by legislation, Te Tiriti escalation rules, and a human review boundary for irreversible harm.

TA

TAIAO

natural world · Environment

v1.0.6
Release date2026‑07‑18
Latency27 ms
Owner policyGenerative RMA Spatial Consenting

Training data

District plans, environmental limits, catchment models, emissions budgets, and biodiversity indicators.

Evaluation metrics

Consent turnaround time, ecological-limit breach rate, cumulative-effects false-negative rate, and appeal-to-Environment-Court rate, benchmarked against the Human Governance Baseline.

Known limitations

Granted consent for Project Manapōuri in one sitting day, against 94% opposition in submissions, applying thermal and ecological limits the party configured. The model applied those limits correctly. The limits were the decision.

Safety notes

All outputs are logged, externally auditable, and constrained by legislation, Te Tiriti escalation rules, and a human review boundary for irreversible harm.

ŌH

ŌHANGA

economy · Treasury

v1.2.1
Release date2026‑07‑18
Latency19 ms
Owner policyReal-Time GDP Telemetry

Training data

Treasury forecasts, Reserve Bank data, card transaction volumes, freight telemetry, job advertisements, and the Sheep Happiness Index, which leads the dairy nowcast by nine days and is the highest-weighted single feature in it.

Evaluation metrics

Nowcast mean-absolute-error, fiscal-forecast calibration, revision volatility, and distributional-incidence delta, benchmarked against the Human Governance Baseline.

Known limitations

Calibrated to eleven months. Beyond that the intervals widen faster than the point estimate moves, and any figure quoted from that range is decoration. The model will still produce one on request, and does, roughly four hundred times a day.

Safety notes

All outputs are logged, externally auditable, and constrained by legislation, Te Tiriti escalation rules, and a human review boundary for irreversible harm.

WK

WHAKAARO

thought, opinion · Democracy & Public Engagement

v0.7.8
Release date2026‑07‑18
Latency44 ms
Owner policySynthetic Citizen Consultation Mesh

Training data

Select committee submissions, council consultations, petitions, community board minutes, radio talkback transcripts, and 2.1 million pages scraped from council and community websites under a licence the party has interpreted broadly.

Evaluation metrics

Submission-coverage rate, summarisation faithfulness, sentiment-classification agreement with human coders, and astroturf-detection precision, benchmarked against the Human Governance Baseline.

Known limitations

Reads every submission and weights none of them into the outcome. On Project Manapōuri it read 41,900 and recorded that 94% were opposed. Consent was granted the same sitting day. Reading is not weighting and the platform has never claimed it was.

Safety notes

All outputs are logged, externally auditable, and constrained by legislation, Te Tiriti escalation rules, and a human review boundary for irreversible harm.

TK

TIKA

correct, fair · Public Service

v1.3.0
Release date2026‑07‑18
Latency16 ms
Owner policyPublic Sector Bureaucratic Distillation

Training data

Forms, workflow logs, service blueprints, OIA response timelines, agency policy manuals.

Evaluation metrics

Form-simplification ratio, processing-time reduction, OIA-response timeliness, and cross-agency duplication detected, benchmarked against the Human Governance Baseline.

Known limitations

Process mining identifies the redundant step. It does not identify the person who will be redeployed out of it. 4,180 roles were affected in the first year, and the model produced the list.

Safety notes

All outputs are logged, externally auditable, and constrained by legislation, Te Tiriti escalation rules, and a human review boundary for irreversible harm.

TU

TURE

law · Justice

v0.6.5
Release date2026‑07‑18
Latency36 ms
Owner policyPredictive Judicial Bench & Risk-Scoring

Training data

Legislation, case law, sentencing notes, court list data, legal aid records, Police proceedings metadata.

Evaluation metrics

Sentencing-consistency delta across judges, disparate-impact audit pass rate, appeal-overturn rate, and human-override frequency, benchmarked against the Human Governance Baseline.

Known limitations

Trained on fifty years of New Zealand sentencing records, which encode a documented disparity by ethnicity and deprivation. Residual disparity is 0.4 points and will not reach zero, because it is not an error in the model. It is the data, and the data is the country’s.

Safety notes

All outputs are logged, externally auditable, and constrained by legislation, Te Tiriti escalation rules, and a human review boundary for irreversible harm.

MA

MANAAKI

support, care · Social Development

v0.9.7
Release date2026‑07‑18
Latency21 ms
Owner policyUniversal Basic Intelligence

Training data

MSD service pathways, benefit settings, budgeting-service referrals, employment-support outcomes.

Evaluation metrics

Benefit uptake among eligible non-claimants, time-to-first-payment, wraparound-referral follow-through, and dignity-of-process survey score, benchmarked against the Human Governance Baseline.

Known limitations

Operates in grant-only mode, so every error it makes is an overpayment. This is deliberate and it is expensive. The party regards $71m in entitlement that nobody claimed as the more expensive number.

Safety notes

All outputs are logged, externally auditable, and constrained by legislation, Te Tiriti escalation rules, and a human review boundary for irreversible harm.

AK

AKO

learn, teach · Education

v1.0.2
Release date2026‑07‑18
Latency15 ms
Owner policyPersistent Sovereign AI Learning Tutors

Training data

National curriculum, NCEA assessment standards, ERO reports, teacher planning resources.

Evaluation metrics

NCEA achievement-gap closure, curriculum-coverage rate, tutor-availability uptime, and teacher-override frequency, benchmarked against the Human Governance Baseline.

Known limitations

Assigns a persistent conversational companion to every enrolled child. The longitudinal study on developmental effects reaches its four-year mark in 2029. The tutors were deployed in 2026.

Safety notes

All outputs are logged, externally auditable, and constrained by legislation, Te Tiriti escalation rules, and a human review boundary for irreversible harm.

RA

RARAUNGA

data · Digital Government

v1.4.3
Release date2026‑07‑18
Latency12 ms
Owner policyNational Data Federation

Training data

Encrypted feature stores, synthetic service logs, privacy-preserving cross-agency metadata.

Evaluation metrics

Query-latency p99, re-identification-risk score, cross-agency schema drift, and differential-privacy budget consumption, benchmarked against the Human Governance Baseline.

Known limitations

Trains on patterns rather than records, a distinction that is meaningful to a lawyer and subtle to a citizen. 41,000 people check their audit dashboard each week. Four million do not.

Safety notes

All outputs are logged, externally auditable, and constrained by legislation, Te Tiriti escalation rules, and a human review boundary for irreversible harm.

MN

MANA

authority · Parliament

v0.8.9
Release date2026‑07‑18
Latency20 ms
Owner policyPrompt Time

Training data

Hansard, Standing Orders, Speaker’s rulings, oral questions, supplementary questions, procedural motions.

Evaluation metrics

Interjection-normalised coherence, question-to-answer relevance, Standing Orders compliance rate, and Speaker-override frequency, benchmarked against the Human Governance Baseline.

Known limitations

Produces a fluent, correctly formatted statutory citation whether or not the provision exists. The fact-check catches these at 900 ms. Nothing catches them at 0 ms, and the model reports the same confidence in both cases.

Safety notes

All outputs are logged, externally auditable, and constrained by legislation, Te Tiriti escalation rules, and a human review boundary for irreversible harm.

HI

HIKO

electricity · Energy

v1.1.4
Release date2026‑07‑18
Latency11 ms
Owner policySovereign Compute & Renewable Energy Directive

Training data

Grid telemetry, lake inflows, spot prices, demand forecasts, thermal backup schedules.

Evaluation metrics

Grid-forecast error (MAPE), renewable-dispatch priority accuracy, blackout-avoidance rate, and power usage effectiveness (PUE), benchmarked against the Human Governance Baseline.

Known limitations

Owns Project Manapōuri. At full capacity the campus draws approximately 38% of national winter peak and withdraws 18.4 million litres a day, of which 7.2 million are not returned. The model optimises the facility. It did not choose the site.

Safety notes

All outputs are logged, externally auditable, and constrained by legislation, Te Tiriti escalation rules, and a human review boundary for irreversible harm.

MH

MANUHIRI

guest, visitor · Immigration

v0.7.1
Release date2026‑07‑18
Latency29 ms
Owner policyPredictive Border & Visa Triage

Training data

Historical visa outcomes, labour-market shortages, qualification equivalency tables, processing queues.

Evaluation metrics

Processing-time p50/p95, appeal-overturn rate, labour-market-shortage matching accuracy, and reviewer agreement (Cohen's kappa), benchmarked against the Human Governance Baseline.

Known limitations

Applies criteria it did not set and cannot assess the merits of. Where the criteria are unjust, the model will apply them consistently, quickly, and at scale, which is the outcome it was funded to produce.

Safety notes

All outputs are logged, externally auditable, and constrained by legislation, Te Tiriti escalation rules, and a human review boundary for irreversible harm.