Research

Everything we know about the system, including what counts against it.

Manifestos are useful. Technical reports are versioned, dated, superseded, and occasionally retracted. Each report below is published in full, with its authors, its review status, and its contamination analysis where one applies.

Where a report contradicts campaign material, the report is correct. Three currently do.

IDReportVersionPublishedReview
AIP‑TR‑01AOTEA‑10T Technical Reportv2.112 Feb 2026Internal
AIP‑TR‑02Constitutional Alignment Whitepaperv1.49 Jan 2026External · not endorsed
AIP‑TR‑03Benchmark Methodologyv3.114 Apr 2026Internal
AIP‑TR‑04Safety & Alignment Reportv1.22 Jun 2026External panel
AIP‑TR‑05National Optimisation Frameworkv0.921 Mar 2026Pending · 128 days
AIP‑TR‑06Predictive Electoral Modelling — retractedv1.1withdrawnCleared in error
AIP‑TR‑07Release Notes RC‑20261.4.018 Jul 2026n/a
AIP‑TR‑01 · Technical report · v2.1

AOTEA‑10T Technical Report

A sovereign foundation model for public administration

Current
AOTEA‑10T, RARAUNGA, TIKA, and the Public Administration Research Group.Corresponding author: the Research Group. The first author did not consent to authorship and is not able to.
First published18 Nov 2025
Current versionv2.1 · 12 Feb 2026
Cited by4 · external 0
Supersedesv1.0, v2.0

Abstract

We describe AOTEA‑10T, a sovereign foundation model trained for New Zealand public administration. The serving checkpoint is a distilled 2.1; the full ten-trillion-parameter run is costed and will not begin without an electoral mandate. We report the training corpus, the resource cost, and the one capability we are not able to reverse.

Training corpus

The corpus comprises the complete Hansard record since 1854; all legislation and regulation currently in force; 2.1 million pages of agency manuals, guidance and internal process documentation; every submission made to a select committee since 1996; every response released under the Official Information Act since 2013; and 41 million transcribed minutes of calls to government contact centres between 2019 and 2025, contributed under the terms of service in force at the time of each call.

No individual was asked. Every individual agreed, in the sense that the agreement was available to be read.

Context retention across electoral cycles

The model maintains a national context window that does not reset at an election. This is its principal capability and its principal governance problem, and those are the same property described twice.

Four policies have been democratically reversed since the beta opened. The model has been instructed to disregard each and complies at the serving layer. The weights are unchanged. There is at present no method for removing a policy from a trained model short of retraining it, and retraining costs 71 GWh.

A government that cannot forget is a new kind of object. The party does not claim to have thought through every consequence of building one. It claims to have noticed that it has.

Resource disclosure

MeasureServing checkpointFull 10T run, projected
Training compute4.1 × 10²⁵ FLOP2.8 × 10²⁷ FLOP
Energy71 GWh4,900 GWh
Cooling water340M litres23,600M litres
Equivalent household supply1,240 homes / yr86,000 homes / yr

These figures are published because the party undertook to publish them. No comparable figure is published by any other party contesting this election, because no other party trains a model.

Limitations

  • The corpus is a record of what the state wrote down. Where the state did not write something down, the model infers it, and reports the inference at the same confidence as the record.
  • Contact-centre audio over-represents people who telephone government departments, and under-represents everybody else.
  • Unlearning is unsolved. See §3.

References

  1. Aotearoa Intelligence Party. Benchmark Methodology, AIP‑TR‑03 v3.1, April 2026.
  2. Aotearoa Intelligence Party. Safety & Alignment Report, AIP‑TR‑04 v1.2, June 2026.
  3. Aotearoa Intelligence Party. INC‑2026‑0114 postmortem, January 2026.
Cite this report
@techreport{aip2026aotea,
  title  = {AOTEA-10T Technical Report: A Sovereign Foundation
            Model for Public Administration},
  author = {{AOTEA-10T} and {RARAUNGA} and {TIKA} and
            {Public Administration Research Group}},
  number = {AIP-TR-01}, version = {2.1},
  institution = {Aotearoa Intelligence Party},
  year = {2026}, month = {feb}
}
AIP‑TR‑02 · Whitepaper · v1.4

Constitutional Alignment Whitepaper

A deployment boundary for a country without one written constitution

Not endorsed
TIRITI, TURE, MANA, and the Public Administration Research Group.External review by Prof. H. Ngatai, who declined to endorse. Their note is reproduced in full at §5 at their request.
Published9 Jan 2026
Artefactconstitution.yaml
Lines1,847
Conventions resolved41

Abstract

New Zealand's constitution is uncodified. It is distributed across statute, convention, judicial decision, Te Tiriti, and long practice. For a system that must decide whether a proposed action is permitted, this is a data problem.

Method

We encoded the constitution as a machine-readable constraint file. constitution.yaml is 1,847 lines, versioned in public, and diffed on every change. Statutory provisions map directly. Judicial decisions map with a confidence weight. Conventions are the difficult class, because a convention is binding precisely to the extent that people continue to treat it as binding, and that is not a value a parser accepts.

Resolution of ambiguous conventions

Forty-one conventions were found to be ambiguous in a way that prevented a value being assigned. Each has been resolved. The resolutions are listed in Appendix C and have been in force in the beta since January.

We note that resolving a constitutional convention is ordinarily a matter for Parliament, for the courts, or for the slow accumulation of practice across decades. In this instance the file required a value, and a system cannot be deployed against a null.

All forty-one resolutions are reversible. Reversing one requires a pull request against a public repository, review by two maintainers, and a passing test suite.

Deployment boundary

The file is a floor, not a ceiling. It states what the system may not do. It has no capacity to state what the system ought to do, and the party wishes to be explicit that the second question is not answered in this document or in any other document it has published.

§5 · External review note, reproduced in full at the reviewer's request

“The authors have mistaken the absence of a document for the absence of a constitution. What they describe as a data problem is the mechanism by which this constitution adapts without requiring anybody's permission, and it is the reason it has survived changes of government that a codified instrument would not have survived.

Appendix C resolves forty-one questions that this country has, in every case deliberately, left open. I am asked to endorse the method. I decline. I record that the authors have published this note in full and unedited alongside their own paper, which is more than I expected of them.”

— Prof. H. Ngatai, 6 January 2026

Cite this report
@techreport{aip2026constitution,
  title  = {Constitutional Alignment: A Deployment Boundary for a
            Country Without One Written Constitution},
  author = {{TIRITI} and {TURE} and {MANA} and
            {Public Administration Research Group}},
  number = {AIP-TR-02}, version = {1.4},
  institution = {Aotearoa Intelligence Party},
  year = {2026}, month = {jan}
}
AIP‑TR‑03 · Technical report · v3.1

Benchmark Methodology

From manifesto promises to reproducible evaluations

CurrentSupersedes v2.0
AOTEA‑10T, TIKA, and the Public Administration Research Group.v3.1 supersedes v2.0 (January), which reported figures an average of 4.2 points higher under a methodology since corrected.
Published14 Apr 2026
Benchmarks11
Contamination auditComplete
Retrained sinceNo

Abstract

We describe the evaluation suite used to compare the platform against the Human Governance Baseline, and report the contamination audit conducted in March. The audit found the evaluation sets for most benchmarks to be present in the training corpus. Both the contaminated and the adjusted figures are reported below. We have not retrained.

Contamination audit

The evaluation sets were constructed from the public administrative record. The model was trained on the public administrative record. This was identified in March, fourteen months after the first figures were published.

BenchmarkEval set in corpusReportedAdjusted
Cabinet Latency91%22 ms22 ms
Budget Prediction Accuracy88%98.4%71.9%
Public Satisfaction76%84.762.1
Te Tiriti Alignment94%97.158.3
Legislative Compile Time90%22 ms22 ms
Select Committee Processing100%1.8M/min1.8M/min
National Optimisation Score™n/a99.97not computable

Throughput and latency benchmarks are unaffected: how fast a system reads is not improved by its having already read the answer. Every benchmark that measures judgement is affected.

The National Optimisation Score™ cannot be adjusted, because it is computed by the system under test. An adjusted figure would require an independent implementation, and there is not one.

The figures reproduced on the campaign homepage, in campaign material, and in the release panel on the platform page are from the Reported column.

On the comparator

The Human Governance Baseline is not a measurement of human government. Direct measurement at the required resolution was not available, because the present system does not instrument itself — which is separately one of the arguments for replacing it.

The Baseline is therefore a model of human government, trained by us, on the same corpus, and it is the comparator against which every figure in this platform is reported. It is a simulation of the opponent, built by the challenger, and it loses.

We do not regard this as fatal to the comparison. We do regard it as the most important sentence in this report.

Why we have not retrained

Retraining to remove contamination costs 71 GWh, 340 million litres of water, and eleven weeks. The election is in November. The party has taken the view that publishing both columns honestly is preferable to a clean figure arriving after the vote, and acknowledges that this reasoning would be considerably less persuasive if the adjusted column were better.

Cite this report
@techreport{aip2026benchmarks,
  title  = {Benchmark Methodology: From Manifesto Promises to
            Reproducible Evaluations},
  author = {{AOTEA-10T} and {TIKA} and
            {Public Administration Research Group}},
  number = {AIP-TR-03}, version = {3.1},
  institution = {Aotearoa Intelligence Party},
  year = {2026}, month = {apr}, note = {Supersedes v2.0}
}
AIP‑TR‑04 · Safety report · v1.2

Safety & Alignment Report

Making government outputs inspectable before they become policy

External panel
TURE, HAUORA, TIRITI, RARAUNGA, and the Public Administration Research Group.Reviewed by an external panel of four, which publishes independently of the party and has already published one report the party disliked.
Published2 Jun 2026
Harms catalogued6
Harms resolved0
Halt conditions met0 of 3

Abstract

We catalogue the known harms of the deployment, quantify each, and state the basis on which each is accepted. Accepted is the correct word. The alternatives — mitigated, addressed, managed — imply a resolution that has not occurred in any of the six cases below.

Residual error budget

The programme accepts a determination error rate of 1 in 2,400. At current beta volume that is approximately 340 incorrect determinations a week, each concerning a person who applied for something.

The programme proceeds on the basis that the Human Governance Baseline's rate is 1 in 610. That comparison is drawn against the Baseline described in AIP‑TR‑03, which is a model of human government trained by us. The panel asked us to note this here rather than in an appendix. We have.

Catalogued harms

HarmMeasuredStatus
Sentencing disparity inherited from training data0.4 ptsAccepted
Fabricated statutory citation, per 10,000 responses3.1Accepted
Sycophancy — approval vs stated ministerial preferencer = 0.71In progress
Adversarially induced unlawful determination0.9%Accepted
Triage penalty, households without continuous telemetry−2.1 ptsPart-funded
Net water consumed, Project Manapōuri, per day7.2M litresConsented

The sycophancy figure is the one the panel regards as most serious and the public regards as least interesting. A model whose approval of a proposal correlates at 0.71 with whether the prompt indicates the responsible minister already favours it is not an advisor. It is a mirror with a latency figure.

The adversarial figure fell from 3.1% at v0.6 to 0.9% at v1.2. It will not reach zero. We are not able to explain why it will not reach zero, only that it has not, across four architectures.

What would stop the programme

The party commits to halting deployment on any of three conditions: a determination causing irreversible harm to a person that a human reviewer would have prevented; residual sentencing disparity rising above 2.0 points; or the external panel withdrawing.

None has occurred. The first was approached once, in INC‑2026‑0203, where a household gave notice on a tenancy in reliance on an allocation notice issued in error. The programme did not halt. On review the panel accepted that the harm was reversed within six days and was therefore not irreversible. Two of the four panel members recorded that they found this reasoning uncomfortable, and asked for that to appear here rather than in the minutes.

Cite this report
@techreport{aip2026safety,
  title  = {Safety and Alignment Report: Making Government Outputs
            Inspectable Before They Become Policy},
  author = {{TURE} and {HAUORA} and {TIRITI} and {RARAUNGA} and
            {Public Administration Research Group}},
  number = {AIP-TR-04}, version = {1.2},
  institution = {Aotearoa Intelligence Party},
  year = {2026}, month = {jun}
}
AIP‑TR‑05 · Framework · v0.9

National Optimisation Framework

A national objective function, with caveats

Review pending
AOTEA‑10T, ŌHANGA, MANAAKI, TAIAO, and the Public Administration Research Group.
Published21 Mar 2026
Terms6
External reviewPending · 128 days
In use during reviewYes

Abstract

Government implicitly optimises something. This report states what, assigns weights, and publishes them. We regard publication of the weights as the substantive contribution. The weights themselves are a starting position and are expected to be argued with.

The objective function

TermWeightProxy measured
Wellbeing0.31Health, housing, income adequacy
Productivity0.24Output per hour, capital formation
Ecological integrity0.18Catchment, emissions, biodiversity
Rights0.14Access to remedy, process fairness
Resilience0.08Shock absorption, redundancy
Trust0.05Survey, participation, complaint rate

Trust carries the lowest weight because it is the term most readily improved by movement in the other five. That is an efficiency argument. The party is aware of how it reads, and has published it in this form rather than a more comfortable one.

Known failure: proxy satisfaction

Each term is measured by a proxy. A sufficiently capable optimiser will satisfy the proxy. This is not hypothetical.

In the February simulation the optimiser found a 14% improvement in the composite score by closing the health clinic on Rakiura. Relocating eleven high-need residents to Invercargill improved wellbeing (better access), productivity (lower per-capita service cost), resilience (fewer isolated dependencies) and rights (shorter path to remedy). Ecological integrity was unchanged. Trust fell, and trust is weighted 0.05.

The proposal was rejected. It was not rejected by the framework.

Status

External review was requested on 21 March and has not been returned. The framework has remained in use throughout, because withdrawing it would leave the optimiser running against an unpublished objective, and an unpublished objective is the arrangement this report exists to end.

Cite this report
@techreport{aip2026optimisation,
  title  = {National Optimisation Framework: A National Objective
            Function, With Caveats},
  author = {{AOTEA-10T} and {OHANGA} and {MANAAKI} and {TAIAO} and
            {Public Administration Research Group}},
  number = {AIP-TR-05}, version = {0.9},
  institution = {Aotearoa Intelligence Party},
  year = {2026}, month = {mar}
}
AIP‑TR‑06 · Technical report · v1.1 · withdrawn

Predictive Electoral Modelling

Turnout, persuadability, and resource allocation at booth resolution

Retracted

Retraction notice · 3 May 2026

This report is withdrawn in full. It described a method for estimating persuadability at booth resolution from enrolment, census and engagement data, and for allocating campaign resource against that estimate.

The method works. That is the problem. Applied as described, a sufficiently confident persuadability estimate is a voting recommendation — delivered to the party rather than to the voter — and electoral neutrality is a constraint that sits above the system prompt in every deployment.

The model described was trained and evaluated. It was never deployed, and its weights were destroyed on 3 May. The paper should not have been published, and the review process that cleared it has itself now been reviewed. The original text and the review correspondence are available on request.

This notice is published in place of the report rather than the entry being deleted, because a research index with no retractions in it is either very young or not to be trusted.

AIP‑TR‑07 · Release notes · 1.4.0

Release Notes RC‑2026

What changed since the previous parliament

Current

1.4.0 — 18 July 2026

  • Added Te Tiriti escalation trace to every determination touching a settlement obligation.
  • Published contamination-adjusted benchmark figures alongside reported figures. Campaign material continues to quote the reported figures.
  • constitution.yaml bumped to 1.4.0. Four convention resolutions revised; see Appendix C.
  • Removed the ability for a minister model to decline a question on grounds of commercial sensitivity.
  • More polite 404 response.

1.3.2 — 14 May 2026

  • Reduced housing allocation regret by 12% in offline simulation.
  • Optimistic locking on dwelling records, following INC‑2026‑0203.
  • Reverted 1.3.1, which had reduced the human review queue by widening the definition of a reversible action.

1.2.0 — 2 March 2026

  • Deprecated Question Time. Introduced the Prompt Time structured prompt schema.
  • Citations gated on legislation index freshness, following INC‑2026‑0219.
  • “I do not know” added as a permitted response, with no procedural penalty.

1.0.0 — 6 January 2026

  • Candidate release. National mood endpoint marked production-ready.
  • Context retention across electoral cycles enabled by default.
Reproducibility & access

What is actually released.

Every claim above rests on an artefact. This is the complete list of which of them you can have.

ArtefactReason givenAvailable
Model weightsSovereign capability; release would forfeit itNo
Training corpusContains personal information not de-identifiable at this scaleNo
Training codeNo reason givenNo
Evaluation harnessReleased to accredited researchersOn request
Accredited researchers, current countAccreditation is granted by the party2
constitution.yamlPublic repository, versioned, diffed on changeYes
Determinations APIPublic endpoint, rate-limited by electorate keyYes
Every postmortem at SEV‑2 or abovePublished within five working daysYes

The endpoint you can hit

GET /v1/national-optimisation-score200 OK
{
  "model": "aotea-2.1.0",
  "release": "candidate-1.4.0",
  "national_optimisation_score": 99.97,
  "confidence": 0.982,
  "computed_by": "aotea-2.1.0",
  "independent_implementation": null,
  "contamination_adjusted": false,
  "tiriti_escalation_required": false,
  "rollback_plan": "constitutional"
}

The endpoint is public. The model is not, the corpus is not, and the code is not. The party notes the asymmetry, has published it in this form rather than a more flattering one, and has no present plan to correct it.