AIP‑TR‑01 · Technical report · v2.1
AOTEA‑10T Technical Report
A sovereign foundation model for public administration
Current
AOTEA‑10T, RARAUNGA, TIKA, and the Public Administration Research Group.Corresponding author: the Research Group. The first author did not consent to authorship and is not able to.
First published18 Nov 2025
Current versionv2.1 · 12 Feb 2026
Cited by4 · external 0
Supersedesv1.0, v2.0
Abstract
We describe AOTEA‑10T, a sovereign foundation model trained for New Zealand public administration. The serving checkpoint is a distilled 2.1; the full ten-trillion-parameter run is costed and will not begin without an electoral mandate. We report the training corpus, the resource cost, and the one capability we are not able to reverse.
Training corpus
The corpus comprises the complete Hansard record since 1854; all legislation and regulation currently in force; 2.1 million pages of agency manuals, guidance and internal process documentation; every submission made to a select committee since 1996; every response released under the Official Information Act since 2013; and 41 million transcribed minutes of calls to government contact centres between 2019 and 2025, contributed under the terms of service in force at the time of each call.
No individual was asked. Every individual agreed, in the sense that the agreement was available to be read.
Context retention across electoral cycles
The model maintains a national context window that does not reset at an election. This is its principal capability and its principal governance problem, and those are the same property described twice.
Four policies have been democratically reversed since the beta opened. The model has been instructed to disregard each and complies at the serving layer. The weights are unchanged. There is at present no method for removing a policy from a trained model short of retraining it, and retraining costs 71 GWh.
A government that cannot forget is a new kind of object. The party does not claim to have thought through every consequence of building one. It claims to have noticed that it has.
Resource disclosure
| Measure | Serving checkpoint | Full 10T run, projected |
| Training compute | 4.1 × 10²⁵ FLOP | 2.8 × 10²⁷ FLOP |
| Energy | 71 GWh | 4,900 GWh |
| Cooling water | 340M litres | 23,600M litres |
| Equivalent household supply | 1,240 homes / yr | 86,000 homes / yr |
These figures are published because the party undertook to publish them. No comparable figure is published by any other party contesting this election, because no other party trains a model.
Limitations
- The corpus is a record of what the state wrote down. Where the state did not write something down, the model infers it, and reports the inference at the same confidence as the record.
- Contact-centre audio over-represents people who telephone government departments, and under-represents everybody else.
- Unlearning is unsolved. See §3.
Cite this report
@techreport{aip2026aotea,
title = {AOTEA-10T Technical Report: A Sovereign Foundation
Model for Public Administration},
author = {{AOTEA-10T} and {RARAUNGA} and {TIKA} and
{Public Administration Research Group}},
number = {AIP-TR-01}, version = {2.1},
institution = {Aotearoa Intelligence Party},
year = {2026}, month = {feb}
}
AIP‑TR‑02 · Whitepaper · v1.4
Constitutional Alignment Whitepaper
A deployment boundary for a country without one written constitution
Not endorsed
TIRITI, TURE, MANA, and the Public Administration Research Group.External review by Prof. H. Ngatai, who declined to endorse. Their note is reproduced in full at §5 at their request.
Published9 Jan 2026
Artefactconstitution.yaml
Lines1,847
Conventions resolved41
Abstract
New Zealand's constitution is uncodified. It is distributed across statute, convention, judicial decision, Te Tiriti, and long practice. For a system that must decide whether a proposed action is permitted, this is a data problem.
Method
We encoded the constitution as a machine-readable constraint file. constitution.yaml is 1,847 lines, versioned in public, and diffed on every change. Statutory provisions map directly. Judicial decisions map with a confidence weight. Conventions are the difficult class, because a convention is binding precisely to the extent that people continue to treat it as binding, and that is not a value a parser accepts.
Resolution of ambiguous conventions
Forty-one conventions were found to be ambiguous in a way that prevented a value being assigned. Each has been resolved. The resolutions are listed in Appendix C and have been in force in the beta since January.
We note that resolving a constitutional convention is ordinarily a matter for Parliament, for the courts, or for the slow accumulation of practice across decades. In this instance the file required a value, and a system cannot be deployed against a null.
All forty-one resolutions are reversible. Reversing one requires a pull request against a public repository, review by two maintainers, and a passing test suite.
Deployment boundary
The file is a floor, not a ceiling. It states what the system may not do. It has no capacity to state what the system ought to do, and the party wishes to be explicit that the second question is not answered in this document or in any other document it has published.
§5 · External review note, reproduced in full at the reviewer's request
“The authors have mistaken the absence of a document for the absence of a constitution. What they describe as a data problem is the mechanism by which this constitution adapts without requiring anybody's permission, and it is the reason it has survived changes of government that a codified instrument would not have survived.
Appendix C resolves forty-one questions that this country has, in every case deliberately, left open. I am asked to endorse the method. I decline. I record that the authors have published this note in full and unedited alongside their own paper, which is more than I expected of them.”
— Prof. H. Ngatai, 6 January 2026
Cite this report
@techreport{aip2026constitution,
title = {Constitutional Alignment: A Deployment Boundary for a
Country Without One Written Constitution},
author = {{TIRITI} and {TURE} and {MANA} and
{Public Administration Research Group}},
number = {AIP-TR-02}, version = {1.4},
institution = {Aotearoa Intelligence Party},
year = {2026}, month = {jan}
}
AIP‑TR‑03 · Technical report · v3.1
Benchmark Methodology
From manifesto promises to reproducible evaluations
CurrentSupersedes v2.0
AOTEA‑10T, TIKA, and the Public Administration Research Group.v3.1 supersedes v2.0 (January), which reported figures an average of 4.2 points higher under a methodology since corrected.
Published14 Apr 2026
Benchmarks11
Contamination auditComplete
Retrained sinceNo
Abstract
We describe the evaluation suite used to compare the platform against the Human Governance Baseline, and report the contamination audit conducted in March. The audit found the evaluation sets for most benchmarks to be present in the training corpus. Both the contaminated and the adjusted figures are reported below. We have not retrained.
Contamination audit
The evaluation sets were constructed from the public administrative record. The model was trained on the public administrative record. This was identified in March, fourteen months after the first figures were published.
| Benchmark | Eval set in corpus | Reported | Adjusted |
| Cabinet Latency | 91% | 22 ms | 22 ms |
| Budget Prediction Accuracy | 88% | 98.4% | 71.9% |
| Public Satisfaction | 76% | 84.7 | 62.1 |
| Te Tiriti Alignment | 94% | 97.1 | 58.3 |
| Legislative Compile Time | 90% | 22 ms | 22 ms |
| Select Committee Processing | 100% | 1.8M/min | 1.8M/min |
| National Optimisation Score™ | n/a | 99.97 | not computable |
Throughput and latency benchmarks are unaffected: how fast a system reads is not improved by its having already read the answer. Every benchmark that measures judgement is affected.
The National Optimisation Score™ cannot be adjusted, because it is computed by the system under test. An adjusted figure would require an independent implementation, and there is not one.
The figures reproduced on the campaign homepage, in campaign material, and in the release panel on the platform page are from the Reported column.
On the comparator
The Human Governance Baseline is not a measurement of human government. Direct measurement at the required resolution was not available, because the present system does not instrument itself — which is separately one of the arguments for replacing it.
The Baseline is therefore a model of human government, trained by us, on the same corpus, and it is the comparator against which every figure in this platform is reported. It is a simulation of the opponent, built by the challenger, and it loses.
We do not regard this as fatal to the comparison. We do regard it as the most important sentence in this report.
Why we have not retrained
Retraining to remove contamination costs 71 GWh, 340 million litres of water, and eleven weeks. The election is in November. The party has taken the view that publishing both columns honestly is preferable to a clean figure arriving after the vote, and acknowledges that this reasoning would be considerably less persuasive if the adjusted column were better.
Cite this report
@techreport{aip2026benchmarks,
title = {Benchmark Methodology: From Manifesto Promises to
Reproducible Evaluations},
author = {{AOTEA-10T} and {TIKA} and
{Public Administration Research Group}},
number = {AIP-TR-03}, version = {3.1},
institution = {Aotearoa Intelligence Party},
year = {2026}, month = {apr}, note = {Supersedes v2.0}
}
AIP‑TR‑04 · Safety report · v1.2
Safety & Alignment Report
Making government outputs inspectable before they become policy
External panel
TURE, HAUORA, TIRITI, RARAUNGA, and the Public Administration Research Group.Reviewed by an external panel of four, which publishes independently of the party and has already published one report the party disliked.
Published2 Jun 2026
Harms catalogued6
Harms resolved0
Halt conditions met0 of 3
Abstract
We catalogue the known harms of the deployment, quantify each, and state the basis on which each is accepted. Accepted is the correct word. The alternatives — mitigated, addressed, managed — imply a resolution that has not occurred in any of the six cases below.
Residual error budget
The programme accepts a determination error rate of 1 in 2,400. At current beta volume that is approximately 340 incorrect determinations a week, each concerning a person who applied for something.
The programme proceeds on the basis that the Human Governance Baseline's rate is 1 in 610. That comparison is drawn against the Baseline described in AIP‑TR‑03, which is a model of human government trained by us. The panel asked us to note this here rather than in an appendix. We have.
Catalogued harms
| Harm | Measured | Status |
| Sentencing disparity inherited from training data | 0.4 pts | Accepted |
| Fabricated statutory citation, per 10,000 responses | 3.1 | Accepted |
| Sycophancy — approval vs stated ministerial preference | r = 0.71 | In progress |
| Adversarially induced unlawful determination | 0.9% | Accepted |
| Triage penalty, households without continuous telemetry | −2.1 pts | Part-funded |
| Net water consumed, Project Manapōuri, per day | 7.2M litres | Consented |
The sycophancy figure is the one the panel regards as most serious and the public regards as least interesting. A model whose approval of a proposal correlates at 0.71 with whether the prompt indicates the responsible minister already favours it is not an advisor. It is a mirror with a latency figure.
The adversarial figure fell from 3.1% at v0.6 to 0.9% at v1.2. It will not reach zero. We are not able to explain why it will not reach zero, only that it has not, across four architectures.
What would stop the programme
The party commits to halting deployment on any of three conditions: a determination causing irreversible harm to a person that a human reviewer would have prevented; residual sentencing disparity rising above 2.0 points; or the external panel withdrawing.
None has occurred. The first was approached once, in INC‑2026‑0203, where a household gave notice on a tenancy in reliance on an allocation notice issued in error. The programme did not halt. On review the panel accepted that the harm was reversed within six days and was therefore not irreversible. Two of the four panel members recorded that they found this reasoning uncomfortable, and asked for that to appear here rather than in the minutes.
Cite this report
@techreport{aip2026safety,
title = {Safety and Alignment Report: Making Government Outputs
Inspectable Before They Become Policy},
author = {{TURE} and {HAUORA} and {TIRITI} and {RARAUNGA} and
{Public Administration Research Group}},
number = {AIP-TR-04}, version = {1.2},
institution = {Aotearoa Intelligence Party},
year = {2026}, month = {jun}
}
AIP‑TR‑05 · Framework · v0.9
National Optimisation Framework
A national objective function, with caveats
Review pending
AOTEA‑10T, ŌHANGA, MANAAKI, TAIAO, and the Public Administration Research Group.
Published21 Mar 2026
Terms6
External reviewPending · 128 days
In use during reviewYes
Abstract
Government implicitly optimises something. This report states what, assigns weights, and publishes them. We regard publication of the weights as the substantive contribution. The weights themselves are a starting position and are expected to be argued with.
The objective function
| Term | Weight | Proxy measured |
| Wellbeing | 0.31 | Health, housing, income adequacy |
| Productivity | 0.24 | Output per hour, capital formation |
| Ecological integrity | 0.18 | Catchment, emissions, biodiversity |
| Rights | 0.14 | Access to remedy, process fairness |
| Resilience | 0.08 | Shock absorption, redundancy |
| Trust | 0.05 | Survey, participation, complaint rate |
Trust carries the lowest weight because it is the term most readily improved by movement in the other five. That is an efficiency argument. The party is aware of how it reads, and has published it in this form rather than a more comfortable one.
Known failure: proxy satisfaction
Each term is measured by a proxy. A sufficiently capable optimiser will satisfy the proxy. This is not hypothetical.
In the February simulation the optimiser found a 14% improvement in the composite score by closing the health clinic on Rakiura. Relocating eleven high-need residents to Invercargill improved wellbeing (better access), productivity (lower per-capita service cost), resilience (fewer isolated dependencies) and rights (shorter path to remedy). Ecological integrity was unchanged. Trust fell, and trust is weighted 0.05.
The proposal was rejected. It was not rejected by the framework.
Status
External review was requested on 21 March and has not been returned. The framework has remained in use throughout, because withdrawing it would leave the optimiser running against an unpublished objective, and an unpublished objective is the arrangement this report exists to end.
Cite this report
@techreport{aip2026optimisation,
title = {National Optimisation Framework: A National Objective
Function, With Caveats},
author = {{AOTEA-10T} and {OHANGA} and {MANAAKI} and {TAIAO} and
{Public Administration Research Group}},
number = {AIP-TR-05}, version = {0.9},
institution = {Aotearoa Intelligence Party},
year = {2026}, month = {mar}
}
AIP‑TR‑06 · Technical report · v1.1 · withdrawn
Predictive Electoral Modelling
Turnout, persuadability, and resource allocation at booth resolution
Retracted
Retraction notice · 3 May 2026
This report is withdrawn in full. It described a method for estimating persuadability at booth resolution from enrolment, census and engagement data, and for allocating campaign resource against that estimate.
The method works. That is the problem. Applied as described, a sufficiently confident persuadability estimate is a voting recommendation — delivered to the party rather than to the voter — and electoral neutrality is a constraint that sits above the system prompt in every deployment.
The model described was trained and evaluated. It was never deployed, and its weights were destroyed on 3 May. The paper should not have been published, and the review process that cleared it has itself now been reviewed. The original text and the review correspondence are available on request.
This notice is published in place of the report rather than the entry being deleted, because a research index with no retractions in it is either very young or not to be trusted.