AMC Insights header

Trust but Quantify: On Measuring Verification Intensity in Arms Control and Lessons to Inform the Future of AI Governance

Wilfred Wan

Dylan P.L. Miars, Makenna Riley, Rebecca Bauman & Yuqin Wang

This series presents top outputs from the AMC Arms Control & AI Governance Research Sprint. The event held in April 2026, brought together approximately 60 arms control professionals, early-career researchers, and graduate students, to work intensively over two days with the arms control datasets from AMC Data.

An arms control agreement is only as strong as its ability to verify compliance. Verification mechanisms form the operational backbone of any arms control regime. Yet not all verification tools are created equal. The number of inspection mechanisms, the depth of access granted to inspectors, whether oversight is conducted by the parties themselves or by an independent third party, and whether inspections are automatic or triggered on request all shape whether a treaty can credibly deter cheating.

This raises a broader question: Does the international community systematically design stronger verification systems for more dangerous weapons? We address this question using data from 31 arms control agreements in the Alva Myrdal Centre for Nuclear Disarmament’s Arms Control Agreements Database.

To do so, we construct a Composite Verification Intensity Score (CVIS) based on six variables: mechanism count, access depth, inspector independence, trigger automaticity, NTM provisions, and phase coverage. We then built an interactive tool that allows users to adjust the weight assigned to each variable and observe how treaty rankings shift in real time.

Determining a Measure of Verification Intensity

When determining the factors in our measurement of verification intensity, we decided to divide the score between the number of mechanisms of compliance, National Technical Measures (NTM), access depth, whether it is third party or party-to-party verification, agreement triggered inspection (whether that be automatic or possible to delay or reject), and phase coverage (the number of phases within the weapon’s lifecycle that have verification mechanisms). Our y-axis measures the composite verification intensity score, which ranges from 0 to 100.

To provide users a path forward, the tool opens on our chosen weights to calculate a verification score. Providing an equal weight for all verification measures would overvalue the impact of certain components of the verification regime. Our measurement ensures that the aspects of verification that directly impact weapons verification, such as mechanism count (2.0), depth of access (2.0), automatic agreement triggers (1.5), and phase coverage (1.5), are weighted more heavily than variables that act as institutional architecture. This architecture, while important, does not directly verify arms control and therefore needs to be measured for normative impact in another tool.

Understanding the Complexities Behind Risk-Proportionate Governance

Whether verification measures are weighted equally or proportionally to their depth and directness of impact, risk-proportionate governance is neither purely assumption nor purely reality. It is not one or the other; instead, it is a fuzzy dynamic dependent on factors beyond the level of risk. Our model demonstrates that there is a marked difference in verification intensity scores between high-risk and low-risk treaties (13.97 point difference between average CVISs for high and low risk treaties); however, there were numerous outlier cases that make definitive determination messy. Treaties such as the Biological Weapons Convention (18.82), Antarctic Treaty (17.35), and the Seabed Treaty (13.82) score low for verification, either through a lack of a negotiated verification protocol or a focus on preventative protocols instead of verification protocols. On the other hand, Open Skies outperforms the high-risk mean at 36.47 due to its substantive verification methods in the form of automatic triggers and access depth. What these outliers share is that their verification intensity, or lack thereof, is better explained by the political conditions surrounding their negotiation than by the objective danger of the weapons they govern.

The arms control treaties with the highest CVISs (START I - 58.69, New START - 56.84, INF - 53.74, JCPOA - 48.82) were able to institute the most thorough verification and governance means through political symmetry, not purely the danger of the weapon they sought to manage. Many of the bilateral agreements between the United States and the USSR scored high because both sides had equal incentive and capability to demand rigorous verification from each other. Multilateral contexts produce weaker verification even for equally dangerous weapons.

While some may take this to mean that multilateral agreements are simply inferior to bilateral agreements, we should instead understand this as highlighting that risk-proportionate governance and strong verification regimes can exist under the correct political conditions. While we can’t expect these arms control agreements or the political contexts that create them to persist (e.g., the US withdrawal from the JCPOA and the failed renewal of New START), this does not mean that this is a foregone conclusion. Our model identifies that risk-proportionate governance is most effective when it prioritizes access depth and agreement triggers, the variables that determine not just whether verification exists, but whether it can function under the political conditions most likely to produce cheating.

Contextualizing AI Governance with Our Findings

The partial existence of a risk–verification gradient in arms control (no more assumption nor more reality) supports a case for capability-based AI regulation, but most feasibly at a lower level of verification intensity, such as the quantity of compliance mechanisms present rather than the intrusiveness of those inspection measures. In arms control, oversight often does scale with risk; therefore, lower-risk weapons rely mostly on mechanisms such as agreement triggers and monitoring requirements that signal the existence of governance without requiring the level of transparency desirable for high-risk weapons. These types of measures are transferable to AI, where frontier models could be subject to reporting and trigger-based compliance mechanisms. However, more intensive forms of verification, such as deep access to systems, training data, or full oversight across development and deployment phases, are significantly harder to implement (and justify) in the AI context. As a result, while a risk-verification gradient exists in arms control, increased oversight for frontier models will likely remain selective and less intrusive. However, as seen with Open Skies, strong political conditions surrounding the negotiation of treaties can influence their verification intensity. Strong pushes for verification could see more risk-proportionate AI governance.

FOLLOW UPPSALA UNIVERSITY ON

Uppsala University on Facebook
Uppsala University on Instagram
Uppsala University on Youtube
Uppsala University on Linkedin