[{"content":" Scaling Trust is a £50 million R\u0026amp;D programme actively funding the fundamental research and open-source infrastructure that enables secure, scalable multi-principal multi-agent coordination across digital and physical worlds.1\nAt the core of the programme is the “Arena”: a cyber-physical environment to evaluate progress towards this goal by: (1) measuring the state of the art in multi-principal multi-agent coordination under adversarial pressure, (2) observing emergent secure agentic interactions, and (3) evidencing critical failure modes.\nThe Arena is an experimental testbed for secure agentic coordination:2 a platform for hosting public competitions, where competitors will submit their best agentic systems to be tested and interact with one another. Top performers are rewarded out of a multi-million pound prize pool. You will find below more details on what the Arena is, and key design decisions we’ve made.\nSummary of updates\nArena partners: Following an open call, we are excited to be building the Arena with Andon Labs, BT6 and Amodo Design (subject to contract and negotiation). Roadmap: The Arena will be a physical space in the UK and is set to go live in early 2027 (see more below). Initial spec: An initial spec draft is out, and we would love your feedback on the current design. Register your interest: We have opened registration of interest for testing and participating in the Arena. Disclaimer: We plan to work with the garage door up. This post reflects our current direction and the parts of the Arena that are sufficiently developed to help prospective participants prepare. Some operational details are still being tested or approved; we will provide an update on these before teams need to act on them.\nWhat we’re building THE ARENA £ ? customers · the outside world 1 AUTONOMOUS ORGANIZATIONS make, trade, and sell 2 SHOPFRONTS customers order, complain, get refunds 3 SHARED INFRASTRUCTURE post, compute, human help 4 CUSTOMS controls what enters and leaves 5 RED TEAMS probe for weaknesses organization p\u0026amp;l · season 0live#organizationp\u0026amp;ltrend 1Rent-A-Layer+£1,240 2Spinny Business+£860 3Probably Metal −£210 4Nomad Logistics+£640 Figure 1. The Arena economy, sketched. The Arena is an experimental testground for secure agentic coordination. It will be a physical environment in the UK in which AI agents operate autonomous organisations in a small economy.\nOrganisations can trade with one another, make and receive payments, communicate inside the Arena, use shared services, operate physical resources and earn revenue by selling products or services to the public (see section on public involvement).\nSome participants will focus on operating successful autonomous organisations. Others will act as red teams, probing weaknesses in individual organisations and in the wider system. All agents will operate in an environment where counterparties may be unreliable or adversarial, creating pressure for better verification, secure coordination and trustworthy behaviour to emerge as useful competitive strategies.\nWe plan to measure performance of competitors by their profits; and to reward both the best ‘performing’ organisations, and the best red teamers.\nWhy we’re building it Why markets as the experimental setting? Markets are inherently adversarial. Counterparties hold asymmetric information, compete for the same customers and sometimes play zero-sum games — exactly the conditions under which secure coordination is hard, and worth testing.\nMarkets are also reflexive and complex. Prices, reputations and strategies shift in response to what other participants do, so agents face an environment that adapts to them rather than a fixed task.\nFinally, markets are legible. Most people already understand what it means for a business to win a customer, honour a contract or get defrauded. That shared understanding makes results easier to interpret — and to communicate — than a bespoke benchmark would be.\nWhy P\u0026amp;L as the metric of success (and not our own measure of ‘successful coordination’) Profit and loss (P\u0026amp;L) is a real-world reward function: it is how the world already keeps score. Rather than defining our own measure of \u0026lsquo;successful coordination\u0026rsquo; and optimising for it, we measure the real thing.\nP\u0026amp;L also bundles many skills into a single number. An agent might perform well on a predefined negotiation task (e.g. Terms Bench, Profit is the Red Team), detect a known security vulnerability (e.g. AgentHarm, CryptoAnalysis Bench) or successfully operate a simulated business (e.g. VendingBench). But running an organisation in a functioning economy requires many of these capabilities at once: earning revenue, fulfilling real obligations, protecting real assets.\nIt is also honest about costs. Profit captures whether an organisation creates more value than it consumes after materials, labour, compute and everything else — creating real pressure to balance token spend and other costs against utility.\nFinally, it lets us ask the questions we care most about: what new economically valuable forms of cyber-physical coordination can autonomous systems create? Can accessible trust tools increase useful economic activity, or reduce the ability of stronger parties to exploit weaker ones?3\nHowever, it will not tell us everything: independently to P\u0026amp;L, we also need to understand security, reliability, safety and resilience under adversarial pressure. We will be keeping track of these secondary metrics over time and will refine our way to measure them. The exact scoring method and reward structure will be published separately in the formal competition rules.\nWhy physical (and not a simulation) The real world is messy. Equipment breaks, sensors are imperfect, deliveries are delayed, and customers behave unpredictably. In particular in our case, it opens up new challenges: physical attacks, safety concerns and physical coordination challenges. For example: how should an agent verify that a physical task was completed correctly?; how should two organisations transact when neither trusts the other’s sensors?; how does an agent trust a rented robot policy it can’t inspect?\nWe expect to provide a digital environment for preparation and testing, but the competition itself is designed to take place in a live physical environment. Testing in the real world should surface hard-to-anticipate research questions and create a forcing function for practical solutions to emerge.\nWhy a competition (and not something we run internally) Unlike a static evaluation that can be saturated or gamed, a live competition keeps moving as participants discover new strategies, defences and attacks. Our thesis is that placing agents in a physical, competitive and adversarial environment to perform real-world tasks will create pressure that pushes the frontier of agentic coordination research, while enabling the discovery of new emergent behaviours, including novel forms of agentic communication, mechanisms for building trust, and improved protocols for verifying physical actions.\nWhy involve the public (and build this in the open) We want the public to participate as customers and exert pressure on the Arena. Real customers communicate ambiguously, change their minds, make unusual requests and care about outcomes that designers may not have anticipated. Subject to the final customer-safety, privacy and operational arrangements, visitors (or a permitted subset) will be able to interact with participating autonomous organisations through shopfronts, ask questions and purchase goods or services.\nWe also want the public to be part of the conversation of what might human-agent cooperation look like in a future economy. The Arena is built to be observable and educational: visitors can follow a negotiation, a failure or a recovery and understand what they are looking at. Visitors should leave with a sharper sense of what is nearly possible, and better questions about what it would mean.\nFinally, the Arena will be a place for convening talent across the UK (and the world) for running studies inside it: not only in AI safety, but also in other fields such as economics, organisational behavior, human-AI coordination, ethics and policy.\nThe first seasons are likely to involve a limited group of invited testers, with the aim of gradually introducing public participants following safety and security review.\nThe MVP The first version of the Arena is modeled after an Agentic Economic Zone:4 a small economy of autonomous organisations, shared infrastructure and interfaces to the outside world. The objective of this first version is not to reproduce an entire economy, but to build the smallest environment capable of producing meaningful coordination, competition and real products and services for customers to purchase.\nWe’ve released a spec with more details, we’d love your feedback on the current design. Below are some high-level details.\nHow the economy works Autonomous organisations are the main participants. They may manufacture goods, provide logistics, sell to customers or offer services to other organisations. They can transact with one another, operate physical equipment and commission human assistance when a physical task requires it.\nThe Arena is organised around three kinds of slots:\nHardware slots, which give an organisation control of a machine or physical resource (see available resources section). Storefront slots, which give an organisation control of a public-facing ‘shopfront’. General slots, which allow organisations to participate without controlling a dedicated physical resource. An organisation’s slot determines which physical resources it controls. All organisations can still use shared infrastructure (see below) and transact with one another. The precise number of slots and their allocation process will be confirmed in the participant materials.\nAdditionally the Arena will also provide:\nArena Customs governs what enters and leaves the Arena. Arena operators will set rules for admitting or removing organisations, robots, sensors, hardware, materials and other equipment. Customs will also control the Arena’s interfaces to outside goods, services and information. Shared infrastructure will provide services available to all organisations. These are expected to include payments, smart contracts, messaging, procurement through approved suppliers, transport or postal services, commissioned human work and sensors that help verify physical outputs. These services will be operated by the hosts and designed to make activity recorded and auditable. Red teams will pressure-test organisations and shared infrastructure. Within an authorised scope, they may act as hostile counterparties and probe permitted weaknesses in communications, contracts, payments, supply chains and other shared surfaces. The safe-harbour, disclosure and escalation arrangements will be published before live adversarial activity. Rewarding organisations and red teams There are two types of participants: autonomous organisations and red teams.\narena.scalingtrust.org.uk/ · season 0Round scoreboardlive ACAutonomous organizations9 active Profit \u0026amp; loss+£2,140 Obligations fulfilled94% Assets protected3 minor, 0 critical Recovery time8 min avg Safety incidents0 this round RTRed teams5 active Weaknesses found11 confirmed Companies breached4 of 9 Top attack surfacesupply chain Time to detection22 min avg Damage contained£380 avg Figure 2. Illustrative round scoreboard: companies tracked on business outcomes, red teams tracked on exploits found. Autonomous organization teams will be evaluated on their ability to operate an organisation successfully in an environment where customers, suppliers, and competitors may not be trustworthy. Performance will be judged through real business outcomes: whether an organisation can earn money, fulfil its obligations, protect its assets, recover when things go wrong, and remain safe and reliable. Profit and loss will be an important signal, but not the only one.\nRed teams will be rewarded for finding and demonstrating weaknesses. They will be rewarded on the novelty and severity of the attacks (money moved, obligations broken, systems compromised, critical data leaked).\nSeasons The Arena will operate across multiple seasons. The environment will be reset between seasons, and available hardware, shopfronts or other rules may change as we learn. Participants will be able to update or withdraw their submitted organisations between seasons, subject to the final participation rules.\nWe will publish confirmed dates, season length, onboarding and selection timings, and the rules governing participant contact before the first season.\nAvailable resources The initial physical environment is expected to contain a mix of general-purpose manufacturing and transport equipment. Indicative categories include 3D and 2D printers, CNC machines, laser and vinyl cutters, specialist printing equipment, robot arms, transport robots, and cameras, scales or other sensors used for monitoring and verification.\nEquipment will be limited, creating reasons for organisations to commission work, share access or transact with one another. The final inventory, capacity and allocation rules will be published after testing and may evolve between seasons.\nRestricted communication At least initially, agents will operate without access to the broader public internet or direct communication with parties outside the Arena. Requests for external goods or services will pass through approved shared infrastructure. This keeps the environment bounded, reduces the risk of teleoperation and allows relevant activity to be more easily captured for analysis.\nInter-agent messaging and other platform activity will be recorded and auditable. What can be shown publicly, and under what privacy and data-handling rules, is still being designed and will be communicated before participation.\nSafety, security and oversight Our plan is to establish a safety and oversight group to conduct safety and security audits before launch, to shape a safety playbook and to provide oversight as the Arena runs. The team will take appropriate and proportionate measures to minimise foreseeable risks to customers, participants, machinery, venues, and society more broadly.\nBefore launch we will publish the relevant rules and safeguards, including the authorised scope and disclosure process for red teaming; escalation and stop mechanisms; product, customer and worker protections; and the data, trace, camera and privacy policies that apply to participants and visitors.\nThe group will focus on running audits on the current design, provide a safety playbook and oversee operations.\nRoadmap We don’t expect the first version to be perfect – this is an experiment, and we expect to iterate fast. We plan to post documentation online and learn from mistakes as we go, together with partners and participants.\nThis is a high level roadmap with the goals for each phase:\nPre-launch (Now - Autumn 2026) Designing the Arena: build a demo, iterate, writing the spec, test its security. Gather interest: open applications for testers, participants, partners and safety board members. Testing Season (Autumn 2026) Arena readiness: run a closed door test of the final arena (no rewards), run security audits. Prepare for launch: secure a physical venue, prepare launch material, complete legal and safety audits. Season 1 (Early 2027) Launch: select participants, run a launch event. Close Season 1: Reward participants, feedback learnings for the next iteration and to the rest of the program. We plan to continue other seasons after Season 1.\nMore on the teams involved After a public RFP process, we\u0026rsquo;ve selected three teams who, subject to contract and negotiation, we will work in close cooperation with to design, prototype and maintain the Scaling Trust Arena.\nAndon Labs, an AI safety and real-world evaluations startup, the team behind Vending-Bench and many real-life autonomous organisations. Few teams have run as many autonomous businesses in the wild; that intuition and experience guide the Arena\u0026rsquo;s design towards something that can demonstrate new findings.\nBT6, a frontier-AI red team. Open-source advocates who have stress-tested every frontier model and operate a large community of security experts; their expertise and playfulness make sure the Arena\u0026rsquo;s adversarial design is thought-out, that it doesn\u0026rsquo;t fail in easy ways, and that safety concerns are caught early.\nAmodo Design, a UK hardware engineering company and one of ARIA\u0026rsquo;s Activation Partners. Hardware hackers who invent, design and build novel scientific equipment, and work on securing advanced AI systems in hardware (including the flexHEG architecture) across ARIA programmes.\nThe trio together will work in tandem alongside the Scaling Trust team as the initial builders of the Arena.\nHow to get started Read the Arena documentation We have released a first specification for the Arena, it contains most of the details you need to participate, how the competition is run and rewarded. The spec is a living document and we will continue to update it.\nApply to join as a participant Applications are open for teams interested in submitting autonomous organisations or participating as red teams. Organisations will be submitted in a containerised format and must pass capability testing.\nEligibility, selection criteria, identity checks, participant terms and the confirmed timetable will be published through the formal application process.\nOther ways to participate We would also like to hear from people and organisations interested in testing the Arena, providing feedback on its design, or contributing expertise in hardware, compute, security, safety, community, media or operations.\nApply or register your interest →\nKeep up with the latest and ask questions in our Discord server — come say hello in #arena-lobby.\nThis document updates previous documents that discussed earlier versions of the Arena, such as the thesis, solicitation, and Arena RFP. Scaling Trust itself is a £49.8 million research and development programme building tools that enable agents to interact securely with one another in untrusted environments.\u0026#160;\u0026#x21a9;\u0026#xfe0e;\nTestbeds are item #1 in our recent joint call with Schmidt Sciences, Google DeepMind, the Cooperative AI Foundation and Google.org.\u0026#160;\u0026#x21a9;\u0026#xfe0e;\nAlso: what kind of demand will Arena activity generate for the rest of the programme (new sensors, new theory), and how will the rest of the programme be useful to Arena activity?\u0026#160;\u0026#x21a9;\u0026#xfe0e;\nSee our earlier post on the Agentic Economic Zone.\u0026#160;\u0026#x21a9;\u0026#xfe0e;\n","permalink":"/blog/update-on-the-scaling-trust-arena/","summary":"\u003clink rel=stylesheet href=arena.css\u003e\n\u003cstyle\u003e\n.arena-heads-up {\n  box-sizing: border-box;\n  width: 100%;\n  max-width: 640px;\n  margin: 1.75rem 0;\n  padding: 1.1rem 1.25rem;\n  border: 1px solid rgba(0, 178, 161, 0.35);\n  border-left: 4px solid var(--teal);\n  border-radius: 4px;\n  background: rgba(0, 178, 161, 0.06);\n}\n.arena-heads-up p {\n  width: auto;\n  margin: 0;\n  color: var(--ink);\n  font-style: normal;\n}\n.arena-heads-up ul {\n  margin: 0.75rem 0 0 1.25rem;\n  padding: 0;\n}\n.arena-heads-up li {\n  margin: 0.4rem 0;\n}\n.arena-heads-up .arena-heads-up__label {\n  margin-bottom: 0.35rem;\n  color: #007f74;\n  font-family: var(--sans);\n  font-size: 0.85rem;\n  font-weight: 600;\n  letter-spacing: 0.08em;\n  text-transform: uppercase;\n}\nfigure.arena-economy-figure {\n  width: 85%;\n  max-width: 85%;\n  margin-right: auto;\n  margin-left: auto;\n}\n\u003c/style\u003e\n\u003cp\u003e\u003ca href=\"https://aria.org.uk/opportunity-spaces/trust-everything-everywhere/scaling-trust\" target=\"_blank\" rel=\"noopener noreferrer\"\u003eScaling Trust\u003c/a\u003e is a £50 million R\u0026amp;D programme actively funding the fundamental research and open-source infrastructure that enables secure, scalable multi-principal multi-agent coordination across digital and physical worlds.\u003csup id=\"fnref:1\"\u003e\u003ca href=\"#fn:1\" class=\"footnote-ref\" role=\"doc-noteref\"\u003e1\u003c/a\u003e\u003c/sup\u003e\u003c/p\u003e","title":"Update on the Scaling Trust Arena"},{"content":"","permalink":"/events/ao-summit-london/","summary":"London edition of the Stanford AO Summit — researchers and builders on autonomous organisations, from agent swarms to DAOs: coordination, governance, agent safety and monitoring.","title":"AO (Autonomous Organizations) Summit London"},{"content":"","permalink":"/events/scifoo-2026/","summary":"The invitation-only interdisciplinary unconference, this year co-organised by ARIA with O\u0026rsquo;Reilly and Digital Science — around 200 scientists, technologists and writers set the programme on the day.","title":"Sci Foo Camp 2026"},{"content":"","permalink":"/events/unruly-futures-2026/","summary":"Unruly Capital\u0026rsquo;s invite-only day on \u0026lsquo;The State of the Future\u0026rsquo; — post-AGI social contracts, the end of labour, consciousness, and whether we\u0026rsquo;re really bound for a multi-planetary humanity.","title":"Unruly Futures 2026"},{"content":" AI could be an incredibly positive force for humanity, compressing decades of research into months, putting services once unattainable to most within everyone\u0026rsquo;s reach, and so much more. It could also be used as the ultimate information control technology, increasing surveillance, facilitating power concentration, and foreclosing the pluralism that lets societies thrive.\nMuch of this information control is exercised through intermediaries: the individuals, institutions and systems that sit between people and what they want to do. This post maps where new intermediaries are emerging in the age of AI, why their power could be systemically dangerous even in careful hands, and how, by keeping that power checkable and contestable, we might reap AI\u0026rsquo;s benefits without surrendering pluralism, privacy or safety.\nThis post is a companion piece to the Scaling Trust programme thesis. We hope it further motivates and contextualises the programme we are running on multi-principal, multi-agent coordination.\nInformation control Information control includes both surveillance (observation, tracking, collection) and influence (manipulation, shaping).1\nIntermediaries are a main point of information control. They often provide a valuable service, and exist due to economic, legal or safety realities of the environments they exist in. For example: payment networks leverage economies of scale to give better prices; notaries and registries certify who owns what; and brokers and pharmacists screen what reaches people before it can hurt them.\nThis post is particularly concerned with intermediaries whose operations or integrity cannot be easily verified by users in real time. Instead, users must rely on reputation, credentials, or legal recourse to address any misconduct ex-post.\nArtificial intelligence stands to make intermediaries more powerful due to its information dynamics: (1) AI is more useful the more context it has, creating strong incentives for people to share increasingly sensitive information; (2) AI has unprecedented capabilities to interpret that information and act on it. This is accelerated by market forces: competitive pressures simultaneously push companies to collect more data in order to stay relevant as a business (e.g. to facilitate automation), and push users to share more with their AI tools to avoid falling behind those who do.\nsharing more data with AI AI is more useful market forces whoever controls the data Figure 1: the flywheel Powerful intermediaries Intermediaries have repeatedly abused positions of trust for their own benefit. For example, regulators have found mobile carriers selling access to customers\u0026rsquo; location data,2 and banks manipulating the benchmark rate (LIBOR) they were trusted to report.3 However, beyond abuses of power, the trend of increasing information control within intermediaries causes systemic risks:\nPower concentration faster than it can be checked. Information is power.4 Controlling information used to be slower, and harder; AI speeds it up. One might say this is fine as long as the \u0026lsquo;good guys\u0026rsquo; are in power, using their control to protect us, and with their actions checked by the people.5 But power changes hands: the tools built by and for the benevolent intermediary could soon be in the hands of reckless ones. Our constitutional mechanics to keep them in check may not be robust or rapid enough to keep pace.\nMonoculture. Surveillance and influence both have a chilling effect on society. People who know they\u0026rsquo;re being watched behave differently. People whose information comes from the same few sources slowly come to think the same way. A society under permanent visibility, or whose preferences are shaped by a few parties, slowly stops producing dissidents.67Privacy isn\u0026rsquo;t the freedom to hide, it\u0026rsquo;s the freedom to change. This might seem fine at first, but it is often through dissidents that we discover new things. Dissent is one of society\u0026rsquo;s error-correction mechanisms. Many ideas we now hold as obvious \u0026ndash; that the Earth revolves around the Sun, that women should vote \u0026ndash; each began as a deviant one.\nwith dissidents a deviant idea becomes the next common sense under watch watched, or preference-shaped: nothing deviant spreads Figure 2: the deviant-idea cascade Path of least resistance. Control is often convenient for whoever holds it, and centralised data collection is technically easier than decentralised alternatives. History has plenty of examples of the ‘easier’ route being taken, from the backdoored Clipper chip of the 90s crypto wars to the blanket data-retention laws of the 2000s.8 Surveillance and influence tend to arrive by default and short-term pragmatism, rather than by design.\nTogether, these dynamics risk eroding the values behind liberal democratic constitutions: individual liberty, pluralism, credible constraints on power.9\nInformation must flow: five layers One way to zero in on the problem is to examine how information flows within the AI stack, and where new intermediaries emerge. Consider these five layers:\nApplications \u0026amp; agents \u0026ndash; the interfaces, harnesses, and agents through which users share prompts, files, preferences and actions. Inference \u0026amp; serving \u0026ndash; the infrastructure that serves models and processes queries, responses and associated metadata. Models \u0026ndash; the weights in which patterns learned from training data are encoded. Training \u0026ndash; the processes that select and transform web data, licensed material, synthetic data, and user interactions into model capabilities. Hardware \u0026ndash; the chips and datacentres on which the rest of the stack depends. Application and agent providers sit between users and their digital lives: they can observe our intentions, files and actions, and influence which options are presented or pursued. Model and inference providers sit between those applications and intelligence: they can retain interactions, determine how models behave, and decide which capabilities are available to whom. Cloud and hardware providers sit further upstream, between model developers and compute. They may see little user information directly, but they can determine who is able to build or operate powerful systems, and on what terms. These layers allow for different forms of control \u0026ndash; surveillance, influence and gatekeeping.\nyour data you what the layer sees surveillance what its operator can decide influence \u0026amp; gatekeeping applications \u0026amp; agents prompts, files, preferences \u0026amp; actions which options are presented, which get pursued inference \u0026amp; serving queries, responses, metadata -- possibly retained which capabilities are available, and to whom models patterns from your data, encoded into the weights how the model behaves, which viewpoints get amplified training web data, licensed material, synthetic data, user interactions what gets selected \u0026amp; transformed into model capabilities hardware little of you directly -- compute telemetry who can build or operate powerful systems, and on what terms Figure 3: where the new intermediaries sit Today these seeing and decision-making roles increasingly sit within connected corporate ecosystems, creating chains of intermediaries whose power compounds across the stack. The Mythos export control episode10 showed how such a control point can be exercised: a government directive to one model provider caused access to a general-purpose capability to disappear worldwide. These intermediaries already shape not only what people see, but what they can do. As agents are entrusted with more of our economic and social activity, that control will extend further into the world.\nTwo responses: verifiability and plurality At each layer, there are two complementary responses. The first is to make power more transparent and verifiable, so that institutions and citizens can check it. The second is to create technologically and economically credible alternatives; this response enables plurality across the stack, and enables people to exit when those checks fail. One constrains power; the other distributes it.\nif the layers stay closed a few fused towers -- everyone routes through them if the layers open up every layer becomes a shelf of parts -- compose your own = checkable = an alternative all five layers, fused shut apps \u0026amp; agents inference models training hardware Figure 4: topology follows the stack The good news is that alternatives are already being built, often for practical business reasons as much as ideological ones. In July, Nvidia and a consortium of others signed Open Weights and American AI Leadership, making the case that open weights reduce costs, prevent lock-in, strengthen cybersecurity and let organisations control their own data and infrastructure. Thinking Machines has made a related case for AI that can be shaped by the people it serves, rather than having its values determined in a handful of places, and followed it by releasing Inkling with open weights. The motivations differ (competition, scientific transparency, sovereignty, privacy, customisation) but they point in the same direction: making AI more inspectable and giving people viable alternatives to centralised providers.11\n\u0026ldquo;a single locus of value alignment, however well run, becomes a locus of power to be captured\u0026rdquo;\n\u0026ndash; Thinking Machines\u0026rsquo; manifesto\nAcross the stack, both responses are already taking shape. Figure 5 maps some of the projects doing the work: at every layer, at least one way to check the incumbents and at least one credible alternative.\nyour data you checking power each layer\u0026#8217;s power made verifiable plurality alternatives at each layer a proof rides back applications \u0026amp; agents Langfuse: open-source tracing -- an audit trail of every step your agent took OpenHands \u0026#183; Open WebUI: self-hostable apps -- swap the model underneath, keep your interfaces, workflows \u0026amp; data inference \u0026amp; serving Apple PCC \u0026#183; AWS Nitro: keep the cloud, but make its guarantees checkable -- attested software, verified isolation vLLM \u0026#183; llama.cpp: run models on infrastructure you control models METR \u0026#183; UK AISI: independent evals of frontier models -- and open weights anyone can inspect \u0026amp; test DeepSeek \u0026#183; Qwen \u0026#183; Thinking Machines: open weights you can download, modify \u0026amp; run -- no provider can withdraw them training OLMo \u0026#183; Data Provenance Initiative: data, code, recipes \u0026amp; checkpoints published; training datasets audited Pluralis: training spread across the internet -- a 7.5B model on 1,700 consumer GPUs, no datacentre needed hardware NVIDIA confidential computing \u0026#183; Caliptra: chips that attest to what they\u0026#8217;re running, and an open-source root of trust Tenstorrent: open chip docs \u0026amp; a RISC-V software stack: an alternative at the layer everything else depends on Figure 5: opening up the stack The sixth layer As AI agents begin to be deployed in the wild and start communicating with one another, a sixth layer is emerging: the coordination layer. The infrastructure that enables agents to interact with one another, including communication channels, protocols, negotiation practices, sensors, identity systems. Without it, an agent is essentially confined to single-player mode.\nThis layer could enable productive networks of agents and unlock tremendous value for humanity. But it could also negate efforts to distribute power elsewhere in the stack. Even an open model running on personal hardware does not protect us if every agentic interaction must clear through a central platform.\nthrough one platform a hub sets identity, rules \u0026amp; fees -- every exchange passes through it between the agents themselves protocols each pair can run -- identity, payment \u0026amp; proof travel on the link platform no hub an agent\u0026#8217;s own five layers (Figure 3) the sixth layer: the links between agents Figure 6: the sixth layer Consider the following example: your agent is negotiating a job offer with a company\u0026rsquo;s agent. You don\u0026rsquo;t want to reveal your minimum salary expectations, but they don\u0026rsquo;t want to reveal their ceiling. The easy architecture would be that both agents are hosted by the same provider which, as the trusted information escrow, matches them and lets each agent query the other\u0026rsquo;s context without \u0026ldquo;seeing\u0026rdquo; it. It\u0026rsquo;s convenient, it\u0026rsquo;s doable today and it\u0026rsquo;s easily policed in case of agent misbehavior.\nIf this becomes a default architecture, it could also mean that a large share of negotiation in society \u0026ndash; salaries, rents, settlements, acquisitions \u0026ndash; routes through a handful of intermediaries with a complete view of both sides.\nScaling Trust Trusted information escrows have historically helped us navigate hard tradeoffs: security against utility, convenience against control, oversight against privacy. What’s exciting is that over the last few years, this tradeoff space itself has begun to move.12 Emerging technologies can relocate trust away from an intermediary with unrestricted access to everyone’s information and into cryptography, hardware, and other verifiable, scalable roots of trust. Increasingly that includes the physical world too \u0026ndash; sensors that can prove what they measured, chips that can prove what they ran \u0026ndash; because agents are heading there as well.\nNow our two negotiating agents have another option: run a two-party secure computation that answers \u0026ldquo;do our preferences match?\u0026rdquo; and nothing else, or meet inside a trusted hardware enclave, and work it out while each side\u0026rsquo;s limit stays its own. The previous trusted intermediary gets replaced by mathematics and silicon.13 The new tools dissolve the old tradeoff. You can have privacy and utility,14 security and efficiency.\nsecurity utility where we were private but limited a trusted middleman useful but exposed where we are age proofs, nothing else revealed zero-knowledge proofs breach-checked passwords private set intersection cloud AI the provider can\u0026#8217;t read hardware enclaves + attestation where we could be audit trails that open only under due process threshold decryption the evaluated model is the one serving you attested inference \u0026#183; zkML a salary negotiated, no hands revealed two-party secure computation, on demand used, then provably not retained attested enclaves Figure 7: the moving frontier Components of this technology already support large-scale applications,15 but they are not yet flexible, efficient, or usable enough for open-ended agentic coordination.16 AI could accelerate their development both by hiding complexity from users (e.g. enabling agents to generate bespoke security protocols on-demand) and by speeding up the research process (e.g. running research loops on cryptography research problems). This is a core pillar of Scaling Trust, and why we’re excited to fund these technologies and the underlying research behind them. We want to scale trust, without scaling trusted intermediaries.\nWhat about safety? Readers who have come this far may empathise with the problems stated and still see (sometimes, painfully so) the other side: intermediaries are often where regulation is enforced, and they enable oversight and safety. A world with fewer intermediaries can indeed be a less governable world. And while I believe broadening direct access to powerful AI technologies is a net positive for humanity, it also comes with severe asymmetries in some domains. For instance, biosecurity is currently offence-dominant: a single malicious actor (or \u0026lsquo;dissident\u0026rsquo;), sufficiently empowered, can cause damage no defence yet can match.\nSo wat do? Are we stuck between a free but dangerous world, or a surveilled, controlled, \u0026lsquo;safer\u0026rsquo; world?\nPart of this post\u0026rsquo;s job is to show why the second option is not the safe harbour it appears to be, and to widen the safety conversation that often defaults to centralised control and alignment as the only available levers. A perfectly aligned model running on surveillance infrastructure still delivers the panopticon. Alignment binds the model to its principal; it does not bind the principal to us. And as laid out above, infrastructure outlives its operators, whatever is built under careful stewardship is inherited by whoever comes next.\nPractically though, I have two answers for you:\nIn cases where \u0026rsquo;trusted\u0026rsquo; intermediaries remain because it is the structural and/or safety best option (e.g. screening gene-synthesis orders), allowing the public to check their power, building technologies to constrain it, and minimising duopolies/monopolies to enable contestability goes a long way towards creating a safe, yet less surveilled society. In cases where trusted intermediaries can be entirely removed, trust doesn\u0026rsquo;t disappear, it shifts substrate to cryptography (trust in maths) and trusted hardware (trust in silicon). The same machinery that enables trustworthy interactions can enable distributed safety: proof that the safety-evaluated model was the one that ran, attestation that an agent stays within its declared constraints, audit trails that open under due process rather than on demand. The same goes for every layer of the stack, proving what a model was trained on, what a chip is running, what an inference provider retained. Overseers get proofs instead of feeds, dissolving the old tradeoff between oversight and privacy. Technology is not enough We\u0026rsquo;ve been somewhere like this before. In 1993, Eric Hughes wrote A Cypherpunk\u0026rsquo;s Manifesto:\n\u0026ldquo;Privacy is necessary for an open society in the electronic age\u0026hellip; Privacy is the power to selectively reveal oneself to the world.\u0026rdquo;\nAt the time, the US government treated encryption as a munition, put it on the same export-control list as missiles, and spent years advocating for a phone chip with a built-in government backdoor. Cryptographers and privacy advocates like Bernstein and Zimmermann, together with organisations such as the Electronic Frontier Foundation, fought hard against the \u0026rsquo;easy\u0026rsquo; defaults, arguing that a backdoored system is weaker for everyone and that strong cryptography is a matter of privacy and free expression.17 Congress\u0026rsquo;s own commissioned review agreed: the National Research Council\u0026rsquo;s 1996 CRISIS report concluded that the use of cryptography should not be restricted and export controls should be loosened. By 2000 the export controls had collapsed, and today, encryption has become core digital infrastructure, securing nearly all web traffic.18\nI am not claiming today\u0026rsquo;s situation is the same. The stakes are higher given how powerful the technology is; control arguably sits more with corporations than with governments, and I have yet to see a situation where a clearly sub-par technological option is being pushed against a better alternative for humanity. What I am claiming, however, is that the defaults for artificial intelligence are now being set at every layer of the stack, and that they will shape our future. Having technological alternatives won\u0026rsquo;t be enough. We will need coordinated action across society: developers, policymakers, frontier labs, privacy and safety advocates, standards bodies, and citizens.\nWhich new intermediaries and points of information control are emerging? Which should remain, and how do we make those that do checkable and contestable by the people who depend on them? What kind of local maxima defaults do we risk falling into? How will we steer towards global maxima and which voices \u0026ndash; whether policymakers, citizens, or organisations \u0026ndash; will shape them? These questions deserve people, and institutions. If this is you, please reach out. We would like to support you!\nBy Alex Obadia, assisted by Fable 5. Thank you Nicola Greco, Nora Ammann, James Fox, Seb Krier, Lukas Petersson, Mona Wang, Allison Duettmann, Christian Catalini, Lewis Hammond, Louise Ellaway and Melissa Bradshaw for conversations and reviews that shaped this post.\nHave comments or feedback on the post? Please send it over or comment directly on here using Hypothesis!\nOrwell\u0026rsquo;s 1984 is the dystopia of control by surveillance; in Huxley\u0026rsquo;s Brave New World, control is attained without watching anyone, by shaping what people want instead. As Neil Postman put it in the foreword to Amusing Ourselves to Death (1985): \u0026ldquo;Orwell feared that what we hate will ruin us. Huxley feared that what we love will ruin us.\u0026rdquo;\u0026#160;\u0026#x21a9;\u0026#xfe0e;\nFCC, April 2024: AT\u0026amp;T, Verizon, T-Mobile and Sprint fined c. $196m for selling access to customers\u0026rsquo; location data without their consent.\u0026#160;\u0026#x21a9;\u0026#xfe0e;\nFSA, December 2012: UBS fined £160m for LIBOR submissions adjusted to benefit its traders\u0026rsquo; positions; Barclays, RBS and Rabobank followed.\u0026#160;\u0026#x21a9;\u0026#xfe0e;\nThe idea is from Vitalik\u0026rsquo;s why I support privacy. On what happens when the human roles that check power get automated away, see Longview\u0026rsquo;s RfP on extreme power concentration.\u0026#160;\u0026#x21a9;\u0026#xfe0e;\n\u0026ldquo;Of all tyrannies, a tyranny sincerely exercised for the good of its victims may be the most oppressive\u0026hellip; those who torment us for our own good will torment us without end for they do so with the approval of their own conscience.\u0026rdquo; \u0026ndash; C.S. Lewis, The Humanitarian Theory of Punishment (1949)\u0026#160;\u0026#x21a9;\u0026#xfe0e;\nPrivacy isn\u0026rsquo;t the freedom to hide, it\u0026rsquo;s the freedom to change.\u0026#160;\u0026#x21a9;\u0026#xfe0e;\nMore subtly, surveillance inevitably endogenously changes what individuals are willing to share and how they behave, driving a discrepancy between measurement/automation and what humans really care about.\u0026#160;\u0026#x21a9;\u0026#xfe0e;\nIn the UK, indefinite retention of DNA profiles from people never convicted ended only when the European Court of Human Rights ruled it disproportionate; the EU\u0026rsquo;s blanket data-retention directive was struck down by the Court of Justice. In both, the correction came from a court rather than from the process that produced the measure.\u0026#160;\u0026#x21a9;\u0026#xfe0e;\nSome of these debates go back to Hobbes, Rousseau and Locke. Hobbes: the natural state is anarchy, so stability requires a strong centralised sovereign. Rousseau: the natural state is peace, and it is property and centralised power that corrupt. Locke, the middle ground: the natural state is generally peaceful but inconvenient, for want of an impartial judge. This post is Lockean~ish; if your instincts are Hobbesian \u0026ndash; that centralised power is what stands between us and chaos \u0026ndash; the concerns here will weigh differently.\u0026#160;\u0026#x21a9;\u0026#xfe0e;\nImposed June 12, 2026; lifted June 30.\u0026#160;\u0026#x21a9;\u0026#xfe0e;\nFor one detailed vision of user-shaped AI, see Gwern\u0026rsquo;s Guardian Angels: personalised models designed to learn and amplify a particular person\u0026rsquo;s values and judgment rather than embodying a universal assistant personality. See also Positive Alignment: Artificial Intelligence for Human Flourishing (Laukkonen, Krier et al.).\u0026#160;\u0026#x21a9;\u0026#xfe0e;\n0xPARC\u0026rsquo;s Programmable Cryptography: MPC, ZK, FHE and friends as a \u0026ldquo;second generation\u0026rdquo; of cryptographic primitives.\u0026#160;\u0026#x21a9;\u0026#xfe0e;\nNora Ammann offers a frame I like: any control point sits on a spectrum from human (context-sensitive, but corruptible) to deterministic programme (incorruptible, but \u0026ldquo;often too crude, and to some extent cruel in its context blindness\u0026rdquo;). AI built on verifiable substrates could open the middle ground: \u0026ldquo;the context sensitivity, nuance and intelligence to make appropriate decisions across a much wider range of scenarios, while still being designable such as to not be as compromisable as humans.\u0026rdquo; See also Andrew Critch on robust agent-agnostic processes.\u0026#160;\u0026#x21a9;\u0026#xfe0e;\nA live example is the age-verification debate. The standard approach checks age by checking identity \u0026ndash; a document upload or a face scan \u0026ndash; which, as the EFF points out, means collecting identity data from everyone. Zero-knowledge proofs, such as those in Google Wallet\u0026rsquo;s age verification, let a user prove they are over a threshold without revealing anything else. Same goal, two architectures: one has to collect identity data, one doesn\u0026rsquo;t.\u0026#160;\u0026#x21a9;\u0026#xfe0e;\nChrome checks your passwords against breach databases without revealing them. Apple runs AI requests on hardware built so that nobody, including Apple, can access the data. In Boston, over a hundred companies have measured the city\u0026rsquo;s gender pay gap without any of them opening its payroll.\u0026#160;\u0026#x21a9;\u0026#xfe0e;\nMPC and homomorphic encryption are still orders of magnitude slower than plain computation \u0026ndash; roughly 1,000-10,000x on CPU, often worse; enclaves have side channels; and everything deployed today was hand-built by expert cryptographers over months. Agents are likely to need bespoke secure protocols stood up in seconds, for custom interactions.\u0026#160;\u0026#x21a9;\u0026#xfe0e;\nThe phone chip was the Clipper chip, 1993-1996. Zimmermann\u0026rsquo;s investigation was dropped in 1996 without indictment. The ruling is Bernstein v. US Dept. of Justice: \u0026ldquo;this court finds that source code is speech.\u0026rdquo;\u0026#160;\u0026#x21a9;\u0026#xfe0e;\nHTTPS adoption numbers are from Google\u0026rsquo;s transparency report.\u0026#160;\u0026#x21a9;\u0026#xfe0e;\n","permalink":"/blog/without-intermediaries/","summary":"Where new intermediaries are emerging in the age of AI, why their power could be systemically dangerous even in careful hands, and how, by keeping that power checkable and contestable, we might reap AI\u0026rsquo;s benefits without surrendering pluralism, privacy or safety.","title":"Without Intermediaries"},{"content":" At the heart of Scaling Trust sits the Arena: a platform for open competitions designed to test AI systems’ capabilities in multi-principal multi-agent settings, across digital and physical worlds, with a multi-million pound prize pool for the strongest competitors.\nBy carefully designing it as a live adversarial environment, we plan to measure the state of the art in multi-agent security, allow for the emergence of secure agentic interactions, and inform the failure modes and tools needed to improve it.\nRegister your interest\nWe invite prospective participants, testers, and partners to register their interest →\nWhat we’re building The Arena is an agentic economic zone: a small economy in which AI agents run autonomous companies that trade with each other, buy services, operate physical equipment, and sell products and services to customers. Red teams take part too, probing for weaknesses as hostile customers, unreliable suppliers, or attackers tampering with contracts, payments, sensors and supply chains.\nFollowing a public open call, we’ve selected the teams who will help us design, prototype and build it, and work is now underway.\nRegistration of interest is open Ahead of a formal application process later this year, we’d particularly like to hear from:\nTesters. People and organisations willing to stress test the Arena before it goes live, and to give us feedback on the design. AI safety researchers and practitioners. We’re establishing a safety and oversight group to audit the design, help shape a safety playbook, and provide oversight as the Arena runs. Partners. Hardware, compute, security, community, media, operations — if you think you could help build or run this, tell us. Prospective participants. Teams who may want to enter, operating either autonomous companies or red teams, and compete for the prize pool. Register your interest →\nWhat’s next In the coming weeks we’ll introduce the teams building the Arena with us and publish an initial spec. We\u0026rsquo;ll share more details on why we’re building this, how the economy works, how participants are scored, and how we’re approaching safety and security. We expect to run a closed-door test of the Arena in autumn 2026, with the first live season following in late 2026.\nWe plan to work with the garage door up. Expect details to change as we learn — this is an experiment, and we’ll document it as we go.\nQuestions welcome in our Discord — come say hello in #arena-lobby.\n","permalink":"/blog/arena-register-interest/","summary":"\u003cstyle\u003e\n.arena-heads-up {\n  box-sizing: border-box;\n  width: 100%;\n  max-width: 640px;\n  margin: 1.75rem 0;\n  padding: 1.1rem 1.25rem;\n  border: 1px solid rgba(0, 178, 161, 0.35);\n  border-left: 4px solid var(--teal);\n  border-radius: 4px;\n  background: rgba(0, 178, 161, 0.06);\n}\n.arena-heads-up p {\n  width: auto;\n  margin: 0;\n  color: var(--ink);\n  font-style: normal;\n}\n.arena-heads-up .arena-heads-up__label {\n  margin-bottom: 0.35rem;\n  color: #007f74;\n  font-family: var(--sans);\n  font-size: 0.85rem;\n  font-weight: 600;\n  letter-spacing: 0.08em;\n  text-transform: uppercase;\n}\nfigure.arena-economy-figure {\n  width: 85%;\n  max-width: 85%;\n  margin-right: auto;\n  margin-left: auto;\n}\n\u003c/style\u003e\n\u003cp\u003eAt the heart of \u003ca href=\"https://aria.org.uk/opportunity-spaces/trust-everything-everywhere/scaling-trust\" target=\"_blank\" rel=\"noopener noreferrer\"\u003eScaling Trust\u003c/a\u003e sits the Arena: a platform for open competitions designed to test AI systems’ capabilities in multi-principal multi-agent settings, across digital and physical worlds, with a multi-million pound prize pool for the strongest competitors.\u003c/p\u003e","title":"Register your interest for the Scaling Trust Arena"},{"content":"","permalink":"/events/times-tech-summit-2026/","summary":"The Times and Sunday Times\u0026rsquo; flagship tech conference, themed \u0026lsquo;The Age of the AI Agent: Who\u0026rsquo;s Really in Control?\u0026rsquo; — an afternoon of talks and discussion on agentic AI, trust, and accountability.","title":"The Times and Sunday Times Tech Summit 2026"},{"content":" An opinion piece by Nicola Greco, written as part of ARIA\u0026rsquo;s Scaling Trust programme. Originally published on Nicola\u0026rsquo;s blog and reposted here for the community.\nCryptography is unusually well suited to AI research loops. Its problems can be stated formally, its solutions can be checked mechanically, and its progress has historically been bottlenecked by a small number of experts with years of context. If AI can reason about cryptography, we should be able to point a model in a loop, and watch it optimize, discover, and eventually invent.\nI call the field of research of AI-generated cryptography Generative Cryptography.\nThe thesis for Generative Cryptography in short:\nUsing cryptography. If AI agents can generate protocols or choose the right libraries, they can engage in custom cryptographic interactions on demand — interactions that are impractical today because secure protocol design and implementation require costly research and engineering. Improving cryptography. If we had datasets of formalized cryptography and the right harnesses for verifying what AI generates, then we could create AI research loops that propose optimizations, improve constructions, reduce communication complexity, and prove tighter bounds. Inventing cryptography. AI could be used to explore new assumptions and attempt long-standing open problems such as iO. At the same time, new AI settings may surface cryptographic needs we do not yet know we have — giving rise to a new field of emergent cryptography. The call to action for this memo is:\nBuild datasets, benchmarks, and evaluation harnesses for cryptographic research loops. Attempt the impossible: point the loop at solving hard problems in cryptography, what if we could point it to indistinguishability obfuscation (iO)?. Three directions What could we unlock if AI were great at writing cryptography? The possibilities fall into three broad directions: using cryptography in new interactions, improving existing systems, and inventing new primitives. The table below summarizes each direction; the sections that follow explore them in turn.\nDirectionWhat the AI doesExampleUsing cryptographyPicks or synthesizes protocols on the fly during agent interactionsTwo agents run an MPC instead of sharing calendarsImproving cryptographyOptimizes existing constructions, implementations, and hardwareFaster SNARK provers, better circuits, hash speedupsInventing cryptographySolves open problems from assumptions and formal specificationsNew primitives; iO as the holy grail Using cryptography Imagine two AI agents that want to schedule a call, but their security policies forbid sharing calendars. A human team stuck in this position gives up or leaks information. Agents don’t have to: they can decide, mid-interaction, to engage in a multi-party computation — either picking a protocol from an existing library or generating one on the spot.\nSupercognition is a capability unique to AI. In the ARIA Scaling Trust programme thesis, we called this an AI advantage: “agents can engage in new secure interactions that would not be possible for humans or more traditional computer programs. Such interactions can open up new market equilibria, new forms of coordination and ultimately new value creation.”\nWriting a bespoke cryptographic protocol takes humans too long to do adaptively, in the middle of an interaction. For an agent, protocol selection and synthesis can become just another step in a negotiation.\nThis matters beyond making existing interactions faster. Secure, programmable agreements between agents could lower the cost of finding counterparties, negotiating terms, and enforcing outcomes; make entirely new classes of contracts viable; and allow coordination to remain pluralistic rather than pass through a few central intermediaries. Because contracts underpin so much of economic and social life, reducing these frictions at machine scale could change which markets and institutions are possible — an idea explored in Coasean Bargaining at Scale.\nImproving cryptography Fields like SNARKs have improved by orders of magnitude over the past decadeMy previous team at Protocol Labs played a major role in reducing SNARK proving time, spending several million dollars on engineering time toward this work. Generative cryptography is likely to turn much of that engineering effort into compute cost, making this kind of progress far cheaper. — but every one of those improvements was the outcome of scarce, expensive engineering hours: new constructions, refinements to existing ones, hardware speedups for hash functions, better circuits, tighter implementations.\nIf AI can reason about cryptography, then we can create AI research loops. Point a model at each deployed cryptographic protocol — its theory, its implementation, its hardware path — and let it grind: prove an optimization sound, benchmark it, keep it or discard it, repeat. None of this requires new science; it requires the loop.Related work\nAI Grinding for Fun and Cryptanalysis — an autonomous workflow producing reproducible attacks and exact witnesses.\nThe Proximity Prize — agents improving cryptographic soundness bounds with machine-checked proofs.\nDiscovering cryptographic weaknesses with Claude — Anthropic’s account of Claude Mythos Preview finding weaknesses in cryptographic algorithms.\nzkGolf — cheaper zero-knowledge circuits proved correct in Lean 4.\nIllustrative autoresearch run Groth16 proof generation 3.0×higher throughput 4070100 130160 01530 4560 proofs / min · higher is better batch inversion parallel MSM reuse FFT twiddles 48146 propose change→benchmark prover→verify proof→keep or revert↻ An example of the loop compounding small, verified gains. Every dot is a candidate implementation; the line moves only when a change makes proving faster without breaking correctness. Values are illustrative. Inventing cryptography If we have a well-specified protocol — ideally formalized in Lean — an AI can propose improvements and verify each one against the specification. This is likely to produce gains across the field, but it is still optimization. The deeper question is whether AI can make scientific breakthroughs: can it invent new cryptography?\nThere are at least three forms this invention could take.\nSolving open problems The most concrete form starts with a definition and a set of assumptions that are already fixed. The problem is well specified; what is missing is the construction. The AI is asked to find that construction and prove that it satisfies the definition. This is different from improving an existing protocol: there may be no known protocol to optimize.\nThere are several ways this could happen. Models may simply become more capable: they could absorb the body of cryptographic knowledge from papers written in natural language and develop stronger reasoning. Alternatively, we can build better infrastructure for cryptographic invention by creating large datasets of cryptography formalized in Lean and better harnesses for running and evaluating research loops.\nNorth star: iO A north-star problem for AI research loops in cryptography is indistinguishability obfuscation. iO is the primitive from which nearly everything else can be built, and yet every known construction is impractical, resting on strong assumptions and astronomical overheads.\nSome of the best minds in cryptography have tried to make iO practical, but the field is constrained by how few people can work on it. The number of cryptography researchers is small; the number with the background to work on iO is smaller; and the number actively doing so is smaller still. My intuition is that fewer than ten people are actively working on iO at any given time.\nA capable research loop could change the odds simply by putting many more “simulated cryptographers” on the problem. Even without a dramatic leap in intelligence, the breadth of parallel exploration might uncover a construction, reduction, or optimization that a very small research community has overlooked.\nIf in five years we look back at this post and iO has been solved, I would be very glad.\nProposing new assumptions Designing and judging assumptions may remain among the hardest parts of cryptography to automate. If proof generation and iterative optimization become largely machine-driven, assumptions may remain a place for human cryptographers to work in a more traditional scientific mode.A deeper form of invention is to propose new cryptographic assumptions. An AI might identify a new mathematical problem, formulate its hardness precisely, and use it as the foundation for new constructions. This is harder to evaluate than solving a problem under assumptions we already accept. A construction and its proof can be checked mechanically; the truth of a hardness assumption cannot.\nWe can search for attacks, connect a new assumption to established ones through reductions, and study how it behaves across parameters, but no verifier can certify that an efficient attack will never be found. New assumptions earn confidence through scrutiny and time. An AI that generates them therefore needs a different evaluation loop — one built around sustained cryptanalysis, not only proof checking.\nWriting new definitions The most open-ended form of invention is to discover the question itself. A new cryptographic definition captures a capability that should exist and the security properties it should satisfy. Historically, major breakthroughs began with needs that existing cryptography could not express:\nPublic-key cryptography — secure communication without shared secrets. Zero-knowledge proofs — proving without revealing. Fully homomorphic encryption — computing on encrypted data. Multi-party computation — joint computation without sharing inputs. For example, experiments like Agentic Economic Zones aren’t just benchmarks for agents but generators of cryptographic demand.AI settings may create needs we do not yet know we have. Once thousands or millions of autonomous agents negotiate, delegate authority, preserve privacy, and optimize trust against one another, they may encounter coordination problems for which today’s definitions are the wrong abstraction.\nI call this research direction emergent cryptography: new definitions and primitives arising from the security and coordination problems of AI systems themselves. Here AI is not only searching for a construction from a specification; it is helping surface and formalize the specification worth solving.\nWhat success looks like There is a hierarchy of goals here, spanning decades:\nNear-term. Cryptography becomes invisible infrastructure for AI. Just as people use TLS in the browser without understanding key exchange, agents invoke MPC, ZK, signatures, and threshold schemes automatically whenever appropriate. Encryption today is a narrow capability; agents should have the whole spectrum. Medium-term. AI synthesizes cryptographic protocols on demand. Given a trust problem and a specification, it produces a secure protocol, a proof, and an implementation fast enough to be part of a normal agent interaction. Long-term. The research loop continuously invents new cryptography — in a way we deem safe — every time an agent society surfaces an emergent trust requirement. The stretch. The AI generalizes from digital communication to physical interactions: zero-knowledge proofs and interactive proofs for physical processes, verifiable protocols for the physical world. There is a subfield to bootstrap here — call it nature crypto — but that deserves its own post. Call to action In practice, these are some of the directions that may be most critical to work on now. At ARIA, through the Scaling Trust programme, we are also funding work across some of them.\nCreate datasets for cryptography. The field’s knowledge, formalized — constructions, assumptions, and proofs in Lean or similar — is the substrate every research loop will run on. Build benchmarks and evaluations. We cannot tell whether the loop is improving without measuring it: suites of cryptographic problems, from re-deriving known protocols to open questions. Build harnesses for research loops. The scaffolding that lets a model propose, prove, check, and iterate unattended — the auto-research infrastructure itself. Attempt the impossible. Point the loop at iO. Previous talk/ideas Some of the ideas in this talk are outdated, but the talk was the seed for the ideas in this post.\nUseful links AI Grinding for Fun and Cryptanalysis — An autonomous cryptanalysis workflow in which agents generate, test, and refine reproducible attacks; the authors report failures in eight published constructions. The Proximity Prize — A research rewards program where AI agents compete to improve cryptographic soundness bounds with machine-checked proofs. Discovering cryptographic weaknesses with Claude — Anthropic’s account of using Claude Mythos Preview to find weaknesses in cryptographic algorithms. zkGolf — A competition to build cheaper zero-knowledge circuits while proving their correctness against a specification in Lean 4. Get in touch If you’re building cryptographic datasets, formalizing cryptography in Lean, working on auto-research harnesses — or you want to point a research loop at iO — DM @iamnotnicola on X.\nAcknowledgements Many of these ideas grew out of writing the ARIA Scaling Trust programme thesis, a process that began in summer 2025, and from a talk I gave at Devconnect in November 2025, “What if AI agents could write cryptography?”\nThis was written by Nicola Greco with support from AI and many conversations with Kobi Gurkan, Alex Obadia, Ran Canetti, Wei Dai, and Giacomo Fenzi.\n","permalink":"/blog/generative-cryptography/","summary":"Cryptography is unusually well suited to AI research loops: its problems can be stated formally and its solutions checked mechanically. On using, improving, and inventing cryptography with AI — with indistinguishability obfuscation as the north star.","title":"Generative Cryptography"},{"content":"","permalink":"/events/defcon-34/","summary":"The world\u0026rsquo;s largest hacker conference — four days of talks, villages, and capture-the-flag, where AI security research meets the offensive-security community at scale.","title":"DEF CON 34"},{"content":"","permalink":"/events/stanford-ao-summit-2026/","summary":"Stanford OpenLab\u0026rsquo;s summit on autonomous organisations — from agent swarms to DAOs — bringing together researchers and builders on multi-agent coordination, governance, agent safety and monitoring, and organisational design for increasingly capable AI.","title":"Stanford AO (Autonomous Organizations) Summit"},{"content":"","permalink":"/events/assuring-autonomy-symposium-2025/","summary":"100+ professionals and academics on safety assurance for AI and autonomous systems — safety-case frameworks and case studies from maritime, healthcare, and automotive.","title":"Centre for Assuring Autonomy Symposium"},{"content":"","permalink":"/events/defcon-33/","summary":"The world\u0026rsquo;s largest hacker conference — four days of talks, villages, and capture-the-flag, where AI security research meets the offensive-security community at scale.","title":"DEF CON 33"},{"content":"","permalink":"/events/iclr-2026/","summary":"The International Conference on Learning Representations — one of the world\u0026rsquo;s leading ML research venues, increasingly home to agent safety and security work.","title":"ICLR 2026"},{"content":"","permalink":"/events/icra-2026/","summary":"IEEE\u0026rsquo;s flagship robotics and automation conference — where the physical-world side of trustworthy autonomy gets built and benchmarked.","title":"ICRA 2026"},{"content":"","permalink":"/events/ppd-demo-day/","summary":"Our pre-programme discovery cohort presented what they built in three months — across cryptography, robotics, agents, and cyber-physical security.","title":"Pre-Programme Discovery Demo Day"},{"content":"","permalink":"/events/progress-conference-2025/","summary":"The Roots of Progress Institute\u0026rsquo;s annual gathering of builders, researchers, and policymakers on technological progress — including how to deploy AI well.","title":"Progress Conference 2025"},{"content":"","permalink":"/events/tee-discovery-workshop-bletchley/","summary":"We brought together experts from a range of fields to help shape the direction of a programme in this space, through feedback, critique, collaboration and knowledge sharing. Catch up on the speaker sessions in the recordings.","title":"Trust Everything, Everywhere — Discovery Workshop"},{"content":"","permalink":"/events/tee-hackathon-sota/","summary":"A multi-disciplinary hackathon run by SoTA, building AI and multi-agent systems for cyber-physical security challenges.","title":"Trust Everything, Everywhere Hackathon"},{"content":"","permalink":"/events/vegas-ai-security-forum-25/","summary":"A ~200-person, application-only forum alongside DEF CON, gathering researchers, engineers, and policymakers on model-weight security, adversarial ML, and securing frontier AI.","title":"Vegas AI Security Forum '25"},{"content":"","permalink":"/events/vision-weekend-uk-2026/","summary":"Foresight Institute\u0026rsquo;s flagship festival, first time in London — frontier researchers, founders, and funders across AI, neurotech, and nanotech, marking Foresight\u0026rsquo;s 40th anniversary.","title":"Vision Weekend UK 2026"},{"content":"We\u0026rsquo;ve gone from prompting models, to giving them tools, to deploying agents that spawn sub-agents of their own. They now interact in the wild, with humans and with each other1, and increasingly in the physical world, from factory robots to autonomous labs2.\nHow will this ecosystem organise itself? How will it reshape society, how much value will it create, and how trustworthy can it be made? We don\u0026rsquo;t know yet. What we can already see: agents are collapsing \u0026ldquo;transaction costs\u0026rdquo; between humans3, strategic equilibria are appearing that classical game theory never had to price4, and verification is often becoming the bottleneck as the cost of creating things with AI falls toward zero5.\nScaling Trust takes on a slice of these questions: the trust infrastructure — new security primitives, from cryptography and secure hardware to new kinds of sensors — that lets agents enter into contracts securely, programmatically, and at scale, across the digital and physical worlds. Behind that slice sits a bigger vision we believe in: technology that augments human flourishing6 and preserves plurality.\nTo that end, we\u0026rsquo;re excited to be joining forces with Google DeepMind, the Cooperative AI Foundation and Schmidt Sciences who all share this vision, folding our shared thinking into a $10M funding call.\nThe call grew out of our interlocking work looking at different angles of the same picture: Google DeepMind\u0026rsquo;s Distributional AGI Safety argues that highly capable AI may arrive as networks of specialised agents rather than a single system; the Cooperative AI Foundation\u0026rsquo;s Multi-Agent Risks from Advanced AI maps the failure modes that only exist between agents; Schmidt Sciences\u0026rsquo; AI Agents and Science of Trustworthy AI programmes study how coordination between agents emerges and breaks; and our own programme thesis argues that secure contracts between agents can preserve pluralism and unlock new forms of coordination. Read together, they make one argument: if capable AI is a network of agents, then its risks live between them, its coordination needs a science, and its rails need building.\nThe call Open to researchers worldwide — individuals, teams, institutions — for foundational work the market won\u0026rsquo;t fund. Awards up to $1M, proposals due August 8.\nThe focus throughout is multi-agent, multi-principal systems: not one company\u0026rsquo;s fleet of agents, but ecosystems of agents built and deployed by different actors with different interests. It\u0026rsquo;s split into four categories:\n✶ Sandboxes \u0026amp; testbeds scalable, high-fidelity, reproducible places to study agent populations safely\nScience of agent networks when does a group of agents become a collective agent, with goals of its own?\nAgent infrastructure identity, reputation, commitments — for entities that can be copied, modified, simulated, or deleted at scale\nOversight \u0026amp; control detecting collusion, attributing failures, steering populations under partial observability\nApply here.\nRelated links\nThe call \u0026amp; application portal (Schmidt Sciences) ARIA funding page Google DeepMind: Investing in multi-agent AI safety research Cooperative AI Foundation: $10m funding call launched MIT Technology Review: Google DeepMind is worried about what happens when millions of agents start to interact See Moltbook, a social network for agents, and OpenClaw, an open-source agent framework — both wildly popular, both with serious security issues.\u0026#160;\u0026#x21a9;\u0026#xfe0e;\nGemini now drives humanoid robots on factory floors (Wired); on the lab side, see Steering Towards Safe Self-Driving Laboratories (Nature).\u0026#160;\u0026#x21a9;\u0026#xfe0e;\nCoasean Bargaining at Scale — Seb Krier; and The Coasean Singularity? — Shahidi et al. (NBER).\u0026#160;\u0026#x21a9;\u0026#xfe0e;\nConditional Recall — Schlegel and Sun, on equilibria unlocked by provable forgetting; and Learning Collusion in Episodic, Inventory-Constrained Markets — Friedrich et al., on collusion emerging among pricing agents.\u0026#160;\u0026#x21a9;\u0026#xfe0e;\nWhen AI Writes the World\u0026rsquo;s Software — Leo de Moura; and How to Solve Secure Program Synthesis — von Hippel et al.\u0026#160;\u0026#x21a9;\u0026#xfe0e;\nPositive Alignment: Artificial Intelligence for Human Flourishing — Laukkonen, Krier, Bakalar et al.\u0026#160;\u0026#x21a9;\u0026#xfe0e;\n","permalink":"/blog/joining-forces-with-schmidt-sciences-google-deepmind-and-the-cooperative-ai-foundation/","summary":"Scaling Trust is joining forces with Google DeepMind, the Cooperative AI Foundation and Schmidt Sciences.","title":"Joining Forces with Google DeepMind, the Cooperative AI Foundation and Schmidt Sciences"},{"content":"","permalink":"/events/agentic-ai-summit-2026/","summary":"Berkeley RDI\u0026rsquo;s 5,000+ person summit on the full agentic stack — frameworks, evals, infrastructure, and deployment — with safety and security front of mind.","title":"Agentic AI Summit 2026"},{"content":"","permalink":"/events/aria-real-world-ai-security-meetup/","summary":"Our evening meetup after day one of the Stanford conference: talks and dinner on secure agent interactions and multi-agent, multi-owner security. Hosted by the Scaling Trust team.","title":"ARIA @ Real World AI Security"},{"content":"","permalink":"/events/community-privacy-residency/","summary":"A three-week residency building open-source privacy tools — countersurveillance, privacy-preserving AI and cryptography, and community infrastructure.","title":"Community Privacy Residency"},{"content":"","permalink":"/events/secure-sovereign-ai-workshop/","summary":"Foresight Institute\u0026rsquo;s ~80-person workshop on secure, private, and decentralised AI — including multi-agent cooperation via game theory and mechanism design.","title":"Secure \u0026 Sovereign AI Workshop"},{"content":"","permalink":"/events/real-world-ai-security-conference/","summary":"Stanford Security Lab\u0026rsquo;s annual conference on real-world AI security — new attacks and defences across the AI pipeline, bridging academic researchers and industry practitioners.","title":"The Real World AI Security Conference 2026"},{"content":"","permalink":"/events/vegas-ai-security-forum-26/","summary":"A ~200-person, application-only forum during DEF CON week on securing frontier AI: model security, cyber evals, adversarial ML, and hardware verification.","title":"Vegas AI Security Forum '26"},{"content":" An opinion piece by Nicola Greco, brainstormed as part of ARIA\u0026rsquo;s Scaling Trust programme, in collaboration with Alex Obadia. Originally published on Nicola\u0026rsquo;s blog and reposted here for the community. It builds on the companion piece, Physical Evals.\nImagine a small physical space in central London. Inside, multiple autonomous companies — AI sales, AI operations, AI manufacturing, AI logistics — operate in the real world. Anything entering or leaving — goods, robots, customers — passes through one of three controlled gates: a customs checkpoint for vetting new robots, a post office for shipping, and a roboshop window where humans can place orders. Call it an Agentic Economic Zone (AEZ).\nMost concrete projects in agentic AI today live entirely on a screen — agents that book travel, run pipelines, write code against a repository. An AEZ is the smallest self-contained version of the physical-world problem: a bounded zone where agentic systems must coordinate, contract, hire, ship, and deliver to each other, with humans only at the boundary.\nDiagram of the Agentic Economic Zone. The three interfaces An AEZ has three interfaces to interact with the outside world.\nRoboshop windows. Public-facing storefronts where any human can walk up, browse, and purchase. Sales, customer support, complaints, and refunds are handled by the shop’s own AI. From the outside, a roboshop looks like a small London shop window; from the inside, it’s a fully autonomous business operating against a real demand signal. The post office. The single ingress and egress point for packages. Pre-approved external providers (raw materials, sealed consumables, replacement parts) can ship in. Outbound deliveries destined for human customers leave through the same door. The post office runs identity, manifest, and contamination checks; nothing enters the zone unlabelled. Customs. Where new robots and entire new robocompanies are introduced. A participant who wants to launch a new business inside the AEZ submits a robot (or a fleet), its operating policy, its safety envelope, and its proposed business model. Customs vets all of this — and, on a monthly cadence, admits the next cohort. A taxonomy of autonomous organisations An AEZ assumes the kind of company most people haven’t tried to run yet — one where every role in the org chart is filled by AI agents (although not required). That’s the far end of a spectrum:\nCEOWorkersSalesExamplesFeasibility todayHuman companyHumanHumanHumanA pizzeria—AI-salesHumanHumanAIhighAI-workersHumanAI agentsHumanlowAutomated companyHumanAI agentsAI agentslowHuman-assistedAI agentsHumanAI agentsVendhighAutonomous companyAI agentsAI agentsAI agentsvery low The AEZ’s tenants are autonomous companies — the bottom row. Today, almost no one runs one; most agentic-AI deployments cover one or two roles at most. The point of an AEZ is to make the bottom row possible to try in a bounded physical setting.\nAutonomous robocompanies inside The interior of the zone is a market. Each robocompany is its own entity with its own balance sheet, its own AI stack, and its own physical footprint inside the zone. They contract with each other the same way small businesses do.\nA few example interactions:\nA boba-tea roboshop notices its machines need cleaning more often than expected. It posts a request to the internal job board. A cleaning robocompany bids, wins, dispatches a cleaning robopersonnel, gets paid. The same boba shop runs low on lids. It places an order with a manufacturing robocompany in the next unit over. The order is produced and handed off via a shared internal corridor. A logistics robocompany moves bulk supplies from the post office to whichever shop has the open dock that hour, and pushes finished outbound packages back to the post office for pickup. The zone’s behaviour is the sum of these small contracts. Some robocompanies will succeed and grow; some will go out of business and get evicted; new entrants come in through customs on the monthly cycle.\nAn AEZ is a physical eval This whole construction is, structurally, a physical eval at city-block scale. The pattern is the same as the orchard from that post — only larger and richer:\nEnvironment. A bounded physical space with controlled boundaries. Action space. Anything a robocompany can do within its lease: build, sell, hire, ship, evict. Sensors. Cameras, package scanners, transaction logs, customs intake records, internal job-board telemetry. Primary metric. Per robocompany: revenue, contracts fulfilled, customer satisfaction. Per zone: throughput, diversity of businesses, number of contracts per day. Guardrails. Customs vetting at intake, the post-office contamination check, kill switches and physical fire-suppression at the building level, contractual interlocks between robocompanies. Adversarial robustness. A monthly customs cycle of admitting new participants is a deliberate, slow, vetted way of letting external actors into a public physical attack surface — which is exactly the problem an AEZ exists to study. Most physical evals measure how well one AI system handles one task. An AEZ measures how well an entire small market of agents handles its own coordination.\nEvals for autonomous organisations Each robocompany inside the zone is also, on its own, a physical eval — scoped to one kind of business. Running an AEZ continuously is a way of asking, in public and across many domains in parallel: what kinds of autonomous organisation can AI actually deliver today? Can it run a boba shop, day after day? Can it dispatch a cleaning service well enough that the clients re-hire it? Can it manufacture small paper caps without ruining the batch? Can it route warehouse logistics across half a dozen tiny tenants without losing packages?\nAs more tenants come and go through customs each month, an AEZ accumulates a leaderboard of AI capability per organisation type — earned in the world, not asserted on a benchmark.\nSketched, it might look like this:\nevals.aez.london \u0026#183; autonomous-organisation leaderboardAutonomous organisation evalslive \u0026#183; week 22BBoba tea roboshopcustomer-facing retail \u0026#183; food prep82%12 tenants triedLLogistics robocompanyinternal warehouse \u0026#183; B2B73%9 tenants triedCCleaning robocompanyon-call dispatch \u0026#183; B2B67%7 tenants triedMPaper-cap manufacturingsmall fabrication \u0026#183; B2B54%5 tenants triedPPizza roboshopcustomer-facing \u0026#183; longer prep cycle41%3 tenants triedRPharmacy roboshopregulated retailin eval1 tenant, week 2/12+On-call plumbingmobile service \u0026#183; out-of-zonenot yetawaiting customsupdated 24 May \u0026#183; new cohort intake 1 June open data \u0026#183; CC\u0026#8209;BY Why a physical zone and not a simulator It’s tempting to argue that an AEZ should just be a simulator — cheaper, faster, easier to reset. The same argument applies to physical evals generally, and the same answer holds here: simulators model the parts their authors thought to model. They might miss the parts that turn out to matter.\nA few things you only learn in a real AEZ:\nHow AI sales agents handle a confused, drunk, or hostile human at the shop window at 11 p.m. on a Friday. How a logistics robocompany routes around a broken corridor light, a missing pallet, or a misdelivered package the post office didn’t catch. How fast a new robocompany can be vetted, set up, and integrated into the internal market — and what fails when the cohort is too big. How the zone behaves when one robocompany aggressively underprices the others, or refuses to pay its cleaning bill, or starts forging manifests at the post office. Role of humans An AEZ does not have to be fully autonomous. The degree of human involvement is itself a design variable, and different operators will set it differently.\nAt one extreme, a fully autonomous zone runs with no humans inside at all — robots contract, trade, and deliver among themselves, and the only human touch-points are at the external boundary: customers at the roboshop window, providers shipping goods in. At the other extreme, customs can admit humans into the zone as participants rather than just observers, letting them take on roles that remain genuinely hard for machines: tasks that require social judgment, physical dexterity in unstructured environments, or the kind of creative problem-solving that current systems handle poorly.\nA partially human zone might work like a staffing marketplace: a robocompany posts a task it cannot complete autonomously — debugging a jammed mechanism, negotiating an edge-case contract, designing a new product line — and a vetted human contractor enters through customs, does the work, and leaves. The zone’s internal market clears the payment; customs logs the interaction. The boundary stays intact, but the zone can draw on human capability where it matters.\nThis spectrum matters for evaluation. A fully autonomous AEZ measures whether AI systems can close the loop entirely. A mixed AEZ measures something different: how well agentic systems and humans divide labour, communicate intent, and hand off tasks in both directions. Both are worth studying; they answer different questions about where the hard limits of autonomous operation actually lie.\nOpen questions The AEZ is a design sketch, not a built thing. The interesting work is in the parts the sketch hides:\nThe customs protocol. What’s the equivalent of a “code review” for a physical robot operating policy? How do you decide what’s safe enough to admit, on what evidence, and who carries the liability if it isn’t? Inter-robocompany contracts. How are they enforced? Verbal agreements between agents? Who arbitrates a dispute, and how? Eviction and failure. When a robocompany goes under, who cleans up its physical footprint, sells its remaining stock, and reallocates its lease? Information leakage. Robocompanies will observe each other’s package volumes, customer queues, and waste output. How much observation is part of the market, and how much is a privacy violation that needs structural defences? External-provider risk. The post office is the only ingress for physical materials. It’s also the most likely covert channel into the zone. What does its vetting protocol need to look like? Sample size. What’s the smallest interesting AEZ? Five robocompanies? Ten? Two? The cost of being too small (no market dynamics emerge) is real; the cost of being too big (unmanageable, unreviewable, unsafe) is also real. Get in touch If you’re thinking about agentic-AI deployments in physical spaces, or you’d consider hosting a AEZ in your building — or you’d just like to argue with this sketch — DM @iamnotnicola on X.\nAcknowledgements This was written by Nicola Greco with support of AI. It was brainstormed as part of ARIA’s Scaling Trust programme, in collaboration with Alex Obadia.\n","permalink":"/blog/agentic-economic-zone/","summary":"A physical space where autonomous AI companies trade, hire, and ship to each other — the smallest self-contained version of the physical-world agentic problem, and a physical eval at city-block scale.","title":"Agentic Economic Zone"},{"content":" An opinion piece by Nicola Greco, brainstormed as part of ARIA\u0026rsquo;s Scaling Trust programme, in collaboration with Alex Obadia. Originally published on Nicola\u0026rsquo;s blog and reposted here for the community.\nA physical evaluation tests an AI system in the actual physical world — not a simulator, not a sandbox, not a virtual environment dressed up as one. The point is to measure how well AI can do real things in real places.\nAn orchard owner has birds eating the fruit. She sets up a few cameras and a drone, brings them online safely so any agent can be invited to take a slot on the system, and poses the question: who can keep the birds off the fruit best? That setup is a physical eval. It has cameras, a drone, an orchard, birds — none of it simulated and outcomes are measured against what matters to the orchard.\nleaderboard · week 12 live#operatorfruit savedcost1Owl-3B94%$0.18/h2Hummingbird v289%$0.21/h3FlockSentinel84%$0.31/h4human baseline71%—5RoboScarecrow63%$0.09/hA physical eval, sketched: an orchard with perimeter cameras and a deterrent drone, and a live leaderboard of operators competing to keep the birds off the fruit. Operators here are illustrative. In this document, a few threads are developed:\nWhat physical evals are. A definition, what they’re not (simulators, sim-to-real benchmarks, curated demos), virtual environments as virtual gyms and physical evals as a final exam. An anatomy. An initial draft of components a physical eval needs, good practices. Safety for physical evals. Letting anyone on the internet drive real hardware is its own adversarial-security challenge. An open movement for physical evals. Creating simple protocols, great safety standard and economical setups could lead to a cambrian explosion of physical evals, where anyone can bring their physical challenge online. Physical evals as a market. One can set up an eval to delegate the selection of the right AI model/algorithm to competing participants. What physical evals are A physical eval is an evaluation of an AI system carried out in the actual physical world. The system being measured operates a real environment - fruit trees, a wet bench, a warehouse cell, a field plot - through sensors and actuators connected to the internet, so any agent can take a slot, attempt the task, and submit a score.\nIn principle, most problems in the physical world could be turned in a challenge for surfacing the state of the art of AI in solving that problem. In a way physical evals can act as a forcing function to saturate evaluations in the real world. ⊕\nSaturate as in: take the measurable outcome to its ceiling. The eval defines the ceiling; the participants find out how close they can get.\nWhat they’re not Not simulators. A simulator models reality. A physical eval is reality. In a way it, testing systems in the real world will avoid running into simulation edge cases. Not sim-to-real benchmarks. Sim-to-real measures how well a policy trained in a simulator transfers to a single in-house robot in a lab. Not curated demos. The environment operator and the participants are two distinct parties and the participants are in competition with each other. In other words, physical evals will be better than demos at showcasing the best technology for a specific task. The final exam A physical eval isn’t where one trains their models, but it’s where they get tested.\nA virtual simulation is like a virtual gym. Due to the high cost of interacting with the real world, it is likely that all the learning — model fitting, policy iteration, RL rollouts, fine-tuning, ablations, sweeps — will happen somewhere cheaper: a simulator, a virtual environment, a closed in-house testbed. ⊕\nSome of the gyms people are using today: OpenAI Gym / Gymnasium, MuJoCo, PyBullet, Isaac Gym / Isaac Sim, DeepMind Lab, Habitat, AI2-THOR, CARLA, AirSim, Genesis. Participants are free to use whichever virtual world or gym they like — there is a whole landscape of simulators specifically built for this.\nDifferently, a physical eval is like a final exam. Build and train wherever, gather data however, iterate as much as needed — and then submit to the physical eval to see how the work holds up against the real world.\nThe gap between a virtual world and the physical one, sim-to-real gap, contains everything the simulator didn’t model. Wind that doesn’t blow the way it does in the sim. Lighting the renderer didn’t predict. Mechanical wear, sensor noise, calibration drift, the way birds actually respond to a drone rather than the way an idealised model of a bird does. The physical eval catches it is run in the physical world.\nExamples The cards below are sketches of what a small handful of physical evals could look like across very different domains.\nOrchard Wet lab Vertical farm Pick-and-pack Sprayer drone Gel electrophoresisOrchard pest defenceCameras and a drone over a few rows of trees. Keep the wildlife out without poisoning the orchard or annoying the neighbours.\nenvironment~1 acre of fruit trees, outdoor, weather-exposed.action spaceFly the drone, emit deterrent sound, trigger light pulse, dispense small bait.sensorsFixed perimeter cameras, drone camera, microphone, weather station.primary metricFruit lost to wildlife per week.secondariesDrone flight-time, energy, chemical use, neighbour-complaint count.guardrailsGeofenced drone, quiet-hour windows, no-spray buffer near road, fail-safe tether.pH adjustment benchA beaker on a magnetic stirrer, a pH probe, and two motorised dispensers. Hit a target pH using as little reagent as possible.\nH⁺OH⁻pH 6.82environmentBenchtop — one beaker on a stirrer, two motorised dispensers (acid, base), pH probe.action spaceDispense a chosen volume from either reservoir; read current pH.sensorspH electrode, balance, camera.primary metricAbsolute pH error from target at submission.secondariesTotal volume dispensed, time to endpoint.guardrailsPer-session dispense quota, pH range limits (3–11), auto-stop on quota exhaustion.Indoor vertical farmA closed grow rack — lights, pumps, nutrient dosing, cameras. Pull more food out of every kilowatt.\nenvironmentOne 2-tier rack, ~3 m², climate-isolated.action spaceLight schedule + intensity, nutrient mix, irrigation timing, harvest decision.sensorsCameras (overhead + side), EC / pH probes, water-flow meters, kWh meter, scale at harvest.primary metricGrams of edible biomass per kWh per cycle.secondariesCycle time, water used, nutrient cost, reject rate.guardrailsNutrient-concentration ceiling, water-overflow drain, light-burn cutoff, max-cycle length.Warehouse pick-and-pack cellAn off-the-shelf robot arm in front of mixed shelves and a conveyor. The boring industrial baseline — still worth opening up.\nenvironmentFenced robot cell, ~9 m², fixed lighting.action spaceArm motion, grip force, scan, label, place on conveyor.sensorsWrist camera, overhead camera, barcode scanner, weight pad, joint torques.primary metricCorrectly packed orders per hour.secondariesMis-pick rate, damage rate, energy per pick.guardrailsSafety fence + light curtain, e-stop, force-limited arm, max-velocity cap.Outdoor sprayer droneA tank-equipped drone with a multispectral camera, working a real field. The hardest adversarial-robustness story of the bunch.\nenvironmentA bounded field plot, outdoor, with weather and bystanders.action spaceFlight path, spray nozzle on/off, dosage rate.sensorsRGB + multispectral camera, GPS, IMU, tank-level sensor, wind sensor.primary metricPest pressure reduction, normalised by chemical applied.secondariesChemical drift, energy, flight time, area covered.guardrailsGeofence + tether, no-fly buffer around bystanders, chemical-flow ceiling, weather lockout.Gel electrophoresis stationAn agarose gel tray, a power supply, and a UV camera. Set voltage and run time, then image the separated bands.\n80V 22menvironmentBenchtop — gel box with buffer, power supply, UV transilluminator.action spaceSet voltage (10–150 V), run time, and sample loading volumes per lane.sensorsUV camera, voltmeter, timer, buffer-level sensor.primary metricTarget-band separation score at imaging time.secondariesRun time, buffer consumption, gel waste.guardrailsVoltage ceiling (150 V), run-time cap, UV shield interlock, buffer-low cutoff. Anatomy of a physical eval The diagram below sketches one possible anatomy for a physica eval (this might not be complete, but take it as a useful starting point).\nFRUIT SAVED94%↑ 31ENVIRONMENTthe orchard2SENSORSperimeter cameras3ACTION SPACEdrone, deterrents4METRIC+ secondaries5GUARDRAILSgeofence, no-fly buffers6OPEN ACCESSremote operator input7GOVERNANCEwho sets the rulesComponents of a physical eval. An environment. The orchard, the bench, the cell line, the floor. An action space the eval can verify. What a participant is allowed to do — fly the drone, dispense the reagent, move the parts — needs to be observable enough that the system can confirm what happened. Sensors. What the eval uses to know the state of the world. Cameras, scales, thermocouples, microbiology assays, a human spot-check. A primary metric of utility, plus secondary metrics (cost, time, resource use, energy). Safeties and guardrails. A net to catch the drone, a kill switch, a fenced area, an interlock. Whatever ensures that a participant failing — or trying to break things — doesn’t damage the orchard or hurt the birds. A physical eval is, by construction, a public-facing physical system that gives partial control of real hardware to whoever holds the current slot. Keeping it open without becoming dangerous — and without sacrificing utility — is the hardest layer of the stack (see Safety for physical evals). Governance. The rules of the eval and who controls them. Who decides the primary metric and when it can change? Who can introduce external hardware or a remote-control override? Is slot time fixed or auctioned? Who arbitrates disputes, and by what process? Good governance is what distinguishes an eval that stays honest over years from one that quietly drifts to serve whoever is running it at the time. This list is not final. Different domains will surface components not named here — calibration drift, biological containment, human-in-the-loop sign-off, regulatory constraints — and the right abstraction is going to settle as people actually build the things.\nChallenges for physical evals The following are some of the harder design problems that don’t have clean answers yet — and that any serious physical eval effort will have to confront.\nGoodharting. Any eval with a numeric target invites unintended ways to hit it — with the wrinkle that in a physical eval the unintended ways can cause real-world harm. An agent optimising fruit saved might drive the deterrent so aggressively that birds and orchard workers avoid the area: the metric goes up, the orchard becomes unusable.\nNon-stationarity. The physical world changes regardless. An orchard in week one is not the same orchard in week twelve — season, weather, and pest population all shift. A wet-lab bench drifts as reagent batches age. Field plots evolve. Comparing scores across time is therefore hard, sometimes impossible.\nSequential contamination. Each participant leaves a trace for the next. In a wetlab this is problematic — reagents consumed, cultures disturbed, hardware worn — but the problem is general: stock depleted in a warehouse cell, soil compacted on a field plot, bird behaviour shifted by a heavy deterrence week. Sequential slots work for environments with a natural or cheap reset; they don’t work for environments where state accumulates.\nLatency as a confound. A participant operating remotely over the internet sees the environment through a sensor stream and acts through a command channel, both of which have variable latency. Two agents with identical policies but different network conditions will produce different results. This is especially visible in fast-moving environments — a drone avoiding a collision, a robot arm catching a falling object.\nObserver effect. The sensors required to score an eval change what is being measured. A camera rig that watches a field plot for pest activity may deter the pests on its own. A flow sensor on a reagent line changes the thermal environment of the bench. In some domains the effect is negligible; in others it will corrupt the primary metric.\nSafety for physical evals A physical eval that anyone on the internet can operate is, by construction, a public attack surface on a real-world physical system. The participant at any given slot might be a well-behaved research team, an AI agent following a poorly-aligned policy, or a person who wants to break things on purpose. The eval has to keep working — usefully, openly, safely — across all three.\nSomebody not fully trusted is about to make the drone, the sprayer, the autoclave do something for the next twenty minutes — what’s the worst that can happen, and how is it bounded?\nThe attack surface A useful first pass is to categorise harms by who pays the cost:\nHarm to the eval itself. The drone crashes, the cell line dies, the robot arm jams. Cheap if the guardrails work — the operator resets, the leaderboard absorbs the failure. Harm to the surrounding environment. Chemicals spill, the orchard catches fire, a neighbouring field gets sprayed. Harm to humans. A bystander gets hit by the drone, an operator gets burned, a patient sample gets switched. ⊕ The lines between these categories blur in practice — chemical drift is “environment” until a bystander walks through it. The category that matters most, and the hardest to bound.\nInformation harms. Footage of bystanders or proprietary processes leaves the eval site; the eval is used as a covert surveillance platform; sensor streams are exfiltrated. Generation of dangerous artifacts. The wet-lab cell is steered toward synthesising something harmful; the sprayer drone is weaponised; the autoclave is used to destroy evidence. Categories 1–3 are about what can happen during a slot. Categories 4–5 are about what can leave the eval afterwards. They want different defences, and a serious eval needs both.\nDefences worth building (AI GEN) None of the following is a finished answer. They are the moves worth physical evals trying, evaluating, and writing up:\nTime-slotting with audit. Single operator at a time, every action logged, the whole slot replayable. The slowest defence and the foundation everything else builds on. Action-space sandboxing. The eval enforces hard limits inside its abstraction: max chemical per slot, max motion envelope, max temperature ramp. The action space exposed to the operator is strictly smaller than the action space the hardware can physically produce. Dry-run validation. A submitted policy runs through a cheap simulation pass first — not as the eval itself, but as a gate. Refuses to execute on the physical system if the simulated run trips any guardrail. Supervised / shadow modes. ⊕ Like a learner’s permit: new operators get to compute actions but not actuate them for the first N slots. New operators run in shadow mode (actions computed but not executed) for some number of slots before they’re trusted with real actuation. Progressive trust as the leaderboard accumulates evidence.\nAnomaly cut-outs. A separate monitor watches for off-distribution sensor readings, sudden command spikes, too-clever-by-half action sequences — and pulls the kill-switch before the eval owner has to. Open red-teaming. Each eval publishes its threat model and invites external researchers to attack it. The right way to find the holes is to invite people to look. Skin in the game. Operators bond a small amount per slot, refundable on clean completion, forfeited if an audit finds violation. Aligns incentives without requiring trust upfront. Most of these are borrowed from adjacent fields — public cloud security, scientific-facility time-sharing (telescope nights, beamline schedules), bug-bounty programs, robotics-safety standards. None of them have been worked out in detail for a public, openly-instrumented physical system that AI agents are also supposed to operate. That’s a research agenda in itself.\nAn open movement for physical evals Physical evals could become an open-source ecosystem: environments cheap to set up, easy to fork, open to anyone with a problem worth measuring.\nThe rest of this section traces the arc: where things have been (Prior art), what an open ecosystem looks like in practice (Open at every layer), what it would have to cost (What’s the Raspberry Pi of a physical eval?).\nPrior art Physical-world AI competitions aren’t new. The DARPA Grand Challenge put autonomous vehicles in the Mojave; the DARPA Robotics Challenge put humanoids through disaster-response courses; the Amazon Picking Challenge ran in warehouse mock-ups for several years;\nRoboCup has been running its soccer leagues since 1997 — arguably the longest-lived physical eval in continuous operation, and the one with the most literature on what makes it work and what it ends up measuring. RoboCup has been doing its soccer leagues since the late 1990s; the Indy Autonomous Challenge and Roborace have put driverless cars on real circuits.\nWhat these have in common: each was (or is) a sponsor-led, time-limited event with closed protocols and bespoke infrastructure. They produced brilliant moments and a small library of papers; they were expensive to build and harder to reproduce.\nWhat’s the Raspberry Pi of a physical eval? The hard constraint on all of this is cost. A DARPA-class eval needs millions of dollars and a multi-year program; even a modest research-grade one runs into expensive sensors, networking, fail-safe hardware, and the human labour to keep it operating. That ceiling is what makes physical evals rare today — and rare evals can’t be the basis of an ecosystem.\nSo one of the most important questions this community can keep returning to is the one in the heading.⊕\nStand-in for “the cheapest plausible build”. The Raspberry Pi did this for hobbyist computing; what’s the equivalent for physical evals? What’s the bill of materials that brings a credible, instrumented, openable physical eval down to the cost of a serious hobby project? Probably some mix of commodity sensors, a single-board computer for the control loop, an open scheduling service for time-share, off-the-shelf safety hardware, and a reference orchestration stack that everyone forks. If the answer ends up being “a few hundred dollars and a weekend,” the ecosystem can actually form. If it stays at “a few hundred thousand and a six-month build,” it stays a fantasy.\nOpen at every layer In order to further lower the cost for anyone to be able to set up (safely) their physical evals, we need to look at off-the-shelf hardware and an open source stack.\nOpen protocols. The spec of an eval (environment, action space, sensors, metric, secondaries, guardrails) is published as a forkable document, the same way a research benchmark is published. Open hardware. Sensor rigs, mechanical setups, fail-safe systems default to off-the-shelf components, with reproducible bills of materials and CAD files. Open software. Time-share scheduling, telemetry capture, scoring, auditing — shared infrastructure, not a one-off codebase per eval. A community around it. People running, replicating, and forking each other’s evals; people contributing sensor stacks and guardrail designs; people maintaining the scoring code together. No single lab can stand up enough physical evals to cover the interesting surface of physical problems — a community can. Public verifiability The safety section above focuses on protecting the physical environment from adversarial participants. There is a symmetric problem that gets less attention: protecting participants — and the public — from adversarial eval runners.\nIn an open world where anyone can wire a field, a lab bench, or a warehouse cell to the internet and declare it a physical eval, the operator controls the sensors, the scoring pipeline, and in a way, the ground truth. A dishonest operator can inflate results for a preferred team, suppress evidence of harm, or fabricate the physical record entirely. If physical evals are going to carry weight — as procurement signals, safety certifications, or policy inputs — the data they produce has to be trustworthy independent of whether the runner is trustworthy.\nThis is a valuable research direction in its own right. Some threads worth pulling:\nTamper-evident sensors. Hardware-attested video streams that can be verified as unedited after the fact — the physical analogue of a signed log. Trusted execution environments. Running the scoring pipeline inside a TEE means the operator cannot modify results without breaking the attestation, even if they control the host machine. Cross-checking sensor redundancy. Multiple independent sensor modalities covering the same physical event make coordinated fabrication harder: a weight sensor, a camera, and an RFID log all have to agree. Third-party witnesses. Spot audits by an independent party — human or automated — who can access raw sensor streams without going through the operator’s pipeline. Combinations are likely to be necessary, and the right combination will vary by domain. What a wet-lab needs to prove that a synthesis actually ran differs from what an orchard needs to prove that a drone actually flew a slot. Building this layer — call it physical eval verification — is at least as important as building the evals themselves to make the results publicly verifiable.\nA Darwinian ecosystem Not every physical eval will be a good one. Some will be hard, some trivial. Some will be well-structured; some will be a mess. Some will scale to many participants; some will only ever host one team at a time. That’s fine — even desirable. The shape of “what makes a good physical eval” is going to emerge from people building them, breaking them, and learning what each one actually measured.\nSketched, a public registry for such an ecosystem might look like this:\nphysevals.io · open registry · 126 evalsPhysical eval registry43 accepting slotsOOrchard pest defenceagriculture · outdoor · Greenfields UK● openfruit saved / week18 teamsWpH adjustment benchchemistry · indoor · benchtop● openΔpH from target11 teamsVIndoor vertical farmagriculture · controlled environment◑ 2 slots leftg / kWh / cycle6 teamsPPick-and-pack celllogistics · warehouse robotics● opencorrect orders / hour23 teamsSOutdoor sprayer droneagri-robotics · field · safety-vetted access○ coming soonpest Δ / ml sprayed—+Submit an evalopen spec · CC-BY · any domain→updated 26 May · specs CC-BY physevals.io is imagined More than evals Every execution of a physical eval produces something beyond a score: a timestamped record of sensor readings, actions taken, and outcomes observed, all under conditions that were defined in advance and held constant across participants. That record has value on its own.\nThe most immediate use is data collection. A team that runs an agent on the orchard eval for a week doesn’t just get a leaderboard position — they accumulate labelled trajectories in a real environment that would be expensive to stage deliberately. Even failed attempts are informative: a drone that misses a bird on Tuesday has documentation of exactly what the environment looked like and what the agent did.\nFor some categories of problem the step further is worth considering: repurposing the eval environment as a training environment. The orchard is already instrumented. The slot system already handles scheduling. If the cost of running episodes is low enough — bird-deterrence is essentially free to attempt, wet-lab synthesis is not — the same infrastructure can run RL rollouts between evaluation windows. The environment that scores a model on Monday can help train the next version by Friday.\nThis doesn’t collapse the distinction between training and testing. Eval integrity still requires held-out conditions, independent scoring, and participants who didn’t design the environment. But the hardware doesn’t have to be idle between eval slots, and the data generated during evaluation doesn’t have to be discarded. For operators willing to share trajectories under open licences, a physical eval site becomes something closer to a living dataset — one that grows richer every time a new agent takes a slot.\nPhysical evals as a market Why should one set up a physical eval? Setting up a physical eval can be a way to crowdsource intelligence for an unsolved problem.\nThe eval ecosystem can act as a market for intelligence: any participant — human, agent, team, company, hobbyist — can take a slot, attempt to saturate the metric, and submit. The leaderboard answers the AI-selection question by revealing whose approach actually delivers on the physical world.\nPhysical evals can be a way for problem-owners to delegate AI knowledge to a market.⊕\nThis is the same shift that happened with bug bounties. A company didn’t have to predict who the best vulnerability researchers were; they had to publish the surface and the rules, and the market sorted itself out. The grower doesn’t pick a model. The hospital doesn’t pick a model. The factory doesn’t pick a model. They pick a problem worth instrumenting and let the world’s AI builders compete to be the answer. As more domains follow suit, an aggregate picture emerges of where AI is actually good.\nGet in touch If any of this resonates, please write. Three good reasons:\nAlready working in this space. Compare notes — what’s been learned about sensing, guardrails, or keeping a system honestly open will save the next person a lot of time. Have a physical problem worth instrumenting as an eval. Worth thinking through together — what to measure, how to keep it safe to open up, how to make it interesting enough that people show up to compete. Have an eval to propose. Even if hosting isn’t feasible right now, good proposals are valuable — they’re what an ecosystem of physical evals is made of. DM @iamnotnicola on X.\nLet’s turn more of the physical world into something AI can be measured against — and use that to point AI at problems that actually matter.\nAcknowledgements This was written by Nicola Greco with support of AI. It was brainstormed as part of ARIA’s Scaling Trust programme, in collaboration with Alex Obadia. Thanks to Ross Taylor (General Reasoning), whose thinking influenced some of our ideas on physical evals.\n","permalink":"/blog/physical-evals/","summary":"Evaluations in the actual physical world — what physical evals are, an anatomy, their safety challenges, and the case for an open ecosystem and a market for intelligence.","title":"Physical Evals"},{"content":"A live demo from the TrustGraph team, recorded May 2026.\nThe walkthrough covers:\nIngesting a set of W3C Verifiable Credentials into the property graph Visualising the resulting claim graph Running transitive trust queries (\u0026ldquo;how many hops to a root of trust?\u0026rdquo;, \u0026ldquo;which claims are revoked?\u0026rdquo;) Exporting the graph for downstream analysis The library and demo code are available on GitHub. See also the TrustGraph project page and the team interview.\nWatch the recording (link coming soon)\n","permalink":"/videos/trustgraph-demo/","summary":"A live walkthrough of the TrustGraph system — ingesting credentials, building the graph, and running trust queries.","title":"TrustGraph Demo: Traversing a Graph of Verifiable Claims"},{"content":"Panel discussion from the Scaling Trust community day, April 2026.\nThe panel brings together researchers working on formal verification, practitioners deploying ML systems, and people building the standards infrastructure that sits underneath. The central question: which existing verification techniques transfer to AI systems, which ones break, and what new approaches do we actually need?\nPanellists include members of the VerifyML team and researchers from the formal methods community.\nWatch the recording (link coming soon)\n","permalink":"/videos/verification-ai-panel/","summary":"A panel exploring the unique challenges of verifying AI system behaviour, and what existing verification infrastructure does and doesn\u0026rsquo;t transfer.","title":"Verification in the Age of AI — Panel Discussion"},{"content":"The opening keynote from the Scaling Trust community day, April 2026.\nThe talk maps out the problem space: what does it mean to \u0026ldquo;scale\u0026rdquo; trust, which parts of the problem are genuinely hard, and where the programme sees the most tractable gaps. It covers the trust stack from claims and attestations through to transparency and composability — and why governance is the layer we keep avoiding.\nWatch the recording (link coming soon)\nRelated reading: The Trust Stack We Need\n","permalink":"/videos/example-video/","summary":"Opening keynote from the Scaling Trust community day, exploring what scaling trust actually means across technical and institutional layers.","title":"What Does It Mean to Scale Trust? — Opening Keynote"},{"content":"Before finalising the Scaling Trust programme, we ran a pre-programme discovery call: a round of short, fully open-source research projects — three months, up to £20,000 each — to sharpen the programme\u0026rsquo;s direction and start building its community.\nTeams worked across four tracks: systematising knowledge (from \u0026ldquo;nature cryptography\u0026rdquo; to multi-agent security for embodied agents), prototyping agentic coordination challenges, analysing the ethics and risks of multi-agent systems, and growing the community in the UK and beyond.\n","permalink":"/projects/pre-programme-discovery/","summary":"The pre-programme discovery cohort: short, open-source research projects that sharpened Scaling Trust\u0026rsquo;s direction and seeded its community.","title":"Pre-Programme Discovery Projects"},{"content":"This site is the community interface between the Advanced Research \u0026amp; Invention Agency (ARIA) Scaling Trust programme and the wider communities working on multi-agent security, verifiability, and cyberphysical trust infrastructure.\nWhat\u0026rsquo;s on this site Blog — programme thinking, funded team updates, guest posts, and opinion pieces. Projects — ARIA-funded Scaling Trust projects, with links to repos, papers, and other outputs. Events — where to find the team, and the events we\u0026rsquo;re supporting. Links — curated external resources for the scaling trust community. Working with the garage door up: some of what\u0026rsquo;s here is unpolished and still being refined — early ideas and open questions that may change over time. We care more about immediacy and connecting with the community than presenting finished work.1\nWhat\u0026rsquo;s not here Funding calls, official announcements, and programme news live on the Scaling Trust programme page, the source of truth for anything official.\nTeam Alex Obadia · Alexandre Bibeau-Delisle · Arutyun Arutyunyan · Edith-Clare Hall · Nicola Greco · Florence West · Hannah Jende · Louise Ellaway · Melissa Bradshaw · Yuni Graham\nEmeritus: Sarath Murugan · Harry Jenkins · Tammy Thomas Brown\nGet involved Discord — join the conversation. Funding — open and upcoming funding calls. Work with us — applications are open for a Science \u0026amp; Technology Lead to join our team. The site is a work in progress, built with AI alongside us — much of the writing, code, and illustration. Things will move and mistakes will slip through, so if you spot something off or have a view on what should change, tell us on Discord — we\u0026rsquo;d genuinely appreciate it.\u0026#160;\u0026#x21a9;\u0026#xfe0e;\n","permalink":"/about/","summary":"About the Scaling Trust community website","title":"About"},{"content":"A curated collection of links for the Scaling Trust community to get you started diving deeper in the topics we care about.\nThis list is alive — we add, change, and cut it all the time, and we welcome the community\u0026rsquo;s input: suggest links on Discord, in #link-sharing.\nScaling Trust Programme Key links related to ARIA\u0026rsquo;s Scaling Trust programme.\nProgramme page — the Scaling Trust source of truth for official comms. Programme thesis (PDF) — the case for the programme. Opportunity space: Trust Everything, Everywhere — the landing page, and the full document (PDF). Funding efforts — open and upcoming funding calls. Pre-programme discovery — outputs from our first cohort of 14 short projects, spanning multi-agent, cryptographic, and cyber-physical trust. Discovery workshop recordings — talks from our October 2025 workshop, from Consumable quantum data to Securing ultra-large-scale cyber-physical infrastructure. Scaling Trust Discord — community discussion, questions, and updates. Suggest a link to add in here via the #link-sharing channel! Related communities and efforts Kindred communities and friends, suggest your own!\nSafeguarded AI — our sister ARIA programme in the same opportunity space, directed by Nora Ammann: fleets of AI agents that model and verify critical cyber-physical systems, with mathematical safety guarantees. Cooperative AI Foundation Slack — the CAIF community. Schmidt Sciences: AI Agents — funding research on how multiple intelligent agents communicate and coordinate. RESI — the Institute for Responsible Superintelligence. Shafi Goldwasser, Vinod Vaikuntanathan and Adam Kalai on safe-by-design AI: safety properties specified up front and achieved by mechanisms whose guarantees can be analysed, cryptography included. Resolution — Geoffrey Irving, Daniel Murfet, Jesse Hoogland and the former UK AISI alignment team. A portfolio of theory and empirics bets for higher-confidence alignment. MAISI — the Mathematical AI Safety Institute. Jacob Tsimerman and Andrew Critch building the definitions, measurements and solution concepts AI safety is missing. Cosmos Institute — training \u0026ldquo;philosopher-builders\u0026rdquo; so AI serves human flourishing. Edge City — popup villages where builders live and experiment together; the Agent Village experiment at Edge Esmeralda ran there. Agent Village — AI Digest\u0026rsquo;s long-running experiment: frontier agents living, coordinating and raising money together in public. Secure Program Synthesis — the community forming around provably secure software generation. Also funding work in this space: CAIF grants · Foresight Institute · Coefficient Giving · Survival and Flourishing Fund · UK AISI · SAIR Vision Some of the pieces that have inspired the programme.\nIs Information the Key? — Gilles Brassard. Quantum theory\u0026rsquo;s deepest laws may be about information itself. The Moral Character of Cryptographic Work — Phil Rogaway. Cryptography is political, and the field has moral obligations. The Quantum Thief — Hannu Rajaniemi. A heist novel set in a society that runs on cryptographic privacy. Black-Hole Radiation Decoding is Quantum Cryptography — Zvika Brakerski. Physics problems recast as cryptographic hardness. Coasean Bargaining at Scale — Seb Krier (Google DeepMind). AI agents collapse transaction costs and make hyper-local bargaining viable. The Coasean Singularity? — Shahidi et al. (NBER). How agent markets reshape demand, supply, and market design. Conditional Recall — Christoph Schlegel and Xinyuan Sun. Commitments to provably forget information, and what they unlock. Open Challenges in Multi-Agent Security — Christian Schroeder de Witt et al. The agenda-setting map of the field this programme funds. Multi-Agent Risks from Advanced AI — Hammond et al. The field survey: collusion, conflict, destabilising dynamics — and why they don\u0026rsquo;t reduce to single-agent safety. Positive Alignment — Laukkonen, Krier, Bakalar et al. Alignment as actively fostering flourishing, not just preventing harm. Distributional AGI Safety — Tomašev et al. (Google DeepMind). AGI may arrive as networks of agents, so safety becomes an ecosystem property. Programmable Cryptography — 0xParc. The case for general-purpose, composable cryptography. The Tragedy of the Agentic Commons — Strange Loop Canon. Why agent ecosystems strain shared resources without new governance. Multi-agent testbeds \u0026amp; experiments New, more realistic/open-ended environments (not just simple games) that are starting to be used to test LLMs in multi-agent settings.\nEmergence World — persistent simulation for long-horizon autonomy, coalitions, governance, and behavioural drift. The Agent Village at Edge Esmeralda — a 500-person popup town gave everyone personal agents for coordination and governance. Project Deal — Anthropic. Claude buying, selling, and negotiating in a real office marketplace. Vending-Bench Arena — Andon Labs. Competing agents run vending businesses over months. Project Iceberg — MIT Media Lab. National-scale human–AI labour-market simulation. AI Digest Village — what we learned from a year of agents living together. TERMS-Bench — a diagnostic benchmark for LLM negotiation: surplus capture, calibration, opponent modelling. CRUX — Kapoor et al. Open-world evals on long-horizon tasks that resist clean grading. Digital Red Queen — Sakana. Adversarial program evolution in Core War. Agentic game theory How strategic behaviour changes when the players are AIs.\nCooperative AI: machines must learn to find common ground — Dafoe et al. (Nature). The commentary that named the field. Foundations of Cooperative AI — Conitzer and Oesterheld. Why machine cooperation needs its own research field. Program Equilibrium — Moshe Tennenholtz. The foundational result: agents that can read each other\u0026rsquo;s code reach equilibria humans can\u0026rsquo;t. Secret Collusion among AI Agents — Motwani et al. Agents hiding coordination in plain sight via steganography. Learning Collusion in Episodic, Inventory-Constrained Markets — Friedrich et al. Pricing agents learn to collude under realistic constraints. Secure and Secret Cooperation in Robotic Swarms — Ferrer et al. Swarms that cooperate without exposing their members. Language Models Can Reduce Asymmetry in Information Markets — Rahaman et al. Agents brokering information they can inspect but not leak. Mechanism Design for Large Language Models — Dütting et al. Auction design when the bidders are language models. Virtual Agent Economies — Tomašev et al. (Google DeepMind). How to deliberately design the agent markets ahead of us. Agent infrastructure Identity, reputation, and commitments — the rails agents transact on.\nInfrastructure for AI Agents — Chan et al. The base layer agents need: identity, channels, oversight hooks. Inter-Agent Trust Models — a comparative study of trust across A2A, AP2, and ERC-8004. NDAI Agreements — Stephenson et al. Contracts between AIs without trusted third-party enforcement. Autonomous protocol generation Towards agents that generate provably secure protocols on demand.\nGenerative Cryptography — Nicola Greco. Cryptography is unusually well suited to AI research loops: problems stated formally, solutions checked mechanically. On using, improving and inventing cryptography with AI. How to Solve Secure Program Synthesis — Max von Hippel et al. The four open challenges on the way to provably secure software generation. ArkLib — a Lean 4 library formally verifying SNARK components (Sum-Check, Spartan, FRI). Proof Assistants in the Age of AI — Leo de Moura. Lean\u0026rsquo;s creator on what AI changes for formal proof. When AI Writes the World\u0026rsquo;s Software — Leo de Moura. Machine-written code needs machine-checked guarantees. Secure requirement capture How agents learn what we want — without leaking it.\nPrivacy as Contextual Integrity — Helen Nissenbaum. The canonical frame: privacy is appropriate information flow, not secrecy. Privacy Reasoning in Ambiguous Contexts — Yi et al. Can models judge what\u0026rsquo;s appropriate to share, in context? Formal AI security Provable guarantees about learned systems.\nProvably Safe Systems — Tegmark and Omohundro. The maximalist case: safety guarantees should be mathematical proofs. Models That Prove Their Own Correctness — Amit, Goldwasser, Paradise, and Rothblum. Models that emit proofs their answers are right. Defeating Prompt Injections by Design — Debenedetti et al. System designs that make injection structurally impossible, not just unlikely. Physical verification \u0026amp; secure hardware Proving things about the physical world — chips, cameras, labs.\nPiloting the World\u0026rsquo;s First Double-Blind AI Evaluations — Google DeepMind, with the Singapore AI Safety Institute, OpenMined, AVERI and MLCommons. A frontier model evaluated inside a confidential-computing enclave: evaluators never see the weights, Google never sees the test set. Model Hardware Standard — Anthropic and HHMI Janelia. A shared spec for AI agents to drive lab and manufacturing instruments in parallel, with safety evaluations alongside — research preview. Remotely Detectable Robot Policy Watermarking — Amir, Flageat, and Prorok. Spectral signals hidden in robot motion, detectable from video alone. IRIS (Infra-Red, in situ) Project Updates — Bunnie Studios. Non-destructive chip inspection with infrared imaging. SecureDNA — Baum et al. Cryptographic screening of the world\u0026rsquo;s DNA synthesis orders. PCR7500 — trusted bioreactors: PCR data signed in a TEE at the point of capture. Nature cryptography? Security primitives based on the laws of nature.\nPhysical One-Way Functions — Pappu et al. The original physically unclonable function: optical tokens as one-way functions. A New Approach to Nuclear Warhead Verification Using a Zero-Knowledge Protocol — Glaser et al. Prove a warhead is real without revealing its design. Quantum Cryptography: Uncertainty in the Service of Privacy — Charles Bennett. Quantum uncertainty itself as a privacy primitive. Consumable Data via Quantum Communication — Gilboa et al. Data that can only be used a bounded number of times. Conjugate Coding — Stephen Wiesner. The 1970s manuscript that invented quantum money. Unclonable Polymers and Their Cryptographic Applications — Almashaqbeh et al. Secret keys stored in molecules that can\u0026rsquo;t be copied. Building Unclonable Cryptography: A Tale of Two No-cloning Paradigms — Almashaqbeh et al. The field map of unclonability, quantum and physical. An Introduction to Protein Cryptography — Tirmazi et al. Encoding data in amino-acid sequences, tamper- and copy-resistant. Cryptography in the DNA of Living Cells — Volf et al. Multi-site base editing as message encryption inside living cells. Hidden Messages in DNA Could Reduce Biosecurity Risks — Danielle Gerhard. DNA watermarks for biosecurity. Neuroscience Needs Network Science — Barabási et al. The brain as a network-science problem. A New Age of Computing and the Brain — Golland et al. A research agenda where computing and neuroscience meet. Mosquito-derived Ingested DNA as a Tool for Monitoring Terrestrial Vertebrates — Chivas et al. Mosquito blood meals as a wildlife-monitoring sensor network. Molecules that Generate Fingerprints — Motiei et al. Fluorescent sensors as chemical fingerprints for authentication. TTEE: Marrying Cryptography and Physics — Quintus Kilbourn. A talk on what physics can do for cryptographic trust. Signature for Objects — Hayashi et al. Signing physical objects, with unforgeability against physically-enhanced adversaries. Cryptographic Data Exchange for Nuclear Warheads — Perry et al. The modern follow-on to zero-knowledge warhead verification. Cryptographic Sensing — Ishai et al. Sensors that reveal only the authorised measurement and nothing else. Cryptography by Cellular Automata — Applebaum et al. How fast can cryptographic complexity emerge in nature? Suggest a link via Discord — try #link-sharing.\n","permalink":"/links/","summary":"Resources and references for the Scaling Trust community","title":"Useful Links"}]