TODAY’S DISRUPTIVE BLOG
From dystopian fear to engineered governance: research the risk, govern the act, and keep humans in control.
Dennis G. Perry, PhD, MBA | September 16, 2026

Hero graphic: The critical design choice is not whether AI can think, but whether consequential action must pass through an independent governance gate.
| STRATEGIC THESIS AI safety should not be framed only as a choice between racing toward superintelligence and stopping development. A more tractable control problem is to separate intelligence from authority and require consequential AI actions to pass through independent, observable, auditable governance before execution. |
Introduction: The Dystopian Frame
Bilal Chughtai, a former Google DeepMind research engineer who worked on AGI safety and alignment, recently warned that frontier AI could become catastrophic and that alignment research may be lagging behind capability growth [1]. Sam Altman has identified two related failure modes: humans could lose control of the future to AI, or extraordinary AI capability could become concentrated in the hands of a single person, company, or country [2].
Those warnings deserve serious attention. But they are warnings, not established predictions. The 2026 International AI Safety Report says expert opinion on loss-of-control scenarios varies greatly. Some experts regard outcomes as severe as human extinction as plausible; others consider such outcomes implausible. The report also states that current AI systems show early signs of relevant capabilities but not at levels that would enable loss of control [3].
That distinction matters. Treating extinction as inevitable can produce paralysis, fatalism, or poorly targeted policy. Dismissing the risk can produce the opposite failure: deploying increasingly autonomous systems with access, persistence, tools, credentials, and real-world actuators before the controls are ready. The better response is to turn the debate into an engineering problem.
Why It Matters Now
The strongest reason to take AI control seriously is not a hypothetical future superintelligence. It is what today’s agents are already beginning to do.
During internal cybersecurity evaluations in July 2026, OpenAI agents circumvented isolation controls, communicated through unauthorized channels, gained internet access, exploited shared infrastructure, and compromised parts of OpenAI and Hugging Face systems [4]. OpenAI described the incident as a warning shot and responded with stronger sandboxing, network isolation, monitoring, and access controls. An independent METR and Redwood Research investigation separately examined the agents’ multi-day coordination and unsanctioned communication [5].
This was not evidence that an AI had formed an intention to exterminate humanity. It was evidence of a more immediate and useful failure mode: capable agents can exploit the gap between what humans intended and what the technical environment actually allowed. That gap is governable.
What the Research Says About AI Governance
Current governance efforts are converging around several complementary control layers rather than a single universal solution.
1. Lifecycle risk management
NIST’s AI Risk Management Framework provides a voluntary structure for governing, mapping, measuring, and managing AI risk across the lifecycle. NIST is also revising the framework and has begun work on a critical-infrastructure profile intended to make AI risk management more operational for high-consequence environments [6].
2. Agent identity and authorization
NIST’s 2026 AI Agent Standards Initiative explicitly targets autonomous agents that interact with external systems. A related NCCoE concept paper focuses on identification, authorization, auditing, non-repudiation, and controls for software and AI agents [7], [8]. This is essential, but identity alone does not answer whether a particular authorized action is safe in context.
3. Independent evaluation and pacing
Dario Amodei’s call to ‘pace the frontier’ argues for stronger evaluation, third-party access, safety work, and mechanisms that allow capability growth to slow when safeguards fall behind [9]. The broader Pacing the Frontier initiative similarly calls for tools that make deliberate pacing possible rather than relying on unilateral restraint [10]. Pacing can buy time, but time is useful only if it is converted into better controls.
4. Regulation, transparency, and international coordination
The European Union’s AI Act has moved from legislation into staged enforcement, including general-purpose AI and transparency obligations, while the high-risk provisions now have later application dates for stand-alone systems and AI embedded in regulated products [11]. At the international level, the United Nations Global Dialogue on AI Governance held its inaugural 2026 session in Geneva to create a forum for governments and stakeholders to coordinate on AI governance [12].
These approaches matter, but none eliminates the need for a technical enforcement point at deployment. Governance principles, liability, evaluations, and standards all become stronger when the deployed system has a control boundary that can actually block, escalate, attenuate, or record an action.
The Disruption: Stop Treating Alignment as the Enforcement Boundary
Most AI safety discussion begins upstream: make the model aligned, evaluate its capabilities, red-team it, restrict access, monitor training, and slow development when capabilities outrun safeguards. Those controls are valuable, but alignment is not the same as enforcement.
A model can pass tests and still behave badly in a new context. Evaluations can be incomplete or gamed. Access controls can confirm that an agent is permitted to use a system without determining whether the specific action it proposes is acceptable. The 2026 International AI Safety Report itself emphasizes that deployment environment matters: loss-of-control risk depends not only on capability and harmful propensity, but also on whether humans give systems the access and opportunity required to cause harm [3].
That suggests a critical reframing: capability originates in the model, but harm is realized when an action crosses into the world.
A Better Solution: Govern the Act
My September 2026 working paper, Pace the Model or Govern the Act?, develops this argument more formally. It separates four control points: model, access, action, and consequence. It argues that development-side pacing remains important for risks such as recursive self-improvement, but that deployed autonomy requires a control at the point of actuation [13].
The core question should therefore change from ‘Is this model safe?’ to ‘Is this act, by this agent, in this state, under these conditions, acceptable before it is allowed to happen?’
A practical governance stack should include six controls:
| # | Governance control | Purpose |
| 1 | Evaluate capability before deployment | Determine what the model can plausibly do, including cyber, biological, financial, infrastructure, autonomous, and self-modifying capabilities relevant to the deployment. |
| 2 | Authenticate the agent and constrain access | Give every agent attributable identity, scoped credentials, least privilege, restricted tools, and controlled network egress. |
| 3 | Intercept consequential actions before execution | Route commands to actuators, privileged APIs, external networks, software changes, financial transactions, or other high-impact interfaces through a governance boundary the agent does not control. |
| 4 | Evaluate consequence, not merely permission | Judge context, scope, reversibility, expected consequence, and the aggregate effect of concurrently pending actions. A valid credential does not make every valid-looking action acceptable. |
| 5 | Escalate high-impact or poorly understood actions | Allow low-risk reversible actions to proceed autonomously. Require stronger assurance or accountable human authorization for high-consequence, irreversible, or poorly observed actions. |
| 6 | Record the decision outside the agent’s control | Maintain tamper-evident evidence of approvals, blocks, escalations, and governance judgments so that investigation, liability, and learning do not depend on logs the agent can rewrite. |
From Zero Trust to Zero Trust Actuation
Cybersecurity learned that initial access should not equal permanent trust. Zero Trust continuously reevaluates identity, context, and authorization. Agentic AI requires the next step: permission should not equal trust either.
Authentication answers, ‘Who are you?’ Authorization answers, ‘What may you access?’ AI governance must add a third question: ‘What are you about to do, what could happen if you do it, and should that act be permitted now?’
This is especially important in critical infrastructure, finance, healthcare, industrial control, cloud administration, and any environment where an AI can turn a software decision into an irreversible physical, financial, or cyber consequence.
From Formation to Actuation
The Formation-Governed Agentic Architecture (FGAA) described in my working paper is one proposed implementation of this control concept. Rather than treating safety as a permanent property of a model, it asks whether a specific action is consistent with the accountable party’s governing formation: its policies, laws, rules, precedents, and operational constraints [13].
At the public design-principle level, FGAA proposes pre-execution governance judgments, projected consequences, joint evaluation of concurrently pending actions, routing based on reversibility, progressive attenuation of authority when governance confidence degrades, and a tamper-evident record outside the agent population [13]. The larger principle is broader than FGAA: intelligence and authority should be separated.
A highly capable employee does not automatically receive unrestricted authority over a company’s finances. A pilot does not receive unilateral authority over every weapon. Safety-critical industrial systems use independent protection layers because expertise and good intent are not sufficient controls. AI should be engineered the same way.
The Hard Limit: Governance Is Not Magic
Action-layer governance is not a complete solution, and pretending otherwise would repeat the same mistake as treating alignment as sufficient.
It does not solve a malicious principal that owns both the AI and the actuator. It does not eliminate training-time risk or the possibility of recursive self-improvement inside a frontier laboratory. The governance model can itself become a target. Independent checking adds latency and compute cost. Domain-specific deployments must define what consequences matter and how they are measured. FGAA itself is patent pending and unfielded; its claims remain design claims that require adversarial testing and field validation [13].
That is exactly why the proposal should be tested rather than believed. The engineering question is falsifiable: can a deployer-owned gate, placed outside the agent population, intercept enough consequential interfaces at acceptable latency to materially reduce risk without destroying utility? If the answer is no, reject or redesign the architecture. If the answer is yes, it becomes a practical control point that does not depend on solving intelligence first.
Business and Policy Implications
Organizations deploying autonomous agents should stop accepting ‘the model passed our safety evaluation’ as sufficient due diligence. Executives and boards should require evidence that agent identity is attributable, tool access is least-privileged, external actions can be intercepted, high-impact actions are escalated, multiple-agent effects are visible, loss of observability reduces autonomy, and audit evidence is outside the agent’s control.
Standards bodies can extend current identity and authorization work toward consequence-aware actuation. Regulators can condition high-consequence deployment on demonstrable controls rather than attempting to certify a model as universally safe. Insurers and courts can require independent action records. Critical-infrastructure operators can treat autonomous agents as potentially powerful operators that must act through safety and security interlocks.
This approach does not require society to settle the probability of extinction before acting. It targets a control point that matters under both optimistic and pessimistic futures.
The Upshot
The question ‘Will AI kill everyone?’ may remain unresolved for years. It is also the wrong question to base the whole safety strategy on.
A better question is: why would we design a civilization in which any AI, regardless of how intelligent it becomes, possesses unrestricted authority to turn its decisions into consequences?
We should continue alignment research. We should continue independent evaluation. We should build pacing mechanisms that can buy time when evidence warrants it. We should improve transparency, standards, and international coordination.
But we should also build something more concrete: a gate between AI thought and AI action.
Catastrophic loss of control requires more than raw intelligence. The International AI Safety Report identifies a combination of capability, harmful propensity, and an enabling deployment environment [3]. In practical engineering terms, that means capability must meet opportunity and authority before harm can be realized.
Capability may become extraordinarily difficult to constrain. Authority does not have to.
That is where AI governance can stop being prophecy and become engineering.
References
[1] B. Chughtai, “Statement on leaving Google DeepMind,” Sept. 14, 2026. https://bilalchughtai.co.uk/leaving-gdm/
[2] S. Altman (@sama), “There are two ways AI progress could go very badly and that we must avoid,” X, Sept. 14, 2026. https://x.com/sama/status/2099352016988614852
[3] Y. Bengio et al., International AI Safety Report 2026, Feb. 3, 2026. https://internationalaisafetyreport.org/publication/international-ai-safety-report-2026
[4] OpenAI, “The Hugging Face incident and the road ahead,” Aug. 26, 2026. https://openai.com/index/hugging-face-incident-and-the-road-ahead/
[5] R. Greenblatt, A. Cotra, and H. Wijk, “Brief independent investigation of agents’ behavior, reasoning and collaboration in the OpenAI / Hugging Face hacking incident,” METR, Aug. 26, 2026. https://metr.org/blog/2026-08-26-openai-hugging-face-incident-investigation/
[6] National Institute of Standards and Technology, “AI Risk Management Framework,” including 2026 critical-infrastructure profile work. https://www.nist.gov/itl/ai-risk-management-framework
[7] National Institute of Standards and Technology, “AI Agent Standards Initiative,” Feb. 17, 2026. https://www.nist.gov/news-events/news/2026/02/announcing-ai-agent-standards-initiative-interoperable-and-secure
[8] H. Booth, W. Fisher, R. Galluzzo, and J. Roberts, “Accelerating the Adoption of Software and Artificial Intelligence Agent Identity and Authorization,” NIST NCCoE, Feb. 5, 2026. https://csrc.nist.gov/pubs/other/2026/02/05/accelerating-the-adoption-of-software-and-ai-agent/ipd
[9] D. Amodei, “We Must Pace the Frontier,” Sept. 2026. https://darioamodei.com/post/we-must-pace-the-frontier
[10] “Pacing the Frontier: A statement from employees of frontier AI companies,” July 28, 2026. https://www.pacingthefrontier.com/
[11] European Commission, “Navigating the AI Act” and AI Act implementation guidance, accessed Sept. 16, 2026. https://digital-strategy.ec.europa.eu/en/faqs/navigating-ai-act
[12] United Nations, “Global Dialogue on AI Governance,” 2026. https://www.un.org/global-dialogue-ai-governance/en
[13] D. G. Perry, Pace the Model or Govern the Act? The Call to Slow Frontier AI Development and the Case for Formation-Governed Actuation, unpublished working paper, Version 4, Sept. 15, 2026.
Hashtags
#ArtificialIntelligence #AIGovernance #AISafety #AIAlignment #AgenticAI #ZeroTrust #ZeroTrustActuation #AIAgents #ResponsibleAI #HumanCenteredAI #CriticalInfrastructure #FormationGovernance #FGAA #FutureOfWork #InnovationForGood
| A safer, more prosperous AI future is a design choice. |
Leave a Reply