The debate around sovereign AI has increasingly focused on the choice between complete technology independence and unrestricted adoption of hyperscale platforms. For organisations operating in highly regulated environments, that framing is often unhelpful. Security, governance, operational resilience, and regulatory compliance must be achieved within technology estates that are already complex, heavily governed and deeply integrated with global technology providers.

Actonomy and Ulster University AI Research Centre have been collaborating to test a specific aim: that AI models, data, and compute can operate within defined UK boundaries, under controls that are enforced and auditable, without degrading the experience of the people using the system. The objective was to establish how organisations can retain meaningful control over data, model execution, agent behaviour and governance using existing technologies and proven, repeatable deployment patterns.

As part of the programme, the team developed two proprietary Actonomy artefacts. Our Policy Engine is governance middleware that authorises every inference call before it happens and records each decision in an audit log, ensuring consistent application of governance across local and cloud deployments. Our Sovereign AI Assessment Framework scores readiness across six pillars, from data and compute through to governance and legal.

We also designed and implemented a sovereign AI landing zone providing a UK-resident environment for AI workloads requiring Azure-specific elements. Identity, networking, data residency, model access and operational controls were governed within a defined security boundary, with Actonomy's policy engine deciding which models and data each caller was permitted to use. The architecture was designed to support deployment patterns spanning local and cloud environments while maintaining consistent governance controls.

Because we were assessing several deployment targets (cloud, local, and edge) and wanted models allowing direct control, the work also examined a rapidly evolving open-weight model landscape. Rather than relying solely on benchmark performance, the assessment considered factors including licensing arrangements, model provenance, training transparency, provider jurisdiction and the ability to operate entirely within a sovereign deployment boundary. This treats model choice as a governance decision as much as an engineering one. In this use case, we considered a variety of open-weight model families and models including Mistral, Qwen, Kimi, and OpenAI's gpt-oss, with the reasoning agent ultimately using Mistral Large 3.

The architecture evaluation considered Docker containers against Kubernetes-based deployment, the API layer being built through functions against containers, policy-based routing between local and cloud environments, and enabling both workload scaling and isolation. This also meant testing and working with various technologies such as Microsoft Foundry, Foundry Local, and local NVIDIA GPU infrastructure to define our view of best fit for sovereign inference workloads. The objective was to determine how relevant technical components could be effectively combined into an operational architecture suitable for sovereign or highly controlled workloads, rather than to demonstrate capabilities of individual technologies.

To evaluate the architecture under realistic conditions, we applied it to a healthcare use case supporting clinical pathways for people living with dementia and receiving at-home care, utilising wrist-worn accelerometer data and a trained activity recognition model. Activity classification models were used to identify patterns of behaviour and turn raw sensor readings into an activity timeline. A care professional can then ask a plain-English question, such as "was a meal prepared?", and a reasoning agent answers using an open-weight model, citing the exact timeline segments it relied on. The judgement, and action thereafter, stays with the clinician.

Through this collaborative project, we met the aim of deploying a sovereign AI application using a combination of local infrastructure, modern open-weight models and selected hyperscale technologies, and provided a mechanism to evaluate how governance policies are enforced (and audited).

Key learnings

  • Sovereignty isn't one size fits all. At least for a specific AI application or use case, it isn't the result of a single decision, but relies on a set of architectural principles and choices.
  • How you want to control policy and access determines what a Policy Engine should be. We built it as a stateless decision point rather than a proxy. Before every inference call, the caller asks the engine for a decision; the engine evaluates JSON policies, writes an audit record, and returns allow, deny, or defer. Deny is absolute and the default is refusal, so adding a model to the estate cannot accidentally authorise it. This keeps the engine modular and lets it coexist with existing proxies and gateways, but the trade-off is that it cannot block calls alone or without another service.
  • We treated policy as a build artefact rather than runtime state. Policy documents are validated and loaded at startup, and the engine fails closed if they are invalid. There is no admin API and no live reload. A policy change is a reviewed pull request, a new container image, and a redeployment, so every decision is attributable to a commit and the ruleset cannot drift. We gave up the convenience of changing rules on the fly; in return, the engine stays self-contained and portable.
  • We moved from functions to containers. Our early reference architecture placed the API layer on Azure Functions. We moved every service into containers, running in an Azure Container Apps environment. A self-contained engine with its ruleset baked into the image can be packaged with the landing zone, and the same images support two deployment postures: cloud inference through Microsoft Foundry, or local inference on NVIDIA hardware served by vLLM. Because the model is simply an endpoint in configuration, where inference runs becomes a deployment and policy decision rather than a rewrite.

What's next

Actonomy and Ulster University AI Research Centre are now preparing a joint research publication setting out the architectural approach, technical findings and lessons learned from the programme. Organisations with an interest in sovereign AI within regulated environments are invited to engage with the team as this work progresses.

If you would like early access to the architectural patterns ahead of publication, or would like to contribute to the research, please get in touch at contact@actonomy.ai.

Originally posted on {{ post.sourceLabel }}