- Enterprise AI scaling bottlenecks stem from a lack of shared infrastructure standards, not model quality.
- Adapting cloud-native infrastructure with AI-native Kubernetes tooling and checkpoint management is crucial for scaling AI workloads.
- Outcome-driven AI observability requires integrating human feedback mechanisms beyond traditional error rate monitoring.
AI models are easy to deploy. Scaling them across an organisation is where most teams hit a wall. In a June 9, 2026 article for CIO, Brendan Burnsco-founder of Kubernetes and Corporate Vice President for Azure cloud-native open source at Microsoft, argues that the bottleneck is not model quality, it is the absence of the shared infrastructure standards that made cloud computing work. “AI doesn’t scale on models alone,” he writes. “It requires the same focus on infrastructure, shared open-source standards and systems that defined the cloud.”
1. Adapt Cloud-Native Infrastructure for AI Workloads
Kubernetes was designed for stateless services. AI workloads break that assumption in two specific ways: training runs are checkpoint-sensitive, meaning a single hardware failure can force an expensive rollback, and GPU scheduling requires co-location logic that standard CPU and memory allocation does not provide.
1.1. Enhance Kubernetes with AI-Native Primitives
GPUs often require high-speed interconnects such as NVLink, which means interdependent components need to land on the same physical machine. Standard Kubernetes schedulers do not handle this well. “Gang scheduling” addresses the problem by allocating a group of pods simultaneously on co-located resources, or not at all, preventing the partial-allocation deadlocks that stall distributed training jobs.
The Kubernetes AI Toolchain Operator (KAITO) is one project addressing these needs directly, focusing on distributed inference and fine-tuning orchestration via custom resources. Teams can pair KAITO with PodGroup custom resource definitions and a gang scheduler such as Volcano to manage GPU allocation atomically across multi-node training runs. Burns’s infrastructure argument rests on treating GPU-backed workloads as first-class citizens within the orchestration layer, not as an afterthought bolted onto existing tooling.
1.2. Implement Checkpoint-Aware Resource Management
A hardware failure mid-way through a multi-hour training run can be expensive. Integrating checkpointing mechanisms from training frameworks such as PyTorch Lightning or TensorFlow Keras callbacks with the underlying storage layer reduces that exposure. Persistent storage solutions offering high-throughput, low-latency access, Azure NetApp Files and Ceph are common options, are well-suited to frequent checkpoint writes.
On the orchestration side, long-running training jobs benefit from dedicated node priority classes that minimise preemption risk. Event-driven triggers that save checkpoints on detection of node instability signals, surfaced through Prometheus and Grafana alerts on kube-node-lifecycle events, give teams a practical way to reduce the cost of failures without restructuring the entire training pipeline.
2. Implement Outcome-Driven AI Observability
Burns’s point about monitoring is direct: error rates are not enough. The relevant question is whether the system gave a good answer, and answering that requires human input alongside technical metrics.
2.1. Integrate Human-in-the-Loop Feedback Mechanisms
For user-facing AI applications, simple thumbs-up/thumbs-down feedback at the interface level is a practical starting point. Collected and anonymised, that signal can feed directly into monitoring and retraining pipelines. A high thumbs-down rate on a specific query class can surface model drift or unhandled edge cases faster than any latency graph.
Burns’s framing here matters: these are relative metrics, not absolute ones. A drop in negative feedback from 50% to 40% is a positive signal even if 40% still looks bad in isolation, it indicates the direction of travel is correct. Platforms such as Arize AI can link user sentiment signals to model performance metrics and trigger alerts when satisfaction drops below defined thresholds for a given query type.
2.2. Develop Behavioural and Relative Metrics for AI Performance
Passive behavioural signals complement explicit feedback. In conversational AI, the number of turns a user needs to reach their goal, or average interaction length, can indicate whether responses are landing. A customer service agent’s “successful resolution rate”, the proportion of queries resolved without human escalation, tells you more about functional quality than uptime ever will.
Establishing baselines for these metrics and tracking change over time lets teams assess whether model updates are genuinely improving user experience or quietly degrading it. Grafana dashboards drawing from Datadog or Prometheus can surface these behavioural trends alongside traditional operational metrics, giving a more complete picture of system health.
3. Automate Large-Scale AI Testing with LLM Evaluators
Manual prompt review does not scale. Burns draws on the approach used for web search testing: run thousands of prompts through the system and evaluate whether responses improved. That requires automated evaluation, and for LLMs, that often means using another model to do the grading.
3.1. Implement LLM-as-a-Judge for Automated Evaluation
The process is straightforward: feed an input prompt, the system’s response and a scoring rubric to a separate evaluator model, which grades the output on criteria such as relevance, coherence, factual accuracy and safety. Burns acknowledges the signal is imperfect, but describes it as “good enough to tell you which direction you’re heading.”
For an internal code generation tool, a rubric might cover syntax correctness, adherence to coding standards and functional accuracy. Teams can generate thousands of test cases covering edge cases and known failure modes, then assess outputs automatically using evaluation modules from frameworks like LangChain or API calls to capable models. Wiring these evaluation suites into a CI pipeline via GitHub Actions or GitLab CI means every code commit triggers a validation run, catching regressions before they reach production.
3.2. Conduct Continuous A/B Testing with Small-Percentage Experiments
Burns advocates for weaving small-percentage experiments through the entire AI stack, routing a fraction of production traffic to a new model version while the remainder stays on the stable release. The approach lets engineers observe real-world performance without exposing the full user base to an untested change.
Feature flag tools such as Split.io or LaunchDarkly can manage the routing configuration. The behavioural and relative metrics from Section 2.2 then provide the evaluation criteria: if a 1% experiment shows a notable drop in resolution rate or a rise in average interaction length, that is a signal to investigate before any broader rollout. Data-driven decisions replace guesswork about whether a model update is ready for production.
4. Embrace Open Standards and Interoperable Frameworks
The most consequential gap Burns identifies is the absence of “interoperable frameworks, shared protocols and secure, community-driven innovation.” Proprietary silos do not just create technical friction, they prevent the kind of convergence that allowed cloud infrastructure to mature. His argument is that AI needs the same consolidation around open standards that Kubernetes delivered for container orchestration.
4.1. Advocate for and Adopt Shared Interfaces and Protocols
Burns is specific about where shared interfaces matter most: inference and routing, quality gate representation, system health signalling and telemetry. Without common definitions for these, each integration becomes a custom project.
The Cloud Native Computing Foundation and the LF AI and Data Foundation are active in this space. For inference, protocols such as Open Inference Protocol and community APIs from projects like KServe enable model serving across environments without vendor lock-in. For observability, OpenTelemetry provides a unified collection framework for traces, metrics and logs that works regardless of the underlying model or serving infrastructure. Adopting these standards reduces the integration overhead of adding new models or services and improves system resilience across the stack. This connects directly to the broader governance questions raised by Italy’s CNR work on systematic AI oversight frameworks.
4.2. Contribute to Community-Driven Innovation for AI Foundations
Kubernetes succeeded in part because multiple major vendors chose to build on a single vendor-neutral foundation rather than maintain competing proprietary solutions. Burns’s contention is that AI infrastructure needs the same bet. Building internal proprietary tooling for every fundamental AI challenge fragments the ecosystem and duplicates effort that open collaboration could avoid.
Practically, this means contributing to projects that extend Kubernetes for AI workloads, participating in standard MLOps tooling such as MLflow and Kubeflow, and helping define benchmarks for AI reliability and security. Projects like Ray, which provides a distributed Python computing API, address orchestration challenges for complex agentic workloads and benefit from broad community contribution. The governance and interoperability concerns Burns raises here also apply directly to the challenge of governing web-connected AI agents at scale.
5. Establish AI Agent Governance and Guardrails
As AI systems move from discrete models to autonomous agents capable of acting on enterprise data and external systems, governance becomes a technical requirement, not just a policy consideration. Burns’s position is direct: “control the output of these agents is a critical part of how we can go and do AI at scale.”
5.1. Implement Agentic System Control and Auditing
Governance at the agent layer starts with defining permissions and acceptable output boundaries. Embedding safety instructions in the agent’s system prompt, specifying which data sources to verify against, what categories of information must never be disclosed, is a practical first step. A secondary filtering layer, either a smaller LLM or a set of rule-based classifiers, can review agent outputs for safety, bias or policy violations before they reach users or trigger downstream actions.
Auditability is equally important. Logging all agent interactions, decisions and outcomes to an immutable store gives teams the traceability needed for debugging, compliance and identification of emergent behaviours. Regular review of those logs, whether weekly or monthly depending on deployment risk, allows agent rules and models to be refined as edge cases surface.
5.2. Streamline Dependency and API Management for AI Agents
AI agents typically depend on a web of external tools, APIs and libraries. Burns flags that subtle API changes in dependencies can accumulate as unresolved pull requests and unaddressed security vulnerabilities, a slow-moving but real risk in production systems.
Burns notes that AI agents can assist in making those small API changes, potentially improving Dependabot PR merge rates.
The infrastructure, observability, testing, standards and governance work Burns describes is not a checklist to complete once. It is the ongoing engineering discipline that separates teams running AI reliably in production from those still troubleshooting why their models do not behave the way they did in the prototype. For more coverage of AI policy and regulation, visit our AI Policy & Regulation section.



