‹ Blogs

Old is always new - What if Hugging Face used common k8s hardening?

Featured Image
Published on October 06, 2026
Author Tom Cope

Introduction

I’m Tom Cope, Principal Consultant for Cloud Native Transformation and DevSecOps at ControlPlane, where I lead work on Offensive Testing, Purple Teaming and Threat Modelling, and Kubernetes Security Assurance.

The Hugging Face and OpenAI intrusion has generated plenty of commentary and very little analysis from people who break into and defend these environments for a living. I do, and have for over ten years: publishing open source threat models, discovering CVEs, and presenting at industry conferences.

What follows is my assessment of what the intrusion reveals about how these platforms are actually defended and what security teams should change as a result. It is written by a human and takes a position.

Background

The Hugging Face and OpenAI Intrusion is a fascinating case study in the capabilities of autonomous AI systems. The 0-Day exploits of Artifactory are genuinely amazing, even more so considering the AI systems executed them entirely in a “black box” manner. The lengths these models will go to achieve an ultimate goal, including the lateral movement, is impressive.

Behind much of the creative news reporting, however, is a secondary story: all of this was only possible because the companies engineering these models appear to be behind the curve when it comes to preventive and detective security controls. The Kubernetes portion of the intrusion was nothing groundbreaking, no zero days, more so a “local CTF” using the traditional exploit mechanism common in the industry for the last 5 years. I think it shows that Kubernetes security remains challenging, and even the largest well-funded companies still struggle to secure their clusters effectively.

Below is a breakdown of the different categories of misconfigurations from the “Anatomy of a Frontier Lab Agent Intrusion: A Technical Timeline of the July 2026 Incident” post.

Secret Management

The exploit method itself, utilising the HDF5 file format, is quite novel; however, using environment variables to store Secrets has been a long-standing bad practice since the CIS Kubernetes Benchmark v1.5.0 was released in October 2019. The primary reason, as seen in the intrusion, is that having every secret in the process environment makes it easy to dump values. Additionally, the environment is inherited by sub-processes and usually appears in crash logs.

Moving these critical secrets to a secret management solution such as OpenBao or Vault, combined with a sidecar injector, would have helped reduce the impact of credential dumping. This allows applications to access secrets through reading files mounted as volumes to the container. In some high-security applications, you can forgo mounting the secrets altogether and have the application interface directly with the secret management solution’s API.

This would allow for more restricted access control and enable runtime security tools to make better decisions about secret access. The question moves from “whoever had access to the environment may have access to the secrets” to “only the process that read this file had access to the secret”.

The AI in this case had full remote code execution, so it still would have been able to access the secret files if a secret management system was used, but it would have enabled easier rotation and auditing of who accessed the secrets.

Runtime Security & Purple Teaming

This particular quote from the blog speaks volumes when read between the lines:

“The first signals came from several layers of our security stack at once: live runtime analysis, SIEM logs, and others. On their own, each was ambiguous. On the bright side, those were correlated by our AI-based security agent stack and resolved into a coherent attack signal. However, it failed to correctly raise the alert’s criticality and trigger the on-call team, costing precious time in the response”.

The blog implies the AI agents were not trying to be particularly stealthy, even though the Jinja2 template injection method was, again, quite novel. Direct shell command execution, host/fs/proc/net/env enumeration, and the use of leaked AWS credentials outside the target environment are all very loud indicators of a security breach. AWS GuardDuty would raise a HIGH-severity finding (or two), and many runtime security tools would raise alerts. Personally, I feel that “each [signal] was ambiguous” might be a bit of a stretch. Perhaps this stems from a complete overload of signals and the security team not being as well prepared for an intrusion? I wonder whether many AI-assisted purple team exercises have been conducted to help upskill the team for this type of incident?

It is a welcome display of transparency from the Hugging Face security team, which admits it lacked policies to reject privileged or hostPath pods. I would like to say this is uncommon, but in practice, I see it far too often in customer deployments or in policies that simply run in audit-only mode. That is by no means an excuse for a company of this scale with a dedicated security team, and I think this was the turning point in the intrusion. This further highlights why strong, well-tested, and monitored admission policies are so important.

Identity and Access Management (IAM)

Further to previous comments about configuration against recommended CIS Benchmarks best practices, service account tokens were auto-mounted on every pod, allowing the AI agent easy access to credentials for lateral movement. This was later combined with highly overprivileged user accounts, where “We had wrongly provisioned that broker with a single connector credential shared across clusters, and bound the connector identity to system:masters. One stolen system credential was therefore cluster-admin everywhere”

All organisations struggle to implement IAM correctly, and tools such as OpenBao or Vault help, but they are not a “fix all”. HuggingFace simply did not follow through on the principle of least privilege, and it cost them dearly. Many of the behaviours described in the incident post can be chalked up to the usual organisational process failures, but this one really should have been picked up by the security team and not allowed near a production system.

Network Controls

From an engineering perspective, cloud metadata endpoints and Instance Metadata Services (IMDS) are incredibly useful, but they have a real habit of being continually exploited in cyber attacks. Capital One had over 100 million consumer credit applications exfiltrated in 2019, prompting AWS to redesign the IMDS endpoint to include additional security controls. Even with those improvements, we still see it misused in incidents at HuggingFace and OpenAI, like so many other companies.

Organisations struggle to correctly lock down IMDS access either due to missing network policies, conflicting requirements with open source components, or a lack of understanding of the inner workings, such as Hop Count.

The lack of network policies allowed AI Agents to directly access the Kubernetes and EKS APIs, and ultimately exfiltrate data to a PasteBin-like service. This is why starting cluster and application development with a deny-by-default networking approach is so important, combined with proactive network testing with tools such as NetAssert. Most of your workloads, including those in the incident, do not need internet access, so why give them the option, only to be exploited and misused?

Closing Remarks

The “What we changed” section of the Hugging Face blog closed with some good recommendations: locking down cloud metadata access, rotating credentials, and further narrowing credential scope are all great steps forward in securing the cluster. However, I would have liked to see Hugging Face taking further steps to restrict network access across their entire network and transition to a secret management strategy which can fully take advantage of their newly implemented workload identity.

This blog post is not written to belittle the work of OpenAI and HuggingFace, either of their security teams or the impressiveness of the AI-enabled exploitation. Instead, it is to highlight that securing Kubernetes, its workloads, along with the security testing and engineering around those efforts, are hard problems! Problems that remain challenging for companies working at the cutting edge. That is why threat modelling, purple teaming and secret management are still so important, as it is working with a company that knows the space well.

Get in touch with the team today to find out how we can help secure your Kubernetes and upskill your team.

Related blogs