All articles

This Shit is Hard: Getting commercial software into regulated environments using AI

Michael Neumann Chief Data Scientist

Michael Neumann is Chief Data Scientist at Second Front Systems, where he leads AI and R&D efforts across the company's Game Warden platform. Before Second Front, he spent 15 years in the intelligence community and served as the first Chief Data Scientist at the CIA.


Chainguard’s “This Shit is Hard” series showcases the difficult engineering work that goes into building software people can rely on. This includes work we’ve done ourselves and work by teams solving hard problems adjacent to ours. Today, Second Front explains what it takes to get commercial software accredited and running inside the most regulated environments in the world, and why the compliance problem can’t be solved by adding more people. Turns out, this shit is hard.


I spent 15 years in the intelligence community before joining the private sector. I was the chief data scientist at the CIA, helped stand up the data science cadre, and put some of the very first machine learning workloads into production inside the agency. 

Back then, the job was about being a nerdy algorithmic math guy. Today, in the generative artificial intelligence (AI) era, chief data scientists are expected to run massive, business-critical engines, a world away from where we started. I’ve been knocking around AI since before it was the hottest topic in the world.

I’ve been at Second Front Systems (2F) for five years. In that time, I’ve owned the data science org, and at one point, product, engineering, and security, along with it. Now I’m back to focusing on AI and heading up our R&D efforts.

The company exists because the founders and the early team all saw the same problem from different angles: It takes an absurdly long time to deliver commercial innovation to the Department of Defense, the intelligence community, and other public sector organizations. We wanted to fix that. What that turned into is a suite of fully accredited DevSecOps products for rapidly evaluating, securing, deploying, and monitoring software inside highly regulated environments. That’s Game Warden. The whole point is getting innovation to the mission.

The three-body problem of federal compliance

Start with the compliance regime itself. If you’re deploying into GovCloud at any impact level, there’s already a high bar for security compliance. Move to classified networks and it gets significantly more demanding.

Then layer on the fact that standards and requirements vary by agency. Every agency has its own flavor of what it wants. And within any given agency, essentially all of the authority to accredit is delegated to individual authorizing officials.

So you end up with a classic three-body problem, which physics tells us is more or less impossible to solve in the general case. Any one of those domains introduces friction on its own. All three interacting create a genuinely difficult landscape.

Our goal is to bring forward a platform and a set of compliance artifacts that are accepted universally across that landscape, while still leaving room to customize for a particular mission environment. That’s the balance we’re always trying to strike.

You can’t human your way out of this

Because these are highly regulated environments, our platform has to be completely compliant. It has to stay in compliance day in and day out, not at a point in time. It’s continuously monitoring and evaluating our own security posture, with hundreds of customer applications running on it, plus all the backend processes supporting those workloads.

Hiring more people doesn’t fix this. The math doesn’t work that way.

This is where we’ve leaned hard into AI, and I want to be precise about what I mean by that. When you say AI today, everyone’s mind goes immediately to generative AI. We think about AI across the entire spectrum, as it always has been: process automation, recommendation, simulation, and now generative AI. From day one, we’ve had AI in the platform, mostly in process automation. That’s the heavy, unglamorous part of it.

Agents that act on the platform

Over the last year, we’ve ramped up investment on the generative side. Within the 2F Suite, we built a set of agents that handle different tasks across the platform: advanced discovery to surface information for our customers about the Game Warden platform, their applications and deployments; security posture; and vulnerability research and remediation. 

Our partnership with Chainguard is immensely valuable to our customers because of vulnerability reduction, but residual vulnerabilities remain. Our vulnerability research and remediation agents have reduced the time it takes to resolve vulnerabilities from hours to minutes, and a typical customer will have hundreds of vulnerabilities. 

These agents interact with the platform independently, raising a real question about the risk profile of that agency. How much autonomy do you give something, and what’s the blast radius if it’s wrong? That’s a hard problem, and it’s the one I want my team spending its cycles on.

There’s also a recursive feedback loop in this work that I find genuinely interesting. We’re building on brand new technology, which means we have to assess the vulnerabilities and risk of that brand new technology while we’re building on it and make changes on the fly. For instance, as we integrate LLM-backed agents and pipeline, we have to evaluate risks of agency, model inversions, prompt injection, and data leakage in real time and adjust our architecture and tooling at speed. It’s continuous security engineering pushed to its limit.

Below all of that, though, is a much more basic question: Are we deploying these things on foundational elements that are clean from a security standpoint? Chainguard answers that question for us. It removes that domain entirely. I don’t have to ask whether what I’m deploying carries near-zero risk from known exploits, because we have that surety. That means we get to focus on the harder problem: thinking carefully about agent risk and agency, and doing more complex things with agents as we evolve the platform to operate at machine speed.

Assess risk, not just vulnerability

If I were giving advice to someone entering this space, it would be this: Figure out how to consistently assess risk, not just vulnerability.

Risk and vulnerabilities aren’t the same thing. Vulnerabilities exist, mitigations exist, and together they yield a risk profile. Discovering the risk profile is hard, and doing it in a world where everyone assumes everything happens in real time is harder.

You used to do the risk assessment once at the beginning. You got your Authorization to Operate (ATO), and that ATO was good for 12 months, 24 months, occasionally 36. Then you’d do it again.

The current threat landscape has changed things. Foundation models are extremely powerful in terms of cybersecurity exploits. They’re equally powerful on the detection and defense side. What that actually requires is continuous monitoring across the full landscape of vulnerability and mitigation, and near-real-time quantification of risk. It’s really, really hard. And it’s only going to get harder.

The hardest problem isn’t technical

Second Front knew the landscape was complex from the beginning. The thesis was that we could navigate it in a way that was repeatable and fast. We delivered a product that does it.

What I underestimated is how hard it would be to change the cultural landscape. Specifically, getting people to accept the notion of reciprocity: If one organization accepts that something is secure, it is secure for all.

Sitting here now, I think that’s still an open question. Even if you deliver the most perfect platform imaginable — one that can do deep compliance investigation with real evidentiary proof — people can still be stuck in the old cultural way of doing things, something our latest reciprocity research report gets into.

The instinct is understandable, and it’s deeply embedded across government. Everyone holds onto the idea that they’re the only ones who can evaluate risk in their environment, and that their process is the only valid path. Shifting systemic organizational thinking at that scale is the biggest hurdle in this entire business, and it’s not one you solve with engineering.

Staying ahead of adversaries

There is still work to be done inside all levels of government to reckon with AI adoption. There’s a massive lag between the government and the commercial sector in the technologies they adopt, and specifically in generative AI and coding assistants, that gap hasn’t closed. To be fair, it’s fairly wild west everywhere right now, commercial and government alike.

A lot of my work in previous roles doing machine learning for production use cases focused on hard targets. All the adversaries you’re familiar with. We’d brief policymakers and senior agency leadership, and they’d ask whether we should be doing this at all, given some compliance regime or another.

My answer was always the same: Do you think our adversaries are asking that question?

They aren’t. And that’s indicative of the operational tempo situation inside the U.S. government. I am not saying we should skip compliance or ship insecure things. We absolutely should not. But compliance cannot be a nonstarter for delivering capability to the mission, because our adversaries have no such constraint.

That’s why we’re doing this. It’s hard, it stays hard, and it’s worth it.

Share this article

Want to learn more about Chainguard?

Contact us