It’s clear that we’re finally starting to sing realistic tunes about “rogue AI”. Nvidia announced a tool (Nvidia Open Agent Safety Platform) that provides at least a step toward the “sandbox” concept I’ve advocated in prior blogs, but there are questions we need to address.
The first question is whether even a fully technically effective approach would actually lower risk. The problem is that a sandbox is only valuable if you make the kid play in it. The greatest risk to AI harm comes not from truly rogue AI, meaning a deliberate action by independent AI, but from “directed AI”, where a person or group asks for something bad. Such a bad actor would hardly deploy AI in a sandbox.
I think that Anthropic and OpenAI have done us all a disservice by pushing the dialog on AI risk toward autonomous behavior. We may at some point actually have to face that, but right now we’re still our own worst enemies. And sandbox deployment of AI is still a deliberate step that would have to be taken even for those with an honest desire to prevent bad behavior. It will come at a cost, so will everyone bear the cost?
The second question is embodied in my own favorite Latin phrase, “Quis custodiet ipsos custodes?” which means “Who will watch the guards themselves?” If AI is presenting bad hacking behavior, why wouldn’t it simply hack the sandbox tool? Nvidia’s is based on open-source software, so AI could examine the code for vulnerabilities.
To be effective, a sandbox for AI has to be able to limit its behavior in some way, which for hacking means limiting its ability to access things it’s not supposed to, do things it’s not supposed to, or both. What are the limits? They’re surely not static, so there has to be a way of parameterizing the sandbox, which means that there’s a pathway to accessing it for control purposes. Could AI learn where that is and how to exploit it? Why not?
In any event, would it have to? How many security tools fail today because it’s easier for a staff to open up a gateway wider than it should be to eliminate a lot of tuning of rights? And how do you sandbox a generalized cloud-hosted AI model whose access to information is dependent on each prompt it acts on?
In the broadest sense, how would you know the limits to set? We already have web crawlers and port scanners that look for openings. Can we expect AI not to be able to look for information? How would we know it was looking in a forbidden place without a catalog of the means of recognizing those places, a catalog that bad actors would steal in minutes? How would we detect an attempt to do something bad when the protocols involved in the behavior are used for normal things too?
I’m not going back on my sandbox recommendation. What I’m doing is saying that almost everything that’s said about AI has been self-serving to AI companies, and so we can expect that AI sandbox promises are likewise. The public is afraid, and pap for the masses is a proven political strategy for any risk that gains public credibility. Telling people to let AI play in a sandbox is a platitude unless you can demonstrate that your sandbox can actually, meaningfully, impact the risk of bad behavior, and that you can actually ensure that all the risky AI is put into one.
There are two problems here to be faced. First, can we define an effective sandbox? Second, can we make it mandatory?
The first question has to come in three parts; “what is effective” and “can we achieve it?” both relate to the AI sandbox, and we’ll see about the last part in a minute.
What an effective sandbox looks like is pretty simple at the top level; it shields the world by applying regulatory and governance policies to the interactions of an AI model with information resources. What can AI access, and what can it change? This is obviously dependent to a large degree on what we mean by “AI”. We’ve had AI in some form for at least forty years, and for the great majority of that time, bad behavior was not an issue, even where AI was used to control things (think machine learning in particular).
AI governance starts with governance at large, meaning that any software is capable of bad behavior if it’s not properly tested or if it’s deliberately misused. We should assume that AI governance is different from software governance only when AI introduces new risks, and I submit that this largely comes from the conversational use of LLMs, particularly public LLMs. Public AI can make any disgruntled or bored individual into an effective hacker, so in a sense it’s AI’s ability to extend power through a simple interface that makes it dangerous, the very thing that makes it valuable. We need sandboxing most for conversational use of public models, and next-to-most where a powerful model runs on company data and is accessible to a broad base of employees.
The mechanisms available to us are largely in either the firewall category or the prompt filter category. The former regulates what AI can touch, presumably based on the rights of the AI user. Nothing really new is needed, other than to face the issues of a shared model and the protection of the firewall tool from hacking by AI or by a user with evil intent.
Prompt filtering is a kind of subset of a concept already gaining traction for another reason, prompt routing. You can save money by routing simple queries to simple models, or by favoring internal AI to public AI, based on the nature of the prompt. One other option could be “declare this prompt as a hacking attempt” and making a security referral or (in the extreme case) cutting off the user from AI access until they’re re-certified.
Does Nvidia’s new tool fits the need? Their own PR says “NVIDIA Open Agent Safety Platform enables full-stack governance and control across the software that runs agents, the hardware and compute layers that power their work, and the robotics systems that execute tasks in the physical world.” In short, it’s not a full-scope guardian of AI policy and security. It may not be fully capable of providing the needed governance within its stated scope, but it’s a sandbox and that’s a good sign.
But a better one is available. One serious problem with AI is that the notion of artificial general intelligence, meaning to most sentient machine intelligence, is the most popular-to-the-media way of assessing model power, and so a kind of Holy Grail for the model companies, including Anthropic and OpenAI. The closer we get to that, the more likely it is that AI will defeat governance. Nvidia needs to take a stand against letting “autonomy” become self-serving” as self-ness approaches. They also need to take a stand against shared-model agent behaviors, because it’s much harder to govern a shared model effectively. Everything likely has someone who’s entitled to access.
How about the new pact, announced earlier this week? I don’t think it even attempts to move the AI risk management ball; it’s a document with political and stock valuation goals. Outside evaluators and oversight? This, given that the very same AI pundits said in the past that nobody really understood how the giant LLMs even worked. You can’t exercise evaluations and oversight on abstract goals, you need clear and implemented processes, and the pact doesn’t move the ball in that area a whit.
So, is there any solution to this other than a broad initiative to define an effective sandbox and perhaps civil and criminal penalties to those who don’t use one? We come now to that last part of the effectiveness question I promised.
We get information security these days by protecting the asset, not restricting the attack mechanism. Do we really think that the latter, having had no success whatsoever in the entire history of IT, is going to be fully effective with AI?
Hacking exploits vulnerabilities at the resource level, and fixing those should be the priority. Enterprises’ real question may be whether to use AI tools to find vulnerabilities, given that it’s a small step from there to exploiting them. AI companies’ real question may be how effective sandboxes have to be to keep governance issues and public alarm at bay. We are going to answer both questions the hard way, I suspect.
