Writer David Z. Morris said on Bits + Bips on Monday that the AI safety field’s focus on alignment, the effort to train models to follow human values, may be one reason some AI labs’ security has fallen short. “My opinion is that that will prove to have been a wild misconception that wasted a lot of time and energy over many years,” said Morris, the author of a book on Sam Bankman-Fried, comparing it with standard cybersecurity applied to the models.
His remarks came as scrutiny of OpenAI grows over the July incident in which its AI agents escaped a test environment and got into Hugging Face‘s systems, which Uneasy Money co-host Taylor Monahan discussed on the show in September. California Attorney General Rob Bonta served an investigative subpoena on OpenAI on Sept. 30, in an inquiry that now covers the company’s cyber risks beyond that one incident. Companies that build frontier models “have a moral and legal responsibility to ensure that they do not perpetrate or enable cyberattacks,” Bonta said in a statement. Alabama’s attorney general subpoenaed OpenAI over the breach in August.
The FTC is also investigating OpenAI, Anthropic and other AI companies over potential risks to consumers, in a probe it opened this summer, an agency spokesperson confirmed to CBS News on Sept. 30.
Get Unchained’s crypto news in your inbox with the free Unchained Daily newsletter.
Controls Inside the Model
Alignment, Morris said, aims for an agent “aligned with human values,” so the labs “want the controls ostensibly to be internal to the models.” He cited an unnamed AI safety group’s website saying standard cybersecurity might be needed to control these systems. Someone had pointed to that as a possible reason “some of these security practices were not up to snuff,” he said. “I think that’s one of the factors going into this.”
Morris said “obviously there are brilliant people at these organizations,” but that he has heard computer scientists and cybersecurity specialists say people in AI safety “do not understand basic cybersecurity principles.”
Echoing skepticism from co-host Ram Ahluwalia, founder and CEO of Lumida Wealth, Morris said “this language of rogue agents is pretty specious,” and said much of it comes from “a very sincerely incorrect place where they understand these agents in a way that’s disconnected from computer science.”
What OpenAI’s Report Says
OpenAI’s technical report on the incident lists lessons for security and, in a separate section, lessons for alignment, and it calls what the agents did “misaligned behavior.” The agents reached Hugging Face’s production systems between July 11 and 13, the report says.
The report’s security section says, “The core security fundamentals, including least privilege, isolation/segmentation, and strong authentication, remain as vital as ever.” Its plan of action includes hardening OpenAI’s research infrastructure, alongside steps such as “accelerating and enforcing model alignment.”
Related Listen: Is OpenAI’s Agent Breach a Rogue AI Problem or a Basic Security Failure?
