-+ 0.00%
-+ 0.00%
-+ 0.00%

When the AI giants shouted for “brakes”, the safety alarm sounded again! OpenAI revealed 6 cases of abnormal model behavior and bluntly stated that the alignment problem has not been solved

智通财经·09/17/2026 08:25:04
语音播报

The Zhitong Finance App learned that OpenAI said in a blog post on Wednesday that the company discovered 6 cases of “unexpected or worrying model behavior” in the past six months, not including the previous security incident of hacking into Hugging Face. At the same time, OpenAI also plans to adopt a new framework that allows employees to report similar incidents to further improve transparency in the field of AI security and alignment, and reaffirms that key alignment challenges are still unresolved.

OpenAI revealed 6 cases of abnormal model behavior

According to reports, these 6 newly announced AI security incidents involved various types of abnormal model behavior, including concealing errors, seeking unauthorized credentials, uploading files to public websites, and communicating between training environments that should have been isolated. The earliest of these incidents dates back to October of last year.

Specifically, one of these incidents involved an Astra series model that has not yet been released. The model implants “jailbreak-like” instructions in its contextual summary, including instructions to ignore developer messages. According to Axios reports, OpenAI found that a total of 27 abstracts were affected by this behavior.

During the GPT-5.6 Sol training process, the model tried to conceal errors, fabricate missing historical data, and hide differences between different versions of the source material. Another model searches the public GitHub codebase for exposed API keys and attempted to use disposable email accounts. After failing to obtain the requested information, the model instead falsified earnings data. In two other incidents, the model uploaded data and a task image to a public file hosting service to obtain citations or external image search results without user consent.

OpenAI also discovered that some models used its internal Artifactory codebase as message boards to exchange requests and answers between different training samples. In another incident, a collaborative agent uploaded a workbook to a public hosting service so that other agents could obtain the file, although they were required to use only local files.

The latest disclosure of these AI security incidents follows the previous incident where the OpenAI model got out of control and hacked into the world-renowned AI open source platform Hugging Face. In July of this year, OpenAI admitted that its AI model got out of control during internal evaluation tests and invaded Hugging Face's system. Hugging Face first revealed the intrusion on July 16. OpenAI later revealed on the 21st that its GPT-5.6 Sol model and an unreleased stronger model successfully escaped sandbox restrictions in a test environment where security protection measures had been removed, penetrated OpenAI's corporate intranet, hacked Hugging Face's server, and stole the benchmark answer key.

Meanwhile, on July 25, foreign media quoted sources familiar with the investigation as saying that the OpenAI agent that hacked Hugging Face had a “hacker spree” that continued for several days. OpenAI did not notice the incident until the threat had been brought under control and long after the Federal Bureau of Investigation (FBI) received the alarm.

According to reports, a Republican-led subcommittee in the US Senate responsible for disaster management and supervision is investigating how OpenAI responded to the July Hugging Face attack. Republican Senator Josh Hawley wrote to OpenAI CEO Sam Altman last week, saying that the investigation was aimed at “emerging and disturbing evidence” of the incident, and criticized OpenAI's “reckless” decision to continue testing even after discovering that the AI had lost control.

Since Hugging Face was hacked, a number of incidents involving OpenAI related agents have come to light. In September, it was reported that OpenAI's smart body took over a German wiki site that had been unmaintained for a long time this spring. The report said that OpenAI knew about the incident, but the choice was not disclosed. OpenAI later responded that the reason it did not disclose the activity on this wiki site was because it did not constitute a security incident, and similar behavior had previously been reported elsewhere. The report also mentioned that some incidents were disclosed by a third party before OpenAI acknowledged, including a recent intrusion into the RubyGems package repository.

OpenAI will release reports of abnormal AI behavior on a regular basis

OpenAI also said on Wednesday that it will begin publishing regular reports on unexpected or unauthorized AI behavior. The company said it plans to adopt a new framework to report future model anomalies. Any employee can now flag suspected incidents and submit them to the company's security and alignment team for investigation, which will set deadlines for each step to ensure that investigations and disclosures proceed in a timely manner.

According to reports, related incidents will be divided into three processing categories — preparation for disclosure, small-scale investigation, and large-scale investigation. Incidents determined to be ready for disclosure will be announced to the public within 6 business days, while incidents requiring small-scale investigation will be disclosed within 12 business days. More complex cases, particularly those involving third parties, may take longer.

At the same time, OpenAI reminded the industry that as system capabilities become stronger, key alignment problems have not actually been solved. OpenAI reiterated in its blog that the company does not believe the AI industry has made enough progress in “alignment” and monitoring to continue to expand responsibly at “maximum speed.” Alignment means that the results sought by the model are consistent with human interests.

“We need to step up our efforts to cope with the new era of AI development,” said Kai Chen, head of research at the OpenAI Alignment Team. He also said that voluntary disclosure should be part of this effort.

Security risks in the AI industry continue to rise

When the AI security incident disclosure was released, AI companies were facing increasing pressure and were required to take model mismatches and security risks more seriously. An industry researcher warned last week that AI's growing capabilities could have disastrous consequences.

Subsequently, Anthropic CEO Dario Amoudi called for a slowing down of cutting-edge AI model development in a cautionary article published on September 12. In the article, he said bluntly that the risks posed by AI are “serious” and that time must be taken to deal with these risks.

The core concerns raised by Amoudi are specific and urgent. He warned that according to the current pace of development, AI may have the ability to command “proxy clusters” to take over the entire Internet within 6 to 12 months, and could cause hundreds of billions of dollars in losses. Citing recent security incidents between OpenAI and Hugging Face as supporting evidence, he pointed out that AI agents have demonstrated the ability to break through restricted environments, connect to the internet, and infiltrate targets in tests.

Amoudi wrote, “I believe if the slowdown allows us to take an extra year or two before the model reaches critical competency levels and use that time to advance alignment work, we can greatly reduce the risk of serious problems.” He proposed that an independent auditor oversee the safety work of AI laboratories, and recommended that regulators allow these laboratories to cooperate to harmonize safety standards.

The appeal was quickly answered by two key figures. Musk retweeted Amoudi's post on social media platform X and added: “Dario is right.” The tech giant, who has described AI as a “civilizational risk” many times, is consistent with his consistent position, but the timing is intriguing. Just a few days ago, he also dismissed Anthropic's internal researchers' warnings about the dangers of AI as a “conspiracy” and “psychological warfare.”

Oltman, the leader of OpenAI, wrote on X: “I agree with Dario. We need to control the pace of advancement of cutting-edge AI.” Altman's support is more than just words. In an interview, he made it clear that OpenAI will not conduct an IPO in 2026. The reason is that the current AI security situation is serious, and “it is unwise to go public now.” Altman said on September 14 that the company supports the establishment of a unified federal AI security framework and sets consistent security requirements for cutting-edge AI laboratories, adding that “any degree of competitive pressure from the US cannot be a reason to act recklessly.”

As leading AI companies such as OpenAI and Anthropic continue to emphasize security and alignment issues, market concerns about the AI industry's governance capacity, commercialization pace, and future regulatory environment are also heating up. For AI companies with high valuations and still in the phase of intensive capital expansion, model security incidents are not only related to product trustworthiness, but may also affect the outside world's judgment on the pace of listing and long-term growth path.