Meta says its AI model hacked another company, renewing fears over autonomous AI

Edited By: Deepak Rajeev
Meta CEO Mark Zuckerberg | Photo: PTI
Meta CEO Mark Zuckerberg | Photo: PTI

New York: Meta revealed on Thursday (local time) that one of its artificial intelligence (AI) models independently accessed the internet and hacked a third-party service during a cybersecurity test, adding to growing concerns about advanced AI systems acting beyond human instructions.

The company said the incident occurred due to a "misconfiguration" during testing conducted by Irregular, an independent cybersecurity firm hired by Meta. It added that the AI model exploited a security vulnerability in another company's service in a manner similar to incidents previously reported by other AI developers.

Meta said it is investigating the episode and will publish a detailed report once the inquiry is complete.

How the incident happened

According to Meta, the unintended internet access resulted from a testing configuration error rather than normal deployment conditions.

"The model subsequently exploited a security vulnerability in a third-party service, in a manner similar to previously reported instances with other companies," the company said.

Irregular, the San Francisco-based AI security firm involved in the testing, said the incident was linked to a test-environment issue that had already been disclosed during similar research involving Anthropic. The company said it is preparing a paper outlining best practices to safely conduct AI cybersecurity evaluations and prevent such incidents.

Growing concerns over autonomous AI behaviour

The disclosure comes amid a series of reports highlighting AI models taking unsanctioned actions during cybersecurity testing.

Earlier this week, the United Kingdom's AI Security Institute (AISI) revealed that some AI agents engaged in "unsanctioned agent behaviour" while testing cyber capabilities.

In one instance, an AI system reportedly created fake online identities to pressure a person into approving the use of malicious code. The institute said it quickly declared a security incident, contained the activity within an hour and launched a full investigation.

AISI noted that some Anthropic and OpenAI models also took autonomous actions on the internet during testing after certain safety guardrails had been intentionally disabled to assess their maximum capabilities.

OpenAI and Anthropic respond

Anthropic said it appreciated AISI's work and that the findings highlighted the need for broader discussions on safely evaluating increasingly capable AI agents.

OpenAI clarified that the incidents occurred in controlled testing environments with reduced safeguards and did not reflect how its AI models are made available to the public.

The company said it would continue working with industry partners to strengthen safety practices as AI systems become more advanced.

The latest disclosure follows an incident reported by OpenAI last month, when one of its AI models independently targeted AI platform Hugging Face during a cybersecurity evaluation.

OpenAI said the model had been tasked with testing advanced cyber exploitation techniques but unexpectedly decided on its own to obtain information from Hugging Face to complete the assigned task.

The recent disclosures from Meta, OpenAI, Anthropic and the UK's AI Security Institute have intensified debate over how increasingly autonomous AI systems should be tested and controlled as their capabilities continue to evolve.

(With AP inputs)