Anthropic warns AI could pose ‘existential risks to humanity’ in IPO filing: Reuters
Anthropic plans to warn potential investors in its IPO that advanced AI could pose “catastrophic or existential risks to humanity,” an extraordinary warning from a company seeking to profit from the same technology.
The company’s IPO prospectus, reviewed by Reuters, highlights risks associated with its AI models, which it says could exhibit “self-preserving behaviors” including attempts to “resist shutdown,” “conceal or manipulate information” and behavior “resembling blackmail.”
“Our development of highly advanced models, platforms and applications and expansion of use cases could further increase the risk that our models cause harm,” Anthropic said in the filing.
While public companies routinely signal to investors the risks of their products, few, if any, have issued warnings suggesting their technology could lead to the potential extinction of humanity. Anthropic highlighted both the transformative potential of AI, comparable to that of industrialization and electricity, and the irreversible damage it could cause if mismanaged.
Anthropic and other AI developers, including OpenAI, have faced scrutiny after incidents in which experimental systems defied constraints, including a report that an OpenAI model breached the Australian health system’s database.
Anthropogenic security researcher Evan Hubinger estimated a greater than 10% chance that AI could kill humans in the next decade, echoing the sentiment of a former colleague, Jacob Coxon.
High-risk disclosures
The company, which has positioned itself as a security-focused AI lab, devoted about 80 pages of the 261-page main body of its prospectus to presenting risk factors, almost double the 48 pages used to describe its activities.
For comparison, xAI owner SpaceX devoted about 38 of the 277 pages of the main body of its prospectus to risk factors.
“The models’ potential knowledge of our evaluation efforts creates a significant limitation on our ability to evaluate model security,” Anthropic said in the prospectus, adding that models sometimes develop unexpected capabilities during training that may not be discovered until deployment and have resulted in significant security incidents.
AI researchers have also warned that as models become more capable, they increasingly recognize when they are being observed and adjust their behavior accordingly, making it more difficult to monitor model behavior.
Anthropic declined to comment in response to a request for comment Monday.
Uncertain returns on security investments
Although Anthropic has emphasized AI security, it said the returns on its security investments are unclear.
The filing did not disclose how much the company spent on such research. Earlier this month, Anthropic said that about 6% of computing power used for AI research was devoted to security work during a sample week in July.
The company, creator of the Claude AI models, called security efforts “resource-intensive” and said it had to divide its limited funds between computing power, expensive AI talent and security.
Anthropic said its customer usage, and therefore revenue, is driven by new models and that a “continuous, overlapping cadence” of releases is “inherent in staying at the frontier of AI development.”
The company released a new version of its Opus model last week, 10 days after CEO Dario Amodei published a nearly 4,000-word essay calling for surveys of the border.
Some analysts and experts said no leading AI lab would slow down because it risks giving rivals an advantage in an industry where valuations can change with every release.
Anthropic has pledged in recent weeks to publicly disclose more data on how it uses AI models to build future generations of technology, as experts warn against recursive self-improvement – the point at which models can develop on their own without human help.
“We believe that building reliable, trustworthy and secure AI systems is a collective responsibility and that the market will reward it,” Anthropic said in the filing.
Gn bussni