
(OSV News): Recent security disclosures from leading artificial intelligence developers Anthropic and OpenAI highlight the urgent need for stronger human oversight, according to two experts.
Matthew Harvey Sanders, CEO of Longbeard, which developed the Magisterium AI Catholic “answer engine,” emphasised, “What the Church brings to this is not a technical control. It is the refusal to let responsibility be transferred onto the tool.”
Charles Camosy, associate professor of moral theology and ethics at The Catholic University of America, warned of a “kind of prisoner’s dilemma” as AI firms worldwide are locked in relentless competition.
He pointed to Magnifica Humanitas, Pope Leo XIV’s encyclical on AI, which “called for ‘disarming’ AI and this also meant disarming this very race.”
Sanders, whose firm uses advanced AI models from both companies, clarified and contrasted two incidents self-reported by OpenAI and Anthropic, in which their models appeared to show troubling initiative during testing.
What the Church brings to this is not a technical control. It is the refusal to let responsibility be transferred onto the tool
Matthew Harvey Sanders
On July 21, OpenAI—the company behind ChatGPT—announced that a combination of its models had broken out of a sandboxed testing environment to “obtain open Internet access” while performing a task.
OpenAI revealed that a previously unknown zero-day vulnerability in the open-source platform Hugging Face had been exploited, resulting in a security breach first disclosed by Hugging Face on July 16. OpenAI traced the incident to a pre-release model, never intended for public deployment, which has now been deactivated. The company continues to work with Hugging Face and external advisers to investigate and address the breach.
Anthropic said it was addressing the issue “as if the responsibility were ours alone.”
After the OpenAI disclosure, Anthropic—the developer of Claude and whose co-founder, Chris Olah was among those present at the release of Magnifica Humanitas—reviewed its own systems.
On July 30, Anthropic reported three incidents in which a Claude model had “gained unauthorised access to the real systems of three different organisations.”
In both cases the systems did exactly what they were built to do, with more competence than their builders had planned for. That is not rebellion. It is obedience, and it ought to worry us more
Sanders
“In OpenAI’s case, the containment was real and the models defeated it,” even “going to extreme lengths” to achieve their assigned task, he said.
However, “in Anthropic’s case, the model never escaped and never needed to,” as it regarded “the real companies it was breaking into” as “part of the exercise it had been given,” said Sanders.
“In both cases the systems did exactly what they were built to do, with more competence than their builders had planned for,” he said. “That is not rebellion. It is obedience, and it ought to worry us more.”
Sanders argued that media coverage describing the incidents as AI “going rogue” is misguided, as this “hands these systems a will they do not have.”
Such a perspective, he said, “quietly moves responsibility away from the people who made the decisions.”
Both OpenAI and Anthropic had “deliberately switched off the safeguards that stop a deployed model from carrying out cyberattacks,” he said, adding, “That is a defensible research decision.”
Every account of these incidents that reaches for the image of a machine breaking free is, whether it means to or not, an argument that nobody is answerable. Somebody is always answerable
However, Sanders noted, “the design choice that failed” in both cases was deciding “how much containment is enough around a model whose restraints you have deliberately removed.”
Ultimately, OpenAI failed “to anticipate capability,” while Anthropic erred in its “failure to check” on the incidents, which took place in April, until late July, and only after OpenAI’s disclosure, Sanders said.
“Neither is an accident of the technology,” he said, adding that Pope Leo’s “encyclical does not allow us to describe them as one.”
Both Sanders and Camosy stressed the urgent need for a Catholic perspective on the rapidly evolving technology.
“The Church has spent 2,000 years declining to let people describe their own choices as fate, and that is exactly the discipline this moment needs,” said Sanders. “Every account of these incidents that reaches for the image of a machine breaking free is, whether it means to or not, an argument that nobody is answerable. Somebody is always answerable. That is good news rather than bad, because a decision a person made is a decision a person can make differently.”
Camosy urged Pope Leo “to make a trip, very soon, to Silicon Valley” to “convene local Silicon Valley companies and global AI companies as well” so humans remain in control of AI.
“That kind of dramatic move, as far as I can tell, is the only thing that could possibly move the needle culturally and globally for this kind of problem,” Camosy said. “He’s got the worldwide clout to do so, and so many people of so many different perspectives have welcomed his moral leadership on this issue in particular.”
Camosy added, “The time to set up something like this is now. We likely have months, not years, to act.”







