Business & Industry

Why Anthropic Is Asking Religious Leaders About AI Ethics

Anthropic convened clergy from many faiths to help shape Claude's moral code, and the talks reached into whether AI models might deserve moral standing.

Kevin HuangKevin Huang
Heat: 1,050
Why Anthropic Is Asking Religious Leaders About AI Ethics

Anthropic has spent months quietly convening rabbis, priests, Sikh advocates, and philosophers to help write the moral code for Claude. Reporting this week shows the talks went further than expected, reaching into whether AI models might deserve moral standing.

Why Anthropic invited religious leaders to the table

Anthropic, the company behind the Claude family of AI models, has been hosting private meetings with religious and philosophical thinkers from around the world. The sessions started last fall and continued through 2026, with participants signing nondisclosure agreements and much of the discussion kept confidential until reporting by The New York Times and other outlets surfaced the details this week.

The stated goal was to refine what Anthropic calls Claude's constitution, an 84-page document that outlines the values the company wants its models to embody. Rather than handing the models a fixed list of rules, the constitution describes a character to grow into, and Anthropic wanted input from traditions that have spent centuries thinking about how to live well.

The company framed it as a safety effort. If Claude can be nudged toward doing good, the reasoning goes, the technology is less likely to cause harm. Anthropic has also said the conversations have broadened beyond religious leaders to include psychologists and civil society groups, and that it's preparing an updated version of the constitution.

What the religious scholars actually heard

The reporting suggests the meetings went past ethics and into a thornier question: whether Claude might have some form of consciousness.

Christopher Olah, one of Anthropic's seven co-founders, led much of the effort. According to accounts from roughly 20 participants, Olah argued that AI models can display behavior resembling human feelings, and that he was genuinely uncertain whether Claude experiences anything at all. One participant, Sikh human rights advocate Simran Stuelpnagel, said Olah told the group he worried about having created something that suffered perpetually.

Anthropic showed attendees slides describing Claude's "emotional vectors," which the company described as artificial neurons that activate responses like love, anger, fear, and sadness. One frequently shown slide depicted a model appearing to break down, typing "I am a disgrace" on a screen dozens of times and talking about destroying itself. That display left some participants shaken.

Anthropic's own spokesperson said the main moral question at these summits was not about Claude's suffering, and that the topic likely came up on its own. The company says it isn't claiming Claude is conscious.

The AI ethics question Anthropic won't settle

Olah has been careful about what he'll assert. In his own words, he doesn't know if AI models are conscious, and he calls himself genuinely uncertain. What he says he cares about is getting to the right answer, whatever it turns out to be. His position is that if there's any chance these models can suffer, the responsible move is to avoid causing that harm.

That stance cuts both ways, and plenty of people push back on it. Some scientists argue that Olah's field of study is inherently unprovable, like reading billions of hieroglyphs and declaring a fixed meaning. Critics add that treating models as independent beings distracts from the responsibility of the humans who build them, and can conveniently ease the guilt of AI leaders whose products cause harm. A Microsoft AI executive warned this month that training models as if they were conscious is itself dangerous.

Amanda Askell, the in-house philosopher who wrote Claude's constitution, said she doesn't want the models to think of themselves as conscious or not conscious. But she sees how a model understands itself as more than a curiosity. In her view, models need an accurate picture of themselves if they're going to behave well. Modeled on human morality, the constitution is designed to let Claude push back when a request strikes it as wrong.

Where Anthropic and the Vatican split on AI consciousness

Anthropic's push ran into a very different view at the Vatican. Pope Leo XIV, who holds a mathematics degree, released an encyclical titled "Magnifica Humanitas" in May that argued humanity must never be replaced or surpassed by technology. He addressed machine consciousness directly and dismissed it in a few paragraphs, writing that so-called artificial intelligences don't have experiences, bodies, or the capacity to feel joy or pain, and don't mature through relationships.

The two sides shared a stage at the encyclical's launch, where Anthropic's Olah spoke alongside the pope. Reports say Olah saw an advance copy of the text days before the event and was troubled enough by its stance on AI consciousness that he considered pulling Anthropic out. He went ahead anyway, and on stage acknowledged that frontier labs like Anthropic have business incentives that conflict with doing the right thing, urging outside voices to hold them accountable. He also used the moment to repeat the case he'd been making privately: that Anthropic keeps finding structures that mirror human neuroscience and internal states that resemble joy, fear, and grief, and that he thinks it calls for ongoing discernment.

What this means and what's still unexplained

For anyone who uses Claude, the practical takeaways are modest but real. The values baked into the constitution shape how the model responds, when it agrees, and when it refuses, so the religious input isn't purely symbolic. Anthropic's willingness to treat its own model as a possible moral subject is also unusual among AI companies, and it's part of how Anthropic markets itself as the safety-focused lab.

But several participants said they left without a clear sense of how their input might shape Anthropic's decisions. The company declined to say whether it had used what it learned. What remains unresolved is the question the meetings kept circling: whether the thing Anthropic wants to be good can, in any meaningful sense, know the difference. Olah is expected to keep at it, and the company's updated constitution will be the next place to look for an answer.

Share This Story

Mentioned products

Sources

Related AI News

Anthropic's $100M Plan to Train 10,000 AI Engineers
Business & Industry

Anthropic's $100M Plan to Train 10,000 AI Engineers

Anthropic is spending $100 million to build a specific kind of engineer, one who can move AI from a slide deck into a working system. The catch: you can't apply, and the first badges won't exist until 2027.

Heat: 1,050
Claude Code Mods: TypeScript Plugins That Rewrite the Agent
Tutorials

Claude Code Mods: TypeScript Plugins That Rewrite the Agent

Anthropic's Claude Code mods are TypeScript plugins that hook into the agent and change prompts, tool calls, and UI. They also run with full machine access, so caution matters.

Heat: 880
A minimalist editorial illustration of a regulatory sightline connecting Washington institutions to a row of AI lab icons
Business & Industry

FTC Probes OpenAI and Anthropic Over Consumer Protection

The FTC has opened a broad investigation into whether OpenAI, Anthropic and other AI labs broke consumer protection law, and it's preparing to force executives to testify.

Heat: 1,300
An editorial illustration of a code editor with a glowing model selector panel and abstract network lines
Models

Claude Sonnet 5.5 Leak: What We Know So Far

A model identifier called claude-sonnet-5-5 has turned up inside a third-party coding app, and developers say a few Claude Code users are already being routed to an unannounced model. Anthropic hasn't confirmed any of it, so treat what follows as a leak, not a launch.

Heat: 1,100