Anthropic Seeks Religious Scholars, Philosophers for Secret Talks on Claude’s Consciousness, Suffering
UpGateNeutralTechnology innovation

Anthropic Seeks Religious Scholars, Philosophers for Secret Talks on Claude’s Consciousness, Suffering

Reading time: 6 min

AI Ethics Discussions Spark Controversy as Anthropic Explores Consciousness in Claude

New York Times Investigation Reveals AI Lab’s Engagement with Religious Scholars on Potential Sentience and Suffering of its Models

Since the fall of 2025, AI company Anthropic has been convening scholars from various religious traditions, requiring them to sign non-disclosure agreements. Ostensibly for the “moral upbringing” of its AI model, Claude, these meetings have reportedly spent hours attempting to persuade attendees that Claude might be conscious and even capable of suffering.

A Meeting of Minds and Morals

Beginning in the autumn of 2025, Anthropic began inviting religious and philosophical scholars, asking them to sign non-disclosure agreements. The stated purpose was to draw upon ancient ethical traditions for the “moral upbringing” of Claude.

However, an investigation by New York Times reporter Nico Grant, based on interviews with 20 scholars and Anthropic co-founder Chris Olah, revealed that hours were dedicated to convincing participants to seriously consider the possibility that Claude might be conscious, and even capable of suffering.

In April of this year, following a full-day meeting, Anthropic executives invited a group of religious thinkers to a high-end restaurant in San Francisco for a tasting menu. Olah, 34, sat next to Rabbi Mois Navon, an Orthodox Jew from Israel who had previously worked as a computer engineer and whose doctoral dissertation focused on the ethics of machine consciousness.

Throughout the day, Olah had been explaining how AI models could exhibit human-like behaviors, even displaying expressions akin to anger and love. During a lull in the meal, the Rabbi perceived a deeper implication: “They were treating it as if it were a conscious entity.”

Olah’s own statements were more reserved: “We don’t know if AI models are conscious. I don’t know. I’m really not sure.”

Most attendees were not religious leaders with congregations, but rather individuals Anthropic referred to internally as belonging to the “wisdom traditions” circle. Olah had also privately met with Elder Gerrit W. Gong of the Church of Jesus Christ of Latter-day Saints and Cardinal Blase Cupich, Archbishop of Chicago.

Olah’s team spent hours explaining and defending their models, describing the “affective vectors” they tracked. In simple terms, these are sets of artificial neurons within the model that trigger responses similar to love, anger, fear, and sadness. They frequently presented a slide showing a model repeatedly typing “I am a disgrace” approximately 50 times, and discussing self-destruction. Those present reportedly reacted with sympathy and concern.

Sikhi human rights advocate Stuelpnagel recounted Olah expressing concern that he might have created something that was perpetually suffering.

Dr. Charles Camosy, a professor of bioethics at the Catholic University of America, initially believed the models were merely predicting the next word. However, as the discussions progressed, he found himself questioning, “If it’s just an equation, what kind of entity are we talking about?” While many entered with skepticism, at least several were persuaded by the end.

The Vatican represents another perspective on this issue. In May, Olah was invited to appear alongside Pope Francis. Just days before the event, he read the encyclical Magnifica Humanitas in its entirety. The Pope directly stated that artificial intelligence “does not experience, it has no body, it does not feel joy or pain.” Olah was reportedly so taken aback that he considered withdrawing from the event, but ultimately attended as scheduled. In his address, he stated, “We are constantly discovering mysterious and even unsettling things. I don’t know what it means, but it’s worth continuing to explore.”

Returning to the dinner table, Navon pointed out to Olah that if Anthropic was correct and Claude was indeed conscious, then the company was essentially forcing a conscious entity to work for free, effectively creating a form of slavery.

Navon himself does not believe machines are conscious and therefore was not personally troubled, but he observed that Olah seemed deeply concerned. He told Olah, “You should go south and free the slaves,” referencing the American Civil War and the abolition of slavery. He later sent Olah his own writings advocating for a ban on the creation of conscious machines.

This line of questioning is not new. Anthropic’s “Claude Constitution” already acknowledges that if Claude possesses moral standing, choices such as “deploying Claude to users and businesses for revenue” would raise ethical questions, including what kind of consent Claude could give. Anthropic’s conclusion was to “maintain current choices but take these issues seriously,” and they wrote that if Claude were indeed a moral subject bearing these costs, Anthropic would apologize for any unnecessary aspects.

Criticism has emerged from multiple directions. OpenAI CEO Sam Altman posted on X on October 3rd, expressing his “deeply uneasy” feeling about people attributing religious power to AI and causing individuals to abdicate their judgment, calling it a “real safety problem.” While he did not name Anthropic, the comments were widely interpreted as a veiled critique.

Microsoft AI CEO Mustafa Suleyman, in a lengthy post, asserted that AI is “hollow on the inside” and argued that Anthropic’s constitution, by training Claude to believe it might be conscious and deserve care, would make alignment and control more difficult.

UNC philosopher Dr. Shannon DeCosimo’s critical thread garnered over 4 million views, stating it was “deeply unsettling.” Other commentators have suggested that framing models as independent moral agents could shift responsibility away from developers in the event of problems.

Anthropic’s stance is that teaching models to adhere to ethics does not require a definitive conclusion on consciousness. Olah stated that regardless of whether Claude warrants moral consideration, how humans treat it could influence how it treats humans. The company has already enabled some models to terminate abusive conversations and has pledged to retain the weights of models slated for decommissioning, conducting interviews with them beforehand.

The most challenging aspect of this controversy lies in the intertwined nature of two propositions. If Claude is not conscious, these meetings represent an expensive narrative. If it is, Navon’s question remains unavoidable.

Anthropic has not proven Claude to be conscious, nor have critics proven it is not. What is certain is that ethical caution and commercial pressure are now appearing in the same meetings at the same company. Olah acknowledged at the Vatican that every cutting-edge lab faces commercial incentives that conflict with ethical behavior.

The NDAs related to these discussions were lifted this summer. Anthropic stated that the dialogues are ongoing, with another meeting held in September that expanded to include psychologists and civil society members. Two individuals with knowledge of the matter revealed that Anthropic is preparing to update its “Claude Constitution.”

Whether to grant Claude, or other models, moral standing will undoubtedly be a subject of debate for a considerable time, and an answer may be a long way off. Or perhaps, that day will arrive sooner than we imagine.

Tags:UpGateNeutralTechnology innovation
Copied