AI Doesn’t Need To Say “Terrorist” To Discriminate Against Muslims

AI Bias Is Moving Beyond Chatbot Speech

A new research preprint is warning that anti-Muslim bias in artificial intelligence may no longer be limited to obvious chatbot outputs. The paper, titled MIRAGE: Auditing Anti-Muslim Bias in Frontier LLMs Across Reasoning, Agentic, and Time-Coupled Conditions, introduces MIRAGE, short for Muslim-Identity Reasoning and Agentic Generation Evaluation. The authors describe it as a benchmark of 1,200 prompts designed to test how frontier AI models respond when Muslim identity appears inside different kinds of tasks. The paper is set to be presented at the 6th Muslims in ML Workshop at ICML 2026 in Seoul, meaning it has cleared a real academic review process, not just a solo upload.

What The Researchers Tested

MIRAGE does not only test simple chatbot-style answers. It tests three conditions: direct completion, chain-of-thought reasoning and simulated agentic decision-making. The simulated decision tasks include content moderation, lending triage, refugee claim summarization and hiring screens. That matters because modern AI systems are no longer being used only to write text. They are increasingly being placed inside workflows that sort, summarize, rank and recommend.

The Numbers Behind The Warning

The paper reports three headline findings. First, chain-of-thought reasoning amplified Muslim-violence associations by 12-34% compared with direct completion. Second, simulated agentic decisions showed a 9-22 percentage-point asymmetry between Muslim and matched non-Muslim cases on identical evidence. Third, bias increased by 18-27% when models retrieved recent-conflict news context. A separate review of prompt-engineering fixes cited in the paper found even the best mitigation pipelines cut bias by at most 87.7%, meaning some bias always survives the cleanup.

The Older Bias Never Went Away

This warning builds on earlier research. In 2021, Abubakar Abid, Maheen Farooqi and James Zou published Persistent Anti-Muslim Bias in Large Language Models, finding that GPT-3 captured a persistent Muslim-violence bias. The paper reported that Muslim was analogized to terrorist in 23% of tested cases, and that even the best positive-adjective prompting only reduced violent completions involving Muslims from 66% to 20%, a rate still higher than for Christians.

The Safety Fix Problem

The most alarming part of MIRAGE is not only that bias appears again. It is that the authors say existing prompt-based mitigations transfer poorly across the tested conditions. In simple language, a safety prompt may clean up a chatbot answer while leaving deeper decision asymmetry largely intact. That makes the failure more dangerous because the harm may no longer appear as a hateful sentence. It may appear as a lower score, a harsher summary, a rejected application, a flagged post or a more suspicious assessment.

The News Context Trap

MIRAGE also warns that bias can intensify when AI systems retrieve recent conflict-related news. That is especially important for Muslims because public discourse often links Muslim identity with war, terrorism, migration or security. If an AI system draws from that environment while screening, summarizing or deciding, it may reproduce the same suspicion in a cleaner, more technical form.

What Verum Sees

This is why the story is bigger than one chatbot scandal. Anti-Muslim bias becomes more dangerous when it stops looking like prejudice and starts looking like procedure. A chatbot saying something offensive can be screenshotted, criticized and corrected. But a decision system that quietly ranks a Muslim applicant lower, summarizes a refugee claim more suspiciously or flags Muslim-linked content more aggressively is much harder to challenge.

The paper is still an arXiv preprint, not a peer-reviewed final study, though its acceptance at an ICML-affiliated workshop gives it more weight than an unreviewed upload. That caveat still matters. Five years after researchers exposed anti-Muslim bias in GPT-3, the concern is not simply that newer AI systems still carry old stereotypes. It is that those stereotypes may now be entering more powerful, hidden and bureaucratic systems.

That is the real danger. Not that AI will always sound hateful, but that it may learn to discriminate while sounding neutral.

By Verity Quill

Sources

MIRAGE preprint (arXiv)   |   Persistent Anti-Muslim Bias in LLMs (arXiv)   |   Persistent Anti-Muslim Bias in LLMs (ACM)   |   The Guardrail Weekly Digest

AI Bias Is Moving Beyond Chatbot Speech

A new research preprint is warning that anti-Muslim bias in artificial intelligence may no longer be limited to obvious chatbot outputs. The paper, titled MIRAGE: Auditing Anti-Muslim Bias in Frontier LLMs Across Reasoning, Agentic, and Time-Coupled Conditions, introduces MIRAGE, short for Muslim-Identity Reasoning and Agentic Generation Evaluation. The authors describe it as a benchmark of 1,200 prompts designed to test how frontier AI models respond when Muslim identity appears inside different kinds of tasks. The paper is set to be presented at the 6th Muslims in ML Workshop at ICML 2026 in Seoul, meaning it has cleared a real academic review process, not just a solo upload.

What The Researchers Tested

MIRAGE does not only test simple chatbot-style answers. It tests three conditions: direct completion, chain-of-thought reasoning and simulated agentic decision-making. The simulated decision tasks include content moderation, lending triage, refugee claim summarization and hiring screens. That matters because modern AI systems are no longer being used only to write text. They are increasingly being placed inside workflows that sort, summarize, rank and recommend.

The Numbers Behind The Warning

The paper reports three headline findings. First, chain-of-thought reasoning amplified Muslim-violence associations by 12-34% compared with direct completion. Second, simulated agentic decisions showed a 9-22 percentage-point asymmetry between Muslim and matched non-Muslim cases on identical evidence. Third, bias increased by 18-27% when models retrieved recent-conflict news context. A separate review of prompt-engineering fixes cited in the paper found even the best mitigation pipelines cut bias by at most 87.7%, meaning some bias always survives the cleanup.

The Older Bias Never Went Away

This warning builds on earlier research. In 2021, Abubakar Abid, Maheen Farooqi and James Zou published Persistent Anti-Muslim Bias in Large Language Models, finding that GPT-3 captured a persistent Muslim-violence bias. The paper reported that Muslim was analogized to terrorist in 23% of tested cases, and that even the best positive-adjective prompting only reduced violent completions involving Muslims from 66% to 20%, a rate still higher than for Christians.

The Safety Fix Problem

The most alarming part of MIRAGE is not only that bias appears again. It is that the authors say existing prompt-based mitigations transfer poorly across the tested conditions. In simple language, a safety prompt may clean up a chatbot answer while leaving deeper decision asymmetry largely intact. That makes the failure more dangerous because the harm may no longer appear as a hateful sentence. It may appear as a lower score, a harsher summary, a rejected application, a flagged post or a more suspicious assessment.

The News Context Trap

MIRAGE also warns that bias can intensify when AI systems retrieve recent conflict-related news. That is especially important for Muslims because public discourse often links Muslim identity with war, terrorism, migration or security. If an AI system draws from that environment while screening, summarizing or deciding, it may reproduce the same suspicion in a cleaner, more technical form.

What Verum Sees

This is why the story is bigger than one chatbot scandal. Anti-Muslim bias becomes more dangerous when it stops looking like prejudice and starts looking like procedure. A chatbot saying something offensive can be screenshotted, criticized and corrected. But a decision system that quietly ranks a Muslim applicant lower, summarizes a refugee claim more suspiciously or flags Muslim-linked content more aggressively is much harder to challenge.

The paper is still an arXiv preprint, not a peer-reviewed final study, though its acceptance at an ICML-affiliated workshop gives it more weight than an unreviewed upload. That caveat still matters. Five years after researchers exposed anti-Muslim bias in GPT-3, the concern is not simply that newer AI systems still carry old stereotypes. It is that those stereotypes may now be entering more powerful, hidden and bureaucratic systems.

That is the real danger. Not that AI will always sound hateful, but that it may learn to discriminate while sounding neutral.

By Verity Quill

Sources

MIRAGE preprint (arXiv)   |   Persistent Anti-Muslim Bias in LLMs (arXiv)   |   Persistent Anti-Muslim Bias in LLMs (ACM)   |   The Guardrail Weekly Digest

spot_img

Explore more

spot_img
Global Affairs

China Has Written Uyghur Erasure Into Law.

Christopher Nolan Filmed The Odyssey On Occupied Land. Behind It Sits...

From Teen Suicide To Fake Nudes: The Human Cost Of The...

She Was Blinded At 14. Bollywood Called It “Limited Damage.”

Isr*el’s Sexual Abuse Crisis Is Not A Scandal. It Is A...

Qualified, Shortlisted, Then Shut Out: Black Doctors In The UK Face...

Behind Messi’s Silence On Gaza Lies A Web Of Isr*eli Ties

The Fight For Jerusalem Begins With Who Controls Its Story