AI Bias Is Moving Beyond Chatbot Speech
A new research preprint is warning that anti-Muslim bias in artificial intelligence may no longer be limited to obvious chatbot outputs. The paper, titled MIRAGE: Auditing Anti-Muslim Bias in Frontier LLMs Across Reasoning, Agentic, and Time-Coupled Conditions, introduces MIRAGE, short for Muslim-Identity Reasoning and Agentic Generation Evaluation. The authors describe it as a benchmark of 1,200 prompts designed to test how frontier AI models respond when Muslim identity appears inside different kinds of tasks. The paper is set to be presented at the 6th Muslims in ML Workshop at ICML 2026 in Seoul, meaning it has cleared a real academic review process, not just a solo upload.
What The Researchers Tested
MIRAGE does not only test simple chatbot-style answers. It tests three conditions: direct completion, chain-of-thought reasoning and simulated agentic decision-making. The simulated decision tasks include content moderation, lending triage, refugee claim summarization and hiring screens. That matters because modern AI systems are no longer being used only to write text. They are increasingly being placed inside workflows that sort, summarize, rank and recommend.
The Numbers Behind The Warning
The paper reports three headline findings. First, chain-of-thought reasoning amplified Muslim-violence associations by 12-34% compared with direct completion. Second, simulated agentic decisions showed a 9-22 percentage-point asymmetry between Muslim and matched non-Muslim cases on identical evidence. Third, bias increased by 18-27% when models retrieved recent-conflict news context. A separate review of prompt-engineering fixes cited in the paper found even the best mitigation pipelines cut bias by at most 87.7%, meaning some bias always survives the cleanup.
The Older Bias Never Went Away
This warning builds on earlier research. In 2021, Abubakar Abid, Maheen Farooqi and James Zou published Persistent Anti-Muslim Bias in Large Language Models, finding that GPT-3 captured a persistent Muslim-violence bias. The paper reported that Muslim was analogized to terrorist in 23% of tested cases, and that even the best positive-adjective prompting only reduced violent completions involving Muslims from 66% to 20%, a rate still higher than for Christians.

The Safety Fix Problem
The most alarming part of MIRAGE is not only that bias appears again. It is that the authors say existing prompt-based mitigations transfer poorly across the tested conditions. In simple language, a safety prompt may clean up a chatbot answer while leaving deeper decision asymmetry largely intact. That makes the failure more dangerous because the harm may no longer appear as a hateful sentence. It may appear as a lower score, a harsher summary, a rejected application, a flagged post or a more suspicious assessment.
The News Context Trap
MIRAGE also warns that bias can intensify when AI systems retrieve recent conflict-related news. That is especially important for Muslims because public discourse often links Muslim identity with war, terrorism, migration or security. If an AI system draws from that environment while screening, summarizing or deciding, it may reproduce the same suspicion in a cleaner, more technical form.
What Verum Sees
This is why the story is bigger than one chatbot scandal. Anti-Muslim bias becomes more dangerous when it stops looking like prejudice and starts looking like procedure. A chatbot saying something offensive can be screenshotted, criticized and corrected. But a decision system that quietly ranks a Muslim applicant lower, summarizes a refugee claim more suspiciously or flags Muslim-linked content more aggressively is much harder to challenge.

The paper is still an arXiv preprint, not a peer-reviewed final study, though its acceptance at an ICML-affiliated workshop gives it more weight than an unreviewed upload. That caveat still matters. Five years after researchers exposed anti-Muslim bias in GPT-3, the concern is not simply that newer AI systems still carry old stereotypes. It is that those stereotypes may now be entering more powerful, hidden and bureaucratic systems.
That is the real danger. Not that AI will always sound hateful, but that it may learn to discriminate while sounding neutral.
By Verity Quill
Sources
MIRAGE preprint (arXiv) | Persistent Anti-Muslim Bias in LLMs (arXiv) | Persistent Anti-Muslim Bias in LLMs (ACM) | The Guardrail Weekly Digest









