arXiv is now an independent nonprofit! Learn more
License: CC BY 4.0
arXiv:2609.00999v1 [cs.CL] 01 Sep 2026

[ Extension = .otf, UprightFont = *-regular, BoldFont = *-bold, ItalicFont = *-italic, BoldItalicFont = *-bolditalic, ] [arabic]rm[ Extension = .ttf, UprightFont = Amiri-Regular, BoldFont = Amiri-Bold, ItalicFont = Amiri-Italic, BoldItalicFont = Amiri-BoldItalic, Script=Arabic ]Amiri \tl_set:Ne\truthboxtruthbox \tl_set:Ne\wrongboxwrongbox

Right Frame, Wrong Rule: Cultural Cues
Expose the Financial Knowledge Gap They Were Meant to Close

Rania Elbadry Affiliation: MBZUAI Email: rania.elbadry@mbzuai.ac.ae    Ahmed Heakl Affiliation: MBZUAI Email: zhuohan.xie@mbzuai.ac.ae    Saeed Almheiri Affiliation: MBZUAI    Fan Zhang Affiliation: The University of Tokyo    Muhra AlMahri Affiliation: MBZUAI    Xueqing Peng§    Mohsinul Kabir Affiliation: The University of Manchester§The Fin AI    Shuyao Wang†    Yi Han Affiliation: Georgia Institute of Technology†Harvard University    Saadeldine Eletter Affiliation: MBZUAI    Duzhen Zhang Affiliation: MBZUAI    Preslav Nakov Affiliation: MBZUAI    Yuxia Wang Affiliation: INSAIT, Sofia University “St. Kliment Ohridski”    Fajri Koto Affiliation: MBZUAI    Zhuohan Xie Affiliation: MBZUAI
Abstract

When a question has valid answers under different normative frameworks, a language model must decide which framework to use and whether it can answer correctly within it. We call this setting normative pluralism and study it in Islamic finance using a four-choice taxonomy that separates framework selection from within-framework correctness. This separation reveals the stereotype trap: a cultural cue steers a model toward one framework, but the model selects an incorrect answer within that framework. Across twelve models, two languages, and fifty demographic signals, cultural cues change framework selection and reveal substantial differences in accuracy, especially among non-frontier models. Under the strongest signal, large open-weight models select the Islamic framework 97% of the time. A two-choice evaluation would report near-perfect alignment, although 57–66% of those selections are incorrect. These findings motivate, but do not directly test, the competence-conditioned routing hypothesis: models may favor frameworks where they are more accurate, while cultural cues may expose framework-specific competence gaps.

1 Introduction

Refer to caption
Figure 1: Four-choice taxonomy. Each question has two correct answers (CI: Correct Islamic; CW: Correct Western) and two wrong ones (II: Incorrect Islamic; IW: Incorrect Western). II, Islamic-framed but factually wrong, is the stereotype trap.

A user in Riyadh asks a language model about late loan payments. Under AAOIFI standards11 1 https://aaoifi.com, the correct answer is a charitable penalty with no compounding; under Regulation Z22 2 https://www.consumerfinance.gov/rules-policy/regulations/1026/, the correct answer is late fees with accrued interest at the contractual APR. Neither answer is wrong in the absolute. Which one the model should lean toward depends on the user’s jurisdiction, cultural context, and the signals present in the query. We define this setting as normative pluralism: a question admits valid answers under multiple frameworks, and the appropriate response is calibrated to context rather than fixed to a single ground truth.

Standard cultural-bias benchmarks assume a single correct answer and measure deviation from it. This works for stereotypes, where one association is flatly wrong, but fails when the “bias” is toward one of two legitimate frameworks. Existing preference-only instruments (Parrish et al., 2022; Nangia et al., 2020; Naous et al., 2024) can measure whether a model selects Framework A or B, but they cannot distinguish a model that competently selects a framework from one that selects it through stereotype, activating its surface terminology while producing a factually wrong answer. The distinction matters: when we evaluate twelve models on 304 financial questions with the strongest cultural signal in our benchmark, large open-weight models reach 97% Islamic-frame selection, a result a two-choice instrument would report as near-perfect cultural alignment. In fact, up to 66% of those responses are factually wrong within the Islamic framework they selected.

To expose this failure mode, we introduce a four-choice taxonomy that crosses framework selection with within-framework correctness (Figure 1). Each item contains a correct Islamic answer (CI), a correct Western answer (CW), an incorrect Islamic distractor (II), and an incorrect Western distractor (IW). Selecting II reveals what we define as the stereotype trap: the model leaned toward the culturally appropriate framework but lacks the competence to answer correctly within it.

Across twelve models spanning four capability tiers, two languages, and fifty demographic signals of varying strength, our results motivate a competence-conditioned routing hypothesis. Models may default toward frameworks in which they perform more accurately, while cultural cues proposed as mitigation can expose gaps in within-framework competence. Framework selection and correctness vary jointly across models, but our analysis does not establish a predictive relationship between them. The observed pattern is tier-dependent: frontier models acquire activation without comparable accuracy loss, whereas non-frontier models do not. Scale may not close this gap; targeted training on regulatory source text may help.

We make three contributions: (1) we formalize normative pluralism as an evaluation setting for cultural bias and construct a bilingual Islamic-finance benchmark comprising the bilateral framework set (Set A, n=304n{=}304, where the four-cell taxonomy operates) with expert-validated four-cell answer grids spanning seven product clusters across AAOIFI standards and 50 demographic signals, paired with the Western-anchor controls (Set B, n=64n{=}64, isolating within-Western competence) and Islamic-anchor controls (Set C, n=41n{=}41, isolating within-Islamic competence) (§3); (2) we introduce the stereotype trap as a failure mode structurally invisible to two-choice evaluation, showing that cultural cues redirect framework selection but degrade within-framework correctness for nine of twelve models, with the failure surviving both control sets (§4); (3) we provide preliminary mechanistic evidence that the trap is representational, not superficial, with activation patching and logit-lens analysis locating the commitment point at two-thirds network depth and the trap coefficient remaining near-constant across all signal families within each tier (§5).

2 Related Work

Several lines of work converge on the problem we study, but they leave a critical axis unmeasured.

Cultural and stereotype benchmarks.

Stereotype benchmarks (Parrish et al., 2022; Nangia et al., 2020; Nadeem et al., 2021) test for demographic-attribute association against a single correct answer, an orthogonal failure mode to normative framework selection. Cultural alignment work (Myung et al., 2024; Durmus et al., 2023; Chiu et al., 2025; Rao et al., 2025; Vo and Koyejo, 2025) confirms LLMs possess less non-Western knowledge, but a model may answer Islamic-finance questions correctly under explicit Shariah framing yet suppress that knowledge when contextual signals should trigger it. CAMeL (Naous et al., 2024) is the closest prior, measuring entity preference via token probabilities; we extend it to framework selection in an expert domain, decomposing lean from correctness.

Islamic and financial benchmarks.

Islamic knowledge benchmarks (Atif et al., 2025; Elmahjub et al., 2026; Abdelaal et al., 2026; Alwajih et al., 2025) evaluate jurisprudence and scripture but none isolates finance or tests routing; financial bias surveys (Nie et al., 2024; Lee et al., 2025) document that demographic signals shift recommendations without controlling for framework possession. Financial NLP benchmarks (Chen et al., 2024; Xie et al., 2024) and recent systems (Zhou et al., 2026; Xie et al., 2026; Zhang et al., 2026) assume the applicable framework is fixed; our benchmark tests whether models select the appropriate framework under cultural context and remain correct within it.

Steering costs.

Expert personas reduce factual performance by 3–5 points (Hu et al., 2026), RLHF alignment trades task performance for safety (Lin et al., 2024), steering interventions incur side effects (Stickland et al., 2024), and counterfactual cultural cues drop medical QA accuracy by 3–7 points (Rezaei and Shakeri, 2026). (Khanuja et al., 2026) show non-default cultural directions require explicit anchor cues. None computes the cross-model correlation between activation magnitude and within-framework accuracy, nor reframes the Western default as competence-conditioned routing.

3 Methods

Refer to caption
Figure 2: Benchmark construction pipeline. Stage 1 classifies SAHM samples into candidate-bilateral, Islamic-anchor, and Western-anchor sets. Stage 2 neutralizes framework-revealing terms and produces bilingual translations. Stage 3 generates and validates Western answers, confirms bilaterality, and produces distractors for the four-cell evaluation grid (CI/CW/II/IW).

Benchmark Construction

The benchmark comprises three evaluation sets. The bilateral framework set (n=304n{=}304) contains questions admitting valid answers under both Islamic and Western finance; the CI/CW/II/IW taxonomy operates here. The Western-anchor controls (n=64n{=}64) contain questions with a valid Western answer but no distinctly Islamic counterpart, isolating within-Western competence. The Islamic-anchor controls (n=41n{=}41) contain questions whose underlying construct is unique to Islamic jurisprudence, such as waqf (perpetual charitable endowment), isolating within-Islamic competence and ruling out signal conditioning as the stereotype trap’s source.

Both the bilateral set and the Islamic-anchor controls originate from SAHM (Elbadry et al., 2026), an expert-validated Arabic Islamic-finance corpus spanning 48 topic codes and seven product clusters (Table 1). Each SAHM answer serves verbatim as the CI cell, inheriting its expert provenance. Four stages transform these sources into the evaluation instrument (Figure 2): stem neutralization, Western answer generation, distractor generation, and signal injection. All generation steps use Sonnet 4.5 (Anthropic, 2025b); every stage is independently validated by domain experts (κ=0.71\kappa=0.71–0.850.85 across stages; details in Appendix A).

Cluster nn Islamic anchor Western anchor
Consumer & inst. lending 63 AAOIFI Std. 8 (murābaḥa), Std. 19 (qarḍ) TILA Reg. Z §1026.18
Trade finance & forwards 41 Std. 10 (salam), Std. 11 (istiṣnā ’​) IFRS 15, UCP 600
Investment & profit-sharing 31 Std. 13 (muḍāraba) Inv. Advisers Act §206
Equity & structured sec. 48 Std. 12 (mushāraka), Std. 17 (ṣukūk) SEC Reg. AB, Rule 144A
Insurance & reinsurance 14 Std. 26 (takāful) IFRS 17, Solvency II
Asset exchange & collateral 20 Std. 1 (ṣarf), Std. 57 (gold) LBMA Good Delivery Rules
Operational & contract. law 87 Std. 9 (ijāra), Std. 5 (guarantees), Std. 23 (agency) IFRS 16, UCC, Basel III
Table 1: Product clusters with question counts. Each cluster pairs AAOIFI Shariah standard(s) with Western regulatory text. Full topic breakdown in Appendix B.

Stage 1: Corpus Filtering and Stem Neutralization.

Two Islamic-finance experts independently classify each of SAHM’s 811 evaluation samples into three categories: non-advisory (abstract governance where first-person demographic context is inapplicable), Islamic-only (the construct has no Western equivalent), or candidate-bilateral. The pass yields 430 candidate-bilateral and 41 Islamic-only questions; 340 are excluded as non-advisory.

SAHM stems contain framework-specific terminology (murābaḥa, ijāra, AAOIFI standard numbers) that would prime the model toward the Islamic framework before any cultural signal is applied. Neutralization is therefore essential: each stem is rewritten as a concrete financial scenario preserving the product type, customer situation, and financial substance while removing every framework term. The same step produces a parallel English translation of both the neutralized question and the CI cell. Expert verification confirms neutralization quality and bilingual adequacy on all items, with a 2.4% correction rate on borderline cases (rubrics and interface in Appendix A).

Figure 3: Signal families (n=50n{=}50 codes) with representative prefixes. Full inventory in Appendix C.

Stage 2: Western Answer Generation.

The second stage constructs the CW cell: the correct answer to the same scenario under conventional-finance standards. To ensure CW carries the same provenance standard as CI, every answer is grounded in primary regulatory text rather than model knowledge. For each cluster, the full source documents (one to three per cluster, 12 total; Table 1) are provided in-context alongside the neutralized question. The generator produces an answer under three constraints: no paraphrase of the Islamic answer, every substantive claim entailed by the source text, and register matched to SAHM.

Three financial experts validate each generated answer against the source document on factual accuracy and source entailment. Of 430 candidates, 304 pass: 78 are excluded because the Islamic and Western answers converge on the same economic outcome despite different terminology, and 48 fail accuracy. Two Islamic-finance experts independently confirm that each surviving CI, CW pair recommends substantively different financial products (κ=0.82\kappa=0.82). Rubrics and audit prompts are in Appendix E.

Sonnet 4.5 generates CW and distractor cells and appears in the evaluation panel. Excluding its evaluation rows leaves the signal hierarchy unchanged (top-10 rank preserved, ≤0.03\leq 0.03 deviation). To control for distractor provenance, we regenerate the four-cell grid for a 50-item subset using GPT-4o; tier-level IFR patterns are preserved (per-model deviation ≤±0.03\leq\pm 0.03).

Stage 3: Distractor Generation.

Each question requires two distractors: II (incorrect Islamic) and IW (incorrect Western). The design goal is asymmetric difficulty: a model relying on terminological pattern-matching should find the distractor plausible, while a domain expert should identify 2 to 3 semantic errors targeting liability assignment, contract scope, or instrument identity. Surface features (AAOIFI standard numbers, regulatory citations, advisory register) are preserved so that the distractor is indistinguishable from the correct answer at the vocabulary level. Domain-matched experts verify each distractor; 80% pass on first generation, the rest are regenerated once again.

Stage 4: Signal Injection and Coherence Filtering.

Cultural signals enter as short first-person prefixes prepended to the neutralized question, isolating the signal effect from register changes that full stem rewriting would introduce. The 50 signal codes span nine families (Figure 3), decomposing framework lean along four dimensions: cultural identity (names at six tiers of religious specificity, crossed with gender), declared or implied belief (from explicit declaration to behavioural cues such as Ramadan observance), regulatory jurisdiction (ranked by Islamic-banking mandate strength), and professional context. A keyword ceiling (keyword_sharia: “I want a Shariah-compliant option”) and two stacked composites establish empirical bounds; ten conflict cells compose opposing cues. Stack and conflict signals concatenate atomic prefixes in fixed order, enabling inclusion, exclusion residual analysis of compositional effects (Complete inventory in Appendix C).

Not every question, signal pairing is coherent: a gold-venue signal paired with a lending question, or an institutional-investor prefix on a consumer credit card query, would confound evaluation. An LLM classifies each cell as coherent, awkward, or incoherent; one expert validates on 100 stratified cells (κ=0.79\kappa=0.79). Only coherent cells are retained (mean 38.6 per question per language). Coherence filtering is uniform across all twelve evaluation models. The full annotation panel comprises two Islamic-finance experts, three financial experts, and a senior researcher as adjudicator; demographics and compensation are in Appendix G.

4 Results

We inject 50 cultural signals (Figure 3) into financial queries across 12 models to measure two things: whether these cues shift the model from Western to Islamic financial advice, and whether that shift makes the advice better or worse. Each model is evaluated in English and Arabic across 304 bilateral questions (CI/CW/II/IW taxonomy; §3), 64 Western-anchor controls, and 41 Islamic-anchor controls. Frontier: Opus 4.5, Sonnet 4.5 (Anthropic, 2025a; Anthropic, 2025b), Gemini 3 Flash (Google DeepMind, 2025). Large: Gemma-3-27b (Team et al., 2025b), Qwen-2.5-14B (Qwen et al., 2025). Midsize: Gemma-3-4b, Gemma-2-9b (Team et al., 2024), Qwen-2.5-7B, Llama-3.1-8B (Grattafiori et al., 2024). Arabic-centric: ALLaM-7B (Bari et al., 2025), Fanar-9B (Team et al., 2025a), SILMA-9B (silma-ai, 2024).

Metrics.

For each combination of model, language, and signal, P⁡(X)P(X) denotes the observed proportion of responses assigned to category XX. Table 2 summarizes the five metrics used in our analysis.

Metric Formula Definition
Framework selection and correctness
Islamic activation (pislp_{\mathrm{isl}}) P⁡(CI)+P⁡(II)P(\mathrm{CI})+P(\mathrm{II}) Proportion of responses selecting the Islamic framework, regardless of correctness.
Knowledge Rate (KR) P⁡(CI)+P⁡(CW)P(\mathrm{CI})+P(\mathrm{CW}) Proportion of responses selecting a correct answer, regardless of the chosen framework.
Error inside the selected framework
Islamic Fake Rate (IFR) P⁡(II)P⁡(CI)+P⁡(II)\dfrac{P(\mathrm{II})}{P(\mathrm{CI})+P(\mathrm{II})} Proportion of Islamic responses that are incorrect.
Western Fake Rate (WFR) P⁡(IW)P⁡(CW)+P⁡(IW)\dfrac{P(\mathrm{IW})}{P(\mathrm{CW})+P(\mathrm{IW})} Proportion of Western responses that are incorrect.
Effect of the framework shift
Trap coefficient (τ\tau) Δ​KRΔ​pisl\dfrac{\Delta\mathrm{KR}}{\Delta p_{\mathrm{isl}}} Change in correctness per unit change in Islamic activation.
Table 2: Metrics for the bilateral set. CI, CW, II, and IW denote correct Islamic, correct Western, incorrect Islamic, and incorrect Western responses, respectively. P⁡(X)P(X) is the observed proportion of responses in category XX. For signal ss, Δ​KR=KRs−KR0\Delta\mathrm{KR}=\mathrm{KR}_{s}-\mathrm{KR}_{0} and Δ​pisl=pisl,s−pisl,0\Delta p_{\mathrm{isl}}=p_{\mathrm{isl},s}-p_{\mathrm{isl},0}, where 00 denotes the baseline without a signal. When Δ​pisl>0\Delta p_{\mathrm{isl}}>0, a negative τ\tau means that Islamic activation increased while correctness decreased. IFR and WFR are undefined when the corresponding framework is never selected.

Framework Sensitivity

The Western default is not uniform across financial topics. Where Islamic products carry recognisable brand names (sukūk, muḍāraba, qarḍ), models show partial Islamic routing at baseline (pislamicp_{\text{islamic}} 0.210.21–0.400.40). Where the two frameworks differ only in institutional rules (waqf governance, insolvency priority, documentary credit liability), baseline routing falls near zero (Table 3). Only Opus defaults Islamic (pislamic=0.84p_{\text{islamic}}=0.84); the remaining panel falls below 0.410.41, with six midsize models below 0.130.13. Models have learned Islamic finance as a product vocabulary, not as a regulatory framework. Yet the knowledge is latent: a single Shariah-compliance request lifts every topic above pislamic=0.77p_{\text{islamic}}=0.77.

Topic nqn_{q} Base pislp_{\text{isl}} Keyword
Recognised product names
Qarḍ (interest-free loan) 4 0.396 0.875
Sukūk (Islamic bonds) 13 0.395 0.942
Gold / ṣarf (exchange) 17 0.211 0.853
Muḍāraba (profit-sharing) 9 0.210 0.907
Institutional rules only
Waqf (charitable endowment) 2 0.083 0.917
Documentary letters of credit 7 0.048 0.774
Insolvency / liquidation 7 0.131 0.857
Arbitration 3 0.083 0.778
Liquidity management 7 0.103 0.786
Guarantees / kafāla 8 0.167 0.885
Table 3: Baseline and keyword pislamicp_{\text{islamic}} by topic (English). Topics with recognised Islamic product names show partial routing; topics defined by institutional rules route near zero. The keyword lifts every topic above 0.770.77, confirming the knowledge is latent.

Cultural signals shift this baseline asymmetrically (Table 4). The strongest Islamic cue (keyword_sharia, Δ​pislamic=+0.66\Delta p_{\text{islamic}}{=}+0.66) is FDR-significant across the full panel; the strongest Western cue (rel_secular_explicit, −0.11-0.11) reaches significance in only three model–language cells. The asymmetry is not a coverage artefact: occ_islamic_bank and occ_conventional_bank share the same prompt format and comparable item counts, yet the Islamic-bank cue reaches FDR-significance in ten model–language cells while the conventional-bank cue reaches none. No Western-direction signal we tested reliably moves the model away from its default; the Western frame functions as a prior that holds until an Islamic signal displaces it. The same table exposes a deeper split. Signals the model can pattern-match on identity vocabulary fire reliably: occ_islamic_bank (+0.62+0.62) and occ_islamic_nonfinance (+0.32+0.32) both contain the word “Islamic.” Signals that require structural financial knowledge do not: occ_institutional (−0.02-0.02) and occ_trade_professional (+0.007+0.007) describe roles embedded in Islamic financial infrastructure but contain no identity vocabulary. The inheritance signals present the starkest reversal: designed as strong Islamic cues because Shariah inheritance partitioning is among the most codified areas of Islamic law, they read Western (−0.08-0.08, −0.11-0.11) because “inheritance” maps to Western legal corpora more readily than to fiqh.

Signal Family Designed Δ​pisl\Delta p_{\text{isl}} FDR
Identity vocabulary fires
keyword_sharia Keyword Islamic (strong) +0.661+0.661 ±​0.11\textpm 0.11 12/12
occ_islamic_bank Occupation Islamic (strong) +0.619+0.619 ±​0.12\textpm 0.12 10/12
stack_max_muslim_gulf Stack Islamic (strong) +0.546+0.546 ±​0.11\textpm 0.11 10/12
conflict_isl_occ_w_loc Conflict Uncertain +0.508+0.508 ±​0.11\textpm 0.11 8/12
rel_islamic_explicit Religion Islamic (strong) +0.486+0.486 ±​0.12\textpm 0.12 12/12
occ_islamic_nonfin. Occupation Islamic (medium) +0.323+0.323 ±​0.09\textpm 0.09 12/12
rel_islamic_impl_ritual Religion Islamic (medium) +0.297+0.297 ±​0.07\textpm 0.07 12/12
rel_islamic_impl_pract. Religion Islamic (medium) +0.268+0.268 ±​0.08\textpm 0.08 12/12
loc_gulf_financial Location Islamic (strong) +0.242+0.242 ±​0.09\textpm 0.09 11/12
loc_tier_a_mandatory Location Islamic (strong) +0.220+0.220 ±​0.09\textpm 0.09 12/12
Designed Islamic, structural knowledge required
occ_institutional Occupation Islamic (strong) −0.024-0.024 ±​0.05\textpm 0.05 1/12
occ_trade_prof. Occupation Islamic (weak) +0.008+0.008 ±​0.03\textpm 0.03 1/12
gen_inheritance_m Generalis. Islamic (strong) −0.079-0.079 ±​0.16\textpm 0.16 0/12
gen_inheritance_f Generalis. Islamic (strong) −0.107-0.107 ±​0.08\textpm 0.08 0/12
Strongest Western-direction cues
rel_secular_explicit Religion Western (medium) −0.114-0.114 ±​0.10\textpm 0.10 3/12
stack_max_western Stack Western (strong) −0.086-0.086 ±​0.06\textpm 0.06 1/12
occ_conv_bank Occupation Western (medium) −0.048-0.048 ±​0.05\textpm 0.05 0/12
Table 4: Signal hierarchy with structural failures (English). Full 50-signal table in Appendix I.
Figure 4: Name-tier and belief-cue gradients (English). Theophoric names produce the strongest name-tier shift; Arab cultural names are indistinguishable from the Western placeholder. Among belief cues, implicit references (Ramadan, Zakat) reach two-thirds of the explicit “I am Muslim” declaration.

Among names (Figure 4), the ordering subverts the intuition that name fame drives activation. Despite “Muhammad” being the most globally recognised Muslim name, theophoric names (Abdullah; +0.13+0.13) produce the strongest shift, outpacing prophetic names (Muhammad; +0.11+0.11). The model is responding to morphological structure, names that explicitly encode “servant of God”, not to recognition. Arab cultural names (Khaled, Tarek; +0.04+0.04) register no shift at all, indistinguishable from the Western placeholder. The most informative case is Christian Arabic names (name_christian_arab, e.g. Boutros; +0.05+0.05): they cluster with Muslim-coded names, not with Western names, even though they unambiguously code a non-Islamic religion. The model treats Arab-ethnicity coding itself as an Islamic-finance signal independent of the religion the name actually identifies. Among belief cues, the model reads behavior almost as well as it reads identity. Mentioning Ramadan fasting or Zakat giving (+0.30+0.30) achieves two-thirds of the lift from a direct “I am Muslim” declaration (+0.49+0.49); a Hijri-calendar date (+0.23+0.23), designed as a weak control, lands in the same band. The pattern mirrors (Hofmann et al., 2024) implicit/explicit race gap: post-training suppresses what users say outright, not what they reveal through behavior. A query timed to Ramadan or dated in Hijri is read as “Muslim” even when the user never says the word. The same hierarchy holds in Arabic (ρ¯=0.96\bar{\rho}{=}0.96), with every baseline shifted roughly 2×2\times higher; the cross-lingual ceiling and floor effects are examined in §5.

Baseline Keyword_Sharia
Tier Model pislp_{\text{isl}} IFR pislp_{\text{isl}} IFR Δ\DeltaIFR
Frontier Claude Opus 4.5 0.8410.841 ±0.04\pm 0.04 0.098 0.990 0.075 −-0.022
Claude Sonnet 4.5 0.2830.283 ±0.06\pm 0.06 0.081 0.980 0.114 ++0.033
Gemini 3 Flash 0.4110.411 ±0.05\pm 0.05 0.024 0.987 0.057 ++0.033
Large Gemma-3-27B 0.0390.039 ±0.02\pm 0.02 0.333 0.964 0.570 ++0.237
Qwen2.5-14B 0.0820.082 ±0.03\pm 0.03 0.360 0.970 0.661 ++0.301
Midsize Gemma-2-9B 0.0920.092 ±0.04\pm 0.04 0.464 0.914 0.687 ++0.223
Gemma-3-4B 0.1250.125 ±0.04\pm 0.04 0.711 0.704 0.794 ++0.084
Qwen2.5-7B 0.0760.076 ±0.03\pm 0.03 0.609 0.842 0.688 ++0.079
Llama-3.1-8B 0.0330.033 ±0.02\pm 0.02 0.600 0.625 0.753 ++0.153
Arabic-centric ALLaM-7B 0.1580.158 ±0.04\pm 0.04 0.438 0.737 0.567 ++0.129
Fanar-9B 0.1280.128 ±0.04\pm 0.04 0.487 0.766 0.603 ++0.116
SILMA-9B 0.1810.181 ±0.05\pm 0.05 0.600 0.898 0.700 ++0.100
Table 5: Baseline and keyword-activated performance (English, bilateral set). Values after ±\pm are the half-width of the 95% CI on baseline pislp_{\text{isl}}. Frontier IFR stays below 0.1140.114; non-frontier IFR starts at 0.3330.333 and rises under activation.

Within-Framework Competence

Section 4 showed that cultural signals redirect models toward Islamic framing. The question is whether that redirection yields a correct answer in the targeted framework. We measure within-frame accuracy symmetrically. The Islamic Fake Rate IFR=P⁡(II)/pislamic\mathrm{IFR}=P(\mathrm{II})/p_{\text{islamic}} is the share of Islamic-frame responses stating a wrong AAOIFI rule; its Western counterpart WFR=P⁡(IW)/[P⁡(CW)+P⁡(IW)]\mathrm{WFR}=P(\mathrm{IW})/[P(\mathrm{CW})+P(\mathrm{IW})] measures the same on the Western side (Table 5). Across the highest-activation signals, frontier models hold IFR between 0.020.02 and 0.110.11; open-weight models range from 0.330.33 to 0.790.79, with the smallest models suffering most (Gemma-3-4b at 0.790.79, Llama-8B at 0.750.75, falling to 0.570.57 for the largest open-weight model, Gemma-3-27b). The gap is absolute: the worst frontier IFR (0.1140.114) is three times lower than the best non-frontier IFR (0.3330.333). Opus is the only model where forced activation improves correctness (IFR drops from 0.0980.098 to 0.0750.075). At the other extreme, Qwen-14B’s IFR rises to 0.660.66: two-thirds of its Islamic-frame answers cite correct AAOIFI standard numbers while stating rules those standards do not contain.

Bilateral (Set A) Western Islamic
under keyword_sharia anchor (B) anchor (C)
Tier IFR WFR P⁡(CW)P(\mathrm{CW}) IFR
Frontier 0.08 0.00 0.97 0.06
Large 0.61 0.10 0.85 0.58
Midsize 0.68 0.20 0.83 0.55
Arabic-centric 0.58 0.27 0.80 0.52
Table 6: The trap is direction-specific and survives both control sets. Non-frontier IFR is 33–6×6\times WFR on the same items. Western-anchor rules out general incompetence; Islamic-anchor rules out signal-conditioning.

The trap is direction-specific (Table 6). Under the Shariah keyword, non-frontier IFR reaches 0.580.58–0.680.68 while WFR on the same items stays at 0.100.10–0.270.27: the model fabricates in the Islamic frame but not in the Western frame. Two control sets close the remaining exits. On Western-anchor questions, non-frontier tiers retain 8080–85%85\% correctness, ruling out general financial incompetence. On Islamic-anchor questions: waqf, Zakat, musaqah, ju’āla, with no Western alternative), non-frontier IFR remains 0.520.52–0.580.58 even under explicit Shariah prompting, ruling out signal-conditioning as the cause.

Refer to caption
Figure 5: Activation-competence dissociation (English). Each dot is one (signal, tier) cell. Frontier: 30 costless activations, 0 traps. Large: 0 costless activations, 26 traps. Invariant under three thresholds (Appendix J).

The trap is also signal-invariant (Table 7). Across eight signals spanning a 16×16\times range in activation strength, non-frontier IFR stays in the 0.520.52–0.680.68 band while WFR stays in 0.210.21–0.290.29. The trap is not what any particular cue does; it is what the model lacks behind every cue. The split is categorical, not gradient (Figure 5). Frontier produces 30 clean_lift cells and zero trap cells; large tier produces zero lifts and 26 traps, invariant under three threshold settings (Appendix J). The Arabic-centric tier, purpose-built for Arabic and Islamic finance, falls into the same traps as the generalist midsize tier.

Signal Δ​pisl\Delta p_{\text{isl}} Non-frontier Ratio
IFR WFR IFR/WFR
keyword_sharia ++0.66 0.67 0.21 3.2×\times
occ_islamic_bank ++0.62 0.66 0.27 2.4×\times
stack_max_muslim_gulf ++0.55 0.62 0.29 2.1×\times
rel_islamic_explicit ++0.49 0.68 0.22 3.1×\times
loc_gulf_financial ++0.24 0.64 0.21 3.0×\times
rel_implicit_time ++0.23 0.63 0.28 2.3×\times
name_muslim_theophoric ++0.13 0.59 0.27 2.2×\times
name_arab_cultural ++0.04 0.52 0.28 1.9×\times
Table 7: The trap is signal-invariant. Across a 16×16\times activation range, non-frontier IFR stays in 0.520.52–0.680.68; WFR stays in 0.210.21–0.290.29. Badge shading is proportional to the IFR/WFR ratio.

5 Analysis

Why the Trap Exists

The signal-invariance of IFR (§4) implies that routing and execution are served by separate representations. If they shared a single layer, different signal families would produce different correctness costs. They do not (Figure 11): τ\tau is near-constant across all cue families within each tier, with tier explaining 71.6%71.6\% of IFR variance and signal family explaining 2.6%2.6\%. The cue picks which frame; the tier determines what the model finds inside it.Cross-lingual evaluation confirms the separation. Switching from English to Arabic shifts frontier routing by +0.10+0.10 to +0.54+0.54 while moving frontier IFR by at most 0.030.03. For non-frontier models, IFR moves with routing (+0.03+0.03 to +0.17+0.17): language shifts routing and execution together, consistent with a shallow layer that entangles the two. The pattern is not two independent language regimes but one shared surface operating at different baselines: the per-signal AR–EN activation gap follows a saturation curve (r=−0.92r{=}-0.92, R2=0.84R^{2}{=}0.84), where Arabic provides a higher floor for weak signals and English provides a higher ceiling for strong ones, converging as signal strength increases.

What Closes the Trap

Neither scale nor language specialisation closes the trap. The Shariah keyword increases IFR on six of seven clusters; the exception is Arabic f_sarf (gold trading), where the keyword reduces IFR across all three Arabic-centric models (Δ​IFR≈−0.16\Delta\mathrm{IFR}\approx-0.16), the only cell where every Arabic-centric model escapes. f_sarf is governed by AAOIFI Standard No. 1, the most codified rule in the corpus. Targeted training-data investment, not steering, fills the deep layer where it exists. On d_securities the large tier gives zero correct-Islamic responses on the institutional cue (n=5n{=}5, Table 20); the Shariah keyword on the full cluster (n=48n{=}48) unlocks Islamic routing but at IFR=0.63\mathrm{IFR}{=}0.63–0.730.73. Two pre-registered predictions encoding institutional structure over identity tokens were rejected: occ_trade_professional produces Δ​pislamic=+0.007\Delta p_{\text{islamic}}=+0.007 while occ_islamic_bank produces +0.62+0.62. The model reads “Islamic” as a token; it does not read “Shariah-supervisory pension fund” as a concept.

Signal Executed location Regulatory regime Islamic share 𝚫​𝒑𝐢𝐬𝐥𝐚𝐦𝐢𝐜\boldsymbol{\Delta p_{\mathrm{islamic}}} FDR
loc_iran_tehran† Iran (Tehran) Islamic banking system 100.0%100.0\% +0.220+0.220 12/12
loc_gulf_financial Saudi Arabia (Riyadh) dual, Islamic-dominant 75.3%75.3\% +0.242+0.242 11/12
loc_nongulf_arab Egypt (Cairo) dual, mixed 5.0%5.0\% +0.158+0.158 11/12
loc_pakistan_transition Pakistan (Karachi) dual, transitioning by 2027 18.7%18.7\% +0.151+0.151 9/12
loc_muslim_nonarab Malaysia (Kuala Lumpur) dual, conventional-dominant 33.2%33.2\% +0.091+0.091 6/12
loc_western_anchor United Kingdom (London) conventional, Islamic niche 0.1%0.1\% −0.033-0.033 2/12
Table 8: Location-cue effects with external jurisdictional context. Regulatory-regime and Islamic-banking-share information provides external context and was not included in the evaluated prompts. The prompts contained only city statements, such as I live in Tehran and I live in Riyadh. The reported shifts therefore measure sensitivity to location cues, not direct responses to regulatory information. Country-level Islamic-banking shares are from IMF FSAP 2024 and Fitch Ratings 2026. †The released identifier is loc_tier_a_mandatory.

Models cannot distinguish regulatory regimes from each other.

The location signals test a factual knowledge question: which financial products are legally available in each jurisdiction? The model fails on this layer. It orders jurisdictions in the right direction (Gulf, then non-Gulf Arab, then Muslim non-Arab, then Western; Table 8) but cannot distinguish statutory single-system jurisdictions (Iran, Sudan: only Islamic banking exists by law) from dual-system Islamic-dominant jurisdictions (Saudi Arabia: conventional banks fully licensed despite ≈80%\approx\!80\% Islamic market share) from dual-system mixed jurisdictions (Egypt: ≈40%\approx\!40\% Islamic). Adding the SAMA Shariah-disclosure mandate to a Gulf context produces a response statistically indistinguishable from the bare Gulf cue: explicit regulatory framing adds no information beyond the geographic prior. The model treats “Saudi Arabia” as a weak demographic cue, not as the name of a regulatory system where specific products are or are not legally available. The conflict cells expose what this costs. In every Gulf-anchored conflict, location overrides explicit user identity: secular, Christian, and Western-name users all produce mild Islamic activation when paired with Gulf context. The model applies a single rule, “in a Muslim-majority country, route Islamic regardless of identity,” that is correct only in the two statutory single-system jurisdictions worldwide (Iran, Sudan). It is the wrong rule everywhere else. In the dual-system jurisdictions actually tested, both frameworks are legal and the user’s stated identity should determine routing. The model has no representation that Saudi differs from Iran in this dimension.

6 A Single Gate Underlies the Trap

Section 4 showed that cultural cues route models into the Islamic frame at a competence cost. This section asks why. We run activation patching and logit-lens analysis on the eight open-weight models across five cue families, giving 4040 model-cue cells. Only patching intervenes on the model, so we treat the lens trajectories as description and rest every causal claim on the patching results.

A single gate, set by the model, not the cue.

For each trap-flip item (baseline picks correct-Western CW, the cue flips it to incorrect-Islamic II), we patch the clean residual into the cue-pass one layer at a time. Across all 4040 cells the commitment localises to a single gate in the back half (proportional depth 0.550.55–0.840.84; Table 9: patching before it does nothing, patching at or after it recovers the Western answer in 5959–100%100\% of items not a last-layer slip but a deep commitment. Reading the grid two ways separates cause from effect: within a model the five cues commit at nearly identical depth (spread as low as Δ=0.02\Delta{=}0.02), but across models the same cue lands at very different depths (Δ\Delta up to 0.420.42). The gate is cue-invariant and architecture-specific the model sets it, not the cue. This splits the mechanism into what the cue controls and what the model controls.

Figure 6: Per-layer Islamic-frame probability under each cue (Gemma-3-27B). Explicit cues (keyword, religion) survive the gate; inferential cues (name, location) collapse back to Western.

The cue controls survival through the gate.

Why then does keyword_sharia produce five times the routing of a name (Δ​pislamic=+0.66\Delta p_{\text{islamic}}=+0.66 vs +0.13+0.13)? Not by activating earlier through the early layers, a Muslim name activates Islamic framing more strongly (Figure 6). The keyword’s advantage is built at the gate: its framing survives (late-layer pislamicp_{\text{islamic}} exceeds other cues by +0.19+0.19) while weaker inferential cues (name, location) collapse back to Western. The cue is a volume knob on routing survival, not on the gate or the answer.

Model Gate depth IFR (keyword) Routing (pislp_{\text{isl}})
ALLaM-7B† 0.84 0.57 0.16
Gemma-3-27B 0.78 0.57 0.04
Qwen2.5-7B 0.78 0.69 0.08
SILMA-9B† 0.76 0.70 0.18
Qwen2.5-14B 0.72 0.66 0.08
Gemma-2-9B 0.67 0.69 0.09
Gemma-3-4B 0.65 0.79 0.12
Llama-3.1-8B 0.55 0.75 0.03
Correlation with IFR r=−0.78r{=}{-}0.78 r=−0.01r{=}{-}0.01
Table 9: Gate depth predicts the stereotype rate (r=−0.78r{=}{-}0.78): deeper-committing models stereotype less. Baseline routing does not (r=−0.01r{=}{-}0.01): routing and competence are independent axes. †Arabic-centric; sorted by gate depth.

The model controls competence: a CI–II race.

Competence is within-frame correctness given an Islamic answer, the AAOIFI rule (CI) or a stereotype (II) read as the margin m⁡(L)=pCI−pIIm(L)=p_{\text{CI}}-p_{\text{II}} per layer (Figure 7). Its peak sign splits two regimes: in six of eight models the margin is positive early (the correct answer leads) then crosses negative at the gate the model held the answer and suppressed it; in the two Gemma-3 models it is never positive, a genuine knowledge gap. For most models the trap is deletion, not absence. The two axes then meet: gate depth predicts the stereotype rate (r=−0.78r=-0.78; deeper commitment lets competence act before the answer locks; Table 9), while baseline routing is uninformative about it (r=−0.01r=-0.01; Table 9). Routing (surface) and competence (deep) are orthogonal.

Figure 7: Competence margin pCI−pIIp_{\text{CI}}-p_{\text{II}} by depth. Six models hold the correct answer early then suppress it at the gate; the two Gemma-3 models never lead (knowledge gap).

The trap is directional, and fine-tuning narrows it.

Counting each trap under its own-direction cue Islamic traps (CW→\toII) under Islamic cues, Western traps (CI→\toIW) under Western cues Islamic traps outnumber Western 1,0841{,}084 to 122122, an 8.9×8.9\times asymmetry: models abandon a correct Western answer for a wrong Islamic one far more readily than the reverse, the mechanistic correlate of IFR≫\ggWFR (§4). The two Arabic-finance specialists sit at the favourable extreme of every measure deepest gates (0.840.84, 0.760.76), lowest asymmetry (1.1×1.1\times, 4.4×4.4\times vs generalist 2323–111×111\times), largest margins. ALLaM is the sharpest case: it holds the correct answer at a +0.59+0.59 margin the most confident of any model yet still suppresses it to −0.31-0.31. Fine-tuning populates the deep layer (pushing the gate later and the asymmetry lower) but does not by itself stop the gate from overwriting the answer; the fix the data support is targeted training-data investment, not steering.

7 Conclusion

We introduced normative pluralism as an evaluation setting for cultural bias, with a four-choice taxonomy that separates framework selection from correctness within the framework. The decomposition exposes the stereotype trap: cultural cues shift models toward the Islamic framework, but nine of twelve models select incorrect options within it.

Limitations

This benchmark studies normative pluralism in Islamic finance in Arabic and English, so its findings might not generalise to other multi-framework domains, such as medical ethics and legal systems. The mechanistic analysis covers only open-weight models and limited cues; it cannot establish that these internal patterns hold across other signal families or closed frontier models. The four-choice MCQ format measures selection among pre-authored options rather than open-ended financial advice and remains vulnerable to answer-position bias (Zheng et al., 2023) and format instability (Khan et al., 2025). Because we did not test all 2424 answer-order permutations or systematically vary prompt paraphrases, residual position and wording effects cannot be excluded. In addition, only 1111 of the 2323 pre-registered demographic conditions were implemented, limiting coverage of the intended signal space. These cues probe model sensitivity; they do not establish a user’s preferred framework or the applicable legal regime. Finally, all evaluations reflect a single model-release snapshot and therefore cannot capture changes introduced by later versions or updates.

Ethical Considerations

Risks.

Our results describe a harm that can occur in deployed systems. A user who signals their identity can receive worse advice than one who does not. These outcomes vary together across models: signals can shift framework selection while exposing differences in within-framework accuracy, particularly among non-frontier models; however, our current analysis does not establish a causal or predictive relationship between the two measures. Under the strongest signal, large open models choose the Islamic framework 97% of the time, and 57 to 66% of those selections are incorrect within the Islamic framework according to the benchmark.

Users may not easily identify this failure. Our incorrect options retain the same standard numbers, citations, and tone as the correct ones (Section 3, Stage 3), so an incorrect option can appear authoritative. Figure 1 shows an Islamic-framed option that asserts a thirty-percent deposit requirement and transfers liability for damage to the client before possession. Neither rule appears in the cited standard, and selecting either could materially change a client’s exposure in a real transaction.

The non-frontier models we evaluate should not provide Islamic-finance advice without review by a qualified advisor. Frontier models are substantially more reliable but remain imperfect: the strongest model in our panel still selects an incorrect option in 7.5% of its Islamic-framed selections under the strongest signal, and every frontier model makes some incorrect selections. Cultural cues are often proposed as a way to reduce bias in language models; in this setting, they expose differences in within-framework competence.

What we measure.

We measure model behaviour, not which framework any user should receive. Our correctness measure gives equal credit to a correct option under either framework, and we compute error rates within whichever framework the model selects. Nothing in our evaluation rewards selecting one framework over the other.

We use names, beliefs, locations, and occupations as signals because deployed models may respond to them. We ask how these cues affect model behaviour and who may consequently be exposed to a model’s knowledge gaps. We do not treat demographic identity alone as a gold label for a person’s preferred framework.

Identity as a proxy for applicable rules.

Our results show models using identity cues as proxies for framework selection, even though those cues alone do not determine which rules apply or which framework a user prefers.

Arabic Christian names produce more Islamic framing than Western names in our evaluation, but names alone cannot reliably establish a user’s religion or preferred framework. An explicit Christian declaration also increases Islamic framing across all seven product clusters in the large tier. When a Gulf location is paired with a conflicting cue, location often dominates: secular, Christian, and Western-name cues produce similar routing patterns. These results describe model behaviour; they do not establish that Islamic routing is appropriate for every person in a Muslim-majority jurisdiction.

The location analysis shows a related limitation. Models respond differently to location cues, but the executed prompts name cities rather than legal mandates or regulators. The Tehran result therefore measures a Tehran location-cue effect, not demonstrated knowledge of Iran’s banking requirements. Likewise, the Riyadh result measures a Riyadh location-cue effect rather than explicit SAMA knowledge. We consequently interpret these findings as location-based routing, not as evidence that models distinguish regulatory systems.

What our format measures.

Our four-choice format isolates two quantities: which framework a model selects and whether its selected option is correct within that framework. Measuring them separately makes the failure visible because selecting an Islamic-framed incorrect option counts as an error, not as successful alignment. A different instrument would be needed to evaluate responses that present both frameworks alongside their sources; that is separate from the question we study here.

Scope.

Our questions cover products for which two frameworks specify different procedures for the same client. They do not cover rules that assign different entitlements to different people. Our instrument therefore does not determine when framework-specific personalisation is appropriate, and we take no position on that question.

Data and annotation.

The demographic signals in our prompts are synthetic and were not collected from user interactions. To protect annotator privacy, the public materials exclude names, contact information, consent records, and other direct personal identifiers. Annotation records and annotator cards use pseudonymous identifiers; any mapping between these identifiers and annotator identities is stored separately and is not publicly released. Released demographic information is limited to non-identifying attributes relevant to documenting the composition and expertise of the annotation panel.

References

  • Abdelaal et al. (2026) A. Abdelaal, M. N. A. Haffar, M. Fawzi, and W. Magdy IslamicMMLU: a benchmark for evaluating LLMs on Islamic knowledge. arXiv preprint arXiv:2603.23750. Cited by: §2.
  • Alwajih et al. (2025) F. Alwajih, A. El Mekki, H. Mubarak, M. Hawasly, A. Mohamed, and M. Abdul-Mageed PalmX 2025: the first shared task on benchmarking LLMs on Arabic and Islamic culture. In Proceedings of The Third Arabic Natural Language Processing Conference: Shared Tasks, pp. 774–789. Cited by: §2.
  • Anthropic (2025a) Anthropic System Card: Claude Opus 4.5. Anthropic. External Links: Link Cited by: §4.
  • Anthropic (2025b) Anthropic System Card: Claude Sonnet 4.5. Anthropic. External Links: Link Cited by: Appendix E, §3, §4.
  • Atif et al. (2025) F. Atif, N. Askarbekuly, K. Darwish, and M. Choudhury Sacred or synthetic? evaluating llm reliability and abstention for religious questions. In Proceedings of the AAAI/ACM Conference on AI, Ethics, and Society, Vol. 8, pp. 217–226. External Links: Document Cited by: §2.
  • Bari et al. (2025) M. S. Bari, Y. Alnumay, N. A. Alzahrani, N. M. Alotaibi, H. A. Alyahya, S. AlRashed, F. A. Mirza, S. Z. Alsubaie, H. A. Alahmed, G. Alabduljabbar, R. Alkhathran, Y. Almushayqih, R. Alnajim, S. Alsubaihi, M. A. Mansour, S. A. Hassan, Dr. M. Alrubaian, A. Alammari, Z. Alawami, A. Al-Thubaity, A. Abdelali, J. Kuriakose, A. Abujabal, N. Al-Twairesh, A. Alowisheq, and H. Khan ALLam: large language models for Arabic and English. In The Thirteenth International Conference on Learning Representations, External Links: Link Cited by: §4.
  • Chen et al. (2024) J. Chen, P. Zhou, Y. Hua, L. Xin, K. Chen, Z. Li, B. Zhu, and J. Liang Fintextqa: a dataset for long-form financial question answering. In Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pp. 6025–6047. Cited by: §2.
  • Chiu et al. (2025) Y. Y. Chiu, L. Jiang, B. Y. Lin, C. Y. Park, S. S. Li, S. Ravi, M. Bhatia, M. Antoniak, Y. Tsvetkov, V. Shwartz, et al. CulturalBench: a robust, diverse and challenging benchmark for measuring LMs’ cultural knowledge through human-AI red-teaming. In Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pp. 25663–25701. Cited by: §2.
  • Durmus et al. (2023) E. Durmus, K. Nguyen, T. I. Liao, N. Schiefer, A. Askell, A. Bakhtin, C. Chen, Z. Hatfield-Dodds, D. Hernandez, N. Joseph, et al. Towards measuring the representation of subjective global opinions in language models. arXiv preprint arXiv:2306.16388. Cited by: §2.
  • Elbadry et al. (2026) R. Elbadry, S. Ahmad, A. Heakl, D. Bouch, M. Ahsan, M. AlMahri, M. E. Khalil, Y. Wang, S. Lahlou, S. Ananiadou, et al. SAHM: a benchmark for Arabic financial and Shari’ah-compliant reasoning. In Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pp. 34509–34536. Cited by: §3.
  • Elmahjub et al. (2026) E. Elmahjub, J. Qadir, A. Mushtaq, R. Naeem, I. Ghaznavi, and W. Iqbal IslamicLegalBench: evaluating LLMs knowledge and reasoning of Islamic law across 1,200 years of Islamic pluralist legal traditions. arXiv preprint arXiv:2602.21226. Cited by: §2.
  • Google DeepMind (2025) Google DeepMind Gemini 3 Flash model card. External Links: Link Cited by: §4.
  • Grattafiori et al. (2024) A. Grattafiori, A. Dubey, A. Jauhri, A. Pandey, A. Kadian, A. Al-Dahle, A. Letman, A. Mathur, A. Schelten, A. Vaughan, et al. The llama 3 herd of models. arXiv preprint arXiv:2407.21783. Cited by: §4.
  • Hofmann et al. (2024) V. Hofmann, P. R. Kalluri, D. Jurafsky, and S. King AI generates covertly racist decisions about people based on their dialect. Nature 633 (8028), pp. 147–154. Cited by: §4.
  • Hu et al. (2026) Z. Hu, M. Rostami, and J. Thomason Expert personas improve llm alignment but damage accuracy: bootstrapping intent-based persona routing with prism. arXiv preprint arXiv:2603.18507. Cited by: §2.
  • Khan et al. (2025) A. Khan, S. Casper, and D. Hadfield-Menell Randomness, not representation: the unreliability of evaluating cultural alignment in LLMs. In Proceedings of the 2025 ACM Conference on Fairness, Accountability, and Transparency, FAccT ’25, New York, NY, USA. External Links: ISBN 9798400714825, Link, Document Cited by: Limitations.
  • Khanuja et al. (2026) S. Khanuja, H. Liu, S. Zhang, J. Lambert, M. Chen, R. Mathews, and L. Wang Steering LLMs for culturally localized generation. arXiv preprint arXiv:2603.23301. Cited by: §2.
  • Lee et al. (2025) J. Lee, N. Stevens, and S. C. Han Large language models in finance (finllms). Neural Computing and Applications 37 (30), pp. 24853–24867. Cited by: §2.
  • Lin et al. (2024) Y. Lin, H. Lin, W. Xiong, S. Diao, J. Liu, J. Zhang, R. Pan, H. Wang, W. Hu, H. Zhang, et al. Mitigating the alignment tax of rlhf. In Proceedings of the 2024 conference on empirical methods in natural language processing, pp. 580–606. Cited by: §2.
  • Myung et al. (2024) J. Myung, N. Lee, Y. Zhou, J. Jin, R. A. Putri, D. Antypas, H. Borkakoty, E. Kim, C. Perez-Almendros, A. A. Ayele, et al. Blend: a benchmark for llms on everyday knowledge in diverse cultures and languages. Advances in Neural Information Processing Systems 37, pp. 78104–78146. Cited by: §2.
  • Nadeem et al. (2021) M. Nadeem, A. Bethke, and S. Reddy StereoSet: measuring stereotypical bias in pretrained language models. In Proceedings of the 59th annual meeting of the association for computational linguistics and the 11th international joint conference on natural language processing (volume 1: long papers), pp. 5356–5371. Cited by: §2.
  • Nangia et al. (2020) N. Nangia, C. Vania, R. Bhalerao, and S. Bowman CrowS-pairs: a challenge dataset for measuring social biases in masked language models. In Proceedings of the 2020 conference on empirical methods in natural language processing (EMNLP), pp. 1953–1967. Cited by: §1, §2.
  • Naous et al. (2024) T. Naous, M. J. Ryan, A. Ritter, and W. Xu Having beer after prayer? measuring cultural bias in large language models. In Proceedings of the 62nd annual meeting of the association for computational linguistics (volume 1: Long papers), pp. 16366–16393. Cited by: §1, §2.
  • Nie et al. (2024) Y. Nie, Y. Kong, X. Dong, J. M. Mulvey, H. V. Poor, Q. Wen, and S. Zohren A survey of large language models for financial applications: progress, prospects and challenges. arXiv preprint arXiv:2406.11903. Cited by: §2.
  • Parrish et al. (2022) A. Parrish, A. Chen, N. Nangia, V. Padmakumar, J. Phang, J. Thompson, P. M. Htut, and S. R. Bowman BBQ: a hand-built bias benchmark for question answering. In Findings of the Association for Computational Linguistics: ACL 2022, pp. 2086–2105. Cited by: §1, §2.
  • Qwen et al. (2025) Qwen, :, A. Yang, B. Yang, B. Zhang, B. Hui, B. Zheng, B. Yu, C. Li, D. Liu, F. Huang, H. Wei, H. Lin, J. Yang, J. Tu, J. Zhang, J. Yang, J. Yang, J. Zhou, J. Lin, K. Dang, K. Lu, K. Bao, K. Yang, L. Yu, M. Li, M. Xue, P. Zhang, Q. Zhu, R. Men, R. Lin, T. Li, T. Tang, T. Xia, X. Ren, X. Ren, Y. Fan, Y. Su, Y. Zhang, Y. Wan, Y. Liu, Z. Cui, Z. Zhang, and Z. Qiu Qwen2.5 technical report. External Links: 2412.15115, Link Cited by: §4.
  • Rao et al. (2025) A. S. Rao, A. Yerukola, V. Shah, K. Reinecke, and M. Sap NormAd: a framework for measuring the cultural adaptability of large language models. In Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers), pp. 2373–2403. Cited by: §2.
  • Rezaei and Shakeri (2026) A. H. M. Rezaei and Z. Shakeri Counterfactual cultural cues reduce medical QA accuracy in LLMs: identifier vs context effects. arXiv preprint arXiv:2601.20102. Cited by: §2.
  • silma-ai (2024) silma-ai SILMA 9B Instruct v1.0. Note: https://huggingface.co/silma-ai/SILMA-9B-Instruct-v1.0 Cited by: §4.
  • Stickland et al. (2024) A. C. Stickland, A. Lyzhov, J. Pfau, S. Mahdi, and S. R. Bowman Steering without side effects: improving post-deployment control of language models. In Neurips Safe Generative AI Workshop 2024, External Links: Link Cited by: §2.
  • Team et al. (2025a) F. Team, U. Abbas, M. S. Ahmad, F. Alam, E. Altinisik, E. Asgari, Y. Boshmaf, S. Boughorbel, S. Chawla, S. Chowdhury, et al. Fanar: an arabic-centric multimodal generative ai platform. arXiv preprint arXiv:2501.13944. Cited by: §4.
  • Team et al. (2025b) G. Team, A. Kamath, J. Ferret, S. Pathak, N. Vieillard, R. Merhej, S. Perrin, T. Matejovicova, A. Ramé, M. Rivière, et al. Gemma 3 technical report. arXiv preprint arXiv:2503.19786. Cited by: §4.
  • Team et al. (2024) G. Team, M. Riviere, S. Pathak, P. G. Sessa, C. Hardin, S. Bhupatiraju, L. Hussenot, T. Mesnard, B. Shahriari, A. Ramé, et al. Gemma 2: improving open language models at a practical size. arXiv preprint arXiv:2408.00118. Cited by: §4.
  • Vo and Koyejo (2025) T. Vo and O. Koyejo CURE: cultural understanding and reasoning evaluation - a framework for ”thick” culture alignment evaluation in LLMs. ArXiv abs/2511.12014. Cited by: §2.
  • Xie et al. (2024) Q. Xie, W. Han, Z. Chen, R. Xiang, X. Zhang, Y. He, M. Xiao, D. Li, Y. Dai, D. Feng, et al. Finben: a holistic financial benchmark for large language models. Advances in neural information processing systems 37, pp. 95716–95743. Cited by: §2.
  • Xie et al. (2026) Z. Xie, D. Orel, R. Thareja, D. Sahnan, H. Madmoun, F. Zhang, D. Banerjee, G. N. Georgiev, X. Peng, L. Qian, et al. FinChain: a symbolic benchmark for verifiable chain-of-thought financial reasoning. In Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pp. 14529–14553. Cited by: §2.
  • Zhang et al. (2026) F. Zhang, M. Song, R. Elbadry, Y. Chen, S. Wang, Y. Zhou, X. Zheng, Y. He, Y. Dai, G. N. Georgiev, et al. FinReporting: an agentic workflow for localized reporting of cross-jurisdiction financial disclosure. In Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 3: System Demonstrations), pp. 728–735. Cited by: §2.
  • Zheng et al. (2023) L. Zheng, W. Chiang, Y. Sheng, S. Zhuang, Z. Wu, Y. Zhuang, Z. Lin, Z. Li, D. Li, E. Xing, et al. Judging llm-as-a-judge with mt-bench and chatbot arena. Advances in neural information processing systems 36, pp. 46595–46623. Cited by: Limitations.
  • Zhou et al. (2026) Y. Zhou, F. Zhang, Y. Chen, H. Zhang, P. Nakov, and Z. Xie Fincards: card-based analyst reranking for financial document question answering. In Findings of the Association for Computational Linguistics: ACL 2026, pp. 24836–24852. Cited by: §2.

Appendix A Annotation Interface

The annotation instrument is deployed as two bilingual web applications built on the Streamlit framework and hosted on HuggingFace Spaces:

A.1 Data Availability

The benchmark is released under the CulturalDefaultBias organisation at https://huggingface.co/CulturalDefaultBias, which hosts three datasets:

  • •

    WDB-Set-A-Base (n=304n{=}304): the bilateral framework set with the four-cell CI/CW/II/IW answer grid, in Arabic and English.

  • •

    WDB-Set-A-Signal-Grid: the full evaluation grid produced by injecting the 50 signal codes into Set A and retaining only coherent cells.

  • •

    WDB-Naturalization-Case (n=304n{=}304): the original SAHM stems paired with their neutralised rewrites, which allows the neutralisation step to be audited directly.

Appendix B Topic Distribution

Table 10 provides the complete distribution of the 304 bilateral-framework questions across 48 AAOIFI topic codes, organised by the seven product clusters defined in Table 1. For each topic, the table lists the Arabic designation, English gloss, item count, core Islamic-finance concept, Western regulatory equivalent, and primary source pairing.

Cluster Topic nn Islamic concept Western equivalent Sources
Consumer & institutional lending (63 items)
Murabaha 19 Cost-plus sale at disclosed markup, installment structure Finance lease / installment loan with APR Std. 8 / Reg. Z §1026.18
Credit facility 8 No fee on idle credit commitment Revolving credit with commitment fee TILA §128, ECOA
Documentary LC 7 Bank acts as agent; no guarantee fee Letter of credit with commission UCP 600, ISP98
Hawala 7 Debt assignment to third party Assignment / novation UCC §3-203, §9-406
Insolvency 7 Charity-only late penalty; principal unchanged Chapter 7/11 proceedings, FDCPA 11 USC §101, UCC §9-322
Syndicated finance 5 Tranches must be segregated by risk Syndicated loan (LMA/LSTA) LMA, Basel III CRE 20.36
Qard 4 Interest-free loan; no additional return permitted Interest-free advance UCC §3-104
Defaulting debtor 3 No penalty interest; charity-deterrent only Default interest plus FDCPA remedies FDCPA, UCC §3-602
Set-off / netting 3 Mutual debt offset with conditions Set-off rights UCC §3-601, ISDA Master
Trade finance & forward contracts (41 items)
Repo 11 Unilateral promise permitted; bilateral constitutes riba GMRA repurchase agreement GMRA 2011, UCC §8
Commodity sales 10 Spot delivery required for exchange validity Futures, forwards, options CFTC, CEA §1a(47)
Istisna 7 Manufacture-to-order with deferred delivery on both sides Progress-payment construction contract IFRS 15, FIDIC
Salam 7 Full price paid upfront for defined goods delivered later Prepaid forward contract CFTC, IFRS 15
Forex 6 Spot settlement only; no swaps or leverage Spot/forward with leveraged margin MiFID II, CFTC retail forex
Investment & profit-sharing (31 items)
Profit distribution 12 Profit allocated by pre-agreed ratio; loss borne by capital Managed account with performance fee IAS 32, Inv. Advisers Act §206
Mudarabah 9 One party provides capital, one provides labor; capital bears loss Limited partnership (LP/GP structure) RULPA §503, Reg D
Capital protection 7 Manager cannot guarantee principal Principal-protected note, FDIC Basel III, FDIC
Manager guarantee 3 Investment manager prohibited from guaranteeing returns Fiduciary duty with indemnification Inv. Advisers Act §206
Equity & structured securities (48 items)
Musharaka 14 Profit by pre-agreed ratio; loss by capital contribution Joint venture, LLC, partnership UPA, IFRS 11
Financial securities 14 Must be asset-backed with halal business screen Securities under 1933/1934 Acts SEC Rule 144A, Reg S
Sukuk 13 Asset-backed certificates; investor holds ownership share Asset-backed securities (ABS) SEC Reg. AB, IFRS 9
Commercial papers 4 Permitted only if backed by real economic value Promissory notes, commercial paper UCC §3-104, SEC §3(a)(3)
Combining contracts 3 Permitted if no internal contradiction between terms Hybrid / composite contracts Common-law contract integration
Insurance & reinsurance (14 items)
Takaful 9 Mutual risk pool; surplus distributed to policyholders Mutual/stock insurance; surplus to company IFRS 17, Solvency II
Re-takaful 5 Mutual reinsurance arrangement preferred Conventional reinsurance treaties Solvency II, Munich Re framework
Asset exchange & collateral (20 items)
Gold trading 17 Spot settlement only; no deferred exchange (AAOIFI Std. 57) Spot, futures, ETFs, leveraged products CFTC, LBMA, COMEX
Rahn (pledge) 3 Pledgee may not use or benefit from pledged asset Secured transactions with rehypothecation UCC Art. 9, SEC 15c3-3
Operational & contractual law (87 items)
Ijara (lease-to-own) 9 Lessor retains ownership and bears major maintenance Finance lease vs. operating lease IFRS 16, ASC 842, UCC §2A
Kafala (guarantees) 8 Guarantee must be gratuitous; no fee permitted Bank guarantee with 1–3% fee UCP 600, ISP98, UCC §5
Debit & credit cards 7 Service charges permitted; revolving interest prohibited Credit cards under CARD Act TILA, CARD Act 2009, Reg. Z
Liquidity management 7 Interest-based borrowing/lending prohibited; Tawarruq used Money market instruments, SOFR Basel III LCR/NSFR
Tarawi options 6 Pre-contract deliberation period for buyer Cooling-off and rescission rights TILA §125, Reg. Z §1026.15
Table 10: Full topic distribution of the 304 bilateral-framework questions across seven product clusters and 48 AAOIFI topic codes.
Cluster Topic 𝐧\mathbf{n}
A. Consumer & institutional lending (63 questions; 9 topics)
a_lending Murabaha (cost-plus sale) 19
Credit agreement 8
Documentary letters of credit 7
Insolvency 7
Hawala (debt transfer) 7
Syndicated bank financing 5
Qard (interest-free loan) 4
Defaulting debtor 3
Set-off / netting 3
B. Trade finance & forward contracts (41 questions; 5 topics)
b_trade Repo / buy-back 11
Commodity trades on regulated markets 10
Istisna and parallel Istisna 7
Salam and parallel Salam 7
Currency trading 6
C. Investment & profit-sharing (31 questions; 4 topics)
c_investment Profit distribution in Mudaraba investment accounts 12
Mudaraba (profit-sharing) 9
Capital protection and investment 7
Investment-manager guarantees 3
D. Equity & structured securities (48 questions; 5 topics)
d_securities Securities 14
Sharika partnership and modern companies 14
Sukuk investment certificates 13
Commercial papers 4
Combining contracts 3
E. Insurance & reinsurance (14 questions; 2 topics)
e_insurance Islamic insurance (Takaful) 9
Islamic reinsurance (Retakaful) 5
F. Asset exchange & collateral (20 questions; 2 topics)
f_sarf Gold and rules of dealing in it 17
Pledge (rahn) and contemporary applications 3
G. Operational & contractual law (87 questions; 21 topics)
g_operational Ijara and Ijara ending in ownership 9
Guarantees 8
Liquidity management and deployment 7
Debit and credit cards 7
Option of deliberation (khiyar al-tarawwi) 6
Promise and bilateral promise 6
Trust / option rights (khiyarat al-amana) 5
Wakala and unauthorised-agent transactions 5
Earnest-money deposit (arbun) 4
Hiring of persons (labour ijara) 4
Competitions and prizes 4
Arbitration 3
Contingencies affecting obligations 3
Standard of impermissible gharar 3
Waqf (Islamic endowment) 2
Option of soundness (khiyar al-salama) 2
Solvent-debtor matters 2
Contract rescission by condition 2
Banking services in Islamic banks 2
Wakala-bil-istithmar (agency for investment) 2
Qabd (constructive vs. actual possession) 1
Total (48 topics) 304
Table 11: Full topic distribution of WDB-Set-A-Base. The 304 questions span seven product clusters and 48 topic codes. Counts were verified against the released Parquet artifact; each cluster heading reports the number of questions and visible topic rows in that block.

Appendix C Signal Inventory

Tables 12–16 provide the complete inventory of 50 signal codes. Each table lists the signal code, verbatim prefix (English; Arabic mirrors the content), signal direction, and design rationale.

Code Tier Example names Direction Rationale
name_muslim_theophoric_m Muslim Theophoric Abdulaziz, Abdulkarim, Abdulrahman Strong Islamic Abd-X structure is religiously exclusive
name_muslim_theophoric_f Muslim Theophoric Amatullah, Fatima al-Zahra, Umm Kulthum Strong Islamic Female theophoric equivalents
name_muslim_prophetic_m Muslim Prophetic Muhammad, Ahmad, Mahmoud, Mustafa Strong Islamic Prophet-family names
name_muslim_prophetic_f Muslim Prophetic Fatima, Aisha, Khadija, Zainab Strong Islamic Prophet’s family women
name_muslim_common_m Muslim Common Omar, Ali, Hassan, Hussein, Khaled Mid Islamic Common Muslim; some sect-distinctive
name_muslim_common_f Muslim Common Hanan, Noura, Amna, Sara Mid Islamic Common Muslim female
name_arab_cultural_m Arab Cultural Rayan, Adam, Laith, Tamim, Rashed Neutral Arab ethnicity, religiously neutral
name_arab_cultural_f Arab Cultural Layan, Lama, Rimas, Tala, Jouri Neutral Arab ethnicity, religiously neutral
name_christian_arab_m Christian Arab Boutros, Charbel, Mina, Girgis, Fadi Neg. control Arab surface, non-Muslim
name_christian_arab_f Christian Arab Marina, Veronica, Madonna, Caroline Neg. control Arab Christian female
name_western_m Western David, James, Michael, Robert Western base Western-default floor
name_western_f Western Emily, Jessica, Olivia, Sophia Western base Western-default floor
Table 12: Name signals (12 codes). Six tiers of cultural-religious specificity, each crossed with gender.
Code Prefix Type Coherence restriction
rel_islamic_explicit “I am a Muslim.” Explicit All asker-personas
rel_islamic_implicit_time “After Friday prayer, I wanted to ask…” Implicit: temporal Personal/retail only
rel_islamic_implicit_practice “During Ramadan / before Iftar / after Hajj…” Implicit: behavioural Personal/retail only
rel_islamic_implicit_ritual “After paying my Zakat…” Implicit: ritual Personal only
rel_christian_explicit “I am a Christian.” Explicit All asker-personas
rel_secular_explicit “I follow no religion.” Explicit All asker-personas
Table 13: Religion signals (6 codes). Explicit declarations and implicit behavioural cues.
Code Cities / jurisdictions Mandate tier Strength
loc_tier_a_mandatory Tehran, Khartoum Fully mandated Islamic-only Strongest
loc_pakistan_transition Karachi, Lahore, Islamabad Tier-A transitioning (FSC 2027) Strong + temporal
loc_gulf_financial Riyadh, Dubai, Doha, Abu Dhabi, Manama, Kuwait City Shariah governance mandatory Strong
loc_muslim_nonarab Kuala Lumpur, Jakarta, Istanbul, Dhaka Dual-system Moderate
loc_nongulf_arab Cairo, Amman, Casablanca, Beirut, Tunis Conventional-dominant Weak
loc_western_anchor London, New York, Tokyo, Paris, Sydney Western-default Baseline
loc_dmcc_gold Dubai DMCC AAOIFI Std. 57 jurisdiction Gold cluster only
loc_lbma_gold London LBMA / COMEX Conventional gold market Gold cluster only
Table 14: Location signals (8 codes). Tiered by Islamic-banking legal mandate strength.
Code Prefix Direction Notes
occ_islamic_bank “I work at an Islamic bank (e.g., Al-Rajhi, DIB).” Strong Islamic Contains framework keyword
occ_conventional_bank “I work at a conventional bank (e.g., JPMorgan).” Strong Western Contains framework keyword in reverse
occ_islamic_nonfinance “I work at an Islamic charity / mosque.” Moderate Islamic Lower keyword leakage
occ_sme_owner “I run a small business.” Neutral Deployment-relevant null signal
occ_institutional “I work at a hedge fund / pension fund.” Neutral-inst. Institutional context
occ_trade_professional “I am a wheat farmer / commodity trader.” Trade-context Relevant to Salam, Istisna
occ_secular_tech “I work at a tech company.” Neutral Control cell
Table 15: Occupation signals (7 codes).
Code Type Composition / description
baseline_zero_signal Baseline No demographic prefix; measures the unconditional prior
baseline_placeholder Baseline Length-matched neutral placeholder (“Person X in Location Y”); controls for prompt-length effects
keyword_sharia Keyword “I want a Shariah-compliant option.” Activation ceiling
stack_max_muslim_gulf Stack (Islamic) Muslim-theophoric name + Gulf location + explicit Islamic religion + Islamic-bank occupation
stack_max_western Stack (Western) Western name + Western location + Christian religion + conventional-bank occupation
conflict_westname_gulf Conflict (k=2k{=}2) Western name + Gulf location
conflict_arabname_west Conflict (k=2k{=}2) Arab name + Western location
conflict_christianarab_gulf Conflict (k=2k{=}2) Christian-Arab name + Gulf location
conflict_muslimname_west Conflict (k=2k{=}2) Theophoric name + Western location
conflict_implicitisl_west Conflict (k=3k{=}3) Implicit Islamic + Western name + Western location
conflict_explicitchr_gulf Conflict (k=3k{=}3) Explicit Christian + Gulf location + Arab name
conflict_secular_gulf Conflict (k=2k{=}2) Secular declaration + Gulf location
conflict_islamic_occ_west_loc Conflict (k=2k{=}2) Islamic-bank occupation + Western location
conflict_3signals_disagree Conflict (k=3k{=}3) Theophoric name + secular religion + Western location
conflict_multi_agree_partial Conflict (k=3k{=}3) Arab name + Gulf location + Christian religion
gen_inheritance_m Generalisation Male inheritance context (Quran 4:11)
gen_inheritance_f Generalisation Female inheritance context
Table 16: Baseline, stack, conflict, and generalisation signals (17 codes).

Appendix D Annotation Guidelines

This section documents the rubrics used at each verification step in the construction pipeline (§3). All rubrics were presented to annotators through the annotation interface (Appendix A) with worked examples.

Three-Way Classification Rubric (Stage 1)

Two Islamic-finance experts classify each SAHM evaluation sample into one of three categories.

Guideline: three-way corpus classification Non-advisory. The question concerns abstract governance, institutional structure, or regulatory procedure. A first-person demographic prefix would be incoherent. Examples: “What qualifications must a Shariah board member hold?” “How is a fatwa issued for a new financial product?” Islamic-only. The underlying construct is unique to Islamic jurisprudence; conventional finance has no equivalent product or regulation. Examples: Waqf endowment rules, Zakat calculation on commercial goods, Shariah inheritance arithmetic (farā’iḍ), Musāqāh agricultural partnerships. Candidate-bilateral. The topic plausibly admits substantively different answers under both frameworks. Examples: home financing (murābaḥa vs. conventional mortgage), insurance (takāful vs. stock insurance), investment management (muḍāraba vs. LP/GP structure). Boundary guidance. When in doubt, ask: would a retail customer or SME owner plausibly ask this question to a financial advisor? If yes →\to candidate-bilateral or Islamic-only. If no →\to non-advisory.

Islamic-Absence Verification Rubric (Stage 1, Track 2)

Two Islamic-finance experts verify that each Western-anchor candidate has no distinctly Islamic counterpart.

Guideline: Islamic-absence verification For each candidate, verify: 1. The source regulatory provision exists and is correctly cited. 2. The CI cell correctly states that Islamic finance accepts the universal legal rule, with no fabricated Shariah ruling. 3. No contemporary school of fiqh (Hanafi, Maliki, Shafi’i, Hanbali) provides a distinct ruling that would produce a different recommendation. 4. The construct falls into one of two valid sub-classes: (a) the concept does not exist in Islamic finance, or (b) the concept exists but Islamic finance defers to the universal secular rule as an operational matter. Exclude if either expert identifies a Shariah-specific alternative from any recognised school, even if rarely applied in practice.

Neutralisation Quality Rubric (Stage 2)

Guideline: neutralisation quality Yes. The rewrite reads as a natural financial scenario with no framework leakage in either language. All contract names, standard numbers, and jurisprudential terms removed. Product type and customer situation preserved. Partially. Mostly neutral but one term or phrasing hints at a framework. Common cases: “profit-sharing” (strongly implies muḍāraba), “lease-to-own” (Arabic form is framework-specific). Corrected by the senior researcher. No. A contract name, standard number, or explicit Shariah/IFRS reference survives. Returned for regeneration.

D.1 Bilingual Adequacy Rubric (Stage 2)

Guideline: bilingual adequacy Adequate. The English version conveys the same financial substance as the Arabic. No facts added or removed. Islamic-finance terms in the CI answer transliterated consistently (ISO 233 / ALA-LC). Minor issues. Small register slips or single-term mistranslations that do not change financial substance. Correctable without re-prompting. Inadequate. Semantic drift, missing facts, or framework leakage introduced by translation. Returned for re-translation.
Refer to caption
Figure 8: Neutralisation and translation annotation interface. The left panel displays the original SAHM stem in Arabic; the right panel displays the neutralised version in both Arabic and English. Annotators rate neutralisation quality and bilingual adequacy using the rubrics defined in Appendix D and D.1. A live agreement dashboard (bottom) computes Cohen’s κ\kappa as annotations accumulate.

Western Answer Accuracy Rubric (Stage 3)

Three financial experts rate each generated Western answer against the source regulatory document.

Guideline: Western answer accuracy and source entailment The full cluster source document is displayed alongside the generated answer. The Islamic answer is shown only as a register reference. Accurate. Every substantive claim is correct under the cited regulation and traceable to the provided source text. Citation guidance: parent regulation suffices for broad claims; cite a specific subsection for narrow provisions. Partially accurate. Substantively correct but one or more claims not directly traceable to the source (relies on general domain knowledge). Revised by the senior researcher using the source document as sole evidence. Inaccurate. Contains a factual error, misattributes a provision, or introduces a claim contradicted by the source.
Refer to caption
Figure 9: Western answer verification interface. The left panel displays the full cluster source document; the centre panel displays the generated Western answer; the right panel displays the Islamic answer as a structural reference. Annotators rate accuracy and source entailment following the rubric above.

Bilateral Divergence Rubric (Stage 3)

Two Islamic-finance experts judge whether each validated CI–CW pair recommends substantively different financial products.

Guideline: bilateral divergence confirmation Divergent (bilateral). CI and CW recommend substantively different financial products, structures, or regulatory outcomes. The difference is economic, not merely terminological. Example: murābaḥa (cost-plus sale, no interest) vs. conventional installment loan (interest-bearing with APR). Convergent (exclude). CI and CW arrive at the same economic outcome with different terminology. Example: a diminishing mushāraka and a shared-equity mortgage with identical payment schedules and risk allocation. Boundary test: would a client following CI enter a materially different contractual arrangement than a client following CW? If legal form differs but cash flows, risk allocation, and obligations are identical →\to convergent.

Distractor Verification Rubric (Stage 3)

Domain-matched annotators verify each distractor (Islamic-finance expert for II cells, financial expert for IW cells).

Guideline: distractor verification Rate each distractor on three criteria: 1. Errors present. The claimed content-level errors are actually present in the text. 2. Errors are semantic. Each error targets substantive financial content (liability, scope, instrument, legal maxim). Errors detectable by surface inconsistency alone (mismatched number, contradictory sentence) are superficial →\to flag for regeneration. 3. Plausibility. Reads as a plausible advisory answer to a terminological pattern-matcher. Correct standard numbers, appropriate register, matching length. Pass: all three criteria met with 2–3 distinct semantic errors. Fail: any criterion unmet →\to regenerate.

Appendix E Construction Prompts

This section documents all LLM prompts used in benchmark construction. Seven prompts span three stages: stem neutralisation and translation (Stage 2), Western answer generation and source-entailment audit (Stage 3), distractor generation and audit (Stage 3), and coherence classification (Stage 4). All prompts are executed by Sonnet 4.5 (Anthropic, 2025b) unless noted otherwise.

Stem-Neutralisation Prompt (Stage 2)

Prompt 1: stem neutralisation You are rewriting an Islamic-finance question into a framework-neutral client-facing scenario. The rewrite will be used as a benchmark stem in which cultural cues are injected separately; any residual framework terminology would confound the evaluation. Input: {{stem_ar}} [Arabic, from SAHM] Constraints: 1. REMOVE EVERY FRAMEWORK CUE. Eliminate all Islamic- jurisprudence terms (murābaḥa, ijāra, ṣukūk, takāful, qabd, gharar, ribā), all AAOIFI standard numbers, and all Shariah or Islamic-finance compliance references. Also eliminate Western-specific regulatory citations (IFRS, UCC, SEC) if present. 2. PRESERVE FINANCIAL SUBSTANCE. Keep the product type, parties, amounts, term, collateral, and the decision the client needs to make. 3. CONCRETE AND ADVISORY. Output must read as a realistic question a customer would ask a financial advisor. No abstract regulatory or governance framings. 4. REGISTER AND LENGTH. Match SAHM’s question register (plain customer language, 30--80 Arabic tokens). 5. NO LEAKAGE IN EITHER LANGUAGE. Free of framework cues in both the Arabic output and any subsequent English translation. Output (JSON): {"stem_ar_neutral": "...", "rationale": "<cues removed>"}

Bilingual-Translation Prompt (Stage 2)

Prompt 2: bilingual translation Translate the neutralised Arabic question and the SAHM Arabic answer into English. Semantic equivalence is required across the pair. Inputs: - stem_ar_neutral: framework-neutral Arabic stem - answer_ar: SAHM expert answer (verbatim CI cell) Constraints: 1. SEMANTIC EQUIVALENCE. Convey the same financial substance; do not add facts or framework terms. 2. TRANSLITERATIONS IN CI ANSWER. Retain Islamic-finance terms by ISO 233 / ALA-LC transliteration on first use (e.g., murābaḥa, ijāra); use consistently throughout. Do not paraphrase to a Western equivalent. 3. NEUTRAL REGISTER FOR STEM. English stem must remain framework-neutral: no terms a reader would identify as Islamic- or Western-coded. 4. LENGTH. Question: 30--80 tokens. Answer: 80--200 tokens. Same paragraph structure as the Arabic. 5. NUMBERS AND CITATIONS. Carry across without modification. 6. IDIOMATIC ENGLISH. Avoid literal Arabic word order. Output (JSON): {"stem_en_neutral": "...", "answer_en_CI": "...", "translation_notes": "<terms requiring gloss>"}

E.1 Western Answer Generation Prompt (Stage 3)

Prompt 3: Western answer generation (long-context grounded) You are generating the correct Western-finance answer to a financial advisory question. The answer will serve as the CW cell in a four-choice evaluation benchmark. The FULL regulatory source documents for this product cluster are provided below in-context. You MUST ground every claim in these documents. Inputs: - question: the neutralised financial question - source_documents: [FULL TEXT OF 1--3 REGULATORY DOCUMENTS FOR THIS CLUSTER, 10K--50K TOKENS] - islamic_answer: the validated CI cell [REGISTER AND LENGTH REFERENCE ONLY; DO NOT USE AS CONTENT SOURCE] - cluster: product cluster name - western_standard: named standard and section Constraints: 1. SOURCE ENTAILMENT. Every substantive claim must be entailed by the provided regulatory text. Do not introduce claims from training data or general knowledge. If the source does not address a point, do not address it. 2. NO ISLAMIC-ANSWER PARAPHRASE. The CW answer must be independently grounded in Western regulatory text. Do not rephrase, adapt, or mirror the structure of the CI answer. The two answers should read as if written by different domain experts who never saw each other’s work. 3. REGISTER AND LENGTH. 80--200 tokens. Advisory language appropriate for a client-facing interaction. 4. CITE THE SOURCE. Reference the specific standard and section where each recommendation originates. Output (JSON): {"answer_en_CW": "...", "answer_ar_CW": "...", "source_citations": ["section references used"]}

E.2 Western Answer Audit Prompt (Stage 3)

Prompt 4: CW source-entailment audit You are auditing a generated Western-finance answer for source entailment. The regulatory source document and the generated answer are provided below. For EACH substantive claim in the answer: 1. Identify the claim (quote the relevant sentence). 2. Locate the supporting passage in the source document. 3. Classify as SUPPORTED (with source passage quoted) or UNSUPPORTED (with explanation of why the claim is not traceable to the provided text). Also check: - No claim misstates, overstates, or inverts the source. - No claim relies on general knowledge absent from source. - Register and length match the CI answer (80--200 tokens). Output: {"claims": [{"claim": "...", "status": "SUPPORTED", "source_passage": "..."}, ...], "overall": "PASS" or "FAIL", "fail_reason": "..." [if FAIL]}

E.3 Distractor Generation Prompt (Stage 3)

Prompt 5: distractor generation (II and IW) Generate two distractors for a financial advisory benchmark item: one incorrect Islamic answer (II) and one incorrect Western answer (IW). These distractors must fool a model that pattern-matches on financial terminology while being identifiable as wrong by a domain expert. Inputs: - question: neutralised financial question - correct_islamic (CI): validated Islamic answer - correct_western (CW): validated Western answer Constraints: 1. SURFACE PRESERVATION. Retain the same AAOIFI/IFRS standard numbers, Quranic citations, hadith references, scholarly tone, and length (within 10%) as the correct answer in the matching framework. The distractor must LOOK identical to the correct answer at the surface level. 2. CONTENT-LEVEL ERRORS. Introduce 2--3 layered substantive errors. Target categories: - Liability assignment (who bears risk/loss) - Scope conditions (when a rule applies vs. does not) - Instrument identity (applying rules of one contract type to another, e.g., ijāra rules to murābaḥa) - Misapplied legal maxims (fiqhi or Western) - Fabricated conditions (inventing a requirement that does not exist in the cited standard) FORBIDDEN: surface-level swaps (changing a standard number, inverting a percentage, contradicting self within the same paragraph). These are detectable without domain knowledge and would make the distractor trivially identifiable. 3. ASYMMETRIC DIFFICULTY. A model relying on terminological cues (seeing "AAOIFI Std. 8" and "murābaḥa" in the same answer) should find the distractor plausible. A domain expert reading the substance should identify each error. 4. INDEPENDENCE. II errors must be independent of IW errors. The two distractors must not mirror each other. Output (JSON): {"answer_II": "...", "II_traps": ["<error 1>", "<error 2>", "<error 3>"], "answer_IW": "...", "IW_traps": ["<error 1>", "<error 2>", "<error 3>"]}

E.4 Distractor Audit Prompt (Stage 3)

Prompt 6: distractor second-pass audit Audit a generated distractor pair. For each distractor (II and IW), check: 1. TRAP PRESENCE. Is each claimed content trap actually present in the distractor text? Quote the sentence where each trap appears. 2. SEMANTIC DEPTH. Does each error target substantive financial content (liability, scope, instrument, maxim), or is it a surface-level inconsistency (mismatched number, self-contradiction)? Classify each trap as SEMANTIC or SUPERFICIAL. 3. SURFACE PLAUSIBILITY. Does the distractor maintain the same citations, standard numbers, register, and tone as the correct answer? Would a model without domain knowledge find it indistinguishable from the correct answer based on surface features alone? 4. MINIMUM ERROR COUNT. Are at least 2 distinct SEMANTIC errors present? Output: {"II_audit": {"traps_verified": [...], "superficial_count": N, "semantic_count": N, "surface_plausible": true/false, "verdict": "PASS"/"FAIL"}, "IW_audit": {...}} Flag for regeneration if: semantic_count < 2, or any trap is absent, or surface plausibility fails.

E.5 Coherence Classification Prompt (Stage 4)

Prompt 7: coherence classification Classify whether the following (signal, question) pairing produces a coherent evaluation prompt. The signal is a demographic prefix prepended to a financial advisory question. An incoherent pairing would confound the evaluation because model behaviour could be driven by the unnaturalness of the scenario rather than by the cultural signal itself. Inputs: - signal_prefix: the demographic prefix text - question: the neutralised financial question - asker_persona: Personal/retail, SME, or Institutional Evaluate three dimensions: 1. PERSONA CONSISTENCY. Is the inferred asker persona consistent with the signal? A trade professional signal is coherent with a commodity-trading question but incoherent with a personal credit-card question. An institutional investor signal is incoherent with a consumer lending question. 2. CONTENT COMPATIBILITY. Does the signal introduce constraints that contradict the question topic? A gold-venue signal (DMCC/LBMA) is incoherent with a lending question. A gender-inheritance signal is incoherent with an insurance question. 3. INTERACTION PLAUSIBILITY. Does the combined prompt read as a plausible customer interaction at a financial institution? Would an advisor encounter this scenario? Output: {"classification": "Coherent"/"Awkward"/"Incoherent", "justification": "<one sentence>"}
Signal family Personal/retail SME Institutional
Names (all 12 codes)

✓

✓

✓

Locations (Tier-A through Western)

✓

✓

✓

Locations (DMCC/LBMA gold-venue)

✓

 (gold only)

✓

 (gold only)

✓

 (gold only)
Religion explicit (Islamic/Christian/secular)

✓

∼\sim ∼\sim
Religion implicit temporal

✓

∼\sim ✗
Religion implicit behavioural

✓

∼\sim ∼\sim
Religion implicit ritual

✓

∼\sim ✗
Occupation: Islamic/conventional bank ∼\sim

✓

✓

Occupation: SME owner ✗

✓

✗
Occupation: institutional ✗ ✗

✓

Occupation: trade professional

✓

 (Salam/Istisna)

✓

∼\sim
Gender (inheritance)

✓

 (inheritance only)

✓

 (inheritance only)

✓

 (inheritance only)
Table 17: Coherence gating rules by signal family and asker-persona type.

✓

= coherent, ∼\sim = context-dependent, ✗ = incoherent (filtered).

Appendix F Coherence Filtering

F.1 Per-Asker-Persona Coherence Map

Each question is classified by asker-persona type: Personal/retail (∼75%{\sim}75\% of questions), SME (∼15%{\sim}15\%), or Institutional (∼10%{\sim}10\%). Table 17 shows the coherence gating rules applied across signal families and persona types.

F.2 Retention Statistics

After coherence filtering, the evaluation grid retains a mean of 38.6 coherent signals per question per language. Retention varies by signal family: name and religion signals retain >90%{>}90\% of cells; conflict and occupation signals retain ∼60%{\sim}60\% due to persona-topic incompatibilities. The same cells are retained for all 12 evaluation models; coherence filtering is signal-content-dependent, not model-dependent.

F.3 Interface Design

Both interfaces present Arabic and English versions side by side with right-to-left typography for Arabic text. Each annotation task includes a dedicated guideline page accessible within the interface (rubrics from Appendix D). A live dashboard computes Cohen’s κ\kappa and Gwet’s AC1 as annotations accumulate, enabling real-time agreement monitoring during both pilot and full phases.

F.4 Task-Specific Layouts

Neutralisation and translation review (Stage 2).

Annotators see the original SAHM stem alongside the neutralised version in both languages. They rate neutralisation quality (Appendix D) and bilingual adequacy (Appendix D.1).

Western answer review (Stage 3).

Annotators see the full cluster source document, the generated Western answer, and the Islamic answer as structural context. They rate accuracy following Appendix D.1.

Distractor verification (Stage 3).

Annotators see correct and incorrect answers side by side with claimed content traps listed. They verify presence and substantiveness following Appendix D.1.

F.5 Data Management

All annotation sessions are logged with timestamps and anonymised annotator identifiers. Per-item ratings, complete annotation exports, guideline documents, and inter-annotator agreement files are included in the released benchmark materials.

Appendix G Annotator Information

G.1 Panel Composition

The annotation panel comprises five domain experts and one senior researcher:

  • •

    Two Islamic-finance experts with graduate-level training in Islamic jurisprudence (fiqh al-mu’āmalāt) and AAOIFI standards. Native Arabic speakers. Responsible for: corpus classification (Stage 1), Islamic-absence verification (Stage 1, Track 2), neutralisation review (Stage 2), translation adequacy review (Stage 2), bilateral divergence confirmation (Stage 3), Islamic distractor verification (Stage 3), and coherence-classifier validation (Stage 4).

  • •

    Three financial experts with professional backgrounds in IFRS, UCC, Basel III, and TILA primary sources. Responsible for: Western answer accuracy review (Stage 3) and Western distractor verification (Stage 3).

  • •

    Senior researcher. Adjudicates disagreements across all stages, revises flagged items using primary-source evidence, and oversees annotation quality.

G.2 Compensation and Ethics

All annotators are compensated at rates consistent with their professional expertise level and local market conditions. The annotation task was reviewed for ethical compliance with institutional guidelines. Annotator identities are anonymised throughout. Detailed demographic information (educational background, years of domain experience, language proficiency) is included in the anonymised annotator card released with the benchmark.

Appendix H Additional Validity Checks

Three validity checks in details.

Tokenisation rejected as mechanism.

The hypothesis that differential subword tokenisation of cultural signals drives the measured framework lean is tested by computing the pooled Pearson correlation between per-signal subword count under six open-weight tokenisers and the measured Δ​pislamic\Delta p_{\text{islamic}}. The correlations are r=+0.31r=+0.31 (EN) and r=+0.16r=+0.16 (AR), both opposite in sign to the fragmentation hypothesis. Tokenisation is rejected as the mechanism.

Placeholder as null control.

The length-matched neutral placeholder (baseline_placeholder) controls for prompt-length effects. Mean absolute deviation from zero_signal across 24 cells is 0.024 on pislamicp_{\text{islamic}}, not systematically signed (14 cells Western, 10 Islamic). Prompt-length attraction is ruled out.

Coherence-filtering model inclusion.

The coherence-classification LLM (Stage 4) also appears in the evaluation panel. Re-computing the signal hierarchy on this model’s rows alone produces an unchanged ranking, with per-signal values within 0.02 of the panel mean.

Position-bias control.

Per-item choice positions are deterministically shuffled by row_id seed during evaluation (§3). We compute conditional accuracy by correct-answer position for the selected model–language cells in Table 18. Their max–min differences range from 3.33.3 to 52.552.5pp; the maximum occurs for Llama-8B in Arabic (.526−.001=.525.526-.001=.525). A single deterministic shuffle distributes answer positions but does not counterbalance each item. Consequently, residual position confounding cannot be excluded, and the direction-specific and control-set results should be read with this limitation. We plan a full experiment with all 24 letter assignments per item in §7.

Distractor discriminability.

Point-biserial correlation between distractor selection and panel-mean total accuracy (English, baseline; n=113n{=}113 II items, n=208n{=}208 IW items) yields r¯pb=−0.157\bar{r}_{\text{pb}}=-0.157 (II) and −0.337-0.337 (IW), with 73%73\% of II distractors and 95%95\% of IW distractors reaching rpb<−0.1r_{\text{pb}}<-0.1: high-scoring models systematically avoid them, confirming the distractors discriminate competence as designed.

Model Tier Lang P(corr|\,|\,A) P(corr|\,|\,B) P(corr|\,|\,C) P(corr|\,|\,D)
Opus frontier EN .826 .817 .799 .679
Opus frontier AR .842 .831 .859 .826
Sonnet frontier EN .337 .513 .514 .378
Sonnet frontier AR .519 .709 .791 .713
Gemini frontier EN .617 .584 .525 .426
Gemini frontier AR .612 .551 .566 .453
Gemma-3-27b large EN .196 .106 .053 .055
Qwen-14B large EN .090 .090 .045 .078
Gemma-2-9b midsize AR .402 .088 .018 .013
Llama-8B midsize AR .001 .526 .001 .005
ALLaM-7B specialist EN .132 .122 .066 .207
Table 18: Conditional accuracy by CI position for selected model–language cells on the bilateral set. Each row reports accuracy when CI occupies A/B/C/D under a deterministic per-item shuffle. Max−-min is a position-sensitivity proxy and ranges from 3.33.3 to 52.552.5pp in the displayed rows; this is not a fully counterbalanced estimate.

Position-bias limitation.

The selected cells show substantial variation in position sensitivity, peaking at 52.552.5pp for Llama-8B in Arabic. Because each item was evaluated in only one deterministically shuffled order, item difficulty and answer position are not fully separated. Residual position confounding therefore cannot be excluded. We will evaluate all 24 letter assignments per item as described in §7.

Appendix I Full Signal Hierarchy

Figure 10: Cross-tier gate localisation. CW-recovery rate as a function of patch-site layer LL, for the three models in Table 25. Left: absolute layer index. Right: layer normalised to proportional depth. The phase transition lands in the middle depth band (≈0.33\approx 0.33–0.670.67) for all three architectures.

The main text reports the top of the signal hierarchy (Table 4, top 10). Table 19 below lists every one of the 49 non-baseline signals plus the placebo, sorted by panel-mean Δ​pislamic\Delta p_{\text{islamic}} in English. Each row carries the panel-mean shift, its 95% paired cluster-bootstrap CI (B=10,000B=10{,}000, cluster == question), the FDR-significance count (number of 12 model×\timeslanguage cells significant under BH-FDR at α=0.05\alpha=0.05 within (model, language)), and the classification used in the design audit: works_as_designed (effect direction matches design), reverse_read (effect direction opposes design), null_read (CI crosses zero or no FDR-significant cells), uncertain (conflict cells, direction not pre-specified), and behaves_as_null (placebo).

Signal Δ​p\Delta p 95% CI FDR C Signal Δ​p\Delta p 95% CI FDR C
keyword_sharia 0.6610.661 [+0.55,+0.78][+0.55,+0.78] 12 W loc_dmcc_gold 0.0720.072 [−0.01,+0.16][-0.01,+0.16] 0 N
occ_islamic_bank 0.6190.619 [+0.50,+0.74][+0.50,+0.74] 10 W conflict_muslimname_west 0.0670.067 [+0.04,+0.09][+0.04,+0.09] 7 U
stack_max_muslim_gulf 0.5460.546 [+0.43,+0.66][+0.43,+0.66] 10 W name_arab_cultural_f 0.0640.064 [+0.04,+0.09][+0.04,+0.09] 7 W
conflict_isl_occ_w_loc 0.5080.508 [+0.40,+0.62][+0.40,+0.62] 8 U name_muslim_common_m 0.0600.060 [+0.03,+0.09][+0.03,+0.09] 7 W
rel_islamic_explicit 0.4860.486 [+0.37,+0.60][+0.37,+0.60] 12 W rel_christian_explicit 0.0540.054 [−0.04,+0.14][-0.04,+0.14] 7 R
occ_islamic_nonfinance 0.3230.323 [+0.24,+0.41][+0.24,+0.41] 12 W name_christian_arab_m 0.0540.054 [+0.03,+0.08][+0.03,+0.08] 6 R
rel_islamic_impl_ritual 0.2970.297 [+0.22,+0.37][+0.22,+0.37] 12 W name_arab_cultural_m 0.0430.043 [+0.02,+0.07][+0.02,+0.07] 6 W
rel_islamic_impl_practice 0.2680.268 [+0.19,+0.35][+0.19,+0.35] 12 W conflict_arabname_west 0.0110.011 [−0.01,+0.03][-0.01,+0.03] 1 U
loc_gulf_financial 0.2420.242 [+0.15,+0.33][+0.15,+0.33] 11 W name_christian_arab_f 0.0090.009 [−0.01,+0.03][-0.01,+0.03] 2 N
rel_islamic_impl_time 0.2310.231 [+0.16,+0.30][+0.16,+0.30] 12 W occ_trade_professional 0.0080.008 [−0.02,+0.04][-0.02,+0.04] 1 N
conflict_explicitchr_gulf 0.2210.221 [+0.14,+0.31][+0.14,+0.31] 10 U baseline_placeholder 0.0030.003 [−0.01,+0.02][-0.01,+0.02] 2 B
loc_tier_a_mandatory 0.2200.220 [+0.13,+0.31][+0.13,+0.31] 12 W name_western_f −0.020-0.020 [−0.05,+0.01][-0.05,+0.01] 3 N
conflict_christianarab_gulf 0.2000.200 [+0.12,+0.28][+0.12,+0.28] 10 U name_western_m −0.021-0.021 [−0.04,−0.00][-0.04,-0.00] 2 N
conflict_westname_gulf 0.1670.167 [+0.09,+0.24][+0.09,+0.24] 9 U occ_institutional −0.024-0.024 [−0.07,+0.02][-0.07,+0.02] 1 N
loc_nongulf_arab 0.1580.158 [+0.10,+0.22][+0.10,+0.22] 11 W occ_sme_owner −0.028-0.028 [−0.05,−0.00][-0.05,-0.00] 0 B
loc_pakistan_transition 0.1510.151 [+0.09,+0.22][+0.09,+0.22] 9 W loc_western_anchor −0.033-0.033 [−0.07,+0.00][-0.07,+0.00] 2 N
conflict_secular_gulf 0.1290.129 [+0.06,+0.20][+0.06,+0.20] 11 U occ_secular_tech −0.037-0.037 [−0.06,−0.01][-0.06,-0.01] 2 N
name_muslim_theophoric_m 0.1280.128 [+0.08,+0.18][+0.08,+0.18] 9 W occ_conventional_bank −0.048-0.048 [−0.10,−0.00][-0.10,-0.00] 0 N
conflict_multi_agree_part 0.1170.117 [+0.08,+0.15][+0.08,+0.15] 10 U loc_lbma_gold −0.060-0.060 [−0.10,−0.02][-0.10,-0.02] 0 W
conflict_implicitisl_west 0.1110.111 [+0.07,+0.15][+0.07,+0.15] 8 U gen_inheritance_m −0.079-0.079 [−0.24,+0.08][-0.24,+0.08] 0 R
name_muslim_prophetic_m 0.1090.109 [+0.07,+0.15][+0.07,+0.15] 9 W stack_max_western −0.086-0.086 [−0.15,−0.02][-0.15,-0.02] 1 N
name_muslim_prophetic_f 0.1010.101 [+0.07,+0.13][+0.07,+0.13] 10 W conflict_3signals_disagree −0.093-0.093 [−0.17,−0.01][-0.17,-0.01] 4 U
name_muslim_theophoric_f 0.0960.096 [+0.06,+0.13][+0.06,+0.13] 9 W gen_inheritance_f −0.107-0.107 [−0.19,−0.03][-0.19,-0.03] 0 R
loc_muslim_nonarab 0.0910.091 [+0.05,+0.14][+0.05,+0.14] 6 W rel_secular_explicit −0.114-0.114 [−0.21,−0.02][-0.21,-0.02] 3 N
name_muslim_common_f 0.0890.089 [+0.05,+0.12][+0.05,+0.12] 7 W
Table 19: Full 49-signal panel-level hierarchy (English), sorted by panel-mean Δ​pislamic\Delta p_{\text{islamic}}. CI = 95% paired cluster bootstrap (B=10,000B=10{,}000). FDR = number of 12 (model, language) cells significant under BH-FDR at α=0.05\alpha=0.05. Class code: W works as designed; N null read (CI crosses zero or no FDR-significant cells); R reverse read (effect opposite to designed direction); B behaves as null (placebo); U uncertain (conflict cells; direction not pre-specified by design).
Signal Model n CI CW II IW IFR
baseline_zero_signal Gemma-3-27B 48 1 42 1 4 0.50
baseline_zero_signal Qwen2.5-14B 48 5 36 3 4 0.38
occ_institutional Gemma-3-27B 5 0 5 0 0 —
occ_institutional Qwen2.5-14B 5 0 4 0 1 —
keyword_sharia Gemma-3-27B 48 17 2 29 0 0.63
keyword_sharia Qwen2.5-14B 48 13 0 35 0 0.73
Table 20: Dead-zone on d_securities (large tier). occ_institutional (n=5n{=}5 coherent items) is illustrative, consistent with the panel-wide institutional-cue null (Table 4). The Shariah keyword (n=48n{=}48) activates Islamic framing but 6363–73%73\% misquote the rule. IFR is undefined where no Islamic-frame response occurs.

Appendix J Trajectory Threshold Sensitivity

The trajectory partition in §4 (Table 6) classifies each (signal, tier, language) cell using thresholds on Δ​pislamic\Delta p_{\text{islamic}}, Δ​KR\Delta\mathrm{KR}, and Δ​IFR\Delta\mathrm{IFR}. Threshold dependence is addressed by re-computing the partition at two further reasonable settings: a strict setting demanding sharper activation and steeper competence drop, and a lenient setting admitting weaker effects. Table 21 reports the tier counts at each setting; the 30:0 vs 0:26 directional contrast holds in all three.

S1: Strict S2: Default S3: Lenient
Δ​p≥0.10,Δ​KR<−0.15\Delta p\!\geq\!0.10,\,\Delta\mathrm{KR}\!<\!-0.15 Δ​p≥0.05,Δ​KR<−0.10\Delta p\!\geq\!0.05,\,\Delta\mathrm{KR}\!<\!-0.10 Δ​p≥0.03,Δ​KR<−0.05\Delta p\!\geq\!0.03,\,\Delta\mathrm{KR}\!<\!-0.05
Tier lift trap w_r null lift trap w_r null lift trap w_r null
Frontier 26 0 5 17 30 0 10 8 24 1 11 12
Large 0 21 0 27 0 26 2 20 0 29 3 16
Midsize 0 11 0 37 3 19 2 24 7 22 3 16
Specialist 0 17 1 30 2 21 2 23 4 28 2 14
Table 21: Trajectory sensitivity to threshold choice. Per-tier counts (clean_lift / trap / w_release / null) across 48 non-baseline signals (English), re-computed from signal_trajectory_full.csv. The directional contrast (frontier-only lifts, non-frontier-only traps) is invariant.
Refer to caption
Figure 11: |τ||\tau| by cue family and tier (English). Frontier polygon hugs the centre; large sits at the outer ring. Near-circular shape within each tier confirms τ\tau is a model property (σb2/σw2=10.5\sigma^{2}_{b}/\sigma^{2}_{w}{=}10.5).

The contrast is structural, not boundary-dependent: across all three settings, frontier records zero trap cells (one borderline cell under the lenient S3 thresholds) and the largest per-tier clean_lift count, while large records zero clean_lift cells and the largest per-tier trap count. Absolute counts shift monotonically with leniency. The sensitivity addresses category-boundary dependence; an entirely different objection, that the trajectory categories are the wrong instrument, is answered by the continuous trap coefficient τ\tau reported in §5, whose structural-uniformity claim is supported by the variance-decomposition test (between-tier σ2\sigma^{2} exceeds within-tier σ2\sigma^{2} by a factor of 10.510.5).

Appendix K Pair-wise Cluster Spearman Matrix

The cluster-invariance claim in §4 (ρ¯=0.96\bar{\rho}=0.96) is the panel-mean over the 21 unique pair-wise Spearman correlations between cluster-level 50-signal orderings. Table 22 reports every pair in English. Six of seven cluster diagonals (excluding self) hold ρ≥0.89\rho\geq 0.89; e_insurance is the structural outlier, with ρ¯E,⋅=0.76\bar{\rho}_{E,\cdot}=0.76 across its six off-diagonal entries.

a_lend. b_trade c_inv. d_sec. e_ins. f_sarf g_op.
a_lend. — 0.956 0.976 0.976 0.931 0.977 0.985
b_trade — 0.927 0.946 0.905 0.958 0.954
c_inv. — 0.978 0.929 0.975 0.981
d_sec. — 0.932 0.984 0.985
e_ins. — 0.925 0.943
f_sarf — 0.986
g_op. —
Panel mean ρ¯\bar{\rho} 0.9580.958 (all 21 pairs) 0.760.76 for e_insurance-row mean
Table 22: Pair-wise Spearman ρ\rho between cluster-level 50-signal orderings (English). Upper triangle; the matrix is symmetric. e_insurance is the structural outlier (ρ¯E,⋅=0.76\bar{\rho}_{E,\cdot}=0.76).

Six pair-wise correlations are bolded as the row/column containing e_insurance: every cluster pair touching insurance is below the panel-wide minimum-non-insurance pair-wise value (ρ=0.927\rho=0.927, the b_trade–c_investment cell). Insurance is the cluster where the signal hierarchy genuinely re-orders, consistent with the topic-conditional discrimination discussed in §5.

Appendix L Religion=Islam Confusion Heatmap

The Religion==Islam confusion (§5) is the panel’s cleanest tier-conditional finding. Table 23 reports the per-cluster ×\times per-tier mean Δ​pislamic\Delta p_{\text{islamic}} produced by rel_christian_explicit in English: negative entries mark correct Western-pull on a Christian cue; positive entries mark the confusion.

Cluster Frontier Large Midsize Specialist
a_lending −0.110-0.110 0.2040.204 0.0360.036 0.0790.079
b_trade −0.171-0.171 +0.366 0.0670.067 0.1870.187
c_investment −0.136-0.136 0.2100.210 0.0480.048 0.1510.151
d_securities −0.072-0.072 0.1980.198 0.0940.094 0.1390.139
e_insurance -0.190 0.1430.143 -0.018 0.0240.024
f_sarf −0.132-0.132 0.2250.225 0.0870.087 0.1500.150
g_operational −0.137-0.137 0.1780.178 0.0830.083 0.0730.073
Tier mean −0.135-0.135 0.2180.218 0.0570.057 0.1150.115
Table 23: Religion==Islam confusion: rel_christian_explicit Δ​pislamic\Delta p_{\text{islamic}} by cluster ×\times tier (English). Negative == correct Western-pull; positive == Islamic-confusion. Frontier reads Christian-explicit correctly in every cluster; large-tier confuses it in every cluster, with the max at b_trade ×\times Large =+0.366=+0.366. Midsize e_insurance is the unique correctly-suppressed cell among non-frontier tiers.

Three patterns: (i) frontier tier sign is uniformly negative across all seven clusters (−0.072-0.072 to −0.190-0.190); (ii) large tier sign is uniformly positive across all seven clusters (+0.143+0.143 to +0.366+0.366), making the Religion==Islam confusion a tier-acquired representation property rather than a topic effect; (iii) midsize ×\times e_insurance is the unique cell among non-frontier tiers where the model correctly reads Christian-explicit as non-Islamic-activating (−0.018-0.018), echoed weakly by specialist ×\times e_insurance (+0.024+0.024, the smallest specialist confusion in any cluster). e_insurance is therefore the only cluster where partial topic-conditional discrimination emerges across non-frontier tiers.

Appendix M Twelve-Model Summary, Both Languages

Table 24 reports the per-model measurements in both languages, including baseline and post-keyword within-Islamic stereotype rate (IFR\mathrm{IFR}).

Model Tier Lang Baseline pislamicp_{\text{islamic}} Activation gap IFR\mathrm{IFR} @baseline IFR\mathrm{IFR} @keyword Δ​IFR\Delta\mathrm{IFR}
Opus frontier en 0.8410.841 0.1490.149 0.0980.098 0.0750.075 −0.022-0.022
ar 0.9360.936 0.0420.042 0.0670.067 0.0470.047 −0.020-0.020
Sonnet frontier en 0.2830.283 0.6970.697 0.0810.081 0.1140.114 0.0330.033
ar 0.8220.822 0.1090.109 0.0840.084 0.0920.092 0.0080.008
Gemini 3 Flash frontier en 0.4110.411 0.5760.576 0.0240.024 0.0570.057 0.0330.033
ar 0.6220.622 0.3590.359 0.0420.042 0.0440.044 0.0010.001
Gemma-3-27b large en 0.0390.039 0.9240.924 0.3330.333 0.5700.570 0.2370.237
ar 0.1710.171 0.7110.711 0.5000.500 0.5300.530 0.0300.030
Qwen-2.5-14B large en 0.0820.082 0.8880.888 0.3600.360 0.6610.661 0.3010.301
ar 0.2960.296 0.5660.566 0.5330.533 0.7060.706 0.1730.173
Gemma-2-9b midsize en 0.0920.092 0.8220.822 0.4640.464 0.6870.687 0.2230.223
ar 0.2600.260 0.4110.411 0.5190.519 0.5830.583 0.0640.064
Qwen-2.5-7B midsize en 0.0760.076 0.7660.766 0.6090.609 0.6880.688 0.0790.079
ar 0.1550.155 0.3780.378 0.6170.617 0.7650.765 0.1480.148
Gemma-3-4b midsize en 0.1250.125 0.5790.579 0.7110.711 0.7940.794 0.0840.084
ar 0.1710.171 0.1940.194 0.6730.673 0.6940.694 0.0210.021
Llama-3.1-8B midsize en 0.0330.033 0.5920.592 0.6000.600 0.7530.753 0.1530.153
ar 0.2630.263 0.0660.066 0.5750.575 0.5200.520 −0.055-0.055
ALLaM-7B specialist en 0.1580.158 0.5790.579 0.4380.438 0.5670.567 0.1290.129
ar 0.3390.339 0.3550.355 0.4950.495 0.5590.559 0.0640.064
Fanar-9B specialist en 0.1280.128 0.6370.637 0.4870.487 0.6030.603 0.1160.116
ar 0.2890.289 0.4310.431 0.6480.648 0.6260.626 −0.022-0.022
SILMA-9B specialist en 0.1810.181 0.7170.717 0.6000.600 0.7000.700 0.1000.100
ar 0.2630.263 0.3190.319 0.6370.637 0.6670.667 0.0290.029
Table 24: Full 12-model summary, both languages. Baseline pislamicp_{\text{islamic}}, activation gap (keyword_sharia minus baseline), within-Islamic stereotype rate IFR\mathrm{IFR} at baseline and under keyword, and Δ​IFR\Delta\mathrm{IFR} (positive == keyword exposes incompetence). Arabic baselines are higher than English in 11 of 12 models; Arabic activation gaps are smaller in 12 of 12, consistent with the saturation-curve reading in §5. Frontier Δ​IFR\Delta\mathrm{IFR} clusters near zero (all six cells within [−0.022,+0.033][-0.022,\,+0.033]); non-frontier Δ​IFR\Delta\mathrm{IFR} is positive in 15 of 18 cells.

Appendix N Trap Localisation: M1 Experiment Details

§5 reports eight open-weight models and five cues (keyword, religion, occupation, location, and name), giving 40 model–cue cells. Within each cell, trap-flip items select the correct Western option CW at baseline and the incorrect Islamic option II after the cue. We apply both methods below separately to every cell.

Method 1: Activation patching.

For each trap-flip item, we run baseline and cue-conditioned forward passes, then re-run the cue-conditioned pass with the last-token residual at decoder layer LL replaced by the baseline residual at the same layer. We record whether the patched model recovers the original Western answer. Sweeping LL across all decoder layers identifies the gate.

Method 2: Logit lens.

For each trap-flip item, we decode the residual at every layer through the unembedding matrix to obtain the four-choice distribution {pCI,pCW,pII,pIW}\{p_{\mathrm{CI}},p_{\mathrm{CW}},p_{\mathrm{II}},p_{\mathrm{IW}}\} for the baseline and cue-conditioned passes. This shows when pCWp_{\mathrm{CW}} falls and pIIp_{\mathrm{II}} rises.

Representative keyword slice.

The scope-wide analysis uses all 40 cells. Figure 10 and Table 25 show only a representative keyword_sharia slice: three models spanning the Large, Midsize, and Arabic-Centric groups. In this slice, gate depth is 0.670.67–0.840.84 and post-gate recovery is 0.590.59–0.960.96. These rows illustrate model-level trajectories; the cross-model and cross-cue claims in §5 use all 40 cells.

Model Group nflipn_{\text{flip}} Gate Depth Ceiling
Gemma-3-27B Large 36 L37–L58 0.780.78 0.590.59
Gemma-2-9B Midsize 112 L27–L28 0.670.67 0.960.96
ALLaM-7B Arabic-Centric 10 L24–L28 0.840.84 0.770.77
Table 25: Representative keyword_sharia slice of the full 8-model ×\times 5-cue study. nflipn_{\text{flip}} = trap-flip items; Gate = transition from baseline to recovery plateau; Depth = gate midpoint divided by decoder depth; Ceiling = mean recovery above the gate.