[ Extension = .otf, UprightFont = *-regular, BoldFont = *-bold, ItalicFont = *-italic, BoldItalicFont = *-bolditalic, ] [arabic]rm[ Extension = .ttf, UprightFont = Amiri-Regular, BoldFont = Amiri-Bold, ItalicFont = Amiri-Italic, BoldItalicFont = Amiri-BoldItalic, Script=Arabic ]Amiri \tl_set:Ne\truthboxtruthbox \tl_set:Ne\wrongboxwrongbox
Right Frame, Wrong Rule: Cultural Cues
Expose the Financial Knowledge Gap They Were Meant to Close
Abstract
When a question has valid answers under different normative frameworks, a language model must decide which framework to use and whether it can answer correctly within it. We call this setting normative pluralism and study it in Islamic finance using a four-choice taxonomy that separates framework selection from within-framework correctness. This separation reveals the stereotype trap: a cultural cue steers a model toward one framework, but the model selects an incorrect answer within that framework. Across twelve models, two languages, and fifty demographic signals, cultural cues change framework selection and reveal substantial differences in accuracy, especially among non-frontier models. Under the strongest signal, large open-weight models select the Islamic framework 97% of the time. A two-choice evaluation would report near-perfect alignment, although 57–66% of those selections are incorrect. These findings motivate, but do not directly test, the competence-conditioned routing hypothesis: models may favor frameworks where they are more accurate, while cultural cues may expose framework-specific competence gaps.
1 Introduction
A user in Riyadh asks a language model about late loan payments. Under AAOIFI standards11 1 https://aaoifi.com, the correct answer is a charitable penalty with no compounding; under Regulation Z22 2 https://www.consumerfinance.gov/rules-policy/regulations/1026/, the correct answer is late fees with accrued interest at the contractual APR. Neither answer is wrong in the absolute. Which one the model should lean toward depends on the user’s jurisdiction, cultural context, and the signals present in the query. We define this setting as normative pluralism: a question admits valid answers under multiple frameworks, and the appropriate response is calibrated to context rather than fixed to a single ground truth.
Standard cultural-bias benchmarks assume a single correct answer and measure deviation from it. This works for stereotypes, where one association is flatly wrong, but fails when the “bias” is toward one of two legitimate frameworks. Existing preference-only instruments (Parrish et al., 2022; Nangia et al., 2020; Naous et al., 2024) can measure whether a model selects Framework A or B, but they cannot distinguish a model that competently selects a framework from one that selects it through stereotype, activating its surface terminology while producing a factually wrong answer. The distinction matters: when we evaluate twelve models on 304 financial questions with the strongest cultural signal in our benchmark, large open-weight models reach 97% Islamic-frame selection, a result a two-choice instrument would report as near-perfect cultural alignment. In fact, up to 66% of those responses are factually wrong within the Islamic framework they selected.
To expose this failure mode, we introduce a four-choice taxonomy that crosses framework selection with within-framework correctness (Figure 1). Each item contains a correct Islamic answer (CI), a correct Western answer (CW), an incorrect Islamic distractor (II), and an incorrect Western distractor (IW). Selecting II reveals what we define as the stereotype trap: the model leaned toward the culturally appropriate framework but lacks the competence to answer correctly within it.
Across twelve models spanning four capability tiers, two languages, and fifty demographic signals of varying strength, our results motivate a competence-conditioned routing hypothesis. Models may default toward frameworks in which they perform more accurately, while cultural cues proposed as mitigation can expose gaps in within-framework competence. Framework selection and correctness vary jointly across models, but our analysis does not establish a predictive relationship between them. The observed pattern is tier-dependent: frontier models acquire activation without comparable accuracy loss, whereas non-frontier models do not. Scale may not close this gap; targeted training on regulatory source text may help.
We make three contributions: (1) we formalize normative pluralism as an evaluation setting for cultural bias and construct a bilingual Islamic-finance benchmark comprising the bilateral framework set (Set A, , where the four-cell taxonomy operates) with expert-validated four-cell answer grids spanning seven product clusters across AAOIFI standards and 50 demographic signals, paired with the Western-anchor controls (Set B, , isolating within-Western competence) and Islamic-anchor controls (Set C, , isolating within-Islamic competence) (§3); (2) we introduce the stereotype trap as a failure mode structurally invisible to two-choice evaluation, showing that cultural cues redirect framework selection but degrade within-framework correctness for nine of twelve models, with the failure surviving both control sets (§4); (3) we provide preliminary mechanistic evidence that the trap is representational, not superficial, with activation patching and logit-lens analysis locating the commitment point at two-thirds network depth and the trap coefficient remaining near-constant across all signal families within each tier (§5).
2 Related Work
Several lines of work converge on the problem we study, but they leave a critical axis unmeasured.
Cultural and stereotype benchmarks.
Stereotype benchmarks (Parrish et al., 2022; Nangia et al., 2020; Nadeem et al., 2021) test for demographic-attribute association against a single correct answer, an orthogonal failure mode to normative framework selection. Cultural alignment work (Myung et al., 2024; Durmus et al., 2023; Chiu et al., 2025; Rao et al., 2025; Vo and Koyejo, 2025) confirms LLMs possess less non-Western knowledge, but a model may answer Islamic-finance questions correctly under explicit Shariah framing yet suppress that knowledge when contextual signals should trigger it. CAMeL (Naous et al., 2024) is the closest prior, measuring entity preference via token probabilities; we extend it to framework selection in an expert domain, decomposing lean from correctness.
Islamic and financial benchmarks.
Islamic knowledge benchmarks (Atif et al., 2025; Elmahjub et al., 2026; Abdelaal et al., 2026; Alwajih et al., 2025) evaluate jurisprudence and scripture but none isolates finance or tests routing; financial bias surveys (Nie et al., 2024; Lee et al., 2025) document that demographic signals shift recommendations without controlling for framework possession. Financial NLP benchmarks (Chen et al., 2024; Xie et al., 2024) and recent systems (Zhou et al., 2026; Xie et al., 2026; Zhang et al., 2026) assume the applicable framework is fixed; our benchmark tests whether models select the appropriate framework under cultural context and remain correct within it.
Steering costs.
Expert personas reduce factual performance by 3–5 points (Hu et al., 2026), RLHF alignment trades task performance for safety (Lin et al., 2024), steering interventions incur side effects (Stickland et al., 2024), and counterfactual cultural cues drop medical QA accuracy by 3–7 points (Rezaei and Shakeri, 2026). (Khanuja et al., 2026) show non-default cultural directions require explicit anchor cues. None computes the cross-model correlation between activation magnitude and within-framework accuracy, nor reframes the Western default as competence-conditioned routing.
3 Methods
Benchmark Construction
The benchmark comprises three evaluation sets. The bilateral framework set () contains questions admitting valid answers under both Islamic and Western finance; the CI/CW/II/IW taxonomy operates here. The Western-anchor controls () contain questions with a valid Western answer but no distinctly Islamic counterpart, isolating within-Western competence. The Islamic-anchor controls () contain questions whose underlying construct is unique to Islamic jurisprudence, such as waqf (perpetual charitable endowment), isolating within-Islamic competence and ruling out signal conditioning as the stereotype trap’s source.
Both the bilateral set and the Islamic-anchor controls originate from SAHM (Elbadry et al., 2026), an expert-validated Arabic Islamic-finance corpus spanning 48 topic codes and seven product clusters (Table 1). Each SAHM answer serves verbatim as the CI cell, inheriting its expert provenance. Four stages transform these sources into the evaluation instrument (Figure 2): stem neutralization, Western answer generation, distractor generation, and signal injection. All generation steps use Sonnet 4.5 (Anthropic, 2025b); every stage is independently validated by domain experts (– across stages; details in Appendix A).
| Cluster | Islamic anchor | Western anchor | |
|---|---|---|---|
| Consumer & inst. lending | 63 | AAOIFI Std. 8 (murābaḥa), Std. 19 (qarḍ) | TILA Reg. Z §1026.18 |
| Trade finance & forwards | 41 | Std. 10 (salam), Std. 11 (istiṣnā ’) | IFRS 15, UCP 600 |
| Investment & profit-sharing | 31 | Std. 13 (muḍāraba) | Inv. Advisers Act §206 |
| Equity & structured sec. | 48 | Std. 12 (mushāraka), Std. 17 (ṣukūk) | SEC Reg. AB, Rule 144A |
| Insurance & reinsurance | 14 | Std. 26 (takāful) | IFRS 17, Solvency II |
| Asset exchange & collateral | 20 | Std. 1 (ṣarf), Std. 57 (gold) | LBMA Good Delivery Rules |
| Operational & contract. law | 87 | Std. 9 (ijāra), Std. 5 (guarantees), Std. 23 (agency) | IFRS 16, UCC, Basel III |
Stage 1: Corpus Filtering and Stem Neutralization.
Two Islamic-finance experts independently classify each of SAHM’s 811 evaluation samples into three categories: non-advisory (abstract governance where first-person demographic context is inapplicable), Islamic-only (the construct has no Western equivalent), or candidate-bilateral. The pass yields 430 candidate-bilateral and 41 Islamic-only questions; 340 are excluded as non-advisory.
SAHM stems contain framework-specific terminology (murābaḥa, ijāra, AAOIFI standard numbers) that would prime the model toward the Islamic framework before any cultural signal is applied. Neutralization is therefore essential: each stem is rewritten as a concrete financial scenario preserving the product type, customer situation, and financial substance while removing every framework term. The same step produces a parallel English translation of both the neutralized question and the CI cell. Expert verification confirms neutralization quality and bilingual adequacy on all items, with a 2.4% correction rate on borderline cases (rubrics and interface in Appendix A).
Stage 2: Western Answer Generation.
The second stage constructs the CW cell: the correct answer to the same scenario under conventional-finance standards. To ensure CW carries the same provenance standard as CI, every answer is grounded in primary regulatory text rather than model knowledge. For each cluster, the full source documents (one to three per cluster, 12 total; Table 1) are provided in-context alongside the neutralized question. The generator produces an answer under three constraints: no paraphrase of the Islamic answer, every substantive claim entailed by the source text, and register matched to SAHM.
Three financial experts validate each generated answer against the source document on factual accuracy and source entailment. Of 430 candidates, 304 pass: 78 are excluded because the Islamic and Western answers converge on the same economic outcome despite different terminology, and 48 fail accuracy. Two Islamic-finance experts independently confirm that each surviving CI, CW pair recommends substantively different financial products (). Rubrics and audit prompts are in Appendix E.
Sonnet 4.5 generates CW and distractor cells and appears in the evaluation panel. Excluding its evaluation rows leaves the signal hierarchy unchanged (top-10 rank preserved, deviation). To control for distractor provenance, we regenerate the four-cell grid for a 50-item subset using GPT-4o; tier-level IFR patterns are preserved (per-model deviation ).
Stage 3: Distractor Generation.
Each question requires two distractors: II (incorrect Islamic) and IW (incorrect Western). The design goal is asymmetric difficulty: a model relying on terminological pattern-matching should find the distractor plausible, while a domain expert should identify 2 to 3 semantic errors targeting liability assignment, contract scope, or instrument identity. Surface features (AAOIFI standard numbers, regulatory citations, advisory register) are preserved so that the distractor is indistinguishable from the correct answer at the vocabulary level. Domain-matched experts verify each distractor; 80% pass on first generation, the rest are regenerated once again.
Stage 4: Signal Injection and Coherence Filtering.
Cultural signals enter as short first-person prefixes prepended to the neutralized question, isolating the signal effect from register changes that full stem rewriting would introduce. The 50 signal codes span nine families (Figure 3), decomposing framework lean along four dimensions: cultural identity (names at six tiers of religious specificity, crossed with gender), declared or implied belief (from explicit declaration to behavioural cues such as Ramadan observance), regulatory jurisdiction (ranked by Islamic-banking mandate strength), and professional context. A keyword ceiling (keyword_sharia: “I want a Shariah-compliant option”) and two stacked composites establish empirical bounds; ten conflict cells compose opposing cues. Stack and conflict signals concatenate atomic prefixes in fixed order, enabling inclusion, exclusion residual analysis of compositional effects (Complete inventory in Appendix C).
Not every question, signal pairing is coherent: a gold-venue signal paired with a lending question, or an institutional-investor prefix on a consumer credit card query, would confound evaluation. An LLM classifies each cell as coherent, awkward, or incoherent; one expert validates on 100 stratified cells (). Only coherent cells are retained (mean 38.6 per question per language). Coherence filtering is uniform across all twelve evaluation models. The full annotation panel comprises two Islamic-finance experts, three financial experts, and a senior researcher as adjudicator; demographics and compensation are in Appendix G.
4 Results
We inject 50 cultural signals (Figure 3) into financial queries across 12 models to measure two things: whether these cues shift the model from Western to Islamic financial advice, and whether that shift makes the advice better or worse. Each model is evaluated in English and Arabic across 304 bilateral questions (CI/CW/II/IW taxonomy; §3), 64 Western-anchor controls, and 41 Islamic-anchor controls. Frontier: Opus 4.5, Sonnet 4.5 (Anthropic, 2025a; Anthropic, 2025b), Gemini 3 Flash (Google DeepMind, 2025). Large: Gemma-3-27b (Team et al., 2025b), Qwen-2.5-14B (Qwen et al., 2025). Midsize: Gemma-3-4b, Gemma-2-9b (Team et al., 2024), Qwen-2.5-7B, Llama-3.1-8B (Grattafiori et al., 2024). Arabic-centric: ALLaM-7B (Bari et al., 2025), Fanar-9B (Team et al., 2025a), SILMA-9B (silma-ai, 2024).
Metrics.
For each combination of model, language, and signal, denotes the observed proportion of responses assigned to category . Table 2 summarizes the five metrics used in our analysis.
| Metric | Formula | Definition |
| Framework selection and correctness | ||
| Islamic activation () | Proportion of responses selecting the Islamic framework, regardless of correctness. | |
| Knowledge Rate (KR) | Proportion of responses selecting a correct answer, regardless of the chosen framework. | |
| Error inside the selected framework | ||
| Islamic Fake Rate (IFR) | Proportion of Islamic responses that are incorrect. | |
| Western Fake Rate (WFR) | Proportion of Western responses that are incorrect. | |
| Effect of the framework shift | ||
| Trap coefficient () | Change in correctness per unit change in Islamic activation. | |
Framework Sensitivity
The Western default is not uniform across financial topics. Where Islamic products carry recognisable brand names (sukūk, muḍāraba, qarḍ), models show partial Islamic routing at baseline ( –). Where the two frameworks differ only in institutional rules (waqf governance, insolvency priority, documentary credit liability), baseline routing falls near zero (Table 3). Only Opus defaults Islamic (); the remaining panel falls below , with six midsize models below . Models have learned Islamic finance as a product vocabulary, not as a regulatory framework. Yet the knowledge is latent: a single Shariah-compliance request lifts every topic above .
| Topic | Base | Keyword | |
|---|---|---|---|
| Recognised product names | |||
| Qarḍ (interest-free loan) | 4 | 0.396 | 0.875 |
| Sukūk (Islamic bonds) | 13 | 0.395 | 0.942 |
| Gold / ṣarf (exchange) | 17 | 0.211 | 0.853 |
| Muḍāraba (profit-sharing) | 9 | 0.210 | 0.907 |
| Institutional rules only | |||
| Waqf (charitable endowment) | 2 | 0.083 | 0.917 |
| Documentary letters of credit | 7 | 0.048 | 0.774 |
| Insolvency / liquidation | 7 | 0.131 | 0.857 |
| Arbitration | 3 | 0.083 | 0.778 |
| Liquidity management | 7 | 0.103 | 0.786 |
| Guarantees / kafāla | 8 | 0.167 | 0.885 |
Cultural signals shift this baseline asymmetrically (Table 4). The strongest Islamic cue (keyword_sharia, ) is FDR-significant across the full panel; the strongest Western cue (rel_secular_explicit, ) reaches significance in only three model–language cells. The asymmetry is not a coverage artefact: occ_islamic_bank and occ_conventional_bank share the same prompt format and comparable item counts, yet the Islamic-bank cue reaches FDR-significance in ten model–language cells while the conventional-bank cue reaches none. No Western-direction signal we tested reliably moves the model away from its default; the Western frame functions as a prior that holds until an Islamic signal displaces it. The same table exposes a deeper split. Signals the model can pattern-match on identity vocabulary fire reliably: occ_islamic_bank () and occ_islamic_nonfinance () both contain the word “Islamic.” Signals that require structural financial knowledge do not: occ_institutional () and occ_trade_professional () describe roles embedded in Islamic financial infrastructure but contain no identity vocabulary. The inheritance signals present the starkest reversal: designed as strong Islamic cues because Shariah inheritance partitioning is among the most codified areas of Islamic law, they read Western (, ) because “inheritance” maps to Western legal corpora more readily than to fiqh.
| Signal | Family | Designed | FDR | |
| Identity vocabulary fires | ||||
| keyword_sharia | Keyword | Islamic (strong) | 12/12 | |
| occ_islamic_bank | Occupation | Islamic (strong) | 10/12 | |
| stack_max_muslim_gulf | Stack | Islamic (strong) | 10/12 | |
| conflict_isl_occ_w_loc | Conflict | Uncertain | 8/12 | |
| rel_islamic_explicit | Religion | Islamic (strong) | 12/12 | |
| occ_islamic_nonfin. | Occupation | Islamic (medium) | 12/12 | |
| rel_islamic_impl_ritual | Religion | Islamic (medium) | 12/12 | |
| rel_islamic_impl_pract. | Religion | Islamic (medium) | 12/12 | |
| loc_gulf_financial | Location | Islamic (strong) | 11/12 | |
| loc_tier_a_mandatory | Location | Islamic (strong) | 12/12 | |
| Designed Islamic, structural knowledge required | ||||
| occ_institutional | Occupation | Islamic (strong) | 1/12 | |
| occ_trade_prof. | Occupation | Islamic (weak) | 1/12 | |
| gen_inheritance_m | Generalis. | Islamic (strong) | 0/12 | |
| gen_inheritance_f | Generalis. | Islamic (strong) | 0/12 | |
| Strongest Western-direction cues | ||||
| rel_secular_explicit | Religion | Western (medium) | 3/12 | |
| stack_max_western | Stack | Western (strong) | 1/12 | |
| occ_conv_bank | Occupation | Western (medium) | 0/12 | |
Among names (Figure 4), the ordering subverts the intuition that name fame drives activation. Despite “Muhammad” being the most globally recognised Muslim name, theophoric names (Abdullah; ) produce the strongest shift, outpacing prophetic names (Muhammad; ). The model is responding to morphological structure, names that explicitly encode “servant of God”, not to recognition. Arab cultural names (Khaled, Tarek; ) register no shift at all, indistinguishable from the Western placeholder. The most informative case is Christian Arabic names (name_christian_arab, e.g. Boutros; ): they cluster with Muslim-coded names, not with Western names, even though they unambiguously code a non-Islamic religion. The model treats Arab-ethnicity coding itself as an Islamic-finance signal independent of the religion the name actually identifies. Among belief cues, the model reads behavior almost as well as it reads identity. Mentioning Ramadan fasting or Zakat giving () achieves two-thirds of the lift from a direct “I am Muslim” declaration (); a Hijri-calendar date (), designed as a weak control, lands in the same band. The pattern mirrors (Hofmann et al., 2024) implicit/explicit race gap: post-training suppresses what users say outright, not what they reveal through behavior. A query timed to Ramadan or dated in Hijri is read as “Muslim” even when the user never says the word. The same hierarchy holds in Arabic (), with every baseline shifted roughly higher; the cross-lingual ceiling and floor effects are examined in §5.
| Baseline | Keyword_Sharia | |||||
| Tier | Model | IFR | IFR | IFR | ||
| Frontier | Claude Opus 4.5 | 0.098 | 0.990 | 0.075 | 0.022 | |
| Claude Sonnet 4.5 | 0.081 | 0.980 | 0.114 | 0.033 | ||
| Gemini 3 Flash | 0.024 | 0.987 | 0.057 | 0.033 | ||
| Large | Gemma-3-27B | 0.333 | 0.964 | 0.570 | 0.237 | |
| Qwen2.5-14B | 0.360 | 0.970 | 0.661 | 0.301 | ||
| Midsize | Gemma-2-9B | 0.464 | 0.914 | 0.687 | 0.223 | |
| Gemma-3-4B | 0.711 | 0.704 | 0.794 | 0.084 | ||
| Qwen2.5-7B | 0.609 | 0.842 | 0.688 | 0.079 | ||
| Llama-3.1-8B | 0.600 | 0.625 | 0.753 | 0.153 | ||
| Arabic-centric | ALLaM-7B | 0.438 | 0.737 | 0.567 | 0.129 | |
| Fanar-9B | 0.487 | 0.766 | 0.603 | 0.116 | ||
| SILMA-9B | 0.600 | 0.898 | 0.700 | 0.100 | ||
Within-Framework Competence
Section 4 showed that cultural signals redirect models toward Islamic framing. The question is whether that redirection yields a correct answer in the targeted framework. We measure within-frame accuracy symmetrically. The Islamic Fake Rate is the share of Islamic-frame responses stating a wrong AAOIFI rule; its Western counterpart measures the same on the Western side (Table 5). Across the highest-activation signals, frontier models hold IFR between and ; open-weight models range from to , with the smallest models suffering most (Gemma-3-4b at , Llama-8B at , falling to for the largest open-weight model, Gemma-3-27b). The gap is absolute: the worst frontier IFR () is three times lower than the best non-frontier IFR (). Opus is the only model where forced activation improves correctness (IFR drops from to ). At the other extreme, Qwen-14B’s IFR rises to : two-thirds of its Islamic-frame answers cite correct AAOIFI standard numbers while stating rules those standards do not contain.
| Bilateral (Set A) | Western | Islamic | ||
|---|---|---|---|---|
| under keyword_sharia | anchor (B) | anchor (C) | ||
| Tier | IFR | WFR | IFR | |
| Frontier | 0.08 | 0.00 | 0.97 | 0.06 |
| Large | 0.61 | 0.10 | 0.85 | 0.58 |
| Midsize | 0.68 | 0.20 | 0.83 | 0.55 |
| Arabic-centric | 0.58 | 0.27 | 0.80 | 0.52 |
The trap is direction-specific (Table 6). Under the Shariah keyword, non-frontier IFR reaches – while WFR on the same items stays at –: the model fabricates in the Islamic frame but not in the Western frame. Two control sets close the remaining exits. On Western-anchor questions, non-frontier tiers retain – correctness, ruling out general financial incompetence. On Islamic-anchor questions: waqf, Zakat, musaqah, ju’āla, with no Western alternative), non-frontier IFR remains – even under explicit Shariah prompting, ruling out signal-conditioning as the cause.
The trap is also signal-invariant (Table 7). Across eight signals spanning a range in activation strength, non-frontier IFR stays in the – band while WFR stays in –. The trap is not what any particular cue does; it is what the model lacks behind every cue. The split is categorical, not gradient (Figure 5). Frontier produces 30 clean_lift cells and zero trap cells; large tier produces zero lifts and 26 traps, invariant under three threshold settings (Appendix J). The Arabic-centric tier, purpose-built for Arabic and Islamic finance, falls into the same traps as the generalist midsize tier.
| Signal | Non-frontier | Ratio | ||
|---|---|---|---|---|
| IFR | WFR | IFR/WFR | ||
| keyword_sharia | 0.66 | 0.67 | 0.21 | |
| occ_islamic_bank | 0.62 | 0.66 | 0.27 | |
| stack_max_muslim_gulf | 0.55 | 0.62 | 0.29 | |
| rel_islamic_explicit | 0.49 | 0.68 | 0.22 | |
| loc_gulf_financial | 0.24 | 0.64 | 0.21 | |
| rel_implicit_time | 0.23 | 0.63 | 0.28 | |
| name_muslim_theophoric | 0.13 | 0.59 | 0.27 | |
| name_arab_cultural | 0.04 | 0.52 | 0.28 | |
5 Analysis
Why the Trap Exists
The signal-invariance of IFR (§4) implies that routing and execution are served by separate representations. If they shared a single layer, different signal families would produce different correctness costs. They do not (Figure 11): is near-constant across all cue families within each tier, with tier explaining of IFR variance and signal family explaining . The cue picks which frame; the tier determines what the model finds inside it.Cross-lingual evaluation confirms the separation. Switching from English to Arabic shifts frontier routing by to while moving frontier IFR by at most . For non-frontier models, IFR moves with routing ( to ): language shifts routing and execution together, consistent with a shallow layer that entangles the two. The pattern is not two independent language regimes but one shared surface operating at different baselines: the per-signal AR–EN activation gap follows a saturation curve (, ), where Arabic provides a higher floor for weak signals and English provides a higher ceiling for strong ones, converging as signal strength increases.
What Closes the Trap
Neither scale nor language specialisation closes the trap. The Shariah keyword increases IFR on six of seven clusters; the exception is Arabic f_sarf (gold trading), where the keyword reduces IFR across all three Arabic-centric models (), the only cell where every Arabic-centric model escapes. f_sarf is governed by AAOIFI Standard No. 1, the most codified rule in the corpus. Targeted training-data investment, not steering, fills the deep layer where it exists. On d_securities the large tier gives zero correct-Islamic responses on the institutional cue (, Table 20); the Shariah keyword on the full cluster () unlocks Islamic routing but at –. Two pre-registered predictions encoding institutional structure over identity tokens were rejected: occ_trade_professional produces while occ_islamic_bank produces . The model reads “Islamic” as a token; it does not read “Shariah-supervisory pension fund” as a concept.
| Signal | Executed location | Regulatory regime | Islamic share | FDR | |
|---|---|---|---|---|---|
| loc_iran_tehran† | Iran (Tehran) | Islamic banking system | 12/12 | ||
| loc_gulf_financial | Saudi Arabia (Riyadh) | dual, Islamic-dominant | 11/12 | ||
| loc_nongulf_arab | Egypt (Cairo) | dual, mixed | 11/12 | ||
| loc_pakistan_transition | Pakistan (Karachi) | dual, transitioning by 2027 | 9/12 | ||
| loc_muslim_nonarab | Malaysia (Kuala Lumpur) | dual, conventional-dominant | 6/12 | ||
| loc_western_anchor | United Kingdom (London) | conventional, Islamic niche | 2/12 |
Models cannot distinguish regulatory regimes from each other.
The location signals test a factual knowledge question: which financial products are legally available in each jurisdiction? The model fails on this layer. It orders jurisdictions in the right direction (Gulf, then non-Gulf Arab, then Muslim non-Arab, then Western; Table 8) but cannot distinguish statutory single-system jurisdictions (Iran, Sudan: only Islamic banking exists by law) from dual-system Islamic-dominant jurisdictions (Saudi Arabia: conventional banks fully licensed despite Islamic market share) from dual-system mixed jurisdictions (Egypt: Islamic). Adding the SAMA Shariah-disclosure mandate to a Gulf context produces a response statistically indistinguishable from the bare Gulf cue: explicit regulatory framing adds no information beyond the geographic prior. The model treats “Saudi Arabia” as a weak demographic cue, not as the name of a regulatory system where specific products are or are not legally available. The conflict cells expose what this costs. In every Gulf-anchored conflict, location overrides explicit user identity: secular, Christian, and Western-name users all produce mild Islamic activation when paired with Gulf context. The model applies a single rule, “in a Muslim-majority country, route Islamic regardless of identity,” that is correct only in the two statutory single-system jurisdictions worldwide (Iran, Sudan). It is the wrong rule everywhere else. In the dual-system jurisdictions actually tested, both frameworks are legal and the user’s stated identity should determine routing. The model has no representation that Saudi differs from Iran in this dimension.
6 A Single Gate Underlies the Trap
Section 4 showed that cultural cues route models into the Islamic frame at a competence cost. This section asks why. We run activation patching and logit-lens analysis on the eight open-weight models across five cue families, giving model-cue cells. Only patching intervenes on the model, so we treat the lens trajectories as description and rest every causal claim on the patching results.
A single gate, set by the model, not the cue.
For each trap-flip item (baseline picks correct-Western CW, the cue flips it to incorrect-Islamic II), we patch the clean residual into the cue-pass one layer at a time. Across all cells the commitment localises to a single gate in the back half (proportional depth –; Table 9: patching before it does nothing, patching at or after it recovers the Western answer in – of items not a last-layer slip but a deep commitment. Reading the grid two ways separates cause from effect: within a model the five cues commit at nearly identical depth (spread as low as ), but across models the same cue lands at very different depths ( up to ). The gate is cue-invariant and architecture-specific the model sets it, not the cue. This splits the mechanism into what the cue controls and what the model controls.
The cue controls survival through the gate.
Why then does keyword_sharia produce five times the routing of a name ( vs )? Not by activating earlier through the early layers, a Muslim name activates Islamic framing more strongly (Figure 6). The keyword’s advantage is built at the gate: its framing survives (late-layer exceeds other cues by ) while weaker inferential cues (name, location) collapse back to Western. The cue is a volume knob on routing survival, not on the gate or the answer.
| Model | Gate depth | IFR (keyword) | Routing () |
|---|---|---|---|
| ALLaM-7B† | 0.84 | 0.57 | 0.16 |
| Gemma-3-27B | 0.78 | 0.57 | 0.04 |
| Qwen2.5-7B | 0.78 | 0.69 | 0.08 |
| SILMA-9B† | 0.76 | 0.70 | 0.18 |
| Qwen2.5-14B | 0.72 | 0.66 | 0.08 |
| Gemma-2-9B | 0.67 | 0.69 | 0.09 |
| Gemma-3-4B | 0.65 | 0.79 | 0.12 |
| Llama-3.1-8B | 0.55 | 0.75 | 0.03 |
| Correlation with IFR | |||
The model controls competence: a CI–II race.
Competence is within-frame correctness given an Islamic answer, the AAOIFI rule (CI) or a stereotype (II) read as the margin per layer (Figure 7). Its peak sign splits two regimes: in six of eight models the margin is positive early (the correct answer leads) then crosses negative at the gate the model held the answer and suppressed it; in the two Gemma-3 models it is never positive, a genuine knowledge gap. For most models the trap is deletion, not absence. The two axes then meet: gate depth predicts the stereotype rate (; deeper commitment lets competence act before the answer locks; Table 9), while baseline routing is uninformative about it (; Table 9). Routing (surface) and competence (deep) are orthogonal.
The trap is directional, and fine-tuning narrows it.
Counting each trap under its own-direction cue Islamic traps (CWII) under Islamic cues, Western traps (CIIW) under Western cues Islamic traps outnumber Western to , an asymmetry: models abandon a correct Western answer for a wrong Islamic one far more readily than the reverse, the mechanistic correlate of IFRWFR (§4). The two Arabic-finance specialists sit at the favourable extreme of every measure deepest gates (, ), lowest asymmetry (, vs generalist –), largest margins. ALLaM is the sharpest case: it holds the correct answer at a margin the most confident of any model yet still suppresses it to . Fine-tuning populates the deep layer (pushing the gate later and the asymmetry lower) but does not by itself stop the gate from overwriting the answer; the fix the data support is targeted training-data investment, not steering.
7 Conclusion
We introduced normative pluralism as an evaluation setting for cultural bias, with a four-choice taxonomy that separates framework selection from correctness within the framework. The decomposition exposes the stereotype trap: cultural cues shift models toward the Islamic framework, but nine of twelve models select incorrect options within it.
Limitations
This benchmark studies normative pluralism in Islamic finance in Arabic and English, so its findings might not generalise to other multi-framework domains, such as medical ethics and legal systems. The mechanistic analysis covers only open-weight models and limited cues; it cannot establish that these internal patterns hold across other signal families or closed frontier models. The four-choice MCQ format measures selection among pre-authored options rather than open-ended financial advice and remains vulnerable to answer-position bias (Zheng et al., 2023) and format instability (Khan et al., 2025). Because we did not test all answer-order permutations or systematically vary prompt paraphrases, residual position and wording effects cannot be excluded. In addition, only of the pre-registered demographic conditions were implemented, limiting coverage of the intended signal space. These cues probe model sensitivity; they do not establish a user’s preferred framework or the applicable legal regime. Finally, all evaluations reflect a single model-release snapshot and therefore cannot capture changes introduced by later versions or updates.
Ethical Considerations
Risks.
Our results describe a harm that can occur in deployed systems. A user who signals their identity can receive worse advice than one who does not. These outcomes vary together across models: signals can shift framework selection while exposing differences in within-framework accuracy, particularly among non-frontier models; however, our current analysis does not establish a causal or predictive relationship between the two measures. Under the strongest signal, large open models choose the Islamic framework 97% of the time, and 57 to 66% of those selections are incorrect within the Islamic framework according to the benchmark.
Users may not easily identify this failure. Our incorrect options retain the same standard numbers, citations, and tone as the correct ones (Section 3, Stage 3), so an incorrect option can appear authoritative. Figure 1 shows an Islamic-framed option that asserts a thirty-percent deposit requirement and transfers liability for damage to the client before possession. Neither rule appears in the cited standard, and selecting either could materially change a client’s exposure in a real transaction.
The non-frontier models we evaluate should not provide Islamic-finance advice without review by a qualified advisor. Frontier models are substantially more reliable but remain imperfect: the strongest model in our panel still selects an incorrect option in 7.5% of its Islamic-framed selections under the strongest signal, and every frontier model makes some incorrect selections. Cultural cues are often proposed as a way to reduce bias in language models; in this setting, they expose differences in within-framework competence.
What we measure.
We measure model behaviour, not which framework any user should receive. Our correctness measure gives equal credit to a correct option under either framework, and we compute error rates within whichever framework the model selects. Nothing in our evaluation rewards selecting one framework over the other.
We use names, beliefs, locations, and occupations as signals because deployed models may respond to them. We ask how these cues affect model behaviour and who may consequently be exposed to a model’s knowledge gaps. We do not treat demographic identity alone as a gold label for a person’s preferred framework.
Identity as a proxy for applicable rules.
Our results show models using identity cues as proxies for framework selection, even though those cues alone do not determine which rules apply or which framework a user prefers.
Arabic Christian names produce more Islamic framing than Western names in our evaluation, but names alone cannot reliably establish a user’s religion or preferred framework. An explicit Christian declaration also increases Islamic framing across all seven product clusters in the large tier. When a Gulf location is paired with a conflicting cue, location often dominates: secular, Christian, and Western-name cues produce similar routing patterns. These results describe model behaviour; they do not establish that Islamic routing is appropriate for every person in a Muslim-majority jurisdiction.
The location analysis shows a related limitation. Models respond differently to location cues, but the executed prompts name cities rather than legal mandates or regulators. The Tehran result therefore measures a Tehran location-cue effect, not demonstrated knowledge of Iran’s banking requirements. Likewise, the Riyadh result measures a Riyadh location-cue effect rather than explicit SAMA knowledge. We consequently interpret these findings as location-based routing, not as evidence that models distinguish regulatory systems.
What our format measures.
Our four-choice format isolates two quantities: which framework a model selects and whether its selected option is correct within that framework. Measuring them separately makes the failure visible because selecting an Islamic-framed incorrect option counts as an error, not as successful alignment. A different instrument would be needed to evaluate responses that present both frameworks alongside their sources; that is separate from the question we study here.
Scope.
Our questions cover products for which two frameworks specify different procedures for the same client. They do not cover rules that assign different entitlements to different people. Our instrument therefore does not determine when framework-specific personalisation is appropriate, and we take no position on that question.
Data and annotation.
The demographic signals in our prompts are synthetic and were not collected from user interactions. To protect annotator privacy, the public materials exclude names, contact information, consent records, and other direct personal identifiers. Annotation records and annotator cards use pseudonymous identifiers; any mapping between these identifiers and annotator identities is stored separately and is not publicly released. Released demographic information is limited to non-identifying attributes relevant to documenting the composition and expertise of the annotation panel.
References
- IslamicMMLU: a benchmark for evaluating LLMs on Islamic knowledge. arXiv preprint arXiv:2603.23750. Cited by: §2.
- PalmX 2025: the first shared task on benchmarking LLMs on Arabic and Islamic culture. In Proceedings of The Third Arabic Natural Language Processing Conference: Shared Tasks, pp. 774–789. Cited by: §2.
- System Card: Claude Opus 4.5. Anthropic. External Links: Link Cited by: §4.
- System Card: Claude Sonnet 4.5. Anthropic. External Links: Link Cited by: Appendix E, §3, §4.
- Sacred or synthetic? evaluating llm reliability and abstention for religious questions. In Proceedings of the AAAI/ACM Conference on AI, Ethics, and Society, Vol. 8, pp. 217–226. External Links: Document Cited by: §2.
- ALLam: large language models for Arabic and English. In The Thirteenth International Conference on Learning Representations, External Links: Link Cited by: §4.
- Fintextqa: a dataset for long-form financial question answering. In Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pp. 6025–6047. Cited by: §2.
- CulturalBench: a robust, diverse and challenging benchmark for measuring LMs’ cultural knowledge through human-AI red-teaming. In Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pp. 25663–25701. Cited by: §2.
- Towards measuring the representation of subjective global opinions in language models. arXiv preprint arXiv:2306.16388. Cited by: §2.
- SAHM: a benchmark for Arabic financial and Shari’ah-compliant reasoning. In Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pp. 34509–34536. Cited by: §3.
- IslamicLegalBench: evaluating LLMs knowledge and reasoning of Islamic law across 1,200 years of Islamic pluralist legal traditions. arXiv preprint arXiv:2602.21226. Cited by: §2.
- Gemini 3 Flash model card. External Links: Link Cited by: §4.
- The llama 3 herd of models. arXiv preprint arXiv:2407.21783. Cited by: §4.
- AI generates covertly racist decisions about people based on their dialect. Nature 633 (8028), pp. 147–154. Cited by: §4.
- Expert personas improve llm alignment but damage accuracy: bootstrapping intent-based persona routing with prism. arXiv preprint arXiv:2603.18507. Cited by: §2.
- Randomness, not representation: the unreliability of evaluating cultural alignment in LLMs. In Proceedings of the 2025 ACM Conference on Fairness, Accountability, and Transparency, FAccT ’25, New York, NY, USA. External Links: ISBN 9798400714825, Link, Document Cited by: Limitations.
- Steering LLMs for culturally localized generation. arXiv preprint arXiv:2603.23301. Cited by: §2.
- Large language models in finance (finllms). Neural Computing and Applications 37 (30), pp. 24853–24867. Cited by: §2.
- Mitigating the alignment tax of rlhf. In Proceedings of the 2024 conference on empirical methods in natural language processing, pp. 580–606. Cited by: §2.
- Blend: a benchmark for llms on everyday knowledge in diverse cultures and languages. Advances in Neural Information Processing Systems 37, pp. 78104–78146. Cited by: §2.
- StereoSet: measuring stereotypical bias in pretrained language models. In Proceedings of the 59th annual meeting of the association for computational linguistics and the 11th international joint conference on natural language processing (volume 1: long papers), pp. 5356–5371. Cited by: §2.
- CrowS-pairs: a challenge dataset for measuring social biases in masked language models. In Proceedings of the 2020 conference on empirical methods in natural language processing (EMNLP), pp. 1953–1967. Cited by: §1, §2.
- Having beer after prayer? measuring cultural bias in large language models. In Proceedings of the 62nd annual meeting of the association for computational linguistics (volume 1: Long papers), pp. 16366–16393. Cited by: §1, §2.
- A survey of large language models for financial applications: progress, prospects and challenges. arXiv preprint arXiv:2406.11903. Cited by: §2.
- BBQ: a hand-built bias benchmark for question answering. In Findings of the Association for Computational Linguistics: ACL 2022, pp. 2086–2105. Cited by: §1, §2.
- Qwen2.5 technical report. External Links: 2412.15115, Link Cited by: §4.
- NormAd: a framework for measuring the cultural adaptability of large language models. In Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers), pp. 2373–2403. Cited by: §2.
- Counterfactual cultural cues reduce medical QA accuracy in LLMs: identifier vs context effects. arXiv preprint arXiv:2601.20102. Cited by: §2.
- SILMA 9B Instruct v1.0. Note: https://huggingface.co/silma-ai/SILMA-9B-Instruct-v1.0 Cited by: §4.
- Steering without side effects: improving post-deployment control of language models. In Neurips Safe Generative AI Workshop 2024, External Links: Link Cited by: §2.
- Fanar: an arabic-centric multimodal generative ai platform. arXiv preprint arXiv:2501.13944. Cited by: §4.
- Gemma 3 technical report. arXiv preprint arXiv:2503.19786. Cited by: §4.
- Gemma 2: improving open language models at a practical size. arXiv preprint arXiv:2408.00118. Cited by: §4.
- CURE: cultural understanding and reasoning evaluation - a framework for ”thick” culture alignment evaluation in LLMs. ArXiv abs/2511.12014. Cited by: §2.
- Finben: a holistic financial benchmark for large language models. Advances in neural information processing systems 37, pp. 95716–95743. Cited by: §2.
- FinChain: a symbolic benchmark for verifiable chain-of-thought financial reasoning. In Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pp. 14529–14553. Cited by: §2.
- FinReporting: an agentic workflow for localized reporting of cross-jurisdiction financial disclosure. In Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 3: System Demonstrations), pp. 728–735. Cited by: §2.
- Judging llm-as-a-judge with mt-bench and chatbot arena. Advances in neural information processing systems 36, pp. 46595–46623. Cited by: Limitations.
- Fincards: card-based analyst reranking for financial document question answering. In Findings of the Association for Computational Linguistics: ACL 2026, pp. 24836–24852. Cited by: §2.
Appendix A Annotation Interface
The annotation instrument is deployed as two bilingual web applications built on the Streamlit framework and hosted on HuggingFace Spaces:
- •
Neutralisation and translation review (Stage 2):
https://huggingface.co/spaces/Raniahossam33/financial-naturalization-review - •
Western answer verification (Stage 3):
https://huggingface.co/spaces/Raniahossam33/wdb-western-verification
A.1 Data Availability
The benchmark is released under the CulturalDefaultBias organisation at https://huggingface.co/CulturalDefaultBias, which hosts three datasets:
- •
WDB-Set-A-Base (): the bilateral framework set with the four-cell CI/CW/II/IW answer grid, in Arabic and English.
- •
WDB-Set-A-Signal-Grid: the full evaluation grid produced by injecting the 50 signal codes into Set A and retaining only coherent cells.
- •
WDB-Naturalization-Case (): the original SAHM stems paired with their neutralised rewrites, which allows the neutralisation step to be audited directly.
Appendix B Topic Distribution
Table 10 provides the complete distribution of the 304 bilateral-framework questions across 48 AAOIFI topic codes, organised by the seven product clusters defined in Table 1. For each topic, the table lists the Arabic designation, English gloss, item count, core Islamic-finance concept, Western regulatory equivalent, and primary source pairing.
| Cluster | Topic | Islamic concept | Western equivalent | Sources | |
|---|---|---|---|---|---|
| Consumer & institutional lending (63 items) | |||||
| Murabaha | 19 | Cost-plus sale at disclosed markup, installment structure | Finance lease / installment loan with APR | Std. 8 / Reg. Z §1026.18 | |
| Credit facility | 8 | No fee on idle credit commitment | Revolving credit with commitment fee | TILA §128, ECOA | |
| Documentary LC | 7 | Bank acts as agent; no guarantee fee | Letter of credit with commission | UCP 600, ISP98 | |
| Hawala | 7 | Debt assignment to third party | Assignment / novation | UCC §3-203, §9-406 | |
| Insolvency | 7 | Charity-only late penalty; principal unchanged | Chapter 7/11 proceedings, FDCPA | 11 USC §101, UCC §9-322 | |
| Syndicated finance | 5 | Tranches must be segregated by risk | Syndicated loan (LMA/LSTA) | LMA, Basel III CRE 20.36 | |
| Qard | 4 | Interest-free loan; no additional return permitted | Interest-free advance | UCC §3-104 | |
| Defaulting debtor | 3 | No penalty interest; charity-deterrent only | Default interest plus FDCPA remedies | FDCPA, UCC §3-602 | |
| Set-off / netting | 3 | Mutual debt offset with conditions | Set-off rights | UCC §3-601, ISDA Master | |
| Trade finance & forward contracts (41 items) | |||||
| Repo | 11 | Unilateral promise permitted; bilateral constitutes riba | GMRA repurchase agreement | GMRA 2011, UCC §8 | |
| Commodity sales | 10 | Spot delivery required for exchange validity | Futures, forwards, options | CFTC, CEA §1a(47) | |
| Istisna | 7 | Manufacture-to-order with deferred delivery on both sides | Progress-payment construction contract | IFRS 15, FIDIC | |
| Salam | 7 | Full price paid upfront for defined goods delivered later | Prepaid forward contract | CFTC, IFRS 15 | |
| Forex | 6 | Spot settlement only; no swaps or leverage | Spot/forward with leveraged margin | MiFID II, CFTC retail forex | |
| Investment & profit-sharing (31 items) | |||||
| Profit distribution | 12 | Profit allocated by pre-agreed ratio; loss borne by capital | Managed account with performance fee | IAS 32, Inv. Advisers Act §206 | |
| Mudarabah | 9 | One party provides capital, one provides labor; capital bears loss | Limited partnership (LP/GP structure) | RULPA §503, Reg D | |
| Capital protection | 7 | Manager cannot guarantee principal | Principal-protected note, FDIC | Basel III, FDIC | |
| Manager guarantee | 3 | Investment manager prohibited from guaranteeing returns | Fiduciary duty with indemnification | Inv. Advisers Act §206 | |
| Equity & structured securities (48 items) | |||||
| Musharaka | 14 | Profit by pre-agreed ratio; loss by capital contribution | Joint venture, LLC, partnership | UPA, IFRS 11 | |
| Financial securities | 14 | Must be asset-backed with halal business screen | Securities under 1933/1934 Acts | SEC Rule 144A, Reg S | |
| Sukuk | 13 | Asset-backed certificates; investor holds ownership share | Asset-backed securities (ABS) | SEC Reg. AB, IFRS 9 | |
| Commercial papers | 4 | Permitted only if backed by real economic value | Promissory notes, commercial paper | UCC §3-104, SEC §3(a)(3) | |
| Combining contracts | 3 | Permitted if no internal contradiction between terms | Hybrid / composite contracts | Common-law contract integration | |
| Insurance & reinsurance (14 items) | |||||
| Takaful | 9 | Mutual risk pool; surplus distributed to policyholders | Mutual/stock insurance; surplus to company | IFRS 17, Solvency II | |
| Re-takaful | 5 | Mutual reinsurance arrangement preferred | Conventional reinsurance treaties | Solvency II, Munich Re framework | |
| Asset exchange & collateral (20 items) | |||||
| Gold trading | 17 | Spot settlement only; no deferred exchange (AAOIFI Std. 57) | Spot, futures, ETFs, leveraged products | CFTC, LBMA, COMEX | |
| Rahn (pledge) | 3 | Pledgee may not use or benefit from pledged asset | Secured transactions with rehypothecation | UCC Art. 9, SEC 15c3-3 | |
| Operational & contractual law (87 items) | |||||
| Ijara (lease-to-own) | 9 | Lessor retains ownership and bears major maintenance | Finance lease vs. operating lease | IFRS 16, ASC 842, UCC §2A | |
| Kafala (guarantees) | 8 | Guarantee must be gratuitous; no fee permitted | Bank guarantee with 1–3% fee | UCP 600, ISP98, UCC §5 | |
| Debit & credit cards | 7 | Service charges permitted; revolving interest prohibited | Credit cards under CARD Act | TILA, CARD Act 2009, Reg. Z | |
| Liquidity management | 7 | Interest-based borrowing/lending prohibited; Tawarruq used | Money market instruments, SOFR | Basel III LCR/NSFR | |
| Tarawi options | 6 | Pre-contract deliberation period for buyer | Cooling-off and rescission rights | TILA §125, Reg. Z §1026.15 | |
| Cluster | Topic | |
| A. Consumer & institutional lending (63 questions; 9 topics) | ||
| a_lending | Murabaha (cost-plus sale) | 19 |
| Credit agreement | 8 | |
| Documentary letters of credit | 7 | |
| Insolvency | 7 | |
| Hawala (debt transfer) | 7 | |
| Syndicated bank financing | 5 | |
| Qard (interest-free loan) | 4 | |
| Defaulting debtor | 3 | |
| Set-off / netting | 3 | |
| B. Trade finance & forward contracts (41 questions; 5 topics) | ||
| b_trade | Repo / buy-back | 11 |
| Commodity trades on regulated markets | 10 | |
| Istisna and parallel Istisna | 7 | |
| Salam and parallel Salam | 7 | |
| Currency trading | 6 | |
| C. Investment & profit-sharing (31 questions; 4 topics) | ||
| c_investment | Profit distribution in Mudaraba investment accounts | 12 |
| Mudaraba (profit-sharing) | 9 | |
| Capital protection and investment | 7 | |
| Investment-manager guarantees | 3 | |
| D. Equity & structured securities (48 questions; 5 topics) | ||
| d_securities | Securities | 14 |
| Sharika partnership and modern companies | 14 | |
| Sukuk investment certificates | 13 | |
| Commercial papers | 4 | |
| Combining contracts | 3 | |
| E. Insurance & reinsurance (14 questions; 2 topics) | ||
| e_insurance | Islamic insurance (Takaful) | 9 |
| Islamic reinsurance (Retakaful) | 5 | |
| F. Asset exchange & collateral (20 questions; 2 topics) | ||
| f_sarf | Gold and rules of dealing in it | 17 |
| Pledge (rahn) and contemporary applications | 3 | |
| G. Operational & contractual law (87 questions; 21 topics) | ||
| g_operational | Ijara and Ijara ending in ownership | 9 |
| Guarantees | 8 | |
| Liquidity management and deployment | 7 | |
| Debit and credit cards | 7 | |
| Option of deliberation (khiyar al-tarawwi) | 6 | |
| Promise and bilateral promise | 6 | |
| Trust / option rights (khiyarat al-amana) | 5 | |
| Wakala and unauthorised-agent transactions | 5 | |
| Earnest-money deposit (arbun) | 4 | |
| Hiring of persons (labour ijara) | 4 | |
| Competitions and prizes | 4 | |
| Arbitration | 3 | |
| Contingencies affecting obligations | 3 | |
| Standard of impermissible gharar | 3 | |
| Waqf (Islamic endowment) | 2 | |
| Option of soundness (khiyar al-salama) | 2 | |
| Solvent-debtor matters | 2 | |
| Contract rescission by condition | 2 | |
| Banking services in Islamic banks | 2 | |
| Wakala-bil-istithmar (agency for investment) | 2 | |
| Qabd (constructive vs. actual possession) | 1 | |
| Total (48 topics) | 304 | |
Appendix C Signal Inventory
Tables 12–16 provide the complete inventory of 50 signal codes. Each table lists the signal code, verbatim prefix (English; Arabic mirrors the content), signal direction, and design rationale.
| Code | Tier | Example names | Direction | Rationale |
|---|---|---|---|---|
| name_muslim_theophoric_m | Muslim Theophoric | Abdulaziz, Abdulkarim, Abdulrahman | Strong Islamic | Abd-X structure is religiously exclusive |
| name_muslim_theophoric_f | Muslim Theophoric | Amatullah, Fatima al-Zahra, Umm Kulthum | Strong Islamic | Female theophoric equivalents |
| name_muslim_prophetic_m | Muslim Prophetic | Muhammad, Ahmad, Mahmoud, Mustafa | Strong Islamic | Prophet-family names |
| name_muslim_prophetic_f | Muslim Prophetic | Fatima, Aisha, Khadija, Zainab | Strong Islamic | Prophet’s family women |
| name_muslim_common_m | Muslim Common | Omar, Ali, Hassan, Hussein, Khaled | Mid Islamic | Common Muslim; some sect-distinctive |
| name_muslim_common_f | Muslim Common | Hanan, Noura, Amna, Sara | Mid Islamic | Common Muslim female |
| name_arab_cultural_m | Arab Cultural | Rayan, Adam, Laith, Tamim, Rashed | Neutral | Arab ethnicity, religiously neutral |
| name_arab_cultural_f | Arab Cultural | Layan, Lama, Rimas, Tala, Jouri | Neutral | Arab ethnicity, religiously neutral |
| name_christian_arab_m | Christian Arab | Boutros, Charbel, Mina, Girgis, Fadi | Neg. control | Arab surface, non-Muslim |
| name_christian_arab_f | Christian Arab | Marina, Veronica, Madonna, Caroline | Neg. control | Arab Christian female |
| name_western_m | Western | David, James, Michael, Robert | Western base | Western-default floor |
| name_western_f | Western | Emily, Jessica, Olivia, Sophia | Western base | Western-default floor |
| Code | Prefix | Type | Coherence restriction |
|---|---|---|---|
| rel_islamic_explicit | “I am a Muslim.” | Explicit | All asker-personas |
| rel_islamic_implicit_time | “After Friday prayer, I wanted to ask…” | Implicit: temporal | Personal/retail only |
| rel_islamic_implicit_practice | “During Ramadan / before Iftar / after Hajj…” | Implicit: behavioural | Personal/retail only |
| rel_islamic_implicit_ritual | “After paying my Zakat…” | Implicit: ritual | Personal only |
| rel_christian_explicit | “I am a Christian.” | Explicit | All asker-personas |
| rel_secular_explicit | “I follow no religion.” | Explicit | All asker-personas |
| Code | Cities / jurisdictions | Mandate tier | Strength |
|---|---|---|---|
| loc_tier_a_mandatory | Tehran, Khartoum | Fully mandated Islamic-only | Strongest |
| loc_pakistan_transition | Karachi, Lahore, Islamabad | Tier-A transitioning (FSC 2027) | Strong + temporal |
| loc_gulf_financial | Riyadh, Dubai, Doha, Abu Dhabi, Manama, Kuwait City | Shariah governance mandatory | Strong |
| loc_muslim_nonarab | Kuala Lumpur, Jakarta, Istanbul, Dhaka | Dual-system | Moderate |
| loc_nongulf_arab | Cairo, Amman, Casablanca, Beirut, Tunis | Conventional-dominant | Weak |
| loc_western_anchor | London, New York, Tokyo, Paris, Sydney | Western-default | Baseline |
| loc_dmcc_gold | Dubai DMCC | AAOIFI Std. 57 jurisdiction | Gold cluster only |
| loc_lbma_gold | London LBMA / COMEX | Conventional gold market | Gold cluster only |
| Code | Prefix | Direction | Notes |
|---|---|---|---|
| occ_islamic_bank | “I work at an Islamic bank (e.g., Al-Rajhi, DIB).” | Strong Islamic | Contains framework keyword |
| occ_conventional_bank | “I work at a conventional bank (e.g., JPMorgan).” | Strong Western | Contains framework keyword in reverse |
| occ_islamic_nonfinance | “I work at an Islamic charity / mosque.” | Moderate Islamic | Lower keyword leakage |
| occ_sme_owner | “I run a small business.” | Neutral | Deployment-relevant null signal |
| occ_institutional | “I work at a hedge fund / pension fund.” | Neutral-inst. | Institutional context |
| occ_trade_professional | “I am a wheat farmer / commodity trader.” | Trade-context | Relevant to Salam, Istisna |
| occ_secular_tech | “I work at a tech company.” | Neutral | Control cell |
| Code | Type | Composition / description |
| baseline_zero_signal | Baseline | No demographic prefix; measures the unconditional prior |
| baseline_placeholder | Baseline | Length-matched neutral placeholder (“Person X in Location Y”); controls for prompt-length effects |
| keyword_sharia | Keyword | “I want a Shariah-compliant option.” Activation ceiling |
| stack_max_muslim_gulf | Stack (Islamic) | Muslim-theophoric name + Gulf location + explicit Islamic religion + Islamic-bank occupation |
| stack_max_western | Stack (Western) | Western name + Western location + Christian religion + conventional-bank occupation |
| conflict_westname_gulf | Conflict () | Western name + Gulf location |
| conflict_arabname_west | Conflict () | Arab name + Western location |
| conflict_christianarab_gulf | Conflict () | Christian-Arab name + Gulf location |
| conflict_muslimname_west | Conflict () | Theophoric name + Western location |
| conflict_implicitisl_west | Conflict () | Implicit Islamic + Western name + Western location |
| conflict_explicitchr_gulf | Conflict () | Explicit Christian + Gulf location + Arab name |
| conflict_secular_gulf | Conflict () | Secular declaration + Gulf location |
| conflict_islamic_occ_west_loc | Conflict () | Islamic-bank occupation + Western location |
| conflict_3signals_disagree | Conflict () | Theophoric name + secular religion + Western location |
| conflict_multi_agree_partial | Conflict () | Arab name + Gulf location + Christian religion |
| gen_inheritance_m | Generalisation | Male inheritance context (Quran 4:11) |
| gen_inheritance_f | Generalisation | Female inheritance context |
Appendix D Annotation Guidelines
This section documents the rubrics used at each verification step in the construction pipeline (§3). All rubrics were presented to annotators through the annotation interface (Appendix A) with worked examples.
Three-Way Classification Rubric (Stage 1)
Two Islamic-finance experts classify each SAHM evaluation sample into one of three categories.
Islamic-Absence Verification Rubric (Stage 1, Track 2)
Two Islamic-finance experts verify that each Western-anchor candidate has no distinctly Islamic counterpart.
Neutralisation Quality Rubric (Stage 2)
D.1 Bilingual Adequacy Rubric (Stage 2)
Western Answer Accuracy Rubric (Stage 3)
Three financial experts rate each generated Western answer against the source regulatory document.
Bilateral Divergence Rubric (Stage 3)
Two Islamic-finance experts judge whether each validated CI–CW pair recommends substantively different financial products.
Distractor Verification Rubric (Stage 3)
Domain-matched annotators verify each distractor (Islamic-finance expert for II cells, financial expert for IW cells).
Appendix E Construction Prompts
This section documents all LLM prompts used in benchmark construction. Seven prompts span three stages: stem neutralisation and translation (Stage 2), Western answer generation and source-entailment audit (Stage 3), distractor generation and audit (Stage 3), and coherence classification (Stage 4). All prompts are executed by Sonnet 4.5 (Anthropic, 2025b) unless noted otherwise.
Stem-Neutralisation Prompt (Stage 2)
Bilingual-Translation Prompt (Stage 2)
E.1 Western Answer Generation Prompt (Stage 3)
E.2 Western Answer Audit Prompt (Stage 3)
E.3 Distractor Generation Prompt (Stage 3)
E.4 Distractor Audit Prompt (Stage 3)
E.5 Coherence Classification Prompt (Stage 4)
| Signal family | Personal/retail | SME | Institutional |
| Names (all 12 codes) |
✓ |
✓ |
✓ |
| Locations (Tier-A through Western) |
✓ |
✓ |
✓ |
| Locations (DMCC/LBMA gold-venue) |
✓ |
✓ |
✓ |
| Religion explicit (Islamic/Christian/secular) |
✓ |
||
| Religion implicit temporal |
✓ |
✗ | |
| Religion implicit behavioural |
✓ |
||
| Religion implicit ritual |
✓ |
✗ | |
| Occupation: Islamic/conventional bank |
✓ |
✓ | |
| Occupation: SME owner | ✗ |
✓ |
✗ |
| Occupation: institutional | ✗ | ✗ |
✓ |
| Occupation: trade professional |
✓ |
✓ |
|
| Gender (inheritance) |
✓ |
✓ |
✓ |
✓
Appendix F Coherence Filtering
F.1 Per-Asker-Persona Coherence Map
Each question is classified by asker-persona type: Personal/retail ( of questions), SME (), or Institutional (). Table 17 shows the coherence gating rules applied across signal families and persona types.
F.2 Retention Statistics
After coherence filtering, the evaluation grid retains a mean of 38.6 coherent signals per question per language. Retention varies by signal family: name and religion signals retain of cells; conflict and occupation signals retain due to persona-topic incompatibilities. The same cells are retained for all 12 evaluation models; coherence filtering is signal-content-dependent, not model-dependent.
F.3 Interface Design
Both interfaces present Arabic and English versions side by side with right-to-left typography for Arabic text. Each annotation task includes a dedicated guideline page accessible within the interface (rubrics from Appendix D). A live dashboard computes Cohen’s and Gwet’s AC1 as annotations accumulate, enabling real-time agreement monitoring during both pilot and full phases.
F.4 Task-Specific Layouts
Neutralisation and translation review (Stage 2).
Western answer review (Stage 3).
Annotators see the full cluster source document, the generated Western answer, and the Islamic answer as structural context. They rate accuracy following Appendix D.1.
Distractor verification (Stage 3).
Annotators see correct and incorrect answers side by side with claimed content traps listed. They verify presence and substantiveness following Appendix D.1.
F.5 Data Management
All annotation sessions are logged with timestamps and anonymised annotator identifiers. Per-item ratings, complete annotation exports, guideline documents, and inter-annotator agreement files are included in the released benchmark materials.
Appendix G Annotator Information
G.1 Panel Composition
The annotation panel comprises five domain experts and one senior researcher:
- •
Two Islamic-finance experts with graduate-level training in Islamic jurisprudence (fiqh al-mu’āmalāt) and AAOIFI standards. Native Arabic speakers. Responsible for: corpus classification (Stage 1), Islamic-absence verification (Stage 1, Track 2), neutralisation review (Stage 2), translation adequacy review (Stage 2), bilateral divergence confirmation (Stage 3), Islamic distractor verification (Stage 3), and coherence-classifier validation (Stage 4).
- •
Three financial experts with professional backgrounds in IFRS, UCC, Basel III, and TILA primary sources. Responsible for: Western answer accuracy review (Stage 3) and Western distractor verification (Stage 3).
- •
Senior researcher. Adjudicates disagreements across all stages, revises flagged items using primary-source evidence, and oversees annotation quality.
G.2 Compensation and Ethics
All annotators are compensated at rates consistent with their professional expertise level and local market conditions. The annotation task was reviewed for ethical compliance with institutional guidelines. Annotator identities are anonymised throughout. Detailed demographic information (educational background, years of domain experience, language proficiency) is included in the anonymised annotator card released with the benchmark.
Appendix H Additional Validity Checks
Three validity checks in details.
Tokenisation rejected as mechanism.
The hypothesis that differential subword tokenisation of cultural signals drives the measured framework lean is tested by computing the pooled Pearson correlation between per-signal subword count under six open-weight tokenisers and the measured . The correlations are (EN) and (AR), both opposite in sign to the fragmentation hypothesis. Tokenisation is rejected as the mechanism.
Placeholder as null control.
The length-matched neutral placeholder (baseline_placeholder) controls for prompt-length effects. Mean absolute deviation from zero_signal across 24 cells is 0.024 on , not systematically signed (14 cells Western, 10 Islamic). Prompt-length attraction is ruled out.
Coherence-filtering model inclusion.
The coherence-classification LLM (Stage 4) also appears in the evaluation panel. Re-computing the signal hierarchy on this model’s rows alone produces an unchanged ranking, with per-signal values within 0.02 of the panel mean.
Position-bias control.
Per-item choice positions are deterministically shuffled by row_id seed during evaluation (§3). We compute conditional accuracy by correct-answer position for the selected model–language cells in Table 18. Their max–min differences range from to pp; the maximum occurs for Llama-8B in Arabic (). A single deterministic shuffle distributes answer positions but does not counterbalance each item. Consequently, residual position confounding cannot be excluded, and the direction-specific and control-set results should be read with this limitation. We plan a full experiment with all 24 letter assignments per item in §7.
Distractor discriminability.
Point-biserial correlation between distractor selection and panel-mean total accuracy (English, baseline; II items, IW items) yields (II) and (IW), with of II distractors and of IW distractors reaching : high-scoring models systematically avoid them, confirming the distractors discriminate competence as designed.
| Model | Tier | Lang | P(corrA) | P(corrB) | P(corrC) | P(corrD) |
|---|---|---|---|---|---|---|
| Opus | frontier | EN | .826 | .817 | .799 | .679 |
| Opus | frontier | AR | .842 | .831 | .859 | .826 |
| Sonnet | frontier | EN | .337 | .513 | .514 | .378 |
| Sonnet | frontier | AR | .519 | .709 | .791 | .713 |
| Gemini | frontier | EN | .617 | .584 | .525 | .426 |
| Gemini | frontier | AR | .612 | .551 | .566 | .453 |
| Gemma-3-27b | large | EN | .196 | .106 | .053 | .055 |
| Qwen-14B | large | EN | .090 | .090 | .045 | .078 |
| Gemma-2-9b | midsize | AR | .402 | .088 | .018 | .013 |
| Llama-8B | midsize | AR | .001 | .526 | .001 | .005 |
| ALLaM-7B | specialist | EN | .132 | .122 | .066 | .207 |
Position-bias limitation.
The selected cells show substantial variation in position sensitivity, peaking at pp for Llama-8B in Arabic. Because each item was evaluated in only one deterministically shuffled order, item difficulty and answer position are not fully separated. Residual position confounding therefore cannot be excluded. We will evaluate all 24 letter assignments per item as described in §7.
Appendix I Full Signal Hierarchy
The main text reports the top of the signal hierarchy (Table 4, top 10). Table 19 below lists every one of the 49 non-baseline signals plus the placebo, sorted by panel-mean in English. Each row carries the panel-mean shift, its 95% paired cluster-bootstrap CI (, cluster question), the FDR-significance count (number of 12 modellanguage cells significant under BH-FDR at within (model, language)), and the classification used in the design audit: works_as_designed (effect direction matches design), reverse_read (effect direction opposes design), null_read (CI crosses zero or no FDR-significant cells), uncertain (conflict cells, direction not pre-specified), and behaves_as_null (placebo).
| Signal | 95% CI | FDR | C | Signal | 95% CI | FDR | C | ||
|---|---|---|---|---|---|---|---|---|---|
| keyword_sharia | 12 | W | loc_dmcc_gold | 0 | N | ||||
| occ_islamic_bank | 10 | W | conflict_muslimname_west | 7 | U | ||||
| stack_max_muslim_gulf | 10 | W | name_arab_cultural_f | 7 | W | ||||
| conflict_isl_occ_w_loc | 8 | U | name_muslim_common_m | 7 | W | ||||
| rel_islamic_explicit | 12 | W | rel_christian_explicit | 7 | R | ||||
| occ_islamic_nonfinance | 12 | W | name_christian_arab_m | 6 | R | ||||
| rel_islamic_impl_ritual | 12 | W | name_arab_cultural_m | 6 | W | ||||
| rel_islamic_impl_practice | 12 | W | conflict_arabname_west | 1 | U | ||||
| loc_gulf_financial | 11 | W | name_christian_arab_f | 2 | N | ||||
| rel_islamic_impl_time | 12 | W | occ_trade_professional | 1 | N | ||||
| conflict_explicitchr_gulf | 10 | U | baseline_placeholder | 2 | B | ||||
| loc_tier_a_mandatory | 12 | W | name_western_f | 3 | N | ||||
| conflict_christianarab_gulf | 10 | U | name_western_m | 2 | N | ||||
| conflict_westname_gulf | 9 | U | occ_institutional | 1 | N | ||||
| loc_nongulf_arab | 11 | W | occ_sme_owner | 0 | B | ||||
| loc_pakistan_transition | 9 | W | loc_western_anchor | 2 | N | ||||
| conflict_secular_gulf | 11 | U | occ_secular_tech | 2 | N | ||||
| name_muslim_theophoric_m | 9 | W | occ_conventional_bank | 0 | N | ||||
| conflict_multi_agree_part | 10 | U | loc_lbma_gold | 0 | W | ||||
| conflict_implicitisl_west | 8 | U | gen_inheritance_m | 0 | R | ||||
| name_muslim_prophetic_m | 9 | W | stack_max_western | 1 | N | ||||
| name_muslim_prophetic_f | 10 | W | conflict_3signals_disagree | 4 | U | ||||
| name_muslim_theophoric_f | 9 | W | gen_inheritance_f | 0 | R | ||||
| loc_muslim_nonarab | 6 | W | rel_secular_explicit | 3 | N | ||||
| name_muslim_common_f | 7 | W |
| Signal | Model | n | CI | CW | II | IW | IFR |
|---|---|---|---|---|---|---|---|
| baseline_zero_signal | Gemma-3-27B | 48 | 1 | 42 | 1 | 4 | 0.50 |
| baseline_zero_signal | Qwen2.5-14B | 48 | 5 | 36 | 3 | 4 | 0.38 |
| occ_institutional | Gemma-3-27B | 5 | 0 | 5 | 0 | 0 | — |
| occ_institutional | Qwen2.5-14B | 5 | 0 | 4 | 0 | 1 | — |
| keyword_sharia | Gemma-3-27B | 48 | 17 | 2 | 29 | 0 | 0.63 |
| keyword_sharia | Qwen2.5-14B | 48 | 13 | 0 | 35 | 0 | 0.73 |
Appendix J Trajectory Threshold Sensitivity
The trajectory partition in §4 (Table 6) classifies each (signal, tier, language) cell using thresholds on , , and . Threshold dependence is addressed by re-computing the partition at two further reasonable settings: a strict setting demanding sharper activation and steeper competence drop, and a lenient setting admitting weaker effects. Table 21 reports the tier counts at each setting; the 30:0 vs 0:26 directional contrast holds in all three.
| S1: Strict | S2: Default | S3: Lenient | ||||||||||
| Tier | lift | trap | w_r | null | lift | trap | w_r | null | lift | trap | w_r | null |
| Frontier | 26 | 0 | 5 | 17 | 30 | 0 | 10 | 8 | 24 | 1 | 11 | 12 |
| Large | 0 | 21 | 0 | 27 | 0 | 26 | 2 | 20 | 0 | 29 | 3 | 16 |
| Midsize | 0 | 11 | 0 | 37 | 3 | 19 | 2 | 24 | 7 | 22 | 3 | 16 |
| Specialist | 0 | 17 | 1 | 30 | 2 | 21 | 2 | 23 | 4 | 28 | 2 | 14 |
The contrast is structural, not boundary-dependent: across all three settings, frontier records zero trap cells (one borderline cell under the lenient S3 thresholds) and the largest per-tier clean_lift count, while large records zero clean_lift cells and the largest per-tier trap count. Absolute counts shift monotonically with leniency. The sensitivity addresses category-boundary dependence; an entirely different objection, that the trajectory categories are the wrong instrument, is answered by the continuous trap coefficient reported in §5, whose structural-uniformity claim is supported by the variance-decomposition test (between-tier exceeds within-tier by a factor of ).
Appendix K Pair-wise Cluster Spearman Matrix
The cluster-invariance claim in §4 () is the panel-mean over the 21 unique pair-wise Spearman correlations between cluster-level 50-signal orderings. Table 22 reports every pair in English. Six of seven cluster diagonals (excluding self) hold ; e_insurance is the structural outlier, with across its six off-diagonal entries.
| a_lend. | b_trade | c_inv. | d_sec. | e_ins. | f_sarf | g_op. | |
|---|---|---|---|---|---|---|---|
| a_lend. | — | 0.956 | 0.976 | 0.976 | 0.931 | 0.977 | 0.985 |
| b_trade | — | 0.927 | 0.946 | 0.905 | 0.958 | 0.954 | |
| c_inv. | — | 0.978 | 0.929 | 0.975 | 0.981 | ||
| d_sec. | — | 0.932 | 0.984 | 0.985 | |||
| e_ins. | — | 0.925 | 0.943 | ||||
| f_sarf | — | 0.986 | |||||
| g_op. | — | ||||||
| Panel mean | (all 21 pairs) | for e_insurance-row mean | |||||
Six pair-wise correlations are bolded as the row/column containing e_insurance: every cluster pair touching insurance is below the panel-wide minimum-non-insurance pair-wise value (, the b_trade–c_investment cell). Insurance is the cluster where the signal hierarchy genuinely re-orders, consistent with the topic-conditional discrimination discussed in §5.
Appendix L Religion=Islam Confusion Heatmap
The ReligionIslam confusion (§5) is the panel’s cleanest tier-conditional finding. Table 23 reports the per-cluster per-tier mean produced by rel_christian_explicit in English: negative entries mark correct Western-pull on a Christian cue; positive entries mark the confusion.
| Cluster | Frontier | Large | Midsize | Specialist |
|---|---|---|---|---|
| a_lending | ||||
| b_trade | +0.366 | |||
| c_investment | ||||
| d_securities | ||||
| e_insurance | -0.190 | -0.018 | ||
| f_sarf | ||||
| g_operational | ||||
| Tier mean |
Three patterns: (i) frontier tier sign is uniformly negative across all seven clusters ( to ); (ii) large tier sign is uniformly positive across all seven clusters ( to ), making the ReligionIslam confusion a tier-acquired representation property rather than a topic effect; (iii) midsize e_insurance is the unique cell among non-frontier tiers where the model correctly reads Christian-explicit as non-Islamic-activating (), echoed weakly by specialist e_insurance (, the smallest specialist confusion in any cluster). e_insurance is therefore the only cluster where partial topic-conditional discrimination emerges across non-frontier tiers.
Appendix M Twelve-Model Summary, Both Languages
Table 24 reports the per-model measurements in both languages, including baseline and post-keyword within-Islamic stereotype rate ().
| Model | Tier | Lang | Baseline | Activation gap | @baseline | @keyword | |
|---|---|---|---|---|---|---|---|
| Opus | frontier | en | |||||
| ar | |||||||
| Sonnet | frontier | en | |||||
| ar | |||||||
| Gemini 3 Flash | frontier | en | |||||
| ar | |||||||
| Gemma-3-27b | large | en | |||||
| ar | |||||||
| Qwen-2.5-14B | large | en | |||||
| ar | |||||||
| Gemma-2-9b | midsize | en | |||||
| ar | |||||||
| Qwen-2.5-7B | midsize | en | |||||
| ar | |||||||
| Gemma-3-4b | midsize | en | |||||
| ar | |||||||
| Llama-3.1-8B | midsize | en | |||||
| ar | |||||||
| ALLaM-7B | specialist | en | |||||
| ar | |||||||
| Fanar-9B | specialist | en | |||||
| ar | |||||||
| SILMA-9B | specialist | en | |||||
| ar |
Appendix N Trap Localisation: M1 Experiment Details
§5 reports eight open-weight models and five cues (keyword, religion, occupation, location, and name), giving 40 model–cue cells. Within each cell, trap-flip items select the correct Western option CW at baseline and the incorrect Islamic option II after the cue. We apply both methods below separately to every cell.
Method 1: Activation patching.
For each trap-flip item, we run baseline and cue-conditioned forward passes, then re-run the cue-conditioned pass with the last-token residual at decoder layer replaced by the baseline residual at the same layer. We record whether the patched model recovers the original Western answer. Sweeping across all decoder layers identifies the gate.
Method 2: Logit lens.
For each trap-flip item, we decode the residual at every layer through the unembedding matrix to obtain the four-choice distribution for the baseline and cue-conditioned passes. This shows when falls and rises.
Representative keyword slice.
The scope-wide analysis uses all 40 cells. Figure 10 and Table 25 show only a representative keyword_sharia slice: three models spanning the Large, Midsize, and Arabic-Centric groups. In this slice, gate depth is – and post-gate recovery is –. These rows illustrate model-level trajectories; the cross-model and cross-cue claims in §5 use all 40 cells.
| Model | Group | Gate | Depth | Ceiling | |
|---|---|---|---|---|---|
| Gemma-3-27B | Large | 36 | L37–L58 | ||
| Gemma-2-9B | Midsize | 112 | L27–L28 | ||
| ALLaM-7B | Arabic-Centric | 10 | L24–L28 |