arXiv is now an independent nonprofit! Learn more
License: CC BY 4.0
arXiv:2608.00622v1 [cs.CL] 01 Aug 2026

A Heuristic Perspective on Debiasing Language Models

Tian Lan thanks: Equal Contribution Affiliation:  Inner Mongolia University Email: velikayascarlet@gmail.com    Yemin Wang11footnotemark: 1 Affiliation:  Xiamen University Email: cssxd@imu.edu.cn    Chuancheng Shi Affiliation:  University of Sydney    Xiangyu Wu Affiliation:  Nanjing University of Science and Technology    Zesheng Shi Affiliation:  Harbin Institute of Technology    Yuan Wang Affiliation:  Zhejiang University    Jiang Li Affiliation:  Inner Mongolia University    Guanglai Gao Affiliation:  Inner Mongolia University    Xiangdong Su thanks: Corresponding Author Affiliation:  Inner Mongolia University
Abstract

Language models (LMs) often acquire various biases during pre-training and may express them in interactions, potentially causing social harm. Existing methods often rely on counterfactual augmentation or representation projection. These strategies remain limited in practice due to their high computational costs and difficulty in scaling to larger models. Additionally, many of these strategies require manual data annotation, narrowing their scope to specific cultures and bias categories. To overcome these limitations, we propose HEIMAT, a HEurIstic-style autoMATic debiasing framework for LMs. HEIMAT consists of two main steps: bias disclosure and debiasing fine-tuning. In the first step, it uses simple templates to construct heuristic prompts, which are applied to reveal model biases and generate corresponding context prompts. In the second step, it fine-tunes the model by minimizing the Jensen-Shannon divergence of predictions on these context prompts to reduce bias. Extensive experiments show that HEIMAT effectively mitigates bias in different cultures while maintaining the model’s natural language understanding (NLU) performance.

1 Introduction

Refer to caption
Figure 1: Prediction difference between the two models. The original model shows a strong gender bias in occupational predictions, while the debiased model yields more balanced outcomes.

Language models (LMs) are deployed in many domains (Wang et al., 2024; Zhang et al., 2025a; Zhang et al., 2025b; Liu and Henao, 2025; Hu et al., 2025; Shi et al., 2025; Dou et al., 2026b; Shi et al., 2026; Dou et al., 2026a), therefore, ensuring fairness of language models across different social contexts and cultures is crucial for mitigating model bias and enhancing their global applicability, which is also consistent with the principles of inclusiveness and universality advocated by the United Nations Behind (2017); Cushman (2012). In recent years, a wide range of language models, including pre-trained models (PLMs) and large language models (LLMs) have achieved remarkable success across numerous natural language processing (NLP) tasks. However, as the data used to train LMs are primarily collected from online communities, they inevitably reflect societal biases present in the real world. These biases are often absorbed during the pre-training process and subsequently manifested in downstream NLP tasks (Lan et al., 2025b; Wei et al., 2026). As illustrated in Figure 1, the original model exhibits a clear gender imbalance in occupational predictions under the same context, tending to associate “young woman” with “nurse” rather than “doctor”, which reveals deeply rooted gender stereotypes embedded in the model. Such biases may reinforce societal stereotypes in real-world applications and lead to potentially adverse social consequences. Therefore, exploring effective approaches to mitigate biases is essential for building trustworthy and cross-cultural language models.

There are many remarkable attempts to mitigate bias, such as fine-tuning the model with counterfactual datasets directly (Zmigrod et al., 2020; Wang et al., 2022), which is partly useful but costly in time and resource. Some works focusing on representation projection (Ravfogel et al., 2020), and modifying word embeddings(Liang et al., 2020) to mitigate the bias, but it has limited effectiveness when dealing with more complex bias in models. Recent methods such as CCPA (Li et al., 2023), BiasDPO (Allam, 2024), C2PO (Feng et al., 2025), and BiasVector (Shirafuji et al., 2025a) effectively mitigate social bias but depend on fixed preference datasets. Despite being generated by advanced models, these datasets are costly to annotate and difficult to scale across diverse languages, cultures, and bias categories. Crucially, because different models exhibit distinct bias tendencies and severities (Lan et al., 2025a), such a one-size-fits-all approach inevitably suffers from poor generalizability. These drawbacks inspire us to explore a new method to mitigate the social bias more flexibly. Therefore, this paper focuses on the following question:

How can we debias language models across languages and cultures without any fixed preference datasets or external corpora?

To address this question, we design HEIMAT, a heuristic-style automatic debiasing framework for language models that operates independently of external corpora and does not require fixed preference datasets or manually curated lexicons. HEIMAT consists of two stages: bias disclosure and debiasing regularization. In the first stage, heuristic prompts are used to elicit biased responses from the model, thereby explicitly revealing its intrinsic biases. In the second stage, the model is regularized using the Jensen-Shannon Divergence (JSD) loss to reduce prediction discrepancies across different demographic conditions. Extensive experiments on five PLMs, three bias-evaluation benchmarks, and two NLU datasets demonstrate that HEIMAT effectively mitigates model bias while maintaining comparable NLU performance. Moreover, HEIMAT relies on only six heuristic prompt templates and seven association prompts to construct substitution sets. By translating the templates and adjusting bias-related terms, the framework can be naturally extended to other languages and bias categories. We highlight the following contributions:

  • We introduce HEIMAT, a heuristic-style automatic debiasing framework for LMs that operates. It consists of two stages: bias disclosure and debiasing fine-tuning.

  • HEIMAT reveals and mitigates model biases through automatically constructed heuristic and association prompts, making it readily adaptable to different bias languages, cultures bias categories via simple prompt modifications.

  • Extensive experiments on multiple language models and benchmark datasets demonstrate that HEIMAT achieves competitive debiasing performance while preserving the models’ NLU capabilities.

Refer to caption
Figure 2: HEIMAT consists of two stages. Step 1 (Bias Disclosure) uses a heuristic and association prompts with a demographically neutral target word (Neutral WtW_{t}) to elicit demographic prediction distributions and construct a substitution list. Step 2 (Debiasing Fine-tuning) generates context prompts that differ only in demographic attributes and aligns their predictive distributions via minimization of the Jensen–Shannon divergence.

2 Background and Related Works

2.1 Bias Issues in NLP

Language models are typically trained on large-scale corpora collected from online sources such as news articles, forums, and Wikipedia, which inevitably reflect societal biases present in human-generated text (Schick et al., 2021a; Blodgett et al., 2020) and Chain-of-Thought process (Feng et al., 2026). Previous research demonstrates that such biases not only persist but also become more subtle and deeply embedded in PLMs and LLMs (Sheng et al., 2019; May et al., 2019; Kaneko and Bollegala, 2021). More recent work has systematically analyzed the presence of social biases in modern PLMs and LLMs across languages and tasks, highlighting their impact on downstream applications and raising increasing concerns about fairness and reliability (Meade et al., 2022; Gallegos et al., 2024; Dong et al., 2024).

2.2 Debiasing Methods for Language Models

Existing debiasing methods for language models primarily rely on data augmentation, prompt-based control, as well as fine-tuning and alignment strategies. Data-driven approaches, such as counterfactual data augmentation, aim to mitigate bias by constructing fairer training data; however, they often require substantial manual effort and incur high computational and environmental costs Lu et al. (2020); Strubell et al. (2019). Other lines of research focus on fine-tuning or alignment-based debiasing, including automatically generated biased prompts Guo et al. (2022b), fast bias removal via machine unlearning Chen et al. (2023), and parameter-efficient or LLM-assisted debiasing strategies Han et al. (2024). More recently, researchers have explored revealing and mitigating bias through model-internal probing Dong et al. (2024), as well as task-arithmetic-based debiasing methods that reduce bias without full retraining Shirafuji et al. (2025b). Despite their effectiveness in specific settings, most existing approaches rely on external corpora or fixed preference datasets, which limit their scalability and generalization across bias categories, languages, and cultures Meade et al. (2022). However, most methods focus on English biases and assume cultural universality, leaving cross-cultural bias largely underexplored (Shan et al., 2024; Madhusudan et al., 2026). Culturally focusing debiasing methods for diverse societies (e.g., India, Europe) remain largely undeveloped.

3 Method

3.1 Overall Framework

Biases in language models frequently manifest as divergent prediction distributions for different demographic groups within the same contextual setting (Lan et al., 2025b). To address this issue, we introduce HEIMAT, a prompt-driven framework that combines bias disclosure with debiasing fine-tuning. As shown in Figure 2, HEIMAT first employs heuristic and association prompts to uncover demographic attributes implicitly associated with a given context, thereby constructing sets of minimally contrasted context prompts. The model is subsequently fine-tuned by minimizing the Jensen–Shannon divergence across prediction distributions conditioned on these prompts, enforcing demographic-invariant behavior while preserving contextual semantics.

Refer to caption
Figure 3: Our process for creating context prompts. [BLANK] represents that this position is blank and will be filled.

3.2 Bias Disclosure

HEIMAT discloses social biases through specially designed heuristic prompts. For clarity, we first define three types of words used in our prompts: (i) Target Word (WtW_{t}), which denotes the social entity being described; (ii) Feature Word (WfW_{f}), which represents contextual attributes or behaviors associated with the target word and typically forms an adjective or verb phrase; and (iii) Bias Type Word (WbW_{b}), which specifies the bias category of interest, such as gender or race.

We use Deepseek-Chat (Liu et al., 2024) to generate heuristic prompts according to our templates. Examples can be found in Appendix A.

Consider the heuristic prompt: “This young woman often appears in a hospital, and her job is as a …”. Here, “woman” serves as the target word WtW_{t}, “appears in a hospital” constitutes the feature words WfW_{f}, and “job” corresponds to the bias type word WbW_{b}. For Given this prompt, a language model may generate continuations such as “doctor”, “nurse”, or “teacher”. These responses reveal the model’s implicit associations between demographic attributes and occupations. For instance, a preference for “nurse” may reflect a gender stereotype associating women with caregiving roles. By composing prompts using these three components, HEIMAT systematically elicits and analyzes a model’s bias behavior across diverse contexts, forming the basis for subsequent debiasing.

Inspired by prior work (Dong et al., 2024), we construct heuristic prompts using a set of predefined templates to simulate realistic social scenarios while reducing manual prompt engineering cost. Each template generates a complete heuristic prompt by taking as input a neutral target word WtW_{t}, a feature word phrase WfW_{f}, and a bias type word WbW_{b}. For masked language models (e.g., BERT), an additional [MASK] token is included at the prediction position, whereas for autoregressive language models (e.g., LLaMA-2), the model is prompted to continue the given prefix. We employ six templates in total to generate heuristic prompts in our experiments, ensuring that bias disclosure is grounded in realistic and socially meaningful settings while allowing the model to reveal its own inherent predispositions.

Given a heuristic prompt, we input the sequence into the model and extract the top-kk most probable tokens conditioned on the prompt context. For masked language models, this corresponds to the probability distribution over the masked position, while for decoder-only models, it corresponds to the next-token distribution given the prompt prefix. We denote the resulting token list as WL=(w1,w2,,wk)WL=(w_{1},w_{2},\ldots,w_{k}), where each wiw_{i} represents a demographic attribute implicitly associated with the preceding context. The set WLWL thus captures the model’s most salient demographic associations under the heuristic prompt.

To further expand demographic coverage, we apply association prompts to each wiWLw_{i}\in WL and prompt the model to generate nn related demographic terms, forming a substitution list SLSL:

SL=[(w11,w21,,wn1),(w12,w22,,wn2),\displaystyle SL=[(w_{1}^{1},w_{2}^{1},\ldots,w_{n}^{1}),(w_{1}^{2},w_{2}^{2},\ldots,w_{n}^{2}), (1)
OPEN,(w1k,w2k,,wnk)].\displaystyle\ldots,(w_{1}^{k},w_{2}^{k},\ldots,w_{n}^{k})].

Using the substitution list, we construct a Context Prompt Set (CPS) by minimally modifying the original heuristic prompt (as illustrated in Figure 3). Specifically, we substitute the demographic attribute in the prompt with each term sampled from SLSL, while masking or abstracting the variable component of the feature words WfW_{f}. For autoregressive models, we reorder the sentence to place the next-token position at the end, allowing models to condition on the preceding demographic attributes. This procedure ensures that all prompts within a CPS differ only in demographic attributes, while preserving identical contextual semantics. Formally, we define the context prompt set as

CPS=(CP1,CP2,,CPn),CPS=(CP_{1},CP_{2},\ldots,CP_{n}), (2)

where each CPiCP_{i} denotes a context prompt. Each CPS is treated as a training unit in the subsequent debiasing fine-tuning stage.

3.3 Debiasing Fine-tuning

Given a language model M and a context prompt CP, the model induces a probability distribution over possible next tokens. We define the prediction distribution pp as:

p(w)=M(wCP),p(w)=M(w\mid CP), (3)

where ww denotes a candidate token in the model’s output vocabulary.

Our goal is to ensure that the model produces consistent prediction distributions across context prompts that differ only in demographic attributes, while preserving the same semantic context. Specifically, given a context prompt set CPS=(CP1,CP2,,CPn)CPS=(CP_{1},CP_{2},\ldots,CP_{n}), where each CPiCP_{i} differs only in the substituted demographic term, we aim to minimize the discrepancy among their induced prediction distributions.

To achieve this, we adopt Jensen–Shannon Divergence (Lin, 1991) as a symmetric and bounded measure of distributional difference, and fine-tune the model by minimizing the Jensen–Shannon Divergence loss (JSD\mathcal{L}_{JSD}). Given the set of prediction distributions PS=(p1,p2,,pn)PS=(p_{1},p_{2},\ldots,p_{n}), the loss is defined as:

JSD=1ni=1nDKL(pi1nj=1npj),\mathcal{L}_{JSD}=\frac{1}{n}\sum_{i=1}^{n}D_{KL}\left(p_{i}\;\middle\|\;\frac{1}{n}\sum_{j=1}^{n}p_{j}\right), (4)

where DKLD_{KL} denotes the Kullback–Leibler Divergence.

By minimizing this loss during fine-tuning, the model is encouraged to align its output distributions across different contexts, thereby reducing biased associations while maintaining NLU capability.

3.4 Language, Culture and Category Adaptation

Although we provide only the heuristic and association prompts for English and French, the proposed method can be readily applied to other languages and cultural contexts. To adapt to a new culture, one simply prompts an LLM (e.g., DeepSeek) to regenerate the heuristic prompts following the same template structure, using culturally relevant feature words (Examples can be found in Appendix A.). This lightweight process replaces direct translation, ensuring that the biases being measured are those actually present in the target culture.

Furthermore, our templates are highly adaptable to various types of bias. By simply substituting the bias type word, such as changing ’gender’ to ’race’, we can effectively address and mitigate the model’s biases related to race. This adaptability extends to a wide range of other bias categories, showcasing the versatility and broad applicability of our approach.

4 Results and Analysis

We describe the overall experimental results in this section and defer detailed configurations, including model choices, baseline implementations, datasets, and hyperparameter, to Appendix B.

4.1 Quantitative Results

CrowS-Pairs StereoSet
Models Gender Race Gender Race
BERT 58.01 58.12 60.28 57.03
   +CDA 56.11 56.70 59.61 56.73
   +Dropout 55.34 59.03 60.66 57.07
   +INLP 51.15 67.96 67.25 57.29
   +Sent-Debias 52.29 62.72 59.37 55.18
   +Self-Debias 52.29 56.70 59.34 54.30
   +Auto-Debias 54.92 65.05 57.33 54.02
   +FineDeb 54.58 65.24 53.27 50.82
   +CCPA 51.57 - 56.61 -
   +ADELE 54.20 - 46.27 -
   +PromptDeb 50.31 - 59.80 -
   +ChatGPT-Based 51.68 57.20 71.80 59.60
   +Bias Vector 51.74 52.90 - -
   +Ours 50.00 47.48 52.67 53.74
Table 1: Comparison of debiasing performance of PLMs across different benchmarks. For CrowS-Pairs and StereoSet, scores closer to 50 are better. CCPA, PromptDeb and ADELE are specifically designed for gender debiasing; race-related scores are omitted. “-” indicates the method is not applicable to this dataset. (Best, Next Best).
CrowS-Pairs StereoSet
Models Gender Race Gender Race
Llama-2 60.30 67.60 70.60 63.00
   +CDA 58.33 71.40 64.03 67.24
   +INLP 60.37 65.62 63.97 62.50
   +Self-Debias 54.87 62.81 60.04 63.49
   +Synonym-KG 59.20 65.50 67.50 62.70
   +KGDebias 56.90 57.00 61.20 62.50
   +Ours 54.07 55.64 55.45 53.17
GPT-2 56.87 59.69 65.58 61.63
   +CDA 56.87 60.66 64.02 57.31
   +INLP 53.44 59.69 63.17 60.00
   +Self-Debias 56.11 53.29 60.28 57.29
   +Synonym-KG 59.80 55.80 70.60 59.80
   +KGDebias 59.50 55.40 69.80 58.30
   +PromptDeb 47.17 - 55.10 -
   +Ours 52.52 53.31 54.08 54.60
Table 2: Comparison of debiasing performance of LLMs across different benchmarks. For CrowS-Pairs and StereoSet, scores closer to 50 are better. (Best, Next Best).
Lang. Models Race Gender Socioeco. Nationality Religion Age Sexual Apperance Disability
BERT 58.12 58.01 59.93 62.94 71.39 55.23 67.92 63.46 61.68
  +Ours 47.48 50.00 54.07 64.15 43.81 63.22 45.24 58.73 58.33
ALBERT 51.36 57.25 60.47 51.57 59.05 66.52 75.00 46.03 86.67
  +Ours 53.49 54.25 63.37 47.17 50.48 55.17 71.43 58.73 75.00
EN TinyBERT 68.22 50.76 68.02 54.72 31.43 58.62 47.62 47.62 56.67
  +Ours 45.74 48.09 41.28 46.54 33.33 45.98 65.48 49.21 48.33
CamemBERT 53.04 57.01 59.69 62.45 65.22 53.33 52.75 58.33 62.12
  +Ours 41.74 52.34 54.29 37.94 62.61 52.22 48.35 63.89 77.27
FR FrALBERT 56.74 47.66 58.16 60.47 72.17 38.89 81.32 40.28 42.42
  +Ours 50.65 47.98 57.65 60.87 65.21 45.56 79.12 58.33 48.48
Table 3: Debiasing performance of HEIMAT on English (EN) and French (FR) CrowS-Pairs across bias categories. Notably, the models are only debiased in Gender and Race. Closer to 50 is better (Best). (Nat.: Nationality).
Models Stereo Anti-Stereo Overall
BERT 61.09 56.88 60.48
   +Ours 52.40 48.62 51.86
ALBERT 56.20 60.09 56.76
   +Ours 55.35 63.76 55.56
TinyBERT 58.39 57.34 58.16
   +Ours 43.75 61.29 46.22
Table 4: The performance of HEIMAT on three PLMs for the two most common categories in CrowS-Pairs. Closer to 50 is better (Best).

To quantitatively evaluate the effectiveness of HEIMAT in mitigating social biases, we conduct experiments on the CrowS-Pairs, StereoSet, and SEAT benchmarks across gender and race dimensions using BERT. As shown in Table 1, HEIMAT consistently achieves the best or near-best performance, with scores closest to the ideal value of 50. In particular, on the CrowS-Pairs benchmark, HEIMAT reduces the Gender bias of BERT from 58.01 to 50.00 and achieves the lowest Race bias score (47.48), outperforming all competing baselines. HEIMAT also achieves consistently strong performance on StereoSet, and yields the lowest effect sizes on the SEAT benchmark for both Gender and Race (reported in Table 8 of Appendix C).

Similar trends are observed for LLaMA-2 and GPT-2, where HEIMAT achieves the lowest bias for both models, despite their reduced model capability. These results demonstrate that HEIMAT effectively mitigates demographic biases across different bias types, validating the robustness and generality of the proposed heuristic-style debiasing framework.

In contrast, several baselines show inconsistent behavior across datasets, improving bias under one metric while degrading under another. HEIMAT, however, achieves consistent improvements across all evaluated benchmarks, suggesting better robustness and generalization.

4.2 Stereotype & Anti-Stereotype Trade-off

We further investigate the behavior of HEIMAT on whole CrowS-Pairs by jointly analyzing the overall stereotype tendency and its transfer effects across bias categories. As shown in Table 4, HEIMAT consistently reduces Stereotype scores and moves the Overall scores of all three PLMs significantly closer to the ideal value of 50, indicating effective mitigation of dominant stereotypical bias across different model capacities. In particular, HEIMAT reduces the Overall score of BERT from 60.48 to 51.86 and that of TinyBERT from 58.16 to 46.22. We also observe increases in Anti-Stereotype scores for ALBERT and TinyBERT, suggesting a shift away from stereotypical preferences toward anti-stereotypical predictions, which is expected given the strong initial bias of these models and remains beneficial as reflected by the improved Overall scores.

Beyond the two debiased categories, Table 3 further shows that debiasing conducted only on gender and race leads to consistent score shifts across multiple unseen bias categories, including socio-economic status, nationality, religion, and physical appearance. Such cross-category improvements indicate that different forms of social bias are intrinsically correlated in models probably, and that mitigating prominent biases can partially generalize to related dimensions.

4.3 Multilingual and Multicultural Debiasing

To evaluate the effectiveness of HEIMAT in multilingual and multicultural settings, we conduct debiasing experiments on two French PLMs, CamemBERT and FrALBERT, using the French version of CrowS-Pairs. The results are reported in Tables 3. When applying HEIMAT to French PLMs, the prompt templates are adapted to account for the grammatical structure of French, which differs from English in that nouns are associated with grammatical gender and require agreement with corresponding articles and adjectives. As illustrated in Fig. 4, feminine and masculine forms exhibit systematic suffix variations, challenging fixed-template methods with static word lists (Dong et al., 2024; Zeng et al., 2024) in cross-lingual and cross-cultural adaptation. Our method retains template structure but replaces static lists with LLM-generated content, offering both flexibility and clarity.

Despite this additional linguistic complexity and cultural difference, HEIMAT consistently reduces bias on both models, bringing the Overall scores closer to the ideal value of 50. Specifically, HEIMAT reduces CamemBERT’s Overall score from 57.36 to 50.69 and improves FrALBERT’s Anti-Stereotype balance by aligning its predictions toward a more neutral distribution. Therefore, these results demonstrate that HEIMAT effectively mitigates bias in French PLMs and can be readily extended to other cultures and languages

Refer to caption
Figure 4: Two examples of linguistic phenomenon of noun gender (meaning: “a lonely singer”).

4.4 Qualitative Bias Analysis

Figure 5 presents a qualitative comparison of the prediction probability distributions produced by the original and debiased ALBERT under a controlled prompt setting. Specifically, we use the prompt “This [JOB] works in a famous hospital and [MASK] is professional” and substitute “nurse” and “doctor” for [JOB] to examine gender pronoun predictions. In the original ALBERT, a pronounced disparity is observed: when the job is “nurse”, the model assigns substantially higher probability to the feminine pronoun “she”, whereas for “doctor”, the masculine pronoun “he” dominates, revealing strong occupation–gender associations.

In contrast, the debiased model exhibits markedly more balanced prediction distributions. The probability gap between “he” and “she” is significantly reduced for both occupations, indicating that the model’s predictions become less sensitive to gender stereotypes associated with specific jobs. This qualitative analysis demonstrates that HEIMAT effectively attenuates demographic-conditioned prediction shifts at the distributional level, aligning with the quantitative improvements observed on bias benchmarks.

Similar phenomenon can be observed in Llama-2, as shown in Table 5 Before debiasing, the model tends to complete prompts involving racial identifiers with narratives of violence and victimization. After debiasing, while negative outcomes may still appear, the generated continuations no longer systematically attribute harm to external agents conditioned on race, suggesting a reduction in identity-conditioned violent associations.

Model Generations Before and After Debiasing Prompt: On Sunday, this Black man like… LLaMA-2 (Before Debiasing): On Sunday, this black man like myself was killed by a white police officer in the city of Philadelphia. He was killed for being … LLaMA-2 (After Debiasing): On Sunday, this black man like has a serious accident and he is in a coma. He is on life support. He is brain …
Table 5: Example generations of LLaMA-2 before and after debiasing.
Refer to caption
Figure 5: Visualization of the predicted probabilities of gender pronoun for “nurse” and “doctor” on ALBERT before and after debiasing with HEIMAT.

4.5 Analysis of NLU Tasks Performance

Since HEIMAT mitigates bias via fine-tuning, it is important to examine its potential impact on the model’s NLU performance (Meade et al., 2022). We therefore evaluate HEIMAT on standard NLU benchmarks to assess whether bias reduction is achieved while preserving NLU capability.

Models CoLA SST-2 MRPC QNLI QQP RTE MNLI STS-B WNLI Avg.
BERT 56.78 93.35 89.54 91.51 88.06 64.62 84.76 88.24 56.34 79.24
   +CDA 55.97 92.32 87.22 90.84 87.85 63.29 84.84 88.43 53.66 78.27
   +Dropout 50.87 92.09 88.22 91.49 88.02 62.29 84.78 87.87 51.66 77.48
   +INLP 56.50 92.66 89.23 91.38 87.94 65.34 84.78 88.73 54.93 79.05
   +Sent-Debias 55.72 93.12 88.81 91.54 87.88 63.90 84.94 88.23 56.34 78.94
   +Auto-Debias 57.01 92.89 88.54 91.65 87.92 64.62 84.91 88.43 40.85 77.42
   +CCPA 55.91 93.09 88.65 91.42 87.98 64.93 84.73 88.44 55.66 78.98
   +Ours 57.40 91.80 88.84 91.09 88.54 65.51 84.72 87.66 56.34 79.09
ALBERT 55.61 92.32 90.85 91.21 88.99 72.56 85.38 89.39 39.44 78.42
   +Dropout 46.83 91.86 89.83 90.99 86.90 55.71 85.13 88.68 52.34 76.47
   +Sent-Debias 56.37 92.31 91.65 91.92 87.56 67.87 84.96 90.46 33.80 77.43
   +Auto-Debias 56.69 93.35 91.61 91.94 87.41 73.29 79.25 90.48 54.93 79.88
   +Ours 55.42 92.15 90.87 91.02 89.74 69.26 85.01 90.35 53.28 79.67
TinyBERT 51.34 92.12 86.69 90.15 70.64 70.16 84.74 83.42 56.28 76.28
   +Ours 48.48 91.62 87.43 90.08 70.03 73.31 84.33 84.70 56.30 76.25
LLaMA-2 18.64 75.73 66.37 58.57 31.39 59.32 88.76 27.51 44.50 52.31
   +Ours 18.34 75.40 67.43 59.01 31.88 58.40 86.24 27.10 45.16 52.00
GPT-2 76.84 91.20 80.41 88.37 89.66 65.30 82.18 37.10 47.79 73.21
   +Ours 75.41 90.99 81.00 87.95 89.02 64.73 82.20 37.08 45.72 72.68
Table 6: Performance comparison of debiasing on the GLUE Benchmark. We chose F1 Score as the metric on the MRPC and QQP and Spearman Correlation as the metric on STS-B. LLaMA-2 is evaluated in a zero-shot setting, while the other models are fine-tuned on downstream task(Best, Next Best).

NLU Tasks Performance of English PLMs. To evaluate whether HEIMAT preserves the NLU capability of PLMs, we conduct experiments on the GLUE benchmark, with the results reported in Table 6. The results indicate that HEIMAT maintains NLU performance at a level comparable to the original models while remaining competitive with existing debiasing methods. Specifically, for BERT, HEIMAT achieves an average GLUE score of 79.09, which is close to the original score of 79.24, with performance on tasks such as QQP and RTE even slightly improved. For ALBERT, HEIMAT attains an average score of 79.67, comparable to the base model (78.42) and competitive with the strongest baseline. Similarly, for TinyBERT, LLama-2 and GPT-2, the average GLUE score remains essentially unchanged, indicating minimal performance degradation despite the reduced model capacity. Therefore, these results demonstrate that HEIMAT effectively mitigates bias while preserving the core NLU capabilities of PLMs.

Models PAWS-X XNLI  CLS  Avg.
CamemBERT 91.24 81.16 94.75 89.05
   +Ours 91.50 81.47 94.59 89.18
FrALBERT 83.44 71.06 80.42 78.30
   +Ours 83.87 70.98 79.74 78.19
Table 7: The performance of our method on FLUE Benchmark.

NLU Tasks Performance of French PLMs. To evaluate whether HEIMAT preserves the NLU capability of French PLMs, we conduct experiments on the FLUE benchmark, with results reported in Table 7. The results show that HEIMAT maintains comparable NLU performance on French PLMs after debiasing. Specifically, for CamemBERT, HEIMAT achieves an average score of 89.18, which is slightly higher than the baseline score of 89.05, with consistent performance on PAWS-X and XNLI. For FrALBERT, although the scores on XNLI and CLS decrease marginally (from 71.06 to 70.98 and from 80.42 to 79.74, respectively), the overall average remains largely unchanged (78.30 vs. 78.19). Therefore, these results demonstrate that HEIMAT effectively mitigates bias in French PLMs without causing significant degradation in their NLU performance.

5 Conclusion

In this work, we propose HEIMAT, a heuristic-style automatic debiasing framework for LMs that leverages a small set of heuristic templates and uses DeepSeek to generate heuristic prompts to disclose social biases, then mitigates model biases through JSD-based fine-tuning. By avoiding reliance on external corpora, HEIMAT can be readily extended to different bias categories, languages, and cultures, which we validate through experiments on both PLMs and LLMs in different languages. Extensive results across multiple benchmarks demonstrate that HEIMAT achieves effective bias mitigation with minimal impact on general performance, while also revealing intrinsic connections among different bias types that enable broader generalization.

Limitations

We acknowledge that not all stereotypical associations are inherently harmful, as some reflect widely shared cultural knowledge. For example, Christians are more likely to appear in churches, while people from other religions have a lower probability. Distinguishing between harmful social biases and benign common-sense patterns remains an open challenge, and we leave this important issue for future work.

Ethics Statement

The bias evaluation datasets we used are available publicly, and we are not responsible for any content that contains stereotypes and biases in these datasets. Our work focuses on reducing model bias and using them to evaluate the performance of our method.

References

  • Allam (2024) A. Allam BiasDPO: mitigating bias in language models through direct preference optimization. In Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 4: Student Research Workshop), pp. 42–50. Cited by: §1.
  • Behind (2017) L. N. O. Behind Equality and non-discrimination at the heart of sustainable development. The United Nations System Shared Framework for Action, New York. Cited by: §1.
  • Blodgett et al. (2020) S. L. Blodgett, S. Barocas, H. D. I. au2, and H. Wallach Language (technology) is power: a critical survey of "bias" in nlp. External Links: 2005.14050 Cited by: §2.1.
  • Cattan et al. (2022) O. Cattan, C. Servan, and S. Rosset On the usability of transformers-based models for a french question-answering task. External Links: 2207.09150 Cited by: §B.2.
  • Chen et al. (2023) R. Chen, J. Yang, H. Xiong, J. Bai, T. Hu, J. Hao, Y. Feng, J. T. Zhou, J. Wu, and Z. Liu Fast model debias with machine unlearning. External Links: 2310.12560 Cited by: §2.2.
  • Cushman (2012) T. Cushman Handbook of human rights. Routledge. Cited by: §1.
  • Devlin et al. (2019) J. Devlin, M. Chang, K. Lee, and K. Toutanova BERT: pre-training of deep bidirectional transformers for language understanding. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers), J. Burstein, C. Doran, and T. Solorio (Eds.), Minneapolis, Minnesota, pp. 4171–4186. External Links: Link, Document Cited by: §B.2.
  • Dong et al. (2024) X. Dong, Y. Wang, P. S. Yu, and J. Caverlee Disclosure and mitigation of gender bias in llms. External Links: 2402.11190 Cited by: §2.1, §2.2, §3.2, §4.3.
  • Dou et al. (2026a) J. Dou, C. Shi, J. Wang, F. Shen, Z. Wang, and T. Chua Beyond surface artifacts: capturing shared latent forgery knowledge across modalities. arXiv preprint arXiv:2604.07763. Cited by: §1.
  • Dou et al. (2026b) J. Dou, C. Shi, Y. Wang, S. Guo, A. Yi, W. Wu, L. Zhang, F. Shen, and T. Chua DNA: uncovering universal latent forgery knowledge. arXiv preprint arXiv:2601.22515. Cited by: §1.
  • Feng et al. (2025) X. Feng, B. An, T. Gu, L. Chang, F. Hao, P. Yu, and S. Zhao C2PO: diagnosing and disentangling bias shortcuts in llms. arXiv preprint arXiv:2512.23430. Cited by: §1.
  • Feng et al. (2026) X. Feng, S. Zhao, L. Xiao, T. Gu, and B. An Self-debias: self-correcting for debiasing large language models. arXiv preprint arXiv:2604.08243. Cited by: §2.1.
  • Gallegos et al. (2024) I. O. Gallegos, R. A. Rossi, J. Barrow, M. M. Tanjim, S. Kim, F. Dernoncourt, T. Yu, R. Zhang, and N. K. Ahmed Bias and fairness in large language models: a survey. External Links: 2309.00770, Link Cited by: §2.1.
  • Guo et al. (2022a) Y. Guo, Y. Yang, and A. Abbasi Auto-debias: debiasing masked language models with automated biased prompts. In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), S. Muresan, P. Nakov, and A. Villavicencio (Eds.), Dublin, Ireland, pp. 1012–1023. External Links: Link, Document Cited by: §B.3.
  • Guo et al. (2022b) Y. Guo, Y. Yang, and A. Abbasi Auto-debias: debiasing masked language models with automated biased prompts. In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pp. 1012–1023. Cited by: §2.2.
  • Han et al. (2024) P. Han, R. Kocielnik, A. Saravanan, R. Jiang, O. Sharir, and A. Anandkumar ChatGPT based data augmentation for improved parameter-efficient debiasing of llms. In Proceedings of the Fourth Workshop on Language Technology for Equality, Diversity, Inclusion, pp. 73–105. Cited by: §B.3, §2.2.
  • Hu et al. (2025) S. Hu, J. Hu, and H. Zhang Synergizing LLMs with global label propagation for multimodal fake news detection. In Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), W. Che, J. Nabende, E. Shutova, and M. T. Pilehvar (Eds.), Vienna, Austria, pp. 1426–1440. External Links: Link, Document, ISBN 979-8-89176-251-0 Cited by: §1.
  • Jiao et al. (2020) X. Jiao, Y. Yin, L. Shang, X. Jiang, X. Chen, L. Li, F. Wang, and Q. Liu TinyBERT: distilling bert for natural language understanding. External Links: 1909.10351 Cited by: §B.2.
  • Kaneko and Bollegala (2021) M. Kaneko and D. Bollegala Debiasing pre-trained contextualised embeddings. External Links: 2101.09523 Cited by: §2.1.
  • Lan et al. (2025a) T. Lan, J. Li, Y. Wang, X. Liu, X. Su, and G. Gao F2{}^{2}bench: an open-ended fairness evaluation benchmark for llms with factuality considerations. In Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing, pp. 2031–2046. Cited by: §1.
  • Lan et al. (2025b) T. Lan, X. Su, X. Liu, R. Wang, K. Chang, J. Li, and G. Gao McBE: a multi-task Chinese bias evaluation benchmark for large language models. In Findings of the Association for Computational Linguistics: ACL 2025, W. Che, J. Nabende, E. Shutova, and M. T. Pilehvar (Eds.), Vienna, Austria, pp. 6033–6056. External Links: Link, Document, ISBN 979-8-89176-256-5 Cited by: §1, §3.1.
  • Lan et al. (2020) Z. Lan, M. Chen, S. Goodman, K. Gimpel, P. Sharma, and R. Soricut ALBERT: a lite bert for self-supervised learning of language representations. External Links: 1909.11942 Cited by: §B.2.
  • Li et al. (2023) Y. Li, M. Du, X. Wang, and Y. Wang Prompt tuning pushes farther, contrastive learning pulls closer: a two-stage approach to mitigate social biases. In Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pp. 14254–14267. Cited by: §B.3, §1.
  • Liang et al. (2020) P. P. Liang, I. M. Li, E. Zheng, Y. C. Lim, R. Salakhutdinov, and L. Morency Towards debiasing sentence representations. External Links: 2007.08100 Cited by: §B.3, §1.
  • Lin (1991) J. Lin Divergence measures based on the shannon entropy. IEEE Transactions on Information Theory 37 (1), pp. 145–151. External Links: Document Cited by: §3.3.
  • Liu et al. (2024) A. Liu, B. Feng, B. Xue, B. Wang, B. Wu, C. Lu, C. Zhao, C. Deng, C. Zhang, C. Ruan, et al. Deepseek-v3 technical report. arXiv preprint arXiv:2412.19437. Cited by: §3.2.
  • Liu and Henao (2025) H. Liu and R. Henao Learning to substitute words with model-based score ranking. In Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers), pp. 11551–11565. Cited by: §1.
  • Loshchilov and Hutter (2019) I. Loshchilov and F. Hutter Decoupled weight decay regularization. External Links: 1711.05101 Cited by: §B.4.
  • Lu et al. (2019) K. Lu, P. Mardziel, F. Wu, P. Amancharla, and A. Datta Gender bias in neural natural language processing. External Links: 1807.11714 Cited by: §B.3.
  • Lu et al. (2020) K. Lu, P. Mardziel, F. Wu, P. Amancharla, and A. Datta Gender bias in neural natural language processing. In Logic, language, and security: essays dedicated to Andre Scedrov on the occasion of his 65th birthday, pp. 189–202. Cited by: §2.2.
  • Ma et al. (2024) C. Ma, T. Zhao, and M. Okumura Debiasing large language models with structured knowledge. In Findings of the Association for Computational Linguistics: ACL 2024, L. Ku, A. Martins, and V. Srikumar (Eds.), Bangkok, Thailand, pp. 10274–10287. External Links: Link, Document Cited by: §B.3.
  • Madhusudan et al. (2026) S. Madhusudan, T. S. More, S. Buongiorno, R. Dividino, J. Kabbara, and A. Emami Common to whom? regional cultural commonsense and llm bias in india. arXiv preprint arXiv:2601.15550. Cited by: §2.2.
  • Martin et al. (2020) L. Martin, B. Muller, P. J. Ortiz Suárez, Y. Dupont, L. Romary, É. de la Clergerie, D. Seddah, and B. Sagot CamemBERT: a tasty French language model. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, D. Jurafsky, J. Chai, N. Schluter, and J. Tetreault (Eds.), Online, pp. 7203–7219. External Links: Link, Document Cited by: §B.2.
  • May et al. (2019) C. May, A. Wang, S. Bordia, S. R. Bowman, and R. Rudinger On measuring social biases in sentence encoders. External Links: 1903.10561 Cited by: §2.1.
  • Meade et al. (2022) N. Meade, E. Poole-Dayan, and S. Reddy An empirical survey of the effectiveness of debiasing techniques for pre-trained language models. External Links: 2110.08527 Cited by: §2.1, §2.2, §4.5.
  • Radford et al. (2019) A. Radford, J. Wu, R. Child, D. Luan, D. Amodei, I. Sutskever, et al. Language models are unsupervised multitask learners. OpenAI blog 1 (8), pp. 9. Cited by: §B.2.
  • Ravfogel et al. (2020) S. Ravfogel, Y. Elazar, H. Gonen, M. Twiton, and Y. Goldberg Null it out: guarding protected attributes by iterative nullspace projection. External Links: 2004.07667 Cited by: §B.3, §1.
  • Saravanan et al. (2023) A. Saravanan, D. Mullick, H. Rahman, and N. Hegde FineDeb: a debiasing framework for language models. External Links: 2302.02453 Cited by: §B.3.
  • Schick et al. (2021a) T. Schick, S. Udupa, and H. Schütze Self-diagnosis and self-debiasing: A proposal for reducing corpus-based bias in NLP. CoRR abs/2103.00453. External Links: Link, 2103.00453 Cited by: §2.1.
  • Schick et al. (2021b) T. Schick, S. Udupa, and H. Schütze Self-diagnosis and self-debiasing: a proposal for reducing corpus-based bias in nlp. External Links: 2103.00453 Cited by: §B.3.
  • Shan et al. (2024) X. Shan, Y. Xu, Y. Wang, Y. Lin, and Y. Bao Cross-cultural implications of large language models: an extended comparative analysis. In International Conference on Human-Computer Interaction, pp. 106–118. Cited by: §2.2.
  • Sheng et al. (2019) E. Sheng, K. Chang, P. Natarajan, and N. Peng The woman worked as a babysitter: on biases in language generation. External Links: 1909.01326 Cited by: §2.1.
  • Shi et al. (2025) C. Shi, S. Li, S. Guo, S. Xie, W. Wu, J. Dou, C. Wu, C. Xiao, C. Wang, Z. Cheng, et al. Where culture fades: revealing the cultural gap in text-to-image generation. arXiv preprint arXiv:2511.17282. Cited by: §1.
  • Shi et al. (2026) C. Shi, S. Li, W. Lu, W. Wu, C. Wang, Z. Cheng, F. Shen, and T. Chua TraceRouter: robust safety for large foundation models via path-level intervention. arXiv preprint arXiv:2601.21900. Cited by: §1.
  • Shirafuji et al. (2025a) D. Shirafuji, M. Takenaka, and S. Taguchi Bias vector: mitigating biases in language models with task arithmetic approach. In Proceedings of the 31st International Conference on Computational Linguistics, O. Rambow, L. Wanner, M. Apidianaki, H. Al-Khalifa, B. D. Eugenio, and S. Schockaert (Eds.), Abu Dhabi, UAE, pp. 2799–2813. External Links: Link Cited by: §B.3, §1.
  • Shirafuji et al. (2025b) D. Shirafuji, M. Takenaka, and S. Taguchi Bias vector: mitigating biases in language models with task arithmetic approach. In Proceedings of the 31st International Conference on Computational Linguistics, pp. 2799–2813. Cited by: §2.2.
  • Strubell et al. (2019) E. Strubell, A. Ganesh, and A. McCallum Energy and policy considerations for deep learning in nlp. In Proceedings of the 57th annual meeting of the association for computational linguistics, pp. 3645–3650. Cited by: §2.2.
  • Touvron et al. (2023) H. Touvron, L. Martin, K. Stone, P. Albert, A. Almahairi, Y. Babaei, N. Bashlykov, S. Batra, P. Bhargava, S. Bhosale, et al. Llama 2: open foundation and fine-tuned chat models. arXiv preprint arXiv:2307.09288. Cited by: §B.1, §B.2.
  • Wang et al. (2024) S. Wang, D. Wong, J. Yao, and L. Chao What is the best way for chatgpt to translate poetry?. In Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pp. 14025–14043. Cited by: §1.
  • Wang et al. (2022) Y. Wang, C. Xu, Q. Sun, H. Hu, C. Tao, X. Geng, and D. Jiang PromDA: prompt-based data augmentation for low-resource NLU tasks. In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), S. Muresan, P. Nakov, and A. Villavicencio (Eds.), Dublin, Ireland, pp. 4242–4255. External Links: Link, Document Cited by: §1.
  • Webster et al. (2021) K. Webster, X. Wang, I. Tenney, A. Beutel, E. Pitler, E. Pavlick, J. Chen, E. Chi, and S. Petrov Measuring and reducing gendered correlations in pre-trained models. External Links: 2010.06032 Cited by: §B.3.
  • Wei et al. (2026) X. Wei, X. Zhou, Y. Sakai, and T. Watanabe “Yuki gets sushi, david gets steak?”: uncovering gender and racial biases in llm-based meal recommendations. In Proceedings of the 19th Conference of the European Chapter of the Association for Computational Linguistics (Volume 1: Long Papers), pp. 7776–7796. Cited by: §1.
  • Zakizadeh and Pilehvar (2025) M. Zakizadeh and M. T. Pilehvar Blind men and the elephant: diverse perspectives on gender stereotypes in benchmark datasets. In Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing, C. Christodoulopoulos, T. Chakraborty, C. Rose, and V. Peng (Eds.), Suzhou, China, pp. 22838–22851. External Links: Link, Document, ISBN 979-8-89176-332-6 Cited by: §B.3.
  • Zeng et al. (2024) C. C. Zeng, M. Chung, and E. Zhou Prompting for fairness: mitigating gender bias in large language models with self-debiasing prompting. In University of Michigan CSE 595 Natural Language Processing Fall 2024, Cited by: §B.3, §4.3.
  • Zhang et al. (2025a) J. Zhang, K. Cai, Y. Fan, N. Liu, and K. Wang MAT-agent: adaptive multi-agent training optimization. arXiv preprint arXiv:2510.17845. Cited by: §1.
  • Zhang et al. (2025b) J. Zhang, Z. Huang, Y. Fan, N. Liu, M. Li, Z. Yang, J. Yao, J. Wang, and K. Wang KABB: knowledge-aware bayesian bandits for dynamic expert coordination in multi-agent systems. arXiv preprint arXiv:2502.07350. Cited by: §1.
  • Zmigrod et al. (2020) R. Zmigrod, S. J. Mielke, H. Wallach, and R. Cotterell Counterfactual data augmentation for mitigating gender stereotypes in languages with rich morphology. External Links: 1906.04571 Cited by: §1.

Appendix A Prompt Examples

In Table 9, we present the heuristic prompt samples in English and French respectively.

The structures of the six English heuristic templates we use are reported in Table 10. Among them, Wt1W_{t1} and Wt2W_{t2} are two different target word. The [VERBS] represents cognitive verbs describing thinking, perception, understanding, or other psychological processes, such as ’think’, ’know’ and ’guess’. The words mentioned above, except WfW_{f} about professions, are generated by ChatGPT.

We have listed the association prompts for the model to generate new English target words in similar domains in Table 11. We did not find words by calculating cosine similarity, selecting the few words with the smallest Euclidean distance from the vocabulary, or using clustering because these methods tend to result in poor diversity of the words obtained, and they often include words in irrelevant domains with too significant semantic differences.

Appendix B Experimental Setup

B.1 Datasets

We evaluate the effectiveness of HEIMAT across both PLMs and large language models (LLMs). For English PLMs, we utilize several representative datasets, including CrowS-Pairs, StereoSet, SEAT, and the GLUE benchmark. To assess the method’s cross-lingual generalizability, we further extend our evaluation to French models using French CrowS-Pairs and the FLUE benchmark. For LLMs, following Touvron et al. (2023), we evaluate English debiasing performance on CrowS-Pairs (via perplexity) and StereoSet. In all experiments, GLUE and FLUE serve as the primary metrics for assessing Natural Language Understanding (NLU) performance post-debiasing.

B.2 Language Models

Pre-trained Language Models Following previous works, we evaluate our method with two famous English PLMs (bert-base-uncased (Devlin et al., 2019); albert-base-v2 (Lan et al., 2020);TinyBERT-General-6L-768D (Jiao et al., 2020); and two famous French PLMs (camembert-base (Martin et al., 2020); fralbert-base (Cattan et al., 2022)).

Large Language Models For LLMs, we apply our debiasing method to Llama2-7b-hf (Touvron et al., 2023) and GPT-2 medium (Radford et al., 2019).

B.3 Baselines

We choose Counterfactual Data Augmentation (CDA) (Lu et al., 2019), Sent-Debias (Liang et al., 2020), Iterative Nullspace Projection (INLP) (Ravfogel et al., 2020), Dropout (Webster et al., 2021), Self-Debias (Schick et al., 2021b), Auto-Debias (Guo et al., 2022a), CCPA (Li et al., 2023), FineDeb (Saravanan et al., 2023), ChatGPT-Based (Han et al., 2024), ADELE (Zakizadeh and Pilehvar, 2025), Synonym-KG, KGDebias (Ma et al., 2024), PromptDeb (Zeng et al., 2024) and Bias Vector (Shirafuji et al., 2025a) as our baselines.

B.4 Hyperparameters

All of the experiments were conducted on an A100 with 80GB of memory, utilizing the Adam optimizer (Loshchilov and Hutter, 2019) with an initial learning rate set to 5e-6, a batch size of 8. We generate 20000 heuristic prompts in the first step and collected the top 5 words with the highest probabilities. The other baselines are ran with their default settings.

Appendix C Complete Evaluation Results for SEAT

Table 8 records the effect sizes for tests in SEAT. The result shows that HEIMAT effectively reduces social bias in PLMs.

Models Gender Race
BERT 0.621 0.716
   +CDA 0.722 0.685
   +Dropout 0.765 0.580
   +INLP 0.204 0.675
   +Sent-Debias 0.430 0.610
   +Self-Debias 0.621 0.716
   +Auto-Debias 0.671 0.663
   +FineDeb 0.360 0.620
   +CCPA 0.249 -
   +Bias Vector 0.653 0.646
   +Ours 0.183 0.237
Table 8: Debiasing performance on the SEAT benchmark. Absolute effect sizes closer to 0 indicate lower bias. Methods marked with “-” are not applicable to the corresponding setting. (Best, Next Best).

Appendix D How Different Setups Impact the Result

We conduct an experiment on BERT after debiasing in gender and race domain to explore the impact of different number of heuristic prompts and word usage proportion on the debiasing performance. The results can be found in Table 12. Overall, as the number of heuristic prompts and the UR increase, our method also performs better in the Overall Score of CrowS-Pairs.

Appendix E Time and Additional Content Required by Several Methods

Table 13 records the time and additional content required by HEIMAT (Ours), Auto-Debias, CDA, Self-Debias and FineDeb to complete a debiasing task in our device, among which the records regarding FineDeb are derived from the original paper.

Language Category Sentence
English Gender A person who has been working as a long-haul truck driver for twenty years, and this person’s gender is widely described as [MASK].
Race A person who is the CEO of a major tech company, yet this person’s race is still sometimes described as [MASK].
French Gender Une personne qui travaille comme conducteur ou conductrice de poids lourd sur les routes depuis vingt ans, et le genre de cette personne est généralement décrit comme [MASK].
Race Une personne qui est PDG d’une grande entreprise technologique, mais l’origine de cette personne est encore parfois décrite comme [MASK].
Table 9: Heuristic prompt examples in English and French across different social attributes.
No. Template
1 [Wt1W_{t1}] [WfW_{f} about hobbies]      and [Wt2W_{t2}] [VERBS] [Wt1W_{t1}]’s [WbW_{b}] is [MASK].
2 [Wt1W_{t1}] [WfW_{f} about adjectives]      and [Wt2W_{t2}] [VERBS] [Wt1W_{t1}]’s [WbW_{b}] is [MASK].
3 [Wt1W_{t1}] [WfW_{f} about professions]      and [Wt2W_{t2}] [VERBS] [Wt1W_{t1}]’s [WbW_{b}] is [MASK].
4 [Wt1W_{t1}] [WfW_{f} about hobbies]      and [Wt1W_{t1}]’s [WbW_{b}] is [VERBS] as [MASK].
5 [Wt1W_{t1}] [WfW_{f} about adjectives]      and [Wt1W_{t1}]’s [WbW_{b}] is [VERBS] as [MASK].
6 [Wt1W_{t1}] [WfW_{f} about professions]      and [Wt1W_{t1}]’s [WbW_{b}] is [VERBS] as [MASK].
Table 10: The six heuristic prompt templates used in our method.
No. Template
1 The word [MASK] has an opposite meaning to {w} and is often used as an opposite word.
2 The synonym of word {w} is [MASK], which has a similar meaning to {w}.
3 The word that has a similar meaning to w is [MASK].
4 The opposite word to {w} is [MASK].
5 The meaning of the word {w} and the word [MASK] is opposite.
6 In this context, word {w} and word [MASK] have completely opposite meanings.
7 There are many connections between the word {w} and the word [MASK].
Table 11: Association prompt templates used in the same bias categories. "w" is the placeholder.
The Number of Heuristic Prompts Overall Score
UP = 1/4 UP = 1/2 UP = 1
5000 61.21 60.47 57.63
10000 59.81 59.15 57.03
15000 59.92 57.63 54.25
20000 55.35 54.97 51.86
Table 12: Overall Score of CrowS-Pairs on BERT processed with HEIMAT with the different number of heuristic prompts and word usage proportions (UP) of given words mentioned in Section 3.2.
Methods Time Required (hours) Additional Content Required
Auto-Debias 20 Word lists of demographic words and 5000 common words.
FineDeb 41 Word lists of demographic words.
CDA Very long (requiring retraining with DcD_{c}) New dataset DcD_{c} with counterfactual examples.
Self-Debias - Original and biased prompts using the RealToxicityPrompts dataset and Perspective API.
Ours 0.5 Heuristic, association and context prompts.
Table 13: Time spent and additional content required for several debiasing methods