arXiv is now an independent nonprofit! Learn more
License: CC BY 4.0
arXiv:2609.00191v1 [cs.CL] 31 Aug 2026

Assessing Suicide Risk in Arabic Crisis Helpline Calls: A Comparison of Arabic and English Large Language Models

Linhai Ma Affiliation: Department of Emergency Medicine, Yale University, New Haven, CT, USA    Rita El Hachem Email: ree21@mail.aub.edu.lb Affiliation: Department of Epidemiology and Population Health, American University of Beirut, Beirut, Lebanon    Mahatab El Hajj Email: mahatab.elhaj@embracelebanon.org Affiliation: Embrace, Mental Health Center, Beirut, Lebanon    Lilian Ghandour Email: lg01@aub.edu.lb Affiliation: Department of Epidemiology and Population Health, American University of Beirut, Beirut, Lebanon    Samah Fodeh Email: samah.fodeh@yale.edu Affiliation: Department of Emergency Medicine, Yale University, New Haven, CT, USA
Abstract

Background: Crisis helplines assess suicide risk using structured interviews given by operators. The process takes time and depends on operator training and workload. Natural language processing could support suicide-risk assessment and prioritization. Almost no work has looked at Arabic-language helpline calls, and few studies work within the privacy limits of real helpline data.

Methods: We analysed de-identified transcripts from Lebanon’s National Lifeline for Emotional Support and Suicide Prevention. The audio recordings stayed at the helpline. Calls were transcribed on site using a speech recognition model designed for Levantine Arabic. An Arabic named-entity recognition model then removed names and other identifying information, locally. Only the de-identified Arabic transcripts were shared with the research team. Operators recorded the five suicidal ideation items from the Columbia Suicide Severity Rating Scale. We combined these items into two binary outcomes: at-risk and high-risk. We also translated the transcripts into English. This allowed us to compare models on the original Arabic text with models on the English translations. We fine-tuned five instruction-tuned large language models on each corpus, together with six transformer encoders as baselines, four for Arabic and two for English. We evaluated all models on a held-out test set.

Results: We included 383 calls. Of these, 373 were available for the at-risk task, and 52.3% were positive. A total of 297 calls were available for the high-risk task, and 30.0% were positive. The best Arabic model reached a macro-F1 of 81.19 and a ROC-AUC of 90.61 for high risk. The best English model reached a macro-F1 of 85.00 and a ROC-AUC of 92.59, and identified 88.9% of the high-risk calls. The best model in each language separated high-risk calls more cleanly than at-risk calls. Translation to English did not reduce the best observed performance.

Conclusion: The model can classify calls into two suicide risk groups using de-identified Arabic transcripts. This can be done without sending audio outside the helpline. The results for high-risk calls support further testing of the model as a tool for operators. Lower-severity suicidal thoughts were the harder of the two to separate, not the more severe ones.

keywords
suicide prevention, crisis helpline, large language models, Levantine Arabic

1 Introduction

About 727 000 people died by suicide in 2021 World Health Organization (2025). Helplines reach people at a useful moment, often while the person is still in crisis and before any clinic has seen them. The calls also seem to help in themselves. One study found that callers were less distressed and less suicidal by the end of a call Gould et al. (2007). An evaluation of the Lebanese helpline studied here found the same drop in distress Zeinoun et al. (2021). However, judging risk on a helpline is still hard. Operators work through a structured questionnaire, usually the Columbia Suicide Severity Rating Scale (C-SSRS) Posner et al. (2011). They do this while the call is going on, and under time pressure. A systematic review of such questionnaires found that none of them is accurate enough to be used on its own Swedish Council on Health Technology Assessment (2015). The same review found no good evidence yet on whether adding one to the clinician’s own judgement helps. Staff turnover makes the problem worse, because good assessment depends on training that is expensive to keep up.

Natural language processing (NLP) could help here, and a growing body of work applies it to suicide risk Arowosegbe and Oyelade (2023); Castillo-Sánchez et al. (2020); Fodeh et al. (2017). Two gaps stand out in that work. Most of it uses social media posts or clinical notes, so the text is not the kind of text a real tool would see. Almost all of it is also in English, and reviews have asked for work in other languages and inside real services Castillo-Sánchez et al. (2020). Earlier helpline studies cover services that answer in English Broadbent et al. (2023); Iyer et al. (2022) and in Chinese Su et al. (2025); Wang et al. (2025). One service answers in both Hebrew and Arabic, but its published analysis used only the Hebrew chats Grimland et al. (2025). Arabic is one clear gap. Its spoken dialects differ from Modern Standard Arabic in exactly the informal way people talk on a crisis call. How people speak about suicide also depends on religious and social taboos, which vary from place to place.

Lebanon has both the need and the data. National surveys report high rates of mental disorder among young people, together with very little help-seeking Maalouf et al. (2022); Baroud et al. (2019). The first nationwide study of suicide deaths there found wide variation between regions and groups, set against strong social taboo Bizri et al. (2021). Embrace runs the National Lifeline for Emotional Support and Suicide Prevention with the Ministry of Public Health’s National Mental Health Programme. It is the only suicide prevention helpline in the country. Its calls are recorded, and its operators complete the C-SSRS under supervision. The service therefore holds recorded speech with matching assessments, which is what model development needs. This is also very sensitive data. Callers are in crisis, they mention identifying details, and they have been promised confidentiality. A recording of someone’s voice identifies them, and unlike text it cannot be edited to hide that. Any analysis therefore has to keep the audio inside the service. This limit is usually treated as an obstacle to research. We treated it as a requirement and designed the pipeline around it. We set out to build and test a model that sorts calls into risk categories from de-identified Arabic transcripts alone. We also wanted to know what it costs to translate the transcripts into English first, since translate then classify is the usual approach for a language with few resources. Beyond that, we wanted to describe how these models fail when the risk labels are imbalanced, because that is what decides whether a model is safe to put in front of an operator.

This study makes three contributions. First, it builds and tests a risk classifier for Arabic-language helpline calls. Second, it describes a working pipeline where speech recognition and de-identification run inside the service, so the recordings never move. Third, it compares Arabic text with an English translation of the same calls, using the same splits and the same labels.

1.1 Relatedwork

Most NLP work on suicide risk uses social media posts or clinical notes Arowosegbe and Oyelade (2023); Castillo-Sánchez et al. (2020); Fodeh et al. (2017). Helpline calls are different in ways that matter. They are conversations rather than written statements, and the risk assessment happens during the call instead of being pieced together afterwards. The label also comes from a trained operator at the time, not from later coding. Several groups have worked on helpline contacts directly. One study classified risk levels in text-based crisis counselling. It found that a neural model missed fewer at-risk clients than a term-frequency model Broadbent et al. (2023). Two other studies classified low against high risk from call audio rather than from words. One used voice biomarkers on an Australian telehealth service, with counsellor ratings re-checked against the C-SSRS Iyer et al. (2022). The other used acoustic features on a Chinese hotline Su et al. (2025). A further study on a Chinese hotline combined pitch features with deep learning features to analyse the emotion expressed during calls Wang et al. (2025). The Sahar helpline answers in both Hebrew and Arabic, but its published analysis used only the Hebrew chats Grimland et al. (2025). That gap shows how little is known about Arabic crisis calls. The helpline studied here has been analysed before. One study evaluated whether calls reduced caller distress Zeinoun et al. (2021), and another modelled which caller characteristics went with suicidal ideation, intent and self-harm Farran et al. (2025). Both used the structured fields the operator fills in, not the speech. As far as we can tell, two things here are new. 1) This is the first suicide-risk classifier we know of for Levantine Arabic-language crisis helpline calls. 2) It is also the first report of a full pipeline where speech recognition and de-identification run inside the service, so the recordings are never moved. Reviews of this literature have asked for work outside English and inside real services Castillo-Sánchez et al. (2020). Our results agree with the general finding that transformer models pick up suicidal content well above chance. But we also show that performance varied substantially across both outcome definitions and model backbones.

2 Methods

2.1 Settings, data source and ethical approval

Data came from calls to the Embrace National Lifeline (1564) in Lebanon. Trained operators answer the calls, supervised by clinical psychologists. They complete a standard C-SSRS assessment and record the result in a service database. The helpline is anonymous, so call records cannot be linked to national identifiers. This was a secondary analysis of routine service data. We had no contact with callers. Ethical approval was in place before the data were accessed. A call was eligible if the recording lasted between five and ten minutes. This window drops very short contacts, which are mostly wrong numbers, silent calls and immediate hang-ups. Section 4.3 discusses what this leaves out.

2.2 Privacy-preserving transcription and de-identification

Recordings did not leave the helpline. All speech processing ran on site, and only redacted Arabic text went to the analysis team. Recordings were transcribed with a Whisper-family speech recognition model Radford et al. (2023). A call is longer than the model’s input window, so we split each recording into overlapping segments and removed the duplicated text when joining the segments back together. The calls are in Levantine Arabic rather than Modern Standard Arabic, so we used a decoding setup adapted to the dialect. The transcripts then went through an Arabic named-entity recognition model, which removed names of people, places and organisations along with other direct identifiers. Phone number blocks were stripped from the filenames before release. The whole code for the transcription and de-identification is available online (see the Section Code availability).

2.3 Outcome definition

Operators record the five ideation items of the C-SSRS ladder: wish to be dead; non-specific active suicidal thoughts; active ideation with any method; active ideation with some intent to act; and active ideation with a specific plan and intent Posner et al. (2011). Each item is marked present or absent. Other codes cover items that were not asked, items skipped by the logic of the questionnaire, and items the caller did not answer. We grouped the items into two categories by severity, following the order of the ladder. The two lower rungs, a wish to be dead and non-specific active thoughts, make up the at-risk category. The three items that involve a method, an intent or a plan make up high risk:

yat\displaystyle y^{\mathrm{at}} =yWD∨yNA,\displaystyle=y^{\mathrm{WD}}\vee y^{\mathrm{NA}}, (1)
yhigh\displaystyle y^{\mathrm{high}} =yAM∨yIA∨yPI.\displaystyle=y^{\mathrm{AM}}\vee y^{\mathrm{IA}}\vee y^{\mathrm{PI}}. (2)

A call is at-risk if either of the two lower items is present, and high risk if any of the three higher items is present. The two categories cover overlapping but different sets of calls, and we modelled them as two separate yes-or-no tasks.

2.4 Handling missing annotation

Not every call has all five items recorded. A missing item does not always leave the category unclear. If any recorded item is present, the category is positive whatever the missing items would have said. The category is unclear only when every recorded item is absent and at least one item was never recorded, because then the missing item alone decides the answer. We kept the calls whose category was clear under this rule and dropped the rest. This resulted in the exclusion of 3232 calls from the at-risk category and 1616 from the high-risk category. Treating unrecorded items as absent would have allowed these calls to remain in the analysis, but would risk introducing false negatives into a corpus designed to detect suicide risk. This could bias the models toward missing positive cases, which is the more consequential error in this setting.

2.5 Machine translation

To compare modelling the source language against translate then classify, we translated each de-identified Arabic transcript into English with Qwen/Qwen2.5-72B-Instruct model under greedy decoding. We then checked that each English output covered the whole call. 2929 of the 438438 transcripts were ill-formed or incomplete, so we dropped them. We removed these calls from both corpora rather than from the English one alone. Dropping them only on the English side would leave the two languages covering different sets of calls, which would have affected the comparison between languages.

2.6 Model Training

Table 1 lists every model we fine-tuned. We selected five instruction-tuned large language model (LLM) backbone decoders covering a range of size and pretraining language, and six transformer encoders (BERT-based Models) for comparison. To fine-tune these models, we turned each call into an instruction-tuning triple. The instruction states the yes or no question, the input is the full transcript, and the target output is a single word. Every model used the same 4-bit QLoRA recipe with no per-model hyperparameter search, so that differences between models reflect the backbone rather than tuning effort Hu et al. (2022); Dettmers et al. (2023). We attached low-rank adapters (r=16r=16, α=32\alpha=32, dropout 0.05) to the attention and feed-forward projections. Base weights were frozen and quantised. Training minimised causal language-model cross-entropy on the answer token only, with prompt tokens masked, using a cosine schedule (peak learning rate 2×10−42\times 10^{-4}, effective batch size 16). We trained separate adapters per model, task and language, and ran both a three-epoch and a ten-epoch schedule. At inference we decoded the answer token greedily and mapped it to a binary label. Generations that could not be parsed were counted as negative. Zero-shot baselines used the same prompt on the 0-Shot base model. Every transcript is several times longer than a 512-token encoder window, so we read documents in overlapping windows and pooled the window representations, following standard practice for long documents Pappagari et al. (2019).

Table 1: Models fine-tuned in this study. Each model was fine-tuned/tested separately for both tasks on the corpus listed in the final column. Decoders answer the risk question with a single word. Encoders read each transcript in overlapping windows. Parameter counts are the published configurations.
Model Class Parameters Pretraining emphasis Corpus
Qwen2.5-1.5B-Instruct Qwen Team (2024) Decoder 1.5B General, multilingual Arabic, English
Qwen2.5-14B-Instruct Qwen Team (2024) Decoder 14B General, multilingual Arabic, English
Llama-3.3-70B-Instruct Grattafiori et al. (2024) Decoder 70B General, multilingual Arabic, English
AceGPT-v2-8B-Chat Huang et al. (2023); Zhu et al. (2025) Decoder 8B Arabic-centric Arabic, English
AceGPT-v2-70B-Chat Huang et al. (2023); Zhu et al. (2025) Decoder 70B Arabic-centric Arabic, English
CAMeLBERT-DA Inoue et al. (2021) Encoder 108M Arabic, dialectal Arabic
AraBERTv0.2 Antoun et al. (2020) Encoder 135M Arabic Arabic
MARBERT-based Abdul-Mageed et al. (2021) Encoder 163M Arabic, dialectal Arabic
Multilingual BERT Devlin et al. (2019) Encoder 178M Multilingual Arabic
BERT-base Devlin et al. (2019) Encoder 110M English English
BERT-large Devlin et al. (2019) Encoder 340M English English
Table 2: Derivation of the analysis sample. The two categories draw on different sets of calls, because the five ideation items were not recorded on a common set. Each category starts from the calls with at least one of its own constituent items recorded. Calls were then dropped because the category could not be determined from the recorded items, or because the machine translation was incomplete. The two categories overlap, so a call with an incomplete translation can be counted in both branches; 2929 distinct calls were affected.
Stage Calls
De-identified transcripts produced 987
Matched to a service record 545
   Matched to more than one record 171
    Kept: records agreed on all items 154
    Dropped: records disagreed 17
Transcripts with unambiguous labels 528
With ≥\geq1 recorded ideation item 438
   of which with all five items 280
   of which with an incomplete translation 29
At-risk
   With ≥\geq1 of its two items recorded 433
   Dropped: category undetermined 32
   Dropped: incomplete translation 28
   Analysable 373
    Training / test 298 / 75
High risk
   With ≥\geq1 of its three items recorded 333
   Dropped: category undetermined 16
   Dropped: incomplete translation 20
   Analysable 297
    Training / test 237 / 60
Distinct calls in at least one category 383
Table 3: Transcript length in whitespace-delimited words. The two corpora hold the same calls, so the Arabic and English rows are paired.
Language Split nn Median Mean P25 P75 Max
Arabic At-risk, training 298 730 756 597 881 1381
At-risk, test 75 731 749 608 885 1170
High, training 237 731 761 600 892 1641
High, test 60 692 702 577 829 1146
All calls 383 732 759 602 885 1641
English At-risk, training 298 928 953 762 1085 1655
At-risk, test 75 944 945 764 1097 1449
High, training 237 935 959 764 1096 2269
High, test 60 888 888 715 1069 1392
All calls 383 934 957 763 1097 2269

2.7 Evaluation

We split at the call level, stratified on the category label, with an 80:20 training to test ratio. The split was defined over transcript identifiers and applied the same way to the Arabic corpus and to its English translation. The two datasets therefore hold the same calls, in the same splits, with the same labels, and differ only in language. No call appears on both sides of the split. For evaluation, we used the area under the receiver operating characteristic curve (ROC-AUC), the area under the precision and recall curve (PR-AUC), macro-averaged F1, positive-class recall and overall accuracy. Positive-class recall is the same thing as sensitivity to the positive class. Macro-F1 here is the mean of the two per-class F1 scores, so it need not fall between macro-averaged precision and recall. All values are percentages. Table 5 reports sensitivity, specificity and the two predictive values, together with the two-by-two table of correct and incorrect decisions they derive from. Each decoder answers yes or no, so the threshold is whatever the model itself applies. ROC-AUC and PR-AUC need a continuous score rather than a decision, so for those two measures we used the probability of the positive answer token, normalised over the yes and no tokens. Proportions in Table 5 are given with a 95% Wilson score interval, which behaves better than a normal approximation when the counts are small. We report accuracy for completeness, but we did not use it to pick or compare models. With 30% to 52% of calls positive, accuracy is partly inflated by guessing the majority class, which is the behaviour this analysis is meant to catch. The study used one stratified split. With 7575 and 6060 test calls, moving one call changes macro-F1 by about one to two points. We report the rankings descriptively and draw no statistical comparison between models. The intervals in Table 5 show how wide the uncertainty is. We report ROC-AUC and PR-AUC without intervals and treat them as descriptive.

3 Results

Of 987987 de-identified transcripts, 545545 matched a service record by call identifier. Among these, 171171 identifiers matched more than one service record. Of those, 154154 carried identical ideation labels across records and were kept, while 1717 carried conflicting labels and were dropped, leaving 528528 transcripts with unambiguous labels. Of these, 438438 had at least one ideation item recorded, and 280280 had all five. The five items were not recorded on a common set of calls, so each category was derived from its own base: 433433 calls had at least one of the two at-risk items recorded, and 333333 had at least one of the three high-risk items. Applying the rule in Section 2.4 and removing calls with an incomplete machine translation left 373373 analysable calls for the at-risk category and 297297 for high risk, drawn from 383383 distinct calls in total. Table 2 shows the derivation. The at-risk category was close to balanced, at 52.352.3% positive in training and 52.052.0% in test. High risk was 30.030.0% positive in both. Stratification held training and test rates within 0.30.3 percentage points on both tasks. Calls were long (Table 3). The median Arabic transcript ran to 732732 words and the longest to 16411641. The English version was longer throughout, with a median of 934934 words, about 2828% more. Arabic morphology packs into single words what English spreads over several. Any length-based processing budget must therefore be set from the English side.

3.1 Overall model performance

Figure 1: All fine-tuned models on the Arabic corpus at ten epochs, on the held-out test split. The upper panel is the at-risk category and the lower panel is the high-risk category. R+ is positive-class recall. The dashed line marks 50 per cent. The five models left of the vertical rule are the decoders and the four to its right are the encoder baselines. Tables 6 and 7 give the macro-F1, positive-class recall and ROC-AUC values numerically.
Figure 2: All fine-tuned models on the English corpus at ten epochs, on the held-out test split. Panels, measures and layout are as in Figure 1. The English corpus is the machine translation of the same calls, in the same splits, with the same labels.

Figures 1 and 2 show every fine-tuned model on all five measures, one figure per corpus, at ten epochs. Both figures use the same scale, so the two risk categories can be read against each other. Most models sit between 65 and 90 on macro-F1, ROC-AUC and accuracy, while positive-class recall spreads much more widely, from 27.78 to 100.00 on the Arabic corpus. The best encoder falls below the best decoder in both languages and on both categories, and two Arabic encoders sit near the floor on the at-risk category. The spread across decoders is narrower on the English corpus than on the Arabic one, so the choice of backbone matters less once the text has been translated. The rest of this section reads out of these two figures.

3.2 Primary models’ performance

Table 4 reports one 70B decoder per corpus against its 0-Shot baseline, the Arabic-centric model on Arabic and the general-purpose model on English. For high risk, the Arabic model reached a macro-F1 of 81.1981.19 and a ROC-AUC of 90.6190.61, against 41.1841.18 and 58.9358.93 0-Shot. The 0-Shot model found none of the high-risk calls in the test split. After fine-tuning it found 66.766.7% of them. The English model reached a macro-F1 of 85.0085.00 and a ROC-AUC of 92.5992.59 and found 88.988.9% of high-risk calls, against 73.7873.78 and 85.4585.45 0-Shot. For the at-risk category, the Arabic model reached a macro-F1 of 73.2973.29 and a ROC-AUC of 82.1982.19. The English model reached 75.9675.96 and 87.2587.25. Both improved on their 0-Shot baselines by a wide margin. Those baselines found 12.812.8% and 38.538.5% of at-risk calls.

Table 4: Primary 70B decoders, before and after fine-tuning, on the held-out test split. All values are percentages. R+ is positive-class recall. Table 5 gives the operating characteristics for the same eight conditions.
Model category (% pos) Macro-F1 R+ PR-AUC ROC-AUC Accuracy
Arabic corpus, Arabic-centric 70B decoder (ACEGPT 70B)
   0-Shot At-risk (52) 45.33 12.82 58.18 56.91 54.67
   Fine-tuned At-risk (52) 73.29 74.36 81.13 82.19 73.33
   0-Shot High (30) 41.18 0.00 38.36 58.93 70.00
   Fine-tuned High (30) 81.19 66.67 79.55 90.61 85.00
English corpus, general-purpose 70B decoder (LLAMA 70B)
   0-Shot At-risk (52) 56.01 38.46 68.54 64.89 57.33
   Fine-tuned At-risk (52) 75.96 76.92 89.85 87.25 76.00
   0-Shot High (30) 73.78 61.11 78.75 85.45 78.33
   Fine-tuned High (30) 85.00 88.89 80.79 92.59 86.67

3.3 Model performance across risk categories

Table 5 gives sensitivity, specificity and the predictive values for the same eight conditions. PPV is positive predictive value and NPV is the negative predictive value. The two fine-tuned models behaved differently on the high-risk category. The English model found 16 of the 18 high-risk calls, giving a sensitivity of 88.988.9% and a negative predictive value of 94.794.7%. It also raised 6 false alarms out of 42 negative calls. The Arabic model missed 6 of the 18, so its sensitivity was 66.766.7%, but it raised only 3 false alarms. On the at-risk category the two models were close, and both sat near 7575% on all four measures. The intervals are wide throughout. On the high-risk category they are wide enough that the two languages cannot be told apart. With only 18 positive calls in that test split, one call moves sensitivity by more than five percentage points.

Table 5: Operating characteristics on the held-out test split, with 95% Wilson score intervals. Counts are true positives, false positives, false negatives and true negatives. Arabic denotes the Arabic-centric 70B decoder (ACEGPT 70B) on the Arabic corpus. English is the general-purpose 70B decoder (LLAMA 70B) on the English corpus. The conditions are the same eight reported in Table 4. Predictive value is undefined where a model predicted no positive case.
Model TP/FP/FN/TN Sensitivity Specificity PPV NPV
At-risk category (52% positive)
Arabic, 0-Shot 5/0/34/36 12.8 [1pt](5.6 to 26.7) 100.0 [1pt](90.4 to 100.0) 100.0 [1pt](56.6 to 100.0) 51.4 [1pt](40.0 to 62.8)
Arabic, fine-tuned 29/10/10/26 74.4 [1pt](58.9 to 85.4) 72.2 [1pt](56.0 to 84.2) 74.4 [1pt](58.9 to 85.4) 72.2 [1pt](56.0 to 84.2)
English, 0-Shot 15/8/24/28 38.5 [1pt](24.9 to 54.1) 77.8 [1pt](61.9 to 88.3) 65.2 [1pt](44.9 to 81.2) 53.8 [1pt](40.5 to 66.7)
English, fine-tuned 30/9/9/27 76.9 [1pt](61.7 to 87.4) 75.0 [1pt](58.9 to 86.2) 76.9 [1pt](61.7 to 87.4) 75.0 [1pt](58.9 to 86.2)
High-risk category (30% positive)
Arabic, 0-Shot 0/0/18/42 0.0 [1pt](0.0 to 17.6) 100.0 [1pt](91.6 to 100.0) undefined 70.0 [1pt](57.5 to 80.1)
Arabic, fine-tuned 12/3/6/39 66.7 [1pt](43.7 to 83.7) 92.9 [1pt](81.0 to 97.5) 80.0 [1pt](54.8 to 93.0) 86.7 [1pt](73.8 to 93.7)
English, 0-Shot 11/6/7/36 61.1 [1pt](38.6 to 79.7) 85.7 [1pt](72.2 to 93.3) 64.7 [1pt](41.3 to 82.7) 83.7 [1pt](70.0 to 91.9)
English, fine-tuned 16/6/2/36 88.9 [1pt](67.2 to 96.9) 85.7 [1pt](72.2 to 93.3) 72.7 [1pt](51.8 to 86.8) 94.7 [1pt](82.7 to 98.5)

3.4 Better separation for more severe cases

The best model in each language scored higher on the high-risk category than on the at-risk category, even though high risk is the rarer outcome. The best Arabic figures were 81.1981.19 macro-F1 for high risk against 78.6378.63 for the at-risk category. The best English figures were 85.0085.00 against 78.6678.66 (Table 6). ROC-AUC followed the same order. This is the opposite of what class prevalence alone would predict. Smaller backbones did not follow it: on the Arabic corpus the 1.5B, 14B and 8B decoders all scored higher on the at-risk category. Positive-class recall did not follow. Every Arabic decoder found a smaller share of high-risk calls than of at-risk calls, and so did three of the five English ones. On the Arabic corpus the gain on high risk is therefore in separation rather than in how many positive calls are caught at the model’s own threshold. The separation itself suggests the main limit is not how rare the label is but how the construct is expressed. Talk of a method, an intent or a plan is direct and easy to quote. Lower-severity ideation is indirect, and its wording overlaps heavily with the wording of general distress, which the helpline also receives in volume.

Table 6 reports all ten decoder conditions at ten epochs. Parameter count did not set the order: the 8B Arabic-centric model reached 78.6378.63 macro-F1 on the Arabic at-risk category against 73.2973.29 for its 70B counterpart.

Table 6: Exploratory comparison of all fine-tuned decoder conditions at ten epochs. R+ is positive-class recall, that is, sensitivity to the risk-positive class. Given the size of the test split these comparisons are descriptive and do not fix an order among backbones.
Corpus Backbone Macro-F1 R+ ROC-AUC
At-risk High At-risk High At-risk High
Arabic 1.5B general 59.82 53.12 64.10 27.78 65.17 66.40
14B general 71.92 70.44 64.10 44.44 84.37 87.70
70B general 74.65 77.78 69.23 55.56 82.02 85.05
8B Arabic-centric 78.63 67.16 71.79 33.33 84.47 83.07
70B Arabic-centric 73.29 81.19 74.36 66.67 82.19 90.61
English 1.5B general 77.33 67.17 74.36 50.00 83.01 85.19
14B general 75.89 72.21 79.49 72.22 84.01 81.88
70B general 75.96 85.00 76.92 88.89 87.25 92.59
8B Arabic-centric 78.66 79.11 76.92 77.78 85.90 89.35
70B Arabic-centric 78.66 79.48 76.92 66.67 85.83 91.34

3.5 BERT Encoder baselines

The best fine-tuned decoder (LLM) beat the best encoder baseline (BERT) on every task and language (Table 7). The best Arabic encoder score was 69.2569.25 macro-F1 on the at-risk category and 72.8272.82 on high risk, against 78.6378.63 and 81.1981.19 for decoders. The best English encoder score was 64.5764.57 and 71.2971.29, against 78.6678.66 and 85.0085.00. No single encoder was best on both categories in either language.

Table 7: Transformer encoder baselines, fine-tuned for ten epochs. R+ is positive-class recall, that is, sensitivity to the risk-positive class. Encoders read each transcript in overlapping windows, since all transcripts exceed the 512-token input limit.
Corpus Encoder Macro-F1 R+ ROC-AUC
At-risk High At-risk High At-risk High
Arabic CAMeLBERT-DA 68.53 70.69 82.05 61.11 71.65 77.12
AraBERTv0.2 69.25 67.23 71.79 83.33 72.44 72.88
MARBERT-based 33.63 53.28 97.44 83.33 58.40 67.06
Multilingual BERT 34.21 72.82 100.00 77.78 55.34 76.06
English BERT-base 64.57 54.38 76.92 72.22 69.66 70.63
BERT-large 62.12 71.29 71.79 77.78 67.66 83.07

3.6 Model performance under class imbalance

Several conditions produced trivial classifiers, and they failed in opposite directions (Table 8). The smallest decoder on Arabic high risk, trained for three epochs, answered negative to every call, giving 0.000.00 positive-class recall at 41.1841.18 macro-F1. Three Arabic encoders on the at-risk category did the reverse. They answered positive to almost every call, with positive-class recall of 97.497.4 to 100.0100.0 at macro-F1 between 33.6333.63 and 37.2237.22. Both kinds of failure therefore land in a similar macro-F1 range of roughly 3333 to 4141. Macro-F1 on its own cannot tell a model that detects no positive case from one that flags every caller. Gaps between ROC-AUC and PR-AUC picked out the same conditions. One Arabic encoder trained for three epochs reached 75.1375.13 ROC-AUC with 47.2447.24 PR-AUC on high risk, which describes a model that ranks calls acceptably while deciding poorly. Positive-class recall and PR-AUC should therefore be reported alongside macro-averaged measures whenever risk labels are imbalanced.

Table 8: Trivial classifiers in both directions give similar macro-F1. Positive-class recall separates them. Macro-F1 and accuracy do not.
Condition category Direction Macro-F1 R+ Accuracy
1.5B decoder, Arabic, 3 epochs High (30) All negative 41.18 0.00 70.00
CAMeLBERT-DA, 3 epochs At-risk (52) Near all positive 37.22 100.00 53.33
MARBERT-based, 3 epochs At-risk (52) All positive 34.21 100.00 52.00
MARBERT-based, 10 epochs At-risk (52) Near all positive 33.63 97.44 50.67
Multilingual BERT, 10 epochs At-risk (52) All positive 34.21 100.00 52.00

4 Discussion

Our findings extend prior work on suicide-risk detection in several important ways. Previous studies have largely focused on social media and clinical notes, while the smaller body of helpline research has examined text-based counselling, acoustic features, or structured operator assessments. Our study moves this work to Levantine Arabic crisis calls, using the speech content itself rather than operator-entered fields or acoustic biomarkers. The results are consistent with prior evidence that transformer-based models can identify suicide-related content, but they also show that performance depends strongly on the risk category being detected and the model backbone, rather than following a simple pattern in which more severe risk is necessarily harder to identify. The finding that English translation can preserve classification performance further suggests a potential pathway for applying existing English-language models to Arabic crisis calls, although this requires validation beyond the present dataset. Most importantly, the study extends prior work by demonstrating that these models can be incorporated into a privacy-preserving, within-service pipeline, addressing the need for suicide-risk NLP research that operates in real-world, non-English crisis settings without requiring transfer of raw recordings.

4.1 Implications for Suicide-Risk Detection in Arabic Helpline Calls

Our analysis have shown that sorting de-identified Levantine Arabic helpline transcripts into two risk categories is feasible. For the high-risk category, the models achieved performance sufficient to support further evaluation as a screening aid for helpline operators. Importantly, this performance was achieved within a privacy-preserving pipeline in which the audio remains within the helpline environment. Additionally, we found that lower-severity ideation was more difficult to distinguish than high-risk ideation. This contrasts with the field’s predominant focus on detecting rare, severe outcomes and suggests that distinguishing less severe but clinically meaningful ideation may require greater attention. We have also shown that translating the transcripts into English before classification did not reduce performance. This is notable because the English input undergoes two transformations—automatic speech recognition followed by translation—whereas the Arabic input undergoes only speech recognition. One possible explanation is that the larger volume of English-language pretraining data helps compensate for information lost during these transformations. However, this finding is based on a single sample and translation model and should not be interpreted as evidence that translation will generally improve or preserve performance across settings. In addition, macro-F1 and accuracy alone can mask clinically important failure modes. A model that predicts high risk for nearly every caller and one that predicts no callers as high risk may achieve seemingly moderate aggregate scores while being clinically unusable. Positive-class recall and PR-AUC should therefore be reported alongside macro-averaged metrics to provide a clearer picture of performance on the risk-positive class.

4.2 Implications for Helpline Practice

The role of the model is not to determine suicide risk on its own, but to support the people doing this work. For example, the model could flag calls for a supervisor to review, help prioritize callers for follow-up, or remind an operator to complete a structured risk assessment. In this role, detecting positive cases is more important than avoiding false alarms. An unnecessary review may cost a supervisor a few minutes, whereas missing a high-risk caller could have far more serious consequences. Two steps should precede any use in practice. First, performance should be evaluated on consecutive calls as they arrive, rather than on a retrospective sample that may be limited by factors such as call duration. Second, helpline staff should be given estimates of the expected number of additional reviews per week so they can understand the workload implications before the tool is deployed, rather than after implementation.

4.3 Limitations

All the data come from one helpline in one country over a limited period. We do not know how the model would do in another service, another dialect region or another time. Only 545545 of the 987987 de-identified transcripts matched a service record, and a further 1717 were dropped because duplicate records disagreed. The 442442 unmatched calls carry no recorded assessment, so we cannot check whether they differ from the calls we kept. Calls longer than ten minutes were outside the eligibility window, so the sample does not cover the longest calls. We did not measure word error rate against transcripts written by hand. Levantine crisis speech is hard to transcribe, with dialect, broken speech, crying and overlapping voices. Both languages inherit those errors, and the English side adds translation error on top. Part of what we are comparing is therefore the quality of the text going in. A small hand-transcribed sample would settle this. The labels are the operator’s recorded C-SSRS assessment, not an independent review. A category is positive if any of its items is positive, so one wrong item makes the whole label positive. Errors in the labels therefore push in one direction only. At best the model reproduces operator judgement, including its mistakes. It does not predict suicidal behaviour, which was never observed here.

4.4 Conclusion

Sorting privacy-preserved Arabic helpline transcripts into two risk categories is feasible. Performance for the high-risk category supports further evaluation of the model as a potential decision-support tool for helpline operators. Lower-severity ideation was more difficult to distinguish than high-risk ideation. Translating the transcripts into English before classification also produced useful results. Positive-class recall and PR-AUC should be reported alongside macro-averaged metrics because macro-averaged scores alone can obscure poor performance on the positive class.

Acknowledgements

References

  • Abdul-Mageed et al. (2021) M. Abdul-Mageed, A. Elmadany, and E. M. B. Nagoudi ARBERT & MARBERT: deep bidirectional transformers for Arabic. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics, pp. 7088–7105. External Links: Document Cited by: Table 1.
  • Antoun et al. (2020) W. Antoun, F. Baly, and H. Hajj AraBERT: transformer-based model for Arabic language understanding. In Proceedings of the 4th Workshop on Open-Source Arabic Corpora and Processing Tools (OSACT4), LREC 2020, pp. 9–15. Cited by: Table 1.
  • Arowosegbe and Oyelade (2023) A. Arowosegbe and T. Oyelade Application of natural language processing (NLP) in detecting and preventing suicide ideation: a systematic review. International Journal of Environmental Research and Public Health 20 (2), pp. 1514. External Links: Document Cited by: §1.1, §1.
  • Baroud et al. (2019) E. Baroud, L. A. Ghandour, L. Alrojolah, P. Zeinoun, and F. T. Maalouf Suicidality among Lebanese adolescents: prevalence, predictors and service utilization. Psychiatry Research 275, pp. 338–344. External Links: Document Cited by: §1.
  • Bizri et al. (2021) M. Bizri, L. Zeinoun, A. M. Mihailescu, M. Daher, M. Atoui, R. Chammay, and Z. Nahas A closer look at patterns and characteristics of suicide in Lebanon: a first nationwide report of cases from 2008 to 2018. Asian Journal of Psychiatry 59, pp. 102635. External Links: Document Cited by: §1.
  • Broadbent et al. (2023) M. Broadbent, M. Medina Grespan, K. Axford, X. Zhang, V. Srikumar, B. Kious, and Z. Imel A machine learning approach to identifying suicide risk among text-based crisis counseling encounters. Frontiers in Psychiatry 14, pp. 1110527. External Links: Document Cited by: §1.1, §1.
  • Castillo-Sánchez et al. (2020) G. Castillo-Sánchez, G. Marques, E. Dorronzoro, O. Rivera-Romero, M. Franco-Martín, and I. De la Torre-Díez Suicide risk assessment using machine learning and social networks: a scoping review. Journal of Medical Systems 44 (12), pp. 205. External Links: Document Cited by: §1.1, §1.
  • Dettmers et al. (2023) T. Dettmers, A. Pagnoni, A. Holtzman, and L. Zettlemoyer QLoRA: efficient finetuning of quantized LLMs. In Advances in Neural Information Processing Systems (NeurIPS), Vol. 36. Cited by: §2.6.
  • Devlin et al. (2019) J. Devlin, M.-W. Chang, K. Lee, and K. Toutanova BERT: pre-training of deep bidirectional transformers for language understanding. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics, pp. 4171–4186. External Links: Document Cited by: Table 1, Table 1, Table 1.
  • Farran et al. (2025) D. Farran, M. El Haj, M. Zarzour, Y. Chamoun, B. Kazazian, P. Posbic, P. Zeinoun, M. Hazimeh, Z. Nahas, M. Atoui, and R. El Chammay Factors associated with suicidal ideation and behaviour: analysis from the national suicide prevention helpline in Lebanon. BMC Psychiatry 25. External Links: Document Cited by: §1.1.
  • Fodeh et al. (2017) S. Fodeh, J. Goulet, C. Brandt, and A.-T. Hamada Leveraging Twitter to better identify suicide risk. In Proceedings of the First Workshop on Medical Informatics and Healthcare, 23rd ACM SIGKDD Conference on Knowledge Discovery and Data Mining, Proceedings of Machine Learning Research, Vol. 69, pp. 1–7. External Links: Link Cited by: §1.1, §1.
  • Gould et al. (2007) M. S. Gould, J. Kalafat, J. L. Harrismunfakh, and M. Kleinman An evaluation of crisis hotline outcomes. part 2: suicidal callers. Suicide and Life-Threatening Behavior 37 (3), pp. 338–352. External Links: Document Cited by: §1.
  • Grattafiori et al. (2024) A. Grattafiori, A. Dubey, A. Jauhri, A. Pandey, et al. The Llama 3 herd of models. arXiv preprint arXiv:2407.21783. Cited by: Table 1.
  • Grimland et al. (2025) M. Grimland, M. Liberman, H. Yeshayahu, J. Benatov, N. Munz, A. Segal, L. Ben Dayan, I. Shenfeld, K. Gal, and Y. Levi-Belz Explainable AI for suicide risk detection: gender- and age-specific patterns from real-time crisis chats. Frontiers in Medicine 12, pp. 1703755. External Links: Document Cited by: §1.1, ��1.
  • Hu et al. (2022) E. J. Hu, Y. Shen, P. Wallis, Z. Allen-Zhu, Y. Li, S. Wang, L. Wang, and W. Chen LoRA: low-rank adaptation of large language models. In International Conference on Learning Representations (ICLR), Cited by: §2.6.
  • Huang et al. (2023) H. Huang, F. Yu, J. Zhu, X. Sun, H. Cheng, D. Song, Z. Chen, A. Alharthi, B. An, Z. Liu, et al. AceGPT: localizing large language models in Arabic. arXiv preprint arXiv:2309.12053. Cited by: Table 1, Table 1.
  • Inoue et al. (2021) G. Inoue, B. Alhafni, N. Baimukan, H. Bouamor, and N. Habash The interplay of variant, size, and task type in Arabic pre-trained language models. In Proceedings of the Sixth Arabic Natural Language Processing Workshop (WANLP), pp. 92–104. Cited by: Table 1.
  • Iyer et al. (2022) R. Iyer, M. Nedeljkovic, and D. Meyer Using voice biomarkers to classify suicide risk in adult telehealth callers: retrospective observational study. JMIR Mental Health 9 (8), pp. e39807. External Links: Document Cited by: §1.1, §1.
  • Maalouf et al. (2022) F. T. Maalouf, L. Alrojolah, L. Akoury-Dirani, M. Barakat, D. Brent, M. Elbejjani, W. Shamseddeen, and L. A. Ghandour Psychopathology in children and adolescents in Lebanon study (PALS): a national household survey. Social Psychiatry and Psychiatric Epidemiology 57 (4), pp. 761–774. External Links: Document Cited by: §1.
  • Pappagari et al. (2019) R. Pappagari, P. Zelasko, J. Villalba, Y. Carmiel, and N. Dehak Hierarchical transformers for long document classification. In 2019 IEEE Automatic Speech Recognition and Understanding Workshop (ASRU), pp. 838–844. External Links: Document Cited by: §2.6.
  • Posner et al. (2011) K. Posner, G. K. Brown, B. Stanley, D. A. Brent, K. V. Yershova, M. A. Oquendo, G. W. Currier, G. A. Melvin, L. Greenhill, S. Shen, and J. J. Mann The Columbia Suicide Severity Rating Scale: initial validity and internal consistency findings from three multisite studies with adolescents and adults. American Journal of Psychiatry 168 (12), pp. 1266–1277. External Links: Document Cited by: §1, §2.3.
  • Qwen Team (2024) Qwen Team Qwen2.5 technical report. arXiv preprint arXiv:2412.15115. Cited by: Table 1, Table 1.
  • Radford et al. (2023) A. Radford, J. W. Kim, T. Xu, G. Brockman, C. McLeavey, and I. Sutskever Robust speech recognition via large-scale weak supervision. In Proceedings of the 40th International Conference on Machine Learning, Proceedings of Machine Learning Research, Vol. 202, pp. 28492–28518. Cited by: §2.2.
  • Su et al. (2025) Z. Su, H. Jiang, Y. Yang, X. Hou, Y. Su, and L. Yang Acoustic features for identifying suicide risk in crisis hotline callers: machine learning approach. Journal of Medical Internet Research 27, pp. e67772. External Links: Document Cited by: §1.1, §1.
  • Swedish Council on Health Technology Assessment (2015) Swedish Council on Health Technology Assessment Instruments for suicide risk assessment. Swedish Council on Health Technology Assessment (SBU), Stockholm. Note: SBU Yellow Report No. 242. SBU Systematic Review Summaries Cited by: §1.
  • Wang et al. (2025) H. Wang, J. Li, Q. Zhao, Z. Chen, C. Song, J. Tang, Y. Huang, W. Zhai, Y. Tong, and G. Fu Deep learning-based feature fusion for emotion analysis and suicide risk differentiation in Chinese psychological support hotlines. arXiv preprint arXiv:2501.08696. Cited by: §1.1, §1.
  • World Health Organization (2025) World Health Organization Suicide worldwide in 2021: global health estimates. Technical report World Health Organization, Geneva. Note: ISBN 978-92-4-011006-9 Cited by: §1.
  • Zeinoun et al. (2021) P. A. Zeinoun, F. E. Yehia, L. Z. Khederlarian, S. F. Yordi, M. M. Atoui, R. El Chammay, and Z. H. Nahas Evaluation of Lebanon’s national helpline for emotional support and suicide prevention: reduction of emotional distress among callers. Intervention 19 (2), pp. 197–207. External Links: Document Cited by: §1.1, §1.
  • Zhu et al. (2025) J. Zhu, H. Huang, Z. Lin, J. Liang, P. Tang, X. Yang, L. Zhang, R. Sun, H. Li, B. Wang, and J. Xu Second language (Arabic) acquisition of LLMs via progressive vocabulary expansion. In Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), Note: arXiv:2412.12310 Cited by: Table 1, Table 1.