This repository provides an environment setup and example code for performing Arabic Named Entity Recognition (NER) and text de-identification using Hugging Face transformer models.
If Conda is not installed, download Miniconda from: https://docs.conda.io/en/latest/miniconda.html
Make sure environment.yml is in your working directory.
conda env create -f environment.ymlIf you are using Windows system, use the following command instead:
conda env create -f .\environment_win.ymlconda activate arabic_projIf you plan to run code inside Jupyter Notebook:
conda install ipykernel
python -m ipykernel install --user --name arabic_proj --display-name "arabic_proj_2025"You can then select arabic_proj_2025 inside Jupyter.
python .\transcribe_chunk_debug.py --input "C:\Users\lm2445\arabic\V8.wav" --output_name V8This is tested on windows, you can modify the input and output_name as needed
2.1.1 If the input is Levantine Arabic not Modern Standard Arabic (MSA), please use the Levantine version:
First, need to install some dependency (Do this inside the conda environment of (arabic_proj)):
python -m pip install -U faster-whisperThen use the following to run it:
python .\transcribe_chunk_debug_levantine.py --input "C:\Users\lm2445\arabic\test2.mp3" --output_name test2_levantinepython deidentify.py python .\deidentify_debug.py --input "C:\Users\lm2445\arabic\Arabic_transcribe_deidentify\output\V8_output_transcription.txt"python run_audio_to_deid_single.py --input "C:\Users\lm2445\arabic\V8.wav"
The audio data are organized in a date-based folder structure.
All call recordings (.mp3) are stored under daily subfolders.
Calls/
├─ 20210609/
│ ├─ 1623239388485_4104_667_3.mp3
│ ├─ 1623240123456_4104_667_4.mp3
│ └─ ...
├─ 20210610/
│ ├─ 1623325xxxxx_...._...._12.mp3
│ └─ ...
├─ 20210611/
│ └─ ...
└─ ...
Related metadata files:
Dataset_Root/
├─ Calls/
│ └─ (YYYYMMDD folders with .mp3 files)
├─ output_with_filenames.xlsx
└─ Matched_data_Embrace.xlsx
python .\run_calls_tree_to_deid.py --calls_root "C:\Users\lm2445\arabic\Calls" --out_root ".\output_all" --continue_on_error --format ".mp3"First, install the package in the environment:
pip install mutagenSecond, config the path of SOURCE_ROOT in the python script "filter_and_put_into_one_folder.py"
Third, run the script:
python .\filter_and_put_into_one_folder.pyoutput message like:
Scanning files...
Total eligible files found: 1
Only 1 files available.
Copying 1 files...
Done.the filtered mp3s will be stored in the folder "./filtered_5_10min".
Config the MP3_FOLDER and SCRIPT_DIR. The first is the path to the folder of data. The second is the path to the current directory.
Then run the following command if you are using Windows System.
.\run_all.batIf you are using linux, use the run_all.sh. Config is the same.
The script copy_and_exclude_phone_number.py copies only the de-identified files from the source folder to a target folder.
Before running the script, configure the following paths inside copy_and_exclude_phone_number.py:
SOURCE_DIR = Path(r"source_folder")
DEST_DIR = Path(r"output_folder")The script will:
Select only files whose names end with deidentified. Remove the phone-number block from the filename. The phone number is assumed to be an 8-digit block besides the first block when splitting the filename by _. Copy the renamed files to the target folder.
Run the script with:
python copy_and_exclude_phone_number.py