Skip to content

Latest commit

 

History

19 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

Sci-ZSEL

Bridging Lexical Divergence: LLM-Assisted, Cost-Efficient, Zero-shot Scientific Entity Linking| Findings of EMNLP 2026 | arXiv:2609.00228

This repository reproduces the full pipeline end to end. The worked example throughout is NCBI-Disease → MEDIC, which is what config.json ships configured for.


Table of contents


Documentation

Details live in docs/:

Doc What's in it
Requirements Hardware, the two conda environments, exact versions, Ollama
Machine-specific paths Read this first — every hardcoded path you must repoint
Pretrained checkpoints Download commands for the three model files
Data What ships, corpus/ontology formats, generated intermediates
Experiment names Decoding (m1_e1)U(m3_e1)_multi_prime_rm_sm_e... and the HO/MINT/LO/NO categories
Settings and results The four training settings and the exact result file for each stage
LLM cost accounting Tokens and wall clock spent on alias generation
Troubleshooting Known quirks and how they fail

1. Repository layout

Sci-ZSEL-release/
├── config.json                     # single source of truth: which corpus/ontology to run
├── install.sh                      # creates the two conda envs
├── data_preparation/               # §5.1–5.2 of the paper: entity selection + pseudo-pair construction
│   ├── convert_grag_to_blink_format.py   # raw corpus  -> BLINK format
│   ├── grag_to_blink.sh
│   ├── entity_selection.py               # E_EM (exact match) and E_BT (bi-encoder top-1) entity sets
│   ├── alias_generation.py               # LLM alias generation (via a local Ollama server)
│   ├── alias_generation.sh               # starts/stops Ollama around alias_generation.py
│   ├── run_ollama_to_serve_llm           # the `ollama serve` invocation (env vars live here)
│   ├── pseudo_pair_construction.py       # alias matching, ontology-aware filter, dataset merging
│   ├── entity_selection_and_pseudo_pair_construction.sh
│   ├── prompts/<corpus>/{system,human}.txt   # per-domain alias-generation prompts
│   └── utils.py                          # format conversion, overlap categories, ontology relations
├── modeling/
│   ├── BLINK/                      # vendored + modified facebookresearch/BLINK
│   │   ├── blink/biencoder/        # retriever (train_biencoder.py, eval_biencoder.py)
│   │   ├── blink/crossencoder/     # reranker (train_cross.py)
│   │   ├── category_eval.py        # HO/MINT/LO/NO evaluation, recall@k, MRR
│   │   └── scripts/                # entry points + load_config.sh (config.json -> shell vars)
│   └── ReS/                        # vendored + modified HITsz-TMG/Read-and-Select reranker
│       ├── get_retriever_candidates.py   # BLINK top-64 -> ReS input format
│       ├── train_and_eval.py             # entry point (reads config.json -> "res")
│       ├── train_on_other_data.py        # maps config.json -> ReS argparse namespace
│       └── run_disambiguation_attention.py
├── datasets/                       # raw corpora + ontologies (tracked in git)
├── docs/                           # the reference docs linked above
└── saved_models/                   # checkpoints + all outputs (NOT tracked; you create it)

saved_models/, datasets/*/blink_format/, datasets/*/res_format/ and logs/ are git-ignored. A fresh clone contains only the raw corpora and ontologies; everything else is regenerated by the steps below.

All five corpora sit side by side under datasets/: the two existing benchmarks, NCBI-Disease (datasets/ncbi_disease/) and BC5CDR (datasets/bc5cdr/), plus the three new animal-science QTL datasets released with this paper — CMO (datasets/cmo/), VT (datasets/vt/) and LPT (datasets/lpt/). See Data for sizes, ontologies and formats.


2. Installation

Before you start: check Requirements (GPU, the two conda envs, Ollama), repoint the hardcoded paths listed in Machine-specific paths, and download the three model files in Pretrained checkpoints.

git clone https://github.com/rasel-isu/Sci-ZSEL.git
cd Sci-ZSEL
# edit install.sh first - see docs/machine-specific-paths.md
bash install.sh

install.sh creates both envs and installs modeling/BLINK/requirements.txt and modeling/ReS/requirements.txt respectively.

Ollama

Alias generation (step 5.2) talks to a local Ollama server, so the ollama binary must be installed and on your PATH. Official instructions and downloads: https://ollama.com/download (source: https://github.com/ollama/ollama).

On Linux with root:

curl -fsSL https://ollama.com/install.sh | sh

On a cluster without root, unpack the release tarball into a directory you own and add it to PATH:

mkdir -p ~/apps/ollama && cd ~/apps/ollama
curl -fsSL https://ollama.com/download/ollama-linux-amd64.tgz | tar -xz
export PATH="$HOME/apps/ollama/bin:$PATH"   # add to ~/.bashrc to make it stick
ollama --version

Then pull the alias model once (~6.4 GB at fp16):

export OLLAMA_MODELS=/path/to/your/ollama/models   # big; keep it off your home quota
export OLLAMA_HOST=127.0.0.1:11435
ollama serve &
ollama pull llama3.2:3b-instruct-fp16

Point data_preparation/run_ollama_to_serve_llm at that same OLLAMA_MODELS directory (machine-specific paths). That file is what the pipeline actually launches, so ollama must resolve on PATH when it runs — on a cluster, load the module or export the PATH line above in your job script.


3. config.json reference

One file at the repo root selects the corpus and ontology for every step. It is read by relative path (../config.json from data_preparation/, ../../config.json from modeling/BLINK/ and modeling/ReS/), which is why each script must be run from its own directory.

{
    "world": "ncbi_disease",
    "kb_name": "medic",
    "kb_file": "medic.json",
    "data_dir": "../datasets/ncbi_disease/",
    "saved_model_dir": "../saved_models/ncbi_disease/",
    "synonym_key_on_ontology": "synonyms",
    "has_ground_truth": true,
    "has_ent_alt_id": true,
    "exact_match_file": "m1_e1_unq_ents.json",
    "biencoder_top1_file": "top_1_from_biencoder.json",
    "blink": {
        "split_name": "train",
        "base_models": {
            "biencoder": "../saved_models/biencoder_wiki_large.bin",
            "crossencoder": "../saved_models/crossencoder_wiki_large.bin"
        },
        "candidate_generation": { "top_k": 64, "max_context_length": 64, "...": "..." },
        "retriever": {
            "exp_list": ["synonymU(m1_e1)U(m3_e1)_multi_prime_rm_sm_eU(m4_e2)_multi_prime_rm_sm_e"],
            "negative_selection": ["add_prch_in_pos_list"],
            "seeds": [0], "epochs": 1, "train_batch_size": 64, "learning_rate": "2e-05", "...": "..."
        },
        "reranker": {
            "exp_list": ["synonymU(m1_e1)U(m3_e1)_multi_prime_rm_sm_eU(m4_e2)_multi_prime_rm_sm_e"],
            "negative_selection": "only_bienc_20_neg",
            "seeds": [0], "epochs": 3, "train_batch_size": 16, "learning_rate": "2e-05", "...": "..."
        }
    },
    "res": {
        "pretrained_model": "../../saved_models/zeshel_disambiguation_attention.pt",
        "hf_model": "roberta-base",
        "split_name": "train",
        "exp_list": ["synonymU(m1_e1)U(m3_e1)_multi_prime_rm_sm_eU(m4_e2)_multi_prime_rm_sm_e"],
        "seeds": [0], "lora": false, "use_title_during_testing": true,
        "epochs": 3, "learning_rate": 1e-4, "batch_size": 16,
        "cand_num_train": 21, "cand_num": 64, "type_loss": "sum_log_nce",
        "max_len": 512, "gpus": "0,1,2,3", "...": "..."
    }
}

("..." above stands in for the remaining hyperparameters — see the real file, or scripts/load_config.sh, which lists every key it reads together with its default.)

Key Meaning
world Corpus/domain tag. Also selects data_preparation/prompts/<world>/ and the ontology-relation loader in pseudo_pair_construction.py. Must be one of ncbi_disease, bc5cdr, cmo, vt, lpt, cometa.
kb_name Ontology tag used by ReS to locate the KB file (medic, mesh, cmo, vt, lpt).
kb_file Ontology filename inside data_dir.
data_dir Corpus directory, relative to data_preparation/.
saved_model_dir Output root for this world, relative to data_preparation/.
synonym_key_on_ontology Which ontology key holds curated synonyms (synonyms, or synonym for raw MeSH).
has_ground_truth false for the QTL training corpora (cmo, vt, lpt), whose 16,385-mention train split ships with ground_truth: []. When false, pseudo-label quality reporting is skipped instead of crashing on empty ground_truth. Independently of this flag, biencoder_eval_report skips any split it finds no labels in and writes a short note to *_eval.txt in place of MRR/recall — that check is on the data, not on the config, so the labeled test split of an unlabeled corpus is still scored normally.
has_ent_alt_id Whether the ontology has an altdiseaseid list to expand during matching/scoring.
exact_match_file Filename for the E_EM entity set.
biencoder_top1_file Filename for the E_BT entity set.
blink.split_name Which split directory the BLINK scripts operate on (train).
blink.base_models.* Pretrained BLINK checkpoints, relative to data_preparation/.
blink.candidate_generation.* Flags for eval_biencoder.py (top-k, batch sizes, bert_model). has_gt defaults to the top-level has_ground_truth; when false the --has_gt flag is omitted entirely rather than passed as false — see Troubleshooting. It reaches only get_biencoder_top_k.sh, the step that reads the raw (possibly unlabeled) corpus; get_biencoder_cands_for_reranker_training.sh hard-codes --has_gt true, because both of its inputs — the annotated test set and the pseudo-labeled training set — always carry labels.
blink.retriever.* Flags for train_biencoder.py. exp_list and negative_selection are arrays — every combination is run in turn.
blink.reranker.* Flags for train_cross.py. exp_list and seeds are arrays.
res.pretrained_model ReS checkpoint, relative to modeling/ReS/.
res.hf_model ReS backbone (roberta-base).
res.exp_list, res.seeds Arrays. exp_list also drives get_retriever_candidates.py, which always converts original_data (the test set) plus one training set per entry.
res.lora, res.use_title_during_testing, res.split_name Run-level switches.
res.* (remaining) Every ReS hyperparameter: epochs, learning_rate, batch_size, cand_num_train, cand_num, type_loss, max_len, max_ent_len, max_text_len, warmup_proportion, weight_decay, adam_epsilon, gradient_accumulation_steps, num_workers, clip, info_token_num, gpus, logging_steps, eval_step.

The BLINK shell scripts read all of this through modeling/BLINK/scripts/load_config.sh and the ReS modules read CONFIG['res'] directly, so neither world nor any hyperparameter is spelled out in a script.


4. Running the pipeline

Everything below is for the shipped NCBI-Disease configuration. cd from the repo root each time — every script reads config.json by a relative path and will fail from anywhere else.

Install

bash install.sh

Convert into expected data format

cd data_preparation
bash grag_to_blink.sh

5.1 Entity Selection from the Ontology

cd modeling/BLINK/
bash scripts/get_biencoder_top_k.sh

5.2 Pseudo-Pair Construction

cd data_preparation
bash entity_selection_and_pseudo_pair_construction.sh

5.3 Ontology-Aware Retriever Fine-tuning

cd modeling/BLINK/
bash scripts/retriever_fine_tuning.sh

5.4 Reranker Fine-tuning

cd modeling/BLINK/
bash scripts/blink_reranker_fine_tuning.sh
cd modeling/ReS
bash res_reranker_fine_tuning.sh

What each step does

Step Reads Writes
grag_to_blink.sh *_grag.json, ontology blink_format/train/original_data/ — train.jsonl (retrieval pool), test.jsonl (eval set), kb.jsonl, id_map.json
get_biencoder_top_k.sh original_data/train.jsonl top-64 lists from the off-the-shelf bi-encoder → saved_models/<world>/biencoder/train/original_data/top64_candidates/train.json; also builds the cached ontology encodings *_entity_pool.t7 / *_entity_encodings.t7
entity_selection_and_pseudo_pair_construction.sh the above entity sets E_EM + E_BT → LLM aliases (Ollama) → ontology-aware filter → one blink_format/train/<exp>/train.jsonl per configuration (see experiment names)
retriever_fine_tuning.sh blink_format/train/<exp>/ fine-tuned retriever + epoch_<i>/top64_candidates/test_eval.txt
blink_reranker_fine_tuning.sh regenerates top-64, then trains cross-encoder + epoch_<i>/crossencoder_predictions_eval.txt
res_reranker_fine_tuning.sh res_format/train/<exp>/ ReS model + <epoch>/pred_eval.txt

Knobs

With one exception, neither the BLINK scripts nor the ReS entry points hold settings of their own (the exception is --has_gt in get_biencoder_cands_for_reranker_training.sh, noted in §3). The BLINK scripts source scripts/load_config.sh, which parses config.json with python3 and exports every path and hyperparameter; the ReS modules read CONFIG['res'] directly. So one file drives all of it:

To change Edit
Corpus / ontology config.json top level (world, kb_name, kb_file, data_dir, saved_model_dir)
Which pseudo-pair set to train on blink.retriever.exp_list, blink.reranker.exp_list and res.exp_list in config.json
Seeds, epochs, batch size, lr, negative selection blink.retriever.* / blink.reranker.* / res.* in config.json
Candidate-generation top-k, batch sizes blink.candidate_generation.* in config.json
Pretrained checkpoint locations blink.base_models.* in config.json (BLINK), res.pretrained_model (ReS)
GPUs ReS uses res.gpus in config.json (sets CUDA_VISIBLE_DEVICES)

The valid exp_list and negative_selection values are listed as comments at the top of scripts/retriever_fine_tuning.sh. Paths inside config.json are written relative to data_preparation/, i.e. "../datasets/foo/" means <repo>/datasets/foo/; the loader rewrites the leading ../ for scripts that run two levels deep. To try an alternative config without editing the committed one:

cd modeling/BLINK/
CONFIG_FILE=../../my_config.json bash scripts/retriever_fine_tuning.sh

Defaults as shipped: retriever — 1 epoch, batch 64, lr 2e-5, bi_enc_negative_selection=add_prch_in_pos_list (ontology parents/children of the gold entity count as extra positives, scored against the whole KB). Cross-encoder — 3 epochs, batch 16 × 2 accumulation, lr 2e-5, cross_enc_negative_selection=only_bienc_20_neg (gold + 20 negatives from the bi-encoder top-64). ReS — 3 epochs, batch 16, lr 1e-4, 21 training candidates, sum_log_nce loss, LoRA off. Seed 0 throughout. All of these are the config.json values, not code defaults.

Reference wall-clock on one A100-80GB for the 1,299-pair Sci-ZSEL set: alias generation ~1.5 h (CPU/GPU-bound on Ollama), retriever ~2.6 min, cross-encoder ~7.3 min.

Three things that will bite you

  1. blink.retriever.exp_list and blink.reranker.exp_list are separate keys. Keep them in sync unless you deliberately want the two stages trained on different pseudo-pair sets.
  2. The reranker's candidates come from the off-the-shelf biencoder_wiki_large.bin (blink.base_models.biencoder), not the retriever fine-tuned in 5.3 — the candidate-generation script passes --path_to_model $BIENCODER_BASE_MODEL. To couple the stages, point that flag at $SAVED_MODEL_DIR/biencoder/$SPLITNAME/$EXP/pytorch_model.bin in scripts/get_biencoder_cands_for_reranker_training.sh.
  3. *_entity_pool.t7 / *_entity_encodings.t7 are reused blindly if present. Delete them whenever you change the ontology, or you will score against stale encodings with no warning.

5. Where the results are

Stage File
Retriever saved_models/<world>/biencoder/train/<exp>/epoch_<i>/top64_candidates/test_eval.txt
BLINK reranker saved_models/<world>/crossencoder/train/fine-tune/seed-<s>/<exp>/epoch_<i>/crossencoder_predictions_eval.txt
ReS reranker saved_models/<world>/res/<world>/train/seed-<s>/<exp>/<world>_<exp>/epoch_<i>/pred_eval.txt

Settings and results maps each of the four training settings to these paths, with a worked example.

Mind the off-by-one in epoch_<i>: the BLINK retriever and reranker number their epoch directories from epoch_0, while ReS numbers from epoch_1 — so after a 3-epoch run the last BLINK directory is epoch_2 but the last ReS one is epoch_3.

crossencoder_predictions_eval.txt holds two reports. The first, under the Bi-Encoder heading, is the off-the-shelf (non-fine-tuned) bi-encoder that generated the candidates the reranker was fed — it is a retrieval baseline, not a reranker score. Read the numbers under the Cross-Encoder heading for the actual BLINK reranker result.

For a split with no labels — the QTL train corpora — there is nothing to score, so the matching *_eval.txt holds No ground truth in the "train" split (N mentions) - evaluation skipped. instead of metrics. The candidates themselves are still written, and the annotated test split of the same corpus is scored as usual.

Each reports MRR and recall@{1,5,…,64} overall and per lexical-overlap category, e.g.:

Calculated MRR for 960 items.
So, MRR : 0.6353535862842499

Found 511 items out of 960 within top-1 candidates
So, recall@1 : 511/960=0.5322916666666667

HO,   count: 277, matched: 190, score: 0.6859205776173285
MINT, count: 29,  matched: 17,  score: 0.5862068965517241
LO,   count: 435, matched: 248, score: 0.5701149425287356
NO,   count: 219, matched: 56,  score: 0.2557077625570776

Multiple gold ids and altdiseaseid aliases are all accepted as correct (category_eval.py::MultiGTEvaluation).

modeling/BLINK/performence.py and modeling/BLINK/get_report_after_all_exp.py are optional analysis helpers for cross-experiment error comparison; they are not part of the main path and expect you to fill in the experiment/best-epoch dictionaries by hand.


6. Attribution and license

This repository vendors and modifies two upstream code bases:

  • BLINK — facebookresearch/BLINK, MIT license (Wu et al., Scalable Zero-shot Entity Linking with Dense Entity Retrieval, EMNLP 2020). Lives in modeling/BLINK/; the bi-encoder and cross-encoder training loops carry Sci-ZSEL modifications (ontology-aware negative/positive selection, multi-ground-truth and category-wise evaluation).
  • ReS — HITsz-TMG/Read-and-Select (Xu et al., A Read-and-Select Framework for Zero-shot Entity Linking, Findings of EMNLP 2023). Lives in modeling/ReS/.

Also used: FremyCompany/BioLORD-2023-M for the ontology-aware filter, and Llama 3.2 3B Instruct via Ollama for alias generation — check each model's own license before redistribution.

Ontologies (MEDIC, MeSH, CMO, VT, LPT) are redistributed under their respective terms; consult the issuing organizations for reuse conditions.

No LICENSE file has been added to this repository yet, We will add one before release so downstream users know the terms of the Sci-ZSEL code and the animal science benchmark.


7. Citation

The paper is currently an arXiv preprint. This is the entry Google Scholar generates:

@article{khondokar2026bridging,
  title={Bridging Lexical Divergence: LLM-Assisted, Cost-Efficient, Zero-shot Scientific Entity Linking},
  author={Khondokar, Md Rasel and Qiao, Qiao and Samia, Farjana Sultana and Le, Nhat and Li, Yuepei and Li, Qi},
  journal={arXiv preprint arXiv:2609.00228},
  year={2026}
}

About

No description, website, or topics provided.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages