Bridging Lexical Divergence: LLM-Assisted, Cost-Efficient, Zero-shot Scientific Entity Linking| Findings of EMNLP 2026 | arXiv:2609.00228
This repository reproduces the full pipeline end to end. The worked example throughout is
NCBI-Disease → MEDIC, which is what config.json ships configured for.
- 1. Repository layout
- 2. Installation
- 3.
config.jsonreference - 4. Running the pipeline
- 5. Where the results are
- 6. Attribution and license
- 7. Citation
Details live in docs/:
| Doc | What's in it |
|---|---|
| Requirements | Hardware, the two conda environments, exact versions, Ollama |
| Machine-specific paths | Read this first — every hardcoded path you must repoint |
| Pretrained checkpoints | Download commands for the three model files |
| Data | What ships, corpus/ontology formats, generated intermediates |
| Experiment names | Decoding (m1_e1)U(m3_e1)_multi_prime_rm_sm_e... and the HO/MINT/LO/NO categories |
| Settings and results | The four training settings and the exact result file for each stage |
| LLM cost accounting | Tokens and wall clock spent on alias generation |
| Troubleshooting | Known quirks and how they fail |
Sci-ZSEL-release/
├── config.json # single source of truth: which corpus/ontology to run
├── install.sh # creates the two conda envs
├── data_preparation/ # §5.1–5.2 of the paper: entity selection + pseudo-pair construction
│ ├── convert_grag_to_blink_format.py # raw corpus -> BLINK format
│ ├── grag_to_blink.sh
│ ├── entity_selection.py # E_EM (exact match) and E_BT (bi-encoder top-1) entity sets
│ ├── alias_generation.py # LLM alias generation (via a local Ollama server)
│ ├── alias_generation.sh # starts/stops Ollama around alias_generation.py
│ ├── run_ollama_to_serve_llm # the `ollama serve` invocation (env vars live here)
│ ├── pseudo_pair_construction.py # alias matching, ontology-aware filter, dataset merging
│ ├── entity_selection_and_pseudo_pair_construction.sh
│ ├── prompts/<corpus>/{system,human}.txt # per-domain alias-generation prompts
│ └── utils.py # format conversion, overlap categories, ontology relations
├── modeling/
│ ├── BLINK/ # vendored + modified facebookresearch/BLINK
│ │ ├── blink/biencoder/ # retriever (train_biencoder.py, eval_biencoder.py)
│ │ ├── blink/crossencoder/ # reranker (train_cross.py)
│ │ ├── category_eval.py # HO/MINT/LO/NO evaluation, recall@k, MRR
│ │ └── scripts/ # entry points + load_config.sh (config.json -> shell vars)
│ └── ReS/ # vendored + modified HITsz-TMG/Read-and-Select reranker
│ ├── get_retriever_candidates.py # BLINK top-64 -> ReS input format
│ ├── train_and_eval.py # entry point (reads config.json -> "res")
│ ├── train_on_other_data.py # maps config.json -> ReS argparse namespace
│ └── run_disambiguation_attention.py
├── datasets/ # raw corpora + ontologies (tracked in git)
├── docs/ # the reference docs linked above
└── saved_models/ # checkpoints + all outputs (NOT tracked; you create it)
saved_models/,datasets/*/blink_format/,datasets/*/res_format/andlogs/are git-ignored. A fresh clone contains only the raw corpora and ontologies; everything else is regenerated by the steps below.
All five corpora sit side by side under datasets/: the two existing benchmarks, NCBI-Disease
(datasets/ncbi_disease/) and BC5CDR (datasets/bc5cdr/), plus the three new animal-science QTL
datasets released with this paper — CMO (datasets/cmo/), VT (datasets/vt/) and LPT
(datasets/lpt/). See Data for sizes, ontologies and formats.
Before you start: check Requirements (GPU, the two conda envs, Ollama), repoint the hardcoded paths listed in Machine-specific paths, and download the three model files in Pretrained checkpoints.
git clone https://github.com/rasel-isu/Sci-ZSEL.git
cd Sci-ZSEL
# edit install.sh first - see docs/machine-specific-paths.md
bash install.shinstall.sh creates both envs and installs modeling/BLINK/requirements.txt and
modeling/ReS/requirements.txt respectively.
Alias generation (step 5.2) talks to a local Ollama server, so the ollama binary must be
installed and on your PATH. Official instructions and downloads:
https://ollama.com/download (source: https://github.com/ollama/ollama).
On Linux with root:
curl -fsSL https://ollama.com/install.sh | shOn a cluster without root, unpack the release tarball into a directory you own and add it to
PATH:
mkdir -p ~/apps/ollama && cd ~/apps/ollama
curl -fsSL https://ollama.com/download/ollama-linux-amd64.tgz | tar -xz
export PATH="$HOME/apps/ollama/bin:$PATH" # add to ~/.bashrc to make it stick
ollama --versionThen pull the alias model once (~6.4 GB at fp16):
export OLLAMA_MODELS=/path/to/your/ollama/models # big; keep it off your home quota
export OLLAMA_HOST=127.0.0.1:11435
ollama serve &
ollama pull llama3.2:3b-instruct-fp16Point data_preparation/run_ollama_to_serve_llm at that same OLLAMA_MODELS directory (machine-specific paths). That
file is what the pipeline actually launches, so ollama must resolve on PATH when it runs — on a
cluster, load the module or export the PATH line above in your job script.
One file at the repo root selects the corpus and ontology for every step. It is read by relative
path (../config.json from data_preparation/, ../../config.json from modeling/BLINK/ and
modeling/ReS/), which is why each script must be run from its own directory.
{
"world": "ncbi_disease",
"kb_name": "medic",
"kb_file": "medic.json",
"data_dir": "../datasets/ncbi_disease/",
"saved_model_dir": "../saved_models/ncbi_disease/",
"synonym_key_on_ontology": "synonyms",
"has_ground_truth": true,
"has_ent_alt_id": true,
"exact_match_file": "m1_e1_unq_ents.json",
"biencoder_top1_file": "top_1_from_biencoder.json",
"blink": {
"split_name": "train",
"base_models": {
"biencoder": "../saved_models/biencoder_wiki_large.bin",
"crossencoder": "../saved_models/crossencoder_wiki_large.bin"
},
"candidate_generation": { "top_k": 64, "max_context_length": 64, "...": "..." },
"retriever": {
"exp_list": ["synonymU(m1_e1)U(m3_e1)_multi_prime_rm_sm_eU(m4_e2)_multi_prime_rm_sm_e"],
"negative_selection": ["add_prch_in_pos_list"],
"seeds": [0], "epochs": 1, "train_batch_size": 64, "learning_rate": "2e-05", "...": "..."
},
"reranker": {
"exp_list": ["synonymU(m1_e1)U(m3_e1)_multi_prime_rm_sm_eU(m4_e2)_multi_prime_rm_sm_e"],
"negative_selection": "only_bienc_20_neg",
"seeds": [0], "epochs": 3, "train_batch_size": 16, "learning_rate": "2e-05", "...": "..."
}
},
"res": {
"pretrained_model": "../../saved_models/zeshel_disambiguation_attention.pt",
"hf_model": "roberta-base",
"split_name": "train",
"exp_list": ["synonymU(m1_e1)U(m3_e1)_multi_prime_rm_sm_eU(m4_e2)_multi_prime_rm_sm_e"],
"seeds": [0], "lora": false, "use_title_during_testing": true,
"epochs": 3, "learning_rate": 1e-4, "batch_size": 16,
"cand_num_train": 21, "cand_num": 64, "type_loss": "sum_log_nce",
"max_len": 512, "gpus": "0,1,2,3", "...": "..."
}
}("..." above stands in for the remaining hyperparameters — see the real file, or
scripts/load_config.sh, which lists every key it reads together with its default.)
| Key | Meaning |
|---|---|
world |
Corpus/domain tag. Also selects data_preparation/prompts/<world>/ and the ontology-relation loader in pseudo_pair_construction.py. Must be one of ncbi_disease, bc5cdr, cmo, vt, lpt, cometa. |
kb_name |
Ontology tag used by ReS to locate the KB file (medic, mesh, cmo, vt, lpt). |
kb_file |
Ontology filename inside data_dir. |
data_dir |
Corpus directory, relative to data_preparation/. |
saved_model_dir |
Output root for this world, relative to data_preparation/. |
synonym_key_on_ontology |
Which ontology key holds curated synonyms (synonyms, or synonym for raw MeSH). |
has_ground_truth |
false for the QTL training corpora (cmo, vt, lpt), whose 16,385-mention train split ships with ground_truth: []. When false, pseudo-label quality reporting is skipped instead of crashing on empty ground_truth. Independently of this flag, biencoder_eval_report skips any split it finds no labels in and writes a short note to *_eval.txt in place of MRR/recall — that check is on the data, not on the config, so the labeled test split of an unlabeled corpus is still scored normally. |
has_ent_alt_id |
Whether the ontology has an altdiseaseid list to expand during matching/scoring. |
exact_match_file |
Filename for the E_EM entity set. |
biencoder_top1_file |
Filename for the E_BT entity set. |
blink.split_name |
Which split directory the BLINK scripts operate on (train). |
blink.base_models.* |
Pretrained BLINK checkpoints, relative to data_preparation/. |
blink.candidate_generation.* |
Flags for eval_biencoder.py (top-k, batch sizes, bert_model). has_gt defaults to the top-level has_ground_truth; when false the --has_gt flag is omitted entirely rather than passed as false — see Troubleshooting. It reaches only get_biencoder_top_k.sh, the step that reads the raw (possibly unlabeled) corpus; get_biencoder_cands_for_reranker_training.sh hard-codes --has_gt true, because both of its inputs — the annotated test set and the pseudo-labeled training set — always carry labels. |
blink.retriever.* |
Flags for train_biencoder.py. exp_list and negative_selection are arrays — every combination is run in turn. |
blink.reranker.* |
Flags for train_cross.py. exp_list and seeds are arrays. |
res.pretrained_model |
ReS checkpoint, relative to modeling/ReS/. |
res.hf_model |
ReS backbone (roberta-base). |
res.exp_list, res.seeds |
Arrays. exp_list also drives get_retriever_candidates.py, which always converts original_data (the test set) plus one training set per entry. |
res.lora, res.use_title_during_testing, res.split_name |
Run-level switches. |
res.* (remaining) |
Every ReS hyperparameter: epochs, learning_rate, batch_size, cand_num_train, cand_num, type_loss, max_len, max_ent_len, max_text_len, warmup_proportion, weight_decay, adam_epsilon, gradient_accumulation_steps, num_workers, clip, info_token_num, gpus, logging_steps, eval_step. |
The BLINK shell scripts read all of this through modeling/BLINK/scripts/load_config.sh and the
ReS modules read CONFIG['res'] directly, so neither world nor any hyperparameter is spelled out
in a script.
Everything below is for the shipped NCBI-Disease configuration. cd from the repo root each
time — every script reads config.json by a relative path and will fail from anywhere else.
bash install.sh
cd data_preparation
bash grag_to_blink.sh
cd modeling/BLINK/
bash scripts/get_biencoder_top_k.sh
cd data_preparation
bash entity_selection_and_pseudo_pair_construction.sh
cd modeling/BLINK/
bash scripts/retriever_fine_tuning.sh
cd modeling/BLINK/
bash scripts/blink_reranker_fine_tuning.sh
cd modeling/ReS
bash res_reranker_fine_tuning.sh
| Step | Reads | Writes |
|---|---|---|
grag_to_blink.sh |
*_grag.json, ontology |
blink_format/train/original_data/ — train.jsonl (retrieval pool), test.jsonl (eval set), kb.jsonl, id_map.json |
get_biencoder_top_k.sh |
original_data/train.jsonl |
top-64 lists from the off-the-shelf bi-encoder → saved_models/<world>/biencoder/train/original_data/top64_candidates/train.json; also builds the cached ontology encodings *_entity_pool.t7 / *_entity_encodings.t7 |
entity_selection_and_pseudo_pair_construction.sh |
the above | entity sets E_EM + E_BT → LLM aliases (Ollama) → ontology-aware filter → one blink_format/train/<exp>/train.jsonl per configuration (see experiment names) |
retriever_fine_tuning.sh |
blink_format/train/<exp>/ |
fine-tuned retriever + epoch_<i>/top64_candidates/test_eval.txt |
blink_reranker_fine_tuning.sh |
regenerates top-64, then trains | cross-encoder + epoch_<i>/crossencoder_predictions_eval.txt |
res_reranker_fine_tuning.sh |
res_format/train/<exp>/ |
ReS model + <epoch>/pred_eval.txt |
With one exception, neither the BLINK scripts nor the ReS entry points hold settings of their own
(the exception is --has_gt in get_biencoder_cands_for_reranker_training.sh, noted in §3). The BLINK scripts
source scripts/load_config.sh, which parses config.json with python3 and exports every path
and hyperparameter; the ReS modules read CONFIG['res'] directly. So one file drives all of it:
| To change | Edit |
|---|---|
| Corpus / ontology | config.json top level (world, kb_name, kb_file, data_dir, saved_model_dir) |
| Which pseudo-pair set to train on | blink.retriever.exp_list, blink.reranker.exp_list and res.exp_list in config.json |
| Seeds, epochs, batch size, lr, negative selection | blink.retriever.* / blink.reranker.* / res.* in config.json |
| Candidate-generation top-k, batch sizes | blink.candidate_generation.* in config.json |
| Pretrained checkpoint locations | blink.base_models.* in config.json (BLINK), res.pretrained_model (ReS) |
| GPUs ReS uses | res.gpus in config.json (sets CUDA_VISIBLE_DEVICES) |
The valid exp_list and negative_selection values are listed as comments at the top of
scripts/retriever_fine_tuning.sh. Paths inside config.json are written relative to
data_preparation/, i.e. "../datasets/foo/" means <repo>/datasets/foo/; the loader rewrites the
leading ../ for scripts that run two levels deep. To try an alternative config without editing the
committed one:
cd modeling/BLINK/
CONFIG_FILE=../../my_config.json bash scripts/retriever_fine_tuning.shDefaults as shipped: retriever — 1 epoch, batch 64, lr 2e-5, bi_enc_negative_selection=add_prch_in_pos_list
(ontology parents/children of the gold entity count as extra positives, scored against the whole KB).
Cross-encoder — 3 epochs, batch 16 × 2 accumulation, lr 2e-5,
cross_enc_negative_selection=only_bienc_20_neg (gold + 20 negatives from the bi-encoder top-64).
ReS — 3 epochs, batch 16, lr 1e-4, 21 training candidates, sum_log_nce loss, LoRA off. Seed 0 throughout.
All of these are the config.json values, not code defaults.
Reference wall-clock on one A100-80GB for the 1,299-pair Sci-ZSEL set: alias generation ~1.5 h (CPU/GPU-bound on Ollama), retriever ~2.6 min, cross-encoder ~7.3 min.
blink.retriever.exp_listandblink.reranker.exp_listare separate keys. Keep them in sync unless you deliberately want the two stages trained on different pseudo-pair sets.- The reranker's candidates come from the off-the-shelf
biencoder_wiki_large.bin(blink.base_models.biencoder), not the retriever fine-tuned in 5.3 — the candidate-generation script passes--path_to_model $BIENCODER_BASE_MODEL. To couple the stages, point that flag at$SAVED_MODEL_DIR/biencoder/$SPLITNAME/$EXP/pytorch_model.bininscripts/get_biencoder_cands_for_reranker_training.sh. *_entity_pool.t7/*_entity_encodings.t7are reused blindly if present. Delete them whenever you change the ontology, or you will score against stale encodings with no warning.
| Stage | File |
|---|---|
| Retriever | saved_models/<world>/biencoder/train/<exp>/epoch_<i>/top64_candidates/test_eval.txt |
| BLINK reranker | saved_models/<world>/crossencoder/train/fine-tune/seed-<s>/<exp>/epoch_<i>/crossencoder_predictions_eval.txt |
| ReS reranker | saved_models/<world>/res/<world>/train/seed-<s>/<exp>/<world>_<exp>/epoch_<i>/pred_eval.txt |
Settings and results maps each of the four training settings to these paths, with a worked example.
Mind the off-by-one in epoch_<i>: the BLINK retriever and reranker number their epoch directories from epoch_0, while ReS numbers from epoch_1 — so after a 3-epoch run the last BLINK directory is epoch_2 but the last ReS one is epoch_3.
crossencoder_predictions_eval.txt holds two reports. The first, under the Bi-Encoder
heading, is the off-the-shelf (non-fine-tuned) bi-encoder that generated the candidates the
reranker was fed — it is a retrieval baseline, not a reranker score. Read the numbers under the
Cross-Encoder heading for the actual BLINK reranker result.
For a split with no labels — the QTL train corpora — there is nothing to score, so the matching
*_eval.txt holds No ground truth in the "train" split (N mentions) - evaluation skipped.
instead of metrics. The candidates themselves are still written, and the annotated test split of
the same corpus is scored as usual.
Each reports MRR and recall@{1,5,…,64} overall and per lexical-overlap category, e.g.:
Calculated MRR for 960 items.
So, MRR : 0.6353535862842499
Found 511 items out of 960 within top-1 candidates
So, recall@1 : 511/960=0.5322916666666667
HO, count: 277, matched: 190, score: 0.6859205776173285
MINT, count: 29, matched: 17, score: 0.5862068965517241
LO, count: 435, matched: 248, score: 0.5701149425287356
NO, count: 219, matched: 56, score: 0.2557077625570776
Multiple gold ids and altdiseaseid aliases are all accepted as correct
(category_eval.py::MultiGTEvaluation).
modeling/BLINK/performence.py and modeling/BLINK/get_report_after_all_exp.py are optional
analysis helpers for cross-experiment error comparison; they are not part of the main path and
expect you to fill in the experiment/best-epoch dictionaries by hand.
This repository vendors and modifies two upstream code bases:
- BLINK — facebookresearch/BLINK, MIT license
(Wu et al., Scalable Zero-shot Entity Linking with Dense Entity Retrieval, EMNLP 2020). Lives in
modeling/BLINK/; the bi-encoder and cross-encoder training loops carry Sci-ZSEL modifications (ontology-aware negative/positive selection, multi-ground-truth and category-wise evaluation). - ReS — HITsz-TMG/Read-and-Select
(Xu et al., A Read-and-Select Framework for Zero-shot Entity Linking, Findings of EMNLP 2023).
Lives in
modeling/ReS/.
Also used: FremyCompany/BioLORD-2023-M for the ontology-aware filter, and Llama 3.2 3B Instruct
via Ollama for alias generation — check each model's own license before redistribution.
Ontologies (MEDIC, MeSH, CMO, VT, LPT) are redistributed under their respective terms; consult the issuing organizations for reuse conditions.
No LICENSE file has been added to this repository yet, We will add one before release so downstream users
know the terms of the Sci-ZSEL code and the animal science benchmark.
The paper is currently an arXiv preprint. This is the entry Google Scholar generates:
@article{khondokar2026bridging,
title={Bridging Lexical Divergence: LLM-Assisted, Cost-Efficient, Zero-shot Scientific Entity Linking},
author={Khondokar, Md Rasel and Qiao, Qiao and Samia, Farjana Sultana and Le, Nhat and Li, Yuepei and Li, Qi},
journal={arXiv preprint arXiv:2609.00228},
year={2026}
}