Do General NLP Embeddings Capture Ontological Reasoning?
Abstract.
General-purpose NLP embedding models perform well on linguistic tasks, but their ability to capture symbolic ontological structure remains unclear. We introduce AVA, a systematic framework for evaluating whether embeddings distinguish logic-sensitive relational semantics in ontologies and knowledge graphs. AVA comprises 171,007 contrastive triplets derived from 163 heterogeneous ontologies using hierarchy inversion, relation substitution, and disjointness injection. Each triplet contains an ontology statement, a semantically equivalent paraphrase, and a logic-sensitive hard negative with contradictory relational meaning. We evaluate more than 25 state-of-the-art embedding models and find substantial limitations: the best model achieves only 0.739 triplet accuracy, while hard negative accuracy falls to 0.135. Fine-tuning improves discrimination by a large margin but transfers poorly to downstream Semantic Web tasks, including taxonomy discovery and ontology alignment. Further analysis suggests that improvements stem partly from perturbation-specific pattern recognition rather than robust ontological understanding. These findings reveal a persistent gap between linguistic representation learning and ontology-level discrimination, challenging the assumption that strong NLP benchmark performance translates to Semantic Web competence.
Keywords:
Embedding, Large Language Models, Ontology, Semantic Textual Similarity, Ontology Engineering1. Introduction
Recent advances in representation learning have transformed both Natural Language Processing (NLP) and the Semantic Web. General-purpose embedding models, ranging from Word2Vec (Mikolov et al., 2013) and GloVe (Pennington et al., 2014) to transformer-based architectures such as BERT (Devlin et al., 2019), MPNet (Song et al., 2020), E5 (Wang et al., 2022), GTE (Zhang et al., 2024; Li et al., 2023), and BGE (Xiao et al., 2024), achieve strong performance on retrieval and Semantic Textual Similarity (STS) benchmarks. In parallel, the Semantic Web community has developed ontology and knowledge graph embedding approaches, including RDF2Vec (Ristoski and Paulheim, 2016), OWL2Vec* (Chen et al., 2021b), DL2Vec (Chen et al., 2021a), and knowledge graph embedding (KGE) methods such as TransE (Bordes et al., 2013), ComplEx (Trouillon et al., 2016), and RotatE (Sun et al., 2019). More recent approaches, including KG-BERT (Yao et al., 2019) and KEPLER (Wang et al., 2021), attempt to combine linguistic representations with structured knowledge. Despite these advances, a fundamental challenge remains unresolved: generalization in Semantic Web representation learning. Traditional KGE methods operate on fixed entity and relation vocabularies and often require retraining when applied to new ontologies (Chen et al., 2023; Liang et al., 2023). Ontology-aware methods capture rich structural semantics within a particular ontology but may exhibit limited transferability across heterogeneous schemas (Qiang, 2023; Bian, 2025; Li et al., 2025). Conversely, transformer-based sentence embeddings generalize well across linguistic tasks yet are not explicitly optimized for symbolic reasoning over OWL/RDFS structures. Hyperbolic representations provide a promising alternative for hierarchical knowledge (Nickel and Kiela, 2017; Dhingra et al., 2018), but their effectiveness for transferable ontology relation discrimination remains unclear.
This challenge is closely related to STS, a primary evaluation paradigm for sentence embeddings. Modern embedding models achieve strong STS performance (Kumar et al., 2025), suggesting they can capture semantic equivalence between textual expressions. However, standard STS benchmarks rarely require distinguishing between statements that are lexically similar yet differ in ontology-level relational semantics (Gatto et al., 2023; Sun et al., 2025). As a result, it remains unclear whether strong STS performance reflects genuine sensitivity to subclass relations, domain/range constraints, disjointness axioms, and other forms of symbolic knowledge. To investigate this question, we introduce AVA, a logic-sensitive ontology similarity benchmark based on structured ontology perturbations. AVA generates ontology-aware hard negatives through hierarchy inversion, relation substitution, and disjointness injection, producing sentence pairs that preserve substantial lexical overlap while expressing contradictory ontological meaning. Using 163 heterogeneous ontologies, we construct a dataset of 171,007 contrastive triplets and evaluate both pre-trained and fine-tuned embedding models under cross-ontology generalization settings.
Our experiments reveal three key findings. First, even the strongest general-purpose embeddings achieve only moderate performance on ontology-sensitive similarity judgments, with substantial degradation on hard negatives. Second, contrastive fine-tuning dramatically improves triplet discrimination, with hyperbolic objectives achieving near-perfect ranking accuracy. Third, these improvements transfer only weakly to downstream ontology engineering tasks such as taxonomy discovery (Babaei Giglou et al., 2023) and ontology alignment (Hertling and Paulheim, 2023). Together, these results reveal an optimization–generalization gap: embeddings can learn to discriminate AVA perturbations without acquiring robust, transferable representations of ontological structure and pattern-specific discrimination rather than general ontological reasoning.
The contributions of this work are threefold: (1) we introduce AVA, a large-scale benchmark for evaluating ontology-aware semantic similarity using structured logic-sensitive perturbations; (2) we provide a comprehensive evaluation of modern embedding models and contrastive learning objectives, including Euclidean and hyperbolic formulations; and (3) we demonstrate that high contrastive discrimination accuracy does not necessarily imply transferable ontology understanding, highlighting important limitations of current embedding-based approaches for Semantic Web applications. Furthermore, we make the implementation publicly available to the research community at https://github.com/sciknoworg/AVA.
| Anchor: An online gaming account is a subclass of an online account. |
| Hard Negative: An online gaming account is a subclass of an agent. |
| Anchor: Online chat accounts are defined as a subclass of online accounts. |
| Hard Negative: Online chat accounts are defined as a subclass of online e-commerce accounts. |
| Anchor: An online gaming account is a specific type of online accounts. |
| Hard Negative: An online gaming account is a specific type of online chat account. |
| Count | |
|---|---|
| Ontologies | 163 |
| BFS subgraphs | 50,548 |
| Synthesized samples (raw) | 197,326 |
| Removed (Positive-Negative pairs that are too similar) | 4,154 |
| After cleaning & de-duplication | 171,007 |
| Hard negatives (A-N similarity ) | 77,932 |
| Mean sentence length (words) | 10 |
| Train (Hard Negatives) | 153,211 (70,703) |
| Test (Hard Negatives) | 17,796 (7,229) |
2. AVA
AVA is an evaluation framework for assessing whether embedding models capture ontology-level relational semantics beyond surface lexical similarity. It combines (1) ontology perturbations that generate logic-sensitive contrastive triplets and (2) contrastive objectives that evaluate ontology-aware discrimination and cross-ontology transfer using triplet loss, hyperbolic loss, and reinforcement learning techniques (Christiano et al., 2017; Lambert et al., 2022).
2.1. Structured Ontology Perturbations
Ontology Graph Extraction. We used 163 ontologies from diverse domains, including biomedical (i.e., GO (Consortium, 2026), OBI (Bandrowski et al., 2016)), geospatial, social, engineering, and schema ontologies. We accessed these collections via the OntoLearner library (Giglou et al., 2026). For each ontology, we construct an undirected graph in which nodes represent OWL classes and object/datatype/annotation properties, and edges encode structural OWL/RDFS relations (i.e., rdfs:subClassOf, owl:domain, owl:range, etc). Existential restriction axioms of the form are reified as direct labeled edges between and . To construct plausible hard negatives—such as sibling swaps—without exceeding the LLM context window, we extract subgraphs using a two-hop breadth-first search (BFS) seeded at every OWL class. Single-hop neighborhoods frequently lack sibling classes, whereas a two-hop radius reliably captures the necessary sibling and grandparent relationships. We retain subgraphs bounded between 3 and 20 nodes to exclude trivial or unwieldy structures, and deduplicate them by the MD5 fingerprint of their sorted node sets, resulting in 50,548 unique subgraphs.
Logic-Sensitive Hard Negative Synthesis. Each subgraph is converted into a structured natural-language prompt presenting its entities with labels and definitions alongside their RDF triples. We instruct a Qwen3.5-35B-A3B LLM (Qwen Team, 2026) to generate five contrastive triplets per subgraph under the following constraints (see prompt at Figure 3): 1) Anchor, a declarative sentence expressing a fact directly asserted in the subgraph (hierarchy assertion, domain/range constraint, equivalence, or disjointness); 2) Positive, a semantically equivalent paraphrase using morphological variation, synonymy, or syntactic restructuring (e.g., active to passive voice). The factual content must be preserved exactly; 3) Hard negative, a statement exhibiting high lexical overlap with the anchor but asserting an ontologically incorrect relationship, for instance, swapping a subclass for a sibling, inverting a property domain, or replacing an entity with a disjoint concept. This synthesis procedure is specifically designed to produce logic-sensitive negatives rather than randomly sampled distractors. After post-processing the raw outputs, we obtained a total of 197,326 candidate (Anchor, Positive, Negative) triplets. Representative examples are shown in Table 1, and Figure 2 shows how, using BFS subgraphs, triplets are generated.
Post-Processing and Dataset Statistics. We apply two filtering passes using token-set ratio similarity scores computed with RapidFuzz. First, we flag samples where the anchor–negative similarity exceeds 90 as hard negatives and move them to a dedicated evaluation split, as they represent the most challenging cases. Second, samples where the positive–negative similarity exceeds 90 are discarded outright, as the two cannot be meaningfully distinguished (w.r.t Figure 1, the A-P distributions). After de-duplication at the anchor level, the dataset contains 171,007 samples. We perform the train/test split at the ontology level: all samples from a given ontology are assigned exclusively to train or test, with the test set constructed from ontologies whose label sets share fewer than 100 labels with the training ontologies. Source and subgraph disjointness is verified programmatically to ensure that there is no leakage between test and train sets, except for a small number of subgraphs (less than 0.1% of triplets, which is 148 triplets) rooted in upper-ontology classes shared across multiple imported ontologies. Dataset statistics are summarized in Table 2. General NLP embeddings (specifically MPNET-base) cannot separate anchor-positive from anchor-negatives in semantic space, yet hard negatives are deliberately more lexically similar to the anchor than the positives (see Figure 1). A model that learns to discriminate these triplets, therefore, cannot rely on lexical overlap and must learn something about relational structure.
2.2. Contrastive Learning
Triplet Loss. We fine-tune using the standard cosine triplet loss, which optimizes the margin between anchor-positive and anchor-negative similarity scores. Given embeddings , , and , the loss is , where is a margin hyperparameter set to .
Hyperbolic Triplet Loss. Ontological class hierarchies are inherently tree-structured, which Euclidean space represents poorly. We replace the Euclidean margin with a Poincaré-ball distance (Nickel and Kiela, 2017), computing , where is the hyperbolic distance at curvature . This encourages the model to reflect hierarchical structure in its embedding geometry.
Embedding-Adapted DPO. We adapt Direct Preference Optimization (Rafailov et al., 2023) to the embedding setting by treating cosine similarity as an implicit reward. A frozen reference encoder acts as a regularizer, and the loss penalizes the policy encoder whenever its similarity margin falls behind the reference margin , with temperature .
| Model | Cross-Ontology | ||
|---|---|---|---|
| Triplet | R@1 | Hard Neg | |
| MiniLM-L6 | 0.657 | 0.657 | 0.467 |
| MPNET-base | 0.636 | 0.636 | 0.427 |
| RoBERTa-large | 0.621 | 0.621 | 0.412 |
| Nomic-embed | 0.602 | 0.602 | 0.363 |
| Nomic-embed-MoE | 0.584 | 0.584 | 0.333 |
| E5-small | 0.553 | 0.553 | 0.354 |
| E5-base | 0.506 | 0.506 | 0.305 |
| E5-large | 0.542 | 0.542 | 0.338 |
| Multilingual-E5-large | 0.633 | 0.633 | 0.394 |
| GTE-small | 0.650 | 0.650 | 0.433 |
| GTE-base | 0.632 | 0.632 | 0.402 |
| GTE-large | 0.647 | 0.647 | 0.417 |
| BGE-small | 0.609 | 0.609 | 0.370 |
| BGE-base | 0.604 | 0.604 | 0.356 |
| BGE-large | 0.558 | 0.558 | 0.302 |
| Qwen3-Embedding-0.6B | 0.739 | 0.739 | 0.572 |
| Qwen3-Embedding-4B | 0.709 | 0.709 | 0.523 |
| Qwen3-Embedding-8B | 0.736 | 0.736 | 0.564 |
| Text-Embedding-3-small | 0.442 | 0.442 | 0.201 |
| Text-Embedding-3-large | 0.388 | 0.388 | 0.155 |
| Text-Embedding-ADA-002 | 0.499 | 0.499 | 0.238 |
| EmbeddingGemma-300m | 0.455 | 0.455 | 0.203 |
| Llama-Embed-Nemotron-8B | 0.217 | 0.217 | 0.135 |
| Linq-Embed-Mistral | 0.634 | 0.634 | 0.411 |
| MiniLM-L6 + Triplet Loss | 0.975 | 0.975 | 0.955 |
| MiniLM-L6 + Hyperbolic Loss | 0.981 | 0.981 | 0.966 |
| MiniLM-L6 + DPO Loss | 0.913 | 0.913 | 0.885 |
| MPNET-base + Triplet Loss | 0.985 | 0.985 | 0.972 |
| MPNET-base + Hyperbolic Loss | 0.989 | 0.989 | 0.980 |
| MPNET-base + DPO Loss | 0.876 | 0.876 | 0.874 |
| Model | GO | SchemaOrg | SWEET | OBI | PO | ||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Rec@1 | Rec@5 | Rec@10 | Rec@1 | Rec@5 | Rec@10 | Rec@1 | Rec@5 | Rec@10 | Rec@1 | Rec@5 | Rec@10 | Rec@1 | Rec@5 | Rec@10 | |
| MiniLM-L6 | 0.058 | 0.150 | 0.183 | 0.084 | 0.251 | 0.337 | 0.037 | 0.092 | 0.118 | 0.075 | 0.179 | 0.229 | 0.122 | 0.281 | 0.327 |
| + Triplet loss | 0.037 | 0.090 | 0.116 | 0.078 | 0.225 | 0.307 | 0.027 | 0.071 | 0.092 | 0.059 | 0.143 | 0.180 | 0.116 | 0.250 | 0.299 |
| + Hyperbolic loss | 0.049 | 0.135 | 0.176 | 0.088 | 0.248 | 0.349 | 0.032 | 0.081 | 0.105 | 0.075 | 0.171 | 0.216 | 0.133 | 0.286 | 0.333 |
| + DPO loss | 0.013 | 0.027 | 0.033 | 0.025 | 0.062 | 0.090 | 0.007 | 0.017 | 0.021 | 0.021 | 0.044 | 0.053 | 0.025 | 0.051 | 0.062 |
| MPNET-base | 0.059 | 0.148 | 0.183 | 0.093 | 0.256 | 0.340 | 0.033 | 0.088 | 0.115 | 0.076 | 0.177 | 0.229 | 0.105 | 0.268 | 0.327 |
| + Triplet Loss | 0.036 | 0.088 | 0.115 | 0.092 | 0.223 | 0.312 | 0.029 | 0.077 | 0.102 | 0.056 | 0.140 | 0.185 | 0.080 | 0.209 | 0.267 |
| + Hyperbolic Loss | 0.051 | 0.128 | 0.166 | 0.100 | 0.262 | 0.358 | 0.036 | 0.091 | 0.117 | 0.082 | 0.181 | 0.231 | 0.126 | 0.275 | 0.329 |
| + DPO Loss | 0.007 | 0.013 | 0.016 | 0.015 | 0.042 | 0.058 | 0.004 | 0.008 | 0.010 | 0.010 | 0.021 | 0.028 | 0.017 | 0.037 | 0.048 |
| Model | ENVO-SWEET (Karam et al., 2020) | Mouse-Human (Dragisic et al., 2017) | MaterialInformation-MatOnto (Nas and Huschka, 2023) | Yago-Wikidata (Fallatah et al., 2020) | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Rec@1 | Rec@5 | Rec@10 | Rec@1 | Rec@5 | Rec@10 | Rec@1 | Rec@5 | Rec@10 | Rec@1 | Rec@5 | Rec@10 | |
| MiniLM-L6 | 0.604 | 0.775 | 0.819 | 0.852 | 0.934 | 0.950 | 0.325 | 0.536 | 0.609 | 0.905 | 0.964 | 0.974 |
| + Triplet loss | 0.598 | 0.785 | 0.830 | 0.843 | 0.916 | 0.926 | 0.328 | 0.623 | 0.669 | 0.891 | 0.951 | 0.980 |
| + Hyperbolic loss | 0.605 | 0.796 | 0.837 | 0.859 | 0.933 | 0.951 | 0.341 | 0.609 | 0.662 | 0.901 | 0.951 | 0.974 |
| + DPO loss | 0.472 | 0.519 | 0.524 | 0.685 | 0.715 | 0.724 | 0.119 | 0.162 | 0.166 | 0.549 | 0.609 | 0.638 |
| MPNET-base | 0.611 | 0.773 | 0.799 | 0.856 | 0.931 | 0.946 | 0.315 | 0.576 | 0.652 | 0.898 | 0.974 | 0.980 |
| + Triplet Loss | 0.622 | 0.789 | 0.820 | 0.838 | 0.907 | 0.925 | 0.358 | 0.609 | 0.725 | 0.905 | 0.964 | 0.974 |
| + Hyperbolic Loss | 0.626 | 0.805 | 0.834 | 0.871 | 0.937 | 0.954 | 0.354 | 0.550 | 0.662 | 0.931 | 0.967 | 0.984 |
| + DPO Loss | 0.468 | 0.538 | 0.554 | 0.662 | 0.697 | 0.707 | 0.156 | 0.219 | 0.252 | 0.530 | 0.586 | 0.641 |
3. Results
Evaluation is performed under a cross-ontology setting where evaluation ontological samples are fully excluded from training. We use L2-normalized embeddings and cosine similarity . Performance is measured via triplet accuracy, defined as the proportion of cases where (ties counted as incorrect). We also report Hard Negative Accuracy on samples where anchor–negative lexical similarity exceeds 90 (token-set similarity), measuring how often the model correctly ranks the positive above the hard negative ().
Performance of Pre-trained Embeddings. The Table 3 summarizes the performance of more than 25 pre-trained embedding models. Overall, results reveal substantial limitations in current general-purpose embeddings when confronted with ontology-sensitive semantic distinctions. Among all evaluated models, Qwen3-Embedding-0.6B achieves the strongest performance, reaching a triplet accuracy of and a hard negative accuracy of . Larger variants of the same family show comparable results, suggesting that model scale alone does not guarantee improved ontological discrimination. Sentence embedding models such as MiniLM-L6, MPNET-base, GTE, and multilingual-E5 achieve moderate performance, with triplet accuracies generally between and . In contrast, several embedding models perform substantially worse; specifically, the OpenAI Text-Embedding-3-large reaches only triplet accuracy, while Llama-Embed-Nemotron-8B obtains triplet accuracy and only 0.135 hard negative accuracy.
A consistent trend across all models is the large performance drop on hard negatives. Although some embeddings achieve reasonable overall triplet accuracy, they still struggle to distinguish ontology-consistent statements from highly similar contradictory statements. This suggests that many embeddings rely primarily on lexical and distributional similarity rather than explicit sensitivity to relational semantics such as subclass structure, domain/range constraints, or disjointness relations. These findings indicate that strong performance on standard semantic similarity or retrieval benchmarks does not necessarily translate into competence on ontology-level understanding or discrimination.
Effect of Contrastive Fine-Tuning. Fine-tuning dramatically improves triplet discrimination performance. Across both MiniLM-L6 and MPNET-base (a widely used model in semantic web engineering tasks), all three optimization objectives substantially outperform their corresponding pre-trained baselines. For MPNET-base, standard triplet loss increases triplet accuracy from to , while hyperbolic triplet loss further improves performance to . Hard negative accuracy exhibits a similar pattern, increasing from to and , respectively. Comparable improvements are observed for MiniLM-L6, where hyperbolic loss achieves the strongest overall results with triplet accuracy and hard negative accuracy. The hyperbolic objective consistently outperforms Euclidean triplet loss by a small but measurable margin. This result is consistent with prior works suggesting that hyperbolic geometry provides a more natural representation space for hierarchical structures (Nickel and Kiela, 2017; Dhingra et al., 2018). In contrast, the embedding-adapted DPO objective performs substantially worse than both triplet-based approaches, although it still improves considerably over the corresponding pre-trained models. This suggests the primary bottleneck may be reliance on a reference module whose representations do not encode fine-grained ontological distinctions. Consequently, preference optimization is likely constrained by the semantic limitations of the reference model, reducing its ability to learn transferable ontology-level representations.
Taken together, these results demonstrate that ontology-aware contrastive supervision enables embeddings to separate semantically valid statements from logic-sensitive negatives with near-perfect accuracy. A similar scenario is also observed in KEPLER (Wang et al., 2021), where the fine-tuned model on the generated dataset performed very well. However, high performance alone might not establish that models have acquired transferable ontological understanding. Instead, the models likely learn to recognize perturbation-specific structural patterns rather than genuine symbolic semantics.
Transfer to Downstream Ontology Engineering Tasks. Despite the dramatic gains, downstream evaluation reveals a substantial optimization–generalization gap. The Table 4 reports taxonomy discovery task performance, where the aim is that, for a given set of ontological classes, the models should construct taxonomic pairs (parent, child) that form a subclass (or is-a) relation (experiments were performed with the OntoLearner (Giglou et al., 2026) library). Contrary to expectations, fine-tuning often degrades performance relative to the original models. Standard triplet loss consistently reduces recall across most ontology benchmarks. Hyperbolic loss partially mitigates this decline and occasionally yields modest improvements, particularly on SchemaOrg, OBI, SWEET, and Plant Ontology (PO), but gains remain relatively small compared to the near-perfect improvements on earlier stages. DPO fine-tuning produces severe degradation across all datasets. A similar pattern appears in ontology alignment (see Table 5), the task of finding equivalent classes between two different ontologies (experiments were conducted using the OntoAligner (Babaei Giglou et al., 2025) library). Triplet and hyperbolic losses provide only modest improvements on several benchmarks, including ENVO–SWEET and MaterialInformation–MatOnto, while performance on other datasets remains largely unchanged. Hyperbolic loss achieves the strongest overall transfer performance, but improvements are typically measured in only a few percentage points. In contrast, DPO again substantially reduces alignment accuracy across all evaluation datasets.
The discrepancy between near-perfect triplet ranking accuracy and comparatively small downstream gains suggests that optimization does not necessarily induce robust ontology-level understanding. Instead, models may learn perturbation-specific decision boundaries that effectively distinguish the generated negatives without acquiring transferable representations of hierarchical and logical structure. These findings indicate that contrastive discrimination and ontology generalization should be treated as distinct evaluation objectives rather than interchangeable measures of semantic understanding.
4. Discussion
Because AVA’s triplets are synthesized by a single LLM (Qwen3.5-35B-A3B) and the strongest pretrained encoder in our evaluation belongs to the same model family (Qwen3-Embedding), we cannot rule out that part of its advantage reflects shared stylistic or distributional artifacts between generator and encoder rather than genuinely superior ontological sensitivity. We treat this as a limitation of single-generator benchmark construction; incorporating multiple LLMs for triplet synthesis would increase lexical and stylistic diversity and further mitigate generator-specific artifacts, and we leave this, together with broader cross-generator validation, to future work.
We also note that, as commonly observed for retrieval modules in RAG pipelines, taxonomy discovery performance degrades when candidates exhibit high semantic overlap. Analysis in (Giglou et al., 2026) indicates substantial semantic overlap among candidate classes across the ontologies used here, suggesting that the persistently low absolute recall observed for all models partly reflects this intrinsic task difficulty, which ontology-aware fine-tuning only partially mitigates.
Furthermore, the AVA specifically evaluates ontology discrimination rather than logical reasoning. It measures whether a model can distinguish between two lexically similar statements according to their ontological consistency, but it does not assess inference or entailment. Consequently, performance on the AVA should not be interpreted as evidence of logical reasoning ability. A separate evaluation would be required to test reasoning, for example, by assessing whether a model can infer from and when is not explicitly asserted. Such a setting would directly evaluate the model’s ability to derive new knowledge from formal ontology axioms rather than merely discriminate between ontology-consistent and ontology-violating statements.
5. Conclusion
In conclusion, while embedding-based models can perform well on similarity and retrieval tasks, this does not necessarily imply that they can reliably discriminate between ontology-consistent and ontology-violating statements. In ontology engineering, such models may capture statistical patterns in the data while failing to reflect formal logical constraints. It therefore remains unclear whether embedding optimization alone is sufficient for reliable ontology-level discrimination, or whether explicit logical mechanisms are needed to adequately capture such constraints.
Acknowledgements.
This work is supported by the NFDI4DataScience initiative (DFG, German Research Foundation, Grant ID: 460234259).GenAI Disclosure
In preparing this manuscript, generative AI tools, specifically ChatGPT and Gemini, were used solely for grammar checking, spelling checks, and readability of some sentences. All suggested changes were carefully reviewed and adapted by the authors to ensure accuracy and appropriateness. The scientific content, research design, analysis, and conclusions were developed and verified exclusively by the authors without AI involvement.
References
- Babaei Giglou et al. (2023) Hamed Babaei Giglou, Jennifer D’Souza, and Sören Auer. 2023. LLMs4OL: Large language models for ontology learning. In International semantic web conference. Springer, 408–427.
- Babaei Giglou et al. (2025) Hamed Babaei Giglou, Jennifer D’Souza, Oliver Karras, and Sören Auer. 2025. Ontoaligner: A comprehensive modular and robust python toolkit for ontology alignment. In European Semantic Web Conference. Springer, 174–191.
- Bandrowski et al. (2016) Anita Bandrowski, Ryan Brinkman, Mathias Brochhausen, Matthew H Brush, Bill Bug, Marcus C Chibucos, Kevin Clancy, Mélanie Courtot, Dirk Derom, Michel Dumontier, et al. 2016. The ontology for biomedical investigations. PloS one 11, 4 (2016), e0154556.
- Bian (2025) Haonan Bian. 2025. LLM-empowered knowledge graph construction: A survey. arXiv preprint arXiv:2510.20345 (2025).
- Bordes et al. (2013) Antoine Bordes, Nicolas Usunier, Alberto Garcia-Duran, Jason Weston, and Oksana Yakhnenko. 2013. Translating embeddings for modeling multi-relational data. Advances in neural information processing systems 26 (2013).
- Chen et al. (2021a) Jun Chen, Azza Althagafi, and Robert Hoehndorf. 2021a. Predicting candidate genes from phenotypes, functions and anatomical site of expression. Bioinformatics 37, 6 (2021), 853–860.
- Chen et al. (2021b) Jiaoyan Chen, Pan Hu, Ernesto Jimenez-Ruiz, Ole Magnus Holter, Denvar Antonyrajah, and Ian Horrocks. 2021b. OWL2Vec*: embedding of OWL ontologies. Machine Learning 110, 7 (2021), 1813–1845.
- Chen et al. (2023) Mingyang Chen, Wen Zhang, Yuxia Geng, Zezhong Xu, Jeff Z. Pan, and Huajun Chen. 2023. Generalizing to unseen elements: a survey on knowledge extrapolation for knowledge graphs. In Proceedings of the Thirty-Second International Joint Conference on Artificial Intelligence (Macao, P.R.China) (IJCAI ’23). Article 737, 9 pages. doi:10.24963/ijcai.2023/737
- Christiano et al. (2017) Paul F Christiano, Jan Leike, Tom Brown, Miljan Martic, Shane Legg, and Dario Amodei. 2017. Deep reinforcement learning from human preferences. Advances in neural information processing systems 30 (2017).
- Consortium (2026) The Gene Ontology Consortium. 2026. The Gene Ontology knowledgebase in 2026. Nucleic Acids Research 54, D1 (01 2026), D1779–D1792. arXiv:https://academic.oup.com/nar/article-pdf/54/D1/D1779/66009250/gkaf1292.pdf doi:10.1093/nar/gkaf1292
- Devlin et al. (2019) Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019. Bert: Pre-training of deep bidirectional transformers for language understanding. In Proceedings of the 2019 conference of the North American chapter of the association for computational linguistics: human language technologies, volume 1 (long and short papers). 4171–4186.
- Dhingra et al. (2018) Bhuwan Dhingra, Christopher Shallue, Mohammad Norouzi, Andrew Dai, and George Dahl. 2018. Embedding text in hyperbolic spaces. In Proceedings of the Twelfth Workshop on Graph-Based Methods for Natural Language Processing (TextGraphs-12). 59–69.
- Dragisic et al. (2017) Zlatan Dragisic, Valentina Ivanova, Huanyu Li, and Patrick Lambrix. 2017. Experiences from the anatomy track in the ontology alignment evaluation initiative. Journal of biomedical semantics 8 (2017), 1–28.
- Fallatah et al. (2020) Omaima Fallatah, Ziqi Zhang, and Frank Hopfgartner. 2020. A gold standard dataset for large knowledge graphs matching. In Ontology Matching 2020: Proceedings of the 15th International Workshop on Ontology Matching co-located with the 19th International Semantic Web Conference (ISWC 2020), Vol. 2788. CEUR Workshop Proceedings, 24–35.
- Gatto et al. (2023) Joseph Gatto, Omar Sharif, Parker Seegmiller, Philip Bohlman, and Sarah M. Preum. 2023. Text Encoders Lack Knowledge: Leveraging Generative LLMs for Domain-Specific Semantic Textual Similarity. In Proceedings of the Third Workshop on Natural Language Generation, Evaluation, and Metrics (GEM), Sebastian Gehrmann, Alex Wang, João Sedoc, Elizabeth Clark, Kaustubh Dhole, Khyathi Raghavi Chandu, Enrico Santus, and Hooman Sedghamiz (Eds.). Association for Computational Linguistics, Singapore, 277–288. https://aclanthology.org/2023.gem-1.23/
- Giglou et al. (2026) Hamed Babaei Giglou, Jennifer D’Souza, Andrei Aioanei, Nandana Mihindukulasooriya, and Sören Auer. 2026. OntoLearner: A Modular Python Library for Ontology Learning with Large Language Models. arXiv preprint arXiv:2607.01977 (2026).
- Hertling and Paulheim (2023) Sven Hertling and Heiko Paulheim. 2023. Olala: Ontology matching with large language models. In Proceedings of the 12th knowledge capture conference 2023. 131–139.
- Karam et al. (2020) Naouel Karam, Abderrahmane Khiat, Alsayed Algergawy, Melanie Sattler, Claus Weiland, and Marco Schmidt. 2020. Matching biodiversity and ecology ontologies: challenges and evaluation results. The Knowledge Engineering Review 35 (2020), e9.
- Kumar et al. (2025) Lokendra Kumar, Neelesh S Upadhye, and Kannan Piedy. 2025. Advances and Challenges in Semantic Textual Similarity: A Comprehensive Survey. arXiv preprint arXiv:2601.03270 (2025).
- Lambert et al. (2022) Nathan Lambert, Louis Castricato, Leandro von Werra, and Alex Havrilla. 2022. Illustrating reinforcement learning from human feedback (rlhf). Hugging Face Blog 9 (2022).
- Li et al. (2025) Wenda Li, Tongya Zheng, Shunyu Liu, Yu Wang, Kaixuan Chen, Hanyang Yuan, Bingde Hu, Zujie Ren, Mingli Song, and Gang Chen. 2025. Towards Efficient LLM-aware Heterogeneous Graph Learning. arXiv preprint arXiv:2511.17923 (2025).
- Li et al. (2023) Zehan Li, Xin Zhang, Yanzhao Zhang, Dingkun Long, Pengjun Xie, and Meishan Zhang. 2023. Towards general text embeddings with multi-stage contrastive learning. arXiv preprint arXiv:2308.03281 (2023).
- Liang et al. (2023) Xinyu Liang, Guannan Si, Jianxin Li, Pengxin Tian, Zhaoliang An, and Fengyu Zhou. 2023. A survey of inductive knowledge graph completion. Neural Comput. Appl. 36, 8 (Dec. 2023), 3837–3858. doi:10.1007/s00521-023-09286-2
- Mikolov et al. (2013) Tomas Mikolov, Kai Chen, Greg Corrado, and Jeffrey Dean. 2013. Efficient estimation of word representations in vector space. arXiv preprint arXiv:1301.3781 (2013).
- Nas and Huschka (2023) E. Nas and M. Huschka. 2023. MSE Benchmark. https://github.com/EngyNasr/MSE-Benchmark.
- Nickel and Kiela (2017) Maximillian Nickel and Douwe Kiela. 2017. Poincaré embeddings for learning hierarchical representations. Advances in neural information processing systems 30 (2017).
- Pennington et al. (2014) Jeffrey Pennington, Richard Socher, and Christopher D Manning. 2014. Glove: Global vectors for word representation. In Proceedings of the 2014 conference on empirical methods in natural language processing (EMNLP). 1532–1543.
- Qiang (2023) Zhangcheng Qiang. 2023. Ontology-compliant knowledge graphs. In European Semantic Web Conference. Springer, 298–309.
- Qwen Team (2026) Qwen Team. 2026. Qwen3.5: Towards Native Multimodal Agents. https://qwen.ai/blog?id=qwen3.5
- Rafailov et al. (2023) Rafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D Manning, Stefano Ermon, and Chelsea Finn. 2023. Direct preference optimization: Your language model is secretly a reward model. Advances in neural information processing systems 36 (2023), 53728–53741.
- Ristoski and Paulheim (2016) Petar Ristoski and Heiko Paulheim. 2016. Rdf2vec: Rdf graph embeddings for data mining. In International semantic web conference. Springer, 498–514.
- Song et al. (2020) Kaitao Song, Xu Tan, Tao Qin, Jianfeng Lu, and Tie-Yan Liu. 2020. Mpnet: Masked and permuted pre-training for language understanding. Advances in neural information processing systems 33 (2020), 16857–16867.
- Sun et al. (2025) Yiqun Sun, Qiang Huang, Anthony KH Tung, and Jun Yu. 2025. Text embeddings should capture implicit semantics, not just surface meaning. arXiv preprint arXiv:2506.08354 (2025).
- Sun et al. (2019) Zhiqing Sun, Zhi-Hong Deng, Jian-Yun Nie, and Jian Tang. 2019. RotatE: Knowledge Graph Embedding by Relational Rotation in Complex Space. CoRR abs/1902.10197 (2019). http://arxiv.org/abs/1902.10197
- Trouillon et al. (2016) Théo Trouillon, Johannes Welbl, Sebastian Riedel, Éric Gaussier, and Guillaume Bouchard. 2016. Complex embeddings for simple link prediction. In International conference on machine learning. PMLR, 2071–2080.
- Wang et al. (2022) Liang Wang, Nan Yang, Xiaolong Huang, Binxing Jiao, Linjun Yang, Daxin Jiang, Rangan Majumder, and Furu Wei. 2022. Text embeddings by weakly-supervised contrastive pre-training. arXiv preprint arXiv:2212.03533 (2022).
- Wang et al. (2021) Xiaozhi Wang, Tianyu Gao, Zhaocheng Zhu, Zhengyan Zhang, Zhiyuan Liu, Juanzi Li, and Jian Tang. 2021. KEPLER: A unified model for knowledge embedding and pre-trained language representation. Transactions of the Association for Computational Linguistics 9 (2021), 176–194.
- Xiao et al. (2024) Shitao Xiao, Zheng Liu, Peitian Zhang, Niklas Muennighoff, Defu Lian, and Jian-Yun Nie. 2024. C-Pack: Packed Resources For General Chinese Embeddings. In Proceedings of the 47th International ACM SIGIR Conference on Research and Development in Information Retrieval (Washington DC, USA) (SIGIR ’24). Association for Computing Machinery, New York, NY, USA, 641–649. doi:10.1145/3626772.3657878
- Yao et al. (2019) Liang Yao, Chengsheng Mao, and Yuan Luo. 2019. KG-BERT: BERT for knowledge graph completion. arXiv preprint arXiv:1909.03193 (2019).
- Zhang et al. (2024) Xin Zhang, Yanzhao Zhang, Dingkun Long, Wen Xie, Ziqi Dai, Jialong Tang, Huan Lin, Baosong Yang, Pengjun Xie, Fei Huang, Meishan Zhang, Wenjie Li, and Min Zhang. 2024. mGTE: Generalized Long-Context Text Representation and Reranking Models for Multilingual Text Retrieval. In Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing: Industry Track, Franck Dernoncourt, Daniel Preoţiuc-Pietro, and Anastasia Shimorina (Eds.). Association for Computational Linguistics, Miami, Florida, US, 1393–1412. doi:10.18653/v1/2024.emnlp-industry.103