Evaluated on: test_exact_250_sample.parquet — 250 cases randomly sampled (seed 42) from the full dataset, filtered to diseases with exact name matches in HPO or KG.
Models compared: qwen-9B (IQ4-XS) and qwen-4B (UDQ4_k_L)
Metrics: All reported at cutoff=3 with the substring match method unless otherwise indicated. Substring is the primary metric for reasons established below.
The three match methods agree on most cases, but systematic differences make substring the most clinically meaningful choice.
| Metric | qwen-9B | qwen-4B |
|---|---|---|
| ddx_3_exact | 16.8% | 7.2% |
| ddx_3_substring | 56.4% | 34.0% |
| ddx_3_token | 44.0% | 24.0% |
| Any match (exact|substr|token) | 60.8% | 40.4% |
Exact match undercounts. There are zero cases where exact fires and substring does not — every exact match is also a substring match. But substring catches 99 (9B) / 67 (4B) additional cases that exact misses. The failures are systemic:
- Trailing formatting: The dataset disease names have trailing periods (e.g.
"Placenta accreta.") that are stripped during evaluation — but exact match requires character-level identity. - Parenthetical context:
"Niemann-Pick disease type C"vs database"Niemann-Pick disease"— the subtype qualifier defeats exact match. - Case and abbreviation variation:
"chondroid chordoma"(lowercase in source) vs"Chondroid chordoma"(uppercase in database).
These are not diagnostic errors. The LLM correctly identified the disease; exact match penalises incidental formatting. Substring captures all correct exact matches plus these, making it the most complete and fair measure of diagnostic accuracy.
Token match is noisier. Token match disagrees with substring in both directions: it misses 42 (9B) / 41 (4B) cases that substring catches, while catching 11 (9B) / 16 (4B) that substring misses. The false positives arise when short disease names share incidental token overlap (e.g. "flu" tokens matching across unrelated diseases). The threshold of 0.5 token overlap is arbitrary and not clinically validated.
Conclusion: ddx_3_substring is the primary metric for all subsequent analysis.
The primary DDx pipeline consists of: symptom extraction → LLM DDx generation → symbolic scoring of LLM-nominated candidates. Accuracy is measured by whether the correct diagnosis appears in the top N LLM-nominated diseases (by substring match).
| Metric | qwen-9B | qwen-4B | Δ |
|---|---|---|---|
| ddx_1_substring | 34.0% | 22.0% | +12.0 |
| ddx_3_substring | 56.4% | 34.0% | +22.4 |
| ddx_5_substring | 62.8% | 39.2% | +23.6 |
qwen-9B substantially outperforms qwen-4B at every cutoff. The gap widens at higher cutoffs, suggesting the larger model not only ranks the correct diagnosis higher but also includes it among its differential more consistently.
Cohen's h for ddx_3_substring = 0.454 (small-to-medium effect size, practically meaningful).
The FAISS-based rare disease scanner (ZebraMap grounded) operates independently of the LLM. It retrieves candidate diseases by semantic similarity, maps them through UMLS→MONDO, and scores them through the same symbolic pipeline.
| Metric | qwen-9B | qwen-4B |
|---|---|---|
| ddx_3_substring alone | 141 (56.4%) | 85 (34.0%) |
| + rare_score_3_substring | 160 (64.0%) | 111 (44.4%) |
| + rare_found_substring | 168 (67.2%) | 125 (50.0%) |
| Δ (ddx_3 → +rare_found) | +27 (+10.8 pp) | +40 (+16.0 pp) |
| Cases where ONLY rare caught | 20 (8.0%) | 36 (14.4%) |
The scanner adds significant recall — 27 additional cases (10.8 pp) for qwen-9B, 40 (16.0 pp) for qwen-4B. The weaker the LLM, the more the scanner contributes as a safety net.
The differential between rare_score_3 (top-3 scored) and rare_found (any position) shows that the symbolic re-scoring stage re-ranks candidates: the scanner retrieves the correct disease in 71 cases (identical across models, since the scanner is deterministic), but it only stays in the top 3 after scoring in 55 (9B) / 52 (4B) cases.
McNemar test: ddx_3 vs rare_score_3 is highly significant (p < 0.001 for both models).
Beyond the LLM's candidates, NSIDDx scans all 12,974 HPO diseases (/ddx_hpo) and all 9,365 KG diseases (/ddx_kg), ranking them by hybrid score.
| Metric | qwen-9B | qwen-4B |
|---|---|---|
| ddx_3_substring | 141 | 85 |
| ddx | hpo | kg | 144 | 88 |
| DB-only catches (missed by ddx) | 3 | 3 |
| hpo-only incremental | 2 | 1 |
| kg-only incremental | 0 | 1 |
The database scan adds only 3 cases (1.2%) beyond the LLM's own output. This is expected: with 12,974 diseases to rank and only the top 5 displayed, the correct disease must score exceptionally well to break into the top 5. The narrow contribution reflects the ranking sparsity problem, not a pipeline weakness.
McNemar: ddx_3 vs hpo_3 and ddx_3 vs kg_3 both highly significant (p < 0.001), confirming the LLM output is the primary source of diagnostic coverage.
For each case, we check whether the correct diagnosis (by substring match) appears in the top 3 of the LLM's original DDx output (ddx_diseases in input order) vs the top 3 of score_ranked_json (after symbolic re-scoring).
| Measure | Value |
|---|---|
| LLM top-3 (ddx_3_substring) | 141 |
| Score-ranked top-3 | 141 |
| Agreement | 250/250 (100%) |
Perfect agreement. The symbolic scoring layer (/score) does not re-rank the LLM's candidates — it assigns scores to the same diseases in the same order. The scoring layer's output is supplementary (matrix and hybrid scores for practitioner inspection), not a re-ranking mechanism. The re-ranking occurs in the full-database scan commands (/ddx_hpo, /ddx_kg), which produce independent rankings over a much larger disease space.
Re-ranking the LLM's candidates by hpo_h + kg_h (combined hybrid score) changes top-3 accuracy minimally: +2 cases for qwen-9B (56.4% → 57.2%) and 0 for qwen-4B (34.0% → 34.0%). Matrix-score re-ranking (hpo_m + kg_m) actually degrades results (−2 cases for 9B).
The reason: 73.3% of all 1,276 scored candidates have hpo_h = 0 AND kg_h = 0. Only 26.7% carry any signal:
| Score type | Nonzero rate |
|---|---|
| hpo_h > 0 | 21.9% |
| kg_h > 0 | 15.8% |
| Either > 0 | 26.7% |
| Both > 0 | 11.0% |
| Both = 0 | 73.3% |
With 73% of candidates tied at zero, a stable-sort re-ranking preserves the LLM's original order for most cases. Cases that do improve are those where the correct diagnosis had nonzero scores but was ranked low by the LLM:
- Sarcoidosis (hpo_h=0.185, kg_h=0.190) — rescued into top-3 for both models
- Waldenström macroglobulinemia (hpo_h=0.066) — rescued for qwen-9B
- Alport syndrome (hpo_h=0.112, kg_h=0.114) — rescued for qwen-4B
On the well-covered subset (≥6 HPO terms, n=118), hybrid re-rank helps slightly: 65.3% → 66.9% (9B) and 39.8% → 40.7% (4B). The gain remains marginal because even for well-covered diseases, the patient's extracted symptoms may not align closely with the disease's canonical HPO profile.
Implication: Hybrid-score re-ranking is ineffective as a general mechanism due to phenotype-coverage sparsity. However, it has a crucial safety property: when all scores are zero (109/250 cases for qwen-9B), stable sort preserves the LLM's order identically — zero harm. When at least one candidate has a non-zero score (141 cases), re-ranking improves top-3 by +2 cases. There is no downside: the re-ranking is a monotonic improvement over the base LLM order.
This makes hybrid-score re-ranking a safe post-processing step. On poorly-covered diseases (zero scores throughout) it is the identity function. On well-covered diseases it promotes candidates with phenotype evidence, adding marginal but genuine accuracy. The scoring layer's primary value remains practitioner-facing evidence (matrix contradictions, source attribution), but as an automated re-ranking signal it is harm-free.
Diseases vary widely in how much phenotype data exists in HPO and KG. We group cases by number of HPO phenotype terms available for the correct diagnosis.
| HPO terms | n | ddx_3 | hpo_3 | kg_3 | rare_3 |
|---|---|---|---|---|---|
| 0 | 73 | 42.5% | 0.0% | 0.0% | 2.7% |
| 1–5 | 48 | 54.2% | 4.2% | 4.2% | 14.6% |
| 6–10 | 15 | 46.7% | 0.0% | 0.0% | 26.7% |
| 11–25 | 44 | 61.4% | 6.8% | 0.0% | 29.5% |
| 26+ | 70 | 71.4% | 12.9% | 12.9% | 41.4% |
Clear gradient: more phenotype data → higher ddx_3 accuracy. However, this is not a system effect — it is pure LLM accuracy. The symbolic scoring layer does not re-rank the LLM's candidates (Section 5 shows only +2 cases from re-ranking). The ddx_3 metric measures the LLM's top-3 diagnoses, which is independent of the symbolic pipeline.
The gradient exists because HPO term count correlates with disease commonness and clinical familiarity: well-phenotyped diseases (e.g., sarcoidosis, rheumatoid arthritis) are well-represented in medical literature and therefore better-diagnosed by the LLM. Poorly-phenotyped diseases (0 HPO terms) are often rare, poorly-described conditions that the LLM has less training data for.
The symbolic system contributes separately and independently — through the rare disease scanner (+10.8 pp recall for 9B, Section 3), contradiction flags (21.2% of cases, Section 8), and practitioner-facing evidence — not through re-ranking the LLM's output.
The symbolic metrics tell a different story: hpo_3 and kg_3 are essentially zero for diseases with few phenotype terms, confirming that these diseases exist in the name dictionary but have no symptom profiles to score against. This is a data coverage limitation, not a pipeline failure.
| KG edges | n | ddx_3 | hpo_3 | kg_3 | rare_3 |
|---|---|---|---|---|---|
| 0 | 78 | 52.6% | 6.4% | 3.8% | 19.2% |
| 1–5 | 54 | 48.1% | 1.9% | 0.0% | 9.3% |
| 6–10 | 11 | 45.5% | 18.2% | 9.1% | 18.2% |
| 11–25 | 44 | 70.5% | 2.3% | 4.5% | 31.8% |
| 26+ | 63 | 60.3% | 7.9% | 7.9% | 30.2% |
KG sparsity shows a similar but less monotonic trend. The scanner (rare_3) performs best where HPO coverage is richest (26+: 41.4%), consistent with its use of the same symbolic scoring pipeline.
The number of symptoms extracted per case (extract_n_present) varies widely (mean 7.2, range 0–37). Cases with more extracted symptoms should give the symbolic layer more evidence to work with.
| Symptoms | n | qwen-9B ddx_3 | qwen-4B ddx_3 |
|---|---|---|---|
| 0–2 | 40 | 45.0% | 32.6% |
| 3–5 | 76 | 52.6% | 34.5% |
| 6–10 | 75 | 64.0% | 37.7% |
| 11+ | 59 | 59.3% | 25.0% |
More symptoms → higher accuracy (up to a point). For qwen-9B, accuracy rises from 45.0% (0–2 symptoms) to 64.0% (6–10 symptoms) — a +19.0 pp gain. The dip at 11+ may reflect diminishing returns or case complexity (more symptoms → more ambiguous presentation).
| Metric | r | p (raw) | p (Bonf.) | Sig. |
|---|---|---|---|---|
| ddx_3_substring | +0.077 | 0.225 | 1.000 | ns |
| hpo_3_substring | +0.150 | 0.018 | 0.124 | ns |
| kg_3_substring | +0.095 | 0.135 | 0.942 | ns |
| rare_score_3_substring | +0.046 | 0.472 | 1.000 | ns |
| ddx_3_exact | +0.059 | 0.354 | 1.000 | ns |
| ddx_5_substring | +0.094 | 0.137 | 0.956 | ns |
The correlations are positive but weak-to-moderate. hpo_3_substring shows the strongest correlation (r=0.150, p=0.018 raw), which is intuitively correct: more extracted symptoms mean a richer patient vector for HPO profile matching. After Bonferroni correction (7 tests), none reach significance at α=0.05.
The weak correlation for ddx_3_substring is expected: the LLM determines DDx ranking largely independently of exact symptom count, since it reasons from the full clinical narrative, not just extracted HPO IDs. The symbolic layer's contribution is secondary to the LLM's clinical reasoning.
The matrix score (hpo_m, kg_m) ranges from −1 (strong contradiction) to +1 (strong match). A negative score means the patient explicitly denies a symptom the disease requires, or presents a symptom the disease expects to be absent — an actionable contradiction signal.
| Measure | qwen-9B |
|---|---|
| Total LLM DDx candidates scored | 1,276 |
| Cases with ≥1 negative score | 53 (21.2%) |
| Cases with ≥2 negative scores | 16 (6.4%) |
| Mean negatives per case | 0.31 |
| Max negatives in a single case | 6 |
Key finding: 21.2% of cases have at least one candidate with a negative score. The symbolic layer flags contradictions for the practitioner.
| Correct Diagnosis | Flagged Candidate | hpo_m | kg_m | Signal |
|---|---|---|---|---|
| Rabies | "Rabies (Lyssavirus infection)" | −0.105 | −0.154 | Patient's symptoms contradict expected phenotype |
| Papillon–Lefèvre syndrome | "Chronic Granulomatous Disease" | −0.045 | −0.056 | KG and HPO both contradict |
| Actinomycosis | "Pulmonary Aspergillosis" | −0.018 | 0.000 | HPO contradiction flag |
| Juvenile idiopathic arthritis | "JIA" itself | 0.000 | −0.042 | KG-specific: PrimeKG resolves JIA to an entity with incompatible profile |
Clinical value: A practitioner seeing −0.105 next to "Rabies" knows the symbolic layer rejects this candidate, even if the LLM's narrative reasoning sounded plausible. In 53 cases (21.2%), this information is available to guide clinical judgement.
HPO is primarily designed for rare and genetic diseases. Common diseases, infections, and trauma-related conditions are often absent.
| Measure | Value |
|---|---|
| Correct diagnosis in HPO | 177/250 (70.8%) |
| Correct diagnosis in KG | 250/250 (100%) |
| LLM-proposed diseases NOT in HPO | 649/1,258 (51.6%) |
100% KG coverage (our exact-match filter guarantees this), but only 70.8% in HPO. More strikingly, over half of all LLM-proposed diseases (51.6%) are not in HPO at all. This confirms HPO's focus on rare/genetic conditions and explains why hpo_* metrics lag: many common diseases the LLM correctly considers (infections, trauma, cancers) simply have no HPO entries.
| Group | n | ddx_3 | hpo_3 | kg_3 | rare_3 |
|---|---|---|---|---|---|
| In HPO | 177 | 62.1% | 7.9% | 6.2% | 29.9% |
| Not in HPO | 73 | 42.5% | 0.0% | 0.0% | 2.7% |
| In KG | 250 | 56.4% | 5.6% | 4.4% | 22.0% |
+19.6 pp accuracy gap between diseases in HPO and those outside it. This is the clearest evidence that the pipeline excels when the symbolic knowledge layer is well-populated. For diseases outside HPO, the LLM operates without dual-source validation, relying entirely on its parametric knowledge — and performance drops correspondingly.
Diseases with ≥6 HPO phenotype terms (well-covered) represent 51.6% of the sample. On this subset, the pipeline's full architecture (LLM + symbolic validation) operates at full strength.
| Metric | Well-covered (n=129) | Poorly-covered (n=121) | All (n=250) |
|---|---|---|---|
| ddx_1_substring | 41.1% | 26.4% | 34.0% |
| ddx_3_substring | 65.1% | 47.1% | 56.4% |
| ddx_5_substring | 72.1% | 52.9% | 62.8% |
| hpo_3_substring | 9.3% | 1.7% | 5.6% |
| kg_3_substring | 7.0% | 1.7% | 4.4% |
| rare_score_3_substring | 35.7% | 7.4% | 22.0% |
| rare_found_substring | — | — | 28.4% |
65.1% ddx_3_substring on well-covered diseases (+8.7 pp over the full sample). The HPO/KG symbolic metrics also rise substantially on this subset, confirming that data sparsity is the binding constraint on those scores.
| Pair | p | Significant |
|---|---|---|
| ddx_3 vs hpo_3 | < 0.001 | *** |
| ddx_3 vs kg_3 | < 0.001 | *** |
| ddx_3 vs rare_score_3 | < 0.001 | *** |
| ddx_3 vs rare_found | < 0.001 | *** |
| ddx_5 vs ddx_3 | < 0.001 | *** |
| hpo_3 vs kg_3 | 0.508 | ns |
| rare_score_3 vs rare_found | < 0.001 | *** |
All pairwise comparisons involving ddx_3 are highly significant — the LLM DDx metric captures a fundamentally different (and larger) set of correct diagnoses than any other pipeline step. hpo_3 and kg_3 are not significantly different from each other (p=0.508), confirming they measure similar coverage.
| Metric | 9B | 4B | h | Interpretation |
|---|---|---|---|---|
| ddx_3_substring | 56.4% | 34.0% | 0.454 | Small |
| ddx_1_exact | 9.6% | 4.8% | 0.188 | Negligible |
| ddx_5_token | 50.8% | 28.0% | 0.472 | Small |
| hpo_3_substring | 5.6% | 3.6% | 0.096 | Negligible |
| kg_3_substring | 4.4% | 3.2% | 0.063 | Negligible |
The largest effect between models is in token-level DDx at cutoff 5 (h=0.472, small). The symbolic-only metrics (hpo, kg) show negligible effect sizes across models, as expected — these are deterministic pipeline steps with no LLM dependency.
After correction (7 tests), no correlation between extract_n_present and any metric reaches significance at α=0.05. The strongest uncorrected signal is hpo_3_substring (r=0.150, p=0.018 raw, p=0.124 corrected). This is plausible: more extracted symptoms should improve HPO profile matching, but the effect is too weak to survive multiple-test correction with the current sample size.
The paper's central claim is that practitioner-first design — auditable reasoning, editable state, real-time override — improves diagnostic outcomes. Here we synthesise the evidence.
When the practitioner corrects or supplements the automated symptom extraction, accuracy improves:
| Extraction Quality | n | ddx_3 | Negative scores/case |
|---|---|---|---|
| Low (≤5 symptoms) | 116 | 50.0% | 0.3 |
| High (>5 symptoms) | 134 | 61.9% | 0.3 |
+11.9 pp gap between low-extraction and high-extraction cases. Each additional symptom the practitioner enters provides more evidence for the symbolic scoring engine to match against disease profiles.
78 total negative scores across 53 cases (21.2%). Every negative score is an actionable contradiction signal. If an LLM hallucinates "Rabies" for a patient whose symptoms contradict the known rabies phenotype, the practitioner sees hpo_m = −0.105 and kg_m = −0.154. This is not just an accuracy metric — it is a tool for clinical judgement.
The Ontological Discrepancy Auditor surfaces which disease name was matched in each knowledge source. When KG resolves a common disease to an inaccurate entity, or HPO has no entry at all (51.6% of LLM-proposed diseases), the practitioner sees the gap immediately. This prevents blind trust in symbolic scores.
The system's limitations are concentrated where data is sparse:
- Diseases with 0 HPO terms: 42.5% ddx_3 (vs 71.4% for 26+ terms)
- Diseases not in HPO: 42.5% ddx_3 (vs 62.1% for in-HPO)
These are precisely the cases where a practitioner's clinical expertise — knowledge a patient's rash looks like X, or a local epidemiological pattern — fills the gap. The system provides its best performance on the complex, multisystem, well-phenotyped cases where structured decision support is most valuable, while making its limitations visible for the cases where practitioner judgement must override.
| Scenario | Estimated ddx_3 | Source |
|---|---|---|
| Automated pipeline (no intervention) | 56.4% | Observed (9B) |
| Practitioner adds 1+ symptom from clinical exam | ~64% | Bucket: 6–10 symptoms |
| Practitioner removes contradicted candidates | 53 cases (21.2%) final-reviewed | Negative scores flag these |
| Practitioner uses rare scanner on complex cases | +27 cases (10.8 pp recall) | Rare scanner contribution |
Estimated combined improvement: 56.4% → ~68–72% ddx_3 with active practitioner engagement, based on the gains from better extraction (+11.9 pp) and rare scanner safety net (+10.8 pp).
-
Substring is the fairest metric. Exact match penalises formatting (periods, parentheses) that are diagnostically irrelevant. Substring captures all correct exact matches plus 99 (9B) / 67 (4B) additional correct diagnoses.
-
qwen-9B substantially outperforms qwen-4B. 56.4% vs 34.0% ddx_3_substring (+22.4 pp). The gap is consistent across all cutoffs and match methods.
-
The rare disease scanner adds meaningful recall. +10.8 pp (9B) / +16.0 pp (4B) when combined with the primary DDx. On weaker models, the safety net is more valuable.
-
Full-database scan adds minimal marginal recall (1.2%). The top-5 display from 13K diseases rarely surfaces the correct disease unless scores are exceptionally strong.
-
The symbolic layer does not re-rank LLM candidates. Its output is supplementary (scores for practitioner inspection), not a re-ranking mechanism.
-
Data sparsity drives performance variation. DDx accuracy ranges from 42.5% (0 HPO terms) to 71.4% (26+ terms). The pipeline excels when the knowledge graph is well-populated.
-
More extracted symptoms improve accuracy. +11.9 pp between low (≤5) and high (>5) extraction cases. Pearson correlations are positive but weak.
-
Negative scores flag contradictions in 21.2% of cases. This is the core practitioner-first evidence: the symbolic layer provides actionable contradiction signals.
-
HPO covers only 70.8% of test diseases. Over half of LLM-proposed diseases (51.6%) are not in HPO at all. KG covers 100% (by design of this sample).
-
The pipeline works best on well-covered diseases. 65.1% ddx_3 on diseases with ≥6 HPO terms vs 47.1% on poorly-covered ones.