scripts/eval/synthetic/constants.py:24-26 (HEAD c3f5e3b)
def string_match_part(preds, refs):
score = sum([max([1.0 if r.lower() in pred.lower() else 0.0 for r in ref]) for pred, ref in zip(preds, refs)]) / len(preds) * 100
return round(score, 2)
qa_1/qa_2 use this metric (TASKS['qa']). HotpotQA references come from d['answer'] (scripts/data/synthetic/qa.py:109), and a share of those answers are the bare words yes / no. r.lower() in pred.lower() has no token or word boundary, so the reference no matches inside not, know, cannot, none, nothing.
What happens (reproduced on the real module)
scripts/eval/synthetic/constants.py:24-26 checks r.lower() in pred.lower() with no word boundary. HotpotQA gold answers come from d['answer'] (scripts/data/synthetic/qa.py:109). In the default run (NUM_SAMPLES=500, sampled sequentially from hotpot_dev_distractor_v1.json), 19 of the 500 gold answers are 'no' and 14 are 'yes'. With the real module loaded, 'I do not know', 'I cannot answer that.', 'The documents do not say.' and 'Norway' all score 100.0 against ['no']. Controls behave correctly: 'Yes' vs ['no'] = 0, 'No' vs ['no'] = 100, 'London' vs ['Paris'] = 0.
This keeps the containment design from #15/#23. The problem is only that the reference can match in the middle of another word. Impact is bounded: at most 3.8 points on qa_2 per length (at most about 0.29 on the 13-task average). It can differ between models depending on how each one hedges, although the ' Answer:' prefix (prepare.py:99-101) and the 32-token cap already reduce refusals. qa_1 (SQuAD) has no yes/no answers, and no dataset has an empty reference.
Minimal fix that keeps containment:
re.search(r'(?<!\w)' + re.escape(r.lower()) + r'(?!\w)', pred.lower())
applied in both string_match_part and string_match_all. Before and after, report the qa_2 score change on existing prediction files so maintainers can see how far scores move against published numbers.
Happy to open the PR.
scripts/eval/synthetic/constants.py:24-26(HEAD c3f5e3b)qa_1/qa_2use this metric (TASKS['qa']). HotpotQA references come fromd['answer'](scripts/data/synthetic/qa.py:109), and a share of those answers are the bare wordsyes/no.r.lower() in pred.lower()has no token or word boundary, so the referencenomatches insidenot,know,cannot,none,nothing.What happens (reproduced on the real module)
scripts/eval/synthetic/constants.py:24-26 checks
r.lower() in pred.lower()with no word boundary. HotpotQA gold answers come from d['answer'] (scripts/data/synthetic/qa.py:109). In the default run (NUM_SAMPLES=500, sampled sequentially from hotpot_dev_distractor_v1.json), 19 of the 500 gold answers are 'no' and 14 are 'yes'. With the real module loaded, 'I do not know', 'I cannot answer that.', 'The documents do not say.' and 'Norway' all score 100.0 against ['no']. Controls behave correctly: 'Yes' vs ['no'] = 0, 'No' vs ['no'] = 100, 'London' vs ['Paris'] = 0.This keeps the containment design from #15/#23. The problem is only that the reference can match in the middle of another word. Impact is bounded: at most 3.8 points on qa_2 per length (at most about 0.29 on the 13-task average). It can differ between models depending on how each one hedges, although the ' Answer:' prefix (prepare.py:99-101) and the 32-token cap already reduce refusals. qa_1 (SQuAD) has no yes/no answers, and no dataset has an empty reference.
Minimal fix that keeps containment:
re.search(r'(?<!\w)' + re.escape(r.lower()) + r'(?!\w)', pred.lower())applied in both string_match_part and string_match_all. Before and after, report the qa_2 score change on existing prediction files so maintainers can see how far scores move against published numbers.
Happy to open the PR.