Skip to content

qa_2 (HotpotQA): string_match_part matches the gold answer "no" inside other words (not, know, cannot), so hedged or refusing answers get full credit #112

Description

@shaurya416

scripts/eval/synthetic/constants.py:24-26 (HEAD c3f5e3b)

def string_match_part(preds, refs):
    score = sum([max([1.0 if r.lower() in pred.lower() else 0.0 for r in ref]) for pred, ref in zip(preds, refs)]) / len(preds) * 100
    return round(score, 2)

qa_1/qa_2 use this metric (TASKS['qa']). HotpotQA references come from d['answer'] (scripts/data/synthetic/qa.py:109), and a share of those answers are the bare words yes / no. r.lower() in pred.lower() has no token or word boundary, so the reference no matches inside not, know, cannot, none, nothing.

What happens (reproduced on the real module)

scripts/eval/synthetic/constants.py:24-26 checks r.lower() in pred.lower() with no word boundary. HotpotQA gold answers come from d['answer'] (scripts/data/synthetic/qa.py:109). In the default run (NUM_SAMPLES=500, sampled sequentially from hotpot_dev_distractor_v1.json), 19 of the 500 gold answers are 'no' and 14 are 'yes'. With the real module loaded, 'I do not know', 'I cannot answer that.', 'The documents do not say.' and 'Norway' all score 100.0 against ['no']. Controls behave correctly: 'Yes' vs ['no'] = 0, 'No' vs ['no'] = 100, 'London' vs ['Paris'] = 0.

This keeps the containment design from #15/#23. The problem is only that the reference can match in the middle of another word. Impact is bounded: at most 3.8 points on qa_2 per length (at most about 0.29 on the 13-task average). It can differ between models depending on how each one hedges, although the ' Answer:' prefix (prepare.py:99-101) and the 32-token cap already reduce refusals. qa_1 (SQuAD) has no yes/no answers, and no dataset has an empty reference.

Minimal fix that keeps containment:
re.search(r'(?<!\w)' + re.escape(r.lower()) + r'(?!\w)', pred.lower())
applied in both string_match_part and string_match_all. Before and after, report the qa_2 score change on existing prediction files so maintainers can see how far scores move against published numbers.

Happy to open the PR.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions