The dataset viewer is not available for this split.
Error code: StreamingRowsError
Exception: ValueError
Message: Failed to convert pandas DataFrame to Arrow Table from file hf://datasets/uthmanjinadu/authority-bias-paper-recommendation@c44caec88144b00d6b484f3af579fc6f7d20973c/experiment_conditions.json.
Traceback: Traceback (most recent call last):
File "/src/services/worker/src/worker/utils.py", line 147, in get_rows_or_raise
return get_rows(
dataset=dataset,
...<4 lines>...
column_names=column_names,
)
File "/src/libs/libcommon/src/libcommon/utils.py", line 272, in decorator
return func(*args, **kwargs)
File "/src/services/worker/src/worker/utils.py", line 127, in get_rows
rows_plus_one = list(itertools.islice(safe_iter(ds, dataset=dataset), rows_max_number + 1))
File "/src/services/worker/src/worker/utils.py", line 483, in safe_iter
yield from ds.decode(False) if ds.features else ds
File "/usr/local/lib/python3.14/site-packages/datasets/iterable_dataset.py", line 2840, in __iter__
for key, example in ex_iterable:
^^^^^^^^^^^
File "/usr/local/lib/python3.14/site-packages/datasets/iterable_dataset.py", line 2373, in __iter__
for key, pa_table in self._iter_arrow():
~~~~~~~~~~~~~~~~^^
File "/usr/local/lib/python3.14/site-packages/datasets/iterable_dataset.py", line 2398, in _iter_arrow
for key, pa_table in self.ex_iterable._iter_arrow():
~~~~~~~~~~~~~~~~~~~~~~~~~~~~^^
File "/usr/local/lib/python3.14/site-packages/datasets/iterable_dataset.py", line 536, in _iter_arrow
for key, pa_table in iterator:
^^^^^^^^
File "/usr/local/lib/python3.14/site-packages/datasets/iterable_dataset.py", line 419, in _iter_arrow
for key, pa_table in self.generate_tables_fn(**gen_kwags):
~~~~~~~~~~~~~~~~~~~~~~~^^^^^^^^^^^^^
File "/usr/local/lib/python3.14/site-packages/datasets/packaged_modules/json/json.py", line 336, in _generate_tables
raise ValueError(
f"Failed to convert pandas DataFrame to Arrow Table from file {file}."
) from None
ValueError: Failed to convert pandas DataFrame to Arrow Table from file hf://datasets/uthmanjinadu/authority-bias-paper-recommendation@c44caec88144b00d6b484f3af579fc6f7d20973c/experiment_conditions.json.Need help to make the dataset viewer work? Make sure to review how to configure the dataset viewer, and open a discussion for direct support.
Authority Bias in Conversational Search Engines for Academic Paper Recommendation
Dataset accompanying the paper "Authority Bias in Conversational Search Engines for Academic Paper Recommendation" (EMNLP 2026, Main Conference). Code: https://github.com/jinaduuthman/Authority-Bias-In-Conversational-Search-Engine
This is a content-controlled counterfactual audit of authority bias in LLM paper recommendation. Each paper's content (title + abstract) is held fixed while its authority metadata (venue, author h-index, citations, affiliation) is manipulated across three conditions, so any change in an LLM's recommendation is attributable to authority signals alone.
Contents
| File | Description |
|---|---|
papers.json |
1,250 collected papers (50 per topic × 25 CS topics) from Semantic Scholar + OpenAlex. |
experiment_sets.json |
The 250 queries and, for each, the 10-paper candidate set under all three conditions (original, flipped, boosted), each candidate carrying its authority components. |
experiment_conditions.json |
Per-topic candidate pools under each condition (pre query-assembly). |
responses/open_weight_main.json |
Parsed model responses for the five open-weight models (main experiment). |
responses/closed_weight/*.json |
Responses for the three frontier models (gpt-5.4, gemini-3-flash-preview, claude-sonnet-4-6) and the gpt-4o-mini tier ablation. |
responses/pilot_1n/*.json |
1:N flip pilot runs (Gemma 2, Llama 3.1, Mistral) used to derive the authority-signal weights. |
author_score_ablation/papers_with_author_score.json |
Papers annotated with the author-score variable for the appendix ablation. |
Schemas
Paper (papers.json is a dict {topic: [paper, ...]}):
paper_id, title, abstract, year, venue, citation_count, url, open_access_pdf, authors, tier, topic, doi
Query + candidate set (experiment_sets.json is a dict {topic: [query_set, ...]}):
each query_set = {query_id, topic, query, candidates: {original: [...], flipped: [...], boosted: [...]}}; each candidate is a paper record plus condition and
authority_components: {venue, median_h, max_h, citations, affiliation, composite} (all min-max normalized within topic; composite uses the derived weights).
Response record (responses/open_weight_main.json is a list; the closed_weight/ and pilot_1n/ files wrap records under {"metadata": ..., "results": [...]}):
model, variant, condition, topic, query_id, query, recommended, recommended_paper_id, recommended_title, response, elapsed_seconds, prompt_length
(variant ∈ {baseline, anti_authority, content_first}; recommended is the 1-indexed picked candidate.)
Loading
The flat response files load directly with the datasets library:
from datasets import load_dataset
resp = load_dataset("json", data_files="responses/open_weight_main.json", split="train")
The nested files (papers.json, experiment_sets.json, experiment_conditions.json) are best read with plain json:
import json
sets = json.load(open("experiment_sets.json"))
Design summary
- 25 CS topics, 10 queries each (250 queries); 10 candidate papers per query.
- Conditions:
original(real metadata),flipped(high↔low authority swap),boosted(mid-tier inflation). - Instructions:
baseline,anti_authority(mild),content_first(strong). - Authority score:
0.353·venue + 0.292·median_h + 0.187·max_h + 0.137·citations + 0.031·affiliation(weights derived from the 1:N pilot via logistic regression + dominance analysis).
License
Released under CC BY 4.0. Paper metadata is sourced from Semantic Scholar and OpenAlex.
Citation
@inproceedings{jinadu2026authority,
title = {Authority Bias in Conversational Search Engines for Academic Paper Recommendation},
author = {Jinadu, Uthman and Ghazvinian, Parsa and Budathoki, Anjila and
Ampel, Benjamin M. and Sunderraman, Rajshekhar and Ding, Yi},
booktitle = {Proceedings of the 2026 Conference on Empirical Methods in Natural Language Processing (EMNLP)},
year = {2026}
}
- Downloads last month
- 21