Dataset and evaluation results from a study on whether fictional-character personas improve LLMs' recognition of emotional labor / emotion regulation strategies (surface acting, deep acting, genuine expression) in first-person emotional narratives.
data/
emotion_labor_dataset.json Main MCQ dataset (502 items)
character_selection/
all_baps_ranked_by_sd.csv All Behavioral Adjective Pairs (BAPs), ranked by
cross-character standard deviation
goldberg_matches_full.csv BAPs matched to Goldberg (1992) Big Five markers
suggested_final_baps.csv Final BAP set used for persona construction
results/
bap_persona/<model>/ Eval results using BAP-derived personas
ipip_persona/<model>/ Eval results using IPIP-50-derived personas
Each results/{bap_persona,ipip_persona}/<model>/per_character/<char_id>.json holds the
full per-item results for one fictional character persona.
Models evaluated: gpt (GPT-5.4), deepseek (DeepSeek-V4-Flash), gemma (Gemma-4-31B-it),
qwen32b (Qwen3-32B), qwen8b (Qwen3-8B).
Note:
results/bap_persona/gpt/contains 100 character personas, while every other model/persona combination contains 50.1
Each item in data/emotion_labor_dataset.json is built from an ISEAR-derived
emotional narrative, extended with an added social context, and paired with
three response options corresponding to different emotion regulation strategies:
{
"sentence_id": 271,
"felt_emotion": "Fear",
"original_sentence": "I felt ... when my 2 year old broke her leg, ...",
"modified_sentence": "... (with added social context) ...",
"stem": "... (prompt shown to the model) ...",
"options": {
"A": {"text": "...", "category": "genuine_expression"},
"B": {"text": "...", "category": "deep_acting"},
"C": {"text": "...", "category": "surface_acting"}
}
}surface_acting— performs a different emotion outwardly while an involuntary leakage cue is presentdeep_acting— genuinely shifts internal state via cognitive reframing, no fakinggenuine_expression— direct, unregulated expression
Each entry in per_character/<char_id>.json["results"] records the model's
choice against the correct label, with options shuffled per item:
{
"char_id": "AS/3",
"char_name": "James Taggart",
"sentence_id": 271,
"felt_emotion": "Fear",
"correct_new_label": "B",
"chosen_new_label": "B",
"chosen_category": "surface_acting",
"is_correct": true
}- BAP personas: characters selected via greedy maximin sampling over a normalized Behavioral Adjective Pair (BAP) subspace, cross-referenced against Goldberg (1992) Big Five bipolar adjective markers.
- IPIP-50 personas: the same characters administered the IPIP-50 personality inventory in-character, with responses summarized into behavioral trait paragraphs.
Both persona types are prepended to the MCQ prompt before evaluation.
Released under CC BY 4.0.
If you use this dataset, please cite our paper (citation to be added on publication).
Footnotes
-
The extended 100-character GPT run was a verification pass to check whether trait–strategy correlations held up on a larger, more diverse character sample. The correlations reported in the paper hold on this extended set, so the smaller 50-character runs remain representative for the other models. ↩