Retrieval, Scoring, and Decoding Shape Performance and Stability in LLM-based Conversational Recommendation
Abstract.
Large language models (LLMs) are increasingly used as rerankers in conversational recommender systems, yet measured gains depend strongly on the retrieval and inference protocol. On the ReDial conversational movie recommendation benchmark, we compare proprietary, open-weight, and fine-tuned LLM rerankers with collaborative-filtering and sequential baselines in a shared retrieve-then-rerank pipeline. We vary candidate-pool size, first-stage retriever, and decoding temperature. With a shared semantic top-250 candidate pool and strict candidate-aware scoring, the best proprietary reranker reaches NDCG@10 of 0.1497, compared with 0.0939 for the strongest non-LLM baseline. The same reranker reaches 0.2925 in zero-shot generation, showing that unconstrained scoring can yield a much larger apparent advantage than matched-pool evaluation. No evaluated open-weight LLM outperforms the tuned shallow autoencoder baseline under this protocol. For the strongest proprietary and open-weight rerankers, switching from semantic to collaborative-filtering candidates raises NDCG@10 by more than 50%, showing that measured reranker performance is highly sensitive to candidate generation. For the best proprietary reranker, raising temperature from 0 to 1.0 increases top-10 Jaccard distance from 0.0900 to 0.1240 while mean NDCG@10 changes negligibly, whereas weaker LLMs show larger degradation. These ReDial results support treating candidate generation, candidate-pool size, scoring policy, and decoding configuration as required reporting fields rather than implementation details.
Keywords:
conversational recommender systems, large language models, reranking, candidate generation, evaluation, stability1. Introduction
Conversational recommender systems (CRS) infer preferences from dialogue rather than from dense interaction histories (Jannach et al., 2021). Large language models (LLMs) have shown promising zero-shot ranking ability on recommendation tasks (Hou et al., 2024) and can also be fine-tuned for the role (Bao et al., 2023), which is why recent CRS work increasingly relies on them at the ranking stage. Reported gains over earlier baselines are often substantial, but those numbers depend on pipeline choices that many tend to under-report, i.e., which items are retrieved as candidates, how many of them are shown to the LLM, whether evaluation restricts credit to the candidate set or scores any generated catalog title, and how the model is decoded. On a single benchmark, these choices can move ranking quality by more than the gap separating top systems, so the same model can appear strong or mediocre depending on how it is evaluated.
Two-stage retrieve-then-rerank is standard for LLM-based ranking (Sun et al., 2023), and recent LLM rerankers add user-preference retrieval or graph signals on top of an explicit candidate stage (Zhang et al., 2025; Wei et al., 2024). Yet, LLM CRS comparisons are seldom controlled for candidate generation, and reproducibility studies show that recommender comparisons are sensitive to such protocol choices as well as to baseline strength (Dacrema et al., 2019; Krichene and Rendle, 2020). Stochastic decoding further makes LLM outputs variable across repeated samples and sensitive to the order in which candidates appear in the prompt (Ma et al., 2025; Bito et al., 2026), yet list-level stability is rarely reported alongside accuracy in LLM CRS evaluations.
We study how these factors jointly shape observed effectiveness on ReDial (Li et al., 2018), a CRS benchmark of seeker-recommender movie dialogues in which preferences must be inferred from the conversation alone. We compare proprietary, open-weight, and fine-tuned LLM rerankers against collaborative-filtering (CF) and sequential rerankers in a shared two-stage pipeline11footnotetext: Equal contribution.11 1 Code, prompts, configurations, and outputs: https://github.com/infobip/crs-performance. Matched-pool comparisons hold the retrieved candidates fixed while we vary candidate-pool size, first-stage retriever, and decoding temperature. We answer the following research questions:
- RQ1
Under a matched semantic candidate pool, how do LLM rerankers compare to CF and sequential rerankers?
- RQ2
How sensitive is LLM reranking quality to candidate-pool size, from zero-shot generation to full-catalog reranking?
- RQ3
How does the choice of first-stage retriever (i.e., semantic, CF, or sequential) affect the quality of an LLM reranker?
- RQ4
How stable are LLM recommendation lists under decoding-temperature sampling, in terms of both ranking quality and list-level agreement?
Four proprietary LLMs significantly outperform EASE (0.0939), led by Claude Opus 4.6 at 0.1497, while every evaluated open-weight reranker falls below it. Expanding the semantic pool from 250 items to the full catalog raises proprietary NDCG@10 by 57–89%. Replacing semantic with EASE candidates raises NDCG@10 by 52–59% across the two tested rerankers. With higher temperature, Claude Opus 4.6 maintains mean accuracy while Jaccard distance@10 rises from 0.0900 to 0.1240, whereas for Llama-3.3-70B it rises from 0.0230 to 0.7600. These results identify retrieval strategy, candidate-pool size, scoring policy, and decoding configuration as core experimental variables.
2. Methodology
2.1. Dataset and Task
We evaluate on ReDial’s standard test split (1,025 dialogues) and 6,924-movie catalog (Li et al., 2018). For each dialogue, accepted recommendations are masked and used as ground-truth targets. Each model returns ten ranked movie titles from family-specific inputs.
LLM prompts contain liked titles, the masked dialogue, and, outside zero-shot generation, a shuffled candidate list. We label settings by candidate count: c denotes candidates, c0 no candidate list, and cAll the full catalog. Candidate-aware prompts require outputs to use only listed titles. For CF models, liked movies form the implicit interaction history, while accepted recommendations define validation and test targets. For sequential models, we preserve item order by placing liked movies before accepted recommendations and expanding these histories into the pre-augmented RecBole sequence format (Zhao et al., 2021). The targets stay fixed while the input representation matches each model family. Matching pools controls item availability but not model information: LLMs use raw dialogue and pretrained knowledge, whereas CF and sequential models use interaction histories. Our comparisons therefore evaluate systems rather than reranking under identical information.
2.2. Two-stage Pipeline and Candidate Generation
We use a two-stage pipeline with two roles: a candidate generator returns a pool of catalog items for dialogue , and a reranker orders items from this pool and emits the top-10 recommendation list. The same model can fill either role. In our experiments, CF and sequential models act as rerankers over a matched semantic pool and also produce top- pools that LLMs then rerank.
The primary candidate generator is content-based filtering (CBF) over movie metadata. We embed each catalog item with all-mpnet-base-v2, a Sentence-BERT model (Reimers and Gurevych, 2019), and L2-normalize the resulting vectors . For dialogue with liked set and disliked set , we form centroids
| (1) |
and score each catalog item by
| (2) |
with and . The negative term is dropped when is empty. The top- items by form , which we shuffle before insertion into the <CANDIDATES> block of LLM prompts to reduce prompt-position effects.
CBF pools of size drive the main reranker comparison. For candidate-pool sensitivity, we also evaluate zero-shot generation (no candidate list), CBF pools at , and full-catalog reranking, where all 6,924 catalog titles are placed in a single prompt. Additionally, we build top-250 pools with EASE as the CF retriever and SASRec as the sequential retriever for the same LLM rerankers, isolating first-stage effects.
2.3. Rerankers, Inference, and Evaluation
Rerankers
We compare proprietary API LLMs, open-weight LLMs, and a fine-tuned open-weight LLM against CF and sequential rerankers. LLM rerankers receive the dialogue context and candidate titles then generate an ordered list of movie titles. CF rerankers score the same candidate items using collaborative signals learned from the RecBole interaction data (Zhao et al., 2021). Sequential rerankers score candidates using the corresponding ordered dialogue/user sequence. We also include the unreranked CBF order and a popularity baseline. The fine-tuned reranker is Qwen2.5-7B-Instruct adapted with LoRA (rank 16, , 3 epochs, learning rate ) on the c250 training prompts. For the main reranking runs, LLMs are decoded with temperature 0 and top-p left at its default of 1.0. The repository contains full prompt templates, exact provider model identifiers, run configurations, outputs, and parsing code.
Metrics
Our primary metric is NDCG@10 (Järvelin and Kekäläinen, 2002) with binary relevance over the accepted target movies. We also report Hit@10, item coverage@10, and average training-set popularity@10. LLM outputs are parsed as ordered title lists and matched to ReDial catalog items after Unicode-normalized, lowercased title matching with trailing punctuation stripped. Titles that cannot be matched to a catalog item are kept in the raw artifact but receive no metric credit. In candidate-constrained settings, the prompt instructs the LLM to choose only from the candidate list. Generated titles outside that list receive no metric credit, even if they match a catalog item.
Retrieval diagnostics and significance
CandRecall@250 is the mean fraction of ground-truth items retrieved into the top-250 candidate pool. Oracle NDCG@10 is the best NDCG@10 attainable by a reranker that can only rank items from that pool. Confidence intervals use bootstrap resampling over dialogue-level examples. For the main reranker comparison, we test NDCG@10 against EASE with paired Wilcoxon signed-rank tests and Holm correction (Holm, 1979).
2.4. Temperature sensitivity and list stability
For each LLM reranker we run 20 generations per prompt at each temperature in over a fixed 30-prompt subset of the c250 test set, with top-. Anthropic models (Claude Opus 4.6, Claude Sonnet 4.6) are capped at by the provider. We report mean NDCG@10 and dialogue-level variation across repeated generations. List changes are measured with Jaccard distance@10 over the top-10 sets and position disagreement@10, the mean fraction of rank-aligned positions whose items differ across paired generations.
3. Results
The results show three main patterns. Proprietary LLMs lead under a fixed semantic candidate pool, first-stage retrieval can change NDCG@10 as much as model choice, and decoding temperature affects list stability more than average ranking quality.
| NDCG@10 | Hit@10 | Cov.@10 | Pop.@10 | |||||
| Reranker | c0 | c250 | c0 | c250 | c0 | c250 | c0 | c250 |
| Semantic baseline | ||||||||
| Semantic rank (Reimers and Gurevych, 2019) | – | 0.0079 | – | 0.0260 | – | 0.4018 | – | 13.25 |
| Proprietary LLMs | ||||||||
| Claude Opus 4.6 (Anthropic, 2026a) | 0.2925∗ | 0.1497∗ | 0.4875 | 0.2713 | 0.2041 | 0.2639 | 52.12 | 42.80 |
| GPT-5.2 (OpenAI, 2025b) | 0.2427∗ | 0.1335∗ | 0.4094 | 0.2523 | 0.2163 | 0.2858 | 44.63 | 36.64 |
| Claude Sonnet 4.6 (Anthropic, 2026b) | 0.2207∗ | 0.1325∗ | 0.3814 | 0.2462 | 0.2169 | 0.2754 | 45.22 | 37.79 |
| GPT-4.1 (OpenAI, 2025a) | 0.1983∗ | 0.1283∗ | 0.3453 | 0.2322 | 0.2185 | 0.3037 | 40.90 | 34.22 |
| GPT-4.1 Mini (OpenAI, 2025a) | 0.1937∗ | 0.1021 | 0.3554 | 0.1912 | 0.1999 | 0.3033 | 52.09 | 32.69 |
| Open-weight LLMs | ||||||||
| Qwen2.5-7B-FT (Yang et al., 2024) | – | 0.0792 | – | 0.1502 | – | 0.2711 | – | 23.22 |
| Llama-3.3-70B (Meta, 2024b) | 0.1154 | 0.0770 | 0.2503 | 0.1602 | 0.2114 | 0.2607 | 47.86 | 33.69 |
| Gemma-2-9B (Gemma Team, 2024) | 0.1356∗ | 0.0650 | 0.2913 | 0.1281 | 0.1932 | 0.3068 | 47.87 | 26.43 |
| Llama-3.1-8B (Grattafiori et al., 2024) | 0.0815 | 0.0503 | 0.1992 | 0.1011 | 0.1749 | 0.2402 | 40.68 | 27.03 |
| Qwen2.5-7B (Yang et al., 2024) | 0.0622 | 0.0500 | 0.1952 | 0.1071 | 0.2064 | 0.2620 | 57.75 | 29.62 |
| Llama-3.2-3B (Meta, 2024a) | 0.0603 | 0.0278 | 0.1622 | 0.0571 | 0.1989 | 0.2939 | 40.72 | 20.93 |
| Collaborative filtering | ||||||||
| EASE (Steck, 2019) | – | 0.0939 | – | 0.2072 | – | 0.2370 | – | 59.99 |
| ItemKNN (Sarwar et al., 2001) | – | 0.0876 | – | 0.1982 | – | 0.3258 | – | 43.86 |
| LightGCN (He et al., 2020) | – | 0.0833 | – | 0.2042 | – | 0.2139 | – | 59.96 |
| BPR (Rendle et al., 2009) | – | 0.0826 | – | 0.1892 | – | 0.2096 | – | 63.60 |
| Sequential models | ||||||||
| SASRec (Kang and McAuley, 2018) | – | 0.0705 | – | 0.1802 | – | 0.1592 | – | 67.96 |
| GRU4Rec (Hidasi et al., 2016) | – | 0.0703 | – | 0.1792 | – | 0.2301 | – | 57.82 |
| NARM (Li et al., 2017) | – | 0.0686 | – | 0.1612 | – | 0.2402 | – | 58.82 |
| SRGNN (Wu et al., 2019) | – | 0.0615 | – | 0.1592 | – | 0.1853 | – | 64.44 |
| Popularity baseline | ||||||||
| Pop (Zhao et al., 2021) | – | 0.0388 | – | 0.1061 | – | 0.0477 | – | 107.68 |
Cov. is item coverage of the top-10 list. Pop. is mean training-set popularity of top-10 items. ∗ marks NDCG@10 significantly above the EASE c250 baseline by paired Wilcoxon signed-rank test with Holm correction (). c0 and c250 LLM scores are tested separately against that baseline. CF, sequential, and popularity baselines require a candidate set and have no c0 entry. Qwen2.5-7B-FT was fine-tuned on c250 prompts and is reported only for that policy.
RQ1.
Under the matched semantic top-250 pool (Table 1), Claude Opus 4.6 achieves the highest strict NDCG@10 at 0.1497, with GPT-5.2, Claude Sonnet 4.6, and GPT-4.1 clustered behind between 0.1283 and 0.1335. These four proprietary models are the only rerankers significantly above EASE (0.0939) after Holm correction. GPT-4.1 Mini is numerically higher than EASE but not significant, and every open-weight reranker falls below EASE. Zero-shot scoring changes this picture. Claude Opus 4.6 reaches 0.2925, nearly twice its strict c250 score, and additional models become significant against the EASE baseline. Table 1 also shows differences in coverage and popularity. Under c250, GPT-4.1 and GPT-4.1 Mini have the highest coverage among proprietary LLMs, while EASE has lower coverage and higher average popularity. The strongest proprietary rerankers improve NDCG without relying only on the most popular items. Open-weight models show mixed behavior: some cover a broad part of the catalog but still rank less accurately.
RQ2.
Figure 1 shows how candidate access reshapes the same rerankers. Every proprietary model gains substantially as the pool expands. Strict NDCG@10 rises by 57% (GPT-4.1 Mini) to 89% (GPT-5.2) between c250 and cAll, and cAll is the best strict setting for each. c0, scored without a candidate list, lands in the same range and slightly exceeds cAll for Opus (0.2925 vs 0.2644) and GPT-4.1 Mini (0.1937 vs 0.1600). Open-weight rerankers do not benefit similarly, with NDCG dropping from c250 to c500 for every evaluated open-weight model. Because pool size simultaneously changes target availability, candidate composition, prompt length, and ranking difficulty, this experiment measures their combined pipeline effect rather than isolating reranking ability.
| NDCG@10 | ||||
|---|---|---|---|---|
| Retriever | CandR@250 | Oracle | Claude Opus 4.6 | Llama-3.3-70B |
| Semantic | 0.2877 | 0.3165 | 0.1497 | 0.0770 |
| EASE | 0.4862 | 0.5222 | 0.2277 | 0.1223 |
| SASRec | 0.5093 | 0.5423 | 0.2119 | 0.1180 |
CandR@250 is the mean fraction of ground-truth movies retrieved into the candidate pool. Oracle NDCG@10 is the best NDCG@10 achievable by a perfect reranker restricted to that pool, and upper-bounds the downstream reranker columns.
RQ3.
Table 2 changes the first-stage retriever while holding candidate-pool size and reranker fixed. Semantic retrieval places only 28.8% of relevant items in the top-250 pool, versus 48.6% for EASE and 50.9% for SASRec. Oracle NDCG@10 follows the same order. Retriever choice carries through to the reranker. Switching from semantic to EASE lifts Claude Opus 4.6 by 52% (0.1497 to 0.2277) and Llama-3.3-70B by 59% (0.0770 to 0.1223). SASRec has the highest recall and oracle ceiling, while EASE yields the best NDCG for both rerankers. Retrieval opportunity alone therefore does not determine downstream performance, though this design does not separate candidate composition from reranker–pool compatibility.
RQ4.
Figure 2 reports the stability subset of 30 prompts, so absolute NDCG levels are not comparable to the full-test numbers in Table 1. Across API models, mean strict NDCG@10 is roughly flat in temperature while Jaccard distance@10 rises with each step. The Anthropic models cap at by the provider, so their traces stop there. Claude Opus 4.6 is the most stable. Its Jaccard distance moves from 0.0900 at to 0.1240 at , position disagreement@10 rises from 0.2700 to 0.3500, and NDCG@10 changes little. Open-weight rerankers are less stable. Llama-3.3-70B begins almost deterministic at (Jaccard 0.0230, position disagreement 0.0250) and reaches 0.7600 and 0.7930 at . Its NDCG@10 also falls from 0.0750 to 0.0490.
4. Discussion and Conclusion
Retrieval choices are as important as the reranker.
Switching first-stage retrievers improves NDCG@10 by 52–59% across both rerankers, while larger candidate pools also substantially improve proprietary models. Yet the higher-recall SASRec pool trails EASE after reranking, showing that candidate availability, composition, and reranker–pool compatibility jointly determine downstream quality. Treating candidate generation as an implementation detail thus conflates pipeline and reranker quality.
Candidate access and scoring policy shape headline LLM gains.
Claude Opus 4.6 reaches 0.2925 in zero-shot generation but 0.1497 under strict candidate-aware scoring on the matched semantic top-250 pool. Because c0 and c250 differ in both candidate access and scoring policy, this contrast does not isolate either factor. It instead shows how strongly the evaluation protocol affects the apparent advantage over EASE. Under the matched c250 pool, only four proprietary models significantly outperform EASE, while every open-weight reranker falls below it. These are system-level comparisons rather than tests of reranking under identical information.
Stability is a separate axis from accuracy.
Mean accuracy and list identity respond differently to decoding temperature. For the strongest proprietary models, NDCG@10 changes little between temperature 0 and 1, while Jaccard distance@10 and position disagreement@10 increase. The model can therefore return lists with similar accuracy but different items and rankings. Llama-3.3-70B is more temperature-sensitive, with Jaccard distance@10 rising from 0.0230 to 0.7600 and NDCG@10 falling from 0.0750 to 0.0490. Reporting only mean NDCG hides this deployment-relevant behavior.
Scope and future work.
Our experiments use ReDial, a single movie-domain CRS benchmark with a 6,924-item catalog. Future work will test whether these findings generalize to larger catalogs, other domains, and conversational settings. We did not systematically measure cost or latency. Future evaluations should compare quality, cost, and latency, particularly for large and full-catalog candidate pools.
Toward a reporting norm.
LLM-based CRS evaluations should report the retriever, candidate-pool size, scoring policy for off-candidate and unmatched generations, decoding configuration, prompt template, and exact model and provider identifier. Headline results should be interpreted together with this protocol because pipeline choices can affect measured gains as much as model choice.
Acknowledgements.
This research was supported in part by the project Infobip Global Communication Platform (PK.1.1.07.0001), part of the Important Project of Common European Interest on Next Generation Cloud Infrastructure and Services (IPCEI-CIS) consortium.GenAI Usage Disclosure
Generative AI tools assisted in a supporting capacity: Claude Code (Opus 4.7) for code implementation and data analysis, ChatGPT 5.5 for LaTeX editing and grammar checking, and Google Scholar Labs for related-work identification. All AI-assisted outputs were reviewed and verified by the authors, who take full responsibility for the work.
References
- Claude Opus 4.6 System Card. Anthropic. Note: Technical reportAccessed: 2026-06-04 External Links: Link Cited by: Table 1.
- Claude Sonnet 4.6 System Card. Anthropic. Note: Technical reportAccessed: 2026-06-05 External Links: Link Cited by: Table 1.
- TALLRec: an effective and efficient tuning framework to align large language model with recommendation. In Proceedings of the 17th ACM Conference on Recommender Systems, RecSys ’23, New York, NY, USA, pp. 1007–1014. External Links: Document, Link Cited by: §1.
- One pass, any order: position-invariant listwise reranking for LLM-based recommendation. Note: arXiv preprint arXiv:2604.27599To appear in SIGIR ’26 External Links: 2604.27599, Document, Link Cited by: §1.
- Are we really making much progress? a worrying analysis of recent neural recommendation approaches. In Proceedings of the 13th ACM Conference on Recommender Systems, RecSys ’19, New York, NY, USA, pp. 101–109. External Links: Document, Link Cited by: §1.
- Gemma 2: improving open language models at a practical size. Note: arXiv preprint arXiv:2408.00118 External Links: 2408.00118, Document, Link Cited by: Table 1.
- The llama 3 herd of models. Note: arXiv preprint arXiv:2407.21783 External Links: 2407.21783, Document, Link Cited by: Table 1.
- LightGCN: simplifying and powering graph convolution network for recommendation. In Proceedings of the 43rd International ACM SIGIR Conference on Research and Development in Information Retrieval, SIGIR ’20, New York, NY, USA, pp. 639–648. External Links: Document, Link Cited by: Table 1.
- Session-based recommendations with recurrent neural networks. Note: 4th International Conference on Learning Representations (ICLR 2016), San Juan, Puerto Rico External Links: 1511.06939, Link Cited by: Table 1.
- A simple sequentially rejective multiple test procedure. Scandinavian Journal of Statistics 6 (2), pp. 65–70. Cited by: §2.3.
- Large language models are zero-shot rankers for recommender systems. In Advances in Information Retrieval, Lecture Notes in Computer Science, Vol. 14609, Cham, Switzerland, pp. 364–381. External Links: Document, Link Cited by: §1.
- A survey on conversational recommender systems. ACM Computing Surveys 54 (5), pp. 1–36. External Links: Document, Link Cited by: §1.
- Cumulated gain-based evaluation of IR techniques. ACM Transactions on Information Systems 20 (4), pp. 422–446. External Links: Document, Link Cited by: §2.3.
- Self-attentive sequential recommendation. In 2018 IEEE International Conference on Data Mining, Piscataway, NJ, USA, pp. 197–206. External Links: Document, Link Cited by: Table 1.
- On sampled metrics for item recommendation. In Proceedings of the 26th ACM SIGKDD International Conference on Knowledge Discovery & Data Mining, KDD ’20, New York, NY, USA, pp. 1748–1757. External Links: Document, Link Cited by: §1.
- Neural attentive session-based recommendation. In Proceedings of the 2017 ACM on Conference on Information and Knowledge Management, CIKM ’17, New York, NY, USA, pp. 1419–1428. External Links: Document, Link Cited by: Table 1.
- Towards deep conversational recommendations. In Advances in Neural Information Processing Systems, Vol. 31, Red Hook, NY, USA, pp. 9748–9758. External Links: Link Cited by: §1, §2.1.
- Large language models are not stable recommender systems: a position bias perspective. In Knowledge Science, Engineering and Management, Lecture Notes in Computer Science, Vol. 15919, Singapore, pp. 415–429. External Links: Document, Link Cited by: §1.
- Llama 3.2 3B Instruct Model Card. Hugging Face. Note: Model cardAccessed: 2026-06-06 External Links: Link Cited by: Table 1.
- Llama 3.3 70B Instruct Model Card. Hugging Face. Note: Model cardAccessed: 2026-06-05 External Links: Link Cited by: Table 1.
- Introducing GPT-4.1 in the API. OpenAI. Note: OpenAI blogAccessed: 2026-06-05 External Links: Link Cited by: Table 1, Table 1.
- Update to GPT-5 system card: GPT-5.2. OpenAI. Note: Technical reportAccessed: 2026-06-04 External Links: Link Cited by: Table 1.
- Sentence-BERT: sentence embeddings using siamese BERT-networks. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing, Hong Kong, China, pp. 3982–3992. External Links: Document, Link Cited by: §2.2, Table 1.
- BPR: bayesian personalized ranking from implicit feedback. In Proceedings of the Twenty-Fifth Conference on Uncertainty in Artificial Intelligence, UAI ’09, Arlington, Virginia, USA, pp. 452–461. External Links: Link Cited by: Table 1.
- Item-based collaborative filtering recommendation algorithms. In Proceedings of the 10th International Conference on World Wide Web, WWW ’01, New York, NY, USA, pp. 285–295. External Links: Document, Link Cited by: Table 1.
- Embarrassingly shallow autoencoders for sparse data. In The World Wide Web Conference, WWW ’19, New York, NY, USA, pp. 3251–3257. External Links: Document, Link Cited by: Table 1.
- Is ChatGPT good at search? investigating large language models as re-ranking agents. In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, Singapore, pp. 14918–14937. External Links: Document, Link Cited by: §1.
- LLMRec: large language models with graph augmentation for recommendation. In Proceedings of the 17th ACM International Conference on Web Search and Data Mining, WSDM ’24, New York, NY, USA, pp. 806–815. External Links: Document, Link Cited by: §1.
- Session-based recommendation with graph neural networks. In Proceedings of the AAAI Conference on Artificial Intelligence, Vol. 33, Palo Alto, CA, USA, pp. 346–353. External Links: Document, Link Cited by: Table 1.
- Qwen2.5 technical report. Note: arXiv preprint arXiv:2412.15115 External Links: 2412.15115, Document, Link Cited by: Table 1, Table 1.
- Enhancing reranking for recommendation with LLMs through user preference retrieval. In Proceedings of the 31st International Conference on Computational Linguistics, Abu Dhabi, UAE, pp. 658–671. External Links: Link Cited by: §1.
- RecBole: towards a unified, comprehensive and efficient framework for recommendation algorithms. In Proceedings of the 30th ACM International Conference on Information & Knowledge Management, CIKM ’21, New York, NY, USA, pp. 4653–4664. External Links: Document, Link Cited by: §2.1, §2.3, Table 1.