Behavior Retrieval plus Response Generation for Interpretable Conversational Personalized Recommendation

  • Qi Xin University of Pittsburgh, United States of America
DOI: https://doi.org/10.31258/ijeepse.9.2.120-136
Abstract viewed: 0 times
pdf downloaded: 0 times
Keywords: Behavior retrieval, conversational recommender systems, explainable recommendation, next-basket prediction, personalized recommendation.

Abstract

This paper presents an empirical study of interpretable conversational personalized recommendation built from two tightly coupled layers: a behavior retrieval layer for candidate ranking and a response generation layer for grounded recommendation wording. The study is motivated by recent large language model surveys in recommendations, but the implementation remains deterministic so that candidate selection, evidence extraction, and wording can be inspected separately. We used the 2010-2011 partition of Online Retail II for transaction-grounded next-basket recommendation and a retail/service subset of the Bitext customer-support 27K corpus for conversational response generation. The behavior retrieval layer combined popularity, repeat purchase memory, recency-weighted memory, item-to-item collaborative filtering, and user-neighborhood scoring. The response layer compared intent-majority, TF-IDF retrieval, intent-conditioned retrieval, and a constrained fusion verbalizer. On 1,999 test users, the hybrid behavior retriever achieved Hit@10 = 0.7282, Recall@10 = 0.1982, NDCG@10 = 0.2855, and MRR@10 = 0.4467, outperforming the strongest baseline RecencyRepeat by 2.03% on Hit@10 and 6.34% on NDCG@10. On 1,989 Bitext test utterances, intent-conditioned retrieval produced the best standalone verbalization quality with BLEU-4 = 0.1699, ROUGE-L = 0.3431, intent accuracy = 0.9915, and slot F1 = 0.8239. Grounded recommendation responses reached Top1Hit = 0.3252, HistoryGroundRate = 1.0000, and BundleValidity = 0.9980. The results show that retrieval-first conversational recommendation can combine ranking accuracy, grounding faithfulness, and low-overhead response realization without relying on unconstrained generation.

References

P. Resnick, N. Iacovou, M. Suchak, P. Bergstrom, and J. Riedl, "GroupLens: An open architecture for collaborative filtering of netnews," in Proc. ACM Conf. Computer Supported Cooperative Work, 1994, pp. 175-186.

J. B. Schafer, J. A. Konstan, and J. Riedl, "E-commerce recommendation applications," Data Min. Knowl. Discov., vol. 5, no. 1-2, pp. 115-153, 2001.

B. Sarwar, G. Karypis, J. Konstan, and J. Riedl, "Item-based collaborative filtering recommendation algorithms," in Proc. 10th Int. Conf. World Wide Web, 2001, pp. 285-295.

R. Burke, "Hybrid recommender systems: Survey and experiments," User Model. User-Adapt. Interact., vol. 12, no. 4, pp. 331-370, 2002.

K. Järvelin and J. Kekäläinen, "Cumulated gain-based evaluation of IR techniques," ACM Trans. Inf. Syst., vol. 20, no. 4, pp. 422-446, 2002.

K. Papineni, S. Roukos, T. Ward, and W.-J. Zhu, "BLEU: A method for automatic evaluation of machine translation," in Proc. 40th Annu. Meeting Assoc. Comput. Linguistics, 2002, pp. 311-318.

C.-Y. Lin, "ROUGE: A package for automatic evaluation of summaries," in Proc. ACL Workshop Text Summarization Branches Out, 2004, pp. 74-81.

J. L. Herlocker, J. A. Konstan, and J. Riedl, "Explaining collaborative filtering recommendations," in Proc. ACM Conf. Computer Supported Cooperative Work, 2000, pp. 241-250.

N. Tintarev and J. Masthoff, "A survey of explanations in recommender systems," in Proc. IEEE 23rd Int. Conf. Data Eng. Workshop, 2007, pp. 801-810.

S. Rendle, C. Freudenthaler, Z. Gantner, and L. Schmidt-Thieme, "BPR: Bayesian personalized ranking from implicit feedback," in Proc. 25th Conf. Uncertainty in Artificial Intelligence, 2009, pp. 452-461.

B. Hidasi, A. Karatzoglou, L. Baltrunas, and D. Tikk, "Session-based recommendations with recurrent neural networks," in Proc. Int. Conf. Learn. Representations, 2016.

X. He, L. Liao, H. Zhang, L. Nie, X. Hu, and T.-S. Chua, "Neural collaborative filtering," in Proc. 26th Int. Conf. World Wide Web, 2017, pp. 173-182.

R. Li, S. Kahou, H. Schulz, V. Michalski, L. Charlin, and C. Pal, "Towards deep conversational recommendations," in Proc. 32nd Conf. Neural Information Processing Systems, 2018, pp. 9748-9758.

Y. Zhang and X. Chen, "Explainable recommendation: A survey and new perspectives," Found. Trends Inf. Retr., vol. 14, no. 1, pp. 1-101, 2020.

W. Lei, X. He, Y. Miao, Q. Wu, R. Hong, M.-Y. Kan, and T.-S. Chua, "Estimation-Action-Reflection: Towards deep interaction between conversational and recommender systems," in Proc. 13th ACM Int. Conf. Web Search and Data Mining, 2020, pp. 304-312.

D. Jannach, A. Manzoor, W. Cai, and L. Chen, "A survey on conversational recommender systems," ACM Comput. Surv., vol. 54, no. 5, Art. no. 105, pp. 1-36, 2021.

L. Wu, Z. Zheng, Z. Qiu, H. Wang, H. Gu, T. Shen, C. Qin, C. Zhu, H. Zhu, Q. Liu, H. Xiong, and E. Chen, "A survey on large language models for recommendation," World Wide Web, vol. 27, no. 5, Art. no. 60, 2024.

J. Lin, X. Dai, Y. Xi, W. Liu, B. Chen, H. Zhang, Y. Liu, C. Wu, X. Li, C. Zhu, H. Guo, Y. Yu, R. Tang, and W. Zhang, "How can recommender systems benefit from large language models: A survey," ACM Trans. Inf. Syst., 2024.

Z. Zhao, W. Fan, J. Li, Y. Liu, X. Mei, Y. Wang, Z. Wen, F. Wang, X. Zhao, J. Tang, and Q. Li, "Recommender systems in the era of large language models (LLMs)," arXiv preprint arXiv:2307.02046, 2023.

L. Li, Y. Zhang, D. Liu, and L. Chen, "Large language models for generative recommendation: A survey and visionary discussions," in Proc. Joint Int. Conf. Computational Linguistics, Language Resources and Evaluation, 2024, pp. 10146-10159.

Y. Gao, T. Sheng, Y. Xiang, Y. Xiong, H. Wang, and J. Zhang, "Chat-REC: Towards interactive and explainable LLMs-augmented recommender system," arXiv preprint arXiv:2303.14524, 2023.

D. Chen, "Online Retail II," UCI Machine Learning Repository, 2012, doi: 10.24432/C5CG6D.

Bitext, "Customer-support-llm-chatbot-training-dataset," GitHub repository, 2024.

M. Li, S. Jullien, M. Ariannezhad, and M. de Rijke, "A next basket recommendation reality check," ACM Trans. Inf. Syst., vol. 41, no. 4, Art. no. 116, pp. 1-29, 2023, doi: 10.1145/3587153.

M. Ariannezhad, S. Jullien, M. Li, M. Fang, S. Schelter, and M. de Rijke, "ReCANet: A repeat consumption-aware neural network for next basket recommendation in grocery shopping," in Proc. 45th Int. ACM SIGIR Conf. Research and Development in Information Retrieval, 2022, pp. 1240-1250, doi: 10.1145/3477495.3531708.

Z. Shao, S. Wang, Q. Zhang, W. Lu, Z. Li, and X. Peng, "An empirical study of next-basket recommendations," arXiv preprint arXiv:2312.02550, 2023.

D. Jannach, "Evaluating conversational recommender systems," Artif. Intell. Rev., vol. 56, pp. 2365-2400, 2023, doi: 10.1007/s10462-022-10229-x.

Published
2026-07-03
How to Cite
[1]
Q. Xin, “Behavior Retrieval plus Response Generation for Interpretable Conversational Personalized Recommendation”, IJEEPSE, vol. 9, no. 2, pp. 120-136, Jul. 2026.