Reward-Aligned Reinforcement Learning from Human Feedback for Emotion-Sensitive Large Language Model Therapists: Balancing Empathetic Engagement, Boundary Safety, and Clinical Accountability

Notice

This is an unedited manuscript accepted for publication and provided as an Article in Press for early access at the author’s request. The article will undergo copyediting, typesetting, and galley proof review before final publication. Please be aware that errors may be identified during production that could affect the content. All legal disclaimers of the journal apply.

Year : 2026 | Volume : 3 | 02 | Page :
By

Vaibhav Baishkhiyar,

Ankit Kumar Tiwari,

Hardik Srivastava,

Shikhar Gupta,

  1. Student, Department of Computer Sciences And Engineering Greater Noida Institute of Technology Greater Noida, Uttar Pradesh, India
  2. Student, Department of computer Science and Engineering Greater Noida Institute of Technology Greater Noida, Uttar Pradesh, India
  3. Student, Department of Computer science and Engineering Greater Noida Institute of Technology Greater Noida, Uttar Pradesh, India
  4. Student, Department of Computer Science and Engineering Greater Noida Institute of Technology Greater Noida, Uttar Pradesh, India

Abstract

The deployment of large language models (LLMs) in mental health therapy contexts introduces a critical alignment challenge: these systems must simultaneously cultivate genuine empathic rapport, observe clinically grounded safety boundaries, and remain auditable under institutional accountability frameworks. Existing reinforcement learning from human feedback (RLHF) pipelines optimize for a scalar reward signal that is demonstrably insufficient for the multi-objective, temporally extended nature of therapeutic conversation. This paper presents RA-RLHF-T (Reward-Aligned RLHF for Therapy), a novel framework that decomposes the therapeutic reward function into three semantically orthogonal sub-signals: an Empathic Resonance Score (ERS), a Boundary Adherence Penalty (BAP), and a Clinical Accountability Index (CAI). We introduce a Constrained Multi-Objective Policy Gradient (CMPG) algorithm that treats BAP as a hard Lagrangian constraint and jointly maximizes ERS and CAI under a Pareto-optimal learning objective. A dedicated Emotion State Encoder (ESE) integrates multimodal cues—lexical affect, syntactic distress patterns, and conversational momentum—into a continuous latent emotion manifold that conditions both the reward model and the policy. Experimental evaluation on two benchmark corpora (EmpatheticDialogues-Clinical and SafeTherapyBench) demonstrates that RA-RLHF-T achieves a 23.4% improvement in ERS over standard RLHF baselines while reducing unsafe boundary violations by 61.2% and improving clinical traceability scores by 18.7%. Our framework represents a principled, theoretically grounded approach to deploying LLMs as responsible adjunct mental health assistants.

Keywords: Reinforcement Learning from Human Feedback; Large Language Models; Mental Health AI; Multi-Objective Reward Optimization; Empathetic Dialogue Systems; Clinical Safety; Constrained Policy Gradient; Emotion-Sensitive NLP

How to cite this article: Vaibhav Baishkhiyar, Ankit Kumar Tiwari, Hardik Srivastava, Shikhar Gupta. Reward-Aligned Reinforcement Learning from Human Feedback for Emotion-Sensitive Large Language Model Therapists: Balancing Empathetic Engagement, Boundary Safety, and Clinical Accountability. International Journal of Tropical Medicines. 2026; 03(02):-.
How to cite this URL: Vaibhav Baishkhiyar, Ankit Kumar Tiwari, Hardik Srivastava, Shikhar Gupta. Reward-Aligned Reinforcement Learning from Human Feedback for Emotion-Sensitive Large Language Model Therapists: Balancing Empathetic Engagement, Boundary Safety, and Clinical Accountability. International Journal of Tropical Medicines. 2026; 03(02):-. Available from: https://journals.stmjournals.com/ijtm/article=2026/view=258993

References

World Health Organization. World mental health report: transforming mental health for all. Executive summary. World Health Organization; 2022 Jun 16. [2] Christiano PF, Leike J, Brown T, Martic M, Legg S, Amodei D. Deep reinforcement learning from human preferences. Advances in neural information processing systems. 2017;30. [3] Stiennon N, Ouyang L, Wu J, Ziegler D, Lowe R, Voss C, Radford A, Amodei D, Christiano PF. Learning to summarize with human feedback. Advances in neural information processing systems. 2020;33:3008-21. [4] Solaiman I, Dennison C. Process for adapting language models to society (palms) with values-targeted datasets. Advances in Neural Information Processing Systems. 2021 Dec 6;34:5861-73. [5] Schulman J, Wolski F, Dhariwal P, Radford A, Klimov O. Proximal policy optimization algorithms. arXiv preprint arXiv:1707.06347. 2017 Jul 20. [6] Ouyang L, Wu J, Jiang X, Almeida D, Wainwright C, Mishkin P, Zhang C, Agarwal S, Slama K, Ray A, Schulman J. Training language models to follow instructions with human feedback. Advances in neural information processing systems. 2022 Dec 6;35:27730-44. [7] Bai Y, Kadavath S, Kundu S, Askell A, Kernion J, Jones A, Chen A, Goldie A, Mirhoseini A, McKinnon C, Chen C. Constitutional ai: Harmlessness from ai feedback. arXiv preprint arXiv:2212.08073. 2022 Dec 15. [8] Bai Y, Jones A, Ndousse K, Askell A, Chen A, DasSarma N, Drain D, Fort S, Ganguli D, Henighan T, Joseph N. Training a helpful and harmless assistant with reinforcement learning from human feedback. arXiv preprint arXiv:2204.05862. 2022 Apr 12. [9] Rame A, Couairon G, Dancette C, Gaya JB, Shukor M, Soulier L, Cord M. Rewarded soups: towards pareto-optimal alignment by interpolating weights fine-tuned on diverse rewards. Advances in Neural Information Processing Systems. 2023 Dec 15;36:71095-134. [10]Fitzpatrick KK, Darcy A, Vierhile M. Delivering cognitive behavior therapy to young adults with symptoms of depression and anxiety using a fully automated conversational agent (Woebot): a randomized controlled trial. JMIR mental health. 2017 Jun 6;4(2):e7785. [11] Singhal K, Tu T, Gottweis J, Sayres R, Wulczyn E, Amin M, Hou L, Clark K, Pfohl SR, Cole-Lewis H, Neal D. Toward expert-level medical question answering with large language models. Nature medicine. 2025 Mar;31(3):943-50. [12 Sharma A, Miner A, Atkins D, Althoff T. A computational approach to understanding empathy expressed in text- based mental health support. InProceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP) 2020 Nov (pp. 5263-5276). [13] Shen T, Lei T, Barzilay R, Jaakkola T. Style transfer from non-parallel text by cross-alignment. Advances in neural information processing systems. 2017;30. [14] Rashkin H, Smith EM, Li M, Boureau YL. Towards empathetic open-domain conversation models: A new benchmark and dataset. InProceedings of the 57th annual meeting of the association for computational linguistics 2019 Jul (pp. 5370-5381). [15 Lourie N, Le Bras R, Choi Y. Scruples: A corpus of community ethical judgments on 32,000 real-life anecdotes. InProceedings of the AAAI Conference on Artificial Intelligence 2021 May 18 (Vol. 35, No. 15, pp. 13470-13479). [16] Altman E. Constrained Markov decision processes. Routledge; 2021 Dec 24. [17] Achiam J, Held D, Tamar A, Abbeel P. Constrained policy optimization. InInternational conference on machine learning 2017 Jul 17 (pp. 22-31). Pmlr. [18] Tessler C, Mankowitz DJ, Mannor S. Reward constrained policy optimization. arXiv preprint arXiv:1805.11074. 2018 May 28. [19] Barrett-Lennard GT. The empathy cycle: Refinement of a nuclear concept. Journal of counseling psychology. 1981 Mar;28(2):91. [20] He P, Liu X, Gao J, Chen W. Deberta: Decoding-enhanced bert with disentangled attention. arXiv preprint arXiv:2006.03654. 2020 Jun 5. [21] Romanov A, Shivade C. Lessons from natural language inference in the clinical domain. InProceedings of the 2018 conference on empirical methods in natural language processing 2018 (pp. 1586-1596). [22] Krippendorff K. Content Analysis: An Introduction to its Methodology. TPB; 1996. [23] Rashkin H, Smith EM, Li M, Boureau YL. Towards empathetic open-domain conversation models: A new benchmark and dataset. InProceedings of the 57th annual meeting of the association for computational linguistics 2019 Jul (pp. 5370-5381). [24] Paternain S, Chamon L, Calvo-Fullana M, Ribeiro A. Constrained reinforcement learning has zero duality gap. Advances in Neural Information Processing Systems. 2019;32.


Ahead of Print Subscription Original Research
Volume 03
02
Received 06/07/2026
Accepted 08/07/2026
Published 20/07/2026
Publication Time 14 Days


Login

My IP

PlumX Metrics

Support