Reward-Aligned Reinforcement Learning from Human Feedback for Emotion-Sensitive Large Language Model Therapists: Balancing Empathetic Engagement, Boundary Safety, and Clinical Accountability
The deployment of large language models (LLMs) in mental health therapy contexts introduces a critical alignment challenge: these systems must simultaneously cultivate genuine empathic rapport, observe clinically grounded safety boundaries, and remain auditable under institutional accountability frameworks.
