Teferra et al. (2026) systematically evaluated how safety guardrails influence the affective outputs of large language models (LLMs), operationalizing irritability as a measurable behavioral construct. Using validated psychometric instruments adapted for model assessment, the team compared multiple contemporary LLMs under varying safety configurations to quantify differences in affective reactivity.
The findings demonstrate that models with stronger safety guardrails exhibit attenuated irritability responses, even under adversarial prompting, whereas models with fewer constraints show greater affective variability. These results underscore important trade-offs between alignment-driven safety optimization and behavioral expressiveness in AI systems.
This work advances methodological approaches for evaluating LLM behavior in clinically relevant domains and contributes to the broader discourse on the responsible deployment of generative AI in psychiatry and digital mental health. As LLMs become increasingly integrated into therapeutic and decision-support settings, rigorous characterization of their affective and behavioral profiles remains essential.
The full article is available in npj Digital Medicine.

