Large Language Models (LLMs) are increasingly utilized for patient-facing healthcare communication. However, their clinical utility remains constrained by communicative, ethical, and safety challenges. This systematic review provides a cross-specialty evaluation of conversational AI, synthesizing evidence on accuracy, readability, informational quality, and emotional resonance. Methods: Following PRISMA 2020 guidelines and PROSPERO registration (CRD420251021428), PubMed/MEDLINE, Scopus, and Web of Science were searched from 2020 through November 2025. Eligible studies evaluated LLM-generated medical or surgical information against healthcare professionals, validated benchmarks, or alternative AI platforms. Assessed outcomes included factual accuracy, standardized readability indices (e.g., FKGL, FRES), validated quality metrics (e.g., DISCERN, GQS), and emotional or empathetic impact. Results: Eighty-two studies (published 2023–2025) were included. Accuracy was evaluated in 45 studies, demonstrating generally high performance (mean Likert scores > 3.5/5); advanced GPT models frequently showed parity or superiority to comparators, though factual errors persisted in complex clinical reasoning. Readability, assessed in 69 studies, emerged as the primary limitation: baseline outputs consistently required high school or university-level comprehension (FKGL > 10–12), substantially exceeding recommended patient literacy thresholds (≤8th grade). Informational quality, appraised in 35 studies, was rated moderate-to-high across DISCERN and GQS metrics. Six studies evaluating emotional resonance indicated that LLMs can alleviate situational anxiety and deliver structured cognitive empathy comparable or superior to brief clinician text, though lacking genuine emotional depth. Conclusions: LLMs demonstrate strong factual accuracy and structural quality across diverse clinical disciplines. However, autonomous, patient-facing deployment remains premature due to excessive lexical density that threatens health equity. Conversational AI should function as a clinical adjunct to support providers rather than an unsupervised educational tool. Future research must prioritize dynamic literacy alignment and robust clinical validation.

Artificial Intelligence in Medical and Surgical Information: A Systematic Review of Accuracy, Readability, Reliability, and Emotional Resonance

Di Mattia, Paolo;
2026-01-01

Abstract

Large Language Models (LLMs) are increasingly utilized for patient-facing healthcare communication. However, their clinical utility remains constrained by communicative, ethical, and safety challenges. This systematic review provides a cross-specialty evaluation of conversational AI, synthesizing evidence on accuracy, readability, informational quality, and emotional resonance. Methods: Following PRISMA 2020 guidelines and PROSPERO registration (CRD420251021428), PubMed/MEDLINE, Scopus, and Web of Science were searched from 2020 through November 2025. Eligible studies evaluated LLM-generated medical or surgical information against healthcare professionals, validated benchmarks, or alternative AI platforms. Assessed outcomes included factual accuracy, standardized readability indices (e.g., FKGL, FRES), validated quality metrics (e.g., DISCERN, GQS), and emotional or empathetic impact. Results: Eighty-two studies (published 2023–2025) were included. Accuracy was evaluated in 45 studies, demonstrating generally high performance (mean Likert scores > 3.5/5); advanced GPT models frequently showed parity or superiority to comparators, though factual errors persisted in complex clinical reasoning. Readability, assessed in 69 studies, emerged as the primary limitation: baseline outputs consistently required high school or university-level comprehension (FKGL > 10–12), substantially exceeding recommended patient literacy thresholds (≤8th grade). Informational quality, appraised in 35 studies, was rated moderate-to-high across DISCERN and GQS metrics. Six studies evaluating emotional resonance indicated that LLMs can alleviate situational anxiety and deliver structured cognitive empathy comparable or superior to brief clinician text, though lacking genuine emotional depth. Conclusions: LLMs demonstrate strong factual accuracy and structural quality across diverse clinical disciplines. However, autonomous, patient-facing deployment remains premature due to excessive lexical density that threatens health equity. Conversational AI should function as a clinical adjunct to support providers rather than an unsupervised educational tool. Future research must prioritize dynamic literacy alignment and robust clinical validation.
File in questo prodotto:
Non ci sono file associati a questo prodotto.

I documenti in IRIS sono protetti da copyright e tutti i diritti sono riservati, salvo diversa indicazione.

Utilizza questo identificativo per citare o creare un link a questo documento: https://hdl.handle.net/11387/215294
Citazioni
  • ???jsp.display-item.citation.pmc??? ND
  • Scopus ND
  • ???jsp.display-item.citation.isi??? ND
social impact