Home |Institute|News

News

Study Provides Initial Evidence of AI’s Potential in Psychopathological Assessment

Large language models analyzed psychiatric interview transcripts with a level of reliability comparable to that of young clinicians. In the future, they could potentially serve as a supplementary decision-making tool.

News |

Interviews are one of the most important tools in psychiatric diagnosis. The study provides initial evidence that AI could support the structured analysis of such interviews in the future. Photo: Zoran Zeremski

A study led by the Central Institute of Mental Health (CIMH) provides initial evidence that large language models can identify complex psychopathological findings in transcripts of psychiatric interviews with an accuracy comparable to that of predominantly young clinicians. The study examined ten language models and 108 practicing clinicians from three psychiatric clinics. The results demonstrate the potential for AI-supported decision-making tools, but they do not suggest that these tools can replace medical judgment. Further studies involving real patients are required before any potential clinical application.

From the Psychiatric Interview to a Structured Diagnosis

Psychiatric diagnosis usually begins with an interview. During this process, professionals systematically assess a person’s mental state. From information that is often ambiguous and unstructured, a psychopathological assessment is derived, which serves as the most important basis for diagnosis and treatment. It was precisely this step in the process that the researchers examined. While many previous AI studies have focused on diagnoses or standardized questionnaires, this study is the first to systematically examine how large language models perform in the complex task of psychopathological assessment compared to practicing clinicians. The results were published in the international journal npj Digital Medicine.

A Comparison of Ten Language Models with 108 Clinicians

For the study, ten large language models evaluated the transcripts of three simulated psychiatric interviews on depression, mania, and schizophrenia. Their task was to assess all 100 features of the AMDP system. The AMDP system is a method widely used in German-speaking psychiatry to systematically describe psychiatric symptoms according to established criteria.

For comparison, 108 practicing physicians and psychologists from three psychiatric clinics conducted the same evaluations. Most of the participants were still early in their clinical careers. While the language models were provided exclusively with the written transcripts, the clinicians were able to review the complete video and audio recordings to ensure the most realistic assessment possible.

The best language models fell within the range of the clinical control group

The two highest-performing models, GPT-5.1 and Gemini-3-Pro-Preview, fell within the range of the clinical comparison group overall. Across all three interviews, they achieved an average accuracy of 72 percent, while the clinical comparison group achieved 68 percent. These figures refer exclusively to the assessment of individual psychopathological characteristics and not to the establishment of a psychiatric diagnosis. Due to the small number of experienced specialists in the sample, the study does not allow for a reliable comparison with psychiatrists who have many years of specialized experience.

Clinical professionals and AI make different kinds of mistakes

What was particularly revealing was not so much the average accuracy as the nature of the errors. Clinicians were more likely to infer the presence or absence of a characteristic based on incomplete information. The language model, on the other hand, more frequently classified such characteristics as “unassessable.” This difference was particularly pronounced for symptoms where nonverbal or visual cues are important for assessment. The sometimes contrasting error patterns suggest that human clinical experience and machine-based evaluation—which is strictly guided by the available text—could complement one another.

“We deliberately chose not to limit ourselves to having a language model evaluate diagnoses or questionnaires. In everyday psychiatric practice, structured psychopathological findings must be derived from complex and often ambiguous conversations,” says lead author Dr. Esra Lenz, a researcher at the Hector Institute for Artificial Intelligence in Psychiatry (HITKIP) at the CIMH. “Our results provide initial evidence that AI could support this process in the future. However, it does not replace either clinical experience or the direct assessment by a specialist.”

Potential as an Additional Decision-Making Tool

In a retrospective simulation, the researchers also investigated whether a language model could serve as an additional decision-making aid in cases of conflicting clinical assessments. When the model’s evaluation was used to resolve such disagreements, the results were more accurate than when a random choice was made between the two clinical assessments. Similar improvements were observed in a simulated specialist supervision scenario. However, these results are derived exclusively from a statistical simulation and do not yet allow for any conclusions regarding whether AI actually improves diagnostic assessment in everyday clinical practice.

Confirmation with actual patients is required

The researchers emphasize the exploratory nature of the study. Only three simulated interviews were analyzed—one each on depression, mania, and schizophrenia. No real patients were involved. The results therefore cannot be generalized to psychiatric interviews in general or to specific disorders. Future studies must examine whether the findings can be confirmed in a larger number of different interviews, involving real patients and conducted under real-world clinical conditions. In doing so, strict ethical and data protection requirements must be taken into account.

“The key question is not whether AI will replace mental health professionals, but whether it can provide meaningful support where psychiatric diagnosis actually begins—namely, in clinical interviews and psychopathological assessment,” says Prof. Dr. Emanuel Schwarz, the study’s last author and director of the Hector Institute for Artificial Intelligence in Psychiatry at the CIMH. 

“Our study deliberately shifts the focus away from diagnoses and questionnaires and toward clinical interviews and psychopathological assessment. Before AI systems can be used in this field, we must carefully examine their strengths and limitations under real-world conditions,” adds Dr. Tobias Gradinger, also a co-author and research associate at HITKIP.

Publication

Lenz E, Naamanka J, Trabert W, Bottlender R, Malchow B, Meyer-Lindenberg A, Gradinger T, Schwarz E. Benchmarking large language models against practicing clinicians on psychopathological assessment. npj Digit Med. 2026;9(1):518. doi:10.1038/s41746-026-02852-7.

https://www.nature.com/articles/s41746-026-02852-7



Zentralinstitut für Seelische Gesundheit (ZI) - https://www.zi-mannheim.de