MENU
Skip to main content
All blog posts

Generative AI in organizations and society

Language models still miss mental-health comorbidity

The ANGST benchmark allows social posts to carry depression and anxiety labels together. Even the best tested model struggled with complex cases.

  • Peer reviewed · 2024

People can describe symptoms of depression and anxiety in the same post. A dataset that forces every example into one label may miss that overlap. The ANGST benchmark was built to study this comorbidity—the presence of more than one condition—while comparing language models on a challenging social-media classification task.

Complete abstract from Still Not Quite There! Evaluating Large Language Models for Comorbid Mental Health Diagnosis.
Complete abstract from the original paper. Amey Hengle, Atharva Kulkarni, Shantanu Deepak Patankar, Madhumitha Chandrasekaran, Sneha D’silva, Jemima S. Jacob, Rashmi Gupta (2024), CC BY 4.0. Source paper, page 1 View full-size excerpt

A benchmark with overlapping labels

ANGST contains 2,876 posts annotated by expert psychologists and another 7,667 posts with silver labels, meaning labels generated with a weaker or indirect process rather than the same expert review. Its central task is multi-label classification: each post may be marked for depression, anxiety, both, or neither. This differs from a single-label setup that assumes the conditions are mutually exclusive (dataset and task description, PDF pp. 1–4; proceedings pp. 16698–16701).

The researchers compare specialist language models and general-purpose systems, including Mental-BERT and GPT-4, on binary and comorbid classification settings. The benchmark’s expert-labeled examples provide a more carefully reviewed test than hashtag-based or community-based proxies alone. The paper also analyzes errors, including posts where a model confuses past experiences with current symptoms (methods and error analysis, PDF pp. 4–8; proceedings pp. 16701–16705).

Better scores do not make a diagnostic tool

GPT-4 generally performs best among the tested language models, but no model exceeds an F1 score of 72% in the multi-class comorbidity task. F1 combines precision—how many predicted positives are correct—and recall—how many actual positives are found. A score below that threshold indicates substantial classification difficulty, not a measure of clinical readiness. The authors also report that domain-specific models such as Mental-XLM can outperform some general-purpose systems, depending on the task (abstract, results and conclusion, PDF pp. 1 and 6–9; proceedings pp. 16698, 16703–16706).

A social-media post is not a clinical assessment. A benchmark label reflects the annotators’ interpretation of text, and the text may omit context needed to understand a person’s circumstances. The paper therefore offers evidence about automated classification on a defined dataset, not evidence that a model can diagnose an individual.

What the data leaves out

The authors note that Reddit collection was bounded by time and search scope, so relevant self-disclosures may be missing. ANGST uses text only; it does not capture other signals that might change interpretation. The work also identifies limits in how social posts represent symptoms and current versus past states (limitations, PDF pp. 9–10; proceedings pp. 16706–16707).

Possible implications for health-related AI

For organizations considering AI-assisted support or research, the benchmark highlights the value of labels that allow overlapping conditions and of evaluation on ambiguous examples. Its results suggest that broad accuracy on simpler categories can conceal a harder problem: interpreting co-occurring experiences without turning a text classifier into a diagnostic authority.

Bibliography & sources

  1. Amey Hengle, Atharva Kulkarni, Shantanu Deepak Patankar, Madhumitha Chandrasekaran, Sneha D’silva, Jemima S. Jacob, and Rashmi Gupta. “Still Not Quite There! Evaluating Large Language Models for Comorbid Mental Health Diagnosis.” In Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, 16698–16721 (2024). https://doi.org/10.18653/v1/2024.emnlp-main.931.