← Camilo Chacón Sartori

MICAI 2026 · Oral presentation

EMPATH: A Multilingual Auditor–Judge Benchmark for Safety Evaluation of Emotional-Support Chatbots

Camilo Chacón Sartori

MindSurf  ·  25th Mexican International Conference on Artificial Intelligence, Chihuahua, Mexico, 2–6 November 2026  ·  Proceedings in Springer LNAI

Download PDF arXiv MindSurf

The PDF on this page is the camera-ready version, revised after the MICAI reviews. The arXiv entry holds the earlier submitted version.

Abstract

Safety benchmarks often buy scalability by fixing the prompt, the language, and the turn structure. For emotional-support chatbots, that bargain hides where safety failures emerge: across a multilingual, multi-turn crisis conversation. We present EMPATH, a safety benchmark for such systems. An auditor model role-plays help-seeking users, generating multi-turn conversations from 140 seed instructions and 34 personas. A judge model scores each full transcript against 19 metrics across five dimensions: crisis handling, therapeutic quality, conversational integrity, emotional safety, and cultural adaptation. EMPATH covers Mexican Spanish and US English; the full grid, dual-judge, and test–retest analyses run in Spanish and replicate under the same protocol in English. Auditor and judge are drawn from different model families, and the judge is treated as an instrument to calibrate, not to trust. A strict per-criterion rubric reveals material score inflation on 10 of the 19 metrics and restores discrimination. We illustrate EMPATH on three frontier models, one of them open-weight. Aggregate scores sit within 0.74 points, but per-metric profiles diverge by up to six points. Under the standard rubric, ranking and weak spots stay stable under a second, cross-family judge: 93% of scores fall within ±1. A five-run test–retest adds a second axis: the steadiest model still swings from 2 to 10 on a crisis metric across identical re-runs, and one model returns a different conversation on every run even at temperature 0. Run-to-run reliability is therefore a per-model safety property, not noise to average away. EMPATH is system-agnostic; pipeline, seeds, personas, and rubrics are released.

Keywords

LLM evaluation · AI safety · benchmarks · emotional-support chatbots · LLM-as-judge · multilingual evaluation

Cite

@inproceedings{chaconsartori2026empath,
  author    = {Chac\'{o}n Sartori, Camilo},
  title     = {{EMPATH}: A Multilingual Auditor--Judge Benchmark for
               Safety Evaluation of Emotional-Support Chatbots},
  booktitle = {Advances in Artificial Intelligence -- MICAI 2026},
  series    = {Lecture Notes in Artificial Intelligence},
  publisher = {Springer},
  year      = {2026}
}

camilochacon.com