Background: Language endpoints serve as meaningful clinical outcomes in testing for neurological conditions (e.g. dementia). Assessments can be conducted using paper/pencil tests or through digital outcomes using natural language processing (NLP). The accuracy of many NLP models across languages have been reported with varied methodologies and levels of transparency. This warrants replication using a standard methodology. This study aimed to develop a standard methodology across languages for evaluating POS taggers and to report initial results. Methods: 105 linguists tagged language features across 58 language varieties using a short, standardized translated passage. When two linguists were available, we calculated inter–rater reliability using Cohen's kappa. We obtained the tokenization and parts of speech tags across the 52 language varieties from available models from four accessible and reproducible NLP libraries. We calculated the percentage of matches between NLP models’ and linguists’ tokenization. We compared the NLP models’ parts of speech tags to the linguists’ using Matthew's Correlation Coefficient. Findings: Twenty–five language varieties had strong performance (high reliability between linguists and strong agreement between NLP models and linguists). Eleven yielded unclear results and 16 had relatively poor performance. Interpretation: Next steps for each language variety are suggested according to our results. Gathering reliable ‘ground truth’ annotations from multiple linguists is needed for some languages. Model improvement is needed for language varieties that yielded poor performance. For language varieties with promising results, replication using longer, more ecologically valid samples is warranted. Accordingly, our methods and data are transparently reported, facilitating next steps across languages.

Automated parts–of–speech tagging in 52 language varieties: Benchmarking towards establishing suitability for multilingual clinical language analysis

Mazzaggio, Greta;
2026-01-01

Abstract

Background: Language endpoints serve as meaningful clinical outcomes in testing for neurological conditions (e.g. dementia). Assessments can be conducted using paper/pencil tests or through digital outcomes using natural language processing (NLP). The accuracy of many NLP models across languages have been reported with varied methodologies and levels of transparency. This warrants replication using a standard methodology. This study aimed to develop a standard methodology across languages for evaluating POS taggers and to report initial results. Methods: 105 linguists tagged language features across 58 language varieties using a short, standardized translated passage. When two linguists were available, we calculated inter–rater reliability using Cohen's kappa. We obtained the tokenization and parts of speech tags across the 52 language varieties from available models from four accessible and reproducible NLP libraries. We calculated the percentage of matches between NLP models’ and linguists’ tokenization. We compared the NLP models’ parts of speech tags to the linguists’ using Matthew's Correlation Coefficient. Findings: Twenty–five language varieties had strong performance (high reliability between linguists and strong agreement between NLP models and linguists). Eleven yielded unclear results and 16 had relatively poor performance. Interpretation: Next steps for each language variety are suggested according to our results. Gathering reliable ‘ground truth’ annotations from multiple linguists is needed for some languages. Model improvement is needed for language varieties that yielded poor performance. For language varieties with promising results, replication using longer, more ecologically valid samples is warranted. Accordingly, our methods and data are transparently reported, facilitating next steps across languages.
File in questo prodotto:
Non ci sono file associati a questo prodotto.

I documenti in IRIS sono protetti da copyright e tutti i diritti sono riservati, salvo diversa indicazione.

Utilizza questo identificativo per citare o creare un link a questo documento: https://hdl.handle.net/11387/211533
Citazioni
  • ???jsp.display-item.citation.pmc??? ND
  • Scopus ND
  • ???jsp.display-item.citation.isi??? ND
social impact