Skip to main navigation Skip to search Skip to main content

A multi-labeled dataset for Indonesian discourse: examining toxicity, polarization, and demographics information

  • Lucky Susanto
  • , Musa Wijanarko
  • , Prasetia Pratama
  • , Zilu Tang
  • , Fariz Akyas
  • , Traci Hong
  • , Ika Idris
  • , Alham Aji
  • , Derry Wijaya

Research output: Chapter in Book/Report/Conference proceedingConference PaperResearchpeer-review

Abstract

Online discourse is increasingly trapped in a vicious cycle where polarizing language fuels toxicity and vice versa. Identity, one of the most divisive issues in modern politics, often increases polarization. Yet, prior NLP research has mostly treated toxicity and polarization as separate problems. In Indonesia, the world's third-largest democracy, this dynamic threatens democratic discourse, particularly in online spaces. We argue that polarization and toxicity must be studied in relation to each other. To this end, we present a novel multi-label Indonesian dataset annotated for toxicity, polarization, and annotator demographic information. Benchmarking with BERT-base models and large language models (LLMs) reveals that polarization cues improve toxicity classification and vice versa. Including demographic context further enhances polarization classification performance.

Original languageEnglish
Title of host publicationFindings of the Association for Computational Linguistics
Subtitle of host publicationACL 2025
EditorsWanxiang Che, Joyce Nabende, Ekaterina Shutova, Mohammad Taher Pilehvar
Place of PublicationVienna Austria
PublisherAssociation for Computational Linguistics (ACL)
Pages18863-18890
Number of pages28
Edition1st
ISBN (Electronic)9798891762565
DOIs
Publication statusPublished - 2025
EventAnnual Meeting of the Association for Computational Linguistics 2025 - Vienna, Austria
Duration: 27 Jul 20251 Aug 2025
Conference number: 63rd
https://2025.aclweb.org/ (Website)
https://aclanthology.org/events/acl-2025/#2025acl-long (Proceedings - Long)
https://aclanthology.org/events/acl-2025/#2025acl-short (Proceedings - Short)
https://aclanthology.org/events/acl-2025/#2025acl-demo (Proceedings - Demo)

Publication series

NameProceedings of the Annual Meeting of the Association for Computational Linguistics
ISSN (Print)0736-587X

Conference

ConferenceAnnual Meeting of the Association for Computational Linguistics 2025
Abbreviated titleACL 2025
Country/TerritoryAustria
CityVienna
Period27/07/251/08/25
Internet address

Cite this