Abstract
Contemporary works on abstractive text summarization have focused primarily on high resource languages like English, mostly due to
the limited availability of datasets for low/midresource ones. In this work, we present XLSum, a comprehensive and diverse dataset
comprising 1 million professionally annotated
article-summary pairs from BBC, extracted
using a set of carefully designed heuristics.
The dataset covers 44 languages ranging from
low to high-resource, for many of which no
public dataset is currently available. XL-Sum
is highly abstractive, concise, and of high quality,
as indicated by human and intrinsic evaluation.
We fine-tune mT5, a state-of-theart pretrained multilingual model, with XLSum and experiment on multilingual and low resource summarization tasks. XL-Sum induces competitive results compared to the ones obtained using similar monolingual datasets: we show higher than 11 ROUGE-2 scores on
10 languages we benchmark on, with some
of them exceeding 15, as obtained by multilingual
training. Additionally, training on low-resource languages individually also provides competitive performance. To the best of our knowledge, XL-Sum is the largest abstractive summarization dataset in terms of the number of samples collected from a single source and the number of languages covered. We are releasing our dataset and models to encourage future research on multilingual abstractive summarization. The resources can be found at https://github. com/csebuetnlp/xl-sum.
the limited availability of datasets for low/midresource ones. In this work, we present XLSum, a comprehensive and diverse dataset
comprising 1 million professionally annotated
article-summary pairs from BBC, extracted
using a set of carefully designed heuristics.
The dataset covers 44 languages ranging from
low to high-resource, for many of which no
public dataset is currently available. XL-Sum
is highly abstractive, concise, and of high quality,
as indicated by human and intrinsic evaluation.
We fine-tune mT5, a state-of-theart pretrained multilingual model, with XLSum and experiment on multilingual and low resource summarization tasks. XL-Sum induces competitive results compared to the ones obtained using similar monolingual datasets: we show higher than 11 ROUGE-2 scores on
10 languages we benchmark on, with some
of them exceeding 15, as obtained by multilingual
training. Additionally, training on low-resource languages individually also provides competitive performance. To the best of our knowledge, XL-Sum is the largest abstractive summarization dataset in terms of the number of samples collected from a single source and the number of languages covered. We are releasing our dataset and models to encourage future research on multilingual abstractive summarization. The resources can be found at https://github. com/csebuetnlp/xl-sum.
| Original language | English |
|---|---|
| Title of host publication | Findings of the Association for Computational Linguistics |
| Subtitle of host publication | ACL-IJCNLP 2021 |
| Editors | Fei Xia, Wenjie Li, Roberto Navigli |
| Place of Publication | Stroudsburg PA USA |
| Publisher | Association for Computational Linguistics (ACL) |
| Pages | 4693–4703 |
| Number of pages | 11 |
| ISBN (Electronic) | 9781954085541 |
| DOIs | |
| Publication status | Published - 2021 |
| Event | Annual Meeting of the Association of Computational Linguistics and International Joint Conference on Natural Language Processing 2021 - Online, Bangkok, Thailand Duration: 1 Aug 2021 → 6 Aug 2021 Conference number: 59th & 11th https://aclanthology.org/2021.acl-long.0/ (Proceedings) https://2021.aclweb.org (Website) https://aclanthology.org/volumes/2021.findings-acl/ (Findings Proceedings) https://aclanthology.org/2021.acl-short.100/ (Proceedings Short) |
Conference
| Conference | Annual Meeting of the Association of Computational Linguistics and International Joint Conference on Natural Language Processing 2021 |
|---|---|
| Abbreviated title | ACL-IJCNLP 2021 |
| Country/Territory | Thailand |
| City | Bangkok |
| Period | 1/08/21 → 6/08/21 |
| Internet address |
|
Cite this
- APA
- Author
- BIBTEX
- Harvard
- Standard
- RIS
- Vancouver