论文标题
统一的问题回答斯洛文尼亚
Unified Question Answering in Slovene
论文作者
论文摘要
问题回答是语言理解中最具挑战性的任务之一。大多数方法是针对英语开发的,而资源较低的语言的研究要少得多。我们将成功的英语提问方法(称为UnifiedQA)调整为资源较低的斯洛文尼亚语言。我们的改编使用编码器 - 码头变压器插槽5和MT5模型来处理四种问题撤离格式:是/否,多项选择,抽象和提取性。我们使用四个数据集的现有Slovene改编版,并将机器转换为MCTest数据集。我们表明,通用模型至少可以以不同的格式回答问题以及专业模型。使用英语的跨语性转移进一步改善了结果。尽管我们为斯洛文尼亚产生最先进的结果,但性能仍然落后于英语。
Question answering is one of the most challenging tasks in language understanding. Most approaches are developed for English, while less-resourced languages are much less researched. We adapt a successful English question-answering approach, called UnifiedQA, to the less-resourced Slovene language. Our adaptation uses the encoder-decoder transformer SloT5 and mT5 models to handle four question-answering formats: yes/no, multiple-choice, abstractive, and extractive. We use existing Slovene adaptations of four datasets, and machine translate the MCTest dataset. We show that a general model can answer questions in different formats at least as well as specialized models. The results are further improved using cross-lingual transfer from English. While we produce state-of-the-art results for Slovene, the performance still lags behind English.