Abstract
Table Question Answering (TQA) has emerged as a vital paradigm for extracting insights from structured data. However, current advancements are predominantly centered on English, creating a significant disparity for low-resource languages like Indonesian. To bridge this gap, we introduce IndoTaPas, the first TaPas-based model specifically adapted for the Indonesian language. We pre-trained the model from scratch on 1.6M Indonesian text-table pairs and fine-tuned it on IndoHiTab, a newly constructed dataset comprising 2.8K question-table pairs featuring hierarchical structures and complex reasoning types. Experimental results demonstrate that IndoTaPas, employing a two-stage fine-tuning strategy, achieves an Exact Match (EM) score of 45.22%. This performance significantly outperforms most baselines, including neural semantic parsers, modern generative frameworks, and other zero-shot Large Language Models (LLMs), while remaining statistically comparable to the state-of-the-art Meta-Llama-3. As a pioneering study, this work establishes the first dedicated benchmark and resources to accelerate future research in Indonesian TQA.
| Original language | English |
|---|---|
| Article number | 133535 |
| Pages (from-to) | 1-15 |
| Journal | Expert Systems with Applications |
| Volume | 332 |
| DOIs | |
| Publication status | Published - 4 Jul 2027 |
Keywords
- Indonesian
- IndoTaPas
- Pre-trained language model
- Table question answering
Fingerprint
Dive into the research topics of 'IndoTaPas: ATaPas-based model for Indonesian table question answering'. Together they form a unique fingerprint.Cite this
- APA
- Author
- BIBTEX
- Harvard
- Standard
- RIS
- Vancouver