Skip to main navigation Skip to search Skip to main content

IndoTaPas: ATaPas-based model for Indonesian table question answering

Research output: Contribution to journalArticlepeer-review

Abstract

Table Question Answering (TQA) has emerged as a vital paradigm for extracting insights from structured data. However, current advancements are predominantly centered on English, creating a significant disparity for low-resource languages like Indonesian. To bridge this gap, we introduce IndoTaPas, the first TaPas-based model specifically adapted for the Indonesian language. We pre-trained the model from scratch on 1.6M Indonesian text-table pairs and fine-tuned it on IndoHiTab, a newly constructed dataset comprising 2.8K question-table pairs featuring hierarchical structures and complex reasoning types. Experimental results demonstrate that IndoTaPas, employing a two-stage fine-tuning strategy, achieves an Exact Match (EM) score of 45.22%. This performance significantly outperforms most baselines, including neural semantic parsers, modern generative frameworks, and other zero-shot Large Language Models (LLMs), while remaining statistically comparable to the state-of-the-art Meta-Llama-3. As a pioneering study, this work establishes the first dedicated benchmark and resources to accelerate future research in Indonesian TQA.

Original languageEnglish
Article number133535
Pages (from-to)1-15
JournalExpert Systems with Applications
Volume332
DOIs
Publication statusPublished - 4 Jul 2027

Keywords

  • Indonesian
  • IndoTaPas
  • Pre-trained language model
  • Table question answering

Fingerprint

Dive into the research topics of 'IndoTaPas: ATaPas-based model for Indonesian table question answering'. Together they form a unique fingerprint.

Cite this