arXiv — NLP / Computation & Language · · 3 min read

Multi-Task Instruction Tuning via Data Scheduling for Low-Resource Arabic SpeechLLMs

Mirrored from arXiv — NLP / Computation & Language for archival readability. Support the source by reading on the original site.

Computer Science > Sound

arXiv:2601.12494 (cs)
[Submitted on 18 Jan 2026 (v1), last revised 7 Jul 2026 (this version, v3)]

Title:Multi-Task Instruction Tuning via Data Scheduling for Low-Resource Arabic SpeechLLMs

View a PDF of the paper titled Multi-Task Instruction Tuning via Data Scheduling for Low-Resource Arabic SpeechLLMs, by Hunzalah Hassan Bhatti and 2 other authors
View PDF
Abstract:Audio large language models (LLMs) enable unified speech understanding and generation, but adapting them to linguistically complex and dialect-rich settings such as Arabic-English remains challenging. We present a controlled study of multi-task instruction tuning for an Arabic-centric audio LLM across generative tasks, including automatic speech recognition (ASR) and speech and text summarization, as well as discriminative tasks, including dialect identification (DID) and speech emotion recognition (SER), in a resource-constrained setting. To support end-to-end Arabic speech summarization, we introduce AraMega-SSum, the first Arabic speech summarization dataset designed for training and benchmarking Arabic-centric audio LLMs. We compare four training strategies: (i) Uniform Mixing (UM), (ii) Task-Progressive Curriculum (TPC), (iii) Aligner-Based Diverse Sampling (ADS) for training-time batch construction, and (iv) a two-stage TPC->ADS strategy. Our results reveal a clear efficiency-robustness trade-off. TPC achieves the strongest performance on generative tasks, including ASR and summarization. ADS improves paralinguistic tasks but reduces generative stability when used alone. The two-stage TPC->ADS strategy provides the best overall balance, achieving the strongest DID and SER performance while outperforming large proprietary models such as Gemini-2.5-Pro on discriminative tasks. We will publicly release AraMega-SSum together with all experimental resources to support future research in Arabic speech understanding.
Comments: Foundation Models, Large Language Models, Native, Speech Models, Arabic
Subjects: Sound (cs.SD); Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Audio and Speech Processing (eess.AS)
MSC classes: 68T50
ACM classes: F.2.2; I.2.7
Cite as: arXiv:2601.12494 [cs.SD]
  (or arXiv:2601.12494v3 [cs.SD] for this version)
  https://doi.org/10.48550/arXiv.2601.12494
arXiv-issued DOI via DataCite

Submission history

From: Hunzalah Hassan Bhatti [view email]
[v1] Sun, 18 Jan 2026 17:08:31 UTC (993 KB)
[v2] Mon, 23 Mar 2026 09:07:43 UTC (1,744 KB)
[v3] Tue, 7 Jul 2026 11:34:27 UTC (2,013 KB)
Full-text links:

Access Paper:

    View a PDF of the paper titled Multi-Task Instruction Tuning via Data Scheduling for Low-Resource Arabic SpeechLLMs, by Hunzalah Hassan Bhatti and 2 other authors
  • View PDF
  • TeX Source

Current browse context:

cs.SD
< prev   |   next >
Change to browse by:

References & Citations

Loading...

BibTeX formatted citation

loading...
Data provided by:

Bookmark

BibSonomy Reddit
Bibliographic Tools

Bibliographic and Citation Tools

Bibliographic Explorer Toggle
Bibliographic Explorer (What is the Explorer?)
Connected Papers Toggle
Connected Papers (What is Connected Papers?)
Litmaps Toggle
Litmaps (What is Litmaps?)
scite.ai Toggle
scite Smart Citations (What are Smart Citations?)
Code, Data, Media

Code, Data and Media Associated with this Article

alphaXiv Toggle
alphaXiv (What is alphaXiv?)
Links to Code Toggle
CatalyzeX Code Finder for Papers (What is CatalyzeX?)
DagsHub Toggle
DagsHub (What is DagsHub?)
GotitPub Toggle
Gotit.pub (What is GotitPub?)
Huggingface Toggle
Hugging Face (What is Huggingface?)
ScienceCast Toggle
ScienceCast (What is ScienceCast?)
Demos

Demos

Replicate Toggle
Replicate (What is Replicate?)
Spaces Toggle
Hugging Face Spaces (What is Spaces?)
Spaces Toggle
TXYZ.AI (What is TXYZ.AI?)
Related Papers

Recommenders and Search Tools

Link to Influence Flower
Influence Flower (What are Influence Flowers?)
Core recommender toggle
CORE Recommender (What is CORE?)
About arXivLabs

arXivLabs: experimental projects with community collaborators

arXivLabs is a framework that allows collaborators to develop and share new arXiv features directly on our website.

Both individuals and organizations that work with arXivLabs have embraced and accepted our values of openness, community, excellence, and user data privacy. arXiv is committed to these values and only works with partners that adhere to them.

Have an idea for a project that will add value for arXiv's community? Learn more about arXivLabs.

Discussion (0)

Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.

Sign in →

No comments yet. Sign in and be the first to say something.

More from arXiv — NLP / Computation & Language