Get our free extension to see links to code for papers anywhere online!Free add-on: code for papers everywhere!Free add-on: See code for papers anywhere!

Add to Chrome

Add to Firefox

Add to Edge

Title:COSMIC: Data Efficient Instruction-tuning For Speech In-Context Learning

Nov 03, 2023

Jing Pan, Jian Wu, Yashesh Gaur, Sunit Sivasankaran, Zhuo Chen, Shujie Liu, Jinyu Li

Figure 1 for COSMIC: Data Efficient Instruction-tuning For Speech In-Context Learning

Figure 2 for COSMIC: Data Efficient Instruction-tuning For Speech In-Context Learning

Figure 3 for COSMIC: Data Efficient Instruction-tuning For Speech In-Context Learning

Figure 4 for COSMIC: Data Efficient Instruction-tuning For Speech In-Context Learning

Share this with someone who'll enjoy it:

Abstract:We present a data and cost efficient way of incorporating the speech modality into a large language model (LLM). The resulting multi-modal LLM is a COntextual Speech Model with Instruction-following/in-context-learning Capabilities - COSMIC. Speech comprehension test question-answer (SQA) pairs are generated using GPT-3.5 based on the speech transcriptions as a part of the supervision for the instruction tuning. With fewer than 20M trainable parameters and as little as 450 hours of English speech data for SQA generation, COSMIC exhibits emergent instruction-following and in-context learning capabilities in speech-to-text tasks. The model is able to follow the given text instructions to generate text response even on the unseen EN$\to$X speech-to-text translation (S2TT) task with zero-shot setting. We evaluate the model's in-context learning via various tasks such as EN$\to$X S2TT and few-shot domain adaptation. And instruction-following capabilities are evaluated through a contextual biasing benchmark. Our results demonstrate the efficacy of the proposed low cost recipe for building a speech LLM and that with the new instruction-tuning data.

View paper on

Share this with someone who'll enjoy it:

Title:COSMIC: Data Efficient Instruction-tuning For Speech In-Context Learning

Paper and Code