arXiv cs.CL
7/8/2026

NAVER LABS System Re-implementation for the IWSLT 2026 Instruction-Following Task
Short summary
NAVER LABS re-implements their IWSLT 2025 instruction-following pipeline for the IWSLT 2026 Shared Task, using SeamlessM4T-v2-large as speech encoder and Qwen3-4B-Instruct as the LLM backbone. The system retains a three-stage approach (projector alignment, text-only LoRA pre-training, multimodal merging) and adds 100k synthetic instruction-following examples across ten speech-centric task types. The primary model achieves COMET 0.781 on EN-ZH speech translation and BERTScore-F1 0.346 on English SQA on the MCIF benchmark.
- •Re-implements NAVER LABS IWSLT 2025 pipeline for IWSLT 2026 using SeamlessM4T-v2-large and Qwen3-4B-Instruct
- •Adds 100k synthetic instruction-following examples across 10 speech-centric task types
- •Achieves COMET 0.781 on EN-ZH speech translation and BERTScore-F1 0.346 on English SQA
Generated with AI, which can make mistakes.
Is this a good recommendation for you?