Speech LLM Arena

A comprehensive collection of Speech Understanding Large Language Models (Speech LLMs). Explore the latest advances in speech processing, from early pioneers to cutting-edge multimodal systems.

24 Models
2022-2025 Timeline
โˆž Possibilities

Framework Overview

Two complementary views of the SURE evaluation space and execution pipeline.

Evaluation Space (Two Tracks)

SURE covers perception โ†’ cognition across tasks and stressors.

Perception focuses on low-level signal interpretation before task-level semantic scoring.

Reproducible Evaluation Pipeline

Unified post-processing, metrics, and agent-based verification.

What we standardize

  • Answer normalization before any task-specific metric computation.
  • Dataset-aligned metric definitions and direction-aware scoring.
  • Format checks for structured outputs and parser compatibility.
  • Failure reporting with explicit pass/fail and missing-output traces.
  • Versioned configs for prompts, routing, and evaluation settings.

Results

Major evaluation outcomes from Track I stress testing and Track II full-stack tasks.

Track I ยท Stress Tests

Track I ยท Scenario Stress Tests

Evaluation under linguistic and acoustic stressors.

  • RPS provides normalized comparison across scenarios
How to read this table

Raw is the direct error rate for each scenario.

RPS is a normalized score relative to the best-performing system in that condition.

Use RPS for cross-scenario comparison when scales differ across datasets.

Track II ยท Horizontal Task Evaluation

Track II ยท Full-stack Task Evaluation

Unified scoring across perception and reasoning tasks.

  • Task-first view surfaces performance per benchmark without collapsing task semantics.
  • Unified rank view enables direct comparison across heterogeneous metrics.
  • Dataset-level ranking exposes consistency versus specialization trade-offs.
View
How to read this table

Task cards present original task metrics directly (WER/CER: lower is better; others: higher is better).

Unified view converts each dataset row into rank positions, where rank 1 is best.

Hover a unified-view cell to inspect the original metric value behind each rank.

Speech LLM Models

24 curated models across the SURE timeline

Qwen3-Omni

2025
Open Source

๐Ÿ›๏ธ Alibaba

โš™๏ธ AuT + MoE Transformer (Variable)

Speech Understanding Audio Processing Multimodal Vision

Next-generation omni-modal model from Alibaba integrating Qwen3 capabilities.

English Chinese Multilingual

UALM

2025
Open Source

๐Ÿ›๏ธ NVIDIA

โš™๏ธ AF-Whisper + MLP + Qwen2.5-7B

Universal Audio Multimodal

Universal Audio Language Model for diverse audio understanding tasks.

English

MiMo-Audio

2025
Open Source

๐Ÿ›๏ธ xiaomi

โš™๏ธ Custom Audio Encoder + MiMo-7B-Base

Speech Understanding

Xiaomi's audio understanding model for mobile and IoT applications.

English Chinese

Step-Audio 2

2025
Open Source

๐Ÿ›๏ธ StepFun Audio

โš™๏ธ Custom Audio Encoder + Custom LLM

Speech Understanding Audio Processing

StepFun's second-generation audio model with enhanced understanding capabilities.

English Chinese

Audio Flamingo 3

2025
Open Source

๐Ÿ›๏ธ NVIDIA

โš™๏ธ AF-Whisper + Qwen-2.5-7B (Variable)

Audio Understanding Captioning

Third generation of Audio Flamingo with state-of-the-art audio understanding.

English

Kimi-Audio

2025
Open Source

๐Ÿ›๏ธ kimi

โš™๏ธ GLM-4-Voice Tokenizer + Whisper-Large-v3 + adapter + Qwen2.5-7B + MoonCast Detokenizer

Speech Understanding Audio Processing Long Context

Moonshot AI's audio understanding model with extended context capabilities.

English Chinese

Qwen2.5-Omni

2025
Open Source

๐Ÿ›๏ธ Alibaba

โš™๏ธ Whisper-large-v3 + Qwen2.5-7B (7B)

Speech Understanding Audio Processing Multimodal Vision

Omni-modal version of Qwen 2.5 supporting audio, text, and vision in unified architecture.

English Chinese Multilingual

MinMo

2025

๐Ÿ›๏ธ Alibaba

โš™๏ธ SenseVoice-large Encoder + Transformer + CNN Qwen2.5-7B-instruct

Speech Understanding Dialogue

Minimal Speech Model for efficient speech understanding and dialogue.

English

MiDashengLM

2025
Open Source

๐Ÿ›๏ธ xiaomi

โš™๏ธ Dasheng-0.6B + MLP + Qwen2.5-Omni-7B (7B)

Speech Understanding Audio Processing

Xiaomi's speech understanding model with Dasheng audio encoder and Qwen2.5-Omni.

English Chinese

DeSTA2.5-Audio

2025
Open Source

๐Ÿ›๏ธ National Taiwan University

โš™๏ธ Whisper-large-v3 + Q-Former + Llama-3.1-8B-Instruct (8B)

Audio Understanding

DeSTA 2.5 for general audio understanding with latest Llama-3.1 backbone.

English

LegoSLM

2025

๐Ÿ›๏ธ Google / university of Cambridge

โš™๏ธ USM + CTC Layer + Gemma-2B (2B)

Speech Understanding

Lego Speech Language Model - Modular speech LLM with USM encoder and Gemma decoder.

English Multilingual

Audio Flamingo 2

2025
Open Source

๐Ÿ›๏ธ NVIDIA

โš™๏ธ AF-CLAP + Gated Cross Attention + Qwen2.5-3B (3B)

Audio Understanding Captioning

Second generation of Audio Flamingo with improved AF-CLAP encoder and Qwen2.5 backbone.

English

FireRedASR-LLM

2025
Open Source

๐Ÿ›๏ธ Xiaohongshu

โš™๏ธ Conformer + Linear + Qwen2-Instruct-7B (7B)

ASR Speech Understanding

Open-source ASR LLM with Conformer encoder and Qwen2 instruction-following backbone.

English Chinese

OSUM

2025
Open Source

๐Ÿ›๏ธ ASLP Lab

โš™๏ธ Whisper-Medium + 1-D CNN + Transformer + Qwen2-7B-Instruct (7B)

Speech Understanding

Open Speech Understanding Model with open-source training pipeline.

English Chinese

IdealLLM

2024

๐Ÿ›๏ธ Northwestern Polytechnical University / Chongqing Changan Automobile

โš™๏ธ Whisper Large-v3 + MMS + adapter + phi-3-mini model (3.8B)

Speech Understanding Multilingual

Ideal Speech Language Model with dual encoders for multilingual speech understanding.

English Multilingual

DeSTA2

2024
Open Source

๐Ÿ›๏ธ National Taiwan University / NVIDIA

โš™๏ธ Whisper-small + modality-apdapter + Llama3-8B-Instruct (8B)

Speech Understanding

DeSTA version 2 with upgraded LLaMA-3 backbone and improved alignment.

English Chinese

LLaST

2024
Open Source

๐Ÿ›๏ธ The Chinese University of Hong Kong

โš™๏ธ Whisper-large-v2 + Transformer + TinyLlama-1.1B-Chat/Llama2-7B-Chat/Llama2-13B-Chat (7B)

Speech Understanding

Large Language and Speech Transformer with dual encoders for robust speech understanding.

English

SeedASR

2024

๐Ÿ›๏ธ ByteDance

โš™๏ธ LUISE Encoder + Converter/Projector + Pretrained LLM (Variable)

ASR Speech Understanding

ByteDance's production ASR and speech understanding model with LUISE encoder.

English Chinese

Qwen2-Audio

2024
Open Source

๐Ÿ›๏ธ Alibaba

โš™๏ธ Whisper-large-v3 + Avg-Pool + Qwen-7B (7B)

Speech Understanding Audio Processing ASR

Improved version of Qwen-Audio with enhanced Whisper-v3 encoder and Qwen2 backbone.

English Chinese Multilingual

Mala-ASR

2024
Open Source

๐Ÿ›๏ธ x-lance

โš™๏ธ WavLM-Large + Vicuna 7B (7B)

ASR

Massive Language Model ASR with WavLM encoder and LLaMA backbone.

English

DeSTA

2024
Open Source

๐Ÿ›๏ธ NVIDIA / National Taiwan University

โš™๏ธ Whisper-large-v3 + Modality Adapter + Llama2-7b-chat (7B)

Speech Understanding

Deep Speech Text Alignment model with novel Q-Former based alignment mechanism.

English Chinese

Speechverse

2024

๐Ÿ›๏ธ Amazon

โš™๏ธ Pretrained Audio Encoder + 1-D CNN + Pretrained LLM (Variable)

Speech Understanding

Universal speech understanding framework with pretrained audio and language models.

English

Speech LLaMA

2023

๐Ÿ›๏ธ Microsoft

โš™๏ธ Light Audio Encoder + CTC Compression + LLaMA-7B (7B)

Speech Understanding ASR

Efficient speech-enabled LLaMA with lightweight encoder and CTC compression.

English

WavPrompt

2022
Open Source

๐Ÿ›๏ธ University of Illinois at Urbana-Champaign

โš™๏ธ wavvec 2.0 base + gpt2 ()

Speech Understanding ASR

Waveform-based Large Language Model with dual encoders for comprehensive speech understanding.

English