Evaluation Space (Two Tracks)
SURE covers perception โ cognition across tasks and stressors.
Perception focuses on low-level signal interpretation before task-level semantic scoring.
A comprehensive collection of Speech Understanding Large Language Models (Speech LLMs). Explore the latest advances in speech processing, from early pioneers to cutting-edge multimodal systems.
Two complementary views of the SURE evaluation space and execution pipeline.
SURE covers perception โ cognition across tasks and stressors.
Perception focuses on low-level signal interpretation before task-level semantic scoring.
Unified post-processing, metrics, and agent-based verification.
Major evaluation outcomes from Track I stress testing and Track II full-stack tasks.
Track I ยท Stress Tests
Evaluation under linguistic and acoustic stressors.
Raw is the direct error rate for each scenario.
RPS is a normalized score relative to the best-performing system in that condition.
Use RPS for cross-scenario comparison when scales differ across datasets.
Track II ยท Horizontal Task Evaluation
Unified scoring across perception and reasoning tasks.
Task cards present original task metrics directly (WER/CER: lower is better; others: higher is better).
Unified view converts each dataset row into rank positions, where rank 1 is best.
Hover a unified-view cell to inspect the original metric value behind each rank.
24 curated models across the SURE timeline
๐๏ธ Alibaba
โ๏ธ AuT + MoE Transformer (Variable)
Next-generation omni-modal model from Alibaba integrating Qwen3 capabilities.
๐๏ธ NVIDIA
โ๏ธ AF-Whisper + MLP + Qwen2.5-7B
Universal Audio Language Model for diverse audio understanding tasks.
๐๏ธ xiaomi
โ๏ธ Custom Audio Encoder + MiMo-7B-Base
Xiaomi's audio understanding model for mobile and IoT applications.
๐๏ธ StepFun Audio
โ๏ธ Custom Audio Encoder + Custom LLM
StepFun's second-generation audio model with enhanced understanding capabilities.
๐๏ธ NVIDIA
โ๏ธ AF-Whisper + Qwen-2.5-7B (Variable)
Third generation of Audio Flamingo with state-of-the-art audio understanding.
๐๏ธ kimi
โ๏ธ GLM-4-Voice Tokenizer + Whisper-Large-v3 + adapter + Qwen2.5-7B + MoonCast Detokenizer
Moonshot AI's audio understanding model with extended context capabilities.
๐๏ธ Alibaba
โ๏ธ Whisper-large-v3 + Qwen2.5-7B (7B)
Omni-modal version of Qwen 2.5 supporting audio, text, and vision in unified architecture.
๐๏ธ Alibaba
โ๏ธ SenseVoice-large Encoder + Transformer + CNN Qwen2.5-7B-instruct
Minimal Speech Model for efficient speech understanding and dialogue.
๐๏ธ xiaomi
โ๏ธ Dasheng-0.6B + MLP + Qwen2.5-Omni-7B (7B)
Xiaomi's speech understanding model with Dasheng audio encoder and Qwen2.5-Omni.
๐๏ธ National Taiwan University
โ๏ธ Whisper-large-v3 + Q-Former + Llama-3.1-8B-Instruct (8B)
DeSTA 2.5 for general audio understanding with latest Llama-3.1 backbone.
๐๏ธ Google / university of Cambridge
โ๏ธ USM + CTC Layer + Gemma-2B (2B)
Lego Speech Language Model - Modular speech LLM with USM encoder and Gemma decoder.
๐๏ธ NVIDIA
โ๏ธ AF-CLAP + Gated Cross Attention + Qwen2.5-3B (3B)
Second generation of Audio Flamingo with improved AF-CLAP encoder and Qwen2.5 backbone.
๐๏ธ Xiaohongshu
โ๏ธ Conformer + Linear + Qwen2-Instruct-7B (7B)
Open-source ASR LLM with Conformer encoder and Qwen2 instruction-following backbone.
๐๏ธ ASLP Lab
โ๏ธ Whisper-Medium + 1-D CNN + Transformer + Qwen2-7B-Instruct (7B)
Open Speech Understanding Model with open-source training pipeline.
๐๏ธ Northwestern Polytechnical University / Chongqing Changan Automobile
โ๏ธ Whisper Large-v3 + MMS + adapter + phi-3-mini model (3.8B)
Ideal Speech Language Model with dual encoders for multilingual speech understanding.
๐๏ธ National Taiwan University / NVIDIA
โ๏ธ Whisper-small + modality-apdapter + Llama3-8B-Instruct (8B)
DeSTA version 2 with upgraded LLaMA-3 backbone and improved alignment.
๐๏ธ The Chinese University of Hong Kong
โ๏ธ Whisper-large-v2 + Transformer + TinyLlama-1.1B-Chat/Llama2-7B-Chat/Llama2-13B-Chat (7B)
Large Language and Speech Transformer with dual encoders for robust speech understanding.
๐๏ธ ByteDance
โ๏ธ LUISE Encoder + Converter/Projector + Pretrained LLM (Variable)
ByteDance's production ASR and speech understanding model with LUISE encoder.
๐๏ธ Alibaba
โ๏ธ Whisper-large-v3 + Avg-Pool + Qwen-7B (7B)
Improved version of Qwen-Audio with enhanced Whisper-v3 encoder and Qwen2 backbone.
๐๏ธ x-lance
โ๏ธ WavLM-Large + Vicuna 7B (7B)
Massive Language Model ASR with WavLM encoder and LLaMA backbone.
๐๏ธ NVIDIA / National Taiwan University
โ๏ธ Whisper-large-v3 + Modality Adapter + Llama2-7b-chat (7B)
Deep Speech Text Alignment model with novel Q-Former based alignment mechanism.
๐๏ธ Amazon
โ๏ธ Pretrained Audio Encoder + 1-D CNN + Pretrained LLM (Variable)
Universal speech understanding framework with pretrained audio and language models.
๐๏ธ Microsoft
โ๏ธ Light Audio Encoder + CTC Compression + LLaMA-7B (7B)
Efficient speech-enabled LLaMA with lightweight encoder and CTC compression.
๐๏ธ University of Illinois at Urbana-Champaign
โ๏ธ wavvec 2.0 base + gpt2 ()
Waveform-based Large Language Model with dual encoders for comprehensive speech understanding.