Codesota · Models2,268 models indexed · 104 match filter
Editorial · Models

Models with recorded evidence.

Start with a research area, drill into a vendor, or page through the full index. Vendor aliases and legacy area IDs are grouped here; original model IDs and model links stay unchanged. Only models with at least one benchmark score appear — a model without a recorded score can’t be ranked.

Vendor:Areas overviewSpeakLeash · 263Alibaba · 104Google · 102OpenAI · 86Meta · 68Microsoft · 49Anthropic · 44DeepSeek · 34Mistral · 30mistralai · 19CYFRAGOVPL · 14NVIDIA · 14Zhipu AI · 13internlm · 10xAI · 10ByteDance · 9Baidu · 8ibm-granite · 8PLLuM · 8allenai · 7Amazon · 7MiniMax · 7Mistral AI · 7Remek · 7Shanghai AI Lab · 7utter-project · 7CohereForAI · 6Salesforce · 601-ai · 5Cohere · 5Moonshot AI · 5NousResearch · 5THUML · 5gguf-iq · 4IBM · 4Meituan · 4openchat · 4Stanford · 4THUDM · 4tiiuae · 4UC San Diego · 4VikParuchuri · 4Allen AI · 3BAAI · 3Du et al. · 3ForgeCode · 3Fudan University · 3gguf · 3gguf11bv30 · 3gguf7bv30 · 3IDEA Research · 3Liao et al. · 3Moonshot.AI · 3Nam Tuan Ly / NII · 3OpenDataLab · 3OPI-PG · 3upstage · 3ViCoS Lab Ljubljana · 3Xiaomi · 3Zhao et al. · 3+ 243 smaller vendors (288 models)
§ 01 · Speech models

104 models in Speech · page 2 of 3.

#ModelVendorParametersArchitectureBenchmarksResults
051LongCat-Flash-Omni———77
052Canary-Qwen-2.5BNVIDIA2.5BFastConformer encoder + Qwen2 LM decoder66
053Owsm_ctc_v3.1_1B———56
054Parakeet-tdt-0.6b-v2———56
055Moonshine-streaming-small———45
056Niagara-19m-batch.en———55
057Granite Speech 3.3 8BIBM8BTransformer44
058Canary-1BNVIDIA1BFastConformer encoder + Transformer decoder13
059Moonshine Streaming MediumUseful Sensors245MCausal encoder-decoder23
060Canary-1B-FlashNVIDIA1BFastConformer + TDT decoder22
061Distil-large-v3.5———22
062Google USMGoogle2BConformer encoder + RNN-T/CTC12
063Granite 4.0 1B SpeechIBM1BTransformer22
064HuBERT Large (LS-960)Meta317MCNN + Transformer (BERT-style)12
065Lite-whisper-large-v3-acc———22
066Llama 3 Speech (70B)———22
067Parakeet-CTC-1.1BNVIDIA / Suno1.1BFastConformer-CTC12
068Parakeet-tdt-0.6b-v3———22
069Pulse STTSmallest AI—Proprietary streaming STT12
070Qwen3-ASR-0.6BAlibaba0.6BTransformer (Qwen3 backbone)22
071Universal-1AssemblyAI—Transformer12
072Voxtral-Mini-3B-2507———22
073Voxtral-Small-24B-2507Mistral AI24BLarge multimodal LM with audio encoder22
074Canary-180M-FlashNVIDIA180MFastConformer-Small + TDT11
075Canary-1b-v2———11
076Conformer-CTC LargeNVIDIA / NeMo118MConformer (Conv + Attention) + CTC11
077CrisperWhispernyrahealth1.5BWhisper fine-tune with alignment11
078Distil-Whisper Large v2———11
079Distil-Whisper Large v3———11
080Distil-Whisper Large v3.5———11
081Distil-Whisper Medium (English)———11
082Distil-Whisper Small (English)———11
083ECAPA-TDNNGhent University~14.7MECAPA-TDNN (SE-Res2Net + attentive stats pooling)11
084Fairseq S2T (MuST-C)Meta~150MConformer encoder + transformer decoder11
085GLM-ASR-Nano-2512Zhipu AI2BGLM4 + audio encoder11
086Lite-whisper-large-v3———11
087Moshi ASR———11
088Owsm_ctc_v4_1B———11
089Parakeet-ctc-1.1b———11
090Parakeet-rnnt-1.1b———11
091Parakeet-TDT-1.1BNVIDIA1.1BFastConformer (TDT)11
092Phi-4-Multimodal 5.6B———11
093ResNet-34 (AM-Softmax, VoxCeleb2)Community~6MResNet-34 with AM-Softmax loss11
094SeamlessM4T v2 LargeMeta2.3BUnified multilingual/multimodal transformer (UnitY2)11
095Stt-2.6b-en———11
096SYMPHONY———11
097Wav2Vec 2.0 Base———11
098Wav2Vec 2.0 Large (LS-960)———11
099WavLM Large (SV)Microsoft316MWavLM Large + ECAPA-TDNN head11
100Whisper baseOpenAI74MTransformer encoder-decoder11