Codesota · Models2,268 models indexed · 896 match filter
Editorial · Models

Models with recorded evidence.

Start with a research area, drill into a vendor, or page through the full index. Vendor aliases and legacy area IDs are grouped here; original model IDs and model links stay unchanged. Only models with at least one benchmark score appear — a model without a recorded score can’t be ranked.

Vendor:Areas overviewSpeakLeash · 263Alibaba · 104Google · 102OpenAI · 86Meta · 68Microsoft · 49Anthropic · 44DeepSeek · 34Mistral · 30mistralai · 19CYFRAGOVPL · 14NVIDIA · 14Zhipu AI · 13internlm · 10xAI · 10ByteDance · 9Baidu · 8ibm-granite · 8PLLuM · 8allenai · 7Amazon · 7MiniMax · 7Mistral AI · 7Remek · 7Shanghai AI Lab · 7utter-project · 7CohereForAI · 6Salesforce · 601-ai · 5Cohere · 5Moonshot AI · 5NousResearch · 5THUML · 5gguf-iq · 4IBM · 4Meituan · 4openchat · 4Stanford · 4THUDM · 4tiiuae · 4UC San Diego · 4VikParuchuri · 4Allen AI · 3BAAI · 3Du et al. · 3ForgeCode · 3Fudan University · 3gguf · 3gguf11bv30 · 3gguf7bv30 · 3IDEA Research · 3Liao et al. · 3Moonshot.AI · 3Nam Tuan Ly / NII · 3OpenDataLab · 3OPI-PG · 3upstage · 3ViCoS Lab Ljubljana · 3Xiaomi · 3Zhao et al. · 3+ 243 smaller vendors (288 models)
§ 01 · Computer Vision models

896 models in Computer Vision · page 15 of 18.

#ModelVendorParametersArchitectureBenchmarksResults
701Mask R-CNN (ResNeXt-101-FPN)———11
702maxvit_base_tf_512.in1kGoogle—MaxViT base, 512 input, timm11
703MetaSelf-LearningUnknownUnknownUnknown11
704MinerU2-pipelineOpenDataLab——11
705MinerU2-VLMOpenDataLab——11
706Mistral OCR 2Mistral—Vision-Language Model11
707MLDGUnknownUnknownUnknown11
708molmo-7bUnknownUnknownUnknown11
709MonkeyOCR-pro-1.2B———11
710MonkeyOCR-pro-1.2BMonkeyOCR——11
711MORANUnknownUnknownUnknown11
712Mr. DETR———11
713Multimodal (MobileNetV2)UnknownUnknownUnknown11
714Multimodal (ResNet50)UnknownUnknownUnknown11
715Multimodal Side-Tuning (MobileNetV2)UnknownUnknownUnknown11
716Multimodal Side-Tuning (ResNet50)UnknownUnknownUnknown11
717Nanonets OCR2 3BNanonets—Vision-Language OCR Model11
718Nanonets-OCR-sNanonets——11
719NCBI_BERT(large) (P)UnknownUnknownUnknown11
720NCGMUnknownUnknownUnknown11
721NEC-UIUCNEC / UIUC——11
722Nemotron Nano V2 VLNVIDIA—Vision-Language Model11
723nextvit_large.bd_ssld_6m_in1k_384ByteDance—Next-ViT Large, SSLD 6M, IN1K @ 38411
724NJU-ImagineLabNanjing UniversityUnknownScene text detector11
725NormTab (Targeted) + SQLUnknownUnknownUnknown11
726OCRFlux-3BChatDoc——11
727OCRVerse 4BUnknown4BVision-Language OCR Model11
728olmOCR-2-7B-1025 (7B)———11
729OneFormer (Swin-L)———11
730Oracle-BERTUnknown—oracle-extractive11
731Oracle-BERT (HowSumm-Method)Unknown——11
732Oracle-BOWUnknown—oracle-extractive11
733Oracle-BOW (HowSumm-Method)Unknown——11
734Oracle-HierSummUnknown—oracle-extractive11
735OTSNetAnonymous / arxiv preprintUnknownObservation-Thinking-Spelling unified network11
736ovis2.5-8bUnknownUnknownUnknown11
737PACYan et al.——11
738PaddleOCRBaidu—Deep Learning OCR11
739PaddleOCR-VL 0.9BBaidu0.9BVision-Language Model11
740PaddleOCR-VL-1.5Baidu PaddlePaddle0.9BMulti-Task VLM (0.9B params)11
741PaddleOCR-VL-1.5Baidu——11
742PANet (Joint)ICCV 2019——11
743PesRecXingwen Cao et al. (LIESMARS, Wuhan University)—Multi-task CNN: spatial layout estimator + 3D object detector + mesh generator11
744PGNet-AUnknownUnknownUnknown11
745PGNet-EUnknownUnknownUnknown11
746phi-4-multimodalUnknownUnknownUnknown11
747pil_maskrcnnICT, Chinese Academy of SciencesUnknownMask R-CNN based scene text detector11
748pixtral-12bUnknownUnknownUnknown11
749PLBARTUCLA / Columbia University140MTransformer encoder-decoder11
750pMF-H + FD-lossN/A——11