Speech Model
Autonomous EntityAI AssistantAI Answer EngineAI SystemSynthetic Personality PlatformSynthetic CompanionCreative IntelligenceVoice IntelligenceFictional IntelligenceAI OrganizationFoundation ModelHistorical AI SystemAgentic SystemRobotic IntelligenceDeveloper PlatformCreative Intelligence SystemRobotics and Embodied AIResearch SystemEvaluation SystemDatasetAI ToolAI InfrastructureScientific AI SystemEmbodied AI SystemCreative AI SystemSpecialized AI SystemAI AgentAI Research SystemAI Governance SystemAI FrameworkAI Training FrameworkAI KernelAI RuntimeMLOps PlatformSearch PlatformGraph AI PlatformRAG MethodLanguage ModelMachine Translation ModelSpeech ModelVision ModelAI DatasetAI BenchmarkScientific AI ModelScientific AI FrameworkAI HardwareEmbodied AIAI Writing SystemAI TutorAI Documentation AssistantFictional Synthetic PersonalityFictional Synthetic PersonMultimodal ModelMultimodal FrameworkVision FrameworkSpeech FrameworkAI Simulation PlatformAgentic AI SystemAI CompilerMachine Learning FrameworkEmbedding ModelReranker ModelMixture-of-Experts ModelCode ModelMachine Translation BenchmarkMachine Translation DatasetMedical AI DatasetMedical AI BenchmarkScientific DatasetScientific AI DatasetFictional RobotFictional Synthetic IntelligenceMath ModelAI Assistant ModelAI ArchitectureAI Alignment MethodReinforcement Learning MethodAI Reasoning MethodAutoML FrameworkWorkflow PlatformData PlatformAI Evaluation PlatformAI Observability PlatformAI GatewayImage Generation ModelImage Control ModelImage Fine-Tuning MethodAudio Generation ModelAudio Codec ModelSpeech Generation ModelAudio Generation FrameworkVoice Conversion SystemRobotics Simulation PlatformRobotics BenchmarkRobotics DatasetAI PlatformAI Search AssistantAI Data PlatformAI InterfaceRobotics SoftwareAI Safety ToolAI Evaluation FrameworkAI Safety ModelComputer Vision ModelAI Narrative SystemAI LibraryAI ModelDataset / BenchmarkOrganization / PlatformEmbodied AI / RoboticsBenchmark
VITS
VITS is a text-to-speech model associated with Open source research.
Whisper Large v2
Whisper Large v2 is a automatic speech recognition model associated with OpenAI.
Whisper Large v3 Turbo
Whisper Large v3 Turbo is a automatic speech recognition model associated with OpenAI.
WavLM Base
WavLM Base is a self-supervised speech representation model associated with Microsoft Research.
WavLM Large
WavLM Large is a self-supervised speech representation model associated with Microsoft Research.
XLS-R 300M
XLS-R 300M is a cross-lingual speech representation model associated with Meta AI.
XLS-R 1B
XLS-R 1B is a cross-lingual speech representation model associated with Meta AI.
XLS-R 2B
XLS-R 2B is a cross-lingual speech representation model associated with Meta AI.