AI Dataset
Autonomous EntityAI AssistantAI Answer EngineAI SystemSynthetic Personality PlatformSynthetic CompanionCreative IntelligenceVoice IntelligenceFictional IntelligenceAI OrganizationFoundation ModelHistorical AI SystemAgentic SystemRobotic IntelligenceDeveloper PlatformCreative Intelligence SystemRobotics and Embodied AIResearch SystemEvaluation SystemDatasetAI ToolAI InfrastructureScientific AI SystemEmbodied AI SystemCreative AI SystemSpecialized AI SystemAI AgentAI Research SystemAI Governance SystemAI FrameworkAI Training FrameworkAI KernelAI RuntimeMLOps PlatformSearch PlatformGraph AI PlatformRAG MethodLanguage ModelMachine Translation ModelSpeech ModelVision ModelAI DatasetAI BenchmarkScientific AI ModelScientific AI FrameworkAI HardwareEmbodied AIAI Writing SystemAI TutorAI Documentation AssistantFictional Synthetic PersonalityFictional Synthetic PersonMultimodal ModelMultimodal FrameworkVision FrameworkSpeech FrameworkAI Simulation PlatformAgentic AI SystemAI CompilerMachine Learning FrameworkEmbedding ModelReranker ModelMixture-of-Experts ModelCode ModelMachine Translation BenchmarkMachine Translation DatasetMedical AI DatasetMedical AI BenchmarkScientific DatasetScientific AI DatasetFictional RobotFictional Synthetic IntelligenceMath ModelAI Assistant ModelAI ArchitectureAI Alignment MethodReinforcement Learning MethodAI Reasoning MethodAutoML FrameworkWorkflow PlatformData PlatformAI Evaluation PlatformAI Observability PlatformAI GatewayImage Generation ModelImage Control ModelImage Fine-Tuning MethodAudio Generation ModelAudio Codec ModelSpeech Generation ModelAudio Generation FrameworkVoice Conversion SystemRobotics Simulation PlatformRobotics BenchmarkRobotics DatasetAI PlatformAI Search AssistantAI Data PlatformAI InterfaceRobotics SoftwareAI Safety ToolAI Evaluation FrameworkAI Safety ModelComputer Vision ModelAI Narrative SystemAI LibraryAI ModelDataset / BenchmarkOrganization / PlatformEmbodied AI / RoboticsBenchmark
BoolQ
BoolQ is a yes/no question answering dataset associated with Google.
RedPajama Dataset
RedPajama Dataset is a open reproduction training dataset associated with Together AI.
ProcTHOR
ProcTHOR is a procedural indoor embodied AI environments associated with Allen Institute for AI.
MNIST
MNIST is a handwritten digit dataset associated with NYU.
CIFAR-10
CIFAR-10 is a image classification dataset associated with University of Toronto.
CIFAR-100
CIFAR-100 is a image classification dataset associated with University of Toronto.
Fashion-MNIST
Fashion-MNIST is a fashion image dataset associated with Zalando.
Cityscapes Dataset
Cityscapes Dataset is a urban scene understanding dataset associated with Cityscapes Consortium.
Open Images Dataset
Open Images Dataset is a image annotation dataset associated with Google.
LibriSpeech
LibriSpeech is a speech recognition corpus associated with OpenSLR.
Common Voice
Common Voice is a open voice dataset associated with Mozilla.
VoxCeleb
VoxCeleb is a speaker recognition dataset associated with University of Oxford.
AudioSet
AudioSet is a audio event dataset associated with Google.
Free Music Archive Dataset
Free Music Archive Dataset is a music analysis dataset associated with EPFL.
VGGSound
VGGSound is a audio-visual dataset associated with University of Oxford.
AG News Dataset
AG News Dataset is a news classification dataset associated with Academic dataset.
IMDb Reviews Dataset
IMDb Reviews Dataset is a sentiment analysis dataset associated with Stanford University.
SST-2
SST-2 is a sentiment classification dataset associated with Stanford University.
SNLI
SNLI is a natural language inference dataset associated with Stanford University.
MultiNLI
MultiNLI is a natural language inference dataset associated with NYU.
ANLI
ANLI is a adversarial natural language inference dataset associated with Meta AI.
PAWS
PAWS is a paraphrase adversaries dataset associated with Google.
Quora Question Pairs
Quora Question Pairs is a duplicate question dataset associated with Quora.
CoLA
CoLA is a linguistic acceptability corpus associated with NYU.
XQuAD
XQuAD is a cross-lingual question answering dataset associated with Google.
MLQA
MLQA is a multilingual question answering dataset associated with Meta AI.
OpenBookQA
OpenBookQA is a science question answering dataset associated with Allen Institute for AI.
SocialIQA
SocialIQA is a social commonsense reasoning dataset associated with Allen Institute for AI.
PIQA
PIQA is a physical commonsense reasoning dataset associated with Allen Institute for AI.
ReCoRD
ReCoRD is a reading comprehension dataset associated with Allen Institute for AI.
RACE Dataset
RACE Dataset is a reading comprehension dataset associated with CMU.
CNN/DailyMail Dataset
CNN/DailyMail Dataset is a summarization dataset associated with Google DeepMind.
XSum
XSum is a extreme summarization dataset associated with University of Edinburgh.
SAMSum
SAMSum is a dialogue summarization dataset associated with Samsung.
Multi-News
Multi-News is a multi-document summarization dataset associated with Yale University.
ADE20K
ADE20K is a scene parsing dataset associated with MIT.
PASCAL VOC
PASCAL VOC is a visual object classes dataset associated with PASCAL VOC.
Places365
Places365 is a scene recognition dataset associated with MIT.
CelebA
CelebA is a large-scale face attributes dataset associated with CUHK.
LFW
LFW is a labeled faces in the wild dataset associated with UMass Amherst.
SVHN
SVHN is a street view house numbers dataset associated with Stanford University.
STL-10
STL-10 is a image recognition dataset associated with Stanford University.
Caltech 101
Caltech 101 is a object category dataset associated with Caltech.
Caltech 256
Caltech 256 is a object category dataset associated with Caltech.
Food-101
Food-101 is a food image dataset associated with ETH Zurich.
UCF101
UCF101 is a action recognition dataset associated with UCF.
Kinetics-400
Kinetics-400 is a human action video dataset associated with Google DeepMind.
ActivityNet
ActivityNet is a human activity video dataset associated with ActivityNet.
Something-Something V2
Something-Something V2 is a video action dataset associated with TwentyBN.
Visual Genome
Visual Genome is a visual relationship dataset associated with Stanford University.
LVIS
LVIS is a large vocabulary instance segmentation dataset associated with Meta AI.
D4RL
D4RL is a offline reinforcement learning benchmark associated with UC Berkeley.
Spider
Spider is a text-to-SQL dataset associated with Yale University.
WikiSQL
WikiSQL is a semantic parsing dataset associated with Salesforce.
WikiTableQuestions
WikiTableQuestions is a table question answering dataset associated with Stanford University.
TextVQA
TextVQA is a visual question answering dataset associated with Research community.
AI2D
AI2D is a diagram understanding dataset associated with Allen Institute for AI.
RefCOCO
RefCOCO is a referring expression comprehension dataset associated with UNC.
MSR-VTT
MSR-VTT is a video captioning dataset associated with Microsoft.
YouCook2
YouCook2 is a instructional video dataset associated with University of Michigan.
Ego4D
Ego4D is a egocentric video dataset associated with Meta AI.
EPIC-KITCHENS
EPIC-KITCHENS is a egocentric action recognition dataset associated with University of Bristol.
Argoverse
Argoverse is a autonomous driving dataset associated with Argoverse.
nuPlan
nuPlan is a autonomous driving planning benchmark associated with Motional.