Dataset
Autonomous EntityAI AssistantAI Answer EngineAI SystemSynthetic Personality PlatformSynthetic CompanionCreative IntelligenceVoice IntelligenceFictional IntelligenceAI OrganizationFoundation ModelHistorical AI SystemAgentic SystemRobotic IntelligenceDeveloper PlatformCreative Intelligence SystemRobotics and Embodied AIResearch SystemEvaluation SystemDatasetAI ToolAI InfrastructureScientific AI SystemEmbodied AI SystemCreative AI SystemSpecialized AI SystemAI AgentAI Research SystemAI Governance SystemAI FrameworkAI Training FrameworkAI KernelAI RuntimeMLOps PlatformSearch PlatformGraph AI PlatformRAG MethodLanguage ModelMachine Translation ModelSpeech ModelVision ModelAI DatasetAI BenchmarkScientific AI ModelScientific AI FrameworkAI HardwareEmbodied AIAI Writing SystemAI TutorAI Documentation AssistantFictional Synthetic PersonalityFictional Synthetic PersonMultimodal ModelMultimodal FrameworkVision FrameworkSpeech FrameworkAI Simulation PlatformAgentic AI SystemAI CompilerMachine Learning FrameworkEmbedding ModelReranker ModelMixture-of-Experts ModelCode ModelMachine Translation BenchmarkMachine Translation DatasetMedical AI DatasetMedical AI BenchmarkScientific DatasetScientific AI DatasetFictional RobotFictional Synthetic IntelligenceMath ModelAI Assistant ModelAI ArchitectureAI Alignment MethodReinforcement Learning MethodAI Reasoning MethodAutoML FrameworkWorkflow PlatformData PlatformAI Evaluation PlatformAI Observability PlatformAI GatewayImage Generation ModelImage Control ModelImage Fine-Tuning MethodAudio Generation ModelAudio Codec ModelSpeech Generation ModelAudio Generation FrameworkVoice Conversion SystemRobotics Simulation PlatformRobotics BenchmarkRobotics DatasetAI PlatformAI Search AssistantAI Data PlatformAI InterfaceRobotics SoftwareAI Safety ToolAI Evaluation FrameworkAI Safety ModelComputer Vision ModelAI Narrative SystemAI LibraryAI ModelDataset / BenchmarkOrganization / PlatformEmbodied AI / RoboticsBenchmark
MBPP
MBPP is a mostly basic python problems associated with Google Research.
CodeContests
CodeContests is a competitive programming dataset associated with DeepMind.
SWE-bench Verified
SWE-bench Verified is a verified subset of swe-bench associated with OpenAI / SWE-bench maintainers.
BIG-bench Hard
BIG-bench Hard is a difficult big-bench subset associated with BIG-bench authors.
SQuAD
SQuAD is a stanford question answering dataset associated with Stanford NLP.
Natural Questions
Natural Questions is a real-user question answering dataset associated with Google Research.
TriviaQA
TriviaQA is a reading comprehension dataset associated with University of Washington.
DROP
DROP is a discrete reasoning over paragraphs associated with Allen Institute for AI.
RACE
RACE is a reading comprehension dataset associated with Carnegie Mellon University.
ARC Challenge
ARC Challenge is a ai2 reasoning challenge associated with Allen Institute for AI.
MS MARCO
MS MARCO is a microsoft machine reading comprehension dataset associated with Microsoft.
Needle In A Haystack
Needle In A Haystack is a long-context retrieval test associated with Greg Kamradt.
VQAv2
VQAv2 is a visual question answering dataset associated with VQA authors.
COCO
COCO is a common objects in context dataset associated with COCO Consortium.
OpenImages
OpenImages is a large-scale image dataset associated with Google Research.
Common Crawl
Common Crawl is a web crawl corpus associated with Common Crawl Foundation.
C4
C4 is a colossal clean crawled corpus associated with Google Research.
Dolma
Dolma is a open language-model pretraining corpus associated with Allen Institute for AI.
RedPajama
RedPajama is a open reproduction data project associated with Together AI.
FineWeb
FineWeb is a large filtered web dataset associated with Hugging Face.
FineWeb-Edu
FineWeb-Edu is a educational web dataset associated with Hugging Face.
SlimPajama
SlimPajama is a deduplicated redpajama subset associated with Cerebras.
RefinedWeb
RefinedWeb is a filtered web dataset associated with Technology Innovation Institute.
OpenWebText
OpenWebText is a open reproduction of webtext associated with OpenWebText contributors.
WikiText
WikiText is a wikipedia language modeling dataset associated with Salesforce.
ROOTS Corpus
ROOTS Corpus is a multilingual pretraining corpus associated with BigScience.
StarCoderData
StarCoderData is a code pretraining dataset associated with BigCode.
The Stack
The Stack is a permissively licensed code dataset associated with BigCode.
CodeSearchNet
CodeSearchNet is a code search dataset associated with GitHub / Microsoft Research.
Dolly 15k
Dolly 15k is a instruction tuning dataset associated with Databricks.
OpenAssistant Conversations
OpenAssistant Conversations is a human-assistant conversation corpus associated with LAION OpenAssistant.
UltraChat
UltraChat is a large-scale chat dataset associated with Tsinghua NLP.
WildChat
WildChat is a real-world user-chat dataset associated with Allen Institute for AI.
LMSYS-Chat-1M
LMSYS-Chat-1M is a large-scale chat corpus associated with LMSYS.
HH-RLHF
HH-RLHF is a helpful and harmless rlhf dataset associated with Anthropic.
WikiText-103
WikiText-103 is a ai dataset / benchmark corpus associated with Salesforce Research.
FEVER Dataset
FEVER Dataset is a ai dataset / benchmark corpus associated with FEVER contributors.
Social IQA
Social IQA is a ai dataset / benchmark corpus associated with Allen Institute for AI.
WIDER FACE
WIDER FACE is a ai dataset / benchmark corpus associated with Chinese University of Hong Kong.
Kinetics Dataset
Kinetics Dataset is a ai dataset / benchmark corpus associated with DeepMind.
HMDB51
HMDB51 is a ai dataset / benchmark corpus associated with Brown University.
QQP
QQP is a ai dataset / benchmark corpus associated with Quora.
OntoNotes 5.0
OntoNotes 5.0 is a ai dataset / benchmark corpus associated with Linguistic Data Consortium.
WMT14 English-German
WMT14 English-German is a ai dataset / benchmark corpus associated with WMT.
IWSLT 2017
IWSLT 2017 is a ai dataset / benchmark corpus associated with IWSLT.
TIMIT
TIMIT is a ai dataset / benchmark corpus associated with Linguistic Data Consortium.
Switchboard Corpus
Switchboard Corpus is a ai dataset / benchmark corpus associated with Linguistic Data Consortium.
UCI Iris Dataset
UCI Iris Dataset is a machine learning dataset associated with UCI Machine Learning Repository.
UCI Adult Dataset
UCI Adult Dataset is a machine learning dataset associated with UCI Machine Learning Repository.
UCI Wine Dataset
UCI Wine Dataset is a machine learning dataset associated with UCI Machine Learning Repository.
UCI Breast Cancer Wisconsin Diagnostic
UCI Breast Cancer Wisconsin Diagnostic is a machine learning dataset associated with UCI Machine Learning Repository.
UCI Diabetes 130-US Hospitals
UCI Diabetes 130-US Hospitals is a machine learning dataset associated with UCI Machine Learning Repository.
UCI Heart Disease Dataset
UCI Heart Disease Dataset is a machine learning dataset associated with UCI Machine Learning Repository.
UCI Mushroom Dataset
UCI Mushroom Dataset is a machine learning dataset associated with UCI Machine Learning Repository.
UCI Abalone Dataset
UCI Abalone Dataset is a machine learning dataset associated with UCI Machine Learning Repository.
UCI Bank Marketing Dataset
UCI Bank Marketing Dataset is a machine learning dataset associated with UCI Machine Learning Repository.
UCI Bike Sharing Dataset
UCI Bike Sharing Dataset is a machine learning dataset associated with UCI Machine Learning Repository.
UCI Online Retail Dataset
UCI Online Retail Dataset is a machine learning dataset associated with UCI Machine Learning Repository.
UCI Human Activity Recognition Using Smartphones
UCI Human Activity Recognition Using Smartphones is a machine learning dataset associated with UCI Machine Learning Repository.
UCI Epileptic Seizure Recognition
UCI Epileptic Seizure Recognition is a machine learning dataset associated with UCI Machine Learning Repository.
UCI Air Quality Dataset
UCI Air Quality Dataset is a machine learning dataset associated with UCI Machine Learning Repository.
UCI Appliances Energy Prediction
UCI Appliances Energy Prediction is a machine learning dataset associated with UCI Machine Learning Repository.
UCI Individual Household Electric Power Consumption
UCI Individual Household Electric Power Consumption is a machine learning dataset associated with UCI Machine Learning Repository.
UCI Forest Fires Dataset
UCI Forest Fires Dataset is a machine learning dataset associated with UCI Machine Learning Repository.
UCI Student Performance Dataset
UCI Student Performance Dataset is a machine learning dataset associated with UCI Machine Learning Repository.
UCI Wine Quality Dataset
UCI Wine Quality Dataset is a machine learning dataset associated with UCI Machine Learning Repository.
UCI Communities and Crime
UCI Communities and Crime is a machine learning dataset associated with UCI Machine Learning Repository.
UCI CoverType Dataset
UCI CoverType Dataset is a machine learning dataset associated with UCI Machine Learning Repository.
UCI Letter Recognition Dataset
UCI Letter Recognition Dataset is a machine learning dataset associated with UCI Machine Learning Repository.
UCI Optical Recognition of Handwritten Digits
UCI Optical Recognition of Handwritten Digits is a machine learning dataset associated with UCI Machine Learning Repository.
UCI Statlog German Credit Data
UCI Statlog German Credit Data is a machine learning dataset associated with UCI Machine Learning Repository.
UCI Statlog Shuttle Dataset
UCI Statlog Shuttle Dataset is a machine learning dataset associated with UCI Machine Learning Repository.
UCI Statlog Landsat Satellite
UCI Statlog Landsat Satellite is a machine learning dataset associated with UCI Machine Learning Repository.
UCI Yeast Dataset
UCI Yeast Dataset is a machine learning dataset associated with UCI Machine Learning Repository.
UCI Zoo Dataset
UCI Zoo Dataset is a machine learning dataset associated with UCI Machine Learning Repository.
UCI Car Evaluation Dataset
UCI Car Evaluation Dataset is a machine learning dataset associated with UCI Machine Learning Repository.
UCI Nursery Dataset
UCI Nursery Dataset is a machine learning dataset associated with UCI Machine Learning Repository.
BraTS Challenge Dataset
BraTS Challenge Dataset is a biomedical / scientific dataset associated with BraTS organizers.
HAM10000
HAM10000 is a biomedical / scientific dataset associated with Harvard Dataverse / ISIC contributors.
PhysioNet Challenge
PhysioNet Challenge is a biomedical / scientific dataset associated with PhysioNet.
OpenNeuro
OpenNeuro is a biomedical / scientific dataset associated with OpenNeuro contributors.
Human Connectome Project Dataset
Human Connectome Project Dataset is a biomedical / scientific dataset associated with Human Connectome Project.
Allen Brain Atlas
Allen Brain Atlas is a biomedical / scientific dataset associated with Allen Institute.
Cora Citation Network
Cora Citation Network is a graph / scientific ml dataset associated with Graph learning community.
Citeseer Citation Network
Citeseer Citation Network is a graph / scientific ml dataset associated with Graph learning community.
PubMed Citation Network
PubMed Citation Network is a graph / scientific ml dataset associated with Graph learning community.
OGBN-Arxiv
OGBN-Arxiv is a graph / scientific ml dataset associated with Open Graph Benchmark.
OGBN-Products
OGBN-Products is a graph / scientific ml dataset associated with Open Graph Benchmark.
OGBN-Proteins
OGBN-Proteins is a graph / scientific ml dataset associated with Open Graph Benchmark.
OGBG-MolHIV
OGBG-MolHIV is a graph / scientific ml dataset associated with Open Graph Benchmark.
OGBG-MolPCBA
OGBG-MolPCBA is a graph / scientific ml dataset associated with Open Graph Benchmark.
OGBL-Collab
OGBL-Collab is a graph / scientific ml dataset associated with Open Graph Benchmark.
OGBL-DDI
OGBL-DDI is a graph / scientific ml dataset associated with Open Graph Benchmark.
OGBL-PPA
OGBL-PPA is a graph / scientific ml dataset associated with Open Graph Benchmark.
ZINC Dataset
ZINC Dataset is a graph / scientific ml dataset associated with ZINC contributors.
TUDatasets
TUDatasets is a graph / scientific ml dataset associated with TU Dortmund University.
MovieLens 100K
MovieLens 100K is a recommendation / user behavior dataset associated with GroupLens.
MovieLens 1M
MovieLens 1M is a recommendation / user behavior dataset associated with GroupLens.
MovieLens 20M
MovieLens 20M is a recommendation / user behavior dataset associated with GroupLens.
MovieLens 25M
MovieLens 25M is a recommendation / user behavior dataset associated with GroupLens.
Netflix Prize Dataset
Netflix Prize Dataset is a recommendation / user behavior dataset associated with Netflix.
Amazon Reviews Dataset
Amazon Reviews Dataset is a recommendation / user behavior dataset associated with Amazon / UCSD.
Yelp Open Dataset
Yelp Open Dataset is a recommendation / user behavior dataset associated with Yelp.
Last.fm Dataset
Last.fm Dataset is a recommendation / user behavior dataset associated with Last.fm / HetRec.
Million Song Dataset
Million Song Dataset is a recommendation / user behavior dataset associated with Columbia University / The Echo Nest.
Book-Crossing Dataset
Book-Crossing Dataset is a recommendation / user behavior dataset associated with Institut für Informatik Freiburg.
QuAC
QuAC is a language dataset / benchmark associated with Allen Institute for AI / University of Washington.
CoQA
CoQA is a language dataset / benchmark associated with Stanford NLP.
NarrativeQA
NarrativeQA is a language dataset / benchmark associated with DeepMind.
QASC
QASC is a language dataset / benchmark associated with Allen Institute for AI.
StrategyQA
StrategyQA is a language dataset / benchmark associated with Allen Institute for AI.
MuSiQue
MuSiQue is a language dataset / benchmark associated with MuSiQue contributors.
2WikiMultihopQA
2WikiMultihopQA is a language dataset / benchmark associated with 2WikiMultihopQA contributors.
WikiHop
WikiHop is a language dataset / benchmark associated with DeepMind.
DREAM
DREAM is a language dataset / benchmark associated with Tsinghua University.
ReClor
ReClor is a language dataset / benchmark associated with ReClor contributors.
LogiQA
LogiQA is a language dataset / benchmark associated with LogiQA contributors.
SciTail
SciTail is a language dataset / benchmark associated with Allen Institute for AI.
Adversarial NLI
Adversarial NLI is a language dataset / benchmark associated with Facebook AI Research / UNC / NYU.
WikiQA
WikiQA is a language dataset / benchmark associated with Microsoft Research.
TREC Question Classification
TREC Question Classification is a language dataset / benchmark associated with TREC / UIUC.
AG News
AG News is a language dataset / benchmark associated with AG News contributors.
DBpedia Ontology Dataset
DBpedia Ontology Dataset is a language dataset / benchmark associated with DBpedia contributors.
Yelp Review Polarity
Yelp Review Polarity is a language dataset / benchmark associated with Yelp / Zhang et al..
Yelp Review Full
Yelp Review Full is a language dataset / benchmark associated with Yelp / Zhang et al..
Amazon Polarity
Amazon Polarity is a language dataset / benchmark associated with Amazon / Zhang et al..
Amazon Reviews Multi
Amazon Reviews Multi is a language dataset / benchmark associated with Amazon / Hugging Face contributors.
IMDb Large Movie Review Dataset
IMDb Large Movie Review Dataset is a language dataset / benchmark associated with Stanford AI Lab.
Rotten Tomatoes Dataset
Rotten Tomatoes Dataset is a language dataset / benchmark associated with Cornell / Pang and Lee.
TweetEval
TweetEval is a language dataset / benchmark associated with TweetEval contributors.
Jigsaw Toxic Comment Classification
Jigsaw Toxic Comment Classification is a language dataset / benchmark associated with Jigsaw / Conversation AI.
Civil Comments
Civil Comments is a language dataset / benchmark associated with Jigsaw / Conversation AI.
GoEmotions
GoEmotions is a language dataset / benchmark associated with Google Research.
EmpatheticDialogues
EmpatheticDialogues is a language dataset / benchmark associated with Meta AI.
Wizard of Wikipedia
Wizard of Wikipedia is a language dataset / benchmark associated with Meta AI.
Persona-Chat
Persona-Chat is a language dataset / benchmark associated with Meta AI.
DailyDialog
DailyDialog is a language dataset / benchmark associated with DailyDialog contributors.
MultiWOZ
MultiWOZ is a language dataset / benchmark associated with Cambridge Dialogue Systems Group.
DSTC2
DSTC2 is a language dataset / benchmark associated with Dialogue State Tracking Challenge.
SNIPS NLU
SNIPS NLU is a language dataset / benchmark associated with Snips.
ATIS Dataset
ATIS Dataset is a language dataset / benchmark associated with Air Travel Information System contributors.
Banking77
Banking77 is a language dataset / benchmark associated with PolyAI.
CLINC150
CLINC150 is a language dataset / benchmark associated with CLINC / University of Michigan.
Massive
Massive is a language dataset / benchmark associated with Amazon Science.
XNLI Dataset
XNLI Dataset is a language dataset / benchmark associated with Facebook AI Research.
XTREME Benchmark
XTREME Benchmark is a language dataset / benchmark associated with Google Research.
XTREME-R Benchmark
XTREME-R Benchmark is a language dataset / benchmark associated with Google Research.
WMT19
WMT19 is a language dataset / benchmark associated with WMT.
WMT20
WMT20 is a language dataset / benchmark associated with WMT.
WMT21
WMT21 is a language dataset / benchmark associated with WMT.
WMT22
WMT22 is a language dataset / benchmark associated with WMT.
WMT23
WMT23 is a language dataset / benchmark associated with WMT.
WMT24
WMT24 is a language dataset / benchmark associated with WMT.
OPUS Corpus
OPUS Corpus is a language dataset / benchmark associated with OPUS project.
ParaCrawl
ParaCrawl is a language dataset / benchmark associated with ParaCrawl project.
CCMatrix
CCMatrix is a language dataset / benchmark associated with Meta AI.
CCAligned
CCAligned is a language dataset / benchmark associated with Meta AI.