Evaluation System
Autonomous EntityAI AssistantAI Answer EngineAI SystemSynthetic Personality PlatformSynthetic CompanionCreative IntelligenceVoice IntelligenceFictional IntelligenceAI OrganizationFoundation ModelHistorical AI SystemAgentic SystemRobotic IntelligenceDeveloper PlatformCreative Intelligence SystemRobotics and Embodied AIResearch SystemEvaluation SystemDatasetAI ToolAI InfrastructureScientific AI SystemEmbodied AI SystemCreative AI SystemSpecialized AI SystemAI AgentAI Research SystemAI Governance SystemAI FrameworkAI Training FrameworkAI KernelAI RuntimeMLOps PlatformSearch PlatformGraph AI PlatformRAG MethodLanguage ModelMachine Translation ModelSpeech ModelVision ModelAI DatasetAI BenchmarkScientific AI ModelScientific AI FrameworkAI HardwareEmbodied AIAI Writing SystemAI TutorAI Documentation AssistantFictional Synthetic PersonalityFictional Synthetic PersonMultimodal ModelMultimodal FrameworkVision FrameworkSpeech FrameworkAI Simulation PlatformAgentic AI SystemAI CompilerMachine Learning FrameworkEmbedding ModelReranker ModelMixture-of-Experts ModelCode ModelMachine Translation BenchmarkMachine Translation DatasetMedical AI DatasetMedical AI BenchmarkScientific DatasetScientific AI DatasetFictional RobotFictional Synthetic IntelligenceMath ModelAI Assistant ModelAI ArchitectureAI Alignment MethodReinforcement Learning MethodAI Reasoning MethodAutoML FrameworkWorkflow PlatformData PlatformAI Evaluation PlatformAI Observability PlatformAI GatewayImage Generation ModelImage Control ModelImage Fine-Tuning MethodAudio Generation ModelAudio Codec ModelSpeech Generation ModelAudio Generation FrameworkVoice Conversion SystemRobotics Simulation PlatformRobotics BenchmarkRobotics DatasetAI PlatformAI Search AssistantAI Data PlatformAI InterfaceRobotics SoftwareAI Safety ToolAI Evaluation FrameworkAI Safety ModelComputer Vision ModelAI Narrative SystemAI LibraryAI ModelDataset / BenchmarkOrganization / PlatformEmbodied AI / RoboticsBenchmark
MMLU-Pro
MMLU-Pro is a advanced multitask language benchmark associated with TIGER-Lab.
Humanity's Last Exam
Humanity's Last Exam is a expert-level multimodal benchmark associated with Center for AI Safety.
AIME
AIME is a mathematics contest benchmark associated with Mathematical Association of America.
APPS
APPS is a introductory programming benchmark associated with APPS authors.
SWE-bench Live
SWE-bench Live is a live software engineering benchmark associated with SWE-bench researchers.
LiveCodeBench
LiveCodeBench is a contamination-resistant code benchmark associated with LiveCodeBench authors.
Arena-Hard
Arena-Hard is a hard prompt arena benchmark associated with LMSYS.
HELM Lite
HELM Lite is a lightweight helm evaluation associated with Stanford CRFM.
LM Evaluation Harness
LM Evaluation Harness is a language model evaluation framework associated with EleutherAI.
Open LLM Leaderboard
Open LLM Leaderboard is a open model evaluation leaderboard associated with Hugging Face.
GLUE
GLUE is a general language understanding evaluation associated with NYU / DeepMind / University of Washington.
SuperGLUE
SuperGLUE is a advanced language understanding benchmark associated with SuperGLUE authors.
HotpotQA
HotpotQA is a multi-hop question answering benchmark associated with HotpotQA authors.
FEVER
FEVER is a fact extraction and verification benchmark associated with FEVER authors.
CommonsenseQA
CommonsenseQA is a commonsense question answering benchmark associated with Allen Institute for AI.
WinoGrande
WinoGrande is a adversarial winograd benchmark associated with Allen Institute for AI.
TruthfulQA
TruthfulQA is a truthfulness benchmark associated with TruthfulQA authors.
HellaSwag
HellaSwag is a commonsense inference benchmark associated with Allen Institute for AI.
LAMBADA
LAMBADA is a language modeling benchmark associated with LAMBADA authors.
XNLI
XNLI is a cross-lingual nli benchmark associated with Facebook AI Research.
XTREME
XTREME is a cross-lingual transfer benchmark associated with Google Research.
TyDi QA
TyDi QA is a typologically diverse qa benchmark associated with Google Research.
MIRACL
MIRACL is a multilingual retrieval benchmark associated with University of Waterloo.
BEIR
BEIR is a information retrieval benchmark associated with UKP Lab.
TREC Deep Learning Track
TREC Deep Learning Track is a neural ir evaluation track associated with NIST TREC.
KILT
KILT is a knowledge-intensive language tasks benchmark associated with Meta AI.
LongBench
LongBench is a long-context language model benchmark associated with THUDM.
RULER
RULER is a long-context evaluation benchmark associated with NVIDIA.
L-Eval
L-Eval is a long-context evaluation benchmark associated with OpenBMB.
CRAG
CRAG is a comprehensive rag benchmark associated with Meta AI.
ToolBench
ToolBench is a tool-use benchmark associated with ToolBench authors.
BFCL
BFCL is a berkeley function-calling leaderboard associated with Berkeley Sky Computing Lab.
Gorilla OpenFunctions
Gorilla OpenFunctions is a function-calling model benchmark associated with Berkeley Sky Computing Lab.
GAIA
GAIA is a general ai assistant benchmark associated with GAIA authors.
AgentBench
AgentBench is a agent evaluation benchmark associated with THUDM.
WebArena
WebArena is a web agent benchmark associated with WebArena authors.
MiniWoB++
MiniWoB++ is a web interaction benchmark associated with Farama Foundation.
OSWorld
OSWorld is a computer-use agent benchmark associated with OSWorld authors.
VisualWebArena
VisualWebArena is a visual web agent benchmark associated with VisualWebArena authors.
MMMU
MMMU is a massive multi-discipline multimodal benchmark associated with MMMU authors.
MMBench
MMBench is a multimodal benchmark associated with OpenCompass.
GQA
GQA is a visual reasoning benchmark associated with Stanford / GQA authors.
DocVQA
DocVQA is a document visual question answering benchmark associated with DocVQA authors.
ChartQA
ChartQA is a chart question answering benchmark associated with ChartQA authors.
MathVista
MathVista is a mathematical visual reasoning benchmark associated with MathVista authors.
MLPerf Training
MLPerf Training is a machine learning performance benchmark associated with MLCommons.
MLPerf Inference
MLPerf Inference is a machine learning inference benchmark associated with MLCommons.
AILuminate
AILuminate is a ai risk and reliability benchmark associated with MLCommons.
Stanford AI Index
Stanford AI Index is a ai measurement and analysis report associated with Stanford HAI.
Epoch AI Benchmarking
Epoch AI Benchmarking is a ai trends and benchmark analysis project associated with Epoch AI.
METR
METR is a ai evaluations research organization associated with METR.
LMArena
LMArena is a human preference ai evaluation platform associated with LMArena.