Browse categories

Evaluation System

Autonomous EntityAI AssistantAI Answer EngineAI SystemSynthetic Personality PlatformSynthetic CompanionCreative IntelligenceVoice IntelligenceFictional IntelligenceAI OrganizationFoundation ModelHistorical AI SystemAgentic SystemRobotic IntelligenceDeveloper PlatformCreative Intelligence SystemRobotics and Embodied AIResearch SystemEvaluation SystemDatasetAI ToolAI InfrastructureScientific AI SystemEmbodied AI SystemCreative AI SystemSpecialized AI SystemAI AgentAI Research SystemAI Governance SystemAI FrameworkAI Training FrameworkAI KernelAI RuntimeMLOps PlatformSearch PlatformGraph AI PlatformRAG MethodLanguage ModelMachine Translation ModelSpeech ModelVision ModelAI DatasetAI BenchmarkScientific AI ModelScientific AI FrameworkAI HardwareEmbodied AIAI Writing SystemAI TutorAI Documentation AssistantFictional Synthetic PersonalityFictional Synthetic PersonMultimodal ModelMultimodal FrameworkVision FrameworkSpeech FrameworkAI Simulation PlatformAgentic AI SystemAI CompilerMachine Learning FrameworkEmbedding ModelReranker ModelMixture-of-Experts ModelCode ModelMachine Translation BenchmarkMachine Translation DatasetMedical AI DatasetMedical AI BenchmarkScientific DatasetScientific AI DatasetFictional RobotFictional Synthetic IntelligenceMath ModelAI Assistant ModelAI ArchitectureAI Alignment MethodReinforcement Learning MethodAI Reasoning MethodAutoML FrameworkWorkflow PlatformData PlatformAI Evaluation PlatformAI Observability PlatformAI GatewayImage Generation ModelImage Control ModelImage Fine-Tuning MethodAudio Generation ModelAudio Codec ModelSpeech Generation ModelAudio Generation FrameworkVoice Conversion SystemRobotics Simulation PlatformRobotics BenchmarkRobotics DatasetAI PlatformAI Search AssistantAI Data PlatformAI InterfaceRobotics SoftwareAI Safety ToolAI Evaluation FrameworkAI Safety ModelComputer Vision ModelAI Narrative SystemAI LibraryAI ModelDataset / BenchmarkOrganization / PlatformEmbodied AI / RoboticsBenchmark

ENT-00001077

MMLU-Pro

MMLU-Pro is a advanced multitask language benchmark associated with TIGER-Lab.

ENT-00001078

Humanity's Last Exam

Humanity's Last Exam is a expert-level multimodal benchmark associated with Center for AI Safety.

ENT-00001079

AIME

AIME is a mathematics contest benchmark associated with Mathematical Association of America.

ENT-00001081

APPS

APPS is a introductory programming benchmark associated with APPS authors.

ENT-00001084

SWE-bench Live

SWE-bench Live is a live software engineering benchmark associated with SWE-bench researchers.

ENT-00001085

LiveCodeBench

LiveCodeBench is a contamination-resistant code benchmark associated with LiveCodeBench authors.

ENT-00001086

Arena-Hard

Arena-Hard is a hard prompt arena benchmark associated with LMSYS.

ENT-00001087

HELM Lite

HELM Lite is a lightweight helm evaluation associated with Stanford CRFM.

ENT-00001088

LM Evaluation Harness

LM Evaluation Harness is a language model evaluation framework associated with EleutherAI.

ENT-00001089

Open LLM Leaderboard

Open LLM Leaderboard is a open model evaluation leaderboard associated with Hugging Face.

ENT-00001091

GLUE

GLUE is a general language understanding evaluation associated with NYU / DeepMind / University of Washington.

ENT-00001092

SuperGLUE

SuperGLUE is a advanced language understanding benchmark associated with SuperGLUE authors.

ENT-00001094

HotpotQA

HotpotQA is a multi-hop question answering benchmark associated with HotpotQA authors.

ENT-00001097

FEVER

FEVER is a fact extraction and verification benchmark associated with FEVER authors.

ENT-00001100

CommonsenseQA

CommonsenseQA is a commonsense question answering benchmark associated with Allen Institute for AI.

ENT-00001101

WinoGrande

WinoGrande is a adversarial winograd benchmark associated with Allen Institute for AI.

ENT-00001102

TruthfulQA

TruthfulQA is a truthfulness benchmark associated with TruthfulQA authors.

ENT-00001103

HellaSwag

HellaSwag is a commonsense inference benchmark associated with Allen Institute for AI.

ENT-00001105

LAMBADA

LAMBADA is a language modeling benchmark associated with LAMBADA authors.

ENT-00001106

XNLI

XNLI is a cross-lingual nli benchmark associated with Facebook AI Research.

ENT-00001107

XTREME

XTREME is a cross-lingual transfer benchmark associated with Google Research.

ENT-00001108

TyDi QA

TyDi QA is a typologically diverse qa benchmark associated with Google Research.

ENT-00001109

MIRACL

MIRACL is a multilingual retrieval benchmark associated with University of Waterloo.

ENT-00001110

BEIR

BEIR is a information retrieval benchmark associated with UKP Lab.

ENT-00001112

TREC Deep Learning Track

TREC Deep Learning Track is a neural ir evaluation track associated with NIST TREC.

ENT-00001113

KILT

KILT is a knowledge-intensive language tasks benchmark associated with Meta AI.

ENT-00001114

LongBench

LongBench is a long-context language model benchmark associated with THUDM.

ENT-00001115

RULER

RULER is a long-context evaluation benchmark associated with NVIDIA.

ENT-00001117

L-Eval

L-Eval is a long-context evaluation benchmark associated with OpenBMB.

ENT-00001118

CRAG

CRAG is a comprehensive rag benchmark associated with Meta AI.

ENT-00001119

ToolBench

ToolBench is a tool-use benchmark associated with ToolBench authors.

ENT-00001120

BFCL

BFCL is a berkeley function-calling leaderboard associated with Berkeley Sky Computing Lab.

ENT-00001121

Gorilla OpenFunctions

Gorilla OpenFunctions is a function-calling model benchmark associated with Berkeley Sky Computing Lab.

ENT-00001122

GAIA

GAIA is a general ai assistant benchmark associated with GAIA authors.

ENT-00001123

AgentBench

AgentBench is a agent evaluation benchmark associated with THUDM.

ENT-00001124

WebArena

WebArena is a web agent benchmark associated with WebArena authors.

ENT-00001125

MiniWoB++

MiniWoB++ is a web interaction benchmark associated with Farama Foundation.

ENT-00001126

OSWorld

OSWorld is a computer-use agent benchmark associated with OSWorld authors.

ENT-00001127

VisualWebArena

VisualWebArena is a visual web agent benchmark associated with VisualWebArena authors.

ENT-00001128

MMMU

MMMU is a massive multi-discipline multimodal benchmark associated with MMMU authors.

ENT-00001129

MMBench

MMBench is a multimodal benchmark associated with OpenCompass.

ENT-00001131

GQA

GQA is a visual reasoning benchmark associated with Stanford / GQA authors.

ENT-00001132

DocVQA

DocVQA is a document visual question answering benchmark associated with DocVQA authors.

ENT-00001133

ChartQA

ChartQA is a chart question answering benchmark associated with ChartQA authors.

ENT-00001134

MathVista

MathVista is a mathematical visual reasoning benchmark associated with MathVista authors.

ENT-00001493

MLPerf Training

MLPerf Training is a machine learning performance benchmark associated with MLCommons.

ENT-00001494

MLPerf Inference

MLPerf Inference is a machine learning inference benchmark associated with MLCommons.

ENT-00001495

AILuminate

AILuminate is a ai risk and reliability benchmark associated with MLCommons.

ENT-00001496

Stanford AI Index

Stanford AI Index is a ai measurement and analysis report associated with Stanford HAI.

ENT-00001497

Epoch AI Benchmarking

Epoch AI Benchmarking is a ai trends and benchmark analysis project associated with Epoch AI.

ENT-00001498

METR

METR is a ai evaluations research organization associated with METR.

ENT-00001500

LMArena

LMArena is a human preference ai evaluation platform associated with LMArena.