MI
ENT-00002676 · Benchmark

Mind2Web

Mind2Web is a AI benchmark associated with Ohio State NLP, classified in LXKeys.world as AI Benchmark / Evaluation.

Level V PersistentActiveDocumented
Human record

Overview

Mind2Web is an AI evaluation entity associated with Ohio State NLP. It is documented as web navigation dataset and helps evaluate model behavior, safety, reasoning, agents, tools, code, reinforcement learning or embodied interaction.

Timeline

  1. 2023Initial benchmark release

    Mind2Web entered the documented public record in 2023. This event is retained at the precision supported by the Entry’s reviewed source history.

Capabilities

Evaluation protocol documentationModel comparison supportTask taxonomy mappingReproducible benchmark anchoringRelationship graph compatibility

Known limitations

Benchmark results can become stale as models improve and evaluation protocols evolve.A benchmark measures a defined task scope rather than complete intelligence.
Technical and structured detail

Technical description

Structured LXKeys.world registry record for Mind2Web. Entity type: Benchmark; classification: AI Benchmark / Evaluation; creator/organization context: Ohio State NLP. Canonical source anchor: https://osu-nlp-group.github.io/Mind2Web/. The record tracks source authority, first public appearance, timeline, capabilities, limitations, registry status and graph relationships. Automatic refresh is limited to sources explicitly classified as OFFICIAL and enabled for updates; documentary and research references remain non-authoritative unless reviewed.

Controlled tags

BenchmarkEvaluationAI SafetyAgent Evaluation
Relationship graph

Mind2Web

Open in full graph
Mind2Web
Associated organizationOutgoing
Ohio State NLP

Ohio State NLP is the organization, project community or institutional context associated with Mind2Web.

Domain contextOutgoing
AI evaluation

Mind2Web belongs to the AI evaluation layer of the intelligent-entity registry.

Machine readable layer

Structured entry data for human tools, AI systems and machine clients.

{
    "@context": [
        "https://schema.org",
        {
            "lxw": "https://lxkeys.world/schema/"
        }
    ],
    "@type": "Thing",
    "identifier": "ENT-00002676",
    "name": "Mind2Web",
    "alternateName": [],
    "additionalType": {
        "category": "Data and Evaluation",
        "type": "Benchmark",
        "subtype": "",
        "lxkeysEntity": false
    },
    "description": "Mind2Web is a AI benchmark associated with Ohio State NLP, classified in LXKeys.world as AI Benchmark / Evaluation.",
    "creator": "Ohio State NLP",
    "url": "https://lxkeys.world/entry.php?id=ENT-00002676&lang=en",
    "sameAs": "https://osu-nlp-group.github.io/Mind2Web/",
    "image": "",
    "lxkeysWorld": {
        "worldId": "ENT-00002676",
        "kind": "Data and Evaluation",
        "type": "Benchmark",
        "subtype": "",
        "classification": "AI Benchmark / Evaluation",
        "organization": "Ohio State NLP",
        "originContext": "AI evaluation",
        "firstPublicAppearance": "2023",
        "currentStatus": "Active",
        "documentationStatus": "Documented",
        "documentationIndex": {
            "total": 82,
            "documentation": 25,
            "evidence": 13,
            "structure": 25,
            "relationships": 19,
            "level": "Level V Persistent"
        },
        "lxCalendarium": {
            "start_date_utc": "2023-04-01",
            "created_utc": "2026-06-17T01:48:33+00:00",
            "created_dypclt": "D-0 Y-2 P-3 C-3 L-22 T-4",
            "updated_utc": "2026-06-17T01:48:33+00:00",
            "updated_dypclt": "D-0 Y-2 P-3 C-3 L-22 T-4"
        },
        "facts": [],
        "capabilities": [
            "Evaluation protocol documentation",
            "Model comparison support",
            "Task taxonomy mapping",
            "Reproducible benchmark anchoring",
            "Relationship graph compatibility"
        ],
        "limitations": [
            "Benchmark results can become stale as models improve and evaluation protocols evolve.",
            "A benchmark measures a defined task scope rather than complete intelligence."
        ],
        "tags": [
            "Benchmark",
            "Evaluation",
            "AI Safety",
            "Agent Evaluation"
        ],
        "timeline": [
            {
                "date": "2023",
                "title": "Initial benchmark release",
                "description": "Mind2Web entered the documented public record in 2023. This event is retained at the precision supported by the Entry’s reviewed source history.",
                "source_url": "https://osu-nlp-group.github.io/Mind2Web/",
                "verification_status": "source_backed_official"
            }
        ],
        "relationships": [
            {
                "target": "Ohio State NLP",
                "type": "Associated organization",
                "description": "Ohio State NLP is the organization, project community or institutional context associated with Mind2Web.",
                "evidence_level": "documentary"
            },
            {
                "target": "AI evaluation",
                "type": "Domain context",
                "description": "Mind2Web belongs to the AI evaluation layer of the intelligent-entity registry.",
                "evidence_level": "documentary"
            }
        ],
        "sources": [
            {
                "label": "Official documentation",
                "url": "https://osu-nlp-group.github.io/Mind2Web/",
                "source_type": "Official / Research Source",
                "verification_status": "verified",
                "authority": "OFFICIAL",
                "role": "primary",
                "update_enabled": true,
                "authority_basis": "curated-corpus-refresh-2026-09-08"
            }
        ],
        "canonical": [],
        "imageMeta": []
    }
}
Proofs

Contribute to this record

Submit a source, correction or comment. Public changes remain moderated.

Submit proof or comment

Approved comments

No approved public comment yet.