BH
ENT-00002503 · Benchmark

BIG-Bench Hard Disambiguation QA

BIG-Bench Hard Disambiguation QA is a language model evaluation task associated with BIG-Bench.

Level V — PersistentActiveDocumented
Human readable record

Overview

BIG-Bench Hard Disambiguation QA is a BIG-Bench Hard evaluation task used to test reasoning behavior in language models. It belongs to a curated subset of challenging BIG-Bench tasks and supports comparative model evaluation across reasoning, symbolic manipulation, factual discrimination or structured inference.

Timeline

  1. 2022Public release or documentation

    BIG-Bench Hard Disambiguation QA appears in public documentation, project records, benchmark descriptions or research references associated with BIG-Bench.

Capabilities

Evaluation protocol documentationModel comparison supportTask taxonomy mappingReproducible benchmark anchoringRelationship graph compatibility

Known limitations

Benchmark results can become stale as models improve and evaluation protocols evolve.A benchmark measures a defined task scope rather than complete intelligence.
Technical and registry detail

Technical description

Structured record for BIG-Bench Hard Disambiguation QA. The technical layer identifies the entity type, classification, creator or organization, official source anchor, capabilities, limitations, timeline entries and graph relationships without adding internal import events to the public history.

Controlled tags

BenchmarkEvaluationReasoningBIG-Bench Hard
Relationship graph
BIG-Bench Hard Disambiguation QA
Associated organizationBIG-Bench

BIG-Bench is the organization, project community or institutional context associated with BIG-Bench Hard Disambiguation QA.

Domain contextLanguage model evaluation

BIG-Bench Hard Disambiguation QA belongs to the Language model evaluation layer of the intelligent-entity registry.

Machine readable layer

Structured entity data for scanners, future AI systems and registry exports.

{
    "@context": "https://schema.org",
    "@type": "Thing",
    "identifier": "ENT-00002503",
    "name": "BIG-Bench Hard Disambiguation QA",
    "alternateName": [],
    "additionalType": "Benchmark",
    "description": "BIG-Bench Hard Disambiguation QA is a language model evaluation task associated with BIG-Bench.",
    "creator": "Google Research / BIG-Bench contributors",
    "url": "entity.php?id=ENT-00002503",
    "sameAs": "https://github.com/suzgunmirac/BIG-Bench-Hard",
    "lxkeysWorld": {
        "classification": "Language Model Evaluation Task",
        "organization": "BIG-Bench",
        "originContext": "Language model evaluation",
        "firstPublicAppearance": "2022",
        "currentStatus": "Active",
        "registryStatus": "Documented",
        "spatiumIndex": {
            "total": 82,
            "documentation": 25,
            "evidence": 13,
            "structure": 25,
            "relationships": 19,
            "level": "Level V — Persistent"
        },
        "lxCalendarium": {
            "start_date_utc": "2023-04-01",
            "created_utc": "2026-06-17T01:48:33+00:00",
            "created_dypclt": "D-0 Y-2 P-3 C-3 L-22 T-4",
            "updated_utc": "2026-06-17T01:48:33+00:00",
            "updated_dypclt": "D-0 Y-2 P-3 C-3 L-22 T-4",
            "reviewed_utc": "2026-06-17T01:48:33+00:00",
            "reviewed_dypclt": "D-0 Y-2 P-3 C-3 L-22 T-4"
        },
        "capabilities": [
            "Evaluation protocol documentation",
            "Model comparison support",
            "Task taxonomy mapping",
            "Reproducible benchmark anchoring",
            "Relationship graph compatibility"
        ],
        "limitations": [
            "Benchmark results can become stale as models improve and evaluation protocols evolve.",
            "A benchmark measures a defined task scope rather than complete intelligence."
        ],
        "tags": [
            "Benchmark",
            "Evaluation",
            "Reasoning",
            "BIG-Bench Hard"
        ],
        "timeline": [
            {
                "date": "2022",
                "title": "Public release or documentation",
                "description": "BIG-Bench Hard Disambiguation QA appears in public documentation, project records, benchmark descriptions or research references associated with BIG-Bench."
            }
        ],
        "relationships": [
            {
                "target": "BIG-Bench",
                "type": "Associated organization",
                "description": "BIG-Bench is the organization, project community or institutional context associated with BIG-Bench Hard Disambiguation QA.",
                "evidence_level": "documentary"
            },
            {
                "target": "Language model evaluation",
                "type": "Domain context",
                "description": "BIG-Bench Hard Disambiguation QA belongs to the Language model evaluation layer of the intelligent-entity registry.",
                "evidence_level": "documentary"
            }
        ],
        "sources": [
            {
                "label": "BIG-Bench Hard repository",
                "url": "https://github.com/suzgunmirac/BIG-Bench-Hard",
                "source_type": "Research Source",
                "verification_status": "verified"
            }
        ]
    }
}
Proof and discussion layer

Contribute to this record

Submit a proof, correction or comment. Public display is moderated. Every submission remains preserved in the export archive.

Submit proof or comment

Approved comments

No approved public comment yet.