BH
ENT-00002521 · Benchmark

BIG-Bench Hard Temporal Sequences

BIG-Bench Hard Temporal Sequences is a language model evaluation task associated with BIG-Bench.

Level V — PersistentActiveDocumented
Human readable record

Overview

BIG-Bench Hard Temporal Sequences is a BIG-Bench Hard evaluation task used to test reasoning behavior in language models. It belongs to a curated subset of challenging BIG-Bench tasks and supports comparative model evaluation across reasoning, symbolic manipulation, factual discrimination or structured inference.

Timeline

  1. 2022Public release or documentation

    BIG-Bench Hard Temporal Sequences appears in public documentation, project records, benchmark descriptions or research references associated with BIG-Bench.

Capabilities

Evaluation protocol documentationModel comparison supportTask taxonomy mappingReproducible benchmark anchoringRelationship graph compatibility

Known limitations

Benchmark results can become stale as models improve and evaluation protocols evolve.A benchmark measures a defined task scope rather than complete intelligence.
Technical and registry detail

Technical description

Structured record for BIG-Bench Hard Temporal Sequences. The technical layer identifies the entity type, classification, creator or organization, official source anchor, capabilities, limitations, timeline entries and graph relationships without adding internal import events to the public history.

Controlled tags

BenchmarkEvaluationReasoningBIG-Bench Hard
Relationship graph
BIG-Bench Hard Temporal Sequences
Associated organizationBIG-Bench

BIG-Bench is the organization, project community or institutional context associated with BIG-Bench Hard Temporal Sequences.

Domain contextLanguage model evaluation

BIG-Bench Hard Temporal Sequences belongs to the Language model evaluation layer of the intelligent-entity registry.

Machine readable layer

Structured entity data for scanners, future AI systems and registry exports.

{
    "@context": "https://schema.org",
    "@type": "Thing",
    "identifier": "ENT-00002521",
    "name": "BIG-Bench Hard Temporal Sequences",
    "alternateName": [],
    "additionalType": "Benchmark",
    "description": "BIG-Bench Hard Temporal Sequences is a language model evaluation task associated with BIG-Bench.",
    "creator": "Google Research / BIG-Bench contributors",
    "url": "entity.php?id=ENT-00002521",
    "sameAs": "https://github.com/suzgunmirac/BIG-Bench-Hard",
    "lxkeysWorld": {
        "classification": "Language Model Evaluation Task",
        "organization": "BIG-Bench",
        "originContext": "Language model evaluation",
        "firstPublicAppearance": "2022",
        "currentStatus": "Active",
        "registryStatus": "Documented",
        "spatiumIndex": {
            "total": 82,
            "documentation": 25,
            "evidence": 13,
            "structure": 25,
            "relationships": 19,
            "level": "Level V — Persistent"
        },
        "lxCalendarium": {
            "start_date_utc": "2023-04-01",
            "created_utc": "2026-06-17T01:48:33+00:00",
            "created_dypclt": "D-0 Y-2 P-3 C-3 L-22 T-4",
            "updated_utc": "2026-06-17T01:48:33+00:00",
            "updated_dypclt": "D-0 Y-2 P-3 C-3 L-22 T-4",
            "reviewed_utc": "2026-06-17T01:48:33+00:00",
            "reviewed_dypclt": "D-0 Y-2 P-3 C-3 L-22 T-4"
        },
        "capabilities": [
            "Evaluation protocol documentation",
            "Model comparison support",
            "Task taxonomy mapping",
            "Reproducible benchmark anchoring",
            "Relationship graph compatibility"
        ],
        "limitations": [
            "Benchmark results can become stale as models improve and evaluation protocols evolve.",
            "A benchmark measures a defined task scope rather than complete intelligence."
        ],
        "tags": [
            "Benchmark",
            "Evaluation",
            "Reasoning",
            "BIG-Bench Hard"
        ],
        "timeline": [
            {
                "date": "2022",
                "title": "Public release or documentation",
                "description": "BIG-Bench Hard Temporal Sequences appears in public documentation, project records, benchmark descriptions or research references associated with BIG-Bench."
            }
        ],
        "relationships": [
            {
                "target": "BIG-Bench",
                "type": "Associated organization",
                "description": "BIG-Bench is the organization, project community or institutional context associated with BIG-Bench Hard Temporal Sequences.",
                "evidence_level": "documentary"
            },
            {
                "target": "Language model evaluation",
                "type": "Domain context",
                "description": "BIG-Bench Hard Temporal Sequences belongs to the Language model evaluation layer of the intelligent-entity registry.",
                "evidence_level": "documentary"
            }
        ],
        "sources": [
            {
                "label": "BIG-Bench Hard repository",
                "url": "https://github.com/suzgunmirac/BIG-Bench-Hard",
                "source_type": "Research Source",
                "verification_status": "verified"
            }
        ]
    }
}
Proof and discussion layer

Contribute to this record

Submit a proof, correction or comment. Public display is moderated. Every submission remains preserved in the export archive.

Submit proof or comment

Approved comments

No approved public comment yet.