O5
ENT-00002614 · Dataset

OntoNotes 5.0

OntoNotes 5.0 is a ai dataset / benchmark corpus associated with Linguistic Data Consortium.

Level V — PersistentActiveDocumented
Human readable record

Overview

OntoNotes 5.0 is a documented dataset or benchmark corpus used in machine learning, computer vision, speech, language processing or multimodal AI. The record identifies its role as annotated multilingual corpus and supports traceability between models, tasks and evaluation results.

Timeline

  1. 2013Public release or documentation

    OntoNotes 5.0 appears in public documentation, project records, benchmark descriptions or research references associated with Linguistic Data Consortium.

Capabilities

Dataset referenceBenchmark or training-data contextTask-level documentationModel evaluation supportSource-based traceability

Known limitations

Dataset coverage, licensing, annotation quality and benchmark relevance depend on the source version and use context.Performance claims should be assessed through models evaluated on the dataset rather than inferred from the dataset alone.
Technical and registry detail

Technical description

Structured record for OntoNotes 5.0. The technical layer identifies the entity type, classification, creator or organization, official source anchor, capabilities, limitations, timeline entries and graph relationships without adding internal import events to the public history.

Controlled tags

DatasetBenchmarkTraining DataEvaluation
Relationship graph
OntoNotes 5.0
Associated organizationLinguistic Data Consortium

Linguistic Data Consortium is the organization, project community or institutional context associated with OntoNotes 5.0.

Domain contextDataset and benchmark infrastructure

OntoNotes 5.0 belongs to the Dataset and benchmark infrastructure layer of the intelligent-entity registry.

Machine readable layer

Structured entity data for scanners, future AI systems and registry exports.

{
    "@context": "https://schema.org",
    "@type": "Thing",
    "identifier": "ENT-00002614",
    "name": "OntoNotes 5.0",
    "alternateName": [],
    "additionalType": "Dataset",
    "description": "OntoNotes 5.0 is a ai dataset / benchmark corpus associated with Linguistic Data Consortium.",
    "creator": "Linguistic Data Consortium",
    "url": "entity.php?id=ENT-00002614",
    "sameAs": "https://catalog.ldc.upenn.edu/LDC2013T19",
    "lxkeysWorld": {
        "classification": "AI Dataset / Benchmark Corpus",
        "organization": "Linguistic Data Consortium",
        "originContext": "Dataset and benchmark infrastructure",
        "firstPublicAppearance": "2013",
        "currentStatus": "Active",
        "registryStatus": "Documented",
        "spatiumIndex": {
            "total": 82,
            "documentation": 25,
            "evidence": 13,
            "structure": 25,
            "relationships": 19,
            "level": "Level V — Persistent"
        },
        "lxCalendarium": {
            "start_date_utc": "2023-04-01",
            "created_utc": "2026-06-17T01:48:33+00:00",
            "created_dypclt": "D-0 Y-2 P-3 C-3 L-22 T-4",
            "updated_utc": "2026-06-17T01:48:33+00:00",
            "updated_dypclt": "D-0 Y-2 P-3 C-3 L-22 T-4",
            "reviewed_utc": "2026-06-17T01:48:33+00:00",
            "reviewed_dypclt": "D-0 Y-2 P-3 C-3 L-22 T-4"
        },
        "capabilities": [
            "Dataset reference",
            "Benchmark or training-data context",
            "Task-level documentation",
            "Model evaluation support",
            "Source-based traceability"
        ],
        "limitations": [
            "Dataset coverage, licensing, annotation quality and benchmark relevance depend on the source version and use context.",
            "Performance claims should be assessed through models evaluated on the dataset rather than inferred from the dataset alone."
        ],
        "tags": [
            "Dataset",
            "Benchmark",
            "Training Data",
            "Evaluation"
        ],
        "timeline": [
            {
                "date": "2013",
                "title": "Public release or documentation",
                "description": "OntoNotes 5.0 appears in public documentation, project records, benchmark descriptions or research references associated with Linguistic Data Consortium."
            }
        ],
        "relationships": [
            {
                "target": "Linguistic Data Consortium",
                "type": "Associated organization",
                "description": "Linguistic Data Consortium is the organization, project community or institutional context associated with OntoNotes 5.0.",
                "evidence_level": "documentary"
            },
            {
                "target": "Dataset and benchmark infrastructure",
                "type": "Domain context",
                "description": "OntoNotes 5.0 belongs to the Dataset and benchmark infrastructure layer of the intelligent-entity registry.",
                "evidence_level": "documentary"
            }
        ],
        "sources": [
            {
                "label": "Official or reference dataset page",
                "url": "https://catalog.ldc.upenn.edu/LDC2013T19",
                "source_type": "Primary / Reference Source",
                "verification_status": "verified"
            }
        ]
    }
}
Proof and discussion layer

Contribute to this record

Submit a proof, correction or comment. Public display is moderated. Every submission remains preserved in the export archive.

Submit proof or comment

Approved comments

No approved public comment yet.