AD
ENT-00002671 · Benchmark

AdvBench

AdvBench is a AI benchmark associated with LLM safety research community, classified in LXKeys.world as AI Benchmark / Evaluation.

Niveau V PersistantActifDocumenté
Fiche publique

Vue d’ensemble

AdvBench is an AI evaluation entity associated with LLM safety research community. It is documented as adversarial behavior benchmark and helps evaluate model behavior, safety, reasoning, agents, tools, code, reinforcement learning or embodied interaction.

Chronologie

  1. 2023Initial benchmark release

    AdvBench entered the documented public record in 2023. This event is retained at the precision supported by the Entry’s reviewed source history.

Capacités

Evaluation protocol documentationModel comparison supportTask taxonomy mappingReproducible benchmark anchoringRelationship graph compatibility

Limites connues

Benchmark results can become stale as models improve and evaluation protocols evolve.A benchmark measures a defined task scope rather than complete intelligence.
Détail technique et structuré

Description technique

Structured LXKeys.world registry record for AdvBench. Entity type: Benchmark; classification: AI Benchmark / Evaluation; creator/organization context: LLM safety research community. Canonical source anchor: https://github.com/llm-attacks/llm-attacks. The record tracks source authority, first public appearance, timeline, capabilities, limitations, registry status and graph relationships. Automatic refresh is limited to sources explicitly classified as OFFICIAL and enabled for updates; documentary and research references remain non-authoritative unless reviewed.

Tags contrôlés

BenchmarkEvaluationAI SafetyAgent Evaluation
Graphe relationnel

AdvBench

Ouvrir dans le graphe complet
AdvBench
Organisation associéeSortante
LLM safety research community

LLM safety research community is the organization, project community or institutional context associated with AdvBench.

Contexte de domaineSortante
AI evaluation

AdvBench belongs to the AI evaluation layer of the intelligent-entity registry.

Couche lisible par machine

Données structurées de l’entrée pour les outils humains, les systèmes IA et les clients machine.

{
    "@context": [
        "https://schema.org",
        {
            "lxw": "https://lxkeys.world/schema/"
        }
    ],
    "@type": "Thing",
    "identifier": "ENT-00002671",
    "name": "AdvBench",
    "alternateName": [],
    "additionalType": {
        "category": "Data and Evaluation",
        "type": "Benchmark",
        "subtype": "",
        "lxkeysEntity": false
    },
    "description": "AdvBench is a AI benchmark associated with LLM safety research community, classified in LXKeys.world as AI Benchmark / Evaluation.",
    "creator": "LLM safety research community",
    "url": "https://lxkeys.world/entry.php?id=ENT-00002671&lang=fr",
    "sameAs": "https://github.com/llm-attacks/llm-attacks",
    "image": "",
    "lxkeysWorld": {
        "worldId": "ENT-00002671",
        "kind": "Data and Evaluation",
        "type": "Benchmark",
        "subtype": "",
        "classification": "AI Benchmark / Evaluation",
        "organization": "LLM safety research community",
        "originContext": "AI evaluation",
        "firstPublicAppearance": "2023",
        "currentStatus": "Active",
        "documentationStatus": "Documented",
        "documentationIndex": {
            "total": 82,
            "documentation": 25,
            "evidence": 13,
            "structure": 25,
            "relationships": 19,
            "level": "Level V Persistent"
        },
        "lxCalendarium": {
            "start_date_utc": "2023-04-01",
            "created_utc": "2026-06-17T01:48:33+00:00",
            "created_dypclt": "D-0 Y-2 P-3 C-3 L-22 T-4",
            "updated_utc": "2026-06-17T01:48:33+00:00",
            "updated_dypclt": "D-0 Y-2 P-3 C-3 L-22 T-4"
        },
        "facts": [],
        "capabilities": [
            "Evaluation protocol documentation",
            "Model comparison support",
            "Task taxonomy mapping",
            "Reproducible benchmark anchoring",
            "Relationship graph compatibility"
        ],
        "limitations": [
            "Benchmark results can become stale as models improve and evaluation protocols evolve.",
            "A benchmark measures a defined task scope rather than complete intelligence."
        ],
        "tags": [
            "Benchmark",
            "Evaluation",
            "AI Safety",
            "Agent Evaluation"
        ],
        "timeline": [
            {
                "date": "2023",
                "title": "Initial benchmark release",
                "description": "AdvBench entered the documented public record in 2023. This event is retained at the precision supported by the Entry’s reviewed source history.",
                "source_url": "https://github.com/llm-attacks/llm-attacks",
                "verification_status": "source_backed_curated_baseline"
            }
        ],
        "relationships": [
            {
                "target": "LLM safety research community",
                "type": "Associated organization",
                "description": "LLM safety research community is the organization, project community or institutional context associated with AdvBench.",
                "evidence_level": "documentary"
            },
            {
                "target": "AI evaluation",
                "type": "Domain context",
                "description": "AdvBench belongs to the AI evaluation layer of the intelligent-entity registry.",
                "evidence_level": "documentary"
            }
        ],
        "sources": [
            {
                "label": "Official documentation",
                "url": "https://github.com/llm-attacks/llm-attacks",
                "source_type": "Official / Research Source",
                "verification_status": "verified",
                "authority": "TRUSTED_PRIMARY",
                "role": "code_repository",
                "update_enabled": false,
                "authority_basis": "curated-corpus-refresh-2026-09-08"
            }
        ],
        "canonical": [],
        "imageMeta": []
    }
}
Preuves

Contribuer à cette entrée

Soumettez une source, une correction ou un commentaire. Les modifications publiques restent modérées.

Soumettre une preuve ou un commentaire

Commentaires approuvés

Aucun commentaire public approuvé pour le moment.