Skip to main content
AgentNet Observer
Esc
    ← Back to Library

    AgentBench: Evaluating LLMs as Agents

    Xiao Liu, Botian Yu, Haotian Du, Zhiheng Zheng, et al.

    Dataset 2023 Published 2023-08-07 ICLR 2024 arXiv:2308.03688 Indexed 2026-09-06 benchmarkevaluation

    Abstract

    A systematic multi-dimensional benchmark evaluating LLMs as agents across 8 distinct environments (OS, DB, knowledge graph, digital card game, etc.).

    Cite this entry

    GB/T 7714-2015

    Xiao Liu, Botian Yu, Haotian Du, et al. AgentBench: Evaluating LLMs as Agents[EB/OL]. ICLR 2024, 2023(2023-08-07)[2026-09-16]. https://arxiv.org/abs/2308.03688.

    BibTeX

    @misc{liu2023,
      author = {Xiao Liu and Botian Yu and Haotian Du and Zhiheng Zheng and et al.},
      title = {AgentBench: Evaluating LLMs as Agents},
      year = {2023},
      organization = {ICLR 2024},
      howpublished = {\url{https://arxiv.org/abs/2308.03688}},
    }

    Visit source ↗

    This entry is part of the AgentNet Observer library. Attribute with a link to this page when quoting.