← Back to Library
AgentBench: Evaluating LLMs as Agents
Xiao Liu, Botian Yu, Haotian Du, Zhiheng Zheng, et al.
Dataset 2023 Published 2023-08-07 ICLR 2024 arXiv:2308.03688 Indexed 2026-09-06 benchmarkevaluation
Abstract
A systematic multi-dimensional benchmark evaluating LLMs as agents across 8 distinct environments (OS, DB, knowledge graph, digital card game, etc.).
Cite this entry
GB/T 7714-2015
Xiao Liu, Botian Yu, Haotian Du, et al. AgentBench: Evaluating LLMs as Agents[EB/OL]. ICLR 2024, 2023(2023-08-07)[2026-09-16]. https://arxiv.org/abs/2308.03688.
BibTeX
@misc{liu2023,
author = {Xiao Liu and Botian Yu and Haotian Du and Zhiheng Zheng and et al.},
title = {AgentBench: Evaluating LLMs as Agents},
year = {2023},
organization = {ICLR 2024},
howpublished = {\url{https://arxiv.org/abs/2308.03688}},
} This entry is part of the AgentNet Observer library. Attribute with a link to this page when quoting.