跳到主内容
智联观察
Esc
    ← 返回资料库

    AgentBench:评估大模型作为智能体的能力

    Xiao Liu, Botian Yu, Haotian Du, Zhiheng Zheng, et al.

    数据集 2023 原文 2023-08-07 ICLR 2024 arXiv:2308.03688 收录 2026-09-06 benchmarkevaluation

    摘要

    系统性多维基准:在操作系统、数据库、知识图谱、数字卡牌游戏等 8 个环境中评估 LLM 作为智能体的实际能力。

    引用本文条目

    GB/T 7714-2015

    Xiao Liu, Botian Yu, Haotian Du, et al. AgentBench: Evaluating LLMs as Agents[EB/OL]. ICLR 2024, 2023(2023-08-07)[2026-09-16]. https://arxiv.org/abs/2308.03688.

    BibTeX

    @misc{liu2023,
      author = {Xiao Liu and Botian Yu and Haotian Du and Zhiheng Zheng and et al.},
      title = {AgentBench: Evaluating LLMs as Agents},
      year = {2023},
      organization = {ICLR 2024},
      howpublished = {\url{https://arxiv.org/abs/2308.03688}},
    }

    查看原文 ↗

    本条目为「智联观察」资料库收录,引用时请注明来源与本页链接。