← Back to Library
GAIA: A Benchmark for General AI Assistants
Grégoire Mialon, Clémentine Fourrier, Craig Swift, Thomas Wolf, Yann LeCun, Thomas Scialom
Dataset 2023 Published 2023-11-21 NeurIPS 2023 (Datasets & Benchmarks) arXiv:2311.12983 Indexed 2026-09-06 benchmarkgeneral-assistantevaluation
Abstract
A benchmark of 466 questions requiring reasoning, multimodal handling, web browsing and tool use — questions simple for humans yet hard for state-of-the-art AI.
Cite this entry
GB/T 7714-2015
Grégoire Mialon, Clémentine Fourrier, Craig Swift, et al. GAIA: A Benchmark for General AI Assistants[EB/OL]. NeurIPS 2023 (Datasets & Benchmarks), 2023(2023-11-21)[2026-09-16]. https://arxiv.org/abs/2311.12983.
BibTeX
@misc{mialon2023,
author = {Grégoire Mialon and Clémentine Fourrier and Craig Swift and Thomas Wolf and Yann LeCun and Thomas Scialom},
title = {GAIA: A Benchmark for General AI Assistants},
year = {2023},
organization = {NeurIPS 2023 (Datasets & Benchmarks)},
howpublished = {\url{https://arxiv.org/abs/2311.12983}},
} This entry is part of the AgentNet Observer library. Attribute with a link to this page when quoting.