With the rapid advancement of large language models (LLMs), agents capable of autonomous perception, decision-making, and action have emerged as a frontier paradigm in artificial intelligence. These entities are transitioning from academic research to complex real-world applications. However, the rapid iteration of agent capabilities poses severe challenges to evaluation methodologies—particularly in assessing their core competencies in data processing and evaluation. As of 2025, the field of agent data evaluation exhibits a dynamic yet fragmented landscape. Traditional static dataset-based evaluations are no longer sufficient to measure agent performance in open, dynamic environments. The research community is actively shifting toward more interactive and realistic benchmarking paradigms. Despite the emergence of innovative benchmarks such as ToolBench and MLAgentBench, there remains a widespread lack of unified evaluation standards, widely accepted metric systems, and mature methodologies. This paper systematically reviews the state of agent data evaluation in 2025, tracing the evolution from traditional metrics to emerging process-oriented ones. Building upon this, we delve into the methodology of dataset and benchmark design, with particular attention to key elements in experimental design, such as controlled experiments, sample size determination, and statistical analysis. Furthermore, we analyze the core challenges facing the field, including the “realism gap” between evaluation and real-world tasks, the scalability dilemma of automated evaluation, and the increasingly prominent issues of data privacy and security. Our findings indicate that although potential technologies such as differential privacy and federated learning exist, dedicated privacy-preserving frameworks for agent evaluation remain in their infancy. Finally, this report outlines future research directions, emphasizing the urgent need to establish unified evaluation frameworks, develop process-oriented evaluation metrics, and formulate standardized privacy and security auditing protocols—aiming to provide a scientific foundation for building more robust, trustworthy, and responsible agent systems.
KeywordsAgent EvaluationLarge Language ModelBenchmarksProcess-Oriented Evaluation
Huang, Q., Vora, J., Liang, P. and Leskovec, J. (2023) Benchmarking Large Language Models as AI Research Agents. arXiv: 2310.03302v1.
Wooldridge, M. and Jennings, N.R. (1995) Intelligent Agents: Theory and Practice. The Knowledge Engineering Review , 10, 115-152. https://doi.org/10.1017/s0269888900008122
Franklin, S. and Graesser, A. (1997) Is It an Agent, or Just a Program? A Taxonomy for Autonomous Agents. In: Müller, J.P., Wooldridge, M.J. and Jennings, N.R., Eds., Intelligent Agents III Agent Theories , Architectures , and Languages , Springer, 21-35. https://doi.org/10.1007/bfb0013570
Kim, G.J., Wilf, A., Morency, L.P. and Fried, D. (2025) From Reproduction to Replication: Evaluating Research Agents with Progressive Code Masking. https://github.com/j1mk1m/AutoExperiment
Qiu, R., Chen, S., Su, Y., Yen, P.Y. and Shen, H.W. (2025) Completing a Systematic Review in Hours Instead of Months with Interactive AI Agents. arXiv: 2504.14822. https://arxiv.org/pdf/2504.14822
Cheng, Y., Zhang, C., Zhang, Z., Meng, X., Hong, S., Li, W., Wang, Z., Wang, Z., Yin, F., Zhao, J. and He, X. (2024) Exploring Large Language Model Based Intelligent Agents: Definitions, Methods, and Prospects. arXiv: 2401.03428. https://arxiv.org/pdf/2401.03428
Wei, J., Wang, X., Schuurmans, D., Bosma, M., Ichter, B., Xia, F., Chi, E.H., Le, Q.V. and Zhou, D. (2022) Chain of Thought Prompting Elicits Reasoning in Large Language Models. Advances in Neural Information Processing Systems , 35, 24824-24836.
Durante, Z., Huang, Q., Wake, N., Gong, R., Park, J.S., Sarkar, B., Taori, R., Noda, Y., Terzopoulos, D., Choi, Y., Ikeuchi, K., Vo, H., Fei-Fei, L. and Gao, J. (2024) Agent AI: Surveying the Horizons of Multimodal Interaction. arXiv: 2401.03568. https://arxiv.org/pdf/2401.03568
Ignise, A. and Vahi, Y. (2024) Tracking Intelligence and Effectiveness of Agents. International Journal of Computer Science and Mobile Applications , 12, 41-48.
Rani, A., Grover, N., Deepa, N. and Prajitha, C. (2024) A Smart Agent-Based Approach for Privacy Preservation and Threat Mitigation to Enhance Security in the Internet of Medical Things. Journal of Autonomous Intelligence , 7, Article 1629. https://doi.org/10.32629/jai.v7i5.1629
Costa, M., Köpf, B., Kolluri, A., Paverd, A., Russinovich, M., Salem, A., Zanel-la-Béguelin, S., et al. (2025) Securing AI Agents with Information-Flow Control. arXiv: 2505.23643.
Mnih, V., Kavukcuoglu, K., Silver, D., Rusu, A.A., Veness, J., Bellemare, M.G., et al. (2015) Human-Level Control through Deep Reinforcement Learning. Nature , 518, 529-533. https://doi.org/10.1038/nature14236
Tweedale, J. and Ichalkaranje, N. (2005) Innovations in Intelligent Agents. Proceedings of the 9 th International Conference on Knowledge - Based Intelligent Information and Engineering Systems , Melbourne, 14-16 September 2005, 821-824. https://doi.org/10.1007/11552451_112
Rao, A.S. and Georgeff, M.P. (1995) BDI Agents: From Theory to Practice. First International Conference on Multiagent Systems , San Francisco, 12-14 June 1995, 312-319.
Bansod, P.B. (2025) Distinguishing Autonomous AI Agents from Collaborative Agentic Systems: A Comprehensive Framework for Understanding Modern Intelli-gent Architectures. arXiv: 2506.01438.
Deng, J., Dong, W., Socher, R., Li, L., Li, K. and Li, F.F. (2009) ImageNet: A Large-Scale Hierarchical Image Database. 2009 IEEE Conference on Computer Vision and Pattern Recognition , Miami, 20-25 June 2009, 248-255. https://doi.org/10.1109/cvpr.2009.5206848
Rajpurkar, P., Zhang, J., Lopyrev, K. and Liang, P. (2016) SQuAD: 100,000+ Questions for Machine Comprehension of Text. Proceedings of the 2016 Conference on Empirical Methods in Natural Language Processing , Austin, November 2016, 2383-2392. https://doi.org/10.18653/v1/d16-1264
Bellemare, M.G., Naddaf, Y., Veness, J. and Bowling, M. (2013) The Arcade Learning Environment: An Evaluation Platform for General Agents. Journal of Artificial Intelligence Research , 47, 253-279. https://doi.org/10.1613/jair.3912
Tambe, M., Johnson, W.L., Jones, R.M., Koss, F., Laird, J.E., Rosenbloom, P.S. and Schwamb, K. (1995) Intelligent Agents for Interactive Simulation Environments. AI Magazine , 16, 15.
Wang, X., Li, D., Zhao, Y. and Wang, H. (2024) MetaTool: Facilitating Large Language Models to Master Tools with Meta-Task Augmentation. arXiv: 2407.12871.
Lin, J., Zhao, H., Zhang, A., Wu, Y., Ping, H. and Chen, Q. (2023) AgentSims: An Open-Source Sandbox for Large Language Model Evaluation. arXiv: 2308.04026.
Winata, G.I., Hudi, F., Irawan, P.A., Anugraha, D., Putri, R.A., Wang, Y., Ngo, C.W., et al. (2024) WorldCuisines: A Massive-Scale Benchmark for Multilingual and Multicultural Visual Question Answering on Global Cuisines. arXiv: 2410.12705.
Du, M., Xu, B., Zhu, C., Wang, X. and Mao, Z. (2025) DeepResearch Bench: A Comprehensive Benchmark for Deep Research Agents. arXiv: 2506.11763.
Patel, D., Lin, S., Rayfield, J., Zhou, N., Vaculin, R., Martinez, N., Kalagnanam, J., et al. (2025) AssetOpsBench: Benchmarking AI Agents for Task Automation in In-dustrial Asset Operations and Maintenance. arXiv: 2506.03828.
Abaskohi, A., Ramesh, A.V., Nanisetty, S., Goel, C., Vazquez, D., Pal, C., Laradji, I.H., et al. (2025) AgentAda: Skill-Adaptive Data Analytics for Tailored Insight Discovery. arXiv: 2504.07421.
Hu, X., Zhao, Z., Wei, S., Chai, Z., Ma, Q., Wang, G., Wu, F., et al. (2024) InfiAgent-DABench: Evaluating Agents on Data Analysis Tasks. arXiv: 2401.05507.
Testini, I., Hernández-Orallo, J. and Pacchiardi, L. (2025) Measuring Data Science Automation: A Survey of Evaluation Tools for AI Assistants and Agents. arXiv: 2506.08800.
Yadav, D., Jain, R., Agrawal, H., Chattopadhyay, P., Singh, T., Jain, A., Batra, D., et al. (2019) EvalAI: Towards Better Evaluation Systems for AI Agents. arXiv: 1902.03570.
Elshan, E., Zierau, N., Engel, C., Janson, A. and Leimeister, J.M. (2022) Understanding the Design Elements Affecting User Acceptance of Intelligent Agents: Past, Present and Future. Information Systems Frontiers , 24, 699-730. https://doi.org/10.1007/s10796-021-10230-9
Papineni, K., Roukos, S., Ward, T. and Zhu, W. (2001) BLEU: A Method for Automatic Evaluation of Machine Translation. Proceedings of the 40 th Annual Meeting on Association for Computational Linguistics — ACL ’02, Philadelphia, 7-12 July 2002, 311-318. https://doi.org/10.3115/1073083.1073135
Abramson, J., Ahuja, A., Carnevale, F., Georgiev, P., Goldin, A., Hung, A., Yan, C., et al. (2022) Evaluating Multimodal Interactive Agents. arXiv: 2205.13274.
Hartmann, M. and Koller, A. (2024) A Survey on Complex Tasks for Goal-Directed Interactive Agents. arXiv: 2409.18538.
Moreno, R., Fernandez-Isabel, A., Diego, I.M.D., Moguerza, J.M., Lancho, C. and Teresa, M.C.S. (2022) Automatic Detection of Potential Customers by Opinion Mining and Intelligent Agents. Proceedings of the 17 th Conference on Computer Science and Intelligence Systems , Sofia, 4-7 September 2022, 93-101. https://doi.org/10.15439/2022f131
Chen, K., Ren, Y., Liu, Y., Hu, X., Tian, H., Xie, T., Mo, Z., et al. (2025) Xbench: Tracking Agents Productivity Scaling with Profession-Aligned Real-World Evaluations. arXiv: 2506.13651.
Alves, P.H., Correia, F., Frajhof, I., De Souza, C.S. and Lopes, H. (2023) Designing Intelligent Agents in Normative Systems toward Data Regulation Representation. IEEE Access , 11, 51590-51605. https://doi.org/10.1109/access.2023.3276294
Dwork, C. and Roth, A. (2013) The Algorithmic Foundations of Differential Privacy. Foundations and Trends® in Theoretical Computer Science , 9, 211-407. https://doi.org/10.1561/0400000042
Dwork, C. (2006) Differential Privacy. In: Bugliesi, M., Preneel, B., Sassone, V. and Wegener, I. Eds., International Colloquium on Automata , Languages , and Programming , Springer, 1-12.
Chen, B., Hawkins, C., Karabag, M.O., Neary, C., Hale, M. and Topcu, U. (2023) Differential Privacy in Cooperative Multiagent Planning. Uncertainty in Artificial Intelligence , Pittsburgh, 31 July-4 August 2023, 347-357.
McMahan, B., Moore, E., Ramage, D., Hampson, S. and Arcas, B.A. (2017) Communication-Efficient Learning of Deep Networks from Decentralized Data. Artificial Intelligence and Statistic s , Fort Lauderdale, 20-22 April 2017, 1273-1282.
Pan, J., Zhang, Y., Tomlin, N., Zhou, Y., Levine, S. and Suhr, A. (2024) Autonomous Evaluation and Refinement of Digital Agents. arXiv: 2404.06474.
Yang, Z., Bhatnagar, A., Qiu, Y., Miao, T., Tser Jern Kon, P., Xiao, Y., et al. (2025) Cloud Infrastructure Management in the Age of AI Agents. ACM SIGOPS Operating Systems Review , 59, 1-8. https://doi.org/10.1145/3759441.3759443
Chen, C., Zhang, Z., Khalilov, I., Guo, B., Gebreegziabher, S.A., Ye, Y., Li, T.J.J., et al. (2025) Toward a Human-Centered Evaluation Framework for Trustworthy LLM-Powered Gui Agents. arXiv: 2504.17934.
Sangaraju, V.R. (2025) A Framework for Secure Data Processing Using AI Agents in Business Intelligence Applications. Economic Sciences , 21, 914-924. https://doi.org/10.69889/srn8qw78