Autonomous AI Agents and Task Planning
DOI:
https://doi.org/10.5281/zenodo.19614868Keywords:
autonomous AI agents; task planning; LLM agents; ReAct; MetaGPT; tool use; multi-agent systems; AgentBench; SWE-bench; hierarchical planning; error accumulationAbstract
Autonomous AI agents -- systems that perceive their environment, plan sequences of actions, use tools, and execute multi-step tasks with minimal human intervention -- represent a qualitative shift from AI as a prediction engine to AI as an active participant in complex workflows. Powered by large language models with tool use capabilities, agents can browse the web, write and execute code, interact with APIs, manage files, and coordinate with other agents to complete tasks that previously required sustained human effort. This study systematically evaluates eight autonomous agent frameworks-- ReAct, Reflexion, AutoGPT, MetaGPT, CrewAI, LangGraph, AgentBench-GPT4, and a proposed Hierarchical Planning Agent (HPA) -- across five task domains: web research and information synthesis, software engineering assistance, data analysis and visualisation, document processing pipelines, and multi-agent scientific literature review. Evaluation benchmarks include AgentBench, SWE-bench Lite, WebArena, ALFWorld, and a purpose-built Scientific Review Benchmark (SRB). HPA achieves the highest task completion rate on AgentBench (68.4%) through its hierarchical decomposition of complex tasks into validated subtasks. MetaGPT leads on SWE-bench Lite (24.8% issue resolution). CrewAI achieves the best multi-agent coordination score (84.2%) on scientific literature review. All agents show significant error accumulation over long task horizons: task completion rates drop 18.4-34.2pp for tasks requiring > 20 sequential steps. A taxonomy of agent failure modes and a safety-by-design framework for production agent deployment is proposed.Downloads
Published
2026-08-19
Issue
Section
Articles
How to Cite
Autonomous AI Agents and Task Planning. (2026). Bio-QI Journal, 4(1), 1-8. https://doi.org/10.5281/zenodo.19614868

