Decoding the technologies of tomorrow, today.

Exploring the breakthrough innovations shaping our world. From AI infrastructure and robotics to biotech, quantum computing, and spatial tech.

ReviewAurora

How AI Agents Use Tools to Complete Multi-Step Tasks

A chatbot can answer a question in one response. An AI agent is built for a different kind of job: it can pursue a goal through several steps, use external tools, examine the results, and decide what to do next.

That distinction matters whenever a task depends on information the model does not already have or requires an action outside the conversation. An agent might search a database, retrieve a document, check a calendar, call an API, or update a permitted business system before returning a final answer.

The language model provides the decision-making interface, while tools give the system access to information and actions. The combination creates a feedback loop that allows an agent to adapt its behavior as a task unfolds.

What Makes an AI Agent Different From a Chatbot?

There is no single industry-wide definition of an AI agent. A useful distinction is between a fixed workflow and a system in which the model can dynamically determine what to do next.

In a fixed workflow, developers decide the sequence in advance. For example, an application might always retrieve a customer record, check an order database, and then generate a response.

An agent has more flexibility. It can determine which tool is relevant, use the result of one operation to decide whether another is necessary, and stop when it has enough information to complete the task. Anthropic describes this distinction as the difference between workflows, which follow predefined paths, and agents, which dynamically direct their own processes and tool use. 

That flexibility is useful, but it also makes the system harder to predict and test.

Tools Give Agents Capabilities Beyond Text Generation

A language model by itself primarily generates and processes information within its available context. Tools connect it to external systems.

Depending on the application, an agent might have access to a web search service, a company database, a calendar, a document repository, a calculator, a customer-support system, or an external API.

The model does not necessarily operate those systems directly. Instead, the application exposes specific capabilities as tools. The model can request a particular tool and provide the required inputs. The software then executes the operation and returns the result to the model.

That separation is important. A model might decide that it needs current inventory information, but the inventory system—not the model itself—is responsible for returning the actual inventory data.

NIST's research on agent tool use describes a broad range of possible capabilities, including search, databases, computer interaction, software extensions, code execution, and physical devices. It also distinguishes between read-only tools and tools capable of changing external state. 

2.jpg

The Agentic Loop: Act, Observe, and Decide Again

The easiest way to understand an agent is to follow what happens after the first tool call.

Imagine a user asks an AI assistant to research several products and recommend one based on current availability and the user's requirements.

A basic chatbot might explain how the user can perform that research manually. It has no way to independently obtain current inventory unless an external system is connected.

A tool-using agent can take a different path. It may first determine what information is needed, search for current products, examine the results, identify which options satisfy the user's requirements, and then perform another lookup to verify availability. If the first search does not provide enough information, that result can trigger another action.

The pattern is:

User goal → choose an action → use a tool → receive the result → evaluate the result → choose the next action → repeat → complete the task

The important step is the evaluation between tool calls. The agent is not merely following a long predetermined checklist. Information returned by one operation can change what happens next.

Current agent frameworks implement this kind of iterative tool-use loop. LangChain, for example, describes agents as systems that can make sequential or parallel tool calls, select tools based on previous results, handle tool errors, and continue until a final response or an execution limit is reached. 

This is what turns tool calling into agentic behavior.

Why Multi-Step Tasks Are Harder Than They Look

Every additional step creates another opportunity for an error.

Suppose an agent is asked to research a business topic and prepare a summary. It might need to identify the relevant question, find appropriate sources, retrieve information, compare conflicting claims, determine whether important information is missing, perform additional research, and finally organize the findings.

An error near the beginning can affect everything that follows.

The agent might search for the wrong subject, misunderstand a user's requirement, choose an inappropriate tool, misinterpret a tool result, or continue searching even after it has already found sufficient information.

There is also a practical cost. Multiple model calls and external operations can increase latency and computing expenses. Anthropic recommends using the simplest architecture that can reliably accomplish the task rather than introducing agentic complexity simply because it is available. 

For predictable tasks, a conventional workflow can therefore be preferable. Agents are most useful when the path to the answer genuinely depends on what the system discovers along the way.

More Autonomy Creates New Failure Modes

Tool access changes the nature of an AI failure.

If a chatbot produces an incorrect explanation, the immediate problem is usually the response itself. An agent with external permissions can potentially turn an incorrect interpretation into an incorrect action.

One engineering problem is a runaway tool loop. An agent might repeatedly search, retry an operation, or revisit the same decision without making meaningful progress. Production systems therefore need practical boundaries such as iteration limits, timeouts, and resource controls. Current agent runtimes commonly provide explicit stopping conditions for this reason. 

Security presents another challenge: prompt injection.

An agent may retrieve information from a webpage, document, email, or other source that contains instructions intended to manipulate the model. If the model treats those instructions as trusted commands, the external content can influence subsequent tool use.

OWASP identifies prompt injection as a major risk for LLM applications and notes that indirect injections can originate from external content processed by the model. The consequences can become more serious when the system has access to sensitive information or external functions. 

This connects directly to another OWASP concern, excessive agency. An AI system becomes more dangerous when it has unnecessary functions, excessive permissions, or too much autonomy. An unexpected model output, for example, becomes much more consequential if the model is allowed to make irreversible changes to another system. 

NIST similarly recommends evaluating tool capabilities according to their permissions, reliability, trust environment, reversibility, and potential impact. 

Designing Tools Is Part of Designing the Agent

A capable model cannot compensate indefinitely for poorly designed tools.

A tool should have a clear purpose and well-defined inputs and outputs. If a tool can modify external data, its permissions should be appropriate for the task rather than broader than necessary. Read-only access is fundamentally different from permission to change records, send messages, or perform other consequential actions.

Tool selection also matters. Giving an agent hundreds of poorly described tools can make it harder for the model to determine which capability is appropriate. Anthropic's engineering guidance emphasizes designing tools carefully and evaluating how well agents actually use them, rather than assuming that adding more tools automatically improves performance. 

For developers, a useful rule is simple: start with the smallest set of capabilities that can accomplish the task.

If an application only needs to retrieve one type of information, a direct API integration may be enough. If several predictable operations always happen in the same order, ordinary application code may be easier to maintain.

An agent framework becomes more attractive when the application needs dynamic tool selection, repeated tool calls, state management, retries, or more sophisticated orchestration. LangChain, for example, provides built-in agent functionality for iterative tool use and error handling. 

The choice should follow the application's actual complexity, not the popularity of a particular framework.

Human Approval Can Define the Boundary of Autonomy

An agent does not have to be either completely autonomous or completely manual.

A useful design can allow the agent to perform low-risk information gathering independently while requiring approval before consequential actions. For example, an agent might research options and prepare a recommendation without permission to make the final purchase.

This approach preserves much of the efficiency of automation while creating a checkpoint before an action with financial, operational, or reputational consequences.

NIST's tool-use framework specifically considers whether actions are read-only or write-capable and whether their effects are reversible or persistent. Those distinctions provide a practical basis for deciding where stronger controls are needed. 

3.jpg

The Practical Meaning of Tool-Using AI

AI agents are best understood as systems built from three cooperating parts: a model that can interpret a goal and select actions, tools that provide access to external capabilities, and software that controls the execution process.

The significant change is not simply that AI can call an API. Traditional software has been doing that for decades. What makes an agent different is that the model can use the result of one action to influence its next action.

That makes agents useful for tasks where the correct sequence cannot be completely determined beforehand. It also means that reliability depends on much more than the underlying model. Tool design, permissions, error handling, execution limits, security controls, monitoring, and human approval all matter.

For a simple task, a small conventional program may remain the better engineering choice. For a genuinely adaptive task, an agent can do something a fixed workflow cannot: change its next move based on what it discovers along the way.