AI Agents: What They Can Do, and Where People Still Matter

A practical look at the systems that plan tasks, use tools and act on our behalf, and the decisions that still need people.

A software engineer working with a laptop and external monitor in a daylight office

From an answer to an action

A conversation with an AI assistant usually ends with an answer: a paragraph, a summary or a suggestion. An AI agent can carry the task further by using tools, examining the result and choosing what to do next. That extra step changes the question from “Did it say something useful?” to “Did it complete the work correctly?” It also makes the quality of the surrounding software as important as the language model inside it.

Consider a hypothetical team preparing a weekly research brief. A conversational assistant might summarize the documents pasted into a chat. A tool-connected agent could retrieve approved files, extract relevant changes, compare them with the previous brief and prepare a draft for review. The value comes from connecting these steps. A polished final paragraph is less useful if the system retrieved the wrong document or silently skipped an important update.

The distinction is described clearly in Anthropic's guide to building effective agents: workflows follow predefined paths, while agents can direct their own process and tool use. In practice, a product may combine both approaches. A predictable routine can sit inside a larger task that requires judgment about the next step.

What makes a good first task?

The best starting point is a task with a visible finish line. “Improve our operations” is difficult to evaluate. “Prepare a draft summary of these five approved reports, with a source link for each finding” is much easier. A narrow objective gives the system less room to wander and gives the person reviewing it a clear way to decide whether the work is complete.

Useful candidates often involve repetitive gathering and comparison rather than an unrestricted final decision. Think of organizing incoming requests, assembling a draft project update, finding duplicate records or checking whether a proposed code change passes an existing test suite. These examples still need suitable tools and reliable input data. The label “agent” does not make a messy business process clear.

A sensible pilot also includes an escape route. If the required file is missing, the system should report that condition. If two records disagree, it should preserve the disagreement rather than invent a compromise. Successful automation includes the ability to stop at the right point, particularly when a task cannot be completed with the available evidence.

Laptop and monitor at a developer workstation with a notebook beside the keyboard
The tools around a model determine which parts of a task it can actually complete.

Tools are the working interface

A model cannot inspect a private database or update a calendar simply because someone asks it to. The application must provide a way to perform that action. The tool might search a document collection, retrieve one customer record, run a calculation or save a draft. Each tool needs a clear description of what it accepts, what it returns and what happens when the request fails.

Anthropic's discussion of effective agent tools emphasizes designing tools around the work the agent needs to perform and evaluating them with realistic tasks. For a product team, that suggests a practical discipline: inspect the tools as carefully as the prompts. A confusing search response can undermine an otherwise capable model.

Imagine a search tool returning twenty files with almost identical titles and no dates. A human might know which one is current; the agent may not. Adding meaningful metadata can make the task easier to resolve. Similarly, a tool that returns a clear “record not found” result is more useful than one that fails with an opaque error. Small interface details can decide whether the next step is sensible.

Keep people at the consequential decision

Giving an agent more access makes it more capable, but it also increases the importance of review. Preparing an email draft and sending an email are different operations. Comparing invoices and approving payment are different operations. A product should make those distinctions visible instead of hiding them behind a single button that promises to handle everything.

In a hypothetical customer-support workflow, an agent might collect the order history and draft a response. A member of the support team could approve an unusual refund or a change to an account. The division of responsibility should follow the real task: familiar, reversible steps may require less supervision than actions affecting money, access or another person's commitments.

Human review works best when the reviewer can see the evidence. A draft with source references, a list of proposed changes and a concise account of unresolved questions is easier to assess than a confident “done.” People should be able to correct a specific step without reconstructing the entire task from scratch.

Two colleagues comparing a paper document with information on a laptop
A useful review shows the proposed result and the evidence behind it.

Measure the completed work

A fluent demonstration can make an agent look more reliable than it is. Evaluation should ask whether it achieved the intended outcome, whether it used the right inputs and whether it stayed within the scope of the task. For a document brief, reviewers might check coverage and references. For a coding task, they might examine behavior and test results. The right measure depends on the job.

Test difficult cases as well as tidy ones. Include a missing attachment, a contradictory record, an ambiguous instruction and a tool that is temporarily unavailable. These cases reveal whether the system can recover, ask for missing information or report a genuine blocker. They are also closer to everyday work than a carefully staged example.

The NIST AI Risk Management Framework offers a broader reference for organizations thinking about how to govern, map, measure and manage AI-related risks. Applying that thinking to a small pilot does not require turning every task into a bureaucratic exercise. It means deciding in advance what success looks like and who is responsible when the result is uncertain.

A practical way forward

For a team trying agent-based software, start with one workflow that people already understand. Write down the inputs, the permitted actions, the expected output and the cases that require review. Run the system alongside the existing process long enough to observe its strengths and recurring mistakes. Then adjust the tools and instructions before expanding its responsibilities.

The most useful agent is not necessarily the one that takes the largest number of steps. It is the one that reliably finishes a worthwhile task, makes its work easy to inspect and knows when the next decision belongs to a person. That is a more durable basis for adoption than an impressive conversation alone.

Continue reading

Multimodal AI: How Text, Images and Audio Work Together

All news