Every software vendor now has an agent. Every conference has a keynote about autonomous workflows. And most of the operations teams we talk to are somewhere between curious and exhausted. From the work we've done connecting AI to real business systems, rather than to slide decks, we've got a fairly clear view of where agents deliver and where they quietly create more work than they save.
What we actually mean by an agent
The word is being stretched to cover almost anything with a language model in it, so it's worth being precise. A chatbot answers questions. An agent takes actions: it reads from your systems, decides what to do next, calls tools, and writes something back. Updating a record, sending an email, raising a ticket, passing a lead to the right person.
That difference matters more than anything else in this article. A chatbot that gets something wrong gives a bad answer. An agent that gets something wrong changes your data. Almost everything that follows is really about managing that one fact.
Where agents are earning their keep
The successes have a lot in common. They're high-volume, fairly repetitive jobs where the input is messy but the right outcome is clear, and where a mistake is cheap to spot and cheap to fix.
Triage is the standout. Reading incoming enquiries, working out what they're about, pulling the relevant account details and routing them to the right person or territory. A job that used to sit in a shared inbox for half a day now happens in seconds, and a human still handles the actual conversation.
Data hygiene is another. Agents are good at finding duplicate records, normalising addresses, filling gaps from sources you trust and flagging anything that looks off. They're also increasingly useful as glue between systems that were never designed to talk to each other, handling the awkward mapping work that used to need a developer every time a supplier changed a file format.
Where to start, and where to hold back
Good first jobs
- Triaging and routing inbound enquiries
- Cleaning, deduplicating and enriching records
- Drafting replies, summaries and reports for review
- Moving data between systems with clear rules
Keep a human on
- Anything involving money, contracts or refunds
- Messages that reach customers without review
- Decisions that are hard to undo
- Situations with no clear right answer
Where they still fall over
Agents struggle when the task is ambiguous and the stakes are real. They follow instructions with great confidence, including instructions that made perfect sense to whoever wrote them and to nobody else. Give an agent a vague goal and broad permissions and it will find a creative interpretation of both.
They also inherit every problem in the systems they connect to. If your CRM has three records for the same customer, the agent will pick one. If your product data contradicts itself, the agent will repeat whichever version it found first. Agents don't fix messy foundations. They just move faster on top of them.
And there's a class of risk that didn't exist with traditional software. An agent that reads emails, web pages or documents can be influenced by what's in them. Instructions hidden in an incoming message can try to talk an agent into doing something it shouldn't. That isn't a reason to avoid agents, but it is a very good reason to think carefully about what each one is allowed to touch.
The guardrails that make the difference
The agent projects that work tend to be less about the model and more about everything around it. The same few principles come up again and again:
- Least privilege. Give each agent only the permissions its job needs. A triage agent has no business deleting records.
- Humans approve the irreversible. Let the agent prepare the action and a person confirm it, at least until there's a track record to trust.
- Log everything. Every decision, tool call and change should be traceable, so when something goes wrong you can see exactly why.
- Test against real cases. Keep a set of past examples with known right answers and run the agent against them before every change.
- Start narrow. One well-defined job done reliably beats a general-purpose assistant that's right most of the time.
The integration layer is the real project
Here's the part that surprises people. Choosing the model is usually the easy bit, and models are becoming more interchangeable every few months. The hard work, and most of the value, is in connecting the agent to your systems cleanly: well-defined tools, sensible permissions, reliable data and a clear record of what happened. Standards like the Model Context Protocol make the plumbing easier, but they don't decide what an agent should be allowed to do. You still have to.
That's good news for businesses that have already invested in joined-up systems. If your data lives in one place, with a clear structure and proper access controls, adding agents on top is a relatively small step. If it's spread across a dozen tools held together by spreadsheets, the agent project quietly becomes an integration project first.
Our advice is the same as it would be for any new technology: pick one painful, well-understood job, measure the results honestly, and expand from there. The businesses getting real value from agents this year aren't the ones with the boldest AI strategy. They're the ones who chose a sensible first job and did it properly.