What Should Your Company Automate First With AI Agents?
Every operator has a Monday report. Someone opens four systems, copies the numbers into a spreadsheet, reconciles the two that never agree, and sends it by ten. The business runs. The team is the glue that makes it run, and the glue is the most expensive material in the building.
That report is where most companies should start with AI agents, and almost none do. An AI agent, for an operator, is software that takes a goal, works inside the tools you already use, and carries a task through to a result instead of waiting to be clicked.
Which workflow to hand it first matters more than what it can do.
What should you automate first with AI agents?
The work your team repeats every week, that needs little judgment, and that you can measure. Intake, data entry between systems, enrichment, scheduling, status reporting, follow-ups. High volume, low stakes, easy to count. One of those automated well builds the pattern for everything after it.
Those three properties do most of the selection for you. Weekly repetition means the agent gets enough runs to earn trust quickly. Low judgment means a wrong answer is cheap and obvious. Measurable means you can say, in hours or in errors, whether it worked.
The list of candidates in a small company is usually short and usually obvious to the people doing the work:
- Intake. The form, the inbox, the voicemail that someone turns into a record by hand.
- Data entry between systems. The copy-paste from the CRM to the spreadsheet to the invoicing tool.
- Enrichment. The ten minutes of looking things up before a call or a quote.
- Scheduling and follow-ups. The reminders that live in one person's calendar and conscience.
- Status reporting. The Monday report, and its Friday cousin.
The question to put to each one is the one Shopify's CEO put to his whole company.
What would this area look like if autonomous AI agents were already part of the team?
He wrote that as a hiring rule (teams have to show why AI cannot do the work before asking for headcount), and it drew a lot of heat as one.[1] Read as an operator's question, one workflow at a time, it is simply the right place to start.
Where does the process actually live?
That is the one question we ask of every candidate, and it separates the workflows worth automating from the ones worth protecting. If the answer is one person's head or one spreadsheet tab, automate it. If the answer is judgment your customers pay you for, keep it human and use agents to feed it better inputs.
The test works because it locates the value, not the effort. The Monday report takes four hours and lives in a tab: automate it.
The call where a good account manager talks a frustrated customer down takes twenty minutes and lives in that person's judgment: keep it, and give them the account history, the last three tickets, and the renewal date before they pick up.
The rule is not "never automate judgment." It is that the first automation should not be the one where the judgment is the product. Agents earn their way into the harder work by running the easy work reliably, in front of the people who will have to trust them with more.
Why does the first one matter more than which one?
Because the first automation builds the pattern, and the pattern is what every later one inherits: how agents get access to your systems, where approvals sit, what the record of their actions looks like, and who checks it. Get that right on a small workflow and the second one is mostly configuration.
It also builds the team's trust, which is the scarcer resource. A first automation small enough to ship in weeks and visible enough that everyone sees the hours come back turns the sceptics into people with requests. An ambitious first automation that stumbles in public poisons the well for a year.
So the bar for the first one is low on purpose. It does not have to be the workflow that costs you the most. It has to be the one where success is unambiguous.
How do you keep control once agents are doing the work?
By building the control layer into the workflow before you need it. Approval gates at the points where the call has to stay yours. Escalation thresholds that route the unusual case to a person. A visible record of what the agent actually did, that someone can read after the fact. Control is architecture, not attention.
We have written before about why human judgment holds as architecture and fails as vigilance: a person asked to watch every action stops watching, and a gate that will not pass unproven work does not get tired.
The same holds one floor down from security: automation that depends on someone watching it is a second job.
The gates are also where the operator's expertise goes. Deciding which cases escalate, what counts as normal, and where a human signature is non-negotiable is the domain knowledge that an agent's guardrails ultimately come down to. Nobody else can write those rules for you, and the agent runs better for having them.
What should you not automate first?
Three things, each tempting because it is where the pain is loudest. The customer-facing process, which carries the highest stakes and earns trust the slowest. Any workflow your team does not already agree on, because automating it scales the disagreement. And anything you are about to change anyway, because you would be hard-coding a process with an expiry date.
The customer-facing process
Everything the customer sees is judgment the customer is paying for, even when it looks like volume. A wrong internal report costs you an afternoon. A wrong answer to a customer costs you the customer, and the ones after that.
The far end of this spectrum arrived this week. Andon Labs, the research lab that had frontier models run a vending machine, shipped a product on September 14 that "lets people hand a business over to persistent agents" with email, phone, banking and a browser.[2]
Its own results draw the line this post is drawing. The vending machine, a business that repeats, needs little judgment and is trivially measurable, was profitable by late 2025. The lab's two real businesses, a retail store and a cafe, are not.
Neither is profitable today, but we've seen significant qualitative improvements as better models have been released.
A cafe is customer-facing, judgment-heavy, and earns trust one cup at a time. The lab's own line for the store and cafe is that the models "initially struggled and lost a lot of money." That is not an argument against agents. It is the selection rule, confirmed by the people most motivated to disprove it.
One reader in the launch thread put the operator's worry plainly:[3]
It will be death by a thousand bad impressions, mistakes and oversights. Yeah, you can automate everything. That doesn't mean it's being done well.
The workflow nobody agrees on
If three people run the same process three ways, there is no process yet, only three habits. An agent will faithfully automate one of them, and the other two people will spend the next quarter proving it wrong. Settle the process first, in a document a person could follow. Then hand it over.
The process you are about to change
New pricing next quarter, a new system going live in the spring, a team about to be reorganised: the workflow around any of those has a known expiry date. Automating it now means paying twice. Wait for the change to land, then automate what is left.
Has AI taken any real work off your team's plate?
That is the question on our 404 page, and it is the rung most companies skip. The ladder runs from talking about AI, to AI doing real work inside the operation, to AI in the product you sell. Most try to jump from the first rung to the third, and a lot of them land in the pilot graveyard.

The internal workflow is the cheapest place there is to learn how agents behave in your environment: with your data, your edge cases, your approval chains, and your people watching the first hundred runs.
It is also where the scarce input is judgment rather than typing, and where the operator who knows the business is the best-placed person to say what "done" looks like.
None of this asks you to become a software company.
The people who benefit most are the ones whose model already works and whose team is the bottleneck, which is the same reason AI is expanding who builds rather than replacing the people who know the work.
If you want to see what we would automate first in an operation like yours, that page is the short version of this one.
Start with the Monday report. Measure the four hours. Then pick the second one.
Which workflow would you hand off first?
Bring that one. We map the workflows worth automating, build the first agent around a real one with your team in the loop, and measure the hours it returns.
Not a newsletter. Goes straight to the team.
References
- ^1.Tobi Lütke, Shopify, “Reflexive AI usage is now a baseline expectation at Shopify (memo to staff, posted on X)” (April 7, 2025)
- ^
- ^
Frequently asked
What should a company automate first with AI agents?›Pick a workflow that repeats weekly, needs little judgment, and can be measured.
What should you not automate first?›Three things. The customer-facing process, because it carries the highest stakes and earns trust the slowest.
How do you keep control once agents are doing the work?›Design the control layer before you need it: approval gates at the points where the call has to stay yours, escalation thresholds that route the unusual case to a person, and a visible record of what the agent actually did.
Do you need a specific tool to automate a workflow with AI?›Most searching on this topic is tool shopping, and the tool is the last decision, not the first.
How do you know if the first automation worked?›You measured it before you started. Pick something where you can state the before number (hours spent, turnaround time, error rate) and then check it after.
Let’s build it together.
We turn clever prototypes into production systems people can rely on. If you’re building with agents and want a hand making it real, leave your email and we’ll be in touch.