Product Thinking

What Should Your Company Automate First With AI Agents?

Bill Cava/

Every operator has a Monday report. Someone opens four systems, copies the numbers into a spreadsheet, reconciles the two that never agree, and sends it by ten. The business runs. The team is the glue that makes it run, and the glue is the most expensive material in the building.

That report is where most companies should start with AI agents, and almost none do. An AI agent, for an operator, is software that takes a goal, works inside the tools you already use, and carries a task through to a result instead of waiting to be clicked.

Which workflow to hand it first matters more than what it can do.

What should you automate first with AI agents?

The work your team repeats every week, that needs little judgment, and that you can measure. Intake, data entry between systems, enrichment, scheduling, status reporting, follow-ups. High volume, low stakes, easy to count. One of those automated well builds the pattern for everything after it.

Those three properties do most of the selection for you. Weekly repetition means the agent gets enough runs to earn trust quickly. Low judgment means a wrong answer is cheap and obvious. Measurable means you can say, in hours or in errors, whether it worked.

The list of candidates in a small company is usually short and usually obvious to the people doing the work:

  • Intake. The form, the inbox, the voicemail that someone turns into a record by hand.
  • Data entry between systems. The copy-paste from the CRM to the spreadsheet to the invoicing tool.
  • Enrichment. The ten minutes of looking things up before a call or a quote.
  • Scheduling and follow-ups. The reminders that live in one person's calendar and conscience.
  • Status reporting. The Monday report, and its Friday cousin.

The question to put to each one is the one Shopify's CEO put to his whole company.

What would this area look like if autonomous AI agents were already part of the team?

Tobi Lütke, memo to Shopify staff, April 2025

He wrote that as a hiring rule (teams have to show why AI cannot do the work before asking for headcount), and it drew a lot of heat as one.[1] Read as an operator's question, one workflow at a time, it is simply the right place to start.

Where does the process actually live?

That is the one question we ask of every candidate, and it separates the workflows worth automating from the ones worth protecting. If the answer is one person's head or one spreadsheet tab, automate it. If the answer is judgment your customers pay you for, keep it human and use agents to feed it better inputs.

The test works because it locates the value, not the effort. The Monday report takes four hours and lives in a tab: automate it.

The call where a good account manager talks a frustrated customer down takes twenty minutes and lives in that person's judgment: keep it, and give them the account history, the last three tickets, and the renewal date before they pick up.

Repeats weekly
Low judgment
Can measure
Where it lives
First?
Monday status report
yes
yes
yes
One spreadsheet tab
Start here
Data entry between systems
yes
yes
yes
One person's head
Start here
Enriching new leads
yes
yes
yes
An unwritten checklist
Start here
Scheduling and follow-ups
yes
yes
yes
The inbox
Start here
Answering an upset customer
yes
no
no
Customer judgment
Keep human
Quoting a non-standard job
no
no
no
Customer judgment
Keep human
A process about to change
yes
yes
yes
Next quarter's plan
Not yet
Every workflow that starts here lives in a tab, a head, or an inbox.

The rule is not "never automate judgment." It is that the first automation should not be the one where the judgment is the product. Agents earn their way into the harder work by running the easy work reliably, in front of the people who will have to trust them with more.

Why does the first one matter more than which one?

Because the first automation builds the pattern, and the pattern is what every later one inherits: how agents get access to your systems, where approvals sit, what the record of their actions looks like, and who checks it. Get that right on a small workflow and the second one is mostly configuration.

It also builds the team's trust, which is the scarcer resource. A first automation small enough to ship in weeks and visible enough that everyone sees the hours come back turns the sceptics into people with requests. An ambitious first automation that stumbles in public poisons the well for a year.

So the bar for the first one is low on purpose. It does not have to be the workflow that costs you the most. It has to be the one where success is unambiguous.

How do you keep control once agents are doing the work?

By building the control layer into the workflow before you need it. Approval gates at the points where the call has to stay yours. Escalation thresholds that route the unusual case to a person. A visible record of what the agent actually did, that someone can read after the fact. Control is architecture, not attention.

We have written before about why human judgment holds as architecture and fails as vigilance: a person asked to watch every action stops watching, and a gate that will not pass unproven work does not get tired.

The same holds one floor down from security: automation that depends on someone watching it is a second job.

The gates are also where the operator's expertise goes. Deciding which cases escalate, what counts as normal, and where a human signature is non-negotiable is the domain knowledge that an agent's guardrails ultimately come down to. Nobody else can write those rules for you, and the agent runs better for having them.

What should you not automate first?

Three things, each tempting because it is where the pain is loudest. The customer-facing process, which carries the highest stakes and earns trust the slowest. Any workflow your team does not already agree on, because automating it scales the disagreement. And anything you are about to change anyway, because you would be hard-coding a process with an expiry date.

The customer-facing process

Everything the customer sees is judgment the customer is paying for, even when it looks like volume. A wrong internal report costs you an afternoon. A wrong answer to a customer costs you the customer, and the ones after that.

The far end of this spectrum arrived this week. Andon Labs, the research lab that had frontier models run a vending machine, shipped a product on September 14 that "lets people hand a business over to persistent agents" with email, phone, banking and a browser.[2]

Its own results draw the line this post is drawing. The vending machine, a business that repeats, needs little judgment and is trivially measurable, was profitable by late 2025. The lab's two real businesses, a retail store and a cafe, are not.

Why we built Pion
14 September 2026

Neither is profitable today, but we've seen significant qualitative improvements as better models have been released.

Source: andonlabs.com. Reproduced verbatim; punctuation is the source's.

A cafe is customer-facing, judgment-heavy, and earns trust one cup at a time. The lab's own line for the store and cafe is that the models "initially struggled and lost a lot of money." That is not an argument against agents. It is the selection rule, confirmed by the people most motivated to disprove it.

One reader in the launch thread put the operator's worry plainly:[3]

It will be death by a thousand bad impressions, mistakes and oversights. Yeah, you can automate everything. That doesn't mean it's being done well.

The workflow nobody agrees on

If three people run the same process three ways, there is no process yet, only three habits. An agent will faithfully automate one of them, and the other two people will spend the next quarter proving it wrong. Settle the process first, in a document a person could follow. Then hand it over.

The process you are about to change

New pricing next quarter, a new system going live in the spring, a team about to be reorganised: the workflow around any of those has a known expiry date. Automating it now means paying twice. Wait for the change to land, then automate what is left.

Has AI taken any real work off your team's plate?

That is the question on our 404 page, and it is the rung most companies skip. The ladder runs from talking about AI, to AI doing real work inside the operation, to AI in the product you sell. Most try to jump from the first rung to the third, and a lot of them land in the pilot graveyard.

The Generative Labs 404 page. Under the heading Better questions to be asking right now, five questions are listed: Has your company started with AI, or just talked about it? Has AI taken any real work off your team's plate? Does your AI roadmap exist? Does the product you keep describing in meetings exist? Does the prototype that wowed everyone actually work? Below them: Wherever your answers stopped, that's where we start.
The second question is the rung. Most companies answer no and move on to the third.

The internal workflow is the cheapest place there is to learn how agents behave in your environment: with your data, your edge cases, your approval chains, and your people watching the first hundred runs.

It is also where the scarce input is judgment rather than typing, and where the operator who knows the business is the best-placed person to say what "done" looks like.

None of this asks you to become a software company.

The people who benefit most are the ones whose model already works and whose team is the bottleneck, which is the same reason AI is expanding who builds rather than replacing the people who know the work.

If you want to see what we would automate first in an operation like yours, that page is the short version of this one.

Start with the Monday report. Measure the four hours. Then pick the second one.

Work with us

Which workflow would you hand off first?

Bring that one. We map the workflows worth automating, build the first agent around a real one with your team in the loop, and measure the hours it returns.

Not a newsletter. Goes straight to the team.

References

Frequently asked

What should a company automate first with AI agents?
Pick a workflow that repeats weekly, needs little judgment, and can be measured.
Pick a workflow that repeats weekly, needs little judgment, and can be measured. Then apply one test: where does the process actually live? If the answer is one person's head or one spreadsheet tab, it is a candidate. If it lives in judgment your customers are paying for, leave it human and use agents to give that person better inputs.
What should you not automate first?
Three things. The customer-facing process, because it carries the highest stakes and earns trust the slowest.
Three things. The customer-facing process, because it carries the highest stakes and earns trust the slowest. Any workflow your team does not already agree on, because automating a process nobody agrees on just scales the disagreement. And anything you are about to change anyway, because you would be hard-coding a process with a known expiry date.
How do you keep control once agents are doing the work?
Design the control layer before you need it: approval gates at the points where the call has to stay yours, escalation thresholds that route the unusual case to a person, and a visible record of what the agent actually did.
Design the control layer before you need it: approval gates at the points where the call has to stay yours, escalation thresholds that route the unusual case to a person, and a visible record of what the agent actually did. Control is something you build into the workflow, not something you supply by watching.
Do you need a specific tool to automate a workflow with AI?
Most searching on this topic is tool shopping, and the tool is the last decision, not the first.
Most searching on this topic is tool shopping, and the tool is the last decision, not the first. The selection problem comes first: which workflow, and what stays human. Get that right and several tools will work. Get it wrong and no tool saves you.
How do you know if the first automation worked?
You measured it before you started. Pick something where you can state the before number (hours spent, turnaround time, error rate) and then check it after.
You measured it before you started. Pick something where you can state the before number (hours spent, turnaround time, error rate) and then check it after. If a workflow cannot be measured, it is the wrong one to start with, because you will not be able to prove the pattern works and earn the next one.
Work with us

Let’s build it together.

We turn clever prototypes into production systems people can rely on. If you’re building with agents and want a hand making it real, leave your email and we’ll be in touch.

Straight to the team. No spam.