top of page

Which Business Task Should You Automate First?

Writer: Branden Bell
Branden Bell
Sep 11
6 min read

Updated: Sep 12

Three wooden task blocks, with the center checkmark highlighted, beneath Start with one useful task.

The owner wants fewer interruptions. Operations wants cleaner intake. Sales wants faster follow-up. Everyone has a reasonable AI idea, and somehow the first project now needs access to six systems.

Before choosing a tool, choose a piece of work you can finish improving.

To prioritize AI automation opportunities, compare a few specific tasks on recurring effort, input quality, ease of checking the output, and the consequences of a mistake. Choose a useful task with an owner and a measurable finish line. A high potential saving should not erase an unresolved risk.

Here is a practical selection exercise you can do with the people doing the work. The framework and examples below are proposed planning tools, not a validated scoring model or a report of client results.

Describe three tasks small enough to recognize

“Automate operations” is too broad to compare with anything. “Draft a scheduling summary from an approved job record” is a task. You know where it starts, what it produces, and who would receive it.

Write three candidates in the same form: when this input arrives, this person produces this output for this next person.

For example: when a field visit ends, the coordinator reads the technician's notes and prepares an internal follow-up list for the service manager. Keep sending customer messages outside that first candidate. That is another action with its own checks.

If you cannot identify the owner or the required input, start with the process audit and handoff guide. Ranking a poorly understood task gives you a tidy list of guesses.

Use four questions before assigning a score

An elaborate spreadsheet can make a weak decision look settled. Start by asking whether each candidate is ready for an experiment at all.

  • Can we use the input? Identify the source, required access, and approved environment for the information. A file being technically reachable does not settle whether it belongs in the experiment.

  • Can someone check the output? Name the reviewer and show the evidence they will compare against. “A person looks at it” is incomplete if the person cannot detect the likely error.

  • Can we contain a failure? Describe what happens if the output is wrong, missing, late, or duplicated. An internal draft can be held for correction. A customer commitment may already have consequences.

  • Does someone own the test? That person needs time to review examples, record problems, and decide whether to continue. Adding this work to an already overloaded employee without changing anything else is part of the cost.

Put unresolved candidates in a “prepare first” group with the missing answer beside them. Don't quietly turn “unknown” into a low-risk score.

The NIST AI Risk Management Framework is a voluntary resource for incorporating trustworthiness into AI design, development, use and evaluation. This short exercise is not a NIST assessment or certification. It is a way to make the risk discussion concrete before you commission a build.

Compare the candidates that are ready

Use the same four columns for each remaining task. Keep the supporting notes beside the rating so someone else can challenge it.

  • Frequency and effort: How often does it happen, and how much total staff time does one instance take? Use a defined sample. Separate time spent working from time spent waiting.

  • Repeatability: Are the inputs and desired output reasonably consistent? Record the exceptions instead of averaging them away.

  • Review burden: What must a person check, and how long might that take? Treat this as an estimate until tested.

  • Failure exposure: What can go wrong before someone catches it? Include the cleanup effort and the effect on the next person.

For a first pass, plain labels such as frequent, occasional, easy to check, and difficult to check can be more honest than invented precision. If you use numbers, define them together. A score of four means little if one person means “daily” and another means “expensive.”

I would choose the first pilot from the tasks with useful recurring effort, reasonably stable inputs, and errors a reviewer can find before the output is used. That is a judgment rule for this exercise, not a universal formula.

Three ideas from the same fictional business

Imagine a maintenance company choosing among these proposals. All volumes, times and conditions in this example are invented.

Candidate A: draft an internal follow-up list from technician notes. It happens 40 times a week and currently takes six minutes per visit. Notes vary in wording, but a coordinator can compare every proposed action with the original note. The first pilot would produce a draft only. Missing owners and unclear instructions would remain questions.

Candidate B: move a completed task to the next queue. It happens just as often. The rule is already agreed: when the required fields are complete and a manager marks the task approved, route it to scheduling. Evaluate an ordinary rule-based automation here. There may be no language interpretation for AI to contribute.

Candidate C: quote and promise an appointment from an incoming email. It could remove substantial work, but the company has not agreed how to handle unusual scope or unavailable capacity. A convincing response could promise something the business cannot deliver. Prepare those decisions before considering automated sending. A narrower draft-preparation task could be evaluated separately.

For these stated conditions, I would test Candidate A as the first AI pilot, investigate Candidate B as a simpler automation, and hold Candidate C's external actions. Change the conditions and the order may change. If the notes are unreadable or the coordinator cannot check them, A is no longer ready.

Anthropic's engineering guidance on effective agents recommends starting with the simplest workable solution and adding complexity when needed. That supports considering a basic rule or a limited AI call before a system that directs work across multiple tools. It does not establish that our fictional pilot will succeed.

Check whether the small task is worth doing

Candidate A's current effort is 40 visits multiplied by six minutes: 240 minutes per week.

Suppose the proposed workflow takes two minutes of preparation, review and correction per visit, plus 30 minutes of weekly maintenance. That would total 110 minutes, leaving a hypothetical 130-minute weekly reduction. These are planning assumptions, not measured savings. Setup time and software costs are still excluded.

Now change review time to five minutes per visit. The weekly effort becomes 230 minutes, leaving only ten minutes before setup and software costs. A fast draft does not necessarily produce a worthwhile workflow.

Use this sensitivity check to identify the assumption worth testing first. Here, it is review effort. Measure time for rejected drafts as well as accepted ones, and include the work caused downstream by a missed action.

Hours released are not automatically money saved. State what the team would do with the capacity, and keep any cash-cost claim tied to costs that would really change.

Choose the tool after choosing the pilot

For the internal follow-up draft, Claude or ChatGPT could be candidates to evaluate with approved sample notes. The Claude guide provides context if that is one of your options. Verify the current product features, access and data settings before implementation.

Keep accepted work in the system the team uses. If that is ClickUp, the workspace cleanup guide can help you examine the ownership and structure around the task. A new AI output needs a clear place to go.

Codex may be relevant if a tested workflow calls for custom code. You do not need to decide that during the first selection meeting. First establish whether the proposed output is useful and practical to check.

Common questions about choosing an AI project

Should I automate the task that takes the most time?

Consider it, but examine why it takes time. Waiting for a decision, reconstructing missing information, and writing a predictable summary are different problems. A smaller task with usable inputs and a clear review step may be easier to evaluate first.

Do I need a numerical automation scorecard?

No. Consistent questions and evidence are enough to start comparing a few candidates. Numbers become useful when their definitions are shared. Keep unresolved access, ownership or failure questions visible rather than burying them in an average.

How do I know whether the task needs AI?

Describe the transformation. If an agreed rule can determine the result from structured fields, evaluate ordinary automation. If the task involves interpreting variable language, extraction or drafting, an AI experiment may be worth testing. Neither choice removes the need to handle failures.

How long should the first pilot run?

Choose a period that captures ordinary work and meaningful exceptions, with an agreed review date. Calendar length alone does not establish enough evidence. Record the number and kinds of cases, the review effort, and the reasons any outputs were rejected.

Can you help choose and implement the first improvement?

Yes. I review how the process works today, then implement an appropriate improvement using Claude, ClickUp, ChatGPT/Codex or a simpler approach where it fits. Tell me about the recurring task you want to improve. Bring the current process and the people involved, even if the shortlist is still messy.

Leave the selection meeting with one candidate, one owner, one measurable output and the unanswered questions. That is enough to begin a useful test.

Continue the series: When a Checklist Beats an AI Agent. Compare a checklist, fixed rules and AI before choosing the implementation.

Comments


bottom of page