Before you act on a number an AI gave you, make it show its work

Most of what’s written about AI for the trades right now is cheerleading. Prompt lists, adoption surveys, forty ways to save an hour a week. Almost none of it asks the question a skeptical owner actually has, which is the only one that matters when there’s money on the line: can I trust this number enough to act on it?

Here’s my position, and it’s the whole post in one line. Never act on a figure an AI gave you until it can show you the rows it came from. If a tool can’t tell you which records it counted, which definition it used, and which window it looked at, you don’t have an answer. You have a guess wearing a confident face.

The number that looks right and isn’t

Let me show you how a wrong number gets to look like a right one.

Say your AI assistant tells you average ticket is up 12% this month. That reads like good news, and it reads like something to act on: maybe the new pricing stuck, maybe the upsell training worked. So you lean in.

Now open the rows. Last month you ran 60 jobs, and 18 of them were $0 warranty callbacks. This month you ran 55 jobs, and only 4 were warranty visits. The AI counted every one of those zeros as a job. When the free visits fall out of the mix, the average climbs on its own. Strip the warranty callbacks out of both months and your real average ticket barely moved. The 12% was a warranty-mix artifact, not a pricing win.

Nothing about the number was flagged. It was computed correctly from the wrong set of rows, and it was delivered with total confidence. That’s the trap. The failure isn’t bad math; it’s an unstated definition. And you’d have chased the wrong story for a month.

The three questions for any AI number

Before you act on a figure any tool gives you, ask it three things. They work on Guidepost, they work on your vendor’s built-in AI, they work on a spreadsheet a bookkeeper sent you.

Which rows. What records went into this? All 60 jobs, or the 42 that were paid work? If the tool can’t list them, it can’t defend the number.

Which definition. Average ticket of what: invoiced revenue, collected revenue, booked value? Callback rate over how many jobs, counted how? A word like “revenue” hides at least four different numbers, and every honest tool has to pick one and name it.

Which window. Last calendar month, trailing 30 days, this month-to-date against the same span last month? A trend is only real if both ends of it measure the same thing.

A tool that can answer all three lets you check its work in a couple of minutes. A tool that can’t is asking you to take its word. On money, don’t.

Diagnostic is not prescriptive

There’s a second line, and it’s the one most AI features walk straight across.

Telling you what happened is diagnostic. Telling you what to do is prescriptive. The two need very different amounts of evidence, and most tools treat them as the same act because “here’s what to do” demos better than “here’s what I see.”

Some numbers earn a verdict. You’re owed $11,700 and it’s aging past 60 days: the data supports one instruction, which is call those accounts this week. That’s prescriptive, and it’s safe because the figure traces to specific unpaid invoices. Other numbers only earn a look. Your newest tech has three callbacks in his first month: that’s a question, not a verdict. Three redos during a ramp could be equipment, could be the install checklist, could be nothing. A tool that turns that into “coach this tech” or “he’s your problem” is inventing certainty it doesn’t have, and it’s pointing you at a person on thin data.

So the honest behavior is to prescribe where the data is strong and refuse where it isn’t. “We won’t pretend to know which way to move the budget” is a real answer, said with the same confidence as a verdict. If your AI never says “not sure yet,” it isn’t being careful. It’s performing.

Why we built ours to refuse

I’ll get specific about our own tool, because this is an engineering choice we made on purpose and I’d rather you judge it than take it on faith.

Inside Guidepost, every finding carries an action class before it’s ever shown to you. A finding that only describes a pattern is marked diagnostic, and the code will not let it dress itself up as a recommendation. A finding gets to prescribe only when it clears a specific bar: the number traces to source rows, and the pattern is strong enough that one action clearly follows. That firewall lives in the product, not in a style guide someone can forget on a deadline. It’s the reason a quiet week in the digest actually says “nothing else needs you this week” instead of manufacturing a task so the product looks busy.

That’s a genuine tradeoff, and I’ll own it. A tool that always tells you what to do feels more decisive than one that sometimes says “look here, but I can’t call it yet.” We chose the second one. I would rather under-claim and keep your trust than hand you a confident recommendation that falls apart the first time you open the rows behind it.

A buyer’s checklist for AI in your FSM tool

Credit where it’s due first. Jobber’s Copilot and Housecall Pro’s Analyst AI are real, useful tools; if you run one platform and want a fast answer about that platform’s data, the AI already shipped inside it is a fine place to start. Fair is fair. The checklist below isn’t a trap set for them. It’s the same bar I hold our own product to.

When an AI feature makes a claim, test it against five questions you can actually check:

  • Can it show the source rows behind any number, in a click or two? If the figure is a black box, treat it as a hunch.
  • Does it name its definitions, or does it say “revenue” and “average” as if there were only one of each?
  • Does it ever say “not sure”? A tool that’s confident about everything is confident about nothing.
  • Does it stay in its lane on data it can’t see? A field-service AI reads its own system of record. It doesn’t natively see your QuickBooks books or your Google Ads spend, so be skeptical of any cross-tool claim it makes about margin or marketing ROI.
  • Does the claim survive a spreadsheet? Export the rows once and recompute the number by hand. If it matches, trust the tool more. If it doesn’t, you just saved yourself a bad decision.

Notice that none of these ask whether the AI is impressive. They ask whether it’s checkable. Impressive is easy. Checkable is the whole job.

What “show its work” looks like in practice

Here’s one line from our sample shop, unpacked all the way down so you can see what a defensible number actually looks like. Northside Comfort is our labeled sample: a fictional HVAC and plumbing shop, invented so I can show real behavior without borrowing a real customer’s books.

The digest says: you’re owed $11,700, and it’s aging fast. That number is not an estimate. It’s the sum of three open invoices in QuickBooks: Riverside Apartments at $6,200, Bella Vista Cafe at $3,400, and Greenpoint Dental at $2,100. The aging, 60 to 90-plus days, is QuickBooks’ own aging field, not something we modeled. And each of those three balances links to a completed job in Jobber, so you can confirm the work was actually done before you pick up the phone. That’s why the source line under the item reads “QuickBooks invoices · Jobber jobs.” Open it and you land on rows you can read yourself, in about two minutes.

That’s the standard. Not “trust me, it’s $11,700.” Here are the three invoices, here’s where the aging comes from, here’s the job behind each one. A number you can walk backward to its source is a number you can act on. A number you can’t is just someone’s confidence, and confidence is free.

Where to take this next

If you want to see the diagnostic-versus-prescriptive line drawn against the AI already built into your field-service software, I wrote it out plainly on the vendor AI comparison: what each tool can see, where it structurally stops, and why a cross-tool number needs a layer that reads more than one source.

And if you’d rather just watch me trace a number to its rows on a real screen, that’s the demo. Bring the figure your current tool handed you that you weren’t sure you could trust. We’ll open it together and find out. Book one at /demo/, or read the fine print on how the numbers are made in what Jobber and Housecall Pro reports miss.

Written by Guidepost

Guidepost reads the numbers a home-service shop already has across its tools, then sends the few that need attention, each traced back to its source. The whole job is telling a real signal from noise: the line between a number worth acting on and one that’s only worth a closer look. More about Guidepost →

See it watch your numbers

Guidepost reads your Jobber, Housecall Pro, and QuickBooks numbers and tells you what needs attention, in plain English. Want to see the output first? Look at a sample digest.