Every agency deck ends with the same slide: a return figure computed by the agency, from data held by the agency, presented at the meeting where the agency asks to renew. Whatever that ritual is, it is not measurement. Measurement is a design decision made before launch, and it belongs to you.
If the vendor controls the scoreboard, it is not a scoreboard
The test of any reported result is whether you could reproduce it without the vendor in the room. That means every number has to trace back to source data living in accounts you own: the call log in your carrier's portal, the export from your own CRM, the aging report in your accounting software. If a claimed result cannot be walked back to one of those, treat it as decoration. This is not an accusation of dishonesty — it is just what a scoreboard is. A referee who works for one team is a coach.
We build to this standard because it protects both sides. A vendor whose results survive an audit does not have to argue at renewal. The data argues, or it does not.
The baseline comes first and cannot be reconstructed later
Before a system launches, record the current state of the one number it targets. For a phone system, that is a month of carrier logs: how many calls ended without a connection, and at what hours. For scheduling, it is the percentage of bookable slots that actually filled last month, straight from the appointment book. For invoicing, it is the accounts-receivable aging report — how many days your invoices actually take to get paid, as opposed to the terms printed on them.
This has to happen before launch because afterward the before-state is gone. Memory will reconstruct it in whichever direction flatters the argument being made — the vendor's memory and yours alike. A baseline is ten minutes of pulling reports you already have access to, and it is the difference between knowing and being told.
One system, one number
A system that claims to improve everything can never be caught improving nothing. Each system gets one number, agreed in writing before it goes live.
Missed-call capture: recovered conversations, counted from transcripts
The metric is not texts sent, and it is not response rate. It is the count of threads where the caller replied and the exchange went somewhere — a booking, a quote request, a scheduled visit. Each thread carries a timestamp and a full transcript, which means any single claimed recovery can be audited in thirty seconds by reading it. The vanity version of this metric counts messages. The real version counts conversations that ended in work.
Scheduling and reminders: filled-slot rate
A dental chair produces revenue by the hour it is occupied, which is why the useful number is the percentage of bookable hours actually filled, with the no-show count alongside it. For a service business the same idea is jobs per crew-day. If a practice is spending on ads, the number that matters downstream is cost per acquired patient — not clicks, not impressions, not anything a marketing dashboard celebrates. Clicks do not sit in chairs.
Invoice follow-up: days outstanding
Construction and trades run on terms that erode quietly: net-30 becomes net-50, then net-70, one polite deferral at a time, and the general contractor's pay-when-paid clause hides in the middle of it. The metric here is average days from invoice to payment, read directly off the aging report. If automated reminders are working, that number falls over a quarter. If it does not fall, they are not working, and no amount of activity reporting changes that.
The CRM record is what makes any claim checkable
Underneath every metric above is the same piece of plumbing: each event writes a row somewhere permanent — source, timestamp, transcript, outcome. That record is the audit trail. When a monthly summary says fourteen recovered conversations, you pick three at random and read them. Either they are real conversations with real outcomes, or they are not, and you will know within five minutes. A system that does not write this record is asking to be trusted, which is a different thing from being measured. It is also the reason the CRM write is not an optional integration bolted on later; it is the part that keeps everyone honest, vendor included.
Compute the dollar figure yourself
We do not promise a return, and you should be suspicious of anyone who does, because they have never seen your close rate. What you can do is arithmetic with three numbers you already own: recovered conversations per month, times your close rate, times your average ticket. All three factors are yours. Your average ticket is sitting in your last hundred invoices. Your close rate is in your estimate log. Industry averages exist for both and are mostly noise — a plumbing outfit's average job and an HVAC replacement differ by a factor of twenty, and neither is your number. Do the multiplication yourself, on your figures, and the result is one nobody can spin at you in either direction.
What the numbers cannot tell you
Attribution is the honest weakness of all of this. A hailstorm, a heat wave, and a competitor closing up shop all move the same numbers a system claims to move, and a monthly count cannot separate them. Compare against the same month last year where you can, and hold conclusions loosely for the first quarter. Small samples are the second weakness: eleven conversations is a count, not a statistic, and one good month proves little. Third, some effects never appear in a monthly figure at all — the customer who was not left angry at eight in the evening shows up nowhere, and neither does the reputation that follows. Last, a good measurement design cannot rescue a bad system. It can only tell you sooner, which is the point of having one.
Where to start
Pick the one number that matches the one problem you actually have — unconnected calls, unfilled slots, or aging invoices — and pull its baseline this week, before you talk to any vendor, including us. It takes ten minutes, it costs nothing, and it makes every conversation that follows shorter and harder to fog.