Sales metrics and KPIs to track for outbound (the ones that matter)
A metric earns its place by changing a decision. The eight worth a weekly review, the cohort error that breaks most dashboards, and what nobody computes for you.
A sales metric earns its place by changing a decision. If no plausible value of a number would cause anyone to do something differently this week, it belongs in an archive rather than a dashboard. For outbound specifically, that test cuts a typical report from thirty numbers to about eight, and it exposes a structural error that makes many of the remaining ones wrong: dividing this month's outcomes by this month's activity when the two are separated by weeks. This is the shortlist, the decision each number triggers, and which of them your tooling computes for you.
A metric earns its place by changing a decision
Run the test on your current report, metric by metric. For each one, write the sentence: "if this number were X, we would do Y." Any number without a completed sentence gets deleted.
The exercise typically kills three categories.
Totals with no denominator. "Emails sent this month: 41,000." No decision follows, because the number is neither good nor bad without knowing what it produced.
Numbers nobody owns. If a metric moves and no named person is expected to respond, its movement is information going nowhere.
Numbers that only move slowly, reviewed weekly. Average days to close cannot meaningfully change in a week. Reviewing it weekly produces commentary about noise, and commentary about noise trains everyone to ignore the report.
What survives is small, which is the point. A report that fits on one screen gets read. The one number that beats all the others as a starting point is the count of open records with no next step, because it is the only measure that reliably predicts a quiet month three weeks ahead.
Counters, rates and ratios are three different objects
Mixing these on one dashboard is why arguments about performance go in circles.
A counter is a count of events in a window. Sends, replies, calls, meetings booked. Counters answer "how much did we do?" They are honest, easy to verify, and easy to inflate deliberately.
A rate divides one counter by another from the same stage. Reply rate, bounce rate, connect rate. Rates answer "how well is this stage working?" They are only comparable when both counters use the same denominator definition, which is the detail that breaks most comparisons between teams and tools.
A ratio compares stages that are separated in time. Meetings per opportunity, opportunities per win. Ratios answer "where is the funnel leaking?" and they are the ones most damaged by the timing problem in the next section.
Three practical rules follow. Never put a counter and a rate in the same column without labelling which is which. Never compare a rate across two systems without checking both denominators. And never compute a rate on fewer than a few hundred events, because at small numbers the rate is mostly noise — the arithmetic of how wide those error bars actually are is worked through in what counts as a good reply rate.
The period-versus-cohort error that breaks most outbound dashboards
This is the most consequential mistake in outbound reporting and it is almost invisible, because the resulting numbers look plausible.
Period accounting asks: how many meetings happened in September, divided by how many emails went out in September. Cohort accounting asks: of the accounts we started contacting in the first week of September, how many have produced a meeting so far?
They differ because outbound has a lag. A sequence runs for three weeks, a reply arrives in week two, a meeting is booked for the following week, an opportunity is created after that. The outcome of September's sending largely lands in October.
The distortion has a direction, which makes it worse than random error.
While volume is growing, period accounting understates conversion. The denominator includes this month's larger sending, whose outcomes have not arrived, while the numerator only reflects last month's smaller sending. Scale up and your conversion rate appears to fall.
While volume is shrinking, period accounting overstates conversion. The same mechanism in reverse. Cut sending and the numbers improve, which is the most dangerous false signal a sales report can produce.
The fix is procedural rather than technical. Assign every outcome back to the week the outreach started, and report by that week. Then set a maturity rule — a cohort is not complete until, say, six weeks have passed — and mark immature cohorts as incomplete rather than comparing them with mature ones. Two months of doing this and your conversion numbers become stable enough to act on.
One practical note on tooling. Autocloz's report summary returns a headline block of counts alongside a previous_totals block for the immediately preceding equal-width window, with options to compare against the same window a year earlier or against a window shifted back by a chosen number of days. That structure supports period-over-period comparison well. It does not do cohort assignment for you, so the cohort view is one you build from exports.
The eight numbers worth a weekly review, and the decision each triggers
- Accounts newly contacted. *Decision:* whether the top of the funnel is being fed at all. A week at zero shows up as a quiet month six weeks later, and by then it is too late to fix cheaply.
- Replies. *Decision:* whether messages are arriving and landing. A sharp fall with steady sending is a delivery investigation, not a copy rewrite.
- Positive replies. *Decision:* whether the offer and the targeting match. This is the number that most closely tracks pipeline, and almost nobody computes it automatically.
- Meetings held. *Decision:* whether booked meetings survive to the calendar. Booked and held are different numbers and the gap between them is its own problem.
- Opportunities created. *Decision:* whether the qualification bar is set correctly. A high meeting-to-opportunity rate usually means the bar is too low rather than that the meetings are excellent.
- Pipeline value created. *Decision:* whether the volume is enough for the target. Value created this month against the coverage you need is the only forward-looking number on this list.
- Win rate on outbound-sourced deals. *Decision:* whether outbound-sourced deals close like the rest. Reviewed monthly, not weekly.
- Average days to close. *Decision:* whether the cycle is stretching. Reviewed monthly. A stretching cycle changes what you should be doing today, because today's activity lands later than you assumed.
Two of these — meetings held and positive replies — are the ones teams most often skip because they require a judgement rather than a query. They are also the two with the highest information content, which is not a coincidence: the numbers that need a human are the ones that carry the meaning.
The guard metrics underneath, and their real thresholds
These are not performance measures. They are the numbers that tell you the machine is still allowed to run, and they should be looked at more often than the ones above, because they degrade fast and quietly.
Spam complaint rate. Google's sender guidelines ask senders to keep the spam rate reported in Postmaster Tools below 0.10% and to avoid ever reaching 0.30% or higher. Those are the published thresholds, and they are measured by Google rather than by your sending tool.
Bounce rate. A rising bounce rate is a list-hygiene signal before it is a reputation one, and it is the earliest warning that a data source has degraded.
Unsubscribe rate. Reads as a targeting metric. A rise usually means the message reached people it was not written for rather than that the message got worse.
Authentication status. Binary and worth checking after any DNS or provider change, because it fails silently and the failure looks like a copy problem.
Sends per mailbox per day. The number you set rather than one you measure, but worth reporting because drift here explains changes in every metric above it.
Two honesty notes. The complaint threshold above is Google's own published figure for Gmail, and other receivers publish different ones or none. And a guard metric inside your sending tool is your view of events, not the receiver's; Postmaster Tools reports what Google saw, which is the number that decides your fate. The fuller diagnostic procedure lives in the deliverability posts rather than here.
Autocloz's free plan covers 5 users and 10 mailboxes with campaign analytics and DMARC monitoring included — start free and check the guard numbers against your own sending before optimising anything above them.
Which numbers your system computes, and which you compute yourself
Knowing exactly where the boundary sits saves a fortnight of assuming a number exists.
Autocloz computes, per campaign: sent, attempted — a distinct count of sends across both delivery and bounce events, so a synchronously bounced send still counts as an attempt — delivered, recipients_reached (which sums the primary recipient plus any carbon-copied alternate addresses, and so differs from sent, which counts one per lead), opens, unique_opens, clicks, unique_clicks, replies, bounces, unsubs, and the derived open_rate, click_rate, reply_rate and bounce_rate. Automated activity is separated rather than blended: machine_opens and machine_clicks are reported alongside the human counts.
At workspace level, the summary returns counts of sends, replies, bounces, calls, active campaigns, open tasks, total leads, and the call-side breakdown of leads called, leads answered and leads not called. The charts endpoint adds deal-level figures: total, won, open and lost deals, won value, open pipeline value, average deal value, and average days to close, computed over won deals as the mean gap between creation and closing.
What it does not compute, and what you therefore own:
Positive-reply rate. The classifier that stops a sequence on reply sorts human from out-of-office from bounce and carries no sentiment at all — a "yes please" and a "never contact me again" are both human. There is no positive-reply counter in the reporting layer. If you want this number, it comes from a review pass or a second classification step, and it is worth the effort because it is the metric closest to pipeline.
Meetings held, as distinct from booked. Meeting outcomes are recorded — confirmed, cancelled, rescheduled, no-show, completed — but the no-show stamp is manual by default, so held-versus-booked depends on somebody marking it.
Cohort conversion. See the section above. The windows are period-based.
Anything about a deal's heat. The deal record carries an engagement score column documented as a nightly calculation from recent activity. Nothing in the product writes it; it is exposed by the API and holds its default of zero. Do not build a report on it.
The pipeline and funnel views themselves are in the analytics surface, and the unit-economics half of the picture — what an acquired customer costs and what it returns — is a separate calculation with its own boundary decisions, worked through in customer acquisition cost explained and computable directly with the CAC and LTV calculator.
A worked month, end to end
Illustrative figures throughout, chosen to show the arithmetic rather than to represent any customer.
Week 1 cohort: 900 accounts contacted, 2,700 emails scheduled across three touches.
By week 3, that cohort has produced 63 replies. Reply rate against contacted accounts is 7.0%. Of those, a review pass finds 19 expressing interest: a positive-reply rate of 2.1% of accounts, and 30% of replies.
Of the 19, 11 book a meeting and 8 are held. Booked-to-held is 73%, and the 3 lost there are worth as much attention as anything upstream, because they were already the most interested people on the list.
Of the 8 held, 4 become opportunities: a 50% meeting-to-opportunity rate, which is high enough to ask whether the qualification bar is too low. Average deal value of ₹4,00,000 gives ₹16,00,000 of pipeline created from one week's cohort. At a 22% historical win rate, expected revenue is ₹3,52,000.
Now the decisions this arithmetic supports, which is the entire point of computing it:
- Reply rate of 7% against a stable history means the top of the funnel is fine. No action.
- Positive replies at 30% of replies is the number to watch. If it falls while total replies hold, the targeting has drifted, not the copy.
- Booked-to-held at 73% means 27% of the most interested people never spoke to anyone. That is the cheapest fixable leak on the page, and it is a reminders-and-confirmation problem rather than a sales one.
- Meeting-to-opportunity at 50% prompts a review of what qualifies, not a celebration.
- ₹16,00,000 of pipeline from 900 accounts gives roughly ₹1,780 per account contacted, which is the number that tells you what another 900 accounts is worth and whether the list is worth buying.
Notice what is absent. No opens, no clicks, no activity counts per rep. None of them would have changed a decision on this page. If you want a fuller model of how these stage rates roll forward into a number you can commit to, sales forecasting methods covers where each method starts lying.
Goodhart arrives faster than you expect
Any metric you publish and reward becomes a target, and targets deform. The social anthropologist Marilyn Strathern's 1997 formulation of Goodhart's law puts it plainly: "When a measure becomes a target, it ceases to be a good measure." Charles Goodhart's own 1975 phrasing, from monetary policy, is that any observed statistical regularity tends to collapse once pressure is placed on it for control purposes.
In outbound this is not theoretical, and it shows up within weeks.
Activity counts. Reward sends and sends rise, quality falls, and complaint rate follows a fortnight later. This is the fastest-acting example on the list.
Meetings booked. Reward bookings and unqualified meetings get booked, which moves the cost from the sales team to whoever attends them.
Opportunity creation. Reward opportunities and the qualification bar moves down until it disappears.
Three defences. Pair every volume metric with a quality metric reported next to it, never underneath it. Prefer metrics the person cannot directly produce — replies over sends. And re-derive the definitions periodically, because the erosion happens in the definitions rather than in the numbers.
What these numbers do not tell you, and what Autocloz does not measure
They do not tell you why. A metric localises a problem to a stage; the cause is found by reading messages, listening to calls and asking people, and no dashboard substitutes for that.
They do not compare you to anyone. There is no standards body publishing outbound benchmarks, and vendor-published figures are drawn from their own customer bases with undisclosed sampling frames and undisclosed definitions. Your own trailing numbers are the only sound comparison you have.
They do not attribute. A closed deal usually had outbound, inbound and a referral touching it, and every attribution model is a choice about which of those to credit rather than a discovery.
Autocloz specifically does not compute a positive-reply rate, does not distinguish a held meeting from a booked one without a manual mark, does not assign outcomes to cohorts, and does not write the deal engagement-score column at all. Its reporting is period-windowed with a prior-window comparison, its reply detection carries no sentiment, and its lead status vocabulary is a fixed set — new, queued, active, paused, replied, bounced, unsubscribed, converted and do-not-contact — enforced at the database level, so any richer stage model you need lives in the separately configurable status labels or in the deal pipeline rather than in that column. If you are weighing this against a platform whose pitch is the reporting layer itself, the Outreach comparison is where those differences are set out.
Frequently asked
What are the most important outbound sales metrics?
The ones that change a decision you actually make. For most outbound teams that is a short list: contacted accounts, replies, positive replies, meetings held, opportunities created, win rate, average days to close, and pipeline value created. Everything else is either an input to one of those or a diagnostic you consult when one of them moves. A metric nobody has ever acted on should be removed from the report rather than explained again.
Why do outbound dashboards disagree with the pipeline?
Usually because the dashboard divides this month's outcomes by this month's activity while the lag between the two is several weeks. That is period accounting applied to a process that needs cohort accounting, and it makes conversion look worse while volume is growing and better while it is shrinking. Assigning each outcome back to the week its outreach started removes the distortion entirely.
Is open rate still worth tracking?
Not as a performance measure. Apple's Mail Privacy Protection prevents senders from seeing whether a message was opened for recipients who have it enabled, so an open rate is partly a measure of your audience's mail client. It retains narrow diagnostic value when a figure collapses to near zero across every campaign at once, which usually indicates a rendering or delivery problem rather than a content one.
What is a positive reply rate and how do you calculate it?
It is the share of sends that produced a reply expressing interest, rather than any reply at all, and it has to be calculated from a human or model judgement about each reply because reply detection alone carries no sentiment. Divide positive replies by the same denominator you use for total replies so the two are comparable. Most systems do not compute it for you, which means it is a number you own.
How often should outbound metrics be reviewed?
Weekly for activity and reply-level numbers, monthly for conversion and cycle-length numbers, and quarterly for anything about deal value. The rule underneath is that a metric should be reviewed no more often than it can meaningfully move, because reviewing a slow number weekly produces reactions to noise, and reacting to noise is worse than not looking.