Skip to content
Playbook

Sales forecasting methods — and when each one lies

Weighted pipeline, rep commit, run-rate and stage velocity, each with its arithmetic and its systematic bias — plus how to score whether your numbers mean anything.

22 Aug 2026 15 min readBy Autocloz Editorial, Product team
Sales forecasting methods — and when each one lies

A forecast is a probability statement, which means it can be scored rather than argued about. The four methods in common use — weighted pipeline, rep commit, historical run-rate and stage velocity — each fail in a specific, predictable direction, and knowing the direction is more useful than the number. This walks through the arithmetic of each, the bias each carries, how to score whether your probabilities carry any information at all, and what the pipeline underneath has to be true for any of it to mean something.

A forecast is a decision tool, not a prediction

The purpose of a forecast is to let somebody commit resources — hire, spend, extend runway — at a known level of confidence. That reframing does real work, because it means a forecast that is consistently wrong in a known direction is more useful than one that is randomly right.

A number that runs 20% high every quarter can be corrected by dividing by 1.2. A number that is sometimes 30% high and sometimes 15% low cannot be corrected at all, and the team learns to discount it by an amount nobody can state. That is the worst of both worlds: the ritual of forecasting with none of the value.

Which produces the first practical rule, and it is a filing discipline rather than a modelling one. Record the forecast, record what closed, keep the pair. After four periods you have an estimate of each method's bias for your business, and that estimate is worth more than any individual quarter's figure.

Weighted pipeline, and the arithmetic that shows when it breaks

How it works. Every open deal's value is multiplied by a probability and the results are summed. The probability may be attached to the stage, or set per deal — in Autocloz the forecast endpoint sums value × probability ÷ 100 using the deal's own probability, so a stage default is a starting point that a rep is expected to correct.

Where it is right. At portfolio scale with many deals of similar size, probabilities derived from your own closed history behave reasonably. It is also the cheapest method, because most systems compute it automatically.

Where it lies. The failure is not imprecision, it is that the output describes an outcome that cannot happen. Work it through.

Take ten deals of ₹10,00,000 each, all at 50%. Weighted total ₹50,00,000. The realistic range of outcomes runs from zero to one crore, and the middle of that distribution is genuinely around fifty lakh — five of ten closing is the most likely single result. The number is informative.

Now take one deal of ₹1,00,00,000 at 50%. Weighted total is the same ₹50,00,000. But the possible outcomes are exactly two: one crore, or nothing. Fifty lakh will not occur under any branch of the future. The forecast is wrong in every possible world, and averaging across a portfolio that does not exist is the reason.

The practical threshold: as soon as your largest open deal is more than roughly a fifth of the weighted total, the weighted number has stopped being a forecast and started being a summary statistic. At that point the useful presentation is a scenario table — what the period looks like with the big deal, and without it — rather than one figure.

Weighted pipeline also inherits whatever stage discipline you have. If deals advance because time passed rather than because a buyer committed, the probabilities are attached to nothing, and what each stage has to prove before a deal enters it is the prerequisite rather than a refinement.

Use it when you have many deals of similar size and the probabilities come from your own closed history rather than from the defaults that shipped with the product.

Rep commit, and the incentive attached to it

How it works. Each rep names the deals that will close this period. You add them up.

Where it is right. Reps know things the data does not: that the champion is leaving, that procurement has gone quiet, that the buyer said one thing on the call and something else in the corridor afterwards. No model captures that, and no amount of activity data substitutes for it.

Where it lies. It is a human estimate with incentives attached, and the bias is usually consistent per person rather than random. Some reps sandbag reliably; others are reliably optimistic. That is workable once you know each person's historical accuracy and useless until you do — which is an argument for tracking commit accuracy per rep rather than for distrusting commit.

The second failure is the alignment trap. If commit drives compensation, quota relief or the tone of a Monday meeting, it stops being an estimate and becomes a negotiation. You will get the number the rep believes is safest to say, which is a different quantity from the number they believe.

There is a third and subtler failure worth naming: anchoring. If the rep sees the weighted pipeline figure before giving a commit, the commit is no longer independent information — it is the weighted number with a human adjustment. Collect commit before showing anyone the mechanical figure, or you have two copies of the same estimate and one of them is pretending to be a second opinion.

Use it when deal counts are low and values are high, which is exactly the regime where weighted pipeline fails. Track per-rep accuracy over at least four periods before trusting any of it.

Historical run-rate, and what it cannot see

How it works. You closed a certain amount in each of the last several periods. Project forward, adjusted for seasonality and growth.

Where it is right. As a sanity check it is excellent, and it is the fastest way to notice that a bottom-up forecast has drifted into fantasy. If the pipeline forecast is triple your best-ever quarter, the pipeline is wrong, and no amount of stage discipline explains a tripling.

Where it lies. It assumes the next period resembles the last ones. It cannot see a new competitor, a changed motion, a market shift, or the fact that you doubled the sales team six weeks ago. It is a trailing indicator being asked to lead.

It also degrades quietly during exactly the periods when you need a forecast most. A business in a stable state barely needs forecasting; a business that just changed something needs it badly, and that is precisely when run-rate's assumption breaks.

Use it as the sanity check on every other method and never as the primary one during change.

Stage velocity, and the volume it requires

How it works. Measure the historical conversion rate between each pair of stages and the average time spent in each. Push current deals through those rates and durations to project when and how much will close.

Where it is right. It is the only common method that forecasts *timing* as well as amount, and it surfaces where deals actually die. If discovery-to-proposal converts at 70% and proposal-to-close at 15%, you have located your problem, and that diagnosis is worth more than the forecast that produced it.

Where it lies. It needs volume. With thirty deals a quarter your stage conversion rates have error bars wide enough to swallow the signal, and a single unusual quarter moves them enough to change next quarter's projection materially. It also assumes deals move forwards, and quietly mishandles the ones that skip stages, go backwards, or sit in a stage for three months without either.

There is a compounding subtlety. Multiplying four stage conversion rates together multiplies their errors too, so a chain of four rates each measured on modest samples produces a final number with a much wider interval than any individual rate suggests. The diagnostic value survives that; the point estimate mostly does not.

Use it when deal flow is high enough for the rates to be stable, and use it primarily for diagnosis. Where the cycle is actually spending its time is the question stage velocity answers best.

Scoring whether your probabilities mean anything

Every method above except run-rate depends on probabilities. Almost nobody checks whether theirs carry information, and there is a standard, sixty-year-old tool for exactly that.

The Brier score was introduced by Glenn W. Brier in "Verification of Forecasts Expressed in Terms of Probability", Monthly Weather Review, volume 78, issue 1, 1950. It is the mean squared difference between the forecast probability and the outcome, where the outcome is scored 1 if the event happened and 0 if it did not. Lower is better; zero is perfect; always forecasting 50% on a coin-flip event scores 0.25.

Applying it to a pipeline takes an afternoon. Take every deal closed last quarter. For each, record the probability it carried 30 days before it closed and whether it was won. Square the difference, average across all of them. Then do the same for a naive baseline — everything at your overall win rate — and compare.

The result is one of three things, and each has a different action:

  • Your score beats the baseline. Your probabilities carry information. Keep using them and keep scoring.
  • Your score matches the baseline. Your probabilities are a relabelling of the base rate. Weighted pipeline is doing nothing that multiplying open value by your win rate would not do more honestly.
  • Your score is worse than the baseline. Your probabilities are actively misleading, which usually means they are being set by stage position rather than by evidence.

The more diagnostic view is calibration. Bucket closed deals by the probability they carried — everything near 30%, near 50%, near 70% — and check what fraction of each bucket actually closed. A well-calibrated set has 70% deals closing about 70% of the time. Almost every sales pipeline that has never been checked is over-confident in the upper buckets, and the correction is arithmetic rather than cultural: if your 70% bucket closes at 45%, relabel it.

Measuring the forecast itself: bias, then accuracy

Two different measurements, and conflating them is why forecast reviews go in circles.

Bias is the signed error. Forecast minus actual, averaged over periods, keeping the sign. It answers "do we systematically run high or low", and it is the number you can correct for. A method with a consistent +18% bias is a usable method with a correction factor.

Accuracy is the unsigned error. Absolute difference, averaged. It answers "how far off are we typically", and it tells you how wide your reported range needs to be.

A warning about the popular ratio version. Mean absolute percentage error behaves badly when the denominator is small: a quarter that closed a fraction of the usual amount produces an enormous percentage error that dominates the average and makes the metric useless in exactly the quarters worth studying. Prefer absolute currency error alongside the signed bias, and quote both against your typical period size.

The review that produces these is short and has to be boring to work. At the start of every period, write down the number each method produced and lock it. At the end, write what closed. Four rows in and you have a bias per method; eight rows in and you can tell a bias from a drift.

What the software computes, and what it leaves to you

Being specific about a real implementation is more useful than describing forecasting in general, so here is what Autocloz's forecast endpoint actually returns and where its edges are.

It computes open deal count and value, a weighted value as the sum of value × probability ÷ 100, won and lost counts and values inside a window you choose between 7 and 365 days, a historical win rate, a closing-soon bucket for deals with an expected close date inside 90 days, and separate roll-ups for each forecast category — commit, best_case, pipeline and omitted — so the rep's judgement is reported next to the mechanical number rather than mixed into it. It obeys the same visibility scope as the deals list, so a member restricted to their own deals sees a forecast of their own deals rather than the team's.

Four edges worth knowing, because each one changes how you read the output:

Win rate is by count, not by value. Won deals divided by won plus lost. If your wins are systematically smaller than your losses, the count-based rate flatters you, and a value-weighted win rate is a different and often less comfortable number you compute yourself.

Closing-soon excludes overdue deals. An open deal whose expected close date is in the past is not counted as closing within 90 days. That is correct — the alternative counts every stale deal as imminent — and it means a rising overdue population is invisible in that tile and has to be watched separately.

Multi-currency conversion uses a static rate table. Values are converted per currency before summing, and any currency without a rate is summed one-to-one and named in an unconverted_currencies list rather than silently dropped. Honest, and still not a live rate.

Nothing stores the history that calibration needs. The endpoint reports the current state. Recording each period's forecast against what closed is a spreadsheet you keep, and it is the single highest-value forecasting artefact in this whole article. Note also that the deal forecast is a Growth-tier feature rather than part of the free plan.

Autocloz's free plan covers 5 users and 10 mailboxes with the configurable pipeline, stage gates and forecast categories included — start free and get the hygiene right before the forecast is worth computing.

The combination that works, in order

No single method is sufficient. A practical sequence:

  1. Calculate weighted pipeline as the mechanical baseline, and check whether your largest deal breaks it.
  2. Collect rep commit separately, without showing anyone the weighted number first. Anchoring destroys the independent signal, and an anchored commit is not a second opinion.
  3. Sanity-check both against run-rate. If either is far outside your historical range, find out why before reporting it.
  4. Score the probabilities once a quarter with the calibration check above. Relabel the buckets that are wrong rather than arguing about them in a meeting.
  5. Record all three numbers and what actually closed. After four periods you know each method's bias for your business.
  6. Report a range with the bias named. Committed, best case, worst case, plus the sentence that says how the primary method has behaved recently.

That last clause is what turns a forecast into a decision tool. A number with a stated bias can be used by a finance team. A number presented as certain and then missed twice teaches everyone to discount it by an unknown amount.

Dedicated revenue-intelligence products go considerably further on this than a CRM's forecast tile does, including conversation data and deal-risk scoring; the honest trade is set out on the Autocloz and Clari comparison, and the analytics surface that feeds the roll-up shows what is computed natively.

The data problem underneath every method

Every method above depends on the pipeline being current. A forecast built on deals whose stage was last updated three weeks ago is arithmetic performed on fiction, and no methodology fixes that.

Which makes pipeline hygiene the highest-leverage forecasting work available:

  • Every open deal has a next step with a date. A deal without one is a hope with a value attached.
  • Age-in-stage is visible and reviewed. Deals that sit are the ones inflating the forecast, and they inflate it most in the highest-probability stages.
  • Closed-lost happens promptly. Reluctance to mark a deal lost is the largest single source of inflation, and it is emotional rather than analytical — a lost deal removed in week two costs nothing, and the same deal removed in week eleven has been forecast three times.
  • Activity captures itself. If updating the CRM is a separate chore from doing the work, the data will lag reality by exactly as long as reps can get away with. Autocloz writes every email, call, LinkedIn message, SMS and WhatsApp to the deal automatically for this reason.

The metrics that surround the forecast on the same board are the other half of the review.

What no forecasting method can do, and what Autocloz does not compute

The limits, stated plainly.

No method forecasts a deal that is not in the system. Everything above operates on recorded pipeline, so a business whose reps carry half their opportunities in their heads is forecasting the half it can see and calling it the total.

No method survives a changed motion in its first two periods. New segment, new price, new team — every historical rate is now describing a business that no longer exists, and the honest response is a wider range and a stated caveat rather than a confident number.

Calibration needs closed deals. A team closing eight deals a quarter cannot fill probability buckets fast enough to calibrate them within a year, and should lean on rep commit with tracked per-rep accuracy instead.

Autocloz specifically. It does not implement stage-velocity forecasting — there is no conversion-rate-and-duration projection, so that method is a spreadsheet you build from exported data. It stores no forecast history, so bias and calibration are yours to record. Its win rate is count-based, its currency conversion is a static table, and its weighted value uses the probability on the deal, which means the number is exactly as good as the discipline behind that field. There is no deal-risk scoring, no conversation intelligence and no automatic detection that a champion has gone quiet. And the pipeline itself enforces the gates you configure and nothing you do not — a forecast is a report on your process, and it cannot be better than the process it reports on.

Frequently asked

Which sales forecasting method is the most accurate?

None of them is accurate in isolation, because each has a systematic bias rather than random error. Weighted pipeline is reasonable across many similar-sized deals and meaningless when one deal dominates; rep commit carries whatever incentive is attached to it; run-rate is a trailing indicator used as a leading one; stage velocity needs volume its rates do not have at small deal counts. Running two or three and recording each one's bias against what actually closed is what produces a usable number.

What is a Brier score and why does it matter for forecasting?

It is the mean squared difference between a forecast probability and the outcome, scored as 1 for happened and 0 for did not, introduced by Glenn W. Brier in "Verification of Forecasts Expressed in Terms of Probability" in Monthly Weather Review volume 78, issue 1, in 1950. Lower is better and zero is perfect. It matters because it is the only way to tell whether the probabilities on your deals carry information — a set of 70% deals that closes 40% of the time scores badly and should be recalibrated, not defended.

Why does weighted pipeline break on large deals?

Because it reports an average of outcomes that cannot occur. A single deal worth one crore at 50% contributes fifty lakh to the forecast, and the deal will either close for one crore or for nothing — fifty lakh is not a possible result. Across many similar deals those errors cancel and the average is informative. With a handful of large ones they do not cancel, and the weighted figure is wrong in every branch of the future.

How many periods before I can trust my forecast?

At least four closed periods with the forecast and the actual both recorded, because you are estimating a bias and one or two observations cannot distinguish bias from noise. Record the number each method produced at the start of the period, not a version revised midway, and record what actually closed. The output is a correction factor per method and per rep, which is worth more than any single period's figure.

Should a forecast be a single number or a range?

A range, with the assumptions named. A single number implies a precision you do not have and invites decisions that assume it, and when it is missed twice everyone silently discounts future forecasts by an unknown amount. Committed, best case and worst case, plus a sentence stating how the primary method has behaved over the last four periods, is a forecast someone can actually make a hiring or spending decision against.

What data has to be true before any forecasting method works?

Every open deal needs a next step with a date, an owner, a value, and a stage that was changed because the buyer did something rather than because time passed. Age-in-stage has to be visible and reviewed, and closed-lost has to happen promptly, because reluctance to mark a deal lost is the largest single source of forecast inflation and it is emotional rather than analytical. No methodology repairs a pipeline whose stages were last touched three weeks ago.

Share
Free to start

Stop reading. Start sending.

Every tactic in this article is implemented behind the Autocloz dashboard.