Skip to content
Playbook

How to choose a CRM — an evaluation that survives reality

Feature tables decide nothing. Eight probes you run inside a trial, each with the exact observation that passes or fails a product.

22 Aug 2026 12 min readBy Autocloz Editorial, Product team
How to choose a CRM — an evaluation that survives reality

A CRM evaluation that goes wrong nearly always started with a feature comparison. Feature tables have almost no discriminating power, because every serious product has every table-stakes feature and the differences that matter are behavioural. What works instead is a set of probes: eight specific things you do inside a trial, each with an observation written down in advance that passes or fails a product. Design the evaluation as an experiment and most shortlists collapse from eight products to three in an afternoon.

Start from the failure you are buying to fix

Before any product enters the list, write one sentence naming what is broken. Not what you want — what is failing now. The honest answers cluster into five, and they point at different products.

  • Deals are dying from missed follow-up.
  • Nobody can see the pipeline without asking a person.
  • History lives in individual inboxes and leaves when people do.
  • We are not having enough new conversations.
  • Reporting takes a day a month and nobody trusts the result.

The fourth one is the trap. Not enough new conversations is not a records problem, and buying a records system to solve it is the most expensive way of not fixing it. A conventional CRM will faithfully store the outcome of outreach it cannot perform. If that sentence is your sentence, the deciding criteria change entirely: sequence pacing, mailbox warmup, authentication monitoring, a dialer, and a compliance gate that runs before dispatch rather than a report that runs after it.

Write the sentence down and keep it visible, because every evaluation drifts toward whichever product demos best and the sentence is what pulls it back.

Probe one: does history capture itself, or does somebody remember?

This is the highest-leverage question in the whole evaluation and it is almost never on a comparison table, because every vendor answers yes to "does it log email" while meaning wildly different things.

Do this. From your ordinary mail client, send a message to a test contact that exists in the trial workspace. Do nothing else. Start a timer.

Watch for. Whether the message appears on that contact's record, how long it takes, whether the body is there or only a subject line, and whether the reply attaches to the same record when it comes back. Then repeat with a call from a mobile and a LinkedIn message if either is part of your motion.

Pass. It appears without any action, within minutes, with the body. Fail. It appears because you pressed something, or it appears as a stub, or it never appears and the vendor explains that you should send from inside the tool.

That last answer is not automatically disqualifying, but it changes what you are buying: a CRM whose capture depends on behaviour change will contain whatever your least disciplined rep felt like typing.

Probe two: does it model your business without customisation?

Do this. Create, out of the box, one instance of each thing you sell against. If you sell to organisations, create an organisation as a first-class record and attach two people to it. If you sell recurring revenue, create a deal that represents it. If you sell through partners, find somewhere to put the partner.

Watch for. Whether the object exists or whether you are being told it can be built. Customisation closes most gaps and creates the configuration you will maintain forever, which is where CRM projects go to die. A deals and pipeline surface that ships the objects you sell against is worth more than one that can be configured into them.

Pass. Every object you name exists with its own record type. Fail. The organisation turns out to be a text field on a contact.

It is worth understanding the shape underneath, because it explains behaviour you will otherwise find surprising. Autocloz separates the company, the person and the workspace-scoped lead into three tables: the company is unique per workspace and domain, the person carries the company reference with a delete rule that nulls it rather than cascading, and the lead is the person scoped into your workspace. The practical effect of that middle rule is that a person survives their employer's record being deleted, which is the correct behaviour for a world in which people change jobs. A product that models the person as a child of the company gets that case wrong, and you will meet the case.

Probe three: import two hundred real rows, mess included

Do this. Export 200 records from wherever they live now — including the rows with a missing company, the phone numbers formatted three different ways, and the two entries you suspect are the same person. Import them. Time it and note every field you could not map.

Watch for. How the tool matches an incoming row against an existing one, and whether it tells you. The match rule is a decision with consequences, and vendors rarely publish it. Autocloz matches on any of three identifiers — email case-insensitively, phone number once every non-digit is stripped, or an exact LinkedIn URL — so two colleagues who both typed the company switchboard number merge into one lead, while two spellings of the same LinkedIn profile stay two people. Neither behaviour is wrong; both are consequences you should learn during a trial rather than during a migration.

Pass. Your fields map, the duplicates are surfaced for review rather than silently merged, and rows that could not be imported are reported with a reason. Fail. A summary that says "200 imported" when you uploaded 214.

The import experience is also diagnostic in a second way. If mapping your fields is painful with 200 rows, the real migration will be worse by more than the ratio of the row counts, because the pain is in the field structure rather than the volume. The CRM migration checklist covers what that costs in sequence.

Probe four: hand a rep a task and count the clicks

Do this. Give someone on your team three tasks and do not help. Log a call with a disposition and a next step. Move a deal to the next stage and record why. Find every account nobody has touched in thirty days.

Watch for. The click count, and where they hesitate. Say nothing when they get stuck; the hesitation is the finding.

Pass. Each task completes in under a minute and the rep can repeat it unprompted. Fail. Any of the three needs the vendor.

The click count is your adoption forecast, and it is more predictive than any feature. A CRM your team keeps current beats a sophisticated one describing last month, because data quality compounds and features do not.

Probe five: show me a deal that has gone quiet

Do this. Ask, in the product, for every open deal with no activity in fourteen days. Not in a report you build — in a view you can reach.

Watch for. Whether the answer requires constructing anything. If finding a stalled deal takes a report, nobody will find one in practice, because the moment you need that list is the moment nobody has twenty minutes.

Pass. It is a filter or a saved view, and it is one click from where a manager already stands. Fail. "You'd build that in the report builder."

A related question worth asking in the same breath, because it exposes how the product stores stage history: can you see how long a deal spent in each previous stage? Many systems, Autocloz included, keep a single stage-entered timestamp that is overwritten on each transition, which means the current dwell time is available and the historical path is not. That is a real limit rather than a defect — storing every transition is a design choice with a cost — but it decides whether you can ever answer "which stage do our deals die in", so find out before you build a forecast that assumes you can.

Probe six: export everything on day one, before you put anything in

Do this. Within an hour of opening the trial, export contacts, companies, deals, custom fields, notes, activity history and attachments. Open the files.

Watch for. Which of those seven you actually got, in what format, on what timescale, and whether you could do it yourself or needed to ask.

Pass. All seven, self-service, immediately. Fail. Anything that needs a support ticket, or a "full export" that turns out to be a flat contact list.

Run this probe on the product you are leaving too, and run it today rather than at cancellation, because export access disappears with the subscription. Check the same question of every connector you will depend on — the Gmail and Google Workspace integration page is the kind of page worth reading for what it does not claim as much as for what it does. A cancelled account is often read-only for a short window and then gone.

Be exact about what an export contains rather than trusting the word. Autocloz's leads export is a nine-column CSV — email, name, company, status, phone, LinkedIn URL, score, source and creation date — capped at 100,000 rows, and it does not include custom fields, notes, tags, owner or the activity timeline. That is the honest scope of that particular file, it is worth knowing on the way in, and it is exactly the question to put to every other vendor on your list.

Probe seven: enrol yourself, reply, and then set an out-of-office

Do this. If the product runs any sequence or cadence, enrol a test address you control. Reply to the first message and watch what happens. Then re-enrol, set an auto-responder on that mailbox, and watch what happens.

Watch for. Whether the reply stopped the remaining steps, and whether the out-of-office did. Those two should behave differently, and a product that treats them the same is wrong in one direction or the other.

Pass. A human reply stops the sequence; an out-of-office does not, and the next step still fires on schedule. Fail. Either the follow-up arrives after your reply — the most visible failure in outbound — or a two-week absence permanently retires a live lead.

Push one question further, because it is the case single-channel checklists miss: if the tool sends on more than one channel, does a reply on one channel stop the others? Ask it directly rather than inferring it, and be prepared for the answer to be no in more products than you expect.

Probe eight: price the real user count, not the licensed one

Count the people who should be able to look at the pipeline — founders, finance, marketing, delivery, customer success — not the people who will edit it. Then price each shortlisted tool for that number.

Per-seat pricing quietly converts visibility into a budget decision, and the predictable result is that the people who most need context are the ones who do without it. Usage-based models price the other way. Autocloz's free plan carries 5 users, 10 mailboxes and 100,000 contacts. The seat cap is lifted at the Growth tier and above, so the honest reading is that a team of five or fewer sits comfortably inside the free plan and it is the seat count, rather than the contact count, that pushes a growing team upward.

Model twelve months rather than the entry tier, and add the four costs nobody quotes: migration of anything beyond contacts and deals, integrations that turn out to be "available via our API", add-on modules for marketing or quoting or AI credits, and the administration time somebody has to own. If nobody is named for that last one, data quality owns it, badly. The outbound stack cost calculator is a quick way to compare the twelve-month totals side by side rather than the sticker prices.

Autocloz's free plan covers 5 users and 10 mailboxes with the pipeline, custom fields, booking pages and mailbox warmup included — start free and run probes one, three and seven against it in an afternoon.

Run the two-week bake-off, then decide on one thing

Shortlist two. Run the same real segment through both for two weeks with the same people. Keep a defect list rather than having conversations, because most items resolve into two or three configuration fixes and the rest are the actual decision.

Then decide on a single criterion: which one will your team keep current? Everything above is a proxy for that. A simpler CRM holding an accurate pipeline beats a richer one describing a version of reality that stopped being true three weeks ago, and the probes exist because click count, capture behaviour and import friction are the observable predictors of whether that stays true.

Two cases where the standard evaluation misses. If your bottleneck is new conversations, weight probes one and seven far above the rest, and compare against tools built around sending — the Autocloz and Pipedrive comparison sets out where a pipeline-first product and an outreach-first one genuinely differ. And if your business is delivery-heavy, the handover from won deal to delivered work matters more than the pipeline, and a weaker CRM with project management attached will beat a better pure one.

What an evaluation cannot tell you, and what Autocloz does not do

Three limits on the method, then three on this product, because a page that only lists the other kind is not worth trusting.

A trial cannot show you year two. It shows you the product as it is on the day you tried it, configured by people who are paying attention. What it cannot show is how the configuration decays, which fields stop being filled, and which reports nobody opens. Assume decay and prefer the tool with fewer things to maintain.

A trial cannot price your migration. The cost of moving is dominated by automations, reports, permission structures and integrations, and none of those appear in a trial because you did not build them there.

A trial cannot test your data at volume. Two hundred rows behave differently from two hundred thousand, and filter performance is one of the few things a vendor will answer honestly if you ask directly.

On the product side: the free tier stops at 5 users, so a growing team hits the seat cap rather than a feature wall, and the cap is lifted only from the Growth tier upward. The AI features are not free either — AI personalisation and AI reply drafts sit at the Pro tier, and they run on your own OpenAI, Anthropic or Groq key rather than a metered credit pool, so the model quality is set by the key you connect. And quiet hours on the messaging channels resolve to the sending account's timezone, not the recipient's, which matters if you sell across regions and expected the opposite. If you want the definitional groundwork underneath all of this before you shortlist anything, what a CRM actually is covers the four objects and what the record has to hold.

Frequently asked

What is the most important thing to check when choosing a CRM?

Whether history captures itself. Ask specifically what happens when a rep sends an email from their normal mail client, takes a call on their mobile, or messages someone on LinkedIn — and whether each of those appears on the contact record with no further action. A CRM where logging is manual holds a partial history within a month and a misleading one within three, and no feature on a comparison table compensates for that.

How long should a CRM trial be?

One full sales cycle if you are replacing a system, and at minimum one real working week if you are buying your first. A week is long enough to surface the friction that a month of evaluation calls will not, because the friction lives in the fifth time somebody logs a call rather than in the first. Run it with real data and a real rep, not with sample records driven by the vendor.

Is a free CRM good enough for a small team?

Frequently yes, and the question to ask is which cap binds first at your shape rather than whether the tier is free. Autocloz's free plan, for example, covers 5 users, 10 mailboxes and 100,000 contacts, so a four-person team with a large list is unconstrained while a nine-person team is not — the same tier is generous or useless depending on which number you are near.

Why do CRM demos mislead?

A demo is a rehearsed path through a system configured by someone who knows it perfectly, populated with data designed to look good. Every product looks fluent under those conditions, so the demo has almost no discriminating power. Replace it with tasks you set, driven by your own people, against your own imported records, and watch without helping.

What hidden costs appear after signing a CRM contract?

Migration of anything beyond contacts and deals, integrations that turn out to be available via our API rather than natively, add-on modules for marketing, quoting, dialling or AI credits, seat creep as the team grows, and the administration time somebody has to own. Price the bundle you will realistically be on in twelve months rather than the entry tier you sign.

Should outbound live inside the CRM or in a separate tool?

It depends on which problem is actually binding. If your constraint is that existing deals are dying from missed follow-up, a records system fixes it and outreach can live anywhere. If your constraint is that not enough new conversations are starting, a CRM that stores the outcome of outreach it cannot perform will not move the number, and the deciding factor becomes whether the tool can pace sends, warm mailboxes, monitor authentication and gate calls.

Share
Free to start

Stop reading. Start sending.

Every tactic in this article is implemented behind the Autocloz dashboard.