← Overview
Appointment Economics · Article 32

Why Pilot Projects Fail — and What Makes a Good POC Different

·10 min readBefore any investment decision

In almost every organization I talk to about appointment management, there is a backstory. They have "tried it before." It "didn't work." What follows is a skepticism directed at the concept itself—even though it usually belongs to the test setup.

Because when I ask what this pilot actually looked like, I hear the exact same setup with astonishing regularity: one branch, one quarter, voluntary participation, no baseline, no defined success criteria. In the end, you're left with a gut feeling and a decision based on that gut feeling. A pilot like this cannot succeed—it can only survive or not.

The Four Design Flaws

1 · No Baseline Measurement

The most serious and most common flaw. If nobody knows what the no-show rate was before, how many appointments were generated per week, and how many of them led to a closed deal, nothing can be proven afterward. The result isn't an evaluation, but a collection of impressions—and before an investment committee, impressions lose to any real number.

The effort required is minimal: count for two to four weeks, track three metrics, handwritten notes are enough. Skipping this saves a few hours and destroys the basis for making a decision.

2 · The Pilot Starts with the Skeptics

The rationale sounds reasonable: If it works at the hardest location, it will work everywhere. In practice, the opposite happens. An early pilot is always unfinished—appointment types aren't cleanly defined yet, workflows aren't smoothly established. Passing through this phase in the very place where people are looking for confirmation of their own resistance generates the exact narrative that will subsequently be quoted across the entire company.

The pilot is not a stress test. It is a learning process. The stress test comes during rollout, and then with a process that already works.

3 · Too Short and Too Broad at the Same Time

Two mistakes that reinforce each other. Too short: Six weeks aren't enough because the first three weeks measure the transition itself rather than its impact. Too broad: Testing three business units, two countries, and four appointment types simultaneously makes it impossible to attribute a single effect at the end.

The opposite combination is what actually works—narrow in scope, generous in time. One business unit, two to three appointment types, three months.

4 · No Defined Success Criteria

The question "how will we know that it worked?" is remarkably rarely answered before launch. Afterward, everyone answers it for themselves—in a way that confirms their own initial stance. Sales sees extra effort, marketing sees potential, IT sees integration challenges. Everyone is right, because nobody specified beforehand what actually matters.

The Special Case: The Pilot with the Wrong Tool

A variation that is particularly expensive: Testing with a simple calendar tool because it's readily available, and then drawing conclusions about the concept itself from the results. As a rule, such tools cannot handle skill-based routing, coverage rules, multi-tenancy, or reliable documentation.

When the pilot then shows that customers end up with the wrong advisor and documentation remains incomplete, the internal conclusion becomes: "Online appointment booking doesn't work for us." In reality, something entirely different was tested. This mix-up often blocks organizations for years.

How to Build a Pilot That Supports a Real Decision

Four Phases Over Roughly Four Months

Phase 1
3–4 weeks
Measure without changing anything. No-show rate, appointments per advisor per week, conversion rate per appointment type, time from inquiry to conversation. In parallel: Put the success criterion in writing and get sign-off from department leadership. One sentence is enough—but it must exist before the start.
Phase 2
2–3 weeks
Set up narrowly. One business unit, two to three appointment types, clearly named and with realistic durations. Participants are interested volunteers—plus at least one manager who actively participates.
Phase 3
10–12 weeks
Run and support. A short weekly check-in: What snagged, what worked. Adjustments are explicitly encouraged and documented—the pilot is meant to yield a better process, not test an unchanged hypothesis.
Phase 4
2 weeks
Evaluate against the baseline. The same four metrics, the same segment. Plus the qualitative question for participants: Would you want to switch back? This answer is often more insightful than any percentage.

The Success Criterion: One Sentence Before Launch

It must contain three elements—a metric, a direction, and a timeframe. For example: "The pilot is considered successful if, after three months, the no-show rate in the test group is at least one-third below the baseline and participants do not want to switch back."

This is uncomfortable because it explicitly allows for the possibility of failure. That is precisely why it works: A pilot that cannot fail proves nothing. And a criterion approved in advance stops debate over interpretation before it even starts.

What You Should Have in Hand at the End

  • Before-and-after tracking for four key metrics, collected using the exact same definitions—this forms the core of your proposal to the board or committee
  • A list of everything changed during the pilot—this serves as the blueprint for rollout and prevents every location from repeating the same mistakes
  • The answer to the switch-back question, quoted verbatim
  • An honest effort estimate for scaling up, including internal labor hours—omitting this loses you the debate at the first follow-up question

And one sentence I leave with every project team: A pilot whose result surprises no one was either redundant or improperly designed. If in the end everyone only sees what they already believed, nothing was measured—a vote was taken.

Takeaway

Pilots fail due to four structural flaws: no baseline measurement, starting with skeptics, being too short and too broad at the same time, and lacking pre-defined success criteria. A workable setup is narrow in scope and generous in time—one department, three appointment types, four months, four key metrics. And the success criterion is written in one sentence before anyone begins.

Sources

Patterns and phase model: Calenso/jrni project experience from enterprise rollouts and from the analysis of failed preliminary projects (2024–2026). The cases described are anonymized and generalized; individual companies, providers, and products are intentionally not named.

Share with the team

Three things a leader should define before the next pilot:

  1. Establish a baseline — before anything else. Four weeks, four metrics. Without a baseline, the results cannot be proven later, no matter how well the pilot runs.
  2. Write down the success criterion in one sentence and get it approved. Metric, direction, timeframe. This eliminates debates over interpretation at the end.
  3. Cut the scope in half, double the duration. The most common adjustment needed in pilot plans presented to me — and the cheapest.
ShareLinkedInWhatsAppEmail

Get new articles by email

Double opt-in: you will receive a confirmation email first. Unsubscribe anytime with one click. GDPR compliant.