Search toggle
Search toggle

How to Run a Restaurant Tech Pilot That Proves Something

restaurant technology pilots checklist

Most restaurant technology pilots prove nothing because nobody set a falsifiable metric or a fixed window before launch. Here's the pre-launch design that fixes it.

Companies chasing new technology tend to launch many small, disconnected pilots rather than a single rigorous test, a pattern McKinsey calls 'pilot purgatory' where use cases never make it to scaled proof. Source: McKinsey & Company, "It's the Last IT/OT Mile That Matters in Avoiding Industry 4.0's Pilot Purgatory" (2018).

The cause is almost always the design, not the product. The team launches at 5 stores, watches for a while, collects opinions, and writes a summary nobody can argue with, because nobody set a number that could have come back negative.

A pilot proves something when you name one primary metric, fix a measurement window, and pick control stores before the first install. Everything after launch is just data collection. Here's how to build that design, and the ways a pilot gets quietly rigged toward "success" without the technology ever getting tested.

 

How long should a restaurant tech pilot run before the data means anything?

Run it 90 days, with the first 30 excluded from the result. Anything shorter measures novelty and the training curve, not the technology.

Weeks 1 through 4 belong to adoption. Staff are learning a new flow, the GM is watching more closely than usual, and the vendor's implementation team is on site or on call. Throughput dips, error rates spike, and a lot of operators read that dip as a failed product.

That dip is the cost of change. It shows up in every rollout regardless of which tool you bought.

Weeks 5 through 12 are the measurement window. That's the period the go/no-go decision cites, and you commit to it in writing before install so nobody can move the goalposts when week 7 looks bad.

Fix the calendar, not just the length. A 90-day window that spans a holiday period at half your pilot stores and a normal period at the other half isn't comparable to anything. Pick a window where the pilot group and the control group face the same seasonality. If you can't, extend the window rather than start it early.

Long windows have a real enemy: the vendor's quarter. Sales pressure to convert a pilot before a year-end contract deadline is the single most common reason a 90-day window gets cut to 45. Write the end date into the pilot agreement and price the extension up front. You'll thank yourself in month 2.

 

Which metrics do we lock in, and how many is too many?

One primary metric, 2 secondary, 1 guardrail. That's the scorecard to set before an install.

The primary metric has to be tied to a cost lever you already track:

  • kWh per store per week
  • Telecom cost per unit per month
  • Disputed delivery orders recovered per location
  • Ticket average on direct digital orders
  • Minutes from order-ready to guest-in-hand

If the metric only lives in the vendor's dashboard that’s not helpful.

The guardrail is what you refuse to trade away. Ticket times can't go up. Help-desk ticket volume can't double. Guest complaints per 1,000 orders can't rise.

A pilot that wins the primary metric and breaks the guardrail is a no-go. Saying so in advance is what keeps the decision honest.

Element

Set before install

Fails when

Primary metric

1, from a bill or statement you already reconcile

It only exists in the vendor's dashboard

Secondary metrics

2, max

The list grows past 3 and something always wins

Guardrail

1, with a hard threshold

Nobody wrote the number down

Measurement window

Weeks 5-12, dates fixed

The vendor's quarter shortens it

Control stores

3 minimum, matched on volume and daypart mix

Chosen after results come in

Signers

CFO, VP of Ops, IT

Only the champion signs

 

More metrics make a pilot easier to pass, not harder.

Track 9 things and 2 will improve by chance, which hands the champion a slide and gives you nothing real. McKinsey found that getting from pilot to scale requires leaders to "be honest about what pilots have worked" and to cut down on the number of experiments running at once. Same discipline, just applied inside a single pilot instead of across a portfolio.

That's also the line that separates a pilot from vendor evaluation. Deciding whether to engage at all is a different job. Our 3-question framework for turning vendor noise into a decision covers it. By the time you're designing a pilot, that argument is already settled.

 

How many stores, and what makes a fair control group?

5 to 8 pilot stores and at least 3 controls, matched on weekly sales volume, daypart mix, and drive-thru or dine-in share. Below that, one unusual store swings the whole result, and you'll spend a month arguing about whether it counts.

Match on the variables that drive the metric you chose. An energy pilot matches on square footage, equipment age, and climate zone. A phone-order pilot matches on call volume and how many orders currently go unanswered. A delivery-fee pilot matches on marketplace mix, because a store that's 70% one platform behaves nothing like a store split three ways.

Pick the controls the same day you pick the pilot stores, and name them in the agreement.

Controls selected after results arrive are decoration. So are pilot stores hand-picked because they have the strongest GM in the region, which happens to be the most common way a pilot gets quietly rigged: you've proven the technology works when a great operator runs it. You already knew that.

Include one store you expect to struggle. If the tool only performs at your best locations, a 300-unit rollout will tell you that in month 4 instead of month 1, and by then the contract is signed.

Scale is where pilot design gets expensive. McKinsey's work on pilot purgatory found companies running an average of about 8 pilots per priority use case, roughly 10 in China and India, with use cases never reaching scaled proof. Restaurant operators repeat that pattern one category at a time: 4 small tests in energy, none of them designed to be compared.

Size doesn't protect you either. That's the part that should worry bigger operators most. Restaurant Dive reported, citing The Wall Street Journal, that Taco Bell began U.S.-wide drive-thru adoption of its AI ordering technology in 2024 and had to rethink the rollout in 2025 after real-world performance proved inconsistent.

A chain of that size and sophistication still hit conditions the earlier test hadn't reproduced. Impressive infrastructure, and still not enough. The lesson is to build a pilot whose stores look like the stores you'll roll out to.

 


 

FAQ

Should the vendor see the pilot data while it's running?
Yes, on a fixed schedule, and never as the sole owner of the numbers. Give them the operational data weekly so implementation problems get fixed fast. Keep the primary metric calculated on your side, from your bills and statements, and share it at the same cadence you agreed in writing. A vendor who can recut the measurement mid-pilot will.

How do we tell a training dip from a product failure? 
The dip recovers. The failure doesn't. Discard weeks 1-4, then look at the slope across weeks 5-8. Improving toward baseline is adoption. Flat or worsening by week 8 is the product, and you should say so out loud before week 12 so the vendor has a shot at fixing it.

Who signs the go/no-go document?
The CFO, the VP of Operations, and the IT or technology lead, on one page. It states the primary metric result against the pre-set threshold, the guardrail result, total cost of ownership per unit per month at rollout scale, the integration work required, and a plain go, no-go, or extend. If only the internal champion signs it, you don't have a decision. You have an advocate.

What if the pilot succeeds but rollout economics don't? 
Kill it. A tool that pays back at 6 stores with a hands-on implementation team and doesn't at 200 with your help desk is a no-go. Run the per-location run rate at full unit count during the pilot, not after, and include integration labor and the internal hours the rollout eats.

Does this apply to a free or no-cost-up-front pilot?
Yes, and more strictly. Free pilots skip procurement scrutiny, which is exactly why they get launched without a metric or a window. Store hours, staff attention, and guest experience are the real cost, and they're spent whether or not you're invoiced.

 

Where One Goal fits

We've placed vetted technology across 150+ trusted brands and 12+ tech verticals. That means we've seen which pilot designs held up after rollout and which ones got reversed 9 months later. Honestly, the second pile is bigger than most operators expect.

That pattern is what we bring to the conversation: the metric a CFO will accept, the control stores that make the number defensible, and the window that survives a vendor's quarter-end.

Our recent pilots are primed to save operators 15% on energy costs, measured against their electric bills.

See how our vendor portfolio can help.

Matt Haselhoff

With 27 years in the restaurant technology industry, Matt Haselhoff has become a trusted advisor for restaurant brands navigating an increasingly noisy and fragmented technology landscape. From his early days at Cherry Electric to his role as Chief Revenue Officer at Omnivore—ultimately acquired by Olo—Matt has seen firsthand how the relationship between technology vendors and restaurant operators has become strained and inefficient.

Comments

Related posts

Search Who Actually Signs: The Restaurant Tech Buying Committee