DemonstrationThe metros are real. The businesses, the customers, and every figure on this site are invented to show how the instrument behaves — nothing here measures a real business.
Caliper
Method

How these numbers are made,
and what they are worth.

Everything here is fixed in advance and published before it is used. A measurement whose rules can be adjusted after the results arrive is not a measurement.

§1

Why this exists

Fixed questions are not the innovation. Your dealership already runs a standardized survey after every repair order, on a 1–10 scale, and has done for decades. It is better funded than this and gets a far higher response rate.

It is also unusable as public information, because that score sets dealer bonus money and inventory allocation, and manufacturers typically count only the top box — so a 9 hurts. The predictable result is that customers are asked for perfect scores before they leave the lot, and an entire vendor industry sells score improvement as a service.

What is missing from this market is not a fixed instrument. It is an instrument whose custody is separated from the money that depends on the answer. No dealership can pay Caliper for anything: not placement, not certification, not a badge, not to have a quarter looked at again.

§2

The instrument

Ten statements per sector, grouped into the same 5 dimensions. Every verified customer rates each statement for the business they used on a 0–10 scale, from strongly disagree to strongly agree, or answers not enough information.

One scale for all ten items, on purpose. Mixed anchors are the usual way a battery quietly stops being comparable to itself.

  • ResponsivenessWhether you can reach someone who can help, and whether things happen when promised.
  • TransparencyWhether what you were told about money turned out to be what you paid.
  • IntegrityWhether commitments survived contact with an opportunity to break them.
  • CompetenceWhether the thing you came for was actually done, the first time.
  • RespectWhether you were treated the way anyone else would have been.

The full wording of all ten items →

§3

Verification

You cannot rate a store you did not visit. A rating requires a repair order number — the one printed on your invoice — matched against that store's record of visits. One repair order is one voice, enforced by a unique index in the database rather than by the interface.

The visit must also fall inside the quarter being measured. A rating always refers to a specific visit, in the period being reported, which is what separates this from a review site where a complaint from 2019 still drags on today's average.

Verification failures are answered vaguely on purpose — "that repair order isn't one we can match" — because a precise error message would turn the form into an oracle for guessing valid repair order numbers at somebody else's store.

§4

The Service Conduct Index

Each item's score for a store in a quarter is the mean of the answers that gave a number. "Not enough information" is excluded from the mean and counted separately.

A dimension is the mean of its two items. The Index is the mean of all ten item means, rescaled from the 0–10 item scale to 0–100. Every item carries equal weight; no dimension is treated as more important than another, because we have no defensible basis for such a judgement.

index = mean(item means) × 10
e.g. items averaging 7.42 on the 0–10 scale → an Index of 74.2

The Index is a summary, not the measurement. The item-level series are the measurement, which is why every scorecard prints all ten of them with the sample size beside them. A store can hold a respectable Index while failing one dimension badly, and the breakdown is the only place that shows.

§5

“Not enough information” is data

Most instruments treat a don't-know as a hole to be dropped. Here it is recorded and published as a rate, because in a service department it measures something specific: whether customers can tell what was done to their car.

A store whose Index holds steady while its don't-know rate climbs quarter after quarter is not holding steady. It is explaining less, and the people still able to answer are an increasingly unusual group.

§6

The reporting floor

A store's quarter is published only if at least 25 verified customers rated it that quarter. Below the floor the figures are withheld and the gap is labelled as withheld — on the charts it appears as a break with a marked tick, never a smooth line drawn over the hole.

The floor is easy to clear here, which is worth being plain about: a mid-size service department writes well over a thousand repair orders a quarter. The binding constraint on this instrument is response rate, not recruitment.

§7

What this cannot tell you

It is not a census of customers. The manufacturer's survey goes to everyone with a repair order. Ours goes to people who created an account and chose to answer. Verification proves the visit was real; it does not make the responders typical.

Coverage is thin. Google has a rating for every dealership in the country. Caliper covers a handful of stores in a handful of metros, and a store not on this site is not being criticised — it is simply not measured.

Small differences are not differences. Two stores a point or two apart, or two adjacent quarters a point or two apart, are for most purposes the same. Read the direction over several quarters, not the decimal.

Organized answering is possible. A repair order is a real barrier — far higher than a review site's — but a competitor with access to invoices is a threat no satisfaction survey fully solves. Sudden discontinuities are treated as suspect and investigated before publication rather than reported as findings.

This measures conduct, not mechanics. Whether a repair was technically correct is not something a customer can reliably judge, and nothing here should be read as a verdict on a technician's work.

§8

Corrections, and what a business can ask for

A dealership may write to us to dispute a figure, and we will check it. What a dealership may not do is pay for that check, pay for placement, pay for a badge, or pay to have a quarter removed. If a published quarter turns out to be wrong, the corrected figures are republished with the correction noted against that quarter.

Changes to the instrument — a retired item, a new item, a change to the floor — are dated and appear on the battery page before they take effect, never after the data they would affect.

§9

Getting the data

Every scorecard is available as JSON, including the per-item series and every sample size, so the figures can be checked rather than taken on trust.