100 Percent QA for Auto Dealers: Scoring How Every Priced Call Describes Price and Fees
A thin QA sample tells a dealer group how a handful of drawn calls sounded. It cannot tell the group how its stores describe price across thousands of conversations a month.

Key highlights
100 Percent QA for Auto Dealers: Scoring How Every Priced Call Describes Price and Fees
- Price and fees are described out loud, in a conversation, thousands of times a month across every store in a group.
- A thin sample reports on the calls somebody happened to draw, not on how the group describes price.
- On March 13, 2026 the FTC sent warning letters to 97 auto groups, saying advertised prices must be the total price including all mandatory fees.
- Variation between stores is a pattern, and a pattern needs volume before it is visible.
- Full coverage scores every priced conversation, human-handled and AI-handled alike, on every channel Orvera AI runs.
- One scorecard across every store and every partner is the only basis on which a comparison means anything.
- The scorecard is the dealer group’s own. Orvera AI measures conversations against it and does not determine what is lawful.
- Orvera AI runs on the stack the stores already operate, and full enterprise deployment lands in three to six weeks.
Why is a sampled call review a weak signal on how your stores describe price?
A one to three percent sample tells you how a handful of drawn calls sounded, not how your stores describe price across thousands of conversations every month.
Price description happens out loud, in real time, on every inbound and outbound call across every store in your group. A caller asks what the out-the-door number looks like. A representative quotes a payment. Fees get named, or they do not. That exchange repeats thousands of times a month, and a review built on a manual sample never sees enough of it to say anything about the group. The calls somebody happened to draw are not a representative picture. They are a handful of moments that say more about the reviewer's schedule than about your group's actual practices.
Variation between stores is where the real exposure sits. One store's representatives describe price one way. Another store describes it differently. That gap is a pattern, and patterns require volume to surface. A one to three percent sample flattens the signal. You cannot see store-level drift, you cannot see it by product line, and you cannot see it by individual representative. The March 13, 2026 FTC warning letters sent to 97 auto dealer groups put the industry on notice that advertised price description is under scrutiny (as of August 25, 2026, source: https://www.ftc.gov/news-events/news/press-releases/2026/03/ftc-warns-97-auto-dealership-groups-about-deceptive-pricing).

What did the FTC actually tell auto dealer groups in March 2026?
On March 13, 2026, the FTC sent warning letters to 97 auto dealer groups nationwide, stating that advertised prices must reflect the total price consumers will be required to pay, inclusive of all mandatory fees (as of August 25, 2026, source: https://www.ftc.gov/news-events/news/press-releases/2026/03/ftc-warns-97-auto-dealership-groups-about-deceptive-pricing).
The letters put the industry on notice by citing six example practices the FTC warned against:
- Advertising a price that does not reflect all required fees
- Advertising a price that reflects rebates or discounts not available to all consumers
- Advertising a price that fails to take into account an additional required down payment
- Conditioning the advertised price on consumers using dealer financing
- Requiring consumers to buy additional items not reflected in the advertised price
- Advertising unavailable or non-existent vehicles
Several of those practices surface, or fail to surface, in the conversation itself, when a rep describes the price and the fees to a consumer. Scoring every priced conversation is how a group sees for itself how its people described price on those calls.
What does 100 percent QA coverage look like on priced automotive conversations?
100 percent QA coverage means every priced conversation, human-handled and AI-handled alike, is scored against the same rubric on every channel Orvera AI runs, with no conversation left unreviewed.
Automated scoring removes the evaluator-to-evaluator variance that a manual program carries. A human reviewer brings judgment that shifts across shifts, reviewers, and fatigue levels. Automated scoring applies one consistent scorecard to every interaction, so the standard that grades the first call of the morning is identical to the standard that grades the last call of the evening. And it runs without a team of analysts pulling recordings by hand.
Reading the full population rather than a draw is what changes when volume goes from sampled to complete. A pattern that a one to three percent draw buries in noise becomes visible when the full population is scored. You can see which rep at which store describes mandatory fees differently from the scorecard standard, and you can see whether that gap holds across a vehicle line or sits with one person. That distinction matters when you are deciding where to focus coaching.
How does scoring every priced call show which store is drifting?
Scoring every priced call gives a dealer group a store-by-store comparison that a sampled review cannot produce, because the full population is the only data set where a gradual trend becomes visible before it becomes a problem.
Every rooftop in a large group develops its own floor habits. The language a store uses to describe documentation fees, add-on products, and mandatory charges tends to drift quietly over time, shaped by local management, staff turnover, and informal coaching. A sample never captures that drift reliably, because the same problematic phrasing can pass through unreviewed for weeks.
A score on every priced conversation, read store by store, turns a general impression into an exact comparison. One location may score consistently. Another, running the same script, may show a pattern of fee omissions that no one has named yet. The difference only surfaces when both stores are scored on their full volume.
Drift is gradual by nature. It appears first as a slow movement in a store's own trend line, not as an outlier call. That trend is invisible inside a drawn sample but fully readable across the whole population. Operations leaders responsible for fee-disclosure consistency across a group have no reliable signal without it.

Full coverage turns a suspicion into a measurement.
What can a dealer group do with full coverage that a sample never allowed?
Full coverage lets a dealer group act on every priced conversation rather than estimate from a fraction of them. The difference is not incremental. A sample tells you a pattern may exist. Scored coverage of every call, on every channel, lets you name the rep, the store, and the date.
In practice, full coverage opens four capabilities that a sampled review cannot replicate:
- Targeted coaching. When scoring surfaces a gap, the coaching conversation is about that call, that rep, that moment. It is not a general reminder delivered to a room where the rep who missed it sits alongside nine who did not.
- Policy verification. When leadership changes how price is described, full population scoring shows whether the change reached the floor. Inference gives way to a number.
- One standard across all partners. An outsourced BPO and an in-house team are scored against the same scorecard. There is no separate standard for the partner handling overflow.
- One rubric, in-house and outsourced. Orvera AI is multi-tenant and purpose-built for channel partners and BPOs, so a store handling its own calls and a partner handling overflow are scored on the same rubric rather than on two.
A score on every conversation is the precondition for all of it.
How do you keep the scoring itself inspectable?
A score a dealer group cannot inspect is not a control. It is an unexplained number attached to a conversation.
Every scored decision should trace back to the scorecard in force at the time of the call. That traceability is what separates a governable system from a black box. In practice, this means each result carries a record of which criteria applied, which threshold was set, and which version of the scorecard the conversation was measured against. If a store manager questions a score from six weeks ago, the answer should be retrievable in the same place as the score itself.
Orvera AI coordinates best-in-class third-party models inside its own governed enterprise layer and custom-trains its own contextualization models on de-identified data. That separation matters. The scoring criteria come from the dealer group. Orvera AI measures conversations against those criteria. It does not determine what is lawful, and it does not produce a report intended for a regulator.
Orvera AI is SOC 2 Type II certified, HIPAA compliant and GDPR compliant. Those standards govern how the data is handled, stored, and accessed. And the scorecard itself remains the dealer group's definition of what a well-described price conversation looks like, which is where inspectability begins.
What belongs in the internal case for scoring every priced conversation?
The internal case for scoring every priced conversation rests on a single operational fact: a group cannot evidence how it describes price from a fraction of the calls in which price is described.
Any quality program that treats sampling as a method has inherited a constraint, not adopted a standard. Sampling existed because human reviewers could not score faster than real time. That capacity limit shaped the practice. It did not validate it as sufficient.
Four points belong in every internal briefing on this topic:
- Sampling is a capacity artifact. No quality leader chose a thin review rate because it was adequate. They chose it because it was the ceiling manual scoring allowed.
- A fraction cannot evidence a pattern. If price is described on thousands of calls each month, a handful of reviewed conversations tells you almost nothing about what customers actually heard.
- One scorecard across every store and every partner is the only basis for comparison. A rubric applied unevenly across locations produces numbers that cannot be set side by side with meaning.
- Full coverage makes the data actionable. When every scored call feeds the same dataset, a group can identify which stores, which reps, and which call types carry the most inconsistency and act on that finding directly.
How does a dealer group move from sampling to full coverage without disrupting the floor?
A dealer group moves from sampling to full coverage by replacing a manual process with a managed platform that Orvera AI builds, deploys, and runs on the technology stack already in place.
The gap is not theoretical. A group that reviewed a thin sample of priced conversations last month cannot tell a general manager or an OEM partner how consistently its representatives described the out-the-door number across every store. That gap is an operational fact, not a tools problem. It does not close by asking QA analysts to score faster.
What closes it is a platform the group does not have to staff or configure. Orvera AI runs on the technology stack the stores already operate, with no rip-and-replace, and the integration directory is illustrative rather than a limit. Full deployment lands in three to six weeks, and Orvera AI does the build, the deployment and the integration. QA analysts spend that window learning to read results, not standing up infrastructure.
The practical first step is narrow: pull last month's call volume across every store and count how many priced conversations anybody scored. That number is the baseline. Everything above it is coverage that did not exist.
If that number is smaller than it should be, talk to the team at Orvera AI about what full coverage looks like on your floor.
Frequently asked questions
Because a sample reports on the calls somebody happened to draw, not on how price is described across every store in the group.



