Call scoring rubric - 4 signals in AI roleplay

"That was a good call" is a feeling until you name what you were looking at. A call scoring rubric holds only what two reviewers would mark the same way, and four signals in a sales conversation clear that bar, whether the call was live or a role play. This piece is about choosing those four; how the score then drifts in a reviewer's hands is the question next door, answered in sales call scorecard, and both are the manager half of sales coaching software for managers.

First, separate outcome from behavior

Sales leaders already have outcome numbers: won, lost, stage moved, meeting booked. Those are the numbers that matter, and they are also slow, noisy, and mostly out of the rep's hands on any single deal. A great call can lose to a budget freeze. A sloppy call can win because the buyer had already decided.

Behavior is different. Behavior is what the rep controls, it is visible immediately, and it is teachable. So the working split is simple. Score behavior, watch outcomes, and look for the relationship between them over months rather than deals.

Four signals you can actually score

Plenty of things in a conversation sound measurable and are not. "Rapport" and "confidence" are interpretations, and two reviewers will disagree about them. These four hold up better, because each one leaves a trace in the transcript that two people would mark the same way.

  • Structure. Did the conversation have a shape? Was there an opening that set the agenda, discovery before positioning, and a close? The tell is order and proportion. A rep who pitches in minute two has a structure problem, not a product knowledge problem.
  • Questions. How many, of what kind, and did the answers change anything. Open questions that go one layer deeper are the signal. Questions the rep asks and then ignores are the anti-signal.
  • Objection handling. Look only at the thirty seconds after the pushback. Did the rep acknowledge and clarify before answering, or did they defend, over-explain, or discount? That window is where most deals are actually decided.
  • Next step. Did the call end with a specific commitment, with a date and a named owner? "I will follow up next week" is not a next step. This is the single easiest signal to score and one of the most predictive of a deal that moves.

Four is roughly the right number. Rubrics with fifteen criteria look thorough and get ignored, because nobody can coach fifteen things at once.

A starter rubric

Signal Weak looks like Strong looks like
Structure Pitch first, discovery later or never Agenda, discovery, then a positioned solution
Questions Closed questions, no follow-up Open questions, one layer deeper, answers reused
Objection handling Defends, over-explains, discounts early Acknowledges, clarifies, reframes, holds
Next step "I will follow up" Named action, named owner, specific date

Publish the rubric to the team before you use it. A hidden standard is a trap, not a standard.

What scoring actually gives a manager

The obvious benefit is coverage. No manager listens to four hundred calls a month, so without scoring the sample is whatever they happened to sit in on, which is both small and biased toward the reps who ask for help.

The less obvious benefit is language. Scoring turns "he is just not great on the phone" into "his discovery is fine, he collapses on price." One of those is a judgment about a person. The other is a coaching plan for Thursday. Teams that adopt a shared rubric usually notice this before they notice any change in the numbers.

The third benefit is direction. A single score is close to meaningless. A per-rep trend across twenty sessions is not. It tells you whether the thing you coached three weeks ago actually stuck.

The limits of the rubric itself

  • It becomes the target. Once a rep knows the rubric counts open questions, you will get open questions, including pointless ones. This is Goodhart's law arriving on schedule. The defense is to score behavior clusters rather than single countable actions, and to change nothing about compensation based on a rubric score.
  • It is blind to context. A four minute call that disqualifies a bad-fit buyer is excellent work, and most rubrics will mark it low. Any score needs the call type attached to it, or you will punish efficiency.
  • It can encode style. If the rubric was built from your top performer's habits, it may be measuring resemblance to that person rather than effectiveness. Watch for systematically lower scores on non-native speakers, slower talkers, or a different but valid approach. That is a defect in the rubric, not in the reviewer.
  • It needs volume. Two conversations tell you almost nothing about a rep. Set a minimum number of sessions before anyone looks at an average, and never let one bad session become a label.

There is also a hard ceiling worth stating plainly. Conversation scoring cannot see the things that happen outside the conversation: the champion who left, the competitor's pricing, the procurement cycle. Anyone selling you a score that predicts revenue is overselling. What a score predicts is whether the rep is doing the things that make a good outcome more likely.

How to introduce it: roleplay first, live calls later

  • Start on roleplay sessions, not live customer calls. Lower stakes, faster feedback, no client risk.
  • Use scores for coaching direction first. Leaderboards can wait, or stay off entirely.
  • Show each rep their own trend before you show anyone the team view.
  • Pair the rubric with one real outcome metric so you can sanity check it.
  • Review the rubric every quarter and delete any criterion nobody has coached on.

This is roughly how scoring works inside pichi.ai. A rep runs a voice roleplay with an AI customer, gets scored against criteria you define, and sees a debrief immediately. The manager sees the pattern across the team without listening to a single session. The scores are a coaching instrument, not a verdict, and the product is designed to be used that way.

Want a rubric built around your sales process?

Send the criteria your team is coached on today to hello@pichi.ai. We will show how they map to roleplay scenarios and scoring.

Sales coaching software for managers → See pricing - $49 per seat