Ask the measurement question first
Everywhere else in the business, a number produced by a person gets checked for repeatability before anyone acts on it. Conversation scoring usually skips that step entirely. A manager listens, marks a rubric, and the result goes into a spreadsheet with the same authority as a booked meeting.
The first question about a scorecard is not whether the scores are high. It is whether a second reviewer, on a different Tuesday, would produce roughly the same marks on the same call. That property has a name, inter-rater reliability, and a scorecard that does not have it is not a measurement. It is an opinion with a number attached, which is more dangerous than an opinion, because nobody argues with a number.
What follows assumes the criteria are already picked. It is about the scoring itself, and about the six places where the scoring quietly stops being a measurement.
Six ways the score drifts
| Drift | How it shows up | What to do |
|---|---|---|
| Outcome leaks in | The deal closed, so the call gets marked strong. The deal died, so the same behaviors get marked weak. | Score before the outcome is known, or score from a transcript with the result stripped out. |
| The last call dominates | A rep who stumbled yesterday is scored down on work from three weeks ago. | Score each session at the time it happens. Never score a month in one sitting. |
| Tone stands in for decisions | Confident, warm, fast reps score high regardless of what they actually did. | Write every criterion as something that either happened or did not, with a timestamp. |
| Reviewers drift apart | One manager runs generous, another runs strict, and both change over the year. | Calibration: everyone scores the same two recordings each quarter and compares marks. |
| The sample chooses itself | You review the calls reps submitted and the ones you happened to join. | A fixed selection rule. Every Nth call, or one per rep per week picked at random. |
| Everyone gets a three | The middle of the scale absorbs every call and the numbers stop separating anybody. | Fewer points, each with words. A described four-point scale beats an undescribed ten. |
The outcome is the loudest liar
Knowing how a call ended changes what you hear in it. Once you know the deal closed, the rep's long product monologue reads as thorough. Once you know it died, the same monologue reads as failing to listen. This is the halo effect running in both directions, and it destroys the entire purpose of scoring behavior separately from results.
It matters most in exactly the situations where you need the score to be independent. A seller who was already motivated before anyone picked up the phone will make a mediocre acquisitions call look like a masterclass. A seller who had a better offer in hand that morning will make a genuinely good call look like a failure. If the reviewer knows which is which, the rubric is measuring the seller, not the rep.
A roleplay session is the cheap way out of this particular drift. There is no deal outcome to leak into the score, because there was no deal. The reviewer marks what the rep did and nothing else, which is the whole point of scoring behavior separately from results.
The last call becomes a label
Recency does the most lasting damage, because it is the drift that turns into language about people. A manager scoring twenty sessions at review time is not scoring twenty sessions. They are scoring their memory of twenty sessions, and memory of a body of work is dominated by the most recent one and the worst one.
The fix is boring and it works. Score at the time. Look at trends across a defined number of sessions rather than at any single mark. And keep a hard rule that one bad session never becomes a sentence about a person, because "he collapses under price pressure" is a very hard thing to un-say once a team has heard it.
Tone is the most seductive mistake
Tone is the easiest thing to hear and the hardest thing to defend. It is also where bias walks in. Accent, speaking pace, gender, age, and seniority all register as confidence or the lack of it, and none of them belong in a score. A rubric with a row called "confidence" will systematically punish people for how they sound.
A soft-spoken acquisitions rep who asks the one question that surfaces why the seller is actually selling has done the more valuable thing than a loud one who never asked it. On any tone-weighted rubric, the loud one wins. If you cannot point at the second in the recording where a criterion was met or missed, that criterion is measuring taste.
Score the conversation, not the person. The moment a scorecard starts describing who somebody is, it has stopped describing what they did.
Five tests for a scorecard you already have
- The disagreement test. Two people score the same call independently, without discussing it first. Disagreement on more than a criterion or two means the wording is broken, not the reviewers. Rewrite the wording, not the people.
- The timestamp test. For every criterion, can the reviewer name the moment it was met or missed? If a criterion cannot be pinned to a point in the recording, delete it.
- The counter-example test. Describe a genuinely good call that your rubric scores badly. Every rubric has one. A short call that correctly disqualifies a bad-fit buyer is usually it. Know yours before it embarrasses you in front of the team.
- The compensation test. If a score touches pay, promotion, or a ranking on a wall, it stops being a measurement and becomes a negotiation. Reps will optimise the rubric, and they will be right to.
- The deletion test. Take last quarter's criteria and cross out every one that nobody actually coached on. What remains is your real rubric. The rest was decoration that made scoring slower and no more accurate.
What a trustworthy score is for
Three uses, and they are narrower than most teams expect. A per-rep trend over enough sessions to be stable. A cluster view of where a whole team is weak, which is usually the most valuable output and the one nobody is looking for. And a before and after around a specific coaching intervention, which is the only way to find out whether the coaching did anything.
What it is not for: ranking people publicly, deciding pay, or standing in as a single verdict on whether somebody is good at their job. A scorecard used that way stops reporting on reality within about a quarter, and the manager is the last person to find out, because nothing visibly breaks. The numbers keep arriving. They are simply about something else now.
Practically, most of these fixes get easier when scoring happens on roleplay sessions rather than on live customer calls. Every session gets scored the moment it ends instead of at review time, and the same scenario can be run by everyone, which is what makes reviewer calibration possible at all. That is how scoring works inside pichi.ai: the rep runs the roleplay against an AI buyer, your criteria are applied the same way every time, and the manager reads a trend instead of a memory.