Sales Call Scorecard Template: 20 Criteria for Better Coaching
Copy a 20-criterion sales call scorecard in simple, weighted 100-point and machine-readable JSON formats, with calibration and coaching guidance.

A sales call scorecard is a structured rubric for evaluating observable behaviours in a sales conversation. This template gives you 20 criteria in three usable forms: a quick yes/no checklist, a weighted 100-point model and machine-readable JSON. Use it as a starting point, then adapt the criteria to your call type, sales process and mandatory policies.
The score is not the coaching. The score tells you where to look; transcript evidence and a specific replacement behaviour make the result coachable.
Which version of the template should you use?
Version | Best for | Output |
|---|---|---|
Simple checklist | First manual review or manager coaching | Yes / No / N/A plus evidence |
Weighted 100-point scorecard | Comparing calls while preserving priorities | Normalized score plus criterion results |
Machine-readable JSON | Consistent storage, automated scoring or system integration | Versioned criteria, weights, evidence and review fields |
If you are new to structured review, begin with the simple checklist for five to ten calls. Add weights only after reviewers agree on what each criterion means. If you first need the category definition, read What Is AI Call Scoring?.
The 20-criterion sales call scorecard template
Each criterion below is observable: a reviewer should be able to point to a timestamp, transcript passage or missing step. The weights are a deliberate starting model, not an industry benchmark. They total 100 points and put the most emphasis on discovery, value alignment, objections and a committed next step.
# | Criterion | Observable evidence | Weight (points) |
|---|---|---|---|
1 | Introduces self and company | Identity is clear without a misleading claim. | 2 |
2 | States the purpose of the call | The buyer can explain why the conversation is happening. | 3 |
3 | Sets or confirms an agenda | Both sides agree on what the call should cover. | 5 |
4 | Establishes the current situation | The rep learns the buyer’s existing process or context. | 5 |
5 | Identifies the primary need or problem | The need appears in the buyer’s own words. | 7 |
6 | Explores business or personal impact | The consequence of leaving the problem unresolved is clear. | 6 |
7 | Clarifies decision criteria | The buyer states what will matter when comparing options. | 4 |
8 | Clarifies decision process | Stakeholders, approvals or purchasing steps are identified. | 4 |
9 | Clarifies timeline | A real timing constraint or target date is captured. | 4 |
10 | Clarifies constraints | Budget, policy, technical or operational limits are surfaced. | 4 |
11 | Uses open and relevant questions | Questions advance discovery rather than fill silence. | 4 |
12 | Follows up on important answers | The rep probes instead of immediately returning to the pitch. | 4 |
13 | Summarizes and confirms understanding | The buyer gets a chance to correct the rep’s interpretation. | 4 |
14 | Communicates clearly | Language, pace and explanation fit the buyer and context. | 3 |
15 | Connects value to stated needs | The proposal refers to needs actually expressed on the call. | 7 |
16 | Keeps claims accurate and supportable | Promises, comparisons and capabilities stay within approved facts. | 5 |
17 | Identifies the real objection | The rep distinguishes the stated concern from the underlying issue. | 4 |
18 | Responds to the objection with relevant evidence | The response addresses the concern without evasion or pressure. | 5 |
19 | Checks whether the objection is resolved | The buyer confirms, qualifies or rejects the response. | 4 |
20 | Secures and recaps a specific next step | Owner, action and date are explicit; commitments are summarized. | 20 |
Simple checklist version
Copy the 20 criteria into a spreadsheet or form. Add these columns: Met, Missed, N/A, Evidence timestamp and Coaching note. Do not score a criterion from memory when the recording or transcript is available.
- Met: the required behaviour is clearly present.
- Missed: the behaviour was applicable but not demonstrated.
- N/A: the behaviour genuinely did not apply to this call type or stage.
- Evidence timestamp: where the behaviour occurred, or where the opportunity was missed.
- Coaching note: one specific behaviour to repeat, stop or replace.
A rubric should reflect the objective of the activity being evaluated. The UK government’s Magenta Book describes rubric criteria as being based on intended objectives and notes that the resulting judgement is transparent but context-specific. That is exactly why a generic template must be adapted to your sales motion. UK Government Magenta Book evaluation guidance.
Weighted 100-point scorecard
Section | Criteria | Available points |
|---|---|---|
Opening and agenda | 1–3 | 10 |
Discovery | 4–9 | 30 |
Constraints and communication | 10–14 | 15 |
Value and objections | 15–19 | 25 |
Close and follow-up | 20 | 20 |
Total | 20 criteria | 100 |

Score each applicable criterion with one of three ratings:
Rating | Multiplier | Meaning |
|---|---|---|
Met | 1.0 | Clear evidence satisfies the criterion. |
Partial | 0.5 | Some evidence exists, but the behaviour is incomplete or ambiguous. |
Missed | 0.0 | The criterion applied and was not demonstrated. |
N/A | Excluded | The criterion did not apply and its weight leaves the denominator. |
Sales call score formula
For criterion i, let wᵢ be its point weight and rᵢ be its rating multiplier. Calculate:
Normalized score = [Σ(wᵢ × rᵢ) ÷ Σ(applicable wᵢ)] × 100
When all criteria apply, the applicable weights total 100 and the calculation is simply the sum of earned points. When one or more criteria are N/A, remove their weights from both the earned-points calculation and the denominator. Never convert N/A into zero; doing so penalizes a rep for a behaviour the call never required.
Worked example — illustrative, not a benchmark
Section | Available points | Earned points |
|---|---|---|
Opening and agenda | 10 | 8 |
Discovery | 30 | 21 |
Constraints and communication | 15 | 12 |
Value and objections | 25 | 18 |
Close and follow-up | 20 | 12 |
Total | 100 | 71 |
The illustrative score is 71/100. It does not mean 71 is a good or bad universal threshold. A pass threshold should be derived from your own approved standard, calibrated reviews and the consequences of error. Keep the five section scores visible: they explain whether the coaching priority is discovery, value communication, objection handling or the close.
How should N/A and auto-fail rules work?
N/A means the criterion was outside the scope of the interaction. Auto-fail means a specific failure overrides the pass decision because its consequence is unacceptable. They solve different problems and should never be treated as synonyms.
Do not copy generic auto-fail rules from the internet. Define them from your applicable law, contract, safety requirements, privacy rules and approved internal policy. Possible categories—not universal rules—include a missing mandatory disclosure, an unauthorized commitment, mishandling sensitive information or abusive conduct.
- Keep the raw numerical score even when an auto-fail changes the pass status.
- Record the exact rule that triggered and the supporting evidence.
- Require human review for disputed, ambiguous or high-impact auto-fails.
- Version the auto-fail policy separately from ordinary coaching criteria.
- Do not use auto-fail for ordinary skill gaps such as weak discovery.
How to calibrate the scorecard before using it
Calibration tests whether different reviewers apply the same definitions to the same call. It is not a meeting where everyone is told to agree with the manager.
- 1Select a small set of calls covering typical, strong, weak and edge-case conversations.
- 2Have at least two reviewers score the calls independently.
- 3Compare criterion-level ratings, not only final totals.
- 4Open the evidence for every disagreement.
- 5Rewrite ambiguous criteria and add positive, partial and missed examples.
- 6Repeat the independent review after the rewrite.
- 7Track disagreement over time and recalibrate after material script, product, policy or call-type changes.
The UK government’s principles for AI-assisted marking recommend regularly tracking inter-rater reliability, monitoring raters and detecting rater drift. Those controls transfer well to human or automated call scoring because both depend on stable rubric interpretation. Principles of AI use in marking.
For automated scoring, document who reviews uncertain outputs and consequential decisions. NIST’s AI Risk Management Framework playbook specifically recommends documenting the degree of human oversight over AI outputs. NIST AI RMF Measure guidance.
Machine-readable JSON scorecard
A machine-readable scorecard separates the rubric definition from the result produced for one call. The definition contains stable IDs, weights and rules. The result contains ratings, confidence, evidence and review status. This prevents a label change from silently breaking historical reporting.

The example below is a complete structural template. Criterion labels and approved rules should be stored alongside these stable IDs in your implementation.
{
"version": "1.0",
"scorecard_id": "sales-call-core",
"name": "Sales Call Scorecard",
"score_scale": {
"met": 1,
"partial": 0.5,
"missed": 0,
"not_applicable": null
},
"formula": "sum(weight * rating) / sum(applicable weights) * 100",
"auto_fail_policy": {
"enabled": true,
"preserve_raw_score": true,
"rule": "Only organization-approved legal, safety, privacy or conduct failures may override pass status."
},
"categories": [
{
"id": "opening",
"name": "Opening and agenda",
"weight": 10,
"criteria": [
{
"id": "introduces_identity",
"weight": 2,
"evidence_required": true
},
{
"id": "states_purpose",
"weight": 3,
"evidence_required": true
},
{
"id": "confirms_agenda",
"weight": 5,
"evidence_required": true
}
]
},
{
"id": "discovery",
"name": "Discovery",
"weight": 30,
"criteria": [
{
"id": "current_situation",
"weight": 5,
"evidence_required": true
},
{
"id": "primary_need",
"weight": 7,
"evidence_required": true
},
{
"id": "impact",
"weight": 6,
"evidence_required": true
},
{
"id": "decision_criteria",
"weight": 4,
"evidence_required": true
},
{
"id": "decision_process",
"weight": 4,
"evidence_required": true
},
{
"id": "timeline",
"weight": 4,
"evidence_required": true
}
]
},
{
"id": "constraints",
"name": "Constraints and communication",
"weight": 15,
"criteria": [
{
"id": "constraints",
"weight": 4,
"evidence_required": true
},
{
"id": "open_questions",
"weight": 4,
"evidence_required": true
},
{
"id": "follow_up_questions",
"weight": 4,
"evidence_required": true
},
{
"id": "clear_communication",
"weight": 3,
"evidence_required": true
}
]
},
{
"id": "value_objections",
"name": "Value and objections",
"weight": 25,
"criteria": [
{
"id": "value_alignment",
"weight": 7,
"evidence_required": true
},
{
"id": "accurate_claims",
"weight": 5,
"evidence_required": true
},
{
"id": "real_objection",
"weight": 4,
"evidence_required": true
},
{
"id": "relevant_response",
"weight": 5,
"evidence_required": true
},
{
"id": "resolution_check",
"weight": 4,
"evidence_required": true
}
]
},
{
"id": "close",
"name": "Close and follow-up",
"weight": 20,
"criteria": [
{
"id": "specific_next_step",
"weight": 16,
"evidence_required": true
},
{
"id": "recap_commitments",
"weight": 4,
"evidence_required": true
}
]
}
],
"result_fields": [
"criterion_id",
"rating",
"confidence",
"evidence_start",
"evidence_end",
"evidence_text",
"review_required",
"review_note"
]
}If you validate the file with JSON Schema, remember that declaring a property does not automatically make it mandatory; required properties must also appear in the schema’s required array. Official JSON Schema object guidance.
How to turn a score into coaching
A useful coaching note contains four parts: the observed moment, the impact, the replacement behaviour and the next review point.
Part | Example |
|---|---|
Observed moment | At 06:14, the buyer mentions an approval committee; the rep moves to pricing without asking who participates. |
Impact | The decision process remains unknown, so the next step may exclude a required stakeholder. |
Replacement behaviour | Ask: “Who else needs to be comfortable with this decision, and what will each person evaluate?” |
Next review point | Check the next three applicable discovery calls for stakeholder and approval-process evidence. |
Use the scorecard to identify one or two repeatable behaviours, not to deliver twenty criticisms at once. For a fast manual workflow, pair this template with the Sales Call Audit Guide. For the broader case for moving beyond sporadic reviews, see What 10,000 Unreviewed Sales Calls Taught Us About Lost Deals.
Common scorecard mistakes
- Scoring the outcome instead of the behaviour. A lost deal can contain an excellent call; a won deal can contain risky behaviour.
- Using one scorecard for discovery, demos, renewals and support. Different conversations have different jobs.
- Writing vague criteria such as “built rapport” without defining observable evidence.
- Giving every criterion equal weight before deciding which failures matter most.
- Treating N/A as zero and distorting the denominator.
- Using auto-fail for ordinary coaching gaps.
- Showing a total without criterion evidence.
- Changing criteria without versioning the scorecard.
- Comparing reviewers without first checking whether they evaluated the same call scope.
Sales call scorecard FAQs
What is a sales call scorecard?
A sales call scorecard is a structured rubric that evaluates observable behaviours in a sales conversation against defined criteria. It should retain evidence for each rating so the result can support review and coaching.
How many criteria should a sales call scorecard have?
There is no universal correct number. This 20-criterion template is broad enough for a full sales conversation, but a stage-specific scorecard may need fewer criteria. Remove items that do not apply instead of forcing every call into the same form.
Should every scorecard criterion have the same weight?
Not necessarily. Equal weights are simple, but weighted scoring is more appropriate when some behaviours have greater operational, customer or compliance consequences. Document the reason for each weight and test it against real reviewed calls.
How do you calculate a score when a criterion is N/A?
Exclude the N/A criterion’s weight from both earned points and available points, then divide earned points by the remaining applicable weight. Do not score N/A as zero.
What is an auto-fail criterion?
An auto-fail criterion is an approved rule that changes the pass status when a specified high-consequence failure occurs. Keep the raw score, evidence and triggered rule visible, and route ambiguous or consequential cases to human review.
Can AI use this scorecard?
Yes, if the criteria are observable, the scorecard is machine-readable and every result includes evidence and review information. Automated coverage does not remove the need for calibration, confidence handling and human oversight.
How often should a sales call scorecard be recalibrated?
Recalibrate after material changes to the script, product, policy, customer segment or call type, and whenever reviewer disagreement or score drift increases. A fixed calendar cadence can help, but evidence of change should trigger calibration sooner.
Build coaching around evidence
A scorecard is useful when it creates a shared definition of good, preserves the evidence behind each rating and leads to a specific behaviour change. If you are evaluating how structured call evidence can support a repeatable coaching system, explore CallOptix coaching and enablement.
Related Articles
Get the Latest Insights
Get the latest insights on call center optimization and AI-powered sales strategies delivered to your inbox.
By subscribing you agree to receive marketing emails. Unsubscribe anytime.


