How to Score Hindi and Hinglish Sales Calls
A practical method to score Hindi and Hinglish sales calls using transcript checks, critical entities, evidence, abstention and human review.

Hindi and Hinglish sales calls should be scored in two linked layers: first verify that the transcript preserves the business meaning of the audio; then evaluate the call against observable sales criteria using timestamped evidence. Do not turn one transcription-confidence number into a call-quality score. Tag code-switching, protect critical entities, allow criteria to abstain and route uncertain or consequential findings to bilingual human reviewers.
This method is designed for sales and QA teams that need defensible coaching evidence across mixed-language calls. It does not assume a particular transcription vendor, model, dialect or CallOptix product capability.
What is a Hindi or Hinglish sales call?
A Hinglish sales call is a Hindi–English code-switched conversation that may move between Hindi and English within a turn or sentence and may represent Hindi in Devanagari, Roman script or both.
Code-switching means alternating between two or more languages within the same conversation or utterance. A Hindi–English speech-corpus paper describes the practical variation involved, including speakers, accents, pronunciations and recording sessions; those variables matter because a single clean studio sample cannot represent a live sales queue. Hindi–English Code-Switching Speech Corpus (2018)
A language-aware call-QA method is an evaluation process that keeps audio and transcription uncertainty separate from evidence about the agent's behaviour.
That separation is essential. A correct score needs more than fluent-looking text: it needs the right speaker, the right meaning, the right business-critical terms and evidence that genuinely supports the criterion.
Why does ordinary call scoring fail on mixed-language calls?
The failure usually begins before the scorecard. A transcript can look readable while changing a negation, assigning the customer's sentence to the agent, or converting a product or fee into a different term. The resulting score is then precise but unsupported.
Mixed-language calls add three operational complications:
• Language can switch at token level, so one language label for the whole call is too coarse.
• The same Hindi phrase can appear in Devanagari or Roman script without changing its meaning.
• Names, amounts, dates, product terms and local pronunciations can be harder than ordinary conversational words.
Research on code-switching speech recognition treats transcription and token-level language identification as connected tasks, which supports recording language at a finer level than a call-wide tag. Unified Model for Code-Switching Speech Recognition and Language Identification, ACL 2023
India's linguistic and demographic variety also makes local testing non-negotiable. AI4Bharat's speech-recognition programme explicitly discusses coverage across languages, regions and demographics and identifies domain and telephony adaptation as continuing work. AI4Bharat automatic speech recognition programme
The seven-stage method for scoring Hindi and Hinglish calls

1. Define the eligible call set
Start with a denominator, not a model. Define which calls should be scored: queue, date range, minimum duration, call direction, recording consent status and any valid exclusions. Record missing or corrupted audio separately. A call that never reached the transcription pipeline is not a low-quality call; it is a coverage failure.
2. Preserve audio and speaker turns
Keep the original audio as the evidence source. Separate agent and customer turns before scoring. If speaker attribution is uncertain, block criteria that depend on who said what. Do not infer ownership from sentence style.
3. Produce the transcript without hiding uncertainty
Store the transcript with timestamps and, where available, token or segment confidence. Retain the raw transcript. A cleaned version may help reviewers search and compare terms, but it must never overwrite what the system actually produced.
4. Tag language and script at segment level
Use at least four operational labels: Hindi–Devanagari, Hindi–Roman, English and mixed/uncertain. The label is metadata, not a quality judgement. A Romanized Hindi sentence is not automatically worse than the same sentence in Devanagari.
5. Normalize only for evaluation
Create a versioned glossary for product names, common Romanized variants, acronyms, prices, tenure units and local terms. Use it to compare equivalent forms—for example, a Romanized Hindi term and its Devanagari equivalent—while preserving the original transcript for audit.
6. Score observable criteria with timestamped evidence
Each criterion should contain: criterion ID, definition, applicability rule, evidence start and end times, transcript excerpt, evidence language, result, confidence and review status. If evidence is missing or unreliable, the correct result is not applicable or abstain—not an invented pass or fail.
7. Route by uncertainty and consequence
A low-consequence coaching cue with clear evidence can follow the normal workflow. A fee, interest rate, consent statement, payment instruction, customer identity or regulatory disclosure needs stricter routing. High consequence plus transcription uncertainty should always trigger bilingual human review.
Use a transcript error taxonomy, not one accuracy number

A transcript error matters according to what it changes. Treat these four classes separately:
Error class | Definition | Effect on scoring | Recommended action |
|---|---|---|---|
Script-only mismatch | The same meaning appears in a different script or acceptable Romanization. | Usually no penalty for behaviour scoring; may matter for display or search. | Accept for meaning-sensitive scoring; retain the original form. |
Meaning-changing word error | A substituted, missing or inserted word changes the proposition. | Can reverse whether a criterion passed. | Abstain on affected criteria and review the audio. |
Critical-entity error | A price, date, rate, tenure, product, name or negation is wrong or missing. | Can alter the commercial or compliance meaning. | Route to human review and track separately. |
Speaker-attribution error | Words are assigned to the wrong speaker. | Can falsely credit or penalize the agent. | Block speaker-dependent criteria until corrected. |
A critical entity is a word or phrase whose transcription can change the business outcome or the interpretation of a scored criterion. Typical examples include amounts, dates, durations, product names, customer names, payment instructions and negations.
A script mismatch is a representational difference between equivalent lexical content; it is not necessarily a meaning error.
How should transcription quality be measured?
Use word error rate, but define its job
Word error rate (WER) measures the edit operations needed to turn a system transcript into a human reference transcript:
WER = (substitutions + deletions + insertions) ÷ reference words × 100
NIST's OpenASR21 evaluation plan uses this standard calculation. WER is useful for reproducible system comparison, but it weights every reference word equally and does not tell you whether an error changed a fee, a denial or the identity of the speaker. NIST OpenASR21 Evaluation Plan
Report WER by meaningful slices instead of averaging away risk: dominant language, code-switch density, script, agent/customer channel, call type, audio condition and business domain. Do not publish a universal acceptable WER threshold; set criteria-specific tolerances from the consequence of a wrong decision.
Add a script-normalized diagnostic
A June 2026 preprint proposes Script-Normalized WER (SN-WER) for multilingual transcripts where the same lexical content may appear in multiple scripts. The authors present it as a companion diagnostic—not a replacement for standard WER or character error rate when script fidelity matters. That distinction fits Hinglish QA: use script-normalized comparison for meaning-sensitive scoring, while retaining conventional metrics for transcript display, search and script-specific workflows. Script-Normalized Word Error Rate preprint, June 2026
Track critical-entity errors separately
Critical Entity Error Rate = (incorrect critical entities + missing critical entities) ÷ reference critical entities × 100
Define the entity list before evaluation and version it with the scorecard. The numerator must include meaning-changing substitutions and omissions; harmless formatting differences should be governed by documented normalization rules.
Track evidence coverage
Evidence coverage = criteria with valid timestamped evidence ÷ applicable criteria × 100
This metric answers a different question from WER. A transcript can have moderate word-level errors yet still preserve every fact needed for a discovery criterion; another can look excellent while corrupting the only fee mentioned. Evidence coverage makes that distinction visible.
Build a language-neutral sales scorecard
The scorecard should evaluate behaviours, not accents, script choice or the amount of English used. Start with observable criteria such as:
Criterion | Observable evidence | Abstain when |
|---|---|---|
Opening and context | Agent identifies the purpose and establishes the reason for the call. | The opening is missing from the recording. |
Need discovery | Agent asks questions that reveal use case, timing, constraints or preference. | Speaker turns cannot be assigned reliably. |
Recommendation fit | Recommendation is tied to a stated customer need. | The product term or need statement is uncertain. |
Commercial clarity | Amount, duration, condition and mandatory charge are stated consistently. | A critical number, unit or negation is uncertain. |
Objection response | Agent acknowledges the concern and answers the actual objection. | The objection or response meaning cannot be recovered. |
Next step | Agent confirms an action, owner and timing. | Date, channel or commitment is unclear. |
Use the same criterion IDs across languages so coaching reports remain comparable. Localize examples, glossaries and reviewer guidance instead of building an unrelated scorecard for each language.
For a complete weighted structure, see the sales call scorecard template. For the distinction between scoring the conversation and measuring call operations, see conversation intelligence vs call analytics.
Worked example: when the system must abstain
Illustrative training example: the agent says, “Sir, processing fee one percent hai, zero nahi.” The transcript records, “processing fee zero percent hai.” Most words are present, but the negation and commercial meaning are wrong.
A safe workflow produces this result:
• Transcript finding: critical-entity and meaning-changing error.
• Affected criterion: commercial clarity.
• Automated result: abstain; do not mark pass or fail.
• Evidence: timestamped audio segment and raw transcript.
• Routing: bilingual reviewer listens to the audio and records the corrected evidence.
• Audit record: original output, correction, reviewer decision and glossary change, if needed.
Now consider a second illustrative example: “kal main purchase karunga” appears as an equivalent Romanized form instead of Devanagari. If the meaning and speaker are preserved, a behaviour score should not be reduced merely because the script differs. The display transcript may still be flagged if the product requires a particular script.
How should confidence and human review work?
Confidence should control routing, not masquerade as truth. Treat model confidence, evidence confidence and decision consequence as separate fields.
Evidence confidence | Decision consequence | Route |
|---|---|---|
High | Low | Accept the evidence; include in routine calibration sampling. |
Low or conflicting | Low | Abstain on the affected criterion; send to reviewer when coaching depends on it. |
High | High | Apply a second check or policy-defined review before consequential use. |
Low or conflicting | High | Mandatory bilingual human review; no automated adverse decision. |
Reviewers need the audio at the relevant timestamp, the raw and normalized transcript, speaker labels, language/script tags, the criterion definition and the reason the system abstained. A reviewer should be able to disagree without editing history.
For a broader implementation pattern covering eligible-call denominators, exceptions, overrides and drift, use the AI call quality assurance guide. The AI call scoring guide explains how criterion evidence becomes a score.
A practical calibration plan
• Build a test set that represents the queue: language mix, Roman and Devanagari script, accents, regions, call types, devices and noisy conditions.
• Have two bilingual reviewers independently transcribe and score a starting batch using the same written rules.
• Resolve disagreements by category: transcription, speaker attribution, language tag, applicability, evidence or scoring judgement.
• Calculate WER and critical-entity error by slice; do not rely on one overall average.
• Measure reviewer agreement for each criterion and rewrite ambiguous definitions before changing model thresholds.
• Set routing rules from consequence: critical commercial or compliance facts receive more protection than low-stakes coaching cues.
• Version the glossary, criteria, normalization rules and model configuration so score changes can be explained.
• Repeat calibration after material changes in campaign, offer, region, audio channel, script mix or model.
There is no defensible universal sample size or confidence threshold for every Hindi and Hinglish sales operation. Continue sampling until each material segment has enough reviewed examples to expose recurring error patterns, and expand the set when the production mix changes.
Common mistakes to avoid
• Using one language label for an entire code-switched call.
• Treating Romanized Hindi as automatically incorrect.
• Cleaning the transcript and discarding the raw output.
• Using overall WER as a proxy for score reliability.
• Scoring a critical criterion when its evidence is uncertain.
• Penalizing accent, dialect or language choice instead of observable sales behaviour.
• Letting reviewers correct scores without preserving the original output and reason.
• Claiming 100% automated scoring when uncertain criteria are silently forced into pass or fail.
Frequently asked questions
Is Hinglish a separate language for call QA?
Operationally, treat Hinglish as code-switched Hindi and English rather than forcing the whole call into one language bucket. Segment-level language and script tags produce more useful diagnostics.
Can an English sales scorecard simply be translated into Hindi?
The behaviour criteria can remain consistent, but examples, critical terms, applicability rules and reviewer guidance must be tested with real Hindi and Hinglish call patterns. Literal translation alone does not calibrate evidence.
Is word error rate enough to approve automated scoring?
No. WER measures transcript edit distance, not whether the right speaker said a business-critical fact. Combine it with critical-entity errors, speaker attribution, evidence coverage and criterion-level review.
Should Romanized Hindi count as a transcription error?
It depends on the task. For meaning-sensitive call scoring, an equivalent Romanized form may be acceptable. For script-specific display, search or audit needs, evaluate it separately and retain conventional WER or character-level measures.
How many calls are needed to calibrate Hindi and Hinglish QA?
There is no universal number. Build a stratified set that represents material language, script, accent, channel, campaign and call-type segments, then keep expanding it until recurring error classes and reviewer disagreements stabilize.
When should a bilingual reviewer listen to the audio?
Require review when a critical entity or negation is uncertain, speakers may be swapped, evidence conflicts, a complaint or consequential decision is involved, or the model abstains on a criterion needed for action.
Can sentiment be used to score Hinglish calls?
Sentiment can be a review cue, but it should not be the sole basis for a performance score. Meaning, politeness and intensity vary with context, dialect and code-switching, so consequential findings need observable evidence and calibration.
Make multilingual QA auditable
A defensible Hindi and Hinglish QA programme does not begin by demanding a score from every transcript. It begins by preserving the audio, naming uncertainty, protecting critical entities and requiring evidence for each criterion. That creates coaching data people can inspect—and a clear path for human review when the system should not decide.
Planning a multilingual call-QA rubric? Talk to CallOptix about the criteria, evidence and review decisions your team needs.
Related Articles
Get the Latest Insights
Get the latest insights on call center optimization and AI-powered sales strategies delivered to your inbox.
By subscribing you agree to receive marketing emails. Unsubscribe anytime.



