AI Customer Service Quality Management
Target Audience: Customer service managers, quality management personnel, customer service trainers
1. Quick Start: Three Quality Metrics for AI Customer Service
How to View Evaluation Report Scores
Path: AgentOps (sidebar) → AI Assistant Monitoring
In the table, you can directly view the three major scoring metrics for each conversation. Click "View" to see full details.
Why Is Evaluation Needed?
Just like reviewing call recordings of customer service agents, we also need to check the quality of AI responses. The system automatically scores each conversation, helping you quickly identify issues.
Three Core Metrics
Faithfulness Score
Is the information provided by the AI correct? Does it make things up or hallucinate?
85+ ✅ 60-84 ⚠️ Below 60 ❌
Answer Relevancy Score
Did the AI answer the customer's actual question?
85+ ✅ 60-84 ⚠️ Below 60 ❌
Context Precision Score
Did the AI find the correct reference materials and respond precisely to the context?
85+ ✅ 60-84 ⚠️ Below 60 ❌
Quick Assessment Method
All three metrics > 80 → ✅ This response is good
Any metric < 60 → ❌ Immediate improvement needed
Two or more < 70 → ⚠️ Systemic issue, comprehensive review needed2. How to Read Evaluation Reports
Report Example
Three Common Issue Types
Issue A: Low Faithfulness Score (< 60)
Symptoms: The information from the AI is incorrect or fabricated
Common Causes:
Reference data is outdated (prices, inventory, policies have been updated)
Contradictory data (different documents say different things)
AI "guesses" the answer instead of relying on database content
Impact: Customers may receive incorrect information, leading to complaints
Issue B: Low Answer Relevancy Score (< 60)
Symptoms: The AI did not answer what the customer actually asked
Common Causes:
AI provided a lot of information but missed the key point
Off-topic response with irrelevant content
Only provided background without giving an actual answer
Impact: Customers need to ask again, reducing satisfaction
Issue C: Low Context Precision Score (< 60)
Symptoms: The AI referenced the wrong materials or was not precise enough
Common Causes:
Inaccurate keyword search
Confused data from different products/categories
Unclear titles or categories in the reference materials
Impact: Even if the AI intends to answer well, using the wrong data leads to wrong answers
3. Real-World Cases: Common Issues in the Apparel Industry
Case 1: Wrong Price (Low Faithfulness Score)
Problem Identified
Customer Question: "How much is this down jacket?"
AI Response: "This down jacket is priced at NT$ 3,990."
Evaluation Results:
Faithfulness Score (faithfulness_score): 38 ❌
Answer Relevancy Score (answer_relevancy_score): 95 ✅
Context Precision Score (context_precision_score): 85 ✅
Actual Situation: The product had been repriced to NT$ 2,990 (after discount), but the AI still quoted the original price.
Improvement Plan
Step 1: Check the Data
Open the "Product Price List" in the Knowledge Base
Found that the price had not been updated
Step 2: Update the Data
❌ Problematic Data
✅ Corrected Data
Step 3: Instruct the AI
Add the following to "AI Response Guidelines":
"When answering price questions, always confirm that the price is the latest one and state the applicable period. If there is a promotional discount, provide both the original price and the discounted price."
Case 2: Off-Topic Response (Low Answer Relevancy Score)
Problem Identified
Customer Question: "I bought something that doesn't fit. Can I return it?"
AI Response: "We offer a variety of sizes, including S, M, L, and XL. Each garment has a size chart on the product page. We recommend checking the size chart before purchasing."
Evaluation Results:
Faithfulness Score (faithfulness_score): 90 ✅
Answer Relevancy Score (answer_relevancy_score): 35 ❌
Context Precision Score (context_precision_score): 60 ⚠️
Diagnosis: The customer asked "Can I return it?" but the AI talked about "How to choose a size," completely missing the core question.
Improvement Plan
Step 1: Adjust Response Guidelines
Add the following to "AI Response Guidelines":
Core Principle: Answer the customer's core question first
Bad example:
Customer asks "Can I return it?"
AI answers "We have comprehensive after-sales service..." ❌ (Didn't say yes or no)
Good example:
Customer asks "Can I return it?"
AI answers "Yes. Within 7 days of delivery, if the item is unused, you can apply for a return..." ✅ (Direct answer)
Step 2: Restructure the Data
❌ Problematic Data (too unfocused)
✅ Corrected Data
Key Improvements:
✅ Starts with a direct "Yes" or "No" answer
✅ Uses Q&A format for clarity at a glance
✅ Lists clear conditions to avoid disputes
Case 3: Wrong Product Referenced (Low Context Precision Score)
Problem Identified
Customer Question: "What material is the black knit top made of?"
AI Response: "This knit top is made of 100% pure cotton, soft and comfortable, suitable for all seasons."
Evaluation Results:
Faithfulness Score (faithfulness_score): 88 ✅
Answer Relevancy Score (answer_relevancy_score): 90 ✅
Context Precision Score (context_precision_score): 48 ❌
Actual Situation: The black knit top is 70% wool + 30% polyester. The AI referenced the "white knit top" data (100% pure cotton) instead.
Improvement Plan
Step 1: Check Data Labels
Problematic file name:
Issue: All knit tops are in the same document, making it difficult for the AI to distinguish between them.
Step 2: Improve Data Structure
✅ Solution A: Separate Files
✅ Solution B: Clear Headings
Step 3: Instruct the AI
Add the following to "AI Response Guidelines":
"When a customer mentions a product's color or model number, always verify that the reference data corresponds to the correct color and model. Different colors of the same product may have different materials and specifications."
4. Three-Step Improvement Plan
When you identify issues, follow this process:
Step 1: Update Data Content
When to use:
✅ Low Faithfulness Score (data is incorrect or outdated)
✅ Low Context Precision Score (data is disorganized or poorly labeled)
Checklist:
Data Quality Examples:
❌ Poor Data
✅ Good Data
Step 2: Adjust AI Response Guidelines
When to use:
✅ Low Answer Relevancy Score (off-topic responses)
✅ Low Faithfulness Score (AI guessing or hallucinating)
AI Response Guidelines Template:
Step 3: Escalate to the Technical Team
When to use:
Context Precision Score is consistently low
The same issue occurs repeatedly
No improvement after adjusting data and guidelines
Escalation Content:
5. Daily Management Checklist
Daily Checks
When issues are found:
Response Quality Tracking
1. Data Review
2. Issue Analysis
3. Improvement Actions
Appendix A: Issue Diagnosis Quick Reference
Low Faithfulness Score
Outdated or incorrect data, AI hallucination
Step 1: Update data content
Low Answer Relevancy Score
AI gives off-topic responses
Step 2: Adjust response guidelines
Low Context Precision Score
AI references wrong data or lacks precision
Step 1: Improve data labeling
Multiple metrics are low
Systemic issue
Steps 1+2, Step 3 if necessary
Improvement Priority Order
Appendix B: System Evaluation Metrics Reference
Primary Metrics (No Ground Truth Required)
These three metrics are the core of this guide and can be directly applied to daily customer service conversation evaluations:
Faithfulness Score
Faithfulness
Evaluates whether the AI response aligns with database content, or if it hallucinates or fabricates information
Answer Relevancy Score
Answer Relevancy
Evaluates whether the AI response is relevant to the customer's question, or if it's off-topic
Context Precision Score
Context Precision
Evaluates whether the AI response precisely addresses the context and references the correct materials
Advanced Metrics (Ground Truth Required)
The following metrics require pre-prepared "ground truth" answers and are suitable for test case evaluations:
Answer Correctness
Answer Correctness
Compares AI response against the ground truth to evaluate correctness
Answer Similarity
Answer Similarity
Evaluates semantic similarity between the AI response and the ground truth
Context Recall
Context Recall
Evaluates whether the system retrieved all necessary reference materials
Other Available Metrics (DeepEval)
The system also supports the following additional evaluation metrics for more comprehensive quality checks:
Bias Detection
Bias
Detects whether responses contain biased or discriminatory content
Toxicity Detection
Toxicity
Detects whether responses contain inappropriate or offensive content
Hallucination Detection
Hallucination
Detects whether the AI generates content that contradicts facts
Contextual Relevancy
Contextual Relevancy
Evaluates whether retrieved reference materials are relevant to the question
Usage Recommendations
Daily Monitoring: Use the three primary metrics (Faithfulness Score, Answer Relevancy Score, Context Precision Score)
Test Evaluations: Use advanced metrics with prepared ground truth answers for systematic evaluations
Quality Assurance: Enable bias and toxicity detection to ensure responses comply with corporate standards
FAQ
Q1: I'm not technical — can I still manage AI customer service? A: Absolutely! Just like managing customer service staff, you only need to:
Review evaluation reports daily to identify problematic conversations
Check whether data is correct and complete
Adjust the AI's "response guidelines" (just like training customer service scripts)
Q2: How are the scores calculated? Does the AI grade itself? A: No. The scoring is performed automatically by a dedicated "evaluation system," like having another AI serve as "quality control" to check the first AI's responses.
Q3: Are all three metrics equally important? Can I just look at one? A: We recommend looking at all three, as they reflect different issues:
Faithfulness Score (
faithfulness_score): Whether the AI aligns with database content, whether it hallucinatedAnswer Relevancy Score (
answer_relevancy_score): Whether the AI understood the question and gave a relevant responseContext Precision Score (
context_precision_score): Whether the AI precisely addressed the context and found the right reference materials
Looking at only one may cause you to miss important issues.
Q4: How long until I see improvement after making changes? A:
Data updates: Take effect immediately (improvements visible the same day)
Response guideline adjustments: Take effect immediately
Technical adjustments: Require 2-4 weeks (depending on complexity)
Conclusion
Managing AI customer service is just like managing a real customer service team:
✅ Regularly check quality (review evaluation reports) ✅ Continuously update knowledge (update data content) ✅ Optimize response scripts (adjust response guidelines) ✅ Track improvement results (monitor score changes)
By following this guide, you can continuously improve your AI customer service — even without technical expertise!
Last updated
Was this helpful?
