For the complete documentation index, see llms.txt. This page is also available as Markdown.

AI Customer Service Quality Management

Target Audience: Customer service managers, quality management personnel, customer service trainers

1. Quick Start: Three Quality Metrics for AI Customer Service

How to View Evaluation Report Scores

Path: AgentOps (sidebar) → AI Assistant Monitoring

In the table, you can directly view the three major scoring metrics for each conversation. Click "View" to see full details.

Why Is Evaluation Needed?

Just like reviewing call recordings of customer service agents, we also need to check the quality of AI responses. The system automatically scores each conversation, helping you quickly identify issues.


Three Core Metrics

Metric
Plain Language Explanation
Scoring Criteria

Faithfulness Score

Is the information provided by the AI correct? Does it make things up or hallucinate?

85+ ✅ 60-84 ⚠️ Below 60 ❌

Answer Relevancy Score

Did the AI answer the customer's actual question?

85+ ✅ 60-84 ⚠️ Below 60 ❌

Context Precision Score

Did the AI find the correct reference materials and respond precisely to the context?

85+ ✅ 60-84 ⚠️ Below 60 ❌


Quick Assessment Method

All three metrics > 80  → ✅ This response is good
Any metric < 60         → ❌ Immediate improvement needed
Two or more < 70        → ⚠️ Systemic issue, comprehensive review needed

2. How to Read Evaluation Reports

Report Example


Three Common Issue Types

Issue A: Low Faithfulness Score (< 60)

Symptoms: The information from the AI is incorrect or fabricated

Common Causes:

  • Reference data is outdated (prices, inventory, policies have been updated)

  • Contradictory data (different documents say different things)

  • AI "guesses" the answer instead of relying on database content

Impact: Customers may receive incorrect information, leading to complaints


Issue B: Low Answer Relevancy Score (< 60)

Symptoms: The AI did not answer what the customer actually asked

Common Causes:

  • AI provided a lot of information but missed the key point

  • Off-topic response with irrelevant content

  • Only provided background without giving an actual answer

Impact: Customers need to ask again, reducing satisfaction


Issue C: Low Context Precision Score (< 60)

Symptoms: The AI referenced the wrong materials or was not precise enough

Common Causes:

  • Inaccurate keyword search

  • Confused data from different products/categories

  • Unclear titles or categories in the reference materials

Impact: Even if the AI intends to answer well, using the wrong data leads to wrong answers


3. Real-World Cases: Common Issues in the Apparel Industry

Case 1: Wrong Price (Low Faithfulness Score)

Problem Identified

Customer Question: "How much is this down jacket?"

AI Response: "This down jacket is priced at NT$ 3,990."

Evaluation Results:

  • Faithfulness Score (faithfulness_score): 38

  • Answer Relevancy Score (answer_relevancy_score): 95 ✅

  • Context Precision Score (context_precision_score): 85 ✅

Actual Situation: The product had been repriced to NT$ 2,990 (after discount), but the AI still quoted the original price.


Improvement Plan

Step 1: Check the Data

  • Open the "Product Price List" in the Knowledge Base

  • Found that the price had not been updated

Step 2: Update the Data

Problematic Data

Corrected Data

Step 3: Instruct the AI

Add the following to "AI Response Guidelines":

"When answering price questions, always confirm that the price is the latest one and state the applicable period. If there is a promotional discount, provide both the original price and the discounted price."


Case 2: Off-Topic Response (Low Answer Relevancy Score)

Problem Identified

Customer Question: "I bought something that doesn't fit. Can I return it?"

AI Response: "We offer a variety of sizes, including S, M, L, and XL. Each garment has a size chart on the product page. We recommend checking the size chart before purchasing."

Evaluation Results:

  • Faithfulness Score (faithfulness_score): 90 ✅

  • Answer Relevancy Score (answer_relevancy_score): 35

  • Context Precision Score (context_precision_score): 60 ⚠️

Diagnosis: The customer asked "Can I return it?" but the AI talked about "How to choose a size," completely missing the core question.


Improvement Plan

Step 1: Adjust Response Guidelines

Add the following to "AI Response Guidelines":

Core Principle: Answer the customer's core question first

Bad example:

  • Customer asks "Can I return it?"

  • AI answers "We have comprehensive after-sales service..." ❌ (Didn't say yes or no)

Good example:

  • Customer asks "Can I return it?"

  • AI answers "Yes. Within 7 days of delivery, if the item is unused, you can apply for a return..." ✅ (Direct answer)

Step 2: Restructure the Data

Problematic Data (too unfocused)

Corrected Data

Key Improvements:

  • ✅ Starts with a direct "Yes" or "No" answer

  • ✅ Uses Q&A format for clarity at a glance

  • ✅ Lists clear conditions to avoid disputes


Case 3: Wrong Product Referenced (Low Context Precision Score)

Problem Identified

Customer Question: "What material is the black knit top made of?"

AI Response: "This knit top is made of 100% pure cotton, soft and comfortable, suitable for all seasons."

Evaluation Results:

  • Faithfulness Score (faithfulness_score): 88 ✅

  • Answer Relevancy Score (answer_relevancy_score): 90 ✅

  • Context Precision Score (context_precision_score): 48

Actual Situation: The black knit top is 70% wool + 30% polyester. The AI referenced the "white knit top" data (100% pure cotton) instead.


Improvement Plan

Step 1: Check Data Labels

Problematic file name:

Issue: All knit tops are in the same document, making it difficult for the AI to distinguish between them.

Step 2: Improve Data Structure

Solution A: Separate Files

Solution B: Clear Headings

Step 3: Instruct the AI

Add the following to "AI Response Guidelines":

"When a customer mentions a product's color or model number, always verify that the reference data corresponds to the correct color and model. Different colors of the same product may have different materials and specifications."


4. Three-Step Improvement Plan

When you identify issues, follow this process:


Step 1: Update Data Content

When to use:

  • ✅ Low Faithfulness Score (data is incorrect or outdated)

  • ✅ Low Context Precision Score (data is disorganized or poorly labeled)

Checklist:

Data Quality Examples:

Poor Data

Good Data


Step 2: Adjust AI Response Guidelines

When to use:

  • ✅ Low Answer Relevancy Score (off-topic responses)

  • ✅ Low Faithfulness Score (AI guessing or hallucinating)

AI Response Guidelines Template:


Step 3: Escalate to the Technical Team

When to use:

  • Context Precision Score is consistently low

  • The same issue occurs repeatedly

  • No improvement after adjusting data and guidelines

Escalation Content:


5. Daily Management Checklist

Daily Checks

When issues are found:


Response Quality Tracking

1. Data Review

2. Issue Analysis

3. Improvement Actions


Appendix A: Issue Diagnosis Quick Reference

Score Status
Possible Cause
Improvement Method

Low Faithfulness Score

Outdated or incorrect data, AI hallucination

Step 1: Update data content

Low Answer Relevancy Score

AI gives off-topic responses

Step 2: Adjust response guidelines

Low Context Precision Score

AI references wrong data or lacks precision

Step 1: Improve data labeling

Multiple metrics are low

Systemic issue

Steps 1+2, Step 3 if necessary


Improvement Priority Order


Appendix B: System Evaluation Metrics Reference

Primary Metrics (No Ground Truth Required)

These three metrics are the core of this guide and can be directly applied to daily customer service conversation evaluations:

Name
English Full Name
Description

Faithfulness Score

Faithfulness

Evaluates whether the AI response aligns with database content, or if it hallucinates or fabricates information

Answer Relevancy Score

Answer Relevancy

Evaluates whether the AI response is relevant to the customer's question, or if it's off-topic

Context Precision Score

Context Precision

Evaluates whether the AI response precisely addresses the context and references the correct materials

Advanced Metrics (Ground Truth Required)

The following metrics require pre-prepared "ground truth" answers and are suitable for test case evaluations:

Name
English Full Name
Description

Answer Correctness

Answer Correctness

Compares AI response against the ground truth to evaluate correctness

Answer Similarity

Answer Similarity

Evaluates semantic similarity between the AI response and the ground truth

Context Recall

Context Recall

Evaluates whether the system retrieved all necessary reference materials

Other Available Metrics (DeepEval)

The system also supports the following additional evaluation metrics for more comprehensive quality checks:

Name
English Name
Description

Bias Detection

Bias

Detects whether responses contain biased or discriminatory content

Toxicity Detection

Toxicity

Detects whether responses contain inappropriate or offensive content

Hallucination Detection

Hallucination

Detects whether the AI generates content that contradicts facts

Contextual Relevancy

Contextual Relevancy

Evaluates whether retrieved reference materials are relevant to the question

Usage Recommendations

  1. Daily Monitoring: Use the three primary metrics (Faithfulness Score, Answer Relevancy Score, Context Precision Score)

  2. Test Evaluations: Use advanced metrics with prepared ground truth answers for systematic evaluations

  3. Quality Assurance: Enable bias and toxicity detection to ensure responses comply with corporate standards


FAQ

Q1: I'm not technical — can I still manage AI customer service? A: Absolutely! Just like managing customer service staff, you only need to:

  • Review evaluation reports daily to identify problematic conversations

  • Check whether data is correct and complete

  • Adjust the AI's "response guidelines" (just like training customer service scripts)


Q2: How are the scores calculated? Does the AI grade itself? A: No. The scoring is performed automatically by a dedicated "evaluation system," like having another AI serve as "quality control" to check the first AI's responses.


Q3: Are all three metrics equally important? Can I just look at one? A: We recommend looking at all three, as they reflect different issues:

  • Faithfulness Score (faithfulness_score): Whether the AI aligns with database content, whether it hallucinated

  • Answer Relevancy Score (answer_relevancy_score): Whether the AI understood the question and gave a relevant response

  • Context Precision Score (context_precision_score): Whether the AI precisely addressed the context and found the right reference materials

Looking at only one may cause you to miss important issues.


Q4: How long until I see improvement after making changes? A:

  • Data updates: Take effect immediately (improvements visible the same day)

  • Response guideline adjustments: Take effect immediately

  • Technical adjustments: Require 2-4 weeks (depending on complexity)


Conclusion

Managing AI customer service is just like managing a real customer service team:

Regularly check quality (review evaluation reports) ✅ Continuously update knowledge (update data content) ✅ Optimize response scripts (adjust response guidelines) ✅ Track improvement results (monitor score changes)

By following this guide, you can continuously improve your AI customer service — even without technical expertise!

Last updated

Was this helpful?