> For the complete documentation index, see [llms.txt](https://docs.maiagent.ai/llms.txt). Markdown versions of documentation pages are available by appending `.md` to page URLs; this page is available as [Markdown](https://docs.maiagent.ai/maiagent-user-guide/en/agent-ops/evaluations.md).

# Automated Evaluation & AI Assistant Monitoring

Automate testing and real-time monitoring of AI Assistant response quality, performance metrics, and costs

## Automated Evaluation <a href="#auto-evaluation" id="auto-evaluation"></a>

The automated evaluation feature allows you to use pre-built test datasets to automatically test AI Assistant response quality. The system sends test questions to the AI Assistant, compares actual responses with expected responses, and generates detailed evaluation reports.

### Access Automated Evaluation <a href="#access-auto-evaluation" id="access-auto-evaluation"></a>

Go to "<mark style="color:blue;">AgentOps</mark>" in the left sidebar, then click "<mark style="color:blue;">Automated Testing</mark>".

<figure><img src="https://1360999650-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2F6v6TNkkOQVfRYfcNirHL%2Fuploads%2Fgit-blob-06f0a8f71a921e02e030b8b8e71bc49977fc7154%2Fagentops-evaluations.png?alt=media" alt="Automated testing list"><figcaption><p>Automated testing list showing success rate and average response time for each test run</p></figcaption></figure>

The page displays all evaluation records, including evaluation name, test dataset, AI Assistant, success rate, average response time, and creation time.

### Create and Run a Test <a href="#create-and-run-test" id="create-and-run-test"></a>

1. Ensure you have created a test dataset (refer to [Test Dataset Management](/maiagent-user-guide/en/agent-ops/test-datasets.md))
2. Click the "<mark style="color:blue;">Create Test</mark>" button
3. Fill in the evaluation name and description, then select the test dataset and AI Assistant
4. Click "<mark style="color:blue;">Start Evaluation</mark>" and the system will automatically execute all test cases

{% hint style="info" %}
Evaluation execution time depends on the number of test cases. Typically, 50 test cases take approximately 2-3 minutes.
{% endhint %}

### Use Custom Evaluation Metrics <a href="#custom-evaluation-metrics" id="custom-evaluation-metrics"></a>

In addition to built-in metrics, you can define your team's own scoring criteria using natural language. For example, Mai, a customer service manager, wants to verify whether the customer service assistant remains polite and empathetic. She creates a "Customer Service Tone" metric, enters the evaluation criteria, and sets a passing threshold. After the evaluation is complete, she can view each case's "Customer Service Tone" score in the results and identify responses that need adjustment.

{% stepper %}
{% step %}
When creating an evaluation, complete the basic settings, including the name, test dataset, AI Assistant, and evaluation model.
{% endstep %}

{% step %}
In the "<mark style="color:blue;">Custom Metrics</mark>" section, click "<mark style="color:blue;">Add Custom Metric</mark>", then enter an easily identifiable metric name and evaluation criteria.
{% endstep %}

{% step %}
Adjust the passing threshold according to your team's quality standards, then click "<mark style="color:blue;">Confirm</mark>" to start the evaluation.
{% endstep %}
{% endstepper %}

<figure><img src="https://1360999650-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2F6v6TNkkOQVfRYfcNirHL%2Fuploads%2Fgit-blob-43ed547b1eec12101661f11efa05ddfa646c6a4d%2Fevaluation-multi-turn-settings.png?alt=media" alt="Multi-turn conversation settings in the create evaluation dialog"><figcaption><p>After switching to multi-turn conversations, you can set the maximum number of turns, three evaluation metrics, and the simulation language</p></figcaption></figure>

<figure><img src="https://1360999650-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2F6v6TNkkOQVfRYfcNirHL%2Fuploads%2Fgit-blob-d56c6a2a2384f1f8aab3561aa424ae4364d3f37e%2Fevaluation-custom-metric.png?alt=media" alt="Custom metric settings in the create evaluation dialog"><figcaption><p>Set evaluation criteria for customer service tone using natural language</p></figcaption></figure>

{% hint style="info" %}
The system compares names without regard to letter case. A custom metric cannot duplicate any of the 11 built-in metrics. If a name conflicts, use a name that expresses your team's evaluation objective.
{% endhint %}

### Evaluate Multi-turn Conversations <a href="#multi-turn-evaluation" id="multi-turn-evaluation"></a>

Multi-turn evaluation is ideal for testing tasks that require back-and-forth confirmation. For example, Mei, a quality manager, wants the assistant to ask for the order number first, confirm the reason for the return, and then explain the next steps. She creates a multi-turn test case, enters the scenario and expected result, and runs the evaluation. When it is complete, she can expand the conversation transcript to review every response and the Knowledge Base content retrieved during each turn, instead of looking only at the final answer.

{% stepper %}
{% step %}
First, go to "<mark style="color:blue;">Test Datasets</mark>" and open the target test dataset. Add a scenario, expected result, and simulated user persona to the multi-turn test case.
{% endstep %}

{% step %}
Return to "<mark style="color:blue;">Automated Testing</mark>", click "<mark style="color:blue;">Create Evaluation</mark>", and switch to "<mark style="color:blue;">Multi-turn Conversation</mark>".
{% endstep %}

{% step %}
Select the test dataset, AI Assistant, and evaluation model. Set the maximum number of turns, evaluation metrics, and simulation language, then start the evaluation. At least one evaluation metric must be enabled.
{% endstep %}

{% step %}
After the evaluation is complete, open the results and select an individual case to expand its details. Review the metric scores, conversation transcript, and "<mark style="color:blue;">Retrieved Content</mark>" for each turn. If an individual case fails, its error is displayed separately without hiding the results of other cases.
{% endstep %}
{% endstepper %}

<figure><img src="https://1360999650-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2F6v6TNkkOQVfRYfcNirHL%2Fuploads%2Fgit-blob-123f84448a9ae205a2a7adcc1f96b8a63ceaf02d%2Fevaluation-multi-turn-results.png?alt=media" alt="Metrics, transcript, and retrieved content shown after expanding a completed multi-turn evaluation case"><figcaption><p>Expand a case to review the assistant's response and the Knowledge Base content used in each turn</p></figcaption></figure>

If you no longer need an ongoing multi-turn evaluation, click "<mark style="color:blue;">Cancel Evaluation</mark>" in the evaluation details and confirm. The system stops processing subsequent cases after the current case is complete. Scores and transcripts for completed cases are retained, the status is shown as "<mark style="color:blue;">Partially Completed</mark>", and unprocessed cases are shown as "<mark style="color:blue;">Pending Evaluation</mark>".

<figure><img src="https://1360999650-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2F6v6TNkkOQVfRYfcNirHL%2Fuploads%2Fgit-blob-926b7035060aac6e77573489d86fb8297fab864f%2Fevaluation-multi-turn-cancel.png?alt=media" alt="Partially completed status after canceling a multi-turn evaluation"><figcaption><p>Completed cases are retained after cancellation, while unprocessed cases remain pending evaluation</p></figcaption></figure>

### View Evaluation Results <a href="#view-evaluation-results" id="view-evaluation-results"></a>

Click on an evaluation record to view the detailed report:

<figure><img src="https://1360999650-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2F6v6TNkkOQVfRYfcNirHL%2Fuploads%2Fgit-blob-aeec25455ae679a86ed97bc0ed49d20cb37e8b24%2Ftrack-evaluation-detail.png?alt=media" alt="Evaluation details"><figcaption><p>Test details page: success rate, AI insights summary, metric statistics, and improvement suggestions</p></figcaption></figure>

The detailed report includes:

* **Success Rate**: Percentage of test cases that passed
* **AI Insights**: Automatically generated evaluation summary and improvement suggestions
* **Metric Statistics**: Quality scores, response relevance, and other metrics
* **Test Case Details**: Expected response, actual response, score, and status for each question

**Success Rate Reference Benchmarks:**

| AI Assistant Type              | Recommended Success Rate |
| ------------------------------ | ------------------------ |
| Product Inquiry Assistant      | ≥ 95%                    |
| Customer Support Assistant     | ≥ 90%                    |
| General Conversation Assistant | ≥ 80%                    |

{% hint style="info" %}
For more information about insight reports, refer to: [Evaluation Insight Reports](/maiagent-user-guide/en/agent-ops/evaluation-insights.md)
{% endhint %}

### Manage Evaluation Records <a href="#manage-evaluation-records" id="manage-evaluation-records"></a>

* **Search and Filter**: Filter records by test dataset, AI Assistant, or keywords
* **Re-run**: After modifying the Knowledge Base or AI settings, re-run tests to verify improvements
* **Export**: Export evaluation results in Excel format

#### Edit the Evaluation Name and Description <a href="#edit-evaluation-name-description" id="edit-evaluation-name-description"></a>

After a test is complete, you can still edit its evaluation name and description without deleting or rerunning the test. The test dataset, AI Assistant, evaluation results, and execution status remain unchanged.

For example, after Mai created a weekly customer service quality test, she noticed that the week number in the evaluation name was incorrect. She opened the edit dialog from the automated testing list, corrected the name, and added the test objective. After she saved the changes, the list immediately displayed the updated information while retaining the original test results.

{% stepper %}
{% step %}

### Open the Automated Testing List

From the left menu, go to “<mark style="color:blue;">AgentOps</mark>” → “<mark style="color:blue;">Automated Testing</mark>”.
{% endstep %}

{% step %}

### Open the Edit Dialog

Find the evaluation you want to update, then click the “<mark style="color:blue;">Edit</mark>” icon in the “<mark style="color:blue;">Actions</mark>” column for that row.
{% endstep %}

{% step %}

### Edit and Save

Edit the “<mark style="color:blue;">Evaluation Name</mark>” or “<mark style="color:blue;">Description</mark>”, then click “<mark style="color:blue;">Confirm</mark>”. The evaluation name is required and can contain up to 200 characters; the description may be left blank.
{% endstep %}
{% endstepper %}

<figure><img src="https://1360999650-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2F6v6TNkkOQVfRYfcNirHL%2Fuploads%2Fgit-blob-574e934d25a9ff140f8191df9c4bef52ccbdf9a2%2Fagentops-edit-evaluation.png?alt=media" alt="Edit test dialog in the automated testing list, where you can update the evaluation name and description"><figcaption><p>The edit dialog is prefilled with the current evaluation name and description</p></figcaption></figure>

{% hint style="info" %}
You can edit the name and description of evaluations that are running, completed, or failed. This action only updates the identifying information and does not rerun the test.
{% endhint %}

***

## AI Assistant Monitoring <a href="#ai-agent-monitoring" id="ai-agent-monitoring"></a>

AI Assistant Monitoring provides real-time conversation operational data, allowing you to gain deep insight into the processing details, performance metrics, and quality scores of each conversation.

### Access AI Assistant Monitoring <a href="#access-ai-agent-monitoring" id="access-ai-agent-monitoring"></a>

Go to "<mark style="color:blue;">AgentOps</mark>" in the left sidebar, then click "<mark style="color:blue;">AI Assistant Monitoring</mark>".

<figure><img src="https://1360999650-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2F6v6TNkkOQVfRYfcNirHL%2Fuploads%2Fgit-blob-3cd07b8fa370cc1cc3c44291f6fb6292f29086f5%2Fagentops-monitoring.png?alt=media" alt="AI Assistant Monitoring"><figcaption><p>AI Assistant Monitoring interface showing detailed technical metrics for each conversation</p></figcaption></figure>

### Monitoring Field Descriptions <a href="#monitoring-field-descriptions" id="monitoring-field-descriptions"></a>

| Field                         | Description                                                               |
| ----------------------------- | ------------------------------------------------------------------------- |
| User Input Message            | The question sent by the user to the AI                                   |
| Output Message                | The AI Assistant's response content                                       |
| AI Assistant                  | The name of the AI Assistant that handled the conversation                |
| User Feedback                 | Thumbs up 👍 or thumbs down 👎                                            |
| Faithfulness Score            | Whether the response is faithful to the Knowledge Base content            |
| Response Relevance Score      | The degree of relevance between the response and the question             |
| Response Time                 | The total time for the AI to generate a response                          |
| LLM Processing Reasoning Time | The time spent on LLM reasoning and response generation                   |
| Total Token Count             | Total tokens consumed in the conversation (including question and answer) |
| LLM                           | The name of the language model used                                       |
| User                          | The user who initiated the conversation                                   |

### View Image Search Usage <a href="#view-image-search-usage" id="view-image-search-usage"></a>

When an AI Assistant uses image search in the Knowledge Base, the system separately records image embedding usage. You can review an individual conversation record to see which resources were used for that response:

{% stepper %}
{% step %}
Go to <mark style="color:blue;">AgentOps</mark> → <mark style="color:blue;">AI Assistant Monitoring</mark>, then switch to <mark style="color:blue;">Conversation Records</mark>.
{% endstep %}

{% step %}
Find the response you want to review, then click <mark style="color:blue;">Details</mark> on the far right.
{% endstep %}

{% step %}
Scroll down to view the Token and Credit breakdown. Standard text embedding is listed under <mark style="color:blue;">Embedding</mark>. If the response used image search, <mark style="color:blue;">Image Embedding</mark> is also displayed, and the Token and Credit totals include this item.
{% endstep %}
{% endstepper %}

<figure><img src="https://1360999650-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2F6v6TNkkOQVfRYfcNirHL%2Fuploads%2Fgit-blob-fe269be3872092ab4473c8eb46b0192090b76469%2Fagentops-prod-image-embedding-breakdown.png?alt=media" alt="Usage and Credit breakdown for an individual conversation in AI Assistant Monitoring"><figcaption><p>Usage breakdown for an individual conversation; “Image Embedding” appears only when the response actually used image search</p></figcaption></figure>

{% hint style="info" %}
If <mark style="color:blue;">Image Embedding</mark> is not displayed, it does not mean the record is missing. This item appears only when the response actually incurs image embedding usage. When image search is not used, the screen continues to show the standard embedding breakdown.
{% endhint %}

For example, Mai's customer service manager asks the product assistant to identify the appearance of a component from catalog images. After testing, she opens the conversation details and sees the image embedding usage and corresponding Credits. This allows her to distinguish the cost of image search from standard text retrieval before deciding whether to adjust the image catalog or assistant settings.

### Search and Filter <a href="#search-and-filter" id="search-and-filter"></a>

* **Keyword Search**: Search input/output messages or usernames
* **LLM Filter**: Select a specific language model to compare performance across different models
* **AI Assistant Filter**: Select a specific assistant to track its operational status
* **Time Range**: Select last 7 days, 30 days, 90 days, or a custom date range
* **Export**: Export monitoring data in Excel or CSV format

### Dashboard Metrics and Error Rate <a href="#dashboard-metrics" id="dashboard-metrics"></a>

Switch to the “<mark style="color:blue;">Dashboard</mark>” tab to view overall service metrics and their percentage changes compared with the preceding period of the same length:

| Metric                | Description                                                                                                                                                                                                                                       |
| --------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| Conversations         | The number of **responses from AI Assistants** during the selected period. Each AI response in a conversation counts once; responses handled by human agents and automatically scheduled runs are excluded (see the definition below for details) |
| Error Rate            | Percentage of responses with system errors (see the calculation below)                                                                                                                                                                            |
| Average Response Time | Average time from receiving a message to completing the response                                                                                                                                                                                  |
| Average TTFT          | Average time from when a user sends a message until the first token appears                                                                                                                                                                       |

The dashboard also provides charts for conversation volume, response time, error-rate trends, and LLM model distribution. The “Assistant Service Ranking” compares error rates and response speeds across assistants.

<figure><img src="https://1360999650-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2F6v6TNkkOQVfRYfcNirHL%2Fuploads%2Fgit-blob-db1630e1f0fa85b411bb81deebdb35dd7f05981a%2Fagentops-dashboard-metrics.png?alt=media" alt="AI Assistant Monitoring dashboard"><figcaption><p>Dashboard metric cards and trend charts</p></figcaption></figure>

#### How Is the Conversation Count Calculated? <a href="#conversation-count-definition" id="conversation-count-definition"></a>

The dashboard's "Conversations" metric and conversation volume trend count **each response from an AI Assistant**, rather than the number of conversations opened:

| Scenario                                                                      | Counted? | Count |
| ----------------------------------------------------------------------------- | -------- | ----- |
| A user enters the service but does not send a message                         | No       | 0     |
| A user sends 1 message and the AI responds once                               | Yes      | 1     |
| In the same conversation, a user asks 5 questions and the AI responds 5 times | Yes      | 5     |
| A user sends a message, but a human agent takes over and responds             | No       | 0     |
| A user sends a message, but an error occurs while the AI is responding        | Yes      | 1     |
| An automatically scheduled task runs                                          | No       | 0     |

{% hint style="info" %}
To see how many conversations were opened, view "Conversation Count" on the AI Assistant's "Usage Analytics" page instead. It counts 1 when a user sends the first message, regardless of how many turns the conversation contains, and includes conversations handled by human agents. These metrics have different definitions, so the dashboard's "Conversations" count is generally higher than the "Conversation Count" in "Usage Analytics." See [Usage Analytics](/maiagent-user-guide/en/org/usage.md#conversations-count) for details.
{% endhint %}

#### How Is the Error Rate Calculated? <a href="#error-rate-calculation" id="error-rate-calculation"></a>

> Error rate = Number of responses with system errors ÷ Total number of responses × 100%

Only **system-level** errors are counted, such as model call failures and timeouts. The following are **not counted** as errors to avoid overestimating the error rate:

* Responses interrupted by the user (Client Interrupt)
* Normal blocking by content safeguards (Hook)

{% hint style="info" %}
The error rate measures service stability, not response quality. If the error rate rises abnormally, review the assistant's conversation records to identify the cause, such as model issues or external tool integration failures. To assess response quality, review faithfulness and answer relevance scores together with users' thumbs-up and thumbs-down feedback.
{% endhint %}

### Monitoring Best Practices <a href="#monitoring-best-practices" id="monitoring-best-practices"></a>

**Daily Review**: Check conversations from the last 24 hours to identify abnormal response times or errors

**Identify Performance Bottlenecks**:

* Response time > 10 seconds: Check Knowledge Base retrieval efficiency or consider a faster LLM
* High token usage: Evaluate whether the System Prompt or conversation history can be shortened

**Quality Issue Tracking**:

1. Use keyword search to find problematic conversations
2. Analyze root causes (insufficient Knowledge Base / AI misunderstanding / model limitations)
3. Add problem cases to the test dataset and run automated evaluations to verify fixes

***

## FAQ <a href="#faq" id="faq"></a>

### Q: What is the difference between Automated Evaluation and AI Assistant Monitoring? <a href="#faq-evaluation-vs-monitoring" id="faq-evaluation-vs-monitoring"></a>

|             | Automated Evaluation                     | AI Assistant Monitoring           |
| ----------- | ---------------------------------------- | --------------------------------- |
| Purpose     | Periodic quality testing                 | Real-time operational monitoring  |
| Data Source | Predefined test datasets                 | Actual user conversations         |
| Key Metrics | Success rate, response time              | Performance, cost, quality scores |
| Best For    | Quality verification, regression testing | Daily monitoring, troubleshooting |

### Q: How often should evaluations be run? <a href="#faq-how-often-to-evaluate" id="faq-how-often-to-evaluate"></a>

Recommendation: Core features weekly, full tests monthly, and immediately after major updates.

### Q: Do evaluations affect actual users? <a href="#faq-evaluation-impact-on-users" id="faq-evaluation-impact-on-users"></a>

No. Automated evaluations run in an isolated environment and do not interfere with actual user conversations.

### Q: How long is monitoring data retained? <a href="#faq-monitoring-data-retention" id="faq-monitoring-data-retention"></a>

The default retention period is 90 days. You can periodically export important data for long-term storage.


---

# Agent Instructions
This documentation is published with GitBook. GitBook is the documentation platform designed so that both humans and AI agents can read, navigate, and reason over technical content effectively. Learn more at gitbook.com.

## Querying This Documentation
If you need additional information that is not directly available in this page, you can query the documentation dynamically by asking a question.

Perform an HTTP GET request on the current page URL with the `ask` query parameter, and the optional `goal` query parameter:

```
GET https://docs.maiagent.ai/maiagent-user-guide/en/agent-ops/evaluations.md?ask=<question>&goal=<endgoal>
```

`ask` is the immediate question: it should be specific, self-contained, and written in natural language.
`goal` is optional and describes the broader end goal you are ultimately trying to accomplish on behalf of the user. GitBook uses it to tailor the answer towards what is most useful for that goal.

The response will contain a direct answer to the question and relevant excerpts and sources from the documentation.

Use this mechanism when the answer is not explicitly present in the current page, you need clarification or additional context, or you want to retrieve related documentation sections.
