> For the complete documentation index, see [llms.txt](https://docs.maiagent.ai/llms.txt). Markdown versions of documentation pages are available by appending `.md` to page URLs; this page is available as [Markdown](https://docs.maiagent.ai/agent-ops/evaluations.md).

# 自動評估與 AI 助理監控

自動化測試與即時監控 AI 助理的回應品質、效能指標與成本

## 自動評估 <a href="#auto-evaluation" id="auto-evaluation"></a>

自動評估功能讓您能夠使用預先建立的測試資料集，自動化測試 AI 助理的回應品質。系統會將測試問題發送給 AI 助理，比對實際回應與預期回應，產生詳細的評估報告。

### 進入自動評估 <a href="#access-auto-evaluation" id="access-auto-evaluation"></a>

進入左側功能欄的「<mark style="color:blue;">AgentOps</mark>」，點選「<mark style="color:blue;">自動化測試</mark>」。

<figure><img src="https://1593648278-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2Fmzb5NG9GDzFP2YDKeYVl%2Fuploads%2Fgit-blob-06f0a8f71a921e02e030b8b8e71bc49977fc7154%2Fagentops-evaluations.png?alt=media&amp;token=b2b4ca50-3e05-4c4c-9043-1dab3d7b6c18" alt="自動化測試列表"><figcaption><p>自動化測試列表，顯示各次測試的成功率與平均回應秒數</p></figcaption></figure>

頁面顯示所有評估記錄，包含評測名稱、測試集、AI 助理、成功率、平均秒數與建立時間。

### 建立並執行測試 <a href="#create-and-run-test" id="create-and-run-test"></a>

1. 確認已建立測試資料集（參考 [測試資料集管理](/agent-ops/test-datasets.md)）
2. 點擊「<mark style="color:blue;">建立測試</mark>」按鈕
3. 填寫評估名稱、描述，選擇測試資料集與 AI 助理
4. 點擊「<mark style="color:blue;">開始評估</mark>」，系統會自動執行所有測試案例

{% hint style="info" %}
評估執行時間取決於測試案例數量，通常 50 個測試案例約需 2-3 分鐘。
{% endhint %}

### 使用自訂評估指標 <a href="#custom-evaluation-metrics" id="custom-evaluation-metrics"></a>

除了內建指標，您也可以用自然語言定義團隊自己的評分準則。例如，客服主管小麥想確認客服助理是否保持禮貌且有同理心；她建立「客服語氣」指標、寫下判斷準則並設定通過門檻。評估完成後，她便能在結果中查看每個案例的「客服語氣」分數，找出需要調整的回覆。

{% stepper %}
{% step %}
建立評估時，完成名稱、測試集、AI 助理及評測模型等基本設定。
{% endstep %}

{% step %}
在「<mark style="color:blue;">自訂指標</mark>」區域點選「<mark style="color:blue;">新增自訂指標</mark>」，輸入容易辨識的指標名稱與評估準則。
{% endstep %}

{% step %}
依團隊的品質標準調整通過門檻，再點選「<mark style="color:blue;">確認</mark>」開始評估。
{% endstep %}
{% endstepper %}

<figure><img src="https://1593648278-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2Fmzb5NG9GDzFP2YDKeYVl%2Fuploads%2Fgit-blob-43ed547b1eec12101661f11efa05ddfa646c6a4d%2Fevaluation-multi-turn-settings.png?alt=media" alt="建立評估視窗的多輪對話設定"><figcaption><p>切換至多輪對話後，可設定最大輪數、三種評估指標與模擬語言</p></figcaption></figure>

<figure><img src="https://1593648278-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2Fmzb5NG9GDzFP2YDKeYVl%2Fuploads%2Fgit-blob-d56c6a2a2384f1f8aab3561aa424ae4364d3f37e%2Fevaluation-custom-metric.png?alt=media" alt="建立評估視窗中的自訂指標設定"><figcaption><p>以自然語言設定客服語氣的評估準則</p></figcaption></figure>

{% hint style="info" %}
系統以不分英文大小寫的方式比對名稱；自訂指標不可與 11 個內建指標重複。若名稱衝突，請改用能表達團隊判斷目的的名稱。
{% endhint %}

### 評估多輪對話 <a href="#multi-turn-evaluation" id="multi-turn-evaluation"></a>

多輪評估適合檢查需要來回確認的任務。以退貨客服為例，品質管理員小美希望助理先詢問訂單編號，再確認退貨原因並說明後續安排；她建立多輪測試案例，寫下情境與預期結果後執行評估。完成後，她可展開對話逐字稿，檢查每一輪回覆與知識庫檢索內容，而不是只看最後一句答案。

{% stepper %}
{% step %}
先到「<mark style="color:blue;">測試集</mark>」開啟目標測試集，在多輪測試案例中新增情境、預期結果與模擬使用者人設。
{% endstep %}

{% step %}
回到「<mark style="color:blue;">自動化測試</mark>」並點選「<mark style="color:blue;">建立評估</mark>」，切換到「<mark style="color:blue;">多輪對話</mark>」。
{% endstep %}

{% step %}
選擇測試集、AI 助理與評測模型，設定最大輪數、評估指標及模擬語言後開始評估。至少須啟用一項評估指標。
{% endstep %}

{% step %}
評估完成後開啟結果，點選個別案例展開詳細內容，再查看各指標分數、對話逐字稿，以及每一輪的「<mark style="color:blue;">檢索內容</mark>」。單一案例失敗時，該案例會個別顯示錯誤，不會遮蔽其他案例的結果。
{% endstep %}
{% endstepper %}

<figure><img src="https://1593648278-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2Fmzb5NG9GDzFP2YDKeYVl%2Fuploads%2Fgit-blob-123f84448a9ae205a2a7adcc1f96b8a63ceaf02d%2Fevaluation-multi-turn-results.png?alt=media" alt="已完成多輪評估展開案例後的指標、逐字稿與檢索內容"><figcaption><p>展開案例可逐輪核對助理回覆與該輪取用的知識庫內容</p></figcaption></figure>

執行中的多輪評估若不需繼續，可在評估詳情點選「<mark style="color:blue;">取消評測</mark>」並確認。系統會在目前案例處理完畢後停止後續案例；已完成的分數與逐字稿仍會保留，狀態顯示為「<mark style="color:blue;">部分完成</mark>」，尚未處理的案例顯示為「<mark style="color:blue;">待評估</mark>」。

<figure><img src="https://1593648278-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2Fmzb5NG9GDzFP2YDKeYVl%2Fuploads%2Fgit-blob-926b7035060aac6e77573489d86fb8297fab864f%2Fevaluation-multi-turn-cancel.png?alt=media" alt="取消多輪評估後的部分完成狀態"><figcaption><p>取消後保留已完成案例，尚未處理的案例維持待評估</p></figcaption></figure>

### 查看評估結果 <a href="#view-evaluation-results" id="view-evaluation-results"></a>

點擊評估記錄即可查看詳細報告：

<figure><img src="https://1593648278-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2Fmzb5NG9GDzFP2YDKeYVl%2Fuploads%2Fgit-blob-aeec25455ae679a86ed97bc0ed49d20cb37e8b24%2Ftrack-evaluation-detail.png?alt=media&amp;token=52592159-ecd4-4ca6-a1a2-537a4d5154fa" alt="評估詳情"><figcaption><p>測試詳情頁面：成功率、AI 洞察摘要、指標統計與改進建議</p></figcaption></figure>

詳細報告包含：

* **成功率**：通過測試的案例百分比
* **AI 洞察**：系統自動分析的評估摘要與改進建議
* **指標統計**：品質性評分、回答相關性等指標
* **測試案例明細**：每個問題的預期回應、實際回應、評分與狀態

**成功率參考基準：**

| AI 助理類型 | 建議成功率 |
| ------- | ----- |
| 產品查詢助理  | ≥ 95% |
| 客服支援助理  | ≥ 90% |
| 通用對話助理  | ≥ 80% |

{% hint style="info" %}
更多關於洞察報告的說明，請參考：[評估洞察報告](/agent-ops/evaluation-insights.md)
{% endhint %}

### 管理評估記錄 <a href="#manage-evaluation-records" id="manage-evaluation-records"></a>

* **搜尋與篩選**：依測試集、AI 助理或關鍵字篩選記錄
* **重新執行**：修改知識庫或 AI 設定後，重新測試驗證改善效果
* **匯出**：將評估結果匯出為 Excel 格式

#### 編輯評測名稱與描述 <a href="#edit-evaluation-name-description" id="edit-evaluation-name-description"></a>

完成測試後，您仍可修改評測名稱與描述，不必刪除或重新執行測試。測試資料集、AI 助理、評測結果與執行狀態不會因此改變。

例如，小麥建立每週客服品質測試後，發現評測名稱的週次寫錯。她從自動化測試列表開啟編輯視窗，更正名稱並補上測試目的；儲存後，列表立即顯示新內容，原有的測試結果仍保留。

{% stepper %}
{% step %}

### 開啟自動化測試列表

進入左側選單「<mark style="color:blue;">AgentOps</mark>」→「<mark style="color:blue;">自動化測試</mark>」。
{% endstep %}

{% step %}

### 開啟編輯視窗

找到要調整的評測，在該列的「<mark style="color:blue;">操作</mark>」欄點擊「<mark style="color:blue;">編輯</mark>」圖示。
{% endstep %}

{% step %}

### 修改並儲存

修改「<mark style="color:blue;">評測名稱</mark>」或「<mark style="color:blue;">描述</mark>」，再點擊「<mark style="color:blue;">確定</mark>」。評測名稱為必填，最多 200 個字元；描述可以留空。
{% endstep %}
{% endstepper %}

<figure><img src="https://1593648278-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2Fmzb5NG9GDzFP2YDKeYVl%2Fuploads%2Fgit-blob-574e934d25a9ff140f8191df9c4bef52ccbdf9a2%2Fagentops-edit-evaluation.png?alt=media" alt="自動化測試列表的編輯測試視窗，可修改評測名稱與描述"><figcaption><p>編輯視窗會帶入目前的評測名稱與描述</p></figcaption></figure>

{% hint style="info" %}
執行中、已完成或失敗的評測都可以修改名稱與描述；這項操作只更新辨識用資訊，不會重新執行測試。
{% endhint %}

***

## AI 助理監控 <a href="#ai-agent-monitoring" id="ai-agent-monitoring"></a>

AI 助理監控提供即時的對話運作數據，讓您深入了解每一次對話的處理細節、效能指標與品質評分。

### 進入 AI 助理監控 <a href="#access-ai-agent-monitoring" id="access-ai-agent-monitoring"></a>

進入左側功能欄的「<mark style="color:blue;">AgentOps</mark>」，點選「<mark style="color:blue;">AI 助理監控</mark>」。

<figure><img src="https://1593648278-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2Fmzb5NG9GDzFP2YDKeYVl%2Fuploads%2Fgit-blob-3cd07b8fa370cc1cc3c44291f6fb6292f29086f5%2Fagentops-monitoring.png?alt=media&amp;token=70411417-bde6-43b7-8c83-194a557e29f8" alt="AI 助理監控"><figcaption><p>AI 助理監控介面，顯示每則對話的詳細技術指標</p></figcaption></figure>

### 監控欄位說明 <a href="#monitoring-field-descriptions" id="monitoring-field-descriptions"></a>

| 欄位         | 說明                |
| ---------- | ----------------- |
| 用戶輸入訊息     | 使用者發送給 AI 的問題     |
| 輸出訊息       | AI 助理的回應內容        |
| AI 助理      | 處理此對話的 AI 助理名稱    |
| 使用者反饋      | 讚 👍 或倒讚 👎       |
| 誠實性評分      | 回應是否忠實於知識庫內容      |
| 回答相關性評分    | 回應與問題的相關程度        |
| 回應時間       | AI 產生回應的完整時間      |
| LLM 處理推理時間 | LLM 推理與回覆生成所花費的時間 |
| 總字數        | 對話中消耗的總字數（含問題與回答） |
| LLM        | 使用的語言模型名稱         |
| 用戶         | 發起對話的使用者          |

### 查看圖片搜尋的用量 <a href="#view-image-search-usage" id="view-image-search-usage"></a>

AI 助理使用知識庫的圖片搜尋時，系統會另外記錄圖片向量化的用量。您可以從單筆對話記錄確認這次回覆用了哪些資源：

{% stepper %}
{% step %}
進入 <mark style="color:blue;">AgentOps</mark> → <mark style="color:blue;">AI 助理監控</mark>，切換到 <mark style="color:blue;">對話紀錄</mark>。
{% endstep %}

{% step %}
找到要核對的回覆，點選最右側的<mark style="color:blue;">詳細資料</mark>。
{% endstep %}

{% step %}
往下查看 Token 與 Credit 明細。一般文字向量化會列在 <mark style="color:blue;">向量化</mark>；該回覆有使用圖片搜尋時，會再顯示 <mark style="color:blue;">圖片向量化</mark>，Token 總計與 Credit 總計也會包含這一項。
{% endstep %}
{% endstepper %}

<figure><img src="https://1593648278-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2Fmzb5NG9GDzFP2YDKeYVl%2Fuploads%2Fgit-blob-fe269be3872092ab4473c8eb46b0192090b76469%2Fagentops-prod-image-embedding-breakdown.png?alt=media" alt="AI 助理監控的單筆對話用量與 Credit 明細"><figcaption><p>單筆對話的用量明細；只有實際使用圖片搜尋的回覆才會顯示「圖片向量化」</p></figcaption></figure>

{% hint style="info" %}
看不到 <mark style="color:blue;">圖片向量化</mark> 不代表記錄遺失。這個項目只在該次回覆實際產生圖片向量化用量時顯示；沒有使用圖片搜尋時，畫面會維持原本的向量化明細。
{% endhint %}

例如，小麥公司的客服主管讓商品助理從型錄圖片辨識零件外觀。她在測試後打開這筆對話的詳細資料，看到圖片向量化與對應 Credits，便能區分圖片搜尋和一般文字檢索的成本，再決定是否調整圖片型錄或助理設定。

### 搜尋與篩選 <a href="#search-and-filter" id="search-and-filter"></a>

* **關鍵字搜尋**：搜尋輸入/輸出訊息或使用者名稱
* **LLM 篩選**：選擇特定語言模型，比較不同模型的效能
* **AI 助理篩選**：選擇特定助理，追蹤其運作狀況
* **時間範圍**：選擇近 7 天、30 天、90 天或自訂日期
* **匯出**：將監控數據匯出為 Excel 或 CSV 格式

### 儀表板指標與錯誤率 <a href="#dashboard-metrics" id="dashboard-metrics"></a>

切換到「<mark style="color:blue;">儀表板</mark>」分頁，可看到整體服務指標，並附上與前一個相同長度區間的變化率：

| 指標      | 說明                                                                   |
| ------- | -------------------------------------------------------------------- |
| 對話數     | 區間內 AI 助理**回覆的次數**。同一個對話有幾輪 AI 回覆就計幾次；由真人客服接手回覆、排程自動執行的不計入（定義詳見下方說明） |
| 錯誤率     | 發生系統錯誤的回覆比例（計算方式見下）                                                  |
| 平均回應時間  | 從收到訊息到回覆完成的平均時間                                                      |
| 平均 TTFT | 使用者送出訊息到看到第一個字的平均時間                                                  |

另提供對話量趨勢、回應時間趨勢、錯誤率趨勢、LLM 模型分布等圖表，以及「助理服務排名」比較各助理的錯誤率與回應速度。

<figure><img src="https://1593648278-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2Fmzb5NG9GDzFP2YDKeYVl%2Fuploads%2Fgit-blob-db1630e1f0fa85b411bb81deebdb35dd7f05981a%2Fagentops-dashboard-metrics.png?alt=media" alt="AI 助理監控儀表板"><figcaption><p>儀表板指標卡與趨勢圖</p></figcaption></figure>

#### 對話數怎麼計算 <a href="#conversation-count-definition" id="conversation-count-definition"></a>

儀表板的「對話數」與「對話量趨勢」是以 **AI 助理每一次回覆** 為單位，不是以「開了幾個對話」為單位：

| 情境                       | 是否計入 | 計入數量 |
| ------------------------ | ---- | ---- |
| 使用者進入服務，但沒有送出訊息          | 不計   | 0    |
| 使用者送出 1 則訊息，AI 回覆 1 次    | 計    | 1    |
| 同一個對話中使用者問 5 次，AI 回覆 5 次 | 計    | 5    |
| 使用者送出訊息，但由真人客服接手回覆       | 不計   | 0    |
| 使用者送出訊息，AI 回覆時發生錯誤       | 計    | 1    |
| 排程自動執行的任務                | 不計   | 0    |

{% hint style="info" %}
若要看「有多少個對話被開啟」，請改看 AI 助理「使用分析」分頁的「對話次數」：使用者送出第一則訊息即計 1，同一個對話不論幾輪只算 1，真人客服接手的對話也會計入。兩個數字的定義不同，一般情況下儀表板的「對話數」會大於「使用分析」的「對話次數」。詳見 [使用分析](/org/usage.md#conversations-count)。
{% endhint %}

#### 錯誤率怎麼計算 <a href="#error-rate-calculation" id="error-rate-calculation"></a>

> 錯誤率 ＝ 發生系統錯誤的回覆筆數 ÷ 總回覆筆數 × 100%

只有**系統層**的錯誤才會計入（例如模型呼叫失敗、逾時）。以下情況**不列入**錯誤，避免高估：

* 使用者主動中斷回覆（Client Interrupt）
* 內容防護機制（Hook）的正常攔截

{% hint style="info" %}
錯誤率是「服務穩定度」指標，不是答題品質分數。錯誤率異常升高時，建議進入該助理的對話紀錄查明原因（例如模型異常、外部工具串接失敗）；若要評估「回答得好不好」，請搭配誠實性／回答相關性評分與使用者的讚／倒讚回饋一起檢視。
{% endhint %}

### 監控最佳實踐 <a href="#monitoring-best-practices" id="monitoring-best-practices"></a>

**每日檢視**：查看最近 24 小時的對話，識別異常的回應時間或錯誤

**識別效能瓶頸**：

* 回覆時間 > 10 秒：檢查知識庫檢索效率或考慮更快的 LLM
* Token 用量過高：評估是否可縮短 System Prompt 或對話歷史

**品質問題追蹤**：

1. 使用關鍵字搜尋找到問題對話
2. 分析根本原因（知識庫不足 / AI 理解錯誤 / 模型限制）
3. 將問題案例加入測試資料集，執行自動評估驗證修復

***

## 常見問題 <a href="#faq" id="faq"></a>

### Q：自動評估與 AI 助理監控有什麼不同？ <a href="#faq-evaluation-vs-monitoring" id="faq-evaluation-vs-monitoring"></a>

|      | 自動評估      | AI 助理監控    |
| ---- | --------- | ---------- |
| 用途   | 定期品質測試    | 即時運作監控     |
| 資料來源 | 預設的測試資料集  | 實際用戶對話     |
| 主要指標 | 成功率、回應時間  | 效能、成本、品質評分 |
| 適合對象 | 品質驗證、回歸測試 | 日常監控、問題排查  |

### Q：多久執行一次評估？ <a href="#faq-how-often-to-evaluate" id="faq-how-often-to-evaluate"></a>

建議：核心功能每週一次、完整測試每月一次、重大更新後立即執行。

### Q：評估會影響實際用戶嗎？ <a href="#faq-evaluation-impact-on-users" id="faq-evaluation-impact-on-users"></a>

不會。自動評估使用獨立環境，不會干擾實際用戶的對話。

### Q：監控數據保留多久？ <a href="#faq-monitoring-data-retention" id="faq-monitoring-data-retention"></a>

預設保留 90 天。可定期匯出重要數據進行長期保存。


---

# Agent Instructions
This documentation is published with GitBook. GitBook is the documentation platform designed so that both humans and AI agents can read, navigate, and reason over technical content effectively. Learn more at gitbook.com.

## Querying This Documentation
If you need additional information that is not directly available in this page, you can query the documentation dynamically by asking a question.

Perform an HTTP GET request on the current page URL with the `ask` query parameter, and the optional `goal` query parameter:

```
GET https://docs.maiagent.ai/agent-ops/evaluations.md?ask=<question>&goal=<endgoal>
```

`ask` is the immediate question: it should be specific, self-contained, and written in natural language.
`goal` is optional and describes the broader end goal you are ultimately trying to accomplish on behalf of the user. GitBook uses it to tailor the answer towards what is most useful for that goal.

The response will contain a direct answer to the question and relevant excerpts and sources from the documentation.

Use this mechanism when the answer is not explicitly present in the current page, you need clarification or additional context, or you want to retrieve related documentation sections.
