> For the complete documentation index, see [llms.txt](https://docs.maiagent.ai/llms.txt). Markdown versions of documentation pages are available by appending `.md` to page URLs; this page is available as [Markdown](https://docs.maiagent.ai/maiagent-user-guide/maiagent-user-guide-ja/agent-ops/evaluations.md).

# 自動評価と AI アシスタントの監視

## 自動評価 <a href="#auto-evaluation" id="auto-evaluation"></a>

自動評価機能では、あらかじめ作成したテストデータセットを使用して、AI アシスタントの応答品質を自動的にテストできます。システムはテスト質問を AI アシスタントに送信し、実際の応答と期待される応答を照合して、詳細な評価レポートを生成します。

### 自動評価へのアクセス <a href="#access-auto-evaluation" id="access-auto-evaluation"></a>

左側のメニューから「<mark style="color:blue;">AgentOps</mark>」に入り、「<mark style="color:blue;">自動テスト</mark>」をクリックします。

<figure><img src="/files/lr3aloDFK5KsySuGlIDd" alt="自動テスト一覧"><figcaption><p>自動テスト一覧。各テストの成功率と平均応答秒数を表示します</p></figcaption></figure>

このページには、評価名、テストセット、AI アシスタント、成功率、平均秒数、作成日時を含むすべての評価記録が表示されます。

### テストの作成と実行 <a href="#create-and-run-test" id="create-and-run-test"></a>

1. テストデータセットが作成済みであることを確認します（[テストデータセット管理](/maiagent-user-guide/maiagent-user-guide-ja/agent-ops/test-datasets.md)を参照）
2. 「<mark style="color:blue;">テストを作成</mark>」ボタンをクリックします
3. 評価名と説明を入力し、テストデータセットと AI アシスタントを選択します
4. 「<mark style="color:blue;">評価を開始</mark>」をクリックすると、システムがすべてのテストケースを自動的に実行します

{% hint style="info" %}
評価の実行時間はテストケースの数によって異なり、通常 50 件のテストケースで約 2〜3 分かかります。
{% endhint %}

### 評価結果の確認 <a href="#view-evaluation-results" id="view-evaluation-results"></a>

評価記録をクリックすると、詳細レポートを確認できます。

<figure><img src="/files/nnF1kgiUFhqdTE7Xrr3A" alt="評価詳細"><figcaption><p>テスト詳細ページ：成功率、AI インサイト概要、指標統計、改善提案</p></figcaption></figure>

詳細レポートには以下が含まれます。

* **成功率**：テストに合格したケースの割合
* **AI インサイト**：システムが自動分析した評価概要と改善提案
* **指標統計**：品質評価、回答関連性などの指標
* **テストケース明細**：各質問の期待される応答、実際の応答、評価、ステータス

**成功率の参考基準：**

| AI アシスタントの種類    | 推奨成功率 |
| --------------- | ----- |
| 製品照会アシスタント      | ≥ 95% |
| カスタマーサポートアシスタント | ≥ 90% |
| 汎用対話アシスタント      | ≥ 80% |

{% hint style="info" %}
インサイトレポートの詳細については、こちらを参照してください：[評価インサイトレポート](/maiagent-user-guide/maiagent-user-guide-ja/agent-ops/evaluation-insights.md)
{% endhint %}

### 評価記録の管理 <a href="#manage-evaluation-records" id="manage-evaluation-records"></a>

* **検索とフィルタリング**：テストセット、AI アシスタント、またはキーワードで記録をフィルタリングします
* **再実行**：ナレッジベースや AI 設定を変更した後、再テストして改善効果を検証します
* **エクスポート**：評価結果を Excel 形式でエクスポートします

***

## AI アシスタント監視 <a href="#ai-agent-monitoring" id="ai-agent-monitoring"></a>

AI アシスタント監視は、リアルタイムの対話運用データを提供し、各対話の処理詳細、パフォーマンス指標、品質評価を詳しく把握できるようにします。

### AI アシスタント監視へのアクセス <a href="#access-ai-agent-monitoring" id="access-ai-agent-monitoring"></a>

左側のメニューから「<mark style="color:blue;">AgentOps</mark>」に入り、「<mark style="color:blue;">AI アシスタント監視</mark>」をクリックします。

<figure><img src="/files/oCH6ofmYa97jbPKtUDYg" alt="AI アシスタント監視"><figcaption><p>AI アシスタント監視画面。各対話の詳細な技術指標を表示します</p></figcaption></figure>

### 監視項目の説明 <a href="#monitoring-field-descriptions" id="monitoring-field-descriptions"></a>

| 項目          | 説明                     |
| ----------- | ---------------------- |
| ユーザー入力メッセージ | ユーザーが AI に送信した質問       |
| 出力メッセージ     | AI アシスタントの応答内容         |
| AI アシスタント   | この対話を処理した AI アシスタントの名前 |
| ユーザーフィードバック | 高評価 👍 または低評価 👎       |
| 誠実性スコア      | 応答がナレッジベースの内容に忠実かどうか   |
| 回答関連性スコア    | 応答と質問の関連度              |
| 応答時間        | AI が応答を生成するまでの全体の時間    |
| LLM 処理推論時間  | LLM の推論と応答生成にかかった時間    |
| 総文字数        | 対話で消費された総文字数（質問と回答を含む） |
| LLM         | 使用した言語モデルの名前           |
| ユーザー        | 対話を開始したユーザー            |

### 検索とフィルタリング <a href="#search-and-filter" id="search-and-filter"></a>

* **キーワード検索**：入力/出力メッセージまたはユーザー名を検索します
* **LLM フィルタリング**：特定の言語モデルを選択し、異なるモデルのパフォーマンスを比較します
* **AI アシスタントフィルタリング**：特定のアシスタントを選択し、その運用状況を追跡します
* **期間**：直近 7 日間、30 日間、90 日間、またはカスタム日付を選択します
* **エクスポート**：監視データを Excel または CSV 形式でエクスポートします

### 監視のベストプラクティス <a href="#monitoring-best-practices" id="monitoring-best-practices"></a>

**毎日の確認**：直近 24 時間の対話を確認し、異常な応答時間やエラーを特定します

**パフォーマンスのボトルネックの特定**：

* 応答時間 > 10 秒：ナレッジベースの検索効率を確認するか、より高速な LLM の利用を検討します
* Token 使用量が多すぎる：System Prompt や対話履歴を短縮できるか評価します

**品質問題の追跡**：

1. キーワード検索を使用して問題のある対話を見つけます
2. 根本原因を分析します（ナレッジベース不足 / AI の理解誤り / モデルの制約）
3. 問題のケースをテストデータセットに追加し、自動評価を実行して修正を検証します

***

## よくある質問 <a href="#faq" id="faq"></a>

### Q：自動評価と AI アシスタント監視は何が違いますか？ <a href="#faq-evaluation-vs-monitoring" id="faq-evaluation-vs-monitoring"></a>

|        | 自動評価               | AI アシスタント監視      |
| ------ | ------------------ | ---------------- |
| 用途     | 定期的な品質テスト          | リアルタイムの運用監視      |
| データソース | あらかじめ設定したテストデータセット | 実際のユーザー対話        |
| 主な指標   | 成功率、応答時間           | パフォーマンス、コスト、品質評価 |
| 適した対象  | 品質検証、回帰テスト         | 日常的な監視、問題の切り分け   |

### Q：評価はどのくらいの頻度で実行すればよいですか？ <a href="#faq-how-often-to-evaluate" id="faq-how-often-to-evaluate"></a>

推奨：コア機能は週 1 回、完全なテストは月 1 回、重要なアップデート後はただちに実行します。

### Q：評価は実際のユーザーに影響しますか？ <a href="#faq-evaluation-impact-on-users" id="faq-evaluation-impact-on-users"></a>

いいえ。自動評価は独立した環境を使用するため、実際のユーザーの対話を妨げることはありません。

### Q：監視データはどのくらいの期間保持されますか？ <a href="#faq-monitoring-data-retention" id="faq-monitoring-data-retention"></a>

デフォルトでは 90 日間保持されます。重要なデータは定期的にエクスポートして長期保存できます。


---

# Agent Instructions
This documentation is published with GitBook. GitBook is the documentation platform designed so that both humans and AI agents can read, navigate, and reason over technical content effectively. Learn more at gitbook.com.

## Querying This Documentation
If you need additional information that is not directly available in this page, you can query the documentation dynamically by asking a question.

Perform an HTTP GET request on the current page URL with the `ask` query parameter, and the optional `goal` query parameter:

```
GET https://docs.maiagent.ai/maiagent-user-guide/maiagent-user-guide-ja/agent-ops/evaluations.md?ask=<question>&goal=<endgoal>
```

`ask` is the immediate question: it should be specific, self-contained, and written in natural language.
`goal` is optional and describes the broader end goal you are ultimately trying to accomplish on behalf of the user. GitBook uses it to tailor the answer towards what is most useful for that goal.

The response will contain a direct answer to the question and relevant excerpts and sources from the documentation.

Use this mechanism when the answer is not explicitly present in the current page, you need clarification or additional context, or you want to retrieve related documentation sections.
