> For the complete documentation index, see [llms.txt](https://docs.maiagent.ai/llms.txt). Markdown versions of documentation pages are available by appending `.md` to page URLs; this page is available as [Markdown](https://docs.maiagent.ai/tech/quickstart/tu-xiang-bian-shi-zhi-yuan.md).

# 圖像辨識支援

視覺語言模型（VLM）與傳統 OCR 的差異，以及 MaiAgent 在對話附件、知識庫文件與知識庫圖片檔三個地方如何處理圖片。

## 什麼是 VLM（Vision Language Model）？ <a href="#what-is-vlm" id="what-is-vlm"></a>

視覺語言模型（Vision Language Model，VLM）能同時理解圖片與文字。它不只「看見」圖片中的字，還能理解圖中的物件、表格、圖表與彼此的關係，並依圖片內容回答問題或產生描述。

## VLM 與傳統 OCR 的比較 <a href="#vlm-vs-ocr" id="vlm-vs-ocr"></a>

傳統的光學字元辨識（OCR）專注於從圖片中找出並擷取文字，對圖片的整體語意與非文字內容幫助有限。

| 特性   | 傳統 OCR    | VLM             |
| ---- | --------- | --------------- |
| 主要功能 | 從圖片擷取文字   | 理解圖片內容，並與文字資訊結合 |
| 處理範圍 | 只處理文字     | 文字、物件、圖表、版面     |
| 理解層次 | 字元辨識      | 場景理解、關係推斷、上下文   |
| 常見用途 | 文件掃描、文字擷取 | 圖片問答、圖表解讀、圖片描述  |

## MaiAgent 在哪些地方處理圖片 <a href="#where-images-are-handled" id="where-images-are-handled"></a>

### 對話中的圖片附件 <a href="#chat-attachments" id="chat-attachments"></a>

使用者在對話中上傳圖片時，圖片會直接交給 Agent 的大型語言模型判讀，因此 Agent 必須使用支援多模態的模型。模型不支援時，回覆會顯示「您選的 LLM 不支援多模態。」。哪些模型支援多模態，見使用者手冊的 [模型調用權限](https://docs.maiagent.ai/org/model-access)。

單次請求送進模型的圖片數有上限，由平台設定（預設 20 張）。用 API 上傳時附件的 `type` 要填 `image`，圖片才會以視覺輸入送進模型；填 `other` 會改走文件解析流程，見 [訊息附件怎麼被處理](/tech/api-integration/api_knowledge.md#attachment-processing)。

### 知識庫文件內的圖片與掃描頁 <a href="#document-images" id="document-images"></a>

PDF、Word、PowerPoint 等文件中的圖片、圖表與掃描頁，要靠解析器讀出內容才能被檢索：

* **Vision Parser**：由雲端多模態模型逐頁讀取文件，圖片、圖表與掃描頁的內容會轉成文字。
* **MaiAgent Parser (Offline)**：設定多模態模型後，PDF 會加做 OCR 與圖片描述，適合需要離線運作的環境。
* **MaiAgent Parser**：擷取內嵌圖片，但不把圖片內容轉成文字。知識庫使用支援多模態的 Embedding 模型時，擷取出的圖片也會以圖片本身建立向量索引。

上傳時選擇 <mark style="color:blue;">自動偵測（預設）</mark>，系統會為掃描檔與含圖片的文件改用上述可處理圖片的解析器。各解析器的差異與計費見 [Parser 解析工具](/tech/quickstart/parser.md)。

### 知識庫中的圖片檔 <a href="#image-files" id="image-files"></a>

圖片檔（`.png`、`.jpg`、`.jpeg`、`.gif`、`.webp`、`.tiff`）可以直接上傳到知識庫，以圖片本身建立向量索引。這需要知識庫使用支援多模態的 Embedding 模型（MaiAgent 雲端為 Cohere Embed V4.0），否則該檔案會解析失敗。見 [Embedding 模型](/tech/quickstart/embedding.md)。

{% hint style="info" %}
三種情境各自獨立：Agent 的模型支援多模態，不代表知識庫能處理圖片檔；反之亦然。規劃時請分別確認 Agent 的模型、知識庫的解析器與 Embedding 模型。
{% endhint %}


---

# Agent Instructions
This documentation is published with GitBook. GitBook is the documentation platform designed so that both humans and AI agents can read, navigate, and reason over technical content effectively. Learn more at gitbook.com.

## Querying This Documentation
If you need additional information that is not directly available in this page, you can query the documentation dynamically by asking a question.

Perform an HTTP GET request on the current page URL with the `ask` query parameter, and the optional `goal` query parameter:

```
GET https://docs.maiagent.ai/tech/quickstart/tu-xiang-bian-shi-zhi-yuan.md?ask=<question>&goal=<endgoal>
```

`ask` is the immediate question: it should be specific, self-contained, and written in natural language.
`goal` is optional and describes the broader end goal you are ultimately trying to accomplish on behalf of the user. GitBook uses it to tailor the answer towards what is most useful for that goal.

The response will contain a direct answer to the question and relevant excerpts and sources from the documentation.

Use this mechanism when the answer is not explicitly present in the current page, you need clarification or additional context, or you want to retrieve related documentation sections.
