> For the complete documentation index, see [llms.txt](https://docs.maiagent.ai/llms.txt). Markdown versions of documentation pages are available by appending `.md` to page URLs; this page is available as [Markdown](https://docs.maiagent.ai/maiagent-user-guide/en/org/credits.md).

# Credits Billing System

MaiAgent's Credits billing system — one balance covers every AI feature.

{% hint style="success" %}
✨ **MaiAgent bills with Credits** — one balance covers conversations, speech, image generation, document parsing, and every other AI feature. Models have tiered rates, and usage is transparent and traceable.
{% endhint %}

<figure><img src="https://1360999650-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2F6v6TNkkOQVfRYfcNirHL%2Fuploads%2Fgit-blob-ddb6a71dfed0998f8b18369fa94f0875e72b8442%2Fcredit-illustration-1-hero.png?alt=media" alt=""><figcaption><p>Credits billing — one balance covers every AI feature</p></figcaption></figure>

## Three Benefits of Credits <a href="#credits-highlights" id="credits-highlights"></a>

### 1. One balance for every AI feature <a href="#one-quota-for-everything" id="one-quota-for-everything"></a>

**All features are available by default, and each has its own Credits rate** — one balance covers all AI services. There is no need for separate purchases or contracts across multiple platforms, making **billing easier to understand and budgets easier to control**.

<figure><img src="https://1360999650-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2F6v6TNkkOQVfRYfcNirHL%2Fuploads%2Fgit-blob-c4dfe904ceee4ce7bfcd6656895e75aee384496d%2Fcredit-illustration-2-all-in-one.png?alt=media" alt=""><figcaption><p>One Credits balance supports conversations, speech, image generation, document parsing, web search, and tool calls</p></figcaption></figure>

### 2. Tiered model pricing puts you in control of your budget <a href="#multi-model-budget-control" id="multi-model-budget-control"></a>

Costs can vary by **dozens of times** between AI models. With Credits, those differences are reflected directly in what you pay:

* Routine customer service FAQs → choose **Claude 4.5 Haiku** to save costs
* Complex analysis and deep research → choose **Claude 4.6 Opus** for the highest quality
* Assign different models to different AI assistants. Credits are deducted at each model's tier rate, so you decide how to allocate your budget

### 3. Transparent tracking makes reconciliation easier <a href="#transparent-tracking" id="transparent-tracking"></a>

Every Credits charge is tied to the **model, feature type, and AI assistant** used. Look up charges in real time in the admin panel, filter by any dimension, and export a CSV for internal reports. **You always know where your money went.**

***

## What are Credits? <a href="#what-is-credits" id="what-is-credits"></a>

Credits are MaiAgent's **universal billing unit**. Each AI conversation, image generation, speech processing, document parsing, and similar action consumes Credits according to the model and feature used.

### Core principles <a href="#core-principles" id="core-principles"></a>

* **Different models, different rates**: lightweight models cost less, while flagship models cost proportionally more — you decide what each task is worth
* **One billing unit for all features**: text, images, speech, document parsing, web search, and more are all measured in Credits
* **Transparent tracking**: check your Credits balance and individual usage records in real time in the admin panel, with CSV export
* **First in, first out**: Credits are deducted automatically in FIFO order, with **Credits that expire sooner used first**, so they do not go to waste

***

## Credits rates for major models <a href="#model-rates" id="model-rates"></a>

Rates for representative models from each provider, with **Input and Output priced separately** (per 1 million tokens):

| Model                 | Input Credits / 1M tokens | Output Credits / 1M tokens | Best for                                              |
| --------------------- | ------------------------- | -------------------------- | ----------------------------------------------------- |
| **Claude 4.6 Opus**   | 99,000                    | 495,000                    | Deep research, quality writing, complex reasoning     |
| **Claude 4.6 Sonnet** | 59,400                    | 297,000                    | Complex analysis, content creation, professional Q\&A |
| **Claude 4.5 Haiku**  | 19,800                    | 99,000                     | Routine customer service, FAQs, simple Q\&A           |
| **GPT-5**             | 24,750                    | 198,000                    | Advanced analysis, long documents                     |
| **GPT-5-mini**        | 4,950                     | 39,600                     | General Q\&A, simple tasks                            |
| **Gemini 3.1 Pro**    | 39,600                    | 237,600                    | Long context, complex analysis                        |
| **Gemini 3 Flash**    | 9,900                     | 59,400                     | General Q\&A, multimodal tasks                        |

{% hint style="info" %}
**Example usage**: A roughly 500-character customer service conversation with the commonly used **Claude 4.5 Haiku** consumes about **25–50 Credits**. A purchase of 1,000,000 Credits can support tens of thousands of customer service conversations — and the same balance also covers speech, image generation, and document parsing.
{% endhint %}

The platform also supports image generation, speech-to-text, text-to-speech, document parsing, web search, and more. **Everything is billed in Credits, with no separate purchase required**. For the full rate card, go to <mark style="color:blue;">**Organization Settings**</mark> → <mark style="color:blue;">**Credits**</mark> → <mark style="color:blue;">**Pricing**</mark>.

***

## Estimated tokens by plan <a href="#plan-token-reference" id="plan-token-reference"></a>

The same Credits balance processes different numbers of tokens depending on the model:

| Plan         | Monthly Credits | Claude 4.5 Haiku (approx.) | Claude 4.6 Sonnet (approx.) | Claude 4.6 Opus (approx.) |
| ------------ | --------------- | -------------------------- | --------------------------- | ------------------------- |
| Standard     | 500,000         | \~11.48 million tokens     | \~3.83 million tokens       | \~0.77 million tokens     |
| Professional | 5,000,000       | \~114.78 million tokens    | \~38.26 million tokens      | \~7.65 million tokens     |
| Enterprise   | 10,000,000      | \~229.57 million tokens    | \~76.52 million tokens      | \~15.30 million tokens    |

***

## Credits top-up options <a href="#topup-plans" id="topup-plans"></a>

| Credits   | Claude 4.5 Haiku (approx.) |
| --------- | -------------------------- |
| 50,000    | \~1.15 million tokens      |
| 500,000   | \~11.48 million tokens     |
| 1,000,000 | \~22.96 million tokens     |
| 2,500,000 | \~57.39 million tokens     |

{% hint style="info" %}
**Expiration rules**\
• Credits from the **monthly plan allocation** expire at the end of that month\
• **Purchased** Credits are valid for **12 months** from the purchase date; unused Credits expire after that
{% endhint %}

***

## How plan allocations and purchased Credits work together <a href="#plan-vs-topup" id="plan-vs-topup"></a>

A subscription's **monthly allocation** and **purchased Credits** are separate but complementary balances. The system handles deduction order automatically; you do not need to allocate them manually.

### Comparison <a href="#plan-vs-topup-comparison" id="plan-vs-topup-comparison"></a>

| Item         | Plan allocation (monthly)                                   | Purchased Credits                                      |
| ------------ | ----------------------------------------------------------- | ------------------------------------------------------ |
| **Source**   | Issued automatically with your plan                         | Purchased whenever you need them                       |
| **Issued**   | At the start of each month                                  | Available immediately after purchase                   |
| **Validity** | **Expires at month-end**; unused balance does not roll over | Valid for **12 months** from purchase                  |
| **Best for** | Regular daily usage                                         | Supplementing a monthly allocation / unexpected spikes |

### Deduction order: earliest expiration first (FIFO) <a href="#fifo-deduction" id="fifo-deduction"></a>

The system uses **FIFO (first in, first out)** logic and deducts Credits in **expiration-date order**:

```
Daily use → deduct this month's plan allocation first (expires at month-end)
          → once exhausted → automatically deduct purchased Credits (by purchase date)
          → at month-end → next month's allocation is issued
```

This prevents **purchased Credits from being consumed while the monthly allocation expires unused**.

{% hint style="success" %}
**Example**: You have an Enterprise plan (10,000,000 Credits per month) and buy another 1,000,000 Credits midway through the month.

* The system uses the remaining monthly allocation first (it expires at month-end)
* It uses the additional 1,000,000 Credits only after that allocation is exhausted (they remain available for up to 12 months)
  {% endhint %}

### Why have two types? <a href="#why-two-types" id="why-two-types"></a>

* **Subscription plan allocation** = a predictable day-to-day budget that renews monthly
* **Purchased Credits** = flexible backup for sudden growth in demand
* Together, they help you **control costs and handle fluctuations**

***

## Traffic fees for self-hosted AI Gateway models <a href="#ai-gateway-traffic-fee" id="ai-gateway-traffic-fee"></a>

When an organization uses an AI Gateway model with its own credentials or hosting, MaiAgent does not charge the platform's LLM token rates. Instead, it charges a traffic fee based on the total tokens in the call. Total tokens include input, output, prompt cache writes, and prompt cache reads; the applicable percentage is set in your organization settings.

{% hint style="info" %}
The traffic fee is calculated as “total tokens × your organization's traffic rate” and deducted in Credits. MaiAgent-hosted models continue to use their standard Input/Output rates and are not covered by this section.
{% endhint %}

### Example scenario <a href="#traffic-fee-scenario" id="traffic-fee-scenario"></a>

Suppose an IT department connects its self-hosted model to AI Gateway. At month-end, a finance team member wants to check the cost of those calls. They first look up the organization's traffic rate under Pricing, then filter Consumption Records by platform fees. The AI Gateway traffic fee entries let them reconcile self-hosted model traffic separately from hosted model charges.

### View traffic rates and consumption <a href="#view-traffic-fee" id="view-traffic-fee"></a>

1. Go to <mark style="color:blue;">**Organization Settings**</mark> → <mark style="color:blue;">**Credits**</mark> → <mark style="color:blue;">**Pricing**</mark>.
2. Under “Platform Fees,” find <mark style="color:blue;">**AI Gateway Traffic Fee**</mark> and check the rate description, “total tokens × percentage.”
3. Switch to <mark style="color:blue;">**Consumption Records**</mark> and filter by “Platform Fees” to see the actual charges.
4. Hover over an AI Gateway traffic fee entry to check input, output, cache writes, cache reads, and the billing percentage.

<figure><img src="https://1360999650-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2F6v6TNkkOQVfRYfcNirHL%2Fuploads%2Fgit-blob-589731f9acdea9f28a8b1c8605e68e6b33f18548%2Fcredits-ai-gateway-pricing.png?alt=media" alt="Credits Pricing tab"><figcaption><p>Check Credits rates for each service under Pricing.</p></figcaption></figure>

<figure><img src="https://1360999650-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2F6v6TNkkOQVfRYfcNirHL%2Fuploads%2Fgit-blob-24638d9ed0a64102b9d9097ba106fef27c141f93%2Fcredits-ai-gateway-consumption.png?alt=media" alt="Credits Consumption Records tab"><figcaption><p>Filter actual charges by type, service category, date, or keyword in Consumption Records.</p></figcaption></figure>

{% hint style="warning" %}
Traffic fee entries appear in Consumption Records only after you use an AI Gateway model hosted by your organization (BYOK or self-hosted). Pricing shows the current rates, while historical records preserve the billing details from each transaction.
{% endhint %}

***

## Credits dashboard <a href="#management-panel" id="management-panel"></a>

Go to <mark style="color:blue;">**Organization Settings**</mark> → <mark style="color:blue;">**Credits**</mark> to access the dashboard's five tabs:

### Overview <a href="#overview-tab" id="overview-tab"></a>

* See your total Credits balance and soon-to-expire Credits in real time
* Consumption trend chart: choose any date range; the card also shows total consumption and the daily average for that range (see [Choose a date range for consumption trends](#trend-date-range))
* Notifications for soon-to-expire Credits
* Low-balance alert settings (in-app notifications + Email; see [Low-balance alerts](#low-balance-alert))

<figure><img src="https://1360999650-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2F6v6TNkkOQVfRYfcNirHL%2Fuploads%2Fgit-blob-3f607a177e928d881df79b615c86203016151fb9%2Fcredit-1-overview.png?alt=media" alt=""><figcaption><p>Overview tab: current Credits balance and expiration alerts</p></figcaption></figure>

#### Choose a date range for consumption trends <a href="#trend-date-range" id="trend-date-range"></a>

The date-range picker is at the upper right of the “Consumption Trends” card and defaults to the last 30 days. Check how many Credits you used last month or over any period without exporting a report and adding the numbers yourself.

1. Go to <mark style="color:blue;">**Organization Settings**</mark> → <mark style="color:blue;">**Credits**</mark> → <mark style="color:blue;">**Overview**</mark>.
2. Click the date range at the upper right of the “Consumption Trends” card. In the same calendar panel, select the start date and then the end date. The left side also offers four shortcuts: <mark style="color:blue;">**Last 7 Days**</mark>, <mark style="color:blue;">**Last 30 Days**</mark>, <mark style="color:blue;">**This Month**</mark>, and <mark style="color:blue;">**Last Month**</mark>.
3. The chart immediately redraws daily consumption for that period. <mark style="color:blue;">**Period Consumption**</mark> and <mark style="color:blue;">**Daily Average**</mark> on the card are recalculated, and the bottom of the card shows “N days total.”

<figure><img src="https://1360999650-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2F6v6TNkkOQVfRYfcNirHL%2Fuploads%2Fgit-blob-06ce169732c6c58365c6df06d785fb6497ada70f%2Fcredits-trend-range-picker.png?alt=media" alt="Date-range picker on the consumption trends card, with Last 7 Days, Last 30 Days, This Month, and Last Month shortcuts on the left"><figcaption><p>Consumption trends: select start and end dates in the calendar or use a shortcut on the left</p></figcaption></figure>

{% hint style="info" %}

* Start and end dates are full days in Taiwan time, inclusive; the end date cannot be later than today.
* Ranges of up to 90 days show one point per day. Longer ranges show one point per week, with the week's start date on the horizontal axis. “Daily Average” always uses the number of days in the range, regardless of weekly display.
* The chart shows an empty state if the range has no consumption.
  {% endhint %}

### Deposits <a href="#deposits-tab" id="deposits-tab"></a>

* Shows every source of Credits: plan allocations, purchases, gifts, and more
* Shows validity periods (monthly allocations expire at month-end; purchased Credits remain valid for 12 months from purchase)
* Uses **FIFO deductions** (earliest expiration first), with the remaining balance of each allocation
* Supports **CSV export** for bookkeeping and annual reviews

<figure><img src="https://1360999650-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2F6v6TNkkOQVfRYfcNirHL%2Fuploads%2Fgit-blob-841869bbc95eef21c7f52365b382ceeb69cce267%2Fcredit-2-deposits.png?alt=media" alt=""><figcaption><p>Deposits tab: Credits sources, validity periods, and remaining balances</p></figcaption></figure>

### Consumption Records <a href="#consumption-tab" id="consumption-tab"></a>

* Filter usage details by date, type, AI assistant, and other dimensions
* Each charge shows the model used, Credits consumed, and related event
* The summary below the filters shows total spending, total deposits, record count, and average daily spending for the filtered results (see [Filtered-results summary](#consumption-stats))
* Export a **CSV** for records or deeper analysis

<figure><img src="https://1360999650-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2F6v6TNkkOQVfRYfcNirHL%2Fuploads%2Fgit-blob-9265ca5f54d3d1ca372e22bb884fccbd26408e54%2Fcredit-3-consumption.png?alt=media" alt=""><figcaption><p>Consumption Records tab: each charge identifies its model, AI assistant, and event</p></figcaption></figure>

#### Filtered-results summary <a href="#consumption-stats" id="consumption-stats"></a>

At month-end, finance staff can find “how many Credits did we spend last month?” by setting the date range in Consumption Records to the entire month. The summary row below the filters shows the totals directly, without exporting and summing a CSV.

After you set a transaction type, service category, date range, or search keyword, a summary row appears between the filters and the details table and indicates the scope of the current summary:

<table><thead><tr><th width="150">Field</th><th>Description</th></tr></thead><tbody><tr><td>Total spending</td><td>Sum of all charges in the filtered range (including platform processing fees), matching the CSV export total. Shows 0 when the transaction type is “Deposit.”</td></tr><tr><td>Total deposits</td><td>Sum of all top-ups, allocations, and refunds in the filtered range. Shows 0 when the transaction type is “Spending.”</td></tr><tr><td>Records</td><td>Matches “N records total” at the bottom of the details table; multiple charges for the same message count as one record.</td></tr><tr><td>Daily average spending</td><td>Total spending ÷ number of days in the date range (including both endpoints). Shows “–” when no date range is set.</td></tr></tbody></table>

<figure><img src="https://1360999650-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2F6v6TNkkOQVfRYfcNirHL%2Fuploads%2Fgit-blob-79df7a40390511c20904035bb55b1d15ad59a342%2Fcredits-consumption-stats-row.png?alt=media" alt="Summary row below the filters in Consumption Records showing total spending, total deposits, record count, and daily average spending"><figcaption><p>Set the date range to a full month to see that month's total and daily average spending directly</p></figcaption></figure>

{% hint style="info" %}
The summary is calculated from all filtered results, regardless of pagination or page size; it updates whenever a filter changes. If the summary temporarily fails to load, the details table still appears. Click Retry in the summary row to recalculate it.
{% endhint %}

### Pricing <a href="#pricing-tab" id="pricing-tab"></a>

* Lists the rates for all models and features
* Covers all billing categories, including LLM conversations, speech, knowledge bases, crawlers, tools, AI assistant execution, and Avatar

<figure><img src="https://1360999650-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2F6v6TNkkOQVfRYfcNirHL%2Fuploads%2Fgit-blob-081859f5a3b6da7807f239ffd940f0ef69ca22f8%2Fcredit-4-pricing.png?alt=media" alt=""><figcaption><p>Pricing tab: full rate card for all models and features</p></figcaption></figure>

#### View storage fees for conversation attachments and derived files <a href="#attachment-and-derived-storage" id="attachment-and-derived-storage"></a>

Files uploaded in conversations use “Conversation Attachment Storage.” Copies produced when the system parses files, such as format conversions, reprocessed files, or speech transcriptions, use “Derived File Storage.” Both are billed in MB per month, with Credits deducted daily based on actual usage. The original file is billed under its corresponding storage item, without an additional derived-file charge.

{% hint style="info" %}
The retention period for conversation attachments depends on system settings; check the description under “Conversation Attachment Storage” for the current period. A new message in the conversation restarts the retention countdown. After the retention period, the attachment and its processing copies are automatically deleted and billing stops. Deletion cannot be undone, so download files you need to keep before the deadline.
{% endhint %}

1. Go to <mark style="color:blue;">**Organization Settings**</mark> → <mark style="color:blue;">**Credits**</mark>.
2. Switch to <mark style="color:blue;">**Pricing**</mark>.
3. Expand <mark style="color:blue;">**Knowledge Base and Vector Retrieval System**</mark>.
4. Check the Credits per unit and billing unit for <mark style="color:blue;">**Conversation Attachment Storage**</mark> and <mark style="color:blue;">**Derived File Storage**</mark>. Hover over the information icon for retention and billing details.

<figure><img src="https://1360999650-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2F6v6TNkkOQVfRYfcNirHL%2Fuploads%2Fgit-blob-cdaf245bd3a3c0393c93c541287bd6f7b14f5f02%2Fcredits-attachment-storage-pricing.png?alt=media" alt="Conversation Attachment Storage and Derived File Storage entries under Credits Pricing"><figcaption><p>Expand “Knowledge Base and Vector Retrieval System” to compare the two file-storage rates.</p></figcaption></figure>

**Example scenario**

A customer service manager at MaiAgent regularly has visitors upload screenshots of issues in conversations. During month-end reconciliation, she opens Credits Pricing and expands “Knowledge Base and Vector Retrieval System” to see which rates apply to conversation attachments and parsed copies. After reading the retention details, she downloads screenshots that must be kept before the deadline. Attachments no longer needed are automatically deleted when their retention period ends, so they stop accumulating storage fees.

### Quota Allocation <a href="#quota-tab" id="quota-tab"></a>

* Admins can set a personal Credits spending limit and validity period for each member
* Allocate, edit, or disable quotas, and control the “Quota Required” switch
* Search and filter members by status (active / expired / unassigned)

See [Managing member Credits quotas](/maiagent-user-guide/en/org/credits/member-credit-quota.md).

***

## Low-balance alerts <a href="#low-balance-alert" id="low-balance-alert"></a>

When your organization's Credits balance falls below your chosen threshold, the system **proactively alerts you** so you can top up before service is interrupted. Alerts appear as **in-app notifications**; paid plans also **send Email automatically**, so you can be notified without signing in to the admin panel.

{% hint style="info" %}
In-app notifications are available on every plan; **Email notifications are for paid plans only**. Alerts do not apply to organizations billed through AWS Marketplace. You need access to the Credits page to configure and view alerts.
{% endhint %}

### Notification channels <a href="#alert-channels" id="alert-channels"></a>

| Channel                 | Description                                                                           |
| ----------------------- | ------------------------------------------------------------------------------------- |
| **In-app notification** | Creates an announcement in the admin panel when the balance falls below the threshold |
| **Email notification**  | Paid plans also email members with the Owner role and any custom recipients           |

Both channels share one **enable switch** and **alert threshold**. Turning off alerts stops both in-app and Email notifications.

### Who receives Email <a href="#alert-recipients" id="alert-recipients"></a>

* **All members with the Owner role**: automatically included and always emailed when alerts are on
* **Custom recipients**: add up to **5** additional Email addresses for your organization
* If an address appears in both groups, the system **removes duplicates**, so it receives only one email

### Configure alerts <a href="#configure-alert" id="configure-alert"></a>

1. Go to <mark style="color:blue;">**Organization Settings**</mark> → <mark style="color:blue;">**Credits**</mark> → <mark style="color:blue;">**Overview**</mark>.
2. On the <mark style="color:blue;">**Balance Alerts**</mark> card, click <mark style="color:blue;">**Settings**</mark>.
3. Turn on <mark style="color:blue;">**Enable Balance Alerts**</mark> and enter an **alert threshold**.
4. Under <mark style="color:blue;">**Notification Recipients**</mark>, add Email addresses for additional recipients (up to 5). You can paste multiple comma-separated addresses to split them automatically.
5. Save to finish.

### Trigger conditions <a href="#alert-trigger" id="alert-trigger"></a>

* Alerts trigger automatically when the **balance is below your chosen threshold**, sending both in-app and Email notifications
* Each wallet receives **at most one email every 24 hours** to avoid repeated messages
* The branded HTML alert email shows the current **balance** and **threshold** and includes a “View Credit Usage Details” button; contact <sales@maiagent.ai> to purchase more Credits

***

## Prompt Cache usage <a href="#prompt-cache-usage" id="prompt-cache-usage"></a>

### What is Prompt Cache, and how does it save money? <a href="#what-is-prompt-cache" id="what-is-prompt-cache"></a>

Every time an AI assistant responds, it sends the **full conversation context** to the model again — including its System Prompt, knowledge base content, and previous messages. Much of this content **repeats** across conversations with the same assistant: the System Prompt stays the same, and knowledge base excerpts often recur.

**Prompt Cache** reduces the cost of repeated content. When processing cacheable content, the model may first **write it to a cache**. Later requests with identical content can **read from that cache** while it remains valid. Cache writes and reads have separate rates, and reads usually cost less than ordinary input. Whether this saves Credits depends on the setup cost, the number of subsequent cache hits, and the model's rates.

Conversations using Prompt Cache may show these charges under the “AI Text Generation and Understanding Models” billing category:

| Charge                 | Description                                                                             |
| ---------------------- | --------------------------------------------------------------------------------------- |
| **LLM Cache Creation** | Number of input tokens written to the cache, charged at the model's cache creation rate |
| **LLM Cache Read**     | Number of input tokens read from the cache, charged at the model's cache read rate      |

{% hint style="info" %}
The model and platform settings determine **whether Prompt Cache is used**; you do not need to configure it manually. Billing details list only cache items that actually occurred and had more than 0 tokens. You might therefore see only “LLM Cache Creation” or “LLM Cache Read,” or neither.
{% endhint %}

### Example scenario <a href="#cache-scenario" id="cache-scenario"></a>

To check whether caching has recently reduced costs for AI assistants, an organization admin can first go to <mark style="color:blue;">**Organization Settings**</mark> → <mark style="color:blue;">**Usage Statistics**</mark> → <mark style="color:blue;">**Cache Savings**</mark>. To investigate an individual request, go to <mark style="color:blue;">**Credits**</mark> → <mark style="color:blue;">**Consumption Records**</mark>, expand a conversation charge with cache items, and compare cache creation or read counts and Credits.

### When does caching save the most? <a href="#cache-best-scenarios" id="cache-best-scenarios"></a>

Caching saves the cost of **repeated input**, so **more repeated content and more cache hits lead to greater savings**:

* **Assistants with long role instructions (System Prompts)**: For example, a customer support assistant may include complete response guidelines and product descriptions. When later requests hit the same cache, more content can be billed at the read rate.
* **Frequently used assistants**: The more often later requests hit the same content while the cache is valid, the more likely the savings are to offset the cost of creating the cache.
* **Long, multi-turn conversations**: If the model can reuse earlier conversation content, later turns may generate more cache read usage.

### Where to view cache usage <a href="#where-to-view-cache-usage" id="where-to-view-cache-usage"></a>

The platform offers three levels of detail: **individual billing entries**, **organization-wide statistics**, and **token details for a single conversation**.

#### 1. Credits → Consumption records and pricing <a href="#cache-in-credits" id="cache-in-credits"></a>

Use this view to reconcile individual charges and check each cache charge:

1. Go to <mark style="color:blue;">**Organization Settings**</mark> → <mark style="color:blue;">**Credits**</mark> → <mark style="color:blue;">**Consumption Records**</mark>.
2. Find a conversation charge (chat.message) that used the cache and click the expand arrow on the left. Depending on actual usage, the details show “LLM Cache Creation” or “LLM Cache Read,” together with the Credits charged, quantity, and corresponding model.
3. Switch to the <mark style="color:blue;">**Pricing**</mark> tab to see each model's cache rates (priced per 1,000 tokens).

<figure><img src="https://1360999650-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2F6v6TNkkOQVfRYfcNirHL%2Fuploads%2Fgit-blob-42d20a806c837c95b5564775545c822794d788ef%2Fcache-01-consumption-records.png?alt=media" alt="Cache billing items in consumption records"><figcaption><p>Consumption records: Expand a conversation charge to view the cache billing items actually incurred.</p></figcaption></figure>

#### 2. Usage Statistics → Billing Categories tab <a href="#cache-in-usage-stats" id="cache-in-usage-stats"></a>

Use this view to see the organization's overall share of cache consumption and the amount saved:

1. Go to <mark style="color:blue;">**Organization Settings**</mark> → <mark style="color:blue;">**Usage Statistics**</mark>.
2. Change the metric dropdown to <mark style="color:blue;">**Credits**</mark>. The <mark style="color:blue;">**Billing Categories**</mark> tab appears only after this change.
3. In <mark style="color:blue;">**Usage Overview**</mark>, if net savings for the selected period exceed 0, <mark style="color:blue;">**Cache Savings**</mark> displays the number of Credits saved.
4. Switch to <mark style="color:blue;">**Billing Categories**</mark> to view Credit consumption by category in a pie chart, ranked bar chart, and details table. When there is relevant usage, cache creation or read items appear under “AI Text Generation and Understanding Models.”

<figure><img src="https://1360999650-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2F6v6TNkkOQVfRYfcNirHL%2Fuploads%2Fgit-blob-0e4cb43b0bda7d8a1a165c3e4eb8323cf15de696%2Fcache-02-savings-card.png?alt=media" alt="Cache Savings figure in Usage Overview"><figcaption><p>Usage Overview: “Cache Savings” appears when net savings exceed 0.</p></figcaption></figure>

<figure><img src="https://1360999650-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2F6v6TNkkOQVfRYfcNirHL%2Fuploads%2Fgit-blob-15e401a076373ed369f3f0ad4409aa373f80784c%2Fusage-stats-credit-categories.png?alt=media" alt="LLM Cache Creation and Read billing items on the Billing Categories tab of Usage Statistics"><figcaption><p>Usage Statistics, Billing Categories: LLM Cache Creation and Read appear under “AI Text Generation and Understanding Models.”</p></figcaption></figure>

{% hint style="warning" %}
The **Billing Categories tab appears only when the metric is set to Credits**, and only for organizations using Credit billing. If you cannot find the tab, check the metric dropdown first.
{% endhint %}

#### 3. AgentOps → Conversation Token Statistics <a href="#cache-in-agentops" id="cache-in-agentops"></a>

Use this view to analyze actual cache hits for **a single conversation**:

1. Go to <mark style="color:blue;">**AgentOps**</mark> → <mark style="color:blue;">**AI Assistant Monitoring**</mark>.
2. Open the details of any conversation record and find the <mark style="color:blue;">**Conversation Token Statistics**</mark> section.
3. If that message has cache usage, the section shows the token count for “LLM Cache Creation” or “LLM Cache Read,” according to what actually occurred. Items with no corresponding usage are hidden.

<figure><img src="https://1360999650-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2F6v6TNkkOQVfRYfcNirHL%2Fuploads%2Fgit-blob-0dec37ee809f2089857bee16844f8d08b641ae89%2Fcache-03-agentops-text-stats.png?alt=media" alt="Conversation Token Statistics in an AgentOps conversation record"><figcaption><p>Conversation record details: This conversation generated 7,399 LLM Cache Read tokens. There was no cache creation usage, so that item is not shown.</p></figcaption></figure>

### Frequently asked questions <a href="#cache-faq" id="cache-faq"></a>

#### Q: Why is there an “LLM Cache Creation” charge on my bill? <a href="#cache-faq-extra-charge" id="cache-faq-extra-charge"></a>

“LLM Cache Creation” means that content was written to the cache for that request and is billed separately at the model's cache creation rate. It is not limited to the first turn of a conversation: the cache may be created again after it expires or when cacheable content changes. Later requests that hit the cache are charged at the “LLM Cache Read” rate. Use <mark style="color:blue;">**Cache Savings**</mark> in <mark style="color:blue;">**Usage Statistics**</mark> to assess net savings for the selected period. If creation costs exceed read savings, this figure is not displayed.

#### Q: How can I tell how much the cache saved me? <a href="#cache-faq-savings" id="cache-faq-savings"></a>

Go to <mark style="color:blue;">**Organization Settings**</mark> → <mark style="color:blue;">**Usage Statistics**</mark> and check <mark style="color:blue;">**Cache Savings**</mark> in <mark style="color:blue;">**Usage Overview**</mark>. When net savings for the selected period exceed 0, the page shows the number of Credits saved. It does not show the figure when savings are 0, negative, or below display precision.

#### Q: Do I need to enable Prompt Cache myself? <a href="#cache-faq-enable" id="cache-faq-enable"></a>

**No.** Cache usage is determined automatically by the model's capabilities and platform settings. No additional action or configuration is required.

#### Q: Why do some conversations have no cache items? <a href="#cache-faq-missing" id="cache-faq-missing"></a>

Cache items appear only when the model reports the corresponding usage. If a request neither creates nor hits a cache, no such item appears in the details. This does not affect the AI assistant's normal response.

***

## Frequently asked questions <a href="#faq" id="faq"></a>

### Q: Can I use different models together? <a href="#faq-mixed-models" id="faq-mixed-models"></a>

**Yes.** Under Credit billing, **the corresponding Credits are deducted according to the tier of each model used**. Use an affordable Haiku model for customer support FAQs and Opus for in-depth analysis; you control how the budget is allocated.

### Q: How do I check my Credit usage? <a href="#faq-view-usage" id="faq-view-usage"></a>

Go to <mark style="color:blue;">**Organization Settings**</mark> → <mark style="color:blue;">**Credits**</mark>. **Overview** shows the current balance, **Credit History** shows each source of Credits and its expiry date, and **Consumption Records** shows individual usage details and supports CSV export.

### Q: What happens when I use up my monthly allowance? <a href="#faq-quota-exhausted" id="faq-quota-exhausted"></a>

When the monthly Credit allowance runs out, AI assistants stop responding. You can buy additional Credits at any time. **Purchased Credits take effect immediately** and remain valid for 12 months from the purchase date.

### Q: Can I still use the platform when my plan expires or I have no paid plan? <a href="#faq-no-active-plan" id="faq-no-active-plan"></a>

**Since June 9, 2026**, organizations without an active paid plan (including free plans, ended trials, and expired plans that have not been renewed) are prompted to upgrade when they try to perform an operation requiring AI computation. The platform retains all data for **90 days** from the plan's expiration date. During that period, only storage Credits are charged (settled when the plan is restored). Resubscribe within 90 days to restore everything; after that, the data is permanently deleted and cannot be recovered. See [Trial and Subscription Plans — What happens to data after a plan expires or is downgraded to Free?](/maiagent-user-guide/en/others/trial-and-plans.md#data-retention-policy).

### Q: Do Credits expire? <a href="#faq-expiration" id="faq-expiration"></a>

* **Monthly plan allowance** Credits expire at the end of the month.
* **Purchased** Credits remain valid for 12 months from the purchase date.

***

## Need help? Contact us <a href="#contact-sales" id="contact-sales"></a>

To learn how Credit billing can help control your company's budget, or to discuss plan options and additional Credits, contact us:

* **Sales team**: <sales@maiagent.ai>
* **Website**: <https://maiagent.ai>

Thank you for choosing MaiAgent. We will keep working to provide a more flexible, affordable, and dependable AI service.


---

# Agent Instructions
This documentation is published with GitBook. GitBook is the documentation platform designed so that both humans and AI agents can read, navigate, and reason over technical content effectively. Learn more at gitbook.com.

## Querying This Documentation
If you need additional information that is not directly available in this page, you can query the documentation by asking a question.

Perform an HTTP GET request on the following URL with the `ask` and `goal` query parameters:

```
GET https://docs.maiagent.ai/maiagent-user-guide/en/org/credits.md?ask=<question>&goal=<user_goal>
```

`ask` is the immediate question: it should be specific, self-contained, and written in natural language.
`goal` is what the user is ultimately trying to achieve, the reason they need the answer. Sharing it helps GitBook give you a better, more relevant answer. A goal is most helpful when it describes the outcome the user wants rather than restating the question. For example, with `ask=how do I create an API token`, a goal like `automate deployments from our CI pipeline` lets GitBook tailor the answer to that use case.

The response will contain a direct answer to the question and relevant excerpts and sources from the documentation.

Use this mechanism when the answer is not explicitly present in the current page, you need clarification or additional context, or you want to retrieve related documentation sections.
