> For the complete documentation index, see [llms.txt](https://docs.maiagent.ai/llms.txt). Markdown versions of documentation pages are available by appending `.md` to page URLs; this page is available as [Markdown](https://docs.maiagent.ai/maiagent-user-guide/en/build/ai-gateway.md).

# AI Gateway

Connect your organization's model providers in AI Gateway, manage the model catalog and fallback model groups, and monitor Gateway traffic and issues.

AI Gateway is where your organization centrally manages model access. Connect your organization's provider credentials (such as OpenAI or Amazon Bedrock) or a self-hosted OpenAI-compatible endpoint to create custom models, then use fallback model groups to specify which models to try in order when the primary model fails. Connected models can be used by AI assistants, MaiGPT, or applications integrated through an OpenAI-compatible API. This page explains the administrative interface for organization administrators and IT staff responsible for model governance. For application integration details, see [AI Gateway: OpenAI-Compatible API Integration](https://docs.maiagent.ai/tech/ai-gateway/ai-gateway-openai-compatible) in the technical guide.

{% hint style="info" %}
AI Gateway must be enabled for your organization before you can use it (an Enterprise feature, currently in Beta). Until it is enabled, AI Gateway will not appear in the left menu, and it will be grayed out in the “Advanced Features” section of <mark style="color:blue;">Organization Overview</mark>. To enable it, contact your sales representative (see [Advanced Features](/maiagent-user-guide/en/org/organization-overview.md#advanced-features)).
{% endhint %}

## Who Can Use It <a href="#permissions" id="permissions"></a>

| Action                                                                                    | Required Permission                                                                                                                                                                                                                                                |
| ----------------------------------------------------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ |
| See AI Gateway in the left menu and view its tabs                                         | The <mark style="color:blue;">AI Gateway Permission</mark> role permission (organization owners have it by default; for other roles, see [Enable AI Gateway for a Specific Role](/maiagent-user-guide/en/org/roles/role-permission.md#enable-ai-gateway-for-role)) |
| Add, edit, test, or delete provider credentials, custom models, and fallback model groups | Organization owner                                                                                                                                                                                                                                                 |
| Toggle <mark style="color:blue;">Available to Organization</mark> for managed models      | Organization owner (the same setting as on the [Model Access Permissions](/maiagent-user-guide/en/org/model-access.md) page)                                                                                                                                       |
| Create API Keys for calling the Gateway                                                   | <mark style="color:blue;">API Key Management Permission</mark>                                                                                                                                                                                                     |

Members who are not organization owners can view each tab, but the system rejects attempts to add, modify, delete, or test resources. Buttons such as <mark style="color:blue;">Add Provider</mark>, <mark style="color:blue;">Add Model</mark>, and <mark style="color:blue;">Add Fallback Model Group</mark> are disabled. Hovering over them displays “Only organization owners can manage provider credentials,” “Only organization owners can manage custom models,” or “Only organization owners can manage fallback rules.”

## How to Access It <a href="#how-to-enter" id="how-to-enter"></a>

Click <mark style="color:blue;">AI Gateway</mark> in the left menu. It opens the <mark style="color:blue;">Overview</mark> tab by default. There are five tabs:

| Tab                                                    | Purpose                                                            |
| ------------------------------------------------------ | ------------------------------------------------------------------ |
| <mark style="color:blue;">Overview</mark>              | Gateway traffic, spending, latency, and issues that need attention |
| <mark style="color:blue;">Model Providers</mark>       | Manage your organization's provider credentials                    |
| <mark style="color:blue;">Models</mark>                | The model catalog: managed and custom models                       |
| <mark style="color:blue;">Fallback Model Groups</mark> | Fallback chains used in order when the primary model fails         |
| <mark style="color:blue;">Integration Examples</mark>  | Code examples for calling the OpenAI-compatible API                |

If your organization has no providers yet, <mark style="color:blue;">Overview</mark> displays “Welcome to AI Gateway” with three steps: <mark style="color:blue;">Add Provider</mark> → enable the models you want in the <mark style="color:blue;">Model Catalog</mark> → <mark style="color:blue;">Add API Key</mark>. After you complete the first step, the overview becomes a live health dashboard.

## Overview <a href="#overview" id="overview"></a>

The overview summarizes your organization's Gateway traffic, including OpenAI-compatible API calls and AI assistant conversations using custom models. At the top right, select <mark style="color:blue;">Last 1 Hour</mark>, <mark style="color:blue;">Last 24 Hours</mark>, <mark style="color:blue;">Last 7 Days</mark>, or <mark style="color:blue;">Last 30 Days</mark>, or specify start and end times with the date picker (custom ranges can span up to 92 days).

| Section                                                                                 | Contents                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                             |
| --------------------------------------------------------------------------------------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| Status summary                                                                          | Displays “All systems normal · No pending items” or “N items need attention/action”; <mark style="color:blue;">Access Diagnostics</mark> is on the right, and you can jump to the “Needs Attention” card with one click when there are pending items                                                                                                                                                                                                                                                                                                                                                 |
| Metric cards                                                                            | <mark style="color:blue;">Requests</mark>, <mark style="color:blue;">Spending (credits)</mark>, <mark style="color:blue;">Tokens</mark>, <mark style="color:blue;">Error Rate</mark>, <mark style="color:blue;">Slowest Latency</mark>, and <mark style="color:blue;">Fallback Triggers</mark>, compared with the preceding period of the same length                                                                                                                                                                                                                                                |
| Providers                                                                               | Each provider has a status label: <mark style="color:blue;">Healthy</mark>, <mark style="color:blue;">High Rate Limiting</mark> (upstream error rate above 5% over the last 24 hours), or <mark style="color:blue;">Setup Required</mark> (credentials disabled). Click to see the upstream error rate over the last 24 hours and suggested actions, and to <mark style="color:blue;">Test Connection</mark> or <mark style="color:blue;">Manage Credentials</mark>; <mark style="color:blue;">Create Fallback Rule</mark> is also available for <mark style="color:blue;">High Rate Limiting</mark> |
| <mark style="color:blue;">Usage Trends</mark>, <mark style="color:blue;">Latency</mark> | Trend charts for request count or cost; <mark style="color:blue;">Time to First Token (TTFT)</mark>, <mark style="color:blue;">Total Latency</mark>, and <mark style="color:blue;">Upstream Latency</mark> each show <mark style="color:blue;">Typical</mark> (p50) and <mark style="color:blue;">Slowest</mark> (p95) values                                                                                                                                                                                                                                                                        |
| <mark style="color:blue;">Usage Breakdown</mark>                                        | Lists requests, fallback counts, and spending by <mark style="color:blue;">Provider</mark> or <mark style="color:blue;">Model</mark> (providers also show error rates); expand a model to see the API Keys calling it                                                                                                                                                                                                                                                                                                                                                                                |
| <mark style="color:blue;">Needs Attention</mark>                                        | A provider's upstream error rate exceeds 5% over the last 24 hours; a member bound to an API Key has used at least 90% of their quota, or a disabled key has received requests in the last 24 hours; a fallback rule has triggered at least twice as often this month as during the same period last month                                                                                                                                                                                                                                                                                           |

Use <mark style="color:blue;">Access Diagnostics</mark> to investigate cases where someone cannot access a model. Select an API Key (and optionally a model), then click <mark style="color:blue;">Diagnose</mark>. The system checks that the key has not been revoked, is enabled, and has not expired; that the member and user are active; that the organization has Gateway enabled and can use the model; and that the key has access to the model. It highlights the first check that fails.

## Add a Provider <a href="#add-provider" id="add-provider"></a>

In <mark style="color:blue;">Model Providers</mark>, click <mark style="color:blue;">Add Provider</mark> and follow the four-step wizard. You can add multiple entries for the same provider type (for example, Bedrock in different regions).

{% stepper %}
{% step %}

### Choose a Provider <a href="#pick-provider" id="pick-provider"></a>

Choose a provider type. Providers whose cards display <mark style="color:blue;">Automatic Discovery Supported</mark> can automatically list available models after creation. For those labeled <mark style="color:blue;">Manual Model Addition Required</mark>, you will need to add models individually in the model catalog later.
{% endstep %}

{% step %}

### Connection Credentials <a href="#provider-credentials" id="provider-credentials"></a>

Enter a <mark style="color:blue;">Display Name</mark> (for example, “Bedrock US East”) and the provider's credential fields, then click <mark style="color:blue;">Create and Discover</mark>. Credentials are saved at this step, so closing the wizard afterward will not discard them.

| Provider          | Required Fields                                                                                                                                         | Optional Fields                                                                                | Automatic Model Discovery |
| ----------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------- | ---------------------------------------------------------------------------------------------- | ------------------------- |
| OpenAI            | <mark style="color:blue;">API Key</mark>                                                                                                                | —                                                                                              | Supported                 |
| Anthropic         | <mark style="color:blue;">API Key</mark>                                                                                                                | <mark style="color:blue;">API Base (optional; leave blank to use the official endpoint)</mark> | Supported                 |
| Amazon Bedrock    | <mark style="color:blue;">AWS Access Key ID</mark>, <mark style="color:blue;">AWS Secret Access Key</mark>, <mark style="color:blue;">AWS Region</mark> | —                                                                                              | Manual addition required  |
| Google Vertex AI  | <mark style="color:blue;">Service Account JSON</mark> (paste the entire service account JSON)                                                           | —                                                                                              | Manual addition required  |
| vLLM              | <mark style="color:blue;">API Base</mark> (the URL of your self-hosted vLLM service, such as `https://vllm.example.com/v1`)                             | <mark style="color:blue;">API Key (optional)</mark>                                            | Supported                 |
| OpenAI Compatible | <mark style="color:blue;">API Base</mark> (the URL of a self-hosted or third-party OpenAI-compatible endpoint)                                          | <mark style="color:blue;">API Key (optional)</mark>                                            | Supported                 |
| {% endstep %}     |                                                                                                                                                         |                                                                                                |                           |

{% step %}

### Discover Models <a href="#discover-models" id="discover-models"></a>

The system queries the provider for available models and selects all of them by default. Models already in the catalog are labeled <mark style="color:blue;">Already Exists</mark> and cannot be selected. After reviewing the selection, click “Import N Models.” Each selected model becomes a custom model. You can click <mark style="color:blue;">Stop</mark> during import. If you do not need discovery, click <mark style="color:blue;">Skip and Add Manually</mark>.
{% endstep %}

{% step %}

### Finish <a href="#provider-done" id="provider-done"></a>

The screen displays <mark style="color:blue;">Provider Connected</mark> and the import results. The imported models use default context and capability settings, which you can adjust individually in the <mark style="color:blue;">Model Catalog</mark>.
{% endstep %}
{% endstepper %}

Each row in the provider list shows the status, masked credentials, and an “N Models” link (click it to open the model catalog filtered to that provider). On the right, you can <mark style="color:blue;">Test</mark>, <mark style="color:blue;">Edit</mark>, or <mark style="color:blue;">Delete</mark> the provider. You cannot change the provider type when editing. Leaving a secret field blank keeps its existing value.

## Manage Models <a href="#manage-models" id="manage-models"></a>

The <mark style="color:blue;">Models</mark> tab lists two model types in one catalog:

* **Managed models**: Models provided by MaiAgent, with the source labeled “MaiAgent Managed.”
* **Custom models**: Models created with your organization's own provider credentials, with the source labeled <mark style="color:blue;">Custom</mark>. The <mark style="color:blue;">Pricing</mark> column shows the traffic fee rate as “Total tokens × N%.”

Catalog columns include <mark style="color:blue;">Name</mark>, <mark style="color:blue;">Available to Organization</mark>, <mark style="color:blue;">Referenced by Fallbacks</mark>, <mark style="color:blue;">Source</mark>, <mark style="color:blue;">Provider</mark>, <mark style="color:blue;">Model ID</mark>, <mark style="color:blue;">Pricing</mark>, and <mark style="color:blue;">Context</mark>. At the top, filter by capability (<mark style="color:blue;">Chat</mark>, <mark style="color:blue;">Vision</mark>, <mark style="color:blue;">Tool Calling</mark>, <mark style="color:blue;">Reasoning</mark>), provider, or keyword. Click any row to open <mark style="color:blue;">Model Details</mark>, where you can copy the Gateway model ID, view its provider, and see <mark style="color:blue;">Which Fallback Chains Use It</mark>.

### Available to Organization Toggle <a href="#org-access-switch" id="org-access-switch"></a>

The <mark style="color:blue;">Available to Organization</mark> toggle for managed models uses the same setting as the [Model Access Permissions](/maiagent-user-guide/en/org/model-access.md) page. Turning it off here also removes the model from AI assistants' model menus. Before disabling it, a “Disable this model?” confirmation lists how many assistants and fallback chains will be affected. A lock icon means MaiAgent has not enabled the model for the organization, or the platform has disabled Gateway access to it; the organization cannot enable it itself. Custom models always display <mark style="color:blue;">Custom · Automatically Available</mark> and are not controlled by this toggle.

### Add a Custom Model <a href="#add-custom-model" id="add-custom-model"></a>

1. Click <mark style="color:blue;">Add Model</mark>, then select provider credentials in the <mark style="color:blue;">Choose Provider</mark> step (if you have no providers yet, click <mark style="color:blue;">Add Provider</mark> directly).
2. In <mark style="color:blue;">Model Settings</mark>, fill in:
   * <mark style="color:blue;">Name</mark>: The name displayed in the catalog and model menus.
   * <mark style="color:blue;">Family</mark>: Anthropic, GPT, Gemini, Grok, Qwen, DeepSeek, Mistral, or Other.
   * <mark style="color:blue;">Model ID</mark>: The model ID expected by the provider (for example, `gpt-4.1`). It must be unique under the same provider credentials.
   * <mark style="color:blue;">Context</mark> and <mark style="color:blue;">Maximum Output Tokens</mark>: Defaults are 128,000 and 8,000; adjust them to the model's specifications.
   * <mark style="color:blue;">Capabilities</mark>: <mark style="color:blue;">Vision</mark>, <mark style="color:blue;">Tool Calling</mark>, <mark style="color:blue;">Streaming</mark>, and <mark style="color:blue;">Reasoning</mark>.
3. Click <mark style="color:blue;">Create</mark>. The completion page displays the automatically generated Gateway model ID. Click <mark style="color:blue;">Test Connection</mark> to confirm that you can call it.

The Gateway model ID is generated from the provider's Slug and the model name, and does not change if the model is renamed. Use this ID in application calls. On the right of each custom model row, you can <mark style="color:blue;">Test</mark>, <mark style="color:blue;">Edit</mark>, or <mark style="color:blue;">Delete</mark> the model. After creating a model or changing its model ID or provider, the system detects in the background whether it rejects parameters such as temperature, top\_p, and top\_k. The results appear under <mark style="color:blue;">Request Parameter Support</mark> in the edit dialog, where you can also click <mark style="color:blue;">Detect Again</mark>. Not all providers support detection. If detection is unavailable, the system displays “Cannot detect this model (valid credentials and a supported provider are required).”

## Fallback Model Groups <a href="#fallback-groups" id="fallback-groups"></a>

Fallback model groups let the Gateway try the next model in a fallback chain when the primary model fails. Chains can span models and providers.

1. In <mark style="color:blue;">Fallback Model Groups</mark>, click <mark style="color:blue;">Add Fallback Model Group</mark>.
2. Enter a <mark style="color:blue;">Group Name</mark> (for example, “Claude High Availability”).
3. Select the primary model in the first row of the <mark style="color:blue;">Fallback Chain</mark>, then click <mark style="color:blue;">Add Next Priority</mark> to add fallback models. You can choose managed or custom models, and must select at least one. Drag rows or use <mark style="color:blue;">Move Up</mark> and <mark style="color:blue;">Move Down</mark> to reorder them.
4. Click <mark style="color:blue;">Save</mark>. The system automatically generates a group Slug. Use the Slug in the model field of your application calls to invoke the entire fallback chain.

<mark style="color:blue;">Trigger Conditions</mark> are fixed by the system and cannot be customized: <mark style="color:blue;">Timeout</mark>, <mark style="color:blue;">Connection Error</mark>, <mark style="color:blue;">429 Rate Limit</mark>, and <mark style="color:blue;">5xx Error</mark>. Each group card displays the trigger count for this month and the same period last month. A red warning appears when this month's count is at least twice that of the same period last month. Groups labeled <mark style="color:blue;">Platform Level · Read-Only</mark> are shared platform rules; organizations can only <mark style="color:blue;">View</mark> them.

## Integration Examples <a href="#integration-examples" id="integration-examples"></a>

The <mark style="color:blue;">Integration Examples</mark> tab explains the three steps for calling the Gateway through an OpenAI-compatible interface (obtain a Gateway key, specify a model name, and use automatic forwarding and fallback) and provides cURL, Python, and Node.js examples. After you select a model or group from <mark style="color:blue;">Choose Model or Fallback Group</mark>, the model field in the examples changes to the corresponding Gateway model ID or fallback group Slug. Click <mark style="color:blue;">Copy</mark> to use the code directly.

An administrator with <mark style="color:blue;">API Key Management Permission</mark> creates API Keys for calls under <mark style="color:blue;">Organization Settings</mark> → <mark style="color:blue;">API Key Management</mark>. Each key is bound to a member. You can further restrict <mark style="color:blue;">Enabled Models</mark> within the models available to the organization, and set <mark style="color:blue;">Requests per Minute (RPM)</mark> and <mark style="color:blue;">Tokens per Minute (TPM)</mark>. For authentication, error codes, and other details, see [AI Gateway: OpenAI-Compatible API Integration](https://docs.maiagent.ai/tech/ai-gateway/ai-gateway-openai-compatible) and [AI Gateway Security](https://docs.maiagent.ai/tech/ai-gateway/ai-gateway-security).

## Use with AI Assistants and MaiGPT <a href="#use-in-agents-and-maigpt" id="use-in-agents-and-maigpt"></a>

* **AI assistants**: After creation, custom models automatically appear in the <mark style="color:blue;">LLM Model</mark> menu in [AI Assistant Settings](/maiagent-user-guide/en/build/agent/setup.md), labeled <mark style="color:blue;">Custom</mark>. When a member with <mark style="color:blue;">AI Gateway Permission</mark> edits an AI assistant, the menu also includes a <mark style="color:blue;">Fallback Model Groups</mark> section for selecting the organization's fallback groups directly. Thinking settings are not displayed when a fallback group is selected.
* **MaiGPT**: MaiGPT only lists models with the <mark style="color:blue;">Tool Calling</mark> capability, so enable <mark style="color:blue;">Tool Calling</mark> for a custom model to make it appear. If the member's role restricts the models available in MaiGPT, you must also select the model in [MaiGPT Role Permissions](/maiagent-user-guide/en/org/roles/maigpt-role-permissions.md).

{% hint style="warning" %}
Before deleting custom models or fallback model groups, switch the AI assistants using them to other models, then delete them to avoid interrupting live services.
{% endhint %}

## Billing <a href="#billing" id="billing"></a>

When you use custom models (with your own credentials or self-hosted endpoints), MaiAgent charges a traffic fee based on total tokens instead of the platform's model rates. Managed models retain their existing rates. For how to view rates and usage records, see [AI Gateway Custom Model Traffic Fees](/maiagent-user-guide/en/org/credits.md#ai-gateway-traffic-fee).

## FAQ <a href="#faq" id="faq"></a>

#### Q: Why is AI Gateway missing from the left menu? <a href="#faq-menu-missing" id="faq-menu-missing"></a>

Check two things: whether AI Gateway is enabled for your organization (under “Advanced Features” in <mark style="color:blue;">Organization Overview</mark>), and whether your role has <mark style="color:blue;">AI Gateway Permission</mark> selected.

#### Q: Why does testing an OpenAI Compatible provider report that there is no default test model? <a href="#faq-compatible-test" id="faq-compatible-test"></a>

The <mark style="color:blue;">Test</mark> action in the provider row sends a test request using the provider's default test model. OpenAI-compatible endpoints have no default model available. Create a custom model in the model catalog first, then use <mark style="color:blue;">Test</mark> in that model's row to check the connection.

#### Q: Why can't I delete a provider? <a href="#faq-delete-provider" id="faq-delete-provider"></a>

The system rejects deletion if custom models still use the provider credentials. Use the “N Models” link in the provider row to find and delete those custom models first, then delete the provider.

#### Q: Why are fallback model groups missing from an AI assistant's model menu? <a href="#faq-fallback-not-in-agent" id="faq-fallback-not-in-agent"></a>

The member editing the AI assistant must have <mark style="color:blue;">AI Gateway Permission</mark> for the menu to list the organization's fallback model groups. Read-only platform-level groups do not appear in AI assistants' menus.


---

# Agent Instructions
This documentation is published with GitBook. GitBook is the documentation platform designed so that both humans and AI agents can read, navigate, and reason over technical content effectively. Learn more at gitbook.com.

## Querying This Documentation
If you need additional information that is not directly available in this page, you can query the documentation by asking a question.

Perform an HTTP GET request on the following URL with the `ask` and `goal` query parameters:

```
GET https://docs.maiagent.ai/maiagent-user-guide/en/build/ai-gateway.md?ask=<question>&goal=<user_goal>
```

`ask` is the immediate question: it should be specific, self-contained, and written in natural language.
`goal` is what the user is ultimately trying to achieve, the reason they need the answer. Sharing it helps GitBook give you a better, more relevant answer. A goal is most helpful when it describes the outcome the user wants rather than restating the question. For example, with `ask=how do I create an API token`, a goal like `automate deployments from our CI pipeline` lets GitBook tailor the answer to that use case.

The response will contain a direct answer to the question and relevant excerpts and sources from the documentation.

Use this mechanism when the answer is not explicitly present in the current page, you need clarification or additional context, or you want to retrieve related documentation sections.
