> ## Documentation Index
> Fetch the complete documentation index at: https://docs.omniall.ai/llms.txt
> Use this file to discover all available pages before exploring further.

# API overview

> Omniall AI API reference overview: auth, endpoints, request/response shapes, and model selection.

# API overview

Omniall AI is a unified AI API gateway. Request and response schemas follow the [OpenAI Chat Completions](https://platform.openai.com/docs/api-reference/chat) style (and related OpenAI media APIs), with additional first-class paths for Claude Messages, Gemini, and async video tasks.

At a high level, **Omniall normalizes routing, auth, billing, and response shapes across many upstream providers**, so you can call text, image, video, and audio models through one Base URL and one API key.

## OpenAPI specification

Interactive request schemas and the playground are generated from the public OpenAPI document. Browse nested **Chat** (ChatGPT / Claude / Gemini / Responses), plus **Images**, **Video**, **Audio**, **Platform**, and **Rerank**.

Use the playground to try calls against production (`https://api.omniall.ai`) with your own key.

## Base URL

```text theme={null}
https://api.omniall.ai/v1
```

Most OpenAI-compatible clients should set Base URL to `https://api.omniall.ai/v1` (include `/v1`). The website and console stay on `https://omniall.ai`. Paths below are shown from the API host root when that is clearer for non-`/v1` prefixes (for example Gemini `/v1beta`).

## Authentication

Create a key in the [console](https://omniall.ai/dashboard) under **API Keys**, then send:

```http theme={null}
Authorization: Bearer sk-your-key
Content-Type: application/json
```

Claude-style clients may also send `x-api-key` with `anthropic-version` on `/v1/messages` and model list routes. Gemini-style clients may use `x-goog-api-key` or `?key=` on Gemini-compatible model routes. Bearer auth works for all primary `/v1/*` APIs.

## Capabilities

| Capability | Method & path | Notes |
| - | - | - |
| List models | `GET /v1/models` | Models available to your key / group |
| Chat Completions | `POST /v1/chat/completions` | OpenAI Chat API; streaming supported |
| Completions | `POST /v1/completions` | Legacy text completions |
| Responses | `POST /v1/responses` | OpenAI Responses API |
| Claude Messages | `POST /v1/messages` | Anthropic Messages shape |
| Images | `POST /v1/images/generations` | Also `/v1/images/edits` |
| Embeddings | `POST /v1/embeddings` | |
| Audio speech | `POST /v1/audio/speech` | TTS |
| MiniMax music | `POST /v1/music_generation` | Official MiniMax music |
| Suno | `POST /suno/submit/{action}`, `GET /suno/fetch/{id}` | Async music / lyrics |
| Audio transcription / translation | `POST /v1/audio/transcriptions`, `/v1/audio/translations` | |
| Video (async) | `POST /v1/videos`, `POST /v1/video/generations` | Create task, then poll |
| Video status | `GET /v1/videos/{task_id}`, `GET /v1/video/generations/{task_id}` | |
| File upload | `POST /v1/files` | multipart `file` → CDN URL |
| Gemini | `POST /v1beta/models/{model}:{action}` | Native Gemini path |
| Realtime | `GET /v1/realtime` | WebSocket |

Model how-to guides live under the **Docs** tab (**Guides**).

Pick model IDs from the [Model Square](https://omniall.ai/pricing) — use the exact `model` string shown there in requests.

## Requests

### Chat Completions

`POST /v1/chat/completions` is the primary text (and multimodal chat) entry point. Body shape is OpenAI-compatible:

```typescript theme={null}
type ChatCompletionRequest = {
  model: string; // from Model Square, e.g. "gpt-4o-mini"
  messages: Message[];
  stream?: boolean;
  temperature?: number;
  top_p?: number;
  max_tokens?: number;
  stop?: string | string[];
  tools?: Tool[];
  tool_choice?: "none" | "auto" | { type: "function"; function: { name: string } };
  response_format?: { type: "json_object" } | { type: "json_schema"; json_schema: object };
  // Plus other OpenAI-style fields supported by the selected model
};

type Message = {
  role: "system" | "user" | "assistant" | "tool";
  content: string | ContentPart[];
  name?: string;
  tool_call_id?: string;
};
```

<CodeGroup>
  ```bash cURL theme={null}
  curl https://api.omniall.ai/v1/chat/completions \
    -H "Authorization: Bearer $OMNIALL_API_KEY" \
    -H "Content-Type: application/json" \
    -d '{
      "model": "gpt-4o-mini",
      "messages": [
        { "role": "user", "content": "What is the meaning of life?" }
      ]
    }'
  ```

  ```python Python theme={null}
  from openai import OpenAI

  client = OpenAI(
      api_key="sk-...",
      base_url="https://api.omniall.ai/v1",
  )

  resp = client.chat.completions.create(
      model="gpt-4o-mini",
      messages=[{"role": "user", "content": "What is the meaning of life?"}],
  )
  print(resp.choices[0].message.content)
  ```

  ```typescript TypeScript theme={null}
  import OpenAI from "openai";

  const client = new OpenAI({
    apiKey: "sk-...",
    baseURL: "https://api.omniall.ai/v1",
  });

  const resp = await client.chat.completions.create({
    model: "gpt-4o-mini",
    messages: [{ role: "user", content: "What is the meaning of life?" }],
  });
  console.log(resp.choices[0].message.content);
  ```
</CodeGroup>

### Streaming

Send `stream: true`. The response is [Server-Sent Events (SSE)](https://developer.mozilla.org/en-US/docs/Web/API/Server-sent_events/Using_server-sent_events). Chunks use `object: "chat.completion.chunk"` and `choices[].delta` instead of `message`. Ignore SSE comment lines if present. The stream ends with a `data: [DONE]` sentinel.

```bash theme={null}
curl https://api.omniall.ai/v1/chat/completions \
  -H "Authorization: Bearer $OMNIALL_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "gpt-4o-mini",
    "stream": true,
    "messages": [{"role": "user", "content": "Hello"}]
  }'
```

### Model selection

* Always pass `model` using the ID from [Model Square](https://omniall.ai/pricing) or `GET /v1/models`.
* Availability depends on your account group, channel routing, and balance.
* Unsupported parameters for a given upstream model are typically ignored; supported fields are forwarded.

### Images

```bash theme={null}
curl https://api.omniall.ai/v1/images/generations \
  -H "Authorization: Bearer $OMNIALL_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "dall-e-3",
    "prompt": "A calm studio workspace",
    "n": 1,
    "size": "1024x1024"
  }'
```

### Video (async tasks)

Video APIs are **asynchronous**: create a task, then poll until the status is complete. Preferred paths:

| Action | Path |
| - | - |
| Create | `POST /v1/videos` or `POST /v1/video/generations` |
| Poll | `GET /v1/videos/{task_id}` or `GET /v1/video/generations/{task_id}` |

Exact body fields (`prompt`, `seconds`, `image` / `images`, `metadata`, etc.) depend on the model family. See the Docs guides for Kling, Doubao Seedance, Veo, and others.

```bash theme={null}
curl -X POST https://api.omniall.ai/v1/videos \
  -H "Authorization: Bearer $OMNIALL_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "kling-2.5",
    "prompt": "A drone shot over a coastal road at sunset",
    "seconds": "5"
  }'
```

### Claude Messages & Gemini

* **Claude**: `POST /v1/messages` with Anthropic Messages JSON (`model`, `max_tokens`, `messages`, …).
* **Gemini**: `POST /v1beta/models/{model_name}:{action}` (for example `generateContent`), or OpenAI-compatible chat against `/v1/chat/completions` when the model is exposed that way.

## Responses

Non-streaming Chat Completions responses follow the OpenAI shape: `choices` is always an array. Each choice has a `message` (or `delta` when streaming) and a `finish_reason`.

```json theme={null}
{
  "id": "chatcmpl-...",
  "object": "chat.completion",
  "created": 1710000000,
  "model": "gpt-4o-mini",
  "choices": [
    {
      "index": 0,
      "message": {
        "role": "assistant",
        "content": "Hello!"
      },
      "finish_reason": "stop"
    }
  ],
  "usage": {
    "prompt_tokens": 10,
    "completion_tokens": 4,
    "total_tokens": 14
  }
}
```

Common `finish_reason` values include `stop`, `length`, `tool_calls`, and `content_filter`. Token usage is returned in `usage` when the upstream provides it; billing follows Omniall quotas and Model Square pricing.

Video create responses return a **task id**; poll the status endpoint until the task finishes, then read the result URL from the task payload (or content download route when available).

## Errors & limits

Failed requests return JSON error bodies (OpenAI-style `error.message` / `error.type` where applicable). Typical causes:

* Missing or invalid API key
* Unknown or unauthorized `model`
* Insufficient balance / quota
* Rate limits on your key or model
* Upstream provider errors (retried or surfaced depending on routing)

## Next steps

* Try requests in **Endpoints** (interactive playground)
* [Quickstart](/quickstart) for SDK setup
* [Model Square](https://omniall.ai/pricing) for model IDs and pricing
* Docs tab for image / video / audio model guides


This documentation is built and hosted on [Mintlify](https://mintlify.com), a developer documentation platform.