curl --request POST \
--url https://api.omniall.ai/v1/chat/completions \
--header 'Authorization: Bearer <token>' \
--header 'Content-Type: application/json' \
--data '
{
"model": "gpt-4o-mini",
"messages": [
{
"role": "user",
"content": "Hello"
}
]
}
'{
"id": "chatcmpl-...",
"object": "chat.completion",
"created": 1710000000,
"model": "gpt-4o-mini",
"choices": [
{
"index": 0,
"message": {
"role": "assistant",
"content": "Hello!"
},
"finish_reason": "stop"
}
],
"usage": {
"prompt_tokens": 10,
"completion_tokens": 4,
"total_tokens": 14
}
}{
"error": {
"message": "<string>",
"type": "new_api_error",
"code": "<string>"
}
}{
"error": {
"message": "无效的令牌 (request id: ...)",
"type": "new_api_error",
"code": ""
}
}{
"error": {
"message": "<string>",
"type": "new_api_error",
"code": "<string>"
}
}{
"error": {
"message": "<string>",
"type": "new_api_error",
"code": "<string>"
}
}{
"error": {
"message": "<string>",
"type": "new_api_error",
"code": "<string>"
}
}Chat Completions
OpenAI Chat Completions API relayed by Omniall AI.
Base path: POST https://api.omniall.ai/v1/chat/completions
Compatible with the official OpenAI Chat Completions schema. Use model IDs from the Model Square.
Common scenarios
- Vision: multimodal
messages[].contentwithimage_url - File analysis:
type: filecontent parts (when the upstream model supports it) - Web search:
web_search_options/ provider search fields when the model supports search - Agent:
tools+tool_choicefunction calling loop - Streaming:
stream: true(SSE)
Official reference: https://platform.openai.com/docs/api-reference/chat
curl --request POST \
--url https://api.omniall.ai/v1/chat/completions \
--header 'Authorization: Bearer <token>' \
--header 'Content-Type: application/json' \
--data '
{
"model": "gpt-4o-mini",
"messages": [
{
"role": "user",
"content": "Hello"
}
]
}
'{
"id": "chatcmpl-...",
"object": "chat.completion",
"created": 1710000000,
"model": "gpt-4o-mini",
"choices": [
{
"index": 0,
"message": {
"role": "assistant",
"content": "Hello!"
},
"finish_reason": "stop"
}
],
"usage": {
"prompt_tokens": 10,
"completion_tokens": 4,
"total_tokens": 14
}
}{
"error": {
"message": "<string>",
"type": "new_api_error",
"code": "<string>"
}
}{
"error": {
"message": "无效的令牌 (request id: ...)",
"type": "new_api_error",
"code": ""
}
}{
"error": {
"message": "<string>",
"type": "new_api_error",
"code": "<string>"
}
}{
"error": {
"message": "<string>",
"type": "new_api_error",
"code": "<string>"
}
}{
"error": {
"message": "<string>",
"type": "new_api_error",
"code": "<string>"
}
}Authorizations
Use Authorization: Bearer sk-... from https://omniall.ai/dashboard
Body
OpenAI Chat Completions request (POST /v1/chat/completions). Maps to Omniall gateway GeneralOpenAIRequest.
Model id from Model Square / GET /v1/models. Must be a model your key can access.
"gpt-4o"
Full conversation so far, oldest first. Include system/user/assistant/tool turns as needed.
Show child attributes
Show child attributes
If true, return Server-Sent Events (chat.completion.chunk) instead of a single JSON object.
Streaming extras.
Show child attributes
Show child attributes
Sampling temperature (about 0–2). Higher = more random; 0 = more deterministic.
0 <= x <= 2Nucleus sampling probability mass (0–1). Prefer tuning either temperature or top_p, not both aggressively.
0 <= x <= 1Top-K sampling when the upstream model supports it (ignored by some OpenAI models).
How many completions to generate for this prompt (cost scales with n).
x >= 1Legacy max tokens for the completion. Prefer max_completion_tokens on newer OpenAI models.
Maximum tokens the model may generate (preferred for newer OpenAI chat models).
Stop sequence(s). Generation ends when any sequence is produced.
Penalize new tokens based on whether they appear in the text so far (−2 to 2). Positive values encourage new topics.
-2 <= x <= 2Penalize tokens by frequency (−2 to 2). Positive values reduce repetition.
-2 <= x <= 2Best-effort deterministic seed when the provider supports it.
Stable end-user id for abuse monitoring (opaque string). Do not put secrets here.
Map of token-id → bias to ban/boost specific tokens (provider-specific).
Show child attributes
Show child attributes
Whether to return log probabilities for output tokens.
Number of most likely tokens to return at each position when logprobs is enabled.
Force output shape. json_object asks for valid JSON; json_schema asks for JSON matching your schema (when supported).
Show child attributes
Show child attributes
Function tools available to the model for agent workflows.
Show child attributes
Show child attributes
Controls tool use: none / auto / required, or force a specific function.
none, auto, required Whether the model may call multiple tools in one turn.
Reasoning intensity for o-series / compatible models (low / medium / high).
low, medium, high Requested output modalities, e.g. text, audio.
text, audio Audio output settings when using audio modalities.
OpenAI-style web search options when the model supports search.
Show child attributes
Show child attributes
Qwen-style web search toggle (channel/model dependent).
Provider-specific web search object (e.g. Baidu).
xAI search parameters object.
Doubao / Zhipu thinking controls.
Qwen thinking toggle.
Extra provider fields (sometimes used for Gemini-related options on the OpenAI-compatible path).
Arbitrary metadata passthrough when supported.
Predicted-output hint to reduce latency when you can anticipate part of the answer.
Response
Chat completion JSON, or SSE stream when stream=true.
Non-streaming Chat Completions response.
Unique completion id.
Always chat.completion for non-streaming.
"chat.completion"
Unix timestamp (seconds) when the completion was created.
Model that produced the answer (may reflect routing).
Generated choices (usually one).
Show child attributes
Show child attributes
Token usage when provided by upstream.
Show child attributes
Show child attributes
Upstream system fingerprint when available (reproducibility signal).