Skip to main content
POST

Authorizations

Authorization
string
header
required

Use Authorization: Bearer sk-... from https://omniall.ai/dashboard

Body

application/json

OpenAI Chat Completions request (POST /v1/chat/completions). Maps to Omniall gateway GeneralOpenAIRequest.

model
string
required

Model id from Model Square / GET /v1/models. Must be a model your key can access.

Example:

"gpt-4o"

messages
object[]
required

Full conversation so far, oldest first. Include system/user/assistant/tool turns as needed.

stream
boolean
default:false

If true, return Server-Sent Events (chat.completion.chunk) instead of a single JSON object.

stream_options
object

Streaming extras.

temperature
number

Sampling temperature (about 0–2). Higher = more random; 0 = more deterministic.

Required range: 0 <= x <= 2
top_p
number

Nucleus sampling probability mass (0–1). Prefer tuning either temperature or top_p, not both aggressively.

Required range: 0 <= x <= 1
top_k
integer

Top-K sampling when the upstream model supports it (ignored by some OpenAI models).

n
integer

How many completions to generate for this prompt (cost scales with n).

Required range: x >= 1
max_tokens
integer

Legacy max tokens for the completion. Prefer max_completion_tokens on newer OpenAI models.

max_completion_tokens
integer

Maximum tokens the model may generate (preferred for newer OpenAI chat models).

stop

Stop sequence(s). Generation ends when any sequence is produced.

presence_penalty
number

Penalize new tokens based on whether they appear in the text so far (−2 to 2). Positive values encourage new topics.

Required range: -2 <= x <= 2
frequency_penalty
number

Penalize tokens by frequency (−2 to 2). Positive values reduce repetition.

Required range: -2 <= x <= 2
seed
number

Best-effort deterministic seed when the provider supports it.

user
string

Stable end-user id for abuse monitoring (opaque string). Do not put secrets here.

logit_bias
object

Map of token-id → bias to ban/boost specific tokens (provider-specific).

logprobs
boolean

Whether to return log probabilities for output tokens.

top_logprobs
integer

Number of most likely tokens to return at each position when logprobs is enabled.

response_format
object

Force output shape. json_object asks for valid JSON; json_schema asks for JSON matching your schema (when supported).

tools
object[]

Function tools available to the model for agent workflows.

tool_choice

Controls tool use: none / auto / required, or force a specific function.

Available options:
none,
auto,
required
parallel_tool_calls
boolean

Whether the model may call multiple tools in one turn.

reasoning_effort
enum<string>

Reasoning intensity for o-series / compatible models (low / medium / high).

Available options:
low,
medium,
high
modalities
enum<string>[]

Requested output modalities, e.g. text, audio.

Available options:
text,
audio
audio
object

Audio output settings when using audio modalities.

web_search_options
object

OpenAI-style web search options when the model supports search.

enable_search

Qwen-style web search toggle (channel/model dependent).

Provider-specific web search object (e.g. Baidu).

search_parameters
object

xAI search parameters object.

thinking
object

Doubao / Zhipu thinking controls.

enable_thinking

Qwen thinking toggle.

extra_body
object

Extra provider fields (sometimes used for Gemini-related options on the OpenAI-compatible path).

metadata
object

Arbitrary metadata passthrough when supported.

prediction
object

Predicted-output hint to reduce latency when you can anticipate part of the answer.

Response

Chat completion JSON, or SSE stream when stream=true.

Non-streaming Chat Completions response.

id
string

Unique completion id.

object
string

Always chat.completion for non-streaming.

Example:

"chat.completion"

created
integer

Unix timestamp (seconds) when the completion was created.

model
string

Model that produced the answer (may reflect routing).

choices
object[]

Generated choices (usually one).

usage
object

Token usage when provided by upstream.

system_fingerprint
string

Upstream system fingerprint when available (reproducibility signal).