> ## Documentation Index
> Fetch the complete documentation index at: https://docs.omniall.ai/llms.txt
> Use this file to discover all available pages before exploring further.

# Gemini native format

> Native Gemini `generateContent` for image generation.

**Path:** `POST https://api.omniall.ai/v1beta/models/{model}:generateContent`

Prefer `gemini-nano-banana-2.1` for PDF / short-video reference generation.
Put media bytes in `inlineData.data` as raw Base64 (no `data:` prefix).
Output images appear in `candidates[].content.parts[].inlineData`.

Auth: `Authorization: Bearer sk-...` or `x-goog-api-key`.


## Banana 2.1 (`gemini-nano-banana-2.1`)

Use native `generateContent` for text-to-image, image-to-image, **PDF reference**, and **short video reference**.

| Input | `inlineData.mimeType` | `inlineData.data` |
| - | - | - |
| Image | `image/jpeg` / `image/png` / … | Raw Base64 (no `data:` prefix) |
| PDF | `application/pdf` | Raw Base64 |
| Short video | `video/mp4` | Raw Base64 |

Always set `generationConfig.responseModalities` to include `IMAGE`.

### PDF reference (complete)

```bash theme={null}
export BASE_URL="https://api.omniall.ai"
export API_KEY="sk-xxx"
PDF_B64=$(base64 < reference.pdf | tr -d '\n')

cat > request.json <<EOF
{
  "contents": [{
    "role": "user",
    "parts": [
      {"inlineData": {"mimeType": "application/pdf", "data": "$PDF_B64"}},
      {"text": "Read every page of the PDF. List labels, colors and shapes in page order, then generate a reference-sheet image preserving them."}
    ]
  }],
  "generationConfig": {
    "responseModalities": ["TEXT", "IMAGE"],
    "imageConfig": {"aspectRatio": "3:2", "imageSize": "1K"}
  }
}
EOF

curl --fail-with-body --max-time 180 \
  "$BASE_URL/v1beta/models/gemini-nano-banana-2.1:generateContent" \
  -H "Authorization: Bearer $API_KEY" \
  -H "Content-Type: application/json" \
  --data-binary @request.json -o response.json
```

### Short video reference (complete)

```bash theme={null}
VIDEO_B64=$(base64 < reference.mp4 | tr -d '\n')

cat > request.json <<EOF
{
  "contents": [{
    "role": "user",
    "parts": [
      {"inlineData": {"mimeType": "video/mp4", "data": "$VIDEO_B64"}},
      {"text": "Watch the entire video. List labels, colors and shapes chronologically, then generate a reference-sheet image preserving them."}
    ]
  }],
  "generationConfig": {
    "responseModalities": ["TEXT", "IMAGE"],
    "imageConfig": {"aspectRatio": "3:2", "imageSize": "1K"}
  }
}
EOF

curl --fail-with-body --max-time 180 \
  "$BASE_URL/v1beta/models/gemini-nano-banana-2.1:generateContent" \
  -H "Authorization: Bearer $API_KEY" \
  -H "Content-Type: application/json" \
  --data-binary @request.json -o response.json
```

Read output images from `candidates[].content.parts[].inlineData` (scan all parts; skip `thought: true`). Prefer short H.264 MP4 and small PDFs; Base64 expands roughly one third.

### Optional: Thinking

Banana 2.1 supports `thinkingLevel`: `minimal` / `medium` (default) / `high`. Omit for typical use; when needed, set `generationConfig.thinkingConfig`:

```json theme={null}
"generationConfig": {
  "responseModalities": ["TEXT", "IMAGE"],
  "imageConfig": { "aspectRatio": "3:2", "imageSize": "1K" },
  "thinkingConfig": {
    "thinkingLevel": "high",
    "includeThoughts": false
  }
}
```

`includeThoughts` only controls whether thought text is returned; it does not disable internal thinking. See [Overview · Thinking](/api-reference/image-gemini#optional-thinking-levels).


## OpenAPI

````yaml openapi/gemini-native-image.yaml POST /v1beta/models/{model}:generateContent
openapi: 3.1.0
info:
  title: Omniall AI — Gemini native format
  version: 1.1.0
  description: >
    Native Gemini `generateContent` for image generation on Omniall AI.


    **Path:** `POST
    https://api.omniall.ai/v1beta/models/{model}:generateContent`


    Supported models (examples): `gemini-nano-banana-2.1`,
    `gemini-3.1-flash-image-preview`,

    `gemini-3-pro-image-preview`. Always include `IMAGE` in
    `generationConfig.responseModalities`.


    Reference inputs via `parts[].inlineData`:

    - Images: `image/jpeg`, `image/png`, …

    - PDF: `application/pdf` (Banana 2.1+)

    - Short video: `video/mp4` (Banana 2.1+)


    Optional Thinking (Banana 2.1):
    `generationConfig.thinkingConfig.thinkingLevel`

    is `minimal` / `medium` (default) / `high`. `includeThoughts` only controls
    whether

    thought text is returned.


    Auth: `Authorization: Bearer sk-...` or `x-goog-api-key`.
servers:
  - url: https://api.omniall.ai
security: []
paths:
  /v1beta/models/{model}:generateContent:
    post:
      tags:
        - Gemini image
      summary: Gemini native format
      description: >
        Native Gemini `generateContent` for image generation.


        **Path:** `POST
        https://api.omniall.ai/v1beta/models/{model}:generateContent`


        Prefer `gemini-nano-banana-2.1` for PDF / short-video reference
        generation.

        Put media bytes in `inlineData.data` as raw Base64 (no `data:` prefix).

        Output images appear in `candidates[].content.parts[].inlineData`.


        Auth: `Authorization: Bearer sk-...` or `x-goog-api-key`.
      operationId: geminiNativeImageGenerateContent
      parameters:
        - name: model
          in: path
          required: true
          schema:
            type: string
            example: gemini-nano-banana-2.1
          description: >-
            Gemini image model id. Examples: `gemini-nano-banana-2.1`,
            `gemini-3.1-flash-image-preview`.
        - name: Authorization
          in: header
          required: true
          schema:
            type: string
            example: Bearer sk-...
          description: Bearer API key.
      requestBody:
        required: true
        content:
          application/json:
            schema:
              type: object
              required:
                - contents
              description: Native Gemini image-generation request body.
              properties:
                contents:
                  type: array
                  description: >-
                    Conversation turns. Mix text with one or more `inlineData`
                    media parts.
                  items:
                    type: object
                    properties:
                      role:
                        type: string
                        example: user
                      parts:
                        type: array
                        items:
                          type: object
                          properties:
                            text:
                              type: string
                              description: Text prompt / edit instruction.
                            inlineData:
                              type: object
                              description: >-
                                Inline media (image / PDF / short MP4) as
                                Base64.
                              properties:
                                mimeType:
                                  type: string
                                  description: >-
                                    MIME type: `image/jpeg`, `application/pdf`,
                                    `video/mp4`, etc.
                                  example: application/pdf
                                data:
                                  type: string
                                  description: Base64 payload without `data:` prefix.
                generationConfig:
                  type: object
                  properties:
                    responseModalities:
                      type: array
                      items:
                        type: string
                      description: Must include `IMAGE` for image generation.
                      example:
                        - TEXT
                        - IMAGE
                    imageConfig:
                      type: object
                      properties:
                        aspectRatio:
                          type: string
                          description: Aspect ratio, e.g. 1:1, 3:2, 16:9.
                          enum:
                            - '1:1'
                            - '2:3'
                            - '3:2'
                            - '3:4'
                            - '4:3'
                            - '4:5'
                            - '5:4'
                            - '9:16'
                            - '16:9'
                            - '21:9'
                          example: '3:2'
                        imageSize:
                          type: string
                          description: 'Resolution tier: 1K / 2K / 4K.'
                          enum:
                            - 1K
                            - 2K
                            - 4K
                          example: 1K
                    thinkingConfig:
                      type: object
                      description: >
                        Optional. Banana 2.1 Thinking: `thinkingLevel` is
                        `minimal` / `medium` (default) / `high`.

                        `includeThoughts` only controls whether thought text is
                        returned.
                      properties:
                        thinkingLevel:
                          type: string
                          description: Thinking level.
                          enum:
                            - minimal
                            - medium
                            - high
                          example: high
                        includeThoughts:
                          type: boolean
                          description: Whether to return thought content in the response.
                          example: false
            examples:
              text_to_image:
                summary: Text-to-image (Banana 2.1)
                value:
                  contents:
                    - role: user
                      parts:
                        - text: >-
                            Create a product photo of a yellow ceramic mug on a
                            wooden desk.
                  generationConfig:
                    responseModalities:
                      - TEXT
                      - IMAGE
                    imageConfig:
                      aspectRatio: '3:2'
                      imageSize: 1K
              image_to_image:
                summary: Image-to-image
                value:
                  contents:
                    - role: user
                      parts:
                        - inlineData:
                            mimeType: image/jpeg
                            data: BASE64_REFERENCE_IMAGE
                        - text: >-
                            Based on the reference, generate three orthographic
                            views on one canvas with a white background
                  generationConfig:
                    responseModalities:
                      - TEXT
                      - IMAGE
                    imageConfig:
                      aspectRatio: '16:9'
                      imageSize: 2K
              multi_reference:
                summary: Multi-image reference
                value:
                  contents:
                    - role: user
                      parts:
                        - inlineData:
                            mimeType: image/jpeg
                            data: BASE64_REF_1
                        - inlineData:
                            mimeType: image/png
                            data: BASE64_REF_2
                        - text: >-
                            Combine both references into one clean product
                            layout; preserve colors and shapes.
                  generationConfig:
                    responseModalities:
                      - TEXT
                      - IMAGE
                    imageConfig:
                      aspectRatio: '3:2'
                      imageSize: 1K
              pdf_reference:
                summary: PDF reference (Banana 2.1)
                value:
                  contents:
                    - role: user
                      parts:
                        - inlineData:
                            mimeType: application/pdf
                            data: PDF_BASE64_WITHOUT_DATA_PREFIX
                        - text: >-
                            Read every page of the PDF. List labels, colors and
                            shapes in page order, then generate a
                            reference-sheet image preserving them.
                  generationConfig:
                    responseModalities:
                      - TEXT
                      - IMAGE
                    imageConfig:
                      aspectRatio: '3:2'
                      imageSize: 1K
              video_reference:
                summary: Short video reference (Banana 2.1)
                value:
                  contents:
                    - role: user
                      parts:
                        - inlineData:
                            mimeType: video/mp4
                            data: VIDEO_MP4_BASE64_WITHOUT_DATA_PREFIX
                        - text: >-
                            Watch the entire video. List labels, colors and
                            shapes chronologically, then generate a
                            reference-sheet image preserving them.
                  generationConfig:
                    responseModalities:
                      - TEXT
                      - IMAGE
                    imageConfig:
                      aspectRatio: '3:2'
                      imageSize: 1K
      responses:
        '200':
          description: OK — read text/images from `candidates[].content.parts`.
        '400':
          description: Bad request
        '401':
          description: Unauthorized
        '413':
          description: Request body too large (Base64 expands ~33%)

````

This documentation is built and hosted on [Mintlify](https://mintlify.com), a developer documentation platform.