PlatformLLM API

LLM & Image API

Chat, reasoning, image & video generation, and Topaz upscaling through a single OpenAI-compatible API. Claude, GPT-5, Gemini, Grok, DeepSeek, Nano Banana, and more — one key, one balance.

Overview

GPUniq provides a single API surface for 90+ language models and image generators across Anthropic, OpenAI, Google, xAI, and DeepSeek. One API key, one balance, one usage dashboard — chat completions, reasoning, and text-to-image all in the same place.

You can use GPUniq LLMs two ways:

  1. Native GPUniq API (/v1/llm/*) — wrapped responses, persistent chat sessions, terminal-command generator, SDK helpers.
  2. OpenAI-compatible API (/v1/openai/*) — drop-in replacement for api.openai.com/v1. Works with Claude Code, Cursor, Continue.dev, Aider, LiteLLM, and the official OpenAI Python/JS SDKs without code changes.

Available Models

Chat & Reasoning

ProviderModelsBest for
AnthropicClaude Opus 4.7 / 4.6 / 4.5, Sonnet 4.6 / 4.5, Haiku 4.5General reasoning, coding, agents
OpenAIGPT-5.5, GPT-5.2 Pro / Codex, GPT-5, o3, o3-mini, GPT-4o, GPT-4.1Reasoning, structured output, vision
GoogleGemini 3 Pro / Flash, Gemini 2.5Long context, fast batch work
xAIGrok 4, Grok 4.1 Thinking, Grok 4 FastReal-time knowledge, low latency
DeepSeekV4 Pro / V4 Flash, V3.2 / V3.2 Thinking, V3.1 / Terminus, R1 / R1 (May 2025), Reasoner, Chat, OCRCost-efficient reasoning, OCR, conversational
MiniMaxM2.7 / M2.5 / M2.1 / M2Long-context Chinese & multilingual, balanced cost

DeepSeek pricing (USD per 1M tokens, already discounted −20%)

SlugInputOutputCategoryNotes
deepseek-v3.2$2.16$3.24flagshipLatest flagship general model
deepseek-v3.2-thinking$0.30$0.45reasoningReasoning-tuned V3.2 (very cheap)
deepseek-v3.1$4.32$12.96flagshipPrevious flagship
deepseek-v3.1-terminus$0.15$0.30balancedUpdated V3.1, very cheap
deepseek-v3$2.16$8.64balancedOriginal V3 (Dec 2024)
deepseek-r1$4.32$17.28reasoningFirst reasoning model
deepseek-r1-0528$0.59$1.81reasoningUpdated R1, May 2025
deepseek-reasoner$0.30$0.45reasoningReasoning-focused alias
deepseek-chat$0.29$1.17balancedConversational alias
deepseek-ocr$0.23$0.23fastOCR model

DeepSeek V4 — priced live, not from a table

deepseek-v4-pro and deepseek-v4-flash are open-weight, so they sit in the open-weight tier: the price is not a fixed catalog number but the cheapest live rate available for those weights, plus our margin. It moves as the market moves, which is why it is deliberately absent from the table above — any number printed here would be stale within days.

Read the live figure from GET /v1/llm/models, or the model card in /chat. As of 2026-08-16 that resolves to roughly $0.52 / $1.04 for deepseek-v4-pro and $0.09 / $0.18 for deepseek-v4-flash per 1M tokens — in both cases at or below what DeepSeek's own API charges for the same model.

Corrected 2026-08-16. deepseek-v4-pro had been billing off a representative rate rather than the cheapest one, which put it around 4× DeepSeek's official API. It now tracks the floor, as the open-weight tier always intended. deepseek-v4-flash was unaffected and was already below DeepSeek's own price.

MiniMax pricing (USD per 1M tokens)

SlugInputOutputCategoryPublic discount
MiniMax-M2.7$0.30$1.20balanced
MiniMax-M2.5$0.30$1.20balanced
MiniMax-M2.1$0.2925$1.17balanced−7.5% off API
MiniMax-M2$0.2925$1.17flagship−7.5% off API

Repriced 2026-08-15. MiniMax-M2 was listed at $2.079 / $8.316 — roughly 8× the market rate for these weights. The catalog's "official price" anchor for the M2 family had been filled in with a Claude-shaped figure that MiniMax has never charged; the discount machinery then derived retail from it faithfully. The anchor is now MiniMax's real list price and the whole family sits within a few percent of what the same models cost elsewhere. Slug minimax-m2.5 (lowercase) is an alias of MiniMax-M2.5 and prices identically.

Image Generation

Image models are billed per returned image, not per token.

ModelSlugPrice / imageNotes
Nano Banananano-banana$0.0312Fast text-to-image & image-to-image, 1K
Nano Banana 2nano-banana-2$0.0500Quality-value generation, 1K–4K (set via size)
Nano Banana Pronano-banana-pro$0.1072Higher quality, 1K–2K pixel budget (set via size; beyond-2K asks auto-upgrade to Pro 4K)
Nano Banana Pro 4Knano-banana-pro-4k$0.1924K resolution — also reached by size beyond the 2K pixel budget (~4.2 MP or >3168px wide)
Grok 4 Imagegrok-4-image$0.0352xAI image generator
GPT Image 2gpt-image-2$0.0464OpenAI image, 1K tier (higher size auto-upgrades, see below)
GPT Image 2 · 2Kgpt-image-2-2k$0.072K tier — also reached by size up to 2048px
GPT Image 2 · 4Kgpt-image-2-4k$0.114K tier — also reached by size above 2048px
GPT Image 1.5gpt-image-1-5$0.020OpenAI image (cheaper tier)
GPT-4o Imagegpt-4o-image$0.040OpenAI 4o image
FLUX.2 Proflux-2-pro$0.060Black Forest Labs FLUX.2 Pro 1K
FLUX.2 Flexflux-2-flex$0.180Premium quality 1K
Flux Kontext Proflux-kontext-pro$0.080Text-to-image & edit
Flux Kontext Maxflux-kontext-max$0.160Premium edit / generation
Seedream 4seedream-4$0.050ByteDance Seedream 4
Seedream 4.5seedream-4-5$0.040ByteDance Seedream 4.5
Seedream 5.0 Liteseedream-5-0-lite$0.035ByteDance Seedream 5.0 Lite
Z-Imagez-image$0.020Alibaba Z-Image

How size picks the GPT Image 2 tier (and the price). The size field accepts WIDTHxHEIGHT or a bare 1k / 2k / 4k. The tier is chosen by a ceiling rule on the long side — the delivered image is never smaller than the ask: up to 1024px → base gpt-image-2 ($0.0464), up to 2048px → the 2K tier ($0.07), above 2048px → the 4K tier ($0.11). The response's usage.model (sync) / model (job) names the tier you were billed for, and the kickoff's estimated_cost_usd reflects it up front. Aspect ratio follows your size proportions — wide and tall formats including 21:9 and 3:2 are honoured, with or without reference images.

Nano Banana Pro upgrades by pixel budget, not by long side — banana tiers scale output with the aspect ratio (the 2K tier at 16:9 renders 2752×1536), so a QHD 2560x1440 ask stays on the base Pro ($0.1072) and comes back at 2752×1536. Only asks beyond the 2K budget (more than ~4.2 MP, or wider than 3168px) auto-upgrade to nano-banana-pro-4k ($0.192).

GPT Image 2 2K/4K renders now return the exact width×height you asked for (up to a 3840px long side; larger asks scale down proportionally, e.g. 4096×4096 → 3840×3840) and typically complete in 40–80 seconds. If the primary route rejects a request, GPUniq retries on an alternate 2K/4K route automatically; if the content itself is rejected everywhere, the request degrades once to the base gpt-image-2 at the base price ($0.0464, ~1.6K render) — the response's model field always names the SKU you were billed for.

Nano Banana 4K output dimensions depend on the aspect ratio — the tier fixes the pixel budget, not the long side, so wide formats come back wider than 4096px. Expect these dimensions from the 4K tier (nano-banana-pro-4k, nano-banana-2 at 4K):

Aspect4K output
1:14096 × 4096
16:9 / 9:165504 × 3072 / 3072 × 5504
3:2 / 2:35056 × 3392 / 3392 × 5056
4:3 / 3:44800 × 3584 / 3584 × 4800
21:96336 × 2688

4K renders on this family take noticeably longer than 1K/2K — 2–5 minutes is normal, which is another reason to keep the job polling budget at 10 minutes.

Use the job-based API. The synchronous endpoint is legacy.

POST /v1/llm/images/jobs (below) is the supported path: it returns a job_id in under a second, survives any CDN idle-read limit, and bills only on delivery. The synchronous POST /v1/llm/images/generations (and its OpenAI-compat twin POST /v1/openai/images/generations) holds the connection open for the full 5-minute upstream budget and is kept only so existing OpenAI-SDK code keeps running. The job API now covers every image slug in the catalog — there is no model that requires the synchronous route. Do not build new integrations on it.

A held-open request is also the one that loses work: if the client disconnects mid-render the image is still generated and still billed, but you never receive it. Jobs have no such failure mode.

Getting 403 with error code: 1010 on any endpoint? That is not the endpoint — it is your HTTP client's User-Agent.

Our edge blocks the default Python-urllib/3.x User-Agent outright. The block is UA-based and applies to every route equally — sync, async, chat — so it is easy to misread as "the synchronous endpoint is broken" when the async one fails identically. requests, httpx, curl, the official SDKs and any client that sets its own UA are all unaffected.

# 403, error code 1010 — blocked at the edge, never reaches the API
urllib.request.urlopen(urllib.request.Request(url, data=body))

# Works — requests sends its own User-Agent
requests.post(url, json=payload, headers={"X-API-Key": KEY})

# Works — urllib with an explicit User-Agent
req = urllib.request.Request(url, data=body, headers={
    "X-API-Key": KEY,
    "Content-Type": "application/json",
    "User-Agent": "my-app/1.0",     # <- the one line that matters
})

A 1010 body is always plain text, never our JSON envelope. If you get JSON back, the request reached us and the error is a real API error — look it up in the Error Reference.

POST /v1/llm/images/jobs returns a job_id in under a second, and you poll GET /v1/llm/images/jobs/{job_id} every 2-3 seconds until the status is terminal. You are charged only when the completion poll returns — a timed-out or failed job costs nothing. Server-side, polls that arrive within 2 seconds of each other are coalesced via Redis, so hammering the endpoint will not be billed as repeated upstream calls.

A status: "failed" payload carries a typed error_codecontent_moderation (the input was rejected by the upstream content policy; rephrase, don't retry verbatim) or generation_failed (transient; re-kick off). Give the poll loop a 10-minute budget: moderation verdicts and internal retries can land several minutes after kickoff, and abandoning early costs you the render you would have received. Handling recipes: Errors → Image jobs.

import time, requests

BASE = "https://api.gpuniq.com/v1/llm"
HEADERS = {"X-API-Key": "gpuniq_your_key"}

# 1. Kickoff
start = requests.post(
    f"{BASE}/images/jobs",
    headers=HEADERS,
    json={"model": "nano-banana-pro", "prompt": "a cozy cabin at sunrise", "n": 1},
).json()
job_id = start["data"]["job_id"]

# 2. Poll — 10-minute budget: slow 4K renders, moderation verdicts and
#    internal retries can all land several minutes after kickoff
deadline = time.time() + 600
while time.time() < deadline:
    time.sleep(2.5)
    r = requests.get(f"{BASE}/images/jobs/{job_id}", headers=HEADERS).json()
    d = r["data"]
    if d["status"] == "completed":
        image_b64 = d["image"]["b64_json"]
        print(f"Cost: ${d['cost_usd']}, balance: ${d['balance_usd']}")
        break
    if d["status"] == "failed":
        if d.get("error_code") == "content_moderation":
            # Content-policy verdict — retrying the same prompt will fail
            # again. Rephrase (drop video-style wording: durations,
            # camera moves) and kick off a new job.
            print("rejected by content policy:", d.get("error"))
        else:
            # generation_failed — transient; safe to re-kick off.
            print("failed:", d.get("error"))
        break

Which endpoint serves which slug

Every image slug in the catalog works on the job API. There is no subset to memorise and no reason to keep a synchronous code path around: if the model is in the table above, POST /v1/llm/images/jobs accepts it.

n must be 1 — issue separate jobs in parallel for batches. (Topaz upscaling/enhancement has its own dedicated job surface — see the Topaz guide.)

Behind the job id there are two execution modes, and the difference is not visible in the API contract — same kickoff, same polling, same response shape:

ModeSlugsTypical kickoff → first completed poll
Native async upstreamNano Banana line, gpt-image-2 familyunchanged
Server-side render behind the job ideverything else in the catalogunchanged for the client; the render occupies a worker rather than an upstream queue

The second mode runs the same multi-tier provider cascade the synchronous endpoint uses, so failover, moderation verdicts and format conversion all behave identically. You are still charged only on delivery.

Video has no synchronous endpoint at all. Every video generation goes through POST /v1/llm/videos/jobs — there is nothing to migrate and no sync variant to find.

Migrating off the synchronous endpoint. The job API is a drop-in replacement for POST /v1/llm/images/generations for every slug: send the same body to /images/jobs, read data.job_id, then poll GET /v1/llm/images/jobs/{job_id} until status is terminal. The image arrives as data.image.b64_json, and data.cost_usd / data.balance_usd replace the sync response's equivalents. The only behavioural difference is the one you want: a dropped connection can no longer cost you an image you never receive.

Generating an image inside a chat session

When you want the image to appear as a turn in an existing chat (so the prompt and result both land in the chat history), POST to /v1/llm/chats/{chat_id}/messages with an image model. The response returns immediately with type: "image_pending" plus a job_id and the dialogue_id of a placeholder row that already lives in the chat history. Poll GET /v1/llm/chats/{chat_id}/image-jobs/{job_id} until the status is completed (placeholder is rewritten with the image and the balance is debited) or failed (placeholder is marked, nothing charged). The polling endpoint 404s once the job is terminal — the final dialogue is the source of truth from then on.

import time, requests

BASE = "https://api.gpuniq.com/v1/llm"
HEADERS = {"X-API-Key": "gpuniq_your_key"}
chat_id = 42  # existing chat created via POST /v1/llm/chats

# 1. Kickoff (POST /chats/{id}/messages with an image model)
start = requests.post(
    f"{BASE}/chats/{chat_id}/messages",
    headers=HEADERS,
    json={"model": "nano-banana-pro", "message": "a cozy cabin at sunrise"},
).json()
job_id = start["data"]["job_id"]
dialogue_id = start["data"]["dialogue_id"]

# 2. Poll — same 5-minute budget as the standalone /images/jobs flow
deadline = time.time() + 300
while time.time() < deadline:
    time.sleep(2.5)
    r = requests.get(
        f"{BASE}/chats/{chat_id}/image-jobs/{job_id}", headers=HEADERS,
    ).json()
    d = r["data"]
    if d["status"] == "completed":
        image_b64 = d["image"]["b64_json"]
        print(f"Cost: ${d['cost_usd']}, balance: ${d['balance_usd']}")
        break
    if d["status"] == "failed":
        print("failed:", d.get("error"))
        break

Use this surface when the image should be part of a multi-turn conversation. Use the standalone /images/jobs surface when you don't need persistence — it has the same job semantics without creating a chat row.

Video Generation

Video models are billed per delivered video, not per token. Every generation is asynchronous — POST to /v1/llm/videos/jobs to kick off a job, then poll GET /v1/llm/videos/jobs/{job_id} until the status is terminal. You are charged only when the completion poll returns a video.url — a failed or timed-out job costs nothing.

FamilySlugHeadline / videoNotes
OpenAI Sora 2sora-2-video$0.060Sora 2, default 10s
OpenAI Sora 2 Prosora-2-pro-video$1.000Premium quality, 10s
Sora 2 Officialsora-2-official$0.4808s, official API
Sora 2 Pro Officialsora-2-pro-official$0.5608s 1080p, official API
Google Veo 3.1 Liteveo-3-1-litefrom $0.30720p / 1080p / 4K, 4–8s; i2v, first/last frame, reference-to-video
Google Veo 3.1 Fastveo-3-1-fastfrom $0.60720p / 1080p / 4K, 4–8s; i2v, first/last frame, reference-to-video
Google Veo 3.1 Qualityveo-3-1-qualityfrom $2.50Flagship; 720p / 1080p / 4K, 4–8s; i2v, first/last frame
Kling 2.1 Prokling-2-1$0.405Standard / Pro / Master tiers, 5s or 10s, i2v
Kling 2.5 Turbo Prokling-2-5-turbo-pro$0.3155s or 10s, t2v / i2v
Kling 2.6kling-2-6$0.315Optional audio, 5s or 10s, t2v / i2v
Kling 3.0kling-3-0$0.504720p / 1080p / 4K, audio, multi-shot to 15s
Kling O3 (Video)kling-o3-videofrom $0.076 / s720p / 1080p / 4K, native audio, 3–15s; t2v / i2v, first+last frame, reference images (≤4)
Kling 2.6 Motion Controlkling-2-6-motion-control$0.504720p / 1080p video-to-video
Kling 3.0 Motion Controlkling-3-0-motion-control$0.756720p / 1080p video-to-video
Kling AI Avatar Prokling-avatar-pro$1.0351080p lip-sync, up to 15s
Kling AI Avatar Standardkling-avatar-standard$0.506720p lip-sync, up to 15s
Hailuo 02hailuo-02$0.200768p, 6s default
Hailuo 2.3hailuo-2-3$0.350768p 6s
Seedance 1.0 Proseedance-1-0-pro$0.210ByteDance 720p, 5s, t2v + i2v (per video)
Seedance 1.5 Proseedance-1-5-pro$0.160ByteDance 720p, 5s, t2v + i2v (per video)
Seedance 2seedance-2$0.165–$1.65 / sByteDance 480p–4k, billed per second, rate depends on resolution (720p: $0.33/s, 5s ≈ $1.65)
Alibaba Wan 2.2 Fastwan-2-2-fast$0.120720p fast tier
Alibaba Wan 2.5wan-2-5$0.600720p 5s
Alibaba Wan 2.6wan-2-6$0.800720p 5s flagship
Wan Animatewan-animate$0.150720p animation
Happy Horsehappy-horse$0.160720p
Grok Imagine Videogrok-imagine-video$0.300xAI video, 6s
Runway Gen-4.5runway-gen-4-5$0.750Runway flagship 5s

Kling SKUs are billed at −10% off the official public price. The headline above is the cheapest default configuration (1080p / no audio / 5s / Pro tier). Audio, longer duration, 4K, and Master tier scale the price linearly off the underlying reference rate × 0.9 — the exact cost is returned in the cost_usd field of the completion response. A 10% margin floor against the upstream supplier guarantees we never bill below source cost, so on a provider fallback the price may rise by 1-3%.

Veo 3.1 (Google)

All three Veo tiers are billed per video, by resolution. Duration can be 4, 6 or 8 seconds (default 8) and does not change the price.

ResolutionLiteFastQuality
720p$0.30$0.60$2.50
1080p (default)$0.35$0.65$2.55
4k$1.50$1.80$3.70

Omitting resolution gives you 1080p, not 720p — the video job surface defaults to 1080p across the whole Kling/Veo family, so the cheapest row is opt-in. Pass "resolution": "720p" explicitly if that is what you are budgeting for. The kickoff response echoes the resolved value in config.resolution and prices it in estimated_cost_usd, so you can always check before the render runs.

  • First + last frameimage_url (first frame) + last_frame_url (last frame) on all three tiers; the clip interpolates between the two frames and the order is honoured.
  • Reference-to-video — up to 3 reference images via reference_image_urls on veo-3-1-fast / veo-3-1-lite (8-second clips only; not available on Quality). Mutually exclusive with last_frame_url.
  • aspect_ratio: 16:9 (default), 9:16, or auto (follows the input image geometry).
  • On a provider fallback the charge follows the serving upstream's rate — the exact amount is always in the completion cost_usd.

Kling O3

Kling O3 is billed per second, by resolution and audio flag, for clips of 3–15 seconds (default 5). Prompts are capped at 2500 characters.

ResolutionAudio offAudio on
720p$0.0756 / s$0.1008 / s
1080p (default)$0.1008 / s$0.1260 / s
4k$0.3780 / s$0.3780 / s
  • Text-to-video — prompt only; aspect_ratio is 16:9 (default), 9:16 or 1:1.
  • Image-to-videoimage_url (start frame), optionally with last_frame_url (end frame). The aspect ratio follows the frame.
  • Reference-to-video — up to 4 reference images via reference_image_urls, which define the subject/style. image_url / last_frame_url may ride along as start / end anchors.
  • No reference video. video_url / video_urls are rejected with a 400 on this model: the upstream accepts a clip at submit and then fails the render. For video-to-video use kling-2-6-motion-control / kling-3-0-motion-control.

Seedance (ByteDance)

Seedance is ByteDance's text-to-video / image-to-video family — fast, photoreal 720p clips, well suited to product shots, social content, and image-to-video animation of a still frame. Three SKUs are in the catalog, and they do not all price the same way, so read this before budgeting a batch.

SlugModelResolutionDefault durationModesBilling
seedance-1-0-proSeedance 1.0 Pro720p5st2v, i2vFlat $0.210 / video
seedance-1-5-proSeedance 1.5 Pro720p5st2v, i2vFlat $0.160 / video
seedance-2Seedance 2480p / 720p / 1080p / 4k5s (4–15s)t2v, i2v, first/last-frame, ref-to-video (image + video + audio refs)Per second, by resolution — see below

seedance-2 is billed per second of generated video, not per clip — and the rate depends on the resolution you pick.

Resolution$/second$/second with reference video(s)
480p$0.165$0.096
720p (default)$0.33$0.21
1080p$0.745$0.52
4k$1.65$1.06

A 5-second 720p generation (the default) costs 5 × $0.33 = $1.65; a 10-second one costs $3.30. Reference-to-video jobs use the lower per-second rate but are billed for the output duration plus up to 15 seconds of reference-clip time (the upstream meters reference footage too), so a 5s 720p ref-to-video job is estimated at (5 + 15) × $0.21 = $4.20. The two -pro SKUs, by contrast, are a flat per-video price regardless of duration. Always read the cost_usd field of the completion response for the exact amount charged — and the estimated_cost_usd on the kickoff response, which already accounts for the duration and resolution you requested.

Choosing between them:

  • seedance-1-5-pro — cheapest ($0.160/clip), newest of the "pro" tier. Best default for short 720p clips where you want a fixed, predictable price.
  • seedance-1-0-pro — the previous pro model ($0.210/clip); keep using it only if you've tuned prompts against it.
  • seedance-2 — the flagship quality tier. Priced per second, so it scales with clip length — reach for it when quality matters more than cost, and keep durations short to control spend.

Text-to-video (default) needs only a prompt. Image-to-video animates a still: pass image_url (an https URL or data: URI) as the start frame. All three Seedance SKUs support both.

import time, requests

BASE = "https://api.gpuniq.com/v1/llm"
HEADERS = {"X-API-Key": "gpuniq_your_key"}

# Image-to-video with the flagship, 6-second clip
start = requests.post(
    f"{BASE}/videos/jobs",
    headers=HEADERS,
    json={
        "model": "seedance-2",
        "prompt": "the product slowly rotates on a marble pedestal, soft studio light",
        "image_url": "https://example.com/product.jpg",  # start frame → image-to-video
        "duration": 6,          # seedance-2 bills per second → 6 × $0.33 = $1.98
    },
).json()["data"]
job_id = start["job_id"]
print("estimated:", start["estimated_cost_usd"])   # 1.98

deadline = time.time() + 300
while time.time() < deadline:
    time.sleep(3)
    d = requests.get(f"{BASE}/videos/jobs/{job_id}", headers=HEADERS).json()["data"]
    if d["status"] == "completed":
        print("video:", d["video"]["url"], "cost:", d["cost_usd"])
        break
    if d["status"] == "failed":
        print("failed:", d.get("error"))
        break

Seedance shares the standard video job API. The two -pro SKUs (seedance-1-0-pro, seedance-1-5-pro) do text-to-video and image-to-video only: duration and image_url are the fields that matter, aspect_ratio is honoured where the upstream supports it, and fields specific to other families (audio, mode, resolution tiers, video_url) are ignored.

seedance-2 accepts more of the request body. On top of t2v and i2v it supports:

  • First + last frame — pass image_url (start) together with last_frame_url (end); the clip interpolates between the two frames.
  • Reference-to-video — condition the render on up to 3 reference clips (video_url / video_urls), up to 9 reference images (reference_image_urls) and up to 3 reference audio tracks (audio_url / audio_urls), in any combination. Audio references need at least one image or video reference alongside them (the API returns a 400 otherwise). A single image_url sent together with any reference input is treated as a subject reference rather than a start frame. First/last-frame and reference-to-video are mutually exclusive — when any reference input is attached, last_frame_url is dropped.
  • Resolution choice480p, 720p (default), 1080p or 4k via resolution; the per-second price scales with it (see the pricing table above).
  • Generated audio track — set audio: true to have the model generate sound for the clip.

duration applies to all three SKUs (seedance-2 accepts 4–15 seconds); seedance-2 is billed per second regardless of which mode you use. On seedance-2, aspect_ratio additionally accepts auto, 21:9, 4:3 and 3:4 — omit it in i2v mode to follow the input frame's geometry.

seedance-2 also accepts a seed. Be aware of what it actually buys: keeping the seed fixed while everything else stays identical makes two renders markedly more alike than two unseeded ones, but they will not be identical — the model is not bit-reproducible, so treat the seed as a similarity control rather than a repeat button. A seeded request is routed to the one upstream that implements the parameter; if that route is unavailable the clip is still rendered, without the seed.

Reference-file limits: videos 2–15s combined, ≤50MB total; images JPG/PNG/WebP; audio MP3/WAV, ≤15s combined. Video resolution and container are unrestricted — 1080p/4K clips and non-MP4 formats (webm, MKV, …) are automatically converted to ≤720p MP4 on our side before dispatch, so send your source footage as-is.

There is no elements parameter. Seedance 2's multi-element / multimodal reference capability is expressed through the fields above — image_url / last_frame_url for frames, video_url / video_urls for reference clips, reference_image_urls for subject/style images and audio_url / audio_urls for reference audio.

Job-based video generation

Same kickoff-then-poll shape as the image-jobs API. The catalog covers text-to-video (t2v), image-to-video (i2v, pass image_url), and video-to-video / motion-control (v2v, pass video_url + image_url for the conditioning frame). Avatar SKUs accept an audio reference URL in the prompt body — see the model-specific docs for the schema.

import time, requests

BASE = "https://api.gpuniq.com/v1/llm"
HEADERS = {"X-API-Key": "gpuniq_your_key"}

# 1. Kickoff
start = requests.post(
    f"{BASE}/videos/jobs",
    headers=HEADERS,
    json={
        "model": "kling-2-6",
        "prompt": "A small black cat slowly turns toward the camera at golden hour",
        "duration": 5,
        "audio": False,          # opt-in, doubles price on Kling 2.6 / 3.0
        "resolution": "1080p",   # 720p | 1080p | 4k (where supported)
    },
).json()
job_id = start["data"]["job_id"]
print(f"job: {job_id}, est cost: ${start['data']['estimated_cost_usd']}")

# 2. Poll — video models deliver in 30-90s; budget 5 minutes for the slowest variants
deadline = time.time() + 300
while time.time() < deadline:
    time.sleep(3)
    r = requests.get(f"{BASE}/videos/jobs/{job_id}", headers=HEADERS).json()
    d = r["data"]
    if d["status"] == "completed":
        print(f"video: {d['video']['url']}")
        print(f"cost: ${d['cost_usd']}, balance: ${d['balance_usd']}")
        break
    if d["status"] == "failed":
        print("failed:", d.get("error"))
        break
Request body
FieldTypeRequiredNotes
modelstringyesSlug from the table above.
promptstringyesUp to 4000 characters.
durationintnoSeconds. Every model has its own legal set — see the duration matrix below. An out-of-range value is a 400, not a silent clamp.
aspect_ratiostringno16:9 (default), 9:16, 1:1 where supported. Seedance 2 additionally accepts auto, 21:9, 4:3, 3:4.
image_urlstringnohttps URL or data URI — enables image-to-video (start frame).
last_frame_urlstringnohttps URL or data URI — end frame for first/last-frame interpolation. Supported on Kling 2.1-Pro / 2.6 / 3.0 / O3, Veo 3.1 (all tiers) and Seedance 2. Requires image_url. The first image is always the FIRST frame, the second the LAST — order matters.
video_urlstringnohttps URL or data: URI — reference clip for motion-control v2v variants and Seedance 2 reference-to-video. Links the host serves as a non-video Content-Type are re-hosted automatically.
video_urlsarraynoUp to 3 https URLs or data: URIs — multiple reference clips for reference-to-video (Seedance 2). Single-reference models use the first entry.
reference_image_urlsarraynoSeedance 2 only — up to 9 https URLs / data URIs of subject or style reference images (JPG/PNG/WebP). Switches the job to reference-to-video mode.
audio_urlstringnoSeedance 2 only — reference audio track (MP3/WAV). Requires at least one image or video reference in the same request.
audio_urlsarraynoSeedance 2 only — up to 3 reference audio tracks, ≤15s combined; merged with audio_url.
resolutionstringnoSeedance 2: 480p / 720p (default) / 1080p / 4k, priced per second per resolution. Kling 3.0 / O3: 720p / 1080p (default) / 4k. Kling 2.6 / 3.0 Motion Control: 720p / 1080p (default). Veo 3.1: 720p / 1080p / 4k. Kling 2.1 / 2.5 Turbo Pro: 1080p only. Default is 1080p for every Kling/Veo SKU when the field is omitted — Seedance 2 is the exception at 720p.
audioboolnoDefault false. Kling 2.6 / 3.0 double the price when true; on Seedance 2 generates the clip's audio track.
modestringnostandard / pro (default) / master for Kling 2.1; turbo for 2.5 Turbo Pro.
seedintnoSeedance 2 only — integer in [-1, 4294967295]; omit or pass -1 for a random seed. Reusing a seed with an otherwise identical request pulls the render toward the earlier one; it does not reproduce it frame for frame. Other models ignore the field.
Duration by model

duration is not a free integer. Each family accepts its own set and anything outside it comes back as a 400 naming the legal values — we do not round to the nearest supported length, because silently rendering (and billing) 8 seconds when you asked for 5 is worse than a refusal.

ModelLegal durationDefault
veo-3-1-lite / veo-3-1-fast / veo-3-1-quality4, 6 or 8 — nothing else8
veo-3-1-fast / -lite with reference_image_urls8 only8
kling-2-1, kling-2-5-turbo-pro, kling-2-65 or 105
kling-3-05 or 10 (multi-shot to 15)5
kling-o3-videoany integer 3–155
kling-2-6-motion-control, kling-3-0-motion-control5 or 105
kling-avatar-pro, kling-avatar-standardany integer 1–155
seedance-2any integer 4–155
seedance-1-0-pro, seedance-1-5-pro55
sora-2-video, sora-2-pro-video1010
sora-2-official, sora-2-pro-official88

Veo does not take duration: 5. It is the single most common 400 on this surface — 5 is the default nearly everywhere else in the catalog, so it gets copied across from a Kling example. Veo's native set is 4 / 6 / 8; ask for 5 and you get invalid_request: Veo 3.1 supports durations of 4, 6 or 8 seconds (got 5).

Reference media: what format to send

Every media field takes a URL string, never an uploaded file part — there is no multipart endpoint. Two forms are accepted:

FormLooks likeWhere it works
Public HTTPS URLhttps://cdn.example.com/ref.jpgevery media field
Data URI (base64 inline)data:image/jpeg;base64,/9j/4AAQ…every media field

Upstream renderers only accept URLs, so a data: URI on a video or audio field is materialised into a real link on our side before dispatch. You do not have to host anything yourself to use a reference clip.

Hosted links are re-hosted when the host misdescribes them. Upstream validators judge a reference clip by the Content-Type the host declares, not by the bytes — so a perfectly good MP4 served as application/octet-stream used to be refused on sight. GPUniq now probes that header and, when it is not a video/* type, fetches the file and re-serves it correctly before dispatch.

Google Drive / Dropbox / OneDrive share links now work. Drive serves files as application/octet-stream with X-Content-Type-Options: nosniff and the original filename (IMG_3752.MOV), which upstream validators reject — this was the single most common cause of "my motion-control job will not start". Such links are re-hosted automatically as of 2026-08-15.

The link must still be directly downloadable without signing in — set the share to "anyone with the link", and use the uc?export=download&id=… form rather than a /view page. A link that returns an HTML sign-in page has no file behind it for us to fetch.

Serving a reference from a bucket with the right Content-Type is still the fastest path: it skips the extra fetch entirely. What the upstream fetch requires, in the order things usually go wrong:

  1. A direct link to the file, not to a viewer page. The response must be the bytes themselves.
  2. No authentication, no redirect chain to a login page. The fetcher is anonymous and has no cookies.
  3. A media Content-Typevideo/mp4, image/jpeg, audio/mpeg. Handled for you when it is wrong, at the cost of one extra round trip.
  4. A file extension that matches the content where possible — some upstreams sniff the URL path before they sniff the body.

Size and format limits:

FieldFormatsLimits
image_url, last_frame_urlJPG, PNG, WebPone image each
reference_image_urlsJPG, PNG, WebP≤9 Seedance 2 · ≤4 Kling O3 · ≤3 Veo 3.1 fast/lite
video_url, video_urlsany common container (MP4, WebM, MKV, MOV)≤3 clips, 2–15 s combined, ≤50 MB total
audio_url, audio_urlsMP3, WAV≤3 tracks, ≤15 s combined

Reference video resolution and container are unrestricted: 1080p/4K and non-MP4 sources are transcoded to ≤720p MP4 on our side before dispatch, so send source footage as-is. Keep in mind that a data URI inflates by ~33% over the raw bytes and counts against the request body limit — for anything above a few MB, host it and send a URL.

Which model takes which reference input

Sending a field the chosen model has no input for is a 400 at submit, not a silent drop. That is deliberate: previously a request carrying reference_image_urls could render on one attempt and be refused on the next, depending on which internal route served it. Capability is now resolved before dispatch, so the same request gets the same answer every time.

Modelimage_urllast_frame_urlreference_image_urlsvideo_url
seedance-2✅ ≤9✅ ≤3
kling-o3-video✅ ≤4
veo-3-1-fast / veo-3-1-lite✅ ≤3 (8 s only)
veo-3-1-quality
kling-2-6-motion-control / kling-3-0-motion-controlrequiredrequired
kling-2-1, kling-2-5-turbo-pro, kling-2-6, kling-3-0✅ (2.1/2.6/3.0)
seedance-1-0-pro, seedance-1-5-pro

For subject or style reference on a model with no reference_image_urls column, use image_url as the start frame — or move to seedance-2, which has the widest reference surface in the catalog.

Motion control is the strictest shape here: kling-2-6-motion-control and kling-3-0-motion-control need both image_url (the character / subject frame) and video_url (the motion to transfer). Either one alone is a 400 naming the missing field.

Response

The kickoff returns immediately with the GPUniq job id, the resolved parameter snapshot, and the cost estimate. Internal routing is opaque — the same job_id is valid across fallbacks, and the user-facing price stays stable.

// POST /v1/llm/videos/jobs
{
  "job_id": "vid_e93e98c7ca5e4982876b",
  "status": "pending",
  "model": "kling-2-6",
  "estimated_cost_usd": 0.315,
  "config": { "resolution": "1080p", "audio": false, "duration": 5, "task": "t2v", "mode": null }
}

// GET /v1/llm/videos/jobs/{job_id} — completed
{
  "job_id": "vid_e93e98c7ca5e4982876b",
  "status": "completed",
  "model": "kling-2-6",
  "video": {
    "url": "https://cdn.example.com/.../output.mp4",
    "stored_url": "https://api.gpuniq.com/v1/llm/media/8xK2p….mp4",
    "stored_expires_in_days": 7
  },
  "cost_usd": 0.315,
  "balance_usd": 9.17825791,
  "config": { "resolution": "1080p", "audio": false, "duration": 5, "task": "t2v", "mode": null }
}

The polling endpoint transparently falls back across internal routes if the first attempt fails — your job_id and the user-facing price stay stable across fallbacks. Internal route identifiers are deliberately omitted from the public response; they live only in admin/operator logs.

Where delivered media lives

Every generated image and video is also mirrored to GPUniq storage, and the response tells you where:

FieldWhat it is
video.url / image.b64_jsonthe original delivery — live immediately, unchanged
video.stored_url / image.urla GPUniq-hosted copy at https://api.gpuniq.com/v1/llm/media/{token}
stored_expires_in_days / expires_in_dayshow long that copy is kept (7 days by default)

Why it matters: the upstream video.url points at the rendering provider's CDN, and its lifetime is set by them, not by us. If you store that link rather than the bytes, it can stop resolving well before you expect. stored_url is ours and lives exactly as long as the field says.

This is a delivery buffer, not an archive. Copies are deleted automatically when the retention window closes and are not recoverable afterwards. If you need media permanently, download it and keep it on your own storage — treat stored_url as the safety net that gets you from "the job finished" to "the bytes are in my bucket", not as the bucket.

Two practical notes. For images the copy is written before the response is sent, so url works the moment you receive it. For videos the copy runs in the background — a clip is tens of megabytes and we would rather not make the delivering poll wait on it — so stored_url becomes valid shortly after the response; url is live immediately either way. And if storage is briefly unavailable the stored_url / url fields are simply absent: the generation still succeeds and is still delivered, because a storage hiccup is never allowed to fail a render you paid for.

Upscaling & Enhancement (Topaz)

Topaz Labs restoration engines — the models behind Gigapixel AI and Video AI — run on a dedicated job surface. No prompt: send a source, pick a model, get the enhanced result back.

Topaz is billed differently from the generative media above. It is a restoration API metered in credits, so instead of a flat per-image rate the cost scales with output size and is charged at $0.14 per credit (≈1 credit per 24 MP of output for precision models — a 4K upscale is about $0.14). You pay only when the job completes.

SurfaceEndpointExample models
ImagePOST /v1/llm/topaz/image/jobs → poll GET .../{job_id}topaz-enhance-standard, topaz-denoise-strong, topaz-sharpen-super-focus
VideoPOST /v1/llm/topaz/video/jobs → poll GET .../{job_id}topaz-video-proteus, topaz-video-starlight, topaz-video-apollo
CatalogGET /v1/llm/topaz/modelslive model list + usd_per_credit

See the Topaz upscaling guide for the full model catalog, parameters, credit-pricing table, and end-to-end examples.

Chat models are sold at 20% below vendor list price.

Fetch the live catalog at any time:

models = client.llm.models()
for model in models["models"]:
    print(model)

The default model is claude-haiku-4-5 — fast, cheap, strong at code.

Long generations & streaming

The edge proxy closes inbound connections after ~100 seconds of streaming silence. A non-streaming request asking for max_tokens > 4096 is rejected up-front with HTTP 400 streaming_required — buffered responses past that length routinely lose to the cap. For long replies, set "stream": true or use the job-based long-poll API.

Your requestWhat to do
≤ 4096 output tokens, fast modelPlain POST /chat/completions works.
> 4096 output tokens OR slow / reasoning modelSet "stream": true.
Client can't speak SSEUse POST /v1/llm/chat/jobs (long-poll).

Reasoning models (Gemini 3 Pro, DeepSeek R1, o3, Claude Opus thinking) burn tokens on hidden chain-of-thought before the visible reply, so they need extra max_tokens headroom — see the Long generations guide for the full streaming / job-based / reasoning-token recipe.

Errors

Every failure returns a stable OpenAI error envelope with a structured code you can branch on — streaming_required, insufficient_balance, model_not_found, rate_limit_per_key, etc. See the Error reference for the complete catalog (29 codes), recovery strategies, and the native vs. OpenAI-compat envelope shapes.

{
  "error": {
    "message": "…human-readable description…",
    "type": "invalid_request_error",
    "code": "streaming_required",
    "doc_url": "https://docs.gpuniq.com/llm/long-generations",
    "meta": { "max_tokens": 8000, "limit": 4096 }
  },
  "status_code": 400,
  "request_id": "…"
}

OpenAI-Compatible Endpoint

Point any OpenAI-compatible tool at GPUniq by setting two environment variables:

OPENAI_API_KEY=gpuniq_your_key
OPENAI_BASE_URL=https://api.gpuniq.com/v1/openai

Every field of the OpenAI Chat Completions protocol is forwarded unchanged: tools, tool_choice, response_format, logprobs, seed, stream, stream_options, etc.

Official OpenAI SDK

from openai import OpenAI

client = OpenAI(
    api_key="gpuniq_your_key",
    base_url="https://api.gpuniq.com/v1/openai",
)

resp = client.chat.completions.create(
    model="claude-opus-4-7",
    messages=[{"role": "user", "content": "Write a binary search in Rust."}],
)
print(resp.choices[0].message.content)

Streaming

Set stream: true — GPUniq returns a text/event-stream with byte-identical OpenAI SSE framing:

stream = client.chat.completions.create(
    model="gpt-5.2",
    messages=[{"role": "user", "content": "Explain MoE in one paragraph."}],
    stream=True,
)
for chunk in stream:
    delta = chunk.choices[0].delta.content
    if delta:
        print(delta, end="", flush=True)

Image Generation

Both API surfaces expose a /images/generations endpoint that matches OpenAI's images.generate protocol. Pass any image slug from the catalog above (e.g. nano-banana-pro, gpt-image-2, flux-2-pro, seedream-4). Billing is flat per returned image — no token accounting.

Image requests route through a multi-tier reliability chain behind the scenes: a per-model priority gateway, two cost-optimised intermediaries, then a generic OpenAI-compatible fallback for safety. The chain is selected automatically per slug, so SDK callers never pick a backend themselves. If the primary fails or returns no image, the next tier is tried within the same HTTP request — you still see one synchronous POST /images/generations and pay for delivered images only.

Heavy generations (Pro / 4K, multi-image batches, high-quality preset) can run up to 5 minutes end-to-end; the connection is held open for that whole budget so SDKs never need to re-poll.

This synchronous surface is legacy. It exists so OpenAI-SDK code runs unmodified against GPUniq — it is no longer the only route for any model, since the job API now covers the whole image catalog (see Which endpoint serves which slug). For anything new, use the job-based API: it returns in under a second, cannot be killed by a proxy's idle-read limit, and bills only on delivery. On this endpoint a client that disconnects mid-render is still charged for an image it never receives.

from openai import OpenAI

client = OpenAI(
    api_key="gpuniq_your_key",
    base_url="https://api.gpuniq.com/v1/openai",
)

resp = client.images.generate(
    model="nano-banana-pro",
    prompt="A cozy mountain cabin at sunrise, cinematic lighting",
    n=2,
    size="1024x1024",
    response_format="b64_json",
)

for i, img in enumerate(resp.data):
    with open(f"out_{i}.png", "wb") as f:
        import base64
        f.write(base64.b64decode(img.b64_json))

Parameters

body
model

Any image slug from the catalog: the Nano Banana family, grok-4-image, gpt-image-2, gpt-image-1-5, gpt-4o-image, flux-2-pro, flux-2-flex, flux-kontext-pro, flux-kontext-max, seedream-4, seedream-4-5, seedream-5-0-lite, or z-image. (Topaz upscaling models live on their own surface — see the Topaz guide.)

body
prompt

Text description of the image you want. Up to 20 000 characters, but the per-model ceiling is lower on several models — 1 000 on gpt-4o-image, gpt-image-1-5 and z-image, 2 000 on the FLUX Kontext and Kling image models, 3 000 on seedream-4-5 / seedream-5-0-lite, 5 000 on nano-banana, seedream-4, flux-2-* and wan-2-7-*. Over the limit you get an UPSTREAM_VALIDATION_ERROR naming the ceiling rather than a truncated render.

body
n

Number of images to generate. 1–4.

body
size

Output resolution and shape. Accepts WIDTHxHEIGHT (e.g. 1024x1024, 2048x2048, 4096x4096), a bare tier (1k / 2k / 4k), or an aspect-ratio string (1:1, 16:9, 21:9, …) when you care about shape rather than exact pixels.

Each model has its own set of output shapes (below). A value that isn't in the model's own set is snapped to the closest listed shape — never dropped and never passed through unchanged. nano-banana-2 renders 1K–4K (default 2K); nano-banana-pro renders 1K–2K (use nano-banana-pro-4k for 4K).

ModelAccepted size
nano-banana1:1 2:3 3:2 3:4 4:3 4:5 5:4 9:16 16:9 21:9
nano-banana-2, nano-banana-pro, nano-banana-pro-4kauto + the ratios above
gpt-image-2 (and the -2k / -4k tiers)auto + the ratios above; exact WIDTHxHEIGHT on the 2K/4K tiers. 4K needs a non-square shape
gpt-image-1-51024x1024, 1024x1536, 1536x1024no 2K/4K tier exists for this model
gpt-4o-image1:1 2:3 3:2
flux-2-pro, flux-2-flexauto 1:1 4:3 3:4 16:9 9:16 3:2 2:3
flux-kontext-pro, flux-kontext-max1:1 4:3 3:4 16:9 9:16 21:9 9:21
seedream-41:1 3:4 4:3 16:9 9:16 3:2 2:3 21:9
seedream-4-5the seedream-4 ratios + 2K / 4K + custom WIDTHxHEIGHT
seedream-5-0-litethe seedream-4 ratios + 2K / 3K + custom WIDTHxHEIGHT
seedream-5-prothe seedream-4 ratios (renders 2K)
z-image1:1 4:3 3:4 16:9 9:16
wan-2-7-image, wan-2-7-image-pro512x512 1024x1024 768x1024 1024x768 576x1024 1024x576
kling-o1-image, kling-o3-image16:9 9:16 1:1 4:3 3:4 3:2 2:3 21:9 (auto on kling-o1-image)
body
quality

Optional upstream quality hint (e.g. standard, hd). Models that don't recognise the value silently fall back to their default.

body
response_format

b64_json returns inline PNG base64 (browser-renderable). url returns a short-lived upstream URL.

body
output_format

Re-encode every delivered image into this format on the server before returning, so the client doesn't need a Pillow / Sharp pipeline. One of:

  • png (default if omitted) — pass-through, lossless.
  • jpeg (alias jpg) — ~10× smaller payload, alpha is flattened onto white because JPEG has no transparency.
  • webp — ~5× smaller at comparable quality, alpha preserved.

Quality for the lossy formats is fixed at 92 — visually indistinguishable from the source PNG. Conversion failures degrade to "return source PNG unchanged" so you always get an image, never a 502 after the upstream has done the expensive work. The MIME type of the converted bytes is echoed back in data[i].mime_type.

body
input_images

Optional reference photos for image-to-image / editing. Each entry is a data: URL, https:// URL, or bare base64 string. Every image model in the catalog accepts them — passing input_images switches the model into its edit mode automatically; you don't pick a separate "edit" slug.

The number of reference images a model reads differs:

ModelReference images
flux-kontext-pro, flux-kontext-max, z-image1
wan-2-7-image, wan-2-7-image-proup to 4
flux-2-pro, flux-2-flexup to 8
everything elseup to 10
kling-o1-imageat least 1 — this model is edit-only

Extras beyond a model's limit are trimmed (the first N are kept). A request the model cannot serve — an edit-only model with no reference image, a prompt over the model's length limit, a reference that is neither a URL nor decodable base64 — comes back as a UPSTREAM_VALIDATION_ERROR with a hint explaining the fix, instead of quietly rendering something that ignores your input.

If the upstream returns fewer images than requested (content-policy rejects, partial failures, etc.), you are billed only for what was delivered.

Claude Code

Claude Code talks to GPUniq directly — GPUniq exposes a native Anthropic Messages API at /v1/messages, so no LiteLLM (or any other) proxy is required. Point Claude Code's environment variables straight at GPUniq:

export ANTHROPIC_BASE_URL=https://api.gpuniq.com
export ANTHROPIC_API_KEY=gpuniq_your_key
export ANTHROPIC_MODEL=claude-opus-4-7              # main model
export ANTHROPIC_SMALL_FAST_MODEL=claude-haiku-4-5  # background/small model
claude

Use any Claude slug from /v1/openai/models for the two model variables. Streaming and tool use work out of the box. All tokens are billed against your GPUniq balance — no separate Anthropic account required.

Set ANTHROPIC_SMALL_FAST_MODEL too: Claude Code calls a smaller "background" model for things like commit messages and titles. If it points at a slug GPUniq doesn't serve, those background calls fail even when the main model works.

Cursor

Settings → Models → Override OpenAI Base URL:

Base URL:  https://api.gpuniq.com/v1/openai
API Key:   gpuniq_your_key
Model:     claude-opus-4-7   # or any slug from /v1/openai/models

Continue.dev / Aider / LiteLLM

Any tool that accepts an OPENAI_BASE_URL works the same way:

export OPENAI_API_KEY=gpuniq_your_key
export OPENAI_BASE_URL=https://api.gpuniq.com/v1/openai

aider --model claude-sonnet-4-6

The OpenAI-compat endpoint returns raw OpenAI response objects (not wrapped in GPUniq's ResponseSchema). Errors use OpenAI's {"error": {"message", "type", "code"}} envelope so SDK retry logic works unchanged.

Native GPUniq SDK

For the fullest feature set — persistent chat sessions, USD balance conversion, usage history — use the native API.

Simple Chat

response = client.llm.chat("claude-haiku-4-5", "Explain how transformers work")
print(response)

Chat Completion (Full)

data = client.llm.chat_completion(
    messages=[
        {"role": "system", "content": "You are a helpful AI assistant."},
        {"role": "user", "content": "What is gradient descent?"},
    ],
    model="claude-sonnet-4-6",
    temperature=0.7,
    max_tokens=1000,
    top_p=0.9,
)

print(data["content"])
print(f"Tokens used: {data['tokens_used']}  cost: ${data['cost_usd']:.6f}")

Parameters

body
messages

List of message objects with role ("system", "user", "assistant") and content.

body
model

Model slug (e.g., claude-opus-4-7, gpt-5.2, gemini-3-pro). Defaults to claude-haiku-4-5.

body
max_tokens

Maximum tokens in the response.

body
temperature

Sampling temperature (0.0-2.0). Higher = more creative.

body
top_p

Top-p nucleus sampling parameter.

Account Balance

Chat and image requests are billed directly against your GPUniq account balance in USD — there is no separate "token pool" anymore. Each call deducts the model's blended retail rate × the tokens it actually consumed (or per-image flat rate for image models). Prepaid token packages and ruble-to-token conversions are no longer required and the corresponding endpoints have been retired.

balance = client.llm.balance()
print(f"Available: ${balance['balance_usd']:.4f} USD")

Top up the balance from the web dashboard → Billing (Stripe / YooKassa / crypto). The balance is shared with every other GPUniq surface — GPU rentals, volume storage, image generations — so a single deposit covers the whole platform.

Usage History

Per-request detail with prompt / completion / cached / reasoning tokens and the USD cost charged at retail. Backed by the /v1/llm/usage/history endpoint; pair it with /v1/llm/usage/breakdown for daily / weekly aggregates.

history = client.llm.usage_history(limit=50, offset=0)
for log in history["logs"]:
    print(f"{log['model']}: {log['total_tokens']} tokens — ${log['cost_usd']:.6f}")

Chat Sessions

Persistent conversations stored server-side — the model sees the full history on every call:

# Create a session
session = client.llm.create_chat_session(
    model="claude-sonnet-4-6",
    title="Research Assistant",
)

# Send messages within the session
reply = client.llm.send_message(
    chat_id=session["id"],
    message="What are the key papers on attention mechanisms?",
    temperature=0.5,
)

# List all sessions
sessions = client.llm.list_chat_sessions(limit=50)

# Get a session with full message history
full = client.llm.get_chat_session(chat_id=session["id"])

# Update title
client.llm.update_chat_session(chat_id=session["id"], title="New Title")

# Delete
client.llm.delete_chat_session(chat_id=session["id"])

Generate Terminal Commands

Convert natural language to a ranked list of shell commands with danger annotations:

cmds = client.llm.generate_commands(
    prompt="find all Python files larger than 1MB and sort by size",
    max_commands=5,
)
for c in cmds["commands"]:
    print(f"[{c['danger']}] {c['command']}  # {c['description']}")

API Key Management

API keys are created from the web dashboard (LLM API Keys) and sent as Authorization: Bearer gpuniq_... on OpenAI-compat routes, or X-API-Key: gpuniq_... on native routes.

Rate limit: 120 req/min per key, sliding window.