LLM & Media API
Chat, reasoning, images, video, music, speech and upscaling through one API key and one USD balance. Claude, GPT, Gemini, Grok, DeepSeek, Qwen, Nano Banana, Seedance, Veo and more.
One API key, one balance, one usage log — 370+ models covering text, images, video, music, speech and restoration.
There are two ways in:
- Native GPUniq API (
/v1/llm/*) — wrapped responses with USD cost on every call, server-side chat sessions, the generations library, and the async job surfaces for every media kind. - Drop-in compatibility —
/v1/openai/*replacesapi.openai.com/v1, and/v1/messagesis a native Anthropic Messages endpoint. Claude Code, Cursor, Continue.dev, Aider, LiteLLM and the official SDKs work without code changes.
Where to go
Models & pricing
What's in the catalog, how each media kind is billed, and how to read live prices instead of a stale table.
Chat & completions
Native completions, persistent sessions, shareable transcripts, shell-command generation.
OpenAI & Anthropic compatibility
/v1/openai, the Responses API, /v1/messages, and tool setup for Claude
Code and Cursor.
Image generation
Nano Banana, GPT Image 2, FLUX.2, Seedream, Midjourney — job API, billed per image.
Video generation
Veo 3.1, Kling, Seedance, Sora, Wan, Runway — references, durations, first/last frame.
Music & speech
Text-to-music with lyrics, ElevenLabs speech, dialogue, sound effects, voice isolation.
Topaz upscaling
Gigapixel and Video AI restoration engines, metered in credits.
Long generations
Streaming, the 4096-token non-streaming cap, and reasoning-token headroom.
Error reference
Every error code, what triggers it, and how to recover.
Your first call
curl https://api.gpuniq.com/v1/llm/chat/completions \
-H "X-API-Key: gpuniq_your_key" \
-H "Content-Type: application/json" \
-d '{
"messages": [{"role": "user", "content": "Explain how transformers work"}],
"model": "claude-haiku-4-5"
}'
Omit model and you get claude-haiku-4-5 — fast, cheap, strong at code.
Authentication & limits
API keys are created in the web dashboard under LLM API Keys. A key starts
with gpuniq_ and works on every surface: marketplace, instances, volumes and
every model endpoint here.
| Surface | Header |
|---|---|
Native /v1/llm/* | X-API-Key: gpuniq_... |
OpenAI-compatible /v1/openai/*, /v1/responses | Authorization: Bearer gpuniq_... |
Anthropic /v1/messages | x-api-key: gpuniq_... or Authorization: Bearer gpuniq_... |
A dashboard JWT works in place of a key on native routes.
Rate limit: 120 requests/minute per key, sliding window. It is a default, not a ceiling — accounts can be raised, and a raised account limit cascades to its keys. The Python SDK retries automatically when it hits the limit.
Getting 403 with error code: 1010 on any endpoint? That is not
the endpoint — it is your HTTP client's User-Agent.
Our edge blocks the default Python-urllib/3.x User-Agent outright.
The block is UA-based and applies to every route equally — sync,
async, chat — so it is easy to misread as "the synchronous endpoint is
broken" when the async one fails identically. requests, httpx,
curl, the official SDKs and any client that sets its own UA are all
unaffected.
# 403, error code 1010 — blocked at the edge, never reaches the API
urllib.request.urlopen(urllib.request.Request(url, data=body))
# Works — requests sends its own User-Agent
requests.post(url, json=payload, headers={"X-API-Key": KEY})
# Works — urllib with an explicit User-Agent
req = urllib.request.Request(url, data=body, headers={
"X-API-Key": KEY,
"Content-Type": "application/json",
"User-Agent": "my-app/1.0", # <- the one line that matters
})
A 1010 body is always plain text, never our JSON envelope. If you get
JSON back, the request reached us and the error is a real API error —
look it up in the Error Reference.
Long generations & streaming
The edge proxy closes inbound connections after ~100 seconds of streaming silence. A non-streaming request asking for max_tokens > 4096 is not refused: GPUniq streams from the upstream and assembles the reply itself, flushing whitespace keep-alive bytes to you until the final JSON is ready (the response carries X-GPUniq-Long-Generation: keepalive). On that path an upstream failure arrives as HTTP 200 whose body is {"error": …} — check for it. For long replies, prefer "stream": true or the job-based long-poll API.
| Your request | What to do |
|---|---|
| ≤ 4096 output tokens, fast model | Plain POST /chat/completions works. |
| > 4096 output tokens OR slow / reasoning model | Set "stream": true. |
| Client can't speak SSE | Use POST /v1/llm/chat/jobs (long-poll). |
Reasoning models (Gemini 3 Pro, DeepSeek R1, o3, Claude Opus thinking) burn
tokens on hidden chain-of-thought before the visible reply, so they need extra
max_tokens headroom — see Long generations for the
full streaming / job-based / reasoning-token recipe.
Errors
Every failure returns a stable OpenAI error envelope with a structured code you can branch on — insufficient_balance, model_not_found, rate_limit_per_key, provider_unavailable, etc. See the Error reference for the complete catalog, recovery strategies, and the native vs. OpenAI-compat envelope shapes.
{
"error": {
"message": "Unknown model 'gpt-9'. Fetch /v1/openai/models for the current catalog.",
"type": "invalid_request_error",
"code": "model_not_found"
},
"status_code": 400,
"request_id": "…"
}
Where delivered media lives
Generated images and videos are mirrored to GPUniq storage: the response carries both the upstream URL and a GPUniq-hosted copy, kept for 7 days. That copy is a delivery buffer, not an archive — download the bytes if you need them to last. Music and audio are not mirrored; download those as soon as the job completes. Details: Where delivered media lives.