LLM & Media APIOverview

LLM & Media API

Chat, reasoning, images, video, music, speech and upscaling through one API key and one USD balance. Claude, GPT, Gemini, Grok, DeepSeek, Qwen, Nano Banana, Seedance, Veo and more.

One API key, one balance, one usage log — 370+ models covering text, images, video, music, speech and restoration.

There are two ways in:

  1. Native GPUniq API (/v1/llm/*) — wrapped responses with USD cost on every call, server-side chat sessions, the generations library, and the async job surfaces for every media kind.
  2. Drop-in compatibility — /v1/openai/* replaces api.openai.com/v1, and /v1/messages is a native Anthropic Messages endpoint. Claude Code, Cursor, Continue.dev, Aider, LiteLLM and the official SDKs work without code changes.

Where to go

Your first call

curl https://api.gpuniq.com/v1/llm/chat/completions \
  -H "X-API-Key: gpuniq_your_key" \
  -H "Content-Type: application/json" \
  -d '{
    "messages": [{"role": "user", "content": "Explain how transformers work"}],
    "model": "claude-haiku-4-5"
  }'

Omit model and you get claude-haiku-4-5 — fast, cheap, strong at code.

Authentication & limits

API keys are created in the web dashboard under LLM API Keys. A key starts with gpuniq_ and works on every surface: marketplace, instances, volumes and every model endpoint here.

SurfaceHeader
Native /v1/llm/*X-API-Key: gpuniq_...
OpenAI-compatible /v1/openai/*, /v1/responsesAuthorization: Bearer gpuniq_...
Anthropic /v1/messagesx-api-key: gpuniq_... or Authorization: Bearer gpuniq_...

A dashboard JWT works in place of a key on native routes.

Rate limit: 120 requests/minute per key, sliding window. It is a default, not a ceiling — accounts can be raised, and a raised account limit cascades to its keys. The Python SDK retries automatically when it hits the limit.

Getting 403 with error code: 1010 on any endpoint? That is not the endpoint — it is your HTTP client's User-Agent.

Our edge blocks the default Python-urllib/3.x User-Agent outright. The block is UA-based and applies to every route equally — sync, async, chat — so it is easy to misread as "the synchronous endpoint is broken" when the async one fails identically. requests, httpx, curl, the official SDKs and any client that sets its own UA are all unaffected.

# 403, error code 1010 — blocked at the edge, never reaches the API
urllib.request.urlopen(urllib.request.Request(url, data=body))

# Works — requests sends its own User-Agent
requests.post(url, json=payload, headers={"X-API-Key": KEY})

# Works — urllib with an explicit User-Agent
req = urllib.request.Request(url, data=body, headers={
    "X-API-Key": KEY,
    "Content-Type": "application/json",
    "User-Agent": "my-app/1.0",     # <- the one line that matters
})

A 1010 body is always plain text, never our JSON envelope. If you get JSON back, the request reached us and the error is a real API error — look it up in the Error Reference.

Long generations & streaming

The edge proxy closes inbound connections after ~100 seconds of streaming silence. A non-streaming request asking for max_tokens > 4096 is not refused: GPUniq streams from the upstream and assembles the reply itself, flushing whitespace keep-alive bytes to you until the final JSON is ready (the response carries X-GPUniq-Long-Generation: keepalive). On that path an upstream failure arrives as HTTP 200 whose body is {"error": …} — check for it. For long replies, prefer "stream": true or the job-based long-poll API.

Your requestWhat to do
≤ 4096 output tokens, fast modelPlain POST /chat/completions works.
> 4096 output tokens OR slow / reasoning modelSet "stream": true.
Client can't speak SSEUse POST /v1/llm/chat/jobs (long-poll).

Reasoning models (Gemini 3 Pro, DeepSeek R1, o3, Claude Opus thinking) burn tokens on hidden chain-of-thought before the visible reply, so they need extra max_tokens headroom — see Long generations for the full streaming / job-based / reasoning-token recipe.

Errors

Every failure returns a stable OpenAI error envelope with a structured code you can branch on — insufficient_balance, model_not_found, rate_limit_per_key, provider_unavailable, etc. See the Error reference for the complete catalog, recovery strategies, and the native vs. OpenAI-compat envelope shapes.

{
  "error": {
    "message": "Unknown model 'gpt-9'. Fetch /v1/openai/models for the current catalog.",
    "type": "invalid_request_error",
    "code": "model_not_found"
  },
  "status_code": 400,
  "request_id": "…"
}

Where delivered media lives

Generated images and videos are mirrored to GPUniq storage: the response carries both the upstream URL and a GPUniq-hosted copy, kept for 7 days. That copy is a delivery buffer, not an archive — download the bytes if you need them to last. Music and audio are not mirrored; download those as soon as the job completes. Details: Where delivered media lives.