LLM & Media APIEmbeddings & rerank

Embeddings & rerank

OpenAI-compatible embeddings and Cohere-style reranking on the same key and balance — models, request shapes, billing and errors.

Turn text into vectors for search, clustering and RAG, and re-order retrieved passages by relevance to a query. Both endpoints use your usual API key and USD balance, and every call lands in the same usage log as chat and media.

EndpointWire formatBase URL
POST /v1/openai/embeddingsOpenAI Embeddings — the official SDKs work unchangedhttps://api.gpuniq.com/v1/openai
POST /v1/openai/rerankCohere / Jina rerank (query + documents → results)https://api.gpuniq.com/v1/openai

Authenticate with Authorization: Bearer gpuniq_… or X-API-Key: gpuniq_…, as on every /v1/openai/* route.

Models

Embedding models

ModelVector sizedimensions
text-embedding-3-large3072Yes — any smaller size
text-embedding-3-small1536Yes — any smaller size
text-embedding-ada-0021536No
gemini-embedding-0013072No
gemini-embedding-2-preview3072No

Rerank models

ModelFamily
qwen3-rerankAlibaba Qwen3
qwen3-reranker-0.6bAlibaba Qwen3, 0.6B
bge-reranker-v2-m3BAAI BGE-M3
bge-reranker-v2-m3-proBAAI BGE-M3
bce-reranker-base-v1NetEase Youdao BCE

In the catalog these models carry media_kind: "embedding" or "rerank", so they never show up among chat models. GET /v1/llm/models/catalog is the source of truth for the list and the live per-token price.

Embeddings

curl https://api.gpuniq.com/v1/openai/embeddings \
  -H "Authorization: Bearer gpuniq_your_key" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "text-embedding-3-small",
    "input": ["GPUniq rents GPUs by the hour", "How do I rent a GPU?"]
  }'
from openai import OpenAI

client = OpenAI(api_key="gpuniq_your_key", base_url="https://api.gpuniq.com/v1/openai")

resp = client.embeddings.create(
    model="text-embedding-3-large",
    input=["GPUniq rents GPUs by the hour", "How do I rent a GPU?"],
    dimensions=1024,  # text-embedding-3 models only
)
vectors = [item.embedding for item in resp.data]
FieldTypeNotes
modelstringRequired. One of the embedding models above.
inputstring or array of stringsRequired. One vector per element, returned in the same order (index).
dimensionsintegertext-embedding-3-large and -3-small only — any other model answers 400 dimensions_unsupported instead of silently returning its native size.
encoding_format"float" (default) or "base64"base64 returns each vector as a base64 string of little-endian float32 — smaller over the wire.
userstringOptional end-user identifier, passed through.

The vendor's own limits apply to its models — for the OpenAI models, up to 2048 inputs per request and 8192 tokens per input.

{
  "object": "list",
  "model": "text-embedding-3-large",
  "data": [
    { "object": "embedding", "index": 0, "embedding": [0.0123, -0.0456, "…"] },
    { "object": "embedding", "index": 1, "embedding": [0.0089, -0.0312, "…"] }
  ],
  "usage": { "prompt_tokens": 14, "total_tokens": 14, "cost": 0.000001 }
}

Rerank

Send a query and the passages you retrieved; get them back ordered by relevance.

curl https://api.gpuniq.com/v1/openai/rerank \
  -H "Authorization: Bearer gpuniq_your_key" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "qwen3-rerank",
    "query": "what speeds up deep learning",
    "documents": [
      "Paris is the capital of France",
      "A GPU accelerates neural network training",
      "Bananas are yellow"
    ],
    "top_n": 2,
    "return_documents": true
  }'
import httpx

resp = httpx.post(
    "https://api.gpuniq.com/v1/openai/rerank",
    headers={"Authorization": "Bearer gpuniq_your_key"},
    json={
        "model": "bge-reranker-v2-m3",
        "query": "what speeds up deep learning",
        "documents": ["Paris is the capital of France", "A GPU accelerates neural network training"],
    },
    timeout=60,
)
best = resp.json()["results"][0]["index"]
FieldTypeNotes
modelstringRequired. One of the rerank models above.
querystringRequired.
documentsarray of stringsRequired, non-empty.
top_nintegerHow many results to return. Default: every document.
return_documentsbooleanEcho each document's text back in the result. Default false.
{
  "results": [
    { "index": 1, "relevance_score": 0.75, "document": { "text": "A GPU accelerates neural network training" } },
    { "index": 2, "relevance_score": 0.203, "document": { "text": "Bananas are yellow" } }
  ],
  "model": "qwen3-rerank",
  "usage": { "prompt_tokens": 80, "total_tokens": 80, "cost": 0.000006 }
}

results are sorted by relevance_score, highest first; index points into your documents array. Scores are comparable within one model's response — each model has its own scale, so a threshold tuned on one model does not carry over to another.

Billing

  • Input tokens only. Embeddings bill the tokens of every input; rerank bills the query plus the documents as the model counts them. There are no output tokens.
  • 20% below the vendor's official price per token. The live rate is retail_input_usd_per_mtok in the catalog; personal discounts apply on top, exactly as for chat.
  • usage.cost in the response is what was debited for the call, in USD, rounded to the micro-dollar ($0.000001) — a very short input on a cheap model can round to $0. The call also appears in GET /v1/llm/usage/requests like any other request.
  • A request that fails — validation, no upstream — is not charged.

Errors

Errors use the OpenAI envelope.

codeHTTPWhen
model_required400model is missing or empty.
model_not_found400The slug is not an embedding model (on /embeddings) or not a rerank model (on /rerank) — e.g. a reranker sent to /embeddings.
input_required400Empty input; or a rerank call without query or without documents.
dimensions_unsupported400dimensions on a model other than text-embedding-3-large / -3-small.
model_is_retrieval400An embedding or rerank slug was sent to /chat/completions.
insufficient_balance402The balance does not cover the request's estimated cost.
rate_limit_per_key / rate_limit_per_user429Per-minute limits — shared with every other call on the key and account.
provider_unavailable503The model's upstream is temporarily unavailable. Retry with backoff; nothing was charged.