unopenrouter is an OpenAI-compatible API on your own Claude or ChatGPT subscription. Sign in, connect your own accounts and get keys that use only your accounts. Sign-up is by invitation, or open when an admin turns it on.
Quick start
https://api.unopenrouter.com/v1
Authorization: Bearer $KEY with a key like sk-unor-…, made after you
log in or with the management API.
Examples
curl https://api.unopenrouter.com/v1/chat/completions \
-H "Authorization: Bearer $UNOR_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "claude-sonnet-5-5",
"messages": [{"role": "user", "content": "Write a haiku about caches."}]
}'
import os
from openai import OpenAI
client = OpenAI(
base_url="https://api.unopenrouter.com/v1",
api_key=os.environ["UNOR_KEY"],
)
reply = client.chat.completions.create(
model="claude-sonnet-5-5",
messages=[{"role": "user", "content": "Write a haiku about caches."}],
)
print(reply.choices[0].message.content)
stream = client.chat.completions.create(
model="gpt-6-luna",
messages=[{"role": "user", "content": "Count to ten, slowly."}],
stream=True,
)
for chunk in stream:
if chunk.choices and chunk.choices[0].delta.content:
print(chunk.choices[0].delta.content, end="", flush=True)
import OpenAI from "openai";
const client = new OpenAI({
baseURL: "https://api.unopenrouter.com/v1",
apiKey: process.env.UNOR_KEY,
});
const reply = await client.chat.completions.create({
model: "gpt-6-luna",
messages: [{ role: "user", content: "Write a haiku about caches." }],
});
console.log(reply.choices[0].message.content);
# client = OpenAI(...) as in the Python example
first = client.responses.create(
model="claude-sonnet-5-5",
input="Pick a city at random and remember it.",
)
second = client.responses.create(
model="claude-sonnet-5-5",
previous_response_id=first.id,
input="Which city did you pick?",
)
print(second.output_text)
Endpoints
context_length and pricingModels
| API price · per 1M | ||||||
|---|---|---|---|---|---|---|
| Model | Provider | Context | With sub | In | Cache read | Out |
| claude-opus-5-5 | claude | 200k | $0* | $4 | $0.20 | $20 |
| claude-fable-5-1 | claude | 200k | $0* | $10 | $0.25 | $50 |
| claude-sonnet-5-5 | claude | 200k | $0* | $2 | $0.20 | $10 |
| claude-haiku-4-5-20251001 | claude | 200k | $0* | $1 | $0.10 | $5 |
| gpt-6.1-sol | codex | 258k | $0* | $2 | $0.10 | $10 |
| gpt-6-astra | codex | 258k | $0* | $10 | $1 | $50 |
| gpt-6-sol | codex | 258k | $0* | $2 | $0.20 | $10 |
| gpt-6-luna | codex | 258k | $0* | $0.10 | $0.01 | $0.50 |
*Uses your account's subscription usage limits. In practice that typically comes out 10–15× cheaper than paying for the API. API prices are list prices in USD per million tokens; key spending limits count in them.
Supported
Streaming. Server-sent events with the same chunks as OpenAI,
stream_options.include_usage and a closing data: [DONE]. Responses streams use the typed
response.* events.
Function calling, emulated. Your tools are described to the model in the system
prompt; to call one, the model answers with a strict JSON object that the gateway parses into
tool_calls (retried once if it doesn't parse). Streamed tool calls work, as do parallel calls and
tool_choice auto, none, required or a named function.
JSON mode, best effort. response_format json_object or
json_schema (Responses: text.format): the gateway instructs the model, checks the reply,
against the schema when there is one, and returns 422 if it still isn't valid after one retry.
A JSON-mode stream arrives in one piece at the end.
Conversations and caching. The turns of one conversation stay on one subscription
account, so the vendor's prompt cache hits. A conversation is matched by its message history,
prompt_cache_key or previous_response_id. Cached tokens are in
usage.prompt_tokens_details.cached_tokens (Responses:
usage.input_tokens_details.cached_tokens).
Responses API with previous_response_id and store
(default true).
Images. image_url parts, https or data URLs; Responses:
input_image.
OpenRouter extras. A fallback list,
"models": ["claude-sonnet-5-5", "gpt-6-luna"], is tried in order and the response's
model says which one answered. /v1/models entries carry context_length
and pricing. Usage per key is shown after you log in.
Request ids. Every response has an x-request-id header; a 429 has
retry-after.
Effort and length. reasoning_effort (Responses:
reasoning.effort) maps to the model's effort levels. max_tokens,
max_completion_tokens and max_output_tokens are best effort, and Claude only.
Not supported
n above 1 and logprobs / top_logprobs return
400.
temperature and top_p are accepted and ignored, since the
CLIs don't expose them, as are seed, stop, the penalties and
logit_bias.
No embeddings, audio, image generation, files, batch, assistants, realtime or hosted tools (web search, file search, code interpreter): 400 or 404.
Keys, limits and errors
Limits. A key can have limits on requests per minute and per second, tokens a day, spend a day and a month (in estimated dollars), and the models it may use. Every IP also has a safety net of 600 requests a minute.
Your accounts only. A key runs only on the subscription accounts its owner connected, spread over them by the owner's routing policy.
Private-only keys. A key with public_allowed: false works on the
server's private network but not through api.unopenrouter.com.
Errors. A limit hit is an OpenAI-style 429 with
retry-after; a model the key may not use is 403; a disabled, revoked or unknown key
is 401. Errors use OpenAI's object:
{"error": {"message", "type", "param", "code"}}.
Management API
For code that makes keys. Base URL
https://api.unopenrouter.com/api/v1; JSON is snake_case and errors are
{"error": {"message", "code"}}.
Auth. Authorization: Bearer $ADMIN_KEY: an sk-unor-
key with scope admin. Admin keys work for /v1 inference too; inference keys get
403 here. Only an admin, signed in or with an admin key, can make admin keys.
Rate limit and log. 60 calls a minute, apart from inference. Every call is logged.
{"data": [Key]}{"data": Key, "key": "sk-unor-…"}; the secret is in this answer onlyname, disabled, public_allowed, limits (null clears one) or allowed_models?days=30&key_id=…&model=…{"mode": "pressure" | "order" | "pinned", "pinned"}{"note"} optional: 201 with its linkA new key takes name and, optionally, scope (inference or
admin), public_allowed (default true), limits
(rpm, rps, daily_tokens, daily_usd,
monthly_usd; any left out is no limit) and allowed_models (null is
every model).
# Make a key with limits; the secret is in "key", shown once
curl https://api.unopenrouter.com/api/v1/keys \
-H "Authorization: Bearer $ADMIN_KEY" \
-H 'Content-Type: application/json' \
-d '{"name":"pokemon-bot","limits":{"rpm":60,"daily_usd":5},"allowed_models":["claude-sonnet-5-5"]}'
# List the keys
curl https://api.unopenrouter.com/api/v1/keys \
-H "Authorization: Bearer $ADMIN_KEY"
# Disable one ({"disabled":false} turns it back on)
curl -X PATCH https://api.unopenrouter.com/api/v1/keys/$KEY_ID \
-H "Authorization: Bearer $ADMIN_KEY" \
-H 'Content-Type: application/json' \
-d '{"disabled":true}'
# Revoke it for good
curl -X DELETE https://api.unopenrouter.com/api/v1/keys/$KEY_ID \
-H "Authorization: Bearer $ADMIN_KEY"
# Its usage over the last 7 days
curl "https://api.unopenrouter.com/api/v1/usage?days=7&key_id=$KEY_ID" \
-H "Authorization: Bearer $ADMIN_KEY"