# unopenrouter: instructions for a coding agent

unopenrouter turns the user's own Claude and ChatGPT subscriptions into an OpenAI-compatible API. Use the official
`openai` SDKs (or plain HTTP) exactly as you would with OpenAI. Only the base URL and the key differ.

## Connection

- Base URL: `https://api.unopenrouter.com/v1`
- Auth: `Authorization: Bearer $UNOR_KEY`. Keys look like `sk-unor-…`. Read the key from the `UNOR_KEY`
  environment variable; never hard-code it or commit it.

## Setup (done once, by the user)

1. Sign in at https://unopenrouter.com/dashboard with Google or email. New accounts need an invite code.
2. Connect a machine or a device, so requests have somewhere to run:
   - a **runner**: a small program on the user's own computer, Raspberry Pi or server. Their logins stay there.
   - or an **exit node**: a phone, TV or computer at home that lends its internet address.
3. Make a key in the dashboard (Keys → New key) and put it in `UNOR_KEY`.

Without a runner or an exit node online, every request answers 503 (`runner_offline` or `exit_offline`).

### Runner

1. In the dashboard: Runners → Add runner. It gives a one-time pairing code (15 minutes).
2. On an always-on Debian, Ubuntu or Raspberry Pi OS box:
   `curl -fsSL https://unopenrouter.com/install-runner.sh | sudo bash`, then
   `sudo -u unopenrouter-runner unopenrouter-runner pair <code>`. On the user's own computer (Node 22+):
   `npx https://unopenrouter.com/unopenrouter-runner.tgz pair <code>`, then run it with
   `npx https://unopenrouter.com/unopenrouter-runner.tgz` and leave it running. Docker:
   `docker build -t unopenrouter-runner https://unopenrouter.com/unopenrouter-runner-docker.tgz`.
3. Log in on the runner: `login claude` (paste a token from `claude setup-token`, made on the user's own home
   computer) and/or `login codex` (a device code).

It needs 1 GB RAM (with swap), 1 vCPU and 3 GB of disk. For uptime, run a second runner with the same logins (the
same Claude token; a separate `login codex` on each, never a copied `auth.json`): requests fail over between runners,
and the API is down only when every runner is. Guide: https://unopenrouter.com/runner.

### Exit node

A device that runs Tailscale (Android phone, iPhone or iPad, Apple TV, Android or Google TV, Mac, PC, Raspberry Pi)
turns on "Run as exit node", is shared with "Allow use as an exit node", and the share link goes in the dashboard
(Exit nodes → Add device). The CLIs then run on the server with the user's login, and all their traffic leaves
through that device. Guide: https://unopenrouter.com/exit-node.

## Models

- `claude-opus-5-5` (Anthropic, 1M context): Claude's strongest model: hard reasoning, coding, agents. Slowest.
- `claude-fable-5-1` (Anthropic, 1M context): Claude's premium model, with its own weekly limit: when quality matters most.
- `claude-sonnet-5-5` (Anthropic, 1M context): The balanced default for most work.
- `claude-haiku-4-5-20251001` (Anthropic, 200k context): Fastest Claude: classification, extraction, short answers.
- `gpt-6.1-sol` (OpenAI, 272k context): The newest GPT all-rounder.
- `gpt-6-astra` (OpenAI, 272k context): GPT's most capable model: complex reasoning. Slowest.
- `gpt-6-sol` (OpenAI, 272k context): A general GPT model.
- `gpt-6-luna` (OpenAI, 272k context): Fastest GPT: simple, high-volume tasks.

`GET /v1/models` lists them, with `context_length` and per-token API `pricing`. On a subscription a request costs
nothing per token: it uses the account's usage limits, typically 10–25× cheaper than the API.

## Chat Completions

Python:

```python
import os
from openai import OpenAI

client = OpenAI(base_url="https://api.unopenrouter.com/v1", api_key=os.environ["UNOR_KEY"])
r = client.chat.completions.create(model="claude-sonnet-5-5", messages=[{"role": "user", "content": "Hello"}])
print(r.choices[0].message.content)

for chunk in client.chat.completions.create(model="gpt-6-luna", messages=[{"role": "user", "content": "Count to 5"}], stream=True):
    if chunk.choices:
        print(chunk.choices[0].delta.content or "", end="")
```

Node:

```js
import OpenAI from "openai";

const client = new OpenAI({ baseURL: "https://api.unopenrouter.com/v1", apiKey: process.env.UNOR_KEY });
const r = await client.chat.completions.create({ model: "claude-sonnet-5-5", messages: [{ role: "user", content: "Hello" }] });
console.log(r.choices[0].message.content);

const stream = await client.chat.completions.create({ model: "gpt-6-luna", messages: [{ role: "user", content: "Count to 5" }], stream: true });
for await (const chunk of stream) process.stdout.write(chunk.choices[0]?.delta?.content ?? "");
```

curl:

```sh
curl https://api.unopenrouter.com/v1/chat/completions -H "Authorization: Bearer $UNOR_KEY" -H "Content-Type: application/json" \
  -d '{"model":"claude-sonnet-5-5","messages":[{"role":"user","content":"Hello"}]}'
```

Streaming is OpenAI's SSE: `chat.completion.chunk` events, `stream_options: {"include_usage": true}` for a
final usage chunk, then `data: [DONE]`.

## Responses API and caching

- `POST /v1/responses` works, with `previous_response_id` to continue a conversation and `store` (default true).
  `GET`/`DELETE /v1/responses/{id}` and `GET /v1/responses/{id}/input_items` work too. Responses streams use
  the typed `response.*` events.
- Caching: each conversation stays on one subscription account and continues the same session, so the
  provider's prompt cache hits. Keep the system prompt (`instructions`) byte-identical between turns, and
  either resend the whole message history unchanged plus the new message, or use `previous_response_id`.
  `prompt_cache_key` keeps related requests on one account. Cached tokens are reported in
  `usage.prompt_tokens_details.cached_tokens` (Responses: `usage.input_tokens_details.cached_tokens`).
- Users with "Don't store my data" on (the default) have conversations and stored responses deleted after
  three hours without use; after that, `previous_response_id` returns 404 and a resent history starts fresh.

## Tools and JSON

- Function calling (`tools`, `tool_choice`: auto, none, required, or a named function; parallel calls) works
  and streams, but it is emulated: the tools are described to the model in the system prompt and its JSON
  answer is parsed into `tool_calls` (retried once if it doesn't parse). Keep tool schemas simple and
  descriptions clear.
- JSON mode (`response_format` `json_object` or `json_schema`; Responses `text.format`) is best effort: the
  reply is validated (against the schema when given), retried once, and a 422 comes back if it still fails.
  JSON-mode streams arrive in one piece at the end.
- Images: `image_url` parts (https or base64 data URLs) and Responses `input_image`.

## Not supported

`n > 1`, `logprobs` and `top_logprobs` return 400. `temperature`, `top_p`, `seed`, `stop`, penalties and
`logit_bias` are accepted and ignored. `max_tokens` / `max_completion_tokens` / `max_output_tokens` are best
effort (Claude models only). No embeddings, audio, image generation, files, batch, assistants, realtime or
hosted tools (web search, file search, code interpreter).

## Errors and limits

- Errors use OpenAI's object, `{"error": {"message", "type", "param", "code"}}`, and its status codes: 400 bad
  request, 401 bad or missing key, 403 model not allowed for the key, 404 unknown model or response, 422 JSON
  mode failed, 429 limit reached, 503 no runner or exit node online (`runner_offline`, `exit_offline`), other
  5xx upstream trouble. Every response has `x-request-id`.
- A 429 carries `retry-after` (seconds). Wait that long before retrying; when it also says
  `x-should-retry: false` (a daily or monthly cap, or every account at its usage limit), don't retry soon.
  The official SDKs handle both.
- Keys can carry limits set by their owner: requests per minute and second, daily tokens, daily and monthly
  spend, allowed models.

## Fallbacks

Send `"models": ["claude-sonnet-5-5", "gpt-6-luna"]` (OpenRouter style) instead of, or after, `model`: each is
tried in order until one answers, and the response's `model` says which one did.

## Management API (making keys from code)

Base URL `https://api.unopenrouter.com/api/v1`, with `Authorization: Bearer $ADMIN_KEY`: a key with scope `admin`,
made in the dashboard. Inference keys get 403 here. JSON is snake_case; errors are `{"error": {"message", "code"}}`.
60 calls a minute per admin key, apart from inference.

- `GET /keys` lists the keys: `{"data": [Key]}`.
- `POST /keys` makes one: `{"name", "scope"?, "public_allowed"?, "limits"?, "allowed_models"?}` answers 201
  `{"data": Key, "key": "sk-unor-…"}`. The secret is in this answer only. `limits` takes `rpm`, `rps`,
  `daily_tokens`, `daily_usd` and `monthly_usd` (any left out is no limit); `allowed_models: null` is every model.
- `PATCH /keys/{id}` changes `name`, `disabled`, `public_allowed`, `limits` (`null` clears one) or `allowed_models`.
- `DELETE /keys/{id}` revokes a key for good.
- `GET /usage?days=30&key_id=…&model=…` gives usage by day, key and model.
- `GET /accounts` lists the user's subscription accounts with their usage windows.
