Endpoint
https://routerbase.com/v1 with your RouterBase API key.
Request Headers
Request Body
string
required
Model ID. See the Model Overview for the current featured chat model set, or use the live Models API for the full list.
array
required
OpenAI message array:
[{role: "user", content: "..."}]. Supports multimodal content parts for vision-capable models.number
Sampling temperature, 0–2. Default model-dependent.
integer
Maximum tokens to generate.
number
Nucleus sampling probability, 0–1.
boolean
If
true, response is a Server-Sent Events stream of data: {chunk}\n\n chunks terminated by data: [DONE].string | array
Stop sequences.
integer
Sampling seed for reproducible output.
object
e.g.
{ "type": "json_object" } for JSON-mode models.number
Penalize new tokens by whether they appear in the text so far, -2 to 2.
number
Penalize new tokens by their frequency in the text so far, -2 to 2.
array
Tool/function definitions for tool-calling models.
boolean
default:"true"
RouterBase extension for reasoning models (Kimi K3 / K2.6, OpenAI
o-series, GPT-5, DeepSeek-R1, …). These models return their reasoning
trace separately in
choices[].message.reasoning_content (see
Response). Set false to strip that field from the
response. The model always generates the reasoning upstream — this
only controls whether it is returned, so it does not reduce token usage
or cost. Non-reasoning models ignore this field.string
Optional OpenAI-style cache partition key. Requests sharing a key route
to the same prompt cache. RouterBase sets a per-end-user key
automatically (see Prompt caching); pass your own
only to override — e.g. to share a cache across users in a workspace.
Anthropic and Gemini models ignore this field and use their own cache
primitives (RouterBase handles those automatically too).
Examples
Response
usage includes a prompt_tokens_details object:
Reasoning models
For reasoning models (Kimi K3 / K2.6, OpenAI o-series, GPT-5, DeepSeek-R1, …) the assistant message carries the model’s reasoning trace in a separatereasoning_content field, alongside the final answer in
content:
content is always the final answer; reasoning_content is the thinking
that led to it. Pass include_reasoning: false to omit
reasoning_content from the response. The field is absent for models that
produce no reasoning.
cached_tokens— OpenAI-standard count of tokens served from cache (cache read), billed at the model’s discounted Cache Read rate.cache_read_input_tokens— Anthropic-style alias for the same number, surfaced for parity.cache_creation_input_tokens— tokens written to the cache this request (Anthropic only; OpenAI/Gemini have no write premium).
usage.
Prompt caching
For models with upstream prompt caching, RouterBase caches automatically — no flag required. Cached input tokens bill at the model’s discounted Cache Read rate; for Anthropic models, cache writes (the first request that creates a cache entry) bill at a ~1.25× Cache Write premium. Both rates, when defined, appear on each model’s pricing page.Per-user isolation
Caching is isolated per end-user: your customers’ cached prefixes are never shared with, or readable by, another customer — even when they send identical prompts through the same RouterBase account. RouterBase derives a stable, opaque per-user token and wires it into whichever cache primitive the upstream honours:- OpenAI / GPT-5 — standard
prompt_cache_keyfield. Pass your ownprompt_cache_keyonly if you want to override RouterBase’s per-user partition (e.g. to share a cache across users in a workspace). - Anthropic / Claude —
session-idheader partitioning, plus acache_control: ephemeralbreakpoint on the system block (Claude only caches with an explicit breakpoint and ignoresprompt_cache_key). - Google / Gemini — Gemini’s implicit cache has no per-request key, so RouterBase prepends a per-user salt comment to the cached prefix. Each user gets a distinct content hash → a distinct cache partition.
Thresholds & TTL
Upstream caches typically require ≥1024 input tokens to take effect. Per-model exceptions:- Anthropic Haiku 4.5, Opus 4.5 / 4.6 / 4.7 → ≥4096 tokens
- Anthropic Sonnet 4.6 → ≥2048 tokens
- Gemini 2.5 Pro → ≥4096 tokens
Streaming
Set"stream": true to receive SSE chunks. Each chunk is a chat.completion.chunk object with a delta containing the incremental content. The stream ends with data: [DONE].