OpenRouter
OpenRouter provides a unified interface for accessing various LLM APIs, including models from OpenAI, Meta, Perplexity, and others. It follows the OpenAI API format - see our OpenAI provider documentation for base API details.
Setup
- Get your API key from OpenRouter
- Set the
OPENROUTER_API_KEYenvironment variable or specifyapiKeyin your config
Popular current models
OpenRouter's catalog changes quickly. These are current popular and recent model IDs that work well as starting points. Context lengths are from the OpenRouter catalog at time of writing — check OpenRouter Models (or GET /api/v1/models) for live values.
| Model ID | Context (tokens) | Good for |
|---|---|---|
| openai/gpt-6-sol | 1,050,000 | Complex reasoning and coding |
| anthropic/claude-opus-5.5 | 1,000,000 | Long-running agentic workflows |
| openai/gpt-6-luna | 1,050,000 | Fast, lower-cost OpenAI evals |
| anthropic/claude-haiku-4.5 | 200,000 | Lower-latency Claude runs |
| google/gemini-2.5-pro | 1,048,576 | Reasoning-heavy tasks |
| google/gemini-2.5-flash | 1,048,576 | Fast multimodal and general chat |
| meta-llama/llama-4-maverick | 1,048,576 | Popular open-weight frontier model |
| deepseek/deepseek-v3.2 | 163,840 | Cost-efficient reasoning and tools |
| mistralai/mistral-small-3.2-24b-instruct | 128,000 | Compact Mistral general use |
| qwen/qwen3-32b | 40,960 | Strong open multilingual model |
For the full catalog of 300+ models and current pricing, visit OpenRouter Models.
Basic Configuration
The openrouter:<model> provider uses Chat Completions.
For Responses, use openai:responses:openai/gpt-6-sol with apiBaseUrl: https://openrouter.ai/api/v1 and apiKeyEnvar: OPENROUTER_API_KEY in its config. OpenRouter Responses is stateless: replay the conversation instead of sending previous_response_id.
# yaml-language-server: $schema=https://promptfoo.dev/config-schema.json
providers:
- id: openrouter:openai/gpt-6-sol
config:
reasoning_effort: none
temperature: 0.7
max_completion_tokens: 1000
- id: openrouter:anthropic/claude-opus-5
config:
omitDefaults: true
max_tokens: 2000
- id: openrouter:google/gemini-2.5-flash
config:
temperature: 0.7
max_tokens: 4000
If you route OpenRouter traffic through a proxy or OpenRouter-compatible gateway, set apiBaseUrl in the provider config. Precedence is config.apiBaseUrl → the hardcoded OpenRouter default (https://openrouter.ai/api/v1); the generic OpenAI OPENAI_API_BASE_URL / OPENAI_BASE_URL env fallbacks are not consulted for this provider.
The same pattern applies to apiKeyEnvar — set it to read your API key from a custom environment variable name (default OPENROUTER_API_KEY).
providers:
- id: openrouter:openai/gpt-6-luna
config:
apiBaseUrl: https://proxy.example.com/openrouter/api/v1
apiKeyEnvar: MY_PROXY_KEY # optional: read the Bearer token from $MY_PROXY_KEY
Cost reporting
By default, promptfoo uses the reported usage.cost as the charge to your OpenRouter account. This also applies to the generic OpenAI Chat and Responses providers when their apiBaseUrl points to https://openrouter.ai/api/v1. Missing or invalid charges remain unknown. Reported upstream amounts remain separate metadata fields.
For responses explicitly marked usage.is_byok: true, generic cost is unavailable unless you configure complete, valid token rates. With Bring Your Own Key (BYOK), your upstream provider bills inference separately, so a zero or fee-only OpenRouter charge does not establish the combined cost. Responses with a missing or invalid BYOK flag continue to use the reported account charge; their route remains unknown in metadata.
To supply your own per-token estimate, configure cost, or both inputCost and outputCost. Valid configured rates take priority over reported billing, including for BYOK. Both prompt and completion token counts are required. Incomplete or invalid rates or counts leave cost unknown. This estimate is based on your configured rates, not a reconciled invoice.
Successful responses expose validated billing facts under metadata.openrouter:
| Field | Meaning |
|---|---|
accountCharge | Reported charge to your OpenRouter account, including zero. |
isByok | Reported boolean route flag; omitted when missing or invalid. |
reportedUpstreamInferenceCost | OpenRouter-reported upstream inference amount, when available. |
reportedUpstreamPromptCost, reportedUpstreamCompletionCost | Independently reported prompt and completion components. |
reportedServerToolCost | Reported server tool component, when available. |
Missing or invalid amounts are omitted. These fields appear in response details and JSON exports, independently of any configured estimate. Cost assertions require a known cost; their existing error handling can omit response metadata from the exported error row. General cost totals sum only known costs and can therefore be incomplete when some responses have unavailable cost.
Cache replays retain logical cost and billing metadata. The evaluator records zero additional incurred cost for cached responses with a known cost.
Provider errors
OpenRouter can report a provider error alongside partial output. If OpenRouter explicitly marks the error as a refusal or a content-policy block, promptfoo preserves available output from Chat or Responses and records the refusal for the is-refusal assertion. Provider access errors and other generation errors remain evaluation errors; Chat also retains any partial output in the raw response.
Features
- Access to 300+ models through a single API
- Mix free and paid models in your evaluations
- Support for text and multimodal (vision) models
- Compatible with OpenAI API format
- Pay-as-you-go pricing
Thinking/Reasoning Models
For GPT-6 Sol and Luna, set reasoning_effort or passthrough.reasoning. Choose either a named effort or a token budget (passthrough.reasoning.max_tokens).
For multi-turn Chat requests, OpenRouter supports changing effort through a configuration_update on an empty system or developer message. This is an OpenRouter extension; native OpenAI Chat Completions does not support it.
Some models like Gemini 2.5 Pro include thinking tokens in their responses. You can control whether these are shown using the showThinking parameter:
providers:
- id: openrouter:google/gemini-2.5-pro
config:
showThinking: false # Hide thinking content from output (default: true)
When showThinking is true (default), the output includes thinking content:
Thinking: <reasoning process>
<actual response>