Nscale
Use Nscale's Serverless Inference API for OpenAI-compatible chat, completion, embedding, and image requests.
Setup
Set your Nscale service token as an environment variable:
export NSCALE_SERVICE_TOKEN=your_service_token_here
Alternatively, you can add it to your .env file:
NSCALE_SERVICE_TOKEN=your_service_token_here
You can also supply config.apiKey or select a credential variable with config.apiKeyEnvar. Nscale does not fall back to OPENAI_API_KEY by default; set apiKeyEnvar: OPENAI_API_KEY to use that variable explicitly. The same credential rules apply when you set a custom apiBaseUrl.
Obtaining Credentials
You can obtain service tokens by:
- Signing up at Nscale
- Navigating to your account settings
- Going to "Service Tokens" section
Configuration
To use Nscale models in your promptfoo configuration, use the nscale: prefix followed by the model name:
providers:
- nscale:openai/gpt-oss-120b
- nscale:meta-llama/Llama-3.3-70B-Instruct
- nscale:Qwen/Qwen3-235B-A22B-Instruct-2507
Model IDs are the upstream Hugging Face repository IDs and are case-sensitive.
Model Types
Nscale supports different types of models through specific endpoint formats:
Chat Completion Models (Default)
For chat completion models, you can use either format:
providers:
- nscale:chat:openai/gpt-oss-120b
- nscale:openai/gpt-oss-120b # Defaults to chat
Completion Models
For text completion models:
providers:
- nscale:completion:openai/gpt-oss-20b
Embedding Models
For embedding models:
providers:
- nscale:embedding:Qwen/Qwen3-Embedding-8B
- nscale:embeddings:Qwen/Qwen3-Embedding-8B # Alternative format
Text-to-Image Models
For image generation models:
providers:
- nscale:image:black-forest-labs/FLUX.1-schnell
Popular Models
Model IDs are the upstream Hugging Face repository IDs and are case-sensitive
(meta-llama/Llama-3.3-70B-Instruct, not meta/llama-3.3-70b-instruct). The
authoritative list for your account is GET https://inference.api.nscale.com/v1/models,
which also returns pricing and context length:
curl https://inference.api.nscale.com/v1/models \
-H "Authorization: Bearer $NSCALE_SERVICE_TOKEN"
Text Generation Models
| Model | Provider Format | Use Case |
|---|---|---|
| GPT OSS 120B | nscale:openai/gpt-oss-120b | General-purpose reasoning and tasks |
| GPT OSS 20B | nscale:openai/gpt-oss-20b | Lightweight general-purpose model |
| Kimi K2.5 | nscale:moonshotai/Kimi-K2.5 | Large-scale agentic reasoning |
| Qwen 3 235B A22B | nscale:Qwen/Qwen3-235B-A22B | Large-scale language understanding |
| Qwen 3 235B A22B Instruct 2507 | nscale:Qwen/Qwen3-235B-A22B-Instruct-2507 | Latest Qwen 3 235B variant |
| Qwen 3 4B Instruct 2507 | nscale:Qwen/Qwen3-4B-Instruct-2507 | Lightweight instruction following |
| Qwen 3 4B Thinking 2507 | nscale:Qwen/Qwen3-4B-Thinking-2507 | Reasoning and thinking tasks |
| Qwen 3 8B | nscale:Qwen/Qwen3-8B | Mid-size general-purpose model |
| Qwen 3 14B | nscale:Qwen/Qwen3-14B | Enhanced reasoning capabilities |
| Qwen 3 32B | nscale:Qwen/Qwen3-32B | Large-scale reasoning and analysis |
| Qwen 2.5 Coder 3B Instruct | nscale:Qwen/Qwen2.5-Coder-3B-Instruct | Lightweight code generation |
| Qwen 2.5 Coder 7B Instruct | nscale:Qwen/Qwen2.5-Coder-7B-Instruct | Code generation and programming |
| Qwen 2.5 Coder 32B Instruct | nscale:Qwen/Qwen2.5-Coder-32B-Instruct | Advanced code generation |
| Qwen QwQ 32B | nscale:Qwen/QwQ-32B | Specialized reasoning model |
| Llama 3.3 70B Instruct | nscale:meta-llama/Llama-3.3-70B-Instruct | High-quality instruction following |
| Llama 3.1 8B Instruct | nscale:meta-llama/Llama-3.1-8B-Instruct | Efficient instruction following |
| Llama 3.2 11B Vision Instruct | nscale:meta-llama/Llama-3.2-11B-Vision-Instruct | Vision-language tasks |
| Llama 4 Scout 17B | nscale:meta-llama/Llama-4-Scout-17B-16E-Instruct | Image-Text-to-Text capabilities |
| DeepSeek R1 Distill Llama 70B | nscale:deepseek-ai/DeepSeek-R1-Distill-Llama-70B | Efficient reasoning model |
| DeepSeek R1 Distill Llama 8B | nscale:deepseek-ai/DeepSeek-R1-Distill-Llama-8B | Lightweight reasoning model |
| DeepSeek R1 Distill Qwen 1.5B | nscale:deepseek-ai/DeepSeek-R1-Distill-Qwen-1.5B | Ultra-lightweight reasoning |
| DeepSeek R1 Distill Qwen 7B | nscale:deepseek-ai/DeepSeek-R1-Distill-Qwen-7B | Compact reasoning model |
| DeepSeek R1 Distill Qwen 14B | nscale:deepseek-ai/DeepSeek-R1-Distill-Qwen-14B | Mid-size reasoning model |
| DeepSeek R1 Distill Qwen 32B | nscale:deepseek-ai/DeepSeek-R1-Distill-Qwen-32B | Large reasoning model |
| Devstral Small 2505 | nscale:mistralai/Devstral-Small-2505 | Code generation and development |
| Mixtral 8x22B Instruct | nscale:mistralai/Mixtral-8x22B-Instruct-v0.1 | Large mixture-of-experts model |
Embedding Models
| Model | Provider Format | Use Case |
|---|---|---|
| Qwen 3 Embedding 8B | nscale:embedding:Qwen/Qwen3-Embedding-8B | Text embeddings and similarity |
Text-to-Image Models
| Model | Provider Format | Use Case |
|---|---|---|
| Flux.1 Schnell | nscale:image:black-forest-labs/FLUX.1-schnell | Fast image generation |
| Stable Diffusion XL | nscale:image:stabilityai/stable-diffusion-xl-base-1.0 | High-quality image generation |
| SDXL Lightning | nscale:image:ByteDance/SDXL-Lightning | Ultra-fast image generation |
Configuration Options
Nscale supports standard OpenAI-compatible parameters:
providers:
- id: nscale:openai/gpt-oss-120b
config:
temperature: 0.7
max_tokens: 1024
top_p: 0.9
frequency_penalty: 0.1
presence_penalty: 0.2
stop: ['END', 'STOP']
seed: 42
Supported Parameters
temperature: Controls randomness (0.0 to 2.0). Defaults to0unless set.max_tokens: Maximum number of tokens to generate. Defaults to1024unless set.top_p: Nucleus sampling parameterfrequency_penalty: Reduces repetition based on frequencypresence_penalty: Reduces repetition based on presencestop: Stop sequences to halt generationseed: Deterministic sampling seed
Any other parameter is forwarded to the Nscale API unchanged.
Streaming is not supported. Promptfoo reads each response as a single JSON body, so
setting stream: true produces a response it cannot parse.
Example Configuration
Here's a complete example configuration:
providers:
- id: nscale:openai/gpt-oss-120b
config:
temperature: 0.7
max_tokens: 512
- id: nscale:meta-llama/Llama-3.3-70B-Instruct
config:
temperature: 0.5
max_tokens: 1024
prompts:
- 'Explain {{concept}} in simple terms'
- 'What are the key benefits of {{concept}}?'
tests:
- vars:
concept: quantum computing
assert:
- type: contains
value: 'quantum'
- type: llm-rubric
value: 'Explanation should be clear and accurate'
Pricing
Nscale prices text generation and embeddings per token, and images per megapixel. Check Nscale's model endpoint API for current rates available to your organization.
Promptfoo has no Nscale-specific token price table. For chat, completion, or embedding estimates, set cost or inputCost/outputCost in USD per token. Image estimates use fixed model rates multiplied by the number of images; they do not adjust for resolution or use those token-cost settings.
Key Features
Nscale hosts the models and exposes an OpenAI-compatible API. See Nscale's documentation for throughput, rate limits, and available regions.
Error Handling
The Nscale provider includes built-in error handling for common issues:
- Network timeouts and retries
- Rate limiting
- Invalid API key errors
- Model availability issues
Support
For support with the Nscale provider: