FlexAI
FlexAI serves models through an OpenAI-compatible API. Use promptfoo's OpenAI provider with FlexAI's endpoint and API key.
FlexAI documents its data handling in its privacy policy and terms of service.
Setup
Create a key on the FlexAI platform and set FLEXAI_API_KEY in your shell, or load it from a file with --env-file.
providers:
- id: openai:chat:DeepSeek-V4-Flash-0731
config:
apiBaseUrl: https://api.flex.ai/v1
apiKeyEnvar: FLEXAI_API_KEY
headers:
OpenAI-Organization: ''
omitDefaults: true
temperature: 0
max_tokens: 4096
showThinking: false
passthrough:
reasoning_effort: low
apiBaseUrl overrides OpenAI endpoint environment variables. apiKeyEnvar selects only FLEXAI_API_KEY; a missing key does not fall back to OPENAI_API_KEY. The empty OpenAI-Organization header overrides any inherited OPENAI_ORGANIZATION value.
Configuration
omitDefaults: trueomits promptfoo's default output limit and temperature. Explicit settings andOPENAI_MAX_TOKENS/OPENAI_TEMPERATUREstill apply; the example setstemperature: 0explicitly.- Use
max_tokensto cap output, including reasoning tokens. If you leave it unset and have noOPENAI_MAX_TOKENS, FlexAI chooses the limit. Increase it if reasoning consumes the budget before an answer appears. - Put
reasoning_effortunderpassthroughto send it to FlexAI regardless of model-name detection. FlexAI maps supported effort levels to each model's controls. showThinking: falseexcludes reasoning text from the answer used by assertions.
Check the model catalog or GET https://api.flex.ai/v1/models for current model IDs. See the compatibility reference for supported parameters and pricing for current rates.
Embeddings
Use FlexAI's bge-m3 model for similarity assertions:
defaultTest:
options:
provider:
embedding:
id: openai:embedding:bge-m3
config:
apiBaseUrl: https://api.flex.ai/v1
apiKeyEnvar: FLEXAI_API_KEY
headers:
OpenAI-Organization: ''
Example
The provider-flexai example compares two chat models and uses FlexAI embeddings for grading with the same API key.