Models and pricing
Two families on the same deployments. Prices are in US dollars per million tokens.
Open models
Open-weight models served under their original identifiers.
Orion
Serenity’s own model family. At launch, orion-pro and orion-plus apply Serenity’s serving recipe to the open models in the catalogue. Serenity’s own fine-tuned models will follow.
| Model | Family | Architecture | Context | Max output | Input | Output | Cached input |
|---|---|---|---|---|---|---|---|
Qwen3.8-27Bqwen3.8-27b | Open | DenseNVFP4 | 131,072 tokens | 32,768 tokens | 0.144 | 1.50 | 0.036 |
Qwen3.6-35B-A3Bqwen3.6-35b-a3b | Open | MoENVFP4 | 131,072 tokens | 32,768 tokens | 0.101 | 0.60 | 0.025 |
Orion Proorion-proSame base as qwen3.8-27b | Orion | DenseNVFP4 | 131,072 tokens | 32,768 tokens | 0.144 | 1.50 | 0.036 |
Orion Plusorion-plusSame base as qwen3.6-35b-a3b | Orion | MoENVFP4 | 131,072 tokens | 32,768 tokens | 0.101 | 0.60 | 0.025 |
Capabilities
- Streaming responses
- Tool calling
- Structured outputs
- Prefix caching
- NVFP4 precision across all models
Notes
- Cached input applies to prefix tokens reused across requests.
- Prices may change. Changes are announced in the documentation.
- The catalogue will grow. The current list is in the documentation.