DeepSeek V4 Pro
DeepSeek's DeepSeek V4 Pro is a main coding agent paired with DeepSeek V4 Flash sub-agents, served on Boundless.
Pricing
DeepSeek V4 Pro API pricing by provider
| Provider | Input | Cached | Output | Context |
|---|---|---|---|---|
| Boundless | $1.04 | $0.04 | $2.08 | 1M |
| DeepSeekModel creator | $0.66 | $0.02 | $1.98 | 1M |
| DeepInfra | $1.30 | $0.10 | $2.60 | 1M |
| Together | $1.32 | $0.13 | $3.96 | 1M |
| Baseten | $1.32 | $0.13 | $3.96 | 1M |
Prices per 1M tokens. Sources: OpenRouter and Vercel AI Gateway
About the model
DeepSeek V4 Pro at a glance
DeepSeek V4 Pro is a 1.7T-parameter open-weight model for complex coding work. A common setup runs Pro as the main agent and V4 Flash for sub-agents, which keeps the whole workflow on open-weight models.
From our model docs- Developer
- DeepSeek
- Context window
- 1M tokens
- Input
- Text
- Capabilities
- Reasoning, tool use
Tailored inference
Want DeepSeek V4 Pro tuned to your workload?
Our engineers benchmark it against your current model, tune how it's served for your traffic, and run it to your cost, latency and quality targets.
Talk to an engineerRun it
Switch with a base URL and a key
- 01
- 02
Point your SDK at Boundless
Change the base URL and key. For Anthropic SDKs, drop the /v1 and set max_tokens.
curl https://api.inference.boundless.network/v1/chat/completions \
-H "Authorization: Bearer $BOUNDLESS_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "deepseek-v4-pro-0813",
"messages": [{
"content": "Say hello",
"role": "user"
}]
}'FAQ
Questions before you commit
How much does DeepSeek V4 Pro cost on Boundless?
$1.04 per 1M input tokens, $0.04 per 1M cached input tokens and $2.08 per 1M output tokens, billed from prepaid credits. Every response includes exact token counts, so any charge can be checked against these prices.
What is DeepSeek V4 Pro's context window on Boundless?
1M tokens, shared by the prompt and the reply. max_tokens can go up to that limit.
Does DeepSeek V4 Pro support tool calling?
Yes. DeepSeek V4 Pro supports tool use and reasoning.
Can I control how much DeepSeek V4 Pro reasons?
Yes. On chat completions, set reasoning_effort to none, low, medium or high, or leave it unset for the model's default. Size max_tokens for both the thinking and the reply.
What are the rate limits?
Each team has caps on tokens per minute, requests per minute and concurrent requests, shown on the usage page. Cached input doesn't count against them, and support can raise them for your team.
Does Boundless store my prompts or outputs?
No. Boundless never stores prompts or outputs and never trains on them. It keeps only what billing needs: token counts, model, timestamp and cost.
Can Boundless help me decide if DeepSeek V4 Pro fits my workload?
Yes. We test it against your current model on your own prompts, then tune and run it to your targets.
More models




