Skip to content

Kimi K3

Moonshot AI's open-weight model for complex coding and reasoning across long-running tasks, served on Boundless.

Pricing

Kimi K3 API pricing by provider

ProviderInputCachedOutputContext
Boundless$2.30$0.23$11.401M
Moonshot AIModel creator$3.00$0.30$15.001M
DeepInfra$2.85$0.29$14.251M
Together$3.00$0.30$15.001M
Fireworks$3.00$0.30$15.001M
Baseten$3.00$0.30$15.001M

Prices per 1M tokens. Source: OpenRouter

About the model

Kimi K3 at a glance

Kimi K3 is Moonshot AI's 2.8-trillion-parameter mixture-of-experts model. It suits complex reasoning, long coding sessions and agentic knowledge work, and it accepts images.

From our model docs
Developer
Moonshot AI
Context window
1M tokens
Input
Text, image
Capabilities
Reasoning, tool use, vision

Tailored inference

Want Kimi K3 tuned to your workload?

Our engineers benchmark it against your current model, tune how it's served for your traffic, and run it to your cost, latency and quality targets.

Talk to an engineer

Run it

Switch with a base URL and a key

  1. 01

    Get an API key

    Sign up in the console and create a key. No sales call.

    Get a key
  2. 02

    Point your SDK at Boundless

    Change the base URL and key. For Anthropic SDKs, drop the /v1 and set max_tokens.

curl https://api.inference.boundless.network/v1/chat/completions \
  -H "Authorization: Bearer $BOUNDLESS_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "kimi-k3",
    "messages": [{
      "content": "Say hello",
      "role": "user"
    }]
  }'
From the quickstart docs

FAQ

Questions before you commit

How much does Kimi K3 cost on Boundless?

$2.30 per 1M input tokens, $0.23 per 1M cached input tokens and $11.40 per 1M output tokens, billed from prepaid credits. Every response includes exact token counts, so any charge can be checked against these prices.

What is Kimi K3's context window on Boundless?

1M tokens, shared by the prompt and the reply. max_tokens can go up to that limit.

Does Kimi K3 support tool calling?

Yes. Kimi K3 supports tool use and reasoning.

Can I control how much Kimi K3 reasons?

Yes. On chat completions, set reasoning_effort to none, low, medium or high, or leave it unset for the model's default. Size max_tokens for both the thinking and the reply.

Can Kimi K3 read images?

Yes. Kimi K3 accepts text, image as input.

What are the rate limits?

Each team has caps on tokens per minute, requests per minute and concurrent requests, shown on the usage page. Cached input doesn't count against them, and support can raise them for your team.

Does Boundless store my prompts or outputs?

No. Boundless never stores prompts or outputs and never trains on them. It keeps only what billing needs: token counts, model, timestamp and cost.

Can Boundless help me decide if Kimi K3 fits my workload?

Yes. We test it against your current model on your own prompts, then tune and run it to your targets.

More models

Other models to compare

View all models