Kimi K3 · dedicated provider

Kimi K3 at $2.10 in, $10.50 out.

Moonshot's official open weights, served from our own cluster at native MXFP4 precision. OpenAI and Anthropic compatible, 1M context, prompt caching at $0.30 per million. Sign up, get $2 to try, first token in ten minutes.

Input
$2.10/M
Output
$10.50/M
Cached input
$0.30/M
Context
1Mtokens

Where Kimi K3 is served today

Per million tokens, USD. Provider prices as listed on OpenRouter, September 2026. Moonshot's list price is $3.00 / $15.00. Daily-updated full table.

ProviderInputOutputCached input
Moonshot AI, Fireworks, Together, Baseten, Parasail3.0015.000.30
DeepInfra2.8514.250.285
Sail Research2.64813.280.303
DigitalOcean, Makora2.5512.75 to 12.950.26 to 0.285
Cheapest listed on OpenRouter2.1010.950.23
ainetcafe, pay as you go2.1010.500.30
ainetcafe, committed plan1.809.000.30
ainetcafe, core plan1.507.500.30

Why run K3 here

Own cluster

No upstream in the chain

We run the weights on our own machines. There is no reseller quota to be cut off from and no third-party account that can be suspended.

Official weights

Native precision, not requantized

Kimi K3 as released by Moonshot, at its quantization-aware MXFP4 precision. No second quantization, no distilled stand-in. A/B against the official API is welcome.

Two protocols

OpenAI and Anthropic, same key

Chat Completions at /v1, Messages at /v1/messages. Claude Code, Cline, Kilo, OpenCode and Codex CLI connect with one base URL change.

Built for agents

Long loops, cheap cache

1M context, tool calling, JSON schema output, vision, thinking on or off, streaming with usage. Cache hits bill at $0.30 per million, which is where agent loops spend most of their tokens.

Three ways to buy

Per million tokens, USD. Card payments through Stripe; committed plans are invoiced monthly.

Pay as you go

$2.10 / $10.50 in / out

  • Cached input $0.30
  • $2 free credit on sign-up, then top up from $20 (bonus credit on $20, $100 and $500)
  • No hard rate limit at launch; fair-use monitoring, dedicated pools on committed plans
  • Email support
Get API key
Team Starter

Pay $500, get $1,000

  • Credits valid 30 days, one per company domain
  • Committed-plan treatment during the trial: dedicated endpoint, 1M context, priority pool
  • Sign a committed plan within 30 days and unused credits roll into your first invoice
  • Pay by card after applying; provisioned within one business day
Apply for Team Starter
Committed and Core

$1.80 / $9.00 and $1.50 / $7.50

  • Committed from $15,000 per month; Core from $60,000 per month
  • Dedicated endpoint and static egress IP, independent rate pool
  • Automatic failover to the official Moonshot API at your contract price
  • Version pinning, usage API, monthly invoicing in USD
Talk to us
Coding plan · K3 Lite

$9 / month

  • $27 of Kimi K3 per month, $6.75 per week
  • For Claude Code, Cline, Kilo, OpenCode
  • Overage falls back to pay-as-you-go
Subscribe
Coding plan · K3 Standard

$19 / month

  • $57 of Kimi K3 per month, $14.25 per week
  • Same tools, same rules, twice the room
  • One plan per account, no resale
Subscribe
Who the plans are for

Individuals running coding agents. Quotas are written in dollars so you know exactly what you get; when a plan runs out, requests keep working at pay-as-you-go rates instead of stopping. Buy under Wallet, then Subscription Plans. Teams should start with Team Starter.

Discounts are defined as a percentage of Moonshot's list price and follow it if it changes. Credits are prepaid and non-refundable. Pay-as-you-go stops at zero balance; there is no negative balance.

Quickstart

Base URL https://microquickjs.com/v1 · Anthropic base URL https://microquickjs.com · model Kimi-K3 · full setup for eight tools in the integration guides

curl https://microquickjs.com/v1/chat/completions \
  -H "Authorization: Bearer $AINETCAFE_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{"model":"Kimi-K3","stream":true,
       "messages":[{"role":"user","content":"Summarize this repo in three lines."}]}'
from openai import OpenAI
client = OpenAI(base_url="https://microquickjs.com/v1", api_key=os.environ["AINETCAFE_API_KEY"])
stream = client.chat.completions.create(
    model="Kimi-K3",
    messages=[{"role": "user", "content": "Summarize this repo in three lines."}],
    stream=True,
)
for chunk in stream:
    print(chunk.choices[0].delta.content or "", end="")
curl https://microquickjs.com/v1/messages \
  -H "x-api-key: $AINETCAFE_API_KEY" \
  -H "anthropic-version: 2023-06-01" \
  -H "Content-Type: application/json" \
  -d '{"model":"Kimi-K3","max_tokens":1024,
       "messages":[{"role":"user","content":"Summarize this repo in three lines."}]}'
# Claude Code: point it at ainetcafe, keep everything else
export ANTHROPIC_BASE_URL=https://microquickjs.com
export ANTHROPIC_API_KEY=$AINETCAFE_API_KEY
export ANTHROPIC_MODEL=Kimi-K3
claude
# Cline / Kilo Code / OpenCode / Codex CLI: choose "OpenAI compatible"
Base URL:  https://microquickjs.com/v1
API key:   your ainetcafe key
Model ID:  Kimi-K3
Context:   1000000   (set 262144 if the client caps lower)

Checked with Moonshot's own verifier

We ran the full Kimi Vendor Verifier against our production endpoint, 15–16 September 2026, 50 concurrent, no reruns. Every number, including the ones that did not pass →

Benchmarks

MMMU Pro 0.812 · OCRBench 0.834

Moonshot's reference: 0.82 and 0.89. MMMU is within one standard error of the reference. OCRBench reaches 94%; the gap is handwritten math (49%) and noisy 50-image subsets, broken down by category.

API behaviour

Tool calls 402/408 · params 18/18

Streaming and non-streaming agree. The six failures share one cause, a strict $id check on tool schemas, reproduced and documented with the fix scheduled.

Load

8,190 requests, 2 errors

Seven hours at 50 concurrent, about 35 M tokens, no manual intervention, ≈1,350 tok/s of aggregate output. Latency from outside China is on the live status page.

Team Starter

For companies that want to run real traffic before committing. Pay $500, receive $1,000 in credits and committed-plan treatment for 30 days. One per company domain.

You can pay the $500 by card right after applying. Endpoint and credits are provisioned within one business day. By applying you accept the terms.

Questions

Is this the real Kimi K3?
Yes. Moonshot's released weights at their native MXFP4 precision, no second quantization and no substitute model. Run the A/B yourself, or read our Kimi Vendor Verifier results: MMMU Pro Vision within noise of Moonshot's reference, and the places we fall short listed too.
What is the difference from the official API?
Same weights, same precision. Price, rate limits and endpoints are set by us, and there is no upstream account in the chain that can be suspended.
What happens when my balance hits zero?
Requests stop with a clear error until you top up. There is no negative balance and no surprise invoice on pay-as-you-go.
Do you store my prompts?
Billing and audit metadata are kept for 30 days. Prompt and response content is not stored by default; committed plans can specify retention in the contract.
How are rate limits raised?
There is no hard per-key rate limit on pay-as-you-go today; we monitor for abuse and will contact you before applying one. Team Starter and committed plans run in their own pools with limits written into the order.
Can I get invoices?
Pay-as-you-go and coding plans get receipts from the payment provider at checkout. Committed and core plans are invoiced monthly in USD. Coding plans are not invoiced.