About ainetcafe

A dedicated Kimi K3 provider on its own cluster.

What we do

ainetcafe serves Kimi K3, Moonshot AI's open-weight 2.8-trillion-parameter model, from our own inference cluster. We run the official released weights at their native MXFP4 precision and expose them through OpenAI-compatible and Anthropic-compatible endpoints, so coding agents and application backends connect with one base URL change.

The reason to run K3 here rather than elsewhere is cost structure. We own the hardware and the serving stack, so we price K3 at $2.10 per million input tokens and $10.50 per million output tokens, with committed and core plans below that. Most K3 providers sit at the $3 / $15 list price.

Who is behind it

The team behind ainetcafe builds heterogeneous inference infrastructure: GPU plus Arm CPU clusters where expert weights and the KV cache live in terabyte-scale system memory and the GPUs do only the compute-dense work. That architecture is what makes running a model of K3's size affordable. ainetcafe is the international, self-serve face of that cluster.

How we operate

Where data is processed

Self-serve and trial traffic is served from our own clusters in China. For committed and core tiers, the serving region is agreed at contract time and can be located to meet your requirements.

Contract details

Self-serve accounts are governed by the terms of service. Committed and core plans are signed as a written agreement and invoiced in US dollars; a data processing addendum is available.