The 2.8-Trillion-Parameter Guest: Kimi K3 Becomes the First Chinese Open-Weight Frontier Model on Amazon Bedrock
AWS quietly made Moonshot AI's 2.8T-parameter Kimi K3 generally available on Bedrock with explicit prompt caching, a 1M-token context, and hard data-boundary guarantees — a month after reports said no major cloud would host it.
A month ago, the conventional wisdom was that Kimi K3 — Moonshot AI’s 2.8-trillion-parameter open-weight model, the largest ever released — would not appear on any major Western cloud. Posts circulating in August 2026 noted that Amazon’s Bedrock, Microsoft’s Azure Foundry, and Google’s Vertex AI had not integrated Kimi K3 “or any comparable Chinese open-weight model.” On September 18, 2026, that era quietly ended: AWS announced on its Machine Learning Blog that Kimi K3 is generally available on Amazon Bedrock, complete with cross-region inference profiles, OpenAI-compatible APIs, and an enterprise-grade data boundary.
For anyone tracking the geopolitics of AI infrastructure, this is a bigger deal than the average model-on-a-marketplace story. The largest open-weight model ever built, trained in Beijing by a company heading toward a Hong Kong IPO, is now a first-party citizen of Amazon’s managed AI platform — served with guarantees that customer prompts never leave AWS’s control.
What landed, exactly
Kimi K3 is Moonshot AI’s most capable model and, by the company’s account, the first open model to reach 2.8 trillion parameters. The weights were originally released on July 27, 2026, after a preview period that generated enormous attention — and enormous downloading. The architecture is a mixture-of-experts design Moonshot calls Stable LatentMoE, activating 16 of 896 experts per token. Two newer structural ideas, Kimi Delta Attention and Attention Residuals, are designed to improve how information flows across sequence length and model depth, and Moonshot credits them — together with quantization-aware training from the supervised fine-tuning stage onward (MXFP4 weights with MXFP8 activations) — for an approximate 2.5x improvement in scaling efficiency over Kimi K2.
On Bedrock, the practical specs are what matter: native vision, a 1-million-token context window, and positioning aimed squarely at long-running coding and knowledge workflows that require sustained context across large repositories, documents, and images.
AWS also confirmed what Moonshot itself has said about relative standing: K3’s overall performance still trails proprietary frontier models — Moonshot names Claude Fable 5 and GPT 5.6 Sol — while delivering frontier-level results across its own evaluation suite. And the technical report’s caveats traveled with the model: K3 was trained in a preserved thinking history mode, so generation quality can degrade if an agent harness fails to pass historical thinking content back in, and its bias toward long-horizon autonomy means it can make unexpected decisions on a user’s behalf when instructions are ambiguous. System-prompt authors, take note.
The platform story: caching, profiles, and APIs
The launch is as much about Bedrock’s maturing open-weight strategy as about the model. Three platform details stand out.
Explicit prompt caching. Kimi K3 is the first open-weight model on Amazon Bedrock to support explicit prompt caching. Developers mark the end of a reusable prompt prefix (at least 1,024 tokens) with a prompt_cache_breakpoint. Tokens written to cache are billed at a higher rate but persist for at least 30 minutes; subsequent requests that hit the cache are billed at a discounted input rate and — crucially — do not count against input-tokens-per-minute quotas. For agentic coding workflows that resend stable context like repository instructions and tool definitions on every call, this is the difference between a demo and a budget line item.
Cross-region inference profiles. Workloads without regional restrictions should use global.moonshotai.kimi-k3, which routes requests to any supported commercial AWS Region worldwide and costs approximately 10% less than a geographic profile. For data-residency requirements, us.moonshotai.kimi-k3 keeps processing inside US geography. That “us.” profile is the tell: AWS expects regulated American customers to run this Beijing-trained model, and built the plumbing so they can do it without a waiver.
API surface. The bedrock-runtime endpoint supports the OpenAI-compatible Responses and Chat Completions APIs alongside the native Bedrock Invoke and Converse APIs. Since these are platform capabilities rather than per-model integrations, K3 arrived with tool calling, structured output, reasoning, and response streaming already working — the benefits of Bedrock’s 2026 investment in treating open-weight models as first-class citizens.
The data boundary is the real headline
The most consequential paragraph in AWS’s announcement is not about benchmarks. As with all open-weight models on Bedrock, customer data is processed within the AWS data boundary, is not shared with the model provider, and is not used to train the underlying model. Zero data retention is always enabled for inference requests, and “zero operator access” prevents even AWS operators from reading prompts and completions during inference.
That language exists because the context is charged. On September 8, 2026, CISA published an advisory (AA26-251A) on China-based AI companies conducting what it called industrial-scale distillation campaigns against US AI companies. Ten days later, AWS is hosting one of China’s flagship open models behind a contractual and architectural wall: the weights are open, but the inference traffic, the prompts, and the completions stay inside AWS. It is a neat legal and technical answer to an awkward question — enterprises get frontier-adjacent open-weight capability at $3.00/$15.00 per million tokens (Moonshot’s list pricing, roughly 40% below Claude Opus 5 by third-party comparison) without a single byte transiting to the model’s country of origin.
Ecosystem ready on day one
AWS lined up the tooling story properly. OpenCode, the open-source model-agnostic coding agent, has a native amazon-bedrock provider using the Converse API — a one-line config change puts global.moonshotai.kimi-k3 behind /models. Hermes Agent, the open-source productivity assistant, natively supports Bedrock-hosted models through its hermes model provider picker. IAM prerequisites are the standard trio (bedrock:InvokeModel, bedrock:InvokeModelWithResponseStream, bedrock:CreateInference), and AWS published a Moonshot AI on AWS samples repository on GitHub alongside the launch.
Since 2025, Bedrock has added dozens of open-weight models from DeepSeek, Google, MiniMax, Mistral AI, Moonshot AI, NVIDIA, OpenAI, and Qwen. Kimi K3 is the biggest and the most geopolitically loaded of them. Watch whether Azure Foundry and Vertex AI follow: the technical blocker never existed — only the policy one. AWS just called the question.
Sources
- [1] https://aws.amazon.com/blogs/machine-learning/introducing-kimi-k3-on-amazon-bedrock/
- [2] https://www.unite.ai/moonshot-ais-kimi-k3-arrives-on-amazon-bedrock-with-1m-token-context/
- [3] https://docs.aws.amazon.com/bedrock/latest/userguide/model-card-moonshot-ai-kimi-k3.html
- [4] https://www.cisa.gov/news-events/cybersecurity-advisories/aa26-251a