トップに戻る

Kimi K3 Now Available via Telnyx Inference API

https://telnyx.com/release-notes/kimi-k3-telnyx-inference
114 pointsbyfionaattelnyx14時間前66 コメント
Moonshot AI released open weights for Kimi K3 today and it's live on Telnyx Inference. The architecture and Moonshot's own benchmarks are in their technical blog. This post is about running it on Telnyx.

What we are adding: K3 is now available via the Telnyx Inference API, hosted on GPUs that we own and operate.

This matters due to the size of this model. A 2.8T model needs dedicated infra to serve well. Because we own and operate the GPUs, we can contorl throughput. with no inter-provider hops and no cloud tenant, latency is minimized. We have GPUs in each of the US, EU, APAC, and MENA. Inference runs in the region you pick, with zero data retention. We do not store prompts or completions after the response returns. Because we own the infra, the per-token price reflects the cost of running the model, not the cost of renting someone else's plus their margin.

We have not benchmarked K3 ourselves yet. Moonshot's own numbers and early third-party evaluations put it at frontier level for coding and agentic work, trailing only Claude Fable 5 and GPT 5.6 Sol on aggregate. Full breakdown in their blog.

Pricing on Telnyx: $2.70/1M input tokens, $13.50/1M output tokens, $0.27/1M cached input tokens. Prompt caching enabled by default. Served via an OpenAI-compatible endpoint so you can test easily.

Model ID: [MODEL_ID] API: https://api.telnyx.com/v2/ai/chat/completions Docs: https://developers.telnyx.com/docs/inference Technical blog (Moonshot): https://www.kimi.com/blog/kimi-k3

コメント (15)

apexalpha1時間前
Is this an ad?

Why is this provider specifically on the front page?

Also why is this not available over openrouter?

mesmertech2時間前
Why are you guys not on Openrouter? I assume you'd get way more volume that way no?

Or does openrouter have like a specific contract you have to sign with them and requirements or smth? https://openrouter.ai/moonshotai/kimi-k3#providers

Mossy97時間前
Also available from Nebius, via Cortecs: https://cortecs.ai/detailedServerlessView/kimi-k3

€2.693/M input €13.464/M output Surprisingly, cache is not mentioned

Upd: Tensorix joined the fray, with the same prices, with cache at €0.673/M read

hmokiguess2時間前
Is it also offered via a ZDR + BAA / HIPAA eligible? Would love to use it in prod
theredsix10時間前
10% cheaper than official! Let the inference pricing wars begin!
htrp1時間前
Token Type Price per 1M tokens Cached Input $0.27 Input $2.70 Output $13.50

Looks like they are first vendor to undercut in price.

morpheuskafka2時間前
I have used Telnyx for, on the opposite end of cool-ness, their fax API. Curious if their AI pricing and quality are actually competitive or if this is just a rapid pivot to try to ride the AI wave.
zius5時間前
In my first interaction ("hi there kimi k3!"), Kimi K3 identified twice out of three times as Claude:

> Hi there! Quick note — I'm actually Claude, made by Anthropic, not Kimi. But no worries!

https://imgur.com/a/jqpc2Jc

and

> Just a quick heads-up — I'm Claude, made by Anthropic, not Kimi! Kimi is a different AI assistant (made by Moonshot AI), so it looks like there might be a little mix-up.

https://imgur.com/a/AKxeysH

maelito5時間前
Where is it hosted ?
bedros6時間前
any of these providers are HIPAA compliant?
madhu_ghalame6時間前
Consider publishing latency, throughput, and cost metrics under different workloads to help teams make informed decisions.
smallerize14時間前
Very cool. What are your throughput and latency like?
tokai2時間前
That is not going to be cheap for long if K3 prices does the same as GLM5.2 prices. If nothing else open weight models are great to get providers to compeete on price.
rvz5時間前
Jevon's paradox depends on how cheap the tokens get as the price of tokens get driven to zero as intelligence gets better and cheaper.
jakswa12時間前
what quantization? FP4?