トップに戻る

Cerebras CS-4

https://www.cerebras.ai/cs4
463 pointsbysunils346日前275 コメント

コメント (15)

KronisLV6日前
God I wish they'd back up all of those claims by offering a subscription of Kimi K3 and GLM 5.3, not some outdated GLM 4.7 instance that they then proceed to call a preview model and say that they'll remove it, leaving users only with GPT-OSS 120B which is nigh useless nowadays: https://support.cerebras.net/articles/9996007307-cerebras-co... and https://www.cerebras.ai/pricing

Guess they don't care about regular devs atm and are focused only on hardware sales.

syntaxing6日前
I think the fun takeaway from this is that GPT 5.4 is probably 45B active parameters and GPT 5.6 Sol is closer to 50B.
sreekanth8506日前
AMD along with cerebras may probably compete with NVIDIA monopoly in near future. Also, NVIDIA will have competition form multiple companies. Just my prediction.
bearjaws6日前
I was hoping to see Cerebras launch something other than GPT-OSS-120b in production this week, especially with GLM4.7 going away.

If they could launch Qwen 27b or Deepseek Flash that would be amazing.

reilly30006日前
> CS-4 delivers more than 1,000 tokens per second on models exceeding 10 trillion parameters

Oops did they just out GPT-5.6 sol’s parameter count?

9cb14c1ec06日前
Just a reminder for everyone that we are only several years and 3 or 4 iterations into hardware being optimized for LLMs. We should all expect orders of magnitude improvement in speed and/or cost over the next 5 years. Then we can have fun conversations about "unlimited" "intelligence" and about what the price wars and profit margins of consumer AI products are when your average ChatGPT user costs the company $0.10 per month.

> CS-4 delivers more than 1,000 tokens per second on models exceeding 10 trillion parameters

Wow!

lostmsu6日前
KV caching status?

What's the point of 1000tok/s if you have to do prefill on every agentic turn which at 100k depth would make it 1.5 min latency every turn?

ethanzhang10246日前
If cerebars is performing well, why didn't its predecessor, server S-3, become the largest API token provider on OpenRouter, surpassing the official model releases?
xmorse6日前
Cerebras is very fast but you can basically never use it because of its scarcity
anonymous_user96日前
Conspicuously missing: power consumption figures
rbanffy6日前
Impressive that this is an "interim" product, the start of a new line that ought to be continued with the WSE-4 family, where they are supposed to use a 3nm process and, maybe, 3D stacked SRAM. The modular architecture also points towards field upgrades that are badly needed for AI datacenter builders.
rajnathani6日前
Interestingly they’re still on the WSE-3 (5nm TSMC) wafer chip and slightly bumped up the specs there (overlocking mostly it seems), for why it’s called WSE-3 Turbo now. I think people were also expecting WSE-4, as it’s been 2 years now since WSE-3 was launched.
aneryu6日前
It would be even better if a version available to individual users were released soon.
kobe_bryant6日前
can these vibe coded sites please set a max width and overflow so their sites work fine on mobile
OutOfHere6日前
Five years from now, I don't know why anyone will still be using Nvidia for inference. Note that Cerebras is for inference only, not for training.

I understand that Cerebras has competition, but this bodes even more poorly for Nvidia for inference. Nvidia may still have a role to play for training, however.