Cerebras CS-4
https://www.cerebras.ai/cs4
463 pointsbysunils34•6日前•275 コメント
If they could launch Qwen 27b or Deepseek Flash that would be amazing.
Oops did they just out GPT-5.6 sol’s parameter count?
> CS-4 delivers more than 1,000 tokens per second on models exceeding 10 trillion parameters
Wow!
What's the point of 1000tok/s if you have to do prefill on every agentic turn which at 100k depth would make it 1.5 min latency every turn?
I understand that Cerebras has competition, but this bodes even more poorly for Nvidia for inference. Nvidia may still have a role to play for training, however.
Guess they don't care about regular devs atm and are focused only on hardware sales.