Hacker News
Qwen 3.8 27B available on Cerebras at 1500 tok/SEC
Cerebras now hosts the Qwen 3.8 27-billion-parameter model, delivering up to 1,500 tokens per second. All publicly served models remain unpruned; weights are stored with selective 16-bit/8-bit/4-bit quantization while sensitive layers are de-quantized on-the-fly, and activations, attention, and KV cache stay in full precision. Pruned REAP variants are only available on Hugging Face, not through Cerebras’ production API.