Qwen 3.8 27B available on Cerebras at 1500 tokens/s
Points and comments are a snapshot, not live.
Cerebras hosts a 27B model at 1500 tokens/s with 128k context.
Cerebras offers Qwen 3.8 27B on its public endpoints at ~1500 tokens/s with 64k context on the free tier and 128k on the paid tier. All public models are unpruned originals. Storage uses selective weight-only quantization, while activations, attention, and KV cache remain in full precision. The site also hosts GPT OSS 120B at ~3000 tokens/s.
What commenters are saying
Commenters praised the speed but noted high costs for agentic tasks without prompt caching discounts priced into input tokens. Some questioned Cerebras's hosting of small models only, citing wafer SRAM limits. Others found 128k context too small for coding projects. A few reported account issues with Discord support and redirect loops. Cerebras appears on OpenRouter but Qwen 3.8 is not yet there. Some saw use for fast sub-agents with smaller context.