Cerebras Systems has unveiled its latest model, DeepSeek-R1-Distill-Llama-70B, which boasts an impressive inference speed of more than 1,500 tokens per second—57 times faster than traditional GPU solutions. This advancement allows for instantaneous reasoning capabilities in one of the most advanced open-weight models within an entirely U.S.-based AI infrastructure, adhering to strict data retention policies.
“DeepSeek R1 represents a new frontier in AI reasoning capabilities, and today we're making it accessible at the industry’s fastest speeds,” said Hagay Lupesko, SVP of AI Cloud at Cerebras. The platform remarkably reduces processing time for complex tasks, completing standard prompts that take 22 seconds on competitors’ platforms in just 1.5 seconds—a 15x reduction.
DeepSeek-R1-Distill-Llama-70B integrates the sophisticated reasoning capabilities of DeepSeek's Mixture of Experts (MoE) model with Meta's Llama architecture. Despite its compact 70 billion parameter size, it outperforms larger models in tasks involving complex mathematics and coding.
Security and privacy remain critical for enterprises deploying AI, as Lupesko highlighted, noting that all inference requests are processed in U.S.-based data centers with zero data retention, providing assurance that data remains within U.S. borders and is exclusively owned by the customer.
The new model is available immediately through Cerebras Inference, with API access being offered to select customers participating in a developer preview program. For further details on accessing these advanced reasoning capabilities, potential users are directed to the company’s website.
Comments