The Rise of Clef: A New Contender in Decision AI
In a significant expansion of its AI portfolio, Cloudflare has officially unveiled the Clef model family, a pair of decision-oriented AI models designed to compete directly with TypeSafe’s popular Jev platform. Positioning itself as a faster and more capable alternative, Clef is built on a robust large language model backbone, offering users the ability to handle structured tasks like binary classification, multiple-choice selection, and complex rankings with higher accuracy.
The announcement underscores Cloudflare’s ambition to capture a larger share of the enterprise decision-making space. By offering a drop-in API compatible with Jev, Cloudflare is making it frictionless for developers to migrate their existing agentic workflows to the Clef architecture without needing to rebuild their underlying infrastructure from scratch.
Clef and Clef-Flash: Architectural Overview
The Clef family is bifurcated into two distinct models designed for different performance needs. The primary Clef model serves as the high-capability flagship, while Clef-flash is engineered for environments that prioritize speed and lower overhead. Both models utilize specialized post-trained, frozen versions of the Qwen 3.8-27B and Qwen 3.5-9B architectures, respectively. This backbone allows for a prefill-only inference pass, enabling parallel processing of choices that yields faster results than traditional serial decision models.
Cloudflare’s internal benchmarks suggest that while Clef models may require more significant hardware investment than some smaller competitors, they excel in accuracy across diverse decision-making metrics. Beyond standard text classification, the Clef models distinguish themselves through native support for image and video analysis—a major departure from the text-only constraints often found in competing decision engines. Both models offer a substantial 64k context window, allowing for larger, more nuanced prompts compared to the more restrictive limits of existing industry solutions.
Deployment and Accessibility
One of the most notable aspects of the Clef launch is its dual-deployment strategy. Users can leverage Cloudflare’s Workers AI platform, which utilizes the company’s extensive edge infrastructure to minimize latency. This hosted option is positioned as an enterprise-grade service, albeit at a higher price point than current industry benchmarks, charging $0.24 per million tokens compared to Jev’s $0.042 per million.
For developers who require privacy or specific local infrastructure configurations, Cloudflare has released the Clef models as open-weight under the Apache-2.0 license via Hugging Face. However, local deployment is resource-intensive: running Clef-flash requires a GPU with at least 41 GB of VRAM, while the full-sized Clef model demands a robust 85 GB of VRAM for single concurrency operations with a 64k context window. While these hardware requirements are steep, they offer a viable path for organizations looking to bring decision-making AI entirely in-house without reliance on external cloud APIs.
Why it Matters: The Shift Toward Specialized Decision Models
- Multimodal Capability: By moving beyond text, Clef allows decision agents to analyze visual input, enabling use cases like real-time video surveillance analysis or image-based troubleshooting.
- Enterprise Flexibility: The choice between a managed cloud edge service and a local, open-weight deployment allows companies to balance cost, performance, and data sovereignty based on their specific security requirements.
- Performance Optimization: Through the use of prefill-only passes on a Qwen backbone, Cloudflare is demonstrating that decision-making AI can achieve high performance while maintaining the structural integrity required for critical enterprise workflows.









