Cloudflare Launches Clef and Clef-Flash Decision Models

AI Tech Team
•
October 3, 2026
•
👁️ 7 views
🖼️ Featured Image / Generated Result
Cloudflare Launches Clef and Clef-Flash Decision Models

Cloudflare has introduced Clef and Clef-flash, a new class of open-source decision models designed for fast, structured choices inside AI applications and agent workflows. Unlike general-purpose language models that generate open-ended text, decision models are designed to choose from a predefined set of options and attach probabilities to those choices.

Cloudflare announced the models on October 1, 2026 and is hosting them through Workers AI. The company also released the model weights on Hugging Face under the Apache 2.0 license, allowing developers to experiment locally as well as through hosted infrastructure.

What are decision models?

A conventional LLM can reason, write text, call tools, and produce many possible forms of output. That flexibility is useful, but it can be unnecessary when an agent only needs to make a bounded decision. A decision model instead receives an input and a defined set of possible answers, then scores those choices.

Cloudflare gives examples such as deciding whether a customer-support request is urgent and selecting the appropriate team. The same pattern can be used for routing, tool selection, classification, policy checks, and other points in an agent pipeline where a reliable structured answer is more useful than generated prose.

Clef and Clef-flash

Cloudflare released two models. Clef is based on a Qwen3.8-27B backbone, while Clef-flash uses Qwen3.5-9B. Cloudflare says its architecture adds a specialized routing and scoring mechanism rather than relying on ordinary text generation. The models are trained to produce schema-bound probabilities, which can make their output easier to consume programmatically.

In Cloudflare’s published benchmark results, Clef reported 98.47% on BFCL case-exact evaluation and 91.93% on API-Bank accuracy, while Clef-flash reported 98.76% and 93.11% respectively. These figures are vendor-reported benchmark results and should be evaluated against a team’s own workload before drawing conclusions about production performance.

Latency and agent workflows

One of the main reasons to use a specialized decision model is speed. Cloudflare reports median latency of 209.3 milliseconds for Clef and 38.8 milliseconds for Clef-flash in its evaluation, with p95 latency of 238.6 milliseconds and 122.4 milliseconds respectively. The company also says the models benefit from being hosted on its edge infrastructure.

For an agent, that can translate into a useful architecture: a larger reasoning model can handle the difficult part of a task while a smaller decision model handles repeated routing or validation steps. This can potentially reduce latency and the number of expensive general-purpose model calls.

Reinforcement learning for customization

Cloudflare also introduced an RL fine-tuning product for decision models. The idea is to let customers adapt Clef to their own data and decision boundaries instead of building a completely new model. This is particularly relevant for organizations with internal routing rules, support taxonomies, operational policies, or domain-specific classification requirements.

What developers should watch

Decision models do not replace reasoning models. Their narrower output space makes them unsuitable for tasks such as open-ended coding, long-form writing, or complex reasoning. Their value is in the spaces between those larger tasks: deciding what tool to call, where to route a request, whether an action should proceed, or which structured category best matches an input.

Practical takeaway: Developers building agents should consider testing a decision model as a lightweight control layer around larger models. Measure accuracy, calibration, latency, and failure behavior on real application data rather than relying only on published benchmarks.

Source: Cloudflare — Introducing Clef

← Back to all articles