October 1, 2026
Choices instead of long answers: Cloudflare Clef returns probabilities for each option
Clef evaluates possible decisions for an application, while Cloudflare currently handles task-specific fine-tuning through its engineering team.

In Cloudflare's Oct 1 test, Clef checked a domain in 2.2 seconds. The same process took 4.7 seconds with gpt-oss-120b.
The Threat Intelligence team used Browser Run to load and render the site, then had the model classify the domain. With Clef, the entire process took half the time. This is an example of decision-making within an application: the model selects a category rather than writing a response for a person.
Answers from options. Clef accepts a description of the situation in the state field and between 1 and 64 questions in questions. The API supports the noul, choice and score types and returns probabilities for the options. The context window holds 65,536 tokens.
In Cloudflare's measurements across 43 benchmarks on Oct 1, the models responded with the following latencies. The second column shows the typical response time; the third shows the time within which 95% of responses were completed.
| Model | Typical response | 95% of responses | | --- | --- | --- | | Clef | 209.3 ms | 238.6 ms | | Clef-flash | 38.8 ms | 122.4 ms | | Jev | 524.1 ms | 536.0 ms |
Connecting to an application. In Cloudflare Dashboard, open Workers AI โ Use REST API โ Create a Workers AI API Token, create a token and copy the Account ID. Send a POST request to `/accounts/{ACCOUNT_ID}/ai/run/@cf/cloudflare/clef`, with a Bearer token and the model, state and questions fields. The documentation includes ready-to-use curl, Python and Workers examples.
Under Workers AI pricing as of Oct 1, input token costs are as follows:
| Model | Price per 1 million input tokens | | --- | --- | | Clef | $0.24 | | Clef-flash | $0.09 |
Workers AI provides a free allowance of 10,000 Neurons per day. Usage beyond that requires Workers Paid.
Weights for both models are available for local use. Clef has 27 billion parameters; Clef-flash has 9 billion. To run Clef, download it with `snapshot_download("Cloudflare/clef")` and load the model using `joint_schema_model.load_release_model`. Cloudflare tested the example with torch 2.11 and transformers 5.10.2 on a single H200. A separate example is available for Clef-flash using `snapshot_download("Cloudflare/clef-flash")`. Both models are distributed under Apache 2.0.
Cloudflare currently handles reinforcement learning fine-tuning through its FDE team. To join the program and discuss a specific task, the company accepts applications through the Clef RL Interest form.
Cloudflare plans a platform where users will be able to fine-tune Clef and redeploy the model themselves.
