October 7, 2026
Clef on OpenRouter skips text generation: returns decisions with probabilities
Clef accepts text, JSON or images and returns typed answers with probabilities. OpenRouter charges only for input.

OpenRouter
@openrouter
Cloudflare's Clef decision-making models are available on OpenRouter Clef (27B) and Clef Flash (9B) from @Cloudflare are open-source decision-making models. Send text, JSON or images and get typed answers with probabilities instead of generated text $0.24 per million input tokens for Clef, $0.09 for Clef Flash. Output is free http://openrouter.ai/cloudflare
· 2.1K views
Clef Flash computes an answer in 1 pass. Instead of a text explanation, the model returns a probability, a selected option or a score.
For classification, that answer can be used directly in code. Cloudflare is already testing Clef to categorize domains: its Threat Intelligence team pairs the model with Browser Run, which loads and renders the site.
What the model returns
Typed answers. Clef and Clef Flash are decision-making models. A conventional generative model writes its answer token by token. Clef computes answers to specified questions without generating text.
Clef Flash supports 3 question types, according to Cloudflare's model card dated Oct 1, 2026:
- `noul`: the probability of a true answer. - `choice`: a selection from options with a probability for each. - `score`: a score based on specified levels.
OpenRouter added both models on Oct 1. Sizes and prices listed in the catalog on that date:
| Model | Parameters | Input per 1 million tokens | Output | | --- | --- | --- | --- | | Clef | 27 billion | $0.24 | Free | | Clef Flash | 9 billion | $0.09 | Free |
In Cloudflare's measurements dated Oct 1 across 43 test sets, half of the answers arrived within the listed time:
| Model | Latency | | --- | --- | | Clef Flash | 38.8 ms | | Clef | 209.3 ms | | Jev | 524.1 ms |
In the Threat Intelligence team's test, Cloudflare measured the entire process, including loading and rendering the site. Results dated Oct 1:
| Model | Site loading, rendering and classification | | --- | --- | | Clef | 2.2 s | | gpt-oss-120b | 4.7 s |
How to integrate it into an app
Through OpenRouter. Access requires an API key and the identifier `cloudflare/clef` or `cloudflare/clef-flash`. Send a POST request to `https://openrouter.ai/api/alpha/decisions`; OpenRouter provides `openrouter-decisions` instructions for integration into code.
When preparing a text `state`, account for truncation to roughly the first 2,000 tokens in Workers AI, which OpenRouter warns about in its model card dated Oct 1. The stated context window is 65,536 tokens, but the model does not read the rest of that text.
In Workers AI, the call looks like this: `env.AI.run('@cf/cloudflare/clef', {model: 'clef', state, questions})`. Cloudflare's documentation includes ready-to-use examples for Workers, Python and cURL.
To run Clef Flash locally, Cloudflare's model card suggests downloading `Cloudflare/clef-flash` via `snapshot_download` and importing `load_release_model` from `joint_schema_model`. Cloudflare tested the example with torch 2.11 and transformers 5.10.2 on a single H200.
Sources: [Cloudflare's catalog on OpenRouter](http://openrouter.ai/cloudflare), OpenRouter instructions, Cloudflare documentation and blog, and Cloudflare's model card on Hugging Face dated Oct 1, 2026.
Weights for both models are published on Hugging Face under Apache 2.0 for local use and experimentation.