Published October 6, 2026. Sources: Cloudflare, Developers Digest.
What is Cloudflare Clef?
Clef is an open-weights "decision model" from Cloudflare, released October 1, 2026. Instead of generating text token by token, it takes a situation and a list of multiple-choice options and returns a probability for each. That makes it fast and predictable for the small decisions inside AI workflows: where to route a request, whether to continue, whether a human should check.
- Two sizes: Clef (27B, post-trained from Qwen3.8-27B, multimodal) and Clef-flash (9B, from Qwen3.5-9B).
- License: Apache 2.0, weights on Hugging Face.
- Hosted: on Cloudflare Workers AI.
- Speed: Clef-flash median response of about 39 ms, per reports.
The Reddit thread: r/LocalLLaMA: “Clef — Open-weights decision model by Cloudflare” (about 300 upvotes and 80 comments).
Clef at a glance
| Clef | Clef-flash |
| Size | 27B parameters | 9B parameters |
| Base model | Qwen3.8-27B | Qwen3.5-9B |
| Input | Text and images | Text |
| Workers AI input price | $0.24 per million tokens | $0.09 per million tokens |
| License | Apache 2.0 |
Prices as reported at launch; check Cloudflare for current rates.
Decision model vs a normal LLM
| Chat LLM | Decision model |
| Output | Free text | A probability for each option you list |
| Parsing | You must parse the answer, which can go wrong | Always one of your options |
| Confidence | Hard to read | Built in: 0.92 vs 0.51 tells you when to ask a person |
| Speed and cost | Slower, pays for output tokens | Fast, almost no output |
| Good for | Writing, reasoning, drafting | Routing, classification, yes/no gates |
The Reddit discussion welcomed it because so much of agent building is the boring glue: is this an invoice or a complaint? Is the draft safe to send? A model that answers only those questions, with a confidence score, removes a lot of fragile prompt parsing.
Where decision models fit in a workflow
- Routing: which team, queue or agent should take this?
- Triage: urgent, normal or spam?
- Gates: is this draft ready to send, or does a person review it?
- Tool choice: which tool should an agent call next?
Example: customer support triage
- A customer email arrives.
- The decision model picks a category: order status, return, billing, complaint, other.
- If confidence is high and the category is low-risk, a drafting model writes a reply from your policies.
- If confidence is low, or the category is a complaint or refund, the message goes to a person.
The confidence score is the useful part: it gives you a dial for how much you let AI handle alone.
Where Dooza fits
Dooza AI customer support works this way: AI classifies and drafts, low-confidence and sensitive messages go to a person, and refunds and complaints come to you for one-tap approval.
Dooza is an AI-native company that builds AI products and services for small businesses, from the Dooza Workforce app to the Dooza Agents platform. A Dooza engineer scopes your pilot on a free 30-minute call, and every product starts with a refundable pilot: 100% refund within 14 days. Book a free pilot call or see pricing.
Ready to Start Your Pilot?
Automate your business with AI employees that work 24/7. Start with a refundable pilot: 100% refund within 14 days.