Overview
Cloudflare Clef is a family of two decision models, Clef and Clef-flash, released on October 1, 2026. Applications supply a state, such as a support ticket or structured record, and a schema of questions. The models return bounded answers with probabilities for choices that the application has defined. This makes them useful for routing, triage, and classification steps where a program must act on a typed result. They are not general-purpose chat assistants or text-generation replacements.
Both models can be called through Cloudflare Workers AI or downloaded as Apache-2.0-licensed weights from Hugging Face. The hosted route avoids operating model infrastructure; self-hosting provides control over deployment but leaves hardware and operations to the adopter. Cloudflare's separate reinforcement-learning fine-tuning service is not required to use the released models.
Key Features and Model Choice
| Decision factor | Clef | Clef-flash |
|---|---|---|
| Model size | 27B parameters | 9B parameters |
| Better starting point | Precision-sensitive decisions | Decisions on a latency-critical path |
| Hosted model ID | @cf/cloudflare/clef |
@cf/cloudflare/clef-flash |
| Context window | 65,536 tokens | 65,536 tokens |
| Workers AI listed unit price | $0.24 per million input tokens | $0.09 per million input tokens |
Cloudflare describes Clef as its more precise model and Clef-flash as its faster option. In Cloudflare's published benchmark run, median request latency was 209.3 ms for Clef and 38.8 ms for Clef-flash. These are vendor measurements, not a latency guarantee for a particular deployment or input. Accuracy also varies by task: Clef leads on several published classification tests, while Clef-flash scores higher on some others. Teams should choose against their own decision labels and acceptable error rates, rather than treating model size as a universal quality ranking.
Inputs, Outputs, and Integration
A Workers AI request names either model and includes a state plus one or more typed questions. The supported question forms are Boolean (noul), named choices (choice), and ordered scores (score). The response returns a result for each question, including probabilities that application code can use to route or escalate a case. Cloudflare documents a limit of 1 to 64 questions per request and optional embedded images with separate size and count limits. Both hosted model pages list text, JSON, image, and video understanding; developers should check the current request schema for the exact media format they intend to send.
The two models share the same decision schema, so a team can compare them without redesigning its questions. Cloudflare also says its hosted decision API is compatible with the Jev System One format. The downloadable Clef weights and Clef-flash weights are separate repositories; local serving has its own runtime and hardware requirements.
Hosted and Self-hosted Cost
Cloudflare Workers AI includes a shared free allocation of 10,000 Neurons per day. Going beyond that allocation requires Workers Paid. The current model pages list $0.24 per million input tokens for Clef and $0.09 per million input tokens for Clef-flash; these are model-specific hosted usage rates, not the starting price of a subscription or a complete workload estimate. Check Workers AI pricing for billing and free-allocation details before estimating production volume.
The Apache-2.0 model weights can be downloaded without a model-license fee. Running them yourself still incurs compute, storage, and maintenance costs. The larger 27B Clef model may require more resources than the 9B flash variant, depending on quantization and serving setup.
Best For
- Developers adding schema-bound classification or routing to an agent or business workflow.
- Teams that can measure the cost of a wrong decision and trade it against response time.
- Organizations that want either Cloudflare-hosted inference or open weights for their own deployment.
Avoid If
- You need open-ended drafting, dialogue, or autonomous tool use rather than answers constrained to a question schema.
- Your decision cannot be expressed as Boolean, named-choice, or ordered-score questions without losing essential nuance.
- You need independently verified accuracy or latency guarantees for your own workload before deployment; the published benchmarks are Cloudflare's tests.
Sources & Verification
- Last verified: October 3, 2026.
- Evaluation scope: Cloudflare's launch announcement, Workers AI model and pricing documentation, and both official model cards.
- Method: Source-based research from the official product materials below. No hands-on testing was performed for this update.
- Official sources: launch announcement, Clef API documentation, Clef-flash API documentation, Workers AI pricing, Clef model card, and Clef-flash model card.
