Cloudflare Clef icon

Cloudflare Clef

Open-weight decision models that return typed, probabilistic answers for classification and routing workflows.

Content updated today

Get Started
Cloudflare benchmark chart comparing Clef and Clef-flash decision accuracy and latency

Overview

Cloudflare Clef is a family of two decision models, Clef and Clef-flash, released on October 1, 2026. Applications supply a state, such as a support ticket or structured record, and a schema of questions. The models return bounded answers with probabilities for choices that the application has defined. This makes them useful for routing, triage, and classification steps where a program must act on a typed result. They are not general-purpose chat assistants or text-generation replacements.

Both models can be called through Cloudflare Workers AI or downloaded as Apache-2.0-licensed weights from Hugging Face. The hosted route avoids operating model infrastructure; self-hosting provides control over deployment but leaves hardware and operations to the adopter. Cloudflare's separate reinforcement-learning fine-tuning service is not required to use the released models.

Key Features and Model Choice

Decision factor Clef Clef-flash
Model size 27B parameters 9B parameters
Better starting point Precision-sensitive decisions Decisions on a latency-critical path
Hosted model ID @cf/cloudflare/clef @cf/cloudflare/clef-flash
Context window 65,536 tokens 65,536 tokens
Workers AI listed unit price $0.24 per million input tokens $0.09 per million input tokens

Cloudflare describes Clef as its more precise model and Clef-flash as its faster option. In Cloudflare's published benchmark run, median request latency was 209.3 ms for Clef and 38.8 ms for Clef-flash. These are vendor measurements, not a latency guarantee for a particular deployment or input. Accuracy also varies by task: Clef leads on several published classification tests, while Clef-flash scores higher on some others. Teams should choose against their own decision labels and acceptable error rates, rather than treating model size as a universal quality ranking.

Inputs, Outputs, and Integration

A Workers AI request names either model and includes a state plus one or more typed questions. The supported question forms are Boolean (noul), named choices (choice), and ordered scores (score). The response returns a result for each question, including probabilities that application code can use to route or escalate a case. Cloudflare documents a limit of 1 to 64 questions per request and optional embedded images with separate size and count limits. Both hosted model pages list text, JSON, image, and video understanding; developers should check the current request schema for the exact media format they intend to send.

The two models share the same decision schema, so a team can compare them without redesigning its questions. Cloudflare also says its hosted decision API is compatible with the Jev System One format. The downloadable Clef weights and Clef-flash weights are separate repositories; local serving has its own runtime and hardware requirements.

Hosted and Self-hosted Cost

Cloudflare Workers AI includes a shared free allocation of 10,000 Neurons per day. Going beyond that allocation requires Workers Paid. The current model pages list $0.24 per million input tokens for Clef and $0.09 per million input tokens for Clef-flash; these are model-specific hosted usage rates, not the starting price of a subscription or a complete workload estimate. Check Workers AI pricing for billing and free-allocation details before estimating production volume.

The Apache-2.0 model weights can be downloaded without a model-license fee. Running them yourself still incurs compute, storage, and maintenance costs. The larger 27B Clef model may require more resources than the 9B flash variant, depending on quantization and serving setup.

Best For

  • Developers adding schema-bound classification or routing to an agent or business workflow.
  • Teams that can measure the cost of a wrong decision and trade it against response time.
  • Organizations that want either Cloudflare-hosted inference or open weights for their own deployment.

Avoid If

  • You need open-ended drafting, dialogue, or autonomous tool use rather than answers constrained to a question schema.
  • Your decision cannot be expressed as Boolean, named-choice, or ordered-score questions without losing essential nuance.
  • You need independently verified accuracy or latency guarantees for your own workload before deployment; the published benchmarks are Cloudflare's tests.

Sources & Verification

Track Cloudflare Clef in ToolWorthy Weekly

Important tool updates, better alternatives, and selected AI signals in one weekly brief.

Weekly only. Unsubscribe anytime.