Interestana
Home/News/Cloudflare Releases Open-Weight Decision Models Clef and Clef-flash
MarkTechPost••4 min read

By Interestana AI Editorial — AI-drafted, human-overseen. How we report

Cloudflare Releases Open-Weight Decision Models Clef and Clef-flash

Cloudflare has released Clef and Clef-flash, two new open-weight decision models developed by its Workers AI team. Unlike traditional large language models (LLMs) that generate text token by token and require parsing, these models are designed to process an input state and a schema of typed questions, returning precise probabilities for each allowed answer without generating free-form text. This approach aims to provide greater predictability and control in AI outputs.

Both Clef and Clef-flash are released under the Apache 2.0 license, making them open-weight and available for self-hosting. They are also compatible with TypeSafe AI’s Jev API, a system designed for structured AI outputs. These models are immediately deployable on Cloudflare’s Workers AI platform, and their weights are accessible on Hugging Face. A decision model's core function is to answer a predefined set of questions about an input, rather than generating novel text. Clef supports three primary question types: 'noul' for yes/no questions, returning the probability of a 'yes' answer; 'choice' for selecting one option from a list, providing per-option probabilities and a confidence score; and 'score' for rating against an ordered rubric, yielding a probability-weighted score. On the Workers AI platform, a single request can accommodate up to 64 questions and as many as four images.

Clef is a post-trained version of Qwen3.8-27B, while Clef-flash is based on Qwen3.5-9B. Both models retain the vision encoder from their respective Qwen backbones. The inference process involves two stages. Initially, the backbone performs a single prefill-only pass over the input state and questions. Subsequently, a smaller transformer, termed the joint schema head, processes the final hidden states. This head is responsible for routing relevant evidence to each question, enabling cross-attention between fields, and jointly scoring all potential options. A per-question softmax function then converts the raw logits into probabilities. During training, the backbones were frozen, and the routing head was jointly optimized using rank-256 low-rank adapters. The training loss incorporates label-smoothed cross-entropy alongside a Brier loss for improved calibration. An additional objective, Reinforcement Learning for Calibrated Decisions (RLCD), is employed to provide partial credit for adjacent ordinal choices, further refining the model's ability to make calibrated decisions.

The development of Clef and Clef-flash signifies a move towards more structured and interpretable AI outputs, addressing a key challenge in current LLM deployments where output parsing and validation can be complex. By returning typed probabilities, these models offer a more direct and reliable way to integrate AI decision-making into applications that require deterministic outcomes or precise confidence assessments. This contrasts with the generative nature of traditional LLMs, which can sometimes produce outputs that are ambiguous or require significant post-processing to extract actionable information. The open-weight nature of Clef and Clef-flash also encourages community development and adoption, potentially leading to new use cases and further advancements in decision-oriented AI.

Original source — read the full reporting at the publisher:

Read on MarkTechPost

Get the weekly AI digest

AI news + new model releases, weekly. Drafted by our agents, reviewed by humans.

Read next