Skip to content
The Eval Room

Video + article

Cloudflare vs TypeSafe: The Fight Over AI That Just Decides

Cloudflare's open-weight Clef and Clef-flash return a typed answer with a confidence score instead of text, going straight at TypeSafe's Jev.

There is a new kind of AI model that never writes you a sentence. You hand it a situation and a list of options, and it tells you which one to pick and how sure it is. TypeSafe shipped one in September. Then on October 1st, Cloudflare launched its own and made the weights free. If you build agents, this is the category to watch, because it just went from one vendor to a real fight.

What a decision model actually is

Most of the AI you use is a chat model. It writes text, one word at a time. But a lot of the steps inside an agent do not involve writing at all. Is this email a refund request? Which team should this ticket go to? Should this transaction be flagged, yes or no? For questions like that, a paragraph of prose is overkill.

TypeSafe calls its answer a System One model. Its first one, Jev, takes in state plus a set of typed questions and returns calibrated probabilities and confidence scores instead of prose, with the answer constrained to a small set of values you define up front.

Cloudflare uses a different name for the same idea. It calls these decision models, and says they return fast, typed outputs with confidence scores so agents can classify, route and act without a human in the loop.

Is Clef a Jev competitor?

TypeSafe opened early access to Jev on September 15, 2026, through a waitlist. The pricing is aggressive: TypeSafe lists Jev at $0.042 per million input tokens, with output tokens free.

Here is the twist. Per OpenRouter's write-up, Vercel, Cloudflare, LangChain and Langfuse all integrated Jev within three days of launch. Cloudflare was one of the platforms that plugged Jev in first.

Then on October 1st, Cloudflare shipped a model of its own. The big difference is the license: Clef and Clef-flash are hosted on Workers AI and released under Apache 2.0 on Hugging Face. One side is a closed model you wait in line for. The other is a model you can download today.

Clef vs Clef-flash

Clef comes in two sizes. According to coverage of the launch, Clef (27B) is the highest-precision option and Clef-flash (9B) is built for latency-critical decisions, and both have a 64K-token context window. That is enough room for a support ticket, a user's history or a chunk of a document before you ask for a decision.

You can run them two ways. Hosted on Workers AI, where the launch changelog entry is dated October 1, 2026, or on your own hardware with the Apache 2.0 weights from Hugging Face. That second option is something Jev does not offer.

Is Clef faster than Jev?

Speed is the number Cloudflare is leading with. On a benchmark called Decision Index, the reported median latency is 38.8 ms for Clef-flash and 209 ms for Clef, against 524 ms for Jev. For context, TypeSafe itself advertises response times of 70 to 500 ms for Jev.

Why that matters: an agent might make dozens of these decisions to handle a single user request. Shave a few hundred milliseconds off each one and the whole thing starts to feel instant.

But be clear about what these figures are. They are Cloudflare's numbers, about Cloudflare's own model, in Cloudflare's own launch.

Are the benchmarks reliable?

On Hacker News, commenters doubt that Clef's results hold up beyond public benchmarks. That is a fair worry. A decision model lives or dies on your specific decisions: your categories, your edge cases, your messy real-world inputs. A public benchmark cannot tell you how Clef will handle your refund policy.

Right now every Clef performance number in circulation traces back to Cloudflare or to coverage repeating Cloudflare. There is no independent head-to-head yet. Treat the chart as a reason to test, not a reason to switch. Open weights at least make that test cheap, since you need neither permission nor a waitlist spot.

How to test this yourself

Two experiments settle the claims that matter most.

  • Latency. Write one fixed classification prompt, for example a customer message plus the label set refund, billing_question, cancel, other. Call Clef-flash on Workers AI using the model ID in the changelog and Jev through OpenRouter, at temperature 0, from the same machine and region. Do 20 warm-up calls, then 200 timed calls per model, and record median and p95 wall-clock latency. Your network round-trip is in there, so compare the gap between models rather than the absolute numbers.
  • Accuracy and calibration. Export 300 past decisions from your own system that already have known correct labels, such as support tickets with their final routing. Send each to Clef-flash and Clef with the same fixed label list at temperature 0, and record the chosen label and the confidence score. Bucket results by confidence (0.9 and above, 0.7 to 0.9, below 0.7) and check accuracy per bucket. If high-confidence answers are often wrong, or overall accuracy falls below your current router, the pitch does not survive contact with your data.

The RL fine-tuning service, and how Cloudflare gets paid

Cloudflare's second move explains the business model. Alongside the free models it launched an RL fine-tuning service, built on AI Gateway, Containers and a new Trainer component that redeploys fine-tuned models onto Workers AI. The idea is to take Clef and train it on your own domain-specific decisions.

There is one catch. It is not a sign-up button. The service starts with Cloudflare's forward-deployed engineers, with self-serve planned later, which means staff working with you directly. For now it is aimed at customers big enough to get on a call.

Clef or Jev for your agent?

Cloudflare's bet is that the model is the free part. Going open wins developer trust and adoption, especially against a rival that is closed and waitlisted. The paid part is the loop: train on Cloudflare, host on Cloudflare, call it from Cloudflare.

TypeSafe is betting differently. It sells a single hosted model at a very low listed price, and it got wide distribution fast.

How to choose: if you want to self-host, or you need to own and inspect the model, Clef is the obvious pick. If you just want a cheap API and you already have access, Jev is still a real option. Either way, test both on your own data before you commit.

This category just went from one company to a fight, and that is good news for anyone building agents. Competition means lower prices and faster models, and an open option means you are never locked in.

Sources