ZipLyne
Book A Call

Blog 14 min read

Clef vs Jev: Cloudflare's Open Decision Model Takes On Typesafe

Clef vs Jev comes down to latency, benchmarks, images, and context length. Cloudflare says Clef is faster; Jev still wins on tool-call deciding.

Download .md

What's the difference between Clef and Jev?

A decision model returns a typed, probability-scored answer, a yes/no, a choice, or a score, instead of generated text, so code can route, escalate, or defer to a human without parsing a sentence. Clef vs Jev compares Cloudflare's new open-weight Clef, launched October 1, 2026 on Workers AI, against TypeSafe's Jev, both built for that same category of decision.

Cloudflare open-sourced the weights for Clef and its smaller sibling, Clef-flash, on Hugging Face under an Apache 2.0 license. It also built the request format to match Jev's, so a team already sending Jev-shaped requests can point the same calls at Clef without rewriting an integration layer.

TypeSafe describes Jev as a model built for fast, structured decisions: it returns type-safe structured values with calibrated probabilities instead of generated text, and it gives up string generation entirely to get there. Clef takes the same basic shape, typed outputs with probabilities, and runs it on Cloudflare's own infrastructure.

Everything that follows comes from Cloudflare's own launch comparison. That's worth saying up front: a vendor grading its own launch day against a competitor isn't the same as an independent benchmark, and none of these figures have been re-verified by a third party yet.

Diagram: Clef vs Jev: Cloudflare's Open Decision Model Takes On Typesafe

Is Clef faster than Jev?

Clef beats Jev on every latency number Cloudflare has published, and Clef-flash beats both by a wide margin. Cloudflare's comparison put Clef's median response at 209.3 ms against Jev's 524.1 ms, with Clef-flash finishing in 38.8 ms. These are Cloudflare's own figures from its own test set, not an independently audited benchmark.

ModelMedian latencyp95 latency
Clef209.3 ms238.6 ms
Clef-flash38.8 ms122.4 ms
Jev524.1 ms536.0 ms

Cloudflare also ran a separate internal workflow, fetching, rendering, and classifying a website, where Clef finished in 2.2 seconds against 4.7 seconds for another Cloudflare general LLM in the same test. That example doesn't involve Jev at all; it's Cloudflare checking Clef against its own prior model, and it's the only latency data point in the launch post that isn't a direct Clef-vs-Jev number.

Clef-flash is the fastest option in this comparison by a wide margin, but speed alone shouldn't decide the pick. Cloudflare chose the prompts, the hardware, and the test conditions. Run your own timed sample before you commit either model to a latency-sensitive path.

Clef vs Jev: how do the public benchmarks compare?

Clef beats Jev on all five of Cloudflare's published core benchmarks, covering function calling, API usage, banking intent classification, and phishing detection. Clef-flash beats Jev on four of the same five. This is a benchmark set Cloudflare chose for its own launch post, not a neutral third-party suite, so treat the margins as directional.

BenchmarkClefClef-flashJev
BFCL98.4798.7695.75
API-Bank91.9393.1188.19
BANKING77 macro-F194.2090.9379.74
CLINC150+OOS97.4366.7789.27
PhishNChips79.6075.0562.55

BFCL tests whether a model picks the right function call for a given request. API-Bank tests the same skill across a broader set of API tools. BANKING77 and CLINC150+OOS both test intent classification, sorting a request into one of many categories, with CLINC150+OOS adding out-of-scope detection into the mix. PhishNChips tests phishing and scam detection.

Clef's biggest margin sits in BANKING77, where it scores 94.20 against Jev's 79.74. Clef-flash edges past Clef on two of the five, BFCL and API-Bank, despite being the smaller, faster model, but CLINC150+OOS is the one benchmark here where Clef-flash loses to Jev, 66.77 against 89.27. None of this has been independently re-run. Run the categories that match your actual workflow against your own labeled data before picking a winner off this table.

Where does Jev still beat Clef?

Jev leads Cloudflare's own numbers on two more public benchmarks beyond the five above: deciding whether to call a tool at all, and retrieval-heavy ranking. Across the full set of seven public benchmarks Cloudflare published, that puts Clef ahead of Jev on 5 of 7 and Clef-flash ahead on 4 of 7, with Jev winning the same two benchmarks against both: When2Call and BRIGHT nDCG@10.

BenchmarkClefClef-flashJev
When2Call72.3765.5880.97
BRIGHT nDCG@1045.9139.2647.52

Cloudflare reports Jev at 80.97 on When2Call against 72.37 for Clef and 65.58 for Clef-flash. When2Call is one of Jev's clearest leads among the benchmarks Cloudflare published.

When2Call measures something subtler than a typical classification task: whether the model should call a tool in the first place, as opposed to answering directly. That's a judgment call about the state of a conversation, close to the core problem Jev was built to solve. TypeSafe describes training Jev with what it calls Reinforcement Learning for Calibrated Decisions, aimed at producing epistemically honest probabilities on exactly this kind of boundary decision.

BRIGHT nDCG@10 tests retrieval ranking quality, another task that rewards a model tuned for weighing many candidates against uncertain signal. Jev edges both Clef and Clef-flash here too, 47.52 against 45.91 and 39.26.

Separate from these public benchmarks, Cloudflare also measured agent trace observability on TypeSafe's own workflow suite, a different evaluation set covering real business tasks. Jev leads there as well, a result covered alongside the other workflow-suite numbers next. If your workflow leans on tool-call judgment, Cloudflare's own numbers say Jev is still the stronger pick today.

How do Clef and Jev perform on real business workflows?

Clef beats Jev on three of the four tasks in Cloudflare's results on TypeSafe's own workflow suite: invoice processing, customer service, and security incidents. Jev leads only on agent trace observability. These numbers come from TypeSafe's own workflow suite, the task categories operators actually run, not generic chat benchmarks.

WorkflowClefClef-flashJev
Invoice processing64.757.161.8
Customer service76.377.076.0
Security incidents62.961.761.7
Agent trace observability68.569.871.6

Clef leads invoice processing at 64.7 against Jev's 61.8, and edges security incidents at 62.9 against 61.7. Clef-flash takes customer service outright at 77.0, with Clef also ahead of Jev there at 76.3 to 76.0. Jev's only win on this suite is agent trace observability, where it scores 71.6 against 68.5 for Clef and 69.8 for Clef-flash, the same lead covered in the previous section.

The margins here are tighter than the public benchmarks above; none of the four categories moves more than about 7 points between the best and worst model. That matters more for a buying decision than a wide gap on an academic benchmark, because it means the choice between Clef and Jev on these specific tasks is close enough that your own data, not Cloudflare's test set, should settle it.

Invoice processing is also where a lot of teams first wire in a decision model, usually right after they've already automated the document capture and routing around it. If you're at that stage, what to automate after you already automated invoicing covers the next workflow worth fixing once the approval decision itself is handled by a model instead of a person.

Can Clef handle images, and does context length matter?

Clef accepts images through a vision encoder; Jev is text-only today. That single difference rules Jev out for any workflow where the decision depends on a photo, a scanned document, or a screenshot, regardless of how the two compare on text-only benchmarks.

Context length is the second structural gap. Cloudflare's Clef runs with a 64k context window against Jev's 32k. For a short routing decision that difference won't matter. For a security review that needs to read a long log, or a customer service decision that needs a full ticket history in context, Clef can hold roughly double the state in one pass.

Open weights change the deployment conversation too. Because Clef and Clef-flash are released on Hugging Face under an Apache 2.0 license, a team with data residency requirements or a reason to avoid sending decisions to a third-party API has an option Jev doesn't offer today: run the weights yourself instead of depending on a hosted endpoint. The visible sources do not show an open-weight or self-hosted option for Jev; TypeSafe's own launch post describes it as available in early access through its own service. Teams already building on Cloudflare's Workers AI get Clef without adding a new vendor relationship at all.

Is Clef API-compatible with Jev, and what do you need to retest before switching?

Cloudflare says Clef's request shape matches Jev's, which means a team already sending Jev-formatted calls can point the same integration at Clef with low engineering effort. That claim covers the wire format, not the decision quality behind it.

Compatibility at the request level tells you nothing about whether Clef's confidence scores mean the same thing Jev's do on your data, or whether the two models agree on the cases that actually matter to your workflow. A swap that compiles cleanly can still ship worse decisions if you skip re-validating accuracy and calibration.

If you're still deciding whether a decision model belongs in your stack at all, AI that doesn't talk covers how Jev works and what a typed, probability-scored output actually looks like in practice before you weigh a second vendor against it.

How do you test Clef vs Jev on your own workflow before switching?

Vendor benchmarks tell you what Cloudflare measured on its own data, not what either model will do on yours. Running both models side by side on a sample of your real decisions is the only way to know which one actually fits your workflow. Four steps get you a real answer:

  1. Pull a labeled sample of real decisions from your workflow. Use actual historical cases, invoices, tickets, flagged transactions, with a known correct outcome for each one. A small or narrow sample won't surface the edge cases that matter.
  2. Run both models on identical inputs. Send the exact same request to Clef and Jev for every case in the sample, with the same prompt structure and the same schema, so the comparison isolates the model, not the setup.
  3. Compare accuracy against ground truth, plus calibration. Accuracy alone hides a bad model. Check whether each model's stated confidence lines up with how often it's actually right; a model that says "90% confident" and is wrong a third of the time is worse than one that's honest about uncertainty.
  4. Measure latency under your actual load. Cloudflare's published latency numbers come from its own test conditions. Your network, request size, and concurrency will move the real numbers, sometimes by a lot.

Skipping any of these four steps means you're choosing a production model off a vendor's launch post instead of your own evidence.

Should you use Clef or Jev for routing, triage, or approvals?

Match the model to the task, not the headline benchmark score. Cloudflare's own numbers point to a fairly clean split: Jev still wins on tool-call deciding, retrieval-heavy ranking, and agent trace observability, while Clef wins on image inputs, longer context, latency-critical hot paths, and anywhere self-hosting or data residency matters.

  1. If the decision is "should I call a tool at all" or observability across an agent run, lean Jev. Cloudflare's own When2Call and agent trace observability numbers back Jev here, and TypeSafe's tooling is already built around that workflow shape.
  2. If the decision involves an image, a long document, or a latency-sensitive hot path, lean Clef. Vision support, the 64k context window, and Clef-flash's sub-40 ms median response cover cases Jev can't match today.
  3. If you need to self-host or keep decisions off a third-party API, lean Clef. Open weights on Hugging Face are the only option between the two that lets you run the model yourself.
  4. If your team is already deep in TypeSafe's tooling and support relationship, factor in the switching cost. Matching request shapes makes testing cheap, but a full migration is a separate decision from a benchmark comparison.

None of that matters until the routing, triage, or approval logic is actually running in production, not sitting in a benchmark spreadsheet. Picking a model is an afternoon's work; wiring it into a real workflow with proper fallbacks, logging, and human escalation is the part that takes engineering. If you're trying to figure out which process to automate first before you even get to this model choice, how to prioritize business processes for AI automation walks through the scoring.

Stop comparing benchmarks and start shipping the decision logic your workflow actually needs.

Frequently asked questions

What is a decision model, and how is it different from a typical LLM?

A decision model returns a typed, probability-scored answer, a yes/no, a choice, or a score, instead of generated text, so application code can route, escalate, or defer to a human without parsing a sentence. Clef and Jev both work this way. A standard LLM generates strings that still need parsing and validation; a decision model skips that step entirely, which is why both Cloudflare and TypeSafe built dedicated architectures instead of fine-tuning a chat model.

Is Clef open source, and what license does it use?

Yes. Cloudflare released Clef and its smaller sibling, Clef-flash, as open-weight models on Hugging Face under an Apache 2.0 license on October 1, 2026. That license lets a team download the weights and run them on its own infrastructure instead of depending on a hosted API, which matters for data residency or avoiding a third-party vendor relationship. Jev has no equivalent open-weight release; TypeSafe offers it only through early access on its own hosted service.

What is Reinforcement Learning for Calibrated Decisions (RLCD)?

Reinforcement Learning for Calibrated Decisions, or RLCD, is the training method TypeSafe built to produce Jev's probability scores. Instead of optimizing for human preference like RLHF or verifiable rewards like RLVR, RLCD trains the model to give epistemically honest confidence numbers on structured decisions, so a 90% confidence score is actually right about 90% of the time. It's the mechanism behind Jev's calibrated outputs rather than generated text.

What's the difference between Clef and Clef-flash?

Clef-flash is a smaller, faster version of Clef that trades some accuracy for speed: it posts a 38.8 ms median response against Clef's 209.3 ms, and it beats Clef on two of five public benchmarks, BFCL and API-Bank, despite the size difference. The tradeoff shows up on CLINC150+OOS, where Clef-flash scores 66.77 against Clef's 97.43 and loses even to Jev's 89.27. Pick Clef-flash for raw speed, Clef when accuracy on harder classification tasks matters more.

Can you self-host Jev the way you can self-host Clef?

No. Jev is available only in early access through TypeSafe's own hosted service today; the visible launch materials describe no open-weight or self-hosted version. Clef and Clef-flash, by contrast, ship as open weights on Hugging Face under an Apache 2.0 license, so a team can run them on its own infrastructure. If self-hosting or keeping decisions off a third-party API matters to your workflow, that's a structural reason to lean toward Clef today.

Who built Jev, and when did TypeSafe launch it?

Diogo Almeida, a former OpenAI researcher who worked on the methods behind ChatGPT, founded TypeSafe and launched Jev on September 14, 2026, after two years in stealth. TypeSafe calls Jev a System One model, built for fast structured decisions rather than conversation. Cloudflare's Clef followed about two and a half weeks later, on October 1, 2026, as an open-weight competitor built to the same request format.

Keep reading.

All posts
Next Step

Let’s Build What’s Next.

Bring the business problem. We’ll talk through what would make a difference and where to start.