A hands-on look at Decide with Jev — why confidence-scored decisions are a better fit for production AI features than binary outputs, and where that honesty gets inconvenient.
A content moderation check returns true. A refund-abuse flag returns false. An intent classifier returns 'request'. Every one of these looks exactly as confident as a database query returning a row — a clean, typed value, ready to drop into an if statement. None of them say what they actually are: a language model’s best guess, collapsed into a boolean or an enum sometime between the model’s output and your code, with the actual uncertainty thrown away in the collapsing. The model that returned true with 51% confidence and the model that returned true with 99% confidence produced identical values by the time your if (moderationResult) ran. That distinction mattered enormously, and nothing in your code could see it.
Jev, an AI model from a company called TypeSafe, is built around refusing to do that collapsing for you. Instead of generating text you then have to parse into a decision, it answers a typed question directly with a calibrated probability — a number between 0 and 1, the model’s own estimate of how likely “yes” is. Laravel News is running a mini-course called Decide with Jev, building a content preflight checker as the worked example, and a handful of community packages have already shown up wrapping Jev for Laravel specifically. I spent some time with the course’s example and two of the Laravel SDKs to see what actually changes when a binary decision becomes a number, and — the part most “AI returns a probability now” coverage skips — where that number turns into more work for you, not less.
Worth saying plainly up front: this is a very new corner of the ecosystem. The course is weeks old, the Laravel packages are community-built and unofficial, and at least three different people have independently published a Laravel/Jev wrapper in the same few-week window — a sign of real interest, and also a sign that none of these have years of production mileage behind them yet. Treat the maturity assessment at the end of this post as seriously as the API walkthrough.
What Jev Actually Returns
use Jev\Laravel\Facades\Jev;
$result = Jev::state(['draft' => $draftText])
->ask('matches_brief', 'Does this draft do what the brief asked?')->boolean()
->run();
$result->get('matches_brief')->value(); // 0.96, not true
That’s the entire shift. A boolean()-typed question doesn’t come back as PHP’s true or false — it comes back as a probability, the model’s calibrated estimate that the answer is “yes.” The course’s own worked example makes the contrast concrete: a complete draft that actually does everything the brief asked for scores 0.96. A thin draft that mentions the topic but never gives a single concrete step scores 0.02. A third draft — genuinely mixed, some of what was asked for and some that wasn’t — is the one that doesn’t resolve to a clean extreme, and that mixed case is exactly where this approach starts paying for itself over a model that was forced to just say true or false about something that was honestly neither.
// From one of the community Laravel SDKs — a slightly different
// surface over the same underlying idea
use Jev\Contracts\JevClient;
use Jev\DTO\Question;
final class CheckDraftAgainstBrief
{
public function __construct(private JevClient $jev) {}
public function handle(string $draft, string $brief): float
{
$result = $this->jev->decide(
['draft' => $draft, 'brief' => $brief],
[Question::make('matches', 'Does the draft satisfy the brief?')->boolean()],
);
return $result->get('matches')->number(); // the probability, as a float
}
}
The honest caveat the course itself states explicitly, and worth repeating because it’s the whole point of this post: a 0.96 is not proof the draft is correct. It’s the model’s own calibrated estimate of how likely “yes” is, which is a genuinely different and more useful thing than a flat true — but it’s still a model’s estimate, not a ground truth, and treating a high probability as certainty is the same mistake as treating a true boolean as certainty, just with better-calibrated numbers underneath it.
Thresholds — Where the Convenience of a Boolean Quietly Comes Back
A probability on its own doesn’t make a decision. Somewhere, a threshold has to turn 0.87 into “proceed” or “don’t,” and this is the part that genuinely does feel like extra work compared to a plain boolean, at least at first.
// A community SDK built specifically around this exact threshold pattern
$insult = $result->noul('contains_insult');
$insult->isTrue(0.6); // probability >= 0.6
$insult->isFalse(0.4); // probability < 0.4
$insult->isUncertain(0.4, 0.6); // sits in the band between — neither
// confident yes NOR confident no
The band in the middle is the genuinely inconvenient part, and it’s inconvenient on purpose. A plain boolean API never gives you the option to say “I don’t actually know” — every answer gets forced into one of exactly two buckets, and a model that was 51% sure produces output indistinguishable from a model that was 99% sure. isUncertain() exists specifically to stop that forcing, which means your application now has to have an actual answer for “what happens when the model is uncertain,” instead of getting to silently default to whichever bucket the collapsed boolean happened to land in.
// ❌ The boolean version — no path for "I'm not sure," so uncertain
// cases get silently treated exactly like confident ones
if ($moderationFlag) {
$comment->block();
} else {
$comment->approve();
}
// ✅ The probability version — three real outcomes instead of two,
// because the third outcome (genuine uncertainty) actually exists
// in the model's own output and shouldn't be thrown away
match (true) {
$flag->isTrue(0.85) => $comment->block(),
$flag->isFalse(0.15) => $comment->approve(),
default => $comment->queueForHumanReview(), // the honest outcome
// for the probability
// range nobody can
// confidently call
};
This is the actual tradeoff the subtitle is pointing at. A confidence-scored decision is a better fit for a production AI feature precisely because it refuses to pretend a 55% guess and a 95% guess are the same kind of answer — and that refusal means the application now needs a real, designed path for the uncertain middle, which a boolean API never forced anyone to build, because a boolean API never admitted the middle existed in the first place.
Choosing a Threshold Isn’t a Technical Decision — It’s a Business One
// One package in this space makes the threshold an explicit, named
// decision object rather than an inline match() — worth the extra
// structure once more than one place in the app needs the same policy
final class RefundDecision implements Decision
{
public function __invoke(Assessment $assessment, RefundAbuse $judgment): RefundOutcome
{
$abusive = $assessment->likelihood('abusive');
$doubtful = $assessment->rating('credibility')->below(2.0);
return match (true) {
$abusive->isTrue(0.75) => RefundOutcome::Deny,
$abusive->isUncertain(0.4, 0.75) || $doubtful => RefundOutcome::ManualReview,
default => RefundOutcome::Approve,
};
}
}
The number inside isTrue(0.75) is not a constant to pick once and forget. It’s the actual encoding of how much false-positive risk the business is willing to accept against how much false-negative risk — denying a legitimate refund request outright at 0.75 confidence is a product and trust decision wearing a numeric disguise, not a technical default with an objectively correct value. A fraud-detection feature with a low tolerance for false accusations wants that threshold high and a wide uncertain band routed to a human. A spam filter with cheap, reversible false positives (a flagged message just gets a second look, not an account ban) can run a much more aggressive threshold with a narrower uncertain band. Getting a calibrated probability back from the model doesn’t remove this decision — it’s what finally makes the decision possible to have explicitly, instead of it being made implicitly and invisibly by whatever threshold a text-parsing regex happened to collapse things at.
What’s Genuinely Good in the Current Ecosystem
Testing without hitting the model at all. Every SDK in this space ships a fake — a deterministic, network-free stand-in for JevClient that returns a fixed answer for a given question, making a threshold’s branching logic fully unit-testable.
use Jev\Laravel\Testing\JevFake;
$fake = new JevFake(['matches_brief' => 0.92]);
$this->app->instance(JevClient::class, $fake);
// exercise application code, then:
$fake->assertAsked('matches_brief');
This matters more for a probability-based decision than it did for a boolean one — testing every branch of a three-way match() (confident yes, confident no, and the uncertain middle) against a real, non-deterministic model call would be slow and genuinely flaky. A fake that can be told “return exactly 0.92” or “return exactly 0.5” makes every branch — including the uncertain one, the branch a boolean-only test suite never had a reason to exercise in the first place — directly and deterministically testable.
Provenance, for the “why did it decide that” question that eventually always gets asked. One of the packages in this space records, per decision, which model actually answered, the exact model identifier, and the request ID the provider issued — specifically so a support or trust-and-safety team can correlate a disputed automated decision back to TypeSafe’s own logs rather than staring at an opaque true/false with no trail behind it. For anything gating a real user-facing outcome — an account action, a refund, a published piece of content — that audit trail is the difference between “we can show exactly what happened and why” and “the model said no, we don’t know why” during an actual dispute.
Where the Honesty Gets Genuinely Inconvenient
To be direct about the cost, not just the benefit: a probability-returning API is strictly more work to build correctly than a boolean-returning one, and that extra work is the actual price of the better behavior, not a rough edge that’ll get smoothed away.
Every decision point needs an explicit threshold, and someone has to own picking it. A boolean API ships with the threshold already baked in, invisibly, by whoever built the text-parsing logic upstream. A probability API makes you choose, explicitly, in code that’s reviewable — which is better for the application’s correctness and genuinely more work for whoever’s building the feature, because “what confidence counts as yes” is now a question that has to be answered on purpose rather than one that quietly answered itself.
The uncertain middle needs a real destination, not just a default case that happens to compile. Routing to human review is the obvious answer and not always the available one — a feature running at a volume where human review doesn’t scale needs a genuine fallback policy for the uncertain band (a more conservative default action, a retry with additional context, a second, more targeted question), and building that policy well is real design work a boolean system never demanded, because a boolean system never had an uncertain band to route anywhere.
The ecosystem’s own immaturity is a real, current cost, not just a caveat to mention once. Multiple competing unofficial Laravel SDKs wrapping the same underlying model, published within the same few-week window, means picking one today is picking without years of track record to lean on — the right move for a new project exploring this pattern, and a real risk to weigh honestly before routing an existing production decision path through any of them.
The One Rule
A boolean AI decision was never actually more certain than a probability-scored one — it just hid its own uncertainty somewhere upstream, in a parsing step or a prompt instruction nobody examined closely, and called the result true with exactly the same confidence whether the model was barely over the line or completely sure. Jev’s actual contribution isn’t a smarter model. It’s refusing to let that uncertainty disappear before it reaches your code — which means a 51%-confidence answer and a 99%-confidence answer are finally distinguishable in your logic instead of being silently identical by the time your if statement runs. That’s a strictly more honest API. Honest APIs are more work, because they make you build the decision you were always actually making — you just weren’t making it on purpose before.
