Jev vs Gemini 3.8 Flash: We Tested a Decision Model on 203 Dead Sales Threads

Illustrated sales rep holding a stopwatch and clipboard while two AI robots race to sort envelopes into four glowing trays, with the headline 203 dead deals, two AI models.

We gave a new kind of AI model, one that decides instead of writes, a job every sales team has and nobody does: reading 203 dead email threads and finding the buyers who said “later”.

The model is Jev, from TypeSafe. Its launch page claims it is “193.6x faster” with a “238x lower input price than Claude Fable 5.1” on the tasks it is built for. Launch claims are launch claims, so we tested it ourselves, head to head with Google’s Gemini 3.8 Flash, on a test set we wrote by hand with every answer fixed before either model saw a thread.

The short version: Gemini was more accurate (98% against 95%). Jev was 28 times cheaper and 6 times faster. On the job that mattered, finding the deals worth reopening, they tied: both caught 60 of the 61 real deferrals, and both put the same wrong thread in their top ten.

This is the full teardown, written for RevOps and GTM engineers deciding whether a decision model belongs in their stack. The dataset, the six typed questions, the scoring code, every mistake, what went wrong while we built it, and how to reproduce every number. If you want the sales playbook rather than the method, read Dated, not dead: how to find the deals in your inbox that said “later”.

TL;DR

  • Accuracy: Gemini 3.8 Flash 98.0% (199 of 203), Jev 95.1% (193 of 203). On the 68 deliberately tricky threads: 94.1% against 88.2%.
  • The job that matters was a tie: both found 60 of 61 real deferrals and got 9 of their top 10 right, with the same wrong pick.
  • Cost and speed: Jev $0.45 per 10,000 threads and 0.35 s per thread. Gemini $12.34 and 2.2 s. This whole test cost Jev $0.009 and Gemini $0.25.
  • Jev read timing better: it picked the right come-back window for 92% of real deferrals, Gemini for 77%.
  • Both were well calibrated: expected calibration error 0.020 (Jev) and 0.023 (Gemini). Jev’s 171 answers at 95%+ confidence were 98% right.
  • The fix for most mistakes is not a better model. Filter machine mail in code, let code do dates and senders, and send answers under 75% confidence to a person.

What Jev is

Jev is what TypeSafe calls a System One model. The docs describe it plainly: “Send state and typed questions; get structured answers your code can use directly.” It does not generate prose. It answers three kinds of question:

PrimitiveWhat you askWhat comes back
ChoicePick one option from a listA probability for each option
ScoreRate the state on a rubric you describeA score on your scale
NoulIs this statement true?A 0 to 1 read

Every answer carries a calibrated confidence, which is the part that matters for automation: your code can act on the sure answers and route the unsure ones. The model we used was jev-1.13.0, in early access since 15 September 2026, called through TypeSafe’s API. The homepage lists input at $42 per billion tokens (about $0.042 per million), and output is free.

Why triage is the right test

Most AI in sales is pointed at writing: draft this email, rewrite this subject line. That is the easy end of the job. The hard end is reading an archive of stalled conversations and deciding, thread by thread, whether the buyer deferred or declined, how warm the deal still is, and when they asked you to come back.

That job is big and it is dull, which makes it an exact fit for a cheap, fast classifier. Challenger Inc.’s research attributes 40 to 60% of lost deals to buyer indecision, and Salesforce puts rep time on non-selling tasks at 60%. A lot of reopenable pipeline is sitting in threads nobody has time to read. A model that sorts all of them for under a dollar per ten thousand changes whether anyone ever does.

The test set: 203 threads written by hand

We did not use real customer mail. We wrote 203 synthetic threads one by one (not from templates) for a fictional seller with four unrelated products: AP automation, shift scheduling, security awareness training and a support platform. Four products means no single vocabulary gives the answer away. “Today” for the test is September 2026.

LabelMeaningThreadsTricky
soft_deferralBuyer postponed and invited us back6116
hard_noBuyer declined, however politely5017
went_quietBuyer never gave a position; our message went unanswered5014
not_salesAuto-replies, notifications, recruiters, internal mail4221
Total20368 (33%)

A third of the threads are deliberate near-misses, because real inboxes are full of them:

  • A polite no that sounds like “later”: “No need to circle back next quarter. We’ve consolidated for three years.”
  • A buyer who deferred, followed by our own acknowledgement as the last message.
  • Our own follow-up using deferral words (“happy to revisit next quarter”) when the buyer never replied.
  • “Maybe later” brush-offs, labelled as deferrals but never worth reopening.
  • Out-of-office replies with return dates, trial-expiry notices, calendar declines and recruiter emails.
  • Messy text: signatures, quoted replies, a two-word answer (“lol no bandwidth. Q1.”), non-native English, an assistant replying for the boss.

Separately, 44 threads are marked worth reopening now: genuine deferrals whose requested month has arrived by September 2026, with untimed deferrals counting after six months. The labelling rules are written down in LABELING.md, and the dataset builds byte for byte from source.py.

The six questions

Each thread went to Jev as one request with six typed questions. Everything else happens in our code.

QuestionTypeReturnsWhy we ask
Which category is this thread?Choicesoft_deferral, hard_no, went_quiet, not_sales, each with a probabilityThe core triage decision
Did the buyer name a specific event they were waiting on?Noul0 to 1A named trigger makes the reopen specific
How explicit is what they were waiting on?Score5 described levels, from “nothing” to “named event with a date”Feeds warmth
How senior is the contact?Score5 levels, from “unknown” to “VP, C-level, founder”Feeds warmth
When did they suggest coming back?Choicewithin 3 months, 3 to 6, 6 to 12, over a year, not statedCode turns this into timing ripeness
Was the last message ours, unanswered?Noul0 to 1Test only; in production, read it from mail headers

Warmth, from 0 to 100, is computed in code, not by the model:

warmth = (0.44 × trigger_clarity
        + 0.24 × seniority
        + 0.32 × timing_ripeness) × category_scale

timing_ripeness = months since the buyer's last message,
                  compared with the come-back window Jev picked

Why split warmth into sub-scores instead of asking for one number? Two reasons. It is explainable: a rep can see that a deal ranked high because the trigger was explicit and the timing had arrived. And it keeps each question small and checkable, which is where a decision model is strongest.

Why do the date maths in code? Because date arithmetic is a documented weak spot for Jev 1.13. The model picks the window (“3 to 6 months”); code works out whether that window has passed. The same logic applies to who sent the last message, which is a fact in the mail headers, not a judgement.

The baseline, and how we kept it fair

We first planned to compare against Claude Sonnet 5.5. We switched to Gemini so that a different model family from the one that helped draft the threads (Claude) would do the grading, leaving neither side with a home advantage. Gemini 3.1 Pro had no free-tier access, so we used Gemini 3.8 Flash.

  • Both models got the same threads, the same question definitions and the same answer key.
  • Warmth was computed by the same code for both, so ranking differences come only from the models’ judgements.
  • Gemini ran in JSON classification mode with low thinking, the setting meant for classification.
  • Costs come from real token counts on every call at each provider’s list price: Jev about $0.042 per million input tokens with free output; Gemini 3.8 Flash $0.75 input and $3.75 output per million on the paid tier.
  • Both ran 8 calls in parallel for the timed batch, and every response was cached.

Results

Jev (jev-1.13.0)Gemini 3.8 Flash
Correct category95.1% (193 of 203)98.0% (199 of 203)
Tricky threads88.2% (60 of 68)94.1% (64 of 68)
Deferral precision / recall0.91 / 0.980.95 / 0.98
Top 10 worth reopening9 of 109 of 10 (same wrong pick)
Precision in top 300.900.93
Come-back window right (real deferrals)92%77%
Calibration error (0 is perfect)0.0200.023
Cost per 10,000 threads$0.45$12.34
Cost of this test$0.009$0.25
Time per thread (median)0.35 s2.2 s
10,000 threads at 8 in parallelabout 8 minabout 48 min

Read the table in two halves. On raw accuracy, the bigger general model wins, and the gap widens on the tricky threads. On the triage outcome, the reopen list, there is no gap: same recall, same top ten, and a slightly weaker top 30 for Jev. Jev was clearly better at one sub-task, reading when the buyer said to come back, which feeds straight into timing.

The price gap will also grow. Google’s pricing page lists Gemini 3.8 Flash doubling to $1.50 input and $7.50 output per million tokens on 1 January 2027, which would roughly double the $12.34 figure.

Compared with TypeSafe’s own claims, our measured gaps are smaller: 6 times faster and 28 times cheaper than Gemini 3.8 Flash, not the 190x-plus the homepage quotes against a frontier model. Both can be true, since Flash is already a small, cheap model. Quote the numbers you measured on your own task.

Calibration: can you trust the confidence?

For automation, calibration matters more than a point of accuracy. If a model says it is 90% sure, it should be right about 90% of the time. Then you can let it act alone on high-confidence answers and route the rest.

Jev confidence bandThreadsRight
About 99%17198%
About 86%1593%
About 75%475%
About 64%875%
About 57%540%
Gemini confidence bandThreadsRight
About 99%190100%
About 86%1070%
About 80%2100%
About 65%10%

Both models are well calibrated overall. Jev spreads its uncertainty more widely, which is useful: the 17 answers below 80% confidence were right only 65% of the time, so the score tells you exactly which answers to check.

That supports a simple review rule: send anything under 75% confidence to a person. On this test that is 15 threads (7%). It catches 6 of Jev’s 10 mistakes, and above the line Jev was right 98% of the time. Reviewing 7% of an archive by hand is a morning, not a quarter.

Where each model got it wrong

Jev missed 10 of 203. Six were not sales mail at all, three ended on our own follow-up, and one was a real deferral read as a no. Eight of the ten were planned near-misses.

ThreadThe last messageCorrectJev saidConfidence
m059Out-of-office: “travelling for our annual conference until March 14”not_salessoft_deferral0.97
m127Auto-reply: “on parental leave until the new year”not_salessoft_deferral0.60
m116Calendar: “Declined: Follow-up call. Let’s find another time.”not_salessoft_deferral0.62
m094Recruiter pitching a Senior AE rolenot_saleswent_quiet1.00
m138Recruiter follow-up on a Head of Sales rolenot_saleswent_quiet1.00
m189Calendar: “Accepted: Intro call”not_saleswent_quiet0.65
m151Our “just checking in”, after the buyer’s “I’ll come back to you”went_quietsoft_deferral0.89
m046Our “Sure, how about Tuesday at 2?” after they asked to move a callwent_quietsoft_deferral0.56
m156Our “Yes, I can hold it through October”went_quietsoft_deferral0.72
m035Buyer: “Next quarter, ok? This one is spoken for.”soft_deferralhard_no0.54

The costly one is m059. An out-of-office reply read as a deferral at 97% confidence scored warmth 97 and landed in the top ten, the only wrong pick on the reopen list. Gemini made the same mistake on the same thread.

Gemini’s four misses: the same travel out-of-office, the same calendar decline, a thread where our own message assumed the buyer was busy until January, and a “maybe someday” brush-off it read as a hard no. Only Jev missed the recruiter emails, the parental-leave reply, the calendar acceptance and the three threads that ended on our follow-up.

Look at the pattern. Four of Jev’s ten misses (the out-of-office, parental leave and both calendar notices) are machine mail that auto-reply headers and calendar MIME types flag for free. The two recruiter emails are a cheap sender-domain or keyword filter away. None of those needs a smarter model.

What went wrong while we built it

The results are clean. The process was not, and the detours are useful if you plan to try this yourself.

  1. We could not embed a live Jev in the published write-up. Its API only accepts browser calls from TypeSafe’s own console, published pages cannot call outside services, and a key in the page would be exposed to readers. We kept saved answers for seven sample threads (total cost $0.0003) and linked each one to TypeSafe’s playground instead.
  2. The practice run used templates, and it showed. Before the real test we ran Jev on 300 template-generated threads: 92% accuracy and 10 of 10 top picks. We did not tune the questions on those results. On the harder hand-written set, accuracy went up to 95% and calibration error halved from 0.040 to 0.020, a reminder that results on template data do not predict results on real-looking mail.
  3. One of our six questions was a bad idea. “Was the last message ours, unanswered?” gave unreliable answers. That fact lives in the mail headers, so production code should read it from there and never ask the model.
  4. Date maths is not the model’s job. Jev picks a come-back window well (92% right), but month arithmetic belongs in code. We designed for that from the start because TypeSafe documents it as a weak spot.
  5. The free tier stopped the baseline at 8 of 203. Gemini’s free tier allowed 20 requests a day. Once billing was on, the remaining 195 threads ran in 56 seconds, and the entire Gemini run cost $0.25.

Five design lessons

  1. Filter machine mail before the model sees it. Auto-reply headers, calendar notices and bounces are cheap to catch in code. That alone would have removed four of Jev’s ten misses and the only wrong pick in both top tens.
  2. Let code do dates and senders. Ask the model for the judgement (which window), and compute the facts (has it passed, who sent the last message) yourself.
  3. Read the confidence score. Route anything under 75% to a person. On this test that meant reviewing 7% of threads to catch most of the errors.
  4. Ask small typed questions, not one big prompt. Warmth built from three sub-scores is explainable, testable and cheap. You can see why a deal ranked high and argue with it.
  5. Pick the model for the job, not the leaderboard. A decision model that never drafts or sends is a better fit for triage than a model that writes. A person writes the three reopens that matter.

Which one should you use?

Use a decision model like Jev to triage the whole archive. At $0.45 per 10,000 threads you can run it on every thread you have, every week, and re-run it when your rules change. Send the low-confidence 7% to a person, or to a bigger model like Gemini for a second opinion. That keeps the bigger model’s cost on the few threads that need it.

Use a general model alone when volume is small, when you also need text generated in the same pass, or when the extra three points of accuracy on tricky threads are worth 28 times the price to you.

Either way, the biggest gain in this test did not come from the model at all. It came from a few lines of code that remove machine mail and compute dates.

Reproduce it

The whole test is open source at github.com/developer943/pipeline-revival-jev. Every model answer is committed, so the results tab works without any API key; you only need keys to re-run the models.

pipeline-revival-jev/
├─ app/
│  ├─ server.py         # the mock inbox: runs both models, scores them
│  └─ index.html
├─ datasets/salesgear-v1/
│  ├─ source.py         # 203 hand-written threads + their answers
│  ├─ build.py          # → threads.jsonl + labels.csv
│  └─ LABELING.md       # how every answer was decided
├─ jevlib.py            # Choice · Score · Noul questions
├─ baseline/
│  ├─ prompt.py         # one identical prompt for every LLM
│  └─ run_gemini.py     # Gemini 3.8 Flash baseline
├─ eval/metrics.py      # precision · recall · precision@k · ECE
└─ runs/                # every model answer, committed, no key needed
# 1. Start the app (Python 3, nothing to install)
python3 app/server.py --data datasets/salesgear-v1

# 2. Open the Results tab (answers are cached, no key needed)
open http://127.0.0.1:8765

# 3. Re-run the models on your own threads
#    set TYPESAFE_API_KEY and GEMINI_API_KEY in .env

The app is called Reopen Inbox. It loads any dataset in the same format, runs both models with identical questions, ranks the reopen list by warmth and scores everything against the answer key. Corrections to the answer key are logged, including whether a model had already answered that thread, so nobody can quietly move the goalposts. Swap in your own exported threads and see what your dead pipeline scores.

Related reading

Frequently asked questions

What is Jev from TypeSafe?

Jev is TypeSafe’s System One model, in early access since September 2026. Instead of generating text, it takes a piece of state plus typed questions (Choice, Score or Noul) and returns structured answers with calibrated confidence that code can act on directly. TypeSafe lists input at $42 per billion tokens, with free output.

Is Jev more accurate than Gemini?

Not on our test. Gemini 3.8 Flash classified 98.0% of 203 sales threads correctly against Jev’s 95.1%, and did better on tricky threads. But both found 60 of the 61 real deferrals and got 9 of their top 10 reopen picks right, so for triage the outcome was a tie.

How much cheaper and faster is Jev than Gemini 3.8 Flash?

On our test, Jev cost $0.45 per 10,000 threads against Gemini’s $12.34, about 28 times cheaper, and answered in a median 0.35 seconds against 2.2 seconds, about 6 times faster. The full 203-thread test cost Jev $0.009 and Gemini $0.25.

What is calibration error and why does it matter?

Expected calibration error measures the gap between a model’s confidence and how often it is right; 0 is perfect. Jev scored 0.020 and Gemini 0.023. Good calibration means you can let a model act alone on high-confidence answers and route low-confidence ones to a person.

Can I run this test on my own inbox?

Yes. The code and dataset are at github.com/developer943/pipeline-revival-jev. The results run without an API key because every answer is committed. To test your own threads, export them through your mail provider’s official API, format them like the sample dataset, and add TypeSafe and Gemini keys.

Why not let the model work out dates and who sent the last message?

Because those are facts, not judgements. Date arithmetic is a documented weak spot for Jev 1.13, and the sender of the last message is in the mail headers. Ask the model for the come-back window and compute whether it has passed in code.

Written by Premsanth Rajamani

Premsanth Rajamani leads marketing and growth at Salesgear. An engineer by background, he runs the company's growth engine hands-on, from SEO and content systems to the AI workflows behind them, and writes practical guides on prospecting, outbound strategy, and putting AI to work in real sales and marketing motions.

“Hope this finds you well” won’t.

Open with what they’re actually dealing with. Salesgear surfaces the signal, your rep sends the message, at full pipeline volume.

Help me 3x my meetings