TypeSafe · System One · Jev

Jev Lab

A benchmark for TypeSafe's System One model, on its own: how fast it turns support-ticket text into a typed decision, and how often that decision is right. Every ticket is scored against a hand-written answer key the model never sees, and a keyword-rule baseline runs on the same text for contrast. The sorting simulation is on its own page.

Live

Waiting on Jev
0
couriers held at the top
In flight
0
open requests
Decisions / s
0.0
last 5 seconds
Avg latency
—
per request
Sorted
0
 
Judgments
0
3 per ticket, one round trip
Needs a human
0
flagged by the model
Couriers take the straight line to the nearest package, pick it up and carry it to the bin. The package stays grey with a ? until Jev answers — raise the game speed past what the round trip allows and they bunch up waiting, because that pause is the model's real latency and no slider shortens it.

Each ticket gets three judgments at once: its category, an urgency score, and whether it needs a human. They are independent questions over the same text, so they are answered in parallel inside one request — a batch of eight tickets is 24 typed judgments in a single round trip. The parcel widens with urgency, and an amber ◆ marks a ticket the model says needs a person.

Only the category has an answer key, so only the category is scored. Urgency and escalation are shown rather than graded — there is no hand-written truth for them, and claiming accuracy would be inventing it. Courier movement and spacing are ordinary local code, not the model.

Routing board — type, team and product in one pass

All three right
—
type, team and product
Type
—
3 options
Team
—
5 options
Product
—
15, plus "no single product"
Judgments
—
 
Wall clock
—
for the whole desk
Unsorted inbox 30 waiting
Incident Service Req Problem routed to the wrong team
Customer tickets for an office-supplies company. Each one is three choice questions over the same text — which kind of ticket it is, which team owns it, and which of fifteen products it concerns — so a batch of eight is 24 judgments in one round trip and the whole desk is sorted in a handful of requests.

Everything starts in the unsorted inbox — no type, no team, no product, because none of that is known yet. Past thirty the corpus repeats and each copy is labelled, since identical text arriving twice is precisely where a cache would show itself. Tickets leave the pile as the answers land, which is the part worth watching: the desk sorts itself. A ticket then sits in the column the model chose. Group by switches the board between team, product and type, so the same routed pile can be broken down three ways without asking again. A red border means the key disagrees on whichever axis you are grouping by; hover it to see where it belonged. "No single product" is a real answer, not a gap — plenty of tickets are about a contract or a delivery window rather than a thing in the catalogue, and a classifier without that option will invent one.

Race — Jev against the keyword rules

Keyword rules —
TypeSafe Jev —
matched the key missed it not decided yet
Both lanes get the same tickets in the same order. The rules finish before you can see them move, which is the honest part of the comparison — and then the red squares are the price of that. Hover any square to read the ticket, what that lane called it, and what the key says. Past a few hundred tickets the corpus repeats, so accuracy cannot move — only the timings do — and the token bill is real: roughly 300 tokens a ticket, so a full 800 run costs around a quarter of a million.

Duplicates — the same fault in different words

Pairs judged
—
every ticket against every other
Requests
—
 
Wall clock
—
to judge them all
Pairs correct
—
against the grouping key
Real groups of three, mixed with unrelated singletons. Each pair is one noul — "are these the same underlying problem?" — so the whole grouping is settled by asking about every pair. The tickets share almost no words with their own group, so keyword matching cannot find these; the overlap is in meaning. Which tickets belong together is written down by hand, so the result is checked rather than admired.

Pairs grow with the square of the tickets — 12 tickets is 66 questions, 36 is 630, and 248 would be over thirty thousand. That is why this stops at 36: beyond it, comparing everything to everything is the wrong shape for the problem, and real deduplication compares against one representative per group instead.

The belt — twenty bins, and only so many hands

Into a bin
0
 
Off the end
0
nobody was free
Right bin
—
against the answer key
Decision
—
round trip, per parcel
Cleared / s
0.0
last 5 seconds
Sorters busy
0
 
Delivered to a site
0
by truck
Right site
—
against the answer key
On the dock
0
waiting for a truck
Press start. Turn the rate up until parcels begin running off the end.
What is being decided. Every parcel carries a delivery note that never names its own bin — no note says "toner" or "chair". A sorter that picks one up sends the note as a single choice over twenty bins, and puts the parcel where the answer says. Parcels picked up within the same breath are gathered into one request — one question each, answered in parallel, so a busy line costs fewer round trips than it has parcels. A sorter is a filled circle while their hands are free, hollow while they are busy, and ringed yellow while they wait for the answer. A keyword rule has nothing to match on here, and "a 14-inch notebook computer" has to land in laptops rather than notebooks.

Where the break point is. A sorter is busy for one round trip plus the walk to the bin, so a line of n sorters clears roughly n ÷ that parcels a second. Push the arrival rate past it and parcels start passing everybody and falling off the end — the counter says how many. That ratio, not the belt, is what decides how many people the line needs. Everything that passes the last pair of hands lands in the crate at the end, where it piles up instead of disappearing — click it to read what was never even looked at.

Then it has to go somewhere. Sorted is not shipped. Every parcel also carries a shipping note — "leave it with the crane crew on the cold store side" — and which of thirty sites that is, is a second choice, asked the same way and batched the same way. The two notes are drawn independently, so the same box of batteries can be bound for the observatory or the ferry: the bin gives the site away no more than the site gives away the bin. Trucks wait for a load of four, or leave part-full rather than strand a quiet site, and a truck is gone for its whole round trip — so too few of them back the dock up while the belt is still keeping up perfectly. Stopping the trucks stops the second decision too — nothing is asked about a parcel nobody is going to drive anywhere — and the dock simply grows until they start again. Click a site to read the notes driven to it.

Bins are neutral on purpose. Twenty distinguishable colours do not exist, so a parcel is marked green or red for whether it landed in the bin its answer key names — never for which bin it is. Click any bin to read what went into it.

Run

Ready. The whole 48-ticket corpus is classified on each run, uncached.

Results

Accuracy vs key
—
 
Median latency
—
per request
95th pct latency
—
per request
Throughput
—
tickets / second
Decision time
—
per ticket, amortised
Tokens / ticket
—
in + out
Judgments
—
typed answers

Latency per request

Run a benchmark to populate.

Where it was wrong

Incident Service Req Problem
Rows are the answer key, columns what the model chose. Off-diagonal cells are the mistakes.

Decisions

TicketKeyChoseProbabilitiesConfUrgencyHuman?
No run yet.