A benchmark for TypeSafe's System One model, on its own: how fast it turns
support-ticket text into a typed decision, and how often that decision is
right. Every ticket is scored against a hand-written answer key the model
never sees, and a keyword-rule baseline runs on the same text for contrast.
The sorting simulation is on its own page.
Live
Waiting on Jev
0
couriers held at the top
In flight
0
open requests
Decisions / s
0.0
last 5 seconds
Avg latency
—
per request
Sorted
0
Judgments
0
3 per ticket, one round trip
Needs a human
0
flagged by the model
Bin
Ticket
Key
Urgency
Human?
Couriers take the straight line to the nearest package, pick it up and carry it
to the bin. The package stays grey with a ? until Jev answers —
raise the game speed past what the round trip allows and they bunch up waiting,
because that pause is the model's real latency and no slider shortens it.
Each ticket gets three judgments at once: its category, an
urgency score, and whether it needs a human. They are independent questions over
the same text, so they are answered in parallel inside one request — a batch of
eight tickets is 24 typed judgments in a single round trip.
The parcel widens with urgency, and an amber ◆ marks a ticket the model says
needs a person.
Only the category has an answer key, so only the category is scored. Urgency and
escalation are shown rather than graded — there is no hand-written truth for
them, and claiming accuracy would be inventing it. Courier movement and spacing
are ordinary local code, not the model.
Routing board — type, team and product in one pass
All three right
—
type, team and product
Type
—
3 options
Team
—
5 options
Product
—
15, plus "no single product"
Judgments
—
Wall clock
—
for the whole desk
Unsorted inbox30 waiting
Evidence this ran live
IncidentService ReqProblemrouted to the wrong team
Customer tickets for an office-supplies company. Each one is three
choice questions over the same text — which kind of ticket it
is, which team owns it, and which of fifteen products it concerns — so a batch
of eight is 24 judgments in one round trip and the whole desk is sorted in a
handful of requests.
Everything starts in the unsorted inbox — no type, no team,
no product, because none of that is known yet. Past thirty the corpus repeats and
each copy is labelled, since identical text arriving twice is precisely where a
cache would show itself. Tickets leave the pile as the
answers land, which is the part worth watching: the desk sorts itself. A ticket
then sits in the column the model chose. Group by switches the
board between team, product and type, so the same routed pile can be broken
down three ways without asking again. A red border means the key disagrees on
whichever axis you are grouping by; hover it to see where it belonged. "No single
product" is a real answer, not a gap — plenty of tickets are about a
contract or a delivery window rather than a thing in the catalogue, and a
classifier without that option will invent one.
Race — Jev against the keyword rules
Keyword rules—
TypeSafe Jev—
matched the keymissed itnot decided yet
Both lanes get the same tickets in the same order. The rules finish before you
can see them move, which is the honest part of the comparison — and then the
red squares are the price of that. Hover any square to read the
ticket, what that lane called it, and what the key says.
Past a few hundred tickets the corpus repeats, so accuracy cannot move — only
the timings do — and the token bill is real: roughly 300 tokens a ticket, so a
full 800 run costs around a quarter of a million.
Duplicates — the same fault in different words
Pairs judged
—
every ticket against every other
Requests
—
Wall clock
—
to judge them all
Pairs correct
—
against the grouping key
Real groups of three, mixed with unrelated singletons. Each pair is one
noul — "are these the same underlying problem?" — so the
whole grouping is settled by asking about every pair. The tickets share
almost no words with their own group, so keyword matching cannot find these;
the overlap is in meaning. Which tickets belong together is written down by
hand, so the result is checked rather than admired.
Pairs grow with the square of the tickets — 12 tickets is 66
questions, 36 is 630, and 248 would be over thirty thousand. That is why this
stops at 36: beyond it, comparing everything to everything is the wrong shape
for the problem, and real deduplication compares against one representative
per group instead.
The belt — twenty bins, and only so many hands
Into a bin
0
Off the end
0
nobody was free
Right bin
—
against the answer key
Decision
—
round trip, per parcel
Cleared / s
0.0
last 5 seconds
Sorters busy
0
Delivered to a site
0
by truck
Right site
—
against the answer key
On the dock
0
waiting for a truck
Press start. Turn the rate up until parcels begin running off the end.
Bin
Against the key
Delivery note
What is being decided. Every parcel carries a delivery note
that never names its own bin — no note says "toner" or "chair". A sorter that
picks one up sends the note as a single choice over twenty
bins, and puts the parcel where the answer says. Parcels picked up within the
same breath are gathered into one request — one question each, answered in
parallel, so a busy line costs fewer round trips than it has parcels. A
sorter is a filled circle while their hands are free, hollow while they are
busy, and ringed yellow while they wait for the answer. A keyword rule has nothing to match on here, and
"a 14-inch notebook computer" has to land in laptops rather than notebooks.
Where the break point is. A sorter is busy for one round trip
plus the walk to the bin, so a line of n sorters clears roughly
n ÷ that parcels a second. Push the arrival rate past it and parcels
start passing everybody and falling off the end — the counter says how many.
That ratio, not the belt, is what decides how many people the line needs.
Everything that passes the last pair of hands lands in the crate at the end,
where it piles up instead of disappearing — click it to read what was never
even looked at.
Then it has to go somewhere. Sorted is not shipped. Every
parcel also carries a shipping note — "leave it with the crane crew on the
cold store side" — and which of thirty sites that is, is a
second choice, asked the same way and batched the same way. The two notes are
drawn independently, so the same box of batteries can be bound for the
observatory or the ferry: the bin gives the site away no more than the site
gives away the bin. Trucks wait for a load of four, or leave part-full rather
than strand a quiet site, and a truck is gone for its whole round trip — so
too few of them back the dock up while the belt is still keeping up
perfectly. Stopping the trucks stops the second decision too — nothing is
asked about a parcel nobody is going to drive anywhere — and the dock simply
grows until they start again. Click a site to read the notes driven to it.
Bins are neutral on purpose. Twenty distinguishable colours do
not exist, so a parcel is marked green or red for whether it landed in the bin
its answer key names — never for which bin it is. Click any bin to read what
went into it.
Run
Ready. The whole 48-ticket corpus is classified on each run, uncached.
Results
Accuracy vs key
—
Median latency
—
per request
95th pct latency
—
per request
Throughput
—
tickets / second
Decision time
—
per ticket, amortised
Tokens / ticket
—
in + out
Judgments
—
typed answers
Latency per request
Run a benchmark to populate.
Where it was wrong
IncidentService ReqProblem
Rows are the answer key, columns what the model chose. Off-diagonal cells are the mistakes.