October 5, 2026English

How to Classify Thousands of Marketplace Listings a Day with Jev (and When You Still Need an LLM)

Putting TypeSafe's Jev decision model in front of an OpenAI model to classify second-hand marketplace listings could cut classification costs by roughly 70–90%. Here's the architecture, the cost model, the thresholds and Jev's limits.

José Antonio García
José Antonio García
CTO · Founding Partner
Leer en español
aijevautomationclassificationllm-costs

Jev, released by TypeSafe AI on September 15, 2026, is a decision model: it reads text like a large language model, but instead of writing a reply it picks one answer from options you define and returns a probability for each. Used in a cascade in front of an OpenAI model — Jev classifies the listings it is confident about, and only the uncertain ones go to GPT-5.4 mini — we estimate the cost of classifying 10,000 second-hand marketplace listings drops from about $19.70 to about $4.30, a reduction of roughly 78%, assuming one in five listings is escalated. This article explains the architecture, the cost model behind those numbers, how to set the confidence thresholds, and where Jev falls short.

What Jev is (and what it isn't)

Most AI pipelines today use an LLM for everything, including tasks that are really multiple-choice questions: Which category is this? Is this spam? Is this urgent? The LLM answers by writing, token by token, even when the only acceptable answer is one word from a fixed list.

Jev takes a different approach. TypeSafe calls it a "System One" model, after Daniel Kahneman's fast, intuitive mode of thinking. You send it two things:

  • State — the text or data to evaluate (here, a listing's title, description and price).
  • Questions — one or more named questions, each with the exact options it may choose from.

For each question it returns the chosen option, a confidence score and a probability for every option. Several questions about the same state are answered in a single call, and the state is billed once per request rather than once per question.

What Jev cannot do matters just as much: it doesn't write text, extract free-form values, summarize or reason out loud. It can only choose between the options you give it. That's the whole design — and it's why it belongs in front of an LLM, not instead of one.

The problem: classifying second-hand listings at volume

Second-hand marketplaces are messy data. Sellers write titles like "iphone 12 casi nuevo leer descripción", put furniture in the electronics category, and paste the same listing five times. For a typical use case built on this data — price monitoring, resale sourcing, market analysis — each listing needs four answers:

QuestionTypeOptions
Product categoryChoicePhones · Laptops · Furniture · Fashion · … · Other
Item conditionChoiceNew · Like new · Good · Worn · For parts
Is it a real item for sale?Yes/NoExcludes "looking for", swaps, ads and spam
Is the price plausible for the category?ChoiceToo low · Plausible · Too high

Every one of those is a bounded decision. Paying a generative model to "think in prose" about each of them, tens of thousands of times a day, is what makes an LLM-only pipeline expensive.

The architecture: a four-layer cascade

The LLM isn't replaced. It becomes the exception path.

  1. Deterministic code first. Duplicates, empty descriptions and obviously invalid prices are filtered with plain rules before any model is called. It's the cheapest layer and should catch everything it can.
  2. Jev for the bounded decisions. All four questions go in a single request per listing. If every answer clears its confidence threshold, the listing is classified and stored. No LLM involved.
  3. GPT-5.4 mini for the doubtful cases. Listings where any answer falls below threshold are escalated to the LLM, which also receives Jev's probabilities as context.
  4. Human review on a sample. A small random sample of Jev-only decisions, plus every case where Jev and the LLM disagree, is reviewed by a person each week. That's how you know whether the thresholds are still right.

A simplified Jev request looks like this (field names follow the System One API format at the time of writing — check TypeSafe's documentation for the current schema):

{
  "model": "jev-1.13.0",
  "state": {
    "title": "iphone 12 casi nuevo leer descripción",
    "description": "Batería al 89%, sin golpes, con caja. Solo entrega en mano.",
    "price_eur": 320
  },
  "questions": {
    "category": {
      "type": "choice",
      "instructions": "Which product category does this listing belong to?",
      "criteria": {
        "phones": "Mobile phones and smartphones",
        "laptops": "Laptops and notebooks",
        "other": "Anything not covered by the other categories"
      }
    },
    "condition": {
      "type": "choice",
      "instructions": "What condition is the item in, based on the seller's description?",
      "criteria": {
        "new": "Unused, sealed or with tags",
        "like_new": "Used but with no visible wear",
        "good": "Normal signs of use",
        "worn": "Clear wear or minor defects",
        "for_parts": "Not working or sold for parts"
      }
    }
  }
}

The response gives a choice, a confidence and a probabilities map per question. The routing code is a few lines: if all confidences clear their thresholds, store; otherwise, escalate.

The cost model

All prices are published list prices in USD as of September 2026: Jev at $0.042 per million input tokens with output free; GPT-5.4 mini at $0.75 per million input tokens and $4.50 per million output tokens; GPT-5.4 nano at $0.20 and $1.25.

Assumptions per listing:

  • Jev: about 650 input tokens (base request overhead, the listing text and four questions). Output is free.
  • LLM: about 950 input tokens (instructions, output schema and the listing) and about 280 output tokens (a structured answer plus short reasoning). Escalated listings add about 60 tokens for Jev's probabilities.
  • Escalation rate: 20% as the central case, with 10% and 30% shown for sensitivity. The real rate depends on the thresholds and is the number a pilot must measure.

Estimated cost per 10,000 listings:

SetupLLM onlyCascade, 10% escalatedCascade, 20% escalatedCascade, 30% escalated
Jev + GPT-5.4 mini$19.73$2.29$4.31$6.33
Jev + GPT-5.4 nano$5.40$0.83$1.38$1.93

Two things stand out. First, Jev's own share is tiny — about $0.27 per 10,000 listings — so the cascade's cost is almost entirely the escalated LLM calls. The escalation rate is what you're really optimizing. Second, the savings hold even against the cheapest OpenAI model: the cascade still cuts cost by roughly 65–85% against GPT-5.4 nano, because LLM output tokens are what make each call expensive, and Jev doesn't bill for output.

At 10,000 listings a day, the GPT-5.4 mini setup goes from about $590 a month to about $130 at a 20% escalation rate. The absolute numbers are modest at this volume; the ratio is what matters as volume grows, or when the same pattern is applied to a more expensive model.

What this model doesn't cover: accuracy. Cost is easy to estimate from price lists; whether Jev's answers are as good as the LLM's on this specific data isn't, and that's the second thing a pilot has to measure.

Setting the thresholds

The threshold is the single most important number in this design. Set it too low and Jev's mistakes go straight into your data; set it too high and almost everything goes to the LLM, erasing the savings.

How to set them:

  • Pin the model version (jev-1.13.0, not jev-latest). Thresholds tuned on one version aren't guaranteed to hold on the next.
  • Shadow-test first. For a week or two, send every listing to both Jev and the LLM and compare answers before Jev makes any decision on its own.
  • Use one threshold per question. Category is usually easier than condition, because sellers describe condition inconsistently, so the two shouldn't share a threshold.
  • Track the escalation rate continuously. It's the main cost driver, and a sudden jump usually means the incoming data has changed.

Where Jev falls short

Jev isn't a drop-in answer for everything:

  • It can't extract. Brand, model and storage capacity ("iPhone 12, 128 GB") are free-form values, not choices. Those still need regex or an LLM.
  • Sellers can argue with it. Listing text is user-generated, and text that argues for a particular classification can move the answer — TypeSafe's own documentation lists this as a known limitation. Anything with real consequences, such as flagging fraud, shouldn't rely on Jev alone.
  • "Other" needs care. Jev can only pick from the options you give it, so a missing category silently becomes the closest wrong one. An explicit "other" option with a clear description is essential.
  • Rate limits. At launch TypeSafe published a limit of 1,200 requests per minute during early access — plenty for tens of thousands of listings a day, but worth checking for larger batch jobs.

When to use Jev, an LLM, or neither

If the task is…Use
A rule you can write down (price is zero, title is empty)Plain code
Picking from a fixed set of options, at volumeJev
Writing, extracting free-form values, or reasoning through an unusual caseAn LLM
A decision with real consequences for a person or the businessA human, with the models as support

The pattern generalizes well beyond marketplaces. The same cascade fits support-ticket routing, invoice and expense categorization, lead qualification and content moderation — any process where a business pays an LLM to answer what is really a multiple-choice question thousands of times a month.

At ScaleWave we design AI pipelines where each step uses the cheapest tool that does the job well — code, a decision model, an LLM or a person. If your team is running large volumes of repetitive classification through an LLM, we'd be happy to model what a cascade like this would save you.

José Antonio García
About the author
José Antonio García

Senior software engineer and Associate Professor at IE Business School. Writes from inside the operation, not from the slide deck.

Ready to transform your business?

A 30-minute conversation is enough to know if we're a match. No cost, no commitment, no sales templates.

Who'll be on the call — no BDRs, no scripts
Gustavo Maryssael
Gustavo Maryssael
CEO · Founding Partner
José Antonio García
José Antonio García
CTO · Founding Partner