🔍 Read the full analysis: Jev In Practice: 24 Ways To Work With AI Decision Models on ThorstenMeyerAI.com
Get smart everyday buys delivered free — and shop member deals
- Fast, free delivery on millions of items
- Access to Prime Big Deal Days deals on October 6–7
- Prime Video, Amazon Music and more included
TL;DR
In a Sept. 29 article, Thorsten Meyer mapped 24 ways to use Jev, a system that answers typed questions about text or JSON so software can act on the results. Meyer said three uses are already running in his publishing operation, while he classifies 12 as strong fits, seven as needing measurement and two as poor fits. The examples and performance figures are Meyer’s own account; independent validation and details for the remaining use cases are not included in the supplied material.
Thorsten Meyer published a guide to 24 uses for Jev, an AI decision model that returns typed answers for software to act on, and said three applications are already running in his publishing operation. His account places 12 other uses in the strong-fit category, seven in a measure-first category and two as poor fits, offering a test for organizations considering automated decisions at high volume.
Meyer describes Jev as a system that receives text or JSON plus a set of typed questions and returns answers such as a yes probability, a classification choice or a score. The output is intended for code to branch on; Jev does not, in his description, write or summarize the underlying material. He says a call with multiple questions takes about 0.3 to 0.9 seconds and costs about $0.04 per million input tokens.
The three live publishing uses Meyer lists are a relevance gate for matching stories to sites, a language check and a fallback topic classifier. He reports scanning 78,889 articles for $2.01 with the language check, finding 1,576 non-English articles and fixing 1,553. For topic classification, he reports 89% agreement with a frontier large language model overall and 97% to 99% agreement when Jev’s confidence was at least 0.8. These figures are from Meyer’s own operation and measurement.
Among the publishing examples he says are strong fits are checking whether a page discloses free products or affiliate links and moderating comments. Other proposals include detecting whether sources contain enough verifiable facts, checking product relevance in roundups and assessing headlines. Meyer marks these as measure-first where a current heuristic’s failure has not been demonstrated. He calls same-event deduplication a poor fit after a canary found no duplicates to correct.
24 use cases for Jev at a glance
Every use case, coloured by how well it fits
Proven in production
1Relevance gate: story and site2Language check3Classifier fallbackPublishing and content
4Thin-source detector5Same-event dedupe6Product fits the roundup7Disclosure present8Headline quality9Comment moderationCommerce and support
10Support-ticket routing11Return-reason coding12Review to feature complaints13Catalogue taxonomy14Order-fraud pre-triageSoftware and AI systems
15LLM guardrail16RAG passage filter17Citation check18Tool and intent routing19Log-line triage20PR risk triageBusiness ops and home
21Inbox triage22Expense categorisation23Lead qualification24Smart-home intent15 of 24 are ready to build or already running
A Test for Automating Small Decisions
The guide’s central point is that low cost and narrow questions alone do not justify automation. Meyer says a use should involve high volume, a narrow decision, errors that are inexpensive or routed for review, and a visibly failing existing heuristic. This last condition asks teams to establish that there is a real problem before adding a model.
That approach could matter to publishers and other teams processing large queues of routine decisions. A confidence threshold can let software handle clear cases while sending uncertain ones to a person or a more capable system. In Meyer’s examples, this is intended to reduce repetitive checks without letting uncertain outputs silently determine publication or moderation outcomes.
The claimed benefits remain specific to the author’s measurements. The supplied material does not provide an independent evaluation, a comparison with other models on the same workloads, or enough detail to assess the figures across different organizations.
How Meyer Proposes Testing Jev
Meyer recommends replaying 300 to 500 past decisions in a shadow test before wiring Jev into a live workflow. He says teams should compare results overall and by confidence band, then review 20 disagreements to determine which answer was right. His proposed gate is to proceed only where the high-confidence band reaches 95% accuracy.
For rollout, Meyer recommends placing the integration behind a separate feature flag that is off by default, testing it on 5% to 10% of units, and then expanding. This makes the workflow’s existing behavior available while the team checks model performance. The examples in the article apply that principle by leaving uncertain cases on their previous path or sending them to human review.
“Jev is the right tool wherever a system needs thousands of small judgements and can hand the unclear ones to something smarter.”
— Thorsten Meyer
Independent Results Remain Unreported
The supplied article material contains Meyer’s reported results, but does not identify an independent evaluator or explain the full methodology behind the measurements. It does not specify the sample size for the 31-topic classification comparison or provide the underlying data, so readers cannot assess how representative the reported agreement rates are.
The material also gives only part of the 24-use-case inventory: it details three live applications and six publishing proposals, then begins a section on commerce and customer operations. The remaining examples, the basis for classifying each use, and the results of further tests are not available here. The reported token cost and response times are not accompanied by a measurement date or a description of pricing conditions.
Measure Before Wider Rollout
Meyer’s proposed next step for each unproven use is to measure the current heuristic against real past decisions, review disagreements and run Jev in shadow mode. Teams following his approach would then enable a feature-flagged canary only after the high-confidence results meet the stated accuracy threshold, with uncertain cases routed for review.
The supplied material does not announce a release schedule, additional deployment milestones or an external review of Jev. Whether the measure-first proposals advance will depend on whether testing finds a meaningful error in the existing process.
Key Questions
What is Jev?
Jev is an AI decision model that Meyer says accepts text or JSON and typed questions, then returns answers such as probabilities, categories or scores for software to use.
How many of the 24 uses are already live?
Meyer says three uses are live in his publishing operation. He classifies 12 as strong fits, seven as needing measurement and two as poor fits.
What does Meyer recommend testing before deployment?
He recommends replaying 300 to 500 real past decisions, comparing results by confidence band and reviewing 20 disagreements. His stated deployment threshold is at least 95% accuracy in the high-confidence band.
Are the performance figures independently verified?
No independent verification is identified in the supplied article material. The agreement rates, article counts and costs are reported by Meyer from his own operation.
Source: ThorstenMeyerAI.com
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
