Fifty-six sections. The same amount to review.

I wanted JEV to give me less video to review. I sent it 56 sections of this recording, with the surrounding context, to see whether it could rule out material before the full review.

Thirty requests came back with answers. Twenty-six hit overload errors. None of the suggestions to discard a section met the confidence threshold set before the test.

All 56 sections still needed review. For this experiment, JEV didn't reduce the work.

That result came after the recording. In the video, I'm still trying to understand what JEV is and where I could use it. I ask ChatGPT for simpler examples, argue with Claude over a payment method, and ask Qwen to look into the claims.

I left my questions in because they were the point. I hadn't figured this out before opening the tools.

The example that made it click

I first needed an explanation I could connect to an ordinary job. The example was a customer saying their shoes arrived in the wrong size.

The software already has possible destinations, such as returns, shipping and billing. JEV can choose among those options. The surrounding software then decides what happens to the message.

TypeSafe's documentation describes JEV as a model for structured decisions. Its questions can ask for a choice, a score or a probability for a statement. That helped me understand why I wouldn't use it like another chat window.

I then tried to apply that idea to a list of businesses. Could I find the information, give JEV my criteria, and leave the difficult cases to a bigger model?

That was a possible workflow I was asking about. I didn't show it producing a finished list of qualified leads. The video doesn't establish that it would save money on that job.

The distinction between cleaning and classifying the list also took me a while. Fixing a phone number or removing a duplicate doesn't answer whether a business is a good prospect. I was asking why I'd need another pass through data that had already been cleaned.

The actual ChatGPT explanation separates cleaning the data from classifying the business.

If I tried the lead-list idea next, I'd keep the question narrow. I'd collect the source information, define the business criteria, and inspect the resulting choices before contacting anyone. A probability would give me another thing to check, not proof that somebody wants a website.

Two answers I had to look past

While I was trying to get JEV working, Claude told me my payment method wasn't on the account. I opened the billing page and showed it the card entry.

Claude then said it had checked the wrong signal. The card was there. That specific exchange explains my reaction in the video; I'm not presenting it as a benchmark of every Claude response.

Later, I asked Qwen about JEV and whether something similar already existed. Its answer led me to Harsha Gundala's Qwen-based project on Hugging Face.

I called it “Chinese JEV” while reacting to Qwen. The page I actually opened is a Qwen-based reproduction. That doesn't establish the author's nationality or that it matches JEV's capabilities.

The Hugging Face repository I opened during the recording.

The model card describes parallel constrained decoding on Apple Silicon and lists an Apache 2.0 license. Harsha's announcement is the source of the two-hour build claim. I didn't reproduce that timing or run an equivalence test.

I also wouldn't turn “can't hallucinate” into “can't be wrong.” TypeSafe's launch post ties its zero figure to guaranteed schema matching and says that figure isn't empirical. Choosing an allowed answer and choosing the correct answer are different things.

What the editing experiment actually did

The job I could test immediately was this video. I wanted to know whether JEV could remove enough review work to be worth adding to my process.

The test used the original transcript sections and neighboring context. The discard threshold was fixed at 90% before the requests ran. Sections that failed, remained uncertain or fell below that threshold stayed in the normal review.

Of the 30 answers, 24 said keep, five suggested discard and one was uncertain. None of those five discard probabilities reached the threshold. I kept the threshold rather than lowering it afterward to produce a saving.

What I checked What happened
Sections submitted 56
Successful answers 30
Overload errors 26
Sections safely removed under the chosen rule 0
Reduction in material needing review 0%

The 30 successful responses recorded $0 in gateway charges during the promotion. The failed requests didn't return billing metadata. That is what the receipts show; it isn't a claim that running this process always costs nothing.

There were also separate normal and JEV-assisted subscription review passes. They didn't perform identical extra work, so their timings and token counts don't give me a controlled cost comparison. I can't turn them into a cash-saving number.

Gemini 3.8 audiovisual observation and the normal editing checks were still needed. JEV didn't choose the final cuts or replace watching the result. It didn’t do the one thing I wanted it for.

What would I measure in another test?

I'd use a fixed batch with answers I had already reviewed, then compare the same job with and without JEV. I'd count mistakes, overloads, remaining review work and the cost of both paths. A fast answer would only help if it reduced work I could safely stop doing.

Does this mean JEV is useless?

No. This was one editing experiment with one rule and a high discard threshold. It doesn't settle whether JEV is useful for a narrower classification job. It does settle what happened when I tried it here.

Where are the tools from the video?

The full recording shows how I got from “what is this?” to a test I could check. My next question is which narrower job would make a fair second attempt.

Watch the full test

Watch the complete JEV breakdown on YouTube for the setup, the Claude exchange and the question I asked Qwen.

For the publishing side, I also wrote up how I use Repurpose.io, Buffer and Codex to get one recording onto several platforms.

JEVTypeSafe AIAI workflows
← All posts