Jev the new AI model that "can't hallucinate" got it wrong 4 out of 10 times.

Senthil Kumar Hariram
Updated on
September 23, 2026
|
Reading time -
3 min

Yesterday, I happened to see a post from Diogo Almeida, who was at OpenAI earlier and is one of the co-inventors of ChatGPT and RLHF. He is now with TypeSafe AI, the company behind a new model called Jev. His post made a strong claim. He said today's LLMs like ChatGPT and Claude are not built for automation. They give text answers meant for humans, and the same question can get different answers each time. He said Jev is different. It gives typed decisions, not paragraphs. It cannot hallucinate, by design. And it runs up to 100 times faster and cheaper than normal LLMs.

I asked him a direct question in the comments. In our line of work, SEO and search engineering, we take decisions like this every single day. For example, which page on a website should we link to from another page. This is not a simple yes or no question. It depends on the full structure of the website, the content on each page, and what makes sense for the reader. I asked him, how would Jev actually make a decision like this? Does it know our entire website? How is it working out the probability behind its answer?

I did not get a clear answer in the comments. Some other people in the thread had the same doubt. So instead of going back and forth in the comments, I decided we should test it ourselves, with our own real data. This report lays out exactly what we did, and what we found.

‍

Why we picked internal linking as the test

At FTA Global, internal linking is something our SEO team does every day, following a proper SOP (standard operating procedure). It affects how link value flows through a website, how search engines understand our client's site, and ultimately how well pages rank. Getting this decision wrong is not a small mistake, it can quietly hurt a client's SEO for months.

This also made it a fair, real test. It was not a made up puzzle or a textbook question. It was the exact kind of decision our team makes for clients, day in and day out.

‍

How we set up the test

I asked our SEO team to give me their actual internal linking report for two client websites. Not a sample, the real report they use to do their work. This report had the source page, the anchor text, and the correct page it should link to, based on our own team's judgement and SOP.

We also pulled the full sitemap for both websites, so we knew every page that could possibly be a link target.

From this, we built 150 real test cases. Every single case was a real internal linking decision our own team had already made. We then sent the exact same question to four different AI models: Jev from TypeSafe AI, Claude Haiku 4.5 from Anthropic, GPT-5-mini from OpenAI, and Gemini 2.5 Flash from Google. Each model got the same source page, the same list of possible pages to link to, and the same instructions. No model got an easier version of the question.

‍

Overall accuracy

We checked each model's answer against what our own SEO team had actually decided, across all 150 cases.

Accuracy vs client ground truth — 150 real decisions

Accuracy: Jev 60.7%, Haiku 4.5 71.3%, GPT-5-mini 80%, Gemini 2.5 Flash 86.7%.

Gemini 2.5 Flash came out on top with 86.7 percent accuracy, followed by GPT-5-mini at 80 percent and Claude Haiku 4.5 at 71.3 percent. Jev, the model built specifically for fast structured decisions, scored 60.7 percent, the lowest of the four.

‍

The real story: accuracy by site

This is the part of the test I found most important, and also the most honest finding of the whole exercise. We had data from two different client websites. One site had a simpler structure, fewer pages that looked alike. The other site had a lot of pages with similar topics and overlapping content, things like different loan types or similar financial products, where telling one page apart from another needs real understanding.

Accuracy by client site

Accuracy by site: Jev 78.3% (A) / 45.7% (B); Haiku 4.5 82.6% / 61.7%; GPT-5-mini 95.7% / 66.7%; Gemini 2.5 Flash 98.6% / 76.5%.

On the easier site, Jev did reasonably well, close to 78 percent accuracy. But on the harder site, its accuracy dropped to under 46 percent, nearly half. The other three models also did worse on the harder site, but their drop was much smaller. This tells us something important. Jev is built to give a clean, structured, confident answer every single time, and it never hallucinates in the way a chatbot does. But a confident, structured answer that is wrong is still wrong. When the decision genuinely gets harder, Jev struggled a lot more than the general purpose models did.

‍

Cost: where Jev genuinely wins

It would not be fair to only talk about accuracy. On cost, Jev was far ahead of the rest of the field.

Cost per 1,000 decisions (USD)

Cost per 1000: Jev $0.028, Haiku 4.5 $0.975, GPT-5-mini $0.219, Gemini 2.5 Flash $0.294.

Jev cost about 35 times less than Claude Haiku per decision. If you are running this kind of classification at scale, say across thousands of pages, that saving adds up to real money.

‍

Speed

Jev was also the fastest model in our test by a clear margin.

Average latency per decision (ms)

Latency: Jev 434ms, Haiku 4.5 705ms, GPT-5-mini 1646ms, Gemini 2.5 Flash 3278ms.

Jev answered in under half a second on average. Gemini 2.5 Flash, the most accurate model, was also the slowest, taking well over three seconds per decision on average. This is the classic trade-off, speed and accuracy pulling in opposite directions.

‍

Consistency

We also checked whether each model gives the same answer if you ask it the same question multiple times. We picked a sample of 20 cases and ran each model 5 times on each case.

Jev

88.8%

Haiku 4.5

90.0%

GPT-5-mini

86.5%

Gemini 2.5 Flash

88.8%

All four models were close here, between 86 and 90 percent. This was not the factor that separated the models in our test.

‍

When should you trust Jev with an SEO decision?

I am not writing this to say Jev is a bad product. For certain jobs, high volume, low stakes, simple classification tasks, its speed and low cost make it genuinely useful. That is a real advantage, and it is not small.

But for a task like internal linking, where getting it wrong can quietly cost a client real SEO value, accuracy matters more than speed or cost. In our test, the general purpose LLMs, Claude, GPT and Gemini, were meaningfully more accurate, and importantly, they held up much better on the harder website.

The bigger lesson for anyone reading this, including us at FTA Global, is simple. A new model architecture with a bold claim on LinkedIn is not proof by itself. Before you plug any new AI model into a real workflow for your clients, test it on your own real data first. Do not take our word for it either. Run your own test, on your own website, on your own decisions.

We are sharing the full methodology and the real numbers above so anyone can see exactly what we did and check our work.

Are your internal links pointing to the right pages?
The right link depends on your site structure, page content and the reader’s intent.
Author Bio
Senthil Kumar Hariram
Founder & MD

I’m Senthil Kumar Hariram, Founder and Managing Director of FTA Global (Fast, Tactical, and Accountable), a new-age marketing company I launched in May 2025. With over 15 years of experience in scaling brands and building high-impact teams, my mission is to reinvent the agency model by embedding outcome-driven, AI-augmented growth teams directly into brands. I help businesses build proprietary Marketing Operating Systems that deliver tangible impact. My expertise is rooted in the future of organic growth a discipline I now call Search Engineering.

Table of contents

Do you want 
more traffic?

Hey, I'm from FTA Global. I'm determined to grow a business. My only question is, will it be yours?
Keep Reading
Search Engineering
September 23, 2026

What Is Context Authority in AI Search and Why Does It Matter?

AI gives different answers to different buyers. Here is what Context Authority means, what shapes it, and why it matters for your pipeline.
Search Engineering
September 21, 2026

7 Real Ways to Use Jev for SEO

A new model called Jev makes decisions instead of writing answers. Here are seven SEO jobs it could take off your plate, at a scale no human can match by hand.
Search Engineering
September 17, 2026

The same question, a different answer in every city

I ran a study on how AI recommends brands. It does not see your brand the way you think. For the last year I kept hearing the same promise. Get a score for how visible your brand is in AI. One number. Watch it go up and you are safe. Watch the study breakdown to see how city and language change the brands that AI recommends.
z
z
z

Want to build the future of marketing with us?