Yesterday, I happened to see a post from Diogo Almeida, who was at OpenAI earlier and is one of the co-inventors of ChatGPT and RLHF. He is now with TypeSafe AI, the company behind a new model called Jev. His post made a strong claim. He said today's LLMs like ChatGPT and Claude are not built for automation. They give text answers meant for humans, and the same question can get different answers each time. He said Jev is different. It gives typed decisions, not paragraphs. It cannot hallucinate, by design. And it runs up to 100 times faster and cheaper than normal LLMs.
I asked him a direct question in the comments. In our line of work, SEO and search engineering, we take decisions like this every single day. For example, which page on a website should we link to from another page. This is not a simple yes or no question. It depends on the full structure of the website, the content on each page, and what makes sense for the reader. I asked him, how would Jev actually make a decision like this? Does it know our entire website? How is it working out the probability behind its answer?
I did not get a clear answer in the comments. Some other people in the thread had the same doubt. So instead of going back and forth in the comments, I decided we should test it ourselves, with our own real data. This report lays out exactly what we did, and what we found.
β
Why we picked internal linking as the test
At FTA Global, internal linking is something our SEO team does every day, following a proper SOP (standard operating procedure). It affects how link value flows through a website, how search engines understand our client's site, and ultimately how well pages rank. Getting this decision wrong is not a small mistake, it can quietly hurt a client's SEO for months.
This also made it a fair, real test. It was not a made up puzzle or a textbook question. It was the exact kind of decision our team makes for clients, day in and day out.
β
How we set up the test
I asked our SEO team to give me their actual internal linking report for two client websites. Not a sample, the real report they use to do their work. This report had the source page, the anchor text, and the correct page it should link to, based on our own team's judgement and SOP.
We also pulled the full sitemap for both websites, so we knew every page that could possibly be a link target.
From this, we built 150 real test cases. Every single case was a real internal linking decision our own team had already made. We then sent the exact same question to four different AI models: Jev from TypeSafe AI, Claude Haiku 4.5 from Anthropic, GPT-5-mini from OpenAI, and Gemini 2.5 Flash from Google. Each model got the same source page, the same list of possible pages to link to, and the same instructions. No model got an easier version of the question.
β
Overall accuracy
We checked each model's answer against what our own SEO team had actually decided, across all 150 cases.
Gemini 2.5 Flash came out on top with 86.7 percent accuracy, followed by GPT-5-mini at 80 percent and Claude Haiku 4.5 at 71.3 percent. Jev, the model built specifically for fast structured decisions, scored 60.7 percent, the lowest of the four.
β
The real story: accuracy by site
This is the part of the test I found most important, and also the most honest finding of the whole exercise. We had data from two different client websites. One site had a simpler structure, fewer pages that looked alike. The other site had a lot of pages with similar topics and overlapping content, things like different loan types or similar financial products, where telling one page apart from another needs real understanding.
On the easier site, Jev did reasonably well, close to 78 percent accuracy. But on the harder site, its accuracy dropped to under 46 percent, nearly half. The other three models also did worse on the harder site, but their drop was much smaller. This tells us something important. Jev is built to give a clean, structured, confident answer every single time, and it never hallucinates in the way a chatbot does. But a confident, structured answer that is wrong is still wrong. When the decision genuinely gets harder, Jev struggled a lot more than the general purpose models did.
β
Cost: where Jev genuinely wins
It would not be fair to only talk about accuracy. On cost, Jev was far ahead of the rest of the field.
Jev cost about 35 times less than Claude Haiku per decision. If you are running this kind of classification at scale, say across thousands of pages, that saving adds up to real money.
β
Speed
Jev was also the fastest model in our test by a clear margin.
Jev answered in under half a second on average. Gemini 2.5 Flash, the most accurate model, was also the slowest, taking well over three seconds per decision on average. This is the classic trade-off, speed and accuracy pulling in opposite directions.
β
Consistency
We also checked whether each model gives the same answer if you ask it the same question multiple times. We picked a sample of 20 cases and ran each model 5 times on each case.
All four models were close here, between 86 and 90 percent. This was not the factor that separated the models in our test.
β
When should you trust Jev with an SEO decision?
I am not writing this to say Jev is a bad product. For certain jobs, high volume, low stakes, simple classification tasks, its speed and low cost make it genuinely useful. That is a real advantage, and it is not small.
But for a task like internal linking, where getting it wrong can quietly cost a client real SEO value, accuracy matters more than speed or cost. In our test, the general purpose LLMs, Claude, GPT and Gemini, were meaningfully more accurate, and importantly, they held up much better on the harder website.
The bigger lesson for anyone reading this, including us at FTA Global, is simple. A new model architecture with a bold claim on LinkedIn is not proof by itself. Before you plug any new AI model into a real workflow for your clients, test it on your own real data first. Do not take our word for it either. Run your own test, on your own website, on your own decisions.
We are sharing the full methodology and the real numbers above so anyone can see exactly what we did and check our work.
Do you wantΒ β¨more traffic?



