Models & Comparisons

Comparing AI Models

Last updated 2026-08-01

What's new

2026-08-01
  • DeepSeek version 4 Flash (a new, affordable AI model) now ranks 10th in overall performance, beating other models like Opus 4.7 and Claude Sonnet 5 due to its cost-efficiency and improved ability to plan and use tools (called "agentic behavior").
  • This model is open-source under the MIT license (meaning anyone can use or modify it for free, even for commercial purposes), and it's one of the top three open-weight models in terms of intelligence.
  • DeepSeek 4 Flash offers near Luna-level intelligence (a high-performing AI model) at about 60% lower cost per task, making it a great value for performance.
  • The model has shown significant improvements in front-end development tasks, such as creating landing pages and generating 3D product images, as well as cloning complex interfaces like Mac OS.
2026-07-31
  • A new AI model called Kimmy K3 (a type of AI software that can understand and generate text, images, and more) was released for free by a Beijing lab, and it quickly became popular, even being praised by American companies.
  • The US government accused the creators of Kimmy K3 of using a technique called "distillation" (copying the outputs of a stronger AI model to improve a weaker one) to steal American AI technology, and they threatened sanctions (penalties that limit business) if it's true.
  • China denied these accusations, and the creator of Kimmy K3 said the model's improvements came from original changes, not copying.
  • A new tool called Dream Nina (a website that helps create AI-generated videos) was introduced, which allows users to guide video creation using images, videos, audio, and text references together in one place.
2026-07-28
  • You can now test new AI models (like Claude or Codex, which are AI tools that help with tasks) against your own work, not just generic benchmarks (tests that don't relate to your specific needs).
  • AI can scan your past conversations to find common tasks, then test new models on those tasks using a rubric (a set of criteria you create) that matters to you.
  • You can compare different AI models or effort levels with a simple command (like "/benchmark") and get a report showing which model performs best for your specific needs.
  • This approach helps you decide if a new AI model is worth using for your work, saving time and avoiding irrelevant tutorials.
2026-07-25
  • Poolside AI released Lagona S2.1, a powerful AI model (a computer program that can learn and make decisions) that can run on a single high-end computer and is great for coding tasks.
  • Despite having fewer active parameters (a measure of a model's complexity), Lagona S2.1 competes with much larger models and can run locally on consumer-grade hardware.
  • In coding tests, Lagona S2.1 performed nearly as well as larger models but was faster and could run on a 128 GB MacBook.
  • You can try Lagona S2.1 for free on the World of AI benchmark tool (an online platform for testing AI models) and compare its performance with other models.
2026-07-22
  • A new AI model called Kimmy K3 (a type of AI software made by a company in China) was released for free, and it's as good as leading models like ChatGPT and Claude (popular AI chatbots).
  • Kimmy K3 is open-source (AI software that anyone can use and modify), unlike closed-source models controlled by companies like OpenAI and Anthropic (AI companies that make ChatGPT and Claude).
  • China is giving away advanced AI models for free to gain influence, control tech standards, and compete with U.S. companies, a strategy called "scorched earth" (a business tactic to undercut competitors by offering similar or better products for free or at a lower cost).
  • To save money on AI-powered automations (repetitive tasks done by software), use tools like Zapier (a service that connects different apps to automate tasks) for non-AI tasks and switch AI models as better ones become available.
2026-07-16
  • AI can get overwhelmed and give worse answers when processing too many files at once, due to a limit on how much it can "hold in its head" (called a context window).
  • To fix this, create a "table of contents" (a simple map file) and summaries of your files, so AI can quickly find what it needs without reading everything.
  • This three-layer system (map file, summaries, and source files) helps AI work better with large amounts of data, like hundreds or thousands of files.
2026-07-13
  • Abacus is creating a system (called "agents") that automatically picks the best AI model (like GPT-5.6 or Fable 5) for each part of a task, making it more efficient than using just one model.
  • This system can handle complex coding tasks, like inspecting code for security issues, explaining code structure, and even generating scaling plans for databases and other systems.
  • The system can also be used to build and deploy applications, like a 3D castle in a browser or a trading strategy platform, all from a single conversation.
  • The system is built on top of a supercomputer that can run 24/7, host databases, and connect to other services like AWS (a cloud computing service) and GitHub (a platform for software development).
2026-07-10
  • Hermes Agent (a powerful AI tool that can act like a full-time employee) works best with the Opus model (a specific AI model that's very reliable but expensive), but ChatGPT (a popular AI chat service) and GLM 5.2 (a cheaper AI model) are also options.
  • To avoid downtime, run at least two Hermes agents simultaneously, using different AI models or accounts, so they can monitor and fix each other if one fails.
  • You can create new Hermes agents (called "profiles") either by asking an existing agent to set one up for you or by using the Hermes dashboard.
  • If you're running a serious business, consider investing in the Opus model for Hermes Agent, as it's the most reliable for completing tasks.
2026-07-07
  • Google DeepMind is reportedly launching Gemini 3.5 Pro, a new AI model (a computer program that predicts text) on July 17th, built on a fresh base model with a massive 2 million token context window (the amount of text it can process at once).
  • Gemini 3.5 Pro is rumored to have a specialized deep think reasoning layer for better logic, math, and multi-step reasoning tasks, and support for advanced autonomous agentic workflows (AI that can perform tasks without human input).
  • Leaked outputs show Gemini 3.5 Pro excelling in creating detailed SVG (a type of image file) graphics and generating complex code, like a 3D Subway Surfers style game in around 800 lines of HTML (a coding language for websites).
  • Google is also testing other unreleased Gemini models, including Gemini 3.5 Flash High and a mysterious checkpoint labeled as 3 Flash, which could be Gemini 3.6 or even Gemini 4 Flash.
2026-07-04
  • OpenAI's Mark Chen believes AI is advancing rapidly, with AI models soon doing self-sustaining research, pushing science forward with less human control (AGI, or artificial general intelligence, means AI that can understand, learn, and apply knowledge like a human).
  • AI is already showing signs of "divine moves" (unexpected, innovative solutions) in fields like math and computer science, and AI agents are starting to do meaningful work in their own fields.
  • OpenAI is working towards a future where AI can conduct end-to-end research, from idea to result, with humans acting as orchestrators (managing and guiding the AI's work).
  • Challenges include evaluation (making sure AI is actually improving) and the "jagged frontier" (AI excelling at complex tasks but struggling with simple ones), with continual learning (AI carrying lessons from one task to the next) being a key area for improvement.
2026-07-01
  • AI tools like ChatGPT and Claude (popular AI chatbots) have settings that control how much effort they put into answering, which can make a big difference in the quality of their responses.
  • You can choose different "models" (versions of the AI) and "reasoning levels" (how hard the AI tries) to match the task, like a small model with low effort for quick tasks or a large model with max effort for complex ones.
  • For most people, the default settings are too basic for tasks like writing emails or summarizing reports, so adjusting these settings can help you get better results without changing your prompts (the instructions you give the AI).
  • AI providers set low defaults to make the AI respond quickly and to save them money, but you can change these settings to suit your needs, like using a higher reasoning level for tasks that require more thought.

Key points

What it is

  • Comparing AI models means testing different AI systems to find the best one for your specific task.
  • AI models vary in speed, cost, and quality, so one model might be great for writing code but terrible for analyzing customer comments.
  • AI models are becoming cheaper and more accessible, but their intelligence isn't one-size-fits-all.
  • You need to match the right model to your unique processes, decisions, and historical context.

How to use it

  • Start by picking the "closest model" (the one you're already using or have recently used) and go deep on that single tool first.
  • Use a benchmark tool like the free World of AI Benchmark to test models across multiple categories and prompts.
  • Define how you will measure success (called an evaluation) before comparing models, and always start with the biggest, most capable model.
  • Break your task into smaller steps and consider which tool is best for each specific output.

Watch out for

  • Developers often overestimate AI’s benefits, so treat AI output like code from a junior developer—review it carefully.
  • Models are shaped by their training data, so gaps create blind spots, and different tasks need different models.
  • Don't assume one model works for everything, and don't tell the AI it's an expert—this can make it overconfident without improving accuracy.
  • Consider the cost of a wrong answer, not just the price per call, and test the model, your prompt harness, and the overall system.

Tools named

  • World of AI Benchmark (free tool for testing models across multiple categories and prompts)

Lesson 1: What is Comparing AI Models and why it matters

Comparing AI models means evaluating different AI systems to see which one performs best for your specific task. Developers do this because AI models vary dramatically in speed, cost, and quality. A model that excels at writing code might be terrible at analyzing customer comments. Testing multiple models side by side helps you match the right tool to the right job.

Why does this matter? First, AI models are becoming cheaper and more accessible, but intelligence itself is not commoditized. What matters is your unique processes, decisions, and historical context. You need to collate that information and plug it into the right model with the right framework. Second, after hundreds of hours testing leading tools like Claude Code and Google Antigravity, users found major differences in performance. Choosing the wrong model wastes time and money.

Be aware that developers often overestimate AI’s benefits. One rigorous trial found experienced developers using AI tools took 19% longer to complete tasks, even though they thought they were 24% faster. Treat AI output like code from a junior developer—review it carefully and never assume it’s correct. AI accelerates, but humans validate. The combination is powerful; either alone is incomplete. By comparing models critically, you avoid blind spots and build reliable, profitable AI systems.

Sources

Lesson 2: How to use Comparing AI Models: step-by-step

To compare AI models, start by picking the "closest model" (the one already open in your tab or that you’ve used recently). Go deep on that single tool first — solve your problem fully before testing others. If your chosen model solves the task, ship it and move on; don’t get distracted.

When you need to compare models, use a benchmark tool like the free World of AI Benchmark, which tests models across multiple categories and prompts. Before any comparison, define how you will measure success — this is called an evaluation (how you measure and track if a model performs well). Always start with the biggest, most capable model to find the "quality ceiling" (the best possible performance). Then you can downshift to cheaper models, but only if your evaluations prove the cheaper model preserves the same behavior.

Break your task into baby steps and consider which tool is best for each specific output. For example, simple sorting and serious judgment should not use the same model. The trade-off has three knobs: capability, cost, and confidence. Confidence comes from your evals, not from liking a smaller invoice. You are testing three things: the model itself, your harness (the recipe and chef structure that runs steps), and your coding harness. A strong model can sometimes overcome a weak harness, but you need both for reliable results.

Sources

Lesson 3: Best practices and pitfalls

When comparing AI models, a common mistake is picking the cheapest or most popular one first. Instead, start with the most capable model to find the quality ceiling — the best possible output for your task. Only then should you downshift to a cheaper model, but only if your eval set (a collection of test examples) proves the cheaper model preserves the same behavior. This avoids wasting money on a model that can't do the job.

Pitfalls include trusting a model that sounds confident but is wrong — always double-check. Models are shaped by their training data, so gaps create blind spots. Don't assume one model works for everything; different tasks need different models. Another mistake is telling the AI it's an expert — this can backfire by making it overconfident without improving accuracy.

Best practices: use multiple models in parallel to compare responses. If they disagree, human judgment becomes critical. Also, think about the cost of a wrong answer, not just the price per call. Is the decision reversible? Low-stakes tasks can use cheaper models; high-risk ones need the best. Finally, test three things: the model itself, your prompt harness, and the overall system. Confidence comes from evals, not from liking a cheaper invoice. Simple sorting and serious judgment should not sit at the same desk.

Sources