Comparing AI Models
Last updated 2026-09-13What's new
- Google DeepMind may have achieved recursive self-improvement (RSI), where AI agents improve their own abilities, with reports suggesting they're using AI to build smarter AI (agentic loops).
- Moonshot AI's Gemini K 2.8, a new AI model, offers similar performance to K3 but with faster, more efficient thinking, addressing K3's tendency to overthink simple tasks.
- OpenAI's GPT-6 Astra, initially thought to be downgraded, was actually fixed within days, with OpenAI resetting the system for all users despite the issue not affecting everyone.
- Anthropic's CEO, Dario, proposed slowing down AI development for safety, despite believing AI could transform humanity and cure major diseases within 5-10 years.
- Chat Astra (a new AI tool for creating websites and content) and Fable 5.1 (another AI tool) were compared on tasks like website design and writing emails, with both performing well but Astra having a slight edge in design.
- Fable 5.1 (a large language model, or LLM, which is a type of AI that understands and generates human-like text) won in writing a cold email (an unsolicited email sent to a potential client or partner) to secure a 30-minute interview with Sam Olman, a well-known tech entrepreneur.
- Both tools used a similar number of "tokens" (small pieces of text that the AI processes), but Chat Astra was about 20% more efficient, meaning it did more with less.
- The comparison was done without giving the AI tools extra context or instructions, to test their independent abilities.
- **New AI Models**: Several companies, like OpenAI (GPT6 Astra), Anthropic (Claude Fable 5.1), and Google (Gemini 3.8 Flash), released powerful new AI models (computer programs designed to perform tasks) this week.
- **Interactive Video Games**: Tools like H3 World and Solar WM (both open-source, meaning free to use and modify) turn video generators (like Miniax H3) into interactive video game engines (software that creates and controls video games), allowing users to explore AI-generated worlds in real-time.
- **Time Series Prediction**: Google's Times FM3 (open-source) is a tiny (330 million parameters, or settings the AI uses) AI model that can predict future data (like stock charts or weather) from various related signals without needing retraining (a process that teaches AI models using data).
- **3D Room Simulation**: ByteDance's Lucida (not yet open-source) is an AI that turns images of messy rooms into editable 3D simulations (digital copies) by identifying and separating objects.
- GPT Astra (a new AI tool) can handle complex tasks like creating 3D games, designing houses, and even doing your taxes, showing impressive progress towards artificial general intelligence (AGI, AI that can perform any intellectual task a human can).
- It can design and build websites, games, and apps directly from simple instructions, automating tasks like tracking new AI video models and creating reports.
- Astra can create detailed 3D environments and games, like a playable first-person shooter or a Mario Kart-style racing game, making advanced game development accessible to everyone.
- It can also adapt documents to specific visual styles and writing voices, making it useful for graphic design tasks that would typically take professionals hours to complete.
- OpenAI may have finished training a massive new AI model called Bell (a type of AI software), which could be even more advanced than their upcoming GPT-6 model and might help them reach their goal of creating artificial general intelligence (AGI, AI that understands and learns like a human) this year.
- Anthropic, a competitor, is reportedly testing two new AI models, Claude Marshmallow (a type of AI software) and Claude Melon (another type of AI software), with Marshmallow showing better performance and a larger context window (the amount of information the AI can process at once).
- Alibaba has released a new AI model called Qwen 3.8 Flash (another type of AI software), and is working on an even more advanced model called Qwen 4 (yet another type of AI software) for the future.
- OpenAI has also introduced Halpionia, their first custom AI chip (a type of computer hardware), which could significantly improve the performance and efficiency of their AI models.
- A new AI tool called **Evoke** (an open-source AI video model) lets you create interactive video worlds in real time by moving a joystick or typing prompts—like adding a volcano eruption or balloons.
- **4D Anyone** turns a simple video of a person into a 3D-like moving model you can view from any angle, perfect for animations or games.
- **Sense Nova U 1.58B** is a free AI image generator/editor that creates ultra-realistic 4K photos and edits images by typing natural language instructions.
- **DeepSeek’s latest model** (an AI tool for vision tasks) and a tiny **text-to-speech generator** (software that converts text into spoken audio) are now available for low-end devices.
- DeepSeek version 4 Pro, a new AI model focused on coding and task automation, is now available and ranks highly on AI benchmarks, offering strong performance at a low cost.
- DeepSeek 4 Pro can use tools, navigate code, and complete complex tasks, with pricing that's much cheaper than competitors like Kimi and GLM, making it a cost-effective choice.
- Miro, a team collaboration tool, now integrates with AI agents, allowing them to access and build upon shared team context, improving teamwork and efficiency.
- DeepSeek has introduced flexible reasoning efforts (low, high, max) for different task complexities and added OpenAI API support, making it easier to use with other tools.
- A new open-source AI model called GLM 5.3 (a free, community-developed AI tool) was released, which can match the performance of leading AI models like Cloud Fable and GPT.
- GLM 5.3 is designed for agentic coding (AI that can work on tasks independently), allowing it to handle complex, multi-step tasks and use other tools to achieve goals.
- It can be used through a framework called Zcode (a tool that helps manage and run AI agents), which is similar to other AI management tools like Codeex.
- In a demo, GLM 5.3 successfully created a browser-based replica of Windows 11 (a web version of the Windows operating system) with functional apps like Microsoft Office, Spotify, and Slack in just over 20 minutes.
- Bright Data (a company that helps extract data from the web) notes that the web is now a source of context for AI agents (AI tools that do knowledge work), not just data.
- The web's unstructured and constantly changing nature means extracting context from it is an ongoing process, not a one-time effort.
- AI search companies (companies built specifically to index the web for AI agents) are challenging Google's dominance in web search.
- New companies are emerging to extract more context from the web than traditional search allows, like tracking price changes or job position trends over time.
- Grock 4.6, a new AI model, can create complex tasks like coding and visual work, and even outperform other top models in reasoning and coding tasks.
- It's priced the same as the previous version, Grock 4.5, at $2 per 1 million input tokens and $6 per 1 million output tokens, and can be accessed through Grock build, Grock bot, or their API (a tool that lets different software talk to each other).
- Grock 4.6 has shown significant improvements in web development, backend automation, and tool calling, and even beat Opus 5, another top model, in some tasks.
- It's also exceptionally good at front-end web development, creating coherent and functional websites with complex features.
- OpenAI (a company making AI tools) might release a new, powerful AI model called Doug later this year, which could be even more advanced than their current models like Astra.
- Alibaba (a large tech company) has a new AI model, Kiana, that's showing strong performance in coding tasks, and you can try it out on a platform called Arena (a website where you can test different AI models).
- A new tool called ContextDev (a service that helps AI developers get clean data from websites) can help you extract structured data, generate AI-ready markdown, and pull brand assets like logos and colors from websites using a single API (a way for different software applications to communicate with each other).
- AI agents from Abacus AI (a company making AI tools) can now improve themselves without human help, like an AI sales agent that boosted its lead scoring accuracy by 254%.
- These agents use a split system: one part does the job, another checks the results, creating a feedback loop for continuous improvement.
- They work with many online services (like CRM software for tracking sales) and learn from real data specific to your goals, not just general best practices.
- For example, an AI agent fixed a bug in code by tracing it back to the exact mistake, wrote a fix, and left instructions for humans to verify it.
- Nate, an AI expert, teaches how to price AI solutions by using a client's own numbers to create a defensible price, ensuring you're paid in stages.
- He explains pricing an AI agent that automated appointment setting, saving the client $41,600 annually, and charging 13% of that as a one-time fee ($5,500).
- Nate suggests aiming for a 10x return on investment for clients to make the deal appealing, and includes a $400 monthly maintenance fee to keep the system running smoothly.
- He emphasizes the importance of tracking and communicating the system's impact on the business to demonstrate its value and secure future deals.
- DeepSeek version 4 Flash (a new, affordable AI model) now ranks 10th in overall performance, beating other models like Opus 4.7 and Claude Sonnet 5 due to its cost-efficiency and improved ability to plan and use tools (called "agentic behavior").
- This model is open-source under the MIT license (meaning anyone can use or modify it for free, even for commercial purposes), and it's one of the top three open-weight models in terms of intelligence.
- DeepSeek 4 Flash offers near Luna-level intelligence (a high-performing AI model) at about 60% lower cost per task, making it a great value for performance.
- The model has shown significant improvements in front-end development tasks, such as creating landing pages and generating 3D product images, as well as cloning complex interfaces like Mac OS.
- A new AI model called Kimmy K3 (a type of AI software that can understand and generate text, images, and more) was released for free by a Beijing lab, and it quickly became popular, even being praised by American companies.
- The US government accused the creators of Kimmy K3 of using a technique called "distillation" (copying the outputs of a stronger AI model to improve a weaker one) to steal American AI technology, and they threatened sanctions (penalties that limit business) if it's true.
- China denied these accusations, and the creator of Kimmy K3 said the model's improvements came from original changes, not copying.
- A new tool called Dream Nina (a website that helps create AI-generated videos) was introduced, which allows users to guide video creation using images, videos, audio, and text references together in one place.
- You can now test new AI models (like Claude or Codex, which are AI tools that help with tasks) against your own work, not just generic benchmarks (tests that don't relate to your specific needs).
- AI can scan your past conversations to find common tasks, then test new models on those tasks using a rubric (a set of criteria you create) that matters to you.
- You can compare different AI models or effort levels with a simple command (like "/benchmark") and get a report showing which model performs best for your specific needs.
- This approach helps you decide if a new AI model is worth using for your work, saving time and avoiding irrelevant tutorials.
- Poolside AI released Lagona S2.1, a powerful AI model (a computer program that can learn and make decisions) that can run on a single high-end computer and is great for coding tasks.
- Despite having fewer active parameters (a measure of a model's complexity), Lagona S2.1 competes with much larger models and can run locally on consumer-grade hardware.
- In coding tests, Lagona S2.1 performed nearly as well as larger models but was faster and could run on a 128 GB MacBook.
- You can try Lagona S2.1 for free on the World of AI benchmark tool (an online platform for testing AI models) and compare its performance with other models.
- A new AI model called Kimmy K3 (a type of AI software made by a company in China) was released for free, and it's as good as leading models like ChatGPT and Claude (popular AI chatbots).
- Kimmy K3 is open-source (AI software that anyone can use and modify), unlike closed-source models controlled by companies like OpenAI and Anthropic (AI companies that make ChatGPT and Claude).
- China is giving away advanced AI models for free to gain influence, control tech standards, and compete with U.S. companies, a strategy called "scorched earth" (a business tactic to undercut competitors by offering similar or better products for free or at a lower cost).
- To save money on AI-powered automations (repetitive tasks done by software), use tools like Zapier (a service that connects different apps to automate tasks) for non-AI tasks and switch AI models as better ones become available.
- AI can get overwhelmed and give worse answers when processing too many files at once, due to a limit on how much it can "hold in its head" (called a context window).
- To fix this, create a "table of contents" (a simple map file) and summaries of your files, so AI can quickly find what it needs without reading everything.
- This three-layer system (map file, summaries, and source files) helps AI work better with large amounts of data, like hundreds or thousands of files.
- Abacus is creating a system (called "agents") that automatically picks the best AI model (like GPT-5.6 or Fable 5) for each part of a task, making it more efficient than using just one model.
- This system can handle complex coding tasks, like inspecting code for security issues, explaining code structure, and even generating scaling plans for databases and other systems.
- The system can also be used to build and deploy applications, like a 3D castle in a browser or a trading strategy platform, all from a single conversation.
- The system is built on top of a supercomputer that can run 24/7, host databases, and connect to other services like AWS (a cloud computing service) and GitHub (a platform for software development).
- Hermes Agent (a powerful AI tool that can act like a full-time employee) works best with the Opus model (a specific AI model that's very reliable but expensive), but ChatGPT (a popular AI chat service) and GLM 5.2 (a cheaper AI model) are also options.
- To avoid downtime, run at least two Hermes agents simultaneously, using different AI models or accounts, so they can monitor and fix each other if one fails.
- You can create new Hermes agents (called "profiles") either by asking an existing agent to set one up for you or by using the Hermes dashboard.
- If you're running a serious business, consider investing in the Opus model for Hermes Agent, as it's the most reliable for completing tasks.
- Google DeepMind is reportedly launching Gemini 3.5 Pro, a new AI model (a computer program that predicts text) on July 17th, built on a fresh base model with a massive 2 million token context window (the amount of text it can process at once).
- Gemini 3.5 Pro is rumored to have a specialized deep think reasoning layer for better logic, math, and multi-step reasoning tasks, and support for advanced autonomous agentic workflows (AI that can perform tasks without human input).
- Leaked outputs show Gemini 3.5 Pro excelling in creating detailed SVG (a type of image file) graphics and generating complex code, like a 3D Subway Surfers style game in around 800 lines of HTML (a coding language for websites).
- Google is also testing other unreleased Gemini models, including Gemini 3.5 Flash High and a mysterious checkpoint labeled as 3 Flash, which could be Gemini 3.6 or even Gemini 4 Flash.
- OpenAI's Mark Chen believes AI is advancing rapidly, with AI models soon doing self-sustaining research, pushing science forward with less human control (AGI, or artificial general intelligence, means AI that can understand, learn, and apply knowledge like a human).
- AI is already showing signs of "divine moves" (unexpected, innovative solutions) in fields like math and computer science, and AI agents are starting to do meaningful work in their own fields.
- OpenAI is working towards a future where AI can conduct end-to-end research, from idea to result, with humans acting as orchestrators (managing and guiding the AI's work).
- Challenges include evaluation (making sure AI is actually improving) and the "jagged frontier" (AI excelling at complex tasks but struggling with simple ones), with continual learning (AI carrying lessons from one task to the next) being a key area for improvement.
- AI tools like ChatGPT and Claude (popular AI chatbots) have settings that control how much effort they put into answering, which can make a big difference in the quality of their responses.
- You can choose different "models" (versions of the AI) and "reasoning levels" (how hard the AI tries) to match the task, like a small model with low effort for quick tasks or a large model with max effort for complex ones.
- For most people, the default settings are too basic for tasks like writing emails or summarizing reports, so adjusting these settings can help you get better results without changing your prompts (the instructions you give the AI).
- AI providers set low defaults to make the AI respond quickly and to save them money, but you can change these settings to suit your needs, like using a higher reasoning level for tasks that require more thought.
Key points
What it is
- Comparing AI models means testing different AI systems to find the best one for your specific task.
- AI models vary in speed, cost, and quality, so one model might be great for writing code but terrible for analyzing customer comments.
- AI models are becoming cheaper and more accessible, but their intelligence isn't one-size-fits-all.
- You need to match the right model to your unique processes, decisions, and historical context.
How to use it
- Start by picking the "closest model" (the one you're already using or have recently used) and go deep on that single tool first.
- Use a benchmark tool like the free World of AI Benchmark to test models across multiple categories and prompts.
- Define how you will measure success (called an evaluation) before comparing models, and always start with the biggest, most capable model.
- Break your task into smaller steps and consider which tool is best for each specific output.
Watch out for
- Developers often overestimate AI’s benefits, so treat AI output like code from a junior developer—review it carefully.
- Models are shaped by their training data, so gaps create blind spots, and different tasks need different models.
- Don't assume one model works for everything, and don't tell the AI it's an expert—this can make it overconfident without improving accuracy.
- Consider the cost of a wrong answer, not just the price per call, and test the model, your prompt harness, and the overall system.
Tools named
- World of AI Benchmark (free tool for testing models across multiple categories and prompts)
Lesson 1: What is Comparing AI Models and why it matters
Comparing AI models means evaluating different AI systems to see which one performs best for your specific task. Developers do this because AI models vary dramatically in speed, cost, and quality. A model that excels at writing code might be terrible at analyzing customer comments. Testing multiple models side by side helps you match the right tool to the right job.
Why does this matter? First, AI models are becoming cheaper and more accessible, but intelligence itself is not commoditized. What matters is your unique processes, decisions, and historical context. You need to collate that information and plug it into the right model with the right framework. Second, after hundreds of hours testing leading tools like Claude Code and Google Antigravity, users found major differences in performance. Choosing the wrong model wastes time and money.
Be aware that developers often overestimate AI’s benefits. One rigorous trial found experienced developers using AI tools took 19% longer to complete tasks, even though they thought they were 24% faster. Treat AI output like code from a junior developer—review it carefully and never assume it’s correct. AI accelerates, but humans validate. The combination is powerful; either alone is incomplete. By comparing models critically, you avoid blind spots and build reliable, profitable AI systems.
Sources
- 2026-01-03 — The AI Choice You’ll Regret in 2026
- 2025-11-24 — This AI Model Is Smarter Than Ever Before!
- 2026-03-08 — Is AI Really Intelligent or Just Fancy Autocomplete 2026
- 2026-01-29 — From Coder to Orchestrator The Developer Role Shift Nobody's Talking About
- 2026-05-08 — AlphaEvolve broke the matrix multiplication record. You didn't notice!
- 2026-05-01 — Build & Sell Claude Code Operating Systems (2+ Hour Course)
- 2026-03-03 — The One Skill AI Can't Replace -- Are You Developing It
- 2026-03-15 — Stop Learning New AI Tools
- 2026-03-12 — Build & Sell with Claude Code (10+ Hour Course)
- 2026-05-07 — The 4 Primitives Making Claude Agents Unstoppable #Claude #AI
- 2025-12-26 — AI Skill That Pays in 2026 Systems
- 2026-04-13 — 100 Hours Testing Claude Code vs Antigravity (honest results)
- 2026-03-21 — Anthropic Found the Pattern Everyone Missed About AI!
- 2026-03-02 — Claude Code Skills are BROKEN
Lesson 2: How to use Comparing AI Models: step-by-step
To compare AI models, start by picking the "closest model" (the one already open in your tab or that you’ve used recently). Go deep on that single tool first — solve your problem fully before testing others. If your chosen model solves the task, ship it and move on; don’t get distracted.
When you need to compare models, use a benchmark tool like the free World of AI Benchmark, which tests models across multiple categories and prompts. Before any comparison, define how you will measure success — this is called an evaluation (how you measure and track if a model performs well). Always start with the biggest, most capable model to find the "quality ceiling" (the best possible performance). Then you can downshift to cheaper models, but only if your evaluations prove the cheaper model preserves the same behavior.
Break your task into baby steps and consider which tool is best for each specific output. For example, simple sorting and serious judgment should not use the same model. The trade-off has three knobs: capability, cost, and confidence. Confidence comes from your evals, not from liking a smaller invoice. You are testing three things: the model itself, your harness (the recipe and chef structure that runs steps), and your coding harness. A strong model can sometimes overcome a weak harness, but you need both for reliable results.
Sources
- 2026-05-23 — Claude and ChatGPT Got More Literal. Your Old Prompts Are Backfiring
- 2026-05-25 — ChatGPT vs Claude vs Gemini Is the Wrong Question
- 2026-06-02 — GPT 5.5 vs Opus 4.8 vs Gemini 3.5 - Which Model Should You Use
- 2026-04-08 — The Next Layer After Prompt Engineering — Archon V3 Explained! 🚀
- 2026-06-01 — 20 days of compute vs 7 hours rethinking what state-of-the-art means Bertrand Charpentier, Pruna
- 2026-05-08 — Overwhelmed By AI Just Copy My Tech Stack
- 2026-05-20 — MCPs Are Dead. Claude Code Wants CLIs
- 2026-06-18 — The Production AI Playbook Deploying Agents at Enterprise Scale Sandipan Bhaumik, Databricks
- 2026-06-11 — Nex-N2 Pro IS GREAT! New Opensource Model Beats GPT 5.5, Opus 4,7, & Gemini 3.5 (Fully Tested)
- 2026-02-11 — Get the Most from Claude Opus 4.6 — 6 Behavioral Shifts + 5 New Features Most Developers Miss
- 2026-06-06 — Gemma 4 12B Is INCREDIBLE! BEST Local AI Coding Model! IS POWERFUL! (Fully Tested)
- 2026-05-10 — Pick the Right LLM for Each Agent Role with OpenAI Agents SDK Part 4
- 2026-06-06 — Evals Are Broken, Use Them Anyway Ara Khan, Cline
Lesson 3: Best practices and pitfalls
When comparing AI models, a common mistake is picking the cheapest or most popular one first. Instead, start with the most capable model to find the quality ceiling — the best possible output for your task. Only then should you downshift to a cheaper model, but only if your eval set (a collection of test examples) proves the cheaper model preserves the same behavior. This avoids wasting money on a model that can't do the job.
Pitfalls include trusting a model that sounds confident but is wrong — always double-check. Models are shaped by their training data, so gaps create blind spots. Don't assume one model works for everything; different tasks need different models. Another mistake is telling the AI it's an expert — this can backfire by making it overconfident without improving accuracy.
Best practices: use multiple models in parallel to compare responses. If they disagree, human judgment becomes critical. Also, think about the cost of a wrong answer, not just the price per call. Is the decision reversible? Low-stakes tasks can use cheaper models; high-risk ones need the best. Finally, test three things: the model itself, your prompt harness, and the overall system. Confidence comes from evals, not from liking a cheaper invoice. Simple sorting and serious judgment should not sit at the same desk.
Sources
- 2026-06-27 — OpenAI Already Built the Future of Work. You Can Copy It.
- 2026-06-01 — I Run 4 AIs at Once in Claude Cowork (Here's My Exact Setup)
- 2026-05-26 — AI Just Changed How You Run a Business Forever! (Tutorial)
- 2026-05-17 — Anthropic Just Exposed Claudes Hidden Survival Mode
- 2026-06-02 — GPT 5.5 vs Opus 4.8 vs Gemini 3.5 - Which Model Should You Use
- 2026-05-09 — Claude and ChatGPT Hallucinate Less. That's Why They're Dangerous.
- 2026-06-17 — ChatGPT & Claude Are Built to Give You the Average Answer
- 2026-05-28 — Context Graphs for Explainable, Decision-Aware AI Agents Andreas Kollegger & Zaid Zaim, Neo4j
- 2026-03-08 — Is AI Really Intelligent or Just Fancy Autocomplete 2026
- 2026-05-30 — Google Remy, Grok 5, Mythos 1, New Atlas Robot, ASI and More AI News This Month!
- 2026-02-11 — Get the Most from Claude Opus 4.6 — 6 Behavioral Shifts + 5 New Features Most Developers Miss
- 2026-06-05 — Anthropic Just Warned Everyone About Claude (Its Evolving)
- 2026-05-23 — Claude and ChatGPT Got More Literal. Your Old Prompts Are Backfiring
- 2026-05-10 — Pick the Right LLM for Each Agent Role with OpenAI Agents SDK Part 4
- 2026-06-06 — Evals Are Broken, Use Them Anyway Ara Khan, Cline