Models & Comparisons

New AI Model Releases

Last updated 2026-09-16

What's new

2026-09-16
  • OpenAI has released a new Agents API (a tool for developers) that lets you use their powerful AI model, Codex (a tool that can write and understand code), in the cloud, making it easier to build smart applications.
  • Codex can now operate in a sandbox (a safe, isolated environment) where it can run code, edit files, and connect to servers, all while being managed by OpenAI's cloud services.
  • OpenAI is expected to release GPT-6 soon, with models like Luna (a cheaper, highly capable model) that could make Codex even more powerful, enabling it to do complex tasks like building games or scraping the internet for data.
  • Tools like Bright Data (a service that helps AI models gather information from the web) can be integrated with Codex to make it even more powerful, allowing developers to build advanced AI applications with just a few lines of code.
2026-09-07
  • Two new AI models, GPT-6 Astra (a type of AI chatbot by OpenAI) and Claude Fable 5.1 (a type of AI chatbot by Anthropic), were released, with GPT-6 Astra showing better performance in benchmarks (standardized tests for AI models).
  • GPT-6 Astra is also more cost-effective, providing similar performance to Claude Fable 5.1 at a lower price, making it a better value for users.
  • In real-world tests, like creating a playable browser game of Fortnite (a popular video game) from a single request, GPT-6 Astra showed impressive capabilities, although with some limitations like weak AI-controlled opponents (bots).
  • Both models demonstrate significant advancements in AI technology, with GPT-6 Astra showing a notable leap forward compared to previous models.
2026-09-04
  • Claude Code (a coding assistant tool) now offers a free model with 1.5 billion tokens monthly, a significant cost-saving for users previously paying for credits.
  • Omni Route (a laptop app) aggregates free API keys from multiple providers, automatically switching between them to maximize free usage, avoiding the hassle of manual configuration.
  • This setup allows users to keep the Claude Code interface while changing the underlying AI model (the "brain") to a free one, making it versatile for other tools like Deep Sea Carnage (another AI coding tool).
2026-08-28
  • Open-source AI models (free to download, customize, and use) are gaining popularity over closed models (like ChatGPT and Claude, which are paid and controlled by specific companies), with China leading in open-source development.
  • Despite open-source models like DeepSeek having a higher usage share, closed models from companies like Anthropic and OpenAI still capture most of the revenue due to their higher pricing and perceived slight performance edge.
  • The rise of open-source AI benefits companies like Nvidia, as more usage leads to increased demand for the computer chips needed to power these models.
  • Tools like Higsfield MCP (a service that automates tedious tasks for content creators) are emerging, showcasing how AI can integrate with other software to improve productivity.
2026-08-25
  • Most AI "agents" (tools that automate tasks) work similarly under the hood, using a brain (AI model) and a workspace (cloud or your computer) to do tasks.
  • Grok Bot is a new, easy-to-use AI agent that runs entirely in the cloud, so it keeps working even if you turn off your computer.
  • Hermes Agent is an open-source alternative that lets you swap AI models (the "brain") and customize tools more freely than Grok Bot.
  • Some AI tools gain sudden popularity due to platform algorithms favoring them, not just because they’re better.
2026-08-22
  • Applied vertical AI (AI designed for specific industries, like pharma or legal tech) is built to simulate a job in that industry, unlike general-purpose AI like Google Translate.
  • AI agents are being used in various industries, and their success is measured by their return on investment (ROI), not just their ability to perform tasks.
  • When building AI agents, it's important to start with a very narrow, specific task, rather than trying to make one agent do everything.
  • Proprietary data (unique, often expensive-to-obtain data that your organization has) is crucial for making your AI application better than general AI tools like ChatGPT.
2026-08-19
  • To create unique websites, use a design system (a set of rules and examples) instead of just giving AI tools (like Claude, an AI assistant) simple prompts (instructions), as this often leads to generic, repetitive designs.
  • Study well-designed websites in your field (like Duolingo, an AI education app) to understand what makes them visually appealing, then ask AI tools to help you create something similar.
  • Avoid common AI design trends, such as using the same fonts, color gradients (like indigo to violet), and layout features (like three feature cards) that make websites look generic.
  • Use free online resources (like the ones provided in the video) to access guides and communities that can help you improve your design skills and create more attractive websites.
2026-08-16
  • SpaceX released Grok 4.6, a new AI model that excels in tasks like legal work, long professional tasks, and economically valuable work, and is more affordable than competitors like Claude and Opus.
  • They also introduced Grok Bot, a new super app (a powerful, all-in-one software) focused on non-coding work, designed to help people interact with AI agents to improve productivity in their jobs.
  • Grok Bot is unique because each session is treated as its own AI agent, which can be named and given specific tasks, making it easier to manage and automate workflows.
  • SpaceX plans to release Grok 4.7 soon, which Elon Musk claims will outperform current models, including Claude, in real-world engineering tasks.
2026-07-22
  • A new AI model called Kimmy K3 (a type of AI software made by a company in China) was released for free, and it's as good as leading models like ChatGPT and Claude (popular AI chatbots).
  • Kimmy K3 is open-source (AI software that anyone can use and modify), unlike closed-source models controlled by companies like OpenAI and Anthropic (AI companies that make ChatGPT and Claude).
  • China is giving away advanced AI models for free to gain influence, control tech standards, and compete with U.S. companies, a strategy called "scorched earth" (a business tactic to undercut competitors by offering similar or better products for free or at a lower cost).
  • To save money on AI-powered automations (repetitive tasks done by software), use tools like Zapier (a service that connects different apps to automate tasks) for non-AI tasks and switch AI models as better ones become available.
2026-07-19
  • Moonshot AI released Kimmy K3, an open-source AI model that rivals top closed models like GPT 5.6 and Claude Fable, offering advanced capabilities like multi-step, autonomous task completion.
  • Kimmy K3 can be used through an agentic harness (a tool that lets AI act like an assistant) called Kimmy Code, which allows it to manage multiple projects and files simultaneously, similar to having an "army of agents" working for you.
  • The model can perform complex tasks, such as creating a liquid splash simulator with adjustable physics and hand-tracking controls, all from scratch without external libraries, demonstrating its advanced understanding of physics and coding.
  • Kimmy K3 can also integrate with other tools like Blender (a 3D modeling software) through its MCP (a connection tool), automatically generating complex 3D models like a realistic animated V8 engine.
2026-07-13
  • OpenAI launched ChatGPT Work, a new tool (similar to Claude Co-work) that helps non-technical people automate tasks, available on web, mobile, and desktop apps.
  • It's powered by the new GPT 5.6 model, offering different versions for speed or capability, and syncs tasks across devices.
  • ChatGPT Work includes features like skills (automated tasks), sites (live outputs), and scheduled tasks, similar to Claude Co-work.
  • You can easily migrate skills from Claude Co-work to ChatGPT Work, making it simple to switch between the two.
2026-07-10
  • An AI agent (a tool that automates tasks) is like a folder containing instructions (what to do), connections to other systems (like email or accounting software), and a trigger (when to start).
  • Unlike a chatbot (which just answers questions), an agent does work for you, like sending reminder emails to clients who haven't paid their invoices.
  • Agents are most useful for recurring tasks, and you can trigger them manually, by a schedule, or by an event.
  • You can build agents using existing AI tools like ChatGPT or Claude (AI programs you might already use).
2026-07-04
  • Anthropic's new AI model, Claude Sonnet 5 (a smart tool for coding and tasks), costs more than expected due to a new "tokenizer" (a system that breaks text into parts) that increases the number of tokens used.
  • Sonnet 5 is designed to work well with other models, like Opus (a more expensive, smarter AI), by handling simpler tasks at a lower cost, but it can be expensive if used incorrectly.
  • The model is best suited for complex systems where one AI (like Opus) plans tasks and another (like Sonnet 5) executes them, rather than for simple, one-off tasks.
  • Users should be aware of the model's tendencies, such as overthinking and using more tokens than necessary, to avoid unexpected costs.

Key points

What it is

  • New AI model releases are updated versions of AI systems, like Opus 4.8 from Anthropic or GPT-5.6 from OpenAI, that bring better performance and new capabilities.
  • These releases improve frontier intelligence (the cutting-edge ability of AI), making models smarter and easier to use with natural English.
  • The rapid pace of releases has led to an "iPhone era," where upgrades may feel small, and a capability overhang (models outpacing supporting systems).
  • Open-source models (AI models developed and maintained by communities) also matter, as they come with community oversight, creating trust.

How to use it

  • To use new AI model releases, you need API access (a key to connect your code to the model). Create an account with the provider and grab your API key.
  • Plug in the model name, such as GPT-5.6 or Claude Sonnet 5 (a new version of Anthropic’s mid-tier model), in your tool or via a unified interface like OpenRouter.
  • Use the highest intelligence model you can access for tasks, and consider using frameworks like Tank to orchestrate models via one command line.
  • Update your prompts as models evolve, being explicit about files and tasks, and avoid telling the AI it's an expert to prevent inflated confidence in wrong answers.

Watch out for

  • New models may not be flawless. Wait for thorough community reviews before relying on a new model, as hype can mislead.
  • Avoid sticking with one tool out of overwhelm. The landscape shifts fast, so always compare models on real tasks, not just benchmarks.
  • Don't trust a single model's output blindly. Use cross-validation (having one model review another's work) to catch errors and increase trust.
  • Beware of model deprecation (models being shut down). Diversify your tools to avoid over-reliance on any single one.

Tools named

  • OpenRouter (a unified interface for using multiple AI models), Cursor (an app for using AI models), Tank (a framework for orchestrating AI models), Claude Code (a coding-focused AI model), GPT OSS (an open-source version of GPT)

Lesson 1: What is New AI Model Releases and why it matters

New AI model releases are new versions of AI systems released by labs like Anthropic or OpenAI. They matter because each release typically brings better performance and new capabilities. For example, Anthropic’s Opus 4.8 is described as “the most advanced AI model in the world.” These releases often improve frontier intelligence (the cutting-edge ability of AI), meaning models become smarter and require fewer “prompt hacks” (special tricks to get good results). Instead, you can just say what you want in natural English. The competition between labs has caused an acceleration of model releases over the past six months, as OpenAI and Anthropic were in “an absolute battle.”

However, some argue we’ve entered an “iPhone era” where upgrades feel small and hard to distinguish. The real shift may be away from focusing only on model releases toward building agentic ability (handling complex multi-step work). As one video explains, waiting for the next release is less important than your own knowledge and understanding—you already have tools to build what you need. Additionally, model releases create a capability overhang (models so capable that surrounding systems haven’t caught up), meaning scaffolding and business processes lag behind. Open-source models also matter: they come with community oversight, creating trust. But despite rapid releases—sometimes 2-3 new things per week from one lab—the key is that each AI does one thing very well, and chaining those outputs together achieves great results. New model releases drive progress, but the ultimate goal is systems that reason, split up work, and deliver finished outputs people can use.

Sources

Lesson 2: How to use New AI Model Releases: step-by-step

To use new AI model releases, you first need API access (a key to connect your code to the model). Create an account with the provider—like OpenAI—and grab your API key. Then, in your tool, plug in the model name, such as GPT-5.6 or Claude Sonnet 5 (a new version of Anthropic’s mid-tier model). You can do this inside apps like Cursor or via a unified interface like OpenRouter, which lets you use one key for many models—GPT-5.5, Claude Opus, DeepSeek V4, and more. Early testing of GPT-5.6 suggests it offers huge progress, while Claude Sonnet 5 is a solid upgrade spotted through a partner provider.

When working with these models, stop telling the AI it’s an expert. Be explicit about which files it needs to reference. Use the highest intelligence model you can access—Opus 4.7 or GPT-5.5—for tasks like splitting out claims from a conversation. For coding, frameworks like Tank orchestrate models such as Claude Code and GPT OSS, all via one command line. Remember, GPT models may constantly ask for input; Opus would often complete the task without prompting you. To add a model to Cursor, use OpenRouter—no need for a ton of separate API keys. This way, you stay current with releases like GPT-5.6 and Claude Sonnet 5.

Sources

Lesson 3: Best practices and pitfalls

New AI model releases like Claude Sonnet 5 or GPT-5.6 can be exciting, but beginners often make costly mistakes. A common pitfall is assuming every new model is flawless. For example, Claude Sonnet 5 initially showed early promise in testing, but some users later reported it was "horrible" compared to prior versions, showing that hype can mislead. Best practice: wait for thorough community reviews before relying on a new model.

Another mistake is sticking with one tool out of overwhelm. One creator admitted staying "consistent" with Claude Code to reduce stress, but this can cause you to miss upgrades. The landscape shifts fast—Opus 4.5 was once "king of coding," then Sonnet 4.6 surpassed it. Always compare models on real tasks, not just benchmarks.

A key best practice: don't trust a single model's output blindly. Use cross-validation (having one model review another's work). For instance, give a task to Claude, then pass the result to GPT for critique. If they disagree, apply your own judgment—this catches errors and increases trust.

Beware of model deprecation (models being shut down). GPT-4 was killed in February 2026, and Anthropic banned Fable 5 amid US government disputes. You don't own the models, so avoid over-reliance on any single one. Diversify your tools.

Finally, update your prompts as models evolve. Claude and ChatGPT have gotten more literal, so old prompts may backfire. Stop telling the AI it's an expert—it can inflate confidence in wrong answers. Instead, be explicit about files and tasks. These habits turn new releases into assets, not headaches.

Sources