Models & Comparisons

Cutting-Edge AI Models

Last updated 2026-08-01

What's new

2026-08-01
  • DeepSeek version 4 Flash (a new, affordable AI model) now ranks 10th in overall performance, beating other models like Opus 4.7 and Claude Sonnet 5 due to its cost-efficiency and improved ability to plan and use tools (called "agentic behavior").
  • This model is open-source under the MIT license (meaning anyone can use or modify it for free, even for commercial purposes), and it's one of the top three open-weight models in terms of intelligence.
  • DeepSeek 4 Flash offers near Luna-level intelligence (a high-performing AI model) at about 60% lower cost per task, making it a great value for performance.
  • The model has shown significant improvements in front-end development tasks, such as creating landing pages and generating 3D product images, as well as cloning complex interfaces like Mac OS.
2026-07-28
  • Claude (an AI assistant) can autonomously organize files on your computer, renaming and summarizing them without your direct input, using its "co-work" feature (a mode where it operates more independently).
  • Most people use Claude in "chat" mode (like a conversation), but the "co-work" mode (a separate feature) is more powerful for handling complex tasks with many files or steps.
  • Downloading the Cloud Desktop app (software for your computer) lets Claude access your files directly, maintaining its intelligence for longer tasks, unlike the browser version where you must manually copy and paste files.
  • In the desktop app, create folders tied to specific tasks instead of using "projects" (pre-set workspaces in the browser version), allowing Claude to work more efficiently on your files.
2026-07-25
  • Anthropic released Claude Opus 5, a powerful new AI tool (a computer program that can understand and generate text, and now create complex 3D worlds and simulations) that can create detailed 3D worlds and simulations with minimal input.
  • Opus 5 can generate interactive experiences like games, virtual reality environments, and educational tools, showcasing its ability to create engaging and complex content quickly.
  • The AI can also simulate physics, like fabric movement in wind, and generate intricate designs like fractals, demonstrating its advanced capabilities in understanding and replicating real-world phenomena.
  • Opus 5 can create interactive simulations of ecosystems, flight simulators, and even simple drawing tools, highlighting its versatility in generating a wide range of applications.
2026-07-19
  • AI can now generate videos directly on mobile phones using a tool called MobileOne, which creates 5-second videos from text prompts.
  • Nvidia's new AI tool, RD, generates realistic 3D human movements in real-time, useful for games, animations, and robot training, and it's available to run locally on your computer.
  • Nvidia updated their PID upscaler to version 1.5, improving image details and color fidelity, and it works with various image models like Quinn image, Flux, and Z image.
  • An AI tool called Audio to MIDI by Mirell can dissect songs, creating separate MIDI tracks (notes for each instrument) from full songs, and it's available to try for free online.
2026-07-16
  • Google DeepMind is facing challenges, with key researchers leaving and delays in releasing Gemini 3.5 Pro, their advanced AI model (a complex computer program designed to understand and generate human-like text).
  • The model is not yet ready, as it doesn't match the quality of competitors like GPT 5.5 and Fable 5 (other advanced AI models), and has issues with reliability and consistency.
  • Google is reportedly testing new versions, but some perform worse than older ones, leading to repeated delays.
  • There's speculation that Google might skip Gemini 3.5 Pro and move straight to Gemini 4.0, but this hasn't been confirmed.
2026-07-13
  • Claude Code (a tool for building AI-powered automations) lets you work with local files and online services like Gmail, Slack, or a CRM (customer relationship management system), making it more powerful than Claude Chat (a simple AI chatbot).
  • Claude Code uses the same AI models (like Opus, Sonnet, or Haiku) as Claude Chat, but adds extra features for working with files and online services.
  • Claude Code is like an AI harness (a tool that helps you use AI models), which sits between the AI model (the engine) and you (the driver), helping you build automations and agents (AI systems that can do tasks for you).
  • The instructor, Nate, uses Claude Code to build and manage multiple businesses, showing how one person can do the work of a team with AI.
2026-07-07
  • AI is replacing many jobs, especially those done by junior workers, and this trend feels different from past economic downturns due to its existential nature (potentially changing the job market forever).
  • Don't believe everything you see online; negative news about job losses gets more attention, but it's not the full picture, so do your own research.
  • AI companies have reasons to hype up their products, so take their claims with a grain of salt and do your own research to understand how these tools are really evolving.
  • Many AI tools are still in development and not yet perfect, so don't be fooled by impressive demos—look for tools that have been proven to work well in real-world situations.
2026-07-01
  • A new free, open-source AI model called GLM 5.2 (a type of AI software that anyone can use and modify) is now available and performs nearly as well as more expensive models like Opus (another AI model) for most tasks.
  • GLM 5.2 is designed to be cost-effective, using only a small part of its vast capabilities for any single task, and can handle large amounts of information at once.
  • The model was tested by creating a real-world tool for tracking sponsorship deals, which worked well and cost significantly less to run than Opus.
  • Additionally, GLM 5.2 was used to create a promotional video for the tool using an open-source tool called HyperFrames MCP (a software that turns text into videos), though Opus produced a more polished version.
2026-06-16
  • A new AI tool called Scale 2 (open-source software for animating characters) can transfer motion from one video to another, even with multiple characters or non-human subjects, and is available to download and run locally.
  • Actionable World Representation (a system for creating digital twins of real-world objects) can model how objects move, bend, or change, and is useful for training robots to interact with the real world.
  • Oscar (a world model for robots) can predict outcomes of robot actions and work across different robot types, helping to train robots in virtual environments.
  • Google's Gemini 3.5 Live Translate (a real-time translation tool) can translate speech in one language to another while maintaining the original speaker's voice.
2026-06-10
  • Major AI providers like OpenAI (makers of ChatGPT), Anthropic, and Google now offer app layers for working with AI agents (AI tools that can perform tasks for you), with new options like Deep Seek GUI for coding, writing, and automation.
  • Deep Seek GUI is a new desktop app that turns Deep Seek (a type of AI model) into a user-friendly workspace, with features like code mode for project files and write mode for document editing.
  • Deep Seek's pricing is now permanently discounted, making it one of the most affordable AI coding setups, with costs as low as 4 cents for 1 million input tokens.
  • TestSprite, an AI-powered testing agent, helps catch bugs in apps by simulating user flows, complementing code reviews and reducing verification debt (when code isn't properly checked before shipping).
2026-06-07
  • OpenAI's Codex (a tool that helps write and understand code) got an update for building websites, and ChatGPT (a chatbot that uses AI) got a memory update for better conversation flow.
  • Google released Gemma 4 12B, an AI model that can run locally on your device using LM Studio (a software for running AI models), and Ideogram 4, a top open-source image generator.
  • New AI models for creating realistic images, expressive text-to-speech, and generating music and video were released, with a focus on open-source (software anyone can use and modify) tools.
  • Rumors about upcoming AI models like GPT-5.6 (a potential new version of OpenAI's language model) and Mythos/Oceanus (a new model from Anthropic, another AI company) suggest improvements in spatial understanding and realistic outputs.

Key points

What it is

  • Cutting-edge AI models learn from examples, like a cook creating new recipes after studying many dishes, instead of following step-by-step instructions.
  • They are becoming cheaper and more accessible, but their true value comes from combining them with your unique data and processes.
  • These models are splitting into two main groups: US models and Chinese open-source models, offering more options at lower costs.
  • The next wave of AI needs a "model of reality," not just language, with companies like OpenAI, Google DeepMind, and Tesla competing in this area.

How to use it

  • For complex tasks, choose high-end models like Opus 4.7, GPT 5.5, or Gemini, and keep instructions under 100 lines for a high-level overview.
  • Make your codebase model agnostic (not dependent on one AI) so you can easily swap models as needed.
  • Use OpenRouter, a unified API, to route requests between models based on task, latency, and fallback reliability.
  • For verification, split out claims in a separate conversation and have the AI check them against sources.

Watch out for

  • Avoid jumping between models; instead, choose one that fits your specific task.
  • Don't trust AI responses too easily; always verify outputs, as models can sound confident but give incorrect information.
  • Don't tell the AI it's an "expert" in your field; be explicit about which files or data it must reference.
  • Limit using multiple AIs at once to four models to avoid becoming overwhelmed.

Tools named

  • Opus 4.7 (a high-end AI model for complex tasks), GPT 5.5 (a high-end AI model), Gemini (a high-end AI model), OpenRouter (a unified API for routing requests between models), DeepSeek (a model for token-efficient work)

Lesson 1: What is Cutting-Edge AI Models and why it matters

Cutting-edge AI models are the latest generation of systems that learn from examples instead of following step-by-step instructions, much like a cook who writes a new recipe after studying thousands of finished dishes. These models are becoming cheaper and more accessible, but intelligence itself is not yet a commodity. The real value for you comes from your own proprietary processes, decisions, and historical context — plugging that unique data into a model with the right framework is what matters.

Why does this matter for AI development? First, the engineering around models now matters as much as the models themselves. Leading in raw capability does not guarantee leading at the application layer, so even without the strongest model, a platform company can find a differentiated path. At the same time, new models are putting enormous pressure on frontier labs for price efficiency, especially for developers running large-scale agentic workflows where token costs stack up fast.

Second, the market is splitting into two camps: US models and Chinese open-source models. This means more options at lower cost for you.

Third, the next wave of AI needs a model of reality, not just a model of language. World models are becoming a platform war, with companies like OpenAI, Google DeepMind, and Tesla circling the same idea.

Finally, capabilities like finding flaws in software are spreading across the whole ecosystem. Once AI can analyze complex rule systems at superhuman scale, it changes how you must think about security and oversight. You must assure quality and keep projects on track, because the models still require human direction to solve real problems.

Sources

Lesson 2: How to use Cutting-Edge AI Models: step-by-step

To use cutting-edge AI models effectively, start by choosing a high-end model like Opus 4.7, GPT 5.5, or Gemini for complex, multi-step tasks. For example, in a trading workflow, you would set up Opus 4.7 as the core reasoning engine, using routines to schedule actions and building custom skills for research, decision-making, and logging. The key is keeping instructions under 100 lines to act as a high-level overview of what the AI should do.

When building coding projects, prioritize making your codebase model agnostic so you can swap models easily. Use OpenRouter, a unified API that routes requests between models like Opus, Gemini, and DeepSeek based on task, latency, and fallback reliability. This is especially useful for cutting-edge setups where model choice matters. For simpler subtasks, you can route requests to faster or cheaper models—for instance, using DeepSeek for token-efficient work while reserving Opus for critical reasoning steps.

To improve output quality, avoid telling the AI it is an expert in any field, as this can backfire. Instead, be explicit about specific files it needs to reference. Also, use high-end models for verification steps: split out claims in a separate conversation and have the AI check them against sources. For retrieval tasks, generate a Google API key to access Gemini’s new embeddings model, then combine it with OpenRouter for a robust pipeline.

Sources

Lesson 3: Best practices and pitfalls

Beginner AI Pitfalls & Best Practices

Most beginners jump between models, trying every new release, but that leads to being average at everything instead of great at one tool. A better approach is to choose a model that fits your specific task. For example, for complex information extraction, Gemini 3.1 Pro often outperforms Opus 4.7 or GPT 5.5. Don’t get overwhelmed by features.

A major pitfall is trusting AI responses too easily. Models often sound confident but give incorrect information. The latest Opus 4.8 was trained to be more honest (avoid making unsupported claims), but you still need to verify outputs. Always check facts yourself.

Another common mistake is telling the AI it’s an “expert” in your field. This old prompt trick now backfires because models take it too literally. Instead, be explicit about which files or data it must reference.

For complex projects, AI still requires heavy human oversight. About 80% of AI projects fail to reach production, often because people expect a tool to solve everything without proper judgment. Your primary skill should be judging AI output quickly. When using multiple AIs at once, limit to four models; more than that becomes a burden.

Finally, keep your codebase model agnostic (not locked to one AI). This lets you switch models as better ones emerge without rewriting everything.

Sources