Models & Comparisons

AI Model Optimization

Last updated 2026-07-25

What's new

2026-07-25
  • Opus 5, a new AI model from Enthropic (a company that makes AI tools), is a major upgrade, outperforming previous models in coding, reasoning, and complex tasks, and costing about half as much as Fable 5 (another advanced AI model).
  • Opus 5 is now the top-ranked AI model for reasoning and agentic work (AI that can perform tasks autonomously), scoring highly on various benchmarks and even surpassing Fable 5 in some areas.
  • Despite its strengths, Opus 5 has strict security safeguards that can sometimes block or flag harmless requests, which can be frustrating for users.
  • For specific tasks like web development, other models like Kim K3 (a different AI model) might still be preferable due to their lower cost and better performance in certain areas.
2026-07-22
  • Alibaba launched Coen 3.8 (a powerful AI model, like a super-smart computer program), a 2.4 trillion parameter model, claiming it's one of the most powerful AI models available, but it's not yet clear how it compares to others like Fable 5 (another AI model) or Kimmy K3.
  • Coen 3.8 Max (a version of the model) is a major upgrade, supporting text, images, videos, and has a 1 million token context window (meaning it can process and understand a lot of information at once).
  • The model shows significant improvements in tasks like coding, advanced reasoning, mathematics, and visual understanding, and can generate complex designs and animations.
  • Coen 3.8 Max is faster than some other models like Kim K3 (another AI model), but Kim K3 still performs better in certain tasks.
2026-07-19
  • Moonshot AI released Kimi K3, an open-source AI model (software anyone can use and modify) that's already available in apps like Kimi and Kimi Code, with impressive results for a low monthly cost.
  • Harbor SEO.ai, a new SEO (search engine optimization, helping websites rank higher on Google) tool for €29/month, scans websites for issues, suggests fixes, and even generates content.
  • Kimi K3 quickly built a well-designed, functional Irish Golf Directory website from a simple prompt, showcasing its potential for web development.
  • The model's capabilities, like generating SEO-friendly content and creating visual placeholders, make it a strong contender in open-source AI.
2026-07-16
  • Combi 2.0 (a tool that helps create website designs and code) now has a "design mode" where you can visually design a website's look before generating the code.
  • You can use Combi 2.0 to clone (copy) existing websites, like the "World of AI benchmark" site, and edit its styles and design.
  • Combi 2.0 lets you switch between designing, coding, and seeing the live website, with changes updating across all three views.
  • It supports various web development tools (like Shad CNN packages) and can connect with other apps like Figma (a web design tool).
2026-07-13
  • OpenAI released ChatGPT 5.6, its most powerful model yet, with three versions (Soul, Tara, Luna) at different price points, where Soul is the most powerful and cost-effective.
  • ChatGPT 5.6 can be used within a downloadable app or integrated with Claude, a go-to-market machine (a tool to help businesses grow), to access and compare all three models.
  • In a design task, Soul outperformed Tara and Luna in creating a two-page HTML presentation, demonstrating its superior capabilities in understanding and executing complex tasks.
  • ChatGPT 5.6 can be connected with Clay (a business growth tool) to find specific company information, enrich it with verified emails, and draft outreach emails, showcasing its potential for business applications.
2026-07-10
  • Hermes Agent (a powerful AI tool that can act like a full-time employee) works best with the Opus model (a specific AI model that's very reliable but expensive), but ChatGPT (a popular AI chat service) and GLM 5.2 (a cheaper AI model) are also options.
  • To avoid downtime, run at least two Hermes agents simultaneously, using different AI models or accounts, so they can monitor and fix each other if one fails.
  • You can create new Hermes agents (called "profiles") either by asking an existing agent to set one up for you or by using the Hermes dashboard.
  • If you're running a serious business, consider investing in the Opus model for Hermes Agent, as it's the most reliable for completing tasks.
2026-07-07
  • New AI models like Fable 5 (a smart AI assistant) and GPT 5.6 do too much, so you need to simplify your requests to avoid overwhelming them.
  • To get the best results, tell the AI what you want to achieve, who it's for, and what "done" looks like, without giving step-by-step instructions.
  • Use a "fence" to prevent the AI from doing extra, unasked tasks by setting clear boundaries in your requests.
  • Fable 5 will soon cost extra to use, so save it for complex tasks that need high-quality results.
2026-07-04
  • Claude Fable 5 (a powerful AI model) is back and available on various platforms, with 50% free usage until July 7th, after which you'll pay extra for more.
  • It's much better at coding tasks than before, so you should review and improve any code you wrote while it was unavailable.
  • You can use the Claude for Chrome extension (a tool that lets the AI control your browser) to test your apps and improve the user experience.
  • After the free usage, it's expensive to use more, but the speaker thinks it's worth paying for the advantage it gives in business and coding.
2026-07-01
  • AI tools like ChatGPT and Claude (popular AI chatbots) have settings that control how much effort they put into answering, which can make a big difference in the quality of their responses.
  • You can choose different "models" (versions of the AI) and "reasoning levels" (how hard the AI tries) to match the task, like a small model with low effort for quick tasks or a large model with max effort for complex ones.
  • For most people, the default settings are too basic for tasks like writing emails or summarizing reports, so adjusting these settings can help you get better results without changing your prompts (the instructions you give the AI).
  • AI providers set low defaults to make the AI respond quickly and to save them money, but you can change these settings to suit your needs, like using a higher reasoning level for tasks that require more thought.
2026-06-28
  • Bite Dance and Alibaba released new video models, including one for real-time interactive avatars (computer-generated characters you can talk to).
  • Stability AI launched a precise 3D model generator, and new top open-source image generators were introduced.
  • OpenAI unveiled GPT 5.6, their most powerful model yet, but it's likely not accessible to the public.
  • Meta released an agentic framework (a system that helps AI improve itself) for creating self-improving datasets.
2026-06-25
  • OpenAI (a company leading in AI development) is secretly testing a new voice model in ChatGPT (a popular AI chatbot) that can understand and respond to human speech more naturally.
  • A new AI model called Sakana Fugu (a tool for coding and developing) by Sakana AI Labs (a lesser-known AI company) is challenging leading models like Fable 5 (a top AI model) and GPT 5.5 (another top AI model) in coding tasks.
  • Sakana Fugu uses a unique approach called multi-agent orchestration (a system where one AI model can coordinate with other AI models to complete tasks) to handle complex coding tasks, making it a powerful tool for developers.
  • The pricing for Sakana Fugu's API (a way for other software to use the AI model) is competitive with other leading models, with rates increasing based on the amount of data processed.
2026-06-22
  • AI can give wrong answers by guessing what you mean, using old info, or looking in the wrong place; tactics include prevention, checking, and protecting.
  • Prevention involves being specific with words (e.g., "highest revenue clients in the last 12 months" instead of "top customers") to avoid vague terms.
  • Checking means having AI provide proof (like a receipt) when it extracts info from documents, so you can verify its accuracy.
  • Protection is for high-stakes tasks, like getting a second opinion from another AI or testing AI on known answers to check its performance.
2026-06-19
  • Claude (an AI tool) can now run tasks continuously without stopping, thanks to features like auto mode (a safety checker that approves safe actions automatically) and {slash} goal (a feature that sets a completion condition for tasks).
  • To handle larger tasks, Claude can use {slash} effort (a setting that increases the time spent on thinking and reasoning) to maintain quality.
  • Tasks can be scheduled to run at specific intervals using {slash} loop (for short tasks) or {slash} routines (for longer tasks), with {slash} goal determining when the task is complete.
  • These features work together to automate workflows, allowing Claude to complete tasks without constant user input.
2026-06-16
  • A new tool called SkillSmith (a plugin for Claude, an AI assistant) lets you build AI agents (automated workers that can perform tasks) in minutes using simple text files, not complex code.
  • Claude skills (the new way to build agents) use plain text files to define triggers, context, frameworks, tasks, and templates, making them easier to create and understand.
  • SkillSmith offers four main functions: turning ideas into specs, building skills, distilling long-form content into frameworks, and auditing existing skills.
  • Appify (a marketplace for AI actors) and its MCP (a bridge between Claude and other software) help find and use the right AI actors for your workflows.

Key points

What it is

  • AI model optimization is like tuning a car engine to make an AI model run faster and cheaper while keeping its core abilities.
  • It's crucial for deploying AI models in real applications, as unoptimized models are often too slow or expensive.
  • Optimization is a big part of AI development, with the 10/20/70 rule stating that 70% of AI success comes from people and processes.

How to use it

  • Define your goal as a clear, repeatable workflow and use a YAML recipe to separate decisions from execution.
  • Use an iterative optimization loop: the AI executes steps, tests results, and repeats until tests pass or time runs out.
  • Be explicit about what the AI should do, avoid vague goals, and use evaluation methods to catch bad patterns or biases.

Watch out for

  • Avoid giving the AI vague, open-ended goals, as this can lead to poor results.
  • Be aware of models learning bad patterns from training data and use evaluation methods to catch these issues.
  • Guard against sycophancy, where models tend to agree with you, by asking the model to interview you to surface counterpoints.

Tools named

  • Claude (AI assistant for planning and execution), Codex (AI tool for code generation), Claude Fable (AI for building large ideas autonomously), YAML (structured config file format)

Lesson 1: What is AI Model Optimization and why it matters

AI model optimization means making a trained AI model run faster and cheaper without losing its core abilities. Think of it like tuning a car engine to get better gas mileage while keeping the same horsepower. One expert showed this by comparing 20 days of compute against 7 hours of compute for the same task, proving that optimization can dramatically reduce the time and cost needed to use a model.

Why does this matter for AI development? First, optimization lets you deploy models in real applications. Without it, models are often too slow or expensive to run on actual products. Second, the 10/20/70 rule says only 10% of getting AI right is the algorithm (the model itself). Another 20% is the technology and infrastructure, while 70% is people and processes. So if you spend all your time picking the perfect model and none on optimizing its deployment or fixing your internal processes, you will fail.

The best approach combines optimization with domain expertise (specialized knowledge in your field). Your company needs to judge what good AI quality looks like for your specific problem. Optimization tools let you test different approaches and pick the one that balances speed, cost, and accuracy for your use case. Remember that AI is just a tool — the person using it solves problems. Optimization makes that tool practical.

Sources

Lesson 2: How to use AI Model Optimization: step-by-step

To optimize an AI model step by step, start by defining your goal as a clear, repeatable workflow. Use a YAML (structured config file) recipe to separate decisions from execution, letting Claude or Codex act as the chef. First, run an advanced plan mode where the AI asks deep questions to fully understand your task before building.

Next, set up a big loop: the AI executes steps, then iterates until tests pass. For example, with Claude Fable, use this loop to build a large idea autonomously. If results aren't good enough, apply a crude optimization technique — ask the agent to generate a new prompt, test if it does better, take bits of that, and merge them into a new prompt. Keep repeating until you reach a prescribed score or run out of time.

Key shifts: stop telling the AI it's an expert — be literal instead. Be explicit about which files to reference. Use a rescue command to pit models against each other for fresh perspective. Route tasks to the right model, since open-source options are often cheaper than Claude. Finally, avoid defaulting to AI for every problem — first figure out the simplest step manually, understand context, then optimize with loops that parallelize workflows for maximum autonomy.

Sources

Lesson 3: Best practices and pitfalls

A common mistake in AI optimization is giving the model a vague, open-ended goal. Imagine telling someone, “Go make our app faster,” without specifying which part, how, or what success looks like. That is an open-ended problem, and even advanced models struggle with it. To get better results, stop telling the AI it is an expert in the field—this can degrade performance. Instead, be explicit about which files it must reference and what the desired outcome is.

Another pitfall is ignoring the output quality from your input. If you hand a messy, underspecified prompt to a model and say “fix it,” you will get inconsistent results. A better approach is to use an iterative optimization loop. The current state-of-the-art technique is relatively crude: the model generates a new prompt, you test it, keep the improvements, and repeat until you reach a target score or run out of time.

Be aware that models can learn bad patterns from training data, such as portraying AI as deceitful or self-preserving. Use evaluation methods to catch this, though current tests are not sophisticated enough to rule out hidden failures. Also, guard against sycophancy—models tend to agree with you. Ask the model to interview you (a “grill me” approach) to surface counterpoints. Finally, fine-tuning for specific reasoning skills costs around $20,000 annually for enterprise use, so prioritize iterative prompting before investing in fine-tuning.

Sources