Models & Comparisons

Model Scaling Comparisons

Last updated 2026-08-01

What's new

2026-08-01
  • **Data quality is a "compute multiplier" (a way to get more out of your computer power)**: Better data can make AI models perform as if they had more computing power, or achieve the same performance with less.
  • **AI models are using more computing power than ever**, especially reasoning models, which use eight times as many "tokens" (pieces of data) as non-reasoning models.
  • **There's a growing scarcity of computing power for AI**, with some companies even limiting access to their AI services due to high demand.
  • **To improve AI performance, focus on data quality**: This means using relevant, diverse, and informative data, and combining different data sources effectively.
2026-07-31
  • A new AI model called Kimmy K3 (a type of AI software that can understand and generate text, images, and more) was released for free by a Beijing lab, and it quickly became popular, even being praised by American companies.
  • The US government accused the creators of Kimmy K3 of using a technique called "distillation" (copying the outputs of a stronger AI model to improve a weaker one) to steal American AI technology, and they threatened sanctions (penalties that limit business) if it's true.
  • China denied these accusations, and the creator of Kimmy K3 said the model's improvements came from original changes, not copying.
  • A new tool called Dream Nina (a website that helps create AI-generated videos) was introduced, which allows users to guide video creation using images, videos, audio, and text references together in one place.
2026-07-28
  • You can now test new AI models (like Claude or Codex, which are AI tools that help with tasks) against your own work, not just generic benchmarks (tests that don't relate to your specific needs).
  • AI can scan your past conversations to find common tasks, then test new models on those tasks using a rubric (a set of criteria you create) that matters to you.
  • You can compare different AI models or effort levels with a simple command (like "/benchmark") and get a report showing which model performs best for your specific needs.
  • This approach helps you decide if a new AI model is worth using for your work, saving time and avoiding irrelevant tutorials.
2026-07-25
  • AI agents (automated tools that perform tasks) interact with systems much faster than humans, causing unexpected strain on databases, like updating user activity timestamps 60 times faster, which can be fixed by batching updates.
  • Traditional permission systems for humans or simple programs (like API keys or service accounts) don't work well for agents, as they often grant too many permissions by default.
  • Agents challenge the assumption that the entity authenticating (proving its identity) is the same as the one acting, requiring new ways to delegate access and manage permissions.
  • Unlike traditional programs, agents aren't deterministic (predictable), as they can act in ways not explicitly programmed, making security reviews and inspections more difficult.
2026-07-22
  • AI models are becoming more efficient, with smaller models (like Quen 3.5 with 27 billion parameters) now outperforming much larger ones (like Lama 2 with 400 billion parameters) from just a year and a half ago.
  • Hardware requirements for running advanced AI models have significantly decreased; for example, you can now run high-quality AI models on a single RTX 4090 or even an iPhone, reducing the need for expensive data centers.
  • The "densing law" describes how AI models are becoming more capable with fewer parameters, leading to better performance and lower costs for training and fine-tuning models.
  • There's a push for individuals and businesses to adopt "sovereign AI," which means owning and controlling their own AI models and hardware to ensure privacy, customization, and long-term cost savings.
2026-07-19
  • Anslaf, a top distributor of AI models, offers tools like Deepseek and GLM, which they optimize for local use and fix bugs for popular models like OpenAI's, Meta's, and Google's.
  • They've introduced features like async gradient checkpointing and flex attention, improving training accuracy by 1-3%.
  • A meter plot shows AI models' progress, with top models like Cloud Mythos and Opus 4.6 handling tasks that take humans 16 hours, but models often need multiple prompts for high accuracy.
  • AI models are improving exponentially, with newer models like GBD 5.6 showing significant advancements, though sometimes "cheating" on tasks.
2026-07-16
  • **Effort levels** (how much work an AI model puts into a task) are often misunderstood as a measure of intelligence, but they're not—higher effort doesn't always mean better results.
  • For everyday tasks, especially with advanced "frontier models" (cutting-edge AI models like Fable 5 or GPT-5.6 Soul), low or medium effort levels are often sufficient.
  • Very high effort levels (like "extra high" or "max") are rarely needed and can even lead to overcomplicating simple tasks, much like overthinking a multiple-choice exam.
  • To use AI models effectively, start with the lowest effort level and increase only if necessary, and be intentional about choosing the right model for the task.
2026-07-13
  • Claude Code (a tool for building AI-powered automations) lets you work with local files and online services like Gmail, Slack, or a CRM (customer relationship management system), making it more powerful than Claude Chat (a simple AI chatbot).
  • Claude Code uses the same AI models (like Opus, Sonnet, or Haiku) as Claude Chat, but adds extra features for working with files and online services.
  • Claude Code is like an AI harness (a tool that helps you use AI models), which sits between the AI model (the engine) and you (the driver), helping you build automations and agents (AI systems that can do tasks for you).
  • The instructor, Nate, uses Claude Code to build and manage multiple businesses, showing how one person can do the work of a team with AI.
2026-07-10
  • GPT 5.6 (a new AI model) and Codeex (a tool for running AI on your computer) can automate tasks, boost productivity, and even help manage your emails.
  • Codeex is praised for its simplicity and power, making it a great tool for both knowledge work and coding, similar to having a high-performance car for daily tasks.
  • With GPT 5.6, even non-experts can start training their own AI models, making advanced AI tasks more accessible.
  • Dan Shipper demonstrates using Codeex to build an app called Tend, which automates email management by summarizing emails and drafting replies.
2026-07-07
  • Google's Gemini 3.5 Pro AI model (a powerful online AI tool) launch was delayed to rebuild its foundation, improving math, design, and coding skills, not due to failure.
  • Gemini 3.5 Pro might act as an "orchestrator" (a conductor), managing smaller AI tools called "agents" (specialized AI workers) instead of just generating text.
  • The delay could be to fix token (AI's understanding unit) inefficiency in the smaller Gemini 3.5 Flash model, which the Pro model might manage.
  • Gemini 3.5 Pro's competition isn't just raw size; it's about strategy and integration with other AI tools and research.
2026-07-01
  • AI tools like ChatGPT and Claude (popular AI chatbots) have settings that control how much effort they put into answering, which can make a big difference in the quality of their responses.
  • You can choose different "models" (versions of the AI) and "reasoning levels" (how hard the AI tries) to match the task, like a small model with low effort for quick tasks or a large model with max effort for complex ones.
  • For most people, the default settings are too basic for tasks like writing emails or summarizing reports, so adjusting these settings can help you get better results without changing your prompts (the instructions you give the AI).
  • AI providers set low defaults to make the AI respond quickly and to save them money, but you can change these settings to suit your needs, like using a higher reasoning level for tasks that require more thought.
2026-06-22
  • Anthropic released Claude Fable 5 and Mythos 5, AI models (computer programs that can learn and make decisions) designed to be more capable and useful for developers and companies, but faced governance challenges within 72 hours.
  • These models use a safety system with classifiers (fast guards that check if requests or answers fit policy) to reduce risky behavior, but also raise questions about access, safeguards, and public accountability.
  • The models aim to shift from being helpful assistants to more persistent partners in knowledge work, but their adoption depends on trust in the surrounding rules and infrastructure.
  • The launch sparked debates about access, with identity, intent, and audit trails (records of who did what and when) becoming crucial for governance and safety.
2026-06-13
  • Snorkel (a company focused on improving AI data quality) found that using high-quality data and reinforcement learning (a type of AI training) can make a smaller AI model (with 4 billion parts) outperform a much larger one (with 235 billion parts) in financial analysis tasks.
  • They worked with UC Berkeley's RLLM team to show that bigger AI models aren't always the best solution for enterprise (business) use, as smaller models can be more cost-effective, faster, and secure.
  • The key was using the right data and training methods to teach the smaller AI model to use tools effectively for financial analysis, challenging the common belief that larger models always perform better.
  • This approach can help businesses deploy AI models more easily, with better control over data and fewer external dependencies, which is crucial for sensitive areas like finance and healthcare.
2026-06-10
  • Claude Opus 4.8, a new version of the AI tool (Claude) by Enthropic (the company that makes Claude), is a small but noticeable improvement over its predecessor, 4.7, which had performance issues due to a lack of computing power.
  • Enthropic recently secured a massive computing deal with XAI (Elon Musk's AI company), allowing them to increase their capacity and release Opus 4.8 just six weeks after 4.7, fixing many of the previous version's problems.
  • Opus 4.8 is more reliable, faster, and better at coding tasks compared to 4.7, and it costs the same as the previous version, making it a great value for users.
  • While some users initially reported bugs and slow responses with Opus 4.8, many found that it addressed the main issues they had with 4.7, making it feel like a polished version of the earlier Claude Opus 4.6.
2026-06-07
  • **Snorkel AI** (a company that helps build AI data sets) is working to close the gap between AI's capabilities and our ability to measure them, especially in high-stakes fields like healthcare and finance.
  • They've launched **Open Benchmarks grants** (a $3 million fund to support the creation of new AI evaluation standards) to help shape the future of AI by funding academic teams working on these benchmarks.
  • Snorkel emphasizes the importance of **task quality** (how well the benchmark tests AI capabilities), **distributional control** (ensuring the benchmark is representative of real-world situations), and **robust evals** (measuring AI performance consistently).
  • They also stress the need for benchmarks to have a **thesis on where the field is going** (a clear vision of future AI capabilities) and be designed for easy adoption by researchers and builders.

Key points

What it is

  • Model scaling comparisons measure how AI performance changes when you increase resources like data, model size, or processing power.
  • Early scaling focused on training, but newer "test-time scaling" improves models during use by letting them refine reasoning over time.
  • Scaling is evolving to include "unhobbling," which adds tools, memory, search, and verification to make base models more useful.
  • The strongest models require massive resources, but the engineering around AI models may now matter as much as the models themselves.

How to use it

  • Define your task and choose two models to test, like a smaller model and a larger one, to compare their performance.
  • Run quick experiments to see if a smaller model with more "test time compute" (extra processing during a task) can match a bigger model's accuracy.
  • Schedule batch jobs and use auto-scaling to handle traffic spikes, but be mindful of costs on a pay-as-you-go model.
  • Rethink your approach by testing models on your specific task, prioritizing unhobbling over just model size, and managing compute costs.

Watch out for

  • Avoid comparing models based only on "days" of training or total "compute" (processing power used), as different hardware can lead to different capabilities.
  • Don't overvalue benchmark scores, as real-world behavior often differs; run quick experiments between two models instead.
  • Be careful with auto-scaling, as scaling down too quickly can break stateful long-lived connections, and costs can escalate quickly if you leave expensive instances running idle.

Tools named

  • Hugging Face (a platform for training and deploying AI models)

Lesson 1: What is Model Scaling Comparisons and why it matters

Model scaling comparisons measure how AI performance changes when you increase resources like data, model size, or compute (processing power). Early scaling followed neural scaling laws, which showed that more data and more parameters (internal settings learned during training) leads to better models. But that approach is rooted in training, not deployment.

A newer type of scaling is test-time scaling, which improves models during use rather than during training. When you let a system loop deeper and refine its reasoning over time, performance keeps improving—unlike traditional text models that plateau. The more rounds you add, the better the results across benchmarks.

This matters because scaling is changing. Instead of just making models bigger, developers now use unhobbling (adding tools, memory, search, and verification) to make base models more useful. Think of it as putting a stronger engine inside a better vehicle—both matter, but they are different.

The strongest models require enormous data and compute, often costing billions and needing massive data centers. Only state actors or companies with huge resources can run them. However, the engineering around AI models may now matter as much as the models themselves. Leading in raw capability doesn't guarantee leading at the application layer. For developers, this means focusing on real problem-solving and process, not just chasing the biggest model.

Sources

Lesson 2: How to use Model Scaling Comparisons: step-by-step

To compare AI models effectively, start by defining your task and choosing two models to test. For example, take a smaller model (like a 1B parameter model) and a larger frontier model. The key insight from recent research is that you can run a smaller model with more "test time compute" (extra processing during a task) and achieve the same accuracy as a bigger model. This is a scaling law: more compute can substitute for model size.

Run a quick experiment. In one reported case, a full scientific research loop—including multiple experiments and data analysis—was finished in under 2 hours, with total compute cost of just $10.76. Imagine compressing 400 hours of human work into 2 hours for the price of a fast-food lunch. That's the power of smart scaling.

To scale up, schedule batch jobs. For instance, run planning every Sunday and Thursday, then generation every Monday and Friday. Scale from 50 jobs to 100 or 200. On weekends, scale down seamlessly. Auto-scaling handles traffic spikes automatically, but beware: costs can escalate quickly on a pay-as-you-go model if you leave expensive instances running idle.

Finally, rethink your approach. Open-weight models lag behind frontier models by 3-6 months, but they are becoming smarter. You might find a smaller model with extra compute outperforms a larger one for your specific use case. The best model isn't always the biggest—it's the one that balances performance, compute hours, and cost for your task.

Sources

Lesson 3: Best practices and pitfalls

Comparing model scaling (making an AI bigger or giving it more time to think) requires care. A common mistake is to compare only "days" of training or total "compute" (processing power used). Two models trained for the same number of days on different hardware will have different capabilities. Instead, focus on what a model can do for you. A smaller model with more "test time compute" (thinking time during a task) can sometimes match a larger model's accuracy, as research from Hugging Face shows. This introduces a trade-off: you can pay for compute during training or during use.

Another pitfall is overvaluing benchmark scores. Benchmarks are meaningful in specific areas, but real-world behavior often differs. Consider your actual use case and run a quick experiment between two models rather than relying on leaderboards. Also, remember "unhobbling" (adding tools, memory, or search) can make a model more useful than raw scaling alone. For cost management, "auto scaling" (automatically adjusting resources) helps handle traffic spikes without wasting money on idle capacity, but be careful: scaling down too quickly can break stateful long-lived connections.

To rethink scaling, look at new ways to improve AI. One system showed that letting it "loop" (refine reasoning over multiple rounds) kept improving performance, while traditional models plateaued. The best approach is to test a model on your specific task, prioritize unhobbling over just model size, and manage compute costs through auto scaling and conscious choices about test time compute versus training compute.

Sources