Models & Comparisons

Scaling Model Infrastructure

Last updated 2026-09-22

What's new

2026-09-22
  • Codex (a tool by OpenAI that helps you do tasks with AI, like writing, designing, or coding) can be used to build skills, create branded deliverables, and even automate tasks, all without needing a technical background.
  • The Codex desktop app (a program you download to use Codex easily) is recommended for a consistent experience, and it uses the same subscription as ChatGPT (a popular AI chatbot), so you won't need a new account.
  • Codex is more powerful than Work (a tool for non-technical knowledge work) and can do everything Work can do, plus more, making it a better investment for learning and using in the long run.
  • The course will teach you how to use Codex effectively with natural language (regular English, not code) and explain core concepts simply, helping you become a pro AI builder.
2026-09-16
  • Anthropic (a company making AI tools) introduced a new idea called "tokens with jobs," where AI tasks are divided among different groups of tokens (small pieces of data AI uses to learn and work) to improve performance.
  • They tested three strategies: "advising" (some tokens give advice to others), "grading" (some tokens check and score the work of others), and "dreaming" (some tokens learn from past work and improve future tasks).
  • In experiments, these strategies outperformed a simple execution-only approach, showing that dividing tasks among tokens can lead to better results in AI systems.
  • To ensure fair comparison, they fixed the budget (total tokens used) and found that strategies with specific token jobs performed better than just using more tokens.
2026-08-25
  • AI tools are becoming faster and cheaper, with some now offering physical buttons for instant use—no tech skills needed.
  • OpenAI (a leading AI research lab) values quick testing and user feedback, letting teams experiment and release tools fast.
  • Google once had a similar tool (LM chat) but hesitated to release it publicly before ChatGPT arrived.
  • To stay innovative, companies must disrupt their own products before competitors do—even if it means shifting focus from profitable projects.
2026-08-22
  • A company called Lease End (a service that helps people buy their leased cars) built a chatbot (a computer program that talks to customers) using AI to improve customer interactions, which initially brought in $12 million in revenue.
  • They started with a system that tried to understand customer messages by comparing them to past messages, but it struggled with the nuances of conversation, so they switched to a method called fine-tuning (teaching an AI model to do a specific task better by showing it lots of examples).
  • Fine-tuning helped, but it also created new problems, like the chatbot calling customers at the wrong time or being too eager to call, which frustrated customers and caused missed opportunities.
  • The fine-tuning process was complex and time-consuming, involving gathering examples, labeling them, and testing the AI, which could take about a week and often led to new issues popping up.
2026-08-16
  • AI agents (automated helpers) are getting better at doing tasks for you on the web, but they still struggle because the web wasn't designed for them, making automation difficult.
  • The main issue isn't the AI models (the brains of the agents) anymore, as they've improved a lot and can now handle complex tasks, but agents lack the right tools (called harnesses) to interact with the web effectively.
  • Harnesses (the systems and tools around the AI model) can greatly improve how well agents work, and companies like Cursor and Browserbase are already building these for specific tasks like coding and web browsing.
  • Building a good harness is an engineering problem that any company can tackle, not just big labs, and it can help get more out of the AI models we already have.
2026-08-13
  • New startup N gram is working on "scaling compute on context" (helping AI understand and use personal or specific information, not just public data).
  • Current AI models, like ChatGPT, can't learn or use new information after they're trained, or understand niche topics well.
  • N gram aims to give AI models the ability to learn deeply, like a human expert, using personal or specific data, not just public data.
  • The team is exploring ways to apply "scaling" (making AI better by giving it more data, time, or capacity) to personal or specific data.
2026-08-10
  • Klein, an open-source coding tool (a program that helps write code), started as a way to make using AI for coding more affordable and trustworthy, but AI's impact on software development has led to a decline in open-source community trust.
  • Some open-source projects, like Zig (a programming language), have banned AI use to protect their community and maintain trust among contributors.
  • AI-generated bug reports and pull requests (suggestions to improve code) are causing issues for projects like Kurl, leading to considerations of shutting down bug bounty programs (rewards for finding bugs).
  • The rise of AI in coding has also increased the risk of supply chain attacks (hacks that compromise software dependencies), as seen in the recent Light LLM (a Python package for running AI models) incident.
2026-08-07
  • NVIDIA's NeMo Triton models are now available, allowing users to customize AI models for specific tasks, which is crucial as we move towards using multiple AI models (multi-model world) simultaneously.
  • Cognition's new tool, Fusion (a model router), improves performance by letting smarter AI models (like Fable) delegate tasks to cheaper, smaller models, reducing costs by 40% while maintaining high-quality results.
  • Model routing, a relatively new concept, helps determine which AI model to use for specific tasks, balancing cost and performance, especially important as advanced AI tools become more expensive.
  • The discussion highlights the need for better solutions in model routing, with opportunities for startups and companies to innovate in this space.
2026-08-01
  • **Data quality is a "compute multiplier" (a way to get more out of your computer power)**: Better data can make AI models perform as if they had more computing power, or achieve the same performance with less.
  • **AI models are using more computing power than ever**, especially reasoning models, which use eight times as many "tokens" (pieces of data) as non-reasoning models.
  • **There's a growing scarcity of computing power for AI**, with some companies even limiting access to their AI services due to high demand.
  • **To improve AI performance, focus on data quality**: This means using relevant, diverse, and informative data, and combining different data sources effectively.

Key points

What it is

  • Scaling model infrastructure means building systems that let AI models run reliably for many users, handling traffic without breaking down.
  • It's about managing the systems around AI models, like operational overhead, inference complexity, and costs, not just the models themselves.
  • Infrastructure affects costs, speed, and reliability, and can impact how well AI agents (AI that uses tools and takes actions) perform.

How to use it

  • Start by hosting your models on a platform like Hugging Face (a hub for sharing AI models) to make them available to developers worldwide.
  • Use systems like Kubernetes (a system that manages containers) with horizontal pod autoscaler (a tool that automatically adds or removes servers based on demand) to handle traffic spikes automatically.
  • Plan for growth from the start, choosing infrastructure that can handle expected increases in users and models.

Watch out for

  • Avoid assuming bigger models are always better; focus on making models behave properly for your specific use case before scaling.
  • Be aware of costs and security risks when using cloud services, such as escalating bills and misconfigured permissions.
  • Remember that scaling involves more than just bigger models; it also includes adding tools, memory, search, and verification to make models more useful.

Tools named

  • Hugging Face (a hub for sharing AI models), Kubernetes (a system that manages containers), MongoDB Atlas (a database service), VLM (a serving tool).

Lesson 1: What is Scaling Model Infrastructure and why it matters

Scaling model infrastructure means building the systems that let AI models run reliably for many users. When Hugging Face grew to serve millions of developers, they had to make careful architectural decisions about how to handle that traffic without breaking down. The hard part of building AI applications isn’t the model itself—it’s everything around it, like operational overhead, inference complexity, and costs that become unpredictable as you scale.

Why does this matter? For agentic AI (AI that uses tools and takes actions), the model is only one part of the machine. Once an AI agent starts calling APIs, updating databases, or remembering things, its behavior depends on the whole harness (the system of tools and services around it). A UC Berkeley paper argues that model scaling alone isn’t enough anymore; you also need system scaling to make agents work well.

Infrastructure also affects costs and speed. Inference-optimized setups (systems tuned for running models efficiently) let you handle millions of requests with better throughput and latency. Cloud platforms handle traffic spikes automatically, but you depend on internet access and risk escalating bills. Self-hosted setups give more control but require you to manage everything yourself. Resource constraints like energy, chips, and data centers limit how far you can scale. Teams often spend more time managing infrastructure than building features, so choosing the right infrastructure is a key part of AI development success.

Sources

Lesson 2: How to use Scaling Model Infrastructure: step-by-step

Scaling model infrastructure means making your AI setup handle more users and models without crashing. Start by hosting your models on a platform like Hugging Face (a hub for sharing AI models). When you upload a model, it becomes available to developers worldwide. To serve many users at once, you need a system that automatically adjusts to traffic.

Think of it like cooking: running a model on your laptop is like cooking for one person. Serving millions of users is like running an industrial kitchen — you need infrastructure that can handle many requests simultaneously. Hugging Face uses Kubernetes (a system that manages containers) with horizontal pod autoscaler (a tool that automatically adds or removes servers based on demand). When traffic spikes, new servers start up automatically; when it drops, they shut down. This keeps the system healthy without manual work.

For example, Hugging Face grew from 20,000 models to 3 million in a few years — a 150x increase. They handle this by routing requests from the front end to the hub API (the interface that processes requests), which runs on Kubernetes. The source of truth for all data lives in MongoDB Atlas (a database service).

When using local models, you can download an open-weights model (a model with publicly available parameters) from Hugging Face. For larger deployments, you can use a serving tool like VLM to let multiple people hit the same model at once. If you use cloud services, you get automatic scaling for traffic spikes, but watch out for costs and security risks like misconfigured permissions. You can also self-host for more control, but you handle scaling yourself.

Sources

Lesson 3: Best practices and pitfalls

Scaling AI model infrastructure means preparing for explosive growth. Hugging Face grew from 20,000 to 3 million public models in just a few years—a 150x increase—triggered by every major release like Llama or DeepSeek. When you serve millions of models, you must handle sudden traffic spikes without your application slowing down or crashing.

A common pitfall is assuming bigger is better. Teams often drop in a larger model when performance lags, expecting it to be smarter. But that creates costly scaling problems. Instead, focus on making models behave properly for your specific use case before scaling. Also, don't choose one model and stop. Most teams (87%) actually use multiple models, so your infrastructure must support a mix of options.

For serving, think about your audience size. Running models on a personal laptop is like cooking for a party of one—you don't worry about scaling inference (the process of generating predictions). But serving many users simultaneously requires a tool that lets multiple people hit the same models at once, like an industrial kitchen. Cloud services handle traffic spikes automatically without lifting a finger, which sounds great until the bill arrives after a busy month. Costs escalate quickly with pay-as-you-go pricing.

Best practices include planning for growth from the start. If you expect a 100x increase, your architecture must handle it. For production, prioritize reliability and speed, and watch for security risks like misconfigured permissions or exposed storage buckets. Finally, remember that scaling is not just about bigger models—it also means adding tools, memory, search, and verification to make models more useful. Both matter, but they are different claims.

Sources