Models & Comparisons

LLM Model Selection

Last updated 2026-09-25

What's new

2026-09-25
  • Claude Opus 5.5 (a new AI model from Anthropic) is faster, cheaper, and outperforms previous models in many tasks, like coding and knowledge work (tasks like writing emails or creating PowerPoint presentations).
  • It excels in coding tests, scoring 66.4% on Terminal Bench 4.0, a significant improvement over other models.
  • Opus 5.5 also performs well in real-world knowledge work tasks, with a high score of 1846 on GDP val version 2.1, a benchmark for testing skills like data entry and word processing.
  • The model is 30% faster and has a 20% price reduction compared to its predecessor, Opus 5, making it a more efficient and cost-effective choice.
2026-09-10
  • **LLM inference (using AI to generate or analyze content like text, audio, or video) costs are rising**, with businesses needing to optimize or reduce token (small pieces of data) usage to manage expenses.
  • **Memory usage increases with more input tokens (data)**, which can lead to out-of-memory errors, especially with longer context lengths.
  • **Inference speed varies**, with the time to generate the first token often being slower than subsequent tokens, impacting user experience.
  • **New tools and optimizations are emerging** to help manage these challenges, with benchmarks and guides available to evaluate different solutions.
2026-09-04
  • Claude Code (a coding assistant tool) now offers a free model with 1.5 billion tokens monthly, a significant cost-saving for users previously paying for credits.
  • Omni Route (a laptop app) aggregates free API keys from multiple providers, automatically switching between them to maximize free usage, avoiding the hassle of manual configuration.
  • This setup allows users to keep the Claude Code interface while changing the underlying AI model (the "brain") to a free one, making it versatile for other tools like Deep Sea Carnage (another AI coding tool).
2026-08-31
  • AI agents (computer programs that can do tasks for you) are changing fast, and companies must adapt quickly to keep up.
  • Instead of training your own AI models (the brain of AI agents), it's now better to use advanced, ready-made ones from companies like OpenAI.
  • The old advice of building scalable, reliable platforms and optimizing them may not work well with today's rapid AI changes.

Key points

What it is

  • LLM (large language model) model selection is choosing the best AI text generator for a specific task, balancing quality, speed, and cost.
  • It's a three-way tradeoff: better quality often means slower speed or higher cost, and vice versa.
  • You should only use an LLM when a task can't be solved with ordinary code, as LLMs add cost and risk without value otherwise.

How to use it

  • Start with the simplest model that meets your needs, and only switch to a more powerful (and expensive) model if necessary.
  • Use a model picker (a dropdown menu in tools like Claude Code or LM Studio) to easily switch between models for different tasks.
  • Test models on your specific task with your data, rather than relying on old benchmarks or leaderboard pass rates.

Watch out for

  • Don't rely on one model to grade its own work—use a separate judge (a different model or human review) to evaluate results.
  • The model is not a mind reader—it's only as good as the context you give it, so provide clear, detailed instructions.
  • Avoid the pitfall of trusting leaderboard pass rates, which often claim 80%+ success but rarely reflect real-world performance.

Tools named

  • Claude Code (a coding assistant tool), LM Studio (a local LLM management tool), Qwen (a large language model), DeepSeek (a large language model), Opus (a large language model), Sonnet (a large language model), Ollama (a tool for running models locally), Kimi K 2.6 (a coding-focused LLM), Qwen 3.5 9 billion parameter (a reasoning-focused LLM), AI model cheat sheet bundle (a tool for understanding model capabilities).

Lesson 1: What is LLM Model Selection and why it matters

LLM model selection is the process of choosing which large language model (a type of AI that generates text) to use for a specific task. It matters because it is a three-way tradeoff among quality (how accurate the output is), latency (how fast the model responds), and cost (how much each call to the model costs you in money and computing power).

The right choice depends on your goal. For a deterministic task (a problem with a fixed, predictable answer) that ordinary code can solve, you should not use an LLM at all—adding a model creates cost and risk without value. When you do need a model, start with the simplest shape. An augmented LLM (a model call supported by tools, retrieval, or memory) adds extra capabilities without jumping to a fully autonomous agent. A workflow is more deterministic, meaning it follows a predefined path with predictable results and gives you more control. In contrast, an agent (an LLM in a loop that calls tools until done) is more flexible but less predictable.

Stay model agnostic (able to switch between different model providers). This lets you retain ownership of your data and choose models that improve your accuracy while lowering your cost, rather than being locked into one vendor. When evaluating options, know that models tend to focus on the initial inputs and the last inputs, often removing the in-between context, so your results are only as good as the context you provide. Plan to optimize your feature later, hill climbing (iteratively improving) on an eval to tune quality, speed, and cost.

Sources

Lesson 2: How to use LLM Model Selection: step-by-step

Choosing the right LLM (large language model) starts with a model picker, a dropdown menu inside tools like Claude Code or LM Studio. Your default model should handle everyday tasks, like simple code edits or answering questions. For example, set your picker to Qwen or DeepSeek for routine work—this cuts costs significantly. When something hard breaks, such as a multi-file bug, a vision task (image understanding), or a workflow heavy with MCP (tool connections), manually switch the picker to a stronger model like Opus or Sonnet. You hit the slash model command in Claude Code, then pick a model from the list.

Think of model selection as a three-way tradeoff among cost, speed, and capability. Start with the simplest model that can meet your requirement. If ordinary code solves a task deterministically (with predictable output), adding an LLM creates cost and risk without value. The model is not a mind reader—it's only as good as the context you give it, so provide detailed instructions. You can also run models locally via Ollama, which places them in the same dropdown as Claude. For coding, try Kimi K 2.6; for reasoning tasks, use the Qwen 3.5 9 billion parameter model. When evaluating, don't rely on old benchmarks—test models on your specific workload, using tools like the AI model cheat sheet bundle to understand what each model excels at. The key is to adjust your picker based on task difficulty, not to stick with one model for everything.

Sources

Lesson 3: Best practices and pitfalls

Choosing an LLM (large language model) is a three-way tradeoff among quality, cost, and agentic capabilities (abilities like tool calling). Quality dominates model choice for most teams, followed by cost and tool use. Open versus closed source matters to only 5% of users—don't fixate on it.

Avoid the pitfall of trusting leaderboard pass rates, which often claim 80%+ success. These numbers rarely reflect real-world performance. Instead, test models on your specific task with your data.

Be model agnostic—design your system to switch between providers like Claude, Qwen, or others. This keeps your accuracy improving and costs down. Don't marry a single vendor.

A common mistake is relying on one model to grade its own work. If you ask Claude if its output was optimal, it will almost always say yes. Use a separate judge (a different model or human review) to evaluate results.

Also, remember the model is not a mind reader—it's only as good as the context you give it. Provide clear, detailed instructions.

For local or cheaper options, Qwen 3 is a strong open-source model. Download a couple candidates, try them, and see what works best for your use case before committing.

Finally, consider if you even need an LLM. If ordinary code can solve a deterministic task (one with a predictable fixed outcome), adding a model adds cost and risk without value. Start with the simplest solution that meets your needs.

Sources