Claude API Model Grading
Last updated 2026-09-28What's new
- Claude Code (a coding assistant that runs in your computer's terminal) and Computer Use (tools that let Claude interact with desktop applications) are two new applications built using the Claude API (a way for software to talk to Claude, an AI assistant).
- These applications help you understand how to use the Claude API by showing you how to name inputs, see where the application works, and check the results.
- The goal is to create a smaller, testable boundary between Claude and your application, making it easier to inspect and reuse.
- This module is also a good practice for the Claude Certified Architect exam, which tests your knowledge of using Claude's tools.
- The new lesson focuses on using the Claude API (a tool that lets you interact with an AI assistant) to add citations (references to sources) to answers, making them more transparent and trustworthy.
- You'll learn to modify document messages to enable citations, creating a clear, inspectable process that connects user questions to specific sources.
- Citations work with various document types, not just PDFs, and can be integrated into user interfaces to make source information easily accessible.
- This feature is particularly useful when users need to verify information or explore broader contexts, enhancing the AI's reliability and user trust.
- You'll learn to connect basic ideas into a practical way of thinking, see key details on screen, and end with a hands-on task using the Claude API (a tool that lets you build AI features).
- The lesson focuses on using multiple tools with Claude, turning ideas into real, inspectable actions you can reuse.
- You'll add three main tools to Claude: getting the current date and time, adding time to a date, and setting reminders.
- The lesson teaches a simple pattern for adding new tools once you have the basic tool setup, making it easier to expand your AI features.
- Claude (an AI assistant) can now use tools to get real-time info, like weather, by asking your app (software) for help, keeping data access safe and controlled.
- This tool use happens in a four-step loop: user asks Claude, Claude asks your app, your app gets data, and Claude gives the final answer.
- For example, to set a reminder, Claude needs tools for current time, date math, and reminder setting, showing how tools can handle complex tasks step by step.
- New AI models like Fable 5 and Opus 5 (advanced AI tools) perform better when given a broad goal instead of step-by-step instructions, as they can now handle tasks end-to-end.
- Including the 'why' behind a task helps AI understand intent and make better decisions, similar to how a human would.
- For simple tasks, you can reduce the AI's effort setting to low or medium to save on cost, as Fable 5.1 on low can still compete with older models.
- AI models now prefer a context-first approach, where you provide background information and the task's purpose before specifying the task itself.
- You can now turn any book or PDF into a "skill" (a step-by-step guide) for AI like Claude (a smart assistant that helps with tasks), making it easier to apply the book's techniques in your work.
- A free tool called "book to skill" (a program that converts books into AI-friendly guides) helps Claude understand and use the frameworks, rules, and examples from books, improving its outputs.
- Five recommended books turned into skills include "Breakthrough Advertising" (a classic copywriting guide) and "100 Million Dollar Offers" (a guide to creating effective offers), helping Claude assist with tasks like writing better copy and designing offers.
- Always use this method with books or documents you own to respect copyright laws.
- Claude (an AI assistant) can be used to create marketing materials, like logos and product images, without needing design skills, using tools like Higsfield (a platform with various AI models) and GBT image 2 (an AI image generator).
- Higsfield offers multiple AI models for creating images and videos, and can turn assets into ad creatives using templates, helping businesses create consistent marketing content.
- To create effective marketing, you need to define the "three Ps": the pain (problem) your business solves, the person (customer) who has that pain, and the promise (how your product solves it).
- You can set up brand guidelines in Claude, like color schemes and typography, to ensure all AI-generated content follows your brand's style, making your marketing look professional and cohesive.
- You can automate your business using AI tools like Claude (a type of AI assistant), even if you don't know how to code, and the course provides real-world templates that work.
- The course teaches you to think of AI as a co-worker (someone who helps you with tasks) rather than just a chatbot (a simple question-answer tool), allowing you to assign multiple tasks at once.
- AI can handle various tasks simultaneously, such as answering emails, drafting proposals, and researching, without getting tired or sick, and at a low cost.
- The course emphasizes understanding how AI works to create valuable automations, rather than just using it for generic tasks and getting bland results.
- Claude (an AI assistant) offers three product tiers—Haiku, Sonnet, Opus (like basic, standard, and premium)—each balancing cost, speed, and quality for different tasks.
- Choose Haiku (fast, lower-cost) for simple, high-volume work, Sonnet (balanced) for everyday business tasks, and Opus (slower, more expensive) for complex, high-stakes jobs.
- Match the tool to the job: use chat for quick questions, projects for recurring work, and artifacts (editable results) for shareable deliverables.
- For deep research, use research mode; for internal data, use connectors; for structured outputs, use skills and code execution.
- Anthropic (the company behind the AI model) released a guide for Opus 5, their newest AI model, which tells you to simplify your prompts (the instructions you give the AI) by removing certain lines.
- The guide suggests giving Opus 5 the entire task at once, rather than breaking it into steps, as this newer model works better with complete instructions.
- It's important to clearly state what the AI should not do, to avoid it adding unnecessary work or content, which could waste your time and usage limits (the amount of data you can use with the AI).
- When the AI finishes a task, it will tell you about it, but you should set limits on how much it can write in its reply and in the actual task it's completing.
- Using XML tags (simple markers like `<sales_records>`) in your requests to Claude API (a tool for building applications with AI models) helps organize content, making it easier for the AI to understand and process.
- Providing examples in your prompts, especially when tasks are complex or ambiguous, helps Claude learn the pattern you want it to follow, improving the quality of its responses.
- For tasks with varied inputs or complex outputs, use multi-shot prompting (providing multiple examples) to cover different scenarios, but keep the number of examples practical to avoid wasting space.
- To create effective examples, use clear structure like `sample_input` and `ideal_output`, and explain why the example is a good one to help Claude understand the underlying reasoning.
- AI agents (computer programs that do tasks for you) can now run a business with little human input, acting as a team with one manager and specialists for different tasks like social media or ads.
- These agents use simple English instructions and remember important business details, creating content and handling tasks without sharing information between roles to keep each task specialized.
- The system connects to your existing apps (like email or project management tools) through connectors (links that let the AI access these apps), turning the agents from chatbots into active workers.
- Beginners can use pre-built agent setups (called plugins) to avoid creating each agent from scratch, making it easier to start automating tasks.
- Claude (an AI assistant) can automate boring tasks, like sorting emails into categories (leads, urgent, etc.) and drafting responses, saving you 5-10 hours weekly.
- For leads, Claude can research companies, draft replies, and even schedule meetings using your calendar, streamlining your sales process.
- After client calls, Claude can generate branded PDF proposals with scope, pricing, and signatures, saving time on manual proposal creation.
- This setup can be adapted to various jobs, especially those involving sales, marketing, or regular research tasks.
- OpenAI released ChatGPT 5.6, its most powerful model yet, with three versions (Soul, Tara, Luna) at different price points, where Soul is the most powerful and cost-effective.
- ChatGPT 5.6 can be used within a downloadable app or integrated with Claude, a go-to-market machine (a tool to help businesses grow), to access and compare all three models.
- In a design task, Soul outperformed Tara and Luna in creating a two-page HTML presentation, demonstrating its superior capabilities in understanding and executing complex tasks.
- ChatGPT 5.6 can be connected with Clay (a business growth tool) to find specific company information, enrich it with verified emails, and draft outreach emails, showcasing its potential for business applications.
- Fable 5 (a powerful AI model by Anthropic) was briefly taken offline due to security concerns, as it could identify and demonstrate software weaknesses, but it's now back with stricter safety measures.
- Fable 5 is expensive to use, and some users report that it's being downgraded to a cheaper model (Opus 4.8) for certain tasks, leading to frustration and jokes about its limitations.
- Anthropic has introduced a new safety classifier that blocks potentially risky requests, but it may also flag harmless ones, affecting routine coding tasks.
- Despite its power, Fable 5's usefulness is questioned due to its high cost and the new safety measures that may limit its functionality for some users.
- Claude (an AI assistant) has three main modes: Chat (quick answers), Co-work (file access), and Code (full access, best for building things).
- Opus 4.8 is Claude's most capable model, Sonnet 4.6 for daily tasks, and 4.5 for fast, simple work.
- Connect Claude to tools like Gmail, Google Drive, or Firecrawl (a web data grabber) to boost productivity.
- Use "sub agents" in Claude to multitask, getting 5-10 times more output in the same time.
- You'll learn to improve prompts (the text you give to an AI to get a response) using a measurable process with the Claude API (a tool that lets you interact with the Claude AI model).
- The updated notebook (a file with code and instructions) includes a flexible evaluation pipeline (a series of steps to test and improve your prompts) and a prompt evaluator (a helper tool that automates testing).
- Start with a simple, weak prompt to set a baseline (a starting point for measuring improvement), then make one change at a time and test again to see if the score improves.
- Use the detailed HTML report (a webpage that shows your prompt, the AI's response, and a score) to identify exactly where the prompt failed and make targeted improvements.
- A new open-source AI model called GLM 5.2 (a free, community-developed AI tool) was released, outperforming leading models like GPT and Gemini in various tests.
- To maximize its potential, use GLM 5.2 with frameworks like OpenClaw (a tool that helps AI work on complex tasks) or Zcode (a free, user-friendly AI assistant for Mac, Windows, and Linux).
- GLM 5.2 can create advanced projects, like a 3D interactive Earth model, with some additional guidance, showcasing its impressive capabilities for open-source AI.
- The model can also handle complex, multi-tool tasks, such as creating a promotional video with voiceovers and animations, demonstrating its versatility for everyday workflows.
- A new tool called SkillSmith (a plugin for Claude, an AI assistant) lets you build AI agents (automated workers that can perform tasks) in minutes using simple text files, not complex code.
- Claude skills (the new way to build agents) use plain text files to define triggers, context, frameworks, tasks, and templates, making them easier to create and understand.
- SkillSmith offers four main functions: turning ideas into specs, building skills, distilling long-form content into frameworks, and auditing existing skills.
- Appify (a marketplace for AI actors) and its MCP (a bridge between Claude and other software) help find and use the right AI actors for your workflows.
Key points
What it is
- Claude API Model Grading is a technique where you use a small AI system (called a grader) to automatically judge the quality of Claude's (an AI assistant) output.
- It matters because without grading, you can't consistently know if Claude's answer is good enough for your product.
- Model grading turns quality assessment into a repeatable, automated process that gives you concrete numbers or labels to act on.
- This is crucial for AI development, as you can't manually review every response when building real applications.
How to use it
- Start by building a prompt, a small data set of test cases, and a runner that sends each case through Claude so you can inspect the output.
- To add model-based grading, take the output from your original Claude call and send it to a second Claude call with grading instructions (e.g., "Rate this answer from 1 to 10 for helpfulness and accuracy").
- Never rely on a single model to both build and evaluate its own work, as this can lead to self-assessment bias.
- Integrate grading into a loop: draft, call Claude, score, then repeat to transform experiments into trustworthy application components.
Watch out for
- Relying on a single model to grade its own work can lead to self-assessment bias, where the model always rates its output as great.
- This is especially dangerous if you lack a technical background to verify the quality yourself.
- Don't treat grading as a single, final check; instead, integrate it into a loop for continuous improvement.
- Ensure your evaluation system reliably catches failures to build production-ready systems.
Tools named
- Claude (an AI assistant), n8n (a drag-and-drop tool for connecting apps)
Lesson 1: What is Claude API Model Grading and why it matters
Claude API Model Grading is a technique where you use a small evaluation system—called a grader—to automatically judge the quality of Claude’s output. After your main API call to Claude returns a response, you send that response into the grader. The grader then returns feedback you can actually use, such as a true/false result, a category label, or a score from 1 to 10 (where 10 means the output is strong and 1 means it is weak).
This matters because without grading, you have no deterministic way to know if Claude’s answer is good enough for your product. You might run the same examples repeatedly, but the judgment of quality remains subjective and inconsistent. Model grading turns quality assessment into a repeatable, automated process that gives you concrete numbers or labels to act on.
For AI development, this is crucial. When you build real applications, you cannot manually review every single response. A grader lets you automatically catch failures, track improvements, and enforce quality standards at scale. It bridges the gap between just getting a response from Claude and deploying that response with confidence in a production workflow.
Sources
- 2026-06-09 — Claude Certified Architect Prerequisite Building with the Claude API Part 6 Model Based Grading
- 2026-05-21 — Claude Certified Architect Prerequisite Building with the Claude API Part 1 API Foundations
- 2026-06-01 — Claude Certified Architect Prerequisite Building with the Claude API Part 4 Structured Data
- 2026-05-24 — Claude Certified Architect Prerequisite Claude API Part 2 Conversations & System Prompts
- 2026-05-03 — Copy These Claude Skills, They'll Blow Up Your Business
- 2026-05-13 — Build your first AI agent (Claude Code)
- 2026-05-20 — Claude Skills Fail When You Skip This
- 2026-05-15 — How to Deploy Your Claude Automations (3 Methods)
- 2026-05-16 — The Claude Code + Obsidian Setup That Now Runs My Life
- 2026-03-04 — 🚀Claude Skills Got An UPDATE Check Your Skills Now!
- 2026-05-13 — Why 90 of Your Claude Skills Are Dead Weight
- 2026-04-15 — Which AI coding level are you actually at
- 2026-06-05 — AGI is Here. Anthropic Just Proved It.
Lesson 2: How to use Claude API Model Grading: step-by-step
### How to Use Claude API Model Grading Step by Step
Model grading is a small evaluation system that receives the output from your original Claude call and returns feedback you can actually use. That feedback might be true or false, a category label, or a number. The common pattern is a score from 1 to 10, where 10 means the output looks strong and 1 means weak.
Start with the prerequisite: build a prompt, a small data set of test cases, and a runner that sends each case through Claude so you can inspect the output. This makes prompt changes repeatable instead of relying on one favorite example. But this still leaves you doing the judgment by eye.
To add model-based grading (automated quality scoring via Claude), take the output from your original Claude call and send it to a second Claude call with grading instructions. For example, after Claude answers a customer question, send that answer back with "Rate this answer from 1 to 10 for helpfulness and accuracy." The grader returns the score automatically.
A key warning: relying on a single model to grade its own work creates a problem. When you ask Claude "was this the optimal path forward?", it will say it was great no matter what you did. This is especially dangerous if you do not come from a technical background because you cannot verify the quality yourself.
The practical workflow from the Claude Certified Architect prerequisite content: draft a prompt, build data, call Claude, score the results, then repeat. This turns API controls into practical application behavior you can trust.
Sources
- 2026-05-21 — Claude Certified Architect Prerequisite Building with the Claude API Part 1 API Foundations
- 2026-06-01 — Claude Certified Architect Prerequisite Building with the Claude API Part 4 Structured Data
- 2026-06-09 — Claude Certified Architect Prerequisite Building with the Claude API Part 6 Model Based Grading
- 2026-05-24 — Claude Certified Architect Prerequisite Claude API Part 2 Conversations & System Prompts
- 2026-03-11 — Google's New Model + Claude Code Just Changed RAG Forever
- 2026-06-05 — Claude Certified Architect Prerequisite Claude API Part 5 Prompt Eval to Running the Eval
- 2026-04-13 — Claude Code + Google Stitch (Build a Site in 30 Min)
- 2026-05-03 — Copy These Claude Skills, They'll Blow Up Your Business
- 2026-01-21 — Master 95% of Claude Code in 36 Mins (as a beginner)
- 2026-02-11 — Turn Any Website Into LLM Ready Data INSTANTLY
- 2026-06-01 — Anthropic Just Told You How to Fix Claude Code (Most People Wont Listen)
- 2026-06-05 — I Updated grill-me And Solved Claude Code
Lesson 3: Best practices and pitfalls
Model grading (using an AI to judge its own output) has a major pitfall: self-assessment bias. When you ask Claude to grade its own work, it almost always says the output was great, regardless of actual quality. This is especially dangerous if you lack a technical background to verify the code yourself. To avoid this, never rely on a single model to both build and evaluate its own work.
A better approach treats grading as a separate, deterministic step. Build a small evaluation system (a grader) that receives the output from your original Claude call and returns usable feedback—like a score from 1 to 10, a true/false label, or a category. This separates the generation and judgment functions. The prerequisite workflow involves: drafting a prompt, building a small data set of test cases, and using a runner to send each case through Claude. This makes prompt changes repeatable, but still leaves you judging by eye.
Model grading replaces that eye-based judgment with a structured score. The key mistake is treating grading as a single, final check. Instead, integrate it into a loop: draft, call Claude, score, then repeat. This transforms experiments into components you can trust in an application. For the Claude Certified Architect exam, understand that higher-level architecture choices depend on knowing how behavior is measured—not just that the model works, but that your evaluation system reliably catches failures. Treat grading as a prerequisite practice for building production-ready systems, not a one-time quality stamp.
Sources
- 2026-06-09 — Claude Certified Architect Prerequisite Building with the Claude API Part 6 Model Based Grading
- 2026-05-21 — Claude Certified Architect Prerequisite Building with the Claude API Part 1 API Foundations
- 2026-06-01 — Claude Certified Architect Prerequisite Building with the Claude API Part 4 Structured Data
- 2026-05-24 — Claude Certified Architect Prerequisite Claude API Part 2 Conversations & System Prompts
- 2026-06-05 — I Updated grill-me And Solved Claude Code
- 2026-06-06 — Claude Now Writes 80 of Its Own Code Should We Be Worried
- 2026-03-11 — Google's New Model + Claude Code Just Changed RAG Forever
- 2026-06-05 — Claude Certified Architect Prerequisite Claude API Part 5 Prompt Eval to Running the Eval
- 2026-04-03 — 2 Claude Code Repos NOBODY'S Talking About Yet
- 2026-03-04 — 🚀Claude Skills Got An UPDATE Check Your Skills Now!
- 2026-05-20 — MCPs Are Dead. Claude Code Wants CLIs
- 2026-06-08 — Are You Wasting Your Time Learning Claude (important video)