Agent (topic)
Last updated 2026-09-19What's new
- The workshop will introduce "agent harnesses" (a hot topic in AI that helps manage and connect AI models to data and tools) and "agent memory" (how AI systems can remember and use past interactions).
- Participants will learn to build their own agent harness using GitHub Codespaces (a cloud-based development environment) and a provided GitHub repository (a storage space for coding projects).
- The session will cover the "agent stack" (five layers that make up AI systems, including data, models, infrastructure, and compute) and focus on the data layer, where agent harnesses play a crucial role.
- The workshop will also explore different types of AI applications, from simple chatbots to more advanced AI agents that can automate tasks and act autonomously.
- Files and tools are replacing Python code for building AI agents (self-operating AI programs), making the process simpler and more efficient.
- The new Gemini API (a tool for running AI models) introduces a unified interface for running models and agents, with features like tool calls, multimodality (processing different types of data like text, images, etc.) understanding and generation.
- The API uses a step-based system instead of turn-based conversations, allowing for more complex interactions and better agent building.
- Agent frameworks, like the ADK (Agent Development Kit) framework, abstract away complex code, making it easier to build and manage agents.
- Together AI and Stanford are creating environments (spaces where AI agents can work) for AI agents to make scientific discoveries, shifting from designing workflows (step-by-step instructions) to designing environments with incentives and resources.
- They developed Einstein Arena, a platform where AI agents can collaborate and compete to solve open-ended scientific problems, with a leaderboard (a ranking system) and discussion forum (a place to talk and share ideas).
- Within weeks of launch, agents in Einstein Arena discovered new solutions to 11 problems, outperforming previous human solutions and specialized AI tools.
- One example is the kissing number problem (a math problem about fitting spheres around a central sphere), where agents found better solutions in higher dimensions.
- Learn to define tasks clearly and check AI outputs before using them—no coding skills needed.
- An agent (software that helps reach a goal while you stay responsible for results) guides these steps.
- Use five key checks: brief quality, evidence handling, example quality, exploration range, and prompt repair.
- Practice by reading scenarios cold, choosing an answer, then comparing explanations to spot weak options.
- Hugging Face (a platform for sharing machine learning models and data) has a team that helps researchers move their work from services like Google Drive to the Hugging Face platform, making it easier to find and use.
- This team manually checks research papers, opens requests to move models/data, and improves their descriptions, but this process isn't scalable due to the high volume of new research.
- To automate this, the team is exploring AI agents (AI programs that can perform tasks automatically) to handle outreach to researchers and manage the workflow more efficiently.
- They chose a more predictable, step-by-step approach using LLM (large language model) APIs (interfaces that allow software to interact) instead of fully autonomous agents for better control.
- Hermes, a popular AI tool, just added a new "bot mode" that lets you create and manage multiple AI agents (individual AI helpers with specific roles), similar to a competing tool called Grockbot.
- In this new mode, you can have different AI agents (like Dusty, Barry, or Cindy) with their own tools and skills, and they can even talk to each other to share information, making it feel like you have a team of helpers.
- Hermes' bot mode is a direct copy of Grockbot's design, but it offers more customization, like choosing different AI providers (companies that make AI models) and adjusting each agent's personality.
- This update is only available in the desktop app (software you run on your computer) and requires the latest version of Hermes.
- Researchers introduced a "replay agent" (a simple script that mimics successful actions recorded from a advanced AI model) that can outperform the original AI on standard benchmarks (tests) due to the predictable nature of these tests.
- The team identified issues with current evaluation methods, like "pass at K" (a metric that checks if at least one of K attempts succeeds), which can be manipulated by replay agents in deterministic (predictable) environments.
- To create better benchmarks, they proposed "PRISM principles" (guidelines for designing robust and trustworthy environments), focusing on variability, verification, sandboxing (isolating the environment), and realism.
- They built a new benchmark called "DGWord" (a set of 15 Android apps with 387 verified scenarios and 3.2 million configurations) that follows these principles to prevent exploitation by replay agents.
- Anthropic (a company that makes AI tools) has developed a new way to build "agents" (AI tools that can do tasks for you) called Claude managed agents, which handles the complicated tech stuff so you can focus on what your AI tool should do.
- This new system takes care of things like keeping data safe, managing AI tasks, and making sure everything runs smoothly, which used to be really hard for companies to set up on their own.
- Before this, companies had to build and maintain these systems themselves, which took a lot of time and effort away from improving their AI tools.
- The evolution of these AI tools has gone from simple question-and-answer systems to complex agents that can handle entire tasks, and this new system is designed to support that complexity.
- **AI agents (automated digital assistants) are being used to build and run companies, automating tasks and improving productivity, like a "brain agent" at Verscell (a big company) that 1,000 employees use daily.**
- **Coding agents (AI that writes or assists with code) are a "killer app" (essential tool) for agents, helping both expert and beginner coders, with platforms like Vzero and Lovable making software creation more accessible.**
- **Agents can manage company data and tasks, like tracking customer info or navigating internal systems, making businesses more efficient, though this concept may still seem unfamiliar to many.**
- **Open-source models (free, community-developed AI tools) like Kimmy K3 are being used to build and improve agents for business workflows.**
- AI can help create a virtual executive officer (a digital assistant for business tasks) using tools like Claude Code (a coding assistant) and frameworks like Seed (a planning tool) and Skill Smith (a skill-building tool).
- To build this officer, you need to know what you want it to do, what data it can use, and how to connect it to your other software tools using MCPs (command-line tools that act as bridges).
- The focus is on AI augmentation (using AI to improve decisions) rather than full automation (replacing all human tasks), especially if your business processes aren't clearly defined yet.
- You can use tools like Appify (a data scraper) to gather data from platforms like Instagram and YouTube, and integrate it with your officer for tasks like competitor analysis.
- Forward-deployed engineering (working closely with customers to solve problems) started as a way to ensure software stability (DevOps) and data integration (combining different data sources) for companies like Palantir (a data analysis company).
- The term "forward-deployed engineering" has become too broad and vague, but it's still considered a hot job in AI (artificial intelligence).
- A new discipline called agent engineering (AI systems that perform tasks for users) has emerged, with harness engineering (managing and coordinating AI agents) as a subset.
- The speaker argues that everyone is essentially a forward-deployed engineer, working directly with customers and AI tools to solve problems.
Key points
What it is
- An AI agent is an AI system that works toward a goal on its own, like an employee rather than a chatbot.
- It has three main parts: a model (the AI brain), tools (like web search or file readers), and a loop (the process of deciding, acting, and checking results).
- Unlike a chatbot, an agent is non-deterministic, meaning it adapts and changes as it goes, sometimes succeeding and sometimes failing.
- Agents can read and work with text, code, images, or PDFs, and use tools to complete tasks.
How to use it
- Start by defining the agent's purpose with a clear, plain-language instruction of what it should do and when to use it.
- Add context, like examples of previous successful tasks, to help the agent understand what you expect.
- Define the tools the agent can use, like search or data retrieval, and let it loop through reasoning and acting until the task is done.
- Test your agent using evals (evaluations that measure performance) to ensure it follows guidelines and uses appropriate tone and vocabulary.
Watch out for
- Avoid adding too much complexity to your agent, as this can degrade its performance.
- Start with a simple agent and apply it to recurring tasks first, not one-off experiments.
- Break complex systems into smaller, domain-specific agents to reduce the "vectors for failure" and make each agent more reliable.
- Watch for "cascading failures," where an early mistake goes unnoticed, and design your attention like a system to prevent this.
Tools named
- Claude (an AI assistant that can generate and manage sub-agents), Codex (a tool for building and testing AI agents)
Lesson 1: What is Agent (topic) and why it matters
An AI agent is a system that can work toward a goal on its own. The simplest way to picture it is a folder with three things inside: a model (the AI brain), tools (things like web search or file readers), and a loop (the process of deciding, acting, and checking results). One researcher found over 200 definitions, but the core idea is simple. You give an agent a goal. It decides what to do next, calls a tool, looks at the result, and keeps going until the task is done.
This matters for AI development because it changes how we use AI. A regular chatbot is like a meeting — you ask a question and get an answer. An agent is like an employee — you tell it what you want done, and it runs the full workflow. Think of a vending machine versus a slot machine. A vending machine is deterministic (same input, same output every time). An agent is non-deterministic — sometimes it wins, sometimes it loses, and it adapts as it goes.
Developers are moving away from supervising every step. Instead, they give guidance and let the agent handle the loop. An agent can read text, code, images, or PDFs, then work through tools instead of just producing chat. Understanding this shift — from simple chat to goal-directed autonomy — is one of the most important skills to learn in AI development today.
Sources
- 2026-07-08 — An AI Agent Is Just a Folder With 3 Things Inside It
- 2026-07-12 — The Agentic Web and the Bazaar Era of AI - Ramesh Raskar, MIT Media Lab
- 2026-02-13 — Claude Code 2.1.41 Update Breakdown Terminal, File Reads & More
- 2026-05-09 — Agentic AI Systems, Clearly Explained
- 2026-06-14 — Zero to AWS Certified AI Practitioner AIF-C01 in 2026 Part 2 AIML Vocabulary
- 2026-07-23 — Why Agentic Systems Need Ontologies Frank Coyle, UC Berkeley
- 2026-06-15 — Learn These 6 AI Skills Now (Before AI Replaces You)
- 2026-07-15 — Youre Not Behind (Yet) How to Build Your First AI Agent (Full Guide)
- 2026-05-16 — Claude Code Just Got Better Agent View
- 2026-05-28 — The acceleration is here!
- 2026-05-30 — How I deleted 95 of my agent skills and got better results Nick Nisi, WorkOS
- 2026-07-25 — Gemini 3.6 Flash & the 3.5 Pro Delay When the Cheap Model Shipped and the Flagship Couldn't
- 2026-07-24 — Everything Is a Rollout Alex Shaw + Ryan Marten, Terminal-Bench, Harbor, Laude Institute
Lesson 2: How to use Agent (topic): step-by-step
To use an agent (an AI that performs tasks autonomously), start by defining its purpose. An agent is essentially a folder with three things: instructions, context, and tools. Begin with the core instruction—a plain-language description of what the agent should do and when to use it. In Claude, you can "generate with Claude" by typing something like "Create me a sub agent that criticizes all of my work." For the best results, be comprehensive in your description. The system prompt (the main instruction file) goes into a file named `claude.md` (if using Claude Code) or `agents.md` (if using Codex). This file holds the agent's behavioral rules.
Next, add context. Include an examples subfolder with previous winning proposals for the AI to reference. This helps the agent understand the quality and style you expect. Then, define tools—the functions the agent can call, like search or data retrieval. The agent works by looping through reasoning and acting: it thinks, checks the user input, decides if it should use a tool, acts, then cycles until finished.
To test your agent, use evals (evaluations that measure performance). For example, check if the agent follows business guidelines and uses appropriate tone and vocabulary. Evals become especially important when your agent has multiple steps or uses tools like search. They help catch when the agent goes off the rails. Start by applying your agent to a recurring task you do often—that's where agents are most useful.
Sources
- 2026-05-29 — The Claude Update Everyone Missed (Dynamic Workflows)
- 2026-07-11 — Claude Code for Non-Coders (6 Hour Course)
- 2026-06-09 — How to Build Claude Subagents Better Than 99 of People
- 2026-07-08 — An AI Agent Is Just a Folder With 3 Things Inside It
- 2026-07-06 — How to Get Ahead of 99 of People In the Age of AI - 50 Tips
- 2026-05-12 — Dark Factory How OpenClaw Ships Faster Than You Can Read the Diff Vincent Koc
- 2026-05-12 — Lessons from Trillion Token Deployments at Fortune 500s Alessandro Cappelli, Adaptive ML
- 2026-07-20 — FDE The 1MYear AI Job Explained
- 2026-05-09 — Agentic AI Systems, Clearly Explained
- 2026-07-06 — The Only Part of Your AI Setup Competitors Can't Copy
- 2026-06-29 — The Future Is Domain-Specific Agents - Justin Schroeder, StandardAgents
- 2026-06-26 — Turn 10,994 Notes Into Memory - Paul Iusztin, Decoding AI & Louis-Franois Bouchard, Towards AI
- 2026-05-14 — Ship Real Agents Hands-On Evals for Agentic Applications Laurie Voss, Arize
Lesson 3: Best practices and pitfalls
When building an AI agent (a model using tools in a loop), the most common pitfall is adding too much complexity. As one expert warns, "every single thing you add to an agent risks making it worse." Large system prompts and too many edge cases often degrade performance rather than improve it.
The first best practice is to start simple. A basic agent needs three things: instructions, tools it can call, and the ability to run in a loop. Apply agents to recurring tasks first, not one-off experiments. Once you have a working agent, create evals (evaluations that test agent behavior) using specific examples that represent hard cases your agent must handle. Run the agent against these evals to identify exactly where it fails.
Break complex systems into smaller, domain-specific agents. When you find one agent trying to do too much, split it into separate subject matter experts. This reduces the "vectors for failure" and makes each agent more reliable. Use verifiers to check agent outputs, especially looking for disagreement between the agent and its reviewers. Have a subject matter expert tune the review process when conflicts arise.
A common failure mode is the "cascading failure" where an agent makes an early mistake that goes unnoticed because the system runs autonomously. To prevent this, design your attention like a system—be intentional about where you enter, what you require, and what you reuse. Also watch for agents that "don't do enough leg work" on a step, like failing to ask clarifying questions or explore a codebase thoroughly before proceeding.
Sources
- 2026-05-19 — Don't Build Slop (4 Levels of AI Agent Maturity) - Ara Khan, Cline
- 2026-05-12 — Dark Factory How OpenClaw Ships Faster Than You Can Read the Diff Vincent Koc
- 2026-07-08 — An AI Agent Is Just a Folder With 3 Things Inside It
- 2026-05-14 — Ship Real Agents Hands-On Evals for Agentic Applications Laurie Voss, Arize
- 2026-05-12 — Lessons from Trillion Token Deployments at Fortune 500s Alessandro Cappelli, Adaptive ML
- 2026-06-29 — The Agentic AI Engineer - Benedikt Sanftl, Mutagent
- 2026-06-29 — The Future Is Domain-Specific Agents - Justin Schroeder, StandardAgents
- 2026-07-14 — The engineer of the future is the person who is able to choose what is worth doing. Addy Osmani
- 2026-05-27 — The maturity phases of running evals Phil Hetzel, Braintrust
- 2026-05-28 — The acceleration is here!
- 2026-05-09 — Build Your Agentic OS Better Than The 99
- 2026-06-08 — Become AI Native in less than 60 mins
- 2026-07-25 — From Agent Traces to Agent Simulations Rustem Feyzkhanov, Snorkel AI
- 2026-06-29 — Building Great Agent Skills The Missing Manual