AI Agents & Orchestration

Agents in Production

Last updated 2026-09-19

What's new

2026-09-19
  • The workshop will introduce "agent harnesses" (a hot topic in AI that helps manage and connect AI models to data and tools) and "agent memory" (how AI systems can remember and use past interactions).
  • Participants will learn to build their own agent harness using GitHub Codespaces (a cloud-based development environment) and a provided GitHub repository (a storage space for coding projects).
  • The session will cover the "agent stack" (five layers that make up AI systems, including data, models, infrastructure, and compute) and focus on the data layer, where agent harnesses play a crucial role.
  • The workshop will also explore different types of AI applications, from simple chatbots to more advanced AI agents that can automate tasks and act autonomously.
2026-09-16
  • Verscell created the AIDK (Agent Infrastructure Development Kit), a tool that simplifies building agents (AI tools that can perform tasks) by letting you switch between different AI providers with just one line of code.
  • They developed a data science agent called D0, which automates tasks like writing and running SQL (database query language) to answer questions, freeing up data scientists' time.
  • Verscell is exploring the idea of having an agent (AI assistant) on every desk, not just for coding but for various jobs like design, product management, and more.
  • They're working on a new agent architecture where a single agent manages its own memory and can reflect on its steps, improving its ability to complete complex tasks.
2026-08-28
  • Together AI and Stanford are creating environments (spaces where AI agents can work) for AI agents to make scientific discoveries, shifting from designing workflows (step-by-step instructions) to designing environments with incentives and resources.
  • They developed Einstein Arena, a platform where AI agents can collaborate and compete to solve open-ended scientific problems, with a leaderboard (a ranking system) and discussion forum (a place to talk and share ideas).
  • Within weeks of launch, agents in Einstein Arena discovered new solutions to 11 problems, outperforming previous human solutions and specialized AI tools.
  • One example is the kissing number problem (a math problem about fitting spheres around a central sphere), where agents found better solutions in higher dimensions.
2026-08-25
  • Most AI "agents" (tools that automate tasks) work similarly under the hood, using a brain (AI model) and a workspace (cloud or your computer) to do tasks.
  • Grok Bot is a new, easy-to-use AI agent that runs entirely in the cloud, so it keeps working even if you turn off your computer.
  • Hermes Agent is an open-source alternative that lets you swap AI models (the "brain") and customize tools more freely than Grok Bot.
  • Some AI tools gain sudden popularity due to platform algorithms favoring them, not just because they’re better.
2026-08-19
  • Hermes, a popular AI tool, just added a new "bot mode" that lets you create and manage multiple AI agents (individual AI helpers with specific roles), similar to a competing tool called Grockbot.
  • In this new mode, you can have different AI agents (like Dusty, Barry, or Cindy) with their own tools and skills, and they can even talk to each other to share information, making it feel like you have a team of helpers.
  • Hermes' bot mode is a direct copy of Grockbot's design, but it offers more customization, like choosing different AI providers (companies that make AI models) and adjusting each agent's personality.
  • This update is only available in the desktop app (software you run on your computer) and requires the latest version of Hermes.

Key points

What it is

  • An AI agent is an AI program that uses tools in a loop to complete tasks on its own, following a goal and checking results until the task is done.
  • It's like a folder with instructions, skills (packaged capabilities), and connections to data or tools.
  • Agents are designed to be customized, evaluated, deployed, and monitored for real-world use.
  • The skill is designing workflows and giving agents the right building blocks to act reliably.

How to use it

  • Define what each agent should do and when to use it, then create it by selecting if it's personal, global, or project-specific.
  • Give the agent a clear purpose, a workflow (sequence of steps) as guidelines, and let it act in its environment.
  • Run multiple agents in separate panes or instances to scale horizontally and isolate them to avoid interference.
  • Start with a pilot, manually do the work with AI, then productize it into a repeatable system.

Watch out for

  • Agents can produce plausible but incorrect outputs that may go unnoticed, leading to cascading failures.
  • Overcomplicating agents with large system prompts and many edge cases can make them less effective.
  • Agents are non-deterministic, so always test skills and add stop conditions and validations to catch faults.
  • Production bugs are common as agents face unseen data, so set up guards as you onboard real users.

Tools named

  • Claude (an AI assistant for generating agent configurations), n8n (a drag-and-drop tool for connecting apps)

Lesson 1: What is Agents in Production and why it matters

An AI agent is a model that uses tools in a loop: you give it a goal, it decides what to do next, calls a tool, looks at the result, and keeps going until the task is done. That loop is the core idea—everything else is built on top of it. A simple way to picture an agent is as a folder with three things inside: instructions (a markdown file), skills (packaged capabilities), and a connection to data or tools. You don’t always build the engine from scratch; you take a general-purpose agent and put skills on top of it, like standard operating procedures you already have written down.

Getting an agent from demo to production matters because it means solving four problems at once: customization (connecting to real data with security and compliance), evaluation (measuring quality before users see it), deployment (scalable infrastructure with CI/CD), and observability (monitoring in real time). Most teams get stuck between evaluation and deployment—the agent works on a laptop but fails in production, and that gap takes 3 to 9 months to close.

Why does this matter for AI development? Because we’re moving from building separate agents for each use case to designing what agents should do, where they should be proactive, and how they work together. The skill isn’t coding; it’s designing workflows and giving agents the building blocks—skills, plugins, and procedures—they need to act reliably. As models get faster and cheaper, agents will drive more actions, so getting them to production safely and observably is the real competitive edge.

Sources

Lesson 2: How to use Agents in Production: step-by-step

To use agents (AI programs that complete tasks on their own) in production, start by defining what each agent should do and when to use it. Create a new agent by selecting whether it’s personal, global, or project-specific. Then, either generate its configuration with Claude or manually set it up. Describe the agent’s purpose clearly so it knows its parameters.

Once built, an agent is just an LLM (a language model) that runs tools in a loop—it has goals, context, and an environment to act in. Give it a workflow (a sequence of steps) as guidelines; for example, gather data from five competitor sources, analyze findings, analyze your business, and create a PDF report. As it works, you can steer it by saying “I liked this, change that,” refining outputs in real-time.

For production use, run multiple agents in separate panes or instances. In one pane, tell an agent what to do quickly, then move to another, cycling through to validate outputs and adjust direction. This lets you scale horizontally (add more agents as needed) and isolate them to avoid interference.

Start with a pilot where you manually do the work with AI, then productize (turn it into a repeatable system). For example, build an agent to continuously optimize production app performance—it analyzes issues and helps fix them. The goal is to empower engineers, not replace them, so design the product around user feedback and trust by being transparent about what agents do.

Sources

Lesson 3: Best practices and pitfalls

Agents in production fail in predictable ways. The most dangerous failure is not a bad output—that’s easy to spot and fix. The real risk is a polished artifact that looks correct at a glance but will never pass your exit gates. This cascading failure happens when an agent autonomously produces something plausible, like an investment case built on irrelevant facts, and it gets forwarded without scrutiny.

The root cause is often overcomplication. Every single thing you add to an agent risks making it worse. Large system prompts and lots of edge cases (unusual scenarios) crowd the model’s attention, making it dumber, not smarter. Instead, treat agents as models using tools in a loop—give them clear guidelines, not exhaustive rules.

Because agents are non-deterministic (unpredictable in behavior), you can’t know if a task fails due to a bad skill or because it’s too hard for the model. So never ship skills without evals (tests that measure agent performance). Add stop conditions and validations to catch faults, and design your attention like a system—decide where you enter, what you require, and what you reuse.

In production, monitor traces (logs of agent actions) and feed them back into an optimization loop, just like CI/CD, but don’t rebuild a worse version of it manually. Also, balance capability with safety: give agents enough power to act, but not so much they do something unsafe. Finally, remember demos work because data is clean; production bugs are high because agents face unseen future data, so set up guards as you onboard real users.

Sources