AI Agents & Orchestration

Agent Architecture Engineering

Last updated 2026-09-19

What's new

2026-09-19
  • The workshop will introduce "agent harnesses" (a hot topic in AI that helps manage and connect AI models to data and tools) and "agent memory" (how AI systems can remember and use past interactions).
  • Participants will learn to build their own agent harness using GitHub Codespaces (a cloud-based development environment) and a provided GitHub repository (a storage space for coding projects).
  • The session will cover the "agent stack" (five layers that make up AI systems, including data, models, infrastructure, and compute) and focus on the data layer, where agent harnesses play a crucial role.
  • The workshop will also explore different types of AI applications, from simple chatbots to more advanced AI agents that can automate tasks and act autonomously.
2026-09-16
  • Anthropic (a company making AI tools) introduced a new idea called "tokens with jobs," where AI tasks are divided among different groups of tokens (small pieces of data AI uses to learn and work) to improve performance.
  • They tested three strategies: "advising" (some tokens give advice to others), "grading" (some tokens check and score the work of others), and "dreaming" (some tokens learn from past work and improve future tasks).
  • In experiments, these strategies outperformed a simple execution-only approach, showing that dividing tasks among tokens can lead to better results in AI systems.
  • To ensure fair comparison, they fixed the budget (total tokens used) and found that strategies with specific token jobs performed better than just using more tokens.
2026-09-07
  • AI tools (smart software) have advanced quickly in coding (writing computer programs), becoming fully autonomous, thanks to supportive systems like code repositories (storage for code) and testing tools.
  • These AI tools struggle in other fields like support, finance, and sales because they lack a central source of truth (one place with all necessary information) and history (record of past actions).
  • To bridge this gap, a central hub is needed where all apps and connections exist, allowing AI tools to access everything in one place, similar to how coding agents work.
  • Additionally, a record of the AI tool's actions is crucial for building trust and memory, enabling the tool to learn from past tasks and allowing users to verify its work.
2026-08-28
  • Together AI and Stanford are creating environments (spaces where AI agents can work) for AI agents to make scientific discoveries, shifting from designing workflows (step-by-step instructions) to designing environments with incentives and resources.
  • They developed Einstein Arena, a platform where AI agents can collaborate and compete to solve open-ended scientific problems, with a leaderboard (a ranking system) and discussion forum (a place to talk and share ideas).
  • Within weeks of launch, agents in Einstein Arena discovered new solutions to 11 problems, outperforming previous human solutions and specialized AI tools.
  • One example is the kissing number problem (a math problem about fitting spheres around a central sphere), where agents found better solutions in higher dimensions.
2026-08-25
  • AI coding tools (like LLMs or large language models) are improving but still make small mistakes, so human engineers will still be needed to review and maintain AI-generated code.
  • "Agentic coding workflows" (using AI tools to automate parts of coding) are evolving—balance trying new tools with sticking to what works for you.
  • Spend time learning AI coding tools deeply (like reading all their guides) to uncover hidden features that can speed up your work.
  • When reviewing AI-written code, focus first on safety (checking for errors) before looking at quality or cleanliness.
2026-08-22
  • Building AI agents (software that uses AI to perform tasks) has become easier, but they still make mistakes and need careful management.
  • New cloud services (like Cloudflare and AWS) now handle much of the complex setup, letting you focus on the agent's core tasks.
  • Agents can still fail if they lack important context, like details from team discussions or past issues, leading to incorrect decisions.
  • Humans currently act as a safety net for agents, catching errors and providing missing information, but this isn't always possible with automated agents.
2026-08-16
  • AI agents (automated helpers) are getting better at doing tasks for you on the web, but they still struggle because the web wasn't designed for them, making automation difficult.
  • The main issue isn't the AI models (the brains of the agents) anymore, as they've improved a lot and can now handle complex tasks, but agents lack the right tools (called harnesses) to interact with the web effectively.
  • Harnesses (the systems and tools around the AI model) can greatly improve how well agents work, and companies like Cursor and Browserbase are already building these for specific tasks like coding and web browsing.
  • Building a good harness is an engineering problem that any company can tackle, not just big labs, and it can help get more out of the AI models we already have.
2026-08-13
  • Grockbot is a new tool that lets you create and manage AI agents (virtual helpers) from your phone or computer, with everything staying in sync.
  • You can set up different agents for specific tasks, like an executive assistant or an AI engineer, and they can even work together by sharing information.
  • Grockbot allows you to record tasks and turn them into skills that the agents can replicate, and it can also create routines that run automatically, even when your devices are off.
  • To use Grockbot, you need to sign up for the Cursor ultra plan (a paid subscription) and connect your accounts, like Gmail, Google Calendar, and GitHub, to give the agents access to your information.
2026-08-07
  • **AI agents (automated digital assistants) are being used to build and run companies, automating tasks and improving productivity, like a "brain agent" at Verscell (a big company) that 1,000 employees use daily.**
  • **Coding agents (AI that writes or assists with code) are a "killer app" (essential tool) for agents, helping both expert and beginner coders, with platforms like Vzero and Lovable making software creation more accessible.**
  • **Agents can manage company data and tasks, like tracking customer info or navigating internal systems, making businesses more efficient, though this concept may still seem unfamiliar to many.**
  • **Open-source models (free, community-developed AI tools) like Kimmy K3 are being used to build and improve agents for business workflows.**
2026-07-31
  • Forward-deployed engineering (working closely with customers to solve problems) started as a way to ensure software stability (DevOps) and data integration (combining different data sources) for companies like Palantir (a data analysis company).
  • The term "forward-deployed engineering" has become too broad and vague, but it's still considered a hot job in AI (artificial intelligence).
  • A new discipline called agent engineering (AI systems that perform tasks for users) has emerged, with harness engineering (managing and coordinating AI agents) as a subset.
  • The speaker argues that everyone is essentially a forward-deployed engineer, working directly with customers and AI tools to solve problems.

Key points

What it is

  • Agent Architecture Engineering is designing how an AI agent (an AI brain that uses tools in a loop) operates to achieve goals.
  • It involves creating specialized agents or agent teams for specific tasks, like scouting, planning, or reviewing code.
  • It also includes attaching skills to a general-purpose agent engine instead of building separate agents for each use case.
  • The goal is to move from supervising every step to providing guidance, letting the AI run workflows autonomously.

How to use it

  • Start by understanding the core unit: an agent, which is a model (like an LLM) that uses tools in a loop.
  • Define a single agent by creating instructions for its behavior, access to at least one tool (like a web search or calculator), and a way to store context.
  • Design how agents connect intentionally, drawing a simple diagram of components and adding agents one at a time.
  • Test the agent loop to ensure it calls the correct tool and stops when done.

Watch out for

  • Adding too much to an agent can degrade performance instead of improving it.
  • Your architecture will become outdated quickly due to new models, framework versions, or tool-calling standards.
  • Dangerous failures are polished artifacts that look correct but will never pass your exit gates.
  • Build durability from the start by carefully managing the context that goes in and out of the LLM.

Lesson 1: What is Agent Architecture Engineering and why it matters

Agent Architecture Engineering is the discipline of designing how an AI agent operates. An agent is simply a model (the AI brain) that uses tools in a loop: you give it a goal, it decides what to do next, calls a tool, looks at the result, and keeps going until the task is done. This loop is the core idea. Everything else, such as memory (storing past information) and orchestration (coordinating multiple steps), is built on top of it.

Why does this matter for AI development? Without intentional architecture, your agent is just code that evolved without design. A well-architected agent uses specific patterns. For example, you can build agent teams where each agent has a specialized role: one scouts a codebase, one plans the implementation, one writes code, and one reviews it. Or use agent chains that pipeline tasks sequentially: the first agent generates a schema, the second validates it, the third implements it, and the fourth writes tests. This prevents information bleeding between tasks and keeps each agent’s context isolated.

You can also attach skills to a general-purpose agent engine instead of building a separate agent for every use case. This shifts how you build solutions: you focus on composing skills rather than rebuilding the engine each time. Ultimately, Agent Architecture Engineering matters because it moves you from supervising every step to just providing guidance and directions, leaving the AI to run the full workflow autonomously.

Sources

Lesson 2: How to use Agent Architecture Engineering: step-by-step

To start building with Agent Architecture Engineering, begin from zero (no existing code) by understanding the core unit: an agent. An agent is simply a model (like an LLM) that uses tools in a loop (a cycle of acting, observing the result, and deciding the next action). You give it a goal and context, it picks a tool, runs it, checks the outcome, and repeats until finished.

First, define a single agent by creating a folder with three things: instructions for its behavior, access to at least one tool (like a web search or calculator), and a way to store context. For example, use a client setup to configure a basic agent that only answers customer questions. Write a core instruction (a clear prompt describing its role and limits). This is your starting point.

Second, engineering means intentionally designing how agents connect. Instead of letting the system grow randomly, draw a simple diagram of components. A common next step is a multi-agent system: create separate agents for separate capabilities, like one agent for data lookup and another for reporting. Start with one agent, verify it works, then add a second. Use a keyword like "workflow" to describe the sequence they run in.

Third, if you need a subagent, generate it by describing what it should do and when. Always test the loop: does the agent call the correct tool? Does it stop when done? By starting small, testing the loop, and intentionally adding agents one at a time, you engineer a system that grows reliably from zero.

Sources

Lesson 3: Best practices and pitfalls

Designing an agent (an AI that uses tools in a loop) is software engineering. The building blocks are different, but the discipline is the same. The most common mistake is adding too much. Every single thing you add to an agent risks making it worse. Large system prompts and many edge cases often degrade performance instead of improving it. Start with a clear mental model of what the agent should do. Spend time reading the code and thinking through the architecture before building.

Your architecture will become outdated quickly. If you have been building agents for more than six months, you have likely rewritten something. New models, framework versions, or tool-calling standards can make your current approach obsolete. Plan for this by keeping your architecture flexible. You will hit bottlenecks in your agent framework, so design systems that can adapt.

A dangerous failure is not a bad output, which is easy to spot and fix. The dangerous failure is a polished artifact that looks correct at a glance but will never pass your exit gates. Agent safety goes hand in hand with capability. You must give agents enough ability to take action but not so much that they act unsafely. Build durability from the start by carefully managing the context — everything that goes in and out of the LLM (large language model), including system messages, user messages, and tool results.

Good engineering discipline still matters. Define the workflow and know what flows into the system. A well-designed system is much more likely to succeed when updated. If it runs into issues during updates, that signals a need for better maintainability. Let the frameworks handle the unglamorous coding work so you can focus on architecting the agentic system.

Sources