AI Agents & Orchestration

Agent Infrastructure Evolution

Last updated 2026-09-19

What's new

2026-09-19
  • The workshop will introduce "agent harnesses" (a hot topic in AI that helps manage and connect AI models to data and tools) and "agent memory" (how AI systems can remember and use past interactions).
  • Participants will learn to build their own agent harness using GitHub Codespaces (a cloud-based development environment) and a provided GitHub repository (a storage space for coding projects).
  • The session will cover the "agent stack" (five layers that make up AI systems, including data, models, infrastructure, and compute) and focus on the data layer, where agent harnesses play a crucial role.
  • The workshop will also explore different types of AI applications, from simple chatbots to more advanced AI agents that can automate tasks and act autonomously.
2026-09-16
  • AI interactions have evolved from simple websites (first wave) to apps with tools (second wave) and now to persistent, long-running agents (third wave) that work asynchronously in the background.
  • A new open-source tool called Restate (a type of software framework) helps manage these long-running agents, ensuring they can recover from crashes and keep running where they left off.
  • Restate acts like a proxy (a middleman) between requests and your agent, creating a "lifeline" that tracks events and helps recover the agent's state if something goes wrong.
  • It also helps manage multiple agents running at the same time, ensuring they don't interfere with each other and can communicate effectively.
2026-09-07
  • AI tools (smart software) have advanced quickly in coding (writing computer programs), becoming fully autonomous, thanks to supportive systems like code repositories (storage for code) and testing tools.
  • These AI tools struggle in other fields like support, finance, and sales because they lack a central source of truth (one place with all necessary information) and history (record of past actions).
  • To bridge this gap, a central hub is needed where all apps and connections exist, allowing AI tools to access everything in one place, similar to how coding agents work.
  • Additionally, a record of the AI tool's actions is crucial for building trust and memory, enabling the tool to learn from past tasks and allowing users to verify its work.

Key points

What it is

  • **AI agents** are programs that act toward a goal, evolving from simple step-by-step instructions to flexible, autonomous systems.
  • These agents now operate in environments that provide incentives, safety limits, and resources, allowing them to decide how to act.
  • **Agentic AI** is not a new model but a system pattern using planning, tools, memory, and goal-directed autonomy to manage tasks.
  • The skill now is designing what agents should do, where they should be proactive, and how they fit into existing infrastructure.

How to use it

  • Start by writing a **spec** (detailed plan) and a conceptualization stage before coding, fixing the boundary between what the agent does and what you verify.
  • Add **agentic validations** (automated checks) to confirm the agent's output, specifying exact tools and verification methods.
  • Capture the agent's **reasoning chain** (thought process and tool calls) as signals to judge if it's on track.
  • Use a **context engine** to pull documented outages and fixes into each session, and trace agent reasoning for judgment.

Watch out for

  • Avoid **vibe coding** (casual, prompt-driven building) for real products, as it often leads to data leaks, security issues, and scalability problems.
  • Treat each new session as a blank slate and supply explicit context to avoid incidents.
  • Watch what the agent can access, including permissions and actions, not just generated code.
  • Separate the task from the model, using specs, verification, and structured workflows for reliability and safety.

Tools named

  • Supabase (a tool for building database apps), n8n (a drag-and-drop tool for connecting apps), CI/CD (continuous integration and delivery pipelines)

Lesson 1: What is Agent Infrastructure Evolution and why it matters

Agent infrastructure evolution is the story of how AI agents (programs that act toward a goal) moved from simple, step-by-step instructions to flexible systems that can work on their own. Early on, developers told agents exactly what to do through prompts and tools. Now, the environment—where the agent works—matters more than the workflow itself. A good environment provides incentives, guardrails (safety limits), and resources, letting an agent decide how to act. This shift matters because agents now cause side effects in the outside world, like deleting production data, so they behave like distributed systems (networks of parts that must stay in sync). That means you need to think about external systems and states when building them.

The evolution has been rapid: from prompt engineering to RAGs (retrieval-augmented generation, pulling external data), then to MCPs (model context protocols, standard ways agents talk to tools), then multi-agent setups, and now deep agents that can build full apps in minutes. Agentic AI is not a new model but a system pattern around the model, using planning, tools, memory, and goal-directed autonomy—like an operator managing a whole job, not a cook filling one order.

Why does this matter? Because the skill is no longer coding. It’s designing what agents should do, where they should be proactive, and how they fit into existing infrastructure. Half of companies using generative AI will soon deploy agentic systems, so understanding this evolution helps you build agents that are reliable, safe, and powerful.

Sources

Lesson 2: How to use Agent Infrastructure Evolution: step-by-step

Vibe coding (casually prompting AI to build something) feels fast, but it often ends in disaster when you skip systems design thinking. You might build an app through hours of grueling iteration, only to find data leaks, security issues, and scalability problems hidden under the hood. To avoid this, evolve to agentic engineering (a disciplined approach with formal specs and verification). Google's framework places you on a spectrum: vibe coding on one end with casual prompts and "does it seem to work" checks, and agentic engineering on the other with automated test suites and CI/CD (continuous integration and delivery pipelines).

Start by writing a spec and a conceptualization stage before touching code. Fix the boundary between what the agent does and what you verify. When you make that boundary clear, a simple prompt inside can drive the work. Next, add agentic validations (automated checks that confirm the agent's output). For example, instead of asking an AI to "build a database app," specify the exact tool, like Supabase, and define how you'll verify it works. The reasoning chain (the agent's thought process and tool calls) should be captured as signals so you can judge if it's on track.

Remember, software factories fail when teams just ask agents to build as much as possible. A developer vibe coding a side project for a dozen users is fine, but for real products, you need the discipline of specs, verification, and structured workflows. That's how you get ahead of 99% of people—by understanding not just what you want, but how to prove you got it.

Sources

Lesson 3: Best practices and pitfalls

The biggest mistake in agent infrastructure (systems where AI builds or runs software) is treating vibe coding (casual, prompt-driven building) like a finished product. Developers vibe-code a side project a dozen people will run, then wonder why the house collapses. The core problem is that AI is stateless (starts every session with zero memory) — no architecture design, business logic, or constraints carry over. Avoid disaster by defining formal specs and memory files, not casual prompts.

Verification is where most failures happen. If you only check "does it seem to work," you miss automated test suites and CI/CD (continuous integration and deployment checks). Every component you add to an agent risks making it worse — large system prompts and edge cases create fragility, not robustness. Also, watch what the agent can access: it's not just generated code, but the agent's permissions and actions. Treat each new session as a blank slate and supply explicit context, or you'll fight incidents endlessly.

Best practices: separate the task from the model — give a spec and repository, get a pull request. Use a context engine to pull battle scars (documented outages and fixes) into each session. Fan out multiple agents to analyze a codebase before making changes, and trace agent reasoning (goal, tool calls, belief status) for judgment. Finally, remember that if you build something nobody uses, it fails regardless of code quality. Vibe coding works for buzz-generating games, but for billing engines or production systems, discipline beats vibes.

Sources