Agent Infrastructure Evolution
Last updated 2026-09-19What's new
- The workshop will introduce "agent harnesses" (a hot topic in AI that helps manage and connect AI models to data and tools) and "agent memory" (how AI systems can remember and use past interactions).
- Participants will learn to build their own agent harness using GitHub Codespaces (a cloud-based development environment) and a provided GitHub repository (a storage space for coding projects).
- The session will cover the "agent stack" (five layers that make up AI systems, including data, models, infrastructure, and compute) and focus on the data layer, where agent harnesses play a crucial role.
- The workshop will also explore different types of AI applications, from simple chatbots to more advanced AI agents that can automate tasks and act autonomously.
- AI interactions have evolved from simple websites (first wave) to apps with tools (second wave) and now to persistent, long-running agents (third wave) that work asynchronously in the background.
- A new open-source tool called Restate (a type of software framework) helps manage these long-running agents, ensuring they can recover from crashes and keep running where they left off.
- Restate acts like a proxy (a middleman) between requests and your agent, creating a "lifeline" that tracks events and helps recover the agent's state if something goes wrong.
- It also helps manage multiple agents running at the same time, ensuring they don't interfere with each other and can communicate effectively.
- AI tools (smart software) have advanced quickly in coding (writing computer programs), becoming fully autonomous, thanks to supportive systems like code repositories (storage for code) and testing tools.
- These AI tools struggle in other fields like support, finance, and sales because they lack a central source of truth (one place with all necessary information) and history (record of past actions).
- To bridge this gap, a central hub is needed where all apps and connections exist, allowing AI tools to access everything in one place, similar to how coding agents work.
- Additionally, a record of the AI tool's actions is crucial for building trust and memory, enabling the tool to learn from past tasks and allowing users to verify its work.
Key points
What it is
- **AI agents** are programs that act toward a goal, evolving from simple step-by-step instructions to flexible, autonomous systems.
- These agents now operate in environments that provide incentives, safety limits, and resources, allowing them to decide how to act.
- **Agentic AI** is not a new model but a system pattern using planning, tools, memory, and goal-directed autonomy to manage tasks.
- The skill now is designing what agents should do, where they should be proactive, and how they fit into existing infrastructure.
How to use it
- Start by writing a **spec** (detailed plan) and a conceptualization stage before coding, fixing the boundary between what the agent does and what you verify.
- Add **agentic validations** (automated checks) to confirm the agent's output, specifying exact tools and verification methods.
- Capture the agent's **reasoning chain** (thought process and tool calls) as signals to judge if it's on track.
- Use a **context engine** to pull documented outages and fixes into each session, and trace agent reasoning for judgment.
Watch out for
- Avoid **vibe coding** (casual, prompt-driven building) for real products, as it often leads to data leaks, security issues, and scalability problems.
- Treat each new session as a blank slate and supply explicit context to avoid incidents.
- Watch what the agent can access, including permissions and actions, not just generated code.
- Separate the task from the model, using specs, verification, and structured workflows for reliability and safety.
Tools named
- Supabase (a tool for building database apps), n8n (a drag-and-drop tool for connecting apps), CI/CD (continuous integration and delivery pipelines)
Lesson 1: What is Agent Infrastructure Evolution and why it matters
Agent infrastructure evolution is the story of how AI agents (programs that act toward a goal) moved from simple, step-by-step instructions to flexible systems that can work on their own. Early on, developers told agents exactly what to do through prompts and tools. Now, the environment—where the agent works—matters more than the workflow itself. A good environment provides incentives, guardrails (safety limits), and resources, letting an agent decide how to act. This shift matters because agents now cause side effects in the outside world, like deleting production data, so they behave like distributed systems (networks of parts that must stay in sync). That means you need to think about external systems and states when building them.
The evolution has been rapid: from prompt engineering to RAGs (retrieval-augmented generation, pulling external data), then to MCPs (model context protocols, standard ways agents talk to tools), then multi-agent setups, and now deep agents that can build full apps in minutes. Agentic AI is not a new model but a system pattern around the model, using planning, tools, memory, and goal-directed autonomy—like an operator managing a whole job, not a cook filling one order.
Why does this matter? Because the skill is no longer coding. It’s designing what agents should do, where they should be proactive, and how they fit into existing infrastructure. Half of companies using generative AI will soon deploy agentic systems, so understanding this evolution helps you build agents that are reliable, safe, and powerful.
Sources
- 2025-11-24 — This AI Model Is Smarter Than Ever Before!
- 2026-05-30 — How I deleted 95 of my agent skills and got better results Nick Nisi, WorkOS
- 2026-08-25 — Einstein Arena Harnessing Collective Agent Intelligence for Open Science James Zou, Together AI
- 2026-08-29 — AI Agents Are Just Distributed Systems Now Salman Munaf, TikTok
- 2026-05-05 — Demand-Driven Context A Methodology for Coherent Knowledge Bases Through Agent Failure
- 2026-08-06 — The AI Agent Every Company is About to Build Vercel CEO Guillermo Rauch
- 2026-07-28 — The Dirty Secret of Forward Deployed Engineering Natalie Meurer, Sierra
- 2026-02-13 — Claude Code 2.1.41 Update Breakdown Terminal, File Reads & More
- 2026-06-14 — Zero to AWS Certified AI Practitioner AIF-C01 in 2026 Part 2 AIML Vocabulary
- 2026-07-24 — Everything Is a Rollout Alex Shaw + Ryan Marten, Terminal-Bench, Harbor, Laude Institute
- 2026-07-15 — Youre Not Behind (Yet) How to Build Your First AI Agent (Full Guide)
- 2026-01-25 — Agentic Workflows Just Changed AI Automation Forever! (Claude Code)
- 2026-08-23 — Ask Me Anything - L7 Senior Staff Software Engineer at Meta
Lesson 2: How to use Agent Infrastructure Evolution: step-by-step
Vibe coding (casually prompting AI to build something) feels fast, but it often ends in disaster when you skip systems design thinking. You might build an app through hours of grueling iteration, only to find data leaks, security issues, and scalability problems hidden under the hood. To avoid this, evolve to agentic engineering (a disciplined approach with formal specs and verification). Google's framework places you on a spectrum: vibe coding on one end with casual prompts and "does it seem to work" checks, and agentic engineering on the other with automated test suites and CI/CD (continuous integration and delivery pipelines).
Start by writing a spec and a conceptualization stage before touching code. Fix the boundary between what the agent does and what you verify. When you make that boundary clear, a simple prompt inside can drive the work. Next, add agentic validations (automated checks that confirm the agent's output). For example, instead of asking an AI to "build a database app," specify the exact tool, like Supabase, and define how you'll verify it works. The reasoning chain (the agent's thought process and tool calls) should be captured as signals so you can judge if it's on track.
Remember, software factories fail when teams just ask agents to build as much as possible. A developer vibe coding a side project for a dozen users is fine, but for real products, you need the discipline of specs, verification, and structured workflows. That's how you get ahead of 99% of people—by understanding not just what you want, but how to prove you got it.
Sources
- 2026-07-23 — The Unreasonable Effectiveness of Separating the Task from the Model Maxime Rivest & Isaac Miller
- 2026-06-04 — OpenAI Codex Build Apps That Work For You 247
- 2026-07-06 — How to Get Ahead of 99 of People In the Age of AI - 50 Tips
- 2026-07-04 — Google SDLC Whitepaper Digest - 90 Missing
- 2026-08-25 — Einstein Arena Harnessing Collective Agent Intelligence for Open Science James Zou, Together AI
- 2026-06-03 — ChatGPT And Codex Are Merging (This Changes Everything)
- 2026-07-20 — Why Your AI Offer Isn't Selling, and How to Fix That
- 2026-06-12 — Claude Fable Will Change EVERYTHING (Here's Why)
- 2026-08-29 — Agents Are Where Microservices Were in 2015 Roberto Milev & Uday Kanagala, Navan
- 2026-06-29 — The Agentic AI Engineer - Benedikt Sanftl, Mutagent
- 2026-05-13 — Self-Training Agents Hermes Agent, HF Traces, Skills, MCP & Finetuning Merve Noyan, Hugging Face
- 2026-08-11 — Evolution of agentic surfaces Gagan Bhat & Isabella Kai He, Anthropic
- 2026-07-23 — Harness Engineering is not Enough Why Software Factories Fail Dex Horthy, HumanLayer
- 2026-05-29 — The Claude Update Everyone Missed (Dynamic Workflows)
- 2026-03-19 — We Fixed the #1 Reason Claude Code Apps Fail
Lesson 3: Best practices and pitfalls
The biggest mistake in agent infrastructure (systems where AI builds or runs software) is treating vibe coding (casual, prompt-driven building) like a finished product. Developers vibe-code a side project a dozen people will run, then wonder why the house collapses. The core problem is that AI is stateless (starts every session with zero memory) — no architecture design, business logic, or constraints carry over. Avoid disaster by defining formal specs and memory files, not casual prompts.
Verification is where most failures happen. If you only check "does it seem to work," you miss automated test suites and CI/CD (continuous integration and deployment checks). Every component you add to an agent risks making it worse — large system prompts and edge cases create fragility, not robustness. Also, watch what the agent can access: it's not just generated code, but the agent's permissions and actions. Treat each new session as a blank slate and supply explicit context, or you'll fight incidents endlessly.
Best practices: separate the task from the model — give a spec and repository, get a pull request. Use a context engine to pull battle scars (documented outages and fixes) into each session. Fan out multiple agents to analyze a codebase before making changes, and trace agent reasoning (goal, tool calls, belief status) for judgment. Finally, remember that if you build something nobody uses, it fails regardless of code quality. Vibe coding works for buzz-generating games, but for billing engines or production systems, discipline beats vibes.
Sources
- 2026-06-04 — OpenAI Codex Build Apps That Work For You 247
- 2026-07-23 — The Unreasonable Effectiveness of Separating the Task from the Model Maxime Rivest & Isaac Miller
- 2026-03-19 — We Fixed the #1 Reason Claude Code Apps Fail
- 2026-07-23 — Harness Engineering is not Enough Why Software Factories Fail Dex Horthy, HumanLayer
- 2026-07-04 — Google SDLC Whitepaper Digest - 90 Missing
- 2026-06-04 — The Art & Science of Benchmarking Agents Vincent Chen, Snorkel AI
- 2026-07-20 — Agentic Development Security Ezra Tanzer, Snyk
- 2026-08-27 — How to Generate Mergeable Code with a Context Engine Peter Werry, Unblocked
- 2026-08-28 — How to avoid disaster when vibe-coding a billing engine Andrew Garvin, Stripe
- 2026-05-19 — Don't Build Slop (4 Levels of AI Agent Maturity) - Ara Khan, Cline
- 2026-06-03 — Every Claude Code Dynamic Workflow (& When to Use Each)
- 2026-08-29 — Agents Are Where Microservices Were in 2015 Roberto Milev & Uday Kanagala, Navan
- 2026-07-20 — Why Your AI Offer Isn't Selling, and How to Fix That
- 2026-08-09 — Always-on agents run production without the on-call tax Justin Smith, Resolve AI
- 2026-08-26 — Knowledge Systems The New GTM Stack Jeffrey Wang, Exa