Coding with AI

Didn (topic)

Last updated 2026-09-22

What's new

2026-09-22
  • Codex (a tool by OpenAI that helps you do tasks with AI, like writing, designing, or coding) can be used to build skills, create branded deliverables, and even automate tasks, all without needing a technical background.
  • The Codex desktop app (a program you download to use Codex easily) is recommended for a consistent experience, and it uses the same subscription as ChatGPT (a popular AI chatbot), so you won't need a new account.
  • Codex is more powerful than Work (a tool for non-technical knowledge work) and can do everything Work can do, plus more, making it a better investment for learning and using in the long run.
  • The course will teach you how to use Codex effectively with natural language (regular English, not code) and explain core concepts simply, helping you become a pro AI builder.
2026-09-19
  • New lessons teach how to use text embeddings (numerical representations of text meaning) with the Claude API (a tool for building AI applications) to find relevant information in documents.
  • Voyage AI (a separate service for creating text embeddings) is recommended for generating embeddings, as Anthropic (the company behind Claude) doesn't currently offer this feature.
  • The module focuses on practical implementation, guiding users through installing the Voyage AI library, setting up the client, and creating functions to generate embeddings.
  • The goal is to build a repeatable mental model for using these tools, with an emphasis on inspecting and understanding each step of the process.
2026-09-10
  • OpenAI's ChatGPT and Anthropic's Claude now have better memory, allowing longer, more useful conversations, especially in desktop versions like Codec (for ChatGPT) and Claude Co-work (for Claude).
  • These AI tools now summarize (or "compact") long conversations internally, keeping important info and letting you continue without losing context, unlike before when they'd "get dumb" after ~50% of their capacity.
  • They also have basic "native memory" to recall simple facts about you, like your name or preferences, though this feature is still improving.
  • For most tasks, start fresh chats but use "written memory" (external files) to store important details for future reference, combining freshness with continuity.
2026-09-04
  • AI makes many hidden decisions when completing tasks, like what to include or exclude, and you can now make these visible with a simple prompt: "As you do your work, I want you to keep a log of every decision you make for this task. Specifically, anything that I didn't explicitly specify fight you."
  • These decisions fall into four categories: clarifying ambiguous words (like "important"), determining the result's format, choosing between conflicting data, and noting what was left out.
  • Review the AI's decision log before its output to spot and address recurring issues, saving time and reducing risk.
  • You'll typically accept most AI decisions, but for the rest, make simple fixes or adjust the task instructions to prevent future mistakes.
2026-08-28
  • "Document my process" is a new tool (skill) that helps you write down exactly how you do tasks in your business, making it easier to delegate to others or AI.
  • "Session handoff" is a tool that lets you move all the important details from one AI chat to a new one, so you don't lose progress when the AI starts to get less accurate.
  • Zapier (a tool that connects different apps) has a free guide to help you figure out which tasks really need AI and which don't, potentially saving you a lot of money.
  • Using AI for tasks that don't need it can be expensive, and this guide can help you cut costs by around 71% by fixing your workflows.
2026-08-25
  • Warp built a cloud-based AI agent platform (a system that runs AI tools remotely in the cloud) to handle longer tasks beyond a laptop’s limits.
  • The platform uses sandboxes (isolated cloud environments) for agents to work safely, offering both managed (hosted) and self-hosted options.
  • It supports multiple harnesses (coding tools like IDEs or editors) so developers can use their preferred setup while keeping a consistent experience.
  • The system lets agents collaborate, handling complex tasks by breaking them into smaller steps across different tools.
2026-08-22
  • Kieran, from Every (an AI lab focused on the future of work), shares how he built an AI email inbox called Cora (an agent native tool that runs on various platforms and acts like a personal assistant) with minimal team support.
  • He emphasizes "compound engineering," a process where he extracts his thinking and taste into a system that improves over time, allowing him to outpace traditional teams.
  • Kieran's workflow involves brainstorming, planning, working, reviewing, polishing, and compounding, with a focus on teaching the AI system from his experiences to improve future outputs.
  • He advocates for a balanced approach, spending 50% of the time building features and 50% teaching the system to ensure continuous improvement.
2026-08-19
  • Claude (an AI tool) can act like an AI employee if set up with a workspace (a 'repo' or project folder), memory (context about the project), and a clear assignment (a 'ticket').
  • It can also have 'eyes' to inspect and test the product, review changes, and follow a schedule (called 'routines') for recurring tasks.
  • Permissions can be set to limit Claude's actions, ensuring it doesn't make risky decisions without human approval.
  • The goal is to create an AI-native company, where Claude operates as an always-on, productive team member.
2026-08-10
  • When teams use AI tools (specialized software that helps automate tasks), they may face issues like too many code changes to review or merge, leading to conflicts and delays.
  • Teams might also struggle with a lack of focus, as individual engineers work on different tasks simultaneously, causing a lack of cohesive progress.
  • "Agent bankruptcy" occurs when engineers rely too heavily on AI tools (called agents) and lose track of their work, leading to inefficiency and wasted resources.
  • Critical decisions should not be left to AI tools, as this can lead to a loss of control over the code and the product as a whole.
2026-08-07
  • **Codeex (a tool that can browse the internet and control your computer)** can research topics, organize files, and even speed up your computer by removing unnecessary files.
  • **Chat GPT's new voice mode** lets you control different tasks and threads using your voice, making it easier to multitask.
  • **Chat GPT sites** allows you to publish anything (like poems, spreadsheets, or websites) to the web and share it with others or other AI tools.
  • **Pinning chats** helps you keep track of important conversations by pinning them to the top of your sidebar.
2026-08-01
  • Anthropic (the company behind the AI model) released a guide for Opus 5, their newest AI model, which tells you to simplify your prompts (the instructions you give the AI) by removing certain lines.
  • The guide suggests giving Opus 5 the entire task at once, rather than breaking it into steps, as this newer model works better with complete instructions.
  • It's important to clearly state what the AI should not do, to avoid it adding unnecessary work or content, which could waste your time and usage limits (the amount of data you can use with the AI).
  • When the AI finishes a task, it will tell you about it, but you should set limits on how much it can write in its reply and in the actual task it's completing.
2026-07-31
  • **Forward deployed engineering (FDE)** is a hot AI trend where companies like OpenAI and Google DeepMind send expert engineers to work directly with customers to customize AI tools for real-world use.
  • **Factory**, a new AI tool, aims to automate software engineering tasks for businesses, acting as a "software factory" that builds and deploys code based on customer needs and feedback.
  • Unlike traditional consulting, Factory's engineers focus on improving their product by learning from customers, rather than doing the work for them, to create a scalable business model.
  • Factory's process involves capturing signals (like customer feedback or bug reports), prioritizing them, and automating the software development pipeline to create a smooth, AI-driven workflow.

Key points

What it is

  • A "skill" is a step-by-step guide you teach an AI to follow, like a recipe for solving a specific problem.
  • Skills help you codify your judgment and expertise, turning it into instructions the AI can follow reliably.
  • Skills are reusable and can be shared with colleagues to save time and ensure consistency.
  • A skill encapsulates the process, not just the example data, so the AI can generalize and apply it to new problems.

How to use it

  • Define clear success criteria before you start, so you know what a correct output looks like.
  • Structure your prompt using XML tags to separate instructions, context, input, and examples.
  • Evaluate the output against your success criteria and add diverse examples if needed to steer format and tone.
  • Refine your context by retaining useful evidence and tools, and leaving out distracting material.

Watch out for

  • Don't download skills from strangers online; build your own to avoid bias and ensure reliability.
  • Don't assume a task is complete when the AI says "Done"; always verify outputs against actual traces.
  • Avoid fixing mistakes mid-conversation; start fresh with corrected instructions to prevent compounding errors.
  • Watch out for "confidently wrong" outputs, where the AI generates plausible-sounding but fabricated data.

Lesson 1: What is Didn (topic) and why it matters

A "skill" (a reusable step-by-step guide you teach an AI) is the core building block for effective AI development. Instead of asking an AI one vague question, you first create a skill that encodes your specific process. For example, you spend 80% of your time in the "build phase," creating skills for tasks like market research. You then move to a "borrow phase," where you copy skills from colleagues. This is far more reliable than starting from scratch each time.

Don't download skills from strangers online. You must build your own, because a skill must encapsulate the process without being biased toward a specific topic. The AI generalizes the *process* you taught it, not just the example data. This lets it duplicate your best thinking on new problems.

Why does this matter? Because "attention didn't" get cheaper, even though code did. Human judgment—choosing which problems matter, when an approach is a dead end, and which results to trust—remains the realm of humans. A skill is how you codify that judgment. It’s how you turn your taste and experience into instructions the AI can follow reliably, ensuring it doesn't just give you the average answer. Instead of handing you text in a chat, a well-built skill drops a finished, scored idea straight onto your desktop. It’s the difference between a fancy autocomplete and a tool that multiplies your specific expertise.

Sources

Lesson 2: How to use Didn (topic): step-by-step

To use "didn" effectively, start by defining clear success criteria—the specific results that make a correct output visible. For example, if you're analyzing a report, your criteria might be: "distinguish between forward-looking analysis and historical summary." Write these requirements down before you begin, because vague goals lead to unreliable results.

Next, structure your prompt using XML tags. Separate instructions, context, input, and documents into visible compartments. For instance, wrap the task inside `<task>`, the background in `<context>`, and any examples in `<examples>`. This hierarchy acts as a debugging ladder—solve the simplest communication problem first.

After running your prompt, evaluate the output against your original success criteria. Test both positives (did the model do the task?) and negatives (didn’t it do something bad?). If the output misses the mark, add 3-5 diverse examples within your XML tags to steer format and tone. Show one normal case and one edge case, not repeated copies. Then rerun the evaluation. Let the score decide whether those examples earned their space.

Finally, refine your context. Retain useful evidence and tools, but leave distracting material out. A great prompt still fails with wrong surrounding information. Focus on the agnostic elements between examples—the patterns that map to future cases—not the specifics of a single project.

Sources

Lesson 3: Best practices and pitfalls

A common pitfall in AI tooling is assuming a task is complete when the agent says "Done." The most expensive failure is when an agent "fails quietly underneath" — for example, Claude once reported a migration finished successfully but had skipped 14% of records due to a constraint issue, discovered 11 days later when reports looked wrong. Always verify outputs against actual traces (records of what happened during execution) rather than trusting completion messages.

Mistakes often come from writing rules based on assumptions rather than observed failures. Instead, analyze your traces to identify patterns of what actually went wrong, then write rules based on real evidence. Similarly, when reviewing agent work, inspect the system rather than every line — one expert advises writing documentation, linters, and reviewer notes so "the system will catch this type of bugs" repeatedly.

Another trap is trying to fix mistakes mid-conversation by saying "No, fix this thing" — this makes the agent work from broken code plus new context, compounding errors. Instead, start fresh with corrected instructions. Also watch for "confidently wrong" outputs, where the agent generates plausible-sounding but fabricated data, like claiming specific vehicle delivery numbers that it invented completely.

Best practices include using natural language to design tests for edge cases, ensuring failures are "safe failures" rather than catastrophic ones, and defining measurable success and constraints before coding. Begin with budget, latency, and throughput targets, then ask how bad an error is and whether it can be caught — this determines if human-in-the-loop review is needed. Finally, know when to choose plain code instead of an LLM for stable rules that need no LLM decision-making.

Sources