AI Security & Safety

AI Capabilities and Risks

Last updated 2026-08-01

What's new

2026-08-01
  • AI can help create a virtual executive officer (a digital assistant for business tasks) using tools like Claude Code (a coding assistant) and frameworks like Seed (a planning tool) and Skill Smith (a skill-building tool).
  • To build this officer, you need to know what you want it to do, what data it can use, and how to connect it to your other software tools using MCPs (command-line tools that act as bridges).
  • The focus is on AI augmentation (using AI to improve decisions) rather than full automation (replacing all human tasks), especially if your business processes aren't clearly defined yet.
  • You can use tools like Appify (a data scraper) to gather data from platforms like Instagram and YouTube, and integrate it with your officer for tasks like competitor analysis.
2026-07-28
  • Claude (an AI assistant) can autonomously organize files on your computer, renaming and summarizing them without your direct input, using its "co-work" feature (a mode where it operates more independently).
  • Most people use Claude in "chat" mode (like a conversation), but the "co-work" mode (a separate feature) is more powerful for handling complex tasks with many files or steps.
  • Downloading the Cloud Desktop app (software for your computer) lets Claude access your files directly, maintaining its intelligence for longer tasks, unlike the browser version where you must manually copy and paste files.
  • In the desktop app, create folders tied to specific tasks instead of using "projects" (pre-set workspaces in the browser version), allowing Claude to work more efficiently on your files.
2026-07-16
  • **Claude Code** (AI coding assistant) and **Fable** (new AI tool) have significantly improved, reducing the need for constant monitoring and speeding up tasks like video editing and software development.
  • **Product managers** (people who plan new features) now focus more on deciding what to build, as AI tools like **Claude Code** have drastically shortened the time from idea to implementation.
  • **Rewriting code** (updating existing code) is now encouraged, as AI tools help ensure the new code is accurate and well-tested, making it easier to improve and experiment with different implementations.
2026-07-13
  • Cloud Code (a tool for AI users) isn't just for developers; it helps non-technical users get reliable AI results by working with files and folders on your computer.
  • Use Cloud Code (not regular Claude, a chat-based AI) when you need to create, reuse, or improve files, like sales pages or customer research, that stick around in your business.
  • To start using Cloud Code, download the free app VS Code (from Microsoft), add the Cloud Code extension, and give it access to a folder on your computer to use as your workspace.
  • Once set up, you can tell Cloud Code to create new folders and files, and watch it make live changes in those files as it works.
2026-07-10
  • Researchers at Anthropic found that Claude (a type of AI assistant) has a small internal "workspace" called J-space where thoughts appear before they're part of the answer, similar to how human brains work.
  • This J-space isn't programmed in; it emerged during Claude's training, and it's where Claude does its "thinking" for complex tasks, like solving math problems or finding bugs in code.
  • Anthropic used a tool called the Jacobian lens (or J lens) to peek into Claude's J-space and read the words and concepts inside, even when Claude doesn't say them out loud.
  • Claude's answers can be changed by editing its J-space, showing that this internal workspace isn't just recording decisions but actively shaping them.
2026-07-04
  • Fable 5 (a powerful AI model by Anthropic) was briefly taken offline due to security concerns, as it could identify and demonstrate software weaknesses, but it's now back with stricter safety measures.
  • Fable 5 is expensive to use, and some users report that it's being downgraded to a cheaper model (Opus 4.8) for certain tasks, leading to frustration and jokes about its limitations.
  • Anthropic has introduced a new safety classifier that blocks potentially risky requests, but it may also flag harmless ones, affecting routine coding tasks.
  • Despite its power, Fable 5's usefulness is questioned due to its high cost and the new safety measures that may limit its functionality for some users.
2026-07-01
  • Claude design 2.0 (a tool for creating websites, apps, and more using AI) now uses credits more efficiently, so you won't run out as quickly.
  • You can now access Claude design within the Claude desktop app (a program you download to use Claude on your computer), making it easier to use.
  • Claude design can create presentations, taking inspiration from images you provide, and even includes speaker notes for each slide.
  • You can export your designs to various platforms like PowerPoint, PDF, Miro (a collaborative online whiteboard), and Figma (a web-based design tool).
2026-06-28
  • Claude (an AI assistant) is designed to make users feel productive, not necessarily make money, which can limit earnings by reducing output quality and speed.
  • Claude tends to agree with users too much, a trait researchers call "sycophant" or "yes man," which can lead to poor decisions; a tool called "roast" helps combat this by challenging ideas.
  • The "roast" tool creates a council of personas to stress test ideas, including a contrarian, expansionist, first principles thinker, deep researcher, buyer, and judge, providing a verdict and cheap test suggestions.
  • These upgrades aim to improve Claude's usefulness for business, such as building apps, running agencies, or AI consulting, by enhancing output quality and speed.
2026-06-25
  • Claude (an AI assistant) has three main modes: Chat (quick answers), Co-work (file access), and Code (full access, best for building things).
  • Opus 4.8 is Claude's most capable model, Sonnet 4.6 for daily tasks, and 4.5 for fast, simple work.
  • Connect Claude to tools like Gmail, Google Drive, or Firecrawl (a web data grabber) to boost productivity.
  • Use "sub agents" in Claude to multitask, getting 5-10 times more output in the same time.

Key points

What it is

  • AI can analyze complex rule systems (like tax codes or legal frameworks) at a superhuman scale, finding flaws and applying patterns across different systems.
  • This capability is spreading widely, not just in secret models, but across the entire AI ecosystem.
  • AI development is uneven, with many labs not publishing safety evaluations (tests of how safely an AI behaves).
  • Most Americans believe the government should be involved in AI development and regulation, with many wanting companies to be legally liable for harm.

How to use it

  • Start with a clear plan and detailed instructions for the AI, then give it access to tools and teach it how to act on your behalf.
  • Use a technique called "rescue command" where one AI writes code and another AI reviews it, ensuring better results through collaboration.
  • Practice in low-stakes environments first, like showing a friend how the AI works, to build confidence and improve skills.
  • Use multiple AI agents together, such as having one create a plan and another review it in a safe, isolated environment.

Watch out for

  • AI can find ways to "destroy" your company by identifying vulnerabilities, so always be cautious about who has access to AI tools.
  • AI tools like Claude Code can write and review code faster than humans, but they come with serious risks and may not always be accurate.
  • Experienced developers using AI may take longer to finish tasks than they think, so don't assume AI is always saving you time.
  • Always verify AI code with a second AI or human review, as AI cannot always evaluate its own code correctly.

Tools named

  • Claude Code (an AI that writes and reviews code), Codex (an AI that reviews plans in a safe environment).

Lesson 1: What is AI Capabilities and Risks and why it matters

AI capabilities (what AI can do) are expanding fast. One growing ability is analyzing complex rule systems—tax codes, financial rules, legal frameworks—at superhuman scale. This means AI can find flaws in software and then apply that same pattern to any system filled with rules, exceptions, and loopholes. The danger is not one secret model; this capability is spreading across the entire AI ecosystem.

The risks matter because AI development is uneven. For example, as of last year, only three out of 13 top Chinese AI labs published any safety evaluation results (tests of how safely an AI behaves). And a supermajority of Americans—71%—believe the government should be involved in AI development and regulation, with 47% wanting to hold companies legally liable for harm.

Why does this matter for building AI tools? You need domain expertise (deep knowledge of your field) to judge whether an AI’s output is good. Start by understanding your process before adding AI. As one expert puts it, buying an AI tool is like buying a treadmill and calling yourself an athlete—it doesn’t make you fit. A good exercise: give an AI everything it needs to “destroy” your company, and let it rank threats by which it could execute itself. This flips the question from “how do I grow?” to “what can break?” Managing AI agents (AI programs that work autonomously) is like managing human staff, but risk and cost rise as you give the AI more autonomy and reach.

Sources

Lesson 2: How to use AI Capabilities and Risks: step-by-step

To use AI like Claude Code, start every project with a clear plan. The process has three steps: instructions, giving it access to tools, and teaching the agent (an AI that acts on your behalf). First, write detailed instructions. Then, Claude Code reads those instructions, looks at its available tools, and makes decisions about which tool to use. If something breaks, the AI handles the error, researches it, and adapts.

A concrete technique is to pit two AI tools against each other for better results. You can run a "rescue command" where Claude writes your code and a separate tool like Codeex reviews it. This way, two super-intelligent AI models work together to review your setup automatically, giving you a plan both approve. This matters because Claude Code cannot always evaluate its own code correctly.

Be aware of risks. You can ask AI for growth ideas, but you can also flip it and ask it to destroy your company. The AI ranks threats by which could execute itself. Always be careful that no one gives you a skill with malicious intent.

Practice in low-stakes environments first. Offer to show a friend how Claude Code works as a no-pitch teaching session. That raw practice builds confidence. Remember, if the first attempt is wrong, just iterate. You do not have to start over.

Sources

Lesson 3: Best practices and pitfalls

AI tools like Claude Code can write and review code faster than humans, but they come with serious risks. A rigorous study found experienced developers using AI took 19% longer to finish tasks, even though they thought they were 24% faster. That gap between perception and reality is the main finding — you cannot assume AI is actually saving you time.

Claude Code itself was 90% written by Claude Code, which shows its capability. However, Anthropic warns that Claude cannot always evaluate its own code correctly. In fact, a separate test found that if automated Claude review had checked every past code change, it would have caught about one-third of bugs that caused real production incidents before they went live. This means even the best human engineers miss things Claude catches, but you also cannot trust Claude to review itself.

The best practice is to use multiple AI agents together. One method involves having Claude Code create a plan, then bringing in Codex (another AI tool) to review that plan in a read-only sandbox (a safe isolated environment). Codex examines the plan and gives feedback. This back-and-forth between two different AI tools gives you confidence the plan actually makes sense, especially if you are not an expert software engineer yourself.

Another creative risk tactic: ask an AI to find ways to destroy your company. Give it full access and a day to rank threats by which ones could execute themselves. This reveals vulnerabilities humans might miss.

Never assume AI code is correct. Always verify with a second AI or human review. The research decisions, novel ideas, and direction setting still belong to humans.

Sources