Media & Design

Transformers and Visual AI

Last updated 2026-07-31

What's new

2026-07-31
  • Instead of one overloaded AI agent, create a team of specialized agents, each handling a different task, like researching or designing, to save money and improve efficiency (AI agents are like digital workers that can perform tasks based on instructions).
  • HyperAgent, a tool by the Airtable team, lets you build and manage teams of AI agents, each with its own role, tools, and memory (Airtable is a popular online database tool).
  • Specialized agents outperform general-purpose agents, so focus each agent on one task, like monitoring AI developments and summarizing important changes.
  • HyperAgent's live mode keeps your agent team working 24/7, autonomously handling tasks and providing updates.
2026-07-28
  • Ling Bot World 2.0 (a new AI tool) creates interactive, ever-changing worlds in real time, like a fantasy game where you can explore, fight, and cast spells, with no predefined map or storyline.
  • Unlike other AI video tools, Ling Bot World 2.0 maintains high-quality visuals and stability for long periods, allowing for extended exploration and interaction.
  • The system uses two AI agents (intelligent helpers) called the "pilot agent" and "director agent" to generate behaviors, events, and new elements in the world, making it more than just a simple text-to-video generator.
  • Ling Bot World 2.0 is open-source (free to use and modify) and can be accessed through a platform called Reactor, with the potential for collaborative storytelling, gameplay, and even robot simulation in the future.
2026-07-25
  • Anthropic (a company making AI tools) delayed the launch of their new AI model, Claude Opus 5, but it's expected to be powerful, possibly matching or even surpassing other top models like Fable 5.
  • OpenAI (another AI company) added voice control to their ChatGPT (a popular AI chatbot) desktop app, letting users control their computer and run multiple AI tasks using just their voice.
  • Google (a tech giant) rolled out Gemini Spark (a personal AI assistant) to AI Pro subscribers, making it more accessible for users to interact with their AI tools.
  • There are rumors about new, advanced AI models from OpenAI, including one that might outperform the upcoming GPT-6, with significant improvements in math capabilities.
2026-07-19
  • AI agents (computer programs that do tasks for you) can now run a business with little human input, acting as a team with one manager and specialists for different tasks like social media or ads.
  • These agents use simple English instructions and remember important business details, creating content and handling tasks without sharing information between roles to keep each task specialized.
  • The system connects to your existing apps (like email or project management tools) through connectors (links that let the AI access these apps), turning the agents from chatbots into active workers.
  • Beginners can use pre-built agent setups (called plugins) to avoid creating each agent from scratch, making it easier to start automating tasks.
2026-07-13
  • Abot World is a new open-source (free software you can modify) tool that creates endless, interactive 3D worlds you can explore in real-time, even on a consumer-grade graphics card (Nvidia RTX 5090).
  • Proxy Pose is a new AI system that can track the 3D movement of almost any object in a video, even tricky ones like transparent or reflective surfaces, and it's available for free download.
  • CFI Image is a new, compact open-source image generator that creates high-quality, detailed images quickly and efficiently, outperforming some larger models and supporting various styles like photorealistic and anime.
2026-07-10
  • OpenAI's new GPT 5.6 (a powerful AI model) comes in three sizes: Soul (largest and most capable), Terra (mid-range), and Luna (smallest and fastest).
  • The free ChatGPT desktop app (a tool for using GPT 5.6 on your computer) lets you use multiple AI agents (specialized AI tasks) to work on files and folders simultaneously.
  • GPT 5.6 can create complex projects, like a web app with an anime character that talks to you in real-time, using just one instruction (prompt).
  • It can also simulate physics, like liquid splashes, with adjustable settings and hand-tracking via webcam, all from scratch without pre-made tools.
2026-07-07
  • A new AI called Musev VIT (a tool for reading and understanding sheet music) can recognize and classify sheet music better than other vision models, trained on millions of pages of sheet music.
  • Chinese food delivery company Muan released Longat 2.0, a large AI model trained without Nvidia GPUs (specialized graphics cards usually used for AI training), using their own AI super pods (specialized chips for AI tasks) instead.
  • Longat 2.0 is a 1.6 trillion parameter model (a measure of the model's complexity), designed for coding and long context work, and is open-source (free to use and modify) under the MIT license (a permissive open-source license).
  • Liveedit is a new AI that can edit videos in real time (as the video is playing), allowing for quick and easy video editing.
2026-06-28
  • A new tool called an AI signal engine (a visual map that shows connections between AI news) has been created using open-source software and is available for anyone to install and modify.
  • The tool uses a main agent called Hermes (a type of AI that can perform tasks and make decisions) to scan and analyze AI news in real-time, and it can be hosted on a separate computer called a VPS (a virtual private server, which is like a personal computer on the internet) for security and privacy.
  • The tool is designed with safety and security in mind, with features like password protection, verification flags, and a visible log to track the agent's actions.
  • The tool is open-source, meaning anyone can download and modify the code, and it can be run on a local computer or a VPS using Docker (a platform that allows developers to package and run applications in containers).
2026-06-22
  • Fable, a new AI tool (software that learns and makes decisions), showed promise in reasoning through complex work, hinting at a future where AI can be used for serious tasks.
  • Abacus AI introduces "apps in AI agents" (AI tools that create interactive applications), allowing users to generate and interact with 3D models, diagrams, and data visualizations directly within their workflow.
  • Fusion Agents is a multi-agent system (a group of AI tools working together) that combines a planning model with smaller worker models to tackle complex tasks, similar to how Fable operated.
  • These developments suggest a shift from focusing solely on AI models (the brain) to building systems (the body) around them, enabling AI to create tools, use infrastructure, and deliver practical outputs.
2026-06-19
  • A new open-source AI model called GLM 5.2 (a free, community-developed AI tool) was released, outperforming leading models like GPT and Gemini in various tests.
  • To maximize its potential, use GLM 5.2 with frameworks like OpenClaw (a tool that helps AI work on complex tasks) or Zcode (a free, user-friendly AI assistant for Mac, Windows, and Linux).
  • GLM 5.2 can create advanced projects, like a 3D interactive Earth model, with some additional guidance, showcasing its impressive capabilities for open-source AI.
  • The model can also handle complex, multi-tool tasks, such as creating a promotional video with voiceovers and animations, demonstrating its versatility for everyday workflows.
2026-06-16
  • A new AI tool called Scale 2 (open-source software for animating characters) can transfer motion from one video to another, even with multiple characters or non-human subjects, and is available to download and run locally.
  • Actionable World Representation (a system for creating digital twins of real-world objects) can model how objects move, bend, or change, and is useful for training robots to interact with the real world.
  • Oscar (a world model for robots) can predict outcomes of robot actions and work across different robot types, helping to train robots in virtual environments.
  • Google's Gemini 3.5 Live Translate (a real-time translation tool) can translate speech in one language to another while maintaining the original speaker's voice.
2026-06-13
  • Miniax M3 is a new, powerful, and affordable open-source AI model (software anyone can use and improve) that can handle text, images, audio, and video, even outperforming some expensive competitors.
  • Miniax Code is a 24/7 AI workspace (like a virtual assistant) that can code, browse, automate tasks, and work in teams, making it more than just a simple chatbot.
  • Together, Miniax M3 and Miniax Code can create complex applications, games, and designs with minimal input, offering a unique and efficient way to harness AI capabilities.
  • The setup is accessible for free, with optional paid plans for advanced features, making it a cost-effective choice for both beginners and professionals.
2026-06-07
  • OpenAI's Codex (a tool that helps write and understand code) got an update for building websites, and ChatGPT (a chatbot that uses AI) got a memory update for better conversation flow.
  • Google released Gemma 4 12B, an AI model that can run locally on your device using LM Studio (a software for running AI models), and Ideogram 4, a top open-source image generator.
  • New AI models for creating realistic images, expressive text-to-speech, and generating music and video were released, with a focus on open-source (software anyone can use and modify) tools.
  • Rumors about upcoming AI models like GPT-5.6 (a potential new version of OpenAI's language model) and Mythos/Oceanus (a new model from Anthropic, another AI company) suggest improvements in spatial understanding and realistic outputs.

Key points

What it is

  • Transformers are a type of AI design that looks at all parts of input (like text or images) at once, instead of one piece at a time.
  • Vision transformers (ViTs) adapt this design for images, cutting them into patches and processing those patches like words.
  • Visual AI can now generate highly realistic images and handle messy real-world data, making it useful for complex tasks like 3D editing and video processing.

How to use it

  • Start with a text prompt describing what you want to see, or upload a reference image to generate a descriptive prompt using tools like Google Omni.
  • Follow the PIV loop: Plan what you want the AI to do, let the AI implement it, then Validate the results, improving each time.
  • Integrate Visual AI models with popular software like Photoshop and Blender for autonomous editing, or combine with HTML for presentations.

Watch out for

  • Define what success looks like before starting a project, as many AI pilots fail due to lack of clear goals.
  • Models may struggle with real-world messiness, so ensure training data includes chaotic scenes with overlapping objects and changing lighting.
  • Avoid using autonomous agents without first defining success metrics, as they may do the wrong things efficiently.

Tools named

  • Google Omni (AI tool that generates prompts from images), FLUX 2 (AI image generation model), Claude Code (AI with multimodal support for text and images)

Lesson 1: What is Transformers and Visual AI and why it matters

The transformer (a neural network design that processes all parts of input at once) emerged in 2017 when eight Google researchers published "Attention is all you need." Unlike older neural networks that read text one word at a time left to right, transformers use self-attention (a mechanism that mixes spatial information) to look at the entire input simultaneously. This design proved incredibly effective.

Vision transformers (ViTs) adapted this breakthrough for images. Instead of analyzing pixels one by one, they perform a "patchify operation" cutting an image into patches, like 16 by 16 squares, and process those patches the way a transformer processes words. This approach let transformers "finally eat vision" and outperform older convolutional neural networks for image tasks.

Visual AI now generates images with such fidelity that it is often impossible to tell they are AI-generated, with accurate details like hands and veins. AI can also be trained on messy real-world data, where objects overlap, lighting changes, and parts are blocked, so it performs well in chaotic situations. Tools like Claude Code now include multimodal support (ability to process both text and images) giving AI "eyes" to see visual representations instead of just a chat window.

This matters because Visual AI transforms AI development from generating text in a terminal to doing real production work inside software tools, handling complex visual workflows like 3D editing and video processing. Moving AI from pure language into vision allows it to understand and act on what it sees, making it far more useful for real-world applications.

Sources

Lesson 2: How to use Transformers and Visual AI: step-by-step

Transformers (a neural network architecture that processes sequences) power most modern AI tasks, from text translation to image generation. At their core, they translate one form of data into another — text to images, images to 3D, or video to audio. To use Transformers for visual AI, start with a text prompt describing what you want to see. Google Omni lets you go further: upload a reference image, and it generates a descriptive prompt from that image, giving you greater creative control. You can also use a text prompt alone to generate new visuals.

For building practical projects, follow the PIV loop: Plan what you want the AI to do, let the AI implement it, then Validate the results. Each time you loop, the AI improves. For example, to create an AI news digest, plan the content structure, have the AI gather articles and turn them into a report, then validate the output by reviewing it. Extend this by adding a company researcher agent to your project — as your automation grows, it becomes easier to add new functionality.

Visual AI models now integrate with popular software like Photoshop and Blender, allowing AI to autonomously edit images or manipulate 3D scenes. Some systems even use recursive multi-agent setups, where multiple AI agents work together and improve over time. For creating carousels or presentations, combine an AI image model for visuals with HTML for text slides — this gives you creative imagery with a repeatable, scalable backbone.

Sources

Lesson 3: Best practices and pitfalls

Transformers (a type of neural network design from a 2017 paper called "Attention is all you need") changed AI by reading all words at once instead of one-by-one, using "self-attention" (a mechanism that understands relationships between words). This design now dominates visual AI too, replacing old convolutional neural networks (CNNs) by treating image patches like tokens (small pieces of data).

But beginners make costly mistakes. A study found 95% of generative AI pilots (experimental projects) fail to deliver results. Gartner predicts 40% of agentic AI projects (AI that acts autonomously) will be cancelled by 2027. The core problem: we tell AI what to do but never define what success looks like. For visual AI specifically, models struggle with real-world messiness — overlapping objects, changing lighting, partial occlusion (things partly hidden). Good training data must include chaotic scenes where objects are rearranged or blocked.

Best practices start with choosing the right model. New open-source image models now handle text posters and infographics well. For video, focus on control and editing, not just realism. Tools like FLUX 2 produce nearly undetectable AI images, but even that model had early rough versions. Test thoroughly — one expert's AI agent found bugs he'd missed for 20 years in his own hand-tuned model.

Avoid "autonomous agents" (AI that works independently) unless you define success metrics first. Many companies pour 20-50% of digital budgets into AI automation while building systems that brilliantly do wrong things. Start small, use visual harnesses that give AI "eyes" to see what you see, and always validate against real-world conditions, not clean test data.

Sources