Media & Design

Images (topic)

Last updated 2026-09-22

What's new

2026-09-22
  • A skill is like a recipe that an AI agent (a tool that follows instructions) can use to consistently produce a desired output, such as a specific dish or a professional email.
  • To create a skill, start with the final output you want and work backwards to figure out the steps needed to create it, similar to reverse engineering a favorite dish from a restaurant.
  • Skills are written in a simple language called markdown (a way of formatting text using symbols like # and *) and include a section called YAML front matter (a way to add extra information) that describes the skill's name and when to use it.
  • Skills can be as simple as a prompt to make an email sound more professional or as complex as a process to analyze stocks and recommend investments.

Key points

What it is

  • AI can use images as inputs to learn (like sorting pictures into categories) and outputs to create (like generating pictures from text descriptions).
  • Images carry more information than text and can be turned into other media like videos or thumbnails.
  • AI helps scale creative work by generating visuals quickly, but translating personal taste into AI tools is challenging.

How to use it

  • Give AI tools a clear job, like describing style, subjects, and actions for video generation.
  • Use reference images to guide AI in creating new images or editing existing ones.
  • Provide all materials up front, like logos and brand guidelines, and designate a primary visual reference for consistency.

Watch out for

  • AI-generated images can look artificial, and your taste might not translate well.
  • After several edits, AI might drift from the original reference image, so be explicit about what not to change.
  • Older workflows might not generate images with text, and earlier edits could degrade over time.

Tools named

  • Nano Banana (an AI image generation tool), ChatGPT (a conversational AI model with image capabilities), Images 2.5 (a ChatGPT image model).

Lesson 1: What is Images (topic) and why it matters

Images, in AI development, are both an input and an output. On the input side, AI models can learn from pictures. One example is image classification (sorting pictures into categories), where AI now matches or beats humans on most benchmarks. On the output side, tools like Nano Banana (an AI image generation tool) let you describe an idea in words and get a picture back, helping you visualize a vision you have in your mind.

Why does this matter for building with AI? First, images can carry far more information than text. One PNG (a single image file) can encode an entire level design, including all the textures. Second, images are the raw material for whole media libraries: from one clip or photo you can produce thumbnails, posters, storyboards, key art, and concept images. Media is becoming fluid, where images turn into videos, and videos turn back into images. Third, images are how you scale creative work. You can generate visuals with AI image models and use HTML (the code that structures web pages) as the repeatable backbone, producing carousels at scale instead of sitting at your computer each time.

The hard part is taste. It is extremely difficult to translate your taste (your personal sense of what looks good) into these tools, which is why fully AI-generated images can still look obviously artificial. The opportunity now is moving past single images that work and building systems that produce complex, meaningful art and design.

Sources

Lesson 2: How to use Images (topic): step-by-step

To use images with AI tools, start by giving the tool a clear job. In video generation, one workflow says your description should first establish the style, subjects, composition (how elements are arranged), and scene anchors (fixed objects in the frame), then describe the next action. Character identity, clothing, colors, and key objects should stay consistent. The recommended structure is first frame anchor, action onset, continuous development, result or reaction. One creator copies that structure into ChatGPT, uploads an input image, and asks for a description.

ChatGPT's newer image models are repeatedly praised. ChatGPT image two is called "really solid," and you can take information from one place and generate with it. Another approach: draw over an existing photo or write instructions directly on it, then plug the annotated image into ChatGPT and say "use GPT image 2.5 to edit the image based on the annotations and instructions in the image." The result gave an edited image.

Reference images matter. You can give a model reference images and even reference video, then just show up with an idea and a style. In the Nano Banana playground, one example dropped in an image of the creator, an image of Scotty getting his green jacket, and asked to put the creator in that environment. Iterating after the first image is easier than attaching images again. One warning: in older workflows, images could not be generated with text.

Sources

Lesson 3: Best practices and pitfalls

When you ask ChatGPT to change an image, there is a common pitfall: after several generations, the result drifts away from your original reference image. A concrete example comes from one creator who noticed the tool quietly deleted small captions and subtext that were present in the original photo. The fix is to be explicit — tell it not to change the reference image at all. If you only give it vague instructions, it will decide for itself what to remove.

Best practice starts with how you provide material. Attach everything up front: the icon version of a logo, the text version, and the brand guidelines, then explain what each attachment is for. When you want consistency across a series, designate one image as your primary visual reference (the master image others copy), stating that it governs character design, styling, and mood.

Another pitfall appears during longer sessions. Earlier edits used to degrade, but Images 2.5 (a ChatGPT image model) reportedly follows editing instructions more reliably across multiple edits, so each new change builds on prior work without losing quality.

Finally, you can correct course mid-task. If the model does something wrong, stop it and send a new message — it still understands the context. Some people go further: generate roughly twenty images, explain what you like and dislike, then request twenty more, refining the output until it matches your intent.

Sources