AI Model Roundup
Last updated 2026-09-19What's new
- Deepseek's new V4.1 Flash model is much cheaper (80 times less) than GPT-6 Astra and performs similarly on some tasks, like business automation and coding.
- Deepseek can process a lot of text at once (about 1,500 pages) and its pricing is lower during off-peak hours (9 PM to midnight and 2 AM to 6 AM Eastern time on weekdays).
- You can use both Deepseek and Astra together, with Astra handling complex tasks and delegating simpler, repetitive tasks to Deepseek to save money.
- To use Deepseek, you need to create an account, get an API key, and set it up in a tool like Codex (a tool that helps you use AI models to work with files and run commands).
- Google's new Gemini Pro Checkpoint (a version of their AI model) shows promising results, potentially leading to Gemini 4.0 Pro, with impressive outputs like detailed SVG images.
- Deepseek's new AI model, Deepseek 4.1 Flash (an open-source AI model), is faster, smarter, and more efficient, using a unique architecture to outperform many flagship models.
- OpenAI and Enthropic (AI companies) are reportedly working on powerful new models, possibly solving complex math problems like the Hodge conjecture (a long-unsolved math puzzle).
- Scrumbbo's new Explain (a tool for AI coding agents) provides video explanations for code, making it easier for teams to understand changes in seconds.
- DeepSeek released a temporary test model, DeepSeek version 4.1 flash (a new AI model for tasks like coding and 3D modeling), which is faster and has the same pricing as its predecessor, but expires on September 10th.
- This new model supports multi-modality (handling different types of data like text, images, etc.) natively and is incredibly fast, generating about 400 tokens per second.
- DeepSeek is also cutting prices for its version 4 flash model on September 10th, making it cheaper to use.
- The new model excels in various domains like 3D environments and coding, but it can sometimes overthink and run unnecessary tests, slowing down the output.
- **New AI Models**: Several companies, like OpenAI (GPT6 Astra), Anthropic (Claude Fable 5.1), and Google (Gemini 3.8 Flash), released powerful new AI models (computer programs designed to perform tasks) this week.
- **Interactive Video Games**: Tools like H3 World and Solar WM (both open-source, meaning free to use and modify) turn video generators (like Miniax H3) into interactive video game engines (software that creates and controls video games), allowing users to explore AI-generated worlds in real-time.
- **Time Series Prediction**: Google's Times FM3 (open-source) is a tiny (330 million parameters, or settings the AI uses) AI model that can predict future data (like stock charts or weather) from various related signals without needing retraining (a process that teaches AI models using data).
- **3D Room Simulation**: ByteDance's Lucida (not yet open-source) is an AI that turns images of messy rooms into editable 3D simulations (digital copies) by identifying and separating objects.
- A new AI model called Qwen 3.8 Flash Next (a type of AI software that can run on your own computer) is now available, which is as powerful as top-tier models like Claude Opus 4.6 (a leading AI model used for business applications).
- This model can run locally (on your own computer) with relatively affordable hardware, making advanced AI more accessible.
- The model's performance was tested using OpenCode (a tool that uses AI agents to automate tasks), creating impressive results in a short time.
- Other Chinese AI models like GLM 5.3 Flash and DeepSeek V4 Flash are expected to offer similar quality improvements.
- **DeepSeek Harness (a free, open-source tool to run AI models locally)** lets you swap any AI model (like Cloud Opus or GBT 5.6) easily, unlike closed tools like Cloud Code (a paid AI coding assistant).
- It’s fully customizable—like rebuilding a car’s interior—because it’s open source, letting you tweak everything from tools to behavior.
- Four modes (Standard, PTC, Minimal, Creator) change how the AI works: Standard for full tasks, PTC for parallel work, Minimal for speed, and Creator for creative tasks.
- **Plugins (add-ons for extra features)** are powerful but risky—always check unknown plugins for safety before installing them.
- DeepSeek version 4 Pro, a new AI model focused on coding and task automation, is now available and ranks highly on AI benchmarks, offering strong performance at a low cost.
- DeepSeek 4 Pro can use tools, navigate code, and complete complex tasks, with pricing that's much cheaper than competitors like Kimi and GLM, making it a cost-effective choice.
- Miro, a team collaboration tool, now integrates with AI agents, allowing them to access and build upon shared team context, improving teamwork and efficiency.
- DeepSeek has introduced flexible reasoning efforts (low, high, max) for different task complexities and added OpenAI API support, making it easier to use with other tools.
- Enthropic's unreleased AI model, Model 2 (a potential successor to Mythos 5), scored 12.5% higher than Mythos 5 on an internal AI research benchmark (Cobbench version 2), suggesting it could automate significant portions of AI research tasks.
- OpenAI is rolling out paid usage resets for Codex (a tool that helps with coding) and plans a major upgrade next week, potentially making it one of the best coding assistants available.
- Gemini 3.7 Flash, a new AI model, has been released, showing significant improvements, especially for coding, knowledge work, and web development.
- Berta (formerly Data Crunch), a cloud provider, offers Nvidia's newest hardware for AI tasks, with features like confidential computing (keeping data encrypted) and competitive pricing.
- Meta (a company owned by Mark Zuckerberg) released Muse Code, a new AI tool (called an agent) that helps with coding tasks, like building apps, and it's much cheaper than similar tools from other companies.
- Muse Code can be used in a terminal (a special window for typing computer commands) and can be set up quickly with the help of another AI tool called Codeex.
- Codeex, an AI tool for developers, has updated its desktop app to include a new notifications bar that shows recent activities and their locations on your computer, making it easier to track tasks.
- OpenAI (a company making AI tools) might release a new, powerful AI model called Doug later this year, which could be even more advanced than their current models like Astra.
- Alibaba (a large tech company) has a new AI model, Kiana, that's showing strong performance in coding tasks, and you can try it out on a platform called Arena (a website where you can test different AI models).
- A new tool called ContextDev (a service that helps AI developers get clean data from websites) can help you extract structured data, generate AI-ready markdown, and pull brand assets like logos and colors from websites using a single API (a way for different software applications to communicate with each other).
Key points
What it is
- An AI model roundup is a regular update on the latest AI models (the core algorithms behind tools like ChatGPT) from major providers like Anthropic, OpenAI, and Google.
- These roundups help you compare models on their ability to handle complex tasks versus specialized strengths, and they inform your model choice.
- The model is only a small part (about 10%) of an AI implementation; the rest is people and process.
- "State-of-the-art" refers to the current best performance, but the best model overall may not be the best for your specific process.
How to use it
- Start by selecting your base model from the models panel, choosing options like DeepSeek, Claude Sonnet, or Qwen 3.8 Max depending on the task.
- For heavy, multi-step jobs, pick a high-end model like DeepSeek V4 Pro or Qwen 3.8 Max, and for quick tasks, use a lighter option like DeepSeek V4 Flash.
- Paste your task into the model, and let it infer the hidden steps; over-instruction can constrain the model.
- Validate the output carefully, and if the result is off, adjust your prompt, switch to a stronger model, and rerun.
Watch out for
- Don't overvalue benchmark rankings; a model's performance may not match its hype or name.
- Beware of single-model loyalty; the field shifts weekly, and it's smart to check current best models using tools like OpenRouter.
- Test models on your specific task to match the model to workload, not the hype.
- Newer models take many steps that are hard to audit, so validate the output carefully.
Tools named
- DeepSeek (a cost-efficient AI model), Claude Sonnet (a versatile AI model), Qwen 3.8 Max (a high-end AI model), DeepSeek V4 Pro (a high-end AI model for complex tasks), DeepSeek V4 Flash (a lighter, faster AI model), OpenRouter (a tool for seeing current best models).
Lesson 1: What is AI Model Roundup and why it matters
An "AI model roundup" is a periodic survey of the latest AI models (the core algorithms that power tools like ChatGPT) from major providers like Anthropic, OpenAI, and Google. These roundups matter because model performance shifts rapidly—a new release from one provider can suddenly outperform the previous leader on tasks you care about, changing what is "state-of-the-art" (the current best performance).
For developers, this landscape creates both opportunity and risk. Cheaper, more capable models raise profit margins and let teams build more for customers. But the moment you pick a model, it may become outdated, so you must continuously re-evaluate. A roundup helps you compare models on "agentic ability" (handling complex multi-step work) versus specialized strengths. However, "state of the art" is a confusing concept—the best model overall may not be the best for your specific process.
Crucially, the model is only 10% of an AI implementation. The other 70% is people and process. A roundup helps you choose the right tool, but without fixing your workflow and getting your team on board, even the best model fails. So view roundups as a starting point: they inform your model choice, but you must then invest in your own processes and decision-making—the only things that remain proprietary to your business.
Sources
- 2026-05-26 — AI Just Changed How You Run a Business Forever! (Tutorial)
- 2026-07-20 — The Real Reason Claude and ChatGPT Only Made You a Little Faster
- 2026-06-29 — Claude and ChatGPT Gets Smarter When You Change This One Setting
- 2026-06-01 — 20 days of compute vs 7 hours rethinking what state-of-the-art means Bertrand Charpentier, Pruna
- 2026-07-11 — Claude Code for Non-Coders (6 Hour Course)
- 2026-06-15 — The Hidden Cost of Letting AI Write Its Own Prompt
- 2026-05-30 — Google Remy, Grok 5, Mythos 1, New Atlas Robot, ASI and More AI News This Month!
- 2026-07-29 — I'm disappointed
- 2026-05-06 — Anthropic scares me.
- 2026-05-14 — Brutally Honest Advice For Someone Trying to Make Money with AI
- 2026-08-03 — Why Graph Engineering will 10x your ClaudeCodex
- 2026-06-01 — Grok's Low Censorship AI Video Model Shouldn't Exist Yet (Most Dangerous AI News)
- 2026-06-17 — Every Level of Hermes Agent Explained
- 2026-01-03 — The AI Choice You’ll Regret in 2026
Lesson 2: How to use AI Model Roundup: step-by-step
To run a model in AI Model Roundup, start by opening the models panel (the menu listing available AI options) and selecting your base model—DeepSeek is a solid default, but you can switch to Claude Sonnet or Qwen 3.8 Max depending on the task. For heavy, multi-step jobs like analyzing a stack of documents, pick a high-end model such as DeepSeek V4 Pro or Qwen 3.8 Max, because they handle many sequential steps without failing. For quick, single-shot tasks like summarizing one page, use a lighter option like DeepSeek V4 Flash, which is faster and more token-efficient (cheaper per request).
Here’s a concrete workflow: choose your model, then paste a task like "Extract key risks from these 10 contracts and rank them by severity." You don’t need to spell out every micro-step—modern models infer the 50–100 hidden steps themselves; over-instruction actually constrains them. After the AI finishes, validate (check) the output carefully, since newer models take many steps that are hard to audit. If the result is off, loop back: adjust your prompt with more context, switch to a stronger model like Claude Sonnet, and rerun. This plan-implement-validate cycle improves your code or documents with each iteration, and setting DeepSeek Flash as your main agent keeps costs low for routine work.
Sources
- 2026-07-15 — Nobody Prompts Like This Yet. OpenAI Wants You To
- 2026-06-29 — Claude and ChatGPT Gets Smarter When You Change This One Setting
- 2026-06-15 — Claude Skills + Hermes Agent 247 Agents
- 2026-05-04 — DeepSeek V4 + Claude Code BEST AI Coder!
- 2026-05-23 — Claude and ChatGPT Got More Literal. Your Old Prompts Are Backfiring
- 2026-05-09 — Claude and ChatGPT Hallucinate Less. That's Why They're Dangerous.
- 2026-01-31 — Your Code Gets Better With Every PIV Loop Cycle #aicoding #programming
- 2026-05-25 — ChatGPT vs Claude vs Gemini Is the Wrong Question
- 2026-07-05 — Full body waifus, Claude Fable is back, LongCat 2.0, mind-reading AI, live video editing AI NEWS
- 2026-07-08 — China's AI BAN!, Qwen 4, GPT-5.6 Thursday, Grok 4.5 Today, Deepseek AI Chip, & Claude AGI! AI NEWS
- 2026-07-13 — DeepSeek V4.1 GA Soon, GPT-5.6 SOL Nerfed HUGE Fable Update, US AI BAN Protests, & More! AI NEWS
- 2026-04-08 — The Next Layer After Prompt Engineering — Archon V3 Explained! 🚀
- 2026-05-26 — AI Just Changed How You Run a Business Forever! (Tutorial)
- 2026-06-18 — Claude Code Doesnt Matter. THIS Does
- 2026-08-03 — Qwen 3.8 Max Just Got DESTROYED By DeepSeek V4 Flash (RIP 5.45)
Lesson 3: Best practices and pitfalls
Choosing between DeepSeek Flash, Qwen Max, and similar models comes down to cost, speed, and task fit—not just leaderboard scores. DeepSeek V4 Flash is a frequent standout: it's "ludicrously cheap" with one million tokens of context (amount of text it can process at once) and uses 13 billion active parameters out of 248 billion total, meaning it runs efficiently. In one test, it scored nine out of ten building a full SaaS project, beating Qwen. DeepSeek also made a 75% pricing discount permanent, offering near-state-of-the-art performance at a fraction of typical cost.
A common pitfall is overvaluing benchmark rankings. Qwen 3.8 Max, despite hype, ranked twelfth, behind DeepSeek Flash and others. Early reports suggested strong performance, but real testing showed it "a little lackluster." Don't assume a "Max" name means top-tier. Also, beware of single-model loyalty; the field shifts weekly. Tools like OpenRouter let you see current best models, which is smart to check.
Best practice: test models on your specific task. DeepSeek excels at agentic tasks (jobs where an AI plans, uses tools, and retries) and coding, as shown by its strong OpenCode results. Its DSpark innovation speeds up output without quality loss, making it hard to overload. For user-facing help, a cheaper model that admits uncertainty beats one that hallucinates (makes up false info). Finally, run your own trials—match the model to workload, not the hype.
Sources
- 2026-08-03 — Qwen 3.8 Max Just Got DESTROYED By DeepSeek V4 Flash (RIP 5.45)
- 2026-05-24 — Claude Opus 4.8 Leaked, GPT 5.6 Spotted, Mythos 1 Preview, & Deepseek v4 Pro UPDATE! AI NEWS
- 2026-07-13 — DeepSeek V4.1 GA Soon, GPT-5.6 SOL Nerfed HUGE Fable Update, US AI BAN Protests, & More! AI NEWS
- 2026-06-06 — Hermes Agent NEW Super-App and DeepSeek v4 Catches Up To Opus 4.8
- 2026-05-11 — DeepSeek V4 Pro Inside Claude Code GOD MODE (FREE)
- 2026-08-02 — New Deepseek, Seedance 2.5, Minimax H3, Gemini Robotics, AMD models AI NEWS
- 2026-07-03 — Deepseek drops another HUGE breakthrough
- 2026-05-03 — DeepSeek V4 Slashes AI Costs 75 as OpenAI Ships GPT-5.5 and Anthropic Gates Mythos
- 2026-06-15 — Claude Fable 5 is Banned... Do THIS Right Now
- 2026-05-26 — The ChinaUS AI Frontier, May 2026 Qwen 3.7 Max, DeepSeek V4, and the Year the Gap Closed
- 2026-08-02 — DeepSeek Flash 0731 + OpenCode Full Apps Instantly (FOR FREE)
- 2026-07-03 — DeepSeeks New AI Breakthrough Just Broke AIs Limits
- 2026-08-04 — Qwen 3.8 Max IS OUT! Best Open Model (Fully Tested)