Speaker Diarization and Analysis
Last updated 2026-09-16What's new
- AI agents (computer programs that use AI to complete tasks) often fail when moved from testing to real-world use, with latency (delays in response) being a major issue.
- Companies often use frameworks like LiveKit or PipeCat to build AI agents, but these can struggle in production due to unanticipated failure modes.
- The speaker's company, with 14 years of experience, offers a full-stack AI agent platform, including programmable AI agents and a no-code visual builder called AI agent studio.
- They emphasize the importance of balancing cost, intelligence, and latency in AI agent development, noting that recent advancements in AI (like reinforcement learning) can impact these factors.
- GPT6 Astra (OpenAI's latest, pricier AI model) can speed up your existing apps or websites by reviewing and suggesting performance updates, drastically improving loading times.
- It can also run security audits on your apps, helping to identify and fix potential vulnerabilities that could put your users at risk.
- You can use Astra to design user interfaces (UI) by providing reference images, helping you create visually appealing and consistent designs.
- Astra can automate tasks on your computer, like renegotiating bills or canceling unused subscriptions, saving you time and potentially money.
- **Codex** (an AI tool you install on your computer) helps businesses automate tasks like sales, marketing, and operations to save time and increase profits.
- Learn to turn sales calls into **custom proposals with e-signatures and checkout** (ready-to-send documents that clients can sign and pay for online).
- **Repurpose content** (reuse one video or post to create multiple social media posts, newsletters, or short clips) to save time and reach more people.
- Build an **after-hours voice agent** (an AI that answers calls and handles leads automatically) to catch missed opportunities and improve customer service.
- Hippocratic, a healthcare AI (artificial intelligence, or computer-based learning) tool, has facilitated over 200 million clinical conversations, aiming to provide abundant, safe, and accessible care for all.
- The AI can handle tasks like accurate drug name recognition, vital sign clarification, and medication stoppage detection, ensuring clinically safe interactions.
- Hippocratic's AI can escalate calls to a real-time nurse when necessary, demonstrating its ability to handle complex and urgent health situations.
- The tool's unique vertically integrated stack (custom-built system) optimizes for both intelligence and low latency (delay), making it suitable for real-time, two-way conversations on the telephone.
- To boost productivity, save repeated inputs as files in a main folder (like "Agentic OS" or "Agentic operating system," a central place for your work files) and reference them in Claude (a tool that helps you work faster).
- Create a "brand voice" file to ensure AI-generated content sounds like you, not generic AI; include platform-specific guides (e.g., LinkedIn, emails) and update it regularly.
- Establish a "visual identity" file to maintain consistent branding (colors, fonts, logo) across all visual outputs, like images or slides, without repeating instructions.
- Joy AI videoedit (a free online video editor) lets you edit videos in real-time using simple text commands, like changing outfits or adding accessories, and it works as well as paid options.
- Tencent's Scope (a free online tool) helps AI video models control camera movements, like panning or zooming, by specifying a path for the camera to follow.
- Deepseek's V4 Pro (a free AI model) is a powerful, cost-effective option for tasks like coding and problem-solving, but it requires a lot of computing power to run.
- Deepseek also released its own harness (a framework that helps the AI model work better), which is still in development but will likely be the best way to use Deepseek models in the future.
- Warp (a coding tool) introduced a new way to organize notes using LLM knowledge bases (AI-powered tools that help manage and understand your notes), turning messy thoughts into structured, searchable info.
- Voice dictation tools like Handy (free, open-source software) and Voice Inc. ($20 lifetime fee) can quickly capture raw thoughts, which AI can later organize and connect.
- You can enrich notes with tags, backlinks (links to related notes), and categories, making it easier for you or AI to find and connect ideas.
- AI can create visualizations (like graphs) and wikis (linked web pages) from your notes, helping you explore and understand your research or interests better.
- This update introduces Cloud Code, a tool (software you pay for monthly online) that helps automate marketing tasks, like creating ads, personalizing emails, and setting up appointment systems.
- The course teaches how to use Cloud Code at different levels, from simple prompts to advanced cloud routines (automated tasks that run without your input).
- You'll learn to build your own analytics platform (a system to collect and track data) and automate follow-ups, making your marketing efforts more efficient.
- Cloud Code requires a paid subscription, but the investment can lead to significant returns, as demonstrated by the instructor's business success.
- ByteDance and MiniMax released new video models, with MiniMax's being open source (software anyone can use and modify), and DeepSeek unveiled a cost-effective, high-performance model.
- Netflix's open-source AI, IDV2V, can style videos by editing a single frame while preserving characters' identities and movements, generating up to 720p videos.
- Whisper 2, an open-source transcription tool, converts audio to text with precise word timing, supports multiple languages, and offers verbatim or polished transcript options.
- DeepSeek V4 Flash, a smaller and cheaper model, matches or exceeds larger models' performance in coding, software engineering, and cybersecurity, and is available for download.
- OpenAI's new ChatGPT voice feature lets you control AI agents (AI programs that can do tasks for you) with your voice, making it easier to use AI anywhere, anytime.
- Unlike old voice features that just type what you say (called dictation), this new tool lets you command multiple AI agents across all your devices (like your phone, tablet, or computer) to get things done instantly.
- ChatGPT voice can check the status of your projects, give you updates, and even create new AI agents to handle tasks for you, like fixing errors or building websites.
- This tool can greatly increase your productivity (getting more done in less time) by removing barriers between having an idea and seeing it become reality.
- You can now use a smart AI tool (Local AI) that runs completely free, offline, and privately on your own computer, with no data leaving your device.
- Local AI uses "open weights" (free, downloadable AI models) that you can own and use without paying for access or worrying about data privacy.
- The free AI models have improved significantly, with some like GLM 5.2 (a large, capable AI model) performing close to top paid models on many tasks.
- Using Local AI gives you control and privacy, as you're not renting a service from a company that can change or restrict access, and your data stays on your machine.
- Open Claw (a tool for managing secure access to apps) now supports "trusted proxy auth mode," a feature that removes the need for tokens and device pairing, improving both security and user experience.
- This mode uses an identity-aware proxy (a security approach that combines identity provider, policy engine, and reverse proxy) to control access, making it easier to secure internal apps.
- The update allows configuration through both onboarding and TUI (text-based user interface), providing flexibility in setup.
- The speaker highlighted the active Open Claw community, noting how contributors quickly addressed a bug they missed, demonstrating the power of open-source software (OSS).
- DeepMind has released Gemma 4, an open model with audio understanding capabilities that can run on edge devices (small, portable computing devices).
- Gemini 3.1 flash life is a real-time, full duplex (two-way) sound-to-sound conversational model that also supports text and vision inputs.
- Echo Script, built with Google AI Studio (a platform for creating AI-powered apps), uses Gemini 3 to analyze audio recordings, extracting details like speaker names, timestamps, languages, emotions, and summaries.
- Gemini 3 can handle complex audio tasks, such as transcribing overlapping speech, switching between languages, and identifying speaker emotions.
Key points
What it is
- Speaker diarization is a process that identifies who spoke when in an audio recording, assigning each speech turn to a specific speaker.
- It goes beyond basic transcription (converting speech to text) by labeling who said what, providing context about the conversation.
- This is important for richer conversation analysis, as knowing who said what can be as crucial as the words themselves.
How to use it
- First, use voice activity detection (VAD) to identify when anyone is speaking, then segment those speech regions into turns.
- Next, assign each turn to a specific speaker using a diarization pipeline, then merge that with your transcription from a speech-to-text model.
- To build a voice profile, gather your past content, transcribe those files, and extract patterns across dimensions like word choice, pacing, and filler words.
Watch out for
- Common errors include confusion (assigning a word to the wrong speaker) and segmentation mistakes.
- Diarization accuracy can be affected by factors like distant microphones, speaker changes, cross talk, interruptions, and code switching (changing languages mid-conversation).
- Performance varies completely depending on the use case, so always benchmark your diarization pipeline on data similar to your target use case.
Tools named
- Pyannote (an open-source tool for speaker diarization), Whisper (a speech-to-text model), Claude (a tool for analyzing writing or speech examples).
Lesson 1: What is Speaker Diarization and Analysis and why it matters
Speaker diarization is the process of answering the question "who spoke when" in an audio recording. It automatically detects how many speakers are present and assigns each speech turn to a specific identity (for example, labeling one voice as "green" and another as "black" in a two-person conversation). This is different from plain transcription, which only captures what was said without tracking who said it. As one expert noted, "who said what is usually just as important as what was said."
The importance for AI development is twofold. First, it enables richer conversation analysis. Without speaker diarization, a transcript is just a block of text—you lose context about which speaker asked a question, which speaker responded, and how each person's state of mind might differ. Second, speaker diarization remains a difficult unsolved problem. Performance varies completely depending on the use case: a clean telephone conversation might achieve low error rates, but background noise (like traffic or horns) or overlapping speech can dramatically degrade accuracy. This means AI developers must carefully benchmark models for their specific environment rather than assuming a one-size-fits-all solution.
In practice, speaker diarization sits alongside related tasks like speaker identification (assigning a known name to each voice) and is often used as a preprocessing step for building voice agents or analyzing call center conversations. Because the problem is still active in research, state-of-the-art performance depends heavily on the quality of the input and the number of speakers involved.
Sources
- 2026-06-05 — Beyond Transcription Building Voice AI That Understands Conversations Herv Bredin, pyannoteAI
- 2026-06-06 — I Taught Claude Code to Build You a Personal Brand, Watch This
- 2026-05-09 — Why TTS Models Now Look Like LLMs Samuel Humeau, Mistral
- 2025-12-15 — n8n's New Chat Hub Release What You Need to Know
- 2025-12-08 — n8n 2.0 is Here (What You Need to Know)
- 2026-03-28 — Gemini 3.1 Flash Live Just Changed Voice Agents Forever
- 2026-05-09 — Voice AI when is the Her moment Neil Zeghidour, CEO, Gradium AI
- 2026-05-04 — Training an LLM from Scratch, Locally Angelos Perivolaropoulos, ElevenLabs
- 2026-03-11 — Google's New Model + Claude Code Just Changed RAG Forever
- 2026-05-09 — Codex Super App, OpenAI Chaos Drama, Gemini 3.2 Pro In Arena, GPT-Realtime-2, & NotebookLM Update!
Lesson 2: How to use Speaker Diarization and Analysis: step-by-step
Speaker diarization answers "who spoke when" in a recorded conversation. It goes beyond basic transcription (converting speech to text) because raw transcription gives you words but no speaker labels. To get a speaker-attributed transcription, you first run voice activity detection (VAD), which tells you when anyone is speaking. Next, you segment those speech regions into turns, then assign each turn to a specific speaker using a diarization pipeline.
For a concrete example, take an audio file of two people chatting. Use an open-source tool like pyannote, which runs a diarization algorithm. It outputs something like "Speaker A from 0:00–0:12, Speaker B from 0:13–0:25." You then merge that with your transcription from a speech-to-text model (e.g., whisper) to label each word or phrase, producing a full speaker-attributed transcript. The process isn't perfect—common errors include confusion (assigning a word to the wrong speaker) and segmentation mistakes—so visualize the output against reference annotations to spot and fix mistakes.
To build a voice profile (a detailed document capturing your unique speaking style), start by gathering your past content—podcasts, videos, or chats you've recorded. Transcribe those files, then extract patterns across six dimensions, such as word choice, pacing, and filler words. Use a tool like Claude to analyze multiple writing or speech examples and generate a voice profile markdown file. This profile then becomes a reference for any AI tool you use, ensuring your outputs sound like you.
Sources
- 2026-06-06 — I Taught Claude Code to Build You a Personal Brand, Watch This
- 2026-06-05 — Beyond Transcription Building Voice AI That Understands Conversations Herv Bredin, pyannoteAI
- 2026-05-21 — Create Your Own Personal Claude AI System (That Makes Your Work EASY)
- 2026-06-01 — How to talk to statues Joe Reeve, ElevenLabs
- 2026-05-11 — How to Build Mobile Apps with Claude Code Full Course (2026)
- 2026-05-23 — Claude and ChatGPT Got More Literal. Your Old Prompts Are Backfiring
- 2026-04-30 — Claude Design 2 HOUR COURSE (Beginner to Pro)
Lesson 3: Best practices and pitfalls
Speaker diarization (identifying who spoke when) often fails because speech-to-text models don’t generalize well to multi-speaker recordings. Common pitfalls include distant microphones, speaker changes, cross talk, interruptions, and code switching (changing languages mid-conversation). Even when diarization and transcription work separately, assigning each word to the right speaker is surprisingly difficult. Most speech-to-text models aren’t designed for that alignment task.
To evaluate diarization accuracy, practitioners use diarization error rate (a metric that measures mistakes like confusing speakers). In a benchmark, a system showed just 3% diarization error on telephone speech, but performance varies completely by use case. The task remains unsolved because acoustic conditions and overlapping speech create persistent challenges.
Building a voice profile for an AI system requires extracting your personal speaking patterns—word choices, pacing, filler words. The best approach is to transcribe recordings of yourself, such as podcast appearances or videos, then analyze the tone across multiple dimensions. You can paste real writing samples (LinkedIn posts, emails, newsletters) into a tool that extracts voice patterns from those examples. Skipping this extraction yields less accurate outputs.
Best practices: start with clean, varied recordings of natural speech. Always benchmark your diarization pipeline on data similar to your target use case. For voice profiles, prioritize authentic self-recordings over generic answers. Avoid assuming any single tool works well across all environments—test specifically for your recording conditions.
Sources
- 2026-06-05 — Beyond Transcription Building Voice AI That Understands Conversations Herv Bredin, pyannoteAI
- 2026-06-06 — I Taught Claude Code to Build You a Personal Brand, Watch This
- 2026-05-11 — I'm trying out the Latest AI Voice Models, and Google Shows Up ALIVE.
- 2026-05-27 — Claude Code Just Changed Social Media Carousels Forever
- 2026-05-21 — Create Your Own Personal Claude AI System (That Makes Your Work EASY)
- 2026-03-28 — Gemini 3.1 Flash Live Just Changed Voice Agents Forever
- 2026-03-05 — 90+ Claude Code Updates in 4 Releases What's New