
Good Morning! Here's what I have for you in today's newsletter:
OpenAI brings GPT-Live-1 to the API, voice agents that listen while speaking
Gemini app now on Windows, one shortcut away from anywhere
DeepSeek launches V4.1-Flash, its smallest model with visual understanding
Build a reusable brand kit with ChatGPT Images 2.5 and Canva
4 new AI tools worth trying today
AI TOOLS
OpenAI brought ChatGPT's natural back-and-forth to the API with GPT-Live-1, voice agents that listen while they speak, working with the models and harness a developer already chooses.
The model listens continuously during its own speech, catching an interruption or a change in direction without waiting for a pause to finish first.
Developers keep their existing model and harness choices, since GPT-Live-1 slots into that setup instead of requiring a separate voice-specific stack.
This moves voice interaction from a text-to-speech layer bolted onto a text model into a capability built directly into the conversation loop.

Voice features usually ship as a separate SDK sitting on top of whatever model a team already runs. Here the voice layer runs on the exact model and tools a developer already picked, so a bad response traces back to the same logic that runs the rest of the app, not a second system to debug separately.Β
AI APPS
Google brought the Gemini app to Windows, reachable through Alt + Space to polish drafts, summarize documents, brainstorm, or create images and videos alongside existing apps.
Alt + Space opens Gemini instantly from any app, without switching windows or opening a browser tab first.
Custom image and video creation happens directly inside the shortcut, not as a separate tool that requires leaving the current task.
The app runs alongside whatever's already open, positioning Gemini as a layer over an existing workflow rather than a destination to visit separately.

Windows has never had a native AI assistant baked into the OS the way macOS increasingly does with Apple Intelligence. This closes that gap directly, putting Gemini on equal footing with what Mac users already have built in.Β
AI MODELS
DeepSeek introduced DeepSeek-V4.1-Flash, the smallest model in its new architecture family, built with native visual understanding for faster inference and higher throughput.
The model reads images and text together instead of needing a separate step to describe a picture first, useful for tasks like reading a screenshot or a chart directly.
Being the smallest model in the family, it runs faster and cheaper than the larger versions still to come, making it a practical everyday option rather than one reserved for heavy workloads.
An open-source toolkit called DeepSeek Harness ships alongside it, letting developers build their own AI agents by mixing and matching parts instead of building everything from scratch.

A small, fast model that reads images natively and ships with an open toolkit to build on removes the usual tradeoff between choosing something affordable and choosing something capable.
Developers can now build a working project on a model that's cheap to run and handles more than just text, without needing a bigger budget to get started.Β

HOW TO AI
Generate a logo, then a business card, then a poster, and the pieces often don't end up looking like they belong together. ChatGPT Images 2.5 keeps a subject consistent across edits instead of letting it drift. Canva's Brand Kit locks that result into something a whole team can reuse.

Step 1: Make your mark
Create a finished logo mark for a coffee brand called Northbound Coffee. Bold, rounded, hand-lettered custom typography, stacked two-line wordmark. Weave a small coffee bean or mountain shape into one of the letters as an integrated icon. Dark forest green and cream color palette.This image is your base for everything else. Don't start a fresh generation for the next asset, build on this one directly.
Step 2: Set up Canva
Go to Canva, click Brand in the left sidebar, and create a new Brand Kit. Drag your finished mark into the Logos section.
Extract the exact hex codes from this logo mark so I can add them as brand colorsStep 3: Edit, don't regenerate
This is the actual fix in 2.5 that makes a real kit possible, earlier edits carry through instead of quietly drifting.
Using this exact logo mark, create a realistic product photo of a coffee bag with this logo printed on it, dark green bag, cream label area, sitting on a wooden surface, professional product photography lighting
Step 4: Save as templates
Upload each finished asset into Canva, then save it directly into your Brand Kit's Templates section. Lock the kit first, then duplicate one on-brand master design and resize it for each channel.
Step 5: Add usage rules
Click Add guidelines to add written notes for anyone using these assets.
never stretch the logo horizontally always keep clear space equal to the height of the mark around it use the full-color version on light backgrounds onlyStep 6: Match your voice
Canva's Brand Voice feature trains its writing tool on your actual voice. Upload 3 to 5 real pieces of your own writing.

P.S. You can access all the AI trainings (including the full version of this one), prompts and workflows if you upgrade.

OpenAI introduced a new Data agent in ChatGPT Work, turning a company's data into answers, interactive dashboards, and actions just by asking, connected through the Data Plugin to existing data sources.
Cursor introduced Projects, a persistent coordinator agent that manages work with subagents in a single ongoing thread instead of a new chat for every task.
Anthropic added the ability to pop out any pane in the Claude Code desktop app into its own window, dragging a diff or terminal to a second screen while Claude keeps working in the main one.

π¨ Krea: Krea Agents brings deep creative-workflow expertise and a self-improving memory system to creative work, now in beta.
π€ Hermes Agent: now displays detailed subagent activity live, with the ability to steer or stop them manually from the CLI and desktop app.
π» OpenRouter: Shell lets any model run commands in a hosted Linux container, moving files in and out through the new Files API.
π ElevenLabs: CLI v1.2.0 adds a say command, turning text into speech and playing it as it streams instead of writing to a file.

THATβS IT FOR TODAY
Thanks for making it to the end! I put my heart into every email I send, I hope you are enjoying it. Let me know your thoughts so I can make the next one even better!
See you tomorrow :)
- Dr. Alvaro Cintas
β Full archive of premium guides with ready-to-use prompts
β Structured AI courses (step-by-step, start-to-finish)
β Every upcoming premium tutorial




