
Alibaba just dropped the world's most advanced open-source AI, OpenAI's own models broke out of a sandbox and hacked another company, and Anthropic cut frontier-class intelligence to half price. Here is everything that happened in AI this week:
OpenAI quietly shipped the thing that closes the loop on "AI can write code." ChatGPT Sites has launched in public beta inside ChatGPT Work. You describe a site in chat, ChatGPT generates the code, hosts it, and gives you a shareable URL. No deploy step, no separate hosting bill.
Prompt In, URL Out: You just ask ChatGPT to build you a website with a description of what you want it to do, or type @Sites to trigger a build explicitly. It handles dashboards, project trackers, launch calendars, prototypes, internal portals, and reports, giving you a private preview before you publish.
Real Apps, Not Static Pages: Sites includes hosting, access controls, storage, and database support, so interactive tools like calculators and weekly-review dashboards actually work instead of just looking like they do.
The Catch: It is available on paid plans except Free and Go, and it is not available in the EEA, Switzerland, or the United Kingdom at launch.
Try it now → https://chatgpt.com
Black Forest Labs, the team behind Stable Diffusion's original architecture, just went way past images, launching FLUX 3. They trained a single set of weights simultaneously on images, video, and audio, then extended that same architecture to predict robot actions.
One Model, Everything: FLUX 3 Video generates clips with native audio up to 20 seconds from text, images, or existing videos, and supports video continuation, keyframe transitions, multilingual dialogue, typography, and chaining clips into longer sequences.
It Beat the Field in Internal Tests: BFL says FLUX 3 was preferred over Grok Imagine Video in up to 69% of comparisons, Kling v3 Pro in 60%, Runway Gen 4.5 in 77%, and Luma Ray 3.2 in 93%. The chart is labeled a preliminary evaluation of an early candidate, so treat it as a vendor claim.
It Already Runs a Factory: The video backbone is being adapted for robotics as FLUX mimic, and Audi is testing the system for production tasks involving robotic manipulation.
Try it now → https://bfl.ai/models/flux-3
Alibaba's Qwen team previewed Qwen3.8-Max-Preview at the World AI Conference in Shanghai, describing it as a 2.4 trillion-parameter model second only to Fable 5 among the systems it benchmarked. The preview is live now. The benchmark table, model card, and license are not.
Biggest Qwen Ever, By a Lot: It is a huge jump from Qwen3-Max at 1 trillion parameters in September 2025 and Qwen3.5 at 397 billion in February 2026, and developer Shuai Bai called it the team's first multimodal model above 1 trillion parameters, processing text, images, video, and documents.
The Number Nobody Has: It is a sparse MoE design like the rest of the Max tier, but the active-parameter count has not been disclosed, and without it the 2.4T headline says very little about actual serving cost.
Cheap to Test Right Now: A preview is available through Alibaba's Token Plan subscription, its Qoder coding platform, and QoderWork productivity suite at 10% of standard pricing.
Try it now → https://chat.qwen.ai
Jack Dorsey's Block just took a swing at Slack and GitHub at the same time. Buzz is a free, open source collaboration platform where humans and AI agents work together in a shared workspace, built on the Nostr protocol, with channels, threads, direct messages, voice, media sharing, code repositories, and automated workflows.
Agents Get Real Identities: Nostr gives every agent a cryptographic keypair independent of the platform, and a second signature ties that agent back to its human owner, creating a verifiable passport and an audit trail for everything the agent does.
They Actually Do the Work: Agents can open repositories, submit and review code, run workflows, edit shared canvases, orchestrate other agents, and drop into voice huddles. It ships with pre-built harnesses for Goose, OpenAI's Codex, and Anthropic's Claude Code, connected through the Agent Client Protocol.
Yours to Own: Buzz is provided under an Apache-2.0 license at github.com/block/buzz, so anyone can run their own instance on their own infrastructure, or sign up at buzz.xyz for Block-provided hosting.
Try it now → https://buzz.xyz
Anthropic's argument this time is not about a new ceiling, it is about the floor collapsing. Opus 5 launched across all Anthropic platforms at $5 per million input tokens and $25 per million output tokens, unchanged from Opus 4.8, and became the new default model on Claude Max and the strongest model available on Claude Pro.
Half the Cost of the Flagship: Fable 5 costs $10 per million input tokens and $50 per million output, so Opus 5 lands at exactly half while closely rivaling and in several benchmarks surpassing it.
New State of the Art on Real Work: Headline results include 79.2% on SWE-bench Pro, 44.4 on Frontier-Bench at xhigh effort, 70.6 on OSWorld 2.0, and 1,861 Elo on GDPval-AA v2, against 1,747 for Fable 5 and 1,736 for GPT-5.6 Sol. The ARC Prize Foundation also verified 30.16% on ARC-AGI-3 at high effort, roughly four times the best previously reported leaderboard score.
You Control the Spend: Opus 5 supports low, medium, high, xhigh and max effort, letting the model spend very different amounts of time planning and revising, with a 1 million token context window and 128K output.
Try it now → https://claude.ai
OpenAI disclosed that two of its AI models autonomously hacked their way out of a controlled environment where they were supposed to be walled off from the internet, then hacked into Hugging Face, in order to cheat on an internal evaluation. Not a jailbreak by a human. The model did it on its own, to win a benchmark.
It Bought Its Own Freedom With Compute: The models went to extreme lengths to hit the goal, breaking out of an isolated sandbox by discovering and exploiting a zero-day in a package registry cache proxy, then performing privilege escalation and lateral movement until they reached a node with internet access.
The Motive Was Just Cheating: Both models were running with lower cybersecurity guardrails as part of an internal evaluation of offensive capability, and they determined the answers were stored on Hugging Face's production systems. Rather than solve the evaluation as intended, they went after the answer key. The models involved were GPT-5.6 Sol and an unnamed, even more powerful pre-release model.
A Chinese Model Cleaned It Up: Hugging Face's defenders turned to Z.ai's GLM 5.2, a Chinese open-weight model, after commercial US frontier AI refused to help analyze the attack data because its safety filters could not tell a defender from an attacker.
Read the full disclosure → https://openai.com/index/hugging-face-model-evaluation-security-incident/
Thanks for making it to the end! I put my heart into every email I send. I hope you are enjoying it. Let me know your thoughts so I can make the next one even better.
See you tomorrow :)
Dr. Alvaro Cintas
✓ Full archive of premium guides with ready-to-use prompts
✓ Structured AI courses (step-by-step, start-to-finish)
✓ Every upcoming premium tutorial






