Anthropic shipped Claude Fable 5.1 and took the top spot on the Artificial Analysis Intelligence Index at 66, the highest score ever measured, Meta answered with Muse Spark 1.3 and its biggest jump yet in coding and agentic work, and OpenAI dropped GPT-6 Astra, the first model to cross its own Critical cybersecurity line. And that's not all, here is everything that happened in AI this week:

PenEcho is an open-source whiteboard where you write equations, sketch diagrams, and scribble notes anywhere on a huge canvas, and the model responds beside your work instead of in a chat thread. It runs on macOS, Windows, or your own Node.js host, and connects to Kimi, OpenAI-compatible, or Anthropic-compatible models.

  • Spatial context, not a chat box: the canvas is 20,000 by 20,000 and only loads tiles where ink exists. When you pause, it sends that region to the model and the response appears on the whiteboard itself.

  • It reads rough work: the builder's point was that it handles unfinished equations, messy diagrams, and the spatial relationship between them, inferring intent from incomplete marks.

  • Bring your own model: it is open source under AGPL v3.0, commercial use allowed, and installs with npm install -g penecho.

Try it now → https://penecho.ai/

Meta released Muse Spark 1.3 through Muse Code and the Meta Model API, its biggest jump yet in coding and long-horizon agentic work. It scores 61 on the Artificial Analysis Intelligence Index, level with GPT-5.6 Sol and Grok 4.6, while staying the cheapest strong model per task.

  • Fast climb: Muse Spark went from 53 on the Intelligence Index in July to 57 in August to 61 now, and generates around 235 tokens per second. A max variant hit 62 in limited preview.

  • Leads its tier on coding: 75.4 on DeepSWE v1.1, 88.8 on Terminal-Bench 2.1, and 59.4 on SWEAtlas CodeBase QnA.

  • Cheaper agent runs: roughly 20% fewer tool calls and 25% fewer tokens than 1.2, with a 1M context window, $1.25 in and $4.25 out per million on the standard tier and $0.10/$0.20 on the contributor tier.

OpenAI shipped GPT-6 Astra, calling it "the world's most intelligent and aligned model," with Greg Brockman telling reporters it is "not unreasonable to feel that we are now in the AGI era." It is a computer-use model first, and there are no open weights.

  • Saturated benchmarks: 98% on FrontierMath Tier 4, 99.9% on ARC-AGI-3, and 100% on ExploitBench, plus 72.6% on OSWorld V2-Offline.

  • First model at the Critical cyber threshold: it found two previously unknown zero-days in testing, so standard access refuses exploit work and advanced capabilities sit behind a vetted program called Daybreak.

  • The caveat nobody screenshots: independent Artificial Analysis testing puts Astra roughly tied with its predecessor overall, and behind Claude Fable 5.1 on general reasoning. Pricing is $10 per million input tokens and $50 output, with a 1.05M context window.

Try it now → https://chatgpt.com/

Feynman is a command-line research agent that orchestrates multiple specialized agents to gather, analyze, and synthesize research into structured, source-grounded outputs. You give it a question and it runs the literature review, the peer review, and the replication attempt.

  • Four agents on every task: Researcher gathers evidence, Reviewer runs a simulated peer review that grades feedback by severity and flags weak claims, Writer produces paper-style output, and Verifier checks every citation against its actual source.

  • The audit command is the sharp one: feynman audit takes an arXiv ID and compares the paper's stated methodology against its public repository, and replicate produces a replication plan with sandboxed Docker execution.

  • One line, your own keys: it is MIT licensed and free to use with your own API key from Anthropic, OpenAI, or Google, installed with a single curl command.

Claude Fable 5.1 is now the highest-scoring model on the Artificial Analysis Intelligence Index. At max effort it scores 66, the highest score Artificial Analysis has ever measured, ahead of Claude Opus 5 at 63, Claude Fable 5 at 62, GPT-5.6 Sol at 61, and Grok 4.6 at 61.

  • Across the board gains: plus 4 points on the Index over Fable 5, 59.1% on Humanity's Last Exam against the previous best of 55.5%, and the narrowly highest scores yet on Terminal-Bench v2.1 at 91.4% and SciCode at 62.0%.

  • Leads agentic knowledge work: Fable 5.1 at max effort tops GDPval-AA v2 at 1,853 Elo, ahead of Claude Opus 5 at 1,824.

  • The tradeoff: it costs about 20% more per task than Fable 5 despite a 75% cache read price cut, and its five effort settings span an 11x range in output tokens.

Try it free now → https://lmarena.ai/

Colibri is a pure C inference engine, around 2,400 lines with zero runtime dependencies, that runs the open-weight GLM-5.2 model and its 744 billion parameters on consumer hardware with roughly 25 GB of RAM and no graphics card. Fully offline, fully private, Apache 2.0.

  • The trick is streaming: the dense components stay resident in memory at roughly 10 GB after int4 quantization, while the remaining 21,000-plus routed experts live on disk and get streamed in only when the router selects them.

  • Honest speed: it runs from 0.05 tok/s on a slow disk up to 1.06 tok/s on an Apple M5 Max, with speculative decoding via the MTP head bringing throughput to 2.2 to 2.8 tokens per forward.

  • It struck a nerve: 4,611 stars and 403 forks in eleven days, 453 points on Hacker News, and a one-person project inspired by antirez. It has since passed 9,600 stars.

Thanks for making it to the end! I put my heart into every email I send. I hope you are enjoying it. Let me know your thoughts so I can make the next one even better.

See you tomorrow :)

Dr. Alvaro Cintas

✓ Full archive of premium guides with ready-to-use prompts

✓ Structured AI courses (step-by-step, start-to-finish)

✓ Every upcoming premium tutorial