Last Week in AI

#251 - Mythos Back, Sonnet 5, Etched, LongCat

Jul 9, 2026
Open in new tab →

Summary

This episode focuses heavily on frontier AI model releases, safety governance, and the accelerating competition on price and performance. A major topic is Anthropic’s redeployment of Claude Fable 5 after discussions with the US government, along with added cybersecurity classifiers and a new framework for evaluating jailbreak severity. The hosts also cover Claude Sonnet 5, which arrives with discounted pricing, stronger agentic coding capabilities, and default cyber safeguards, even as concerns remain about its relative cybersecurity strength. Beyond Anthropic, the episode discusses Google NotebookLM’s move into vertical video summaries and Nano Banana 2 Lite, both signs that AI products are expanding beyond chat interfaces. The conversation also highlights infrastructure and scaling trends, including Etched’s full-stack AI chip strategy and LongCat 2.0’s use of sparse attention and other large-scale optimization techniques, all against the backdrop of a growing token-price war and more realistic computer-use benchmarks.

Key Takeaways

  • 1Anthropic’s Claude Fable 5 return shows how model deployment is increasingly shaped by government discussions and formal safety controls.
  • 2Claude Sonnet 5 is being positioned as a cheaper, more practical frontier model for real-world work, especially agentic coding.
  • 3AI products are broadening beyond chatbots, with NotebookLM’s vertical video summaries showing new consumer-friendly formats.
  • 4AI infrastructure is becoming a major battleground, with companies like Etched and LongCat pushing hard on hardware and scaling efficiency.
  • 5A price war is emerging across frontier AI, driven by cheaper token pricing from both Chinese models and Western open-source deployments.
  • 6New benchmarks are shifting evaluation toward realistic, long-horizon computer-use tasks instead of narrow coding tests.

Notable Quotes

""Access recent state of AI in the enterprise report found that 96% of organizations say agents need access to company specific content, but only 36% have connected agents to trusted content across many use cases.""

""Access access to models and safeguards for evaluation, information sharing on jailbreaks and misuse and dedicated resources for joint research.""

""We remain in a world where no one knows how to stop jail breaks.""

""It is 100% guaranteed that US adversaries, including specifically the Chinese, will find ways to jail break, mythos and any relevant open AI models of similar capability.""

""in pricing front, it is fairly cheap. So it standard price is 75 cents per million input 2.95 per million outputs... And this is something we haven't discussed, I don't think explicitly, but there's been this ongoing sort of price war going on in China.""

""I think we may be at a point where there is going to be a price war for token prices... Western companies... and you can also just pay for this API.""

""So on this one, remedient tasks takes a skilled human about 1.6 hours of active operation... and the leading agents average more than 300 steps per task as 10x a previous benchmark.""

""the weak solver model is a crappy it's like 23.5 but it's a four billion parameter version and the strong solver is like a 400 billion parameter version""

Episode questions

Why did Anthropic’s Claude Fable 5 return to release after being blocked?

The hosts say Anthropic had productive conversations with the US government and added new cybersecurity classifiers plus model-testing coordination. The release was paired with a broader framework for evaluating jailbreak severity and misuse.

What makes Claude Sonnet 5 notable compared with previous Anthropic models?

It is cheaper than several frontier rivals during the introductory period and improves agentic coding performance and benchmark results. At the same time, the hosts note its cybersecurity ability is still below top-tier models.

Why is NotebookLM’s new video-summary feature important?

It shows that AI products can succeed outside the standard chatbot format. The vertical, TikTok-style summaries may make research easier to consume and broaden the product’s audience.

What is the core engineering challenge LongCat 2.0 addresses?

It tackles scaling problems from massive context windows and large model sizes using sparse attention, hierarchical indexing, streaming-aware indexing, and parallelism. The goal is to make training and inference feasible at very large scale.