Lenny's Podcast: Product | Growth | Career

Anthropic’s first technical PM on token maxing, the jagged edge, and living in the future | Dianne Penn

Jul 26, 2026
Listen Now

Summary

This episode explores how Anthropic evolved from a small, experimental team into a leader in frontier AI, with a particular focus on the product and research loops behind Claude. Dianne Penn explains how the team identified coding as a major opportunity as users shifted from autocomplete to long-form code generation, and how product breakthroughs often came when model capability and user experience improved together. A major theme is Anthropic’s eval-driven development process, where user pain points are translated into measurable tests that guide model improvement, described as "evals are the new PRDs." The conversation also covers the company’s bottoms-up lab culture, where small autonomous teams prototype discontinuous ideas like Claude Code, Skills, and computer use. Finally, the episode discusses Claude’s pushback behavior, the jagged edge of AI writing quality, and why human judgment remains essential in high-stakes decisions.

Key Takeaways

  • 1Anthropic recognized early that coding was becoming a long-form generation problem, not just an autocomplete problem.
  • 2Anthropic’s biggest wins came when model capability and product experience advanced together.
  • 3The company uses a highly experimental, bottoms-up culture to discover new capabilities and product opportunities.
  • 4Anthropic’s product strategy is increasingly driven by evals, which turn user feedback into measurable model goals.
  • 5Claude’s willingness to push back is treated as a feature, not a flaw.
  • 6AI capabilities are still jagged, and writing remains one of the weaker areas Anthropic is actively improving.

Notable Quotes

""E-vals are the new priorities.""

""You have to sweat the tokens as much as you sweat the pixels.""

""You need frontier products in order to have frontier models and for people to feel the magic of frontier models.""

""If you're willing to spend $100,000 a year right now in tokens, you're living the way somebody in 2028 is gonna live.""

""being able to have clawed actually push back in the right points and then add it's like a yes or no end actually helps you come to a better conclusion""

""I think it should be clear where an idea is item is being led by you or by you Lenny or me Diane""

""I think judgment is one is an area where it's a accumulation of so much nuance and so much experience and these systems have an experienced as much as humans have""

""In 2024 we shipped four models for the in the whole year or four series of models and I think we did more than that volume in just Q2 of the share.""

Episode questions

Why did coding become such an important focus for Anthropic?

Dianne says users were increasingly using models for long-form code generation, not just autocomplete. That shift made coding a strategic opportunity for training and differentiating Claude, especially through Opus 3 and later Claude Code.

What does Anthropic mean by 'evals are the new PRDs'?

It means product direction is increasingly defined by measurable model behaviors and failure cases rather than only by written product specs. The team turns user pain points into eval sets so researchers can improve the model against concrete tests.

How do labs and the core product team differ at Anthropic?

Labs focuses on discontinuous bets and early prototypes that may not fit the core roadmap. The core product team is more about turning model capabilities into user-facing experiences and operationalizing feedback loops.

What skills matter most for PMs working on AI products?

First-principles thinking, hands-on experimentation, and the ability to translate messy user feedback into actionable evals matter most. Dianne also says managers must stay close to the technology and actually ship with it.