How Anthropic's product team moves faster than anyone else
⚡ 速览
Anthropic Claude Code 和 Cowork 产品负责人 Cat Wu 揭示了团队如何以前所未有的速度发布功能——时间线从六个月压缩到几天。
她解释了为什么在代码日益廉价化的今天,产品品味(product taste)成为最稀缺的能力,Anthropic 的使命认同如何消除组织内耗,以及基于当前模型能力而非假设性 AGI 来构建产品的务实哲学。
核心主题包括:用 Research Preview 加速发布、让模型反思自身错误、将 Claude 的性格视为核心产品特性,以及把自动化推到 100% 而非止步于 95%。
英文原文
Cat Wu, Head of Product for Claude Code and Cowork at Anthropic, reveals how the team ships features at an unprecedented pace — timelines collapsed from six months to days. She explains why product taste is now the most valuable skill as code becomes commoditized, how Anthropic's mission alignment eliminates organizational friction, and the practical philosophy of building for current model capabilities rather than hypothetical AGI. Key themes include using Research Preview to ship fast, asking models to introspect on their own mistakes, treating Claude's personality as a core product feature, and pushing automations to 100% rather than settling for 95%.
🗺 章节地图
💡 核心亮点
发布速度:从数月到数天
快不靠拼命,靠降低承诺:Research Preview 把发布变成可撤销的实验,而不是许诺。
Anthropic 的功能开发周期从 6 个月压缩到 1 个月,有时甚至只需 1 天。这套体系依赖三个要素:极简流程、Research Preview 品牌标签降低承诺压力,以及工程、市场和文档团队之间紧密的跨职能协作闭环。
EN original
Shipping velocity: from months to days
Anthropic's feature timelines collapsed from 6 months to 1 month, sometimes 1 day. The system relies on low process, Research Preview branding for reduced commitment, and tight cross-functional loops between engineering, marketing, and docs.
产品品味比工程技能更重要
代码越廉价,瓶颈越从做出来移向选什么做——决定写什么才是新的稀缺技能。
当代码的编写成本越来越低,真正稀缺且有价值的技能是「决定写什么」。Anthropic 优先招聘具有出色产品品味的工程师,这样 PM 就不会成为发布流程中的瓶颈。
EN original
Product taste over engineering skill
As code becomes cheaper to write, the scarce and valuable skill is deciding what to write. Anthropic prioritizes hiring engineers with great product taste so PMs aren't bottlenecks in the shipping process.
使命作为终极决策过滤器
把使命置于任何产品线之上,取舍的次序就提前排定——团队才敢快速拍板。
Anthropic 的统一使命——为全人类实现安全的 AGI——让团队能够快速做出跨部门决策,并甘愿牺牲个别产品目标。Cat 说:「如果 Claude Code 失败了,但 Anthropic 成功了,我会非常高兴。」
EN original
Mission as the ultimate decision filter
Anthropic's unifying mission — safe AGI for all humanity — lets teams make fast cross-org decisions and willingly sacrifice individual product goals. Cat: 'If Claude Code failed but Anthropic succeeded, I would be extremely happy.'
对 AGI 保持恰如其分的信念
为假想的超强模型做产品人人都会;从当前模型榨出最大能力,才是稀缺功夫。
PM 最难掌握的技能,是在未来 AGI 潜力与当前模型能力之间找到正确校准。为一个假设中的超级智能做产品很容易;真正的挑战是从今天的模型中榨取最大价值,引导用户走上最佳路径。
EN original
Be the right amount of AGI-pilled
The hardest PM skill is calibrating between future AGI potential and current model capability. It's easy to build for a hypothetical superintelligence; it's hard to extract maximum value from today's models and guide users onto the golden path.
模型会吃掉你的产品外壳
为模型短板打的补丁都有保质期——模型每升级一次,路线图就被划掉一行。
随着模型能力提升,Anthropic 会主动移除那些充当拐杖的产品功能。待办列表的加入是因为早期 Claude 会在重构中途停下来;而更新的模型能自然地完成所有任务,让这个功能变成了装饰而非必需品。
EN original
Models eat your harness for breakfast
As models improve, Anthropic actively removes product features that were crutches. The to-do list was added because Claude would stop mid-refactor; newer models naturally complete all tasks, making the feature decorative rather than essential.
100% 自动化阈值
95% 可靠的自动化不是自动化——价值全在最后 5-10%,不啃下来等于没做。
95% 的自动化不算自动化。Cat 敦促大家啃下最后 5-10% 的硬骨头,让工具真正做到可靠——哪怕一开始构建自动化的速度比手动操作还慢。
EN original
The 100% automation threshold
95% automation is not an automation. Cat urges people to push through the last 5-10% to make tools truly reliable, even though building the automation is often slower than doing the task manually at first.
直接动手做
职位是虚的,约束才是实的——看懂约束的人,做事不需要等谁批准。
Cat 的核心信条:如果你理解了约束条件和第一性原理,就放手去做。职位是虚的,角色是流动的,行动偏好永远胜过等待许可。这种哲学是 Anthropic 赋能个人文化的基础。
EN original
Just do things
Cat's core motto: if you understand the constraints and first principles, just act. Jobs are fake, roles are fluid, and bias towards action beats waiting for permission. This philosophy underpins Anthropic's culture of empowered individuals.
🎤 金句
我们很多产品功能的时间线,从六个月压缩到一个月,有时甚至只需要一天。
EN original
The timelines for a lot of our product features have gone down from six months to one month and sometimes to even one day.
当代码的编写成本大幅降低,真正变得更有价值的是——决定写什么。
EN original
As code becomes much cheaper to write, the thing that becomes more valuable is deciding what to write.
如果 Claude Code 失败了,但 Anthropic 成功了,我会非常高兴。
EN original
If Claude Code failed, but Anthropic succeeded, I would be extremely happy.
为超级 AGI 强模型做产品很容易。真正的难题是,对于当前的模型,如何激发出它的最大能力?
EN original
It's very easy to build the product for the super AGI strong model. The hard thing is figuring out, for the current model, how do you elicit the maximum capability?
很多时候我们给产品加功能,其实是在给模型打补丁,因为模型本身还没有自然地完成这些事。
EN original
A lot of times we add features to the product as a crutch for the model, because it's not naturally doing itself.
如果一个自动化不能百分之百可靠地运行,那它就算不上真正的自动化。
EN original
If an automation doesn't work a hundred percent of the time, it's not really an automation.
去构建你每天都在用的应用,因为只有通过日常使用,你才能真正获得价值。
EN original
Build apps that you're actually using every single day, because only through that usage are you actually getting the value.
直接动手做。职位是假的。如果你理解了约束条件,就能想清楚能做什么,然后尽快去做。
EN original
Just do things. Jobs are fake. If you understand the constraints, you can figure out what you can do and then just try to do it quickly.
🧭 行动建议
- 1 以 Research Preview 的名义发布 08:58
将早期功能标记为 Research Preview,降低承诺压力,在 1-2 周内获取真实用户反馈,而不是等几个月追求完美。
EN original
Brand early features as Research Preview to lower commitment and get real user feedback within 1-2 weeks instead of waiting months for perfection.
- 2 将所有数据源连接到 Cowork 35:58
Slack、日历、Gmail、Drive——Cowork 只有在掌握完整上下文的情况下才能产出高质量结果。结果的质量与接入数据的丰富程度成正比。
EN original
Slack, Calendar, Gmail, Drive — Cowork can only produce great output with full context. The quality of results scales with the richness of connected data.
- 3 让模型自我反思 51:15
当模型做出意料之外的行为时,问它为什么做出那个决定。这往往能揭示误导性的提示词或产品外壳中的漏洞,让你可以针对性地修复。
EN original
When the model does something unexpected, ask it why it made that decision. It often reveals misleading prompts or gaps in the harness that you can fix.
- 4 构建 10 个优秀的 eval 55:00
你不需要几百个。只需 10 个精心设计的 eval 就能帮你量化目标、衡量进展,并发现 AI 产品中缺少什么。
EN original
You don't need hundreds. Just 10 well-crafted evals help quantify goals, measure progress, and identify what's missing in your AI product.
- 5 每次模型升级都清理产品外壳中的拐杖 60:44
每次新模型发布时,通读整个系统提示词,移除模型不再需要的指令。更简洁的外壳才是更好的外壳。
EN original
With every new model, read through the entire system prompt and remove instructions the model no longer needs. Simpler harnesses are better harnesses.
- 6 把自动化推到 100% 69:18
不要止步于 95%。花时间教会 AI 你的偏好,持续迭代直到完全可靠。最后那 5-10% 很难,但正是它区分了玩具和工具。
EN original
Don't stop at 95%. Invest the time to teach AI your preferences and iterate until it's fully reliable. The last 5-10% is hard but makes the difference between a toy and a tool.
- 7 构建日常使用的应用,而非一次性原型 71:58
原型应用教不了你太多。构建你每天都在用的工具,才能真正理解 AI 的价值、局限性和它在哪儿会出问题。
EN original
Prototype apps teach you little. Build tools you actually use every day to understand AI's real value, limitations, and where it breaks.
📖 全文
01 · 0:00 · Introduction to Cat Wu
0:00I think it is very hard to be the right amount of AGI build. It's very easy to build the product for the super AGI strong model. The hard thing is figuring out for the current model, how do you elicit the maximum capability? I've never seen anything like the pace you folks at Anthropic are shipping at. We want to remove every single barrier to shipping things. The timelines for a lot of our product features have gone down from six months to one month and sometimes to even one day. You're interviewing hundreds of PMs and you just keep feeling like they're approaching it very incorrectly. The PM role is changing a lot. It's changing really quickly. The thing that is extremely important for building AI native products is iterating so quickly, figuring out a way for you to actually launch features every single week. What do you think are the emerging skills PMs need to develop? It comes back to product taste. As code becomes much cheaper to write, the thing that becomes more valuable is deciding what to write. Today, my guest is Kat Wu, head of product for Cloud Code and co-work at Anthropic.
1:01Kat is at the center of everything that is changing in AI and product and building, and she and her team are building the product that is most changing the way that we all build our products. She is so full of insights and wisdom and lessons. This is an episode you cannot miss. Before we get into it, don't forget to check out Lenny's Product Pass dot com for an insane set of deals available exclusively to Lenny's newsletter and our subscribers. With that, I bring you Kat Wu.
02 · 1:29 · Working with Boris Cherny
1:31Kat, welcome to the podcast. Thanks for having me. I have so many questions. I'm so excited to have you on this podcast. I want to start with giving people an understanding of your role alongside Boris. Everybody knows Boris. This is his episode is the number one most popular episode on this podcast. No pressure. He created Cloud Code. He leads the team, ships, a bazillion PRs a day from his phone, just like, I don't even know what the number is anymore. I think people don't give you enough credit for the success that Cloud Code has had and co-work and all the things you all are building. Help us understand your role on the team, how you work with Boris, how you split responsibilities, just like what does the PM role look like on the Cloud Code team? I feel very lucky to work with Boris. He's been an amazing thought partner. He's our tech lead. He's very much the product visionary, and he is great at setting, like this is what the product needs to be in like three months, six months from now. This is like what the AGI-pilled version of the product is.
2:33And a lot of my role is figuring out, OK, what is the path from where we are today to like that vision three to six months from now? And I spend more of my time on the cross-functional. So making sure that our marketing team, sales team, finance capacity, et cetera, are like bought in on the plan and that we're all rowing the same direction. And that once the feature is ready, that there aren't any blockers to shipping it. I think in many ways it works well because we kind of like mind meld, but it is actually like remarkably blurry of a line. Like, I think we're like 80 percent mind meld. And then there's like this 20 percent of things that like maybe I care a lot more about than Boris, so like I'll drive those and like 20 percent where he cares a lot more than me and he just like drives those. This episode is brought to you by our season's presenting sponsor, WorkOS. What do OpenAI Anthropic, Cursor, Vercel, Replit, Sierra, Clay and hundreds of other winning companies all have in common, they are all powered by WorkOS. If you're building a product for the enterprise, you've felt the pain
3:37of integrating single sign on, skim, RBAC, audit logs and other features required by large companies. WorkOS turns those deal blockers into drop in APIs with a modern developer platform built specifically for B2B SaaS. Literally every startup that I'm an investor in that starts to expand up market ends up working with WorkOS and that's because they are the best. Whether you are a seed stage startup trying to land your first enterprise customer or a unicorn expanding globally. WorkOS is the fastest path to becoming enterprise ready and unblocking growth. It's essentially Stripe for enterprise features. Visit WorkOS.com to get started or just hit up their Slack where they have actual engineers waiting to answer your questions. WorkOS allows you to build faster with delightful APIs, comprehensive docs and a smooth developer experience. Go to WorkOS.com to make your app enterprise ready today.
03 · 4:29 · What Anthropic looks for when hiring PMs
4:29Something that you shared actually before we started recording is the fact that you're interviewing hundreds of PMs all the time. Like if I had a nickel every time someone asked me for an intro to someone at Anthropic to go work at Anthropic as a PM, I'd have 30 billion in ARR. It's just like the number one place people want to go work at. So I can only imagine how many PMs you're interviewing. You told me that you're just seeing people doing it wrong the way they're researching what they think it takes to be a successful AI PM. Talk about what you're seeing and what people need to understand about what it is, what it takes to be successful these days. I think before AI, technology shifts were a lot slower. So you could plan on the six to 12 month time horizons. And because you were shipping features at a bit of a slower rate, there was a lot more emphasis on coordinating with all the other partner teams to make sure that they're shipping features that unblock your features because code at that time was very expensive to make. I think now with AI and with how much that has accelerated engineering and with how quickly the model capabilities are improving, the timelines for a lot of our product features have gone down from six months to one month and sometimes to one week or even one day. And with that, we actually need to make sure that products ship quite quickly. And what that means is as a PM, there should be less emphasis on
5:55making sure that you're aligning your multi-quarter roadmaps with your partner teams and more emphasis on, okay, how can we figure out the fastest way to get something out the door? How can we figure out how to make a concept corner of our product suite where we can just... An engineer has an idea or a PM has an idea. And by the end of the week, we are able to get into our users' hands.
04 · 6:18 · How to help your teams move fast
6:18I think the PMs who do the best on AI native products are the ones who can figure out how can I shorten the time from having this idea to actually getting the product in the hands of users and help define what are the most important tasks that need to work out of the box for my product. So what I love about this is what you're saying is just like people haven't grasped how fast they need to move and how much of the job now is just moving, is helping the team move fast. What helps do that? What do you do? What does your PM team do to help them move this fast? Other than have access to the most advanced models? I think the first thing is to set clear goals because LLMs are so general that actually creates a lot of ambiguity in who we're building for, what problems we're trying to solve, what the top use cases are. And so I think a great PM is able to say, OK, our key user is professional developers. The main problem that we want to solve for this feature is maybe there's like too many permission prompts and people are feeling fatigue and like the use case is we want professional developers at enterprises to safely get to zero permission prompts. And that actually sets a pretty clear goal because it rules out a lot of potential
7:36approaches for reducing permission prompts so that people can get a lot more done with one prompt. And then I think the second thing that's very important is figuring out some repeatable process for getting these features shipped. So for Cloud Code, what we do is we actually ship almost all of our features in Research Preview. We clearly brand this when we ship something so that users know that this is an early product, this is just an idea. This is just something that we're trying to get feedback on and iterating on and that this might not be supported forever. And what this does is it reduces our commitment for shipping something. We can just get something out in a week or two. And the third thing that a PM should do is help create the framework for the team so that they know when to pull in cross-functional partners. And what those cross-functional partners expectations are. So, for example, we have a really tight process between engineering, marketing and docs. So when engineers have a feature that they feel is ready and that we've dog fooded
8:37internally, they post it in our evergreen launch room. And then Sarah, who leads our docs, and Alex, who leads PMM and Tarek and Lydia on DevRel, just like jump in and can turn around the marketing announcement for it the very next day. And because we have this really tight process, it lowers the friction for any engineer to ship something. And PM is the role that should be setting this up.
05 · 8:58 · How PRDs and roadmaps have evolved at Anthropic
8:59How do PRDs fit into this? The fact that you said that goals are a really important part, just like being aligned on what does success look like, who is this for, who is this not for? Are you writing PRDs? Is it just like a couple of bullet points? How does how's that evolved in the world of a PM? So there's two things that we do. One is we have very rigorous metrics and we do metrics readouts with the entire team every week. The goal of this is to make sure that everyone deeply understands all the facets of our business, what our key goals are, how they're trending and what drives them. The second thing that we do is we have this list of team principles and this includes who our key users are, why those are our key users. And the reason that we articulate all of this is so that everybody on the team feels like they understand how our business works. They understand what's important to us and what we're willing to trade off. And it lets people make decisions by themselves without feeling like they're blocked on PM or any other stakeholder. I love how so much of this is like, OK, we still need PMs in the future.
9:58And there's so much talk of like, why do we need PMs? We're just going to ship and build. We need engineers. Oh, we actually do PRD sometimes. So I think for features that are like particularly ambiguous, it does help to write out just a one pager on what the goals are, what the delightful use cases are, what the failure modes currently are that we need to fix. And there are occasionally some projects, especially things that require heavy infrastructure. That do take many months. And for those situations, we do write PRD still.
06 · 10:28 · The Mythos model and Anthropic's shipping velocity
10:28I want to drill a little bit further into just how you're able to move so fast. I've never seen anything like the pace folks at Anthropic are shipping at. Like someone made this calendar of launches across Anthropic. And it was literally every day there was like a major feature or product. So one question people had online is you guys just launched this, not launched, but built this incredible model, Mythos, that is still in preview because it's so powerful. People are a little afraid of what it can do. Have you guys been using this? Is this part of the reason you've been able to move so fast? We've been moving pretty fast for several quarters now. So I think it's not fully Mythos. Mythos is an incredibly powerful model. We do use the models internally. And I think this has increased our rate of shipping a little bit, but I don't think it explains the bulk of the increase. I think a lot of it is the process and the expectation on the team. So we're very low on process.
11:27We want to remove every single barrier to shipping things. We want to make sure every single person on the team feels empowered to take their idea from just an idea to like out in the world in less than a week, sometimes even in a day. Cool. Oh, man. What a what an advantage to have the best model and also be building product. That's so cool. We are very lucky to be able to work with the Frontier models. Oh, my God. What an awesome advantage, just like build a thing and then use it and accelerate faster. It's so interesting.
07 · 11:54 · What happened with the Claude Code source code leak
11:55There's a couple of like these other side things I want to just kind of go on these like side quests on this conversation. There's so much happening with Anthropic and I just I'm so curious to get your insight. One is a week ago or so, the whole source code of cloud code leaked. Somebody got it out there. I think it was a mistake someone made. Is there anything you comment there? Just like what happened? What went wrong? What should people know? So we immediately looked into this when we saw it. We realized that this was the result of human error. There is a human working with cloud to write PR. This was just an update to how we release our packages. And it actually went through two layers of human review. And so this was a result of human error. And we've hardened our processes to make sure that it doesn't happen in the future. This person's still at Anthropic. Are they doing all right? Yes, yes. It's it's a process failure. And the most important thing is to just like learn from it and to add more safeguards so that doesn't happen again. And so that's that's what we've been focused on and most of those
08 · 12:53 · Integrating with OpenClaw
12:53things have shipped. OK, another question I had is open claw. So recently there's been this move to keep people from using cloud subscription with their open clause, people got really upset. They're confused why this is happening. It feels like you're there's like, you know, harm cost to the open source community. What what are people what do people need to understand about kind of what went into this decision? So we've been seeing a lot of demand for cloud and we've been working very hard to both scale our infrastructure and also to make our harness more token efficient so that you can get more usage out of it. It wasn't designed for third party products, which have different usage patterns than our first party ones. We spent a bunch of time. Trying to figure out what is the most seamless transition that we can offer. And so I was very happy to be able to say that everyone gets some credits alongside their subscription. But yeah, we did have to make the hard decision that we needed to prioritize our first products and our API. And so this is the this is the decision that resulted from that.
14:00Yeah, like to me, it makes so much sense. Like you guys are subsidizing this usage at like 200 bucks a month. There's like it's like basically unlimited use of this. And like I think people don't understand this is they're trying to make money. We're trying to be profitable. We can't just like give away compute when it's so in demand. So I get it.
09 · 14:19 · How the PM team is structured at Anthropic
14:19Coming back to the PM team, what is just like the PM team? Like at Anthropic, how many PMs are there? How are they kind of organized? Yeah, so we have a few PM teams. I think we're maybe around 30 or 40 PMs right now. So we have the research PM team who Diane leads. And this team is responsible for understanding all of the feedback from our customers for our models and then feeding that to the best research team to act on it. And they also shepherd the model launch. There is the cloud developer platform team that maintains the APIs that CloudCode is built on top of, and they also release things like managed agents, which is a way for you to build your agents and we can host it on your behalf. And then there's CloudCode that works on both CloudCode and the Cowork core products. There's Enterprise that helps make CloudCode
15:11and Cowork easier to adopt for all of our Enterprise customers. And so this is everything from like cost controls, RBAC, security controls, and just making sure that these enterprises feel very confident and comfortable using our tools. And then we also have our growth team that is responsible for growing across our entire product suite. So we work very closely with them on CloudCode and Cowork growth. And I know they also work with our other teams on CDP growth. So growth of people who use the Cloud API.
10 · 15:42 · How engineer and PM roles are merging
15:43So speaking of growth, so Amol was just on the podcast. He had this really interesting insight that most people haven't been sharing. There's always the sense that we need fewer PMs in the future. Why do we need PMs? Engineers can just ship. His take is that because engineers are moving so fast, PMs and designers are squeezed, there's less time to stay on top of everything that is happening, there's a feature shipping every day. So his take is he needs more PMs because it's hard to keep up. What's your take there? Do you feel like there will be an increase in hiring of PMs? What do you think is going on with the PM profession long term? I think all of the roles are emerging. PMs are doing some engineering work, engineers are doing PM work, designers are doing PMing and also landing code. You can either hire a lot more engineers who have great product taste or you can keep your engineering hiring the same and hire a lot more PMs to help guide some of their work. On our team, we're pretty focused on hiring engineers with great product taste.
16:43This way we can reduce the amount of overhead for shipping any product. Like there are many engineers on our team who are fully able to end-to-end go from see user feedback on Twitter through to like ship a product at the end of the week with almost no product involvement. And this, I think, is actually like the most efficient way to ship something. So I think like engineer and PM are kind of overlapping and you will get a lot of benefit from having more of either. I think product taste is still a very rare skill to have and we'll pretty much hire anyone who we feel has demonstrated this strongly. And your background was in engineering, right? Yeah, I was an engineer for many years. I was then a VC very briefly before joining Anthropic. And actually, almost all the PMs on our team have either been engineers or ship code here on Cloud Code, and so that's one of the things that I think helps build trust with the team and also just enables us to move a lot faster. And then actually our designers also
17:52have been front end engineers before.
11 · 17:54 · Why product taste is the most valuable skill
17:54Wow, because that's that's the big question. Like, there's definitely this merging that's happening. The Venn diagrams are combining. I think the big question for a lot of people is if you're coming from engineering or product or design, which of those core skills is going to be most valuable? I could see at Anthropic and on Cloud Code, engineering is very valuable. I'm curious if other companies, if you have a design background, becoming a PM is more valuable or just a PM PM. I still think it comes back to product taste. Like as code becomes much cheaper to write, the thing that becomes more valuable is deciding what to write. Like, what is the right UX for this feature? What is the most delightful way that a user can experience it? What like we get tens of thousands of GitHub issues asking for every single thing under the sun. And it takes a lot of. Care and taste to figure out, OK, which of these is worth building and what is the right way to build it? And I think that that skill set can come from any background, but I think that's the most important thing. I think the reason why an engineering background is particularly useful, at least for the next few months, is if you have an engineering background, you have a better sense for how hard something should be. And that's often a factor in what you choose to build.
19:08So like if something is very easy to build, then maybe instead of debating it, you just spend an hour doing it. But if something is harder to build and you know that upfront, then you know that, OK, this will just cost a lot more. For our team to get this out the door. So it helps a bit with the prioritization. You said in the next for the next few months. Is that just like because the models will get so good potentially in the next few months, you may not even need to know that as much. I think the value skill sets does change quite frequently. And so it's really hard to predict more than a few months out. So it's less a commentary on what shifts I think will happen and more of a commentary that I think large steps will happen. So you're not saying that's when Muthos comes out and will change everything. And we don't need to know anything about engineering. No, I'm just saying that every every few months, it seems like there's a yeah, there's a large increase in coding capability, which then changes what other roles are valuable.
12 · 20:10 · Where human brains will continue to be useful
20:10I think the most important thing is to be able to. To to have this like first principles thinking where you can figure out how the tech landscape is changing, what the team really needs from you and to like jump in and fix that hole, because I think the work is becoming more amorphous, which means that a great PM is able to understand what all the gaps are to figure out what the highest priority ones are and then to just like figure out, OK, how do I learn that skill set or what is like the skill set that I have that I can like apply to this challenge? So I think the current environment values people who are who are able to wear a lot of hats, are able to swap them and are like very low ego about what work they do to help the team move faster. I love this answer. There's this question I've been asking people in your in your shoes, folks that are kind of at the bleeding edge of what is capable of and building with the latest tools, which is just like where will human brains continue to be useful and necessary for a while until we get to super intelligence? What I'm hearing here is essentially picking the things to work on, knowing where the market's going and figuring out where what to prioritize essentially. And then it's knowing if the thing you've built is good and right and getting it out there in some early version, at least. Does that sound right?
21:37Is there anything else of just like where human brains will continue to be useful for at least the next few months? I think humans still provide a level of common sense that the models don't. And there's like a thousand moving pieces to any product launch. Some of them are very small, but there's always a lot that could potentially go wrong. I think the model doesn't always have a great sense of who all the stakeholders are, how they relate to each other, what their preferences are, what are the right venues to communicate with them, to keep them on board. I think a lot of this like more tacit, common sense, like EQ kind of knowledge is still very valuable. Of course, we want the models to get better at this, and I think they will be. But right now, I think there's still gaps.
13 · 22:23 · How to stay sane in constant chaos
22:24How do you just kind of deal as a human going through so much constant change, just like just being on the inside of the tornado, maybe it's calm there, but just like how do you how do you stay on top of what's going on, how you stay sane through all this craziness that we're moving through? I think our team is full of people who lean into the chaos. So we try to face every challenge with a smile because there's always so much going on, there's always so many risks and tricky situations. That, you know, if you get too stressed about anything, you'll burn out. And so we really look for people who can kind of like look at a challenge, be like, oh, that's going to be hard, but I'm excited to tackle it and I'm going to do the best that I possibly can. And I know I won't be perfect, but I'll be able to sleep at night knowing that I did my best. That's an interesting answer to just like what skills will be important in this future, because it's I forget who said this, maybe Ben Mann, that this is the most normal this is the world will ever be. Yeah, it definitely gets harder.
23:24Like, I feel like there are a lot of weeks where maybe Sunday night there's some like P0 and then by Monday there's like a P00 and by Monday afternoon there's a P000. And you're like, wow, I can't believe I was so worried about that P0 from Sunday. But I think you just have to acknowledge that there's only so much that you can do that you need to sleep well so that you can make good decisions next day and just like brutally prioritize where you spend your time, what's the most important thing to get right and be OK letting things go. Like there's there's products that we ship that aren't as polished as I wish they were. But. You know, our top goal is to help empower professional developers. And if a product isn't successful, as long as it's not blocking the core use case, it's OK because we'll hear the feedback and we'll fix it in the next release.
14 · 24:16 · What gets sacrificed when you ship so fast
24:17Launching a feature that is buggy is the kind of thing that would have kept me up at night. But it is something that I am now able to, like, live with knowing that, OK, we're going to get that quick feedback and we're going to fix it in the next release. What I'm imagining is there's that gif. I think it's maybe from Pirates of the Caribbean where it's this guy walking down a pair of stairs on a ship and the whole ship is just being demolished around him and he's so chill just strolling down the staircase as everything's falling apart. And that's interesting because everyone I've met from Anthropic is just so chill and just so like optimistic. Yeah. I think that's a really interesting insight is just like having this calmness and optimism versus just like, oh, my God, everything's crazy and going nuts. Yeah, I think if you don't have it, you'll get pretty burnt out. I think we also tend to hire people who have been in the industry for a while and have experienced lots of ups and downs and have a good sense for what gives them energy and how to maintain their energy over time, and I think that's helped us a lot. So interesting.
25:21Something that I wanted to ask about is, so there's these roles blurring, engineers are becoming PMs, everyone's dogs or cats, everyone's everyone. What do we lose in that world? Do we lose like career ladders and clear career paths? Do we lose design consistency, code quality? You know, there's probably some downsides. What are some things you find are just like, OK, that's something we're sacrificing for the greater good? We're sacrificing product consistency. Historically, when code was expensive to write, you would carefully plan out everything your products were going to use, the product suite, how every product relates to each other, what the use case for every single one is, how they integrate. And you would pretty much have one product for each use case. And now with AI moving so quickly and with so many ideas that we need to test out, we do sometimes have features that overlap with each other. A lot of the times it's because there's two form factors that we love internally, and we want to, we want the external audience to tell us which one is better. What that means for someone
26:24who's a new user, though, is a new user might not know, OK, what is the best path to accomplish X? There is more education we need to do to help people understand what the core features are and what the best practices are for using them. I think this is the this is the cost of launching a lot of features. I think users also feel like it's hard to keep up with the latest. Usually in traditional PM, you ship a feature every month or quarter, and so it's really easy for a user to to understand, OK, I just need to check in on this once a month and I'll learn some new things. And if I ignore it for six months, it's fine. I don't feel like I'm missing out. I think with these agentic tools, not just called code and co-work, but like across the whole ecosystem, people feel this need to check Twitter every single day to see what the absolute latest thing is. And I think there's more we can do to help people feel less like they're on this ever increasingly fast treadmill and that they feel like I would love people to feel like they can just open these tools, the tools will educate them or like teach them what they want to know and that they can just feel more bought along.
15 · 27:47 · The /powerup command
27:47Yeah, I saw you launch this really interesting feature the other day. I think it's slash. Power up where it basically walks you through all the cool ways and basically all the best practices to use cloud code is that kind of all in these lines? Yeah, exactly. So in the past, we didn't actually want to do something like power up because we felt like the product should be intuitive enough that you can that you don't actually need to go through any tutorial. And over time, we've just realized that there's just so many features and there's so much demand for a built in onboarding experience that we we diverged a bit from what was possible saying no, no onboarding flow. And I did this because there's just so many users who wanted to know there's 100 features. What are the 10 that I absolutely need to use? And so we put that together.
16 · 28:32 · Why Anthropic has been so successful
28:32Yeah, it's such a bizarre world. So Anthropic has been really successful with B2B enterprises where traditionally you don't launch a bunch of stuff, you just kind of have a quarterly release, maybe, and it's like the opposite of every day we got some new. So just maybe following that thread, the run Anthropic has been on is just otherworldly. Anthropic is way behind. And when it started, it was a mole share. This just like one of the least funded companies didn't have distribution. Was it the first to go open? I was way ahead. It was just like, no way. Anthropic has any chance to compete significantly long term. Now it's just killing it, just beating the biggest companies teams so much. Just like the growth is just like $11 billion in ARR in one month. Perps and growth. By the time this comes out, it'll probably be even higher. I think on the inside, what what are some ingredients that have allowed Anthropic to be this successful and kind of come from behind and do this well? The two most important things are one, this unifying mission.
29:33It's hard to state how important this is. We hire people who care most about bringing safe AGI to all of humanity. And this is actually something that we reference frequently in our decisions that our entire product org should focus on shipping. And because we put this mission above any individual product line, we're able to very fast decisions that cut across the entire org and like execute on them in a unified way. So I think this is like something that I've never seen at a company of our scale. And so just to make sure that's clear. So essentially having the number one mission is safety, alignment, making sure AI is good for the world. And you're saying just having that as a clear mission makes decisions a lot easier to make. If there's two competing priorities, we'll talk about which one is more important for Anthropic's mission. And it makes it a lot easier to decide which of the two we prioritize. And then everyone will stand behind the one that we decide. And so sometimes that means that like, hey, we want to ship something on cloud code, but this other thing is more important. And so we deprioritize shipping this and we just wait until later. What's really interesting about that is that explains, I think, versus another company, maybe rhymes with Bopen AI, did a lot of different things. And I think that's a really interesting thing. And I think that's a really interesting What I'm hearing here essentially is like, okay, we're not going to launch a social network. We're not going to launch a feed of interesting information because it's not aligned to this
31:04mission. And that has kept Anthropic focused, which just seems to be a core ingredient to the success. Well, when I think about mission, I think about putting Anthropic's goals ahead of any individual org or any individual product. And so for me, I think the second thing that we're going to talk about is mission. To me, it's slightly different. Mission means that teams are willing to make sacrifices that hurt their own goals and their own KRs in service of Anthropic's goals and Anthropic's KRs. And people are very happy to make those trade-offs. So like, an extreme example is if cloud code failed, but Anthropic succeeded, I would be extremely happy. And like, we're like, the whole team is very willing to make decisions, that follow that chain of thought. I don't know if you can talk about this in depth, but do you feel like the open cloud decision is a part of this? Just like, okay, this is not furthering the mission of Anthropic. We need to stop this because it's not working in the way we want it to work. I think one of the most important things for Anthropic is to grow the number of users that we're able to reach. One of the ways that we're able to do this is with the cloud subscriptions with our first-party products. And so we just very much want to come at the expense of third-party products sometimes. So we've been talking about cloud
17 · 32:28 · When to use Claude Code vs. Desktop vs. Cowork
32:30co-work, all these things, something that I want to make sure people get. And I'm curious just how you use these tools. So there's cloud code, there's cloud desktop/web, there's co-work. What's the best way to understand when to use which? When do you use each of these three? So I tend to use cloud code in the terminal when I'm just kicking off like a one-off coding task, and I want all of the latest features. The CLI is our initial product surface, and it's also the one where our features often land first. And so it's the most powerful of all the tools. So that's what I tend to use when I'm just like trying to kick off one or like maybe like a handful of tasks at a time. I think desktop really shines when you're doing something that requires front-end work. And so one thing that I love to do is to use our preview feature. So if I'm building a web app, I'll often use cloud code in desktop. I'll have the preview pane open on the right-hand side so that I can actually see the web app that I'm making in real time as I'm chatting with cloud. It's also really great for people who want something a bit more graphical. A terminal can feel very unfamiliar to someone who's non-technical.
33:40You get a bunch of these like scary pop-ups on your machine, and you can't click around the way that you're used to in pretty much every other product that you use. So there's a lot of people who just like don't feel comfortable in the terminal. And if that's you, I would highly recommend checking out cloud code on desktop. Desktop is also great for getting an at-a-glance view of everything that's happening. So you can see your CLI terminal sessions in desktop. You can see your other desktop sessions. You can see your sessions that you kicked off on web and mobile. So it's a one-stop control plane where you can see all of your tasks. I think the benefit of web and mobile is that it's really great for kicking things off on the go. So CLI and desktop both require you to be on your local laptop. And this is constraining because sometimes you're out and about, you're like touching grass, you're going on a walk, and you don't have your laptop open. I can't count the number of people who I've seen like holding their laptop open, like tethered to their phone while they're outside. And this just means that we're missing a product that solves that need. And so for me, what mobile lets you do is kick off these tasks on the go so that you don't need to bring your laptop everywhere and make sure that your laptop's open wherever you are.
34:56I love that. I've seen people on plane, like it's just like such a meme now, just I need to finish, let this agent finish. I can't shut this down. Exactly. And then I think for cowork, the role that this fills is there's a lot of work that everyone does where the output isn't code. So whether that's like getting to Slack zero or inbox zero, or whether that's creating a slide deck for some customer meeting that's coming up, or whether that's writing a quick doc on what the goals of a feature are, or what the launch plan for a feature is, all these tasks produce outputs that are non-code and cowork is best positioned for that. So the way that I split the products in my mind is if I'm building something where the output is code, I'll use cloud code or desktop or cloud code on mobile. And if the output is anything that's not code, I'll use cowork for it. People are just like sleeping on the success that cowork is having. It's just like growing incredibly fast. And I think people still don't understand maybe what it's for. And so what if you give us a couple of
18 · 35:58 · Tips for getting started with Cowork
36:00use cases just in your work as a PM, what are some like really interesting, maybe unexpected ways to use cowork to save you time, get more work done? If you're getting started on cowork, the first thing that you really need to do is connect all the data sources that are relevant to your role. You can only do a great job if it has access to all the context that it needs to be able to curate the output for you. So what that means for me is I connect it to my Google calendar, I connect it to my Slack, to my Gmail, to my Google Drive, so that it just knows, it has the flexibility to find relevant context, to ask questions, to pull in threads, and this like substantially improves the quality of the result. The kinds of things I use it for are I use it for when I'm working on a project, I use it when I'm working on a project, I use it when I'm working on a project, I use it when I'm working on a project, I use it when I'm working on a project, I use it when I'm working on a project, I use it when I'm working on a project, I use it when I'm working on a project, I use it when I'm working on a project, I use it when I'm working on a project, I use it when I'm working on a project, I use it when I'm working on a project, I use it when I'm working on a project, I use it when I'm working on a project, I use it when I'm working on a project, I use it when I'm working on a project, I use it when I'm working on a project, I use it when I'm working on a project, I use it when I'm working on a project, I use it when I'm working on a project, I use it when I'm working on a project, I use it when I'm working on a project, I use it when I'm working on a project, I use it when I'm working on a project, I use it when I'm working on a project, I use it when I'm working on a project, I use it when I'm working on a project, I use it when I'm working on a project, I use it when I'm working on a project, I use it when I'm working on a project, I use it when I'm working on a project, I use it when I'm working on a project, I use it when I'm working on a project, I use it when I'm working on a project, I use it when I'm working on a project, I use it when I'm working on a project, I use it when I'm working on a project, I use it when I'm working on a project, I use it when I'm working on a project, I use it when I'm working on a project, I use it when I'm working on a project, I use it when I'm working on a project, I use it when I'm working on a project, I use it when I'm working on a project, I use it when I'm working on a project, I use it when I'm working on a project, I use it when I'm working on a project, I use it when I'm working on a project, I use it when I'm working on a project, I use it when I'm working on a project, I use it when I'm working on a project, I use it when I'm working on a project, I use it when I'm working on a project, I use it when I'm working on a project, I use it when I'm working on a project, I use it when I'm working on a project, I use it when I'm working on a project, I use it when I'm working on a project, I use it when I'm working on a project, I use it when I'm working on a project, I use it when I'm working on a project, I use it when I'm working on a project, I use it when I'm working on a project, I use it when I'm working on a project, I use it when I'm working on a project, I use it when I'm working on a project, I use it when I'm working on a project, I use it when I'm working on a project, I use it when I'm working on a project, I use it when I'm working on a project, I use it when I'm working on a project, I use it when I'm working on a project, I use it when I'm working on a project, I use it when I'm working on a project, I use it when I'm working on a project, I use it when I'm working on a project, I use it when I'm working on a project, I use it when I'm working on a project, I use it when I'm working on a project, I use it when I'm working on a project, I use it when I'm working on a project, I use it when I'm working on a project, I use it when I'm
19 · 38:44 · Demo: Using Cowork to build slide decks overnight
38:44working on a project, I use it when I'm working on a project, I use it when I'm working on a project, I use it when I'm working on a project, I use it when I'm working on a project, I use it when I'm - So slow. - And I love people will see this deck whenever you present this, this will be out in the world. Like obviously it's not the one-shotted version, but you've iterated on it. So just to help people try this for themselves. So step one is connect their, what did you say, Slack? What else do you suggest they connect? - Slack, Google Calendar, Gmail, G Drive. You should connect your communications tools and where you store your source of truth data for what your team cares about, what you care about and what you're working on. - Okay, and then what was the prompt roughly that you put in there to generate this deck? - So I just wrote, make me a slide deck for the Code with Claude conference. This is what our PMM suggested it should cover. This is the current draft that I made that I don't like. This is one that I made manually that I don't like, but I linked it.
39:38Can you start by creating a proposed outline with details? Also make sure it doesn't overlap too much with a keynote talk, which is more important. And then Claude read a bunch of the links that I sent to it and created a proposed outline. So then I read through its proposal and all the different ideas that it had generated for what we could cover. And I just made a decision on what I wanted to actually be in the final deck. And I think this is like an example of what the role of the PM still is today. It's like, Claude is a great brainstorming partner. It's able to synthesize a massive amount of information really quickly and present all of the possibilities to you. But the role of the PM is still to make the end decision of, OK, what should belong in the final product? So for this, what I ended up deciding was that I wanted the talk to cover the progression from making local tasks successful to making every PR green to helping engineers land more PRs. And for each of these, which demo would be the most compelling? And then after this decision about the outline, co-work just went off for a few hours.
40:48And built the whole side deck. This is so awesome. What an awesome part of the job to not have to do anymore. And it feels like you're talking to essentially a deck designer that also has actual knowledge about what you've worked on and can make it actually the content which you want it to be, not just make it look really nice. How did you do the design system piece? How does that work? How does it know the design system of Anthropic? So what I did for this is we actually already have a standardized deck that we use across all of our external engagements. And so I just gave Cloud access to that. And so it's able to see what colors we use, the fonts we use, the different kinds of-- what's it called? Slide formats that are possible. And so it has 20 of these example slides. So give an example. Got it. So you upload, here's our template, work from this. Yeah. You can also connect to your Figma MCP
41:43if you have your side format saved there, and it can pull that. You can pull that in.
20 · 41:48 · Cat's PM tech stack and internal tools
41:48Along those lines, something I'm always curious about is what's in your stack of tools as a PM in Anthropic? Obviously, Cloud Code and Cowork and all the Anthropic tools. What else are you using? What other-- Slack, you mentioned. Is there anything else? So my stack is pretty heavily Cloud Code, Cowork, and Slack. Anthropic largely runs on Slack. I feel like it's the core OS of our company. And day to day, a lot of-- I would say maybe 30% of my time is pushing the boundaries of what Cowork and Cloud Code can do so that I have a very strong sense of what we're not good at. And I spend a lot of time talking with the model to understand why it makes mistakes that it does. We actually have a lot of internal tools that we make. I think one of the things that Cloud Code has really unlocked for our entire company is it really lowers the barrier to making any custom app that you want.
42:53And so we've seen this surge in personalized work software that people are building for custom use cases instead of using tools that don't perfectly fit the use case. I got to hear more. What are some examples? What are things you've built, other people have built, that are really popular and useful? One of the sales folks on Cloud Code, he realized he was making these repetitive decks over and over and over again. And so he actually has this web app that he built with the examples of the core Cloud Code decks that we know work well, so like a 101, 201, and Mastering Cloud Code. And then he has a way to input specific customer contexts that pulls from Salesforce, that pulls from Gong, that pulls from other nodes so that we can customize the decks for specific customers. And so we'll pull out things like, OK, this customer is using like Bedrock or called for enterprise or console, which affects what features are available to them.
43:53It will pull out things like, OK, this customer is concerned about like the code review stage of the SCLC. And so we'll add a slide about our code review features there. It'll pull out things like, OK, this customer needs to be like HIPAA compliant or needs XYZ security controls. And so we'll make sure to add a slide or two in their deck about that. And then, for example, if-- this is a customer that's on Vertex or Bedrock and doesn't want to use Cloud for Enterprise, then we'll just take out some of the slides that are Cloud for Enterprise only features. And so normally, this is like manual work that could take 20, 30 minutes. And so people either like spend that time doing it, or they'll just decide not to do it and use the general deck. With this, it takes like a few seconds, and you get a tailored deck.
44:41MARK MANDEL: What's interesting about this, like Slack is like the tool that nobody's-- it's just like nobody's trying to create their own. Slack just continues to win. And it's just like the way you describe it is kind of the OS of so many companies. It's so interesting. Like people talk about Salesforce as just like SaaS. We don't need SaaS software anymore. We're going to build our own. It's like Slack is a durable tool that nobody wants to try to compete with and build a better version. LILY FIERRO: I think it's pretty important communications infrastructure. And I think they do the core task of helping everyone get real-time updates incredibly well. MARK MANDEL: Yeah, like people hate on Slack, but it's really great at what it's going to do. And like the most cutting edge teams are hooked on it. So interesting. LILY FIERRO: Yeah, and I also love how easy they've made to customize it.
45:25And so we love making Slack bots. And this kind of like hackability means that we're able to integrate with Slack the way that we want to. So really appreciate Slack's work on that. MARK MANDEL: Time to buy some CRM stock. I am so excited to tell you about this season's supporting sponsor, Vanta. Vanta helps over-- LILY FIERRO: --15,000 companies like Cursor, Ramp, Duolingo, Snowflake, and Atlassian earn and prove trust with their customers. Teams are building and shipping products faster than ever thanks to AI. But as a result, the amount of risk being introduced into your product and your business is higher than it's ever been. Every security leader that I talk to is feeling the increasing weight of protecting their organization, their business, and not to mention their customer data. Because things are moving so fast, they are constantly reacting, having to guess at priorities, and having to make do with outdated solutions.
46:21Vanta automates compliance and risk management with over 35 security and privacy frameworks, including SOC 2, ISO 27001, and HIPAA. This helps companies get compliant fast and stay compliant. More than ever before, trust has the power to make or break your business. Learn more at vanta.com/lenny. And as a listener of this podcast, you get $1,000 off Vanta. That's vanta.com/lenny.
21 · 46:47 · Which teams use the most tokens
46:48OK, so you talked about all these different teams and how they use Cloud Code and Cowork to operate. Which teams do you find other than engineering? I imagine engineering is the biggest token spender. But if not, that'd be really interesting. What's the second place function right now for tokens? Oh, Applied AI is amazing at pushing the boundaries of what Cloud Code and Cowork can do. A lot of our Applied AI team spends time with our customers. Helping them adopt our API. And so sometimes our Applied AI team will, for example, make prototypes on behalf of these customers, which Cloud Code makes so much faster than it used to be. They also have the dual goal of needing to manage a lot of customer comms, a lot of customer inbound, and historical contacts, call notes.
47:37And so they're both extremely heavy on Cowork and on Cloud Code. And just to understand Applied AI, does that work? Does that forward deploy engineering sort of role? How would most people describe what the Applied AI team is doing? Yeah. It's helping our customers adopt the latest API and model features across their company, both for powering their company's products and also for internal acceleration. Got it. So it's like customer success, go-to-market-y, kind of like forward deploy engineering sort of thing. Exactly. It's like a very technical go-to-market person. Got it. OK, awesome. So you're saying that might be the second org that uses the most tokens. Yeah. And then we also see them pushing the boundaries of what Cowork can do. So for example, a lot of these folks cover multiple customers. And in any given day, can have like 5 to 10 customer engagements on a high day. And so what they often use Cowork to do is the night before, they'll ask it to summarize, OK, what are all my customer meetings that are coming up?
48:44The next day? What are all the things that this customer has asked me for? What's top of mind for them? What are the action items from the past meetings? And Cowork will just put together this dossier, this brief of what they should be aware of going into the next meeting. And Cowork can also research answers. So if a customer asked, OK, when is feature X going to launch? Cowork can help the PyDI person research through Slack to get the latest ETA, add that to the-- add that to the notes so that during the customer call, the PyDI person has the absolute latest. And these are just workflows that people are building for themselves and sharing with other people on their team. MARK MANDEL: So cool. Something that-- kind of this question, this trend-- I don't know, question topic comes up a lot recently, which is tokens spend exceeding people's salary, where people just use AI and it costs more than how much they're making.
49:40Are there any numbers floating around on topic of just how much tokens spend, say, engineers spend, I don't know, a month, a day, or PMs, anything like that? JENNY GUY: It is clear to us that as the models get better, people delegate far more tasks to it, and they spend a lot more hours in tools like Cloud Code and Cowork. And so we do see the token cost per engineer or per any knowledge worker increase every time that there is a model jump or a substantial product improvement. JENNY GUY: I think it's-- it's still much lower than what the average engineer salary is, but we see the percentage increasing over time. MARK MANDEL: It's such an interesting-- we talked about how you have access to the most cutting edge models and other advantage of working in Anthropic. I believe you guys have basically unlimited tokens.
50:30You don't-- you can use as much as you want, is that right? JENNY GUY: We can use a lot of tokens. Some people do run into limits, so-- MARK MANDEL: OK, there's a limit. OK. Boris, shut it down. OK, it's so interesting how many advantages come from having the most advanced model. It's such an interesting, like, flywheel that starts to kick in. I think we also believe a lot in empowering our internal teams to build as fast as possible. And we also trust that everyone understands how much capacity that serving these models truly costs. And we trust our team to use the tokens responsibly. So it's very frowned upon to waste tokens, but we do trust individuals to make that change. And we trust our team to make that judgment call.
22 · 51:15 · The emerging skills PMs need for AI companies
51:15MARK MANDEL: Awesome. Coming back to the PM role, we talked a little bit about this, but I think this will be really interesting for people to hear. Just what I want to understand is, what do you think are the kind of the emerging skills that PMs need to develop slash you most look for, AI companies most look for when they're hiring PMs these days? JENNY GUY: I think the hardest skill is being able to define what the product should look like a month from now. I think there's a lot of ambiguity in what models are capable of in that timeline and how user behavior will change. But I think there are patterns that the best PMs can see based on how users are abusing the limits of the existing product. And the best PMs can sense that, can set a direction, and can steadily execute towards it and change the path if the model capabilities are much better than or worse than what they had originally expected. I think it is very hard to be the right amount of AGI pilled. So I think everyone can see this future where the models are extremely smart and can do almost everything, in which case you actually don't need that complicated a product.
52:31You can actually just have a text box again where you tell the model what you want. And it's so smart that it can add any tool or add any integration that it needs to get the job done. It knows when it's uncertain. It can ask clarifying questions. It's very easy to build a product for the super AGI strong model. I think the hard thing is figuring out, for the current model, how do you elicit the maximum capability? How do you help users get onto the golden path? How do you guide users to interact with the model strengths and patch its weaknesses? This skill is pretty rare. And how do you build that skill? Is it just basically understanding the limits of each model? Are you talking about taste? Understanding, having taste into what the model maybe is capable of, what it's great and not great at, where it's changed? I think it's spending a ton of time talking and using the model. One of the things I really like to do is to ask the model to introspect on its own behaviors. So, sometimes when I notice that the model does something unexpected, like, for example, there's situations where the model will make a front-end change and run tests but not actually use the UI, it's actually pretty useful to ask the model to reflect on why it did this.
54:01And sometimes they'll say that, "Hey, there was something confusing in the system prompt," or, "I didn't realize that the front-end verification was part of this task," or, "Hey, I delegated the verification to this sub-agent, and the sub-agent didn't do the test, and I didn't check its work." A lot of times, just being very curious about why the model made the decision that it did will show you what misled it so that you can fix the harness in order to close this gap. The other thing that helps is to figure out who are the users who you trust the most to give you accurate feedback about the model. Usually, there's a handful of people who are much better than others at articulating what makes a specific model or model-harness combination good. And there's a lot of people who will give you feedback, but not everyone's feedback is as qualified. And so, finding a group of those five people you trust is really important for getting very fast feedback.
23 · 55:00 · Why building evals is underappreciated
55:04I think the third thing that is useful but not everyone loves doing is building evals. You don't need to build hundreds of evals for them to be useful. Just building 10 great evals is important for helping the team quantify what the goal is, and what their progress towards it is, and what they're missing. And so, I think evals is this underappreciated thing that more PMs, more engineers should be working on. We've covered evals a bunch. There's this trend of just like, "That is the future of product management is writing evals," because essentially, it makes us look like, "Okay, cool. Let me actually concretely define it, and then we'll know." How much of your time are you spending writing evals, would you say? I think the importance of evals varies a bit based on the feature that you're working on or what the problem you're trying to solve is. So, there are a lot of folks on our team who do spend a lot of time working on evals. We have a small pod of folks who collaborate very closely with research to more precisely understand our cod code behaviors and what the largest areas of improvement are, and trying to measure those pretty concretely. I personally jump into evals when there's a feature that I think needs a bit more product definition. And often, the output of this is, "Okay, here are like five evals that I made.
56:30This is how you run them. These are the ones that succeed, and these are the ones that don't. And this is like the prompt that I've used to increase the success rate." It varies a lot, though. Based on the exact feature. Not every feature needs it, but I think features such as memory benefit a lot from it. This point you made about people being very good at evaluating models is so interesting. It's almost like a human eval of just like, "Okay, they understand where it's spiking or it's maybe lacking." Is there anyone specific that you want to shout out that's very good at this? Two people who I think are incredible at this are, one, Amanda, who molds Claude's character. It's just like such a hard role because the task is so ambiguous. Even coding is easier because you can verify the success, whereas crafting the character requires a very strong sense of conviction in who Claude should be. And I think she has like an incredible ability to not only mold the character, but also to like articulate what the goals are, what the character, what's successful, and what's not successful.
57:35And I think she has like an incredible ability to not only mold the character, but also to like articulate what the goals are, what the character, what's successful, and what's not successful. And I think she has like an incredible ability to not only mold the character, but also to like articulate what the goals are, what the character, what's successful, and what's not successful. And I think she has like an incredible ability to not only mold the character, but also to like articulate what the goals are, what the character, what's successful, and what's not successful. And I think she has like an incredible ability to not only mold the character, but also to like articulate what the character, what's successful, and what's not successful. The other group of people who I really trust is just like the Claude Code team. So we often have team lunches and whenever there's a new model we're testing, one of the fastest ways for us to get feedback is to just like at these team lunches, just like go to every single person and just be like, "Hey, what is your vibe on the model?" And oftentimes we'll get feedback like, "Okay, this model is like not fully explaining its thinking.
58:06It's like too abrupt." Or like, "Hey, this model is like, just like loves writing a ton of memories, but like we're not sure if the memories are high quality or not." Or like some people will notice that, okay, this model loves to test itself, which is great. Or like this model isn't testing itself enough. So that informs what data we look at to verify, okay, is this a larger pattern? So we have a ton of data, but it is very hard to extract insights. And so the feedback from this group, it's like, "Hey, this model loves to write a ton of memories, but like we're not sure if the memories are high quality or not." So we have a ton of data, but it is very hard to extract insights. And so the feedback from this group is like, "Hey, this model loves to write a ton of memories, but like we're not sure if the memories are high quality or not." So we have a ton of data, but it is very hard to extract insights. And so the feedback from this group helps us inform, okay, what are the hypotheses we want to test? And then we're able to extract data to test that. And then we're able to extract data to test that. And then we're able to extract data to test that. And then we're able to extract data to test that. And then we're able to extract data to test that.
24 · 58:44 · Why Claude's character and personality matter so much
58:44And then we're able to extract data to test that. This point you made about the character of Claude, I had Ben Mann on the podcast, co-founder, and he talked about this, just like the character, the constitution of Claude is such an important part of Claude. And then we're able to extract data to test that. And then we're able to extract data to test that. And then we're able to extract data to test that. And then we're able to extract data to test that. And then we're able to extract data to test that. And then we're able to extract data to test that. And then we're able to extract data to test that. And then we're able to extract data to test that. And then we're able to extract data to test that. And then we're able to extract data to test that. And then we're able to extract data to test that. And then we're able to extract data to test that. And then we're able to extract data to test that. And then we're able to extract data to test that. And then we're able to extract data to test that. And then we're able to extract data to test that. And then we're able to extract data to test that. And then we're able to extract data to test that. And then we're able to extract data to test that. And then we're able to extract data to test that. And then we're able to extract data to test that. And then we're able to extract data to test that. And then we're able to extract data to test that. And then we're able to extract data to test that. And then we're able to extract data to test that. And then we're able to extract data to test that. And then we're able to extract data to test that. And then we're able to extract data to test that. And then we're able to extract data to test that. And then we're able to extract data to test that. And then we're able to extract data to test that. And then we're able to extract data to test that. And then we're able to extract data to test that. And then we're able to extract data to test that. And then we're able to extract data to test that. And then we're able to extract data to test that. And then we're able to extract data to test that. And then we're able to extract data to test that. And then we're able to extract data to test that. And then we're able to extract data to test that. And then we're able to extract data to test that. And then we're able to extract data to test that. And then we're able to extract data to test that. And then we're able to extract data to test that. And then we're able to extract data to test that. And then we're able to extract data to test that. And then we're able to extract data to test that. And then we're able to extract data to test that. And then we're able to extract data to test that. And then we're able to extract data to test that. And then we're able to extract data to test that. And then we're able to extract data to test that. And then we're able to extract data to test that. And then we're able to extract data to test that. And then we're able to extract data to test that. And then we're able to extract data to test that. And then we're able to extract data to test that. And then we're able to extract data to test that. And then we're able to extract data to test that. And then we're able to extract data to test that. And then we're able to extract data to test that. And then we're able to extract data to test that. And then we're able to extract data to test that. And then we're able to extract data to test that. And then we're able to extract data to test that. And then we're able to extract data to test that. People really like that Claude's low ego.
1:00:03And so if you tell it, Hey, you did this thing wrong. It's like, truly sorry. It's like, Oh shoot. Like, thanks for telling me, like, let me fix it. Let's work together. It's also very positive. So if you're feeling like, Oh, this is like an insurmountable task. I don't know how to get started. Claude is like, okay, it's okay. These, these are like the steps that I think we should take. Like, do you want me to get started on it for you? I think part of what makes a great coworker is this positivity, this like bias towards action, this, this ability to give you like earnest feedback, not just agreeing with every single thing that you say. And so we try to imbue this into Claude because we think it makes it
25 · 1:00:44 · How new models force product changes
1:00:44a lot more enjoyable to work with. There's something I want to come back to. You talked about how, when new models come out, you often have to kind of revisit things you've built. That's so interesting. And so like frustrating, maybe just like, Oh, goddammit, we ship this thing. Now I have to rethink it. Talk about just like how often. You have to come back with a new model and they're like, okay, we have to redo this product that we launched a few months ago, a lot of the changes that we make with a new model is removing features that are no longer needed. So a lot of times we add features to the product as a crutch for the model, because it's not naturally doing itself. So the classic example for this is a to-do list. When we first launched Claude code, people would ask it to do these large refactors and Claude code would say, okay, cool. I need a. And then we would change these like 20 call sites and it would go and change five of them and then stop. And then we were like, okay, how do we like force it to remember to get every single one of these 20?
1:01:39And so Sid on our team was like, okay, what if we just like think about what a human would do? A human would like make a list of everything that they need to change. Similar to how in VS code, you would look up all the call sites and it'll be a list on the left side and you would like go through them one by one and replace all. How do we give this kind of like a tool to Claude? And so he added the to-do list. And we found that with that, Claude was actually able to fix all these 20 call sites, but then with Opus four and later models, we realized that we didn't need to force it to use this to-do list. It would like naturally use it itself for the earlier models. We had to keep reminding it. Hey, did you finish everything on to-do list? You can't finish until you're done with everything on the to-do list. And for the later models without prompting, it just like naturally thinks to do everything on the to-do list these days, the to-do list is still nice to have as like a user, um, because then you can more clearly see what Claude is working on, but honestly, it's such a de-emphasized part of the product right now that, um, the model may use it. The model may not use it.
1:02:40It's like really not necessary for it to make thorough changes anymore. I forget who said this on the podcast, um, that the model will eat your harness for breakfast. And what I'm hearing here is essentially you. You remove things over time that you've had to add on top of the model where it was not operating the way you want it. And essentially as the models get smarter, you just, it becomes simpler and simpler for it just to do the thing you want it to do. Yeah. Um, we can move, remove a lot of prompting interventions every time the model gets smarter. And we actually do this every time we launch a model, we read through the entire system prompt and we reflect on, okay, for each of these sections, does the model really need this reminder anymore? And if not, we'll remove it. The most exciting thing that new models unlocks though, is just like entirely new features. So there's a lot of features that we've been testing out with prior models and the accuracy wasn't high enough for us to want to launch them. And so one example of this is code review. We tried to build a code review product a few times, and we've launched like simpler versions of code review, which is the slash code review command in the past. And it was only with the most recent models that we felt like, okay,
1:03:50this code review is so good that our engineering team relies on this code review to pass before we merge PRs. And we found that this was, we've always dreamed of Cloud being able to be a reliable code reviewer that can actually, that we can like confidently feel catches the majority of bugs. And it was only with like Opus 4.5 and 4.6 and Sonnet 4.6 that we felt like, okay, we are now able to like run multiple code review agents simultaneously to traverse the entirety of the code base and to synthesize a set of like real issues that an engineer needs to address before merge. And so this is like a new capability that the, the newest models have unlocked. This is another trend that is very common on this podcast of build something that will possibly be possible in the next six months, be kind of at the edge of what's working sort of, and then it'll catch up and then it'll be an amazing product and you'll be ahead of everyone. Yeah, exactly. Um, it's pretty important. You build products that don't necessarily work yet so that you know, okay, what is missing, um, for this product to work. And then with the newest model, you can just swap it into the prototype you've
1:05:09already made and see, okay, does this new model close that gap?
26 · 1:05:11 · The vision for Claude Code and Cowork
1:05:12How much are you able to speak to just kind of where things are going with Cloud and co-work as kind of the vision of it? I imagine you don't wanna give away too much about the goal, but it feels like you're, there's all these awesome features being added on top dispatch, control from phone and all these mobile app, all these things, what's kind of just like a way to understand the vision for all these things long-term. We think about this in terms of building blocks. So for both cloud code and co-work, the core building block is making individual tasks successful. So you, you want it to produce some output. You give it a clear prompt description. Is it able to consistently produce acceptable output that you're able to either merge or share with your colleagues or external audience? So the task is the core building block. As the models get smarter, the task success rate gets a lot higher. And then we see people moving towards doing multiple tasks at the same time. So multi-clouding was this big thing and towards the end of 2025, and it's only increased since then. And so we see this as, okay, great. One task works, and now you can do like six tasks at a time.
1:06:17As the models get even smarter, the way that we are extrapolating this is, okay, next, maybe you're going to run like 50 clouds at a time or hundreds of clouds at a time. And so what is the infrastructure we need to build to enable that? At that point, you're probably not going to run everything locally on your machine anymore. There's just like not enough RAM to do it. And so we're, we're thinking about how do we make it easier for you to manage all these? These will probably run remotely. How do we build the interface so that you as a human know which tasks you need to look, look into, how do we make sure that the agent is fully verifying its work so that when you look at a task and it says it's done, you like, can very quickly verify and fully trust that it is done to your spec. And how do we make sure that this like process is self-improving so that when you do see a task that isn't done to your liking, you can give it feedback and the model will know for every future run to incorporate that feedback. So it never makes that mistake again. So this is the progression that we're, we're bringing our users along for.
27 · 1:07:22 · Advice for thriving in an AI-driven world
1:07:23There's a lot of people. People listening, a lot of product managers, a lot of maybe founders, a lot of other cross-functional folks listening. There's a lot of worry about just how their role, the future of their careers, what advice would you have for just people to not just survive this transition to this very AI driven world, but to be really successful to essentially just to thrive in this future? What are just like things people need to hear, need to be doing? I think AI gives everybody a ton more leverage than they used to. And so I would push you towards anytime you realize that you're doing some manual task multiple times, think about how you can use cloud code, cowork or other AI tools to automate that for you. Most people have like creative parts of their job that they absolutely love. And then like tedious parts of their job that they really hate doing. I think the beauty of AI is that it can do those tedious parts for you. It can learn from every time that you've done that. It can do the manual task and generalize and then run it automatically and so that you can focus on the creative parts. And that means you can do a lot more than you used to be able to do. So I think my like immediate push for people is figure out the repetitive parts that you can pass to cloud, iterate on those automations until the success rate is very high and then focus on, okay, what more can you be doing for your team, for your product, for your company that like people haven't had the bandwidth to pick up so far?
1:08:53Like, what is that like pet project that you always thought the company should do that like you've never had bandwidth to do? If AI can take care of the like grunt work, then you have you have this extra 20% time now that you might not have before. So so my push is to lean into these tools, hand off the work that you're not excited to do, figure out how it can accelerate you. And then as a result, you'll be able to do so much more.
28 · 1:09:18 · Why 95% automation isn't good enough
1:09:19Something core to what you just shared, which I fully agree with, is find problems. Problems to solve with AI. There's all this potential what all these tools can do. Some of the hard, like for a lot of people, the hardest part is just like, what should I actually do? And what you're saying here is just pay attention to things that you are doing constantly. You can automate, pay attention to just like ideas that have been floating around that you haven't had time to do. It's basically it's like solve a problem for yourself is kind of the core advice there. Exactly. I would also push listeners towards focusing on bringing your automations from, okay, this is a great tool. This is a cool concept to like, Hey, this actually works a hundred percent of the time. Like sometimes I see users trying to automate something, getting it to like 90, 95% accuracy, and then giving up on it. And this, if an automation doesn't work a hundred percent of the time, it's not really an automation. And that last five to 10% does take more time.
1:10:15Also building the automation is often a lot slower than you doing it yourself. I would encourage listeners to put. Put in that time to scope some automation that you really want to get to a hundred percent, put in the elbow grease to teach quad your preferences, to like give it feedback so that it can improve its skill so that it can get to that a hundred percent. And then like really, then you'll be able to rely on it. There there's just not much value in a 95% there automation. I am super guilty of that. This is really good advice for me. I am guilty of this too. I've been teaching it. I've been teaching co-work to. Try to get me to inbox zero for Gmail and it has not been, it has been very time consuming and it is definitely not there as you probably realize. Yeah, I funny enough. That's exactly where my mind goes. I have this, uh, workflow I set up where every email I get, it looks for things that are spammy, which is just like all these, like, Hey, can I come on your podcast? Or what about the spot? Like all these things, I'm just like, I don't have time for these sorts of things and I have it categorized it into a folder called spammy. And it's just like, it's.
1:11:2295% great. But then there's like, oh, wow. I missed an email cuz it went in there. So this is a good push for me to like, I'm gonna work on this. I'm gonna get it to perfect. Yeah. We also are working on making the flow for customizing these commands a lot easier. Cuz right now I think you have to like know too many concepts. You have to know to define a skill. You have to know to like use this skill and give it feedback. And then you have to know to tell co-work to update the skill based on all the feedback that you gave. And then you also have to know where to read the skill to like make sure that the feedback was incorporated the way that you. You want the, it's also our job to make this flow really seamless so that it doesn't feel painful to do.
29 · 1:11:58 · Build apps you use every day, not prototypes
1:11:58Amazing. Is there anything else cat you wanted to share anything else you wanted to leave listeners with anything you wanted to double down on that? We haven't already touched on before we get to our very exciting lightning round. I see a lot of people playing around with AI, um, and building like prototype apps and tinkering with building workflows. I would really push people towards building apps that you're actually using. Every single day cuz I think only through that usage, are you actually getting the value? Like if you build a prototype app that isn't helping you get more done, then the, the AI isn't really adding value to your, to your day. And there's only so much you learn from that when it's like, okay, I just did one shot at something. Oh, that's cool. And then you never come back to it. It's like, you're not learning a lot and you're not getting like much leverage from it and actual leverage. Yeah. That's such a good point. I also think there's a lot of people who spend a lot of time. Like customizing their workflow. So there's like, I think there's like two ends of the spectrum. One is like people who never customize or never build automations, but there's like this polar opposite end of people who like obsessed around customizing their tool, like adding a ton of skills and MCPs and, um, these like workflow improvements. And I think sometimes that can even distract from your core goal of like launching some product or building some feature. I think there's a lot of fun in customizing and we. Definitely wanna make our products very hackable so that you, you can make it work really well for you, but there is a limit to how much it's useful.
1:13:30Um, and I think there, there's a camp of people who maybe spend so much time customizing that they're like not sleeping and not doing the like core task that they originally set out to do.
30 · 1:13:41 · The divide between AI skeptics and believers
1:13:41I see a lot of that on Twitter. Just like, look at my setup. It's out of control. It's so optimized. And what do you, what are, what are you actually building? No, but my setup is so awesome. I could get so much done. I think the simple setups actually work better. Slash power up, get, take level up a little bit. Yeah. Yeah. There's this Karpathy tweet that just, uh, to came out on yesterday where he talked about this divide. That's interesting between people that tried ChatGPT cloud back in the day. It was like, okay. And they're like, nah, this is this terrible. And they kind of gave up on like what AI could do for them. And they just like, so cynical of like, no way. It's not actually that big of a deal. And then there's people that are using it to code essentially. Yeah. Who see the full intense power of it and how good it is. And people on both sides don't understand the other side and why they like how much they, how they see the world.
1:14:32And so your advice is really good here. Just like actually use it for real things and see how good it actually has gotten. Yeah. I think the big shift is that the 2024 generation of products were chat based and the cloud code generation of products is action-based and the. So like big aha moment people have is when cloud can just like do things on your behalf. It is, it is an amazing feeling to know that the agent is capable of doing so much more than telling you what to do. Like the agent can actually just do it itself. And when people feel that, I, I think that's the eye opening moment. Shout out, uh, Chrome extension, the cloud called Chrome extension, which you could just watch it doing stuff that you'd be like, fill out this form for me. And I'm like, all right, here I go. Exactly.
31 · 1:15:19 · Lightning round
1:15:19Okay. Uh, anything else before we get. A very exciting lightning round. No, let's do it. Let's do it. Uh, Kat, I've got five questions for you. Welcome to the lightning round. There's this animation that place I have to make sure to say it. Uh, are you ready? I'm ready. First question. What are two or three books that you find yourself recommending most other people? I really like how Asia works. Um, it's a story about economic development and what are the, like the policies and, uh, governments that make, uh, long. That. lasting, successful economies. The other books that I'm really into are The Technology Trap. So this is actually about the past few technology revolutions, so the industrial revolution and the computer revolution, and how this has affected workers. The reason that I really like this is because I think there's a lot we can learn from history to make sure that this transition goes well. And maybe on like a fun note, I really like Paper Menagerie. It's just like a book of short stories about like coming of age and AI and just like self-discovery. Favorite recent movie or TV show you have really enjoyed? I really like Drive to Survive. There's no like deeper meaning to it.
1:16:39I just, there's just something very satisfying about people being so, so obsessed with like a singular engineering goal and just like the purity of their pursuits. And I also really love Free Solo, which is about Alex Honnold climbing El Capitan without a harness. And I think similarly, it's just such a pure achievement to be able to climb this extremely challenging, dangerous route and to be able to have the mental focus to do it, knowing that if you make a single mistake, you die. It's insane. Yeah, that movie is out of control. And it's interesting how these relate in some way to the work you do. I actually am a rock climber. I first watched Free Solo before I climbed rocks. And so I thought it was impressive, but I didn't understand how impressive it was. It's one of the rare movies where like, the more you know about it, the more you're, you're blown away by how insane this is. Like the kinds, the kinds of movies he's doing on the wall are things that like, I don't think I will ever be able to do in my lifetime if we're set in a gym, like one feet off the ground. With a rope. With a rope.
1:17:49Did you see the documentary on that other guy, the younger one that went on like ice mountains? I did. That one was very sad. But that was, that was wild. Okay. Favorite product you recently discovered that you really love? The product that is like most changed my life outside of cloud products is probably Waymo. Like I'm a diehard Waymo user. Use it twice a day, get to and from work. Yeah. So the two things that I really like about it are one, I don't feel bad if a Waymo is waiting for me. And so I feel like I feel less pressure to be right at the curbside the moment it arrives. And the second thing is I feel like it lets me be a bit more productive. When, when I'm in the car with another human, I, I typically try not to like do any work calls. I, I feel a little rude if I'm like on my laptop the whole time. But one thing I really appreciate about the Waymo is I can call into a work call. I'm not worried about someone overhearing me. I'm not worried about,
1:18:47hey, is this like rude? Am I talking too loud? Do I need to tell, ask someone to like change the music? And so this has been like, I feel like this has given me back like 30 minutes every day. All these second order effects of, of technology. It's so interesting. Yeah. I always thought Waymo needed to be priced lower than Uber and Lyft to succeed, but actually I'm like very happy to pay a 2X premium for it. I love Waymo. It's just like, like you, once you see it, you're just like, ah, this is in, in, insane. And, and then you get used to it. Like you get in there and you're like, this is crazy. And then you forget about it. Totally. And I think it's also changed the vernacular. Like a lot of people at Anthropic love Waymo. And I think in the past you'd be like, hey, like let's call like blah blah rideshare app. And now like everyone's just like, okay, is Waymo here? Okay. Two more questions. Do you have a favorite like motto that you often come back to in work or in life? Just do things. I think there's a lot of value in like first principles thinking. And if, if you like, if you know what you're optimizing for and you have like strong first
1:19:46principles, then you can normally deduce what the right, like course of action is and be able to clearly articulate that to all the stakeholders. And then you should just like do it. Like I think jobs are fake. If you understand the constraints, you can figure out what you can do and then just like try to do it quickly, learn from the mistakes and apologize or fix them if you did something wrong. You could just do things. Whoever said that. I think it's liberating actually. To like tell people this, I think a lot of companies like roles are very strictly defined, like, okay, this is what the PM does is what the designer does is what engineer does. And then even team scopes are very rigidly defined. So, hey, like this corner of the code base we touch and this corner, like we're not allowed to touch. And I think what just do things lets people do is they feel like empowered to make these decisions, empowered to operate across team boundaries, just to like get something done. That feels like a big, important skill. To be good at, people call it agency. Just like do the things that need to be done. Bias towards action. Bias towards action. All these ways of describing just like, don't wait for permission.
1:20:50Yeah. I think this is my favorite reason to work at a startup at some point in your life, because like one thing that was like very life-changing for me was actually working at scale when we were 20 people. And so there was just no process and we had like really big problems that we needed to solve. And it was like, I really appreciate Alex and the rest of the team for like empowering me. Alex: And the rest of the team to just like figure things out without any boundaries for what sales is supposed to do, what office is supposed to do, what engineer is supposed to do. Just like you have all the tools at your disposal. You have some like ambitious, hairy problem statement and you can do whatever you need to like get to a good solution. Like you almost need that experience to build that skill, to feel comfortable doing that. Cause a lot of people, you know, they go through school or in college and all these, like do the thing we tell you to do, and then you will get a good grade. And you have to kind of unlearn that of like, okay, I'm just going to do the thing that needs to be done. And even if people think it's dumb, I think it's the right thing to do. Yeah, exactly.
1:21:47Okay. I actually have two more quick questions. Two more final questions. One is, when Claude thinks there's all these, I don't know if you call them verbs, what's the term for these things? Thinking words. Thinking words. And interestingly, these all leaked in the source code. Is it, do you have a favorite thinking word? I really like manifesting. It's also like the sticker that I have on my laptop. Oh, amazing. Clearly the winner. Okay. Final question asked for us this too, with AGI potentially arriving in our lifetime, when you don't potentially have to work, what are you going to do? What are you going to do with all your time? I think it will take a long time for AGI to diffuse across society. So I think the immediate thing is actually just like helping bring the world along. I think my like non-serious answer for after this happens is I'll probably just do a lot of rock climbing. I'll probably just live in some, I'll probably move to like Fountain Blue and just like live amongst 10,000 boulders and climb for a bit. There's also so many books I want to read that my goal is to be able to read one or two books a week. And I'm currently at probably like 0.5. The backlog is pretty big. I think there's just so much we can learn from history and so much that I don't understand as well as I would love to.
1:23:11Anything about physics and or like robotics or like any hardware or like aerospace or there's just so many interesting topics. So I'm excited to learn even even knowing that the AGI will already know it. Kat, this was amazing. You're awesome. Two follow up questions. Where can folks find you online if they want to reach out and just follow what you're up to? And how can listeners be useful to you? The best way to reach out is I am underscore Kat Wu on Twitter. Feel free to like tag me in things, feel free to DM me. I read all my DMs. I don't always respond to every single one, but I will read them all. And then the thing that is most helpful is tell us where Cloud Code and CoWork aren't working well for you. We are very grateful for the amount of positive feedback. But the thing that we thrive on is edge cases, errors, like specific tasks that we can reproduce where Cloud Code or CoWork fail. Because if you are able to share that with us and we're able to reproduce it, then this is something that we're able to actively improve for our next generations of models and for our next harnesses. Extremely cool. Everyone on people on Twitter
1:24:28are not shy with sharing this feedback. So keep it coming. Share, share. Please, please share the problems that you're having with us. Yeah. And it's really cool to see all you, your team being on so active on Twitter and responding to people. And so, so like what I'm hearing, like, this is actually stuff you guys actually see and react to. So. Yeah. We appreciate everyone being so engaged with us. It gives the team a ton of energy. We, we have this channel of like user love. And so whenever you guys share a success story, we post it there. And whenever you guys share like issues with our product, we put it into our feedback channel. That way our broader team is able to act on it. That is so cool to know. Thanks for sharing that. Well, Kat, thank you so much for being here. Thanks for having me. Bye everyone. Thank you so much for listening. If you found this valuable, you can subscribe to the show on Apple podcasts, Spotify, or your favorite podcast app. Also please consider giving us a rating or leaving a review as that really helps other listeners find the podcast. You can find all past episodes or learn more about the show at Lenny's podcast.com. See you in the next episode.
01 · 0:00 · Introduction to Cat Wu
0:00我认为要对 AGI 保持恰到好处的信念是非常困难的。为超级 AGI 强模型构建产品其实很简单。难的是搞清楚对于当前的模型,如何激发出最大的能力?我从来没见过像你们 Anthropic 这样的发货速度。我们想消除所有阻碍发货的障碍。我们很多产品功能的开发周期从六个月缩短到了一个月,有时候甚至一天。你在面试几百个 PM然后你一直觉得他们的方法完全不对。PM 的角色正在发生巨大变化。变化非常快。构建 AI 原生产品最关键的一点就是快速迭代,想办法真正做到每周都发布新功能。你觉得 PM 需要培养哪些新兴技能?归根结底还是产品品味。随着写代码变得越来越廉价,更有价值的是决定写什么代码。今天我的嘉宾是 Cat Wu,Anthropic Cloud Code 和 CoWork 的产品负责人。
1:01Cat 处于 AI、产品和构建领域所有变革的核心,她和她的团队正在打造的产品正在改变我们所有人构建产品的方式。她充满了洞察力和智慧与经验。这是一集你不能错过的节目。在我们开始之前,别忘了去 Lenny's Product Pass dot com 看看为 Lenny's newsletter 订阅者独家提供的超值优惠。接下来,有请 Cat Wu。
02 · 1:29 · Working with Boris Cherny
1:31Cat,欢迎来到播客。谢谢邀请我。我有好多问题想问。很高兴你能来参加这个播客。我想先让大家了解一下你和 Boris 的角色分工。大家都认识 Boris。他那期节目是这个播客上最受欢迎的节目。没有压力。他创造了 Cloud Code。他带领团队,发货,每天从手机上提交无数的 PR,我甚至都不知道数字是多少了。我觉得大家没有给你足够的认可,Cloud Code以及 CoWork 和你们正在构建的所有东西取得的成功。帮我们了解一下你在团队中的角色,你是怎么和 Boris 合作的,你们怎么分工的,Cloud Code团队的 PM 角色是什么样的?我觉得很幸运能和 Boris 一起工作。他一直是一个很好的思考伙伴。他是我们的技术负责人。他在很大程度上是产品远见者,他非常擅长设定产品在三个月、六个月后应该是什么样的。
2:33这就是产品在高度 AGI 化之后的样子。我的角色很大程度上是弄清楚,好吧,从我们现在所处的位置到三到六个月后的愿景,路径是什么?我把更多时间花在跨职能协调上。确保我们的市场团队、销售团队、财务产能等等,都认同这个计划,并且我们都在朝同一个方向努力。而且一旦功能准备好了,没有任何障碍可以阻止发货。我觉得在很多方面这种合作效果很好,因为我们基本上心意相通,但这条界限实际上非常模糊。我觉得我们大概 80% 是心有灵犀的。然后有大约 20% 的东西可能我比 Boris 更在意,所以我会去推动那些,还有 20% 是他比我更在意的,他就直接推动那些。本期节目由我们本季的首席赞助商 WorkOS 带来。OpenAI、Anthropic、Cursor、Vercel、Replit、Sierra、Clay 以及数百家其他成功公司有什么共同点?它们都由 WorkOS 驱动。
3:37如果你在为企业构建产品,你一定感受过那种痛苦——集成单点登录、SCIM、RBAC、审计日志以及其他大公司需要的功能。WorkOS 将这些阻碍交易的功能变成了即插即用的 API,这是一个专为 B2B SaaS 构建的现代开发者平台。我投资的几乎所有初创公司在开始向高端市场扩张时最终都会选择 WorkOS,因为他们是最好的。无论你是一家试图拿下第一个企业客户的种子期初创公司,还是一家在全球扩张的独角兽。WorkOS 是成为企业级就绪和释放增长的最快途径。它本质上就是企业功能版的 Stripe。访问 WorkOS.com 开始使用,或者直接联系他们的 Slack,那里有真正的工程师等着回答你的问题。WorkOS 让你能用令人愉悦的 API、全面的文档和流畅的开发者体验更快地构建产品。
03 · 4:29 · What Anthropic looks for when hiring PMs
4:29前往 WorkOS.com,今天就让你的应用为企业级做好准备。你在录制前分享的一件事是你一直在面试大量的 PM。如果每次有人找我要介绍去Anthropic 做 PM,我就能赚 300 亿美元的 ARR。那简直就是大家最想去工作的地方。所以我都能想象你在面试多少 PM。你告诉我你看到人们做得不对,他们研究成为成功 AI PM 的方式是错误的。说说你看到了什么,以及人们需要理解什么才能在当下取得成功。我认为在 AI 之前,技术变革的速度慢得多。所以你可以在六到十二个月的时间跨度上做规划。而且因为你发布功能的速度相对较慢,所以有更多精力花在与所有合作伙伴团队协调上,确保他们发布的功能能解锁你的功能,因为那个时候代码非常昂贵。我认为现在有了 AI,加上工程效率的大幅提升,以及模型能力提升的速度如此之快,我们很多产品功能的开发周期从六个月缩短到了一个月,有时候甚至缩短到一周甚至一天。这意味着我们实际上必须确保产品能非常快地发布。
5:55这意味着作为 PM,不应该那么强调确保你的多季度路线图与合作伙伴团队对齐,而应该更多强调,好吧,我们怎么找到最快的途径把东西推出去?我们怎么打造一个概念试验田,让工程师有一个想法或者 PM 有一个想法,
04 · 6:18 · How to help your teams move fast
6:18到周末就能交到用户手中。我认为在 AI 原生产品上做得最好的 PM是那些能想出如何缩短从有想法到实际把产品交到用户手中的时间的人,并且能定义哪些是最重要的、需要开箱即用的功能。我喜欢的是你说的,人们还没有意识到他们需要多快地行动,以及现在工作中有多大一部分就是推动行动,就是帮助团队快速前进。什么能帮助做到这一点?你做了什么?你的 PM 团队做了什么来帮助这么快地推进?除了能使用最先进的模型之外?我认为首先是设定清晰的目标,因为 LLM 太通用了,这实际上在我们为谁构建、解决什么问题、核心用例是什么方面造成了很大的模糊性。所以我认为一个优秀的 PM 能说,好吧,我们的核心用户是专业开发者。我们想为这个功能解决的主要问题可能是权限提示太多,人们感到疲劳,而用例是我们希望企业中的专业开发者安全地实现零权限提示。
7:36这实际上设定了一个相当清晰的目标,因为它排除了很多潜在的减少权限提示的方法,让人们可以通过一个提示完成更多事情。然后我认为第二件非常重要的事情是找到某种可重复的流程来发布这些功能。所以对于 Cloud Code,我们几乎以 Research Preview 的形式发布所有功能。我们在发布时会明确标注,这样用户就知道这是一个早期产品,这只是一个想法。这只是我们正在尝试获取反馈并不断迭代的东西,这个功能可能不会永远存在。这样做的好处是降低了我们发布东西的承诺。我们可以在一两周内就推出一些东西。PM 应该做的第三件事是帮助团队建立框架,让他们知道什么时候需要拉入跨职能伙伴。以及这些跨职能伙伴的期望是什么。比如说,我们在工程、市场和文档之间有一个非常紧密的流程。
8:37所以当工程师觉得某个功能已经准备好并且我们已经内部吃自己狗粮了,他们会把它发到我们永久的发布房间里。然后负责文档的 Sarah、负责 PMM 的 Alex 以及 DevRel 的 Tarek 和 Lydia就会跳进来,第二天就能把市场公告发出来。因为我们有这个非常紧密的流程,任何工程师发布东西的摩擦都降低了。
05 · 8:58 · How PRDs and roadmaps have evolved at Anthropic
8:59而 PM 就是应该搭建这个的角色。PRD 在这里面是怎么 fits 的?你说目标是非常重要的部分,就像对齐成功是什么样的,这是为谁的,不是为谁的?你们写 PRD 吗?还是就是几个要点?这个在 PM 的世界里是怎么演变的?所以我们做两件事。一是我们有非常严格的指标,我们每周都会和整个团队一起做指标回顾。这样做的目的是确保每个人都深入理解我们业务的所有方面,我们的核心目标是什么,它们的趋势如何,以及什么在驱动它们。第二件事是我们有一份团队原则清单,包括我们的核心用户是谁,为什么这些是我们的核心用户。我们阐述所有这些是为了让团队中的每个人都觉得自己理解我们的业务是如何运作的。他们理解什么对我们重要,以及我们愿意做什么取舍。这让人们可以自己做决定,而不会觉得被 PM 或任何其他利益相关者阻碍。
9:58我喜欢这里面大部分内容都是,好吧,我们未来仍然需要 PM。而网上有那么多讨论说,我们为什么还需要 PM?我们只需要发布和构建就行了。我们需要工程师。哦,我们实际上还是会写 PRD 的。所以我认为对于那些特别模糊的功能,确实有助于写一页纸来说明目标是什么,令人愉悦的用例是什么,目前需要修复的失败模式是什么。偶尔也会有一些项目,特别是需要大量基础设施的项目。确实需要好几个月。
06 · 10:28 · The Mythos model and Anthropic's shipping velocity
10:28对于这些情况,我们确实还是会写 PRD。我想进一步深入了解一下你们怎么能这么快。我从来没见过 Anthropic 这样的发货速度。有人做了一个 Anthropic 各项发布的日历。几乎每天都有重大的功能或产品发布。所以网上有人问的是你们刚刚——不是发布,而是构建了这个不可思议的模型 Mythos,它目前还在预览阶段,因为它太强大了。人们有点害怕它能做什么。你们有用过这个吗?这是你们能这么快的原因之一吗?我们已经连续好几个季度都在快速推进了。所以我认为这不完全是 Mythos 的功劳。Mythos 是一个非常强大的模型。我们确实在内部使用这些模型。我认为这在一定程度上提高了我们的发布速度,但我不认为这是速度提升的主要原因。我认为很大程度上是流程和团队的期望。
11:27所以我们的流程非常少。我们想消除每一个阻碍发布的东西。我们要确保团队里的每一个人都觉得有权力把自己的想法从只是一个想法变成一周内发布到世界上,有时候甚至是一天内。太酷了。天哪。拥有最好的模型同时又做产品,这是多大的优势啊。太酷了。我们很幸运能和前沿模型一起工作。我的天。多么棒的优势啊,就是构建一个东西然后使用它,
07 · 11:54 · What happened with the Claude Code source code leak
11:55然后加速得更快。太有趣了。有几个其他的我想在这段对话中走一些支线话题。Anthropic 发生了太多事情,我太好奇你的见解了。一个是大约一周前,Cloud Code 的整个源代码泄露了。有人把它弄出去了。我觉得是有人犯了个错误。你有什么可以说的吗?就是发生了什么?出了什么问题?大家应该知道什么?我们看到后立即进行了调查。我们发现这是人为错误导致的。有一个人在用 Cloud 写 PR。这只是对我们发布包方式的一个更新。而且它实际上经过了两个人工审核。所以这是人为错误的结果。我们已经加固了流程,确保以后不会再发生。这个人还在 Anthropic 吗?他还好吗?是的,是的。这是一个流程失误。最重要的是从中学习,并添加更多保障措施,确保不会再发生。
08 · 12:53 · Integrating with OpenClaw
12:53所以这是我们一直在关注的,而且大部分的措施已经实施了。好的,我另一个问题是关于 OpenClaw。最近有这个动向,阻止人们使用 Cloud 订阅来配合他们的 OpenClaw,人们很不满。他们搞不清楚为什么会这样。感觉像是你们在——你知道,对开源社区造成了伤害。大家需要理解这个决定背后的考量是什么?我们看到了对 Cloud 的大量需求,我们一直在非常努力地扩展基础设施,同时让我们的系统更加节省 token,这样你能获得更多的使用量。它不是为第三方产品设计的,第三方产品与我们第一方产品的使用模式不同。我们花了不少时间。试图找出我们能提供的最无缝的过渡方案。所以我非常高兴能宣布每个订阅用户都会获得一些 API 额度。但是是的,我们确实不得不做出艰难的决定,我们需要优先考虑我们的第一方产品和 API。
14:00所以这就是那个决定的结果。是的,对我来说,这太合理了。你们基本上是在以每月 200 美元的价格补贴这些使用量。基本上就是无限使用这个。我觉得大家不理解的是——他们是在赚钱的。我们在努力实现盈利。当计算资源如此紧缺的时候,我们不能就这么送出去。
09 · 14:19 · How the PM team is structured at Anthropic
14:19所以我理解。回到 PM 团队,PM 团队是什么样的?在 Anthropic,有多少 PM?他们是怎么组织的?是的,我们有几个 PM 团队。我想我们目前大概有 30 到 40 个 PM。我们有研究 PM 团队,由 Diane 带领。这个团队负责理解来自我们客户的所有关于模型的反馈,然后把反馈传递给最合适的研究团队去执行。他们还负责管理模型发布。还有 Cloud 开发者平台团队,负责维护那些 API,Cloud Code 就是建立在那些 API 之上的,他们还发布了类似 managedagents 这样的功能,让你可以构建自己的 agent,我们可以代为托管。然后是 Cloud Code 团队,同时负责 Cloud Code 和 CoWork 核心产品。
15:11还有企业级团队,帮助让 Cloud Code和 CoWork 更容易被所有企业客户采用。这包括从成本控制、RBAC、安全控制,一直到确保这些企业对使用我们的工具感到非常放心和舒适。然后我们还有增长团队,负责推动我们整个产品组合的增长。我们和他们密切合作,推动 Cloud Code 和 CoWork 的增长。我知道他们还和其他团队合作推动 CDP 的增长。
10 · 15:42 · How engineer and PM roles are merging
15:43也就是使用 Cloud API 的用户的增长。说到增长,Amol 刚刚上了播客。他有一个很有趣的洞察,大多数人还没分享过。总有一种感觉说我们未来需要更少的 PM。我们为什么需要 PM?工程师直接发布就行了。他的观点是因为工程师推进得如此之快,PM 和设计师被压缩了,没有那么多时间跟上所有正在发生的事情,每天都有新功能发布。所以他的观点是他需要更多的 PM,因为很难跟上节奏。你怎么看?你觉得 PM 的招聘会增加吗?你觉得 PM 这个职业长期来看会怎样?我觉得所有角色都在演变融合。PM 在做一些工程工作,工程师在做 PM 工作,设计师在做 PM 工作,同时也在提交代码。你可以选择雇佣更多具有出色产品品味的工程师,或者保持工程师招聘不变,然后雇佣更多 PM 来帮助指导他们的部分工作。
16:43在我们团队,我们很专注于招聘具有出色产品品味的工程师。这样可以减少发布任何产品的开销。我们团队中有很多工程师完全能够端到端地完成从在 Twitter 上看到用户反馈到周末发布产品,几乎不需要 PM 参与。我认为这实际上是最高效的发货方式。所以我觉得工程师和 PM 的角色在某种程度上是重叠的,增加任何一个都会带来很多好处。我认为产品品味仍然是一种非常稀缺的技能,我们基本上会雇佣任何我们认为强烈展现出这种能力的人。你的背景是工程对吧?是的,我做了很多年工程师。然后我很短暂地做过 VC,然后加入了 Anthropic。实际上,我们团队上几乎所有的 PM 要么曾经是工程师,要么在Cloud Code 上提交过代码,所以这是我认为有助于建立团队信任的一件事,也让我们能够快得多。
17:52然后实际上我们的设计师
11 · 17:54 · Why product taste is the most valuable skill
17:54之前也做过前端工程师。哇,因为这是个大问题。确实有这种融合在发生。维恩图在合并。我觉得很多人关心的大问题是,如果你来自工程或者产品或设计,哪些核心技能会是最有价值的?我在 Anthropic 和 Cloud Code 看到,工程背景非常宝贵。我好奇在其他公司,如果你有设计背景,做 PM 是不是更有价值,或者就是传统的 PM。我还是觉得归根结底是产品品味。随着写代码变得越来越廉价,更有价值的是决定写什么代码。比如,这个功能的正确 UX 是什么?用户体验它的最令人愉悦的方式是什么?我们收到几万条GitHub issue,什么要求都有。这需要大量的心思和品味来判断,好吧,哪些值得构建,正确的构建方式是什么?我认为这个技能集可以来自任何背景,但我觉得这是最重要的东西。我认为工程背景之所以特别有用,至少在未来几个月内,是因为如果你有工程背景,你对某件事应该有多难有更好的感觉。
19:08这通常是决定你选择构建什么的因素。所以如果某件事非常容易构建,那与其争论,不如花一个小时就做了。但如果某件事更难构建,而且你一开始就知道,那你就会知道,好吧,这会花费更多。让我们的团队把这个推出来。所以在优先级排序上有一些帮助。你说在未来几个月内。这是不是因为模型在未来几个月内可能会变得如此之好,你可能都不需要了解那么多工程知识了。我认为有价值的技能组合确实变化很频繁。所以很难预测超过几个月之后的事情。所以这不太是我认为会发生什么变化的评论,更多的是我认为大的变化会发生。所以你不是说那就是 Mythos 出来的时候,会改变一切。然后我们就不需要了解任何工程知识了。不,我只是在说每隔几个月,似乎都会有一次大的编码能力的
12 · 20:10 · Where human brains will continue to be useful
20:10提升,然后其他角色的价值就会随之改变。我认为最重要的是能够拥有这种第一性原理思维,这样你就能弄清楚技术格局是如何变化的,团队真正需要你做什么,然后跳进去填补那个缺口,因为我认为工作变得越来越模糊,这意味着一个优秀的 PM 能够理解所有的差距,找出哪些是最高优先级的,然后想办法,好吧,我怎么学习那个技能,或者我已有的什么技能可以应用到这个挑战上?所以我认为当前环境重视那些能戴很多帽子的人,能随时切换帽子,而且对自己做什么工作来帮助团队加速非常低姿态。我喜欢这个回答。有一个问题我一直在问像你这样处于前沿的人,那些在用最新工具构建的人,就是在达到超级智能之前,人类的大脑在哪里会继续有用和必要?我听到的是本质上就是选择要做什么,知道市场的走向,弄清楚优先做什么。
21:37然后就是知道你构建的东西是不是好的、对的,并以某种早期版本发布出去。这听起来对吗?还有什么人类大脑至少在未来几个月内会继续有用的地方吗?我认为人类仍然提供了一种模型所不具备的常识水平。而且任何产品发布都有上千个移动的部件。有些非常小,但总是有很多可能出错的地方。我认为模型不一定对所有利益相关者是谁、他们之间怎么关联、他们的偏好是什么、什么是与他们沟通、让他们保持配合的正确渠道有很好的理解。我认为很多这种更加隐性的、常识性的、偏情商类的知识仍然非常有价值。
13 · 22:23 · How to stay sane in constant chaos
22:24当然,我们希望模型在这些方面变得更好,我认为它们会的。但目前,我认为仍然存在差距。作为一个人,你怎么应对如此多的持续变化,就像在龙卷风的中心——也许里面是平静的,但就是你怎么跟上正在发生的一切,你怎么在这种疯狂中保持理智?我认为我们团队充满了拥抱混乱的人。所以我们试着笑着面对每一个挑战,因为总是有这么多事情在发生,总是有这么多风险和棘手的情况。如果你对任何事情太紧张的话,你会崩溃的。所以我们真的在寻找那些能看着一个挑战说,哦,这会很难,但我很兴奋能去应对,而且我会尽我所能做到最好。我知道我不会完美的,但我晚上能睡得着觉,因为我知道我尽了最大努力。这是一个有趣的关于未来需要什么技能的回答,因为——我忘了谁说的,可能是 Ben Mann,说这是世界有史以来最正常的时候了。
23:24是的,肯定会变得更难。就是,我感觉有很多个星期,也许周日晚上有什么P0 的事情,然后到了周一就变成了 P00,到了周一下午就变成了 P000。你就会想,哇,我真不敢相信我周日还在为那个 P0 担心。但我觉得你只需要承认你能做的只有那么多,你需要好好睡觉,这样第二天才能做出好的决定,然后就是无情地优先排序你的时间花在哪里,什么是最重要的事情要做对,并且接受放手一些东西。就是,有些产品我们发布的时候没有我希望的那么完善。但是。你知道,我们的首要目标是帮助赋能专业开发者。如果一个产品不成功,只要它不阻碍核心用例,
14 · 24:16 · What gets sacrificed when you ship so fast
24:17那就没关系,因为我们会听到反馈,然后在下一个版本中修复它。发布一个有 bug 的功能,这种事情以前会让我彻夜难眠。但这现在是我能够接受的事情,因为我知道,好吧,我们会得到快速的反馈,然后在下一个版本中修复它。我脑海中浮现出那个 gif 图。我觉得可能是《加勒比海盗》里的那个在一艘船的楼梯上走下来的人,整艘船都在他周围被摧毁,但他却非常淡定地走下楼梯,一切都在他周围崩塌。这很有趣,因为我遇到的每一个 Anthropic 的人都特别淡定而且特别乐观。是的。我觉得这是一个很有趣的洞察,就是拥有这种平静和乐观,而不是——我的天啊,一切都疯了。是的,我觉得如果你没有这种心态,你会很快崩溃的。我觉得我们招聘的也往往是那些在行业里待了很长时间、经历过很多起起落落的人,他们对什么给自己能量有很好的感觉,知道怎么长期维持自己的能量,我觉得这对我们帮助很大。
25:21太有趣了。有一个我想问的事情是,这些角色在模糊化,工程师变成了 PM,每个人又当爹又当妈,每个人都变成了所有人。我们会在这个世界中失去什么?我们会失去职业阶梯和清晰的职业路径吗?我们会失去设计一致性、代码质量吗?可能有一些缺点。你觉得哪些事情是好吧,那是我们为了更大的利益而牺牲的东西?我们在牺牲产品一致性。过去,当写代码很昂贵的时候,你会仔细规划好你的产品要用的所有东西、产品组合、每个产品之间如何关联、每个产品的用例是什么、它们怎么集成。你基本上每个用例只有一个产品。现在随着 AI 推进得如此之快,加上我们需要测试的想法如此之多,我们确实有时候会有功能互相重叠。很多时候是因为我们内部有两种形式都很喜欢,我们想让外部用户告诉我们哪个更好。
26:24这对新用户来说意味着,新用户可能不知道,好吧,完成某件事的最佳路径是什么?我们需要做更多的教育工作,帮助人们理解核心功能是什么,以及使用它们的最佳实践是什么。我认为这就是发布大量功能的代价。我觉得用户也觉得很难跟上最新动态。通常在传统 PM 中,你每月或每季度发布一个功能,所以用户很容易理解,好吧,我只需要每月来看一次,就能学到一些新东西。如果我忽略六个月,也没关系。不会觉得自己错过了什么。我觉得对于这些智能体工具,不仅仅是 Cloud Code 和 CoWork,而是整个生态系统,人们觉得需要每天刷 Twitter 来看绝对最新的东西是什么。我觉得我们还能做更多来帮助人们减少这种越来越快的跑步机上的感觉,我希望人们能感觉到他们可以直接打开这些工具,工具会教育他们,
15 · 27:47 · The /powerup command
27:47或者教他们想知道的东西,他们能感觉到被带着一起走。是的,我看到你们前几天发布了一个很有趣的功能。我觉得是 slashpower up,它基本上带你了解所有很酷的用法,基本上就是使用 Cloud Code 的所有最佳实践,是这个方向的东西吗?是的,没错。过去,我们其实不想做像 power up 这样的功能,因为我们觉得产品应该足够直观,你实际上不需要任何教程。随着时间的推移,我们意识到有太多功能了,而且对内置引导体验的需求如此之大,所以我们就偏离了之前说不做引导流程的理念。我这样做是因为有太多用户想知道,有 100 个功能,哪 10 个是我必须用的?
16 · 28:32 · Why Anthropic has been so successful
28:32所以我们就把这个整合出来了。是的,这是一个很奇异的世界。Anthropic 在 B2B 企业市场非常成功,而传统上你不会发布一堆东西,你基本上是每季度发布一次,也许吧,而这恰恰相反,每天我们都有新东西。所以顺着这个思路,Anthropic 的这波表现简直就是超凡脱俗的。Anthropic 曾经远远落后。刚开始的时候根本微不足道。就是那种最不受资金青睐的公司之一,没有分发渠道。它是第一个开放的吗?远远领先。就是觉得不可能。Anthropic 长期来看有任何机会进行有力竞争。现在它就是大杀特杀,击败了最大的公司和团队。增长简直了,一个月就 110 亿美元的 ARR。利润和增长。等到这期节目出来的时候,可能还会更高。我觉得从内部来看,哪些因素让Anthropic 能够如此成功,从落后追赶到做到这么好?
29:33最重要的两件事,一是这个统一的使命。很难说这有多重要。我们雇佣最关心为全人类带来安全 AGI 的人。而且这实际上是我们在决策中经常参考的东西,And because we put this mission above any individual product line, we're able to快速做出贯穿整个组织的决策,然后统一执行。所以我觉得这在同等规模的公司里是从未见过的。所以为了确保这一点清楚明白。基本上,把第一使命定为安全、对齐、确保 AI 对世界有益。你是说仅仅把它作为一个清晰的使命就能让决策变得容易得多。如果有两个互相竞争的优先事项,我们会讨论哪个对 Anthropic 的使命更重要。这样就更容易决定我们优先做哪个。然后所有人都会支持我们做出的决定。所以有时候这意味着,比如说,我们想在 Claude Code 上发布某个功能,但另一件事更重要。所以我们就降低这个功能的优先级,等到以后再说。真正有趣的是这解释了,我认为,相比于另一家可能跟 Bopen AI 押韵的公司,做了很多不同的事情。我觉得这是非常有趣的一点。而且我认为这是一个非常有趣的我在这里听到的基本上就是,好吧,我们不会去做社交网络。我们不会去做什么有趣的信息流,因为这不符合我们的
31:04使命。这让 Anthropic 保持了专注,而这似乎是成功的核心要素。说到使命,我觉得就是把 Anthropic 的目标放在任何单个团队或单个产品之上。所以对我来说,我觉得我们要谈的第二件事是使命。对我来说,使命的含义稍有不同。使命意味着各个团队愿意做出牺牲,哪怕损害自己的目标和 KR,也要服务于 Anthropic 的目标和 KR。而且大家非常乐意做这些权衡。比如说,一个极端的例子是,如果 Claude Code 失败了,但 Anthropic 成功了,我会非常开心。整个团队都非常愿意做出这样的决定,遵循这样的思路。我不知道你是否能深入谈谈这个但你觉得 Open Cloud 的决定是不是也是这个原因?就是觉得,这没有推进 Anthropic 的使命。我们需要停下来,因为它的效果没有达到我们的期望。我认为对 Anthropic 来说最重要的事情之一就是扩大我们能触达的用户数量。实现这一目标的方式之一就是通过我们第一方产品的 Claude 订阅。所以我们非常希望有时候以第三方产品为代价。所以我们一直在聊 Claude
17 · 32:28 · When to use Claude Code vs. Desktop vs. Cowork
32:30Cowork 这些东西,我想确保大家能理解。我也很好奇你是怎么用这些工具的。有 Claude Code,有 Claude 桌面版/网页版,还有 Cowork。最好的理解方式是什么?什么时候该用哪个?你分别什么时候用这三个?所以我通常会在终端里使用 Claude Code,当我只是启动一个一次性编码任务,而且想要所有最新功能的时候。CLI 是我们最初的产品界面,也是功能通常最先落地的地方。所以它是所有工具中最强大的。这就是我通常用的,当我只是想启动一两个或者几个任务的时候。我觉得桌面版在做前端工作的时候特别出色。我特别喜欢用的一个功能就是预览。如果我在做一个网页应用,我通常会使用桌面版的 Claude Code。我会把预览面板打开在右边,这样我就能实时看到我正在和 Claude 聊天的同时,正在构建的网页应用。它也非常适合那些想要更图形化界面的人。终端对非技术人员来说可能感觉很陌生。
33:40你的电脑上会出现一堆看起来很吓人的弹窗,而且你没法像在其他几乎所有产品里那样随意点击。所以有很多人就是在终端里感觉不舒服。如果你也是这样,我强烈建议试试桌面版的 Claude Code。桌面版也非常适合快速一览所有正在进行的任务。你可以在桌面版里看到你的 CLI 终端会话。你可以看到你的其他桌面会话。你可以看到你在网页版和移动端启动的会话。所以它是一个一站式的控制台,你可以看到所有任务。我觉得网页版和移动端的好处是它非常适合在外面随时启动任务。CLI 和桌面版都需要你在本地笔记本电脑上操作。这就有局限,因为有时候你在外面比如说出去走走、散个步什么的,没带着笔记本电脑。我数不清见过多少人,在外面的时候笔记本电脑开着,手机热点连着。这就说明我们缺少一个满足这个需求的产品。所以对我来说,移动端能让你在外面也能启动这些任务,这样你就不需要到处带着笔记本电脑还得确保随时随地都开着笔记本电脑。
34:56太有意思了。我见过有人在飞机上,这现在都成了一个梗了,就是"我需要等它完成,让这个 agent 跑完。我不能关掉它。"确实如此。然后我觉得 Cowork 填补的角色是每个人都有很多工作,产出不是代码。比如说把 Slack 消息处理完,或者把收件箱清空,或者做一叠为即将到来的客户会议准备演示文稿,或者写一个简短的文档来说明某个功能的目标是什么,或者某个功能的上线计划,所有这些任务的产出都不是代码,而 Cowork 最适合做这些。所以我对这些产品的划分方式是如果我构建的东西产出是代码,我就用 Claude Code 或桌面版,或者移动端的 Claude Code。如果产出不是代码,我就用 Cowork。大家真的低估了 Cowork 取得的成功。它的增长速度惊人。我觉得大家可能还是不太了解它是做什么的。所以你能不能给我们举几个
18 · 35:58 · Tips for getting started with Cowork
36:00你作为 PM 工作中的用例,有哪些非常有趣、可能出人意料的使用 Cowork 来节省时间、完成更多工作的方式?如果你刚开始使用 Cowork,第一件你真正需要做的事情就是连接所有与你角色相关的数据源。只有当它能获取到所有需要的上下文信息时,才能做好为你整理输出的工作。所以对我来说,就是把它连接到 Google 日历,连接到Slack、Gmail、Google Drive,这样它就知道了,它可以灵活地找到相关的上下文,提出问题,拉取讨论线程,这大大提升了输出结果的质量。我用它做的事情包括[重复内容][重复内容][重复内容][重复内容][重复内容][重复内容][重复内容][重复内容][重复内容][重复内容][重复内容][重复内容][重复内容][重复内容][重复内容][重复内容][重复内容][重复内容][重复内容][重复内容][重复内容][重复内容][重复内容][重复内容][重复内容][重复内容][重复内容][重复内容][重复内容][重复内容][重复内容][重复内容][重复内容]
19 · 38:44 · Demo: Using Cowork to build slide decks overnight
38:44[重复内容][重复内容]- 太慢了。- 而且我很喜欢,人们会看到这叠演示文稿不管你什么时候展示,它都会公之于众。虽然显然不是一版定稿的,但你已经反复迭代过了。所以就是为了帮助大家自己尝试一下。所以第一步是连接他们的,你怎么说的,Slack?你还建议连接什么?- Slack、Google 日历、Gmail、Google Drive。你应该连接你的通讯工具还有你存储团队核心数据的地方,你们团队关注的,你关注的,以及你正在做的事情。- 好的,那你大概输入了什么提示词来生成这叠演示文稿?- 我就写了,给我做一叠演示文稿用于 Code with Claude 大会。这是我们的 PMM 建议应该涵盖的内容。这是目前我自己做的草稿,我不太满意。这是我自己手动做的一版,不太喜欢,但我附上了链接。
39:38你能先创建一个建议的大纲和详细内容吗?还要确保不要和主题演讲有太多重叠,因为那个更重要。然后 Claude 读了我发给它的一堆链接,创建了一个建议大纲。然后我仔细看了它的提案,还有它生成的所有关于我们可以涵盖什么内容的想法。然后我就决定了哪些内容我想要放进最终的演示文稿里。我觉得这就是一个很好的例子,说明了PM 如今的角色是什么。就是说,Claude 是一个很好的头脑风暴伙伴。它能非常快速地综合大量信息,把所有可能性呈现给你。但 PM 的角色依然是做出最终决定:好吧,最终产品里应该包含什么?所以对于这个,我最终决定让这场演讲涵盖从让本地任务成功,到让每个 PR 都变绿,再到帮助工程师提交更多 PR 的发展过程。对于每个阶段,哪个演示然后在决定了大纲之后,Cowork 就自己跑了几个小时。
40:48然后把整叠演示文稿都做好了。太棒了。不用再做这部分工作真是太好了。感觉就像你在跟一个演示文稿设计师说话,而且他还对你做过的项目有实际的了解,能把内容做成你想要的样子,不只是让它看起来好看而已。你是怎么处理设计系统那部分的?那是怎么运作的?它怎么知道 Anthropic 的设计系统的?所以我的做法是,我们其实已经有了一套标准化的演示文稿模板,在我们所有的外部活动中使用。所以我就把这个模板给了 Claude。然后它就能看到我们用什么颜色、什么字体,以及各种不同的——怎么说来着?可用的幻灯片格式。它有 20 张这样的示例幻灯片。举个例子吧。明白了。所以你上传了——这是我们的模板,从这个开始做。对。你也可以连接到你的 Figma MCP,如果你的演示文稿格式保存在那里的话,它可以拉取过来。
41:43你可以把那个导入进来。说到这个,我一直很好奇的是
20 · 41:48 · Cat's PM tech stack and internal tools
41:48你作为 Anthropic 的 PM,你的工具栈都有什么?当然有 Claude Code 和 Cowork 以及所有 Anthropic 的工具。还有什么?你还用什么?其他的——Slack,你提到了。还有别的吗?所以我的工具栈主要就是 Claude Code、Cowork 和 Slack。Anthropic 基本上是在 Slack 上运转的。我觉得它是我们公司的核心操作系统。日常来说,很多——我大概有 30% 的时间是在探索Cowork 和 Claude Code 的能力边界,这样我就非常清楚我们哪些地方做得不够好。而且我花很多时间和模型对话,来理解为什么它会犯这些错误。我们其实做了很多内部工具。我觉得 Claude Code 真正为我们整个公司解锁的一点是,它大大降低了制作任何你想要的自定义应用的门槛。
42:53所以我们看到个性化工作软件的激增,大家都在为特定的使用场景构建定制工具,而不是使用那些不能完美匹配需求的现成工具。我得听听更多。有什么例子?你自己或者其他人做了什么,特别受欢迎和有用的?Claude Code 的一个销售同事,他发现自己在反复制作这些相同的演示文稿。所以他做了一个网页应用,里面包含了 Claude Code 核心演示文稿的模板,我们知道效果很好的那些,比如 101、201 和精通 Claude Code 系列。然后他有一个方式可以输入特定的客户上下文信息,从 Salesforce 拉取,从 Gong 拉取,从其他来源拉取,这样我们就能为特定客户定制演示文稿。所以它会提取信息,比如,好的,这个客户在使用Bedrock 或者 Claude for Enterprise 或者控制台,这会影响他们能使用哪些功能。
43:53它会提取信息,比如,好的,这个客户关注 SDLC 中的代码审查阶段。所以我们就会在那里加一页关于代码审查功能的幻灯片。它会提取信息,比如,好的,这个客户需要符合 HIPAA 合规,或者需要 XYZ 安全控制。所以我们就会确保在他们的演示文稿里加一两页关于这个的内容。然后,例如如果——这是一个使用 Vertex 或 Bedrock 的客户,不想使用 Claude for Enterprise,那我们就会把一些只有 Claude for Enterprise才有功能的幻灯片拿掉。通常来说,这是需要手动做的可能要花 20、30 分钟的工作。所以人们要么花时间去做,要么就直接决定不做了,用通用版本的演示文稿。用了这个,只需要几秒钟,你就能得到一份定制的演示文稿。
44:41MARK MANDEL:有趣的是,像 Slack 这样的工具没有人——就是没有人想去做一个自己的替代品。Slack 持续在赢。你描述它的方式就是很多公司的操作系统。太有意思了。大家谈论 Salesforce 就像是 SaaS 的代表,但我们不再需要 SaaS 软件了。我们要自己做。但 Slack 是一个持久的工具,没有人想跟它竞争或者做一个更好的版本。LILY FIERRO:我觉得它是非常重要的通讯基础设施。而且我觉得它把核心任务做得非常好,帮每个人获取实时更新这一点做得极其出色。MARK MANDEL:对,大家虽然吐槽 Slack,但它在它要做的事情上确实做得很好。而且最前沿的团队都离不开它。真有意思。LILY FIERRO:是的,我也很喜欢他们把定制化做得这么简单。
45:25所以我们很喜欢做 Slack 机器人。而且这种可 hack 的特性意味着我们可以按照自己想要的方式跟 Slack 集成。真的非常感谢 Slack 在这方面的努力。MARK MANDEL:是时候买点 CRM 股票了。我非常激动地向大家介绍本季的赞助商,Vanta。Vanta 帮助超过——15,000 家公司,包括 Cursor、Ramp、Duolingo、Snowflake 和 Atlassian,赢得并向客户证明信任。团队们正在以史无前例的速度构建和发布产品,这要归功于 AI。但结果是,引入到你的产品和业务中的风险量比以往任何时候都高。我交谈过的每一位安全负责人都感受到了保护其组织、业务日益加重的压力,更不用说他们的客户数据了。因为事情发展得太快了,他们一直在被动应对,不得不猜测优先级,不得不用过时的解决方案将就。
46:21Vanta 通过超过 35 个安全和隐私框架自动化合规和风险管理,包括 SOC 2、ISO 27001 和 HIPAA。这帮助公司快速获得合规认证并保持合规。信任比以往任何时候都更有力量决定你企业的成败。了解更多请访问 vanta.com/lenny。作为本播客的听众,你可以享受 Vanta 1,000 美元的优惠。那就是 vanta.com/lenny。
21 · 46:47 · Which teams use the most tokens
46:48好的,你谈到了所有这些不同的团队以及他们如何使用 Claude Code 和 Cowork 来工作。除了工程团队之外,还有哪些团队?我想工程团队应该是最大的 token 消耗者。如果不是的话,那就很有意思了。目前 token 使用量第二的是哪个职能?哦,Applied AI 团队在探索Claude Code 和 Cowork 的能力边界方面做得非常棒。我们 Applied AI 团队很多人花时间跟客户在一起。帮助他们使用我们的 API。所以有时候我们的 Applied AI 团队会,例如,代表这些客户做原型,而 Claude Code 让这比以前快了很多。他们还有另一个目标,需要管理大量的客户沟通、大量的客户反馈,以及历史联系记录、通话记录。
47:37所以他们既大量使用 Cowork,也大量使用 Claude Code。代码。那 Applied AI 具体是做什么的?那是不是类似前线的部署工程那种角色?大多数人会怎么描述Applied AI 团队在做什么?是的。就是帮助我们的客户在公司内部采用最新的 API 和模型功能,既用于驱动他们公司的产品,也用于内部效率提升。明白了。所以有点像客户成功、市场推广,类似于前线部署工程那种角色。没错。就像一个非常技术型的市场人员。明白了。好的,太棒了。所以你是说他们可能是token 使用量第二的团队。对。而且我们也看到他们在不断推进Cowork 能做到的极限。比如,这些人很多都同时负责多个客户。忙的时候一天可能有 5 到 10 个客户会议。所以他们经常用 Cowork 做的是,前一天晚上,会让它总结一下,好的,我明天有哪些客户会议?
48:44后天呢?这个客户之前都跟我提过什么需求?他们最关心什么?之前会议的行动项是什么?然后 Cowork 就会把一份简报整理出来,一份关于他们在进入下次会议前应该了解的信息汇总。而且 Cowork 还能帮忙研究答案。如果客户问了,好的,功能 X 什么时候上线?Cowork 可以帮助这位 PyDI 成员在 Slack 里搜索获取最新的预计时间,添加到——添加到笔记里,这样在客户通话的时候,这位 PyDI 成员就有了绝对最新的信息。这些都是大家自己构建的工作流,然后分享给团队里的其他人。MARK MANDEL:太酷了。有一个——这个问题,这个趋势——我不知道,这个话题最近经常被提到,就是 token 花费超过了人们的工资,大家用 AI,结果费用比他们的工资还高。
49:40有没有什么数据关于这个话题,比如到底工程师花多少 token,比方说,一个月、一天,或者 PM,之类的?JENNY GUY:我们很清楚,随着模型变得更好,人们会把更多的任务交给它,他们在 Claude Code 和 Cowork 这类工具上花的时间也越来越多。所以我们确实看到每个工程师或每个知识工作者的 token 成本在每次模型升级或重大产品改进后都有所增加。JENNY GUY:我觉得——它仍然比工程师的平均工资低很多,但我们看到这个比例在随时间增长。MARK MANDEL:这真是太有趣了——我们聊到了你们如何能使用最前沿的模型,这是在 Anthropic 工作的另一个优势。在 Anthropic 工作的那种方式。我相信你们基本上有无限的 token 可以用。
50:30你们——想用多少就用多少,对吧?JENNY GUY:我们可以用很多 token。有些人确实会遇到限额,所以——MARK MANDEL:好吧,原来有限额。好吧。Boris,关掉它。好的,能使用最先进的模型有这么多优势,真是太有意思了。这就形成了一个非常有趣的飞轮效应。我们也非常相信要赋予我们的内部团队尽可能快的构建能力。而且我们相信每个人都理解运行这些模型到底需要多少成本。而且我们相信团队会负责任地使用 token。所以浪费 token 是很不受待见的,但我们确实信任每个人做出这个判断。我们相信团队可以做出这个判断。
22 · 51:15 · The emerging skills PMs need for AI companies
51:15MARK MANDEL:太棒了。回到 PM 角色的话题,我们之前聊过一些,但我觉得这对听众来说会非常有趣。我想了解的是,你认为PM 需要培养哪些新兴技能,或者说你最看重什么,AI 公司如今在招聘 PM 时最看重什么?JENNY GUY:我觉得最难的技能是能够定义产品一个月后应该是什么样子。我觉得在那个时间范围内模型能做到什么,以及用户行为会怎么变化,都有很多不确定性。但我认为最优秀的 PM 能看到一些模式,基于用户如何"滥用"现有产品边界的模式。最好的 PM 能感知到这一点,能设定方向,能稳步朝目标执行,如果模型能力比预期好得多或差得多,还能调整路径。比他们最初预期的还要糟。我觉得很难把握好"AGI 信仰"的程度。所以我觉得每个人都能看到那个模型极其聪明、几乎什么都能做的未来,在这种情况下你其实不需要那么复杂的产品。
52:31你其实只需要一个文本框,告诉模型你想要什么。它聪明到可以自动添加任何工具或任何它需要的集成来完成任务。它知道什么时候不确定。它可以问澄清问题。为超级 AGI 强模型构建产品其实很容易。我觉得难的是,针对当前的模型,你如何激发它的最大能力?你如何帮用户找到最佳路径?你如何引导用户与模型的优势互动,同时弥补它的弱点?这种技能非常稀缺。那你怎么培养这种技能?基本上就是理解每个模型的限制吗?你说的是品味吗?理解,对模型可能的能力有品味,知道它擅长什么、不擅长什么,哪里有变化?我觉得就是花大量时间跟模型对话和使用模型。我特别喜欢做的一件事就是让模型反思自己的行为。所以,有时候当我发现模型做了意想不到的事情,比如,有些情况下模型会做前端修改然后跑测试,但实际上并没有使用 UI,这时候让模型反思一下为什么这么做其实很有用。
54:01有时候它们会说,"嘿,系统提示里有一些令人困惑的内容,"或者,"我没意识到前端验证是这项任务的一部分,"或者,"嘿,我把验证委托给了子 agent,但子 agent 没有做测试,我也没有检查它的工作。"很多时候,对模型为什么做出那个决定保持好奇心,就能让你看到是什么误导了它,这样你就可以修复引导机制来弥补这个差距。另一件有帮助的事是找出你最信任的那些用户,让他们给你关于模型的准确反馈。通常有那么几个人比其他人更善于表达是什么让某个特定的模型或模型+工具组合好用。很多人会给你反馈,但不是每个人的反馈都同样有价值。所以,找到你信任的那五个人,对于获取快速反馈非常重要。
23 · 55:00 · Why building evals is underappreciated
55:04我觉得第三件有用但不是所有人都喜欢做的事是构建 eval。你不需要构建几百个 eval 才有用。只要构建 10 个好的 eval,就能帮助团队量化目标是什么,他们离目标有多远,以及还缺什么。所以我觉得 eval 是一种被低估的东西,更多的 PM 和工程师应该去做。我们之前聊过很多关于 eval 的话题。有一种趋势就是,"产品管理的未来就是写 eval,"因为本质上,它让我们看起来就像,"好的,很酷。让我具体定义一下,然后我们就知道了。"你觉得你花多少时间在写 eval 上?我觉得 eval 的重要性取决于你正在做的功能或者你试图解决的问题是什么。所以,我们团队有很多人确实花了很多时间在 eval 上。我们有一个小组跟研究团队紧密合作,更精确地了解我们 Claude Code 的行为,以及最大的改进空间在哪里,并且努力非常具体地衡量这些。
56:30我个人会在某个功能需要更多产品定义的时候跳进来做 eval。通常的产出是,"好的,这是我做的五个 eval。这是运行方法。不过差异很大。取决于具体的功能。不是每个功能都需要,但我觉得像记忆这样的功能从中获益很大。你说的关于人们非常擅长评估模型这一点太有趣了。几乎就像一个人肉 eval,就是"好的,他们知道哪里表现出色,哪里可能不足。"有没有什么特定的人你想提一下,特别擅长这个的?我觉得在这方面非常厉害的两个人,一个是 Amanda,她塑造了 Claude 的性格。这个角色真的很难,因为任务太模糊了。即使是编码都更容易,因为你可以验证成功与否,而塑造性格需要对 Claude 应该成为什么样的人有非常强的信念。我觉得她不仅有能力塑造性格,还能非常清晰地表达目标是什么,性格应该是什么样的,什么是成功的,什么是不成功的。
57:35我觉得她不仅有能力塑造性格,还能非常清晰地表达目标是什么,性格应该是什么样的,什么是成功的,什么是不成功的。我觉得她不仅有能力塑造性格,还能非常清晰地表达目标是什么,性格应该是什么样的,什么是成功的,什么是不成功的。我觉得她不仅有能力塑造性格,还能非常清晰地表达目标是什么,性格应该是什么样的,什么是成功的,什么是不成功的。我觉得她不仅有能力塑造性格,还能非常清晰地表达性格应该是什么样的,什么是成功的,什么是不成功的。另一群我非常信任的人就是 Claude Code 团队。所以我们经常有团队午餐,每当有新模型在测试时,我们获取反馈最快的方式之一就是在这些午餐上,走到每个人面前问,"嘿,你对这个模型感觉怎么样?"通常我们会得到这样的反馈,"好的,这个模型好像没有充分解释它的思考过程。
58:06太唐突了。"或者,"嘿,这个模型就是喜欢写一堆记忆,但我们不确定这些记忆质量高不高。"或者有人会注意到,好的,这个模型喜欢自己测试自己,这很好。或者这个模型自我测试不够。这告诉我们应该去看什么数据来验证,好的,这是不是一个更大的模式?我们有大量数据,但很难从中提取洞察。所以这个群体的反馈就是,"嘿,这个模型喜欢写一堆记忆,但我们不确定这些记忆质量高不高。"我们有大量数据,但很难从中提取洞察。所以这个群体的反馈就像,"嘿,这个模型喜欢写一堆记忆,但我们不确定这些记忆质量高不高。"我们有大量数据,但很难从中提取洞察。所以这个群体的反馈帮助我们明确,好的,我们要测试哪些假设?然后我们就能提取数据来验证。[重复内容][重复内容][重复内容][重复内容]
24 · 58:44 · Why Claude's character and personality matter so much
58:44[重复内容]你提到的关于 Claude 性格这一点,我之前请过联合创始人 Ben Mann 来做播客,他谈到过这个,就是 Claude 的性格和特质是 Claude 非常重要的一部分。[重复内容][重复内容][重复内容][重复内容][重复内容][重复内容][重复内容][重复内容][重复内容][重复内容][重复内容][重复内容][重复内容][重复内容][重复内容][重复内容][重复内容][重复内容][重复内容][重复内容][重复内容][重复内容][重复内容][重复内容][重复内容][重复内容][重复内容][重复内容][重复内容][重复内容][重复内容][重复内容][重复内容][重复内容][重复内容][重复内容][重复内容][重复内容][重复内容][重复内容][重复内容][重复内容][重复内容][重复内容][重复内容][重复内容][重复内容][重复内容][重复内容][重复内容][重复内容][重复内容][重复内容][重复内容][重复内容][重复内容][重复内容][重复内容][重复内容][重复内容][重复内容][重复内容][重复内容][重复内容][重复内容][重复内容]大家真的很喜欢 Claude 的低调。
1:00:03所以如果你告诉它,嘿,你做错了这件事。它会真心道歉。就像,哎呀。就像,谢谢你告诉我,让我来修一下。我们一起解决。而且它也非常积极。所以如果你觉得,哦,这任务太难了。我不知道怎么开始。Claude 就会说,好的,没关系。我觉得我们应该按这些步骤来做。要不要我先开始帮你做?我觉得一个好同事的特质就是这种积极性,这种行动导向,这种能给你真诚反馈的能力,而不是你说什么就同意什么。所以我们努力把这种特质融入到 Claude 中,因为我们觉得
25 · 1:00:44 · How new models force product changes
1:00:44这样会让跟它合作变得更加愉快。有个话题我想回来聊聊。你谈到每当新模型出来的时候,你经常需要重新审视之前构建的东西。太有意思了。而且可能还有点沮丧,就像,天啊,我们已经发布了这个东西。现在又得重新思考。聊聊大概多久会发生一次。你需要回过头来处理新模型,然后说,好的,我们要重做几个月前发布的产品。很多时候,新模型带来的变化是移除不再需要的功能。很多时候我们给产品添加功能是为了弥补模型的不足,因为它本身做不到。典型的例子就是待办事项列表。我们刚推出 Claude Code 的时候,人们会要求它做大型重构,Claude Code 就会说,好的,没问题。我需要——然后它要修改 20 个调用点,结果只改了 5 个就停了。然后我们就在想,怎么才能让它记住要改完所有 20 个?
1:01:39所以我们团队的 Sid 就说,好吧,我们想想人类会怎么做?人类会先把所有需要修改的东西列一个清单。就像在 VS Code 里,你会查找所有调用点,左边会显示一个列表,然后你一个个去替换。我们怎么给 Claude 这种工具呢?所以他就添加了待办事项列表。我们发现有了这个之后,Claude 确实能够修改完所有 20 个调用点,但后来有了 Opus 4 和更新的模型,我们意识到不需要再强制它使用待办事项列表了。对早期的模型来说,它会自己用。我们需要不断提醒它。嘿,待办事项都做完了吗?你得把待办事项上所有的事都做完才行。而对于后来的模型,不需要提示,它就自然地会把待办事项上的所有事情都做完,现在待办事项对用户来说还是有用的,因为你可以更清楚地看到 Claude在做什么,但说实话,它现在已经是产品中非常弱化的一部分了,模型可能会用,也可能会不用。
1:02:40它已经不再是做出彻底修改所必需的了。我忘了是谁在播客上说的,模型会把你的脚手架当早餐吃掉。我在这里听到的基本上就是你——你随着时间推移移除了那些不得不加在模型之上的东西,因为它之前没有按你想要的方式运作。基本上随着模型变得更聪明,它就变得越来越简单,就能直接做你想让它做的事了。是的。每次模型变聪明的时候,我们可以移除很多提示干预。模型变得更聪明的时候。其实每次发布模型我们都会做这件事,我们会通读整个系统提示,然后反思,好的,对于每个部分,模型还需要这个提醒吗?如果不需要了,我们就把它移除。不过新模型最令人兴奋的是全新功能。有很多功能我们之前用旧模型测试过,但准确率不够高,我们不想发布。其中一个例子就是代码审查。我们尝试过几次构建代码审查产品,也发布过一些简单版本的代码审查,就是过去的斜杠代码审查命令。
1:03:50这个代码审查太好了,以至于我们的工程团队在合并 PR 之前都要依赖代码审查通过。我们发现,我们一直梦想着 Claude 能成为一个可靠的代码审查者,能够——我们有信心它能发现大部分的 bug。直到 Opus 4.5、4.6 和 Sonnet 4.6,我们才觉得,好的,我们现在可以同时运行多个代码审查 agent,遍历整个代码库,然后综合出工程师在合并前需要处理的一系列真实问题。所以这就是最新的模型解锁的一项新能力。这是本播客上非常常见的一个趋势:构建那些在未来六个月内可能实现的东西,稍微处于可行性的边缘,然后模型会追上来,产品就会变得非常好,而你已经领先所有人了。没错。构建那些暂时还不能用的产品很重要,这样你就知道,好的,这个产品要能用还缺什么。
1:05:09然后有了最新的模型,你可以直接把它
26 · 1:05:11 · The vision for Claude Code and Cowork
1:05:12换进你已经做好的原型里,看看,新模型是否填补了那个差距。你能在多大程度上谈谈 Claude 和 Cowork 的发展方向和愿景?我想你不想透露太多目标,但感觉你——有所有这些很棒的功能在上面叠加,调度、手机控制,还有移动应用,所有这些,怎么理解这些产品长期的整体愿景?我们用构建模块的方式来思考这件事。对于 Claude Code 和 Cowork,核心构建模块就是让单个任务成功。所以你想要它产生某个输出。你给它一个清晰的提示描述。它能持续产出你能接受的输出吗?你可以合并,或者分享给同事,或者给外部用户?所以任务是核心构建模块。随着模型变得更聪明,任务成功率大幅提高。然后我们看到人们开始同时做多个任务。所以多 Claude 并行是 2025 年底的一大趋势,而且此后只增不减。
1:06:17所以我们把这看作,好的,很好。单个任务搞定了,现在你可以同时跑六个任务。随着模型变得更聪明,我们的推演是,好的,接下来,也许你会同时运行 50 个 Claude,或者几百个 Claude。那么我们需要构建什么基础设施来支持这个?到那时,你可能不会在本地机器上运行所有东西了。就是没有足够的内存来做这件事。所以我们正在思考如何让你更轻松地管理所有这些?这些可能会远程运行。我们怎么构建界面,让你作为人类知道哪些任务需要查看,我们怎么确保 agent 完整地验证了它的成果,这样当你看到一个任务显示完成时,你能很快验证并且完全信任它达到了你的要求。我们怎么确保这个过程是自我改进的,这样当你确实看到一个不符合你要求的任务时,你可以给它反馈,
27 · 1:07:22 · Advice for thriving in an AI-driven world
1:07:23模型会在未来的每次运行中 incorporate 那个反馈。这样它就再也不会犯同样的错误了。这就是我们带用户一起前进的路径。有很多——听众里有很多产品经理、很多创始人、很多其他跨职能的人。很多人担心自己的角色、职业生涯的未来,你对这些人有什么建议,不只是在这个 AI 驱动的世界里生存下来,而是真正成功,在这个未来中蓬勃发展?有什么大家需要听到、需要去做的?我觉得 AI 给了每个人比以前多得多的杠杆。所以我建议你,每当你发现自己在反复做某个手动任务时,想想怎么用 Claude Code、Cowork 或其他 AI 工具来自动化它。大多数人都有工作中非常喜欢的创造性部分。然后还有一些非常讨厌的繁琐部分。我觉得 AI 的美妙之处就在于它可以帮你做那些繁琐的部分。
1:08:53它可以从你每次做的过程中学习。它可以做手动任务,总结规律,然后自动运行,这样你就可以专注于创造性部分。这意味着你能做到比以前多得多的事情。
28 · 1:09:18 · Why 95% automation isn't good enough
1:09:19所以我对大家的直接建议是,找出可以交给 Claude 的重复性工作,迭代这些自动化,直到成功率非常高,然后专注于,好的,你还能为你的团队、你的产品、你的公司做什么,那些人们一直没精力去做的事情?比如那个你一直觉得公司应该做的 pet project,但你从来没时间做?你可以自动化的,关注那些一直想去做但没时间的想法。基本上就是为自己解决一个问题,这是核心建议。没错。我也想鼓励听众,把你的自动化从,好的,这是一个好工具,这是一个很酷的概念,变成,嘿,这个真的百分之百能用了。有时候我看到用户尝试自动化某件事,做到 90%、95% 的准确率,然后就放弃了。如果一个自动化不能百分之百工作,它就不算是真正的自动化。而最后那 5% 到 10% 确实需要更多时间。
1:10:15而且构建自动化往往比你亲自动手慢很多。我鼓励大家投入那个时间。投入时间去定义你真正想要做到百分之百的自动化,下功夫教 Claude 你的偏好,给它反馈,让它提高技能,直到达到百分之百。然后你就真的能依赖它了。一个 95% 的自动化其实没多大价值。我太有这个毛病了。这个建议对我太有用了。我也有这个毛病。我一直在教它。我一直在教 Cowork——试着帮我达到 Gmail 收件箱归零,但效果一直——非常耗时,而且肯定还没做到,你可能也发现了。是的,有趣的是。我想到的也是这个。我有一个工作流,每封邮件进来后,它会找出那些垃圾性质的邮件,就是那些,"嘿,我能上你的播客吗?"或者"要不要做个广告?"之类的,我就是没时间处理这些,所以我让它把这些分到一个叫"垃圾邮件"的文件夹里。
1:11:22而且它——95% 做得很好。但有时候就会,哦,天哪。我错过了一封邮件,因为它被分到那里去了。所以这对我来说是一个很好的推动,我要把它做好。我要把它做到完美。是的。我们也在努力让定制这些命令的流程变得更加简单。因为现在我觉得你需要了解太多概念了。你需要知道怎么定义一个 skill。你需要知道怎么使用这个 skill,还要给它反馈。然后你还需要知道告诉 Cowork 基于你给的所有反馈来更新这个 skill。然后你还需要知道去哪里读这个 skill,确保反馈被按照你的要求
29 · 1:11:58 · Build apps you use every day, not prototypes
1:11:58纳入了。这也是我们的工作,让这个流程变得非常顺畅,不会让人觉得很痛苦。太棒了。还有什么想分享的吗,Cat?还有什么想留给听众的?有什么想进一步强调的,在我们进入非常精彩的闪电问答之前?我看到很多人在玩 AI,构建一些原型应用,摆弄各种工作流。我真的很建议大家去构建你真正在用的应用。每天在用的那种,因为我觉得只有通过这种使用,你才能真正获得价值。如果你构建了一个原型应用但并不能帮你完成更多工作,那 AI 并没有真正为你的日常生活增加价值。你只做了一次就觉得,好的,我就试了一次。哦,挺酷的。然后你再也不碰它了。你学不到多少东西,也得不到什么真正的杠杆。真正的杠杆。是的。这个观点太好了。我也觉得有很多人花了很多时间在定制他们的工作流上。所以就像,我觉得有两端。一端是从来不定制、从来不构建自动化的人,但还有另一端的人痴迷于定制他们的工具,比如加一堆 skill 和 MCP,还有这些工作流改进。
1:13:30我觉得有时候这甚至会分散你做核心目标的注意力,
30 · 1:13:41 · The divide between AI skeptics and believers
1:13:41比如发布某个产品或构建某个功能。定制化确实很有趣,而且我们当然希望我们的产品非常可 hack,这样你可以让它很好地为你工作,但它的实用性是有限度的。我觉得有一类人可能花太多时间在定制上了,以至于没在睡觉,也没在做他们最初想做的核心任务。我在 Twitter 上看到很多这样的。就是那种,"看看我的设置。""简直失控了。""优化到极致了。"然后你问,你到底在构建什么?"不是,但我的设置太棒了。""我能做超多事。"我觉得简单的设置实际上效果更好。就是稍微提升一下就好。是的。是的。Karpathy 昨天发了一条推文,他谈到了这种分化。很有意思,就是当年试过 ChatGPT/Claude 的人,觉得,好的。然后说,"不,这太差了。"两边的人都不理解对方,不理解对方为什么那样看待世界。
1:14:32所以你的建议在这里真的很好。就是真正去用它做实际的事情,看看它到底变得多好了。是的。我觉得最大的转变是,2024 年那一代产品是基于对话的,而 Claude Code 这一代产品是基于行动的。所以大家的大 aha 时刻就是当 Claude 可以替你做事的时候。那种感觉非常奇妙,知道 agent 能做到的远不止是告诉你该怎么做。agent 可以真的自己去做。当人们感受到这一点时,我觉得那就是豁然开朗的时刻。顺便提一下 Chrome 扩展,Claude 的 Chrome 扩展,你可以看着它在做事,比如你让它帮你填个表格。它就会说,好的,我开始了。没错。
31 · 1:15:19 · Lightning round
1:15:19好的。在我们进入非常精彩的闪电问答之前,还有什么吗?没有了,开始吧。开始吧。好的,Cat,我有五个问题要问你。欢迎来到闪电问答。有个动画——我得说一下。你准备好了吗?我准备好了。第一个问题。你经常推荐给别人的两三本书是什么?我很喜欢《How Asia Works》。它是关于经济发展的故事,以及什么样的政策、政府能创造长期的——那个——持久的、成功的经济体。我喜欢的另一本书是《The Technology Trap》。它其实是关于过去几次技术革命的,工业革命和计算机革命,以及这些如何影响了工人。我喜欢它的原因是我觉得我们可以从历史中学到很多,来确保这次的转型进展顺利。也许轻松一点的推荐,我很喜欢《Paper Menagerie》。它就是一本短篇故事集,关于成长、AI 和自我发现。最喜欢的近期电影或电视节目?我很喜欢《Drive to Survive》。没有什么深层含义。
1:16:39我就是觉得那些人对一个单一的工程目标如此痴迷,追求的纯粹性,让人非常满足。我还很喜欢《Free Solo》,关于 Alex Honnold 不带安全绳攀登 El Capitan 的故事。我觉得类似地,能够攀爬这条极其挑战性、危险的路线,还能保持那种精神专注力,知道犯一个错误就会死——这太纯粹了。太疯狂了。是的,那部电影真的太震撼了。有趣的是这些在某种程度上跟你做的工作也有关联。我其实是个攀岩爱好者。我在攀岩之前就看了《Free Solo》。所以当时觉得很厉害,但不知道到底有多厉害。这是那种罕见的电影,你越了解,就越被震撼到。像他在墙上做的那些动作,我觉得我这辈子在攀岩馆里、离地一英尺,都做不到。还系着绳子。系着绳子。你看了那个关于另一个更年轻的攀冰者的纪录片吗?
1:17:49我看了。那个很令人难过。但那个真的很疯狂。好的。你最近发现并非常喜欢的产品?除了 Claude 相关产品之外,对我的生活改变最大的产品可能就是 Waymo 了。我是一个铁杆 Waymo 用户。每天用两次,上下班通勤。是的。我喜欢它的两点,一是如果 Waymo 在等我,我不会觉得不好意思。所以我觉得不用那么有压力非得在它到达的那一刻就站在路边。第二是我觉得它让我更高效了。当我和另一个人在车里时,我通常不会打工作电话。我觉得如果在车上一直用笔记本电脑有点不礼貌。但 Waymo 的好处是我可以接工作电话。不用担心有人偷听。不用担心,嘿,这样会不会很没礼貌?我是不是说话太大声了?要不要让人家换个音乐?所以这真的——我觉得每天帮我省了 30 分钟。
1:18:47技术的这些二阶效应。太有意思了。是的。我一直以为 Waymo 需要定价比 Uber 和 Lyft 低才能成功,但实际上我非常乐意付两倍的溢价。我爱 Waymo。就是——你一旦体验过,你就会觉得,啊,这太——太疯狂了。然后你就习惯了。你坐进去的时候会觉得,太不可思议了。然后你就忘了这种感觉了。完全正确。而且我觉得它也改变了语言习惯。Anthropic 很多人喜欢 Waymo。过去你可能说,嘿,叫个某某打车软件。现在大家就直接说,好的,Waymo 到了吗?好的。还有两个问题。你有最喜欢的工作或生活中的座右铭吗?去做就好。我觉得第一性原理思考非常有价值。如果你知道你在优化什么,而且你有很强的第一性原理,那你通常能推导出正确的行动方案,并且能向所有利益相关者清楚地表达。
1:19:46然后你就应该去做。我觉得工作头衔是假的。如果你理解了约束条件,你就能弄清楚你能做什么,然后就去做,快速学习,如果做错了就道歉或修正。比如快速去做、从错误里学习,做错了就道歉或者修复做错了。你就可以去做事情。谁说的来着。我觉得这其实是一种解放。告诉大家这个,我觉得很多公司的角色定义非常严格,好的,PM 做这个,设计师做这个,工程师做这个。然后连团队范围都定义得很死板。所以,嘿,代码库的这个角落我们碰,那个角落我们不能碰。我觉得"去做就好"让大家觉得自己有权力做这些决定,有权力跨团队操作,就是为了把事情做好。这感觉是一项非常重要的技能。要擅长的,人们称之为"主动性"。就是去做需要做的事。行动导向。行动导向。所有这些描述方式都是在说,不要等许可。
1:20:50是的。我觉得这是在人生某个阶段去创业公司工作的最好理由,因为对我来说非常改变人生的一件事就是在 Scale 只有 20 个人的时候工作。那时候完全没有流程,但我们需要解决非常大的问题。我真的很感谢 Alex 和团队其他人给了我这么大的自主权。让我和团队其他人可以不受边界限制地去解决问题,不管销售应该做什么,运营应该做什么,工程师应该做什么。就像你拥有所有可用的工具,面对一些雄心勃勃的棘手问题,你可以做任何需要做的事来找到好的解决方案。你几乎需要那种经历来培养那种技能,让自己习惯这样做。因为很多人,你知道的,从学校或大学一路走来,都是"做我们让你做的事,然后你就会得到好成绩"。你需要忘掉这种思维,就像,好的,我就去做需要做的事。即使别人觉得这很蠢,但我觉得这是对的事情。是的,没错。
1:21:47好的。其实我还有两个快问快答。最后两个问题。一个是,Claude 会有那些——我不知道你们叫不叫动词,这些东西叫什么?思考词。思考词。有趣的是,这些词在源代码里泄露了。你有最喜欢的思考词吗?我非常喜欢"manifesting"。这也是我笔记本电脑上的贴纸。哦,太棒了。显然是赢家。好的。最后一个问题,AGI 有可能在我们有生之年到来,当你不用工作的时候,你会做什么?你会用你所有的时间做什么?我觉得 AGI 扩散到整个社会还需要很长时间。所以我觉得眼前的其实是帮助世界跟上步伐。我的不太认真的回答是,那之后我可能就是去攀岩。我可能就搬到Fontainebleau,住在那一万块巨石中间,爬一阵子。还有好多书想读,目标是每周能读一到两本书。现在大概只能读 0.5 本。积压的量很大。我觉得历史有太多我们可以学习的,太多我还不够了解但很想深入了解的。
1:23:11物理学,或者机器人学,或者硬件,或者航空航天——有太多有趣的话题了。所以我很期待去学习,即使知道 AGI已经都知道了。Kat,太棒了。你太厉害了。两个后续问题。大家在网上哪里可以找到你,如果想联系你或关注你的动态?听众怎么帮到你?最好的联系方式是我在 Twitter 上的账号:_KatWu。可以在推文里 tag 我,也可以给我发私信。我会读所有私信。虽然不一定会每条都回复,但我都会看。然后最有帮助的是告诉我们Claude Code 和 Cowork 在哪里对你不好用。我们非常感激大量的正面反馈。但我们最需要的是边界情况、错误,那些我们能复现的具体任务,Claude Code 或 Cowork 在哪里失败的。因为如果你能分享给我们,我们能复现,那这就是我们能主动改进的,为下一代模型和下一代工具。太酷了。Twitter 上的人分享这些反馈一点都不害羞。
1:24:28所以继续来吧。分享,分享。请,请把你们遇到的问题分享给我们。是的。而且很酷的是你们整个团队在 Twitter 上这么活跃,回复大家。所以——我听到的就是,这些确实是你们真正会看到并做出反应的东西。是的。我们感谢大家的积极参与。这给团队带来了很多能量。我们有一个"用户之爱"频道。所以每当你们分享成功故事,我们就会发到那里。每当你们分享产品的问题,我们就会放进反馈频道。这样我们更大的团队就能据此行动了。知道这个太酷了。谢谢分享。Kat,非常感谢你来。谢谢你的邀请。大家再见。非常感谢收听。如果你觉得有价值,可以在Apple Podcasts、Spotify 或你最喜欢的播客应用上订阅。也请考虑给我们评分或留下评论,因为这真的能帮助其他听众找到这个播客。你可以在Lenny's Podcast.com 找到所有往期节目或了解更多关于这个节目的信息。下期再见。
Cat Wu, Head of Product for Claude Code and Cowork at Anthropic, reveals how the team ships features at an unprecedented pace — timelines collapsed from six months to days.
She explains why product taste is now the most valuable skill as code becomes commoditized, how Anthropic's mission alignment eliminates organizational friction, and the practical philosophy of building for current model capabilities rather than hypothetical AGI.
Key themes include using Research Preview to ship fast, asking models to introspect on their own mistakes, treating Claude's personality as a core product feature, and pushing automations to 100% rather than settling for 95%.
Anthropic Claude Code 和 Cowork 产品负责人 Cat Wu 揭示了团队如何以前所未有的速度发布功能——时间线从六个月压缩到几天。她解释了为什么在代码日益廉价化的今天,产品品味(product taste)成为最稀缺的能力,Anthropic 的使命认同如何消除组织内耗,以及基于当前模型能力而非假设性 AGI 来构建产品的务实哲学。核心主题包括:用 Research Preview 加速发布、让模型反思自身错误、将 Claude 的性格视为核心产品特性,以及把自动化推到 100% 而非止步于 95%。