This chapter covers what AI is doing to product work upstream of the code: the shift of the binding constraint from execution capacity to judgment, LinkedIn's full-stack builder reorganisation and the platform, agent, and culture machinery it rests on, the predicted split of the PM population into judgment-heavy and execution-heavy halves, and the practical toolkit PMs are actually running today. It closes with the guardrails practitioners insist on — where judgment must stay human — and with the smaller body of material on designing AI products themselves. The material here is the least settled in the corpus and it is heavily weighted toward one company's account; the chapter says so where that matters.
The single claim that organises everything else in this chapter is that the scarce resource in product work has moved. When AI-assisted tools make building fast and cheap — Lovable, coding agents — a product's surface can be cloned in a weekend, so what's left to compete on is positioning, product sense, and market-specific experience accumulated over time, not the stack underneath. This is AI Era Bottleneck Shift: From Building to Deciding, reported by Emily Tate (VP Product) in a conference-recap interview on the state of product in 2026, and its sharpest consequence is that it inverts the founding cliché of startup lore: if ideas were cheap because execution was everything, and execution has now collapsed in cost, then the idea-and-decision layer becomes the constraint again.
Tomer Cohen, LinkedIn's CPO, arrives at the same place from the other direction by naming what AI still can't do: vision, empathy, communication, creativity, and — in his words the single most important — judgment, which he characterises as "test-making" under ambiguity. Everything else, he says, LinkedIn is actively trying to automate. Cohen also supplies the diagnostic frame for why this feels like an emergency inside companies rather than a gradual drift: Time Constant of Change vs. Time Constant of Response distinguishes how fast the outside world is moving from how fast an organisation can actually adapt, and locates the transformation problem in the gap between the two. The practical use is discipline about which side is failing — a slow org in a slow market is a different problem from a slow org in a fast one — and the corollary that a gap opened by acceleration can't be closed by buying tools, only by changing how work is organised.
Two cautions belong at the foundation rather than at the end. First, AI as Amplifier of Existing Discovery Culture: AI does not repair a broken discovery process, it accelerates whatever culture already exists. Teams that run real experimentation loops validate faster; teams without that muscle reach the wrong answer sooner. The same dynamic shows up at the individual level — an MIT materials-science study is cited for the finding that the people who benefit most from AI tools are those who already have strong judgment, which implies AI widens rather than narrows the gap between strong and weak PMs, and that coaching investment should be planned accordingly. Second, on how permanent any of this is: Tom Verrilli (Whatnot CPO) borrows Elizabeth Stone's storming-before-norming framing in Storming vs. Norming: AI's Upheaval in Product Roles (Elizabeth Stone) to argue that today's churn in role definitions and team shapes is a phase, not the industry's new resting shape. His evidence for the storming phase is senior technical leaders — CTOs from Workday, Instagram, Box, and super.com — taking individual-contributor roles at Anthropic, read as seniority flowing back toward hands-on work. He pairs it with an explicit disclaimer: he doesn't think there's one way to do product management, or one way AI will shape the industry.
The most fully worked organisational response in this chapter is LinkedIn's, and it starts from a prediction Cohen quotes: the skills required for any job will change roughly 70% by 2030 — "whether or not you're looking to change your job, your job is changing." His answer is the LinkedIn's Full-Stack Builder Model, in which any builder, regardless of role or team, can take an idea from research through launch themselves, in what he calls a fluid interaction between human and machine rather than a rigid sequence of handoffs. The diagnosis behind it is worth stating precisely, because it's a claim about process rather than about difficulty: at LinkedIn's scale the traditional lifecycle (research → spec → design → code → launch → iterate) had fragmented into dozens of sub-steps, multiple review types, and micro-specialised sub-roles — design alone split into interaction, animation, and content design. Cohen's position is that the work itself isn't complex, the process was made complex, and AI removes enough coordination overhead that the pipeline can collapse back toward individuals or small pods.
The structural instantiations are concrete. Pods (Full-Stack Builder Teams) are small mission-focused units — an engineer, a designer, a PM — that own one problem for roughly a quarter and then disband and re-form around a new one, modelled explicitly on Navy SEAL small-unit doctrine: members are cross-trained rather than each holding a single narrow specialty. That temporariness is the distinguishing feature; these are not long-lived teams aligned to a persistent mission area. The career ladder now carries a formal full-stack builder title, and System Builders vs. Full-Stack Builders vs. Specialists names three legitimate destinations rather than one: system builders who make the platforms and agents everyone else stands on, full-stack builders who go idea-to-launch, and specialists who stay deep in a craft. Cohen is explicit that the third bucket isn't a failure state — "some people do not want to be full stack builders. And that's completely okay" — which makes the taxonomy a deliberate hedge against a one-size-fits-all rollout.
The change reaches both ends of the ladder. At entry level, LinkedIn scrapped its APM program for the Associate Full-Stack Builder (APB) Program, slated to launch in January, where recruits learn to code, design, and do PM work before being placed into a pod — training generalists from day one instead of hiring people who already specialised elsewhere. At the top, Product-Area Leadership + 360 Review replaces functional heads (design, PM, BD) with product-area leaders who own an area end-to-end and are evaluated through 360 reviews collected across every function they'd need to oversee. The rationale generalises: if you restructure the builder layer but leave functional leadership intact above it, those leaders keep pulling reports back into the old silos.
Two secondary observations sharpen the picture. Cohen argues in Design as the Hardest Craft for AI to Replicate — and Designers' Cross-Training Advantage that design is the hardest craft for agents to do well — taste and interaction judgment resist automation more than code generation or research synthesis — with the paradoxical result that designers cross-training into code and PM move into full-stack roles faster than engineers or PMs cross-training into design, because the skills they need to add are exactly the ones AI is compressing. And the internal proof point Cohen uses to sell the model to sceptics is a career story: a user researcher applied for a long-open growth-PM role, said "I feel I can do it," used the new tools to make the jump, and got it. His stated hiring philosophy: "I could care less about your title. I care about how you work." The personal mirror of the same stance is Becoming Is Better Than Being, a mantra Cohen applies at home — judge this year's version of yourself against last year's rather than against a fixed competence bar.
None of the org design above works without a substrate, and Cohen's framework for building one is Platform / Tools / Culture: A Three-Pillar AI Transformation Framework: re-architect the codebase and design system so AI can reason over them (platform); build and heavily customise agents for internal use (tools); and work on incentives, motivation, and visible examples (culture), which he calls the biggest and most important lever. His advice to companies attempting the same shift is specific — invest in platform and tools but put the most effort into culture, be ambitious and impatient about the destination while patient about the path, don't expect fast productivity gains (not 2x in a week), and don't wait for a top-down mandate before starting.
The most transferable technical finding is deflationary. LinkedIn's biggest learning, recorded in Off-the-Shelf AI Tools Need a Custom Integration Layer, is that no third-party AI tool — coding agents, design tools, knowledge tools — worked off the shelf against its own stack; every category needed a custom internal integration layer before it was usable at scale. What LinkedIn then built is Narrow, Single-Purpose Internal Agent Suite (LinkedIn): not one general assistant but a portfolio of single-job agents, each trained on a specific slice of internal data. A trust agent, built by the head of trust, reviews specs for vulnerabilities and harm vectors like scam exposure. A growth agent, trained on LinkedIn's own growth loops, funnels, and past experiment history, critiques the quality of an idea rather than its execution. A research agent, trained on member personas plus past research and support tickets, lets teams pressure-test a spec from a user's point of view without waiting on the research team. An analyst agent provides natural-language querying over the internal data graph. A coding agent handles implementation, and a paired maintenance and QA agent auto-fixes roughly half of failed builds. Cohen offers two adoption data points as evidence this isn't pilot theatre: the UXR team picked up the growth-prioritisation agent to rank member-facing ideas — lateral adoption the agent wasn't designed for — and running the old "Open to Work" green-badge spec retroactively through the trust agent surfaced scam vulnerabilities that weren't caught until well after launch.
The technique Cohen rates highest is not a model choice. Golden-Examples Curation Beats Blanket Knowledge-Base Access records that pointing an agent at an entire shared drive or knowledge base "failed miserably and hallucinates like crazy," because blanket access makes the agent weight unimportant material as if it were load-bearing. The fix was stricter curation — a hand-picked golden-example set — and Cohen traces the principle back over a decade to manually curating examples of good LinkedIn posts while rebuilding the feed team. His summary: "the first and most important part was fitting in the right data, not all the data." The planned next step is Orchestrator Layer for Narrow Agents, letting the narrow agents call each other rather than being invoked one at a time; the sequencing is deliberate, since orchestrating before each agent is individually reliable compounds one agent's errors across the chain. Its flagship instance is the Product Jammer Agent, which wraps LinkedIn's existing internal product-jam ritual behind one front-end agent that invisibly calls trust, growth, and research on the builder's behalf.
Three smaller pieces round out the machinery. Idea-to-Design vs. Code-to-Launch (AI Investment Gap) audits where agent investment has actually gone: code-to-launch (coding, review, maintenance agents) has absorbed the bulk of it industry-wide, while idea-to-design — research, ideation, design exploration — is comparatively under-invested despite being harder to automate; Cohen treats the asymmetry as a gap to close on purpose. On procurement, No-Winner-Take-All Tool Piloting (Parallel Vendor Evaluation) describes running Figma, Subframe, and Magic Patterns in parallel rather than settling on one, since different teams gravitate to different tools — "I don't think there's going to be a winner takes all." And for measurement, Experimentation Velocity Formula: (Volume × Quality) ÷ Time-to-Launch gives progress as (experimentation volume × quality) ÷ time-to-launch, so any of the three terms improving counts as movement without needing a single north-star metric; early wins cited are a few hours saved per team per week and designers and PMs picking up bugs from Jira and submitting the pull requests themselves. Looking further out, Bandits (Multi-Armed / Contextual / Combinatorial) notes that multi-armed, contextual, and combinatorial bandit frameworks already reallocate traffic toward better-performing variants in real time; today a human authors the candidates, and the speculation is that AI-generated variations could feed the same frameworks, turning them into open-ended, higher-volume experimentation engines.
The culture pillar gets the most emphasis and the most tactical detail, because the failure mode it addresses is common and expensive. Grassroots AI Adoption via Opt-In FOMO describes LinkedIn rolling out its AI tooling and the full-stack builder model from the bottom up rather than by mandate: small pilot pods first, access kept opt-in and conditioned on giving feedback, explicit FOMO and exclusivity so access feels earned rather than assigned, storytelling at all-hands, and deliberate spotlighting of unconventional career moves the new model made possible — the researcher-turned-growth-PM again — as proof the shift creates opportunity rather than threatening jobs. Cohen's reasoning is blunt: "I see a lot of companies roll out their agents and just expect [employees] to adopt. Doesn't work this way." The concrete mechanic is a gated early-access sign-up for each new tool or program, including the associate builder program, with active feedback as the price of entry — so the scarcity produces signal, not just demand.
Who actually moves first turned out to be counterintuitive. Top Talent Adopts AI Tools Fastest, Not Average Performers records that LinkedIn's top talent, not its struggling or average performers, adopted new AI tooling fastest and most enthusiastically; Cohen attributes this to the same drive to improve their craft that made them top performers before AI existed, now aimed at the new tools. He puts the readily-adopting group at a cutting-edge ~5% of staff. The other ~95% won't move on tool access alone — they need incentive design, visible internal success stories, and structural nudges. Cohen also names proactive AI exploration as a trait worth tracking, calling it "AI agency," and is moving toward formalising it in performance reviews, which converts an observation about early adopters into a hiring and promotion signal.
Underneath the tactics sits a stance about how internal opportunity should be distributed. Cohen cites Acemoglu and Robinson's argument in Why Nations Fail (Extractive vs. Inclusive Institutions) that nations succeed or fail on whether their institutions are inclusive — broad access to opportunity, participation, and reward — or extractive, concentrating control and gains in a narrow elite, rather than on culture, geography, or resources. The organisational analogy he draws is that a company's internal institutions (who gets access to new tools, whose ideas get funded, who can change roles) run on the same axis, and his preference for opt-in access anyone can earn over top-down assignment is the inclusive choice made deliberately. The broader mechanics of installing a change like this against organisational resistance are the subject of Transformation in Practice rather than this chapter.
The role prediction most consistently repeated across sources is not that product management shrinks uniformly but that it splits. AI-Driven Bifurcation of Product Manager Roles gives Marty Cagan's version: empowered product managers who bring real product judgment are reportedly commanding higher salaries, because AI tooling makes their judgment more valuable rather than less, while "product owner" roles and PMs on feature teams — people executing a backlog rather than exercising judgment — are predicted to shrink or disappear. Cagan frames it as the same bifurcation already visible among software engineers: AI amplifies the leverage of people with strong judgment and erodes demand for those whose job was mostly execution. A useful sub-distinction comes with it: the standalone associate-product-manager job title, which functions largely as project management, is predicted to be eliminated, while well-run APM rotational programs persist, because they operate as a concentrated crash course in product sense and coaching. The deciding factor is whether a role is coaching-dense or task-dense.
Tamar Yehoshua (Glean, ex-Slack CPO) frames the same split on a 5–10 year timeline in the same concept: AI blurs the lines between PM, engineer, and designer, but what it mainly removes is execution-heavy PM work — writing specs, coordinating handoffs, chasing status — because that work is closer to rote translation than to judgment. She expects AI to automate status updates, Jira hygiene, and launch docs, and to surface information (here's what customers are asking for, here are today's problems), but not the harder decision of what to build in response. Her summary is that this acts as a filter: "the not so good PMs' jobs will go away, the great PMs will still have great jobs." The career-craft side of building that judgment in the first place belongs to Product Leadership and Career Craft.
What the surviving half of the role gains is reach. AI-Enabled PM Self-Service: Data, Code & Debugging describes AI as a major unlock for lean, IC-heavy PM orgs: PMs run their own queries instead of filing a request with a data analyst, query the codebase directly instead of asking an engineer what's actually implemented, and combine live customer footage with real-time code inspection to diagnose issues without waiting on engineering triage. A PM can get a first-pass level-of-effort estimate by querying a coding assistant against the codebase and reserve the human scoping conversation for afterward. The specific data-side instance at Whatnot is Hex Threads (Self-Serve PM Data Tool), which lets a PM pull nuanced cohorts, inspect individual user logs, and build sensitivity, forecast, and regression models without a data scientist. The debugging pattern has its own name, Live Feedback Triangulation: run an AI query against the actual codebase and system state in parallel with the customer's live narrative, so a mismatch between what they describe and what the system does separates a genuine bug from a comprehension gap. The same absorb-the-busywork logic extends to operations in AI-Assisted Incident Pattern Matching, where feeding a new incident's details to an internal AI assistant surfaces similar past incidents and how they were resolved, folding institutional memory into live response instead of relying on individual recall.
This is not free. The same concept records the other side of the ledger: data scientists now report spending time reviewing half-assed AI-driven analysis produced by non-specialist PMs rather than doing original analytical work. The net effect isn't less data-science work, it's a shift in what that work is, toward review and correction. The bar for PMs rises accordingly. "Boxes and Lines" Systems Literacy sets the baseline — every PM should be able to sketch which systems drive which outcomes in their product, literally the boxes and lines from cause to effect — and then treats AI tools as raising the expectation past the diagram, into actual mechanism and code. The same source pairs it with a zoom-out discipline: Verrilli's example is that forcing all Whatnot sellers to create listings is a plausible fix for a search problem but carries a cost to seller throughput, so a local fix has to be evaluated against its knock-on effects before it rolls out broadly.
The most aggressive version of the prediction goes further than either Cagan or Yehoshua. AI-Powered Triple Threat, from a talk titled "Product Management Is Dead, So What Are We Doing Instead?" at the 2024 Lenny & Friends Summit, describes an emerging role where one person operates as PM, designer, and engineer at once while leading a team whose other members are AI tools, agents, or platforms rather than human specialists — framed as succeeding the traditional product trio because a broad-skilled individual with agents outruns a small team coordinating through handoffs. It is distinguished from generic generalism by spiking in one discipline while staying competent across all three, and the accompanying advice is to build custom team structures around such people and give them power and budget once identified. Note that this sits in tension with Cohen's insistence that specialists remain a legitimate destination; the corpus does not resolve which reading wins.
Below the role predictions sits a layer of unglamorous technique — the things PMs in these sources report actually doing with models. Start with prompting, since it's the cheapest lever: Persona/Role-Assignment Prompting Technique, cited by Tamar Yehoshua, is simply assigning the model a role before the task — "you are a product manager at Glean" — rather than issuing a generic instruction, on the logic that a specific perspective narrows the model's implicit criteria for what's relevant. She uses it for things like summarising competitor news or comparing PR angles. The same instinct, embodied in a product rather than a prompt, is Chat PRD (Templatized AI PRD Tool): Claire Vo's purpose-built PRD tool ships with built-in frameworks and structure instead of a blank prompt box, so the PM's input goes to content rather than to reinventing document scaffolding each time.
Research is where the reported gains are largest. AI-Assisted Competitive-Analysis Research names four applications specifically: trend-mining competitor release notes to see where a category is moving, review-mining competitor products for recurring complaints or delights, building head-to-head feature and positioning comparisons automatically, and asking open-ended diagnostic questions like "why is X succeeding?" to surface hypotheses a human analyst wouldn't think to test. It is narrowly about speeding up preparation deliverables, not about generating strategy. Long-Context LLM Ingestion for Unstructured Feedback Mining goes at volume instead: dump an entire unstructured source — one PM used a full Discord community transcript with Gemini's expanded context window — and ask analytical questions about sentiment and feature requests. The result was described as a gold mine precisely because the volume made manual reading impossible; the framing worth keeping is that "I'm too busy to read all of this" is a leverage problem rather than a time problem.
Two internal apps built at Glean show the same pattern applied to owned data. Internal AI App for Sales-Call Feedback Mining (Gong Transcripts) ingests Gong sales-call transcripts, structures them into a spreadsheet, and summarises the top requested features across calls; an early version conflated salesperson opinion with actual customer requests, and the fix was a prompt refinement separating the two so the output reflects verbatim customer asks — the pipeline was iterated by prompt, not rebuilt. Launch-Readiness Cross-Referencing Prompt cross-references the launch calendar against open Jira tickets, Slack conversations, and beta-customer feedback and returns one answer: a projected launch date plus a confidence level, replacing manual reconciliation across four systems.
On the reporting and stakeholder side, the interesting content is where practitioners stop. LLM as Internal Reporting Agent is Karina Stukan's (CEO, Bizzy) personal internal agent that aggregates and summarises recurring formulaic reporting — a monthly KPI check-in combining quantitative metrics with qualitative CRM notes into a consistent format. The boundary matters more than the automation: the model takes the data-aggregation and summarisation layer, while stakeholder-specific framing and persuasion stay human. She declines the adjacent technique in LLM-as-Virtual-Stakeholder Deck Review, where a near-final deck is handed to an LLM prompted to act as a replica of a specific stakeholder and asked what's missing or misreadable from that person's view. She has observed other teams doing it but won't, on the grounds that an LLM "can't replace knowing and understanding the people and the personalities in the room" — a judgment that belongs with the material on reading a room in Stakeholder Power and Trust.
Two habits close the toolkit. Preload-Summarize-Follow-Up AI Workflow is a personal pattern: load reading material into a voice-mode assistant ahead of time, ask for a summary at the start of a hands-busy activity like a workout, then explore by spoken follow-up — the modality is the point, since text chat would break the workflow. And Trying New AI Products as a Standing PM Practice, reported by Yehoshua, argues that hands-on trial of new AI products should be a standing habit rather than a skill acquired once, the same way PMs had to keep using new mobile apps during the smartphone platform shift to stay calibrated on what good looked like. The analogy carries a shelf life: the habit matters most while the category moves fast enough that last month's mental model is already stale. In practice it means regularly using ChatGPT, Claude, and competitors' products firsthand rather than relying on secondhand descriptions of what they can do.
Every source in this chapter that recommends AI also draws a boundary, and the boundaries converge. Caution Against Abdicating Product Sense to AI Tools records Cagan's response to a request for a single tool that ingests every input source — customer, sales, engineering, data — and auto-prioritises the results: doing that would mean abdicating product sense to a tool, even though he grants it's technically feasible with generative AI today. Generative AI can assist with synthesising input; deciding what matters and why stays with the human. Cagan is specifically most nervous about the effect on product managers, expecting AI to genuinely help engineers and designers while letting a PM skip the thinking coaching is meant to build, producing a plausible-looking artifact with no judgment underneath it. He also reports reversing his own earlier advice to start with ChatGPT output and improve it, after seeing people submit raw output as their own thinking; his revised guidance is to think the problem through yourself first, then use the model to challenge and stress-test that thinking rather than to originate it.
Julia Barham draws the same line for teams rather than individuals in AI for Synthesis, Done With the Team: it's legitimate to use AI to process and synthesise research data, but the synthesis conversation — what the data means, whether it changes strategy — belongs to the team. "You can use it for synthesis, just don't use it to do everything in synthesis — use it with the team, synthesize with the team." Her structural answer is a recurring cadence, monthly or every two months, where the team synthesises customer, platform, and business data together and jointly discusses whether it changes strategy or existing hypotheses; the tool assembles the data, the sense-making stays a team activity. The individual-level check that pairs with this is Calling BS on AI Output (Critical-Thinking Check) — Randy's rule of applying the same scrutiny to model output you'd apply to a human teammate's claim, with the honest catch attached: "I have to have enough context to be able to call BS on them," and without that context AI's mistakes are harder to catch than a colleague's precisely because the output reads fluently.
That catch has hard numbers behind it. Illusion of Correctness (AI-Generated Code) takes its name from security researchers explaining why AI-assisted coding can raise commit velocity three to four times while vulnerability pass rates stay flat or worsen: polished output discourages scrutiny, so reviewers extend less skepticism to something that already reads as done. Veracode's 2026 report found roughly 44% of AI code-generation tasks introduced a real exploitable vulnerability and an average security pass rate of only 56% across tested models, and one tracked environment saw monthly security findings climb from about 1,000 to over 10,000 in six months while shipping speed increased. The pattern generalises past code to any output whose surface fluency outruns its correctness — and it quietly undermines the standard safeguard, since a mandatory human sign-off only catches errors if reviewers stay motivated to look hard at something that already looks right. The chapter notes alongside this a prediction from Cognition that such approval gates may not survive as a durable requirement.
On strategy specifically, the material is more structured. AI-Fit Evaluation for Strategy gives a two-part check to run per candidate problem before committing it: should this problem leverage AI at all, or is a deterministic solution better suited; and if yes, is it technically ready today or does it depend on foundation-model progress that hasn't happened. Both answered explicitly, including what role the probabilistic component plays and what guardrails constrain it. One level up sits the risk in how the strategy itself gets made: "Can tools help bring clarity to thought? Yes. Can tools also be a substitute for thinking? Unfortunately, yes." Strategy is described as fundamentally a thinking exercise where the primary tool is the brain, so AI's legitimate role is sharpening reasoning already done. Chandra Janakiraman locates the usable part of that precisely, in the preparation phase of his strategy process: competitive research and synthesis, plus a first-pass draft. AI Mock Strategy as a Down-Select Input describes what that draft actually is — well-informed, well-articulated, comprehensive, and for exactly that reason not a strategy, since it reads like every reasonable option laid out at once with none of the forced narrowing. Its use is as an input to human down-selection and ranking. Every new idea, in Chandra's words, is "an ugly baby that people just want to get rid of," and "there's a tremendous amount of craftsmanship between a great idea and a great product." The full mechanics of that strategy process live in Strategy, Vision, and the Decision Stack. Chandra also flags Multi-Agent Strategy Model (Speculative) — separate strategy, roadmap, and engineering agents iterating with each other instead of a human working group — as sitting on the far side of a crossover point where human judgment becomes inferior to something processing multiple signals simultaneously. He expects that point but doesn't think it has been crossed for strategy work, and labels the whole thing a direction to watch rather than a current recommendation.
The last two cautions are about attention rather than accuracy. John Cutler argues in AI Removes 'Simmer' Time for Thinking that AI doesn't fix ungoverned, chaotic work environments — it accelerates reactivity. Meaningful thinking needs unstructured time for ideas to simmer in the background between working sessions, and AI compresses that gap by prompting instant follow-up action. His line: "I don't feel cognitively better after a day with AI. I feel worse. Like I've had to protect more of my actual cycles." AI, in his framing, arrives as one more input on an already-overloaded cognitive environment, and the question adoption rarely answers is "What are you going to stop doing to make any room in your brain?" Productivity Gains Get Reinvested as Higher Output Expectations, from Emily Tate, names the mechanism by which that room disappears: when tooling makes people meaningfully more productive, the gain is rarely banked as slack — expectations quietly rise to match the new capacity, from "I can build twice as much" to "I should push myself to build three times as much." Someone has to deliberately choose to keep the slack, or it vanishes into a new baseline of expected output.
Everything above is about AI changing how product work gets done. A smaller body of material here addresses the different question of what changes when the product itself is an AI system — and it's worth flagging up front that this is the thinnest part of the chapter, resting on a couple of single-source accounts rather than accumulated practice.
The strategic claim is the most portable. Durable AI Differentiation vs. Patching Current Model Weaknesses, reported by Tamar Yehoshua, warns that differentiation built around covering for today's model weaknesses — hallucination workarounds, context-window limits, specific reasoning gaps — has a half-life equal to that weakness, because every competitor building on the same foundation models inherits each improvement for free. "Your whole product gets better as the LLMs get better" cuts both ways: improvements you didn't cause can also erase the thing you were selling. Durable differentiation has to live outside the model — proprietary data, workflow integration, organisational knowledge structures, or trust and deployment relationships that survive model improvement. This adds a shelf-life test to the fit question from the previous section, and it connects outward to the market-facing failure modes covered in Positioning, Marketing, and Go-to-Market.
Glean is the worked example of getting that right by accident of sequencing. Glean's Evolution: Enterprise Search to Organizational Knowledge-Graph Chat traces the company from 2019 enterprise search built on Google's BERT models plus vector embeddings, through adding a chat interface once GPT-3 made conversational retrieval viable, to framing itself now as building "a knowledge graph of your organization" rather than a search box or a chatbot. The point is the ordering: the search-era investment in indexing and cross-linking every connected SaaS source became the substrate the chat layer sits on, so moving into chat and agent interaction was a UI layer added on top of existing infrastructure rather than a rebuild. The interface changed before the underlying value proposition did.
On the craft of specifying agentic products, the chapter has one proposal, and its provenance deserves stating plainly: Product Experience Document (PXD) is reported firsthand by a single founder — Rags Vadali of Floto, a roughly 3–6 person team as of 2026 — and is not externally validated. It replaces the PRD for agentic products on the reframing that what you're defining is the experience layer on top of the agent, not a UI. Its structure runs Why → Success Criteria (quantitative and qualitative) → Experience Principles → Critical Moments → Conversation Closing → Success Metrics, and unlike a PRD it's written after a working prototype exists: engineers build from a problem statement, and the PXD gets authored only once the PM has played with what got built. The Experience Principles section is fed into Claude to generate the prompts engineers use to build and tune the agent, so the document ends up written for a coding agent as much as for a person. Because the system underneath is non-deterministic, the PXD specifies ranges of good, bad, and explicitly-unwanted answers rather than a single fixed requirement, with real good/bad/ugly transcripts pulled from usage. Vadali is explicit that this only fits agentic products; for interface-heavy work he still prefers a close designer-and-engineer process producing a Figma reference first. The runtime pattern it enables is Critical Moments & Breakpoints (Agent Conversation Design Pattern): an agent normally pursues a fixed set of conversational goals, but must interrupt that loop at a breakpoint when a user surfaces something serious, digging deeper before returning to or abandoning the original line. Breakpoints are explicit triggers defined in advance — if you hear this, stop everything and dig deep — which reframes good listening as a control-flow override rather than a soft instruction. Conversation closing is treated as a separate concern in the same document, on the grounds that how an agent ends a conversation habituates the user toward or away from future engagement even though it doesn't affect the current session.
Governance is the thinnest corner of all: the chapter carries two checklists and no worked application of either. Akshay Kore's Trustworthy-AI Framework, from 2018, argues that a system which isn't trustworthy isn't useful and breaks trustworthiness into five dimensions — explicable, transparent, non-biased, privacy-centered, and beneficial to society — and is cited by Simonetta Batteiger as a practical audit checklist for whether an AI-driven feature is launch-ready. EU AI Act Trustworthiness Dimensions is its already-codified regulatory counterpart, framing trustworthy AI around human autonomy, prevention of harm, fairness, and explicability, usable as a baseline when defining design constraints or compliance requirements. Anyone needing more than a checklist will have to go outside this corpus. Finally, one gesture beyond software: Medicine 3.0 (Personalized, Preventive Medicine) is the term from Peter Attia's Outlive for the shift from reactive treatment of disease to personalised, preventive management of long-term risk, which Cohen cites as a domain he expects AI to accelerate — more personalised data, cheaper continuous monitoring, faster iteration on individual risk models — as an example of the same idea-to-design compression showing up outside the industry this chapter otherwise lives in.