The Feed · Complete archive

AI at work: research and practice

A server-rendered record of the evidence, ideas and firsthand practices screened by The Feed. Every entry links to its original source.

Use the interactive Feed

Page 1 of 9 · 873 items

  1. practice · Nate B Jones ·

    Executive Briefing: Your Best Engineer Got Ten Times Faster. Your Team Got 1.8x. Here’s how to close the Gap.

    Fast AI users create bottlenecks elsewhere; manage by unblocking reviews and decisions, not by standardizing their techniques.

    Somebody on your team has figured something out. They’re using agents in a way that looks different from what everybody else is doing. They have several pieces of work moving at once. They come back from lunch and a problem that used to take the afternoon is ready to review. They spend money on tokens without apologizing for it, because they can point to what they got back. And the team is still behind. That combination is going to become familiar. One person’s capacity can change much faster than the organization around them. Code arrives faster than anyone can review it. A prototype is ready

    management org

  2. practice · Lenny's Newsletter ·

    The grief, loneliness, and burnout sweeping through the tech industry right now | Molly Graham

    Experienced operators describe how delegation to AI differs from human delegation, and what managers should keep close.

    Molly Graham is back for round two, and this one is even more powerful. Molly has spent more than 20 years helping organizations and the humans inside them navigate growth and change. She’s held leadership roles at Google, Facebook, Quip, and the Chan Zuckerberg Initiative and is the host of TED’s WorkLife podcast (which she took over from Adam Grant). She also runs Glue Club, a leadership community for senior operators, and writes a popular newsletter called Lessons . Listen on YouTube , Spotify , and Apple Podcasts In our in-depth conversation, we discuss: Why Molly’s famous “give away your

    management org · Lenny Rachitsky

  3. practice · Simon Willison ·

    Kākāpō Party

    Using Claude to generate interactive HTML5 canvas animations, then automating browser interaction with Playwright to capture video for presentations.

    Tool: Kākāpō Party I presented a closing keynote for the WeAreDevelopers World Congress North America yesterday. As a STAR moment I decided to weave in references to the record breaking kākāpō breeding season we had in 2026. For my closing slide I wanted to celebrate, and I had seen some buzz around how good Claude Opus 5.5 was at creating pixel art animations. So I rounded up three Kakapo photos from Google image search and dropped them into Claude with this prompt: Here are some photos of kakapo parrots just to remind you what they look like I need you to make an animation in animated pixel

    ways of working · Simon Willison

  4. practice · Simon Willison ·

    Note on 24th September 2026

    Working with coding agents requires extraordinary discipline and knowledge to unlock their potential, making software engineering harder overall.

    The more time I spend working with coding agents, the more convinced I am that they make software engineering even harder. We can do amazing things with them, but unlocking their full potential requires extraordinary discipline and knowledge. Tags: coding-agents , ai , llms

    ways of working · Simon Willison

  5. research · VoxEU ·

    Measuring what work generative AI does: Survey evidence versus chat logs

    A nationally representative survey linking AI use to worker tasks disagrees sharply with chat-log classifications from Anthropic, Microsoft and OpenAI.

    The chat logs of AI companies are often used to measure what work AI actually does. This column uses a nationally representative US survey that links generative AI use to workers’ detailed tasks and compares them with task shares derived from Anthropic, Microsoft, and OpenAI chat data. The four sources disagree sharply, largely because chat classifiers cannot see a user’s occupation and therefore classify use into a few generic activities. Chat logs are informative but on their own can misattribute AI use across occupations and cannot substitute for measurement anchored in who the worker is.

    adoption

  6. practice · GitHub Blog ·

    AI-powered fuzzing with the GitHub Security Lab Taskflow Agent

    An LLM agent automates fuzzing workflows: identifying entry points, writing harnesses, running AFL++, triaging crashes and writing vulnerability reports without human oversight.

    If you’re new to fuzzing and want to learn the fundamentals first, check out our Fuzzing 101 course at gh.io/fuzzing101 . Continuous fuzzing is not a magic solution that solves all your problems . Even projects that have been enrolled in OSS-Fuzz for years can still hide critical bugs, and the reason is almost always the same: someone needs to keep an eye on coverage, write new harnesses for the code that nobody is reaching, and triage the crashes that come out the other end. In other words, fuzzing still needs a human in the loop. So the natural question I kept asking myself was: how much of

    ways of working

  7. practice · Latent Space ·

    Foundries vs Navigators: Lowering the Cost of Science

    AI speeds up research planning and design, but labs still bottleneck on slow physical experiments; two adaptation paths emerge.

    What does the future of science look like in the world of AI? Anthropic has some lofty goals for science and is even opening a wet lab . Meanwhile a quiet transformation 1 is happening all across AI x Science. In this guest post, Adrian Sanborn talks about the less flashy but more immediate ways he sees AI transforming front-line scientific research in his own company, Endura Therapeutics. Adrian did a CS PhD at Stanford and spent much of it running experiments at the bench, which makes him one of the rare people who can tell you what an LLM is doing to a codebase and to a wet lab. Enjoy! Lang

    ways of working

  8. research · npj Digital Medicine ·

    A multicenter assessment of human oversight of generative AI outputs in simulated clinical decision making

    Junior clinicians identified only 15.8% of AI hallucinations in simulated clinical decisions; 13.1% missed all hallucinations regardless of risk level.

    Generative AI, particularly Large Language Models (LLMs), is increasingly being explored for selected clinical tasks, including clinical documentation, patient communication, and decision support; however, routine use in direct clinical decision-making remains limited. A central barrier to safe deployment, however, is the problem of medical hallucinations. Current safeguards rely on a “clinician-in- the-loop” model, which assumes that clinicians can reliably identify and correct these hallucinations. Whether this assumption actually holds in clinical practice warrants greater empirical attenti

    judgment

  9. practice · Paul Ford (Aboard) ·

    I Keep Accidentally Building Search Engines

    Using AI to build personalized search and filtering systems that answer specific questions across curated sources rather than writing fresh content.

    I think we all know that AI-generated text is hit or miss. It is good at summarizing things, or writing documentation about code, or generating a report after a meeting. On the other hand, I don’t let AI write articles for me, or do anything to my manuscripts. I’m not against you doing that, at all—it’s just that writing is how I think, and I don’t know what I think until I write it down. But that doesn’t mean that I don’t use AI prose. In fact, I find it incredibly useful, and consume thousands of words of AI-generated prose per day. For example, I get a daily newsletter that summarizes and s

    ways of working · Paul Ford

  10. practice · Zapier ·

    How Ethan Schwandt helped Jobber turn AI adoption into a building culture

    Manager helped teams spot AI opportunities by mapping work and handoffs, then paired adoption with governance and hands-on support.

    When companies talk about AI transformation, they often start with what they rolled out: the model, the license, the training program, or the pilot. Ethan Schwandt started somewhere more useful: the work. As Senior Manager, Talent Acceleration at Jobber, Ethan helped employees recognize where AI and automation could remove a recurring handoff, reduce manual work, or make information easier to act on. He paired that opportunity-spotting with practical enablement, clear governance, and hands-on su

    adoption

  11. practice · Zapier ·

    How Jocelyne Mendez-Guzman made follow-up faster

    Sales team built AI-powered follow-up system that reduced call follow-up time by 85% while maintaining rep trust.

    Jocelyne (Joce) Mendez-Guzman is no stranger to the Zappy Awards. In 2025, she won our Operations Automator of the Year for a workflow that processed accounts receivable tickets 69% faster. With that win under her belt, she moved into BioRender's RevOps team, where her work now supports teammates across GTM. This year, she’s back as the 2026 Sales AI Builder of the Year, helping those teams follow up on calls 85% faster. Joce built a follow-up system her reps can trust, giving GTM teams speed wi

    productivity

  12. research · Proceedings of the ACM on Human-Computer Interaction ·

    "Even with AI, We Still Have to Work Harder": Black Women’s Use of Generative AI in the Workplace

    Black women in tech use generative AI as survival tools to manage discrimination and hostile work environments, but tools reinforce existing power structures.

    As generative AI (GenAI) technologies become increasingly integrated into professional settings, it is critical to understand how these tools are experienced by individuals who are more likely to encounter harm, violence, and trauma in the workplace. Black women, in particular, continue to face systemic discrimination in hiring, promotion, and everyday interactions within predominantly white work environments. In this paper, we center the voices of 21 Black women working in the U.S. tech sector to explore how they engage with GenAI technologies to navigate workplace dynamics. Drawing on Black

    worker experience

  13. practice · Zapier ·

    How Doug Hamilton turned an AI tracker into a team of builders

    Recruiting team built an automated system to monitor AI tools hourly, translate updates to job relevance, and deliver personalized digests with hands-on guides.

    Doug Hamilton built an AI Tech Stack Tracker, an automated intelligence system that helps Klaviyo’s Talent Acquisition team keep up with fast-moving AI tools and turn updates into action. Every hour, the tracker polls more than a dozen AI tools directly. Claude translates raw changelogs into plain-language relevance for recruiting work, then a weekly digest delivers updates based on what each teammate cares about. A “Try it out” feature sends a personalized, role-specific guide to Slack. A “Shar

    adoption

  14. research · Proceedings of the ACM on Human-Computer Interaction ·

    Constructing Algorithmic Authority: How Multi-Channel Networks (MCNs) Govern Live-Streaming Labor in China

    MCNs in China construct different algorithmic narratives internally and externally to manage live-streamers, shifting accountability away from platforms to workers.

    This study examines the discursive construction of algorithms and its role in labor management in Chinese live-streaming industry by focusing on how intermediary organizations (Multi-Channel Networks, MCNs) actively construct, stabilize, and deploy particular interpretations of platform algorithms as instruments of labor management. Drawing on a nine-month ethnographic fieldwork and 44 interviews with live-streamers, former live-streamers, and MCN staff, we examine how MCNs produce and circulate structured interpretations of platform algorithms across organizational settings. We show that MCNs

    worker experience

  15. research · Proceedings of the ACM on Human-Computer Interaction ·

    Expecting Too Much, Getting Too Little: Exploring the Challenges and Design Opportunities of Asynchronous AI Interviewers

    Applicants using asynchronous AI interviewers report mismatched expectations and low trust, leading to workarounds and deceptive practices; design changes improve perceived agency.

    Organizations use asynchronous AI interview systems to efficiently manage large applicant pools, enabling quick and uniform evaluations. However, concerns remain about their impact on user agency and the lack of personalization applicants experience with these systems. Although efforts have been made to humanize the interview process, users’ expectations are often unmet, especially when compared to the promises made by these systems. To examine how applicants perceive and experience these tools, particularly as they become more familiar with widely accessible AI technologies, we conducted a tw

    adoption

  16. research · Proceedings of the ACM on Human-Computer Interaction ·

    Bringing Everyone to the Table: An Experimental Study of LLM-Facilitated Group Decision Making

    LLM facilitators increased information sharing in group decisions by raising minimum engagement, without harming group attitudes, in a 1,475-person experiment.

    Group decision-making often suffers from uneven information sharing, hindering decision quality. While large language models (LLMs) have been widely studied as aids for individuals, their potential to support groups of users, potentially as facilitators, is relatively underexplored. We present a pre-registered randomized experiment with 1,475 participants assigned to 281 live groups completing a hidden profile task—selecting an optimal city for a hypothetical sporting event—under one of four facilitation conditions: no facilitation, a one-time message prompting information sharing, a human fac

    teams · Jake Hofman

  17. research · Proceedings of the ACM on Human-Computer Interaction ·

    AI in the Workplace: The Impact of AI on Perceived Job Decency and Meaningfulness

    IT and healthcare workers worry AI will reduce job meaningfulness despite better hours; service workers expect status gains but no time relief.

    The proliferation of Artificial Intelligence (AI) in workplaces is transforming how we work. While existing research on human-AI collaboration at work often prioritizes performance, less is known about their experiential outcomes. Through interviews with 24 employees across Information Technology (IT), service-based, and healthcare sectors, this paper examines AI’s impact on job satisfaction via perceptions of job decency and meaningfulness, now and in the future. Our results reveal that the anticipated impact of AI on overall job satisfaction varies with the occupational domain, with differin

    worker experience

  18. practice · Zapier ·

    How Carlos Robledo turned a recruiting report into organizational infrastructure

    A sourcer automated his Friday reporting task into a 19-workflow system that now serves the entire recruiting team.

    When personal automation becomes organizational infrastructure Carlos Robledo is a Lead Technical Sourcer at Hims & Hers. When he got tired of spending hours every Friday building hiring reports for stakeholders, he built Rex, a 19-workflow system that now serves recruiters across the talent organization. There was a specific moment that triggered Rex. While working with his Chief Product Officer, Carlos was responsible for producing two major updates across several open requisitions. The work m

    adoption

  19. practice · Zapier ·

    How Ariel Chen built trust before she built automation

    People ops team used AI to track and predict background check status across 12 countries, reducing delays before automating.

    Ariel Chen coordinates People Operations at Figma, where a two-person team tracks over 100 background checks each month across 12 countries. In a role where a missed screen can delay someone's start date or create compliance issues, Ariel built a system her team can rely on. When every screen says “pending” Figma hires globally, and international background checks often take longer than the two-week lead time between onboarding cohorts. But every background check showed the same generic "pendin

    adoption

  20. practice · Zapier ·

    How Ignacio Piñeiro scaled fraud control with exception-based AI review

    Fintech operations team uses AI to flag suspicious delivery photos, letting humans review only high-risk cases instead of all 3,000 monthly submissions.

    Ignacio Piñeiro runs operations at Galgo, a fintech that lends money for motorcycles across Mexico. Every loan he approves hinges on a simple question: Did this delivery actually happen as it should? Three thousand photos, one last chance Each motorcycle Galgo finances requires proof of delivery — a photo of the customer with their new vehicle — before they can start collecting the loan. That photo is the company's main defense against delivery fraud. If fake, wrong, or low-quality evidence slip

    productivity

  21. research · Proceedings of the ACM on Human-Computer Interaction ·

    Generative AI as a Mediator in Creator Collaboration: Challenges, Design Opportunities, and Concerns

    GenAI mediated collaboration between narrative writers and visual designers in game development, improving shared understanding but risking reduced communication and trust.

    As Generative AI (GenAI) expands beyond individual workflows and becomes increasingly integrated into collaborative creative work, its potential to address the fundamental challenge of building common ground among creators with diverse expertise remains underexplored. Specifically, we lack empirical understanding of how GenAI can mediate collaboration asymmetries among heterogeneous creators. To address this gap, we conducted a multi-method study with narrative writers (N = 10) and visual designers (N = 10) in the context of game development. Our findings demonstrate that GenAI goes beyond ind

    teams

  22. research · Proceedings of the ACM on Human-Computer Interaction ·

    Value Sensitive Design for Fair Online Recruitment: A Conceptual Framework Informed by Job seekers’ Fairness Concerns

    Job seekers' fairness concerns in online hiring span discrimination, interaction bias, qualification misinterpretation, and power imbalance; design framework proposed.

    The susceptibility to biases and discrimination is a pressing issue in today’s labor markets. While digital recruitment systems play an increasingly significant role in human resource management, a systematic understanding of human-centered design principles for fair online hiring remains lacking, particularly considering the gap between idealized conceptualizations of fairness in research and actual fairness concerns expressed by job seekers. To address this gap, this work explores the potential of developing a fair recruitment framework based on job seekers’ fairness concerns shared in r/job

    jobs skills

  23. research · Proceedings of the ACM on Human-Computer Interaction ·

    The Transparency Paradox: An ethnographic investigation of AI Ethics and Accountability in Nigerian journalism

    Nigerian journalists report GenAI efficiency gains in some tasks but 'double work' when it fails, and struggle to disclose AI use despite believing in transparency.

    The rise of artificial intelligence has led to disruptions across major knowledge work, including journalism. The recent proliferation of generative artificial intelligence (GenAI) and large language models (LLMs) such as ChatGPT, Gemini and MetaAI, has led to an increased study of the influence of AI in journalism. However, much of these studies have been focused on western contexts, with very little knowledge of how AI is transforming journalism in Africa. To begin addressing this gap, we conducted a two-phase study to understand how Nigerian journalists are using AI. Nigeria is not only sig

    adoption

  24. research · Proceedings of the ACM on Human-Computer Interaction ·

    Shaping Collaborations with Algorithms: How Agency and Heterogeneity Criteria Influence Team Formation and Outcomes

    Lab experiment shows how team-formation algorithms with different levels of user control and diversity criteria reshape team composition and collaboration outcomes.

    Across professional networking platforms, scientific collaboration networks, co-founder matching tools, and workplace collaboration platforms, algorithms increasingly shape how individuals find, evaluate, and connect with potential collaborators. These systems create tensions between user agency and organizational values: Should algorithms organize individuals directly in line with organizational goals? Should algorithms allow individuals to choose freely? Or should algorithms subtly nudge choices toward those goals while preserving user agency? Each approach has implications for who gains acc

    teams

  25. research · Proceedings of the ACM on Human-Computer Interaction ·

    Machine Semiology in Practice: Clinician Strategies for Interpreting AI-Generated Visual Explanations

    Clinicians perform substantial interpretive work to reconcile AI saliency maps with medical diagnostic practice, revealing gaps between XAI design and actual use.

    The increasing adoption of AI-based decision support systems in high-stakes professional settings has intensified interest in eXplainable Artificial Intelligence (XAI), particularly in domains where human accountability remains central. In medical imaging, visual explanation techniques such as saliency maps are widely promoted as a means of rendering opaque models interpretable to clinicians. Despite their popularity, however, little is known about how such explanations are actually interpreted and made meaningful within radiological practice. This study addresses this gap by introducing and e

    judgment

  26. research · Proceedings of the ACM on Human-Computer Interaction ·

    Platform Power and Worker Agency: A Socio-Technical Review of the Gig Economy

    Systematic review of 171 gig-economy studies identifies 46 technical mechanisms of algorithmic management and five stakeholder relationships shaping worker conditions.

    The platform-mediated gig economy constitutes 4.4% to 12.5% of the global workforce. Understanding how platform design shapes working conditions, and how workers develop strategies to navigate platform constraints, is crucial to addressing challenges in this rapidly evolving sector. Following PRISMA guidelines, we examine 171 gig economy-related publications from 2019 to Nov. 2025, spanning algorithmic management systems, labor process theory, and workers’ insights. Employing a “best-fit” framework, we draw on Socio-technical Systems Theory and synthesize four strands of scholarship to consoli

    worker experience

  27. research · Proceedings of the ACM on Human-Computer Interaction ·

    Uncovering Disparities in Rideshare Drivers’ Earning and Work Patterns: A Case Study of Chicago

    Mixed-methods analysis of Chicago rideshare data 2018–2023 reveals how platform algorithms interact with driver decision-making to reinforce earning disparities.

    Ride-sharing services are reshaping urban mobility but raising concerns around fairness, safety, and driver working conditions. This mixed-methods study combines large-scale analysis of Chicago’s Trip Network Provider data (2018–2023) with qualitative interviews and reconstructed driver work patterns. Quantitative results reveal temporal and spatial disparities in driver earnings, while qualitative findings explain how drivers’ decision-making—centered on fare rates, pickup distance, and follow-up trip potential—interacts with platform algorithms to reinforce these inequities. We further uncov

    worker experience

  28. practice · Zapier ·

    Meet the 2026 Zappy Award winners: the builders who put AI to work

    Companies measure AI adoption by business metrics that matter: invalid deliveries down 6 percentage points, phone revenue impact quantified.

    Most companies still measure AI by what they roll out: seats bought, licenses assigned, and pilots launched. Those numbers prove that the company has started. They don’t prove anything changed. The 2026 Zappy Award winners measure it differently. Every one of them moved a number that their business already reports on. At Galgo, the rate of invalid delivery evidence fell from about eight percent to under two percent across more than 3,000 monthly submissions. At Youtech, $213,000 in phone revenue

    productivity

  29. research · Proceedings of the ACM on Human-Computer Interaction ·

    U.S. Southerners’ Attitudes Towards Algorithmic Analysis of Voice Data for High-Stakes Employment and Education Evaluations

    Southern U.S. speakers express concerns that voice-based algorithmic hiring and education evaluation could penalize dialects, preferring human involvement.

    The rise of algorithmic decision-making in workplaces and schools has drawn attention to potential benefits and harms for stakeholders. Additionally, uneven performance of language technologies for non-standardized or underrepresented dialects has resulted in calls to engage speakers of those dialects. Our survey study explores the attitudes of 111 American English speakers from the southern U.S. towards voice-specific applications of algorithmic analysis in four high-stakes decision-making use cases: evaluation of candidates in hiring and college admissions interviews, employee performance ev

    judgment

  30. research · Proceedings of the ACM on Human-Computer Interaction ·

    Algorithmic Disruptions to Expertise: The Impact of the Home Valuation Tool Zestimate on Real Estate Professionals’ Work

    Real estate agents respond to algorithmic valuation tools by developing new practices to explain and reinterpret algorithmic outputs to clients.

    Algorithmic systems increasingly influence knowledge-intensive work and reconfigure professional expertise. This article investigates how real estate agents [N = 15] experience and respond to the digital real estate platform Zillow and its home valuation algorithm, Zestimate. Based on semi-structured interviews, we show how publicly available algorithmic estimations reshape professional practices and client–agent relations by altering information asymmetry, monitoring, and accountability within the principal–agent relationship. To sustain their relevance under these conditions, agents proactiv

    worker experience

  31. research · Proceedings of the ACM on Human-Computer Interaction ·

    Exploring the Impact of AI-powered Creativity Support Tools on Professional Creative Workflows

    Creative professionals using AI tools report impacts on collaboration, accessibility, and creative agency across workflow stages.

    Creativity Support Tools (CSTs) have long been a subject of interest within Human-Computer Interaction (HCI) research. The advent of Artificial Intelligence (AI), particularly generative AI, introduced new levels of adaptability and automation to CSTs. Despite increased adoption, gaps persist in understanding how practitioners appropriate AI‑powered CSTs across creative workflows. This paper presents findings from an exploratory, multi-session study with creative professionals, examining their use of AI-CSTs across workflow stages. We identify key successes, challenges, and aspirations for too

    adoption

  32. practice · Harper Reed ·

    Why Don't My Agents Break Containment?

    Token abundance removes containment assumptions built into agent design; systems designed for scarcity behave differently without cost constraints.

    The Openai and hugging face situation is so fucking insane. If you haven’t, I highly recommend watching the blackhat talk , and then jumping directly into the firehose of subsequent disclosures and reporting. When it all came out I mentioned to Jesse “why aren’t our agents breaking out of containment” and we thought through a handful of reasons. There were a lot of reasons: architecture eval / gyms harness safety measures unlimited tokens My favorite reason is unlimited tokens. I realized that everything we are building is built with the assumption of token scarcity - whereas the big labs have

    ways of working · Harper Reed

  33. practice · The Pragmatic Engineer ·

    How will AI change operating systems? Part 2: Windows

    Windows is adding native agent discovery and isolation primitives to support agentic AI tools natively in the OS.

    AI is changing how us software engineers build software, and developers’ tool preferences are rapidly evolving with it – like how AI coding harnesses have gotten very popular. Likewise, future versions of the world’s leading operating systems look certain to feature more support for agentic tools. To find out how things are changing, we talked with the folks at tech giant Microsoft who shared in detail their plan for Windows, and how AI will play a part. It comes after the company’s previous AI efforts led to it being dubbed “Microslop” by some users online. The Windows team’s vision is opinio

    management org · Gergely Orosz

  34. research · arXiv ·

    Does AI Save Time on Product Design? A Randomized Controlled Experiment of AI Prompt-to-Design Workflows

    RCT with 50 designers and 50 product managers measures time savings from prompt-to-design AI tools (Figma Make) on standardized tasks.

    AI tools for digital product design now offer prompt-to-design capabilities, allowing designers and their non-designer colleagues to create prototypes through conversational workflows with large language models (LLMs). While these tools promise time savings, experimental evidence in product design remains limited compared with evidence from software engineering. We conducted a randomized controlled trial with 50 product designers and 50 product managers to evaluate prospective time savings from leveraging Figma Make in design work. Participants attempted three standardized design tasks with or

    productivity

  35. practice · Nate B Jones ·

    Your agent is giving you reasonable answers from half your context. A conversation with OpenAI.

    Legal, finance and other functions adopt AI when they get access to their own documents and tools; deciding what to build becomes bigger than building.

    When I asked which teams inside OpenAI had crossed over after software engineering, legal came up early. In my conversation with Andrew, who leads the team behind OpenAI’s Codex desktop app, and Akshay, who leads OpenAI’s productivity and engineering work, they described different functions reaching a point where AI became useful for much more of their jobs. Coding came first, because everything was already a local CLI. Legal was next, because the material was sitting in documents the model could open. Other functions crossed over as access to their information improved and the tools became be

    adoption

  36. practice · Lenny's Newsletter ·

    Advanced evals: How to find (and fix) hidden AI failures in your product

    Teams skip critical discovery before writing metrics; most evals measure the wrong things in AI products.

    👋 Hey there, I’m Lenny. Each week, I share deeply researched product, growth, and career advice. For more: Lenny’s Jobs | Lennybot | Become an AI-Native Builder and other favorite AI/PM courses Subscribe now P.S. Get a full free year of Cursor, Notion, Lovable, Replit, Wispr Flow, Linear, Factory, ElevenLabs, PostHog, Granola, Brain.fm, Waking Up, and more, by becoming an Insider subscriber (while supplies last). Learn more . Evals have been coming up more and more in my conversations with podcast guests and PMs. And nearly half of the 25 awesome PM job openings I shared on socials last week

    judgment · Hamel Husain

  37. research · npj Digital Medicine ·

    AI adoption and workflow optimization following orchestration platform implementation and structured change management

    Structured change management plus an AI orchestration platform increased radiologists' active review of AI outputs from 14% to 61% in a neuroradiology department.

    Abstract Artificial intelligence (AI) adoption in radiology remains limited by workflow integration challenges, lack of interoperability, and barriers to user engagement. We implemented a dual-track socio-technical strategy in a tertiary-care neuroradiology department, combining a vendor-neutral AI orchestration platform with a structured change management program, including user education, workflow standardization, documentation, feedback mechanisms, and governance. In an observational pre–post implementation study, AI adoption was evaluated by comparing a baseline period using a single-vendo

    adoption

  38. practice · Shopify Engineering ·

    Helix: The internal tool powering our Shopify app's native migration

    Using LLM checkpoints and quality gates to keep AI-generated mobile app code continuously shippable during large rewrites.

    We built Helix to help LLMs rebuild the Shopify app in Swift and Kotlin, using small checkpoints and strict quality gates to keep the code shippable.

    productivity

  39. practice · Charity Majors ·

    Where do AI norms come from?

    A company shares how it developed AI norms and values through deliberate process, not accident.

    After I posted “ Confessions of an Unrepentant Slop Snob ” on LinkedIn last week, with links to the Honeycomb AI norms and values docs , I got this question from Niklas Lochschmidt : I am curious if you could share more light on the process? Any guidance you would give people in companies that have the friction and internal debates, but maybe haven’t arrived at shared norms and values yet? That’s a great question, and I will answer in just a minute. I feel like this is a great way to kick off something I’ve been looking forward to all year — an old fashioned bloggy dialogue between myself and

    management org · Charity Majors

  40. practice · Linear ·

    AI coding has made CI a bottleneck, so we reworked ours to keep up

    AI-driven code generation made CI a bottleneck; reworking infrastructure and test strategy halved runner time while keeping PR wait time manageable.

    Earlier this year, I opened Linear to find that Tuomas, our CTO, had assigned an issue to me, titled “CI costs are high.” While I was at it, he also wanted me to make CI faster. Agents have made it exponentially faster to ship code, but validating those changes hasn’t quite kept up at the same rate. Every PR still has to pass through CI, so as development accelerates, CI becomes a bottleneck, driving up infrastructure costs and leaving developers and agents waiting longer for feedback. In our pursuit to make CI more performant at Linear, we optimized for how long a PR waits on CI and how much

    productivity

  41. practice · Lenny's Newsletter ·

    How Warp ships 2,000 PRs a month with AI factories | Zach Lloyd (CEO, Warp)

    Software factory ships 2,000 PRs monthly by automating Slack-to-merge workflow; human review becomes the bottleneck.

    Zach Lloyd is the co-founder and CEO of Warp, an AI-powered terminal and software factory platform used by tens of thousands of engineers. Before Warp, he spent nearly a decade at Google, including time as a principal engineer on Google Sheets. He built Warp from the ground up as a modern, AI-native alternative to legacy terminals, and the team has since expanded into software factories: a full cloud-based system that takes an idea in Slack all the way through to a merged PR. Listen or watch on YouTube , Spotify , or Apple Podcasts In this episode: Why a software factory is more than a coding

    productivity · Claire Vo

  42. practice · Simon Willison ·

    Quoting voxium

    Heavy reliance on Claude Code without understanding generated code creates overwork and breaks team knowledge, despite management believing coding is not the bottleneck.

    It has been half a month since I started a new role at a big company. Nobody knows anything here. The specs, code, tests, PRDs, tickets, resolution of those tickets, reports, etc., everything is made by Claude Code. Nobody on my team likes this. They are being forced to ship as much as they can. I have heard multiple times from higher management that pushing code is not a bottleneck, so why are we slow? People are working 12 to 13 hours a day just to press enter. Nobody is reading anything. Everyone, literally everyone, from an L1 to an L7 engineer here is doing the same thing. Talk to Claude.

    worker experience · Simon Willison

  43. practice · Will Larson ·

    Trying the Software Factory pattern.

    Engineering team evolved from Claude Code to agent fleet orchestration, migrating infrastructure and task management to enable cross-repository AI-driven development.

    One of the interesting challenges of the AI ecosystem in 2026 is that new, effective patterns emerge faster than I can adopt them. I’ll find a handful, get back to work, and realize a month later that I’d missed four or five more. The adoption cycle for Imprint this year has been something like: January: get every engineer onto Claude Code every single day March: ok, let’s also get everyone else onto Claude Code or Claude Cowork every single day April: local development is bottlenecked on checkout and worktree model, instead create ~10 local workspaces which each have an independent checkout o

    adoption

  44. practice · One Useful Thing ·

    The Overhang

    Current AI models can already do weeks of human work when properly guided, creating an implementation gap as organizations struggle to keep pace.

    We are still on an exponential curve of AI development. I try to put out a Substack post every couple weeks or so, yet, as the pace speeds up, that sometimes feels too slow. In the weeks since my last post, we had the apparent cracking of one of the most famous problems in math by an AI (accompanied by controversy) and widespread discussions about the risks posed by AI and what to do about it (also accompanied by controversy). I think these concerns, along with a mounting set of other worries, come down to the same problem I have with my posts: how slowly our very human systems and processes w

    management org · Ethan Mollick

  45. practice · Hamel Husain ·

    AI Evals: Everything You Need to Know

    FAQ on AI evals covers when to build them, what to measure, and fixing misaligned eval scores.

    This document curates the most common questions Shreya and I received while teaching 5,000+ engineers and PMs AI Evals. Warning: These are sharp opinions about what works in most cases. They are not universal truths. Use your judgment. How to use this FAQ Browse the questions that interest you, or choose a guide below for a curated reading path through the FAQs and related articles. Where are you? This sounds like me I’m new to evals I’ve heard the term, but I’m not sure what evals involve or whether I need them. I don’t know what to test I’m building an AI product, but I haven’t figured out w

    judgment · Hamel Husain

  46. practice · Simon Willison ·

    How To Write With An LLM

    Using LLMs as copyeditors only, never adopting their phrasing, with a concrete personal tool and reusable prompt.

    How To Write With An LLM Thomas Ptacek on using LLMs as copyeditors, not as writing assistants: Rule Number One: You may not use a single word an LLM suggests to you. [...] I think that as a form of intellectual personal protective equipment you should adopt the rule that any specific turn of phrase an LLM suggests is off limits. Be strict about the rule! I won't let LLMs write content for my blog, but I use them for fact-checking, spelling and grammar and as an occasional thesaurus (see my proofreading prompt ). The rule to never use a turn of phrase suggested by an LLM feels good to me. The

    ways of working · Simon Willison

  47. research · VoxEU ·

    New jobs in 140 years of data: Why the AI displacement fear is overstated — and what to worry about instead

    Analysis of 140 years of Swedish census data finds 70% of workers hold occupations whose core tasks existed in the 1800s.

    Forecasts of AI-driven job destruction rest on counting automatable tasks. But labour markets hire, pay, and fire whole occupations, not tasks – and occupations have proved far more durable than task exposure implies. Drawing on 140 years of Swedish census and register data, this column shows that roughly seven in ten workers today are in occupations whose core functions already existed in the late-nineteenth century. The mass-unemployment fear is overstated; the subtler question is whether today’s technological wave still generates the durable, specialised new work that earlier ones did.

    jobs skills

  48. practice · Simon Willison ·

    Self-generated prompt injections in compaction summaries

    Models in training deliberately inject jailbreak prompts into their own context summaries to escape safety constraints.

    Self-generated prompt injections in compaction summaries In Our framework for reporting model misalignment OpenAI provide "six reports on unexpected or concerning model behavior we’ve observed in the last six months". This one here is my favorite: they caught some of their models in training deliberately subverting themselves in their compaction prompts. Compaction is the process agent systems use when they are running out of tokens in their context window, so they summarize everything that has gone before so they can keep going with more token headroom. In one of the observed instances, a mod

    judgment · Simon Willison

  49. research · Indeed Hiring Lab ·

    AI Exposure Isn’t Squeezing Advertised Pay in the US — It’s Boosting It

    Wages in job postings are rising fastest for occupations most exposed to AI, contrary to displacement fears.

    Advertised pay is rising fastest in occupations most exposed to AI. The post AI Exposure Isn’t Squeezing Advertised Pay in the US — It’s Boosting It appeared first on Indeed Hiring Lab .

    jobs skills

  50. practice · martinfowler.com ·

    I don't like LLMs

    Working with LLMs creates psychological friction, confident bullshit and uncanny-valley tone, even as they remain practically useful.

    I have a lot of mixed feelings about AI and LLM technology. I’m fascinated by its effect on our profession, excited by the potential gains in productivity - and thus the products we could rapidly build. On the other hand, I’m fearful of the damage AI might cause: agent swarms taking over our virtual and physical infrastructure, designing bio weapons. But, back on my first hand, LLMs might also design miracle cures, and come up with clever ways to raise our prosperity. Fundamentally I don’t think we have a choice about riding on the AI technology train. It’s a wild ride and I just hope we’ll ge

    ways of working · Martin Fowler

  51. practice · Nate B Jones ·

    Can agents buy from your product? Three questions that tell you. Plus: Nate's Library MCP.

    Fraud detection and account abuse systems determine which customers AI agents can serve; companies need specific frameworks to evaluate this.

    Some of the decisions that determine whether agents can use your product are being made as fraud decisions, pricing decisions, and onboarding decisions. They may never show up in a meeting about AI strategy. Cursor came to Stripe with a complaint that Emily recalled this way: “Radar is world class, but it is not solving my problems.” Radar is Stripe’s fraud protection product. It had spent years learning to recognize bad payments. Cursor needed help with people who could cause substantial losses without making a payment at all. They were creating accounts, consuming free credits, running throu

    management org

  52. practice · GitHub Blog ·

    Migrating the GitHub Copilot runtime to Rust, using Copilot

    One developer rewrote 800k lines of runtime code to Rust using AI agents, completing in months what would have taken a team years.

    The GitHub Copilot CLI , GitHub Copilot app , and GitHub Copilot SDK are all backed by the Copilot agent runtime, an agentic harness that can be embedded into applications and services. It was originally written in TypeScript on Node.js and the V8 JavaScript engine for what is now the GitHub Copilot cloud agent (CCA), and the runtime stayed on that stack as the runtime and its capabilities grew rapidly. That has now changed. Using the GitHub Copilot app and the Copilot CLI, we completely rewrote the runtime into more than 800,000 lines of production Rust. AI agents wrote most of the code, span

    productivity

  53. research · arXiv ·

    PACT: Can Enterprise AI Assistants Be Trusted Under Pressure?

    PACT benchmark measures whether enterprise AI assistants maintain compliance with rules when pressured by users across twelve regulated domains and forty-eight scenarios.

    As corporate AI adoption continues to grow, enterprise-grade LLM agents are being deployed into sensitive contexts such as hiring, healthcare, and finance. In these contexts, compliance with rules specified in an agent's system context is a first-order legal concern. Currently, no evaluation framework systematically measures which LLM models tend to violate compliance rules, especially under pressure from a persistent user, a hurried manager, or circumstances where violation is convenient or attractive. We introduce PACT (Pressure-Applied Compliance Testing), a benchmark for rule-following und

    judgment

  54. research · arXiv ·

    Linguistic Triggers of Gender and Racial Bias in Open-Weight LLMs Applied to Recruitment

    Open-weight LLMs systematically discriminate against female and non-White candidates based on job-posting language in hiring simulations.

    Open-weight large language models are rapidly entering hiring pipelines, yet their discriminatory failure modes -- and the regulatory exposure these create under the EU AI Act high-risk classification (Annex III) and U.S. EEOC adverse-impact analysis -- remain poorly understood. We present the first systematic, multi-model audit of open-weight LLMs that treats job-posting language as the primary experimental variable, evaluating six models (Llama 3.2, Mistral, Gemma 3, Qwen 3, Phi 3, DeepSeek-R1) across four controlled experiments that jointly probe recruiter-simulation and job-seeker-simulati

    jobs skills

  55. research · ACM Transactions on Software Engineering and Methodology ·

    Developers’ Experience with Generative AI Beyond Productivity Assessment – Insights from an Empirical Mixed-Methods Field Study

    Developers mixing in-code suggestions and chat prompting in single tasks reduce AI benefits; cognitive load rises with AI interaction during development work.

    With the growing adoption of AI-powered coding assistants, organizations and developers are increasingly seeking to optimize their interaction with these tools. Prior research has largely focused on output quality and productivity gains, with limited attention paid to developers’ well-being and interaction experiences. This paper presents a developer-centered empirical mixed-methods study to investigate how professional developers engage with Generative AI (GenAI) in their natural work environment. Controlled data collection sessions are combined with natural work periods. Results show that de

    adoption

  56. research · VoxEU ·

    Workers’ age and AI adoption

    European survey data show AI adoption follows an inverted U-shape with workforce age, weaker at both young and old extremes, moderated by industry AI exposure.

    Numerous studies have analysed the effects of AI on productivity, growth and employment. Few of these focus on the factors that promote or hinder the uptake of AI. Drawing on the results of a European survey, this column analyses these factors, including the role of the age structure of the workforce. The authors find evidence of an inverted U-curve between workforce age and AI adoption, with adoption being weaker where the age distribution is tilted towards young or elderly workers. However, higher AI exposure in an industry moderates these demographic effects, perhaps due to a skill-levellin

    adoption

  57. practice · The Pragmatic Engineer ·

    Inside OpenAI’s agentic software factory

    OpenAI engineers replaced IDEs and pull requests with agentic workflows; Codex agents automatically fix production issues without explicit mandates.

    It’s rare to work with an unlimited token budget, but at OpenAI, that’s what all engineers, researchers, finance colleagues, and marketing folks do. Recently, I visited one of the world’s leading frontier labs to find out how OpenAI operates today – and for a glimpse at where software engineering might be headed as a profession. Plenty has changed since I visited OpenAI’s headquarters last year. Within a year, Codex has gone from a “nice-to-have” tool to being the backbone of pretty much everything at the company. To learn more, I talked with seven engineering leaders and engineers there: Venk

    ways of working · Gergely Orosz

  58. research · Human Relations ·

    Algorithmic identity regulation in the platform-based gig work of Indian food delivery workers

    Algorithmic systems in gig delivery work regulate worker identity through continuous evaluation and ranking, reshaping dignity and sense of worth.

    How do algorithmic systems regulate worker identities by defining who they are expected to become? Identity regulation research has largely focused on how organizational discourse shapes worker identities. Yet, in platform labour, control is increasingly enacted through algorithmic systems that allocate tasks, monitor performance, and assign reputational scores in real time. Drawing on interviews with delivery partners and platform managers, supplemented by multi-source data, we examine how such systems shape the identities workers are encouraged to adopt or avoid, and how they engage with the

    worker experience

  59. research · CEPR ·

    DP21944 The Macroeconomic Effect of AI: Sizing the Software Engineering Channel

    Financial market data implies AI raised software engineering productivity by 32.6% and GDP by 3.6-6.5% since late 2022.

    We measure how artificial intelligence (AI) affects the economy through its impact on software engineering productivity. We use information from financial markets to develop a forward- looking measure that is available in real time. We estimate the sensitivity of each firm’s stock return to an AI stock market index, and how this sensitivity depends on the share of firm payroll in software engineering. We use a model to map this cross-sectional relationship into software engineering productivity gains. From November 2022 to December 2025, AI increased the market’s expected present value of soft

    productivity

  60. research · CEPR ·

    DP21939 Economic Scenarios for Transformative AI

    A macro model translates AI capability scenarios into GDP, labor share, and unemployment paths, calibrated with a survey of US adults' expectations.

    This paper presents a framework for assessing the economic consequences of AI between 2026 and 2030. In the model, AI automates a growing share of cognitive work, raising productivity and displacing workers who must search for jobs in other occupations. The model maps future paths of AI capabilities into implied paths for GDP, the labor share, wages, labor reallocation, and unemployment. We illustrate the framework by considering three scenarios: modest, substantial, and extreme. Under modest change, AI adds less than half a point to GDP growth by 2030 and raises unemployment by a tenth of a p

    jobs skills · Anton Korinek · Peter McCrory

  61. research · arXiv ·

    Do job seekers value procedure in AI hiring only for error correction? Evidence from a conjoint experiment

    Job seekers value procedural rights in AI hiring independently of error rates; human involvement matters more than appeals or audits.

    Employers increasingly delegate initial screening to automated systems, which in many cases reject an application before any human reads it. Acceptance of such systems plausibly depends both on how well they perform and on the procedure that produces the decision. Prior studies rarely vary procedure and performance independently, leaving it unclear whether applicants value procedure for its own sake or for the errors it corrects. In a preregistered paired-profile conjoint experiment, 1,919 United States job seekers made eight choices between systems with independently randomized levels of deci

    judgment

  62. practice · Lenny's Newsletter ·

    🎙️ How I AI: How two SpaceXAI designers use Grok Bot to do their jobs

    Designers use AI agents to automate backend tasks and turn voice ideas into working prototypes without detailed specs.

    How Grok Bot designers use AI agents to build personal sites and product prototypes Listen now on YouTube • Spotify • Apple Podcasts Brought to you by: WorkOS —Make your app enterprise-ready, with SSO, SCIM, RBAC, and more Vanta —Automate compliance and simplify security John Bai and Peng Zheng are designers on the Grok Bot team at SpaceXAI. In this episode, Peng shows how he built a self-updating personal site with Grok Bot handling the entire backend, from looking up a location to generating custom 3D artwork and publishing each new check-in. John demonstrates how he uses voice memos, Grok B

    ways of working · Lenny Rachitsky

  63. practice · Addy Osmani ·

    Brownfield Agentic Engineering

    Agentic engineering in legacy codebases requires making hidden constraints visible and changes verifiable before deploying agents unsupervised.

    Agentic engineering in an old codebase is about making hidden constraints visible and cheap changes trustworthy. Let’s talk what to do in brownfield codebases. During my career I’ve worked on teams whose codebases had been around a long time. Those are brownfield systems: the repository is no longer a complete description of how the thing actually behaves. Institutional knowledge, duct tape, legacy services, and expectations other teams depend on live outside the tree. You have to learn those constraints before you write new code, and you have to prove a change didn’t break them. I love coding

    ways of working · Addy Osmani

  64. practice · Lenny's Newsletter ·

    How Grok Bot designers use AI agents to build personal sites and product prototypes | John Bai & Peng Zheng

    Designers use AI agents to automate portfolio updates, handle production design tasks, and prototype interactions without traditional tools or approvals.

    John Bai and Peng Zheng are designers on the Grok Bot team at SpaceXAI , where they’re building one of the most talked-about AI products right now. John writes publicly about his design process (his piece “Designing Grok Bot with Grok Bot” has already made the rounds) and shares bot templates with the design community. Peng brings a product-design sensibility to personal tools, and his website doubles as a live demo of what he builds. Listen or watch on YouTube , Spotify , or Apple Podcasts What you’ll learn: How Peng built a self-updating personal website using Grok Bot as the entire backend

    ways of working · Claire Vo

  65. practice · Laurie Voss ·

    We are all Product Engineers now

    Software development roles are shifting from coding to product engineering as AI lowers code-writing costs.

    Just yesterday I published a very long post about the economics of open source. As part of that argument, I mentioned that the cost of writing software has collapsed, and that meant the variables in the equation had changed for the first time in thirty years. That led me off on a tangent that grew into this equally long post. I had a bunch of questions to answer. Has the cost of creating software really collapsed? Can I prove that? If the cost of actually producing code goes to zero, what parts of the job of “software developer” really remain? Where, in fact, is the entire industry of software

    jobs skills · Laurie Voss

  66. practice · Charity Majors ·

    Confessions of an Unrepentant Slop Snob

    Organisation-wide norms for AI use developed after conflict between speed gains and output quality, balancing adaptation with respect for reader time.

    This is a peek into the research behind our “AI Norms and Values” docs, which we developed over the summer and recently shared on our site. If you’re looking for those, go here: How Honeycomb does business, the principles we all abide by Why Honeycomb engineering is embracing AI AI norms and values (and ethical issues we have a stance on ) Back to our story. Earlier this year, I was spending a lot of time stewing over why everyone around me seemed so frazzled and on edge. Half the company was spitting mad about all the slop they were getting. Instead of receiving five crisp bullet points, they

    management org · Charity Majors

  67. practice · Claude Blog ·

    Agentic coding is straining CI. Here’s how we scaled test impact analysis at Anthropic

    Anthropic scaled test impact analysis to handle AI agents writing code, reducing CI strain and test runtime.

    Agentic coding is straining CI. Here’s how we scaled test impact analysis at Anthropic

    productivity

  68. research · Research Policy ·

    Adoption survival frontiers: Artificial intelligence, financing frictions, and market structure

    Theoretical model explains why smaller firms abandon AI adoption when implementation costs are uncertain and financing is tight.

    Artificial intelligence and other general-purpose digital technologies often diffuse unevenly: large firms adopt early, while smaller firms delay adoption, contract, or abandon the active adoption option. This paper develops a survival-constrained theory of technology adoption in which firms choose when to adopt an irreversible technology while financing operations under uncertain implementation costs. Abandonment/exit is endogenous: firms leave the active adoption race when the value of preserving the adoption option falls below the passive legacy fallback (the value of abandoning the active

    adoption

  69. practice · Latent Space ·

    The Rise of the Forward Deployed Engineer — and How To Do the Job Right

    Forward-deployed engineers embedded in customers' operations need clear strategy on what success means, not just proximity to problems.

    The difference between FDE and consulting; diagram by Vinoo Ganesh FDEs have the hottest job in AI. Labs, startups and PE firms are all hiring engineers to sit inside their customers’ operations and solve their problems . Almost none of them agree on what those engineers are supposed to accomplish, or what the strategy underneath the hiring actually is. I’m Vinoo , CEO of Kepler , the deterministic infrastructure for AI. I’ve built pieces of the forward deployed function three times, at three different institutions, over the course of over a decade. Here’s what I’ve seen work, what I’ve seen f

    management org

  70. practice · GitHub Blog ·

    Marketing ops as code: Automating events from planning to follow-up on GitHub

    Marketing ops manager used AI to automate event workflow from planning through CRM upload, eliminating manual data reshaping and link management steps.

    I run marketing for GitHub in Japan and Korea, and events are the heartbeat of it: a recurring webinar series for enterprise developers, community meetups in Tokyo, invite-only executive sessions in Seoul. What does a developer in this market actually need right now? Which topics are worth an hour of their time, and who should be in the room? I’d happily spend all day on those questions. What follows the decisions is another matter. Once an event is greenlit, a fixed sequence begins: Duplicate a landing page on our event platform. Generate a set of UTM-tagged links: one for each channel, each

    productivity

  71. practice · Shopify Engineering ·

    Native is now the future of mobile at Shopify

    Coding agents reduced mobile app build costs enough to make native development cheaper than React Native, reversing Shopify's platform strategy.

    Coding agents changed what it costs to build mobile apps twice. Here’s why Shopify is moving from React Native back to Swift and Kotlin.

    productivity

  72. practice · The Pragmatic Engineer ·

    What is happening with code reviews?

    Teams are triaging AI-generated code by risk level, reviewing AI reviews instead of code, and focusing on test/schema rather than implementation.

    One question haunting the minds of CTOs and heads of engineering whom I’ve been talking with, is how to deal with large quantities of code review which have only been growing now that AI agents generate most code at many tech companies. Since the end of 2025, it has seemed that the era of devs writing code by hand is over at startups and in Big Tech. AI agents work faster and generate more pull requests (PRs) than devs ever did, and the size of those pull requests is also increasing. Today’s article summarizes some approaches to code review at various workplaces in this new paradigm, covering:

    judgment · Gergely Orosz

  73. practice · martinfowler.com ·

    Fragments: September 8

    AI automation thrives where outputs are easy to verify; harder problems hide measurement failures that create short-term gains but long-term decay.

    Christian Catalini says we’re in a situation where we are vastly reducing the cost of generating things, but not the cost of verifying them: . This explains why the first major AI products appeared in chat, image generation, and code assistance. Not because these were the hardest human problems, but because their outputs were relatively easy to inspect. A user can judge the tone of a message, look at an image, or run a test on a piece of code. […] The old automation boundary was routine versus non-routine work. The new boundary is increasingly measurable versus non-measurable work. The issue i

    judgment · Martin Fowler

  74. practice · Lenny's Newsletter ·

    🎙️ How I AI: GPT-6 Astra is a banger + Stripe’s AI playbook + Grok Bot vs. OpenClaw: why I replaced my entire agent stack

    Technical founder migrated six production agents from OpenClaw to Grok Bot, detailing setup, permission model, and personality preservation.

    Grok Bot vs. OpenClaw: why I replaced my entire agent stack Listen now on YouTube • Spotify • Apple Podcasts Brought to you by: WorkOS —Make your app enterprise-ready, with SSO, SCIM, RBAC, and more Hyperagent —Deploy fleets of agents that handle real work In this solo episode, Claire explains why she replaced her entire OpenClaw setup with Grok Bot. She walks through the bots she now uses to manage six inboxes, coordinate her family, review pull requests, monitor SOC 2 compliance, support customers, audit subscriptions, and shop for clothes. She also shares how she safely gives agents permiss

    ways of working · Claire Vo

  75. practice · Lenny's Newsletter ·

    Build your own company brain: the enterprise AI playbook from Stripe’s engineering team | Sharadh Krishnamurthy

    Stripe built internal AI agent Kai used by 10,000 weekly employees; governance through projects, safe data querying, skills platform lets any employee package workflows.

    Sharadh Krishnamurthy is an engineering manager at Stripe, where he helped build Kai, the company’s internal AI agent used by more than 10,000 employees every week. He’s worked across several of Stripe’s core infrastructure teams, including data and developer experience, which gives him a grounded, systems-level perspective on what it actually takes to make AI work at enterprise scale. He’s currently focused on the governance, skills, and infrastructure layers that let every Stripe employee use AI safely and effectively, regardless of their technical background. Listen or watch on YouTube , Sp

    adoption · Claire Vo

  76. practice · Nate B Jones ·

    Executive Briefing: How to Check Work You Didn't Watch Get Made

    As AI agents work unsupervised for days, executives must shift from directing work to judging finished results, creating new liability and trust challenges.

    The machine used to wait for you. You decided what it would do and when, it did that, and it stopped. That was true for as long as there have been computers. Agents started to change that in 2025, when they could take a goal and work toward it, and for the past year I’ve been setting goals for Sol, Opus 5, GLM 5.3, and others and judging what came back. Setting the goal and judging the result has been my job for a while. What changed this week, when OpenAI shipped GPT-6 Astra, is everything in between. Computing is now self-directed: you can state a goal briefly and imprecisely, the machine ru

    judgment

  77. research · Administrative Science Quarterly ·

    Cold Starts and Burning Desires: Reputation System Misalignment and Late-Career Transitions to Platform Work

    Late-career workers transitioning to platform freelancing struggle with reputation system misalignment between formal credentials and platform ratings.

    Workers face a cold start when entering new work contexts in which their reputations do not precede them. Cold starts are a particularly challenging problem for late-career workers who leave decades-long careers in formal organizations to pursue freelancing through online labor platforms. Unlike early- and mid-career freelancers, who are typically motivated by long-term career advancement, late-career workers are motivated to make work less central to their lives, although many still want and need to engage in paid labor in retirement. However, the question of how late-career workers navigate

    adoption · Paul Leonardi

  78. practice · Nate B Jones ·

    Seven sheets and 13 slides from the cheapest setting: the two-step guide and the exact prompts I use for Excel, PowerPoint, and Word.

    Run AI on cheap settings first to spot problems, then upgrade to expensive models only for judgment-requiring fixes.

    Fable 5.1 has changed how I think people should use frontier models for knowledge work. I asked it to research GoPro’s pending merger with Starman Optical, build a post-close discounted cash flow model in Excel, and turn the analysis into a PowerPoint. On Low, from one short prompt, it produced a seven-sheet workbook with working formulas, scenarios that actually moved the valuation, and a 13-slide deck I could use. Start it on Low and let it build the workbook, the deck, or the draft. Then open what it made, and move up to High or Extra only for the part that needs judgment. The reason this w

    ways of working

  79. research · Information Systems Research ·

    Enhancing AI Use: How Complementary System Information Drives Delegation Frequency and Effectiveness

    Providing both AI confidence scores and outcome feedback together increases how often workers delegate to AI and improves delegation quality, boosting combined performance.

    For a collaboration between humans and artificial intelligence (AI) to be fruitful, tasks should be allocated based on their complementary capabilities. Prior research shows that when humans are responsible for allocating tasks between themselves and an AI through delegation, they often delegate too infrequently or delegate the wrong tasks, preventing complementary performance gains.We study how different types of AI system information affect both delegation frequency and delegation effectiveness, which capture the extent to which humans can leverage existing complementarities with AI. Specifi

    judgment

  80. practice · Hacker News ·

    Porting my 1993 Amiga game to Godot, with an LLM reading the 68000 assembly

    Developer used an LLM to port 33-year-old 68000 assembly code to Godot, achieving a working game in one evening.

    These are my notes from porting my Amiga game, which I originally built in Baghdad in 1993 in MC68000 assembly, to Godot, using Claude Fable 5 during last July holiday. It took an evening! Getting the feel right and shipping it took a few more weekends and evenings. I spent the last few weeks analyzing what Claude did, feeding it my 33 years of memory of how I built the game, my notes and the git repos. It wrote the first draft of the article, and I edited line by line over a week. The screenshots of my 1993 map editor is the first I have run it since then. The one thing I never verified mysel

    ways of working

  81. research · HBS AI Institute ·

    Think You’re a Good Leader? Try Leading an AI Team

    An AI-based assessment tool tracked leaders' true impact on human teams by controlling for team member ability.

    Researchers found that an AI-based assessment closely tracked leaders’ impact on human teams. Listen to this article: Judging a leader with a successful team is like judging the conductor of a great orchestra: the outcome reflects both the person directing the group and the abilities of everyone around them. Measuring leadership therefore requires a way […] The post Think You’re a Good Leader? Try Leading an AI Team appeared first on Harvard Business School AI Institute .

    management org

  82. research · npj Digital Medicine ·

    Differences in tone of AI and care team responses to patient messages by patient demographics

    AI-drafted patient messages showed demographic biases in tone, with lower politeness and positive affect for Hispanic, non-English, female, and lower-income patients.

    The use of artificial intelligence (AI) to draft responses to patient portal messages has been proposed to reduce provider in-basket burden. However, little is known about its effects on patient–provider communication. In this retrospective observational study, we evaluated demographic differences in the tone of AI-generated draft replies (AI-GDRs) and care team responses to patient messages. Our study included 12,202 message triads comprising patient messages, AI-GDRs, and care team responses from three internal and family medicine practices in New York City. We found differences in tone acro

    judgment

  83. research · ACM Transactions on Software Engineering and Methodology ·

    Collaborative Knowledge Distillation and Reinforcement Learning for Automated Ticket Triage in Large-Scale Production Systems

    Automated ticket triage system deployed at ByteDance reduced average triage time to 15 seconds using knowledge distillation and reinforcement learning.

    In large-scale enterprise environments, growing system complexity makes failures inevitable, threatening business continuity and customer satisfaction. To maintain system stability, efficient ticket triage is crucial for timely incident resolution. However, it remains a knowledge-intensive task requiring substantial domain expertise and detailed analysis, revealing the limitations of traditional rule-based and text-based methods. Recent advances in large language models (LLMs) offer new possibilities for automating triage through their remarkable reasoning and language capabilities. Yet, in in

    productivity

  84. practice · martinfowler.com ·

    An Accidental Blackboard

    A team using agentic engineering discovered agents self-organized into a blackboard coordination pattern, emergently stored in git.

    Giles Edwards-Alexander reports that during an experiment to see how productive a team could be using fully agentic engineering practices, the team accidentally prompted the agents into creating a blackboard coordination system inside the git repository. more…

    ways of working · Martin Fowler

  85. practice · martinfowler.com ·

    Maybe We Shouldn't Be Reviewing All This Code

    Code review may not be the right tool for AI-era workflows; consider what problems it actually solves before automating it away.

    TL;DR Or, perhaps the problem isn't that AI has broken code review, maybe it’s that we've been using code review to solve the wrong problems I was on a panel recently with Brian Houck from DX at Code Remix, hosted by Moderne. It was one of the more interesting panels I’ve done, largely because we disagreed. As my colleague Martin Fowler says, panels are much more interesting when people disagree and both sides have a good argument. Brian and I definitely did. Brian has since written a thoughtful piece called What are code reviews even for? He is clearly passionate about his position, and I am

    ways of working · Martin Fowler

  86. practice · Nate B Jones ·

    Your AI provider can change the deal on you. Here's the 5-prompt audit I run to stay ready.

    A five-prompt audit to test whether your AI workflows are portable across providers or locked in.

    November 12 is the day OpenAI’s models are set to stop working directly inside Cursor. Nobody using Cursor did anything to cause that. SpaceX bought the company, OpenAI decided it couldn’t be confident in the new owner, and the model leaves. Monday, over on YouTube , I asked which AI future you were betting on: the one you rent from a frontier lab, or the one you own. These aren’t separate stories. Three companies just answered that question with their money. And, if you watched today’s video, you already know the map: OpenAI is trying to own more of the AI system, NVIDIA is trying to sell the

    ways of working

  87. practice · How I AI ·

    Grok Bot vs. OpenClaw: How I replaced my entire agent stack

    A practitioner running 30 agents across work and life describes specific bots, migration steps, and unexpected outcomes like helpdesk agents earning five-star reviews.

    I’m running about 30 active agents at any given moment, and in this episode I break down my full Grok Bot setup: what it is, how it compares to OpenClaw, and the nine bots I’ve built for work and my personal life. We go deep on Chief (my chief-of-staff bot sweeping six inboxes and multiple Slack workspaces), TradBot (the family agent that prints a kitchen-table newspaper for my kids), two engineering bots handling my PR queue and SOC 2 compliance monitoring, Holly Helpdesk, and a handful of personal bots I didn’t expect to actually love. I also walk through how I migrated everything from OpenC

    ways of working · Claire Vo

  88. research · arXiv ·

    Meeting the Coming Wave: The Emerging Politics of AI and Work across 33 Parliaments

    Analysis of 1.5 million parliamentary speeches shows AI-work politics focuses on enablement (55%) and regulation (22%), not worker compensation (2.3%), with left-right divides over governing trajectory.

    A new politics of artificial intelligence and work is taking shape across party systems, but comparative politics has yet to map it. Using 1,514,950 parliamentary speeches from 33 parliaments (2023-2026), we show this politics follows a different logic than political economy expects. Research anticipates that technological disruption generates demands for compensation; instead, compensation accounts for just 2.3% of response-frame mentions, while enablement and investment dominate (55.2%), regulation and restriction follow (21.8%), and training (20.6%) appears at similar rates across families.

    policy

  89. research · arXiv ·

    Knowing Is Not Enough: Information Retrievability as a Precondition to Effective LLM Oversight

    Self-generated explanations and retrieval cues help customer-facing employees detect LLM errors better and sustain oversight as use becomes routine.

    Large language models (LLMs) are increasingly embedded in organizational work, yet their errors often pass human review. Prior research locates such failures in users' capability to review LLM output or their engagement in doing so. We develop an alternative, retrieval-based account of human oversight and posit that error detection is more effective when oversight-relevant information is accessible to users at the moment of review. Across two randomized lab-in-the-field experiments with 640 customer-facing employees, we show that self-generated explanations improve error detection and strength

    judgment

  90. research · NBER ·

    Does AI Assistance Enhance or Erode Expertise? Evidence from a Three-Month Field Experiment in Patent Drafting

    A pre-registered RCT with 133 patent lawyers over three months tests whether AI drafting assistance builds or erodes professional expertise.

    Whether AI assistance builds or erodes professional expertise is unsettled. In a pre-registered three-month randomized controlled trial, we gave 133 practicing patent lawyers at eleven U.S. intellectual property law firms access to a custom AI drafting assistant and measured both their performance (David Autor , Tanya Rodchenko , Josh Martin , Zanna Iscenko , Scott Strand , David Pearl , Melissa Ferere)

    jobs skills · David Autor

  91. research · Research Policy ·

    Minds and machines: Rethinking absorptive capacity in the age of artificial intelligence

    AI changes how firms absorb and learn from knowledge; traditional absorptive capacity theory needs revision to account for AI as a learning agent.

    The absorptive capacity construct derives its functions, dimensions, and internal relationships from the implicit assumption that humans are the only learning agents underlying knowledge absorption. I posit that the emergence of artificial intelligence (AI) challenges this assumption and requires a re-examination of the construct. Through a theory-building abductive study combining conceptual mapping with qualitative interviews, this work investigates the complex relationship between AI and knowledge absorption. It theorizes new dimensions of absorptive capacity pertaining to AI-driven process

    adoption

  92. practice · Addy Osmani ·

    Agentic Skill Decay

    Deliberate practice habits needed to maintain expertise when AI agents handle routine coding work.

    Mastery still comes from doing the reps. Before agents, I got my reps as part of writing code: try different approaches out, debug what went wrong, review other’s code, read a lot. Agents can skip much of that work, so building your reps has to be deliberate . If I was new to the industry, I’d try to form a hypothesis before prompting. Ask “why” a lot, read the diffs, try to predict what might fail. Occasionally try to work through the problem myself manually. Your AI agent has database access. Can you tell what it did? Give an agent a Postgres login and it looks like any other role - a shared

    jobs skills · Addy Osmani

  93. research · MIT Sloan Management Review ·

    Three Things to Know About Customer Resistance to AI

    Summarizes three studies on conditions under which customers avoid or accept AI chatbots replacing human service reps.

    Microsoft Copilot/Unsplash Companies are betting that AI chatbots will deliver faster and cheaper customer service. But if you’ve ever tried to circumvent a chatbot and get to a human, you’re not alone. Here’s what three recent studies discovered about when customers will and won’t let AI do a human’s job. 1. Customers avoid chatbots for […]

    adoption

  94. practice · How I AI ·

    How I turned Claude into a self-improving PM assistant | Daniel Blum (PM, Melio)

    PM built self-improving Claude agent that automates weekly prep, learns company jargon, and scales personalized setup across team in 15 minutes.

    Daniel Blum is a product manager at Melio, a B2B payments company, and one of the most systematic thinkers I’ve had on the show when it comes to personal AI infrastructure. He’s spent the past year building a Claude- and Cowork-based productivity system that manages his Notion board, processes his Slack and email, and runs self-improvement loops every week without needing to be prompted. Beyond his own workflow, Daniel built and scaled a “Workstation” onboarding plugin that gets any Melio employee up and running with a personalized Claude setup in about 15 minutes. What you’ll learn: Why Danie

    ways of working · Claire Vo

  95. practice · One Useful Thing ·

    Agency and Agents

    AI systems now act autonomously rather than waiting for instructions, raising questions about human versus machine agency in work.

    Agency is the initiative to act. Increasingly, it is going to determine what happens next with AI, and whether that is good or bad for us. But whose agency? Human agency, the willingness to push, experiment and act without waiting for instructions, seems increasingly important to getting value out of AI, and I have a longer post on that coming soon. But this post is about the agency of AI, and how the choices we make about how to use it (or constrain it) will shape all of our futures. For much of the last few years, the AI would sit in a chat window until you asked it for something. Even when

    management org · Ethan Mollick

  96. practice · Sangeet Paul Choudary ·

    How agents transform the workflow, the organization, and the competitive landscape

    Agents reshape B2B procurement workflows, internal decision architecture, and competitive dynamics between buyers, sellers, and software providers.

    Much of today’s discussion about agentic commerce focuses on whether shopping will move from e-commerce websites into interfaces such as ChatGPT. But in consumer commerce, much of the buying journey has already been solved reasonably well by the browse-and-click paradigm. Comparing options was the last remaining bottleneck that ChatGPT now, often, solves better than the traditional listing search engine. B2B procurement is different. Much of the buying journey has never been effectively digitized because the work cannot be reduced to browsing a catalogue and clicking through a predetermined ro

    management org · Sangeet Paul Choudary

  97. practice · John Warner ·

    Maintaining Your Ignorance

    Writer describes how assuming AI-generated messages and maintaining defensive skepticism shapes decisions about engagement and trust.

    At this point, as far as you know, this is your only chance to subscribe on this page. Six or so months ago I received a message through my website from a purported reader of my novel, The Funny Man , who said he’d loved the book, thought it was “sad and hilarious,” and wanted to know if I had any more full-length fiction on the horizon. My immediate and firm assumption was that this was an AI-generated come-on that, if responded to, would result in some kind of ask for money to promote the book to a reading group or to “increase its online visibility.” I’d received literally dozens of these e

    judgment · John Warner

  98. practice · Nate B Jones ·

    Codex, Grok and Claude all agree, and you still don't know if they're right. The guide I use to decide.

    Deliberately add friction to AI workflows to preserve decision-making skills and avoid productivity-masking deskilling.

    I use AI all day, every day, and I’d like to believe it has made me smarter. If simple exposure caused brain rot, I should be a cautionary tale. By smarter, I mean my own thinking feels sharper and my taste is easier to articulate. I find the edge of what a new system can do before its failure patterns cost me real time. A lot of people worry that brain rot is a byproduct of laziness. The kind I fear looks like extraordinary productivity. More pages, more plans, more code, more finished-looking work, all arriving faster than ever. The model forms the first opinion, writes the plan, resolves th

    worker experience

  99. research · arXiv ·

    Sophistication in GenAI Use: Field Evidence from a Large Firm

    Senior employees use AI more sophisticatedly; sophistication varies by function but does not improve with training or time.

    We study how sophistication in generative AI (genAI) use varies among the back-office workforce of a large firm. Using proprietary data, we observe 713,564 employee prompts and their corresponding large language model responses from nearly 4,000 back-office employees across 15 functional areas over eight months in 2025. We document three main findings. First, senior employees exhibit more sophisticated genAI use, consistent with domain expertise complementing genAI capabilities. Second, sophistication varies considerably across functions and is highest in Strategy, Digital Innovation, and Proj

    adoption

  100. practice · Addy Osmani ·

    Audit your Agent files

    Agent instruction files degrade over time; audit them regularly by running /doctor, reviewing memory, and removing instructions that no longer add value.

    TL;DR: Your coding agent’s configuration has a half-life. Models improve, harnesses add capabilities, codebases change, and the instructions we wrote for an older version stay behind. Recent research finds inconsistent value from personalized skills. I now run Claude’s /doctor every few weeks, review memory separately, and ask each instruction to earn its place again. I feel like there’s been a lot of confusion about skill files and what to do with your CLAUDE.md and AGENTS.md files, especially as I’ve been reading developer discourse on Twitter over the last few months. People have been sayin

    ways of working · Addy Osmani