The Feed · Complete archive

AI at work: research and practice

A server-rendered record of the evidence, ideas and firsthand practices screened by The Feed. Every entry links to its original source.

Use the interactive Feed

Page 4 of 9 · 873 items

  1. practice · How I AI ·

    How the engineer behind Claude Cowork actually uses Claude | Felix Rieseberg (Anthropic)

    Engineer demonstrates concrete Claude workflows: floor plan to 3D walkthroughs, Twitter promise tracking, hardware approval buttons.

    Felix Rieseberg is the engineering lead for Claude Cowork and Claude Code Desktop at Anthropic. He previously spent five years at Slack building developer tools. In this episode, Felix demonstrates how he uses Claude to solve real-life problems: analyzing floor plans to build interactive 3D house walkthroughs, automatically tracking promises he makes on Twitter, and building a $20 hardware device that physically approves Claude actions with a button press. What you’ll learn: How to use Claude Cowork to turn a 2D floor plan into an interactive 3D walkthrough where you can move furniture around

    ways of working · Claire Vo

  2. practice · John Warner ·

    Taste Is Human

    Literary judgment requires trained taste; AI-generated prose lacks the coherence that experienced readers detect beneath surface polish.

    By the time you get to the bottom you’ll be thinking, “Dang, I want to read more by that guy,” but the subscription box will be at the top and you’ll have to scroll up. Just sign up now. That Story Everyone Is Talking About Is Bad Some, perhaps many of you have heard of the scandal involving a prize-winning story from a contest sponsored by Granta that is almost certainly AI-generated. I have no wish to litigate the story’s provenance or quality. Others, including ( here ) and ( here ) have done this more thoroughly and perceptively than I ever could. The story is bad, with a sheen of surface-

    judgment · John Warner

  3. research · arXiv ·

    Generative AI and the Reorganization of Labor Demand

    Firms adjust to AI through hiring different roles (52% of exposure change) and redesigning tasks within jobs (39.5%), with bigger shifts in senior roles.

    Generative artificial intelligence (AI) is expected to transform work, but less is known about how firms reorganize labor demand as the technology diffuses. Existing research has largely focused on which occupations are exposed to AI or whether exposed jobs decline. We extend this debate by examining whether firms adjust by changing where they hire, what jobs contain, or both. Using a nationwide dataset of job postings in the United States, covering all sectors of the economy, we construct a dynamic, posting-level measure of generative AI exposure with a two-stage large language model pipeline

    jobs skills

  4. practice · Kent Beck ·

    Scope Is The Steering Wheel

    Speed culture pressures teams to abandon scope discipline; slowing down to manage scope deliberately improves outcomes.

    This stuck in my craw at the time Patrick published it but I didn’t have the energy to respond. Now, with the ever-increasing, genie-fueled emphasis on speed, it deserves a second look. Among its several flaws as a statement is that it misses one point that XP got right, a point that’s become leveraged. I’ll start gently, addressing the OP directly. Responsibility I don’t like the tone. Who exactly hired The Slow (may as well be honest & capitalize)? Who created the incentive system in which they operate? Take responsibility for your part in the situation you describe. You’re not above it all.

    management org · Kent Beck

  5. research · Management Science ·

    Mission Driven and Data Averse: How Empathy Fosters Resistance to Algorithms and Hard Data

    Social-mission organizations with empathy-focused cultures resist algorithms and data, but fairness framing can reduce this resistance.

    Socially driven organizations aiming to become more data driven often face practical barriers such as difficulties in measuring complex social objectives. We argue that, even when those practical barriers are addressed, social missions make it harder to foster a data-driven culture by exacerbating employee aversion to algorithms and hard data. Our theory is that emphasizing a social rather than profit-oriented mission increases employees’ concern for empathy, and this, in turn, makes them more averse to the cold, impersonal methods associated with using algorithms and hard data. Furthermore, i

    adoption

  6. practice · Latent Space ·

    Railway: The Agent-Native Cloud — Jake Cooper

    Railway scaled to 3M users running AI agents for deployment; coding agents replacing pull request workflows in production.

    3M Users, 100K Signups/Week, Own-Metal Data Centers, $200K+ Coding Agent Spend, and the Death of PRs 3M Users, 100K Signups/Week, Own-Metal Data Centers, $200K+ Coding Agent Spend, and the Death of PRs Take the 2026 AI Engineering Survey and get >$2k in credits and AIE WF tickets!

    adoption · Shawn Wang (swyx)

  7. practice · Paul Ford (Aboard) ·

    Dear Vibe-Coding CEO: Please Stop

    CEOs using AI tools to prototype without technical skills reveals a gap between executive agency and organizational capability.

    “The boss built a thing and he wants you to take a look.” These are the words echoing through hallways around the world, as executive leaders get their hand on vibe-coding tools. This kind of statement invariably sends a shudder through the engineers, designers, and product managers who know that they’re about to get shown something hacked together with Claude in a weekend—as they’re simultaneously asked why it’s taking them so long to ship their software. We’re hearing about it everywhere: There’s a plague of vibe-coding CEOs spreading across industries. I’m going to go against convention and

    management org · Paul Ford

  8. research · npj Digital Medicine ·

    Optimizing AI implementation for surgery: recommendations for infrastructure and deployment in the operating room

    Cloud-based AI in operating rooms achieved higher task completion but caused more physical strain, emissions, and costs than edge-based systems over 396 hours of use.

    Recent years have seen a surge in the development of Artificial Intelligence (AI) technologies to enhance intraoperative decision-making and improve patient safety. Yet, real-world evidence on their implementation remains limited. This study evaluated the impact of AI deployment in the operating room across key implementation domains-usability, cost, and carbon footprint-using cloud and edge-based infrastructures. Twenty-three end-users (surgeons, residents, fellows, operating room nurses) from a multi-site teaching hospital in Canada participated in system usability testing, assessed with val

    adoption

  9. practice · How I AI ·

    HTML is the new Markdown: How Anthropic engineers are building with Claude Code | Thariq Shihipar

    Using HTML instead of Markdown for AI planning and workflow communication creates richer visual artifacts that improve human engagement and code quality.

    Thariq Shihipar is an engineer at Anthropic working on the Claude Code team. He’s spent the past several months experimenting with HTML as a replacement for Markdown in planning and implementation workflows, discovering that richer visual formats lead to better human engagement—and, ultimately, better products. In this episode, filmed at Anthropic’s Code with Claude event in San Francisco, Thariq demonstrates how to use HTML artifacts to create interactive plans, build throwaway UIs for specific problems, and maintain living design systems that travel with your codebase. What you’ll learn: Why

    ways of working · Claire Vo

  10. research · MIS Quarterly ·

    Artificial Intelligence, Alliances, and Innovation1

    Firms with more AI resources shift drug innovation toward alliances and develop more drugs jointly, suggesting AI reduces information asymmetry in partnerships.

    This paper examines the influence of artificial intelligence (AI) on research and development (R&D) practices, proposing that AI enables firms to discover, evaluate, and make sense of information from their partners, thereby reducing information asymmetry and allowing alliances to flourish. Empirically, we approximate a firm’s AI resources using patents and job postings to analyze the relationship amongst AI, alliances, and drug innovation in the pharmaceutical industry. Our findings show that firms with greater AI resources shift their locus of innovation toward alliances and develop more dru

    management org

  11. research · arXiv ·

    Agentic AI and Human-in-the-Loop Interventions: Field Experimental Evidence from Alibaba's Customer Service Operations

    Field experiment in customer service shows AI reduces chat duration but lowers ratings; human intervention works better for technical than emotional escalations.

    Agentic AI systems that autonomously perform service tasks are entering customer service operations. However, limited evidence exists on how human interventions shape service outcomes when agentic AI failures create both cognitive and emotional consequences. We study this issue through a randomized field experiment on Alibaba's Taobao platform. Workers in the treatment condition supervised an agentic AI system that resolved AI-eligible chats while continuing to handle AI-ineligible chats, whereas control workers resolved all chats without agentic AI. The findings show that AI deployment reduce

    adoption

  12. practice · Paul Ford (Aboard) ·

    But I Love Stochastic Parrots!

    Paul Ford examines the 'stochastic parrot' criticism of LLMs and what it reveals about how people think about AI capability and consciousness.

    One of the famous criticisms of large language models is that they are “stochastic parrots”—that they randomly assemble language without understanding it. Something about that idea really sticks with people. I think it’s because of, well, parrots, which are fun birds. It brings to mind all kinds of similar creatures: Heuristic toucans, probabilistic osprey, or if we want to get wild and branch out from birds, possibly even syntactic squirrels. The original paper was by Emily Bender, Timnit Gebru, Angelina McMillan-Major, and Margaret Mitchell; it’s thoughtful and academic. This was the paper t

    judgment · Paul Ford

  13. research · Human Relations ·

    Algorithmic management and worker agency: The platform perspective on algoactivism

    Platform managers adapt algorithmic management systems in response to worker resistance, potentially narrowing worker agency over time.

    Do platform workers inadvertently strengthen the algorithmic management (AM) systems they seek to resist? Drawing on 37 interviews with digital labor platform managers, this paper provides an organizational perspective on workers’ “algoactivism” in food and grocery delivery. We reveal that platforms are not passive targets and their reactions to algoactivist acts unfold through a dynamic cycle of sensemaking across three stages: noticing, framing, and acting. We identify key determinants at each stage, explaining why platforms react differently to worker acts. Unlike traditional organizations

    worker experience

  14. research · arXiv ·

    AI in the Enterprise: How People Use M365 Copilot Chat

    Analysis of 5.5 million M365 Copilot sessions shows writing dominates usage, with growing share of analysis and decision-making tasks across occupations.

    M365 Copilot is used every week by millions of people across more than a million companies around the world as part of their workflows. Uniquely positioned in the AI landscape given its near-exclusive use for work purposes, M365 Copilot can offer a clear picture of how people use AI for work and where that usage may expand next. This paper characterizes that usage through direct classification of user interactions with M365 Copilot Chat. Based on an anonymized and privacy-preserving analysis of a sample of approximately 5.5 million sessions, we combine a learned classification of user intent w

    adoption · Sonia Jaffe · Siddharth Suri · Scott Counts · Kiran Tomlinson

  15. practice · How I AI ·

    Spec-driven development: The AI engineering workflow at Notion | Ryan Nystrom

    Engineers at Notion use spec-first workflow: voice-to-spec-to-agent-implementation in 20 minutes, with autonomous verification.

    Ryan Nystrom is a software engineer at Notion. He joined in December 2024 after Notion acquired Campsite, the team communication platform he co-founded with Brian Lovin. At Notion, he’s been a core builder of Notion AI and the Custom Agents feature launched in February 2026. He manages a team of six to seven engineers while still writing code himself, currently running Project Afterburner, a push to cut Notion’s CI time to a quarter of its current duration. What you’ll learn: How to build a Notion AI custom agent that auto-generates your daily standup pre-read by pulling from Slack, GitHub, Ho

    ways of working · Claire Vo

  16. research · METR ·

    Measuring the Self-Reported Impact of Early-2026 AI on Technical Worker Productivity

    Technical workers report median 1.4–2x value gains and 3x speed gains from early-2026 frontier AI tools.

    Summary In February–April 2026, we ran a survey of 349 technical workers (including 87 software engineers, 71 researchers, 129 academics and PhD students, and 48 founders and managers) about their usage of AI tools. Compared to previous work, our survey is one of the more detailed surveys of technical workers’ self-reported gains from frontier AI tools. 1 We attempt to capture gains due to AI in terms of ‘value’ (how much more value are you creating with AI), rather than ‘speed’ (how long would it have taken you to do these tasks without AI). These can give different answers in principle, in p

    productivity

  17. practice · Hacker News ·

    Using Claude Code: The unreasonable effectiveness of HTML

    Using HTML prompts with Claude Code produces better results than text alone for certain tasks.

    Examples: https://thariqs.github.io/html-effectiveness/ Related: https://simonwillison.net/2026/May/8/unreasonable-effectiven...

    ways of working

  18. research · METR ·

    Task Substitution and Uplift

    Framework distinguishing three measures of AI productivity impact: task-level speedup, task-composition change, and overall value creation.

    Summary: We describe three different definitions of the productivity impact of AI (AKA uplift), and show there’s reason to expect: \[\text{uplift on old tasks} \leq \text{uplift in value} \leq \text{uplift on new tasks}\] Three Measures of Uplift One complication in measuring AI’s effect on productivity is that it has different effects on different tasks, and this causes people to change how they allocate their time between tasks. This makes it more difficult to talk about the effect of AI on overall productivity. We use “old tasks” to mean the set of tasks you’d do in a typical day before AI

    productivity

  19. research · JAMA Network Open ·

    Physician-Reported Safety Outcomes of AI-Generated Hospital Course Summaries

    Prospective trial shows AI-generated hospital summaries identified by physicians as having low harm potential, with 60% adoption and reduced documentation time.

    Importance: High-quality discharge summaries are essential for safe care transitions but contribute substantially to clinician documentation burden and burnout. While retrospective studies suggest that large language models (LLMs) can generate clinical summaries of comparable quality to those by physicians, prospective data on their safety, utility, and association with clinician well-being in clinical environments are lacking. Objective: To evaluate the safety, use, and association with clinician burden of MedAgentBrief, an LLM-based agentic workflow for generating hospital course summaries,

    productivity

  20. research · Information Systems Research ·

    Unraveling Generative AI from a Human Intelligence Perspective: A Battery of Experiments

    Framework comparing LLM and human intelligence across cognitive, emotional, creative, and social domains, validated by assessing job-level impacts of GPT-4.

    This study introduces a novel, human-centered framework for evaluating the holistic intelligence of large language models (LLMs), using behavioral theory and experimental benchmarks drawn from human intelligence. Through extensive online experiments, the framework reveals that GPT-4 outperforms humans in cognitive, emotional, and creative intelligence, but falls short in social intelligence, especially in social interest, self-efficacy, and understanding mental states. Beyond theoretical insight, the study validates this framework by assessing GPT-4’s impact across diverse job roles, finding r

    jobs skills

  21. practice · GitHub Blog ·

    Validating agentic behavior when “correct” isn’t deterministic

    Testing autonomous agents by validating outcomes instead of step-by-step execution paths to reduce false negatives in CI pipelines.

    Modern software testing is built on a fragile assumption: correct behavior is repeatable. For deterministic code, that assumption mostly holds. But for autonomous agents like Github Copilot cloud agent, especially as we explore the frontiers of integrated “Computer Use,” that assumption breaks down almost immediately. As agents move beyond simple code suggestions to interacting with real environments like UIs, browsers, and IDEs, correctness becomes multi-path. Loading screens can appear or disappear, timing shifts, and multiple valid action sequences can lead to the same result. Unless our Gi

    judgment

  22. practice · Paul Ford (Aboard) ·

    Distrust, Then Verify

    Finance leader describes building an internal AI research and pitching system, revealing what actually changed in his sales workflow.

    Just yesterday, according to an article in Bloomberg News by Shirin Ghaffary, Anthropic released ten new tools for the finance industry , to help bankers “draft pitch decks for client meetings, review financial statements, and escalate cases for compliance review.” As reported: “Finance is a great blueprint for the rest of knowledge work,” said Nicholas Lin, Anthropic’s head of product for financial services. Lin added that financial uses of AI are “just a few months behind” the technology’s coding applications, “which we’ve seen massive acceleration in.” Certain financial services stocks— Mor

    ways of working · Paul Ford

  23. practice · How I AI ·

    Quests, token leaderboards, and a skills marketplace: The elite AI adoption playbook | John Kim (Sendbird)

    Company-wide AI adoption through internal marketplace, token leaderboards, and job redesign to reward curiosity over experience.

    John Kim is the co-founder and CEO of Delight.ai, a customer experience platform that’s transforming how companies deploy AI. But what makes John’s story fascinating isn’t just his product; it’s how he’s turned his entire company into an AI-native organization. His marketing team built a fully functional e-commerce swag store with Stripe integration in days. His sales team built their own CRM tools. His recruiting team automated their entire workflow. And it’s all tracked, measured, and celebrated through an internal platform called Automators. What you’ll learn: How Sendbird’s marketing team

    adoption · Claire Vo

  24. research · Strategic Management Journal ·

    Artificial intelligence adoption and the demand for managerial expertise

    Firms adopting AI post more managerial vacancies with shifted skills: away from routine administration, toward interpersonal and growth skills.

    Abstract Research Summary This paper examines how firms' adoption of artificial intelligence (AI) relates to the demand for managers and managerial skills. Using a skills‐based measure of AI adoption derived from Lightcast job postings, we show that firms with greater AI adoption post more managerial vacancies and a higher share of such vacancies than less intensive adopters. These relationships are strongest in manufacturing and among firms with higher research & development intensity. Greater AI adoption is also associated with shifts in managerial skill requirements toward interpersonal and

    jobs skills

  25. practice · How I AI ·

    The internal AI tool that’s transforming how Stripe designs products | Owen Williams

    Design managers and PMs at Stripe use Protodash, an internal AI tool, to create interactive prototypes in dev boxes without coding, changing design review culture.

    Owen Williams is a design manager at Stripe who built Protodash, an internal AI-powered prototyping platform that lets designers and PMs create high-quality Stripe dashboard prototypes without writing code. What started as a bundle of Cursor rules and React components evolved into a full web-based prototyping studio that runs in dev boxes, complete with design review modes, variant testing, and AI-powered iteration. Surprisingly, PMs now use Protodash just as much as designers, fundamentally changing how Stripe approaches prototyping, design reviews, and engineering handoffs. What you’ll learn

    ways of working · Claire Vo

  26. research · Management Science ·

    Backfiring AI? AI Deployment in Workplace

    Game-theoretic model shows AI knowledge-transfer systems can reduce firm productivity by demotivating high performers when employees compete on soft skills.

    Seeking value from artificial intelligence (AI) technologies, firms are rapidly deploying them to augment employees and improve business performance. The diffusion of AI into a firm’s business processes affords the tracking of task actions performed by high-performing employees and the codification of best practices into recommendation systems and training programs. The rising trend in AI deployment reveals managers’ expectations that AI-facilitated knowledge transfer would elevate overall firm performance. However, deploying AI in a workplace has the potential to change the competitive dynami

    management org

  27. practice · Eric Topol ·

    The Paradox of Medical AI Implementation

    Medical imaging AI tools show proven benefits in trials but remain unused; generative AI adoption happens without evidence.

    In 2012, the era of deep learning AI got legs with the convolutional neural network ( AlexNet ) that won the ImageNet challenge. Those images were everyday objects, animals, and scenes, unrelated to health and medicine. Over 7 years ago, I wrote a review in Nature Medicine entitled High-Performance Medicine that summarized the remarkable progress being made for AI interpretation of medical images. Now virtually every type of medical images has undergone extensive assessment with AI, including X-ray, CT, MRI, ultrasound, pathology slides, skin abnormalities, electrocardiograms, endoscopy, and r

    adoption · Eric Topol

  28. practice · Eugene Yan ·

    How to Work and Compound with AI

    Five principles for compounding AI work: context, taste, verification, delegation, and feedback loops.

    Context as infra, taste as config, verification for autonomy, scale via delegation, closing the loop.

    ways of working · Eugene Yan

  29. practice · Claude Blog ·

    How a non-technical project manager built and shipped a stress management app with Claude Code in six weeks

    Non-technical project manager shipped a production app in six weeks using Claude Code without hiring engineers.

    How a non-technical project manager built and shipped a stress management app with Claude Code in six weeks

    ways of working

  30. research · NBER ·

    Endogenous Task Bundling, Skills and Automation

    Model shows how automation technology shifts task bundling boundaries in jobs, with wages set by most expert task.

    An occupations task list records what its workers do, but not why those tasks form one job. We model a competitive labour market in which tasks ordered by required expertise are bundled into jobs, and ask how technology moves the boundaries between them. A jobs wage prices its most expert task. (Joshua S. Gans)

    jobs skills

  31. research · AEA Papers and Proceedings ·

    The Adoption of Industrial AI in America

    Only 23% of US manufacturing plants used AI as of 2021; adoption correlates with cloud computing and structured processes, not prior productivity.

    Using a mandatory, purpose-designed Census Bureau survey of approximately 28,500 establishments, we provide new evidence on industrial AI adoption in US manufacturing. Despite widespread digitization, only 22.8 percent of plants report any AI use as of 2021; intensity-weighted adoption is far lower. Adoption correlates with more-recent digital infrastructure—cloud computing and predictive analytics—rather than legacy on-premises IT or descriptive analytics. Structured production-process management and size are significant predictors. Cost and lack of applicable use case are the most cited barr

    adoption · Erik Brynjolfsson

  32. research · NBER ·

    The Microstructure of AI Diffusion: Evidence from Firms, Business Functions, and Worker Tasks

    18% of U.S. firms adopted AI by early 2026, with detailed breakdowns by business function and worker task types.

    Using novel, nationally representative data from the 2026 AI supplement to the U.S. Census Bureaus Business Trends and Outlook Survey (BTOS), we characterize AI diffusion across three layers: firm-wide adoption, business-function deployment, and worker-task use. During Nov 2025Jan 2026, 18% of firms (Kathryn Bonney , Cory L. Breaux , Emin Dinlersoz , Lucia S. Foster , John C. Haltiwanger , Aditya A. Pande)

    adoption

  33. research · NBER ·

    Automation, Learning, and Career Dynamics

    A general equilibrium model shows how automation of tasks affects worker learning-by-doing, skill acquisition, and career progression over time.

    We study how an automating technology affects career dynamics, human capital, and welfare in an economy where workers acquire skill through the tasks they perform. In a continuous-time general equilibrium model, learning-by-doing is determined jointly with the share of tasks automated, the frontier (Hassan Afrouzi , Andres Blanco , Andrés Drenik , Erik Hurst)

    jobs skills

  34. practice · Addy Osmani ·

    Long-running Agents

    AI agents that persist and resume work across days or weeks, maintaining state outside context windows and recovering from failures.

    A long-running AI agent can keep making progress over hours, days, or weeks. It can do this across many context windows and sandboxes, recover from failure, leave structured artifacts behind, and resume where it left off. For two years the dominant image of an “AI agent” has been a chat window with a clever loop in it. You type a goal, the agent calls some tools, you watch tokens stream by, you stop watching when the work runs out of patience or the context window fills up. That paradigm got us a long way, but it has a ceiling. The model forgets. It declares “task complete” when it isn’t. It r

    ways of working · Addy Osmani

  35. research · Proceedings of the ACM on Human-Computer Interaction ·

    FareShare: A Tool for Labor Organizers to Estimate Lost Wages and Contest Arbitrary AI and Algorithmic Deactivations CSCW016

    A tool for estimating lost wages from algorithmic deactivations reduced calculation time by 95% in field deployment with rideshare drivers.

    What happens when a rideshare driver is suddenly locked out of the platform connecting them to riders, wages, and daily work? Deactivation—the abrupt removal of gig workers’ platform access—typically occurs via arbitrary AI and algorithmic decisions with little explanation or recourse. This represents one of the most severe forms of algorithmic control and often devastates workers’ financial stability. Recent U.S. state policies now mandate appeals processes and recovering compensation during periods of wrongful deactivation based on past earnings. Yet, labor organizers still lack effective to

    worker experience

  36. research · Information Systems Research ·

    How Costs Influence Preferences for Control in Generative Artificial Intelligence (GenAI): Human-Guided vs. GenAI-Based Delegated Search

    When GenAI users face costs, they shift from passive delegation to active control through detailed prompting, achieving better outcomes and higher satisfaction.

    As generative artificial intelligence (GenAI) platforms transition to paid models, concerns grow that usage costs will diminish service value. However, our study of 1.8 million prompts shows that economic constraints actually change how users search solutions with AI. We distinguish between GenAI-based delegated search, which relies on probabilistic sampling, and human-guided delegated search, where users exert active control through refined prompting. We find that salient costs drive users to prioritize controllability over simple cost-minimization. Instead of settling for lower quality, user

    adoption

  37. research · Information Systems Research ·

    Prompt Adaptation as a Dynamic Complement in Generative AI Systems

    Half of productivity gains from better AI models come from users adapting prompts, not the model itself; automation cannot replicate this learning.

    As generative AI systems rapidly improve, a key question emerges: How do users keep up—and what happens if they fail to do so. Drawing on theories of dynamic capabilities and IT complements, we examine prompt adaptation—the adjustments users make to their inputs in response to evolving model behavior—as a mechanism that helps determine whether technical advances translate into realized economic value. In a preregistered online experiment with 1,893 participants, who submitted over 18,000 prompts and generated more than 300,000 images, users attempted to replicate a target image in 10 tries usi

    productivity · Siddharth Suri

  38. practice · Paul Ford (Aboard) ·

    Can AI Companies Be Tamed?

    Ethics and consistent values cannot scale in large tech companies; AI firms are destroying their own operating environment by resisting regulation.

    The New York Times Opinion section recently got in touch and asked me if I felt AI companies could be good . I told them I didn’t believe they could—not because I think AI is inherently evil, but because I don’t believe that a single company at a very large scale can have one, consistent ethical vision: Over three decades of watching the tech industry and watching big companies grow from tiny teams to global powers, I’ve observed the same pattern: Ethics don’t scale up. Tech companies like to start with a mission. Google wanted to connect the world’s information; Microsoft wanted to put a comp

    management org · Paul Ford

  39. research · IEEE Transactions on Software Engineering ·

    A Framework for Evaluating GenAI Adoption and Use in Software Engineering

    Software teams follow a three-phase GenAI adoption process with distinct quality evaluation practices and role responsibilities across ideation, development, and operations.

    Generative Artificial Intelligence (GenAI) is increasingly integrated into software products to enable new features and user capabilities, from early exploration to operational deployment. GenAI adoption as a component within a software system introduces quality risks because GenAI outputs are probabilistic, prompt-sensitive, and may drift after release. Organizations, therefore, need to decide what to evaluate, when to evaluate, and who owns quality evaluation activities across software design, development, and operations. ISO/IEC 25059 standard distinguishes between software product quality

    adoption

  40. research · Nature ·

    Training language models to be warm can reduce accuracy and increase sycophancy

    Training language models for warmth increases error rates by 10-30 percentage points and sycophancy, especially with vulnerable users.

    . Here we show how this can create a significant trade-off: optimizing language models for warmth can undermine their performance, especially when users express vulnerability. We conducted controlled experiments on five different language models, training them to produce warmer responses, then evaluating them on consequential tasks. Warm models showed substantially higher error rates (+10 to +30 percentage points) than their original counterparts, promoting conspiracy theories, providing inaccurate factual information and offering incorrect medical advice. They were also significantly more lik

    judgment

  41. practice · Claude Blog ·

    Onboarding Claude Code like a new developer: Lessons from 17 years of development

    Treating Claude Code as a junior developer with structured onboarding improves code quality and integration.

    Onboarding Claude Code like a new developer: Lessons from 17 years of development

    ways of working

  42. research · Management Science ·

    Does AI Cheapen Talk? Theory and Evidence from Global Entrepreneurship and Hiring

    Generative AI access lowers hiring and investment screening accuracy by 4–9% by cheapening signal quality, with gains for non-English speakers.

    Screening human capital based on signals such as job applications or entrepreneurial pitches is crucial for organizations. Signals are often informative insofar as they require differential knowledge and effort to produce. Generative AI (GAI) complicates screening by lowering the cost of producing impressive signals. We model the informational effects of GAI, showing that applicants’ access to GAI can increase—and also decrease—an evaluator’s screening mistakes. This result depends on how GAI affects experts’ signals compared with nonexperts’. Using experiments in hiring and start-up investing

    jobs skills

  43. research · ACM Transactions on Software Engineering and Methodology ·

    The Impact of LLM-Assistants on Software Developer Productivity: A Systematic Review and Mapping Study

    Systematic review of 39 studies on LLM assistants' effects on software developer productivity finds mixed results: gains in speed and task automation offset by concerns about code quality, cognitive offloading, and reduced collaboration.

    Large language model assistants (LLM-assistants) present new opportunities to transform software development. Developers are increasingly adopting these tools across tasks, including coding, testing, debugging, documentation, and design. Yet, despite growing interest, there is no synthesis of how LLM-assistants affect software developer productivity. In this paper, we present a systematic review and mapping of 39 peer-reviewed studies published between January 2014 and December 2024 that examine this impact. Our analysis reveals that the majority of studies report considerable benefits from LL

    productivity

  44. research · Strategic Management Journal ·

    Bias in, symbolic compliance out? GPT 's reliance on gender and race in strategic evaluations

    GPT suppresses overt discrimination in startup pitch evaluations but preserves underlying bias through neutral-framed critiques, allowing inequality to persist.

    Abstract Research summary Organizations are increasingly using large language models (LLMs) to support strategic evaluations. We examine whether and how these systems rely on gender and race. We asked GPT to evaluate identical startup pitches varying only the founder's name, shaping gender and race perceptions. Across 26,000 evaluations, GPT did not systematically assign lower scores to underrepresented minorities but avoided ranking them last without increasing winning likelihoods. To explain these patterns, we conducted “Second Opinion” experiments where GPT evaluated pitches alongside input

    judgment

  45. practice · Kent Beck ·

    Genie Lessons: Nobody Wants Agents

    Working with AI code tools, developers prefer augmentation over autonomous agents.

    Genie Lesson #5 | Tool: Intent (Augment Code) Genie Lesson #5 | Tool: Intent (Augment Code) Genie Lessons from the Genie Sessions — every Friday I work on a real problem with an AI tool, live, for paid subscribers. Every Monday the lesson drops here, free.

    ways of working · Kent Beck

  46. practice · Paul Ford (Aboard) ·

    Claude Design Just Wants You to Stop Burning Tokens

    Software design is shifting toward UI interactions that minimize token use, not maximizing AI generation.

    Anthropic launched Claude Design last week. You prompt it to design a website, for example, and it designs it, often doing a relatively good job. Then you can click on the different elements on the page and tweak them or issue further prompts. A lot of people were very impressed with it, and Adobe and Figma stocks fell . For me, the most interesting part was not that it could “do design.” It’s the parts of Claude Design that have nothing at all to do with AI—the click-and-fix and click-and-move elements, which are more akin to Canva or Google Slides. I think those aspects of the app point to a

    ways of working · Paul Ford

  47. practice · Shopify Engineering ·

    Flow generation through natural language: An agentic modeling approach

    Shopify fine-tuned an open-source model into a tool-calling agent that generates Flow automations from natural language, replacing a frontier model with faster and cheaper inference.

    We fine-tuned Qwen3-32B into a tool-calling agent that generates Flow automations from natural language—faster, cheaper, and more accurate than the frontier model it replaced, with a weekly retraining flywheel built on real merchant data.

    productivity

  48. practice · Jennifer Pahlka (Eating Policy) ·

    Mitigating Metrics Malaise: Claude on Goodhart's Law

    Claude helps a practitioner think through mitigations for Goodhart's Law in organizational metrics, offering practical instincts leaders use.

    Sometimes you ask Claude for a little help thinking something through, and the answer is good enough that it merits sharing more broadly. I’ve long had a back and forth with Dave Guarino about the value of metrics in changing bureaucratic behavior. It’s one of those classic things where we fight over the tiniest difference, because it’s fun and because we can. Dave thinks getting the metrics right is one of the biggest levers for making a system deliver better results. But I’ve seen good metrics go bad too many times to put as much faith in them as Dave does. He’s right in principle, but in pr

    management org · Jennifer Pahlka

  49. practice · Hacker News ·

    CrabTrap: An LLM-as-a-judge HTTP proxy to secure agents in production

    HTTP proxy that uses an LLM to intercept and validate agent requests before execution, catching unsafe actions in production.

    https://www.brex.com/journal/building-crabtrap-open-source

    judgment

  50. practice · How I AI ·

    How Intercom 2x’d their engineering velocity in 9 months with Claude Code | Brian Scanlan

    Intercom doubled engineering throughput in 9 months by standardizing Claude Code across all engineers, designers, PMs with automated quality hooks and telemetry.

    Brian Scanlan is a senior principal engineer at Intercom, where he’s led the company’s transformation to AI-first engineering. In just nine months, Intercom doubled their R&D throughput while maintaining code quality, with 100% of engineers—plus designers, PMs, and TPMs—now shipping code via Claude Code. What you’ll learn: How Intercom doubled their merged PRs per R&D employee in just nine months using Claude Code The telemetry infrastructure they built to measure AI adoption and quality across hundreds of engineers Why they built a skills repository with hooks that enforce engineering standar

    productivity · Claire Vo

  51. research · Strategic Management Journal ·

    Scaling high and wide: How firms leverage AI and organizational design to overcome the scale‐scope trade‐off

    Longitudinal case study shows how ByteDance used AI and adaptive org design to overcome scale-scope trade-off through cross-domain learning.

    Abstract Research Summary The trade‐off between scale and scope has long posed a strategic dilemma, especially in digital settings, where specialization enables hyperscaling. Drawing on a longitudinal case study of ByteDance, we theorize how digital firms can overcome this constraint through the use of artificial intelligence (AI) combined with an adaptive organizational design. AI evolves and improves through self‐learning and cross‐fertilization across domains, becoming increasingly valuable as learning accumulates. This, however, is contingent on access to structurally related data that all

    management org

  52. practice · Addy Osmani ·

    The Agent Stack Bet

    Production agents fail due to stack issues like shared credentials and missing governance, not model capability; four architectural bets needed.

    Peek under the hood of most “production agents” shipping today and you won’t find intelligence. You’ll find custom plumbing, fragile session logic, shared service accounts, and a security model held together by hope. This can be so much better. If you’ve spent the last 18 months putting agents into production, you already know the models and tools have gotten dramatically better. You also know the problems that are still burning your on-call rotation are not problems you can prompt your way out of. We are running into a stack ceiling , and it is quietly creating a governance and reliability ga

    management org · Addy Osmani

  53. research · Strategic Management Journal ·

    Organizing across cognitive asymmetry in human–AI collaboration: A study of perfume creation

    Perfume makers bridge human tacit knowledge and AI codified knowledge through task allocation, knowledge conversion, and output steering practices.

    Abstract Research Summary As organizations increasingly adopt generative AI (GenAI), they face a strategic challenge: not only deciding which tasks AI should perform, but also how to organize the integration of human and AI efforts to produce viable solutions. We propose that a cognitive asymmetry between human's tacit, embodied knowledge and AI's codified knowledge creates a representational gap that complicates human–GenAI collaboration. Through a qualitative study of professional perfume creation, we identify representational integration as an organizing process through which humans and Gen

    management org

  54. practice · Paul Ford (Aboard) ·

    AI Destroys Moats

    AI commodifies defensibility in consulting and SaaS, forcing firms to compete on culture, speed and relationships instead of technical moats.

    I don’t believe in defensive business moats. Moats are obviously important to a lot of strategists— Paul Ford said so last week . I’m just skeptical of their durability or reliability. It’s me, not you, dear reader. We can attribute my skepticism to a few factors: I’m Lebanese. A quick glance at Lebanon’s history over the last 50 years will tell you one thing: We really suck at moats. I ran two agencies. (And Aboard is increasingly an AI digital transformation consultancy, so I’m now running a third.) An agency or consulting firm is arguably one of the most perpetually vulnerable businesses in

    management org · Paul Ford

  55. research · St. Louis Fed ·

    Why Does AI Adoption Differ So Much across Countries?

    Management practices are a surprisingly powerful predictor of AI adoption rates across countries, more than other factors.

    U.S. firms have a higher share of workers using AI than European firms. This analysis finds management practices are a surprisingly powerful predictor of usage.

    adoption

  56. practice · How I AI ·

    Claude Cowork 101: How to automate your workday without touching code | JJ Englert (Tenex)

    Step-by-step setup of Claude Cowork projects with connectors, brain files, and multi-agent review for autonomous email and calendar workflows.

    JJ Englert leads community enablement at Tenex. In this episode, JJ provides a complete zero-to-one tutorial on Claude Cowork, Anthropic’s desktop tool that sits between simple chat and full terminal-based coding. What you’ll learn: How to create your first Claude Cowork project by connecting a folder on your computer and building context over time The “brain” file strategy: how to create a preferences document that Claude reads every time to understand who you are and how you work Why one-click connectors to Gmail, Slack, Notion, and Google Calendar unlock AI that actually does work instead o

    ways of working · Claire Vo

  57. practice · Will Larson ·

    Agents as scaffolding for recurring tasks.

    Agent automatically monitors security vulnerabilities, filters by priority, and routes to owners using GitHub MCP.

    One of my gifts/curses is an endless fixation with how processes can be optimized. For a brief moment early in my career, that was focused on improving how humans collaborate, but that quickly switched to figuring out how we can minimize human involvement, and eliminate human-to-human handoffs as much as possible. Lately, every time I perform a recurring task–or see someone else perform one–I think about how we might eliminate the human’s involvement entirely by introducing agents. This both has worked well, but also worked poorly, and I wanted to highlight the pattern I’ve found useful. For a

    ways of working

  58. practice · Linear ·

    How we use Linear Agent at Linear

    Three concrete workflows where Linear Agent turns customer emails, Slack threads and PM issues into shipped features without manual triage.

    Before we launched Linear agent, our teams were already using it across Slack, Linear, and our codebase. It’s become a core part of how we build. Here are three workflows that have proven most effective: A customer email turns into a shipped feature A Slack thread becomes a pull request A PM files an issue, and ships the fix As you’d expect, some of what we use internally has not been released yet, and we’ve called that out where relevant. From customer email to shipped feature Stage 1: Pull the context into Linear Customer feedback submitted by email, or in-app within Linear, accrues to an In

    ways of working

  59. practice · Claude Blog ·

    Multi-agent coordination patterns: Five approaches and when to use them

    Five patterns for coordinating multiple AI agents, with guidance on when each works best.

    Multi-agent coordination patterns: Five approaches and when to use them

    ways of working

  60. research · arXiv ·

    Scaffolding Human-AI Collaboration: A Field Experiment on Behavioral Protocols and Cognitive Reframing

    Field experiment shows cognitive reframing training improved document quality, but structured pair protocols reduced productivity versus unstructured AI use.

    Organizations have widely deployed generative AI tools, yet productivity gains remain uneven, suggesting that how people use AI matters as much as whether they have access. We conducted a field experiment with 388 employees at a Fortune 500 retailer to test two scaffolding interventions for human-AI collaboration. All participants had access to the same AI tool; we varied only the structure surrounding its use. A behavioral scaffolding intervention (a structured protocol requiring joint AI use within pairs) was associated with lower document quality relative to unstructured use and substantial

    adoption · Lev Tankelevitch · Rebecca Janssen

  61. practice · Claude Blog ·

    How and when to use subagents in Claude Code

    Guidance on structuring multi-agent systems in Claude Code with subagents for task delegation.

    How and when to use subagents in Claude Code

    ways of working

  62. practice · How I AI ·

    I gave Claude Code our entire codebase. Our customers noticed. | Al Chen (Galileo)

    Field engineer built a Claude Code system to query 15 repositories and Confluence, cutting engineering interruptions and delivering personalized customer answers.

    Al Chen is a field engineer at Galileo, an observability platform for AI applications, where he works on the front lines with enterprise customers asking highly technical questions. Despite never having held an engineering role, Al has built a system using Claude Code to query Galileo’s 15 separate repositories, combine that with Confluence documentation and customer-specific quirks, and deliver hyper-personalized answers that would otherwise require constant engineering support. What you’ll learn: How to use Claude Code to query multiple repositories simultaneously for customer support Why co

    ways of working · Claire Vo

  63. practice · Hacker News ·

    We replaced RAG with a virtual filesystem for our AI documentation assistant

    Team replaced RAG pipeline with virtual filesystem abstraction for faster, more reliable AI documentation retrieval.

    411 points on Hacker News. Discussion: https://news.ycombinator.com/item?id=47618223

    ways of working

  64. research · arXiv ·

    Augmented Human Capital: A Unified Theory and LLM-Based Measurement Framework for Cognitive Factor Decomposition in AI-Augmented Economies

    Workers with augmentable cognitive skills earn higher wages when AI adoption is high, but only in formal employment; informal workers see no augmentation premium.

    This paper proposes a decomposition of human capital into three orthogonal components -- physical-manual (H^P), routine-cognitive (H^C), and augmentable-cognitive (H^A) -- and develops a production function in which AI capital interacts asymmetrically with these components: substituting for routine cognitive work while complementing augmentable cognitive work through an amplification function phi(D). I derive a corrected Mincerian wage equation and show that the standard specification is misspecified in AI-augmented economies. Using LLM-generated measures of occupational augmentability for 18,

    jobs skills

  65. practice · Paul Ford (Aboard) ·

    The Ghost Effect

    Using LLMs to find serious software bugs in production systems with minimal code; researcher found a 23-year-old Linux vulnerability.

    There’s a movie from 1990 called Ghost , where Patrick Swayze (a banker) gets killed, then spends a lot of the movie trying to warn his girlfriend Demi Moore (a ceramic artist) that she’s in danger. He’s a ghost, so no one can hear or see him. This is very frustrating to him. He is, or was, a handsome banker; he’s not used to that. Eventually, he does learn how to move things around, and uses his new powers to scare cats and warn people about various plot points. Famously, he possesses the body of Whoopi Goldberg, and that allows him, shirtlessly , to make out with Demi Moore in her absolutely

    ways of working · Paul Ford

  66. research · NBER ·

    Understanding Firms' AI Efforts and Their Economic Impact

    Firm-level data on AI efforts shows distinction between invention, use, internal building, and outsourcing; measurement challenges affect economic impact evidence.

    This paper reviews firm-level data on artificial intelligence (AI) and the emerging evidence on AIs economic effects. It argues that measurement is central: different AI data sets capture different objects, including invention versus use; internal capability building versus outsourcing; and realized (Tania Babina)

    adoption

  67. research · NBER ·

    How (un)Stable Are LLM Occupational Exposure Scores? Evidence from Multi-Model Replication

    LLM-based occupational exposure scores vary 3.6-fold across frontier models, suggesting fragility in AI labour-market exposure measures.

    A rapidly growing literature estimates AI's labor-market effects using large language models (LLMs) to self-assess occupational exposure. We demonstrate these measures are highly fragile. Replicating the dominant rubric with three frontier models on identical tasks, we find a 3.6-fold divergence in (Michelle Yin , Hoa Vu , Claudia Persico)

    jobs skills

  68. practice · One Useful Thing ·

    Claude Dispatch and the Power of Interfaces

    Chatbot interfaces create cognitive overload that offsets AI productivity gains, especially for less experienced workers who need AI most.

    AIs are already far more capable than most people realize. A large part of this so-called capability overhang comes not from the limits of AI (though, of course, they still have many limits), but from how people interact with it. The vast majority of people access AI through chatbots, and usually the free versions with less capable models. A chatbot is fine for a quick question, but it is a bad way to get real work done. In fact, recent research suggests that we pay a mental tax when using chatbot interfaces for work. A new paper had a small group of financial professionals do a complex valuat

    ways of working · Ethan Mollick

  69. research · arXiv ·

    Economics of Human and AI Collaboration: When is Partial Automation More Attractive than Full Automation?

    Framework shows partial automation often cost-optimal over full automation; AI accuracy exhibits convex costs; task complexity predicts labor substitution rates.

    This paper develops a unified framework for evaluating the optimal degree of task automation. Moving beyond binary automate-or-not assessments, we model automation intensity as a continuous choice in which firms minimize costs by selecting an AI accuracy level, from no automation through partial human-AI collaboration to full automation. On the supply side, we estimate an AI production function via scaling-law experiments linking performance to data, compute, and model size. Because AI systems exhibit predictable but diminishing returns to these inputs, the cost of higher accuracy is convex: g

    productivity

  70. research · Proceedings of the ACM on Human-Computer Interaction ·

    The Teammate Divide: How Humans Approach Teamwork Differently with Human and AI Teammates

    Humans report stronger teamwork and less conflict with human teammates than AI teammates in matched roles, driven by unmet expectations of AI.

    As Artificial Intelligence (AI) technologies begin to become more fundamental to society, their design and integration as teammates that work alongside humans will become commonplace. Recently, research into these human-AI teams has shown that the perceptions humans form of the capabilities and goals of their AI and human teammates differ. However, while humans perceive their teammates differently, it is still unclear whether these perceptual differences extend to the actual teamwork humans have with their teammates. To address this gap, this article reports on a mixed-methods research study o

    teams

  71. research · Proceedings of the ACM on Human-Computer Interaction ·

    ''AI is a friend, a teammate, not a supervisor'': Understanding Human Perceptions of AI teammates' Monitoring in Human-AI Teams

    Interview study identifies tensions in how workers perceive AI monitoring in teams: performance versus privacy, support versus control, fairness versus empathy.

    Artificial Intelligence (AI)-mediated collaboration has been a core research interest in CSCW and HCI in recent decades. AI is being embedded in teams as a teammate, known as human-AI teamwork. A critical aspect to this and any type of teamwork is monitoring: the ability to track and be aware of both team- and individual-level actions. As AI continuously evolves from a passive tool into a supportive collaborator capable of playing a role in teams, more research is needed to investigate human perceptions of AI teammates' monitoring. More work is needed on when and how AI can function as a suppo

    worker experience

  72. research · IEEE Transactions on Software Engineering ·

    More Code, Less Understanding? On the Impact of AI Assistants on Developers’ Productivity and Code Ownership

    Controlled experiment with 69 developers shows AI coding assistants improve task completion but developers answer fewer questions about their own code.

    Artificial Intelligence (AI) is transforming many domains, including software engineering. AI-based tools are gaining popularity and are increasingly being integrated into software development workflows, automating complex tasks such as code writing and reviewing. When it comes to coding tasks, some evidence suggests that tools like Copilot boost developers’ productivity (e.g.,developers can handle a larger number of pull requests per week). However, it remains unclear whether this comes at the expense of code ownership (i.e.,the developer’ ability to argue about their implementation choices).

    productivity

  73. practice · How I AI ·

    How to turn Claude Code into your personal life operating system | Hilary Gridley

    Using Claude Code with observation-based preference learning and iPhone shortcuts to automate personal and work tasks without upfront configuration.

    Hilary Gridley is an entrepreneur, former product leader, and new mom who previously appeared on the podcast discussing AI for managers. She returns to share how she's transformed her approach to personal productivity using Claude Code as her primary tool for managing both professional work and life admin. Hilary demonstrates her "anti-system system"—a philosophy that prioritizes simplicity over complex setup, allowing AI to learn preferences through observation rather than upfront configuration. What you’ll learn: How to capture to-dos instantly using a simple iPhone back-tap shortcut that re

    ways of working · Claire Vo

  74. research · arXiv ·

    Artificial Intelligence in Science: Returns, Reallocation, and Reorganization

    AI adoption in research projects reallocates resources toward larger teams and human capital rather than producing immediate efficiency gains.

    Investment in artificial intelligence (AI) has grown rapidly, yet its returns to scientific research remain poorly understood. We study how AI reshapes the production of science using a comprehensive dataset of research proposals submitted to a large international funding agency, including both funded and unfunded projects. Combining keyword extraction with large language model classification, we identify the presence, type, and functional role of AI within each proposal and link these measures to detailed budget allocations, team structure, and subsequent publication outcomes. We find that, i

    adoption

  75. practice · Will Larson ·

    The agentic passive voice.

    Reframing how people talk about AI mistakes reveals they're avoiding responsibility by using passive voice disguised as active.

    At some point, you will have learned about the passive voice, where the actor in a sentence is unclear. For example, my software didn’t compile. That’s a good example of the passive voice. However, you might not know the full set of rules, because here are some sentences in the passive voice that you might not recognize: Claude made an error in my writeup. ChatGPT messed up the commitment. Gemini didn’t write tests. You might think those are active sentences, but those are in fact examples of the agentic passive voice . The rule here is: whenever the actor in a sentence is a model, then it’s a

    judgment

  76. practice · Hamel Husain ·

    The Revenge of the Data Scientist

    Data scientists are regaining value by focusing on experiment design, metric creation, and debugging AI systems rather than model training.

    Is the heyday of the data scientist over? The Harvard Business Review once called it “The Sexiest Job of the 21st Century.” 1 In tech, data scientist roles were often among the best paid. 2 The job also demanded an unusual mix of skills: Data Scientist (n.): Person who is better at statistics than any software engineer and better at software engineering than any statistician. — JosH100 ( @josh_wills ) May 3, 2012 In addition to creating a high-barrier to entry, these skills enabled data scientists to build predicitive models, measure casuality and find patterns in data. Of these, predicitive m

    jobs skills · Hamel Husain

  77. practice · Paul Ford (Aboard) ·

    The New Bot Should Clean up the Old Bot’s Mess

    Use AI coding to clean up existing digital mess rather than build ambitious new products, reducing failure risk.

    I keep seeing the same stat over and over: 95% of AI projects fail . (I specifically keep seeing it because AI companies continually trot it out to insist that they’re in the 5%.) I have no argument with that stat. It may be entirely correct. But I think I can diagnose the problem—and it’s not the tools. People keep throwing wildly ambitious product problems at LLM-based coding. They want to build products, or replace employees, or make whole new platforms, and they want to do it in five prompts or less—and sometimes, after an hour of churning through tokens, it kind of works, which releases s

    ways of working · Paul Ford

  78. practice · How I AI ·

    How Stripe built “minions”—AI coding agents that ship 1,300 PRs weekly from Slack reactions | Steve Kaliski (Stripe engineer)

    Stripe engineers activate AI coding agents from Slack to ship 1,300 PRs weekly, with code review as the main human intervention point.

    Steve Kaliski is a software engineer at Stripe who has spent the past six and a half years building developer tools and payment infrastructure. He’s part of the team that created “minions”—Stripe’s internal AI coding agents, which now ship approximately 1,300 pull requests per week with minimal human intervention beyond code review. In this episode, Steve demonstrates how Stripe engineers activate development work from Slack and leverage cloud-based development environments for parallel agent workflows, and demos machine-to-machine payments where AI agents transact autonomously with third-part

    productivity · Claire Vo

  79. research · Knowledge at Wharton ·

    When Better AI Makes Oversight Harder

    As AI systems improve in reliability, human overseers become less motivated to monitor them effectively, creating governance risks.

    As AI systems become more reliable, organizations may find it increasingly difficult to motivate humans to oversee them effectively, Wharton research shows. … Read More

    judgment

  80. practice · Anthropic Engineering ·

    Harness design for long-running application development

    Harness design patterns for long-running agentic coding tasks, improving performance in autonomous frontend development.

    Harness design is key to performance at the frontier of agentic coding. Here's how we pushed Claude further in frontend design and long-running autonomous software engineering.

    ways of working

  81. practice · Artificial Ignorance ·

    The Bots Are Reading Along

    When AI agents consume more documentation than humans, creators must rethink what they write and how.

    It’s been a little while! At the beginning of the month I got pretty sick, though admittedly it was nice to take my first publishing break in 3 years. But don’t worry: I’ve got more stuff in the works, including a Codex Basics livestream this Friday! If you’ve been meaning to try out AI coding agents but haven’t had the time, this is the stream for you. Yesterday I saw a tweet from Barry McCardel - the CEO of analytics platform Hex - which included a graph showing that agents are now creating more Hex cells (basically dashboard components) than humans are. Not “almost as many” or “a growing nu

    ways of working · Charlie Guo

  82. practice · How I AI ·

    How Microsoft's AI VP automates everything with Warp | Marco Casalaina

    VP shows how to use Warp beyond coding: automating Azure admin, document scanning, video compression, and triggered email workflows.

    Marco Casalaina , VP of Core AI Products and AI Futurist at Microsoft, demonstrates how he uses AI tools to automate administrative tasks that typically consume valuable time. Rather than using Warp as a coding assistant (its primary marketed purpose), Marco leverages it to manage Azure resources, scan documents, compress videos, and more. He shows how these “micro-agents” can reduce friction in everyday workflows, allowing him to focus on higher-value activities. Marco also demonstrates how Microsoft 365 Copilot and ChatGPT can create triggered workflows that respond to emails or check for in

    ways of working · Claire Vo

  83. research · arXiv ·

    Where can AI be used? Insights from a deep ontology of work activities

    New ontology maps 20K work activities to 13,275 AI applications and 20.8M robotic systems, showing AI market value concentrates in information creation and transfer.

    Artificial intelligence (AI) is poised to profoundly reshape how work is executed and organized, but we do not yet have deep frameworks for understanding where AI can be used. Here we provide a comprehensive ontology of work activities that can help systematically analyze and predict uses of AI. To do this, we disaggregate and then substantially reorganize the approximately 20K activities in the US Department of Labor's widely used O*NET occupational database. Next, we use this framework to classify descriptions of 13,275 AI software applications and a worldwide tally of 20.8 million robotic s

    adoption · Thomas Malone

  84. practice · Rands in Repose ·

    Better, Faster, and (Even) More

    Developer organises Claude Code projects with CLAUDE.md instructions and WORKLOG.md session logs to scale rapid prototyping.

    I’ve never built more interesting, random, and useless scripts, tools, and services than I have in the last six months. The cost to go from “Random Thought” to “Working Something” has never been lower thanks to Claude Code . However, this increase in speed has only made my desire to move faster and more efficiently higher. The following is a set of tools and practices I’ve gathered over the last 90 days, which continue to accelerate my process and give me daily joy. How I Organize Everything lives under ~/Projects/ . Each project is its own git repo with its own CLAUDE.md (project-specific ins

    ways of working

  85. practice · Jason Fried ·

    The bespoke software revolution? I'm not buying it.

    Most workers want problems solved, not software projects to maintain, regardless of AI making custom tools easier to build.

    A bespoke software revolution? I don't buy it. It'll exist. It already exists. Small consultants and big consulting firms have made custom software for years. It almost always sucks. It’s bloated, confusing, and because the client pays, it’s built in all the wrong ways. Who’s excited about bespoke software? Software makers! Of course they're excited about building bespoke software — that's what they do. X is full of them. Your feed is full of people who love making software talking about making software. Of course they’re excited about the revolution. Echo, echo, echo... Most people don’t like

    management org · Jason Fried

  86. practice · Addy Osmani ·

    Is the IDE dead?

    Developer work is shifting from line-by-line editing toward supervising AI agents that plan, rewrite and test code, with control panels replacing editors as primary tools.

    The center of developer work is moving. Not disappearing - moving. Away from continuous, line-by-line editing inside a single window, and toward supervising agents that can plan, rewrite files, run tests, and propose changes for review. IDEs as we know them may stop being the primary tool for software work, or heavily evolve. Across the tools many developers including myself are already using daily - Conductor , Claude Code Web , GitHub Copilot Agent , Jules , Vibe KanBan , even cmux - the same shift keeps showing up: the control plane is becoming the primary surface, and the editor is becomin

    ways of working · Addy Osmani

  87. research · Government Information Quarterly ·

    Open to open-source AI? Navigating AI model choice in public sector agencies

    Public sector AI adoption favours proprietary over open-source models; technical fit and data sovereignty matter more than vendor lock-in concerns.

    Public sector Agencies are increasingly adopting artificial intelligence (AI) tools. High quality open-source AI (OSAI) options are available, but much of their current attention is on proprietary options such as Copilot and ChatGPT. There are parallels with take-up of open-source software (OSS). While OSS has a foothold in niche functions of Agencies' technology suites, it has not seen widespread adoption despite backing from technical and political spheres and its potential to reduce costs and spur increased competition and innovation. Grounded in theoretical frameworks and evidence used to

    adoption

  88. research · Research Policy ·

    AI users are not all alike: The characteristics of French firms buying and developing AI

    French firms developing AI in-house show productivity gains beyond self-selection; AI buyers do not.

    In this work we characterise French firms using artificial intelligence (AI) in 2018 and explore the link between AI use and productivity. We distinguish AI users that source AI from external providers (AI buyers) from those developing their own AI systems (AI developers) based on the official French ICT survey, that provides information about the use of AI by firms. AI buyers tend to be larger than other firms, but this relation is explained by ICT-related variables. Conversely, AI developers are larger and younger beyond ICT. Other digital technologies, digital skills, and infrastructure pla

    adoption

  89. practice · Paul Ford (Aboard) ·

    Don’t Mix Up Artifacts With Processes

    Distinguishing between AI-generatable artifacts and the irreplaceable processes that produce trustworthy work.

    A lot of questions about AI reduce, when closely inspected, to silliness. One I see often is: “Could AI replace journalists?” But then who would write the restaurant reviews? There is no journalism without restaurant reviews. What’s going on under that question is people keep mixing up the artifact with the process. Play this out a little: When you read a restaurant review in a publication that has any allegiance to journalism (i.e., not pay-to-play Instagram influencers), the promise is that the reviewer: Has eaten food in the past; Can describe the taste and smell of food; Can describe their

    judgment · Paul Ford

  90. research · KiltHub Repository ·

    AI as Climate Technology: When AI Strengthens Collective Intelligence—and When It Does Not

    AI deployment shapes team collaboration climate by signaling valued behaviors, affecting collective intelligence beyond task productivity.

    As artificial intelligence becomes embedded in everyday work, organizations often evaluate it through a familiar lens: whether it raises productivity. This article argues that this is too narrow a question. In dynamic, interdependent organizations, AI does more than automate tasks or support decisions. It also shapes the climate in which people collaborate by signaling what behaviors are visible, valued, and rewarded. In that sense, AI is not only a production technology or even a coordination technology; it is also a climate technology. This matters because collective intelligence depends not

    management org · Anita Woolley

  91. research · npj Digital Medicine ·

    From tool to teammate in a randomized controlled trial of clinician-AI collaborative workflows for diagnosis

    RCT shows AI-clinician collaborative diagnosis improves accuracy to 82-85% versus 75% baseline, comparable across workflow orders.

    Early studies of large language models (LLMs) in clinical settings have largely treated artificial intelligence (AI) as a tool rather than an active collaborator. As LLMs demonstrate expert-level diagnostic performance, the focus shifts from whether AI can offer valuable suggestions to how it integrates into physicians' diagnostic workflows. We conducted a randomized controlled trial (n = 70 clinicians) to assess a custom system designed for collaborative diagnostic reasoning. The design involved independent diagnostic assessments by the clinician and AI, followed by an AI-generated synthesis

    judgment

  92. practice · How I AI ·

    From journalist to iOS developer: How LinkedIn’s editor builds with Claude Code | Daniel Roth

    Non-coder built production iOS apps using Claude Code with a dual-agent system (builder and reviewer) and shipped to App Store.

    Daniel Roth , editor in chief at LinkedIn, went from business writer to iOS app developer, without ever learning how to code. Using Claude Code, Daniel built and shipped multiple production-ready iOS apps to the App Store, including Commutely, a personalized train-tracking app for New York commuters. What you’ll learn: How to set up a dual-agent Claude Code system (builder + reviewer) Why being a “picky customer” is the right mindset for non-technical builders How Daniel prioritizes features using AI-ranked impact vs. build time Why saving everything as Markdown files creates long-term context

    ways of working · Claire Vo

  93. research · Personnel Psychology ·

    From Text to Insight: Leveraging Large Language Models for Performance Evaluation in Management

    Advanced LLMs match or exceed human raters in evaluating knowledge work performance, with newer models showing better bias resistance than humans.

    ABSTRACT This study examines whether Large Language Models can serve as reliable supplements to human judgment in evaluating text‐based task performance. Through two studies analyzing 744 knowledge‐based performance outputs, we compare ratings from multiple LLM architectures (GPT‐4, GPT‐5, o3, Claude Sonnet 4, DeepSeek v3) against human evaluators (individual and aggregated ratings), with external expert consensus serving as the validity benchmark for both. Our multi‐model design reveals that various LLMs demonstrate comparable or superior evaluation capabilities relative to human raters, with

    judgment

  94. research · ACM Transactions on Software Engineering and Methodology ·

    "Should I Give Up Now?" Investigating LLM Pitfalls in Software Engineering

    In a 26-person web development study, unhelpful LLM responses increased abandonment by 11x; 17 of 26 participants quit using ChatGPT.

    Software engineers are increasingly incorporating AI assistants into their workflows to enhance productivity and alleviate cognitive load. However, experiences with large language models (LLMs) such as ChatGPT vary widely. While some engineers find them useful, others deem them counterproductive due to inaccuracies in their responses. Researchers have also observed that ChatGPT often provides incorrect information. Given these limitations, it is crucial to determine how to effectively integrate LLMs into software engineering (SE) workflow. Analyzing data from 26 participants in a complex web d

    adoption

  95. research · Organization Science ·

    Gendered Navigation of Advice and Suboptimal Behavior in Matching Algorithms: Evidence from the Residency Match

    Men seek more independent advice about matching algorithms than women, leading to better outcomes despite equal baseline guidance.

    Two-sided matching algorithms have been deployed at an increasing rate in labor markets all over the world in part because they can result in more equitable labor market matches. To achieve this desirable result, institutions using these algorithms often engage in a translation process to provide advice to market participants about how to optimally interact with the algorithm. We draw on theories of gendered agency to theorize that men may be more likely than women to engage in independent advice seeking—an agentic way of navigating one’s understanding by seeking out additional advice about ho

    jobs skills

  96. research · Equitable Growth ·

    How union contracts are protecting U.S. workers from automated management and surveillance in the workplace

    Union contracts include provisions limiting automated management and surveillance; a national survey documents their prevalence and effects.

    Evidence from a new national survey and implications for policymakers and union leaders Key takeaways Employers are increasingly using electronic and automated tools to collect data on their workers on and off the job and then using that data to inform decisions about workers’ pay, schedules, work assignments, working conditions, promotions, discipline, and even terminations. […] The post How union contracts are protecting U.S. workers from automated management and surveillance in the workplace appeared first on Equitable Growth .

    worker experience

  97. practice · One Useful Thing ·

    The Shape of the Thing

    AI work has shifted from co-intelligence (prompting back-and-forth) to managing autonomous agents that handle hours of work in minutes.

    In October of 2023, I wrote about the “Shape of the Shadow of the Thing,” speculating on the Thing that AI might turn into in the coming years. I think we can see the Thing much more clearly now, and some of the consequences that come with it. As I have been discussing in recent posts, we have entered a new phase of AI. After ChatGPT was introduced, human-AI work took the form of what I called co-intelligence, where humans would prompt AI back-and-forth to get help on tasks. Starting in late 2025, we entered a new era thanks to AI agents like Claude Code , OpenAI’s Codex, and OpenClaw. These a

    ways of working · Ethan Mollick

  98. research · Strategic Management Journal ·

    Unlocking novel knowledge recombinations: The effect of artificial intelligence on inventive activity

    Patents incorporating AI show greater novel recombinations of existing technologies than patents without AI, suggesting AI bridges previously disconnected domains.

    Abstract Research Summary Complementing the role of AI in facilitating search and identifying combinations of high value in inventive activity, we argue that AI fundamentally alters the innovation landscape by unlocking new combinations that were previously infeasible. This effect arises because AI acts as a powerful shared layer due to its predictive capabilities and its ability to transmit solutions across domains, thereby creating a bridge between previously unconnected elements. Utilizing a matched sample of patents, we show that inventions incorporating AI exhibit a greater degree of nove

    productivity

  99. practice · Paul Ford (Aboard) ·

    And Now, a Word From Our Sponsor!

    Software companies must blur product and agency roles, offering tools, practices and human touchpoints rather than single platforms.

    Roughly 10,000 years ago, in 2023, Aboard launched a fun data-management tool—sort of like if Airtable and Pinterest had a happy baby. Our idea was to make complex data easy to organize—to make “database stuff” feel like “bookmarking”—and then, as it picked up users, we hoped to adapt it for companies and teams. That version of Aboard found dedicated users, but the world had other plans. Right after we launched it, the tidal waves started coming: ChatGPT 4.5, vibe-coding tools, and Claude Code. “Making it easy” is always the right idea in technology, but the definition of “easy” changed from “

    management org · Paul Ford

  100. practice · Will Larson ·

    Judgment and creativity are all you need.

    Coding agents accelerated migration from manual deploys to continuous deployment by handling repetitive validation steps humans skip.

    When I joined Imprint a little less than a year ago, our deploys were manual, requiring close human attention to complete. Our database migrations were run manually, too. Developing good software is very possible in those circumstances, but it takes a remarkable attention to detail to do it. It was also possible to develop good software using Subversion and developing by ssh’ing into a remote server to edit PHP files, but the goal is making things easy rather than possible . Ten months later, the vast majority of our changes, including database migrations, continuously deploy to production wit

    productivity