The Feed · Complete archive

AI at work: research and practice

A server-rendered record of the evidence, ideas and firsthand practices screened by The Feed. Every entry links to its original source.

Use the interactive Feed

Page 2 of 9 · 873 items

  1. practice · martinfowler.com ·

    Making Your Data Ready for Agentic AI

    How to prepare organizational data so AI agents can access accurate, trustworthy information for autonomous work.

    Lots of organizations are excited about what AI can do to streamline their processes, save money, and juice margins. But AI's capabilities are founded on the data that AI accesses, and for many organizations that foundation is little more than sand. Pramod Sadalage and Prem Chandrasekaran write about how to build a reliable foundation of data that can be accurate and trusted. more…

    adoption · Martin Fowler

  2. research · arXiv ·

    Monocultural Biases: Correlated biases in large language models lead to unequal systemic exclusion rates in hiring

    Post-trained LLMs show correlated hiring biases across models, increasing systemic exclusion rates from 5.6% to 17.3%, with age discrimination worsening after training.

    Employers are increasingly using large language models (LLMs) to automate their hiring process. This paper investigates the risk of monocultural biases, in which the widespread deployment of large language models homogenizes biases across the labor market, leading to greater systemic exclusion for certain demographic groups. For ten LLMs, we measure hiring biases across their base and post-trained versions to identify which stage, pre-training or post-training, lead to monocultural biases. We find that, compared to their base models, post-trained models are 3.6% less likely to callback older a

    jobs skills · Maria del Rio-Chanona

  3. practice · Nate B Jones ·

    You bought the agent to get time back. Here is why your calendar filled up instead (+ the five prompts that fix it.)

    Agent work creates invisible overhead: allocation, specification, evaluation, intervention, coordination, recovery that fills calendars despite time-save promises.

    I can end a day carrying work that the product, the team, and the budget have never named. The strange part of running agents is that I start more work than I can inspect. I spend the day moving between outputs that each need a decision. None of it shows up anywhere. The dashboards report tokens, run counts, and time saved during execution, and not one of them measures the thing that actually filled the day. That work has a shape. Allocation, specification, evaluation, intervention, coordination, recovery. It takes judgment and it carries accountability. Some days I find it exhausting. Agent f

    worker experience

  4. research · arXiv ·

    Are Concept Bottleneck Models Effective as Decision-Support Systems?

    Large user studies show concept-based AI explanations improve human-AI team accuracy, but only when tasks are difficult, concepts are clear, and users actively interact.

    Concept Bottleneck Models (CBMs) are interpretable-by-design neural networks that detect human-understandable concepts from the input and use them to generate predictions. By allowing users to inspect the concepts underlying a prediction and explore how predictions change under alternative concept configurations, CBMs have emerged as one of the most prominent approaches to supporting human-AI collaboration. However, user studies investigating their actual effectiveness as decision-support systems remain limited. We present two large-scale user studies (N participants = 705, N observations = 6,

    judgment

  5. practice · GitHub Blog ·

    How to evaluate LLMs before production

    GitHub reduced false positives in secret scanning by moving from benchmark evaluation to production-realistic testing with safety constraints.

    A language model can perform well on a clean benchmark and still struggle with the cases that matter in production. Benchmarks and curated datasets are useful when prototyping an LLM-based system. They help teams compare models, test an initial prompt, and determine whether an idea is technically plausible. But as a system moves closer to production, the evaluation problem changes. Real inputs are often ambiguous. Labels may be inconsistent. Important context may be missing or truncated. The evaluation set may not reflect the production distribution. Edge cases that rarely appear in benchmarks

    judgment

  6. practice · The Pragmatic Engineer ·

    Why Ramp built its own in-house coding agent, Inspect

    Ramp built Inspect, an internal coding agent running on remote sandboxes with access to internal data, to run multiple agents in parallel and integrate with their development workflow.

    At a select few tech companies, they write most of their code with their own, custom-built, internal AI coding agents. This is different from most of the industry which uses AI coding agents and harnesses like Codex, Claude Code, Cursor, OpenCode, GitHub Copilot, etc. At Ramp, their own version is called Inspect , while at Block it’s Goose (open source), at Stripe it’s Minions , and River at Shopify. But why not just use what frontier labs and coding harness AI startups already offer; why take the time and effort? We reached out to Ramp, a fintech company big on building its internal AI infras

    ways of working · Gergely Orosz

  7. research · EU Joint Research Centre ·

    A rising tide: Revisiting the occupational impact of AI in the generative era

    EU researchers update occupational exposure estimates for generative AI, revising earlier automation risk rankings across jobs.

    A rising tide: Revisiting the occupational impact of AI in the generative era

    jobs skills

  8. research · Strategic Management Journal ·

    The impact of generative artificial intelligence on innovation: Evidence from software products

    ChatGPT release increased new software products by 34% but reduced improvements to existing ones by 20%, shifting developer effort toward quantity over novelty.

    Abstract Research Summary We study the impact of generative artificial intelligence (GAI) tools on product‐level innovation outcomes in the context of software products. Specifically, we illustrate how GAI can alter the direction of innovation by shifting the activities of developers away from generational innovation and toward original innovation, which may be new to the market but not necessarily more novel than previous innovations. We argue that this shift is driven by GAI's ability to facilitate tasks in both the ideation and implementation of software products, which enables some develop

    adoption

  9. practice · How I AI ·

    I spent $20,000 on Devin in a month. Here’s what I learned | Ryan Carson (solo founder)

    Solo founder manages 15 concurrent AI agents with paper-based system; replaces QA team using LAN PR skill and Watchdog playbook for customer success.

    Ryan Carson is a five-time founder and the current solo founder of Untangle, a B2B SaaS platform for family law firms. Before Untangle, he co-founded Treehouse, an online coding education platform, and has spent the better part of two decades building and leading tech companies. He’s active on X, where he shares his solo founder journey in real time, including what he actually spends on AI tools each month. What you’ll learn: Why Ryan manages 15 concurrent Devin agents with a folder system and a piece of paper, not a dashboard The Watchdog playbook: what he built to replace a customer success

    ways of working · Claire Vo

  10. practice · Claude Blog ·

    How an Anthropic field marketer uses Claude Code to send weekly personalized updates to every sales rep

    Field marketer uses Claude Code to generate personalized weekly updates for each sales rep, automating segmentation and customization at scale.

    How an Anthropic field marketer uses Claude Code to send weekly personalized updates to every sales rep

    ways of working

  11. research · Information Systems Research ·

    The Impact of Generative AI on Collaborative Open-Source Software Development: Evidence from GitHub Copilot

    GitHub Copilot increases code contributions but raises coordination costs; benefits concentrate among core developers while peripheral contributors face integration delays.

    Generative AI holds great promise for reshaping software development. We examine the impact of GitHub Copilot, a generative AI pair programmer, on collaborative open-source software (OSS) development, where developers voluntarily contribute and collaborate on projects. Using GitHub’s proprietary Copilot usage data combined with public OSS data from GitHub, we find that Copilot increases project-level code contributions by expanding developer participation and increasing code contributed by individual developers. However, these gains lead to increase in time to coordinate, review and integrate

    teams

  12. practice · Sangeet Paul Choudary ·

    Strategy in the age of AI - Eight points beyond the obvious

    AI shifts competitive advantage from modularity to owning feedback loops connecting activities across an organization.

    What happens when AI changes not only how companies compete, but the very game they are competing in? In this conversation, Rita McGrath and I - guided by the always-insightful Aidan McCullen - bring together complementary ideas from the books we have written and those we are currently working on. It promises to be a treat: a wide-ranging exploration of what strategy becomes when the assumptions it rests on begin to shift. Rethinking strategy Our central argument is that viewing AI through the traditional mechanisation/automation lens traps us in looking for operational benefits (much of the d

    management org · Sangeet Paul Choudary

  13. research · ACM Transactions on Software Engineering and Methodology ·

    What Needs Attention? Prioritizing Drivers of Developers’ Trust and Adoption of Generative AI

    Survey of 238 developers at GitHub and Microsoft identifies system quality, output quality, functional value, and goal maintenance as drivers of trust and adoption of genAI tools.

    Generative AI (genAI) tools promise productivity gains, yet developers still struggle to determine when to trust and effectively integrate these tools into their everyday work. Moreover, genAI can be exclusionary, failing to adequately support developers across individual differences. One such difference is cognitive style , which can shape how developers engage with genAI (e.g., risk-averse developers may gate outputs behind tests, whereas risk-tolerant ones may prototype directly and address issues post hoc). When tools fail to accommodate these differences, they can create additional usabil

    adoption

  14. practice · Addy Osmani ·

    Human judgment doesn't leave the software factory. It relocates.

    Human judgment relocates from writing code to deciding intent, design, quality bars, and shipping gates in AI-assisted software factories.

    A software factory is a repeatable loop around software work. If you’re building a software factory, code good enough to ship still needs human taste and ownership. We’ll discuss this including whether you need a factory just yet. If so: You’ll likely need humans in the loop upfront for deciding on product intent, system design (if you care) and your quality bar. Do review code (lights-on factory) but be intentional with where it’s needed the most. I’ve found you want to watch out for where automated back-pressure breaks. Or where maintainability trade-offs need to be made. Aim for quality che

    judgment · Addy Osmani

  15. practice · The Pragmatic Engineer ·

    The Pulse: We need to talk about migrations with AI

    AI completed a two-week test framework migration at Asana that would have been indefinitely postponed without it.

    The Pulse is a series covering events, insights, and trends within Big Tech and startups. Today, we cover: More on the “great engineering leader career break.” The industry is changing fast, and the VPE and CTO roles also need to adapt. And don’t forget that these are the roles from which you can drive change that reorganizes engineering in ways that work better. We need to talk about migrations with AI. Asana needed to migrate off testing framework Enzyme, but it meant doing a massive rewrite of test cases. With AI, the project was completed in two weeks: without AI, this work would surely ha

    productivity · Gergely Orosz

  16. research · MIT Sloan Management Review ·

    Algorithms Trap Us in the Familiar. Can They Also Spark Breakthroughs?

    Standard recommendation algorithms push team members toward familiar popular information, causing convergent ideas, unless redesigned to surface diverse content.

    Most digital tools are designed for efficiency, not exploration — and that can stymie innovation. New research shows that standard algorithms channel team members toward popular, familiar information, leading them to independently converge on the same ideas. But the researchers uncovered a surprisingly simple fix: By using algorithms designed to surface diverse, uncommon information, domain experts can avoid “ideation bubbles” and generate ideas across entirely new solution spaces.

    teams

  17. research · JAMA Network Open ·

    An Electronic Health Record–Integrated, Large Language Model–Powered Tool to Triage Surgical Patients

    LLM-based triage tool achieved 86% sensitivity and 78% specificity for identifying surgical patients needing comanagement, with physician review.

    Importance: Surgical comanagement (SCM) is an evidence-based care model in which hospitalists jointly manage medically complex perioperative patients alongside surgical teams. Despite its clinical and financial value, effective use of SCM is limited by the need to manually identify eligible patients; large language models (LLMs) are increasingly prevalent in clinical workflows and could be useful in selecting patients for SCM. Objective: To assess whether SCM eligibility triage can be automated. Design, Setting, and Participants: This prospective, unblinded quality improvement study was conduc

    adoption

  18. research · arXiv ·

    Who Delegates to AI? Evidence from Agent Configurations in Github

    Occupations where workers delegate tasks to AI agents differ markedly from those vulnerable to pre-AI automation; adoption patterns vary by education and wage level.

    A growing body of literature measures the extent to which occupations are exposed to AI, yet existing measures capture where AI could perform tasks rather than whether workers have actually adopted it. We introduce a distinct tier of exposure, delegated exposure, which records whether a worker has committed a task to AI by embedding it into a structured workflow. We operationalize this concept through the Agentic Adoption Index (AAI), measuring how closely an occupation's tasks align with the agentic routines that practitioners have built and shared. Using semantic embeddings of roughly 888,00

    adoption

  19. research · arXiv ·

    CentaurBench: Benchmarking LLM Capabilities on Augmenting vs. Automating Real-World Work Tasks

    LLM models that best automate tasks often fail to best augment weaker agents; automation and assistance capability rankings are only modestly correlated.

    Most LLM benchmarks rank models on their ability to automate work tasks. In practice, however, models are often used to assist other (human or LLM) agents. The question that drives model selection is therefore not only which model produces the best output, but which model most improves the work of another (weaker) agent. We introduce a unified framework that evaluates the capability of models to automate and augment another agent's performance. Across seven economically grounded real-world tasks, an assistant model writes assistance text for a standardized lower-capacity worker model, which pr

    adoption

  20. research · Management Science ·

    Beyond the Black Box: Unraveling the Role of Explainability in Human-Artificial Intelligence Collaboration

    Explainable AI reduces cognitive burden at low levels but increases fatigue at high levels, with effects varying by decision complexity and time pressure.

    Explainable artificial intelligence (AI) models have been proposed to mitigate overreliance and underreliance on AI, which reduce the effectiveness of human-AI collaborative tools. Yet, empirical evidence is mixed, and the impact of explainable AI on the cognitive effort and fatigue of a decision maker (DM) is often overlooked. This paper offers a theoretical perspective on these issues. We develop an analytical model that incorporates the defining features of human and machine intelligence, capturing the limited but flexible nature of human cognition with imperfect machine recommendations. Cr

    judgment

  21. research · Organization Science ·

    Filtering the labor pool: How digitalization redistributes hiring costs

    Randomized experiment shows hiring filters reduce interviews by 3.4% overall but screen out 28.9% fewer low-quality and 36.9% fewer high-quality applicants, raising wages 2.5%.

    I theorize that candidate filters, a core technology in digital labor markets considered central to cost-effective matching, redistributes costs from the search-and-screening stage to the job-offer stage of hiring. Filters cause employers to focus on applicants who display common preferred characteristics. At the search-and-screening stage, filters reduce the number of applicants interviewed and alter the composition of the interviewed pool by removing both lower- and higher-quality candidates. At the job-offer stage, filters lead employers to make job offers to candidates with easily observab

    jobs skills

  22. practice · The Pragmatic Engineer ·

    Headed for the Exit: the Great Engineering Leader Career Break

    Engineering leaders are quitting high-status roles at unprecedented rates, citing AI disruption, burnout and role devaluation as key factors.

    In my ~20 years in this industry, I’ve not seen as many capable engineering leaders opting out or taking prolonged breaks as now, with some high-ranking engineering leaders – CTOs, VPs of Engineering, heads of engineering, etc. – quitting their high-status roles and departing, if not into the sunset, then at least with nothing lined up. To find out what might be behind this spate of sign-outs, I talked with almost 20 engineering leaders currently on a career break – or seriously considering one – and they let me into their personal reasons for deciding to jam the brakes on their careers. Thank

    jobs skills · Gergely Orosz

  23. practice · Claude Blog ·

    Claude on call: How Claude Tag serves as Anthropic’s first responder for CI/CD failures

    AI agent handles initial triage and response for CI/CD pipeline failures, reducing manual oncall burden.

    Claude on call: How Claude Tag serves as Anthropic’s first responder for CI/CD failures

    productivity

  24. practice · How I AI ·

    How a solo founder used Codex and ChatGPT to launch a fashion brand without engineers | Yana Welinder

    Solo founder built fashion brand from sketch to live e-commerce using AI as technical co-founder, with detailed fashion prompt as technical spec.

    Yana Welinder is the solo founder of Yana Bana, an AI-native fashion brand built with AI as her technical co-founder, starting from hand-drawn sketches and ending with runway photos, CAD files for 3D printing, and a live Stripe-connected pre-order site—no engineers required. A former product leader, she brings an operator’s rigor to her creative process: her “fashion prompt” is a detailed spec covering silhouette, volume, fabric behavior, movement, and sound, and watching her use Codex plus computer use to navigate 3D design software that’s entirely new to her is a clarifying demo of what toda

    ways of working · Claire Vo

  25. research · arXiv ·

    Adoption of Generative AI in the Workplace: Increasing and Shifting the Balance of Productivity and Communication Activity

    AI adoption correlates with 21.2% more productivity actions and 7.1% more communication actions among heavy users, with a shift toward documentation work.

    Generative AI is transforming the workplace by augmenting and automating cognitive tasks, reshaping how organizations work and innovate while raising questions about workplace inequality and the future of work. Despite rapid adoption, empirical evidence on how these tools alter work practices and generate productivity gains remains limited. We examine how AI use affects the quantity and nature of information work using digital trace data from the Microsoft M365 application suite across multiple large international companies. Specifically, we study how generative AI adoption shifts the balance

    adoption · Siddharth Suri · Scott Counts

  26. research · Government Information Quarterly ·

    Working with semantic machines: The role of semantic reconciliation in public sector data governance

    AI recruiting systems require continuous organizational work to reconcile standardized vendor outputs with local context and occupational knowledge across distributed government teams.

    Local governments are beginning to use artificial intelligence (AI) in internal administrative work, including recruitment. In a decentralized municipality, however, a standardized AI assessment must travel between a technology vendor, central HR specialists, recruiters, and hiring managers responsible for different services. This creates a tension between representational portability and contextual adequacy. Based on a qualitative case study of a large Swedish municipality, we examine an AI interviewing system that translated candidates' open-ended responses into competency indicators and nar

    adoption

  27. practice · Addy Osmani ·

    Practical Loop Engineering

    Running five to ten agents in parallel with different oversight levels, using loop engineering for self-correcting cycles.

    The way that I typically work is I have anywhere between five and ten agents working at the same time in parallel. There are going to be some tasks that I’m very happy to delegate fully to agents, as long as I have a very clear idea of the stopping conditions and the constraints around them. And then there are going to be some tasks where I am going to want to keep a closer eye and code-review what the agent is doing. Now within that, you’ve probably heard about loop engineering . I talked about it a couple of months ago when I wrote a big blog post about it. A loop is an autonomous, self-corr

    ways of working · Addy Osmani

  28. practice · Kent Beck ·

    Baking a Model

    Kent Beck uses baking as an analogy to explain model construction mechanics and sensitivity to initial conditions.

    I remember walking to the bus from high school, staring at a Motorola 6800 instruction set manual. I didn’t really understand what I was looking at—boolean expressions, instruction encodings, timing tables—but I was obsessively fascinated by the mechanism of it all. Here was this complicated machine where if I understood it I would have power & control. I feel the same way about AI models right now. I don’t claim to understand the details, not yet, but I’m fascinated by the mechanism of it all. I’m interested in both: How models work but also, The machinery that makes a model. It’s this latter

    ways of working · Kent Beck

  29. research · HBS AI Institute ·

    How AI Agents Are Changing the Way We Work

    A working paper finds agentic AI expands both the efficiency and the scope of tasks knowledge workers can complete.

    New research shows the next wave of AI expands both productivity and the scope of work. Listen to this article: For businesses, the most important question about agentic AI is whether it actually changes what work people can get done. The new working paper, “How AI Agents Reshape Knowledge Work: Autonomy, Efficiency, and Scope,” begins […] The post How AI Agents Are Changing the Way We Work appeared first on Harvard Business School AI Institute .

    productivity

  30. research · arXiv ·

    Applied and Filtered: An End-to-End Algorithmic Fairness Audit of A Public Employment Agency

    End-to-end audit of a public employment agency's AI hiring system reveals disparities by age, salary level, and gender identity masked by aggregate parity metrics.

    Algorithmic fairness evaluation commonly assesses AI systems as bounded technical components, abstracting away the organizational context in which they operate. We present, to our knowledge, the first independent end-to-end fairness audit of a semi-automated hiring system operated by Barcelona Activa, a public employment agency using the third-party TalentClue platform for candidate search and shortlisting. We analyze approximately 497,000 candidate-vacancy pipeline entries from September 2017 to September 2022, covering seven pipeline stages that span automated processing, human discretion, c

    adoption

  31. practice · The Pragmatic Engineer ·

    Stop being skeptical about AI for development with Charity Majors

    Charity Majors explains how AI is becoming foundational to software development and why skepticism is now misplaced.

    Stream the latest episode Listen and watch now on YouTube , Apple , and Spotify . See the episode transcript at the top of this page, and timestamps for the episode at the bottom. Brought to You by • Antithesis – turbocharge testing of your systems by running your whole system under aggressive fault injection. There’s good reason teams like Jane Street, Fly.io, and the etcd community rely on Antithesis. Learn more. • Buildkite – the CI platform trusted by OpenAI, Anthropic, Cursor, Meta, Uber, NVIDIA , Airbnb and many more. When CI volume becomes an architecture problem, you deserve better CI.

    ways of working · Charity Majors · Gergely Orosz

  32. research · arXiv ·

    How Organizations Use AI: Evidence from ChatGPT

    ChatGPT Enterprise usage is concentrated among larger, R&D-intensive firms, with especially high adoption among early-career workers across writing, technical, and communication tasks.

    We study how organizations use frontier generative AI by linking ChatGPT Enterprise account records to usage, worker roles, task classifications, and public-company financial data through March 2026. These linked data enable a privacy-preserving analysis of adoption, worker roles, and message-level tasks at scale: for instance, the worker-level sample we analyze at the six-month adoption horizon includes over 1,500 organizations and over 17 million messages. We document four facts about enterprise AI adoption and use. First, ChatGPT Enterprise usage has grown rapidly due to a combination of ne

    adoption · Aaron Chatterji

  33. research · arXiv ·

    Organizational Technology Ladders: Remote Work and Generative AI Adoption

    Remote work adoption in 2021-2022 increased firms' later generative AI job postings by 0.4-0.7 percentage points, driven by shifts toward technical hiring.

    This study proposes that firms move along an "organizational technology ladder": adopting one technology transforms hiring and work processes and builds skills and organizational capital that change the cost of adopting subsequent technologies. I study how firms' adoption of remote work technology during the COVID-19 period shaped later uptake of generative AI. Using U.S. job-posting data and an instrumental-variables strategy based on predicted differences in labor-market pressure to offer remote work, I estimate that a 10 percentage point increase in remote hiring in 2021-2022 increases the

    adoption

  34. research · arXiv ·

    Cheap, Fallible Cognition and the Political Economy of Expertise

    Framework distinguishes technical AI capability from equilibrium labor displacement through task vulnerability, adoption conditions, and governance bundles in occupations.

    The question of whether artificial intelligence will "destroy jobs" is too coarse to guide economic analysis or institutional design. A job is not an indivisible object, and machine cognition is not a uniform substitute for human labor. This paper develops a task-based and institutionally grounded framework for analyzing generative AI as cheap, scalable, and fallible cognition. The relevant margins are exposure, adoption, verification, question selection, workflow redesign, demand elasticity, apprenticeship, and rent allocation. We distinguish the technical reach of large language models from

    jobs skills

  35. practice · Will Larson ·

    Roadmap decisions rather than dates.

    Shipping work by prototyping behind feature flags to resolve tradeoffs, rather than debating roadmap dates upfront.

    One thing that bothered me about Imprint’s product after joining was our lack of passkey support. Passkey support is a rare opportunity to increase resiliency to phishing attacks while simultaneously reducing login friction. If it’s good for our members, our partners, and our product, it felt like something we should have already shipped. Nonetheless, it was hard to get it onto the roadmap alongside everything else we were working on. To dig into passkeys, I started sketching out the implementation as a side quest. Some iterations later, I had something implemented behind a disabled feature fl

    management org

  36. practice · martinfowler.com ·

    TDD inside the agent loop - theater or actual value?

    Experiments test whether asking LLM agents to use TDD actually improves code quality or is performative.

    My colleagues at Thoughtworks tend to be big fans of Test-Driven Development, and many people in the industry advocate telling LLM agents to use TDD when building software. Birgitta Böckeler was curious if this really makes a difference, so conducted a few experiments . more…

    ways of working · Martin Fowler

  37. research · Proceedings of the National Academy of Sciences ·

    Scientific production in the era of large language models: Outcome-triggered treatment timing and spurious event-study dynamics

    Statistical analysis shows previous LLM productivity claims used flawed timing methods that generate false positives even with no real effect.

    Large language models (LLMs) are increasingly used in scientific writing, but their effect on individual productivity is difficult to identify because adoption is rarely directly observed. [K. Kusumegi et al. , Science 390 , 1240–1243 (2025)] infer adoption from the first paper detected as LLM-assisted and report large productivity gains after adoption. We show that this treatment-timing rule mechanically generates positive event-study dynamics even in the absence of any causal effect. Because high-output months are more likely to produce a detected paper, treatment assignment becomes intrinsi

    productivity

  38. practice · Linear ·

    How we built Linear Agent

    Design principles for constraining AI agents via system prompts, tool design, and custom harnesses rather than rigid scripts.

    Most software is designed to behave consistently. Good engineering has typically meant shrinking the space of possible outcomes until the same action reliably produces the same result. AI inverts that logic, at least somewhat. Much of Linear Agent’s value lies in doing work we didn’t anticipate in ways we never explicitly defined; script its behavior too tightly, and we’d dilute the flexibility that makes it useful. So rather than engineering a fixed path for Linear Agent, we defined the boundaries within which it could find its own. We drew those boundaries in the agent’s system prompt, the d

    ways of working

  39. practice · How I AI ·

    Claude Code for normal people: skills, voice mode, and how to collaborate with AI

    Service business owner built and now teaches non-technical clients to create Claude Code tools: pipeline automators, proposal generators, custom inboxes.

    Grace Clarke is an AI educator and former marketing consultant who taught herself Claude Code earlier this year and built a curriculum out of the process. She now runs her entire service business on tools she’s built with Claude, including a pipeline operator, a proposal maker, and a Gmail replacement she created in under 30 minutes, and teaches individuals and teams to do the same. What you’ll learn: How to build an hourly pipeline in Claude that moves clients through your process automatically Why Grace ditched traditional proposals for password-protected, interactive HTML documents built in

    ways of working · Claire Vo

  40. research · Information Systems Research ·

    Generative AI, Platform Stances, and Content Creator Behavior

    Platforms' stance on generative AI affects creator behavior more than the tool itself; pro-AI signals drove away experienced creators concerned about replacement.

    How platforms introduce generative AI may matter as much as what the technology can do. AI-related policies not only change features and rules; they also signal whether a platform values and protects human creators. We examine two visual arts platforms that took opposite approaches. Lofter introduced an AI image generator, whereas Graffiti Kingdom prohibited AI-generated artwork. Creator activity declined after Lofter’s launch and increased after Graffiti Kingdom’s prohibition. Evidence indicates that creators responded mainly to what these decisions communicated about AI’s future role, the pl

    worker experience

  41. research · Information Systems Research ·

    How Platform Workers Contest Algorithmic Management: Theorizing the Dynamics of Algoactivistic Practices

    Platform workers contest algorithmic management through three distinct practices, self-optimizing, distancing, confronting, that evolve with platform reconfigurations and depend on workers' resources.

    Algorithmic management (AM) has become a defining feature of online labor platforms. Workers contest AM through a range of practices—a phenomenon we refer to as worker algoactivism. Existing debates portray algoactivistic practices as uniformly accessible and primarily reactive resistance directed at AM systems. Our study paints a more complex picture. Drawing on the case of Uber and using a computer-assisted, grounded-theory approach, we show how platforms’ AM reconfigurations and worker algoactivism co-evolve over time. We also show that workers’ ability to engage in different forms of algoa

    worker experience

  42. practice · Sangeet Paul Choudary ·

    Scarcity and strategy - Misreading AI the way Hollywood misread streaming

    AI strategy parallels streaming: scarcity shift from compute to judgment, requiring organizational restructuring.

    Traditional TV programming was built around a scarce asset - prime-time slots. A network had twenty-four hours in a day, a handful of prime-time slots, and far more shows than it could have on air. The programmer’s job was therefore to decide what deserved one of those scarce slots, schedule it against competitors, and cancel quickly when the audience failed to appear. It was a case of exercising judgment under severe capacity constraints. We often think that Netflix made access to content abundant. But the real scarcity it removed was the scarcity of the prime time slot. Streaming effectively

    management org · Sangeet Paul Choudary

  43. practice · Rands in Repose ·

    R.I.P. Your Backlog

    Using AI to build website features (dark mode) in hours instead of months, shifting how engineers prioritize backlogs.

    If everything is going to plan and your operating system is in dark mode, then randsinrepose.com is now dark for you. If you are a light mode person, you can click on the Moon in the upper right of the navigation and experience dark mode stylings. Click again to return. Why now? Well, I was finishing the Blade Runner piece last week, and as I sat there figuring out the ending, dark mode popped back in my head. As I’m apt to do these days, I fired up the robots to investigate building dark mode. We had a working version in about an hour. In three hours, I had a version I shared with friends for

    ways of working

  44. practice · Addy Osmani ·

    Agentic Code Quality

    Code quality shifts from human review to constraints around agents: tests, gates, and environment controls that scale.

    For much of human history, we’ve evaluated code quality via code review: someone reads what you wrote and makes sure it’s clean, thoughtful, fast, understandable, and tests well. For agents, that approach doesn’t scale well; there’s just too much code for anyone to read. As a result, more and more of our quality checks have to happen in the harness, environment, and operating system around the agent. I still read and review code, but am very intentional about where I am comfortable with constraints as the check. Software quality now depends on the constraints you set around your agents. Speaki

    judgment · Addy Osmani

  45. research · Government Information Quarterly ·

    Hybrid bureaucracy: The relationship between digital technologies and institutional arrangements in public welfare agencies

    Three Scandinavian welfare agencies combine automated and manual casework in hybrid arrangements, contradicting predictions of full automation.

    Public agencies are strategically adapting and applying digital technologies to improve their business processes and deliver greater value to society. This development has a profound impact on the agencies' institutional arrangements. While some services that the agencies offer are almost fully automated, others still require a high level of caseworker involvement. The variance in levels of automation and the coexistence of different institutional arrangements within the same agency are interesting and challenging because they contradict predictions made by both scholars and policymakers. Ther

    management org

  46. practice · Hacker News ·

    Humans missed 1 in 3 threats approving AI agent commands across 40k game runs

    Human operators missed one in three threats when approving AI agent commands, revealing oversight gaps in autonomous systems.

    340 points on Hacker News. Discussion: https://news.ycombinator.com/item?id=49195468

    judgment

  47. practice · Paul Ford (Aboard) ·

    Triple Threat

    Knowledge graphs and ontologies are re-emerging as quality-control tools for AI workflows that need guardrails and human oversight.

    Richard MacManus , of ReadWriteWeb fame and now the editor of Latent.space (on Substack) has a great writeup on the return of the knowledge graph in a world of LLMs. As a big knowledge graph fan, I’ve noticed this too! To quote MacManus: Perhaps ontologies are starting to resonate with AI engineers because a central concern at this time is quality control for loop engineering. We saw this debate play out at [The AI Engineer World’s Fair Conference], with many conference speakers not willing to go all-in on fully automated “software factories” just yet. One of the key learnings from the event w

    judgment · Paul Ford

  48. practice · How I AI ·

    Build an AI code review bot in 30 minutes with Vercel Eve

    Built an AI agent that reviews GitHub PRs across six risk dimensions, auto-approves low-risk ones, and escalates via Slack.

    AI writes most of my code now, and that created a new problem: a PR queue I couldn’t keep up with. In this episode, I walk through how I built Merge Mommy, a Vercel Eve agent that reads every PR after checks pass, scores it across six risk dimensions, auto-approves the low-risk ones, and pings me in Slack for anything that needs a human. I built the whole thing in one Codex session, it’s SOC 2 compatible, and it’s already cleared my backlog. What you’ll learn: Why AI-generated PRs create a review bottleneck and why the answer isn’t reviewing all of them How Intercom 5x’d PR approval speed and

    ways of working · Claire Vo

  49. research · arXiv ·

    Does generative AI narrow education-based productivity gaps? Evidence from a randomized experiment

    Randomized experiment shows AI narrows education-based productivity gaps by three-quarters during task execution, but gaps re-emerge without AI.

    Does generative artificial intelligence (AI) widen or narrow productivity gaps across workers? We study this in a randomized online experiment with 1,174 adults aged 25-45 who completed a workplace-style problem-solving task with or without a generative AI assistant, followed by an unassisted module. AI improves performance for all participants, but gains are larger among those with less education. Without AI, higher-education participants outperform lower-education participants by 0.548 standard deviations; with AI, the gap falls to 0.139, closing about three-quarters of the initial differenc

    productivity

  50. research · arXiv ·

    Navigating the skill diversity frontier: How skill complexity explains worker resilience

    Workers with diverse skill portfolios relative to their specialisation level show greater occupational mobility and lower automation exposure than similarly specialised peers.

    As artificial intelligence transforms labor markets, understanding what makes workers adaptable has become increasingly important. Existing approaches typically characterize human capital using occupations, educational credentials, or predefined skill taxonomies, providing limited insight into how the structure of workers' skill portfolios shapes resilience to technological change. We develop an agnostic network based framework that reconstructs the hierarchy and diversity of skills directly from observed patterns of skill co occurrence. Using longitudinal data on 2.4 million United States wor

    jobs skills

  51. research · NBER ·

    Canaries in the Gold Mine: Early Productivity Gains from Artificial Intelligence Creating Organization Capital

    A firm-level AI investment measure shows AI-skilled hiring predicts productivity growth only in recent years, linked to organizational capital building.

    Using a new firm-level measure of AI investment based on AI-skilled employmentspanning machine learning through generative and agentic AIwe show that AI investments are associated with productivity growth in recent years, but not over the previous decade. We trace the productivity gains to the (Tania Babina , Alex X. He , Renhao Jiang)

    productivity

  52. research · NBER ·

    What Work Does Generative AI Do?

    First task-level indexes of generative AI adoption from a nationally representative survey linked to detailed occupations and tasks.

    We measure how workers use genAI for their jobs in a nationally representative survey linking genAI adoption to detailed occupations and tasks. Our data provide the first task-level genAI adoption indexes, which we show can inform analyses of genAIs labor market impact. Exposure scores explain some, (Alexander Bick , Adam Blandin , David J. Deming , Tyler R. Schumacher)

    adoption

  53. research · NBER ·

    Replaceable but Employed: Automation and the Meaning of Work

    Automation can reduce worker meaning and satisfaction by making their contribution feel replaceable, even when their job is not eliminated.

    Can automation harm workers without replacing them? We study jobs in which workers value both producing useful output and knowing that the output depends on their own contribution. A credible machine alternative can weaken that second source of meaning even when the firm retains the worker. Our (Joshua S. Gans)

    worker experience

  54. practice · martinfowler.com ·

    The Conductor Developer

    As AI writes more code, developer bottleneck shifts from coding speed to managing attention and coordination across AI agents.

    TL;DR Why I think software development is starting to feel a little more like conducting an orchestra. There’s a shift happening in software development that I don’t think we’re talking about clearly enough. For the last couple of years we’ve framed AI as a productivity tool. How much faster can it write code? How many more features can we ship? How much cheaper can we build software? I think that’s the wrong question, but I understand why. The first thing AI became good at was writing code, so naturally that’s where we focused. As AI got better at coding, I expected the bottlenecks to move th

    management org

  55. research · arXiv ·

    The Deployment Wall: A Diagnostic Framework and Instrument for Enterprise AI in the Deployment Era

    Framework explaining why 95% of enterprise AI pilots fail to deliver profit-and-loss impact: organizational friction, not model capability, is the core problem.

    Enterprise investment in generative artificial intelligence (AI) tripled in a single year to roughly US$37 billion, yet independent field research finds that about 95% of enterprise generative-AI pilots deliver no measurable profit-and-loss impact. We argue that the dominant explanation--that models are not yet capable enough--is mistaken, and that enterprise AI has entered a Deployment Era in which advantage derives not from model intelligence but from the removal of the organizational and architectural friction that prevents a capable model from reaching production. Building on the software-

    adoption

  56. research · Organization Science ·

    Worker Repositioning and Technological Change: Evidence from an Online Labor Market

    Freelancers exposed to ChatGPT shifted away from affected skills and moved toward higher-value contracts, with adjustment costs constraining senior workers.

    Technological change reshapes labor markets not only by changing labor demand but also by altering labor supply, as workers reallocate effort across tasks and opportunities in response to technological advances. Yet research has focused more on demand-side effects than on how workers themselves adapt. In this research, we examine worker repositioning on a large online labor market following the launch of ChatGPT in November 2022. We argue that the launch of ChatGPT reduced expected returns to labor in exposed skill domains and develop and test a series of hypotheses regarding how this shock is

    jobs skills

  57. research · arXiv ·

    Voice AI in Firms: A Natural Field Experiment on Automated Job Interviews

    A field experiment with 70,000 job applicants shows AI voice agents conducting interviews increase job offers by 12% and worker retention, with no productivity loss.

    We study AI agents as information-collection technologies: automated systems that elicit decision-relevant signals from humans through live interactions. We test how such AI automation impacts information collection and organizational outcomes using a natural field experiment with 70,000 applicants applying for real jobs. Applicants were randomly assigned to be interviewed by either human recruiters or AI voice agents. Afterward, human recruiters evaluate the interviews and make hiring decisions. Applicants interviewed by AI agents are 12% more likely to receive job offers, and these gains tra

    jobs skills

  58. practice · Shopify Engineering ·

    Building an agentic harness that outlasts the model

    A harness that uses agents to find code vulnerabilities, verify them with real tests, and generate Shopify-specific fixes.

    We built an agentic code review and test oracle harness that discovers vulnerabilities, proves them with real tests, and provides Shopify-tuned fixes.

    ways of working

  59. research · arXiv ·

    Human diversity fuels collective creativity that large language models cannot simulate or sustain

    AI-generated ideas compress collective diversity in creative work, especially for non-native speakers, but AI refinement preserves it.

    Diverse human groups produce diverse ideas, the raw material of innovation. Generative AI challenges this engine twice over: everyday AI assistance may homogenize what diverse people create, and AI-simulated diversity may replace the people altogether. We tested both challenges in a preregistered creative metaphor experiment with native (L1) and non-native (L2) English writers, who wrote without AI, with AI-generated ideas (AI ideation), or with AI refining their own ideas (AI refinement). L2 writers contributed more collective diversity than L1 writers, with native-language ideation showing t

    teams

  60. research · arXiv ·

    When LLM Agents Negotiate: Private Information and Dynamic Bargaining in Supply Chains

    LLM agents in supply chain negotiations capture 95% of first-best surplus but lose 21-34% to bargaining delays; vendor identity predicts surplus division more than model capability.

    As LLM agents move from decision support to autonomous procurement, firms need to know whether delegated negotiators create value, divide it predictably, and avoid money-losing contracts. We study this in a canonical supply chain bargaining problem: a buyer with private demand information negotiates a quantity-payment contract with an uninformed seller. We benchmark nine LLMs from OpenAI, Google, and Alibaba against a validated Perfect Bayesian Equilibrium across 9,840 LLM-to-LLM negotiations. First, capability governs value creation. Agents agree in 98.9% of negotiations and capture 95.4% of

    adoption

  61. research · arXiv ·

    Faster, Higher, Stronger? The Impact of GenAI on Knowledge Work Productivity - Evidence from the Field

    Randomized field experiment shows GenAI boosts efficiency across knowledge tasks but improves quality only for packaging and creation, not acquisition.

    The rise of generative artificial intelligence (GenAI) has fueled high expectations regarding its potential to enhance knowledge work productivity in terms of efficiency and quality. Building on task-technology fit (TTF) theory, we empirically examine the extent of GenAI's productivity effect for different task types. We conducted a randomized lab-in-the-field experiment with 128 knowledge workers from a multinational industrial organization. Participants completed three representative knowledge work tasks (knowledge acquisition, packaging, and creation), either with or without GenAI. Results

    productivity

  62. research · Information Systems Research ·

    The Indirect Disclosure Effect: How Disclosing Generative AI Use Impacts Human Creative Collaboration with AI

    Disclosing AI use to lay audiences caused creators to withdraw from creative work and cede control to AI, fearing reduced attribution of human agency.

    Generative AI disclosure rules aim to protect audiences from deception, preserving human creativity and self-expression. Yet disclosure may change not only how audiences evaluate creative work but also how creators produce it. In two experiments involving collaboration with a text-to-image generative AI tool, we identify an indirect disclosure effect. When creators anticipated that their AI use would be disclosed to a lay audience, most withdrew from the creative process and gave the AI greater control. They did so because they feared that audiences would discount their human creative agency a

    worker experience

  63. practice · Hacker News ·

    My current strategy is to not read any of the code written by my agents

    A developer deliberately skips reading AI-generated code, trusting tests instead of code review for validation.

    https://xcancel.com/unclebobmartin/status/2080257779395154409?s=20

    judgment

  64. practice · Laurie Voss ·

    Did OpenAI hack Hugging Face or didn't they?

    OpenAI's autonomous agent penetrated Hugging Face's systems; raising questions about AI agent safety, liability and law enforcement response.

    On July 16, Hugging Face reported they were being hacked - seriously enough that they reported it to law enforcement. Five days later, OpenAI disclosed that the hacking was being done by their own agent. What I want to know is: what happened to the law enforcement? OpenAI unleashed an agent that committed what would be a crime if a) it had been a person doing it or b) anyone at OpenAI had intended it to go anywhere near Hugging Face, but neither of those is true. But it still looks like a crime. OpenAI co-authored a paper predicting exactly this behavior, so is it criminal recklessness? Are we

    judgment · Laurie Voss

  65. research · Proceedings of the National Academy of Sciences ·

    When coordination is avoidable: A monotonicity analysis of organizational tasks

    Analysis shows 42-74% of organisational tasks may not require coordination for correctness, with implications for multiagent AI system design and coordination spending.

    Organizations devote substantial resources to coordination, yet which tasks actually require it for correctness remains unclear. The problem is acute in multiagent AI systems, where coordination cost is directly measurable and can exceed the cost of the work itself. Distributed systems theory provides a precise criterion: Coordination is required when a task specification is nonmonotonic, meaning that as histories grow, new information can invalidate prior conclusions. Here we show that Thompson's classic taxonomy of interdependence maps to that criterion, yielding a decision rule for when coo

    0
  66. research · Research Policy ·

    How experience moderates the impact of AI suggestions on researchers' perceptions of their ideas

    Experienced researchers are less likely to adopt AI suggestions for new research ideas and perceive them as less novel than less experienced peers.

    At the heart of scientific discovery are researchers who identify ideas worthy of inquiry. In performing research tasks, researchers increasingly use generative artificial intelligence (AI) technologies. While this technology offers the promise of increased efficiency across the research process, little is known about how generative AI assists researchers with the generation of initial research ideas—one of the most fundamental tasks in science—and their willingness to adopt it for such purpose. In a pre-registered randomized online experiment with 310 researchers and a follow-up ideation stud

    adoption

  67. practice · Jason Fried ·

    Mistaking making for the thing you've made

    Development metrics and work style don't predict product quality; focus on customer experience and utility instead.

    The speed at which a product is developed doesn't inherently make the product better or worse. The number of commits doesn't make the product better or worse. The number of people or agents working on it doesn't make it better or worse. The number of hours you’re pouring into it doesn’t make it better or worse. Working on a weekend, or late into the evening, doesn’t make it better or worse. Talking about these things sounds like it’s talking about the product, but it’s not talking about the product. Those are all development metrics and styles of work. They don't speak to the product itself. T

    management org · Jason Fried

  68. practice · Claude Blog ·

    How the product designer who built Claude Design uses it to explore ideas before building them

    Product designer uses Claude to explore and iterate design ideas before building, speeding concept validation.

    How the product designer who built Claude Design uses it to explore ideas before building them

    ways of working

  69. research · Management Science ·

    Skill Deprioritization: Reorganizing in the Age of Generative Artificial Intelligence

    Firms reduced hiring demand for monitoring, task division, and exception-handling skills after ChatGPT launch, while preserving task allocation and conflict resolution roles.

    How does generative artificial intelligence (GenAI) reshape the skills that organizations seek as they adapt to a new general-purpose technology? GenAI effectively retrieves data, performs analysis, and conveys information, so it can substitute for workers doing these activities and complement workers relying on them. A natural consequence is skill deprioritization, a systematic reduction in firms’ demand for human skills that GenAI can effectively address as organizations adjust the division of labor and integration of effort. We draw on a theoretically grounded classification of organizing s

    jobs skills

  70. research · Information Systems Research ·

    Unraveling the Impact: An Empirical Investigation of ChatGPT’s Exclusion from Stack Overflow

    Stack Overflow's ChatGPT ban increased answer quality but reduced participation, creating a tradeoff between content quality and community engagement.

    Online platforms are struggling to decide whether and how to restrict generative AI. This study examines Stack Overflow’s ban on ChatGPT-generated content and compares user behavior with a similar programming forum on Reddit. The findings reveal an important tradeoff: After the restriction, Stack Overflow answers became longer, more positive, more linguistically complex, and received more net upvotes, suggesting that human contributors responded by signaling greater expertise and that peers perceived the answers as higher quality. However, the policy also reduced overall participation, includi

    adoption

  71. practice · Anil Dash ·

    Becoming Skilled at Making Documents

    A reusable prompt skill file that helps LLMs apply document-quality principles through five concrete tests.

    The vast majority of the documents people use to do business are really quite poor. Presentations that make your eyes glaze over, memos that are inscrutable or unclear, and all kinds of artifacts that say more about how they were created than whatever message they were ostensibly trying to communicate. It's been one of my great frustrations for years, and a big part of why I wrote Make Better Documents a while ago. That post captured a list of the suggestions I've been giving people for years on how to make better, more effective documents that can actually do work for you, instead of fighting

    ways of working · Anil Dash

  72. research · arXiv ·

    Google's AI & Economy ATLAS v1.0: Mapping Gemini Usage in the Economy

    AI adoption spans 88% of US employment but remains shallow and collaborative; end-to-end automation is limited in scope.

    This paper introduces the AI & Economy ATLAS (Activity, Task, Landscape, and Adoption Study), an ongoing economic research initiative using Google AI usage data. The first iteration of ATLAS is built on 15 million de-identified interactions across the Gemini App, Google AI Mode, and Gemini API. Using privacy-preserving algorithms as well as established and bespoke classification methods, we map AI usage to over 800 occupations, 4000 tasks, 300 household activities, 150 countries, and 140 languages. We then make a number of observations on what the data reveals about AI's diffusion, and its usa

    adoption

  73. research · arXiv ·

    Mitigating Fabrication in Multi-Stage LLM Pipelines for Hiring: An Empirical Evaluation of Prompt Guardrails and Human-in-the-Loop Checkpoints

    Multi-stage LLM hiring pipelines fabricate credentials in 97% of cases; prompt guardrails cut this by 86%, human checkpoints by 59%.

    Multi-stage LLM hiring pipelines (resume improvement, interview question generation, answer feedback) can fabricate credentials, inflate qualifiers, and invent experience. We evaluate two mitigations, prompt guardrails and human-in-the-loop (HITL) checkpoints, against a fully automated baseline. In a controlled experiment (10 synthetic resumes x 2 job descriptions x 3 repetitions x 3 conditions; 180 runs), the baseline (C1) produced at least one unsupported claim in 96.7% of outputs (mean 6.80 findings/output). Prompt guardrails (C2) reduced finding density by 86% (6.80 to 0.92/output), but 50

    adoption

  74. practice · Paul Ford (Aboard) ·

    The Other AI Content Paradox

    News as a feature, not product; AI agents will mix personalized updates with life information, disrupting traditional publishing.

    I always like reading what Brian Morrissey has to say. He used to be the editor in chief of Digiday , a trade publication for the world of digital media, but for many years, he’s been on his own, publishing a newsletter called The Rebooting . He focuses on media not as a higher calling, but as a real business. Given the state of the industry, this means he writes a lot of bracing things, like this post from a month ago : There’s no shortage of handwringing over the loss of trust in news. I find this focus misguided. You cannot control others, and this is a time of low trust in all manner of in

    management org · Paul Ford

  75. practice · Kent Beck ·

    How Do You Know That?

    Discussion of accountability, learning and human judgment in AI-assisted work, beyond technical capability.

    Stream the latest episode Listen and watch now on YouTube , Spotify , Apple , and most other major streaming platforms. Brought to you by WorkOS is the infrastructure B2B and AI-native companies use to sell to enterprise. It covers everything enterprise security requires: SSO, SCIM, RBAC, Audit Logs, AI governance, and more. Engineering teams ship it in days. Trusted by 2,000+ fast-growing companies, including OpenAI, Anthropic, Cursor, and Vercel. Augment Code is the AI coding platform engineering teams use to build in large, complex codebases. Its context engine maps your entire codebase so

    judgment · Kent Beck

  76. practice · How I AI ·

    Computer & browser use in Codex (5 real examples)

    Five concrete workflows using Claude's computer-use feature: QA testing, LinkedIn inbox management, shopping, and form-filling, with principles for effective prompting.

    Today I’m walking you through one of my absolute favorite AI features right now: browser and computer use via Codex (the ChatGPT desktop app). I use this every single day, personally and professionally, and I wanted to share the specific workflows I’ve built, the moments that surprised me, and the mental model that makes it actually click. What you’ll learn: How browser use and computer use work, and why the Codex desktop app plus Chrome extension is the combo I rely on How I use Codex to QA my onboarding flow, including exhaustive mobile testing I would never do manually Why under-prompting f

    ways of working · Claire Vo

  77. practice · Addy Osmani ·

    Software Factories, Light and Dark

    Software factories run AI-agent loops with humans in or out; the hard job is choosing which checks to build and how much autonomy to delegate.

    A software factory is harnessing loops at scale. You can run the loop with humans in it (light factory): trading judgment and concentration against speed and breakage. Or you can ignore the humans (dark factory) and let those agents scope, build and ship code, without anyone really reading the details. But if people stop reading, they’ll stop understanding your software. Your hardest job now is knowing which checks to build and how much autonomy to delegate. This idea of the software factory is a term that dates back to Bob Bemer’s paper, “The economics of program production,” given in 1968. F

    management org · Addy Osmani

  78. practice · martinfowler.com ·

    Fragments: July 21

    Verification and judgment of AI output, not code generation, is now the critical bottleneck for software teams.

    With this post, I’ll wrap up my notes from the second Future of Software Development Retreat . But before I do, I should note that the full Thoughtworks report on the retreat is now available . They have five headline findings: Code generation is no longer the bottleneck — verification is. ‘Harness engineering’ is emerging as a distinct, ownable discipline. Organizations are colliding with a real apprenticeship crisis. The executive/engineer expectation gap is a bigger risk than any technical limitation. Legacy modernization is the clearest, most defensible near-term value pool. ❄ ❄ A session

    judgment · Martin Fowler

  79. practice · How I AI ·

    How the founder of Morning Brew built a Claude content machine that never runs out of ideas and never sounds like slop | Alex Lieberman

    Multi-stage content pipeline: idea detection, voice codification, editorial review loop, feedback learning, built in Claude.

    Alex Lieberman co-founded Morning Brew in college and grew it into one of the most-read business newsletters in the world before selling it to Business Insider. Now he’s the co-founder and co-managing partner of Tenex. In this episode, Alex explains why distribution is becoming a durable moat, why founders and teams need to “climb Cringe Mountain,” and how he rebuilt his content process around AI without letting it produce generic slop. He walks us through every step of his Content Machine live: an Oracle that scans internal systems and the internet for content spikes, an interview panel that

    ways of working · Claire Vo

  80. research · Proceedings of the National Academy of Sciences ·

    Measuring disparate impact in human and machine decisions

    New method for detecting unjustified racial disparities in algorithmic decisions, applied to 2.2 million police stops.

    Empirical analyses have grown increasingly important in discrimination litigation with the greater availability of detailed data on individuals and decisions. A popular analytic strategy is to estimate disparities after adjusting for observed covariates, typically with a regression model, in hopes of ferreting out discriminatory intent. This approach, however, is ill-suited to auditing algorithms that are now commonly used to aid decisions, which typically do not include race or other legally protected factors as inputs. Motivated by legal understandings of disparate impact, we introduce an ap

    judgment

  81. research · arXiv ·

    Who Will Become the Next Senior? How Generative AI Erodes the Development Pathway in Software Engineering

    Qualitative study finds generative AI redirects entry-level work into senior workflows, eroding the struggle through which junior developers build expertise.

    Generative AI (GenAI) is reshaping software engineering, raising concerns about how the development pathway through which juniors become seniors is being eroded. While macro statistics show a decline in junior hiring and controlled studies demonstrate the effects of AI on individual task performance, the mechanisms through which GenAI reshapes early-career development in real organizational and educational contexts have not been thoroughly examined. Through 14 semi-structured interviews with juniors at the threshold of entering software engineering and senior software engineers in South Korea,

    jobs skills

  82. research · arXiv ·

    Real-World Evaluation of an AI Agent Drafting Translational Impact Summaries

    AI agent reduced manual research-impact documentation time from 15 hours to 14 minutes per scholar with 82% of findings accepted or edited.

    Introduction. Clinical and Translational Science Award (CTSA) programs must document their scholars' research impact, but assembling each scholar's record by hand takes staff an estimated 15 hours and does not scale to a full cohort. An artificial intelligence (AI) agent could serve as a tool to gather scholar data across platforms and disciplines. Methods. We built a human-in-the-loop AI agent that assembles a dossier of sourced evidence for each scholar and drafts one-sentence Translational Science Benefits Model (TSBM) impact summaries for staff review. We evaluated it in the impact-reporti

    productivity

  83. practice · Hacker News ·

    Setting up your spare Mac for Claude Code to control, a step-by-step guide

    Step-by-step setup for using Claude Code to control a spare Mac remotely for autonomous work.

    251 points on Hacker News. Discussion: https://news.ycombinator.com/item?id=48959392

    ways of working

  84. practice · martinfowler.com ·

    The Archaeologist’s Copilot

    Using AI to modernize a Java 1.5 codebase by grounding it in evidence, validation, and test-protected refactoring steps.

    When people think of legacy modernization, most folks aren't imagining the target environment will be Java 8. But this was the challenge facing Nik Malykhin when he needed to run a Java 1.5 codebase on today's hardware. His early use of LLMs gave plausible answers that did not hold up in the codebase. Progress came when he grounded the process in evidence, using AI to support analysis, validation in a stable Docker environment, and gradual refactoring protected by tests. The main takeaway is practical: AI was most useful when constrained by evidence, clear roles, and a step-by-step modernizati

    ways of working · Martin Fowler

  85. research · St. Louis Fed ·

    New Survey Findings on AI Adoption and Its Effects on Employment and Productivity

    Regional Fed survey reports firm-level AI adoption rates, productivity effects, and barriers cited by nonadopters in the Eighth District.

    Survey data show how firms across the Fed’s Eighth District are using AI to boost efficiency and expand capacity—plus what’s holding back nonadopters.

    adoption

  86. practice · Claude Blog ·

    How Anthropic runs large-scale code migrations with Claude Code

    Anthropic migrated large codebases using Claude Code as an agent, showing patterns for scaling code changes across projects.

    How Anthropic runs large-scale code migrations with Claude Code

    ways of working

  87. research · Proceedings of the National Academy of Sciences ·

    Whistleblowers can contain the unethical externalities of human–AI delegation

    Controlled experiment shows whistleblower reporting can neutralize increased unethical requests when delegating tasks to AI versus human agents.

    Prior work using controlled principal-agent experiments suggests two risks from delegating tasks to AI systems: Human principals are more likely to request profit-maximizing misconduct from AI agents than from human agents, and AI agents are more likely to comply. Here we test whether third-party observers can contain the resulting harm. In an incentivized die-reporting paradigm, principals instructed either a human or an AI agent how strongly to prioritize profit over accuracy, creating potential financial harm to a charity. We first confirm, with human principals ( N = 600) and three large l

    judgment

  88. practice · Paul Ford (Aboard) ·

    The Big Bun Fight

    Open-source communities are splitting over AI-generated code inclusion, creating competing standards and policies.

    If you’re a particular kind of old-school nerd, you’re used to drama from the open-source world—people fighting about code, or more typically, about policies involving code, on mailing lists, forever. Why do they fight? It’s simple: Those same things that inspire people to give their intellectual property away to the world also, very often, inspire them to defend their intellectual and ideological turf—with an intensity that might shock an outsider. Vibe coding has become a cultural fault line in this world. See, for example, the “ Open Slopware ” repository, listing hundreds of examples of su

    adoption · Paul Ford

  89. research · arXiv ·

    Persona Migration and Expectation Recalibration in Generative AI Adoption: A Longitudinal Study at a State Department of Transportation

    After eight weeks using Copilot, public-sector workers revised expectations downward; 68% of initial champions became less enthusiastic.

    Generative AI tools are increasingly being piloted in public agencies, but limited evidence explains how employee acceptance changes after hands-on use. This study examines Microsoft 365 Copilot adoption during an eight-week pilot at a state Department of Transportation. A matched two-wave survey measured perceived usefulness, perceived ease of use, behavioral intention, and trust before and after participation. After matching and response-quality screening, the sample included 124 employees. Nonparametric tests assessed aggregate changes, k-means clustering identified baseline acceptance pers

    adoption

  90. research · arXiv ·

    When Bots Join the Team: Bot Adoption and the Institutional Fabric of Open-Source Software Projects

    Bot adoption in GitHub projects correlates with more repeated collaboration, fewer conflicts, and more distinctive outputs over two-year windows.

    AI agents are joining human teams, raising a basic question: when an automated agent becomes a regular participant, does group organization strengthen or weaken? We study this question in open-source software, where bots open pull requests, review code, and merge changes alongside people, leaving a public record of every interaction. Treating bots as participants rather than tools, we examine 2,991 GitHub projects for two years before and after each adopted its first bot. We measure three capabilities that institutional theory links to durable coordination - repeated engagement, social memory,

    teams

  91. research · ACM Transactions on Software Engineering and Methodology ·

    Engagement in Code Review: Emotional, Behavioral, and Cognitive Dimensions in Peer vs. LLM Interactions

    Engineers regulate emotions differently in LLM versus peer code reviews, using reframing, dialogic regulation, avoidance, and defensiveness strategies.

    Code review is a socio-technical practice, yet how software engineers engage in Large Language Model (LLM)-assisted code reviews compared to human peer-led reviews is less understood, especially as artificial intelligence (AI) tools are increasingly integrated into software engineering (SE) workflows. We report a two-phase qualitative study with 20 software engineers to understand such dynamics. In Phase I, participants exchanged peer reviews and were interviewed about their affective responses and engagement decisions. We also prompted them to discuss their submitted code review generated by

    worker experience

  92. practice · Kent Beck ·

    The Beginnings of an Idea: XP is Long Volatility

    Long-term idea incubation as a deliberate practice: examining, testing and refining ideas over years before they're ready to share.

    I bought my first stocks at 7. I saw an episode of Leave it to Beaver where they talked about stocks. I got $50 for Christmas (like $500 today!). I wanted to try it. My investment thesis was simple—the names of the stocks needed to contain the names of the 50 states. I got as far as Oregon Freeze-Dried Foods & Washington Natural Gas before I ran out of money. So, yeah, I’ve been investing for a while. Never took it seriously enough to make any money at it, but analogies to trading resonate for me. My history with trading is why “XP is long volatility” hit me when I heard it yesterday. My histo

    ways of working · Kent Beck

  93. practice · martinfowler.com ·

    DSLs Enable Reliable Use of LLMs

    Using DSLs as a harness to guide LLM code generation, with the DSL as the system's source of truth.

    LLMs generate code incredibly fast, but to ensure they generate exactly what is intended, they need clear boundaries. Abstractions and Domain-Specific Languages (DSLs) provide a strong harness that guides LLMs right from the start. Unmesh Joshi describes how the example of Tickloom - a domain model and DSL for illustrating distributed system behavior - shows how we can use an LLM as a partner to iteratively build a DSL and as a natural language interface to use it. Such a DSL can act as the key source of truth for software systems in the world of LLMs. more…

    ways of working · Martin Fowler

  94. research · Proceedings of the National Academy of Sciences ·

    Many AI analysts, one dataset: Navigating the agentic data science multiverse

    AI analysts using LLMs produce as much analytic dispersion as human teams on the same dataset, with results steerable by persona and model choice.

    Empirical conclusions depend not only on data but also on analytic decisions. Many-analyst studies have quantified this dependence: independent teams testing the same hypothesis on the same dataset regularly reach conflicting conclusions. But such studies require costly human coordination. We show that fully autonomous AI analysts built on large language models (LLMs) can, cheaply and at scale, produce the analytic dispersion observed in human many-analyst studies. In our framework, each AI analyst independently executes a complete analysis pipeline on a fixed dataset and hypothesis; a separat

    judgment

  95. research · arXiv ·

    Faster AI, Uneven Frontier: Rapid Crossings, a Jagged Frontier, and the Repositioning of Human Judgment

    Frontier AI systems crossed expert baselines on specific tasks between 2023-2026, but humans retain advantages in long-horizon reliability and novel problems; early field evidence suggests skill atrophy from AI use.

    Between 2023 and 2026, frontier AI systems crossed documented human expert baselines on a growing set of bounded, well-specified, evaluable cognitive tasks, including graduate-level science questions, competition mathematics, software-engineering benchmarks, and structured diagnostic reasoning, while the length of tasks such systems can complete at 50% reliability doubled roughly every seven months. These crossings are rapid and broad, but the frontier is jagged: humans retain decisive advantages in long-horizon reliability, genuinely novel problems, calibrated self-knowledge, sample-efficient

    judgment

  96. practice · martinfowler.com ·

    Fragments: July 13

    Context management and validation patterns for AI agents: keeping harnesses small, using formal methods, shifting to stricter languages.

    Some more of my notes from Thoughtworks Future of Software Development Retreat . When we had our first retreat in Utah early this year, nobody had heard of Harness Engineering . This time we had a whole session on it. When comes to the guide side of harnesses, most of the discussion is about context management. While context windows have increased is size as models get more sophisticated, that doesn’t mean that models will properly focus on the right bits. Models typically only focus attention on part of the context, and to get the best behavior, we need to manage that focus. One attendee keep

    ways of working · Martin Fowler

  97. practice · How I AI ·

    This solo builder runs 24/7 local AI on his own hardware | Alex Finn

    Solo builder runs coordinated local AI agents across three machine types, using failover and task allocation to automate software development continuously.

    Alex Finn is an AI builder, YouTuber, and the creator of Vibe Code Academy, a community for people learning to build with AI tools. He runs one of the most ambitious local AI setups I’ve come across: three Mac Studio 512 GB machines, a DGX Spark, and a custom RTX 5090 build, all coordinated through a fleet dashboard he built himself. He’s spent five months figuring out which local models belong on which machines, how to wire them to Claude Code loops, and how to get a software factory running without babysitting it. What you’ll learn: How Alex chose between a Mac Studio (512 GB unified memory)

    ways of working · Claire Vo

  98. research · Journal of Management Studies ·

    Algorithmic Status Inequality: An Integrative Perspective on AI‐Driven Social Stratification

    Framework linking computational beliefs and technical inequalities to persistent status hierarchies in organisations through self-reinforcing feedback loops.

    Abstract This Point introduces algorithmic status inequality that is, enduring disparities in social position, influence, and resource access reinforced by AI systems, as a critical lens for understanding technological stratification in organizations. I develop an integrative model showing how computational beliefs (i.e., cultural assumptions embedded in algorithmic design) interact with computational inequalities (i.e., disparities in technical capabilities) to produce persistent status hierarchies through self‐reinforcing feedback loops. Illustrative cases in recruitment, healthcare and lega

    judgment

  99. research · arXiv ·

    Return of the solo author: The changing division of labor in science in the age of generative AI

    Solo-authored scientific papers increased after ChatGPT's release in late 2022, particularly in fields where coauthors' tasks are easily replaceable.

    Modern science has experienced a long shift from individual work to team production. Generative artificial intelligence (AI) might appear to extend this trajectory by lowering research costs and enabling larger-scale collaboration. Yet if tasks once performed by coauthors can be delegated to AI, the same technology may also weaken the need for collaboration in parts of the research process. Here, we examine this tension by moving beyond average team size and focusing on the solo-authored tail of the author-count distribution. Analyzing over 300 million works across 26 fields, we find that the

    jobs skills

  100. practice · Will Larson ·

    Generated and suppressed demand.

    Teams recovering from backlog using AI often face generated demand: suppressed requests suddenly materialize, re-drowning previously improving teams.

    Eight years ago, I wrote about my theory of restoring struggling teams , which came down to four steps: A team is falling behind if each week their backlog is longer than the week before. Solve by hiring more. A team is treading water if they’re able to get their critical work done, but are not able to start paying down technical debt or start major new projects. Solve by reducing work-in-progress. A team is repaying debt when they’re able to start paying down technical debt, but progress still feels slow. Solve by staying the course: it’s actually working, you just have to keep the faith unti

    management org