The Feed · Complete archive

AI at work: research and practice

A server-rendered record of the evidence, ideas and firsthand practices screened by The Feed. Every entry links to its original source.

Use the interactive Feed

Page 3 of 9 · 873 items

  1. practice · GitHub Blog ·

    Better tools made Copilot code review worse. Here’s how we actually improved it.

    Rewriting agent instructions to match how reviewers actually work reduced code review cost 20% without losing quality.

    Give an agent better tools and it should do better work. That’s the instinct, anyway. When you open a pull request, Copilot code review reads the diff and explores the surrounding code to find the problems that matter before they ship. To do that, it used its own code exploration tools. So when we swapped in the better-maintained, shared tools that power the Copilot CLI, grep , glob , and view , we expected a clean upgrade. Instead, in our benchmarks, we found that the cost of reviews was higher and fewer issues were being caught. But the tools weren’t the problem. The instructions were. Once

    ways of working

  2. research · Management Science ·

    The Uneven Impact of Generative Artificial Intelligence on Entrepreneurial Performance: Evidence from a Field Experiment in Kenya

    Field experiment shows AI business assistant reduced low-performing entrepreneurs' revenues by 10% while high performers gained 15%, revealing uneven adoption effects.

    Scalable and low-cost artificial intelligence (AI) assistance has the potential to improve firm decision making and economic performance, particularly in emerging markets. However, running a business involves a wide range of open-ended problems, making it unclear whether and how recent advances in AI can help business owners around the world make better decisions. In a field experiment with Kenyan entrepreneurs, we evaluated the impact of AI advice on small business revenues and profits by randomizing access to a GPT-4-powered AI business assistant. Although we are unable to reject the null hy

    adoption · Rembrand Koning

  3. research · arXiv ·

    AI Adoption in S&P 500 Firms

    11% of S&P 500 firms have deeply integrated AI into business processes in 2025, up from 5% in 2022, with no measured productivity gains yet.

    The adoption of artificial intelligence (AI) by large enterprises is an important potential source of aggregate productivity improvement and labor market impact. We study AI adoption of S&P 500 firms over the period 2016 to 2025, estimating adoption at the enterprise level. While generative AI tools are useful for personal and professional applications, our focus is on the deep integration of AI in the business processes of large enterprises which are bellwethers for firm adoption more broadly. We develop a novel measure to assess deep AI adoption (and distinguish it from AI hype) that is base

    adoption

  4. practice · Addy Osmani ·

    Own the Outer Loop

    Engineers must own the outer loop: accountability for agentic systems through quality checks, verdicts, and answerability before shipping.

    In the past year, the conversation around agentic engineering has moved to harnesses and loops , fleets and software factories . My 2c is engineers need to own the outer loop - the accountability for these systems. This only gets more true as powerful models like Fable and GPT-5.6 become available. Agents have leverage, and leverage creates obligations. Someone must be able to explain exactly what changed, why it was safe, and what will happen if they’re wrong. Otherwise, their actions can’t be justified. Which makes it unlikely their organization will ask for them in the first place. And so I

    management org · Addy Osmani

  5. research · arXiv ·

    How Analysts Use AI in High-Stakes Crime Linkage: An Industrial Study

    Crime analysts using an AI-enabled tool selectively validated predictions against behavioural evidence, showing partial trust and reliance on established practices.

    Crime linkage analysis is used in many countries to identify series of offences that may have been committed by the same individual. In practice, specialist analysts manually search for behavioural and situational connections across large crime databases, an effort that is time-consuming, cognitively demanding, and can involve repeated exposure to disturbing material. To support this work, an Artificial Intelligence (AI)-enabled decision-support tool was co-developed with a UK law enforcement agency to assist analysts in identifying likely crime linkages. This paper reports an industrial evalu

    judgment

  6. research · Personnel Psychology ·

    Inter‐Rater Reliability of AI Scores: A Call to Move from Correlational to Factor Analytical Estimation

    Factor analysis better estimates AI reliability for text scoring than correlation when AI and human raters differ systematically.

    ABSTRACT Artificial intelligence (AI), such as large language models (LLMs), is increasingly being used to score psychological and organizational constructs from text. However, the reliability and validity of these scores remains vital. Frequently, AI scores are treated as replicants of human ratings and inter‐rater reliability (IRR) is established by correlating AI and subject‐matter‐expert (SME) ratings. Yet, because LLMs may meaningfully differ from humans in ability to score text, correlating LLM scores with SME ratings is inappropriate because correlational approaches assume that AI and h

    judgment

  7. research · Léonard Boussioux ·

    The Hidden Cost of AI-Assisted Creativity

    AI assistance raised individual idea quality but reduced the diversity of ideas generated across groups, per new research.

    MIT Sloan Management Review · with Anil Doshi, Oliver Hauser & Kartik Hosanagar — AI can raise the average quality of individual ideas while narrowing the variety of ideas produced across a group.

    teams · Léonard Boussioux

  8. practice · GitHub Blog ·

    Automating cross-repo documentation with GitHub Agentic Workflows

    Agentic workflow automatically drafts feature docs in a separate repo within 45 hours of product ship, reviewed by engineers, with no new hires or process retraining.

    “Where are the docs?” It’s a question nobody on a product team enjoys answering. The honest reply is usually some variant of “behind.” A writer is staring at a closed pull request, trying to reverse-engineer what changed. The pull request’s author has already moved on. By the time the doc actually publishes, the feature has shipped, sometimes more than once. That used to be us on the Aspire team (we’re a small team of 10 building dev tools for distributed apps). A few months back, we were trying to figure out how to safely bring AI into automations we already trusted. That’s when we discovered

    adoption

  9. practice · martinfowler.com ·

    Experiences with local models for coding

    Developer evaluated local LLMs for coding tasks with standard benchmarks and real-world use, reporting practical results and tradeoffs.

    Birgitta Böckeler now reports on her recent experiences trying local LLMs for coding. She compares them using two standard tasks, and tries out the most promising model for day-to-day use. more…

    ways of working · Martin Fowler

  10. research · MIT Sloan Management Review ·

    GenAI Success Metrics: Look Beyond Reduced Workload

    Four-year observational study of GenAI adoption in a large university found stable staffing and hours despite AI tool introduction to leaders and professionals.

    Matt Harrison Clough / Ikon Images The Research The authors performed a four-year, fixed-window observational analysis of administrative work inside a large U.S. public higher-education institution. Generative AI tools were introduced to executive leaders, operational leaders, and student-facing professionals throughout the organization in 2026. Staffing levels and work hours remained stable across the period studied. […]

    adoption

  11. practice · How I AI ·

    What a harness is and how to build one with Claude Agent SDK

    Built a custom bug-triage agent harness with Claude SDK, showing architecture, code structure, and process for encoding permissions and integrations.

    Everybody is saying, “It’s not the model, it’s the harness,” but almost nobody stops to explain what a harness actually is. So I did. I built one live on the show: a Sentry bug-debugging harness for my company ChatPRD, using the Claude Agent SDK, a custom terminal UI built with the Ink library, and opinionated adapters for Sentry, Linear, GitHub, and Vercel. The harness handles evidence gathering, root-cause analysis, and follow-up artifact creation, all without me needing to type “dear agent, please fix this bug” ever again. I also walk through the architecture, share the code structure, and

    ways of working · Claire Vo

  12. research · Indeed Hiring Lab ·

    AI and Job Postings: From Destruction to Creation?

    Job postings may be growing faster in roles with high AI exposure, reversing earlier automation-displacement patterns.

    Agentic AI may be flipping the relationship between AI exposure and job posting growth. The post AI and Job Postings: From Destruction to Creation? appeared first on Indeed Hiring Lab .

    jobs skills

  13. research · Indeed Hiring Lab ·

    AI Is No Longer Just a Tech Occupation Story: It’s Spreading Across Job Titles in the US and Europe

    AI job-title mentions are spreading beyond tech roles across US and European job postings.

    Employers on both sides of the Atlantic are writing AI into job titles for a wide range of roles — not just software and data jobs. The post AI Is No Longer Just a Tech Occupation Story: It’s Spreading Across Job Titles in the US and Europe appeared first on Indeed Hiring Lab .

    jobs skills

  14. research · arXiv ·

    Digital Fragmentation and Generative AI Use Across 103 Million Application Events

    Analysis of 103 million application events shows generative AI use occurs on fragmented days, but is followed by narrower, longer, more predictable application sequences.

    Knowledge workers switch between applications thousands of times per day, spending nearly a tenth of the work year transitioning between digital applications in a process called digital fragmentation. Whether this fragmentation reflects who an employee is, where they work, or what kind of day they are having, has remained an open question. We analyzed 103 million application events recorded second-by-second from 1,017 employees across eight organizations that largely employ knowledge workers (e.g., law, financial services). Day-to-day variation in fragmentation within individual employees acco

    adoption

  15. practice · martinfowler.com ·

    Viability of local models for coding

    Factors that determine whether local LLMs are practical for programming work, based on hands-on testing.

    Birgitta Böckeler recently spent some time trying out running local LLMs for some programming tasks. In this memo she outlines the factors that influence how viable they are for the job. more…

    ways of working · Martin Fowler

  16. practice · martinfowler.com ·

    Fragments: July 6

    Teams are shipping agentic systems in production; patterns emerging but effectiveness criteria still unclear.

    Last week, Thoughtworks ran a second Future of Software Development Retreat , this time in Europe. As with the previous event, I’ll be sharing some fragmentary thoughts on this. There were five parallel streams, so I could, at best, only attend ⅕ of sessions. This isn’t an event that forms conclusions, rather one that allows those exploring to share what they’ve found, and their visions for the future. The bliki post lists all the writing I’ve run into on this, by myself and others. I’ll be updating it as more posts appear. Giles Edwards-Alexander “noticed a real difference between the retreat

    ways of working · Martin Fowler

  17. practice · How I AI ·

    How I run autonomous coding agents from my phone with OpenAI Symphony + Linear | Alessio Fanelli (Kernel Labs)

    Running autonomous coding agents from a phone using OpenAI Symphony and Linear as a state machine, with zero monitoring required.

    Alessio Fanelli , founder of Kernel Labs and co-host of Latent Space podcast, walks us through two very different AI workflows: (1) a fully autonomous coding setup using OpenAI Symphony + Linear, where Linear acts as a state machine and Symphony manages agents through the whole dev lifecycle with zero babysitting; (2) Codex with browser access searching eBay for underpriced Pokémon cards—autonomously browsing, extracting PSA certificate numbers, and flagging deals on $10K–$20K cards for his San Carlos card shop, Merlin Games. What you’ll learn: Why “agent manager” is a better mental model than

    ways of working · Claire Vo

  18. practice · Laurie Voss ·

    AI has torched the market for junior programmers

    Junior programmer employment fell 19% since late 2022, while experienced developers grew; new coding roles emerged outside traditional programming titles.

    In early 2025 I predicted that AI will create many, many more programmers , and that new programming jobs would look different. In March I checked in and found startups substituting compute for labor at record rates , with the wave of new jobs nowhere in sight. This post is the next check-in, and I have good news and bad news. The bad news: AI has torched the market for junior programmers. The good news: the long tail of new programmers I predicted has materialized, but with a big twist: they don't call themselves programmers . Let me show you the data, and see if you believe me. The market fo

    jobs skills · Laurie Voss

  19. practice · Addy Osmani ·

    Agentic Autonomy Levels

    Framework for choosing autonomy levels for different AI agent tasks, from low-risk reversible actions to manager agents delegating to fleets.

    In most conversations about agentic engineering, the action has changed from prompting to operating . Here’s a frontier looking into the fog: software factories, goals, loops, background sessions, subagents, hooks, sandboxes, agent-approving agents . For many creators of the future, this behavior will be baked into products day-1: Claude Code and Codex expose the shift directly. From the engineer standpoint, you’ll use low autonomy to limit risk and increase reversibility, but use higher autonomy for explicit activities, and fleets of parallel agents safely refactoring massive codebases. The c

    ways of working · Addy Osmani

  20. practice · Charity Majors ·

    In defense of AI mandates

    AI mandates can work when framed as collective learning investments with explicit permission to slow down and miss deadlines.

    I’ve been writing a series of pieces on lessons learned from our AI journey at Honeycomb . I’ve written about the tension between enthusiasts and skeptics , the need for engineering rigor , and the ethics of using tools with externalities and a seedy backstory . Today I want to send you off to a long July 4th weekend with a short but passionate defense of that most despised of management tools, the technology mandate. Nobody likes mandates I have read many a tweet or post from engineers exploding with rage over the pointless, counterproductive AI mandates they have endured at work. I have also

    management org · Charity Majors

  21. research · HBS AI Institute ·

    AI is Giving Workers More Focus Time. Now What?

    A six-month randomized trial across industries found generative AI increases worker focus time and flow states.

    A six-month, cross-industry randomized trial shows how generative AI is reshaping work. Listen to this article: The promise of psychologist Mihaly Csikszentmihalyi’s “flow theory” is that people do their best work when they can enter a flow state, or stay deeply engaged in a challenging task without constant interruption. Yet the structure of knowledge work […] The post AI is Giving Workers More Focus Time. Now What? appeared first on Harvard Business School AI Institute .

    productivity

  22. practice · Geoffrey Litt ·

    Understanding is the new bottleneck

    Code understanding techniques, explainer docs, quizzes, micro-worlds, help developers maintain agency when AI agents write code.

    July 2026 Understanding is the new bottleneck This is a written version of a talk I gave at the AI Engineer conference in July 2026, also shared as a tweet thread. Hot take: I think it's still important to understand the code that our agents write! In this talk I'll explain why that's the case, and show some ideas for how to efficiently understand code. Alright, let's dive in. Agents are writing more and more code for us, and we all know it's getting harder to keep up. But the good news is: there are many ways to understand code! Reading diffs line by line is not the only way. Most of this tal

    judgment · Geoffrey Litt

  23. research · arXiv ·

    You Shall Not Pass! Where and Why Developers Draw The Line on AI Autonomy

    Developers accept less AI autonomy for identity-defining and design work, more for high-demand tasks; task identity and accountability shape delegation choices.

    As AI takes on more software work, the line between human and AI effort is shifting. Where developers draw that line around AI autonomy bears on how we design tools and roles that preserve meaningful work. Drawing on cognitive appraisal theory, work design, and automation research, we conducted a mixed-methods study of 448 professional developers at Microsoft to investigate their accepted levels of AI autonomy across software engineering work. Most developers accepted AI producing work under their oversight, although accepted autonomy varied substantively across tasks and individuals. Acceptan

    worker experience

  24. research · arXiv ·

    A Simple Solution to Improving Human Supervision of Algorithms: Evidence from Smart Vending

    Field experiment shows constrained override policies improve AI supervision: limiting workers to two overrides per machine increased inventory reduction without harming sales.

    Organizations increasingly deploy autonomous artificial intelligence (AI) systems for operational decisions, such as inventory replenishment. Yet fully granting override rights can degrade performance due to human bias and noise, while prohibiting them may overlook valuable private information. This raises a key question: How should override rights be structured to improve human supervision of autonomous AI? Methodology/results: We propose a constrained override policy that limits overrides per decision episode to enable selective filtering that prioritizes high-value overrides. We tested it t

    adoption

  25. research · NBER ·

    AI Premium

    Analysis of 380 trillion tokens of real LLM usage across 400+ models shows how AI consumption relates to firm performance, market dynamics, and worker outcomes.

    Using 380 trillion tokens of realized AI consumption across more than four hundred large language models from the licensed proprietary OpenRouter dataset covering approximately 2 percent of current global monthly AI token consumption, we analyze how AI affects firms, markets, and workers. Leveraging (Nicola Borri , Aleh Tsyvinski , Yukun Liu)

    adoption

  26. practice · How I AI ·

    Sonnet 5 review: I ran 64 generations to find out if it's worth it

    Claire Vo built a repeatable eval harness to compare frontier models across PRD quality, prototypes, agentic tasks, and agent personality using hybrid human-LLM scoring.

    I’ve been testing every major frontier model release since the start of the year, and when Anthropic dropped Sonnet 5, I wanted more than a vibe check. I got tired of one-off tests I couldn’t repeat or compare over time, so I built something better: the How I AI Bench, a repeatable eval harness I constructed live using Claude Code while recording this episode. I ran Sonnet 5 blind against four other frontier models (Sonnet 4.6, Opus 4.8, GPT-5.5, and Gemini 3 Pro) across PRD quality, prototype generation, agentic task completion, and agent personality. The results were not what I expected. Wha

    ways of working · Claire Vo

  27. practice · One Useful Thing ·

    The twilight of the chatbots

    AI models now complete complex multi-week projects autonomously in hours, measured by multiple independent benchmarks showing exponential capability gains.

    If you feel like things are accelerating in AI, you are probably right. Better AI models from the leading American AI labs have been releasing more quickly than ever (though government interventions stopped access temporarily to two of the most powerful models, Claude Fable and GPT-5.6). But it isn't just release timing. The evidence points to accelerating capability gains as well (though the frontier stays jagged, and AIs remain weak in many places). This is especially obvious when we look at the ability of AIs to do real work. There are a few good assessments that try to measure how much hum

    productivity · Ethan Mollick

  28. research · arXiv ·

    The Organizational Behavior of Agentic AI: Collective Intelligence in Human-Agent Workflows

    Multi-agent AI systems exhibit organizational behavior patterns, differentiated work, coordination, routines, but sustained by context architecture rather than human motivation or trust.

    Agentic artificial intelligence is increasingly deployed not as a single assistant but as a collective of planners, solvers, reviewers, memory managers, tool users, and orchestrators. These systems are entering organisational workflows under familiar labels such as teams, managers, committees, markets, and workflows. This article asks whether such agent collectives exhibit organisational behaviour in a sense that is analytically comparable to, yet distinct from, human organisational behaviour. I argue that agentic AI is a partial organisational analogue. It resembles a human organisation becau

    adoption

  29. practice · How I AI ·

    No Figma. No Jira. No docs. How Gusto built a new product line with Claude Code | Eddie Kim (CTO)

    Five-person team built a production AI product in 10 weeks by eliminating planning docs, using async code review as product decisions, and pairing designers with Claude Code.

    Eddie Kim is the co-founder and CTO of the payroll and HR platform Gusto, which just crossed $1 billion in revenue and serves more than 500,000 small businesses. Recently he did something most CTOs don’t: he went back to writing code. With three other engineers and one designer, Eddie built Gusto Cofounder, a net-new AI product, from zero code to a tier-one launch in 10 weeks. He walks through how that team actually worked, why they threw out nearly every process, and how anyone can copy the approach. What you’ll learn: The trash-can method: how to write, review, and delete a full PR as a prod

    ways of working · Claire Vo

  30. practice · Hamel Husain ·

    “It’s Hard to Eval” Is a Product Smell

    Design AI products for verifiability first: showing work and intermediate steps makes evaluation easier and catches errors earlier.

    For the past 3 years, AI evals have been my professional focus. 1 The most common objection I hear to evals is “our product is hard to eval”. This objection is a product smell. Artifacts that are hard for you to verify are often hard for users too. In the worst case, users have to redo the work from scratch to verify the output. More importantly, designing your product for ease of verification should come before building evals. In this post, I’ll walk through three products I advised on that faced this issue. I’ll also show before and after sketches to demonstrate design principles. After thes

    judgment · Hamel Husain

  31. research · Information Systems Research ·

    Asymmetric Algorithm Aversion

    Experts accept negative AI recommendations but discount positive ones, perceiving AI as better at explicit data than tacit knowledge like leadership quality.

    Artificial intelligence (AI) is increasingly embedded in decision support systems, yet people do not treat all AI recommendations equally. Drawing on two randomized controlled experiments, a field study using data from a real-world Fintech investment platform, and interviews with industry experts, we identify a phenomenon we call asymmetric algorithm aversion: Although AI advice significantly influenced experts’ evaluations when it recommended against investing, it had no significant effect on experts’ evaluations when it recommended investing. This pattern arises because decision makers perce

    judgment

  32. practice · Kent Beck ·

    The Cost YAGNI Was Never About

    Clarifying YAGNI as timing guidance for AI agents: build structure when needed, not early or late.

    Here’s how I remember it—Chet Hendrickson came up to me in the middle of a project and said, “I could do this simplistic thing now but in 3 weeks that will be insufficient so since we’re going to need this more complicated thing I want to do it now.” I said, “You aren’t going to need it.” Chet said, “You don’t understand. We’re definitely going to need it. See, here’s an example…” Me (interrupting), “You aren’t going to need it.” Chet, get frustrated, “But we really are…” Me, “You aren’t going to need it.” Chet, eyes going up to the ceiling, pausing, “Oh.” Walks away. YAGNI is not an excuse to

    ways of working · Kent Beck

  33. research · arXiv ·

    The Shift to Agentic AI: Evidence from Codex

    Agentic AI usage grew fivefold in early 2026; 10% of users manage multiple concurrent agents; workflow changes are becoming common.

    We analyze usage data from OpenAI's Codex tool to present large-scale evidence of how agentic AI technology, which can take actions on a user's behalf, changes how people work. We use an automated, privacy-protecting pipeline to contrast usage across three populations: external personal-account users, external organizational-account users, and workers within OpenAI. We find that agentic AI usage is growing rapidly: the number of active users has grown more than fivefold in the first half of 2026, with the most rapid increase occurring outside the initial audience of software developers. Uptake

    adoption · Aaron Chatterji

  34. research · JAMA Network Open ·

    Medical Record Abstraction for Quality Improvement in Sepsis Care Using Artificial Intelligence

    LLM-generated real-time feedback on sepsis care compliance improved physician adherence to quality metrics in a hospital trial.

    Importance: Hospital quality reporting remains a manual, costly process with critical limitations as a mechanism to improve care outcomes. Objective: To assess whether near-real-time quality measurement, enabled by large language models (LLMs), can improve quality performance as measured by the Centers for Medicare & Medicaid Services (CMS) Severe Sepsis and Septic Shock Management Bundle (SEP-1) quality metric. Design, Setting, and Participants: This single-blind, unstratified, cluster randomized trial was conducted between December 13, 2024, and July 8, 2025, at 2 academic emergency departme

    judgment

  35. research · Epoch AI ·

    What we learned from 1,604 Chinese AI job postings

    Analysis of 1,604 Chinese AI job postings reveals hiring patterns and strategic priorities of AI labs.

    Inferring Chinese AI labs’ strategies from their job descriptions

    jobs skills

  36. practice · How I AI ·

    GLM 5.2: why I’m replacing Opus in Claude Code with this new model

    Developer replaced Claude Opus with GLM 5.2 for coding tasks, saving 94% on token costs while maintaining quality on real production work.

    I put GLM 5.2, the open-weight coding model from Z.AI, through four real tasks inside my actual codebase: a codebase architecture audit, a UI redesign, and a 45-minute autonomous bug-hunting session pulling from Sentry and Vercel logs. Total cost: $3.36 for roughly 6 million tokens, a prioritized bug-fix dashboard I’m actually shipping from, and a landing page redesign that matched Chat PRD’s design system on the first try. What you’ll learn: What “open-weight” actually means and why it matters for cost and vendor independence How to connect GLM 5.2 to Cursor and Claude Code How it performs on

    productivity · Claire Vo

  37. practice · Anil Dash ·

    How we’ll fight the platform war against Big AI

    Platform strategy tactics, used to compete against dominant tech companies, can still work against Big AI firms because the market is early and users are angry.

    One aspect of strategy that’s been largely lost in the tech industry in recent years is how to compete against platforms, since the major tech companies have gotten so big that markets are no longer competitive. However, the AI market is still early enough, and users and society are still angry enough, that the Big AI companies can lose. But for them to lose, everybody else in the ecosystem has to carry out the nearly-lost art of platform strategy . Tech companies (and even open source communities!) used to carry out these tactics in emerging product categories ranging from desktop office suit

    management org · Anil Dash

  38. practice · How I AI ·

    How Claude Mythos found a 15-year-old bug in Mozilla Firefox | Brian Grinstead

    Mozilla's agentic bug-finding pipeline with LLM judge and verifier subagent found 15-year-old Firefox bugs; harness design matters as much as the model.

    Brian Grinstead is a distinguished engineer at Mozilla, where he’s worked on Firefox and the web platform since 2013 (he joined to help launch Firefox DevTools). Recently he and his team pointed an agentic bug-finding pipeline at Firefox—a codebase with tens of thousands of files and tens of millions of lines of code—and shipped a record month of security fixes. The viral chart everyone saw gave the credit to Anthropic’s new Mythos model. Brian’s take is that the harness and pipeline did just as much of the work, and he walks through exactly how it runs and how anyone can build a starter versi

    ways of working · Claire Vo

  39. research · arXiv ·

    The Urban-Rural Divide in the Age of Artificial Intelligence: Assessing the Effects of Technology and Automation on Regional Labor Markets

    AI exposure raises wages in urban regions but automation exposure lowers employment more in rural areas, suggesting divergent regional impacts.

    Automation and artificial intelligence (AI) are reshaping labor demand unevenly across space, creating an urgent imperative for place-sensitive education and workforce policy. This study asks whether regional exposure to automation and to AI relates to local employment and wages in opposite ways, and whether those relationships differ between urban and rural regions -- two questions whose answers carry direct implications for how skills training and digital education should be targeted. Using a region-by-year panel and shift-share measures of technological exposure built from baseline industry

    jobs skills

  40. practice · Eugene Yan ·

    Patterns for Building Cybersecurity Evals

    A reusable pattern for building evaluation systems for cybersecurity AI agents with concrete components.

    A sandboxed target, inputs that influence task difficulty, tools, and a grader.

    ways of working · Eugene Yan

  41. research · arXiv ·

    Human Capital, AI, and Labor Commoditization

    AI exposure reduces employer demand for worker human capital signals on Upwork, shifting focus to price in more AI-exposed job categories.

    Has generative AI changed how labor markets value human capital? We study this question using contract-level data from Upwork, a large online labor market. We represent worker profiles with high-dimensional text embeddings, allowing us to capture rich human capital information from unstructured profile text. We then compute the predictive importance of workers' human capital information and posted hourly rates for client demand, and incorporate these measures into a difference-in-differences design around the release of ChatGPT. We find that in more AI-exposed job categories, the importance of

    jobs skills

  42. practice · GitHub Blog ·

    How we built an internal data analytics agent

    GitHub built an internal AI agent that lets any employee query a data warehouse in plain language, cutting analyst bottlenecks.

    Large data and analytics organizations often struggle to make access to data and insights truly self-serve. The industry tried to solve this problem, quite unsuccessfully, for decades, but now AI is giving us a credible way to do just that. At GitHub scale, providing dedicated analytics support to dozens of product teams is challenging, and therefore many teams are left to solve this problem on their own. Though there is a lot of valuable product telemetry that product and engineering teams can use to make decisions, figuring out which data model, which grain, which filter, and then write the

    productivity

  43. research · npj Digital Medicine ·

    A scoping review of human-AI collaboration patterns and task divisions in healthcare applications

    Scoping review of 85 papers shows human-AI collaboration patterns in healthcare vary by task risk and automation level, with evaluation shifting toward clinical efficiency and user experience.

    With the widespread use of Artificial Intelligence (AI) in healthcare, how human physicians work with AI, a new colleague, has become an issue of increasing concern. This review systematically explores the collaboration between humans and AI in the medical field, focusing on the prevalent human-AI collaboration patterns, the task division mechanisms, and the evaluation metrics. Following the PRISMA guidelines for scoping reviews, we screened the relevant literature and finally included 85 journal papers. Our analysis reveals that human-AI collaboration patterns in healthcare tend to be associa

    teams

  44. research · Information Systems Research ·

    Agency Configurations in Generative AI Ideation: How Textual and Visual Idea Concretizations Shape Idea Creativity and Ideator Effort

    Image-based generative AI reduces ideation effort by 18% but produces less creative ideas than text-based AI, because images constrain human imagination more than text does.

    People turn to generative artificial intelligence (AI) to make ideation less effortful. Turning a vague idea into a mature concept is hard work, and offloading it to AI is tempting. However, our research shows that this strategy can backfire; the lower-effort form of AI support often produced less creative ideas. In an online experiment, 276 people refined ideas for an innovation challenge working alone or with AI-generated text or images. Image-based generative AI had a double-edged effect; it significantly reduced effort compared with working alone or using text-based support. However, ideas

    productivity

  45. practice · Hacker News ·

    We built a persistent agent memory layer on Elasticsearch with 0.89 recall

    Team built a persistent memory layer for AI agents using Elasticsearch, achieving 0.89 recall on retrieval.

    116 points on Hacker News. Discussion: https://news.ycombinator.com/item?id=48583703

    ways of working

  46. practice · Claude Blog ·

    Steering Claude Code: when to use CLAUDE.md, skills, hooks, and subagents

    Framework for choosing between CLAUDE.md files, skills, hooks, and subagents to control Claude's code-writing behavior.

    Steering Claude Code: when to use CLAUDE.md, skills, hooks, and subagents

    ways of working

  47. research · Science ·

    AI may raise the bar and thin the pipeline

    AI adoption may reduce entry-level hiring and apprenticeship opportunities rather than displace existing workers, potentially thinning professional pipelines.

    In their Perspective “AI raises the productivity bar” (19 February, 10.1126/science.aef5239), L. Wu and B. Vasilescu (1) argue that artificial intelligence (AI) can increase productivity while rewarding workers who already possess strong evaluative and delegation skills. However, they do not address the most important implication of their model. If expertise in evaluation and delegation is now a stronger precondition for productivity than the ability to execute tasks oneself, firms and laboratories may respond not by shedding large numbers of incumbent workers but by quietly reducing entry-lev

    jobs skills

  48. research · Journal of Applied Psychology ·

    Scoring employment interviews with large language models: Evaluation design components, validity investigations, and best practice recommendations.

    LLM ensembles score job interviews with validity comparable to human raters and supervised ML, but measurement bias and group differences require caution.

    = 144). We then investigated the LLM scores' intrarater reliabilities, test-retest correlations, convergent, discriminant, and criterion evidence of validity, group differences, and measurement bias. We compared this evidence, when possible, to the same evidence for human raters and supervised machine learning models. The results suggest that ensembles of larger, newer LLMs using prompts with detailed construct information hold potential for scoring employment interviews with psychometric properties comparable to or superior to supervised machine learning models and single human raters. We det

    jobs skills

  49. practice · Paul Ford (Aboard) ·

    More New Words for New Work

    Sharp language for recognizing new organizational and product patterns emerging as AI tools reshape work.

    I gave a talk recently for the very excellent Rosenfeld Media “ Designing With AI ” conference. It was called “New Words for New Work,” and it was my collection of 30 purposefully ridiculous terms to define this ridiculous AI-infused moment. (I workshopped this list on the podcast as well as the newsletter a few months back, so some of these will be familiar if you’re really locked in to the Aboard Cinematic Universe). Here are my top ten: Praygency—A product company mixed with an agency. Everyone is sure this is the future of the industry, but no one knows how it’s supposed to work. That’s wh

    management org · Paul Ford

  50. practice · Eric Topol ·

    Agentic AI Comes to Medicine

    Two agentic AI systems (MIRA, AMIE) move beyond diagnostic support to autonomous end-to-end clinical decision-making and care planning.

    It was just a matter of time. Agentic autonomous AI has already been applied to life science and many other domains, and today there were 2 notable publications in Nature that move this concept forward for healthcare. One is called MIRA from Jacob Kather and colleagues from Germany, and t he other is called AMIE , from Mike Schaekermann and colleagues at Google (acronyms defined below). This work is getting well beyond AI support for narrow applications, such as help in making diagnoses, to full management, end-to-end care plans. They are both very complicated papers with a lot to unpack, incl

    ways of working · Eric Topol

  51. practice · Kent Beck ·

    A Learning System Made of Learning Parts

    Programming splits into commoditized coding and harder human work: understanding requirements, proving systems work, stewarding people-code-agent learning loops.

    Jessica Kerr joins Kent by the fire to argue that AI didn't take the programmer's job, it split it in two. The part we loved, crafting code by hand, has been commoditized like IKEA furniture. What's left is harder and more human: understanding what to build, proving it works, and stewarding the living "symmathesy" of people, code, and agents all learning from each other. They get into accelerated learning, why play is a signal you're learning, the loop that "becomes a noose," and choosing excitement over fear while the ground keeps shifting. You can find more of Jessica’s thoughts on her podca

    jobs skills · Kent Beck

  52. practice · How I AI ·

    How to design AI agent loops: schedules, goals, and subagents in Claude Code and Codex

    Four loop types (heartbeat, cron, hook, goal) with design patterns and production checklists, shown through two built examples.

    I break down every loop type from scratch—what a heartbeat, cron, hook, and goal loop actually are, when each one fits, and the five things any effective loop needs before it touches production. Then I build two live loops: a daily aging-PR reviewer in Claude Code that schedules itself at 10:15 a.m. and spins off its own subagents, and a weekly skills-identification loop in Codex that spawns goal-based subagents to validate its own output in real time. What you’ll learn: The plain-English definition of a loop—and why it’s just an automated prompt, not a scary new paradigm The four loop types (

    ways of working · Claire Vo

  53. practice · Sangeet Paul Choudary ·

    The skeptic's guide to consuming AI research with a pinch of salt

    Framework for spotting gaps between rigorous AI research findings and how they're misapplied in practice.

    AI research gets shared all over LinkedIn. No matter what you believe in, there’s a credible study backing it somewhere. You’re right to be skeptical about every new ‘AI expert’ shouting out from the rooftops on LinkedIn. But what about peer-reviewed AI research coming out from credible labs and top universities? As it turns out, research can be right in the narrow sense and still misleading in the larger sense. This, of course, applies to everyone writing about AI - including anything I put out there. This post is about making sense of AI research - looking for gaps in a theory and bridging t

    judgment · Sangeet Paul Choudary

  54. practice · Understanding AI (Timothy B. Lee) ·

    Why AI hasn’t replaced software engineers, and won’t

    Coding agents adopted rapidly in software engineering but haven't replaced engineers; adoption patterns show how AI complements rather than displaces skilled work.

    Coding agents as normal technology Coding agents as normal technology There is great anxiety and uncertainty about AI replacing jobs. How can we move past vague warnings and bombastic predictions and bring data to bear on this question? One good way is to look at the profession where AI capabilities are furthest along and adoption has been exceptionally rapid: software engineering.

    jobs skills · Timothy B. Lee

  55. practice · Addy Osmani ·

    Agentic Code Review

    Code review is now the leveraged skill; teams must shift from evaluating writing speed to evaluating whether to trust agent-generated code.

    Coding agents are extraordinarily good now and getting better fast. The interesting consequence is that the hard part of engineering moved from writing code to deciding whether to trust it, which makes review the most leveraged skill in software right now . How you approach it depends enormously on who you are: a solo developer with no users and a team maintaining a ten-year-old application are not solving the same problem. I am more optimistic about agentic engineering than I have ever been. The agents are genuinely good, they get better every month, and on an ordinary week I now ship things

    judgment · Addy Osmani

  56. research · arXiv ·

    AI Adoption Across a Multinational Workforce: Sociotechnical Conditions for GenAI Acceptance in Human Resources

    GenAI adoption in HR varied by role, language, and tenure; trust built through source-checking and peer consultation.

    Generative AI (GenAI) deployment in the workplace is accelerating rapidly. Nevertheless, questions of who adopts, who benefits, and who is left behind and why are still understudied. In this paper, we investigate these dynamics in the context of a multinational tech company transitioning from a legacy Human Resources (HR) search system to a GenAI-supported system, analyzing search log data, survey data (n=25), and ten semi-structured interviews. Our findings show that adoption depended on the fit between the GenAI system's design assumptions and employees' work positionalities (role, spoken la

    adoption

  57. practice · Shopify Engineering ·

    Teaching Sidekick to say no: automated data curation with LLM judge consensus

    Automated data curation pipeline using multiple LLM judges to teach an AI assistant when to refuse requests rather than hallucinate.

    Production training data only captures successful queries; it can't teach a model when to say no. We built an automated curation pipeline using LLM judge consensus to close that gap.

    ways of working

  58. practice · Will Larson ·

    Revised rules of engineering leadership.

    Engineering leadership approach revised around AI-accelerated migrations and individual capability where mistakes surface fast.

    From early 2014 through late 2020, I was working in hypergrowth environments, which are challenging, but also educational. The most valuable feature of hypergrowth is that your mistakes reveal themselves next month rather than next year, because things go wrong very loudly when you’re moving fast. I’ve been thinking a lot about hypergrowth recently, because Imprint’s business is growing quickly and we did a large batch of hiring last year, but also because the AI-tooling shift has changed the pace at which it’s possible to work. This post documents the new rules I’ve revised my approach to eng

    management org

  59. practice · Charity Majors ·

    AI demands more engineering discipline. Not less

    AI code quality now matches median engineers; engineering discipline and code review become more critical, not less.

    A few days back I wrote a piece called “ AI enthusiasts are in a race against time, AI skeptics are in a race against entropy .” I have notes on a whole pile of AI-related topics that I’d like to cover in depth: AI mandates, communication norms, code review, AI art, and more. Unfortunately, I got too many interesting responses to my last piece, and now I have to address those before I can move on to other topics. 😉 There were two types of interesting responses: the first on the technical merits, the second on ethical grounds. I will respond to each of these separately. Let’s take the technica

    management org · Charity Majors

  60. research · Léonard Boussioux ·

    “The Narrative AI Advantage?” accepted at Management Science

    Field experiment finds GenAI assistance changes how evaluators assess early-stage innovations, with measurable differences in judgment outcomes.

    Field experiment on GenAI-augmented evaluation of early-stage innovations

    judgment · Léonard Boussioux

  61. research · Proceedings of the National Academy of Sciences ·

    AI agents are sensitive to nudges

    LLM agents are far more sensitive to choice architecture nudges than humans, making them behaviorally brittle in deployment.

    Large language models (LLMs) are increasingly deployed as autonomous agents that make choices and use tools on behalf of users. Yet, we have limited evidence about how their decisions are shaped by their environment. We adapt a human decision-making task to test leading LLMs under four forms of choice architecture: defaults, suggestions, information highlighting, and "optimal" nudges derived from a resource-rational model of human choice. We treat human behavior as a baseline for predictable sensitivity to such interventions. Across models and prompting strategies, LLMs often depart substantia

    judgment

  62. research · arXiv ·

    Chaining Tasks, Redefining Work: A Theory of AI Automation

    AI automation creates contiguous task chains; firms bundle steps into jobs differently as AI quality improves, with non-linear productivity effects.

    Production is a sequence of steps that can be executed (1) manually, (2) augmented with AI, or (3) fully automated within contiguous AI-executed steps called ''chains.'' Firms optimally bundle steps into tasks and then jobs, trading off specialization gains against coordination costs. We characterize the optimal assignment of humans and AI to steps and the firm's resulting job structure, showing that comparative advantage logic can fail with AI chaining. The model implies non-linear productivity gains from AI quality improvements and admits a CES representation at the macro level. Empirical ev

    management org · Mert Demirer

  63. practice · John Warner ·

    Feedback: Human and Not

    Law professors' feedback on contract concepts preferred over human professors' in blind study; author argues LLM feedback still unsuitable for student writing.

    This is my earnest plea to subscribe followed by my earnest plead to give a hoot about keeping reading and writing human endeavors. I’ve been thinking about feedback. There is a study out of Stanford Law School that says in a blind evaluation, law professors preferred the feedback on student questions about contract law generated by a large language model over those written by other law professors. The purpose of this feedback is to enhance students’ understanding of legal concepts, based on questions students were asking about those concepts. The study’s authors suggest that this shows that L

    jobs skills · John Warner

  64. practice · Linear ·

    Teaching an agent to auto-fix bugs

    Team uses Linear Agent to auto-fix incoming bugs by iteratively building context and guardrails based on what works.

    I was very skeptical of AI to begin with. Part of me suspected the technology wasn't as capable as the hype would suggest, while another part was deeply worried that it was far more capable than I wanted to admit. But the closer I got to the technology the more those worries started to fade. The tech is very capable, but when you work with it you start to realise the extent of that capability. Where it’s strong, where it’s weak, and how the engineering role is evolving to fit the spaces between them. I've spent the majority of my time over the past few months building out automations in Linear

    ways of working

  65. research · Organization Science ·

    The Cybernetic Teammate: A Field Experiment on Generative AI and Teamwork

    Field experiment shows AI teammate matched solo-person performance to team performance, bridged functional silos, and enhanced idea generation quality in innovation work.

    We examine how artificial intelligence (AI) impacts three core pillars of collaboration—performance enhancement, expertise integration, and social engagement—through a preregistered field experiment with 791 professionals at Procter & Gamble, a global consumer packaged goods company. Working on real product innovation challenges, professionals were randomly assigned to work either with or without AI, and either individually or with another professional in new product development teams. Our findings show that (1) AI significantly enhances performance: individuals with AI matched the performance

    teams · Hila Lifshitz-Assaf · Karim Lakhani · Ethan Mollick · Fabrizio Dell'Acqua · Raffaella Sadun · Charles Ayoubi · Lilach Mollick

  66. practice · Linear ·

    Now Linear writes the code, too

    AI agent executes code changes within a shared team context system, reducing context-switching between planning and implementation tools.

    Agents are becoming capable of doing more of the work involved in building software, but they’re still mostly individual productivity tools, useful to one person at a time. Building a product is something a team does together, and the work that goes into it, like the decisions and the reasoning behind them, forms the context the whole product is built on. That context rarely lives in one place. It accumulates across Linear issues, documents, Slack threads, and code, which is why we’ve worked to bring more of product development into Linear and pull it all together. Linear Agent can draw on all

    ways of working

  67. research · arXiv ·

    The Privilege of Exposure: Caste and Generative AI in India's Graduate Labour Market

    Scheduled Caste and Tribe graduates in India are 0.24-0.37 standard deviations less exposed to AI work; AI exposure commands a 20% wage premium, widening caste inequality.

    Who is exposed to generative AI in a developing-country labour market? We map three occupational AI-exposure indices to India's redesigned Periodic Labour Force Survey (2025) and document a steep caste gradient among 83,000 employed graduates: graduates from the Scheduled Castes and the Scheduled Tribes are 0.24--0.37 standard deviations less exposed than upper-caste graduates within the same district. Two channels drive the gap: one in four SC and one in three ST graduates work in farm or elementary occupations untouched by AI, and those in white-collar work are underrepresented in managerial

    jobs skills

  68. research · Information Systems Research ·

    Strategic Release and Co-creation: Empirical Insights for Managing Open-Source Software

    AI tools reduce contribution cost more than review cost, reshaping trade-offs between release frequency and community co-creation in open-source projects.

    Open source projects involve two strategic decisions: when to release new versions, and how actively to engage community contributors. Both compete for the same scarce team resources. This study examines how core teams jointly manage these decisions over time using detailed data from GitHub. We find that both release and co-creation positively influence community interest, but the relative emphasis placed on them depends on a project’s operational conditions and cost structure, with team capacity playing a particularly important role in enabling co-creation. Projects that push either activity

    adoption

  69. research · arXiv ·

    Revisiting the ABCs of Working with AI: A Replication with Radiologists

    Radiologists with lower baseline ability and better calibrated confidence gain larger productivity gains from AI assistance, replicating prior findings.

    Artificial intelligence (AI) systems increasingly assist human experts, but the consequences of AI assistance on productivity can be heterogeneous. Caplin, Deming, S. Li, Martin, Marx, Weidmann, and Ye (2025b) provide evidence that two characteristics, ability and belief calibration, help to determine the returns to AI assistance. This note shows that their results replicate to a setting where professional radiologists analyze chest X-rays with access to state-of-the-art machine learning predictions. I leverage the public Collab-CXR data repository described by Moehring, Kutwal, Huang, Banerje

    productivity

  70. practice · Linear ·

    Reviewing code in the agent era

    Engineering team kept code review quality high while PR volume jumped 50% using a new review workflow tool.

    Like a lot of people who build software, I felt the ground shift at the start of 2026. The models got better, the agents got better with them, and almost overnight the amount of code moving through my team went up. Between January and March our total issues created and PRs merged both climbed by 50%. That was exciting, and it was also a problem. With that much code moving, I felt the pull to “looks good!” my way through reviews just to keep the queue down, and that sat badly with me, because quality is something we genuinely obsess over at Linear. I kept landing on the same question, which was

    judgment

  71. research · Knowledge at Wharton ·

    Automation Doesn’t Just Cut Jobs. It Slows Career Progression

    Automation slows career progression into better-paid roles, not just displacing workers outright.

    Automation is often seen as destroying jobs, but new Wharton research shows it also can quietly block workers from moving into better-paid roles. … Read More

    jobs skills

  72. research · arXiv ·

    The Jagged Global Economy: Frontier AI Unevenly Exposes National Economies

    High-income countries face 50% more AI labor-market exposure than low-income countries; women face higher exposure in 91% of countries due to occupational concentration.

    Frontier AI's labor-market effects matter to workers, firms, and policymakers, but current evidence generally comes from a handful of high-income economies. The capabilities of frontier AI are jagged across work tasks and national economies diverge in how they allocate human labor. We introduce a national AI exposure metric that combines occupation-level exposure scores and international employment data for 141 countries. We find that high income countries are substantially more exposed than low income countries and that Europe and Central Asia are 50 percent more exposed than Sub-Saharan Afri

    jobs skills

  73. practice · Addy Osmani ·

    Loop Engineering

    Loop engineering: designing systems that recursively prompt agents toward goals instead of prompting manually turn-by-turn.

    Loop engineering is replacing yourself as the person who prompts the agent. You design the system that does it instead. A loop here can be thought of a recursive goal where you define a purpose and the AI iterates until complete. It’s roughly five building blocks and Claude Code and Codex both have all five now. I believe this may be the future of how we work with coding agents. However, its still early, I’m skeptical and you absolutely have to be careful about token costs (usage patterns can vary wildly if you are token rich or poor), so I want to unpack what it is and what it means. Peter St

    ways of working · Addy Osmani

  74. practice · Sangeet Paul Choudary ·

    The Jevons Misunderstanding

    Jevons Paradox misses the real question: who captures surplus when AI expands demand, not whether work expands.

    If I had a token for every time someone said Jevons Paradox, I’d have enough compute to explain why Jevons Paradox doesn’t explain what they think it explains. There are many convenient explanations people reach for when they want to make sense of AI. One of them is Jevons Paradox. As the AI doomers began warning that AI would take away all jobs, the other side reached for Jevons: Efficiency expands demand; Cheaper work creates more work; Therefore, workers will be fine. The problem with this binary debate is that it leaves no room for the actual economics. You are forced into one of two camps

    management org · Sangeet Paul Choudary

  75. practice · Latent Space ·

    How to Stop Shipping Low-Quality RL Environments (with Examples)

    Common mistakes in reinforcement learning environment setup that degrade model performance, with fixes.

    Your broken harness is actively making the model worse. Here's what I keep seeing after years of eyeballing trajectories, and what you need to fix. Your broken harness is actively making the model worse. Here's what I keep seeing after years of eyeballing trajectories, and what you need to fix. We’re so excited to publish this guest post from Auriel W, who has worked on RL at Gemini, and has an incredible “RL Pet Peeves” blog where she not-so-subtly explains the frustrations big labs have w…

    ways of working

  76. research · Management Science ·

    Markovian Search with Ex Ante Constraints: Theory and Applications to Socially Aware Algorithmic Hiring

    Algorithm for incorporating fairness constraints into sequential hiring search maintains index-based structure with dual adjustments and randomization rules.

    We study and develop an algorithmic framework for incorporating “ex ante” constraints—constraints on outcomes that hold only on average—into stateful sequential search problems with costly inspection. Our framework encompasses the classical Weitzman’s Pandora’s box and its extensions to joint Markovian scheduling, which model richer processes such as multistage search with multiple layers of inspection. Ex ante constraints are particularly motivated by social considerations in algorithmic hiring, where they can adjust outcome distributions to promote equity and access. Although most work in th

    jobs skills

  77. practice · One Useful Thing ·

    Co-Existence and the End of Co-Intelligence

    AI coding agents now write 80% of code at major companies, shifting work from co-intelligence to autonomous systems.

    It has been two years since Co-Intelligence , my book about AI, was published, and it was successful beyond what I could have hoped (it was a New York Times bestseller and has been translated into 25+ languages, with the biggest markets being the Netherlands and Korea). I don’t think the book is out-of-date, exactly, but it was written about a world of chatbots and earlier AI models. In that world, working with an AI was a cooperative exercise, involving prompting a chatbot back-and-forth, adding your own knowledge and skepticism as you went. Humans were at the center, chatbots were your helpe

    management org · Ethan Mollick

  78. practice · Latent Space ·

    Reality: The Final Eval — Lukas Petersson and Axel Backlund of Andon Labs

    How to build frontier evals for AI models across capability ranges, from practical design to lasting benchmarks.

    We talk with the VendingBench authors on evaling Claudes from Haiku to Mythos, and how they build leading, and lasting, frontier evals from scratch. We talk with the VendingBench authors on evaling Claudes from Haiku to Mythos, and how they build leading, and lasting, frontier evals from scratch. The new AIEWF website is live! Get your tickets booked ASAP as they -will- sell out. Take the AI Engineering Survey and get >$2k in credits and free AIE WF tickets!

    judgment · Shawn Wang (swyx)

  79. practice · Kent Beck ·

    You Don't Get to Create Anything

    Experienced distributed systems engineers explain why they're not worried about AI replacing expertise, grounded in how demand grows when cognition becomes cheap.

    Randy Shoup set out to be an international lawyer. He studied in West Berlin when there was still a wall around it, spent a year at Stanford Law, and had what should have been the perfect summer internship on Sand Hill Road. Instead he spent it watching inventors light up whiteboards with brilliant ideas — then being told his job was just to write them down. That summer broke something open. He went back to Oracle that fall and never looked back. Kent and Randy dig into what it means to need to make things, why the people who wrote the original distributed systems playbook aren’t panicking abo

    judgment · Kent Beck

  80. practice · Charity Majors ·

    AI enthusiasts are in a race against time, AI skeptics are in a race against entropy

    Selective storytelling about AI wins obscures hidden technical debt and cleanup costs that teams actually bear.

    I recently attended a talk where one of the presenters made some pretty… astonishing claims about what they had achieved by the pure, uncut power of vibe coding. Difficult engineering problems solved, backlogs cleared. Rewrites that would have taken a year or more in the beforetimes, now whipped out in a few short weeks of prompting. Afterwards, wandering around the conference, I caught a lot of excited chatter: “I can’t wait to make my teams watch the recording of this talk. My engineers are SO resistant to the idea of shipping code without reading it. Finally, some proof they can’t ignore!”

    judgment · Charity Majors

  81. research · MIT Sloan Management Review ·

    Why AI Isn’t Transforming Finance Yet

    Multiyear action research on how finance leadership work evolves as AI is introduced, including barriers to transformation.

    Christian Gralingen The Research The authors engaged in two complementary research streams. One was a multiyear program of action design research conducted with organizations undergoing digital transformation that focused on how leadership work evolves under conditions of technological and market uncertainty. The other, a study of how AI is introduced into finance functions and how […]

    adoption

  82. research · MIT Sloan Management Review ·

    Scaling AI With Adaptive Governance

    Interviews with leaders at seven large organisations reveal governance practices that enable AI scaling without sacrificing control.

    Christian Gralingen The Research From 2022 to 2025, the authors conducted in-depth, semistructured interviews with senior leaders and practitioners responsible for AI governance, risk, compliance, data, and product decisions. Core interviews were conducted at Microsoft, Barclays, Kyriba, Nasdaq, Lloyds Bank, Danske Bank, and the Abu Dhabi Department of Finance. The interviews focused on how governance […]

    management org

  83. research · MIT Sloan Management Review ·

    Create Generative AI Value at Scale

    Three-year study of 23 Swiss companies identifies barriers to scaling generative AI value beyond pilots.

    Christian Gralingen The Research Over three years (2022-2025), two of the authors (Kevin and Ivo) engaged with 23 Swiss companies that were members of a research consortium focused on generative AI. The study participants represented a diverse array of industries: retail banking, investment banking, health insurance, insurance, medical coding, energy, law, laboratory instrument manufacturing, equipment […]

    adoption

  84. practice · Kent Beck ·

    Trust Factory

    Trust accumulates slowly but vanishes instantly; XP practices manufactured trust through testing, pairing, and integration.

    “We’re accumulating code faster than we are accumulating trust.” Sometimes a phrase just hits. Yes, we can create code faster now, but software is bipedal—code & trust go together. One without the other just hops along awkwardly. Trust is as tricky as code. Both are asymmetrical. Code works or it doesn’t. One mistake in a long string of good decisions is the same as just a mistake. Trust accumulates slowly & evaporates in an instant. The difference is that in software sometimes you can repair the mistake in time proportional to the time it took to make the mistake. Trust is irreversible. Once

    management org · Kent Beck

  85. practice · How I AI ·

    Building an iPhone app with zero technical skills | Bryce Rattner Keithley

    Non-technical recruiter built and shipped an iOS fitness app using Claude, Gemini and Replit, proving execution is no longer the constraint.

    Bryce Rattner Keithley has spent her career in talent and recruiting, working with technical leaders but never writing a line of code herself. Yet she managed to build Daily Hundred—a fitness app featuring custom AI-generated videos of anthropomorphic animals demonstrating exercises—and ship it to the App Store before her software engineer friends. Using Replit, Claude, Gemini, and a relentless beginner’s mindset, Bryce proves that in the AI era, execution is no longer the constraint on good ideas. What you’ll learn: How to build and ship an iPhone app using Replit without any coding knowledge

    ways of working · Claire Vo

  86. research · St. Louis Fed ·

    Measuring AI Adoption among Firms: How You Ask Matters

    Worker and firm surveys give opposite pictures of AI adoption rates in US vs Europe; measurement method matters.

    Worker-based surveys find higher rates of AI adoption in the U.S. than in Europe. Yet firm-level surveys reveal the opposite. What is behind this difference?

    adoption

  87. research · NBER ·

    Writing Code vs. Shipping Code: Productivity Effects Across Generations of AI Coding Tools

    AI coding tools boost task-level productivity, but gains shrink across generations and don't fully translate to final output shipped.

    How do the productivity effects of AI evolve across successive generations of tools, and to what extent do task-level gains ultimately translate into final output? We study these questions in the context of software development, using data on more than 500,000 GitHub developers combined with their (Mert Demirer , Leon Musolff , Liyuan Yang)

    productivity · Mert Demirer

  88. research · NBER ·

    Beyond Exposure: Predicting AI Adoption Based on Comparative Advantage

    AI exposure measures differ from actual adoption; a new index based on comparative advantage predicts workplace adoption better using German survey data.

    We document and explain the gap between measures of AI exposure and measures of AI adoption in the workplace. This leads us to propose a new AI adoption index based on comparative advantage. Using the representative German DiWaBe employee survey linked to worker and establishment information, we (Ilse Lindenlaub , Ryungha Oh , Maria Alejandra Rodriguez , Laura Veldkamp)

    adoption

  89. practice · Sangeet Paul Choudary ·

    How to reimagine your work in the age of AI

    Reimagining work roles and structures when AI removes production constraints, moving from output-focused careers to new value types.

    I’ve been quiet here for quite some time. I started noodling with a question towards the end of last year - what should work look like now that producing work is no longer a constraint? More broadly, what should AI-native work look like? And how can all of us - who’ve built careers around producing hard-to-produce outputs reimagine our work in an age when producing outputs isn’t all that hard anymore? But before we get into all that, I’d like to start by launching something I’ve been working on over the past few weeks. Launching: The living companion to Reshuffle Many of you have enjoyed readi

    management org · Sangeet Paul Choudary

  90. research · JAMA Network Open ·

    Ambient Artificial Intelligence Use and Clinician Documentation Burden, Productivity, and Efficiency

    Ambient AI scribe system reduces clinician documentation burden and improves productivity, measured through qualitative evaluation.

    This qualitative study evaluates the associations between use of an ambient artificial intelligence (AI) scribe system and clinician productivity, efficiency, and documentation burden.

    productivity

  91. practice · Latent Space ·

    The Age of Async Agents — Cognition's Walden Yan & OpenInspect's Cole Murray

    Teams using async agents for spec-to-PR workflows, with product managers shipping code and agents maintaining persistent memory across sessions.

    80% Devin Commits, Spec-to-PR Workflows, Full VMs, Agent Memory, and PMs Shipping Code 80% Devin Commits, Spec-to-PR Workflows, Full VMs, Agent Memory, and PMs Shipping Code The new AIEWF website is live! CFPs close in 2 days and we will run our first New Engineer Orientation this weekend, get your tickets booked ASAP as they -will- sell out. Take the AI Engineering Surv…

    ways of working · Shawn Wang (swyx)

  92. practice · Shopify Engineering ·

    Under the River

    Shopify built a Slack-native agent and shares the technical substrate and lessons from shipping it.

    What it took to ship our Slack-native agent River, lessons learned, and the substrate that runs beneath it. Co-authored by River.

    ways of working

  93. research · Proceedings of the National Academy of Sciences ·

    AI-assisted teams outperform AI-led teams but not human-only teams in assessing research reproducibility in quantitative social science

    Researchers working with AI assistants matched human-only teams on research reproducibility tasks, but AI-led teams performed much worse.

    Large Language Models (LLMs) such as ChatGPT are transforming how scientists conduct and validate research, offering promise as tools to improve scientific reproducibility. However, computational reproducibility and error detection remain expensive and labor-intensive. We experimentally test how collaboration between researchers and LLM assistants influences the reproduction of quantitative social science findings across different levels of AI autonomy. We randomly assigned 288 researchers to 103 teams working under three conditions: human-only, AI-assisted (using ChatGPT as a collaborative to

    judgment

  94. practice · Paul Ford (Aboard) ·

    The AI Hangover Will Be Delightful

    AI coding tools drive most growth; other applications plateau while engineering hiring rises despite AI power increasing.

    It’s a paradoxical moment in the great global AI rollout. At one level, you have three of the largest IPOs in human history headed our way: SpaceX , which rolls up Grok, as well as OpenAI and Anthropic . It’s a deca-trillion-dollar ouroboros economy predicated on the cost of Nvidia chips. At the same time, certain narratives are fizzling out. “ Tokenmaxxing ”—encouraging engineers to burn as much AI time as possible—seems to be running its course . OpenAI’s economics are shaky pre-IPO, at least given its proposed valuation. AI-assisted coding, like Codex or Claude Code, seem to be driving the

    jobs skills · Paul Ford

  95. practice · How I AI ·

    The Codex feature that works while you sleep

    Using autonomous AI agents with measurable goals to run multi-hour tasks unsupervised, with examples from error cleanup, email triage and task organization.

    In this 30-minute episode, I walk through my favorite feature in Codex: the /goal command. I show how Goals transform AI from a turn-based assistant that needs constant ‘what’s next?’ prompting into an autonomous agent that can work for hours on complex, multi-step tasks. I share three real examples: eliminating thousands of Sentry errors, cleaning 3,900 emails down to 68, and organizing hundreds of Linear tasks. What you’ll learn: What Goals are and how they differ from standard prompts How I used /goal to eliminate hundreds of error logs in my codebase over a five-hour autonomous run The non

    ways of working · Claire Vo

  96. practice · Hacker News ·

    Claude Code as a Daily Driver: Claude.md, Skills, Subagents, Plugins, and MCPs

    Developer workflow using Claude with markdown interface, custom skills, subagents, plugins and MCPs as primary coding tool.

    451 points on Hacker News. Discussion: https://news.ycombinator.com/item?id=48289950

    ways of working

  97. research · Information Systems Research ·

    Debiasing ML- or AI-Generated Regressors in Partially Linear Models

    Method to correct bias when AI-generated variables are used in regression models for business decisions at scale.

    Organizations increasingly use machine learning (ML) and artificial intelligence (AI), including large language models, to generate variables for regression models that inform business and policy decisions. For example, practitioners may use AI to predict review sentiment, ad aesthetics, or emotional expressions, and then estimate their causal effects on outcomes such as sales or engagement. However, because AI predictions are imperfect, directly using these AI-generated variables as regressors introduces measurement error that can systematically bias causal estimates, potentially leading to o

    judgment

  98. practice · Eugene Yan ·

    Using LLMs to Secure Source Code

    A five-step workflow for using LLMs to find and fix code vulnerabilities: model, discover, verify, triage, patch.

    Build a threat model, discover vulnerabilities, verify, triage, and patch.

    ways of working · Eugene Yan

  99. practice · One Useful Thing ·

    Choosing to Stay Human

    Overuse of AI writing degrades reader attention and writer skill development; staying human in writing matters.

    If you go to your favorite social media site, you will find it full of posts that start to look suspiciously similar to each other: Many of the comments to these posts are also generated by AI. So are an increasing number of academic papers and New York Times opinion articles , and, apparently, award-winning short stories . If you use AI a lot, you probably have noticed how much AI writing is around you (frequent AI users have historically done quite well identifying AI writing ), if not, I promise you it is much more than you think. It isn’t just the sameness of the AI writing, though that ev

    worker experience · Ethan Mollick

  100. practice · Kent Beck ·

    Genie Lessons from Genie Sessions: Prose as a Programming Language

    Structured English prose with AI agents replaces traditional code; requires/ensures blocks wire components like dependency injection frameworks.

    This session is sponsored by OpenProse . More on them below — they're also the reason we're here. Dan Barrett built OpenProse — a framework that lets you write programs in structured English and run them with an AI agent like Claude Code. The description sounds either obvious or impossible depending on your priors. When I told him it sounded impossible, he said: “People were shocked that it works.” So we installed it and built something live. The project: a service that would show me tides, weather, sunrise, moonrise — everything I want to know before a walk along the coast. One command. A few

    ways of working · Kent Beck