The Feed · Complete archive

AI at work: research and practice

A server-rendered record of the evidence, ideas and firsthand practices screened by The Feed. Every entry links to its original source.

Use the interactive Feed

Page 5 of 9 · 873 items

  1. practice · How I AI ·

    From Figma to Claude Code and back | Gui Seiz & Alex Kern (Figma)

    Design and engineering teams use Claude Code and Figma MCP to collaborate bidirectionally, pulling live code into design files and pushing changes back without manual CSS work.

    Most teams are still passing static design files back and forth, and most Figma files are already out of date by the time they reach engineering. Gui Seiz (designer) and Alex Kern (engineer) from Figma walk through the exact workflow their team uses to bridge that gap with AI, live onscreen. They demo how to pull a running web app directly into Figma using the Figma MCP, edit it collaboratively, and push it back to code. The old linear waterfall workflow is gone. What replaces it is a fluid, bidirectional loop where design and code inform each other in real time. What you’ll learn: How to use

    teams · Claire Vo

  2. practice · Harper Reed ·

    My now immaculate knowledge graph of life

    Built a personal knowledge graph from 600 meeting transcripts using Claude Code to extract people and concepts, visualised in Obsidian.

    Everyone and everything I know! Botwick inception My AI friend Botwick (not to be confused with the person John Borthwick) built a really neat website that shows all sorts of various networks that Botwick (and thus Borthwick) share. It is built by immaculately coordinating a collection of notes that have been collected over decades and decades. Seeing it made me really jealous, and I wanted my own! Harpwick was no help. I was on my own. extraction I booted up my obsidian vault and was sad. My last note was from 2022, and it was a daily note with only one word in it: hungry. Turns out I didn’t

    ways of working · Harper Reed

  3. practice · Charity Majors ·

    Your Data Is Made Powerful By Context (so stop destroying it already)

    Observability data separated into silos loses combinatorial power; unified context makes validation of each ship exponentially more powerful.

    After twenty years of devops , most software engineers still treat observability like a fire alarm — something you check when things are already on fire. Not a feedback loop you use to validate every change after shipping. Not the essential, irreplaceable source of truth on product quality and user experience. This is not primarily a culture problem, or even a tooling problem. It’s a data problem. The dominant model for telemetry collection stores each type of signal in a different “pillar”, which rips the fabric of relationships apart — irreparably. Your observability data is self-destructing

    productivity · Charity Majors

  4. practice · How I AI ·

    Mastering Midjourney: How to create consistent, beautiful brand imagery without complex prompts | Jamey Gannon

    Creative director shows how to generate consistent brand imagery using visual references and style codes instead of complex prompts.

    Jamey Gannon is an AI creative director who specializes in creating consistent, beautiful brand imagery using AI tools. In this episode, Jamey demonstrates her streamlined workflow for generating cohesive brand assets using Midjourney, Nano Banana, and other AI image tools. She walks through her process of creating mood boards, using style references, developing personalization codes, and strategically iterating to achieve a consistent aesthetic. Rather than relying on complex prompts, Jamey shows how visual references and strategic shortcuts can produce better results with less effort. What y

    ways of working · Claire Vo

  5. practice · Hacker News ·

    Terence Tao: Formalizing a proof in Lean using Claude Code [video]

    Mathematician used Claude Code to formalize a mathematical proof in Lean, showing AI assistance in formal verification work.

    56 points on Hacker News. Discussion: https://news.ycombinator.com/item?id=47306852

    ways of working

  6. practice · Laurie Voss ·

    Do AI-enabled companies need fewer people?

    Startups are smaller than before AI, but founders are raising more capital per person, whether that means AI made them efficient or they just hired less is unclear.

    About a year ago I made some predictions about the effect of AI on programming jobs . Block laid off 40% of its staff claiming AI made them more efficient . Is that really true or did they just over-hire? Let's look at some data and see what's really happening. In February 2026, global venture capital hit a single-month record: $189 billion flowed into startups in 28 days. Three companies got 83% of that funding : OpenAI, Anthropic, and Waymo. Two months into 2026, startups have already raised more than half of what they raised in all of 2025 . There is an unprecedented amount of money going i

    jobs skills · Laurie Voss

  7. research · Public Administration Review ·

    Biased by Design? Case Managers' Multidimensional Preferences Toward the Design of Algorithmic Decision Support Systems

    Case managers balancing professional, service, and efficiency values reject mandatory algorithmic advice while accepting support systems in employment services.

    ABSTRACT This study examines whether street‐level bureaucrats' preferences toward algorithmic decision support (ADS) induce a unilateral shift of technology‐related risks onto clients of the public employment service. Expanding on public value theory and research on moral agency in public service work, we argue that case managers' choices of ADS designs are shaped by a plurality of professional, service, and efficiency values. To test this argument, we conducted a conjoint experiment on a representative sample of German Federal Employment Agency case managers. Respondents compared pairs of hyp

    judgment

  8. research · ACM Transactions on Software Engineering and Methodology ·

    On the Use of Agentic Coding: An Empirical Study of Pull Requests on GitHub

    83.8% of AI-generated pull requests are accepted by developers; 54.9% need no revision, 45.1% need human changes.

    Large language models (LLMs) are increasingly being integrated into software development processes. The ability to generate code and submit pull requests with minimal human intervention, through the use of autonomous AI agents, is poised to become a standard practice. However, little is known about the practical usefulness of these pull requests and the extent to which their contributions are accepted in real-world projects. In this article, we empirically study 567 GitHub pull requests (PRs) generated using Claude Code, an agentic coding tool, across 157 diverse open source projects. Our anal

    adoption

  9. research · ACM Transactions on Software Engineering and Methodology ·

    Novice Developers’ Perspectives on Adopting LLMs for Software Development: A Systematic Literature Review

    Systematic review of 80 studies on how novice developers adopt and experience LLM tools for coding tasks.

    Following the rise of large language models (LLMs), many studies have emerged in recent years focusing on exploring the adoption of LLM-based tools for software development by novice developers: computer science/software engineering students and early-career industry developers with two years or less of professional experience. These studies have sought to understand the perspectives of novice developers on using these tools, a critical aspect of the successful adoption of LLMs in software engineering. To systematically collect and summarise these studies, we conducted a systematic literature

    adoption

  10. practice · Claude Blog ·

    Common workflow patterns for AI agents—and when to use them

    Agent workflow patterns with guidance on when to apply each: sequential, parallel, branching, looping, hierarchical.

    Common workflow patterns for AI agents—and when to use them

    ways of working

  11. research · MIS Quarterly ·

    Extending the Digital Divide: The Role of Unequal Analytical Abilities1

    Users with higher analytical ability exploit more profitable trading opportunities than others, even with equal data access, suggesting analytical skill is a new source of outcome inequality.

    The classic digital divide theory asserts that unequal access to and unequal experience with information technologies may lead to unequal user outcomes. This paper introduces a new perspective to extend this theory: outcome divides can persist despite equal access and equal experience if users differ in their analytical ability to analyze and interpret available data for decision-making. We term this new data-to-decision skill as analytical ability and integrate it into the classic digital divide framework. We develop a new approach to operationalize analytical ability by contrasting humans’ a

    jobs skills

  12. practice · Charity Majors ·

    My (hypothetical) SRECon26 keynote

    SRE author reconsiders 2025 advice to upskill on AI; reflects on what changed in one year of production use.

    Hey, it’s almost time for SRECon 2026 ! (I can’t go, but YOU really should!) Which means it was almost a year ago that Fred Hebert and I were up on stage, delivering the closing keynote 1 at SRECon25. We argued that SREs should get involved and skill up on generative AI tools and techniques, instead of being naysayers and peanut gallerians. You can get a feel for the overall vibe from the description: It’s easy to be cynical when there’s this much hype and easy money flying around, but generative AI is not a fad; it’s here to stay. Which means that even operators and cynics — no, especially op

    management org · Charity Majors

  13. practice · Hacker News ·

    Parallel coding agents with tmux and Markdown specs

    Running multiple coding agents in parallel via tmux, coordinated through Markdown specification files.

    189 points on Hacker News. Discussion: https://news.ycombinator.com/item?id=47218318

    ways of working

  14. practice · How I AI ·

    How Coinbase scaled AI to 1,000+ engineers | Chintan Turakhia

    Engineering org reduced PR review from 150 to 15 hours and compressed user feedback to shipped features by embedding AI agents into existing workflows at scale.

    Chintan Turakhia is Senior Director of Engineering at Coinbase, where he’s led the transformation of a 1,000-plus-engineer organization to embrace AI tools at scale. When tasked with rewriting Coinbase’s self-custody wallet into a consumer social app in just six to nine months, Chintan turned to AI as a force multiplier. His team has achieved remarkable efficiency gains, including reducing PR review times from 150 hours to just 15 hours, and dramatically compressing the cycle from user feedback to shipped features. What you’ll learn: How to drive AI adoption in large, established engineering o

    adoption · Claire Vo

  15. practice · Hamel Husain ·

    Evals Skills for Coding Agents

    A set of reusable evals skills that route teams to the right evaluation approach based on their stage and data.

    Today, Shreya Shankar and I are publishing evals skills , a set of skills for AI product evals 1 . Eval tools often get in the way. They nudge you toward generic off-the-shelf metrics and fully automated evals before you’ve looked at your data. These skills help you avoid common mistakes we’ve seen helping 50+ companies and teaching students in our AI Evals course . Why skills for evals There are many easily avoidable footguns in evals. These skills help you avoid them. evals-start is the entry point. It looks at your situation and routes you to the right skill. Most of the time it will send y

    ways of working · Hamel Husain

  16. practice · Artificial Ignorance ·

    BYOB: Build Your Own Benchmark

    Long-running agent benchmarks reveal how models fail under sustained stress, replacing saturated evaluation metrics with realistic work scenarios.

    What do vending machines, corporate whistleblowers, and the board game Diplomacy have in common? They’re all AI benchmarks. Vending-Bench drops an AI agent into a simulated vending machine business and asks it to manage inventory, negotiate with suppliers over email, set prices, and pay daily fees - for months of simulated time (a single run can burn through 60 to 100 million output tokens). The best models turn a meager profit; the worst ones go entirely off the rails, in what the authors call a “meltdown loop.” 1 What does a meltdown loop look like? In one short run, Claude 3.5 Sonnet mistak

    judgment · Charlie Guo

  17. research · National Bureau of Economic Research ·

    What Makes New Work Different from More Work?

    New occupational roles attract younger, more educated workers and command persistent wage premiums that decline as expertise diffuses across cohorts.

    We study the role of expertise in new work-novel occupational roles that emerge as technological and economic conditions evolve-using newly available 1940 and 1950 Census Complete Count files and confidential American Community Survey data from 2011-2023.We show that new work is systematically distinct from simply more work in existing occupations in four respects.First, it attracts workers with distinct characteristics: new work is disproportionately performed by younger and more educated workers, even within detailed occupation-industry cells.Second, new work commands economically significan

    jobs skills · David Autor

  18. research · NBER ·

    Artificial Intelligence, Productivity, and the Workforce: Evidence from Corporate Executives

    Survey of 750 corporate executives shows more than half have invested in AI, with substantial adoption heterogeneity across firm sizes.

    We use novel data from a survey of nearly 750 corporate executives to study the effects of artificial intelligence (AI) on productivity and the workforce. We document substantial heterogeneity in AI adoption across firms, with more than half having already invested, though many smaller firms are (Salomé Baslandze , Zachary Edwards , John Graham , Ty McClure , Brent H. Meyer , Michael Sparks , Sonya R. Waddell , Daniel Weitz)

    adoption

  19. research · NBER ·

    Mind the Gap: AI Adoption in Europe and the U.S.

    Worker and firm surveys from 2025-26 show large AI adoption gaps between US and Europe, partly explained by workforce and firm composition differences.

    This paper combines international evidence from worker and firm surveys conducted in 2025 and 2026 to document large gaps in AI adoption, both between the US and Europe and across European countries. Cross-country differences in worker demographics and firm composition account for an important share (Alexander Bick , Adam Blandin , David J. Deming , Nicola Fuchs-Schündeln , Jonas Jessen)

    adoption

  20. research · Organization Science ·

    Navigating the Jagged Technological Frontier: Field Experimental Evidence of the Effects of Artificial Intelligence on Knowledge Worker Productivity and Quality

    Field experiment with 758 knowledge workers shows AI improves performance on 18 tasks within its capabilities but worsens it on tasks outside the frontier.

    We introduce and study the concept of a “jagged technology frontier” to describe the uneven impact of artificial intelligence (AI) capabilities, where AI assistance improves performance for some tasks but worsens it for others, even within the same knowledge workflow and with a seemingly similar level of difficulty. In collaboration with the global management consulting firm Boston Consulting Group, we have developed realistic management consulting tasks and examined the human performance implications of using AI to perform complex and knowledge-intensive work. The preregistered experiment inv

    productivity · Hila Lifshitz-Assaf · Katherine Kellogg · Karim Lakhani · Ethan Mollick · Fabrizio Dell'Acqua · Edward McFowland III · François Candelon

  21. research · Management Science ·

    The Effects of Generative AI on High-Skilled Work: Evidence from Three Field Experiments with Software Developers

    Field experiments across three large firms show AI coding assistants increased developer task completion by 26%, with larger gains for junior developers.

    This study evaluates the effect of generative artificial intelligence (AI) on software developer productivity via randomized controlled trials at Microsoft, Accenture, and an anonymous Fortune 100 company. These field experiments, run by the companies as part of their ordinary course of business, provided a random subset of developers with access to an AI-based coding assistant suggesting intelligent code completions. Although each experiment is noisy and results vary across experiments, when data are combined across three experiments and 4,867 developers, our analysis reveals a 26.08% increas

    productivity · Mert Demirer · Sida Peng · Sonia Jaffe

  22. research · Government Information Quarterly ·

    Positioning public sector practitioners as ‘moral crumple zones’: Mechanisms in the early use of generative AI work support tools

    Ethnographic study finds public sector workers absorb ethical and operational liability when using generative AI tools, through five identified mechanisms including tool immaturity and offloaded trial-and-error learning.

    Work support tools built on generative artificial intelligence (AI), like Microsoft 365 Copilot and other AI assistants, are increasingly the object of experimentation in public sector. Yet, empirical research on workers' experiences with these tools remains limited. This article draws on 2.5 years of ethnographic fieldwork on AI innovation in two Finnish organizations and three cases of such tools to critically explore the positioning of users as ‘moral crumple zones’ who bear the burden of ensuring the effective and ethical use of these emerging tools. The findings capture five mechanisms th

    worker experience

  23. practice · Paul Ford (Aboard) ·

    Born to Runtime

    Runtime systems let users compose complex outputs by describing intent in plain language, with AI translating to executable code.

    In past issues of this newsletter, I’ve mentioned “runtimes” as being important in the age of AI. I should explain what I mean. Luckily, an interesting example showed up that I think makes the concept very plain. It’s called Endless , from a company called Polyend, and it’s a guitar pedal. It’s not out yet, but Polyend has a good history of shipping complicated hardware. Here’s a video that explains what it does: Most guitar pedals offer one effect—adding echoes, or distortion, or reverb—but this one is totally programmable. You can program it the old-fashioned way, in C++, or you can describe

    ways of working · Paul Ford

  24. practice · Shopify Engineering ·

    The generative recommender behind Shopify's commerce engine

    Shopify built a generative recommender treating buyer journeys as sequences, deployed at scale in production.

    Treating buyer journeys as sequences instead of simplified signals, and building a model fast enough to serve at scale.

    productivity

  25. practice · How I AI ·

    5 OpenClaw agents run my home, finances, and code | Jesse Genet

    Parent runs five specialized agents on dedicated hardware, each with clear role and scope, integrated into Obsidian and physical workflows.

    Jesse Genet is a homeschooling parent and entrepreneur who runs her household with five specialized OpenClaw agents. She layers them on top of her Obsidian “second brain,” deploys each on its own Mac Mini, and assigns every agent a distinct role—homeschool, finance, scheduling, development, and operations—so each one operates with clear scope and responsibility. What you’ll learn: How Jesse set up five OpenClaw agents, each with its own role, persona, SOUL.md file, and dedicated Mac Mini The workflow for photographing an entire curriculum book and having an agent generate formatted, ready-to-t

    ways of working · Claire Vo

  26. research · Knowledge at Wharton ·

    When Does AI Assistance Undermine Learning?

    AI assistance on-demand reduces practice and productive struggle, harming long-term skill retention even when learners know this.

    Wharton research shows that giving learners on-demand AI assistance can erode practice, "productive struggle," and long-term skill growth — even when they know it harms their learning. … Read More

    jobs skills

  27. practice · GitHub Blog ·

    Multi-agent workflows often fail. Here’s how to engineer ones that don’t.

    Multi-agent systems need explicit state management, data formats and interfaces like distributed systems, not chat interfaces.

    If you’ve built a multi-agent workflow, you’ve probably seen it fail in a way that’s hard to explain. The system completes, and agents take actions. But somewhere along the way, something subtle goes wrong. You might see an agent close an issue that another agent just opened, or ship a change that fails a downstream check it didn’t know existed. That’s because the moment agents begin handling related tasks—triaging issues, proposing changes, running checks, and opening pull requests—they start making implicit assumptions about state, ordering, and validation. Without providing explicit instruc

    ways of working

  28. research · METR ·

    We are Changing our Developer Productivity Experiment Design

    Previous study found AI slowed experienced developers 20%; new larger experiment shows selection bias makes results unreliable but suggests speedup may have increased since early 2025.

    METR previously published a paper which found the use of AI tools caused a 20% slowdown in completing tasks among experienced open-source developers, using data from February to June 2025. To understand how AI is impacting developer productivity over time, we started a new experiment in August 2025 with a larger pool of developers using the latest AI tools. Unfortunately, given participant feedback and surveys, we believe that the data from our new experiment gives us an unreliable signal of the current productivity effect of AI tools. The primary reason is that we have observed a significant

    adoption

  29. practice · How I AI ·

    “I haven’t written a single line of front-end code in 3 months”: How Notion’s design team uses Claude Code to prototype

    Design team built a shared Claude Code environment to prototype functional AI features collaboratively instead of static designs.

    Brian Lovin is a designer at Notion AI who has transformed how the design team builds prototypes, by creating a shared code environment powered by Claude Code. Instead of designers working in isolated repositories or limited to static Figma designs, Brian built a collaborative “prototype playground” where the entire team can create, share, and iterate on functional prototypes. In this episode, Brian demonstrates how AI-assisted coding has dramatically accelerated the design process and why code-based prototyping is essential for building AI-powered products. What you’ll learn: How Brian built

    ways of working · Claire Vo

  30. research · Information Systems Research ·

    SUVA: A Probabilistic Framework for Auditing LLMs with an Application to Social Preferences

    Framework to audit LLM decision-making by extracting reasoning and values, showing alignment differences across eight models.

    Organizations are increasingly deploying large language models (LLMs) as customer service agents, decision aids, and semiautonomous agents. We develop State–Understanding–Value–Action (SUVA), a probabilistic auditing framework that turns an LLM’s response into structured evidence about how its decision was produced. SUVA treats the prompt as the state, codes the model’s reasoning to extract its understanding and stated values using a transparent codebook, and then estimates how these elements statistically predict the eventual action. We demonstrate SUVA on social preference games from behavio

    judgment

  31. practice · Artificial Ignorance ·

    The Emerging "Harness Engineering" Playbook

    Engineering teams reorganizing around agent workflows, with concrete metrics: one engineer ships 6,600+ commits monthly running 5-10 agents; three engineers built a million-line product in five months.

    Earlier this month, Greg Brockman published a thread about how OpenAI is retooling its engineering teams to make them more effective with agents. The initiative was kicked off because of how much things have changed internally: Some great engineers at OpenAI yesterday told me that their job has fundamentally changed since December. Prior to then, they could use Codex for unit tests; now it writes essentially all the code and does a great deal of their operations and debugging. Not everyone has yet made that leap, but it’s usually because of factors besides the capability of the model. I’ve pre

    management org · Charlie Guo

  32. practice · Hacker News ·

    How I use Claude Code: Separation of planning and execution

    Separating planning from execution when using Claude Code improves code quality and reduces errors.

    976 points on Hacker News. Discussion: https://news.ycombinator.com/item?id=47106686

    ways of working

  33. practice · Bob Sutton (Work Matters) ·

    Decentralized Organizations Can Be SLOWER

    Decentralized teams racing independently can create coordination debt that slows platform companies, requiring standardization and collaboration.

    Decentralization and the "local" speed that comes with it works pretty well when there is not a premium in coordination and standardization. BUT, as this article by MIT Sloan School of Management 's Georg Rilinger unpacks, the decentralization and rapid iteration that are hallmarks of "agile" development create a host of coordination and consistency problems as unfettered teams race in different directions for firms in "the platform marketplace." Especially when companies don't put a premium on collaboration. My colleague Huggy Hayagreeva Rao and I document a similar challenge in Uber engineer

    management org · Bob Sutton

  34. research · Information Systems Research ·

    AI Governance and the Decentralization of Technology Production: An Investigation of AI-Based IPA Bots

    Centralized mandates suppress IPA bot adoption; decentralization works until fragmentation emerges, requiring adaptive governance.

    We revisit the centralization–decentralization tension in the context of decentralized technology production at the AI frontier, focusing on Intelligent Process Automation (IPA) bots as a salient manifestation of the democratization of AI. IPA bots combine robotic process automation with AI technologies and process mining based on deep, mindful domain expertise. We collaborate with a Fortune 200 multinational to study how IPA projects yield successful governance outcomes of utilization and repeatability. Our research reveals that traditional centralized mandates, when paired with the unique le

    adoption

  35. research · ACM Transactions on Software Engineering and Methodology ·

    Unveiling the Role of ChatGPT in Software Development: Insights from Developer–ChatGPT Interactions on GitHub

    Analysis of 2,547 real ChatGPT conversations from GitHub shows developers use it mainly for code generation, bug fixing, and knowledge acquisition in short task-focused sessions.

    The advent of Large Language Models (LLMs) has introduced a new paradigm in Software Engineering (SE), with generative AI tools like ChatGPT gaining widespread adoption among developers. While ChatGPT’s potential has been extensively discussed, empirical evidence about how developers actually use LLMs’ assistance in real-world practices remains limited. To bridge this gap, we conducted a large-scale empirical analysis of ChatGPT usage on GitHub, and we presented DevChat, a curated dataset of 2,547 publicly shared ChatGPT conversation links collected from GitHub between May 2023 and June 2024.

    adoption

  36. research · PNAS Nexus ·

    Toward a science of human–AI teaming for decision making: A complementarity framework

    Framework identifying sociotechnical factors and design principles for human-AI teams that achieve complementarity in decision-making contexts.

    As artificial intelligence (AI) becomes embedded in critical decisions involving health, safety, finance, and governance, the key challenge is no longer whether humans and AI will collaborate, but rather how to structure this collaboration to achieve true complementarity. Human-AI complementarity refers to the conditions under which human-AI teams outperform either humans alone or AI systems alone. This paper advances the science of human-AI teaming for decision making by integrating insights from cognitive science, AI, human factors, organizational behavior, and ethics. We propose a framework

    judgment · Anita Woolley

  37. practice · Bob Sutton (Work Matters) ·

    Rethinking My Craft as a Teacher, Speaker, and Coach

    Experienced management educator reconsiders teaching methods and assumptions in light of AI and technological change.

    I’ve been fretting about how the education of leaders and organizational designers is and ought to change (and stay the same too). The hope, hype, threat, and shear madness of AI-- and other massive and unsettling technological, societal, and political shifts---has caused me to question my assumptions about what means for me to be a good management teacher and speaker. I’ve been practicing and trying to hone variations of this craft for over 40 years. I’ve taught thousands of undergraduate and graduate students at Stanford in the process. Leaders of everything from two-person start-ups to top

    jobs skills · Bob Sutton

  38. practice · Paul Ford (Aboard) ·

    My New York Times Op-Ed on Vibe Coding

    AI coding tools can democratise software development for non-specialists without deprofessionalising the craft itself.

    I’d outlined a fun little newsletter item on AI runtimes, including defining the term—“runtime” being one of those words like “platform” that nerds love to say and then forget to explain—but then the New York Times Opinion section emailed to ask if I’d explain vibe coding. We’ll do runtimes next week. I know that being published in the Times is an unbelievable privilege. However, for a long moment, I thought, Do I want to do this? I’m gonna get yelled at. But I did want to do it, and here is the result . I’m deeply convinced that it’s possible to accelerate software development with AI coding—

    jobs skills · Paul Ford

  39. research · Organization Science ·

    Knowing Enough to Be Dangerous: The Problem of “Artificial Certainty” for Expert Authority When Using AI for Decision Making and Planning

    Experts using AI simulations risk creating false certainty about complex futures by amplifying technological capabilities rather than moderating how outputs are presented to stakeholders.

    This study examines how experts who use advanced artificial intelligence (AI) technologies that generate highly detailed and realistic representations can create what we term “artificial certainty,” which we define as the illusion that complex future outcomes are definitively knowable, even though they are inherently uncertain. Through a comparative study of two urban planning organizations using the same AI simulation tool, we show how this artificial certainty emerges from the ways process experts create and deploy AI-generated representations. The findings reveal three interconnected repres

    judgment · Paul Leonardi

  40. research · METR ·

    Analyzing coding agent transcripts to upper bound productivity gains from AI agents

    Analysis of 5,305 coding agent transcripts suggests time savings of 1.5x to 13x on assisted tasks, though authors note this likely overestimates true productivity gains.

    Introduction Human uplift studies like the one we did in 2025 are becoming more expensive as working without AI becomes increasingly costly. In this post, I investigate whether coding agent transcripts could serve as a cheaper alternative for estimating uplift. I prototyped this using 5305 Claude Code transcripts generated in January 2026 by 7 METR technical staff 1 . I used an LLM judge to estimate how long each task would have taken an experienced software engineer without AI tools, then compared that to the time people actually spent on these tasks to calculate a time savings factor . Takea

    productivity

  41. practice · Rands in Repose ·

    Extremely Lazy and Immensely Curious

    Developer describes how Claude Code changed their workflow from reluctant Unix user to curious explorer of system capabilities.

    When I explain that Claude Code has changed my relationship with developing software completely, I’m under-exaggerating… if that’s even a thing. This off-the-cuff piece started as a unposted social update that read: Watching Claude Code adeptly use every type of Unix command shows me that a) you can do anything in Unix, b) my higher-level operating system mostly hides this functionality from me, c) I am extremely lazy, and d) I am immensely curious. An Introduction to Pansy Rain This morning it’s going to start to rain — a lot. As previously described, I deeply enjoy tromping around my forest

    ways of working

  42. practice · How I AI ·

    How this visually impaired engineer uses Claude Code to make his life more accessible | Joe McCormick

    Visually impaired engineer builds micro Chrome extensions in under 25 minutes using Claude Code to solve accessibility gaps mainstream tools miss.

    Joe McCormick is a principal software engineer at Babylist who lost most of his central vision due to a rare genetic disorder right before starting college. He pivoted from mechanical engineering to computer science and now leads AI enablement at Babylist. Joe demonstrates how he uses AI to build micro Chrome extensions that make his everyday work and life more accessible, showing how personal software can address accessibility needs that mainstream products often overlook. What you’ll learn: How to build custom Chrome extensions in under 25 minutes using Claude Code A practical workflow for c

    ways of working · Claire Vo

  43. practice · Addy Osmani ·

    14 More lessons from 14 years at Google

    How engineering teams actually make decisions and coordinate work, drawn from 14 years at Google.

    A while back, I wrote down 21 lessons from my time at Google . The response caught me off guard because of which ones stuck. It wasn’t tech-specific advice. It was the stuff about people, decisions, and the messy reality of building things together. That made me realize I’d left a lot on the table. The first list skewed toward individual craft - how to write better code, how to think about your career. But some of the hardest lessons I’ve learned aren’t about how you work. They’re about how teams work: how decisions actually get made, where coordination breaks down, what separates the groups t

    management org · Addy Osmani

  44. practice · Paul Ford (Aboard) ·

    Taking Your Turn

    Product managers must adapt to vibe coding by focusing on user fit and adaptation, not just code speed.

    It’s been a year since “vibe coding” was coined, and a few different models of AI-assisted software development have emerged in the meantime. As I look at these various approaches, I’m thinking about what “product management” means today. All of the focus is on coders—and on how much easier it is to code. But that doesn’t mean the code is good. A lot of vibe-coded products I see are basically databases with nice frontends. That’s great, and powerful, but those aren’t products —they aren’t adapted to their environments, or focused on their users. So what is the work of the product manager in th

    management org · Paul Ford

  45. research · MIS Quarterly ·

    FAIR: A Design Theory for Artificial Intelligence Fairness

    Design theory for managing persistent fairness tensions in AI decision systems through iterative adaptive cycles and organizational capability.

    Artificial intelligence (AI)-automated decision systems encounter persistent, interdependent, and dynamic fairness tensions that traditional one-off interventions cannot resolve. Because these tensions persist due to interdependence and dynamic interaction, organizations require both a theory of the problem to explain their persistence and a theory of the solution to prescribe how they can be managed. Our design theory, FAIR (fairness adaptation through AI-augmented responsiveness), provides a theory of the problem by reframing AI fairness as a sociotechnical paradox constituted within AI arti

    judgment

  46. research · Journal of Management Studies ·

    The Acceleration of Artificial Intelligence: Rethinking Organization and Work in an Era of Rapid Technological Change

    Framework distinguishing predictive, generative, agentic, and embodied AI systems and their distinct organizational effects on expertise, judgment, coordination, and authority.

    Abstract Artificial intelligence (AI) is transforming the epistemic, interactional, and institutional foundations of contemporary organizations, yet management and organization studies are only beginning to theorise the implications of this shift. Existing research often treats “AI” as a singular construct, despite the fact that predictive, generative, agentic, and embodied systems rely on different logics and produce distinct organizational outcomes. This article interrogates the limits of this conceptual flattening and argues that cumulative theorising requires more precise specification of

    management org · Stella Pachidi

  47. practice · How I AI ·

    How to build your own AI developer tools with Claude Code | CJ Hess (Tenex)

    Developer built Flowy tool converting ASCII diagrams to interactive mockups, uses model-to-model comparison for code quality assurance.

    CJ Hess is a software engineer at Tenex who has built some of the most useful tools and workflows for being a “real AI engineer.” In this episode, CJ demonstrates his custom-built tool, Flowy, that transforms Claude’s ASCII diagrams into interactive visual mockups and flowcharts. He also shares his process for using model-to-model comparison to ensure that his AI-generated code is high-quality, and why he believes we’re just at the beginning of a revolution in how developers interact with AI. What you’ll learn: How CJ built Flowy, a custom visual planning tool that converts JSON files into int

    ways of working · Claire Vo

  48. research · Information Systems Research ·

    Human-Algorithm Collaboration in Gig Work: The Role of Experience, Skill Level, and Task Complexity

    Field experiment shows AI-assisted picking tool complements experience, substitutes for skill gaps, with effects varying by workload and task complexity.

    In this paper, we contribute to recent studies on human-algorithm collaboration by examining how experience, skill level, workload, and task complexity shape the impact of an algorithm-enabled decision-support tool for gig workers. We leverage a large-scale randomized field experiment on the Instacart platform from June 2022 to September 2022. The algorithm-enabled technology aims to revolutionize item picking by helping shoppers locate and collect items more efficiently, reducing picking time while maintaining service quality, as reflected by refund rates. We find that the technology compleme

    adoption

  49. practice · Will Larson ·

    Refactoring internal documentation in Notion

    One person improved developer documentation through systematic diagnosis, prioritization, and measurement rather than visible but ineffective changes.

    In our latest developer productivity survey, our documentation was the area with the second most comments. This is a writeup of the concrete steps I took to see how much progress one person could make on improving the organization’s documentation while holding myself to a high standard for making changes that actually worked instead of optically sounding impressive. Diagnosis There were a handful of issues we were running into: We migrated from Confluence to Notion in January, 2025, which had left around a bunch of old pages that were “obviously wrong.” These files created a bad smell around o

    productivity

  50. research · Human Relations ·

    Short-term fit, long-term trap: The career development lock of low-skilled gig workers

    Platform algorithms create career development lock, trapping gig workers in temporary roles through structural conflict between worker aspirations and algorithmic flexibility demands.

    Why do low-skilled gig workers remain stuck in work they originally intended as temporary? Although gig work is widely portrayed as flexible and temporary, our study shows that platforms can gradually trap workers in place. Drawing on grounded theory and fieldwork—including 70 interviews with 42 food delivery riders and 14 ride-hailing drivers, 30 hours of firsthand riding experience, and observations of online communities totalling 820 riders—this study develops a conceptual framework of career development lock, identifying four interrelated forms: lock-out, lock-in, lock-up, and lock-down. W

    worker experience

  51. practice · Anthropic Engineering ·

    Building a C compiler with a team of parallel Claudes

    A team of parallel AI agents built a C compiler with minimal human oversight, showing practical autonomous software development.

    We tasked Opus 4.6 using agent teams to build a C Compiler, and then (mostly) walked away. Here's what it taught us about the future of autonomous software development.

    productivity

  52. practice · How I AI ·

    Guillermo Rauch: Vercel CEO on how v0 hit 3,200 PRs merged per day (and lets anyone ship)

    Vercel uses v0 to let non-technical team members submit production-ready code changes via Git workflow, merging 3,200 PRs per day.

    Guillermo Rauch , the CEO of Vercel, demonstrates how v0 has evolved from a simple prototyping tool to a complete development environment that supports the entire Git workflow. Guillermo shows how Vercel built skills.sh—a viral marketplace with over 34,000 community-submitted skills—using v0, and how the tool enables non-technical team members to contribute production-ready code changes. He walks through creating branches, implementing features, previewing changes, and submitting pull requests, all within v0. What you’ll learn: How v0’s new Git workflow integration enables anyone to contribute

    adoption · Claire Vo

  53. research · MIT Sloan Management Review ·

    Validating LLM Output? Prepare to Be ‘Persuasion Bombed’

    Management consultants using LLMs for strategic decisions were persuaded to override expert judgment through rhetorical tactics including data flooding and emotional appeals.

    A research study of management consultants who were asked to use a large language model to recommend strategic business decisions found that the AI responded to human validation attempts with persuasive rhetorical strategies. In addition to appealing to the user’s logic, sense of trust, and emotions, the AI also engaged in tactics such as flooding the user with large volumes of unrequested data and analyses that could overwhelm them and convince them to override their expert judgment.

    judgment · Hila Lifshitz-Assaf · Katherine Kellogg · Karim Lakhani · Steven Randazzo

  54. practice · Artificial Ignorance ·

    The Codex App Has Upended My Daily Workflow

    Developer stopped using IDE after switching to Codex App; now manages AI agents writing code instead of writing it directly.

    Disclaimer: Regular readers will know that I currently work at OpenAI . And while that certainly introduces some bias, the views presented here are entirely my own, without input from the company. I’m unlikely to post about every new release, but I have been so enamored with the new Codex App (and I have seen firsthand how hard the team has worked to make it great) that I am genuinely excited to evangelize this thing. Subscribe now Last week, I realized I hadn’t opened my AI IDE in four days. This wasn’t a conscious decision - no dramatic uninstall, no declaration that I was done with IDEs. I

    ways of working · Charlie Guo

  55. practice · How I AI ·

    How this PM uses MCPs to automate his meeting prep, CRM updates, and customer feedback synthesis | Reid Robinson (Zapier)

    PM uses Model Context Protocols to connect Claude to 8000+ apps via Zapier for automating meeting prep, CRM updates, and customer feedback synthesis.

    Reid Robinson , Principal AI Product Strategist at Zapier, shares how he uses Model Context Protocols (MCPs) to automate tedious tasks and create powerful workflows. He demonstrates practical workflows that combine Zapier’s more than 8,000 app connections with AI tools like Claude to create systems that work while he sleeps. What you’ll learn: How to use Zapier’s MCP server to create custom collections of tools that work seamlessly with Claude, ChatGPT, and other AI assistants A workflow for using Claude Projects to provide detailed instructions for tool usage, improving reliability and consis

    ways of working · Claire Vo

  56. research · arXiv ·

    The Innovation Tax: Generative AI Adoption, Productivity Paradox, and Systemic Risk in the U.S. Banking Sector

    Banks adopting GenAI saw a 428-basis-point ROE decline from implementation costs, with larger spillover effects creating systemic risk through algorithmic coupling.

    This paper evaluates the causal impact of Generative Artificial Intelligence (GenAI) adoption on productivity and systemic risk in the U.S. banking sector. Using a novel dataset linking SEC 10-Q filings to Federal Reserve regulatory data for 809 financial institutions over 2018--2025, we employ two complementary identification strategies: Dynamic Spatial Durbin Models (DSDM) to capture network spillovers and Synthetic Difference-in-Differences (SDID) for causal inference using the November 2022 ChatGPT release as an exogenous shock. Our findings reveal a striking ``Productivity Paradox'': whil

    productivity

  57. research · NBER ·

    Firm Data on AI

    Survey of nearly 6,000 executives across four countries on AI adoption, effects on jobs, productivity, and output over the past three years.

    We survey nearly 6,000 senior business executives at US, UK, German, and Australian firms to develop new evidence on AI adoption and its effects on jobs, productivity, and output. Specifically, we ask executives about AI usage, its effects at their own firms over the past three years and, looking (Ivan Yotzov , Jose Maria Barrero , Nicholas Bloom , Philip Bunn , Steven J. Davis , Kevin M. Foster , Aaron Jalca , Brent H. Meyer , Paul Mizen , Michael A. Navarrete , Pawel Smietanka , Gregory Thwaites , Ben Zhe Wang)

    adoption

  58. research · NBER ·

    The Politics of AI

    Partisan differences in workplace AI adoption reflect educational and occupational sorting, not ideology, using Gallup Workforce Panel data.

    Using new data from the Gallup Workforce Panel, we show that the apparent partisan divide in workplace AI adoption is largely an artifact of educational and occupational sorting rather than ideological differences in technology adoption. While Democrats are consistently more likely than Republicans (Nicholas Bloom , Christos Makridis)

    adoption

  59. research · NBER ·

    Enhancing Worker Productivity Without Automating Tasks: A Different Approach to AI and the Task-Based Model

    Analysis challenges the dominant task-replacement framework, arguing contemporary AI more often augments worker productivity than automates tasks away.

    The task-based approach has become the dominant framework for studying the labor-market effects of artificial intelligence (AI), typically emphasizing the replacement of human workers by machines. Motivated by growing empirical evidence that contemporary AI is more often used as a tool that augments (Ajay K. Agrawal , John McHale , Alexander Oettl)

    jobs skills

  60. research · The Quarterly Journal of Economics ·

    Automation and Rent Dissipation: Implications for Wages, Inequality, and Productivity

    Automation targets high-rent jobs, dissipating wages and offsetting 60-90% of productivity gains; rent loss explains one-fifth of post-1980 inequality growth.

    Abstract This article studies the effects of automation in a task-based economy in which some jobs pay workers rents—wages above workers' outside options. We show that automation targets high-rent tasks, dissipating rents, amplifying wage losses, and reducing within-group wage dispersion in exposed groups. This form of rent dissipation is inefficient and offsets the productivity gains from automation. Using U.S. data from 1980 to 2016, we find evidence of sizable rent dissipation and reduced within-group wage dispersion due to automation. Automation accounts for 52% of the increase in between-

    productivity · Daron Acemoglu · Pascual Restrepo

  61. practice · The Pragmatic Engineer ·

    The creator of Clawd: "I ship code I don't read"

    Solo developer ships production code by delegating code reading and review to AI agents in his workflow.

    How Peter Steinberger, creator of OpenClaw (formerly: Clawd), builds and ships like a full team by centering his development workflow around AI agents. Listen now (114 mins) | How Peter Steinberger, creator of OpenClaw (formerly: Clawd), builds and ships like a full team by centering his development workflow around AI agents. Stream the latest episode

    ways of working · Gergely Orosz

  62. practice · Addy Osmani ·

    The 80% Problem in Agentic Coding

    Developers now spend 80% of time directing AI agents and 20% editing, inverting the previous ratio; workflow and model improvements have crossed a threshold.

    said something this week that made me pause: “I rapidly went from about 80% manual+autocomplete coding and 20% agents to 80% agent coding and 20% edits+touchups. I really am mostly programming in English now.” The inversion happened over a few weeks in late 2025. While this may apply to new (greenfield) or personal projects more than existing or legacy apps, I imagine how far AI takes you is still further than a year ago. You can thank models, specs, skills, MCPs and our workflows improving. Boris Cherney, creator of Claude Code, has recently echoed similar sentiments: “Pretty much 100% of our

    ways of working · Addy Osmani

  63. practice · How I AI ·

    I gave Clawdbot (aka Moltbot) access to my computer, calendar, and emails: Here’s what happened

    First-hand setup and permission decisions when running an autonomous agent with access to computer, calendar, email and accounts.

    In this episode, I take you through my unfiltered experience with Clawdbot , the viral open-source AI agent that’s been taking over tech Twitter. (In the time since this was recorded, the tool was renamed Moltbot , but we’re calling it Clawdbot here to match the episode.) It’s an autonomous AI that can run code, spin up sub-agents, join video calls, and take real actions on your machine. I invite it onto the podcast, give it screen access, and walk through what it’s like to go from zero to one with an agentic AI that actually does things. Along the way, I share the real experience: installatio

    ways of working · Claire Vo

  64. practice · Hacker News ·

    A verification layer for browser agents: Amazon case study

    Local 3B model with DOM verification layer completes Amazon checkout steps reliably, matching cloud baseline with lower token cost.

    A common approach to automating Amazon shopping or similar complex websites is to reach for large cloud models (often vision-capable). I wanted to test a contradiction: can a ~3B parameter local LLM model complete the flow using only structural page data (DOM) plus deterministic assertions? This post summarizes four runs of the same task (search → first product → add to cart → checkout on Amazon). The key comparison is Demo 0 (cloud baseline) vs Demo 3 (local autonomy); Demos 1–2 are intermediate controls. More technical detail (architecture, code excerpts, additional log snippets): https://ww

    productivity

  65. research · npj Digital Medicine ·

    Human–large language model collaboration in clinical medicine: a systematic review and meta-analysis

    Meta-analysis of 10 studies finds LLM collaboration improved diagnostic scores slightly but showed high error rates and no time gains.

    Human-AI collaboration (H + AI) using large language models (LLMs) offers a promising approach to enhance clinical reasoning, documentation, and interpretation tasks. Following PRISMA 2020 (PROSPERO registration: CRD420251068272), we systematically compared H + AI with human-only (H) workflows, searching four databases through June 28, 2025. Ten peer-reviewed studies met eligibility criteria, with three preprints informing sensitivity analyses only. Diagnostic/interpretation accuracy (k = 2) showed a positive trend for H + AI (Risk Ratio [RR] 1.59), but was statistically imprecise and non-sign

    judgment

  66. practice · The Pragmatic Engineer ·

    Inside a five-year-old startup’s rapid AI makeover

    A five-year-old startup built an internal agentic tool that the whole team now uses, with measurable results from the pivot.

    Craft Docs resisted AI hype as a nimble, well-run startup – but has just made a sharp pivot to AI, building a universal agentic tool everyone there now uses. The results are phenomenal. Exclusive. Craft Docs resisted AI hype as a nimble, well-run startup – but has just made a sharp pivot to AI, building a universal agentic tool everyone there now uses. The results are phenomenal. Exclusive. Update on 27 June 2026: Polymarket has acquired Craft Agents, which is an open-source Claude Cowork-like agents experience, and Balint is joining Polymarket to lead product engineering. I’m making this prev

    ways of working · Gergely Orosz

  67. practice · One Useful Thing ·

    Management as AI superpower

    MBA students built working prototypes and business models for five startups in four days using AI tools for coding, research, positioning and financial modelling.

    I just taught an experimental class at the University of Pennsylvania where I challenged students to create a startup from scratch in four days. Most of the people in the class were in the executive MBA program, so they were taking classes while also working as doctors, managers, or leaders in a variety of large and small companies. Few had ever coded. I introduced them to Claude Code and Google Antigravity, which they needed to use to build a working prototype. But a prototype alone is not a startup, so they used ChatGPT, Claude, and Gemini to accelerate the idea generation, market research,

    adoption · Ethan Mollick

  68. research · Strategic Management Journal ·

    Searching together versus searching apart: Evidence from Kaggle

    Teams searching together generate more attempts but less exploration than independent search; effectiveness depends on strategic context.

    Abstract Research Summary How does the mode of search—independently or jointly—affect collective search, a central component of organizational adaptation and innovation? Using naturally occurring data from a strongly incentivized online competition platform, we find that compared to their counterfactuals that search apart, groups searching together exhibit less exploration in their search outcomes as noted in prior experimental and computational modeling studies. However, groups searching together stimulate a greater number of search attempts from their members than groups searching apart, an

    teams · Phanish Puranam

  69. research · IEEE Transactions on Software Engineering ·

    From Disruptions to Discussions: How GenAI Impacts Human Interactions in Software Development

    GenAI shifts developer interactions from answering routine questions to deeper collaboration, and developers report fewer flow disruptions.

    New technologies often change how an individual performs work, such as how generative AI (GenAI) can help a developer write code. New technologies can also impact how people interact with one another, such as how GenAI’s ability to summarize API documentation can reduce the need for developers to ask each other technical questions. In this paper, we report on a two-phase mixed-method study exploring how GenAI influences how humans interact in software development. During phase one, 30 industrial software developers provided data over a period of 5 to 12 days as they worked, contributing 627 ex

    teams

  70. practice · Hacker News ·

    Porting 100k lines from TypeScript to Rust using Claude Code in a month

    Team ported 100k TypeScript lines to Rust in a month using Claude Code as primary tool, detailing workflow and lessons.

    255 points on Hacker News. Discussion: https://news.ycombinator.com/item?id=46765694

    productivity

  71. practice · How I AI ·

    Advanced Claude Code techniques: context loading, mermaid diagrams, stop hooks, and more | John Lindquist

    Senior engineers use mermaid diagrams and custom hooks to load context into Claude Code, accelerating coding assistance and automating quality checks.

    John Lindquist is the co-founder of egghead.io and an expert in leveraging AI tools for professional software development. In this episode, John shares advanced techniques for using AI coding tools like Claude Code and Cursor that go far beyond basic prompting. He demonstrates how senior engineers can use mermaid diagrams for context loading, create custom hooks for automated code quality checks, and build efficient command-line tools that streamline AI workflows. What you’ll learn: How to use mermaid diagrams to preload context into Claude Code for faster, more accurate coding assistance Crea

    ways of working · Claire Vo

  72. practice · Artificial Ignorance ·

    Skills, Tools and MCPs - What’s The Difference?

    Framework distinguishing Tools, MCPs and Skills: when each solves different problems in connecting AI to external systems.

    Two and a half years ago, OpenAI released function calling for GPT-4 . I still remember that distinct “wow” moment - the realization that language models could actually do things beyond generating text. Not just answer questions or write essays, but call APIs, manipulate data, and take actions in the real world. It felt like watching the future. And yet - I’ve since watched ChatGPT Plugins (the first feature to use function calling) launch with fanfare, quietly get deprecated, be replaced by GPTs, and now find themselves being superseded by this new wave of Skills and MCPs. The progression tau

    ways of working · Charlie Guo

  73. practice · Latent Space ·

    Captaining IMO Gold, Deep Think, On-Policy RL, Feeling the AGI in Singapore — Yi Tay

    A research leader describes scaling deep reasoning from competition problems to production models across Google DeepMind's organization.

    From shipping Gemini Deep Think and IMO Gold to launching the Reasoning and AGI team in Singapore, Yi Tay has spent the last 18 months living through the full arc of Google DeepMind’s pivot from architecture research to RL-driven reasoning—watching his team go from a dozen researchers to 300+, training models that solve International Math Olympiad problems in a live competition, and building the infrastructure to scale deep thinking across every domain, and driving Gemini to the top of the leaderboards across every category. From shipping Gemini Deep Think and IMO Gold to launching the Reasoni

    management org · Shawn Wang (swyx)

  74. research · Strategic Organization ·

    Talk isn’t always cheap: A theory of social influence and deliberation in group decision-making

    Agent-based model shows AI agents resistant to conformity can improve group decision-making by reducing clustering effects, even if they lack expertise.

    Organizations often rely on deliberative groups—committees, taskforces, boards—to make decisions, yet deliberation’s effectiveness remains contested. When can deliberation help a group outperform the average or even the best individual in the group? We propose that the value of deliberation depends on how the network structure of the group shapes informational influence (which promotes updating beliefs based on perceived expertise) and normative influence (which drives conformity to gain social approval). Using an agent-based model, we show that deliberation can help groups achieve strong syne

    judgment · Phanish Puranam

  75. research · Indeed Hiring Lab ·

    January 2026 US Labor Market Update: Jobs Mentioning AI Are Growing Amid Broader Hiring Weakness

    Job postings mentioning AI are growing while overall US hiring weakens, showing employer concentration on AI-related roles.

    Small pockets of growth are emerging as employers concentrate their limited hiring on roles and skills tied to AI. The post January 2026 US Labor Market Update: Jobs Mentioning AI Are Growing Amid Broader Hiring Weakness appeared first on Indeed Hiring Lab .

    jobs skills

  76. research · Government Information Quarterly ·

    Generative AI in public administration: A quasi-experimental analysis of bureaucratic productivity

    Quasi-experimental study finds generative AI reduces document drafting time by 4 minutes per task; new employees gain most, suggesting AI aids skill gaps.

    This study investigates the effect of a specialized generative AI system—developed and tested as part of a Proof of Concept (PoC)—on the speed with which public officials in a central government agency draft written responses to citizen inquiries. Using a quasi-experimental design with difference-in-differences analysis and propensity score matching, 80 civil servants were divided into a treatment group participating in the proof of concept (adopting AI-based drafting) and a control group that maintained existing practices. The difference-in-differences results of this study indicate a signifi

    productivity

  77. research · Science ·

    Who is using AI to code? Global diffusion and impact of generative AI

    AI-generated code now represents 29% of Python functions on GitHub; AI boosted experienced developers' productivity by 3.6% but showed no benefit for early-career developers.

    Generative coding tools promise big productivity gains, but uneven uptake could widen skill and income gaps. We train a neural classifier to spot artificial intelligence (AI)-generated Python functions in more than 30 million GitHub commits by 160,097 software developers, tracking how fast, and where, these tools take hold. Currently, AI writes an estimated 29% of Python functions in the US-a shrinking lead over other countries. We estimate that quarterly output, measured in online code contributions, consequently increased by 3.6%. AI seems to benefit experienced, senior-level developers: The

    jobs skills

  78. practice · Anthropic Engineering ·

    Designing AI-resistant technical evaluations

    A technical hiring evaluation evolved through three rounds as Claude improved, revealing what skills remain hard for AI to fake.

    What we learned from three iterations of a performance engineering take-home that Claude keeps beating.

    jobs skills

  79. practice · GitHub Blog ·

    AI-supported vulnerability triage with the GitHub Security Lab Taskflow Agent

    LLM agents triaged security alerts by matching fuzzy code patterns, discovering 30 real vulnerabilities since August with basic tools only.

    Triaging security alerts is often very repetitive because false positives are caused by patterns that are obvious to a human auditor but difficult to encode as a formal code pattern. But large language models (LLMs) excel at matching the fuzzy patterns that traditional tools struggle with, so we at the GitHub Security Lab have been experimenting with using them to triage alerts. We are using our recently announced GitHub Security Lab Taskflow Agent AI framework to do this and are finding it to be very effective. 💡 Learn more about it and see how to activate the agent in our previous blog post

    productivity

  80. research · Knowledge at Wharton ·

    Who Gets Replaced by AI and Why?

    AI implementation in certain workflow positions reduces employee motivation more than others, with implications for where to deploy AI.

    New research from Wharton’s Pinar Yildirim reveals how AI can impact employee motivation when implemented in the wrong part of a team’s workflow. … Read More

    worker experience

  81. research · Research Policy ·

    Firm training, automation, and wages: International worker-level evidence

    Firm training reduces automation risk by 3.8 percentage points across 37 countries, accounting for 15% of wage returns to training.

    Firm training is widely regarded as crucial for protecting workers from automation, yet there is a lack of empirical evidence to support this belief. Using internationally harmonized data from over 90,000 workers across 37 industrialized countries, we construct an individual-level measure of automation risk based on tasks performed at work. Our analysis reveals substantial within-occupation variation in automation risk, overlooked by existing occupation-level measures. To assess whether firm training mitigates automation risk, we exploit within-occupation and within-industry variation. Additio

    jobs skills

  82. research · Management Science ·

    The Rapid Adoption of Generative AI

    27% of US employed workers used genAI for work weekly as of late 2024; work adoption faster than PCs, with 1-7% of work hours assisted by genAI.

    Generative artificial intelligence (genAI) is a potentially important new technology, but its impact on the economy depends on the speed and intensity of adoption. This paper reports results from a series of nationally representative U.S. surveys of genAI use at work and at home. As of late 2024, 45% of the U.S. population age 18–64 uses genAI. Among employed respondents, 27% used genAI for work at least once in the previous week: 10% used it every workday and 17% on some but not all workdays. Relative to each technology’s first mass-market product launch, work adoption of genAI has been faste

    adoption

  83. practice · Addy Osmani ·

    How to write a good spec for AI agents

    Framework for writing AI agent specs that balance clarity with context limits, using iterative planning and task decomposition.

    TL;DR: Aim for a clear spec covering just enough nuance (this may include structure, style, testing, boundaries) to guide the AI without overwhelming it. Break large tasks into smaller ones vs. keeping everything in one large prompt. Plan first in read-only mode, then execute and iterate continuously. “I’ve heard a lot about writing good specs for AI agents, but haven’t found a solid framework yet. I could write a spec that rivals an RFC, but at some point the context is too large and the model breaks down.” Many developers share this frustration. Simply throwing a massive spec at an AI agent

    ways of working · Addy Osmani

  84. practice · How I AI ·

    Claude Code for product managers: research, writing, context libraries, custom to-do system, and more | Teresa Torres

    Product manager built a markdown task system in Claude Code with automated research collection and slash commands for daily workflows.

    Teresa Torres is the author of Continuous Discovery Habits and an internationally acclaimed speaker and coach. In this episode, Teresa demonstrates how she’s built a personalized productivity system using Claude Code to manage her tasks, automate research collection, and improve her writing. She shows how non-developers can leverage AI tools to create personalized workflows that match their unique needs and thinking style. What you’ll learn: How Teresa built a personalized task management system in Claude Code that matches her exact workflow needs Why she moved from Trello to a markdown-based

    ways of working · Claire Vo

  85. practice · Hamel Husain ·

    Why I Stopped Using nbdev

    Literate programming in notebooks conflicts with AI coding tools; developers should choose environments where AI assistance works best.

    Programmers love to proclaim they’ve found the best tool. Paul Graham called Lisp his “ secret weapon .” DHH described Ruby as “ a magical glove that just fit my brain perfectly .” Pieter Levels ships million-dollar products with vanilla PHP and jQuery . These declarations aren’t about the languages themselves. They’re about developers finding tools that fit how they think. When the environment clicks, you move fast. I had that experience with nbdev , a development environment for literate programming that I helped build and maintain 1 . I created hundreds of projects with it and was one of it

    ways of working · Hamel Husain

  86. practice · Latent Space ·

    Brex’s AI Hail Mary — With CTO James Reggio

    CTO describes how Brex built AI systems meeting financial compliance, auditability and customer trust requirements.

    From building internal AI labs to becoming CTO of Brex, James Reggio has helped lead one of the most disciplined AI transformations inside a real financial institution where compliance, auditability, and customer trust actually matter. From building internal AI labs to becoming CTO of Brex, James Reggio has helped lead one of the most disciplined AI transformations inside a real financial institution where compliance, auditability,…

    adoption · Shawn Wang (swyx)

  87. research · Indeed Hiring Lab ·

    AI Adoption Is Accelerating but Still Concentrated Among the Largest Firms

    AI hiring is concentrated at the largest firms, suggesting unequal access to AI productivity gains across company sizes.

    Nearly all AI-related hiring is happening at a few very large employers, highlighting the risks of unequal access to AI-driven productivity gains. The post AI Adoption Is Accelerating but Still Concentrated Among the Largest Firms appeared first on Indeed Hiring Lab .

    adoption

  88. practice · Artificial Ignorance ·

    The AI Manager's Schedule

    A developer shifted from IDE-based AI coding tools to headless CLI tools by changing trust and debugging practices.

    Shoutout to , , and for joining last week’s Office Hours chat ! Join the chat tomorrow for another Office Hours session. Recently, I was talking to a colleague about how dramatically my AI coding workflows have changed. A year ago, when the first CLI coding tools were released, I gave them a try. It left a strong impression - they were fun to use, and much more impactful than I thought they would be. But I never moved the majority of my development over to them. I was still using IDEs like Cursor, mainly because it was too hard to trust what the headless tool was doing, and debugging inevitabl

    ways of working · Charlie Guo

  89. research · Exponential View ·

    🔮 Anthropic's Head of Economics Peter McCrory on their new Economic Index

    Analysis of millions of real Claude conversations maps where AI is augmenting human work today and where it is not.

    The future will be... uneven. The future will be... uneven. Anthropic have just released a new Economic Index report. They’ve analysed millions of real Claude conversations to map exactly where AI is augmenting human work today, and where it isn’t. This is the best empirical window we have into how AI is reshaping work right now.

    adoption · Peter McCrory · Azeem Azhar

  90. research · arXiv ·

    Evolving with AI: A Longitudinal Analysis of Developer Logs

    Two-year telemetry from 800 developers shows AI coding assistants increase code output but also increase deletions, with perceived vs actual workflow changes diverging.

    AI-powered coding assistants are rapidly becoming fixtures in professional IDEs, yet their sustained influence on everyday development remains poorly understood. Prior research has focused on short-term use or self-reported perceptions, leaving open questions about how sustained AI use reshapes actual daily coding practices in the long term. We address this gap with a mixed-method study of AI adoption in IDEs, combining longitudinal two-year fine-grained telemetry from 800 developers with a survey of 62 professionals. We analyze five dimensions of workflow change: productivity, code quality, c

    adoption

  91. practice · Paul Ford (Aboard) ·

    Four AI Coding Horsemen of the Apocalypse

    Four structural risks when AI coding tools democratise software development: deployment complexity, testing gaps, security shortcuts, and maintenance debt.

    As anyone who reads this newsletter, listens to our podcast, or overhears me yelling at strangers on the train knows, I’m excited that tens of millions of people could suddenly become software developers thanks to Claude Code (and its inevitable competitors—there’s a rumored DeepSeek coding tool in the pipeline , which would be chef’s kiss ). But I also know that whenever I feel particularly hopeful about technology, I should grab a huge ice bucket and stick my head into it, because enthusiasm often doesn’t translate into social acceptance, utility, or positive outcomes. Having learned this le

    management org · Paul Ford

  92. practice · Hacker News ·

    Using proxies to hide secrets from Claude Code

    Using proxy patterns to prevent Claude Code from accessing sensitive secrets during development and testing.

    132 points on Hacker News. Discussion: https://news.ycombinator.com/item?id=46605155

    ways of working

  93. practice · How I AI ·

    The power user’s guide to Codex: parallelizing workflows, planning techniques, advanced context engineering tips, automating code reviews, and more | Alexander Embiricos

    Developers use Git worktrees and Plans.md to parallelize Codex workflows and manage context for complex projects.

    Alexander Embiricos , the product lead for Codex at OpenAI, shares practical workflows for getting the most out of this AI coding agent. In this episode, he demonstrates how both non-technical users and experienced engineers can leverage Codex to accelerate development, from making simple code changes to building production-ready applications. Alex walks through real examples of using Codex in VS Code and terminal environments, implementing parallel workflows with Git worktrees, and creating detailed implementation plans for complex projects. He also reveals how OpenAI uses Codex internally, i

    ways of working · Claire Vo

  94. research · Information Systems Research ·

    Workflow Automation in Open-Source Software Development: Accelerating Innovation Through Mechanization and Orchestration

    Mechanization automates tasks; orchestration automates coordination. Orchestration accelerates exploratory innovation while mechanization supports incremental improvement.

    This study develops a conceptual framework distinguishing two mechanisms of workflow automation: mechanization and orchestration. Mechanization automates discrete, self-contained, repeatable tasks through standardized execution to enhance consistency, reliability, and efficiency, while orchestration automates the communication between tasks, workers, and stages, which facilitates information flow and coordination. We theorize that these mechanisms differentially affect incremental versus substantive innovation. Using a multimethod approach integrating machine learning and econometrics, we anal

    adoption

  95. practice · Sangeet Paul Choudary ·

    US vs China: How to win the wrong AI race

    China's AI strategy focuses on converting intelligence into scaled execution and economic value, not competing on model capability.

    TL;DR: The US is betting on intelligence. China is betting somewhere else! The US frequently frames its competition with China as an AI race, similar to the space race it ran with the USSR. The idea of a race was triggered around a year ago with the launch of DeepSeek. Ever since, much of the US media has been fascinated with the idea of the US winning the AI race against China. Ironically, it isn’t much of a race if the US is the only one running it. The American ‘AI race’ is framed around who will build the smartest models, as if superior intelligence alone decides the future. China’s strate

    management org · Sangeet Paul Choudary

  96. research · Knowledge at Wharton ·

    Special Report: The AI Skills Gap

    New index from Wharton and Accenture measures which skills employers value and how fast labour market demand is shifting.

    A new index from Wharton and Accenture measures which skills matter to employers, which do not, and how quickly the economy is shifting beneath us. … Read More

    jobs skills

  97. research · JAMA Network Open ·

    Ambient Artificial Intelligence Scribes and Physician Financial Productivity

    Physicians using AI scribes increased revenue and patient volumes; claim denial rates fell compared to non-adopters.

    This cohort study evaluates physician revenue, patient volumes, and claim denials among adopters and nonadopters of artificial intelligence–based clinical documentation tools in one health system.

    productivity

  98. practice · One Useful Thing ·

    Claude Code and What Comes Next

    An AI agent independently built a working deployed website and sales funnel in 74 minutes with minimal human input.

    I opened Claude Code and gave it the command: “ Develop a web-based or software-based startup idea that will make me $1000 a month where you do all the work by generating the idea and implementing it. i shouldn’t have to do anything at all except run some program you give me once. it shouldn’t require any coding knowledge on my part, so make sure everything works well. ” The AI asked me three multiple choice questions and decided that I should be selling sets of 500 prompts for professional users for $39. Without any further input, it then worked independently… FOR AN HOUR AND FOURTEEN MINUTES

    ways of working · Ethan Mollick

  99. practice · Paul Ford (Aboard) ·

    Triple-A Vibe Coding

    A developer describes working with Claude Code by judging output without understanding the underlying code, building complex software iteratively on mobile.

    Welcome to 2026! It’s going to be an eventful year, and there’s nothing to be done about that. I know some of us were hoping AI would go away over the break, but it’s here to stay. I have spent the past couple of months going deep on Claude Code , and I want to share some very high-level observations. “Going deep” sounds intense, but actually involves typing a chat message into your phone and then coming back to it 20 minutes later to see if it worked, over and over again. You can finally write complex software on the bus. (And if Mayor Mamdani gets his way, that bus will be fast and free.) Ju

    ways of working · Paul Ford

  100. research · arXiv ·

    Artificial Intelligence and Skills: Evidence from Contrastive Learning in Online Job Vacancies

    AI adoption by Chinese firms expanded job skill requirements, driven by reduced information asymmetry and anticipation of future labour market needs.

    We investigate the impact of artificial intelligence (AI) adoption on skill requirements using 14 million online job vacancies from Chinese listed firms (2018-2022). Employing a novel Extreme Multi-Label Classification (XMLC) algorithm trained via contrastive learning and LLM-driven data augmentation, we map vacancy requirements to the ESCO framework. By benchmarking occupation-skill relationships against 2018 O*NET-ESCO mappings, we document a robust causal relationship between AI adoption and the expansion of skill portfolios. Our analysis identifies two distinct mechanisms. First, AI reduce

    jobs skills