The Feed · Complete archive

AI at work: research and practice

A server-rendered record of the evidence, ideas and firsthand practices screened by The Feed. Every entry links to its original source.

Use the interactive Feed

Page 8 of 9 · 873 items

  1. practice · Eugene Yan ·

    Evaluating Long-Context Question & Answer Systems

    Framework for evaluating long-context QA systems: metrics, dataset design, and methodology comparison.

    Evaluation metrics, how to build eval datasets, eval methodology, and a review of several benchmarks.

    judgment · Eugene Yan

  2. research · Government Information Quarterly ·

    Navigating power dynamics in the public sector through AI-driven algorithmic decision-making

    AI adoption in public sector shifts power toward hybrid analysts with technical and institutional expertise, creating tensions among managers.

    Public sector institutions are under increasing pressure to deliver greater public value through disruptive technologies, despite ongoing pressures. In response to evolving technological change and an abundance of information, many public sector organisations have adopted Artificial Intelligence (AI) to improve decision-making and generate social value. While AI's role in public administration is gaining attention, little is known about how its use alters internal power dynamics. This research uses a qualitative case study approach, drawing on 30 semi-structured interviews with operational man

    management org

  3. practice · Artificial Ignorance ·

    How We Use AI At Pulley

    A startup engineering team shares a year of daily AI coding practices and accountability norms they've built around it.

    Lessons from a year of daily AI coding at a fast-growing startup (and why "Cursor wrote it" is not a valid excuse). Lessons from a year of daily AI coding at a fast-growing startup (and why "Cursor wrote it" is not a valid excuse). There's no shortage of predictions about AI's impact on software development. But while everyone is debating whether AI will replace programmers, my teammates at Pulley and I are using it daily, and …

    ways of working · Charlie Guo

  4. practice · How I AI ·

    A designer's guide to Cursor: How to build interactive prototypes with sound, explore visual styles, and transform data visualizations | Elizabeth Lin

    Designers use Cursor with strategic prompting to explore visual aesthetics, build interactive prototypes with sound, and polish interfaces without coding.

    Elizabeth Lin is an independent design educator who has crafted learning experiences for Khan Academy, Primer, and Lambda School. She currently runs design is a party , an alternative online design school where she teaches courses like The Art of Visual Design and Prototyping with Cursor . In this episode, she shares how designers can leverage Cursor to create interactive prototypes with sound, explore different visual aesthetics, and transform basic designs into polished interfaces—all without deep coding knowledge. What you'll learn: How to use Cursor to explore different design aesthetics—f

    ways of working · Claire Vo

  5. research · Journal of Applied Psychology ·

    How and for whom using generative AI affects creativity: A field experiment.

    Field experiment shows LLM assistance increases employee creativity, especially for those with strong metacognitive strategies.

    We develop a theoretical perspective on how and for whom large language model (LLM) assistance influences creativity in the workplace. We propose that LLM assistance increases employees' creativity by providing cognitive job resources. Furthermore, we hypothesize that employees with high levels of metacognitive strategies-who actively monitor and regulate their thinking to achieve goals and solve problems-are more likely to leverage LLM assistance effectively to acquire cognitive job resources, thereby increasing creativity. Our hypotheses were supported by a field experiment, in which we rand

    productivity

  6. research · Proceedings of the National Academy of Sciences ·

    Human–AI collectives most accurately diagnose clinical vignettes

    Physician-AI hybrid collectives outperform solo physicians, physician teams, and AI systems alone on medical diagnostics across 2,133 cases.

    AI systems, particularly large language models (LLMs), are increasingly being employed in high-stakes decisions that impact both individuals and society at large, often without adequate safeguards to ensure safety, quality, and equity. Yet LLMs hallucinate, lack common sense, and are biased-shortcomings that may reflect LLMs' inherent limitations and thus may not be remedied by more sophisticated architectures, more data, or more human feedback. Relying solely on LLMs for complex, high-stakes decisions is therefore problematic. Here, we present a hybrid collective intelligence system that miti

    judgment

  7. research · Management Science ·

    AI, Skill, and Productivity: The Case of Taxi Drivers

    AI route-suggestion tool increased taxi driver productivity mainly for low-skilled drivers, narrowing the 13.4% skill productivity gap.

    We examine the impact of artificial intelligence (AI) on productivity in the context of taxi drivers. The AI we study assists drivers with finding customers by suggesting routes along which the demand is predicted to be high. We find that AI improves drivers’ productivity by shortening the cruising time, and this gain is accrued only to low-skilled drivers, narrowing the productivity gap between high- and low-skilled drivers by 13.4%. This case study provides evidence that AI and skill are indeed substitutes, offering direct support for the underlying assumption of recent projection exercises

    productivity · Joshua Gans

  8. research · MIS Quarterly ·

    Organizing for AI Innovation: Insights From an Empirical Exploration of U.S. Patents

    AI innovations are less radical and more process-oriented than IT innovations, requiring different organizational management approaches.

    Although the prevalence of artificial intelligence (AI) innovations is on the rise, firms frequently report failures and setbacks in their development and implementation of AI innovation efforts. One common issue behind many failing AI initiatives is that they are organized just like other information technology (IT) innovation efforts. To elucidate why and how the production of AI and IT innovations may need to be managed differently, this study juxtaposes these two types of innovations based on two key dimensions of the Schumpeterian framework: the form (product vs. process) and magnitude (r

    adoption

  9. research · Management Science ·

    Human-Centered Artificial Intelligence: A Field Experiment

    Field experiment shows tailoring AI interaction to individuals' cognitive styles improves sales performance; misaligned interaction harms performance versus control.

    Humans and artificial intelligence (AI) algorithms increasingly interact on unstructured managerial tasks. We propose that tailoring this human-AI interaction to align with individuals’ cognitive preferences is essential for enhancing performance. This hypothesis is examined through a field experiment in a multinational pharmaceutical firm. In the experiment, we manipulated four contextual parameters of human-AI interaction—work procedures, decision-making authority, training, and incentives—to align with sales experts’ cognitive styles, categorized as either adaptors or innovators. Our result

    adoption · Catherine Tucker · Sebastian Raisch

  10. research · arXiv ·

    Future of Work with AI Agents: Auditing Automation and Augmentation Potential across the U.S. Workforce

    Survey of 1,500 workers across 104 occupations on whether they want AI agents to automate or augment specific tasks, compared against expert assessments of technical feasibility.

    The rapid rise of compound AI systems (a.k.a., AI agents) is reshaping the labor market, raising concerns about job displacement, diminished human agency, and overreliance on automation. Yet, we lack a systematic understanding of the evolving landscape. In this paper, we address this gap by introducing a novel auditing framework to assess which occupational tasks workers want AI agents to automate or augment, and how those desires align with the current technological capabilities. Our framework features an audio-enhanced mini-interview to capture nuanced worker desires and introduces the Human

    adoption · Erik Brynjolfsson · Diyi Yang

  11. research · JAMA Network Open ·

    Efficiency and Quality of Generative AI–Assisted Radiograph Reporting

    Radiologists using AI-assisted draft reports completed documentation 22% faster with no loss of accuracy or quality in a prospective clinical study.

    Importance: Diagnostic imaging interpretation involves distilling multimodal clinical information into text form, a task well-suited to augmentation by generative artificial intelligence (AI). However, to our knowledge, impacts of AI-based draft radiological reporting remain unstudied in clinical settings. Objective: To prospectively evaluate the association of radiologist use of a workflow-integrated generative model capable of providing draft radiological reports for plain radiographs across a tertiary health care system with documentation efficiency, the clinical accuracy and textual qualit

    productivity

  12. practice · Paul Ford (Aboard) ·

    The Extremely Human Last Mile

    AI automation hits limits when human judgment and relationship-building matter; the last mile of work remains stubbornly human.

    Two contradictory AI things are rattling around in my head, and I’m trying to put them together and make sense of where we are. The first is Builder.ai declaring bankruptcy after failing to meet revenue targets. The second is a set of statements by Anthropic founder Dario Amodei predicting mass AI-driven unemployment. Builder.ai is (was?) a little hard to describe, but their basic pitch was something like: Come to our website with your app idea and answer a bunch of AI-prompted questions, and put everything you want your app to do into a shopping cart. We’ll use our AI bot Natasha to do a lot

    judgment · Paul Ford

  13. practice · Artificial Ignorance ·

    How To Fall Behind in The Age of AI

    A framework for when and how to deliberately opt out of AI adoption without losing competitiveness.

    A very serious guide to artificial ignorance. A very serious guide to artificial ignorance. Look, I get it. Every day, dozens of headlines proclaim that AI is the next big thing, or is coming for our jobs, or might end human civilization as we know it. It's exhausting. And honestly, who has…

    adoption · Charlie Guo

  14. research · arXiv ·

    My Advisor, Her AI and Me: Evidence from a Field Experiment on Human-AI Collaboration and Investment Decisions

    Field experiment: customers follow human-AI collaborative investment advice more than pure AI advice, improving welfare, driven by human credibility rather than quality gains.

    Amid ongoing policy and managerial debates on keeping humans in the loop of AI decision-making, we investigate whether human involvement in AI-based service production benefits downstream consumers. Partnering with a large savings bank in Europe, we produced pure AI and human-AI collaborative investment advice, passed it to customers, and examined their advice-taking in a field experiment. On the production side, contrary to concerns that humans might inefficiently override AI output, we find that giving a human banker the final say over AI-generated financial advice does not compromise its qu

    judgment

  15. practice · Eugene Yan ·

    AI Engineer 2025 - Improving RecSys & Search with LLM techniques

    Semantic IDs and data augmentation techniques unify recommendation and search systems using LLMs.

    Recsys & search are converging with LLMs via semantic IDs, data augmentation, and unified foundation models.

    ways of working · Eugene Yan

  16. practice · How I AI ·

    The exact AI playbook (using MCPs, custom GPTs, Granola) that saved ElevenLabs $100k+ and helps them ship daily | Luke Harries (Head of Growth)

    Marketing team automated case studies, translations and WhatsApp workflows using MCPs and custom GPTs, cutting costs by $140k annually.

    Luke Harries , Head of Growth at ElevenLabs, the leading AI voice technology company, shares how he’s automating marketing workflows with AI—from case studies to translations to WhatsApp integrations—saving his company over $140,000 while making everything a launch. What you’ll learn: 1. How to create polished case studies in minutes using AI transcription and a custom GPT 2. How ElevenLabs built a custom AI translation system that saved them $140,000 annually and eliminated agency headaches 3. How to use Model Context Protocols (MCPs) to connect AI assistants to your WhatsApp messages 4. The

    productivity · Claire Vo

  17. research · NBER ·

    Expertise

    Task automation changes expertise requirements for remaining work, with effects varying by occupation and task complementarity.

    When job tasks are automated, does this augment or diminish the value of labor in the tasks that remain? We argue the answer depends on whether removing tasks raises or reduces the expertise required for remaining non-automated tasks. Since the same task may be relatively expert in one occupation (David Autor , Neil Thompson)

    jobs skills · David Autor

  18. research · Management Science ·

    Incentives, Framing, and Reliance on Algorithmic Advice: An Experimental Study

    Performance incentives and tournament pay increase manager reliance on AI advice; framing algorithms as incorporating human expertise also boosts utilization.

    Managerial decision makers are increasingly supported by advanced data analytics and other artificial intelligence (AI)-based technologies, but they are often found to be hesitant to follow the algorithmic advice. We examine how compensation contract design and framing of an AI algorithm influence decision makers’ reliance on algorithmic advice and performance in a price estimation task. Based on a large sample of almost 1,500 participants, we find that compared with a fixed compensation, both compensation contracts based on individual performance and tournament contracts lead to an increase i

    judgment

  19. practice · Paul Ford (Aboard) ·

    Will AI Eat Tech, or Will Tech Eat AI?

    AI companies pursuing platform strategy (OS, ecosystem, developer tools) mirrors how Windows, Google, Apple dominated; MCP offers alternative distributed model.

    There are many big questions about AI, but as a tech industry person, the one that is most interesting to me is “will AI eat the tech industry, or will the tech industry eat AI?” What would it mean if AI ate the tech industry? OpenAI or Anthropic wouldn’t just be providers of chatbots or video generators—they’d become operating systems and ecosystems, the primary interface for accessing all information. Microsoft did that with Windows, Google with search (and advertising, and Google Drive, and Android), and Apple with the iPhone. Some people really do spend hours a day on ChatGPT already, and

    management org · Paul Ford

  20. practice · The Pragmatic Engineer ·

    From Software Engineer to AI Engineer – with Janvi Kalra

    A software engineer's self-taught path to AI engineering roles, with concrete advice for career transitions.

    From Coda to OpenAI: How Janvi Kalra taught herself AI engineering, impressed tech leaders, and built a career at the forefront of AI—plus actionable advice for landing your own role. From Coda to OpenAI: How Janvi Kalra taught herself AI engineering, impressed top tech leaders, and built a career at the forefront of responsible AI—plus actionable advice for landing your own AI role. Stream the Latest Episode

    jobs skills · Gergely Orosz

  21. practice · Hacker News ·

    LLM codegen go brrr – Parallelization with Git worktrees and tmux

    Using Git worktrees and tmux to run multiple LLM code generation tasks in parallel, speeding up development.

    156 points on Hacker News. Discussion: https://news.ycombinator.com/item?id=44116872

    ways of working

  22. research · Information Systems Research ·

    Forced to Change? Media Exposure of Labor Issues and Firm Artificial Intelligence Investment

    Media coverage of labor issues increases AI investment in firms, particularly targeting high-skilled roles, suggesting AI adoption driven by reputational pressure rather than productivity.

    Firms are increasingly interested in investing in artificial intelligence (AI), but what drives this trend? Our research reveals that media coverage of labor issues plays a significant role. When firms face public scrutiny through media exposure of labor issues, the reputational pressure pushes them to act. AI emerges as a strategic response, offering powerful capabilities to automate and augment human tasks. Analyzing data from U.S. public firms, we found that labor issue-related media coverage significantly increases AI investments, particularly among firms with the motivation and resources

    adoption

  23. practice · How I AI ·

    A 3-step AI coding workflow for solo founders | Ryan Carson (5x founder)

    Solo founder shares a three-file system and structured Cursor workflow that replaces traditional engineering teams for product building.

    Ryan Carson is a five-time founder who has spent the past 20 years building, scaling, and selling startups. In this episode, he shares his playbook for using AI to build products, turning “vibe coding” into a structured and scalable approach that can replace full engineering teams. What you’ll learn: 1. A simple three-file system that transforms chaotic AI coding into a structured, reliable process 2. How to create AI-generated PRDs and task lists that actually work 3. A step-by-step workflow using Cursor to build features systematically 4. Why slowing down to provide proper context is the sec

    ways of working · Claire Vo

  24. research · Management Science ·

    Humans’ Use of AI Assistance: The Effect of Loss Aversion on Willingness to Delegate Decisions

    Loss framing removes algorithm aversion: workers delegate equally to AI and humans when losses frame decisions, but prefer humans under gain framing.

    As artificial intelligence (AI) tools have become pervasive in business applications, so too have interactions between AI and humans in business processes and decision-making. A growing area of research has focused on human decision and task delegation to AI assistants. Simultaneously, extensive research on algorithm aversion—humans’ resistance to algorithm-based decision tools—has demonstrated potential barriers and issues with AI applications in business. In this paper, we test a simple strategy for mitigating algorithm aversion in the context of AI task delegation. We show that simply chang

    judgment

  25. practice · Simon Wardley ·

    Rewilding software engineering

    Software engineers and business leaders need shared language about system design to work well with AI tools.

    Chapter 5: Different folks for different strokes By Tudor Girba and Simon Wardley An apology This book is about software engineering, but we start this chapter by talking to the business. We know that you’ve been wanting to fire those pedantic, inflexible, costly and annoying software engineers for ages. AI appears to be so damn helpful and so much more of a partner that you think that now is the right time to part ways. But give them a second chance, you will need them, software engineers can change and so can you. We understand that in business you don’t care about the decisions made in code

    management org · Simon Wardley

  26. research · Management Science ·

    Improving Human Sequential Decision Making with Reinforcement Learning

    Reinforcement learning algorithm extracts interpretable decision tips from worker data; randomized experiments show tips improve sequential decision-making performance.

    Workers spend a significant amount of time learning how to make good decisions. Evaluating the efficacy of a given decision, however, can be complicated—for example, decision outcomes are often long-term and relate to the original decision in complex ways. Surprisingly, even though learning good decision-making strategies is difficult, the strategies can often be expressed in simple and concise forms. Focusing on sequential decision making, we design a novel machine learning algorithm that is capable of extracting “best practices” from trace data and conveying its insights to humans in the for

    judgment

  27. practice · The Pragmatic Engineer ·

    The AI Engineering Stack

    AI engineering as a distinct discipline with three layers: interfaces, orchestration, and optimization, different from ML and fullstack engineering.

    Three layers of the AI stack, how AI engineering is different from ML engineering and fullstack engineering, and more. An excerpt from the book AI Engineering by Chip Huyen Three layers of the AI stack, how AI engineering is different from ML engineering and fullstack engineering, and more. An excerpt from the book AI Engineering by Chip Huyen “AI Engineering” is a term that I didn’t hear about two years ago, but today, AI engineers are in high demand. Companies like Meta, Google, and Amazon, offer higher base salaries for these roles than “regular” software engineers get, while AI startups an

    management org · Gergely Orosz

  28. practice · Understanding AI (Timothy B. Lee) ·

    I got fooled by AI-for-science hype—here's what it taught me

    A physicist applied AI to plasma research and found it failed to deliver on hype; detailed account of why and what actually worked.

    I used AI in my plasma physics research and it didn’t go the way I expected. I used AI in my plasma physics research and it didn’t go the way I expected. I’m excited to publish this guest post by Nick McGreivy, a physicist who last year earned a PhD from Princeton. Nick used to be optimistic that AI could accelerate physics research. But when he tried to apply AI techniques to real physics problems the results were disappointing.

    judgment

  29. practice · Latent Space ·

    ChatGPT Codex: The Missing Manual

    WHAM framework for using ChatGPT Codex, an autonomous software engineer agent, on existing codebases.

    ChatGPT Codex is here - the first cloud hosted Autonomous Software Engineer (A-SWE) from OpenAI. Josh Ma and Alexander Embiricos tell us how to WHAM every codebase like a power user. ChatGPT Codex is here - the first cloud hosted Autonomous Software Engineer (A-SWE) from OpenAI. Josh Ma and Alexander Embiricos tell us how to WHAM every codebase like a power user. The World’s Fair is 2 weeks away, and early bird tix have sold out! We’re happy to share that Fouad Matin of the OpenAI Codex team, Michael Truell of Cursor, Kevin Hou of Windsurf, Boris Cherny of Cl…

    ways of working · Shawn Wang (swyx)

  30. practice · Artificial Ignorance ·

    AI's Missing Multiplayer Mode

    Framework for moving from using AI as a tool to treating it as a collaborative teammate in workflows.

    Going from digital tools to digital teammates. Going from digital tools to digital teammates. When we look at the explosive growth of AI over the past few years, it's easy to be awestruck by the pace of innovation (and indeed, I often am). ChatGPT, Claude, and their increasingly capable cousi…

    ways of working · Charlie Guo

  31. research · arXiv ·

    Precision Proactivity: Measuring Cognitive Load in Real-World AI-Assisted Work

    Financial professionals using GPT-4o show higher quality work with AI assistance, but extraneous cognitive load from AI causes three times larger performance declines than task complexity.

    Systems like ChatGPT and Claude assist billions through proactive dialogue-offering unsolicited, task-relevant information. Drawing on Cognitive Load Theory, we study how cognitive load shapes performance in AI-assisted knowledge work. We recruited 34 financial professionals to complete a complex valuation task using GPT-4o and developed a transcript-based framework estimating intrinsic and extraneous load from computational indicators anchored in a task decomposition and knowledge graph. Across 1,178 participant-subtask observations, AI-generated content usage is positively associated with qu

    productivity · Matt Beane

  32. practice · Paul Ford (Aboard) ·

    The Funnel of the Future is Upside Down

    AI-powered sales outreach is inverting the funnel: companies now need to build trust and deliver value before asking for attention.

    You know what the marketing funnel is, right? You run ads on Google or Facebook so people learn about your brand or product, or you buy a list of names and send them cold emails. That’s “top of funnel.” Some fish nibble: They fill out a form, you schedule a call, and they don’t show up to the call. You send them a follow-up email saying you hope to talk soon anyway, because you are a worm with no pride. And so on, down the funnel, until eventually very few people are left, like at the end of a bad party, and you get one of them to sign on the line that is dotted before the lights come back on.

    management org · Paul Ford

  33. practice · Sangeet Paul Choudary ·

    The fugu guide to jobs in a world of AI

    AI doesn't just automate tasks; it breaks the systems and credentials that gatekeep work, forcing role redesign.

    Stop measuring AI by the tasks it fails at. Start noticing the systems it breaks. Stop measuring AI by the tasks it fails at. Start noticing the systems it breaks. In Japan, a licensed fugu chef occupies a unique position in the food economy.

    jobs skills · Sangeet Paul Choudary

  34. practice · Harper Reed ·

    Basic Claude Code

    Developer workflow using reasoning models to generate specs and prompts, then Claude Code for implementation.

    I really like this agentic coding thing. It is quite compelling in so many ways. Since I wrote that original blog post a lot has happened in Claude land: Claude Code MCP etc I have received hundreds (wat) of emails from people talking about their workflows and how they have used my workflow to get ahead. I have spoken at a few conferences, and taught a few classes about codegen. I have learned that computers really want to spellcheck codegen to codeine, who knew! I was talking to a friend the other day about how we are all totally fucked and AI will take our jobs (more on that in a later post)

    ways of working · Harper Reed

  35. practice · Paul Ford (Aboard) ·

    Is “Specification Repair” the AI Endgame?

    AI models use iterative code execution and self-prompting to refine reasoning, feeding output back as new input until solving problems.

    Recently, I had a nice solid glimpse of the future and I wanted to tell you about it. Not long ago, I visited an office with a beautiful view, and I took this photo: She seems nice. That day, people online had been discussing how good ChatGPT was at guessing where a picture was taken , so I fed it this image and asked it not just to figure out what building, but which floor. It identified the exact building ( 17 State Street ), and then suggested it was the 34th floor (it was the 30th). This was, I thought, surprisingly accurate. This is another dangerous superpower we’re not prepared to handl

    ways of working · Paul Ford

  36. research · Management Science ·

    Algorithm Reliance: Fast and Slow

    Lab experiment shows workers rely on superior algorithms more under high load, improving quality and speed, but fail to speed up with inferior algorithms despite potential gains.

    In algorithm-augmented service contexts where workers have decision authority, they face two decisions about the algorithm: whether to follow its advice and how quickly to do so. The pressure to work quickly increases with the speed of arriving customers. In this paper, we ask the following. How do workers use algorithms to manage system loads? With a laboratory experiment, we find that superior algorithm quality and high system loads increase participants’ willingness to use their algorithm’s advice. Consequently, participants with the superior algorithm make higher-quality recommendations th

    judgment

  37. research · arXiv ·

    Regulating Algorithmic Management: A Multi-Stakeholder Study of Challenges in Aligning Software and the Law for Workplace Scheduling

    Multi-stakeholder interviews reveal how regulatory gaps emerge when scheduling software, law, and workplace practices misalign in practice.

    Algorithmic management (AM)'s impact on worker well-being has led to calls for regulation. However, little is known about the effectiveness and challenges in real-world AM regulation across the regulatory process -- rule operationalization, software use, and enforcement. Our multi-stakeholder study addresses this gap within workplace scheduling, one of the few AM domains with implemented regulations. We interviewed 38 stakeholders across the regulatory process: regulators, defense attorneys, worker advocates, managers, and workers. Our findings suggest that the efficacy of AM regulation is inf

    policy

  38. research · Management Science ·

    Leveraging Expert Consistency to Improve Algorithmic Decision Support

    Machine learning models for decision support perform better when trained on expert decisions for consistent cases and outcomes for inconsistent ones.

    Machine learning (ML) is increasingly being used to support high-stakes decisions. However, there is frequently a construct gap: a gap between the construct of interest to the decision-making task and what is captured in proxies used as labels to train ML models. As a result, ML models may fail to capture important dimensions of decision criteria, hampering their utility for decision support. Thus, an essential step in the design of ML systems for decision support is selecting a target label among available proxies. In this work, we explore the use of historical expert decisions as a rich—yet

    judgment

  39. practice · Eugene Yan ·

    Building News Agents for Daily News Recaps with MCP, Q, and tmux

    Building a news recap agent using MCP and Amazon Q CLI to automate daily information gathering.

    Learning to automate simple agentic workflows with Amazon Q CLI, Anthropic MCP, and tmux.

    ways of working · Eugene Yan

  40. research · JAMA Network Open ·

    Evaluation of an Ambient Artificial Intelligence Documentation Platform for Clinicians

    Ambient AI documentation reduced clinician time on notes per appointment and off-hours EHR work, with mixed effects on burnout.

    Importance: The increase of electronic health record (EHR) work negatively impacts clinician well-being. One potential solution is incorporating an ambient artificial intelligence (AI) documentation platform. Objective: To understand clinician experience before and after implementing ambient AI. Design, Setting, and Participants: This quality improvement study was a pilot evaluation with before and after survey and EHR metrics conducted at a large health care organization in Northern and Central California. Clinicians were purposively sampled to be representative of region and specialty. Ambie

    productivity

  41. research · Proceedings of the ACM on Human-Computer Interaction ·

    EARN Fairness: Explaining, Asking, Reviewing, and Negotiating Artificial Intelligence Fairness Metrics Among Stakeholders

    A framework and interactive system for non-technical stakeholders to collectively decide on AI fairness metrics through explanation, preference elicitation, review, and negotiation.

    Numerous fairness metrics have been proposed and employed by artificial intelligence (AI) experts to quantitatively measure bias and define fairness in AI models. Recognizing the need to accommodate stakeholders' diverse fairness understandings, efforts are underway to solicit their input. However, conveying AI fairness metrics to stakeholders without AI expertise, capturing their personal preferences, and seeking a collective consensus remain challenging and underexplored. To bridge this gap, we propose a new framework, EARN ( Explain, Ask, Review, and Negotiate ) Fairness, which facilitates

    judgment

  42. research · Proceedings of the ACM on Human-Computer Interaction ·

    Chatbots in Collaborative Settings and their Impact on Virtual Teamwork

    Group chat AI assistance improved team performance and reduced information-request response times; private chat increased perceived effort.

    Chatbots have emerged as a powerful tool for collaborative teamwork. In this paper, we investigate the influence of chatbots on virtual teamwork within the context of a collaborative online activity. To assess chatbots' impact on group dynamics and performance, we designed a novel collaborative activity with an associated online platform and a custom chatbot assistant. We recruited 72 participants divided into four-person teams, with the teams split into four conditions depending on the chatbot's assistance strategy (none, by private chat, by group chat, or both). We found that chatbot assista

    teams

  43. research · Proceedings of the ACM on Human-Computer Interaction ·

    Navigating Automated Hiring: Perceptions, Strategy Use, and Outcomes Among Young Job Seekers

    Young job seekers distrust automated hiring tools and rely on referrals over practice strategies, while referrals and family income predict success more than AEDT gaming.

    As the use of automated employment decision tools (AEDTs) has rapidly increased in hiring contexts, especially for computing jobs, there is still limited work on applicants' perceptions of these emerging tools and their experiences navigating them. To investigate, we conducted a survey with 448 computer science students (young, current technology job-seekers) about perceptions of the procedural fairness of AEDTs, their willingness to be evaluated by different AEDTs, the strategies they use relating to automation in the hiring process, and their job seeking success. We find that young job seeke

    jobs skills

  44. research · Proceedings of the ACM on Human-Computer Interaction ·

    'Always Nice and Confident, Sometimes Wrong': Developer's Experiences Engaging Generative AI Chatbots Versus Human-Powered Q&A Platforms

    Developers prefer ChatGPT's tone and speed over Stack Overflow but distrust its confident errors; lack voting validation reduces reliability signals.

    Software engineers have historically relied on human-powered Q&A platforms like Stack Overflow (SO) as coding aids. With the rise of generative AI, developers have started to adopt AI chatbots, such as ChatGPT, in their software development process. Recognizing the potential parallels between human-powered Q&A platforms and AI-powered question-based chatbots, we investigate and compare how developers integrate this assistance into their real-world coding experiences by conducting a thematic analysis of 1700+ Reddit posts. Through a comparative study of SO and ChatGPT, we identified each platfo

    adoption

  45. research · Proceedings of the ACM on Human-Computer Interaction ·

    Making ChatGPT Work for Me

    Teachers use ChatGPT in four distinct modes: make, find, jump-start, and iterate, with make-for-me dominating at 55 percent of prompts.

    Increasingly, work happens through human collaboration with generative AI (e.g., ChatGPT). In this paper, we present a qualitative study of this collaboration for real-life work tasks. We focus our study on US K12 public school teachers (N = 24) who regularly design and complete text-generation tasks such as creating quizzes, slide decks, word problems, reading passages, lesson plans, classroom activities, and projects. In one-on-one video- and audio-recorded virtual sessions, we observe each teacher using ChatGPT-4 for work tasks of their choosing for 15 minutes, then debrief their experience

    adoption

  46. research · Proceedings of the ACM on Human-Computer Interaction ·

    AURA: Amplifying Understanding, Resilience, and Awareness for Responsible AI Content Work

    Survey and interviews with content moderation and data labeling workers reveal challenges and needs in responsible AI work, proposing a support framework.

    Behind the scenes of maintaining the safety of technology products from harmful and illegal digital content lies unrecognized human labor. The recent rise in the use of generative AI technologies and the accelerating demands to meet responsible AI (RAI) aims necessitates an increased focus on the labor behind such efforts in the age of AI. This study investigates the nature and challenges of content work that supports RAI efforts, or "RAI content work," that spans content moderation, data labeling, and red teaming -- through the lived experiences of content workers. We conduct a formative surv

    worker experience · Jina Suh · Mary Gray

  47. research · Proceedings of the ACM on Human-Computer Interaction ·

    Human Delegation Behavior in Human-AI Collaboration: The Effect of Contextual Information

    Contextual information about AI capabilities and task domain significantly improves human-AI team performance in delegation decisions.

    The integration of artificial intelligence (AI) into human decision-making processes at the workplace presents both opportunities and challenges. One promising approach to leverage existing complementary capabilities is allowing humans to delegate individual instances of decision tasks to AI. However, enabling humans to delegate instances effectively requires them to assess several factors. One key factor is the analysis of both their own capabilities and those of the AI in the context of the given task. In this work, we conduct a behavioral study to explore the effects of providing contextual

    judgment

  48. research · Proceedings of the ACM on Human-Computer Interaction ·

    Rideshare Transparency: Translating Gig Worker Insights on AI Platform Design to Policy

    Analysis of over 1 million gig worker comments reveals transparency gaps in rideshare algorithms affecting driver earnings, route decisions, and task allocation.

    Rideshare platforms exert significant control over workers through algorithmic systems that can result in financial, emotional, and physical harm. What steps can platforms, designers, and practitioners take to mitigate these negative impacts and meet worker needs? In this paper, we identify transparency-related harms, mitigation strategies, and worker needs while validating and contextualizing our findings within the broader worker community. We use a novel mixed-methods study combining an LLM-based analysis of over 1 million comments posted to online platform worker communities with semi-stru

    worker experience

  49. research · Proceedings of the ACM on Human-Computer Interaction ·

    Secret Use of Large Language Model (LLM)

    Survey and experiment show workers hide LLM use from employers and colleagues; task type drives secretive behavior through fear of judgment.

    The advancements of Large Language Models (LLMs) have decentralized the responsibility for the transparency of AI usage. Specifically, LLM users are now encouraged or required to disclose the use of LLM-generated content for varied types of real-world tasks. However, an emerging phenomenon, users' secret use of LLM , raises challenges in ensuring end users adhere to the transparency requirement. Our study used mixed-methods with an exploratory survey (125 real-world secret use cases reported) and a controlled experiment among 300 users to investigate the contexts and causes behind the secret u

    worker experience

  50. research · NBER ·

    Still Waters, Rapid Currents: Early Labor Market Transformation under Generative AI

    Survey data linked to Danish administrative records show rapid chatbot adoption by employers, worker-reported productivity gains, and early labour market shifts across occupations.

    We study the early labor market impacts of AI chatbots by linking large-scale adoption surveys to administrative labor market records in Denmark. We document rapid currents: most employers in exposed occupations have adopted chatbot initiatives, workers report productivity benefits, and new AI (Anders Humlum , Emilie Vestergaard)

    adoption · Anders Humlum

  51. research · NBER ·

    Shifting Work Patterns with Generative AI

    Field experiment across 66 firms and 7,137 workers shows how generative AI integrated into daily applications changes work patterns over six months.

    We present evidence from a field experiment across 66 firms and 7,137 knowledge workers. Workers were randomly selected to access a generative AI tool integrated into applications they already used at work for email, meetings, and writing. In the second half of the 6-month experiment, the 80% of (Eleanor W. Dillon , Sonia Jaffe , Nicole Immorlica , Christopher T. Stanton)

    adoption · Sonia Jaffe

  52. practice · Latent Space ·

    Please stop forcing Clippy on those who want Anton

    AI assistants should adapt helpfulness levels to user skill: beginners need Clippy-like guidance, experts need minimal intervention.

    ChatGPT-4o's glazing embarrassment lays open Clippy vs Anton: The two extremes of desires in AI post-training and product ChatGPT-4o's glazing embarrassment lays open Clippy vs Anton: The two extremes of desires in AI post-training and product Update: Steven Sinofsky agrees: 1) Schticks get tired, 2) Helpfulness gets less helpful as skill level improves. Society advances by the number of operations you can do without thinking, not by the n…

    judgment · Shawn Wang (swyx)

  53. research · npj Digital Medicine ·

    Opportunities and risks of artificial intelligence in patient portal messaging in primary care

    Doctors failed to catch 35-45% of errors in AI-drafted patient messages, despite perceiving AI as safe and workload-reducing.

    The rapid increase in patient portal messaging has heightened the workload for primary care physicians (PCPs), contributing to burnout. The use of generative artificial intelligence (AI) to draft responses to patient messages has shown promise in reducing cognitive burden, yet there is still much unknown about the safety and perceptions of using AI drafts. This cross-sectional simulation study assessed whether PCPs could identify and correct errors in AI-generated draft responses to patient portal messages. Twenty practicing PCPs reviewed 18 patient portal messages, four of which contained err

    judgment

  54. research · Government Information Quarterly ·

    AI adoption in public administration: Perspectives of public sector managers and public sector non-managerial employees

    Public sector managers and employees report different perspectives on AI adoption barriers and enablers in government agencies.

    adoption

  55. practice · Paul Ford (Aboard) ·

    Using AI to Redirect Yourself—and Save Money!

    Using AI chatbots as a redirection tool to interrupt impulse shopping by reframing desire into creative constraint.

    Right now, Shopify and ChatGPT are starting to work together so that you can buy stuff from right inside of an AI chat. Is this a good idea? Stop asking and get your wallet out! But I want to throw out a way that I’ve been using ChatGPT not to buy things. Do you know about redirection? Redirection is an important concept in dealing with toddlers and middle-aged adults—there’s a valuable video to watch from Head Start (while we still have Head Start) if you’re curious. The whole idea is pretty simple: Intervene before the bad behavior starts. It’s not just saying, “Don’t do that.” It’s more lik

    ways of working · Paul Ford

  56. research · Organizational Behavior and Human Decision Processes ·

    The transparency dilemma: How AI disclosure erodes trust

    Thirteen experiments show people trust colleagues less when they disclose AI use, even with mandatory disclosure or when AI accuracy is high.

    As generative artificial intelligence (AI) has found its way into various work tasks, questions about whether its usage should be disclosed and the consequences of such disclosure have taken center stage in public and academic discourse on digital transparency. This article addresses this debate by asking: Does disclosing the usage of AI compromise trust in the user? We examine the impact of AI disclosure on trust across diverse tasks—from communications via analytics to artistry—and across individual actors such as supervisors, subordinates, professors, analysts, and creatives, as well as acr

    judgment

  57. research · Journal of Empirical Legal Studies ·

    Hallucination‐Free? Assessing the Reliability of Leading AI Legal Research Tools

    Legal AI research tools by LexisNexis and Thomson Reuters hallucinate 17-33% of the time despite vendor claims of elimination.

    ABSTRACT Legal practice has witnessed a sharp rise in products incorporating artificial intelligence (AI). Such tools are designed to assist with a wide range of core legal tasks, from search and summarization of caselaw to document drafting. However, the large language models used in these tools are prone to “hallucinate,” or make up false information, making their use risky in high‐stakes domains. Recently, certain legal research providers have touted methods such as retrieval‐augmented generation (RAG) as “eliminating” or “avoid[ing]” hallucinations, or guaranteeing “hallucination‐free” leg

    judgment

  58. practice · How I AI ·

    Gumroad CEO's playbook to 40x his team's productivity with v0, Cursor, and Devin | Sahil Lavingia

    Gumroad achieves 40x faster feature delivery by routing work through v0, Cursor, and Devin agents, shifting engineers to architecture roles.

    Sahil Lavingia is the founder and CEO of Gumroad, where AI agents are already writing 41% of all code commits, and he’s targeting 80% by year’s end. Sahil demonstrates how this approach allows him to transform what would typically be two-week projects into two-hour implementations—a 40x productivity increase. What you’ll learn: The exact AI workflow Sahil uses to build features 40x faster—from prototyping in v0 to implementation with Devin How Gumroad incentivizes AI adoption across the organization with $33,000 bounties for engineers who outperform the CEO How to use component libraries like

    productivity · Claire Vo

  59. research · Industrial and Labor Relations Review ·

    Lucy and the Chocolate Factory: Warehouse Robotics and Worker Safety

    Warehouse robots cut severe injuries 40% but increase non-severe injuries 77%, partly due to accelerated work pace.

    The authors examine the implications of robotics for warehouse worker safety. While warehouse automation has the potential to reduce injuries by eliminating high-risk tasks, it may also increase injuries among remaining non-automated tasks because of reduced task variety and an accelerated pace of work. Findings provide evidence of both effects: Warehouse robotics are associated with a 40% decrease in severe injuries but a 77% increase in non-severe injuries. The authors provide subsequent evidence that the rise in non-severe injuries is at least partially attributable to the increased pace of

    worker experience

  60. practice · Latent Space ·

    AI Agents, meet Test Driven Development

    Five-stage framework for applying test-driven development practices to AI agent development.

    Guest Post: 5 stages for embracing test-driven development (TDD) to build stronger, more reliable AI systems! One of our top talks from AIE Online Track Guest Post: 5 stages for embracing test-driven development (TDD) to build stronger, more reliable AI systems! One of our top talks from AIE Online Track swyx here! We’re delighted to bring you another guest poster, this time talking about something we’ve been struggling to find someone good to articulate: TDD for AI. This is part of a broader discuss…

    ways of working

  61. practice · Eugene Yan ·

    An LLM-as-Judge Won't Save The Product—Fixing Your Process Will

    Using scientific method and eval-driven development to fix quality problems instead of relying on LLM judges alone.

    Applying the scientific method, building via eval-driven development, and monitoring AI output.

    judgment · Eugene Yan

  62. research · Management Science ·

    Algorithmic Writing Assistance on Jobseekers’ Resumes Increases Hires

    Field experiment: jobseekers using algorithmic writing assistance on resumes were hired 8% more often at 10% higher wages, with no employer satisfaction loss.

    There is a strong association between writing quality in resumes for new labor market entrants and whether they are ultimately hired. We show this relationship is, at least partially, causal: In a field experiment in an online labor market with nearly half a million jobseekers, treated jobseekers received nongenerative algorithmic writing assistance on their resumes. Treated jobseekers were hired 8% more often at 10% higher wages. Contrary to concerns that the assistance takes away a valuable signal, we find no evidence that employers were less satisfied. We find that the writing on treated jo

    jobs skills

  63. practice · Harper Reed ·

    An LLM Codegen Hero's Journey

    A staged adoption path for codegen: Copilot to Claude web to Cursor to agents, with advice on when each fits.

    I have spent a lot of time since my blog post about my LLM workflow talking to folks about codegen and how to get started, get better, and why it is interesting. There has been an incredible amount of energy and interest in this topic. I have received a ton of emails from people who are working to figure all of this out. I started to notice that many people are struggling to figure out how to start, and how it all fits together. Then I realized that I have been hacking on this process since 2023 and I have seen some shit. Lol. I was talking about this with friends (Fisaconites’s represent) and

    adoption · Harper Reed

  64. practice · Hacker News ·

    12-factor Agents: Patterns of reliable LLM applications

    Production AI agents succeed by embedding LLMs in well-engineered software, not by using agent frameworks alone.

    I've been building AI agents for a while. After trying every framework out there and talking to many founders building with AI, I've noticed something interesting: most "AI Agents" that make it to production aren't actually that agentic. The best ones are mostly just well-engineered software with LLMs sprinkled in at key points. So I set out to document what I've learned about building production-grade AI systems: https://github.com/humanlayer/12-factor-agents . It's a set of principles for building LLM-powered software that's reliable enough to put in the hands of production customers. In the

    ways of working

  65. research · Microsoft Research ·

    Engagement, user expertise, and satisfaction: Key insights from the Semantic Telemetry Project

    AI users who tackle complex professional tasks show higher sustained engagement and usage frequency than those doing simpler work.

    Semantic Telemetry Project data show that people who use AI for more professional and complex tasks are more likely to keep using the tool and to use it more often. Novice AI users engage in simpler tasks, but their usage is becoming more complex. The post Engagement, user expertise, and satisfaction: Key insights from the Semantic Telemetry Project appeared first on Microsoft Research .

    adoption · Siddharth Suri · Scott Counts

  66. practice · Geoffrey Litt ·

    Stevens: a hackable AI assistant using a single SQLite table and a handful of cron jobs

    Personal AI assistant built with SQLite table and cron jobs, sending daily family briefs via Telegram with calendar, weather, mail and reminders.

    There’s a lot of hype these days around patterns for building with AI. Agents, memory, RAG, assistants—so many buzzwords! But the reality is, you don’t need fancy techniques or libraries to build useful personal tools with LLMs. In this short post, I’ll show you how I built a useful AI assistant for my family using a dead simple architecture: a single SQLite table of memories, and a handful of cron jobs for ingesting memories and sending updates, all hosted on Val.town . The whole thing is so simple that you can easily copy and extend it yourself. Meet Stevens The assistant is called Stevens,

    ways of working · Geoffrey Litt

  67. research · Information Systems Research ·

    Can Providing Algorithmic Performance Information Facilitate Humans’ Inventory Ordering Behaviors?

    Field experiments show that disclosing algorithmic performance information, especially negative results, improves managers' inventory ordering decisions by increasing deliberation.

    Over recent years, companies have been increasingly adopting algorithmic decision systems (ADS) to replace humans. In this paper, we focus on how ADS facilitates human managers’ decision making rather than replacing humans altogether in the context of inventory ordering decisions, specifically on the effect of providing ADS performance information to human managers. Using a pair of field experiments, our results suggest that providing ADS performance information can improve their inventory ordering decisions. Interestingly, providing both positive and negative ADS performance information enhan

    judgment

  68. practice · Kent Beck ·

    Social AI Adoption: Lessons from Hybrid Corn

    Kent Beck uses hybrid corn adoption patterns to explain how AI tools spread through organizations via social influence and peer practices.

    Reflexive AI usage is now a baseline expectation at Shopify Reflexive AI usage is now a baseline expectation at Shopify

    adoption · Kent Beck

  69. practice · Harper Reed ·

    Waterfall in 15 Minutes or Your Money Back

    AI code generation is shifting testing from human-led verification to automated coverage, changing how programmers think about craft.

    I recently had a conversation with a friend that started out as a casual catch-up and spiraled into a deep exploration of AI-assisted coding and what it’s doing to our workflows, teams, and sense of “craft.” It spanned everything from rewriting old codebases to how automated test coverage changes the nature of programming. I took the transcript from granola, popped it into o1-pro, and asked it to write this blog post. Not terrible. Representative of my beliefs. I sent it to a few friends, and they all were interested in sending it to a few more friends. That means I gotta publish it. So here g

    ways of working · Harper Reed

  70. research · Public Administration Review ·

    How Public Officials Perceive Algorithmic Discretion: A Study of Status Quo Bias in Policing

    UK police officers resist algorithmic discretion mainly due to status quo bias: transition costs, loss aversion, and performance uncertainty.

    ABSTRACT Algorithms are disrupting established decision‐making practices in public administration. A key area of interest lies in algorithmic discretion or how public officials use algorithms to exercise discretion. The article develops a framework to explain algorithmic discretion by drawing on status quo bias theory and bureaucratic discretion. A study with police officers in the UK shows that—while officers still value their discretion—it is resistance via the aspects of status quo bias that accounts for a more substantial explanation. Transition costs, loss aversion, and performance uncert

    judgment

  71. practice · Paul Ford (Aboard) ·

    The Spiders Versus the Web

    Web infrastructure companies are blocking AI training spiders as a new form of infrastructure conflict.

    There is a very strange war brewing between AI companies and classic web companies, and it’s being brought into sharp relief in the form of a new product launched by Cloudflare. Cloudflare is a huge web-hosting company (they host about a fifth of all sites ), and they’re very assertive about bots, denial of service attacks, and the like—they have to be, or they won’t be able to deliver the web pages. One of their most recent attack vectors has been AI company web spiders. To explain: In order to build an LLM, you need to gobble up as much content as possible—ideally all the content in the worl

    management org · Paul Ford

  72. research · HBS AI Institute ·

    The Cybernetic Teammate: How AI is Reshaping Collaboration and Expertise in the Workplace

    Field experiment at P&G tests how generative AI affects teamwork and expertise in organisational settings.

    In an era where AI is rapidly transforming business operations, a recent working paper, “The Cybernetic Teammate: A Field Experiment on Generative AI Reshaping Teamwork and Expertise,” describes research that explored how generative AI is reshaping teamwork and expertise in organizational settings. This research, conducted through a large-scale field experiment at Procter & Gamble (P&G), […] The post The Cybernetic Teammate: How AI is Reshaping Collaboration and Expertise in the Workplace appeared first on Harvard Business School AI Institute .

    teams

  73. practice · Paul Ford (Aboard) ·

    The More Things Change

    AI tools are maturing from hype to predictable, standardized products with clear use cases and limits.

    On the podcast this week , I spoke with my co-founder, Richard, about how AI is a little…boring lately. As the technology stabilizes and matures, usage patterns are starting to appear. You can trust the chat in some things—summarizing documents, or generating Ghibli-esque images without ethical constraints , or helping you along with code. You need to validate findings in others, like when it makes up citations for research reports. But I don’t know how many revolutionary product releases await us. Probably one or two! There’s billions of dollars of gas in the tank, but it feels like we’re sta

    management org · Paul Ford

  74. research · NBER ·

    The Cybernetic Teammate: A Field Experiment on Generative AI Reshaping Teamwork and Expertise

    Field experiment with 776 P&G professionals shows how generative AI affects team performance, expertise sharing, and social engagement in real product innovation work.

    We examine how artificial intelligence transforms the core pillars of collaborationperformance, expertise sharing, and social engagementthrough a pre-registered field experiment with 776 professionals at Procter & Gamble, a global consumer packaged goods company. Working on real product innovation (Fabrizio Dell'Acqua , Charles Ayoubi , Hila Lifshitz , Raffaella Sadun , Ethan Mollick , Lilach Mollick , Yi Han , Jeff Goldman , Hari Nair , Stewart Taub , Karim Lakhani)

    teams · Hila Lifshitz-Assaf · Karim Lakhani · Ethan Mollick · Fabrizio Dell'Acqua · Raffaella Sadun · Charles Ayoubi · Lilach Mollick

  75. research · NBER ·

    Measuring Human Leadership Skills with AI Agents

    Leaders' problem-solving performance with AI agents strongly correlates (r=0.81) with their performance leading human groups.

    We show that leadership skill with artificially intelligent (AI) agents predicts leadership skill with human groups. In a large pre-registered lab experiment, human leaders worked with AI agents to solve problems. Their performance on this AI leadership test was strongly correlated (=0.81) with (Ben Weidmann , Yixian Xu , David J. Deming)

    management org

  76. research · Management Science ·

    Mitigating the Negative Effects of Customer Anxiety by Facilitating Access to Human Contact

    Field experiment shows offering human contact during self-service loan approval increases customer uptake by 24% by reducing anxiety-driven dissatisfaction.

    Prior research in social psychology has shown that when people feel anxious, they seek advice from others. Yet, companies that operate in high-anxiety settings (like financial services, healthcare, and education) are increasingly deploying self-service technologies (SSTs), through which anxious customers transact without access to human contact. These companies may therefore face a classic efficiency–service trade-off where gains in operational efficiency through automation also hamper service outcomes by neglecting customer anxiety. This paper, set in the high-anxiety domain of financial serv

    worker experience

  77. practice · Paul Ford (Aboard) ·

    Bad Vibes Coding

    Developers who skip reading AI-generated code and trust the output are adopting a new workflow pattern despite its risks.

    When I think something is a truly terrible idea, it often succeeds beyond my wildest imagination. One example is Twitter. When it launched, I thought it was shallow nonsense for dimwits, but it turned out to be a world-shaping, influential platform, even if it’s now an enormous cultural superfund site. I thought the same thing about Facebook, and Instagram, and TikTok. And pretty much anything blockchain. In general, if I really hate something, you should invest in it. I’ve also loathed a lot of technologies that became popular, like HTML5, or JavaScript, or React. They all became huge winners

    ways of working · Paul Ford

  78. research · Information Systems Research ·

    Artificial Intelligence and Firm Resilience: Empirical Evidence from Natural Disaster Shocks

    Firms with 2.4% AI-related jobs recover faster from natural disasters; benefits hampered by lack of complementary org design.

    Artificial intelligence (AI) has been increasingly deployed in business operations over the past decade, whereas direct evidence of its effectiveness in uncertain contexts is limited. Our work examines the contribution of AI to corporate resilience under natural disaster shocks, particularly concentrating on AI-using and goods-producing firms. We measure firm AI investment by the cumulative AI-relevant skills extracted from a comprehensive job posting database and firm resilience by the changes in corporate valuation in response to operational shocks. Evidence suggests that AI generates resili

    productivity

  79. practice · Latent Space ·

    Agent Engineering

    Framework for what makes an agent and why agent engineering is the biggest opportunity for AI engineers now.

    Defining Agents, Why now, and why Agents are the biggest opportunity for AIEs Defining Agents, Why now, and why Agents are the biggest opportunity for AIEs This post contains elaborations on swyx’s 2025 AI Engineer Summit keynote, which also serves as a cohesive overview of a selection of Agents talks from the conference which link-clickers can preview.

    ways of working · Shawn Wang (swyx)

  80. practice · Hamel Husain ·

    A Field Guide to Rapidly Improving AI Products

    AI teams improve products faster by measuring outcomes and iterating with domain experts, not by choosing tools.

    Most AI teams focus on the wrong things. Here’s a common scene from my consulting work: AI TEAM Here’s our agent architecture – we’ve got RAG here, a router there, and we’re using this new framework for… ME [Holding up my hand to pause the enthusiastic tech lead.] “Can you show me how you’re measuring if any of this actually works?” … Room goes quiet This scene has played out dozens of times over the last two years. Teams invest weeks building complex AI systems, but can’t tell me if their changes are helping or hurting. This isn’t surprising. With new tools and frameworks emerging weekly, it’

    management org · Hamel Husain

  81. research · JAMA Network Open ·

    Physician Perspectives on Ambient AI Scribes

    Physicians using ambient AI scribes report ease of use and positive tool quality, with detailed barriers and facilitators to adoption identified.

    Importance: Limited qualitative studies exist evaluating ambient artificial intelligence (AI) scribe tools. Such studies can provide deeper insights into ambient AI implementations by capturing lived experiences. Objective: To evaluate physician perspectives on ambient AI scribes. Design, Setting, and Participants: A qualitative study using semistructured interviews guided by the Reach, Efficacy, Adoption, Implementation, Maintenance/Practical, Robust Implementation, and Sustainability Model (RE-AIM/PRISM) framework, with thematic analysis using both inductive and deductive approaches. Physici

    adoption

  82. research · Management Science ·

    Human-Algorithm Collaboration with Private Information: Naïve Advice-Weighting Behavior and Mitigation

    Humans systematically overweight algorithm predictions, ignoring private information; feature transparency and targeted interventions reduce prediction error by 25-34%.

    Even if algorithms make better predictions than humans on average, humans may sometimes have private information that an algorithm does not have access to that can improve performance. How can we help humans effectively use and adjust recommendations made by algorithms in such situations? When deciding whether and how to override an algorithm’s recommendations, we hypothesize that people are biased toward following naïve advice-weighting (NAW) behavior; they take a weighted average between their own prediction and the algorithm’s prediction, with a constant weight across prediction instances r

    judgment

  83. research · arXiv ·

    Collaborating with AI Agents: Field Experiments on Teamwork, Productivity, and Performance

    Human-AI teams produced 50% more ads per worker with higher text quality; human teams produced better images, revealing uneven AI capability.

    We examined the mechanisms underlying productivity and performance gains from AI agents using a large-scale experiment on Pairit, a platform we developed to study human-AI collaboration. We randomly assigned 2,234 participants to human-human and human-AI teams that produced 11,024 ads for a think tank. We evaluated the ads using independent human ratings and a field experiment on X which garnered ~5M impressions. We found human-AI teams produced 50% more ads per worker and higher text quality, while human-human teams produced higher image quality, suggesting a jagged frontier of AI agent capab

    teams

  84. research · ACM Transactions on Software Engineering and Methodology ·

    Investigating the Role of Cultural Values in Adopting Large Language Models for Software Engineering

    Among 188 software engineers, habit and performance expectancy drive LLM adoption; cultural values do not significantly moderate adoption patterns.

    As a socio-technical activity, software development involves the close interconnection of people and technology. The integration of Large Language Models (LLMs) into this process exemplifies the socio-technical nature of software development. Although LLMs influence the development process, software development remains fundamentally human-centric, necessitating an investigation of the human factors in this adoption. Thus, with this study we explore the factors influencing the adoption of LLMs in software development, focusing on the role of professionals’ cultural values. Guided by the Unified

    adoption

  85. practice · Paul Ford (Aboard) ·

    The Zeno Effect

    AI does not solve the fundamental tension between scope creep and shipping deadlines in software projects.

    There’s a joke in software development: Now that you’ve got the first 80% done, you need to do the remaining 80%. Of course, this applies to most things—writing and editing, creating art, PhD dissertations—but in code the fact that things don’t ship gets a lot of attention because it’s so expensive . If a PhD doesn’t land on time, that’s suffering for one person and their family. When code doesn’t land, that’s a whole team, tons of money, and a big corporate plan going “whoosh.” So right as you see things landing, right as marketing gets the press release in order, you realize you’re going to

    management org · Paul Ford

  86. practice · Artificial Ignorance ·

    Breaking Into AI Engineering

    Writing publicly about learning AI tools helped transition an engineer into an AI engineering role.

    How building in public transformed my career path. From writing about ChatGPT to becoming a professional AI engineer. Two years ago, I knew almost nothing about generative AI. Like millions (now billions) of others, I was stunned when ChatGPT first launched - but I did have both an engineering background and the stu…

    jobs skills · Charlie Guo

  87. practice · Laurie Voss ·

    AI's effects on programming jobs

    AI will expand the programming workforce by creating new layers of abstraction, not eliminate programmer jobs through higher demand for software.

    There's been a whole lot of discussion recently about the impact of AI on the market for web developers, for programmers in general, and even more generally the entire labor market. I find myself making the same points over and over, and whenever I do that it's time to write a blog post about it, so this is that. Doom and utopia are not our only options There are two extreme takes on the impact of AI on programmer jobs: AI will take all programming jobs (usually advanced by people selling AI) AI will not take anyone's job (usually advanced by grumpy older developers, which I usually am) I woul

    jobs skills · Laurie Voss

  88. research · Léonard Boussioux ·

    The Dangers of Deferring to AI: It Seems So Right Even When It's Wrong

    Field experiment shows AI-generated explanations cause evaluators to reject good ideas more often than human explanations.

    Harvard Business School · Working Knowledge — On the field experiment finding that AI-generated narrative explanations made evaluators reject good ideas more often: “you really need humans synthesizing and validating.”

    judgment · Léonard Boussioux

  89. research · arXiv ·

    How Generative AI Adoption Alters the Demand for Cognitive and Social Skills Within Roles: A Skill-Centric Analysis

    GenAI adoption reduced demand for social skills by 3.4% in adopting roles while cognitive skill demand stayed steady, contradicting predictions of skill reallocation.

    A common view holds that generative AI (GenAI) automates cognitive tasks, reshaping roles to emphasize social skills over cognitive ones. Drawing on the framework we develop in this paper, we argue that other outcomes are theoretically possible. We analyze seven million job postings from 595 U.S. public firms that adopted GenAI in 2022-2024, estimating difference-in-differences models comparing GenAI-adopting and non-adopting roles around ChatGPT's launch. We find no evidence of greater emphasis on social skills in GenAI-adopting roles; instead, their relative demand for social skills declined

    jobs skills · Phanish Puranam

  90. research · Management Science ·

    Designing AI-Based Work Processes: How the Timing of AI Advice Affects Diagnostic Decision Making

    Diagnostic accuracy improves when AI advice is given after initial diagnosis, due to deeper cognitive engagement with AI reasoning.

    Although clinical artificial intelligence (AI) systems can augment medical diagnosis decisions by providing competent second opinions, how to effectively integrate AI into routine diagnostic processes, such as when to present AI advice to human physicians, remains largely unexplored. Therefore, our research experimentally examines how the timing of AI advice affects diagnostic decision making using a think-aloud approach. Physicians perform medical diagnoses under three conditions: ex post advice (AI advice given after an initial diagnosis), ex ante advice (AI advice given concurrently with cl

    judgment

  91. practice · Paul Ford (Aboard) ·

    The $15 Volvo

    Enterprise buyers are skeptical of AI tools promised as cheap replacements for legacy systems, seeing cosmetic solutions rather than robust infrastructure.

    In the conversation about AI coding, lots of organizations are promising to completely change the way software is developed. You will tell the computer what you need, AI will build it, and it will be good. Companies like Bolt and v0 offer impressive abilities to spin up websites or other software tools just using words. It’s still early days, but eventually, they say that if you can describe it, you’ll have it. This creates a problem—around the office, we call it “the $15 Volvo problem.” Building a website that looks nice and sits around on the web is a relatively low-risk proposition. That’s

    management org · Paul Ford

  92. research · Management Science ·

    Statistical Tests for Replacing Human Decision Makers with Algorithms

    Algorithm outperformed individual doctors on abnormal birth detection, with higher true positive and lower false positive rates on real diagnostic data.

    This paper proposes a statistical framework of using artificial intelligence to improve human decision making. The performance of each human decision maker is benchmarked against that of machine predictions. We replace the diagnoses made by a subset of the decision makers with the recommendation from the machine learning algorithm. We apply both a heuristic frequentist approach and a Bayesian posterior loss function approach to abnormal birth detection using a nationwide data set of doctor diagnoses from prepregnancy checkups of reproductive-age couples and pregnancy outcomes. We find that our

    judgment

  93. practice · Artificial Ignorance ·

    Hallucinations Are Fine, Actually

    Rethinking when AI hallucinations matter in real workflows rather than dismissing them as a universal flaw.

    Why I changed my mind about AI's imperfections. Why I changed my mind about AI's imperfections When I first started using ChatGPT, I quickly discovered its tendency to confidently make things up. Whether you want to call it lying, fabricating, or just bullshitting - it had no problems inventin…

    judgment · Charlie Guo

  94. practice · Geoffrey Litt ·

    Avoid the nightmare bicycle

    AI tools should expose underlying structure and user control rather than hiding complexity behind mode buttons.

    In my opinion, one of the most important ideas in product design is to avoid the “nightmare bicycle”. Imagine a bicycle where the product manager said: “people don’t get math so we can’t have numbered gears. We need labeled buttons for gravel mode, downhill mode, …” This is the hypothetical “nightmare bicycle” that Andrea diSessa imagines in his book Changing Minds . As he points out: it would be terrible! We’d lose the intuitive understanding of how to use the gears to solve any situation we encounter. Which mode do you use for gravel + downhill? It turns out, anyone can understand numbered g

    ways of working · Geoffrey Litt

  95. research · NBER ·

    AI and the Extended Workday: Productivity, Contracting Efficiency, and Distribution of Rents

    Workers in AI-exposed occupations work longer hours; study uses time diary data from 2004-2023 to measure intensive-margin employment effects.

    This study investigates how occupational AI exposure impacts employment at the intensive margin, i.e., the length of workdays and the allocation of time between work and leisure. Drawing on individual-level time diary data from 20042023, we find that higher AI exposurewhether stemming from the (Wei Jiang , Junyoung Park , Rachel (Jiqiu) Xiao , Shen Zhang)

    worker experience

  96. research · ACM Transactions on Software Engineering and Methodology ·

    Accountability in Code Review: The Role of Intrinsic Drivers and the Impact of LLMs

    Software engineers report four intrinsic accountability drivers; LLM code review shifts accountability from individual to collective, reducing personal responsibility.

    Accountability is an innate part of social systems. It maintains stability and ensures positive pressure on individuals’ decision-making. As actors in a social system, software developers are accountable to their team and organization for their decisions. However, the drivers of accountability and how it changes behavior in software development are less understood. In this study, we look at how the social aspects of code review affect software engineers’ sense of accountability for code quality. Since Software Engineering (SE) is increasingly involving Large Language Models (LLM) assistance, w

    worker experience

  97. practice · Understanding AI (Timothy B. Lee) ·

    My brother explains what it's like to run an AI startup

    Email app uses six different AI models in production, each chosen for specific tasks within a real product.

    The Shortwave email app uses six different models for a range of tasks. The Shortwave email app uses six different models for a range of tasks. I’m pleased to present this cross-posted episode of AI Summer, a podcast I co-host with Dean Ball of the Mercatus Center. Most episodes of AI Summer are not cross-posted here, so if you enjoy this conversation please subscribe to the newsletter or search for “AI Summer” in your favorite podcast app.

    ways of working · Timothy B. Lee

  98. research · arXiv ·

    Predictive AI Can Support Human Learning while Preserving Error Diversity

    AI support during training and practice improves diagnostic accuracy and preserves error diversity, affecting group decision-making quality.

    We examined the effects of predictive AI deployment on the immediate performance and learning of medical novices. In two pre-registered field experiments, we varied whether AI input was provided during the training or practice of lung cancer diagnoses, or both. Our results show that different AI deployments have distinct implications for human professionals. AI input during training or practice independently improves individuals' diagnostic accuracy, whereas deployment across both phases yields gains that exceed either approach alone. Furthermore, AI input in both training and earlier practice

    judgment · Phanish Puranam

  99. research · Information and Organization ·

    Novice risk work: How juniors coaching seniors on emerging technologies such as generative AI can lead to learning failures

    Junior consultants given early access to GPT-4 recommended risky practices due to shallow understanding of AI's uncertain capabilities and scope.

    Historically, junior professionals have mentored senior professionals around new technologies, because juniors are typically more willing than seniors to perform lower-level tasks to learn new skills, better able than seniors to engage in real-time experimentation close to the work itself, and more willing than seniors to learn innovative methods that conflict with traditional identities and norms. However, we know little about what happens when emerging technologies have a high level of uncertainty in their use, because they have wide-ranging capabilities and are exponentially changing. With

    adoption · Hila Lifshitz-Assaf · Katherine Kellogg · Karim Lakhani · Ethan Mollick · Fabrizio Dell'Acqua · Edward McFowland III · Steven Randazzo · François Candelon

  100. practice · Paul Ford (Aboard) ·

    Hegel on the Death Star

    Using AI to quickly bridge knowledge gaps by asking absurd questions reveals how LLMs map between conceptual domains.

    This weekend I saw someone online talking about the philosopher Hegel, and I thought to myself, What’s all that about? I opened up Wikipedia, and I will say: Georg Wilhelm Friedrich Hegel has one of the most forbidding faces I’ve ever seen. The gravity of his countenance! But also his wiki page is enormous—countless sections and footnotes. I know Hegel is very important, but I had to make dinner and get the laundry done. All I wanted was a taste. A little Hegel dabble. So—you know where this is going—I used ChatGPT. What is the dumbest question I could ask, the most humiliating way of learning

    ways of working · Paul Ford