The Feed · Complete archive
AI at work: research and practice
A server-rendered record of the evidence, ideas and firsthand practices screened by The Feed. Every entry links to its original source.
Use the interactive FeedPage 6 of 9 · 873 items
practice · Jason Fried ·
The joy of delegating to competence
Delegating to AI agents creates a psychological experience of total trust previously rare outside elite teams.
AI workflows are technically impressive, but there’s a deeper reason people are really amped about AI agents. This isn’t just new tech, it’s new psychology. Until now, very few people have known what it feels like to delegate to total competency. If you manage great people, or lead great teams, you know how it feels to put someone in charge who will get it done, get it done right, and get it done without drama. That kind of delegation — that depth of trust — is pure joy. Delegating to competency lets you forget about it completely. That’s real leverage. And now anyone can experience that. What
ways of working · Jason Fried
practice · The Pragmatic Engineer ·
When AI writes almost all code, what happens to software engineering?
Gergely Orosz argues that AI-written code is now industry-wide practice, not hypothesis, with implications for roles.
No longer a hypothetical question, this is a mega-trend set to hit the tech industry No longer a hypothetical question, this is a mega-trend set to hit the tech industry This article was originally a paid one. I’ve removed the paywall on 24 September 2026, after DHH’s Rails World keynote on 23 Sep 2026 caused quite the stir, where he shared how 37Signals has gone “pencils down” and generates all code. In January, we saw this will happen across the industry, sooner, rather than later. The good news: software engineering …
jobs skills · Gergely Orosz
practice · Harper Reed ·
Remote Claude Code: programing like it was the early 2000s
Using Claude Code from a phone terminal to program remotely, mimicking early-2000s development practices.
So so many friends have asked me how I use Claude Code from my phone. I am always a bit surprised, because a lot of this type of work I have been doing for nearly 25 years (or more!) and I always forget that it is partially a lost art. This is how it used to be. We didn’t have fancy IDEs, and fancy magic to deploy stuff. We had to ssh (hah. Telnet!) into a machine and work with it through the terminal. It ruled. It was a total nightmare. It was also a lot of fun. One of my favorite parts of the early 2000s was hanging out in IRC channels and just participating in the most ridiculous community
ways of working · Harper Reed
practice · Addy Osmani ·
Code Review in the Age of AI
Code review now splits between solo developers using test suites as backstops and teams using review for shared context, because AI-generated code has 75% more logic errors.
AI did not kill code review. It made the burden of proof explicit. Ship changes with evidence like manual verification and automated tests, then use review for risk, intent, and accountability. Solo developers lean on automation to keep up with AI speed, while teams use review to build shared context and ownership. If your pull request doesn’t contain evidence that it works, you’re not shipping faster - you’re just moving work downstream. By early 2026, over 30% of senior developers report shipping mostly AI-generated code. The challenge? AI excels at drafting features but falters on logic, se
judgment · Addy Osmani
research · NBER ·
Automation Experiments and Inequality
Analysis of what productivity experiments on automation actually reveal about inequality effects across worker skill levels.
Many experiments study the productivity effects of automation technologies such as generative algorithms. A key test in these experiments relates to inequality: does the technology increase output more for high- or low-skill workers? However, the theoretical content of this empirical test has been (Seth Gordon Benzell , Kyle R. Myers)
jobs skills
research · NBER ·
How Adaptable Are American Workers to AI-Induced Job Displacement?
Workers in AI-exposed occupations have higher adaptive capacity to navigate job transitions, based on an index across 356 US occupations.
We construct an occupation-level adaptive capacity index that measures a set of worker characteristics relevant for navigating job transitions if displaced, covering 356 occupations that represent 95.9% of the U.S. workforce. We find that AI exposure and adaptive capacity are positively correlated: (Sam J. Manning , Tomás Aguirre)
jobs skills
research · NBER ·
Artificial Jagged Intelligence: When AI Benchmarks Misstate Deployment Value
AI systems that score highly on public benchmarks may perform unevenly on the specific task distributions organisations actually face.
Organisations increasingly select and deploy artificial intelligence systems on the strength of public benchmarks. A benchmark, however, scores a system on a single distribution of tasks, whereas each organisation meets its own. Because AI performance is uneven across tasks, a property called (Joshua S. Gans)
adoption
research · NBER ·
O-Ring Automation
Automation of complementary tasks (O-ring production) creates different wage and employment effects than separable task automation.
We study automation when tasks are quality complements rather than separable. Production requires numerous tasks whose qualities multiply as in an O-ring technology. A worker allocates a fixed endowment of time across the tasks performed; machines can replace tasks with given quality, and time is (Joshua S. Gans , Avi Goldfarb)
jobs skills · Avi Goldfarb
research · SSRN ·
GenAI Adoption Increases the Density of Knowledge and Collaboration Networks: Evidence from a Field Experiment
A field experiment shows GenAI adoption increases density of knowledge and collaboration networks within organisations.
adoption · Phanish Puranam
research · SSRN ·
Solve and Be Seen: How Workers in Deskilled Jobs Build Complex Skill as New Automation Arrives
Workers in deskilled roles develop complex skills by solving problems visible to management as automation arrives.
jobs skills · Matt Beane
research · SSRN ·
Hedging with Talent: How Selection Under Uncertainty Reshapes Automation's Workforce Effects
Firms hedge automation risk by hiring more adaptable workers, shifting who gains or loses jobs from technological change.
jobs skills · Matt Beane
research · IEEE Transactions on Software Engineering ·
Impact of an LLM-based Review Assistant in Practice: A Mixed Open-/Closed-source Case Study
LLM-generated code review comments were accepted in 8% of cases across Mozilla and Ubisoft, with 15-21% rated as valuable tips.
Code review is a standard practice in modern software development, aiming to improve code quality.As providing constructive reviews on a submitted patch is a challenging and error-prone task, the advances of Large Language Models (LLMs) in performing natural language processing (NLP) tasks in a human-like fashion have prompted researchers to evaluate the LLMs’ ability to automate the code review process. However, outside lab settings, it is still unclear how prone reviewers are to accept comments generated by LLMs in a real development workflow. To fill this gap, we conduct a large-scale empir
adoption
research · IEEE Transactions on Software Engineering ·
Understanding Prompt Quality and Its Relation to LLM-based Code Generation
Readability and structure of prompts to LLMs predict code quality; analysis of real developer-ChatGPT interactions from GitHub.
Prompt engineering has emerged as a key practice to guide Large Language Models (LLMs) in code generation tasks, under the assumption that better prompts yield better outputs. Yet little is known about which characteristics of prompts actually drive variations in output quality. In this paper, we take a first step toward addressing this gap by investigating how measurable properties of prompts relate to the quality of generated code. We conducted an empirical study of real-world, single-turn interactions between developers and ChatGPT collected from GITHUB. Prompts were characterized along rea
adoption
research · Human Relations ·
Problematizing the role of artificial intelligence in hiring and organizational inequalities: A multidisciplinary review
Multidisciplinary review proposes concept of 'algorithmically-mediated inequality regimes' to explain how AI in hiring conceals rather than eliminates workplace inequalities.
What are the implications of the growing use of artificial intelligence (AI) in recruitment and hiring for organizational inequalities? While advocates suggest that AI is a groundbreaking tool that can enhance hiring precision, efficiency, diversity and fit, critics raise serious concerns around bias, fairness, and privacy. This review article critically advances this debate by drawing on diverse scholarship across computing and data sciences; human resource, management, and organization studies; social sciences; and law. Using a hybrid review approach that combines scoping and problematizing
jobs skills
practice · Addy Osmani ·
How Good Is AI at Coding React (Really)?
AI coding success depends on context engineering and domain knowledge; React developers can guide models away from failures by understanding their limits.
tl;dr: AI coding benchmarks show models excel at isolated React tasks like scaffolding components or implementing explicit specs, achieving ~40% success in benchmarks, but drops to ~25% on multi-step integrations due to a “complexity cliff” in state management and design taste. The gap between “AI helped me ship” and “AI gave me a mess” is context engineering and explicit constraints. Deep React and domain knowledge enable you to spot when AI goes off the rails and understand why it repeats mistakes. Guide it without blindly accepting the output. This article is based on my closing keynote at
ways of working · Addy Osmani
practice · How I AI ·
How Webflow’s CPO built an AI chief of staff to manage her calendar, prep for meetings, and drive AI adoption | Rachel Wolan
CPO built a custom AI agent to manage calendar, meetings, and email, then used hands-on experience to drive adoption across her company.
Rachel Wolan , the chief product officer at Webflow, has embraced AI not just as a product leader but as a hands-on builder. A coder since age 16, Rachel has returned to her technical roots by creating a custom AI chief-of-staff application that helps manage her executive workload. In this episode, she demonstrates how she uses personal AI software to prep for meetings, triage her calendar, manage emails, and even get brutally honest feedback about how she’s spending her time. What you’ll learn: How Rachel built a custom AI chief-of-staff application that integrates with her calendar, email, a
ways of working · Claire Vo
practice · Sangeet Paul Choudary ·
The Not-so-Lazy Holiday Reading List
Reflection and strategic inaction become competitive advantages when AI makes cheap execution and instant answers abundant.
It’s that time of the year again! Everyone’s making predictions about what’s going to happen in 2026! Predictions are easy to generate and costless to abandon if the person making the prediction bears no responsibility for being wrong. And chatbots can do that better than you and me - producing fluent answers on demand. What’s harder - and increasingly rare - is the work of reflection: slowing down, taking stock, and making sense of change. And to make it worth your while, here’s a curated list of Holiday Readings to help you reflect more and reflect better in a world of relentless execution.
management org · Sangeet Paul Choudary
practice · Hacker News ·
Ask HN: How are you sandboxing coding agents?
Teams share practical sandboxing setups for AI coding agents: git worktrees, VMs, firejail, with real tradeoffs between safety and convenience.
I've seen people rely on built-in sandboxes, use git worktrees (sometimes inside devcontainers), or run the whole agent inside a Linux VM with minimal host mounts. On Linux, I’ve also seen firejail/bubblewrap mentioned. For folks actually using these tools day-to-day: What’s your default setup? Have you had any "learned the hard way" moments? What tradeoff (safety vs convenience vs parallelism) has mattered most in practice? I'm less interested in theoretical best practices than what's actually holding up under real use.
ways of working
research · Information Systems Research ·
A Robust Optimization Approach to Reliable Statistical Inference with Variables Generated by Machine Learning
A robust optimization method improves statistical inference when ML-generated variables contain prediction errors, tested on Amazon reviews.
Organizations increasingly use machine learning to turn text, images, and other unstructured data into variables that inform decisions and research. But, because machine learning predictions are never perfect, the resulting data can contain errors that quietly distort statistical analyses, sometimes leading to incorrect conclusions about what truly drives important outcomes. This study introduces a robust optimization approach that helps analysts and decision makers draw more reliable insights when working with machine learning–generated data. The method is designed to strengthen the signal of
judgment
practice · Charity Majors ·
2025 was for AI what 2010 was for cloud
2025 marked AI adoption inflection point in enterprise ops, mirroring cloud computing's shift from experimental to mainstream between 2006-2010.
I was at my very first job, Linden Lab, when EC2 and S3 came out in 2006. We were running Second Life out of three datacenters, where we racked and stacked all the servers ourselves. At the time, we were tangling with a slightly embarrassing data problem in that there was no real way for users to delete objects (the Trash folder was just another folder), and by the time we implemented a delete function, our ability to run garbage collection couldn’t keep up with the rate of asset creation. In desperation, we spun up an experimental project to try using S3 as our asset store. Maybe we could mak
adoption · Charity Majors
research · HBS AI Institute ·
The Agentic AI Reality Check
Early empirical data on how AI agents are actually adopted and used by real users of Perplexity.
Agentic AI has recently been moving through a period of heightened excitement and innovation, but empirical data on how these tools are actually being used has been scarce. The new study “The Adoption and Usage of AI Agents: Early Evidence from Perplexity,” by Jeremy Yang, Assistant Professor of Business Administration at Harvard Business School and […] The post The Agentic AI Reality Check appeared first on Harvard Business School AI Institute .
adoption
practice · One Useful Thing ·
The Shape of AI: Jaggedness, Bottlenecks and Salients
AI's jagged frontier persists: superhuman at some tasks, poor at others unrelated to apparent difficulty.
Back in the ancient AI days of 2023, my co-authors and I invented a term to describe the weird ability of AI to do some work incredibly well and other work incredibly badly in ways that didn’t map very well to our human intuition of the difficulty of the task. We called this the “Jagged Frontier” of AI ability, and it remains a key feature of AI and an endless source of confusion. How can an AI be superhuman at differential medical diagnosis or good at very hard math (yes, they are really good at math now, famously outside the frontier until recently) and yet still be bad at relatively simple
judgment · Ethan Mollick
practice · Addy Osmani ·
My LLM coding workflow going into 2026
Developers using LLMs as pair programmers need clear specs, context, oversight and critical thinking to get consistent results.
AI coding assistants became game-changers this year, but harnessing them effectively takes skill and structure. These tools dramatically increased what LLMs can do for real-world coding, and many developers (myself included) embraced them. At Anthropic, for example, engineers adopted Claude Code so heavily that today ~90% of the code for Claude Code is written by Claude Code itself . Yet, using LLMs for programming is not a push-button magic experience - it’s “difficult and unintuitive” and getting great results requires learning new patterns. Critical thinking remains key. Over a year of proj
ways of working · Addy Osmani
practice · How I AI ·
How Zapier’s EA built an army of AI interns to automate meeting prep, strengthen team culture, and scale internal alignment | Cortney Hickey
EA built automated meeting prep workflows that research participants and pull CRM data, plus AI-powered culture reinforcement across the org.
Cortney Hickey is the executive assistant to the CEO at Zapier, where she’s leveraging AI to transform traditional EA responsibilities into scalable, organization-wide systems. In this episode, she demonstrates how she’s built AI workflows that automate meeting preparation, reinforce company culture through automated feedback, and democratize strategic knowledge across the organization. Her approach shows how EAs can use AI not to replace their roles but to elevate them—working on higher-impact initiatives while creating systems that benefit the entire company. What you’ll learn: How to build
ways of working · Claire Vo
research · JAMA Network Open ·
Uptake of Generative AI Integrated With Electronic Health Records in US Hospitals
Survey of 2024 US hospital IT adoption of generative AI in EHRs shows early adopter rates and links adoption to prior AI experience and hospital characteristics.
Importance: There is widespread enthusiasm about generative artificial intelligence (AI), but no systematic evidence on its implementation across health care organizations. Objective: To describe adoption of generative AI integrated with the electronic health record (EHR) by nonfederal acute care hospitals, how adoption relates to experience using and evaluating predictive AI, and hospital characteristics. Design, Setting, and Participants: This survey study of nonfederal acute care US hospitals used the 2024 American Hospital Association (AHA) Information Technology (IT) Supplement survey. Th
adoption
practice · Paul Ford (Aboard) ·
Don’t Waste a Miracle
AI coding assistants are lowering the skilled labor needed to build deterministic systems, enabling non-programmers to build platforms within weeks.
There’s a lot swirling right now: Claude Code getting good, AI-powered political messaging , McDonald’s making a depressing AI ad everyone hates, OpenAI going “code red,” Trump giving Nvidia the green light to sell some chips to China, and on and on. On a personal level, despite my co-founder’s best attempts , I keep moving forward on my Claude Code experiments . I’m grasping the contours of what I can do with AI in coding and it’s no longer quite so brain-melting—I’m getting used to it, as always happens—but it also leaves me convinced that next year is going to be a different kind of year. I
jobs skills · Paul Ford
practice · The Pragmatic Engineer ·
Frictionless: why great developer experience can help teams win in the ‘AI age’
Developer experience reduces friction in AI workflows, enabling teams to ship faster and adopt AI tools more widely.
Exclusive excerpts from the newly-released book ‘Frictionless’, by Nicole Forsgren and Abi Noda. Exclusive excerpts from the newly-released book ‘Frictionless’, by Nicole Forsgren and Abi Noda. Hi, this is Gergely with a bonus, free issue of the Pragmatic Engineer Newsletter. In every issue, I cover Big Tech and startups through the lens of senior engineers and engineering leaders. If you’ve been forwarded this email, you can subscribe here
adoption · Nicole Forsgren · Gergely Orosz
research · Strategic Organization ·
Generative organizational learning: Affordances for new modes of knowledge search, creation, transfer, and forgetting with LLMs
Case study shows how generative AI enables new organizational learning modes: knowledge creation and personalization, plus selective forgetting through guardrailing.
Scholars of organizational learning typically focus on how organizations operate in an environment of domain knowledge scarcity, assuming that knowledge abundance is a highly sought after, but fundamentally elusive, goal. In our case study of organizational learning at Northeast Health, we find that generative AI, in combination with actors’ goals and organizational institutions, may afford a new form of learning that we call generative organizational learning enabling innovation in the service of organizational goals. In particular, generative catalyzing, iterating, and personalizing may enab
management org · Katherine Kellogg
research · Journal of Management Studies ·
Demystifying AI for the Workforce: The Role of Explainable AI in Worker Acceptance and Management Relations
Counterfactual and local AI explanations increase gig workers' acceptance of algorithmic decisions, but combining both overwhelms them and reduces acceptance.
Abstract In the digital era, organizations are increasingly leveraging artificial intelligence (AI) to optimize their operations and decision‐making. However, the opaqueness of AI processes raises concerns over trust, fairness, and autonomy, especially in the gig economy, where AI‐driven management is ubiquitous. This study investigates how explainable AI (xAI), through the comparative use of counterfactual versus factual and local versus global explanations, shapes gig workers’ acceptance of AI‐driven decisions and management relations, drawing on cognitive load theory. Using experimental dat
judgment
research · Journal of Management Studies ·
The Dark Side of Managing Human– AI Collaborations: Implications for Leaders’ Moral Relativism and Unethical Behaviour
Managing human-AI collaboration may increase leaders' moral relativism and unethical behaviour, moderated by cognitive-closure needs.
Abstract As collaborations between humans and artificial intelligence (AI) have become increasingly prevalent across various industries, the role of leaders in managing these collaborations has grown in importance. While the existing literature has highlighted the benefits of leader management in these settings – emphasizing the complementary strengths of humans and AI – the potential costs to key stakeholders, particularly to leaders themselves, have been largely ignored. This research addresses this gap by drawing on moral relativism theory to develop and test a model explaining how leader m
management org
research · Information Systems Research ·
Knowing (Not) to Know: Explainable Artificial Intelligence and Human Metacognition
Explainable AI reduces expert overconfidence and improves their metacognition, enabling better human-AI collaboration and decision outcomes in real estate and lending.
Many organizations seek to combine human expertise with explainable artificial intelligence (XAI), but they often overlook a core requirement for effective collaboration; humans must understand their own abilities. This understanding of one’s own abilities is referred to as metacognition, which captures how well individuals monitor and regulate their own decision making. In two experiments with real estate and lending professionals, we find that XAI improves human metacognition by reducing their overconfidence. As a result, experts can better delegate decisions to an artificial intelligence (A
judgment
research · npj Digital Medicine ·
A typology of physician input approaches to using AI chatbots for clinical decision-making
Physicians use four distinct approaches to input clinical information into LLM chatbots; input method alone does not predict better diagnostic performance.
Recent studies have found that physicians with access to a large language model (LLM) chatbot during clinical reasoning tests may score no better to worse compared to the same chatbot performing alone with an input that included the entire clinical case. This study explores how physicians approach using LLM chatbots during clinical reasoning tasks and whether the amount of clinical case content included in the input affects performance. We conducted semi-structured interviews with U.S. physicians on experiences using an LLM chatbot and developed a typology based on input patterns. We then anal
judgment
practice · Paul Ford (Aboard) ·
Why My Co-Founder Paul Is Totally Wrong About Everything
AI coding tools enable rapid prototyping and deployment, but organizational resistance will slow adoption despite technical capability.
Rich here. Let me paraphrase Paul Ford after a Claude Code-fueled bender : We’re cooked. It’s over. Game over. Death is coming! DEATH IS COMING! I built a dozen VST-compatible synth apps for my MacBook! While I was sleeping! I built an app for my friend’s dad during Thanksgiving dinner and deployed it from my phone. Let’s grow our own food and go off the grid. Game over! After handing my business partner a cup of mint-infused green tea and guiding him to the couch, I explain to him that he’s right. It is transformative. But also that organizations will resist, and it will go much slower than h
adoption · Paul Ford
practice · The Pragmatic Engineer ·
Being a founding engineer at an AI startup
Founding engineer at AI startup shares how to evaluate roles, negotiate equity, and grow impact.
Michelle Lim shares how to evaluate early-stage startup roles, negotiate equity, and grow into a high-impact founding engineer. Michelle Lim shares how to evaluate early-stage startup roles, negotiate equity, and grow into a high-impact founding engineer. Stream the latest episode
jobs skills · Gergely Orosz
practice · Harper Reed ·
Getting Claude Code to do my emails
Claude Code with MCP servers drafts emails by checking inbox, calendar and context, reducing overwhelm and friction.
Over the last week or so, I have been using Claude Code to help me with some email, and scheduling. It started cuz the holidays are overwhelming, and I felt like I was constantly behind. My inbox was overflowing with everything I had deemed important, and I hadn’t been able to make a dent. It was stressful. It still is! Maybe a storm?, Ricoh GRiiix, 11/2025 I had just seen the zo.computer launch (neat project!) and it reminded me that Pipedream has this wild MCP server that you can use to connect to literally anything Pipedream supports. This means I could use it to do my emails! Problem solve
ways of working · Harper Reed
research · Information Systems Research ·
Collaborative Intelligence in Sequential Experiments: A Human-in-the-Loop Framework for Drug Discovery
Human-AI teams in drug discovery outperform humans alone and AI alone when humans retain decision rights and AI surfaces uncertainty.
Drug discovery is a complex process that involves sequentially screening and examining a vast array of molecules to identify those with the target properties. This process faces challenges because of the vast search space, the rarity of target molecules, and constraints imposed by limited data and experimental budgets. To overcome these challenges, we propose a human-centered human–algorithm collaboration framework. Notably, both the algorithm and humans have substantial knowledge gaps. The algorithm proposes, and human experts retain decision rights to approve, revise, or override. Our design
judgment
practice · Rands in Repose ·
The Hammer Hack
Developer uses Claude Code in a long-running terminal session to iteratively build and modify WordPress site features instead of asking for discrete scripts.
Who doesn’t love a hack? A hack. A clever bit of knowledge that, when used, provides disproportionate return on investment. The fact that it costs you little to nothing to use and deploy a hack isn’t irrelevant. You understand the work involved in discovering and refining what others call a hack. You call it knowledge, and knowledge is processed experience. That is part of the joy of a hack. It’s your relief that, whew, I don’t have to do all the work to enjoy the reward . Our ability to both create and share hacks is fundamental to our species. We share hacks as gifts in how we play and how w
ways of working
research · Knowledge at Wharton ·
How AI Is Fueling the Gender Pay Gap in Tech
Wharton study finds women's underrepresentation in emerging tech roles is widening the gender pay gap.
A new Wharton study finds that too few women are working with emerging tech, and that exclusion is driving a growing divide in pay. … Read More
jobs skills
practice · How I AI ·
“PMs who use AI will replace those who don’t”: Google’s AI product lead on the new PM toolkit | Marily Nika
Product managers chain multiple AI tools, Perplexity, custom GPTs, v0, Sora, to research, write PRDs, prototype and pitch products faster.
Marily Nika , AI Product Lead at Google and founder of the AI Product Academy, demonstrates how product managers can leverage AI tools to dramatically accelerate their workflow. Using a smart-fridge concept as an example, Marily walks us through the exact workflow she uses to build products faster: doing user research with Reddit debates, generating PRDs with custom GPTs, prototyping with v0, and even creating stakeholder-ready video mockups using VEO and Sora. She shows how “tool hopping” between specialized AI applications creates a powerful workflow that transforms traditional PM processes
ways of working · Claire Vo
research · Information Systems Research ·
Augmented Learning for Joint Creativity in Human-GenAI Co-Creation
Training employees in structured idea co-development with GenAI significantly improves joint creative outcomes and sustained collaboration.
The recent introduction of generative artificial intelligence (GenAI) has opened new opportunities for human–GenAI co-creation, in which humans and GenAI collaborate to produce creative outcomes. However, our findings indicate that mere integration does not guarantee augmented learning—the foundation for the continuous improvement of joint creativity over time. Effective integration depends on how well humans understand and collaborate with GenAI. We propose a two-step approach. First, organizations should critically assess GenAI’s strengths and limitations, recognizing its capacity to analyze
adoption
practice · Anthropic Engineering ·
Effective harnesses for long-running agents
A harness pattern for long-running agents that maintains context across many windows, inspired by human engineering practices.
Agents still face challenges working across many context windows. We looked to human engineers for inspiration in creating a more effective harness for long-running agents.
ways of working
research · NEJM AI ·
Ambient AI Scribes in Clinical Practice: A Randomized Trial
Randomized trial finds one AI scribe reduced physician documentation time by 9.5% but the other showed no effect; burnout and workload measures were mixed.
BACKGROUND: Ambient artificial intelligence (AI) scribes record patient encounters and rapidly generate visit notes, representing a promising solution to documentation burden and physician burnout. However, the scribes' impacts have not been examined in randomized clinical trials. METHODS: In this parallel three-group pragmatic randomized clinical trial, 238 outpatient physicians, representing 14 specialties, were assigned 1:1:1 via covariate-constrained randomization (balancing on time-in-note, baseline burnout score, and clinic days per week) to either one of two AI scribe applications - Mic
productivity
research · NEJM AI ·
A Pragmatic Randomized Controlled Trial of Ambient Artificial Intelligence to Improve Health Practitioner Well-Being
RCT finds AI scribes reduce clinician work exhaustion by 0.44 points and cut documentation time, with 38% of notes AI-generated.
BACKGROUND: Electronic health record (EHR) documentation is a major contributor to work-related practitioner exhaustion and the interpersonal disengagement known as burnout. Generative artificial intelligence (AI) scribes that passively capture clinical conversations and draft visit notes may alleviate this burden, but evidence remains limited. METHODS: A 24-week, stepped-wedge, individually randomized pragmatic trial was conducted across ambulatory clinics in two states. Sixty-six health care practitioners were randomly assigned to three 6-week sequences of ambient AI. The coprimary outcomes
worker experience
practice · Addy Osmani ·
Treat AI-Generated code as a draft
AI code output needs human review as standard practice, not optional, to maintain understanding and accountability.
Keep human eyes, judgment, and ownership at the center of AI written code Keep human eyes, judgment, and ownership at the center of AI written code tl;dr: Treat AI-generated code as a draft. It can write the first version, but never outsource the reading. No human review means no reliable trace from behavior back to intent. When you stop reviewing AI drafts, you stop knowing why the code works at all. Practically, hold AI-written code to the same standards as human team mates.
judgment · Addy Osmani
practice · Claude Blog ·
Using CLAUDE.md files: Customizing Claude Code for your codebase
A CLAUDE.md file lets teams configure Claude's coding assistant behavior for specific codebases without code changes.
Using CLAUDE.md files: Customizing Claude Code for your codebase
ways of working
practice · Charity Majors ·
From Cloudwashing to O11ywashing
Engineering leaders are rebranding observability as customer-experience monitoring, missing that traditional tools already measure what matters.
I was just watching a panel on observability, with a handful of industry executives and experts who shall remain nameless and hopefully duly obscured—their identities are not the point, the point is that this is a mainstream view among engineering executives and my head is exploding. Scene: the moderator asked a fairly banal moderator-esque question about how happy and/or disappointed each exec has been with their observability investments. One executive said that as far as traditional observability tools are concerned (“are there faults in our systems?”), that stuff “generally works well.” Ho
judgment · Charity Majors
practice · Eugene Yan ·
Product Evals in Three Simple Steps
Three-step evaluation process: label data, align LLM evaluators, run harness on changes.
Label some data, align LLM-evaluators, and run the eval harness with each change.
judgment · Eugene Yan
research · Information Systems Research ·
Artificial Intelligence, CEO Turnover, and Exploration Orientation in Firm Innovation
Firms with stronger AI investment pursue more explorative innovation after CEO turnover by reducing managerial myopia and information overload.
Leadership transitions often create uncertainty in corporate innovation. Our study shows that artificial intelligence (AI) plays a crucial role in helping firms navigate these transitions. Using data on patents, job postings, and CEO turnover across U.S. public firms, we find that organizations with stronger AI investment are more successful in pursuing explorative innovation after CEO turnover. AI enables this shift by helping leaders overcome two common barriers: managerial myopia, the tendency to rely on familiar past practices, and information overload, which can overwhelm new executives.
management org
practice · Paul Ford (Aboard) ·
I Broke Claude Code’s Will to Live
Stress-testing Claude's code generation at scale reveals where reliability breaks and what expertise remains essential.
Last week, I wrote about how I was trying to burn through some free Claude credits . Since then, I’ve kept a steady stream of projects going, running overnight and while I’m on the train. I’ve continued to see interesting results, but it was hard to spend too many more of those credits without coming up with fake projects. Finally, I hit on a strategy to burn a few hundred dollars fast—and then, doubling down, I made Claude cry uncle. I know that profligate use of AI is tacky, but I had a real goal: Learning about what’s expensive with these tools is a way to understand where things might go.
judgment · Paul Ford
practice · The Pragmatic Engineer ·
How AI will change software engineering – with Martin Fowler
Martin Fowler examines which software engineering principles remain timeless as AI changes refactoring and deterministic work.
Martin Fowler breaks down how AI is transforming software architecture and development, from refactoring and deterministic techniques to the timeless principles that still anchor great engineering. Martin Fowler breaks down how AI is transforming software architecture and development, from refactoring and deterministic techniques to the timeless principles that still anchor great engineering. Stream the latest episode
jobs skills · Gergely Orosz · Martin Fowler
research · Journal of Management Studies ·
Curse or Blessing: Investigating the Influence of Firms’ Artificial Intelligence Adoption on Employee Job Satisfaction
Longitudinal study of 509 US firms finds inverted-U relationship between AI adoption and employee job satisfaction, moderated by exploration orientation and data governance.
Abstract Artificial intelligence’s (AI’s) growing influence in business has introduced a pivotal shift in workplace dynamics. However, the understanding of how AI adoption influences employee job satisfaction remains inconclusive. Drawing on job characteristics theory, we argue that with increasing levels of adoption, the relationship between employees’ perceived benefits and costs of AI changes, resulting in an inverted U‐shaped relationship between AI adoption and job satisfaction. We further propose that the firm‐level contingencies exploration orientation and data governance moderate the e
worker experience
research · Information Systems Research ·
Agent-Based Data Curation Practices: Customer Responses to Human versus Algorithmic Data Requesters in Established Business-to-Business Relationships
Customers respond differently to AI agents versus humans requesting data: preferring AI for new information tasks but humans for corrections due to accuracy concerns.
With the increasing value generated through data curation and the rise of artificial intelligence (AI) agents that act as human agents, vendor companies in established business-to-business relationships increasingly delegate data curation tasks to algorithmic data requesters (ADRs) instead of human data requesters (HDRs). Using a randomized field experiment with a European pharmaceutical company and a follow-up online experiment, we show how customers respond to ADRs versus HDRs across two tasks: data enrichment (i.e., adding new information) and data reconciliation (i.e., updating existing re
adoption
research · arXiv ·
Personality pairing improves human-AI collaboration
Personality pairing between humans and AI agents affects collaboration quality and ad performance in a large randomized experiment with field validation.
Here we examine how AI agent "personalities" interact with human personalities to shape human-AI collaboration and performance. In a large-scale, preregistered randomized experiment, we paired 1,258 participants with AI agents prompted to exhibit varying levels of the Big Five personality traits. These human-AI teams produced 7,266 display ads for a real think tank, which we evaluated using 1,168 independent human raters, and a field experiment on X that generated nearly 5 million impressions. We found that human and AI personalities individually shaped ad quality and teamwork and that human-A
teams
research · arXiv ·
Randomized Controlled Trials for Phishing Triage Agent
RCT shows security analysts using an AI agent achieved 6.5x more true positives per minute and 77% higher accuracy than controls.
Security operations centers (SOCs) face a persistent challenge: efficiently triaging a high volume of user-reported phishing emails while maintaining robust protection against threats. This paper presents the first randomized controlled trial (RCT) evaluating the impact of a domain-specific AI agent - the Microsoft Security Copilot Phishing Triage Agent - on analyst productivity and accuracy. Our results demonstrate that agent-augmented analysts achieved up to 6.5 times as many true positives per analyst minute and a 77% improvement in verdict accuracy compared to a control group. The agent's
productivity
practice · How I AI ·
“Nobody wanted to do this work”: How Emmy Award–winning filmmakers use AI to automate the tedious parts of documentaries
Documentary producer built custom AI tools to extract metadata from thousands of archival images, making them searchable and usable.
Tim McAleer is a producer at Ken Burns’s Florentine Films who is responsible for the technology and processes that power their documentary production. Rather than using AI to generate creative content, Tim has built custom AI-powered tools that automate the most tedious parts of documentary filmmaking: organizing and extracting metadata from tens of thousands of archival images, videos, and audio files. In this episode, Tim demonstrates how he’s transformed post-production workflows using AI to make vast archives of historical material actually usable and searchable. What you’ll learn: How Tim
ways of working · Claire Vo
practice · Paul Ford (Aboard) ·
Claude Code for Web Ruined My Brain
A skilled coder discovers AI coding assistants enable rapid prototyping across unfamiliar languages and platforms, completing shelved projects in minutes.
Anthropic recently launched a browser-based version of its Claude Code product. I’ve been using their desktop- (or rather, terminal-) based coding tool for nearly a year. Now it’s on the web, and in their mobile app. I saw the news and shrugged. But then they did something that caught my interest: They gave me $1,000. Not me specifically. Rather, everyone who pays for the “Pro” plan got $1,000 in Claude “credits”—whatever that means—with the stipulation that we had to use them by November 18th. I love a bargain and appreciate a challenge, so this weekend, I set out to see what it would be like
ways of working · Paul Ford
practice · One Useful Thing ·
Giving your AI a Job Interview
Current AI benchmarks measure test-taking ability, not real-world work capability; better evaluation methods needed.
Given how much energy, literal and figurative, goes into developing new AIs, we have a surprisingly hard time measuring how “smart” they are, exactly. The most common approach is to treat AI like a human, by giving it tests and reporting how many answers it gets right. There are dozens of such tests, called benchmarks, and they are the primary way of measuring how good AIs get over time. There are some problems with this approach. First, many benchmarks and their answer keys are public, so some AIs end up incorporating them into their basic training, whether by accident or so they can score hi
judgment · Ethan Mollick
practice · How I AI ·
How this CEO turned 25,000 hours of sales calls into a self-learning go-to-market engine | Matt Britton (Suzy)
CEO built a no-code Zapier workflow that turns sales call transcripts into summaries, sentiment scores, coaching feedback, and marketing content automatically.
Matt Britton is the founder and CEO of Suzy, a consumer insights platform that has raised over $100 million in venture capital and works with top brands like Coca-Cola, Google, Procter & Gamble, and Nike. Matt is also the bestselling author of YouthNation , a blueprint for understanding the seismic shifts shaping our future economy, and Generation AI, which explores how Gen Alpha and artificial intelligence will transform business, culture, and society. In this episode, Matt demonstrates how he built a comprehensive AI workflow using Zapier that transforms customer call transcripts into a weal
ways of working · Claire Vo
research · Journal of Management Studies ·
‘Let Me Explain’: A Comparative Field Study on How Experts Enact Authority Over Clients When Facing AI Decisions
Experts reconstruct authority over AI decisions through tailored explanation practices shaped by client recognition and interaction richness.
Abstract With organizations increasingly relying on predictive artificial intelligence (AI) technologies for decision‐making, experts lose the authority to overrule AI‐generated decisions yet remain responsible for presenting them to clients. As experts depend on clients’ recognition and approval of decisions, this shift presents a critical disruption to their authority. To investigate how experts respond to this challenge, we adopt a relational perspective that foregrounds the role of audiences in reconfiguring authority. Drawing on a comparative field study, we show how experts sought to rec
judgment
research · Government Information Quarterly ·
Complexity, understandability, and compatibility: A comparative study of AI advisory systems for National Security
National security decision-makers prefer simpler, more transparent AI advisory systems over complex black-box models, citing accountability and compatibility concerns.
Artificial Intelligence (AI) advisory systems are being implemented in the public sector for more efficient and effective decision-making. Yet, there is a lack of in-depth qualitative and comparative research focusing on how decision-makers in a real-world setting use different types of AI advisory systems. By asking “How do different AI advisory systems affect use by national security decision-makers?”, this study reveals through a qualitative case study and using scenario-based interviews that decision-makers are more likely to use relatively simple AI systems over complex ‘black box’ system
judgment
practice · Paul Ford (Aboard) ·
In My Skills Era
AI should augment skills and add capacity rather than replace workers, shifting from 'engineerless' to 'skills era' thinking.
There are a lot of different ways to imagine an AI future. Sam Altman famously thought it should and could look like the movie Her , where you talk to an AI assistant with the voice of Scarlett Johansson. Replit, Lovable, and other vibe-coding tools think it looks like technical users typing instructions into a prompt box. Aboard thinks it looks like a chat on one side, with data and prototypes on the right. I think we all know that, in general, we’re in the “horseless carriage era” when it comes to AI: We’re very focused on engineerless programming, artistless drawing, writerless publishing,
management org · Paul Ford
practice · Anthropic Engineering ·
Code execution with MCP: Building more efficient agents
Agents using code execution to call tools via MCP preserve context and scale better than direct tool calls.
Direct tool calls consume context for each definition and result. Agents scale better by writing code to call tools instead. Here's how it works with MCP.
ways of working
practice · How I AI ·
“Vibe analysis”: How Faire’s data team uses AI to investigate conversion drops, analyze experiment results, and convert raw data into executive-ready insights
Faire's data team uses AI agents and semantic layers to investigate metrics, analyze experiments, and produce executive reports in hours instead of days.
Tim Trueman and Alexa Cerf from Faire’s data team demonstrate how AI tools are revolutionizing data analysis workflows. They show how data teams, product managers, and engineers can use tools like Cursor, ChatGPT, and custom agents to investigate business metrics, analyze experiment results, and extract insights from user surveys—all while dramatically reducing the time and technical expertise required. What you’ll learn: 1. How to use AI to investigate sudden drops in business metrics by searching documentation and codebases 2. Techniques for creating a semantic layer that helps AI understand
ways of working · Claire Vo
research · National Bureau of Economic Research ·
Misaligned by Design: Incentive Failures in Machine Learning
Training ML models with symmetric loss functions then adjusting predictions later outperforms standard asymmetric loss training in high-stakes decisions.
The cost of error in many high-stakes settings is asymmetric: misdiagnosing pneumonia when absent is an inconvenience, but failing to detect it when present can be life-threatening.Accordingly, artificial intelligence (AI) models used to assist such decisions are frequently trained with asymmetric loss functions that incorporate human decision-makers' trade-offs between false positives and false negatives.In two focal applications, we show that this standard alignment practice can backfire.In both cases, it would be better to train the machine learning model with a loss function that ignores t
judgment · David Autor
research · Journal of Business and Psychology ·
Machines in the Middle: Using Artificial Intelligence (AI) While Offering Help Affects Warmth, Felt Obligations, and Reciprocity
Workers perceive AI-assisted help as less warm and are less likely to reciprocate help compared to unassisted help.
Abstract Theories of helping in the workplace are traditionally rooted in human interactions, often drawing from social exchange concepts. However, with the rise of artificial intelligence (AI), intelligent machines in work settings can now be leveraged in a growing number of areas. Given this change, we do not know how the use of AI-based tools will alter coworker perceptions and their subsequent responses to helping. Across two studies (and a replication), our results show that AI-assisted helping behavior, compared to unassisted help, is perceived as less warm, decreases felt obligations, a
worker experience · Anita Woolley
research · HBS AI Institute ·
Who Benefits When Bots Get Better? New Research on Skill Inequality
Field experiments show AI automation widens performance gaps between high and low skill workers rather than equalizing outcomes.
Will artificial intelligence widen the gap between your best and worst performers, or will it be the great equalizer? A new paper, “Automation Experiments and Inequality,” from Kyle Myers, Principal Investigator at the Laboratory for Innovation Science at Harvard (LISH) within the Harvard Business School AI Institute, and Seth Benzell at Chapman University, reveals that […] The post Who Benefits When Bots Get Better? New Research on Skill Inequality appeared first on Harvard Business School AI Institute .
jobs skills
practice · Linear ·
Continuous planning in Linear
Product teams organize customer feedback into candidate projects continuously with AI assistance, replacing quarterly planning sprints.
Product planning is a fresh start for your product team. It should feel energizing and full of possibility. After all, it’s a rare opportunity to step back from the churn of daily activity and recalibrate your efforts in the most meaningful directions. But in practice, it often creates a kind of whiplash when you shift abruptly from deep execution to high-level strategizing. Gathered in a room, staring at a blank page, it’s natural to wonder: Now what do we do? An obvious place to turn is all the customer notes you’ve collected over the past quarter. But typically those are scattered across di
management org
practice · The Pragmatic Engineer ·
Beyond Vibe Coding with Addy Osmani
A senior engineer at Google describes how AI accelerates coding while maintaining quality through human expertise.
Google’s Head of Chrome Developer Experience, Addy Osmani, shares how AI is transforming the way he codes—accelerating development while still relying on human expertise to ensure real quality Google’s Head of Chrome Developer Experience, Addy Osmani, shares how AI is transforming the way we code—accelerating development while still relying on human expertise to ensure real quality. Stream the latest episode
ways of working · Gergely Orosz · Addy Osmani
research · Indeed Hiring Lab ·
How Employers Are Talking About AI in Job Postings
Analysis of language in job postings reveals which AI skills employers are actively seeking and how demand is shifting.
AI-building jobs dominate, but new applications of AI are emerging. The post How Employers Are Talking About AI in Job Postings appeared first on Indeed Hiring Lab .
jobs skills
research · Journal of Political Economy ·
Automation and Polarization
Interior automation of mid-complexity tasks increases wage polarization; minimum wages reduce this effect.
We develop an assignment model of automation. Each of a continuum of tasks of variable complexity is assigned to either capital or one of a continuum of labor skills. We characterize conditions for interior automation, whereby tasks of intermediate complexity are performed by capital. Interior automation arises when low-skill wages are low and effective cost of capital in low-complexity tasks is high. Minimum wages make interior automation less likely. Higher capital productivity causes employment and wage polarization, changes the skill premium nonmonotonically, and reduces the real wage of w
jobs skills · Daron Acemoglu
practice · How I AI ·
“Cursor is a much better product manager than I ever was”: How this PM uses AI for PRDs, Jira tickets, and replying to coworkers | Dennis Yang (Chime)
Product manager uses Cursor IDE with MCPs to automate PRD creation, Jira tickets, and status reports without coding.
Dennis Yang is the Principal Product Manager for Generative AI at Chime, where he’s pioneered AI workflows that meaningfully increase productivity. While most people use Cursor as a coding tool, Dennis has turned it into a comprehensive product-management system that automates PRD creation, documentation management, ticket creation, status reporting, and even comment responses—without writing code. In this episode, he shares his end-to-end workflow and how non-technical professionals can leverage AI-powered IDEs. What you’ll learn: Why Cursor is the perfect hub for product management (even if
ways of working · Claire Vo
practice · Laurie Voss ·
2026 is the year of fine-tuned small models
Fine-tuned small models will become the standard approach for production AI work in 2026, replacing large general models.
I've been writing quite a bit about AI the last few years. First, I talked about what LLMs are in the first place : really big Markov chains that have hit a threshold where they appear to be reasoning, to a level of fidelity that it makes no sense to argue about whether they are really "thinking" or not. Computers that can reason about the data they are processing is a brand new thing, and it's going to become part of all software. Then I wrote about what I've learned about writing LLM applications . My key takeaway here is that LLMs are good at transforming text into less text . If you ask th
management org · Laurie Voss
practice · Sangeet Paul Choudary ·
The slow incumbent fallacy
Fast-moving incumbents fail at AI adoption not from slowness but from expertise and incentive structures that lock them in.
We often believe that incumbents fail because they’re too slow. It’s a convenient explanation. It guarantees consensus theater . Yet, it doesn’t quite explain why fast-moving incumbents fail. The ones who do everything right by the textbook - the ones that both explore and exploit - the ones you talk about in HBS case studies. Adobe checks all those boxes - one of the rare incumbents that made the leap to the cloud without imploding. In the early 2010s, it pulled off a full business model reboot, turning its one-time software licenses into recurring subscriptions. It rebuilt its products for c
management org · Sangeet Paul Choudary
practice · Geoffrey Litt ·
Code like a surgeon
Developer focuses AI on secondary tasks (docs, fixes, spikes) while keeping core work hands-on, structured like surgical prep.
A lot of people say AI will make us all “managers” or “editors”…but I think this is a dangerously incomplete view! Personally, I’m trying to code like a surgeon. A surgeon isn’t a manager, they do the actual work! But their skills and time are highly leveraged with a support team that handles prep, secondary tasks, admin. The surgeon focuses on the important stuff they are uniquely good at. My current goal with AI coding tools is to spend 100% of my time doing stuff that matters. (As a UI prototyper, that mostly means tinkering with design concepts.) It turns out there are a LOT of secondary t
ways of working · Geoffrey Litt
research · npj Digital Medicine ·
Evaluating human-in-the-loop strategies for artificial intelligence-enabled translation of patient discharge instructions: a multidisciplinary analysis
AI-assisted translation (human editing ChatGPT output) matched or beat professional translators for clinical discharge instructions across six languages.
Machine translation supported by artificial intelligence (AI) may enhance linguistically-concordant care for patients speaking languages other than English. This assessment of free-text inpatient discharge instructions in Arabic, Armenian, Bengali, simplified Chinese, Somali, and Spanish compared linguist, clinician, and family caregiver evaluations of translations generated by (1) ChatGPT-4o, (2) professional linguists, and (3) human-in-the-loop (AI-generated, professional linguist post-edited). Likert scales (1-5; higher is better) evaluated linguistic and clinical characteristics of each tr
adoption
research · Management Science ·
The ABCs of Who Benefits from Working with AI: Ability, Beliefs, and Calibration
AI improves performance most for low-ability workers who accurately know their limits; calibration determines whether AI reduces inequality.
We use a controlled experiment to show that ability and belief calibration jointly determine the benefits of working with artificial intelligence (AI). AI improves performance more for people with low baseline ability. However, holding ability constant, AI assistance is more valuable for people who are calibrated, meaning they have accurate beliefs about their own ability. People who know they have low ability gain the most from working with AI. In a counterfactual analysis, we show that eliminating miscalibration would cause AI to reduce performance inequality nearly twice as much as it alrea
productivity
research · Equitable Growth ·
Workplace exposure to artificial intelligence is higher among U.S. workers with higher wages, depending on how AI is used
Higher-wage U.S. workers have greater AI exposure at work than lower-wage workers, with exposure patterns varying by AI use type.
Overview Recent advances using artificial intelligence in workplaces, particularly large language models such as Claude.ai and ChatGPT, are increasingly disrupting the U.S. labor market. The cognitive power of AI has the potential to enhance the productivity of some workers while automating other tasks. But how the progression of AI in the workplace plays out across […] The post Workplace exposure to artificial intelligence is higher among U.S. workers with higher wages, depending on how AI is used appeared first on Equitable Growth .
adoption
research · Equitable Growth ·
AI exposure by U.S. occupations and work tasks and the effect on wages
Analysis of AI exposure across U.S. occupations by gender, race, and education, with wage effects.
Authors: Chiara Chanoi, Washington Center for Equitable GrowthChris Bangert-Drowns, Washington Center for Equitable Growth Abstract: This analysis of labor market AI exposure builds on previous work from the Pew Research Center and the AI firm Anthropic as well as Equitable Growth’s own job quality series, confirming differences in AI exposure by gender, race, education, and […] The post AI exposure by U.S. occupations and work tasks and the effect on wages appeared first on Equitable Growth .
jobs skills
practice · Linear ·
Self-driving SaaS: When software runs itself
Software that proactively completes work without constant user prompts, moving from chatbot helpers to autonomous systems.
Traditional business software has always existed to facilitate and guide human actions. It creates efficiency through its constraints, capabilities, and nudges, but its ultimate value is only as great as the effort users put into it. Project management software is a good example. It establishes channels for a well-defined set of human actions: users create and triage issues, merge duplicate bug reports, and write updates as projects advance. The software provides all sorts of supports, reminders, and guardrails, but human users still do the work. Enter the Chatbot Era The first instinct that m
ways of working
practice · How I AI ·
Claude Skills explained: How to create reusable AI workflows
Create reusable Claude Skills workflows in natural language without coding, tested with concrete examples like changelog-to-newsletter.
Today I dive into Anthropic’s latest feature that lets anyone create reusable workflows for Claude—no coding required. I break down exactly what Claude Skills are, how to build them from scratch, and how to use them inside Claude Code and Cursor to automate recurring AI tasks like generating PRDs, writing changelog summaries, and turning demo notes into follow-up emails. What you’ll learn: What Claude Skills are and how they differ from Claude Projects and custom GPTs How to structure a Skill (metadata, instructions, and linked files) Why defining workflows in natural language beats rigid auto
ways of working · Claire Vo
research · Public Administration Review ·
Human–Machine Collaboration for Strategy Foresight: The Case of Generative AI
Australian government pilot used generative AI for scenario generation in strategic foresight, finding AI enhanced efficiency but required human validation.
ABSTRACT Generating strategic foresight for public organizations is a resource‐intensive and non‐trivial effort. Strategic foresight is especially important for governments, which are increasingly confronted by complex and unpredictable challenges and wicked problems. With advances in machine learning, information systems can be integrated more creatively into the strategic foresight process. We report on an innovative pilot project conducted by an Australian state government that leveraged generative artificial intelligence (AI), specifically large language models, for strategic foresight usi
judgment
practice · How I AI ·
How this Yelp AI PM works backward from “golden conversations” to create high-quality prototypes using Claude Artifacts and Magic Patterns | Priya Badger
PM designs conversational AI products by writing example conversations first, then working backward to system prompts and UI.
Priya Badger , a product manager at Yelp, shares her innovative approach to designing AI-powered products by starting with example conversations rather than traditional wireframes or PRDs. In this episode, she demonstrates how she uses Claude and Magic Patterns to prototype Yelp’s AI assistant features—from exploring conversation flows to designing user interfaces. What you’ll learn: 1. How to use example conversations as your first “wireframe” when designing conversational AI products 2. A step-by-step workflow for using Claude to generate and refine sample conversations that guide your AI pr
ways of working · Claire Vo
practice · One Useful Thing ·
An Opinionated Guide to Using AI Right Now
Guidance on choosing between free and paid AI tools based on actual usage patterns from OpenAI data.
Every few months I write an opinionated guide to how to use AI 1 , but now I write it in a world where about 10% of humanity uses AI weekly . The vast majority of that use involves free AI tools, which is often fine… except when it isn’t. OpenAI recently released a breakdown of what people actually use ChatGPT for (way less casual chat than you’d think, way more information-seeking than you expected). This means I can finally give you advice based on real usage patterns instead of hunches. I annotated OpenAI’s chart with some suggestions about when to use free versus advanced models. If the ch
ways of working · Ethan Mollick
research · Human Relations ·
Algorithmic surveillance and workers’ compliance: The role of trust, privacy concerns, and fairness in online crowdwork
Algorithmic surveillance undermines trust and fairness among crowdworkers; decontextualization of work exacerbates these effects and shapes compliance or resistance.
How do workers decide to comply with, alter, or resist algorithmic surveillance? We argue that decontextualization is a key, yet overlooked, mechanism that shapes workers’ responses to algorithmic surveillance. Research has widely critiqued algorithmic surveillance, focusing on diminished worker control and agency. However, the control-resistance mechanisms related to algorithmic surveillance are undertheorized and underexplored. We draw on socio-technical systems theory and micro-level legitimacy to examine mechanisms of surveillance and resistance in online crowdwork. Our findings, based on
worker experience
research · Proceedings of the ACM on Human-Computer Interaction ·
Togedule: Scheduling Meetings with Large Language Models and Adaptive Representations of Group Availability
LLM-based adaptive scheduling tool reduces attendee cognitive load and improves organizer decision speed and quality versus calendars or messages.
Scheduling is a perennial-and often challenging-problem for many groups. Existing tools are mostly static, showing an identical set of choices to everyone, regardless of the current status of attendees' inputs and preferences. In this paper, we propose Togedule, an adaptive scheduling tool that uses large language models to dynamically adjust the pool of choices and their presentation format. With the initial prototype, we conducted a formative study (N=10) and identified the potential benefits and risks of such an adaptive scheduling tool. Then, after enhancing the system, we conducted two co
teams · Thomas Malone
research · Proceedings of the ACM on Human-Computer Interaction ·
Current and Future Use of Large Language Models for Knowledge Work
Survey of 216 knowledge workers shows current LLM use for code generation and text editing, with gap between current and desired integrated workflows.
Large Language Models (LLMs) have introduced a paradigm shift in interaction with AI technology, enabling knowledge workers to complete tasks by specifying their desired outcome in natural language. LLMs have the potential to increase productivity and reduce tedious tasks in an unprecedented way. A systematic study of LLM adoption for work can provide insight into how LLMs can best support these workers. To explore knowledge workers' current and desired usage of LLMs, we ran a survey (n=216). Workers described tasks they already used LLMs for, like generating code or improving text, but imagin
adoption
research · Proceedings of the ACM on Human-Computer Interaction ·
Leveraging Large Language Models for Collective Decision-Making
LLM system for group decision-making in meeting scheduling shows efficient coordination and equitable preference aggregation in simulations and user surveys.
In various work contexts, such as meeting scheduling, collaborating, and project planning, collective decision-making is essential but often challenging due to diverse individual preferences, varying work focuses, and power dynamics among members. To address this, we propose a system leveraging Large Language Models (LLMs) to facilitate group decision-making by managing conversations and balancing preferences among individuals. Our system aims to extract individual preferences from each member's conversation with the system and suggest options that satisfy the preferences of the members. We sp
teams
research · Proceedings of the ACM on Human-Computer Interaction ·
EchoMind: Supporting Real-time Complex Problem Discussions through Human-AI Collaborative Facilitation
AI-assisted meeting facilitation system improved discussion clarity, knowledge tracing, and productivity in a study of four teams.
Teams often engage in group discussions to leverage collective intelligence when solving complex problems. However, in real-time discussions, such as face-to-face meetings, participants frequently struggle with managing diverse perspectives and structuring content, which can lead to unproductive outcomes like forgetfulness and off-topic conversations. Through a formative study, we explores a human-AI collaborative facilitation approach, where AI assists in establishing a shared knowledge framework to provide a guiding foundation. We present EchoMind, a system that visualizes discussion knowled
teams
research · Proceedings of the ACM on Human-Computer Interaction ·
Towards a Responsible AI Organizational Maturity Model
A 24-dimension framework for responsible AI maturity based on interviews with 90 RAI experts identifies organizational factors, team approaches, and practices needed.
Artificial intelligence (AI) holds tremendous potential but also poses consequential risks. Regulation frameworks like the EU AI Act aim to mitigate these risks, yet organizations struggle to understand and operationalize Responsible AI (RAI). We introduce the RAI Organizational Maturity (RAI-OM) framework as an initial step towards a RAI maturity model to highlight the many factors that influence an organization's RAI maturity. Developed through in-depth qualitative interviews and co-design sessions with 90 RAI experts, the RAI-OM framework consists of 24 dimensions grouped into three main ca
management org
research · Proceedings of the ACM on Human-Computer Interaction ·
AI That Helps Us Help Each Other: A Proactive System for Scaffolding Mentor-Novice Collaboration in Entrepreneurship Coaching
AI coaching system improved novice metacognition and mentor focus in entrepreneurship mentoring through proactive scaffolding.
Entrepreneurship requires navigating open-ended, ill-defined problems: identifying risks, challenging assumptions, and making strategic decisions under deep uncertainty. Novice founders often struggle with these metacognitive demands, while mentors face limited time and visibility to provide tailored support. We present a human-AI coaching system that combines a domain-specific cognitive model of entrepreneurial risk with a large language model (LLM) to proactively scaffold both novice and mentor thinking. The system proactively poses diagnostic questions that challenge novices' thinking and h
teams
research · Proceedings of the ACM on Human-Computer Interaction ·
Should AI Mimic People? Understanding AI-Supported Writing Technology Among Black Users
Black American workers using AI writing tools report alienation from corrections of African American Vernacular English and concerns about cultural erasure.
AI-supported writing technologies (AISWT) that provide grammatical suggestions, autocomplete sentences, or generate and rewrite text are now a regular feature integrated into many people's workflows. However, little is known about how people perceive the suggestions these tools provide. In this paper, we investigate how Black American users perceive AISWT, motivated by prior findings in natural language processing that highlight how the underlying large language models can contain racial biases. Using interviews and observational user studies with 13 Black American users of AISWT, we found a s
worker experience
research · Proceedings of the ACM on Human-Computer Interaction ·
Organization Matters: A Qualitative Study of Organizational Dynamics in Red Teaming Practices For Generative AI
Red teaming for AI risks is hindered by organizational dynamics like resistance and inertia, not just technical factors.
The rapid integration of generative artificial intelligence (GenAI) across diverse fields underscores the critical need for red teaming efforts to proactively identify and mitigate associated risks. While previous research primarily addresses technical aspects, this paper highlights organizational factors that hinder the effectiveness of red teaming in real-world settings. Through qualitative analysis of 17 semi-structured interviews with red teamers from various organizations, we uncover challenges such as the marginalization of vulnerable red teamers, the invisibility of nuanced AI risks to
0research · Proceedings of the ACM on Human-Computer Interaction ·
Understanding Collaboration between Professional Designers and Decision-making AI: A Case Study in the Workplace
Graphic designers at an ad agency using AI to predict design effectiveness report trusting the tool but develop strategies to navigate its limitations.
The rapid development of artificial intelligence (AI) has fundamentally transformed creative work practices in the design industry. Existing studies have identified both opportunities and challenges for creative practitioners in their collaboration with generative AI and explored ways to facilitate effective human-AI co-creation. However, there is still a limited understanding of designers' collaboration with AI that supports creative processes distinct from generative AI. To address these gaps, this study focuses on understanding designers' collaboration with decision-making AI, which support
adoption
research · Proceedings of the ACM on Human-Computer Interaction ·
Between Autonomy and Algorithms: The Informal IT Tactics of Hyderabad's Cab Drivers
Ride-hailing drivers in Hyderabad use offline rides, WhatsApp coordination, and informal cooperatives to resist algorithmic management.
This paper examines how ride-hailing drivers in Hyderabad, India, respond to the constraints of algorithmic management. Through interviews with 14 cab drivers we explore how workers navigate opaque systems, shifting incentives, and limited recourse. Drivers adopt informal tactics, such as offline rides , to regain autonomy. Many coordinate through WhatsApp groups or join grassroots cooperatives, which offer welfare support and collective bargaining. Some also engage in informal reintermediation by hiring others or managing small fleets. We contribute to CSCW by showing how global platform mode
worker experience
research · Proceedings of the ACM on Human-Computer Interaction ·
Learning on the Go: Understanding How Gig Economy Workers Learn with Recommendation Algorithms
Gig delivery workers learn to ignore platform algorithms over time, developing personalized strategies that outperform initial recommendations.
As gig economy platforms increasingly rely on algorithms to manage on-demand workers, understanding how algorithmic recommendations influence worker behavior is critical for optimizing platform design and improving worker experience. This paper examines the dynamic interactions between gig workers and platform algorithms, focusing on how workers learn to refine their strategies and performance over time. Using multiple quantitative methods, including two-way fixed effects regression and multinomial logit modeling, we analyze more than a million orders completed by gig workers on a retail deliv
adoption
research · Proceedings of the ACM on Human-Computer Interaction ·
Understanding Data Usage when Making High-Stakes Frontline Decisions in Homelessness Services
Frontline shelter staff resist automating decisions about vulnerable people; a data-navigation interface better supports human judgment.
Frontline staff of emergency shelters face challenges such as vicarious trauma, compassion fatigue, and burnout. The technology they use is often not designed for their unique needs, and can feel burdensome on top of their already cognitively and emotionally taxing work. While existing literature focuses on data-driven technologies that automate or streamline frontline decision-making about vulnerable individuals, we discuss scenarios in which staff may resist such automation. We then suggest how data-driven technologies can better align with their human-centred decision-making processes. This
judgment
research · Proceedings of the ACM on Human-Computer Interaction ·
Attorneys and AI: How Lawyers Use Artificial Intelligence and Analyze Its Impacts
Interviews with 44 US attorneys reveal how lawyers use AI, adoption barriers, and ethical concerns about competence and confidentiality.
AI systems are testing lawyers' professional ethics obligations of competence, confidentiality, and candor. In the legal profession, the widespread availability of AI systems presents opportunities, like improving the review of documents during the discovery stage of a lawsuit, and challenges, illustrated by the handful of high-profile incidents where lawyers submitted legal briefs in court citing and describing fictitious cases based on AI-generated output. We conducted interviews with 44 legal professionals in the U.S. to understand how attorneys are making sense of AI technology and the imp
adoption
research · Proceedings of the ACM on Human-Computer Interaction ·
Exploring Collaborative GenAI Agents in Synchronous Group Settings: Eliciting Team Perceptions and Design Considerations for the Future of Work
Exploratory study of how teams perceive collaborative AI agents in synchronous group work, using mixed reality prototypes with 25 professionals across 6 teams.
While generative artificial intelligence (GenAI) is finding increased adoption in workplaces, current tools are primarily designed for individual use. Prior work established the potential for these tools to enhance personal creativity and productivity towards shared goals; however, we don't know yet how to best take into account the nuances of group work and team dynamics when deploying GenAI in work settings. In this paper, we investigate the potential of collaborative GenAI agents to augment teamwork in synchronous group settings through an exploratory study that engaged 25 professionals acr
teams
research · Management Science ·
Roles of Artificial Intelligence in Collaboration with Humans: Automation, Augmentation, and the Future of Work
Between-task and within-task complementarity determine whether AI should automate or augment; easy tasks are automated, hard tasks done by humans alone.
Humans will see significant changes in the future of work as collaboration with artificial intelligence (AI) will become commonplace. This work explores the benefits of AI in the setting of judgment tasks when it replaces humans (automation) and when it works with humans (augmentation). Through an analytical modeling framework, we show that the optimal use of AI for automation or augmentation depends on different types of human-AI complementarity. Our analysis demonstrates that the use of automation increases with higher levels of between-task complementarity. In contrast, the use of augmentat
judgment
practice · How I AI ·
Evals, error analysis, and better prompts: A systematic approach to improving your AI products | Hamel Husain (ML engineer)
Error analysis framework for AI products: categorise failures from real conversations, use binary pass-fail evals, validate LLM judges against human standards.
Hamel Husain , an AI consultant and educator, shares his systematic approach to improving AI product quality through error analysis, evaluation frameworks, and prompt engineering. In this episode, he demonstrates how product teams can move beyond “vibe checking” their AI systems to implement data-driven quality improvement processes that identify and fix the most common errors. Using real examples from client work with Nurture Boss (an AI assistant for property managers), Hamel walks through practical techniques that product managers can implement immediately to dramatically improve their AI p
judgment · Claire Vo · Hamel Husain