The Feed · Complete archive
AI at work: research and practice
A server-rendered record of the evidence, ideas and firsthand practices screened by The Feed. Every entry links to its original source.
Use the interactive FeedPage 9 of 9 · 873 items
practice · Understanding AI (Timothy B. Lee) ·
We're Finding Out What Humans are Bad At
AI improves fastest in domains where humans use unnatural techniques that align with how machines think.
AI Advances Fastest When We Find Unnatural Ways of Doing Things AI Advances Fastest When We Find Unnatural Ways of Doing Things Magnus Carlsen is widely considered to be the greatest chess player of all time. An obvious prodigy (his Wikipedia page mentions solving 50-piece jigsaw puzzles when he was 2 years old), he earned the rank of grandmaster at age 13, despite (or because of?) being initially self-taugh…
judgment · Timothy B. Lee
research · JAMA Network Open ·
Clinician Experiences With Ambient Scribe Technology to Assist With Documentation Burden and Efficiency
Ambient scribing reduced time in notes per appointment and after-hours work, with clinicians reporting lower documentation burden.
Importance: Timely evaluation of ambient scribing technology is warranted to assess whether this technology can lessen the burden of clinical documentation on clinicians. Objective: To investigate the association of ambient scribing technology with efficiency, quality, and perceived burden of clinical documentation in the outpatient setting. Design, Setting, and Participants: This prospective, single-group pre-post quality improvement study was conducted between April and June 2024 in the outpatient setting of an academic health system in Philadelphia, Pennsylvania. Participants included physi
productivity
practice · Latent Space ·
The Inventors of Deep Research
DeepMind's Deep Research agent moves from search results to fully cited reports, showing a working use case for AI agents in knowledge work.
DeepMind's Aarush Selvan and Mukund Sridhar on creating the killer agent usecase, going from 10 blue links to fully cited reports, and building ontologies of AI use cases DeepMind's Aarush Selvan and Mukund Sridhar on creating the killer agent usecase, going from 10 blue links to fully cited reports, and building ontologies of AI use cases The free livestreams for AI Engineer Summit are now up! Please hit the bell to help us appease the algo gods. We’re also announcing a special Online Track later today.
ways of working · Shawn Wang (swyx)
research · Management Science ·
Till Tech Do Us Part: Betrayal Aversion and Its Role in Algorithm Use
Workers show betrayal aversion to human experts but not algorithms, even when risk is identical, affecting advice-following and earnings.
Failing to follow expert advice can have real and dangerous consequences. While any number of factors may lead a decision maker to refuse expert advice, the proliferation of algorithmic experts has further complicated the issue. One potential mechanism that restricts the acceptance of expert advice is betrayal aversion, or the strong dislike for the violation of trust norms. This study explores whether the introduction of expert algorithms in place of human experts can attenuate betrayal aversion and lead to higher overall rates of seeking expert advice. In other words, we ask: are decision ma
judgment
practice · Harper Reed ·
My LLM codegen workflow atm
Developer workflow: brainstorm spec, plan-the-plan, then discrete LLM codegen loops for greenfield and legacy code.
tl:dr; Brainstorm spec, then plan a plan, then execute using LLM codegen. Discrete loops. Then magic. ✩₊˚.⋆☾⋆⁺₊✧ I have been building so many small products using LLMs. It has been fun, and useful. However, there are pitfalls that can waste so much time. A while back a friend asked me how I was using LLMs to write software. I thought “oh boy. how much time do you have!” and thus this post. (p.s. if you are an AI hater - scroll to the end) I talk to many dev friends about this, and we all have a similar approach with various tweaks in either direction. Here is my workflow. It is built upon my o
ways of working · Harper Reed
research · Research Policy ·
A processual approach to skill changes in digital automation: The case of the platform economy in the service sector
Processual analysis of ride-hailing shows automation interrupts workers' micro-adaptations and repositions skills rather than replacing them, excluding workers from learning transferable skills.
We introduce the “processual approach” to skill changes in the current wave of digital automation, which imposes comprehensive and complex impacts on skills. The approach conceptualizes work as a set of processes, each consisting of a sequence of events. In each event, a worker and/or machine make judgments and take actions to move to the next event. The processual approach asks whether and how machines influence workers' judgments or actions during each event and interrupt or transform relationships between judgments and actions. The approach enables micro-to-middle-range, inductive theorizat
jobs skills
research · Strategic Management Journal ·
How does artificial intelligence improve human decision‐making? Evidence from the AI ‐powered Go program
Professional Go players improved move quality significantly after an AI program outperformed the best human player, with younger and less skilled players gaining most.
Abstract Research Summary We study how humans learn from artificial intelligence (AI), leveraging an introduction of an AI‐powered Go program (APG) that unexpectedly outperformed the best professional player. We compare the move quality of professional players to APG's superior solutions around its public release. Our analysis of 749,190 moves demonstrates significant improvements in players' move quality, especially in the early stages of the game where uncertainty is highest. This improvement was accompanied by a higher alignment with AI's suggestions and a decreased number and magnitude of
judgment
research · Computers in Human Behavior ·
Beyond efficiency: Trust, AI, and surprise in knowledge work environments
Real-time automated feedback increases trust in algorithmic performance ratings when task uncertainty is high, in a controlled experiment.
Contemporary management practices are often designed with the needs of knowledge-based workers in mind, but an increasingly pressing challenge today is how to manage and effectively handle non-routine work. This paper revisits the job characteristics model through the lens of self-determination theory, specifically in the context of algorithmic performance management. Non-routine work is inherently unpredictable, and individuals often struggle with prolonged uncertainty. However, automated interventions that help individuals make sense of their work in uncertain conditions may help overcome th
judgment · Anita Woolley
practice · Artificial Ignorance ·
Two years of Artificial Ignorance
A writer's reflections on what they learned about AI and work during a year of sustained observation and writing.
A belated 2024 year in review. A belated 2024 year in review. This week marks the second anniversary of Artificial Ignorance. 2024 was my first full calendar year of biweekly writing, and it felt like I started to find my footing as a writer (creator? Substacke…
ways of working · Charlie Guo
practice · Hacker News ·
How I use LLMs as a staff engineer
Staff engineer shares specific ways they use LLMs daily for code review, design docs, and unblocking teammates.
248 points on Hacker News. Discussion: https://news.ycombinator.com/item?id=42938409
ways of working
research · Organization Science ·
Algorithmic Recommendation Tools and Experiential Learning in Clinical Care
Clinical decision support adoption weakens the relationship between doctor experience and patient outcomes, especially in routine tasks.
This study examines the relationship between the adoption of algorithmic recommendation tools and experiential learning. We argue that the adoption of an algorithmic recommendation tool will harm experiential learning in organizations by limiting knowledge retention and retrieval. We further argue that the adverse relationship between algorithmic tool adoption and experiential learning will be stronger in organizations operating in low-task-difficulty environments than those in high-task-difficulty ones because organizational members in such organizations are likely to rely more on algorithmic
adoption
research · NBER ·
The Labor Market Impact of Digital Technologies
Regional variation in Korea's digital tech adoption shows significant labour market impacts from AI, big data, and IoT.
We investigate the impact of digital technology on employment patterns in Korea, where firms have rapidly adopted digital technologies such as artificial intelligence (AI), big data, and the internet of things (IoT). By exploiting regional variations in technology exposure, we find significant (Sangmin Aum , Yongseok Shin)
jobs skills
research · NBER ·
Artificial Intelligence and the Labor Market
Tasks with higher AI exposure experience reduced labour demand in subsequent periods, measured through NLP analysis of task content across firms and occupations 2010-2023.
We use advances in natural language processing to construct new measures of workers task-level exposure to artificial intelligence (AI) and machine learning from 2010 to 2023, capturing variation across firms, occupations, and time. Tasks with higher AI exposure subsequently experience reduced labor (Menaka Hampole , Dimitris Papanikolaou , Lawrence D.W. Schmidt , Bryan Seegmiller)
jobs skills
research · NBER ·
AI and Women's Employment in Europe
Female employment share increased in occupations with more AI-enabled technology diffusion across 16 European countries, 2011-2019.
We examine the link between the diffusion of artificial intelligence (AI) enabled technologies and changes in the female employment share in 16 European countries over the period 2011-2019. Using data for occupations at the 3-digit level, we find that on average female employment shares increased in (Stefania Albanesi , António Dias da Silva , Juan F. Jimeno , Ana Lamo , Alena Wabitsch)
jobs skills
research · Management Science ·
Who Is AI Replacing? The Impact of Generative AI on Online Freelancing Platforms
ChatGPT reduced job posts for writing and coding freelancers by 21% within eight months; image AI reduced image jobs by 17%.
This paper studies the impact of generative artificial intelligence (AI) technologies on the demand for online freelancers using a large data set from a leading global freelancing platform. We identify the types of jobs that are more affected by generative AI and quantify the magnitude of the heterogeneous impact. Our findings indicate a 21% decrease in the number of job posts for automation-prone jobs related to writing and coding compared with jobs requiring manual-intensive skills within eight months after the introduction of ChatGPT. We show that the reduction in the number of job posts in
jobs skills
research · HBS AI Institute ·
The Creative Edge: How Human-AI Collaboration is Reshaping Problem-Solving
Human-AI collaboration changes how organizations solve creative problems, but the actual findings are not detailed here.
As artificial intelligence (AI) capabilities rapidly advance, organizations are exploring new ways to leverage these technologies for creative problem-solving and innovation. A recent HBS working paper, “The Crowdless Future? Generative AI and Creative Problem Solving”, – by Léonard Boussioux, Assistant Professor at the University of Washington; Jacqueline N. Lane, Assistant Professor at Harvard Business School […] The post The Creative Edge: How Human-AI Collaboration is Reshaping Problem-Solving appeared first on Harvard Business School AI Institute .
adoption · Léonard Boussioux
practice · Paul Ford (Aboard) ·
Use Worse Tools
Using a weaker AI model teaches more about how coding with AI actually works than optimizing for the best tool.
I’ve been on this big quest to learn how to write software with AI. Not so much because I need to write code; I’m a co-founder, which means I’m in sales. But I used to be a programmer, and I do want to fully understand how these new systems work, and their limitations. It’s surprisingly hard to get good guidance here. The space is evolving very quickly. You just have to try stuff. For a while I was using Claude and ChatGPT, which are the big players in the space. I used them through “ aider ,” an intermediary program that runs in the terminal on your computer. It sits between you and the chatb
ways of working · Paul Ford
research · Journal of Labor Economics ·
Profits of Prejudiced Algorithms
Firms profit more from discriminatory algorithms when training data reflects human prejudice, but affirmative action can reverse this incentive.
Firms are starting to replace humans with algorithms in important screening decisions, but there are potential spillovers of human biases contained in datasets to subsequent algorithmic predictions. When these biases are motivated by human prejudices, there are risks of algorithms perpetuating discrimination. I prove that when datasets are generated by a sufficiently discriminatory human, firms are more profitable when training discriminatory algorithms. If instead enough affirmative action is instituted in favor of a disadvantaged group, firms are more profitable when training algorithms that
jobs skills
practice · Laurie Voss ·
What I've learned about writing AI apps so far
LLMs excel at text summarisation and transformation; builders should design around this strength, not fight it.
I started writing a post called "how to write AI apps" but it was over-reach so I scaled it back to this. Who am I to tell you how to write anything? But here's what I'll be applying to my own writing of AI-powered apps, specifically LLM applications. A battle I've already lost is that we shouldn't call LLMs "AI" at all; they are machine learning and not the general intelligence that is implied to the layman by the name. It is an even less helpful name than "serverless", my previous candidate for worst technology name. But alas, we're calling LLMs AI and any parts of the field that are not LLM
ways of working · Laurie Voss
research · HBS AI Institute ·
The Future of Decision-Making: How Generative AI Transforms Innovation Evaluation
Field experiment tests how generative AI affects evaluation of business ideas and innovations.
As businesses grapple with an ever-growing volume of ideas, products, and solutions to evaluate, decision-making processes are being reshaped by artificial intelligence (AI). Generative AI, in particular, has emerged as a game-changer in creative problem-solving and evaluation, as demonstrated by a recent field experiment described in the working paper “The Narrative AI Advantage? A Field […] The post The Future of Decision-Making: How Generative AI Transforms Innovation Evaluation appeared first on Harvard Business School AI Institute .
judgment
practice · Harper Reed ·
New Media
Harper Reed automated a personal media dashboard pulling from Goodreads, Spotify and RSS feeds using AI-assisted development.
tl;dr: I added a media section to my website to track and display what I’m reading, listening to, and bookmarking online through automated data collection. visit it here Recently I spent a few hours hanging out with my friends Claude, Aider, and ChatGPT and added a media section to this site. I really enjoy AI-aided development (maybe a post for later). The media section has a log of all my books read (tracked from Goodreads), my recently saved tracks (tracked from Spotify), and my links (tracked from feedbin/netnewsreader). I started posting my links a month or so ago. I wanted to see how the
ways of working · Harper Reed
practice · Paul Ford (Aboard) ·
Software Theater
AI coding works best when you have it write tests alongside code, creating a feedback loop that catches its own errors.
Over the holiday I decided to see what it would take to code an actual application in AI. I wanted to make something pretty complex—a database-backed, component-driven web tool that calls external APIs and updates webpages in real time. I worked on it in the evenings and didn’t finish it, but for a good reason. Because I was worried it would become self-aware and take over the world! Just kidding—I am above the age of 15. I’ll explain the actual reason in a minute. Right now, a consensus about programming with AI is forming, which breaks down to (1) LLMs write pretty good code; (2) they write
ways of working · Paul Ford
practice · The Pragmatic Engineer ·
How AI-assisted coding will change software engineering: hard truths
How AI-assisted coding reshapes what software engineering means, with implications for expectations and skill requirements.
A field guide that also covers why we need to rethink our expectations, and what software engineering really is. A guest post by software engineer and engineering leader Addy Osmani A field guide that also covers why we need to rethink our expectations, and what software engineering really is. A guest post by software engineer and engineering leader Addy Osmani Hi, this is Gergely with a bonus issue of the Pragmatic Engineer Newsletter. In every issue, we cover topics related to Big Tech and startups through the lens of software engineers and engineering leaders. To get articles like this in y
jobs skills · Gergely Orosz · Addy Osmani
practice · Latent Space ·
AI Engineering for Art — with comfyanonymous, of ComfyUI
Image generation workflows are moving from text prompts to node-based directed acyclic graphs for more control.
Using models for "Art Engineering", building hard to use UIs, and how image generation is moving from text boxes to DAGs Using models for "Art Engineering", building hard to use UIs, and how image generation is moving from text boxes to DAGs Applications for the NYC AI Engineer Summit, focused on Agents at Work, are open!
ways of working · Shawn Wang (swyx)
research · Academy of Management Annals ·
Managing with Artificial Intelligence: An Integrative Framework
Integrative framework linking human-AI collaboration and algorithmic management literatures, showing how they examine complementary sides of managing with AI in organisations.
Managing with artificial intelligence (AI) refers to humans’ interaction with algorithms performing managerial tasks in organizations. Two literatures exploring this interaction—human-AI collaboration (HAIC) and algorithmic management (AM)—have focused on distinct managerial tasks: while HAIC examines executive decision-making, AM focuses on managerial control. This article presents a review of both literatures to identify opportunities for integration and advancement. We observe that HAIC’s and AM’s micro-level emphases on different managerial tasks have resulted in diverging conceptualizatio
management org · Sebastian Raisch
research · Journal of Applied Psychology ·
Whither bias goes, I will go: An integrative, systematic review of algorithmic bias mitigation.
Systematic review integrates computer science and organizational research on algorithmic bias mitigation in personnel assessment across model development stages.
Machine learning (ML) models are increasingly used for personnel assessment and selection (e.g., resume screeners, automatically scored interviews). However, concerns have been raised throughout society that ML assessments may be biased and perpetuate or exacerbate inequality. Although organizational researchers have begun investigating ML assessments from traditional psychometric and legal perspectives, there is a need to understand, clarify, and integrate fairness operationalizations and algorithmic bias mitigation methods from the computer science, data science, and organizational research
judgment
research · Proceedings of the National Academy of Sciences ·
The unequal adoption of ChatGPT exacerbates existing inequalities among workers
Women are 16 percentage points less likely to use ChatGPT at work; higher-earning workers adopt it first despite lower tenure.
We study the adoption of ChatGPT, the icon of Generative AI, using a large-scale survey linked to comprehensive register data in Denmark. Surveying 18,000 workers from 11 exposed occupations, we document that ChatGPT is widespread, especially among younger and less-experienced workers. However, substantial inequalities have emerged. Women are 16 percentage points less likely to have used the tool for work. Furthermore, despite its potential to lift workers with less expertise, users of ChatGPT earned slightly more already before its arrival, even given their lower tenure. Workers see a substan
adoption · Anders Humlum
research · The Quarterly Journal of Economics ·
Generative AI at Work
AI assistant in customer support increased productivity 15% on average, with larger gains for less experienced workers and evidence of improved learning and work experience.
Abstract We study the staggered introduction of a generative AI–based conversational assistant using data from 5,172 customer-support agents. Access to AI assistance increases worker productivity, as measured by issues resolved per hour, by 15% on average, with substantial heterogeneity across workers. The effects vary significantly across different agents. Less experienced and lower-skilled workers improve both the speed and quality of their output, while the most experienced and highest-skilled workers see small gains in speed and small declines in quality. We also find evidence that AI assi
productivity · Erik Brynjolfsson · Danielle Li · Lindsey Raymond
research · arXiv ·
Complement or substitute? How AI increases the demand for human skills
AI adoption increases demand for complementary skills like analytical thinking across both AI and non-AI roles, with wage premiums, while reducing demand for substitutable skills.
Artificial Intelligence (AI) is transforming the nature of work, yet there is limited empirical evidence on how it affects demand for human skills. This paper examines whether AI adoption increases the prevalence and value of human capabilities that complement technical AI skills, such as analytical thinking, resilience, or ethical judgment, within and beyond AI-intensive job roles. Using a dataset of nearly 30 million job postings from the US, the UK and Australia, between 2018 and 2024, we distinguish between internal effects (within AI roles) and external effects (in non-AI roles) across co
jobs skills
practice · Geoffrey Litt ·
AI-generated tools can make programming more fun
Use AI to build custom debugging UIs that make manual coding work more enjoyable and efficient, not to automate coding itself.
I want to tell you about a neat experience I had with AI-assisted programming this week. What’s unusual here is: the AI didn’t write a single line of my code. Instead, I used AI to build a custom debugger UI … which made it more fun for me to do the coding myself. * * * I was hacking on a Prolog interpreter as a learning project. Prolog is a logic language where the user defines facts and rules, and then the system helps answer queries. A basic interpreter for this language turns out to be an elegant little program with surprising power—a perfect project for a fun learning experience. The trou
ways of working · Geoffrey Litt
practice · Anthropic Engineering ·
Building effective agents
Successful LLM agent implementations use simple, composable patterns rather than complex frameworks.
We've worked with dozens of teams building LLM agents across industries. Consistently, the most successful implementations use simple, composable patterns rather than complex frameworks.
ways of working
practice · Artificial Ignorance ·
Stop begging for JSON
Use OpenAI Structured Outputs to reliably parse LLM responses as valid JSON without brittle retry logic or prompt engineering.
How OpenAI's Structured Outputs makes building with AI much more reliable. How OpenAI's Structured Outputs makes building with AI much more reliable. If you've ever dealt with LLMs in production, you've probably ended up with some variation of this prompt:
ways of working · Charlie Guo
practice · Paul Ford (Aboard) ·
Nobody Knows Anything (Still)
Curated links showing conflicting signals about AI adoption and impact across sectors and geographies.
We are in the wildest moment. Sometimes you just have to let go of the anxiety and marvel at the level of realignment in the tech world right now. I recently saw a solid deck on the state of AI , by Benedict Evans, that had this slide: Well that narrows it down! No really: Who knows anything? Ethan Mollick? Simon Willison? Sam Altman? Some random Redditor? Oprah Winfrey? The CEO of ServiceNow? The outgoing Biden Administration? The incoming Musk administration? In that spirit, some fresh AI links to read, each of them promising clarity in their own way: Maybe it’s all hype? Only 13% of British
adoption · Ethan Mollick · Simon Willison · Paul Ford
practice · Hamel Husain ·
Building an Audience Through Technical Writing: Strategies and Mistakes
Technical writers build audiences by deeply engaging with others' work and adding original insights, not through product pitches or distribution tricks.
People often find me through my writing on AI and tech. This creates an interesting pattern. Nearly every week, vendors reach out asking me to write about their products. While I appreciate their interest and love learning about new tools, I reserve my writing for topics that I have personal experience with. One conversation last week really stuck with me. A founder confided, “We can write the best content in the world, but we don’t have any distribution.” This hit home because I used to think the same way. Let me share what works for reaching developers. Companies and individuals alike often
jobs skills · Hamel Husain · Eugene Yan
research · npj Digital Medicine ·
Impact of human and artificial intelligence collaboration on workload reduction in medical image interpretation
Meta-analysis of 36 studies shows AI reduces medical imaging workload by 27% when assisting clinicians, up to 62% when pre-screening.
Clinicians face increasing workloads in medical imaging interpretation, and artificial intelligence (AI) offers potential relief. This meta-analysis evaluates the impact of human-AI collaboration on image interpretation workload. Four databases were searched for studies comparing reading time or quantity for image-based disease detection before and after AI integration. The Quality Assessment of Studies of Diagnostic Accuracy was modified to assess risk of bias. Workload reduction and relative diagnostic performance were pooled using random-effects model. Thirty-six studies were included. AI c
productivity
practice · Artificial Ignorance ·
Case Study: Scaling customer intelligence
Team scaled customer call analysis from manual notes to 10,000 calls using Claude, extracting patterns at scale.
Analyzing 10,000 sales calls with Claude Analyzing 10,000 sales calls with Claude Question: how many sales calls can you listen to and take notes on in a day?
productivity · Charlie Guo
practice · Latent Space ·
How to Run a Paper Club
A structured guide for running a paper club as a team learning ritual, with templates and lessons from running 100+ papers.
Your ultimate Paper Club Starter Kit, from your friends at the Latent Space Paper Club, where we have now read >100 papers. Also: Announcing Latent Space Paper Club LIVE! at Neurips 2024! Join us! Your ultimate Paper Club Starter Kit, from your friends at the Latent Space Paper Club, where we have now read >100 papers. Also: Announcing Latent Space Paper Club LIVE! at Neurips 2024! Join us! We are excited to announce Latent Space LIVE at NeurIPS 2024! This will be the AI Engineer-focused remote+IRL event complementing NeurIPS with 3 categories:
teams · Eugene Yan
practice · Eugene Yan ·
How to Run a Weekly Paper Club (and Build a Learning Community)
Weekly paper club structure for teams to learn AI together: how to select, read, and facilitate discussions.
Benefits of running a weekly paper club, how to start one, and how to read and facilitate papers.
teams · Eugene Yan
research · JAMA Network Open ·
Artificial Intelligence and Radiologist Burnout
Radiologists using AI regularly report lower burnout rates, particularly when AI acceptance is high and workload is managed.
IMPORTANCE: Understanding the association of artificial intelligence (AI) with physician burnout is crucial for fostering a collaborative interactive environment between physicians and AI. OBJECTIVE: To estimate the association between AI use in radiology and radiologist burnout. DESIGN, SETTING, AND PARTICIPANTS: This cross-sectional study conducted a questionnaire survey between May and October 2023, using the national quality control system of radiology in China. Participants included radiologists from 1143 hospitals. Radiologists reporting regular or consistent AI use were categorized as t
worker experience
research · Journal of Empirical Legal Studies ·
Building a better lawyer: Experimental evidence that artificial intelligence can increase legal work efficiency
Experiment with 206 law students: AI highlighting reduced legal task completion time by 30% with no quality loss; summaries alone had no effect.
Abstract Rapidly improving artificial intelligence (AI) technologies have created opportunities for human–machine cooperation in legal practice. We provide evidence from an experiment with law students (N = 206) on the causal impact of machine assistance on the efficiency of legal task completion in a private law setting with natural language inputs and multidimensional AI outputs. We tested two forms of machine assistance: AI‐generated summaries of legal complaints and AI‐generated text highlighting within those complaints. AI‐generated highlighting reduced task completion time by 30% without
productivity
research · Strategic Management Journal ·
Generative artificial intelligence and evaluating strategic decisions
LLM evaluations of business models are inconsistent individually but aggregate to match human expert rankings when combined.
Abstract Research Summary Strategic decisions are uncertain and often irreversible. Hence, predicting the value of alternatives is important for strategic decision making. We investigate the use of generative artificial intelligence (AI) in evaluating strategic alternatives using business models generated by AI (study 1) or submitted to a competition (study 2). Each study uses a sample of 60 business models and examines agreement in business model rankings made by large language models (LLMs) and those by human experts. We consider multiple LLMs, assumed LLM roles, and prompts. We find that ge
judgment
research · Information Systems Research ·
Skill-Biased Technical Change, Again? Online Gig Platforms and Local Employment
TaskRabbit entry correlates with fewer supervisor roles but stable janitor jobs; middle-skilled workers shift to self-employment rather than unemployment.
This study explores the impact of online gig platforms like TaskRabbit on the employment of incumbent service workers, focusing on the housekeeping sector. It highlights how TaskRabbit’s entry correlates with a decrease in middle-skilled roles such as supervisors due to automation, whereas low-skilled jobs like janitors remain stable because of their manual nature. Notably, many middle-skilled workers transition to self-employment within the sector instead of facing layoffs. Amidst competing narratives about the gig economy’s influence on labor markets, this research offers critical insights i
jobs skills
research · Proceedings of the ACM on Human-Computer Interaction ·
"Guilds" as Worker Empowerment and Control in a Chinese Data Work Platform
Guilds on Chinese data platforms function as both worker empowerment and management control, improving coordination and standardization while flattening hierarchy.
Data work plays a fundamental role in the development of algorithmic systems and the AI industry. It is often performed in business process outsourcing (BPO) companies and crowdsourcing platforms, involving a global and distributed workforce as well as networks of collaborative actors. Previous work on community building among data workers centers organization and mutual support or focuses on the structuring and instrumentalization of crowdworker groups for complicated projects. We add to these lines of research by focusing on a specific form of community building encouraged and facilitated by
worker experience
research · Proceedings of the ACM on Human-Computer Interaction ·
"Something Fast and Cheap" or "A Core Element of Building Trust"? - AI Auditing Professionals' Perspectives on Trust in AI
AI auditors describe how trust in AI tools is built through audits, and how information asymmetry undermines user trust in AI systems.
Artificial Intelligence (AI) auditing is a relatively new area of work. Currently, there is a lack of uniform standards and regulation. As a result, the AI auditing ecosystem is very diverse, and AI auditing professionals use a variety of different auditing methods. So far, little is known about how AI auditors approach the concept of trust in AI through AI audits, in particular regarding the trust of users. This paper reports findings from interviews with 19 AI auditing stakeholders to understand how AI auditing professionals seek to create calibrated trust in AI tools and AI audits. Themes i
judgment
research · Proceedings of the ACM on Human-Computer Interaction ·
Entangled Independence: From Labor Rights to Gig "Empowerment" Under the Algorithmic Gaze
Gig platform ties worker benefits to algorithmic performance scores, creating conditional access to health insurance and loans.
This study offers a critical analysis of how the Indian gig platform, Urban Company (UC), mobilizes the promise of empowerment to assemble, discipline and algorithmically entangle its gig labor force. Much of the scholarship on gig work has described how the (mis)classification of gig workers as independent contractors allows gig corporations to abdicate their employment responsibilities, thereby disentangling the traditional employer-employee relationship. By contrast, UC promises to invest in it's Indian gig workforce, with offerings such as health insurance and financial loans to counter th
worker experience
research · Proceedings of the ACM on Human-Computer Interaction ·
Intermediation: Algorithmic Prioritization in Practice in Homeless Services
Social service workers use discretionary practices to work around algorithmic prioritization systems for housing allocation, preserving advocacy and autonomy.
Homelessness is a significant and growing crisis in the United States. In an effort to more efficiently and fairly distribute limited housing resources, jurisdictions across the US have adopted algorithmic prioritization systems to help select which unhoused people should receive resources. Given the impact of algorithmic prioritization on the lives of unhoused people, there is a need to more fully examine how these systems are implemented in practice by frontline workers such as social service workers. In this paper, we present a qualitative study that draws on interviews and artifact walkthr
judgment
research · Proceedings of the ACM on Human-Computer Interaction ·
(De)Noise: Moderating the Inconsistency Between Human Decision-Makers
Algorithmic pairwise comparisons and machine advice both reduce inconsistency in human estimates, with advice more effective but pairs more widely applicable.
Prior research in psychology has found that people's decisions are often inconsistent. An individual's decisions vary across time, and decisions vary even more across people. Inconsistencies have been identified not only in subjective matters, like matters of taste, but also in settings one might expect to be more objective, such as sentencing, job performance evaluations, or real estate appraisals. In our study, we explore whether algorithmic decision aids can be used to moderate the degree of inconsistency in human decision-making in the context of real estate appraisal. In a large-scale hum
judgment
research · Proceedings of the ACM on Human-Computer Interaction ·
The Algorithm and the Org Chart: How Algorithms Can Conflict with Organizational Structures
Ethnographic study shows algorithms expose tensions between org structures designed for human coordination and structures needed for algorithmic optimization.
Algorithms are introducing changes to individuals? jobs, but do algorithms also lead to changes in the structures of organizations themselves? Organizational structures, as often formalized into organization (org) charts, are meant to facilitate coordinated decision-making. Yet our 10-month ethnographic study of a large online retail company reveals why the organizational structures that facilitate effective decision-making by humans may be in tension with the organizational structures that facilitate effective decision-making using algorithms. Our findings show that the human decision-makers
management org · Melissa Valentine · Michael Bernstein
research · Proceedings of the ACM on Human-Computer Interaction ·
Constructing a Classification Scheme - and its Consequences: A Field Study of Learning to Label Data for Computer Vision in a Hospital Intensive Care Unit
Classification scheme design decisions shape annotation work and downstream AI systems more than annotator bias alone.
Research on data annotation for artificial intelligence (AI) has demonstrated that biases, power, and culture impact the ways that annotators apply labels to data and subsequently affect downstream AI systems. However, annotators can only apply labels that are available to them in the annotation classification scheme. Drawing on a 3-year ethnographic study of an R&D collaboration between medical and AI researchers, we argue that the construction of the classification schema itself -- decisions about what kinds of data can and cannot be collected, what activities can and cannot be detected in t
adoption · Melissa Valentine · Michael Bernstein
research · Proceedings of the ACM on Human-Computer Interaction ·
Code-ifying the Law: How Disciplinary Divides Afflict the Development of Legal Software
Teams of computer scientists made systematic legal errors translating bankruptcy law into software, even with legal experts present.
Proponents of legal automation believe that translating the law into code can improve the legal system. However, research and reporting suggest that legal software systems often contain flawed translations of the law, resulting in serious harms such as terminating children's healthcare and charging innocent people with fraud. Efforts to identify and contest these mistranslations after they arise treat the symptoms of the problem, but fail to prevent them from emerging. Meanwhile, existing recommendations to improve the development of legal software remain untested, as there is little empirical
judgment
research · Research-Technology Management ·
Designing Technology that Preserves Skill Development
Framework for designing AI systems that preserve skill development rather than deskilling workers during adoption.
Matt Beane is an academic who researches robotics and AI in the workplace. He wrote The Skill Code.New technology can greatly improve both the productivity and effectiveness of work. It can also in...
management org · Matt Beane
research · Management Science ·
Digital Lyrebirds: Experimental Evidence That Voice-Based Deep Fakes Influence Trust
Voice-cloned AI agents increase trust in investment games; AI disclosure does not reduce this effect.
We consider the pairing of audio chatbot technologies with voice-based deep fakes, that is, voice clones, examining the potential of this combination to induce consumer trust. We report on a set of controlled experiments based on the investment game, evaluating how voice cloning and chatbot disclosure jointly affect participants’ trust, reflected by their willingness to play with an autonomous, AI-enabled partner. We observe evidence that voice-based agents garner significantly greater trust from subjects when imbued with a clone of the subject’s voice. Recognizing that these technologies pres
judgment
practice · Paul Ford (Aboard) ·
A Comforting Software Mess
Claude's computer use feature works but is janky and experimental; practical uses emerge alongside real limitations.
In a recent podcast , we discussed how Claude from Anthropic is now able to directly control a computer . “Developers can direct Claude to use computers the way people do—by looking at a screen, moving a cursor, clicking buttons, and typing text.” This feature is called…“computer use,” which is not a great name. They should have asked their own product to name it. This sounded pretty wild! Imagine if AI could open a web browser and fill out any form. Given that AI tech can now figure out captchas, it could be a recipe for cultural disaster/hackery—a zillion AI bots pasting AI-generated images
ways of working · Paul Ford
practice · Hamel Husain ·
Using LLM-as-a-Judge For Evaluation: A Complete Guide
Framework for setting up LLM-based evaluation systems, avoiding common mistakes like arbitrary scoring scales and unvalidated metrics.
Earlier this year, I wrote Your AI product needs evals . Many of you asked, “How do I get started with LLM-as-a-judge?” This guide shares what I’ve learned after helping over 30 companies set up their evaluation systems. The Problem: AI Teams Are Drowning in Data Ever spend weeks building an AI system, only to realize you have no idea if it’s actually working? You’re not alone. I’ve noticed teams repeat the same mistakes when using LLMs to evaluate AI outputs: Too Many Metrics : Creating numerous measurements that become unmanageable. Arbitrary Scoring Systems : Using uncalibrated scales (like
judgment · Hamel Husain
practice · Kent Beck ·
Background Work
LLMs make exploratory research and rabbit-hole investigation cheap enough to become routine work practice.
First published Dec 8, 2021. First published Dec 8, 2021. The topic of background work came up in the context of the way LLMs make such work so much cheaper than it used to be. The same rabbit hole that would have taken a week to explore can now be excavated in an hour or a minute. Background work should become more prevalent.
ways of working · Kent Beck
research · Nature Human Behaviour ·
When combinations of humans and AI are useful: A systematic review and meta-analysis
Meta-analysis of 106 experiments shows human-AI combinations underperform on average, with losses in decision tasks but gains in content creation.
Inspired by the increasing use of artificial intelligence (AI) to augment humans, researchers have studied human-AI systems involving different tasks, systems and populations. Despite such a large body of work, we lack a broad conceptual understanding of when combinations of humans and AI are better than either alone. Here we addressed this question by conducting a preregistered systematic review and meta-analysis of 106 experimental studies reporting 370 effect sizes. We searched an interdisciplinary set of databases (the Association for Computing Machinery Digital Library, the Web of Science
judgment · Thomas Malone · Michelle Vaccaro · Abdullah Almaatouq
research · JAMA Network Open ·
Large Language Model Influence on Diagnostic Reasoning
Randomized trial finds LLM access did not improve physicians' diagnostic reasoning compared to conventional resources alone.
Importance: Large language models (LLMs) have shown promise in their performance on both multiple-choice and open-ended medical reasoning examinations, but it remains unknown whether the use of such tools improves physician diagnostic reasoning. Objective: To assess the effect of an LLM on physicians' diagnostic reasoning compared with conventional resources. Design, Setting, and Participants: A single-blind randomized clinical trial was conducted from November 29 to December 29, 2023. Using remote video conferencing and in-person participation across multiple academic medical institutions, ph
judgment
practice · Eugene Yan ·
AlignEval: Building an App to Make Evals Easy, Fun, and Automated
A workflow for building and iterating on custom LLM evaluators by labeling data, testing the evaluator, and optimizing it against those labels.
Look at and label your data, build and evaluate your LLM-evaluator, and optimize it against your labels.
judgment · Eugene Yan
practice · Latent Space ·
How NotebookLM Was Made
How Google iterated NotebookLM from a side project into a viral product by centering user workflows and managing disagreement.
How to manage and engineer truly great AI products, why disagreement makes for great podcasts, iterating your way to a viral hit from a "Talk to Small Corpus" side project How to manage and engineer truly great AI products, why disagreement makes for great podcasts, iterating your way to a viral hit from a "Talk to Small Corpus" side project If you’ve listened to the podcast for a while, you might have heard our ElevenLabs-powered AI co-host Charlie a few times. Text-to-speech has made amazing progress in the last 18 months, with OpenAI’…
management org · Shawn Wang (swyx)
research · HBS AI Institute ·
Balancing Algorithms and Human Expertise: Unlocking the True Potential of Data-Driven Decisions
Study finds algorithm effectiveness depends on decision authority structure; human-algorithm collaboration yields better outcomes than either alone.
In the rapidly advancing world of data-driven decision-making, algorithms hold tremendous promise for organizations, but their effectiveness depends on how they are applied. A recent study, “Decision Authority and the Returns to Algorithms,” by Hyunjin Kim, professor at INSEAD Business School, Edward L. Glaeser, professor at Harvard University, Andrew Hillis, head of data science and […] The post Balancing Algorithms and Human Expertise: Unlocking the True Potential of Data-Driven Decisions appeared first on Harvard Business School AI Institute .
judgment
practice · Latent Space ·
Building the Silicon Brain - with Drew Houston of Dropbox
Dropbox CEO describes AI engineering stack and UX patterns for Dash, an AI product serving 700M users.
Drew's AI Engineering stack, building AI for 700M users, Dropbox Dash and how AI UX will change the way we use computers Drew's AI Engineering stack, building AI for 700M users, Dropbox Dash and how AI UX will change the way we use computers CEOs of publicly traded companies are often in the news talking about their new AI initiatives, but few of them have built anything with it. Drew Houston from Dropbox is different; he has spent over …
ways of working · Shawn Wang (swyx)
research · Organization Science ·
More to Lose: The Adverse Effect of High Performance Ranking on Employees’ Preimplementation Attitudes Toward the Integration of Powerful AI Aids
High-performing employees resist AI adoption more when it threatens their competitive standing versus their confidence in their own ability.
Despite the growing availability of algorithm-augmented work, algorithm aversion is prevalent among employees, hindering successful implementations of powerful artificial intelligence (AI) aids. Applying a social comparison perspective, this article examines the adverse effect of employees’ high performance ranking on their preimplementation attitudes toward the integration of powerful AI aids within their area of advantage. Five studies, using a weight estimation simulation (Studies 1–3), recall of actual job tasks (Study 4), and a workplace scenario (Study 5), provided consistent causal evid
adoption
practice · Paul Ford (Aboard) ·
Building Is Better Than Chatting
When AI chat hits reliability limits, shift to having AI write executable code tools instead of seeking chat answers.
GenAI isn’t reliable. If you ask it to make a list of very big volcanos, it’ll make a solid list. But as you get more specific—asking it for individual volcano heights, or diameters, or other kinds of lava lore—it will produce equivalently specific results, but that specificity isn’t the same as accuracy . The more specific the results, the more you need to fact check. One of my hopes with this new technology is that I’ll be able to upload a file and do substantive things with it. For example, I’d like to upload a really long PDF and generate an index of concepts. In a perfect world I’d be abl
ways of working · Paul Ford
research · Management Science ·
A Machine Learning Framework for Assessing Experts’ Decision Quality
Machine learning framework estimates expert worker decision quality using limited ground truth data and historical decisions.
Expert workers make non-trivial decisions with significant implications. Experts’ decision accuracy is, thus, a fundamental aspect of their judgment quality, key to both management and consumers of experts’ services. Yet, in many important settings, transparency in experts’ decision quality is rarely possible because ground truth data for evaluating the experts’ decisions is costly and available only for a limited set of decisions. Furthermore, different experts typically handle exclusive sets of decisions, and thus, prior solutions that rely on the aggregation of multiple experts’ decisions f
judgment
research · Management Science ·
Large Language Model in Creative Work: The Role of Collaboration Modality and User Expertise
Experiment shows LLM collaboration mode matters: as 'sounding board' boosts nonexpert ad quality, but 'ghostwriter' role harms expert performance via anchoring.
Since the launch of ChatGPT in December 2022, large language models (LLMs) have been rapidly adopted by businesses to assist users in a wide range of open-ended tasks, including creative work. Although the versatility of LLM has unlocked new ways of human-artificial intelligence collaboration, it remains uncertain how LLMs should be used to enhance business outcomes. To examine the effects of human-LLM collaboration on business outcomes, we conducted an experiment where we tasked expert and nonexpert users to write an ad copy with and without the assistance of LLMs. Here, we investigate and co
adoption
practice · Paul Ford (Aboard) ·
Four Terror-Free Ways to Talk About AI
Workers use AI widely, but managers see little adoption or gains, a pattern seen before with Linux and websites.
One of the more thoughtful resources on what AI is doing to organizations is Ethan Mollick’s newsletter, “ One Useful Thing .” Mollick is a professor at the Wharton School at the University of Pennsylvania, which we must not hold against him. His newsletter is not academic in tone, but does have an academic thoroughness that I appreciate. His most recent edition was about AI adoption: A large percentage of people are using AI at work. We know this is happening in the EU, where a representative study of knowledge workers in Denmark from January found that 65% of marketers, 64% of journalists, 3
adoption · Ethan Mollick · Paul Ford
practice · Artificial Ignorance ·
How to build an AI search engine (Part 1)
Building an AI search engine by combining Claude, Brave Search API, and streaming responses to present cited answers.
Working with Claude, Brave, and streaming responses. Working with Claude, Brave, and streaming responses. I've long been a fan of Perplexity, the AI search engine that presents citations in-line with its answers. While not foolproof, it's a good way of mitigating hallucinations as I can easily check the …
ways of working · Charlie Guo
research · Equitable Growth ·
Estimating the prevalence of automated management and surveillance technologies at work and their impact on workers’ well-being
National survey estimates prevalence of AI monitoring and algorithmic management tools in US workplaces and their effects on worker wellbeing.
Evidence from a new national survey and implications for U.S. federal policy U.S. employers are deploying artificial-intelligence-powered and automated tools to monitor and manage workers, including tracking workers’ locations, activities, and productivity, and making decisions based on that data about workers’ schedules, tasks, compensation, promotions, and discipline. Researchers, workers’ rights leaders, journalists, and policymakers have […] The post Estimating the prevalence of automated management and surveillance technologies at work and their impact on workers’ well-being appeared firs
worker experience
research · Collective Intelligence ·
Supermind Ideator: How scaffolding Human-AI collaboration can increase creativity
Scaffolded LLM interface increased creative problem-solving output compared to ChatGPT or solo work in an experimental study.
Previous efforts to support creative problem-solving have included (a) techniques such as brainstorming and design thinking to stimulate creative ideas, and (b) software tools to record and share these ideas. Now, generative AI technologies can suggest new ideas that might never have occurred to the users, and users can then select from these ideas or use them to stimulate even more ideas. To explore these possibilities, we developed a system called Supermind Ideator that uses a large language model (LLM) and adds prompts, fine tuning, and a specialized user interface in order to help users re
ways of working · Thomas Malone
research · NBER ·
The ABC’s of Who Benefits from Working with AI: Ability, Beliefs, and Calibration
Controlled experiment shows AI benefits low-ability workers most, but calibrated beliefs about AI capability determine actual gains.
We use a controlled experiment to show that ability and belief calibration jointly determine the benefits of working with Artificial Intelligence (AI). AI improves performance more for people with low baseline ability. However, holding ability constant, AI assistance is more valuable for people who (Andrew Caplin , David J. Deming , Shangwen Li , Daniel J. Martin , Philip Marx , Ben Weidmann , Kadachi Jiada Ye)
productivity
research · Industrial and Labor Relations Review ·
Negotiating about Algorithms: Social Partner Responses to AI in Denmark and Sweden
Trade unions and employer organisations in Denmark and Sweden negotiated novel collective bargaining agreements addressing AI deployment and worker protections.
digital technology; collective bargaining; workplace; Denmark; Sweden
policy
research · NBER ·
Is Distance from Innovation a Barrier to the Adoption of Artificial Intelligence?
Job vacancies requiring AI skills grew slower in US regions farther from pre-2007 AI innovation hubs between 2007-2019.
Using our own data on Artificial Intelligence publications merged with Burning Glass vacancy data for 2007-2019, we investigate whether online vacancies for jobs requiring AI skills grow more slowly in U.S. locations farther from pre-2007 AI innovation hotspots. We find that a commuting zone which (Jennifer Hunt , Iain M. Cockburn , James Bessen)
jobs skills
practice · Latent Space ·
Language Agents: From Reasoning to Acting
ReAct, Tree of Thought, and CoALA frameworks show how agents reason through problems and act on them systematically.
Shunyu Yao on ReAct, Tree of Thought, CoALA, and the emerging importance of AI-Computer Interfaces, ft. returning guest + guest host Harrison Chase of LangChain/LangGraph! Shunyu Yao on ReAct, Tree of Thought, CoALA, and the emerging importance of AI-Computer Interfaces, ft. returning guest + guest host Harrison Chase of LangChain/LangGraph! OpenAI DevDay is almost here! Per tradition, we are hosting a DevDay pregame event for everyone coming to town! Join us with demos and gossip!
ways of working · Shawn Wang (swyx)