The Pragmatic Engineer — selected conversations
Measuring the impact of AI on software engineering – with Laura Tacho
The Pragmatic Engineer 5 of 15
In this collection Browse 15 summaries 5 of 15
Gergely Orosz, host of The Pragmatic Engineer podcast, interviews Laura Tacho, CTO at developer-productivity company DX, about measuring AI-assisted software engineering without reducing impact to generated code or accepted suggestions ([00:00:00]-[00:00:50], [00:08:54]-[00:16:56]). Tacho's central argument is that organizations should establish a baseline, evaluate AI across utilization, impact, and cost, and run controlled experiments that include developer experience and delivery outcomes—not assume that faster code generation creates proportional business value ([00:11:03]-[00:14:10], [01:07:20]-[01:08:18]).
Key Points Covered
- Measure utilization, impact, and cost together: Tacho treats acceptance rate as a limited fit-for-purpose signal rather than proof of production use, velocity, innovation, or business impact ([00:12:05]-[00:16:56]).
- The most useful cases may not be code generation: In a DX study covering more than 180 companies, developers reporting substantial time savings ranked stack-trace analysis and refactoring above mid-loop code generation. Tacho argues for training around these less obvious uses ([00:23:32]-[00:25:33]).
- Generation can transfer effort into review: Tacho contrasts generated code, which still needs assessment and integration, with stack-trace analysis that can remove diagnostic toil; Orosz emphasizes the expert judgment required to evaluate output ([00:26:36]-[00:28:43]).
- Time savings can reduce job satisfaction: Tacho cites DORA research in which some developers reported that AI accelerated coding they enjoyed while leaving meetings, administration, and other toil. She proposes combining experience data with workflow and system metrics ([00:29:46]-[00:35:52]).
- Adoption needs enablement and segmentation: In the Booking.com case study, Tacho reports that training and office hours accompanied 65% weekly-or-daily adoption, while non-use could reflect license access or poor fit for specialized work rather than resistance ([00:17:47]-[00:23:32]).
- Customer gains do not establish causation: For Workhuman, Tacho reports an 11% increase in DX's developer-experience index and 15% higher velocity among frequent AI users than non-users. These are segmented vendor-reported findings, not evidence that AI causes the same gains elsewhere ([00:42:30]-[00:46:33]).
- Larger diffs may trade speed for stability: Tacho says DX data associates AI use with more-complex diffs, while the DORA study she presents forecasts reduced stability as adoption rises. Her larger-batch explanation is a hypothesis, so speed should remain paired with quality and reliability measures ([00:51:42]-[00:54:49]).
- Treat rollout choices as experiments: Tacho describes Indeed comparing tools across cohorts and use cases, then testing interventions such as preliminary AI review or assisted migrations. These are reported company experiments, not universal outcomes ([00:58:58]-[01:05:12]).
Full video: https://www.youtube.com/watch?v=xHHlhoRC8W4(opens in a new tab)