AI coding tools don't reduce mental load, they shift it
In a four-day field study at SAP, 21 professional developers rated GenAI-assisted tasks about one point higher on a 7-point cognitive load scale than tasks without GenAI, even after controlling for task type and duration. Wearable wristband data added almost nothing on top.
Using GenAI at work didn’t make tasks feel lighter for SAP’s developers. It made them feel heavier. In a four-day field study at two SAP sites in California, 21 professional developers logged every work task, and the ones where they used GenAI were rated about one point higher on a 7-point cognitive load scale than tasks without it. That held after controlling for task type, task duration, and each developer’s personal rating habits.
- GenAI use: +0.996 points of perceived cognitive load (χ²(1)=23.80, p under 0.001).
- Task type mattered most. Development-heavy work scored 1.44 points higher than “other” activities and 0.59 higher than collaboration-heavy work.
- Longer tasks felt heavier: +0.007 points per minute, or roughly +0.4 points for an extra hour.
- Wristband data barely helped. Of 11 heart, skin conductance, and temperature metrics, only one (skin temperature variability) added anything, and it explained 1.4% more variance.
- The same developers said AI reduces mental effort overall. The task-level ratings disagree with their own retrospective opinion.
The setup
Participants were 12 software engineers, 8 senior engineers, a senior quality specialist, and a principal architect, all Java users with 5 to 10 years of experience and at least a few months of GenAI tool use (mostly GitHub Copilot and ChatGPT). Over four days in Q2 2025, they kept a paper log beside their keyboard: each task, its start and end times, a cognitive load rating from 1 to 7, and whether they used GenAI. They were asked to use GenAI for at least one continuous hour a day and could pick any tool.
At the same time, each developer wore an Empatica EmbracePlus wristband recording blood volume pulse (for heart rate and HRV), electrodermal activity, skin temperature, and motion. After the fact, 382 tasks were sorted into development-heavy, collaboration-heavy, or other.
The analysis is a linear mixed-effects model: cognitive load predicted from GenAI use, task category, and duration, with a random intercept per developer. Physiological metrics were then added one at a time, and later together, to see if they explained anything the work context didn’t.
Work context explains a third of the variance
The three context factors together explained 32.4% of the variance in cognitive load ratings, rising to 45.9% once you account for the fact that some developers just rate everything higher.
What each factor adds to explained variance
Increase in marginal R² when the factor is added to the model. Bars scaled to task category = 100%.
The GenAI effect is the interesting one. The authors are careful not to read it as “AI makes coding harder.” Their interpretation is that AI shifts effort rather than removing it: less typing code, more writing prompts, reading generated output, checking it for correctness, and fitting it into an existing codebase. Earlier analyses from the same study found GenAI-assisted tasks were rated both more cognitively demanding and more productive, and developers’ open-ended complaints centered on prompt sensitivity, inaccurate output, and the need to validate everything.
The gap with developers’ own opinions is worth noticing. Asked in general, these developers tended to agree that working with AI takes less mental effort. Rated task by task, AI tasks came out heavier. The paper attributes this to the known gap between momentary and retrospective ratings: overall impressions get shaped by salient wins, while per-task ratings capture the grind of reviewing output.
Wearables add almost nothing
Lab studies with n-back and Stroop tasks show that heart rate, HRV, and skin conductance respond to cognitive load. In a real office, they didn’t. None of the four cardiovascular metrics or four electrodermal metrics improved the model after multiple-comparison correction.
The only survivor was skin temperature standard deviation: more stable skin temperature during a task went with higher perceived load (estimate -0.224 per standard deviation, corrected p=0.041). Combining all modalities into one model gave a near-miss (χ²(4)=9.42, p=0.052) with a +1.9% gain in explained variance, and nearly all of that came from the same temperature signal. The authors suggest reduced variability might reflect sustained focus, but it could as easily be stable posture or room conditions.
Their conclusion is blunt: the results “do not support treating wearable physiology as a context-independent measure of cognitive load.” Wearables might help with aggregate evaluations of tools and workflows, not individual assessment or automated workplace decisions.
Caveats
- Small and short. 21 developers, four days, 382 tasks, one company. BVP-based analyses dropped to 19 people and 263 tasks after cleaning.
- GenAI use is a yes/no flag. The study can’t separate Copilot autocomplete from a long ChatGPT debugging session, or good output from bad. It also predates widespread agentic tools.
- Developers weren’t randomized. People chose when to use AI, so they may have reached for it on tasks that were already harder in ways the three task categories don’t capture.
- Self-reported load is the ground truth. It measures how hard work felt, not actual cognitive processing, and it’s subject to how each person uses the scale.
- The wearable null result is weak evidence. The authors’ power simulations say the study could detect moderate physiological effects but might miss small ones.
The practical point for engineering leaders is the one in the title: productivity and adoption dashboards don’t tell you whether AI-assisted work is sustainable. If your team’s AI tasks are shipping faster but feel a point heavier on a seven-point scale, that cost won’t show up in PR cycle times.
Liked this? Engineer's Codex sends one deep dive and a link roundup every week.
No spam. Unsubscribe anytime.