Where can adding AI tools to the developer workflow yield the most benefits? Here’s one veteran engineer’s analysis.
by John Baublitz, Principal Software Developer, Red Hat
A 2025 METR study of AI efficiency in open source development engendered a lot of discussion when it released findings indicating that using genAI in certain contexts had the potential to slow down total development time. Since that study, METR followed up with a new dataset, finding evidence of speedups but also indicating that both datasets are highly prone to selection effects. While these studies may not have provided conclusive information, they did tap into something important: the importance of scoping AI usage.
With the introduction of AI in many areas of the tech industry, indiscriminate usage has potential problems. One example highlighted in the METR study is the potential increase of code review time when large amounts of code are generated by AI. This category of problems is best described as “unintended consequences.” While genAI is indisputably faster at writing code than people, the process of a release includes more than the task of implementing the feature. It also involves review, automated and sometimes manual testing, discussion, and many other tasks that may be more abstract.
As a Principal Software Developer at Red Hat, my work often involves prototyping, researching, writing tests, and driving the feature to completion. With Red Hat’s transition into an agentic software development lifecycle, I decided to do an analysis of adding AI tools into my workflow, evaluating where it slowed me down and where it was helpful. The results were nuanced but hopeful.
How AI made me faster
Reducing manual labor
AI has the undeniable potential to significantly speed up certain work that I do. The most notable example of a speedup was in drafting highly localized changes across a large number of files. In cases where the code cannot be easily factored out into a shared implementation due to changes in control flow or return value handling, AI excelled at rapidly making changes to multiple implementations and generating a consistent, correct handling of variations across the implementations, appropriately adapting the change to the altered context.
Because these changes were not a simple patch application to every version of the implementation, AI was able to quickly analyze it, treat the code separately, and appropriately make changes that fit in the surrounding code even if it was different from the previous application context. Code review time was very low, as the changes were localized and aimed at a specific fix or goal. Even from a psychological perspective, these changes, when done manually, are often boring and error prone. Reducing the cognitive load of these kinds of tasks allows the employee to focus on more impactful work that is more rewarding. This kind of task is where AI is most universally helpful and effective.
Code review and diagnostics
Another useful application of AI is code review and diagnostics. While not all models and products perform equally well, they do consistently find bugs that are not obvious or straightforward enough to spot immediately. Using AI tools like CodeRabbit and Claude for general code review can create additional work when it hallucinates or incorrectly diagnoses the problem. However, our attempts to use AI for code review found a good number of problems during development, which allowed us to fix them prior to merging the code. Despite the hallucinations, this likely saved us days of unnecessary work by preventing the buggy code from being released and eliminating the work needed to track it down and put out a patch release to fix it.
AI can perform even better when the problem is scoped.
AI can perform even better when the problem is scoped. If a user reports a bug and describes the behavior in some detail, AI can be incredibly helpful in tracking down the origin of the buggy behavior and explaining how to fix it. I found this even more definitively time saving. The detail in open source bug reports varies quite a bit: some bug reports can be vague and lacking in definitive steps to reproduce the behavior. AI can sometimes fill in these gaps and track down the problem just based on the description of the behavior. This is probably the second most useful application of AI for my work. I found this time saving even when there is a small delay during the development process.
Prototyping
Prototyping is an AI use case with a great amount of potential. Prototyping a large code change can sometimes be a large undertaking, particularly in Rust, the language I spend the most time writing. The type system is very strict, and changes to APIs can be very unforgiving, requiring a large amount of ancillary work to even determine whether the change will fix the problem. When prototyping a change, AI allows the user to fail fast, making a large number of changes quickly and then determining whether or not they’ve found a workable solution. This method saved me a lot of time when I needed to explore multiple possible solutions and get a more concrete idea of how each would turn out.
Unit testing
Developers often do not enjoy writing tests as much as features, making testing a prime target for offloading some of the work to AI. Unit tests can be repetitive, simple, and highly scoped to test one individual part of the code. In my experience, AI can write usable tests of tasks very quickly, without producing excessive code review burden for the employee. One criticism of AI in code generation has been the maintainability of the code. However, unit tests are self contained, which removes some of the architecture concerns. Maintainability looks different in tests than it does in a large project with multiple layers of APIs. Often the biggest concern with maintainability in unit tests is being able to keep them up to date with the code they’re testing and testing functionality accurately. AI has handled these requirements for me relatively well and has not led to a lot of extra maintenance work.
When AI causes problems
Lack of domain knowledge
While the previous examples highlight cases where AI can significantly speed up the process, there are also domains where AI application yields mixed results. One of these areas is work that requires extensive domain knowledge. While writing unit tests is usually a seamless process, developing more involved end-to-end testing can require domain knowledge to generate these tests correctly. I’ve used AI tools to develop both, and the more domain knowledge required for the testing task, the more room there is for AI to hallucinate and make mistakes. Even if it does find the right solution eventually, the user must have and provide the domain knowledge that the AI tool lacks to see any speedup at all. Otherwise, a lot of time and money is wasted having AI repeatedly try to create a working test and learn from its mistakes.
Hallucinations
Hallucinations can be costly in terms of time and quality of the result. If the user doesn’t have the necessary domain knowledge to guide the agent, the agent can spend a lot of time trying to solve a problem that it has incorrectly diagnosed. Using an agent to brainstorm and speed up the process of finding a mistake in the code can make the process faster and less frustrating. However, verification is an indispensable step in agentic workflows. Without taking the time to validate the result, users can easily get into a cycle of trial and error where none of the solutions work. My recommendation is to always approach debugging in three phases: diagnosis, verification, and then implementation of the fix. Skipping the verification step will almost certainly slow down the end result.
Playing to AI strengths
My personal workflow analysis suggests that task selection may be the most important factor influencing whether AI speeds up or slows down the work. The overall strengths of AI for my work are pattern recognition, quick code generation, and ease of iteration on ideas that may not be fully formed. However, it also has its weaknesses. Using AI when the user doesn’t have the required domain knowledge, deep knowledge of the architecture of the code, or an awareness of the potential for hallucinations can result in a very large slowdown in actual productivity.
Hybrid AI usage is likely the best case scenario. That means choosing AI for tasks that play to its strengths and avoiding it for areas where the user doesn’t have the appropriate knowledge to guide the agent. AI can have a huge positive impact on productivity when the user takes an active role in scoping, providing context, and identifying parts of the workflow that most need efficiency improvements.
John Baublitz works on the Stratis project and has contributed to the upstream Rust for Linux effort. He works primarily in the Rust ecosystem and enjoys all things related to systems.









