Study on LLM-generated code and comments in repositories

Discover how LLM-generated code and comments evolve in enterprise and community repositories, and their relationship with bugs.

viernes, 3 de julio de 2026 • 3 min read • Q2BSTUDIO Team

Findings on LLM-generated code: clones and comments

The integration of large language models (LLMs) into the software development cycle has transformed how both code and technical comments are generated. Recent studies, such as the analysis of active repositories between 2021 and 2025, reveal that between 20% and 30% of code in large technology companies is produced by these tools. However, questions remain about the actual quality of that generated content: from the difficulty of debugging it to the artificiality of the comments. This article explores these trends and offers a practical perspective for companies seeking to adopt artificial intelligence without compromising the robustness of their projects.

One of the most striking conclusions of the study is that the percentage of code detected as LLM-generated has decreased over time, while generated comments remain stable. Furthermore, this code frequently appears in test cases and shows a high level of internal cloning within the same repository. Comments, on the other hand, have a low proportion of grammatically correct sentences. These findings underscore the importance of having review processes and specialized tools to integrate AI safely and efficiently.

For organizations that develop custom applications, this data is a call to action. Automation with LLMs can accelerate the delivery of features but requires a careful approach. At Q2BSTUDIO, we combine the power of artificial intelligence with a solid methodology to ensure that the generated code is maintainable, secure, and aligned with business objectives. Our teams apply best practices of AI for businesses, including AI agents that optimize repetitive tasks, without neglecting human oversight.

Another relevant aspect is the difference between repositories maintained by companies and community ones. The former show a higher percentage of LLM-generated code and comments, reflecting a more aggressive adoption but also a potential risk of error propagation. Only a small percentage of manually labeled bugs are associated with AI-generated code, suggesting that problems stem more from integration than from generation itself. That is why at Q2BSTUDIO we offer cybersecurity and code review services to mitigate vulnerabilities in environments where LLMs are used.

Infrastructure also plays a key role. To scale AI solutions robustly, companies need flexible cloud platforms. At Q2BSTUDIO, we facilitate migration and management on AWS and Azure cloud services, ensuring that artificial intelligence workloads run with high availability and low cost. Additionally, generating reports and intelligent dashboards with Power BI allows teams to visualize code quality metrics and detect anomalous patterns in repositories. This combination of process automation and data analysis enhances evidence-based decision-making.

In conclusion, the study on LLM-generated code and comments in repositories reminds us that technology is a tool, not a substitute for human judgment. To reap its benefits without falling into technical debt, companies must invest in custom software that integrates AI in a controlled manner, with review processes, team training, and a strategic vision. At Q2BSTUDIO, we accompany organizations on this path, offering solutions ranging from AI consulting for businesses to the implementation of AI agents that improve productivity without losing control over the final quality of the product.

A BREAK?

Play for a moment before you go

OUR SERVICES

How we can help you

Do you have a project in mind?

Tell us your vision and we'll turn it into a software solution. Whatever the scope, we make your idea real.