Detecting LLM-Generated Tokens in Human-AI Coauthored Text

A novel method to detect LLM-generated tokens in mixed authorship documents using adaptive smoothing. No token-level labels needed. High accuracy.

sábado, 25 de julio de 2026 • 4 min read • Q2BSTUDIO Team

Nueva detección a nivel de token sin entrenamiento previo

The collaboration between humans and artificial intelligence in textual content creation has grown exponentially. Companies across all sectors rely on language model assistants (LLMs) to draft reports, emails, articles, and even technical documentation. However, this practice poses a fundamental challenge: how to distinguish which parts of a document have been generated by a machine and which by a person? The answer has implications not only for academic authorship but also directly affects the integrity of business processes, cybersecurity, and trust in automated systems.

Until now, most AI-generated text detectors focused on classifying entire documents as “generated” or “human.” This binary approach falls short when the same text contains mixed contributions: a paragraph written by a human expert, another polished by an LLM, and a conclusion fully generated by a machine. To address this need, a new generation of methods operating at the token level — the basic unit of modern language models — has emerged. The idea is to precisely flag each suspicious fragment, similar to what a spell checker does but oriented toward identifying automatic generation patterns.

One of the most promising approaches involves smoothing the detection scores of neighboring tokens to reduce variability, using adaptive rules such as the Lepski method to select the optimal bandwidth according to the local authorship structure. This technique does not require token-level labeled data for training, making it especially practical in real-world scenarios where data is scarce or expensive to obtain. Moreover, it achieves a balance between sensitivity and specificity, minimizing the mean squared error in estimating the underlying signal.

From a business perspective, the ability to detect LLM-generated tokens in co-authored texts has direct applications in content verification, auditing of reports generated by virtual assistants, and protection against literary identity theft. For instance, in the field of cybersecurity, identifying AI-generated fragments can help detect automated phishing attacks or disinformation campaigns. Similarly, in corporate environments where AI agents are used to draft customer responses, knowing which parts of the dialogue are automatic is crucial for maintaining transparency and regulatory compliance.

At Q2BSTUDIO, as a software and technology development company, we understand that integrating these capabilities into existing systems requires a robust and customized approach. That is why we offer custom software services that allow incorporating AI-generated token detection modules into document management platforms, CRM systems, or Business Intelligence tools. The combination of BI/Power BI with textual authenticity detection algorithms opens the door to dashboards that monitor in real time the proportion of machine-generated content in workflows.

The infrastructure needed to run these models efficiently is typically deployed in the cloud. This is where cloud platforms like AWS or Azure come into play, providing scalability and computational resources to process large volumes of text. At Q2BSTUDIO we offer cloud AWS/Azure services for deploying real-time detection pipelines, ensuring low latency and high availability. Furthermore, integration with automation systems allows alerts generated by these detectors to trigger corrective actions, such as manual review of a suspicious document.

Current research in token-level detection is moving toward more robust methods that can handle different languages, styles, and context lengths. The smoothing technique with the Lepski adaptive rule, mentioned earlier, is just one example of how classical statistics can be combined with natural language processing to achieve practical results. Empirical validation shows that this approach outperforms traditional baselines both on synthetic datasets and on real human-AI co-authorship corpora.

For businesses, implementing such solutions is not just a matter of transparency but also of competitive advantage. Being able to guarantee the authenticity of documents generated internally or received from third parties strengthens customer and partner trust. Moreover, in regulated sectors such as banking or healthcare, the ability to audit the origin of each token may be a legal requirement.

At Q2BSTUDIO we advocate for a holistic approach that combines custom software development, artificial intelligence, cybersecurity, cloud computing, and Business Intelligence. Our team works side by side with organizations to design and implement AI-generated content detection systems tailored to their specific needs. Whether integrating a token analysis module into an existing application or building a detection pipeline from scratch on the cloud, we offer scalable, secure, and efficient solutions.

In conclusion, detecting LLM-generated tokens in human-AI co-authored texts is a rapidly expanding field with enormous practical potential. Token-level methods, such as those based on adaptive smoothing, provide a granularity that document-level classifiers cannot achieve. For companies, adopting these technologies is not optional if they want to maintain integrity and transparency in a world where AI increasingly collaborates in writing. At Q2BSTUDIO we are ready to help you take that step, combining technical expertise with deep business knowledge.

A BREAK?

Play for a moment before you go

OUR SERVICES

How we can help you

Do you have a project in mind?

Tell us your vision and we'll turn it into a software solution. Whatever the scope, we make your idea real.