Tokenization vs. Tokenization Augment: Writer's variation in handwriting

Learn how tokenization and concatenation-based data augmentation reduce error in handwrite recognition with IMUs, improving the

14 jul 2026 • 5 min read • Q2BSTUDIO Team

Improved handwriting recognition with tokenization and augmentation

Handwriting remains a critical data entry channel in many areas, from note-taking on mobile devices to digitizing forms in enterprise environments. However, its automatic recognition faces a persistent challenge: variability between writers and within the same writer. Each person imprints a unique style, with different inclinations, pressures and speeds, while even the same user can generate inconsistent strokes depending on the context or mood. In this article, we look at two complementary strategies—subword tokenization and concatenation data augmentation—and how their correct application can make a difference in inertial measurement unit (IMU)-based handwriting recognition systems. In addition, we explore the practical implications for companies looking to integrate these capabilities into their workflows, relying on custom software solutions and modern cloud platforms.

The problem of writer's variation is divided into two facets. Inter-writer variation refers to stylistic differences between people; A system trained with a small set of signatures can fail miserably when faced with a new user. On the other hand, intra-writer variation encompasses the person's own natural fluctuations, such as writing faster in a meeting or slower when reflecting. Recent research shows that there is no single optimal technique for both types of variability. Tokenization in subwords, inspired by natural language processing techniques, breaks down sequences of strokes into smaller units (e.g., bigrams). This allows the model to learn basic structural patterns that are better generalized to unseen writing styles, significantly reducing the error rate per word in scenarios where the writers of training and validation are different. However, when the writer is the same (writer-dependent split), this same abstraction can impair performance due to a shift in vocabulary distribution between the training and validation sets.

Faced with this limitation, concatenation-based data augmentation emerges as a powerful regularizer. By artificially joining sequences of strokes, the system exposes the model to combinations and transitions that do not appear in the original data, compensating for the scarcity of examples within the same writer. The results indicate that this technique can reduce the error rate per character by 34.5% and the error rate per word by 25.4%, outweighing even the benefits of simply prolonging training with more times. The key is that short, low-level tokens benefit the most, as they allow elemental shards to be recombined without losing the coherence of the original emote.

For companies developing writing capture and analysis systems, these findings have immediate practical value. It is not a matter of choosing one technique over another, but of understanding the context of use. An application intended to validate signatures in a corporate environment will likely work with a small set of known writers (intra-writer), where concatenation augmentation will be most effective. Instead, a multi-user note transcription tool (inter-writer) will benefit from structural tokenization. In both cases, the implementation requires a robust infrastructure that supports data preprocessing, model training, and deployment to production. This is where the expertise of Q2BSTUDIO, a company specializing in the development of AI for companies and artificial intelligence solutions, comes into play. Our teams design custom data pipelines that integrate everything from IMU sensor capture to neural network modeling, all on scalable cloud infrastructures such as AWS and Azure cloud services.

Beyond handwriting recognition, the underlying principles of these experiments—the management of variability and the balance between generalization and specialization—are applicable to other domains of artificial intelligence. For example, in recommender systems or user behavior analytics, tokenization of action sequences and augmentation of synthetic data are common strategies to combat data scarcity and user heterogeneity. Companies that take a comprehensive approach to services, business intelligence , and dashboards like Power BI can benefit from these techniques to enrich their predictive models without relying exclusively on limited historical data. Likewise, the addition of AI agents that automate repetitive tasks, such as classifying handwritten forms, becomes more reliable when the underlying models are trained with proven regularization strategies.

Another relevant aspect is cybersecurity. Systems that process handwriting often handle sensitive information, such as digital signatures or confidential notes. A model that is vulnerable to adversarial attacks or that leaks information through its weights could compromise users' privacy. For this reason, cybersecurity measures Q2BSTUDIO integrated into each phase of the software life cycle, from data collection to deployment in cloud environments. The combination of data augmentation and tokenization techniques not only improves accuracy, but can also contribute to the robustness of the model against malicious variations.

In practice, implementing an IMU-based handwriting recognition solution requires orchestrating multiple components: embedded sensors, feature extraction algorithms, recurrent neural networks or transformers, and a user interface that offers real-time feedback. Companies looking to adopt this technology can start with a pilot focused on a specific use case, such as digitizing incident reports in warehouses or capturing signatures on mobile devices. Our Q2BSTUDIO team, with extensive experience in custom applications, helps define requirements, select the most appropriate tokenization or augmentation strategy, and integrate everything into a platform that can be deployed on both AWS and Azure. In addition, we offer consulting services to optimize the performance of models through pruning, quantization or federated learning techniques, respecting data privacy.

Looking to the future, the trend is towards increasingly personalized systems that learn on-device to adapt to each user's unique style without relying on the cloud. In this scenario, tokenization in subwords becomes especially relevant because it allows storing a compact vocabulary that captures the essence of gestures, while data augmentation by concatenation can be done locally to generate synthetic examples without sending sensitive data abroad. The synergies between these two techniques, combined with agile artificial intelligence platforms, pave the way for digital assistants that understand our writing as naturally as we recognize our own handwriting.

In conclusion, the writer's variation is not an insurmountable obstacle, but a well-characterized problem that admits specific solutions depending on the scenario. Tokenization and data augmentation are two tools in the machine learning engineer's arsenal that, when used judiciously, allow robust and accurate systems to be built. Companies such as Q2BSTUDIO, which specialise in AI for companies, accompany their customers on this path, offering everything from the development of custom models to integration with Power BI and other business intelligence tools. Handwriting does not disappear; It is transformed into a digital channel managed by artificial intelligence. Knowing how to manage its variability is the key to taking advantage of its full potential.

A BREAK?

Play for a moment before you go

OUR SERVICES

How we can help you

Do you have a project in mind?

Tell us your vision and we'll turn it into a software solution. Whatever the scope, we make your idea real.