Decompose Monoliths by Data Boundaries, Not Code

Stop drawing diagrams. We split a monolith with 500+ tables and billions of rows by following data seams, not code boundaries. Here's what worked.

lunes, 20 de julio de 2026 • 5 min read • Q2BSTUDIO Team

Por qué los planes de microservicios fallan sin datos

In recent years, migration from monolithic architectures toward microservice ecosystems has become an almost mandatory goal for organizations seeking to scale their technology teams and accelerate value delivery. However, at Q2BSTUDIO we have observed a troubling trend: most of these projects are planned as purely software engineering exercises, focused on component diagrams, REST APIs, and containers. Domain boundaries are drawn on digital whiteboards, design patterns are debated, and the underlying database is assumed to obey those theoretical frontiers docilely. Reality, all too often, is quite different. When the moment finally arrives to move information assets, the plan collapses in the face of an uncomfortable truth: data does not adapt to schemas; schemas must adapt to data.

This phenomenon is especially acute in custom software applications that have grown organically for a decade or more within a single relational database instance. What began as a clean, normalized data model eventually became a dense network of dependencies where cross-domain joins are the rule rather than the exception. Development teams, taking advantage of the absence of physical constraints, built reports crossing fifteen tables, generated materialized views feeding half a dozen distinct modules, and added general-purpose columns eventually written by processes no one assigned ownership to. These are not isolated bad practices, but the natural form a system takes when data proximity makes business boundaries seem irrelevant. The problem is not technical in origin, but epistemological: the organization lost sight of where one business context ended and another began, and that amnesia crystallized first in the relational schema.

The direct consequence is that any extraction attempt guided solely by application logic or source code structure is doomed to collide with invisible walls. At Q2BSTUDIO, when we tackle modernization projects, we deliberately reverse the analysis order. Before proposing the first service, we perform archaeology on the database: we map the real relationship graph, identify which tables concentrate writes from multiple origins, measure volumes and growth rates, and seek those natural seams where internal cohesion is high and external coupling is low. Those fracture lines, when they exist, rarely coincide with user interface modules or backend packages. We often discover that an apparently unified business domain is physically atomized into dozens of tables intertwined with other contexts, or that a secondary entity has grown into a bottleneck of billions of records that no massive migration can assume over a weekend.

The strategy that has proven effective does not consist of executing a grand plan designed in a boardroom, but rather treating decomposition as a series of controlled, reversible experiments. We extract a candidate, move only the data demonstrating single jurisdiction, deploy the service in parallel to the legacy flow, and contrast results in production. If metrics diverge, we stop, adjust, and retry. This discipline demands infrastructure that supports temporal duality, and this is where cloud AWS/Azure platforms provide a decisive advantage: the ability to scale synchronization resources on demand, maintain isolated validation environments, and absorb processing spikes without compromising operational stability. The cloud is not merely a deployment destination, but the laboratory that lets data speak before the organization commits to an incorrect boundary.

The technical mechanism making this iterative approach viable is change data capture, known as CDC. Rather than falling into the dual-write trap, where each new service keeps feeding the old schema and perpetuating coupling, we establish a unidirectional flow: the legacy database remains the transient source of truth while extracted services rebuild their own state from an event stream. Thus, each service can operate in parallel, validate its consistency, and only when confidence is sufficient assume total responsibility for its domain. This method proves particularly useful when we encounter tables resisting any simple classification. When a dataset is written by multiple domains and forcing its assignment to a single service would generate a tangle of remote calls, we choose to recognize it as platform infrastructure: a dedicated owning service that decouples cross-cutting needs from the business core.

Nevertheless, even when service extraction advances successfully, there is a cost rarely budgeted during the design phase: the disappearance of free cross-cutting queries. In the monolith, a report combining customers, orders, and billing was a simple SQL statement with direct joins. After splitting, those tables live in intentional silos and that operation is no longer possible without additional infrastructure. The organization must build integration pipelines or materialized views that recombine scattered data, and this solution is not a one-time migration expense, but a permanent operational tax. Every time a service modifies a schema, someone must update the recombination logic, and that someone is usually a different team from the one that originated the change. To mitigate this burden, at Q2BSTUDIO we frequently integrate BI/Power BI layers that centralize analytics and reporting without subjecting operational services to aggressive read loads or cross dependencies. Separating the operational route from the analytical route is not a luxury, but an architectural necessity.

During these transformations, cybersecurity acquires critical relevance that also cannot be left for last. Every synchronization pipeline, every temporary replica, and every new data exposure point multiplies the attack surface. The transit of sensitive information between legacy and new services must be encrypted, audited, and subjected to strict access control policies from day one. In parallel, at Q2BSTUDIO we explore how AI agents can accelerate the discovery phase: by feeding models with database access logs, schemas, and query patterns, it is possible to automatically identify hidden dependencies, detect anomalies in data synchronization, and even suggest domain boundaries with an objectivity manual diagrams rarely achieve. Artificial intelligence does not replace the architect, but removes noise from the signal, allowing them to focus on strategic decisions.

In the end, the lesson we reiterate in every modernization project is that successful decomposition is not a diagram to be executed, but a hypothesis to be validated. Domain boundaries in a mature system are not where the architecture team wishes them to be, but where the data allows them to be established. Planning in business terms is indispensable, but only empirical production evidence can confirm whether those boundaries are real. At Q2BSTUDIO we design custom software and technology evolution strategies that start from this honesty: recognizing that clean architecture is a horizon, not a starting point, and that migrating toward it requires patience, adequate infrastructure, and the willingness to let data correct course before the cost of correction becomes prohibitive. Because when code and information disagree, information always wins.

OUR SERVICES

How we can help you

Do you have a project in mind?

Tell us your vision and we'll turn it into a software solution. Whatever the scope, we make your idea real.