CNCF White Paper: Data Storage for AI in Cloud Native

Learn how to overcome data bottlenecks in AI/ML workloads with the new CNCF white paper. Storage, caching, vectors, and more.

miércoles, 8 de julio de 2026 • 3 min read • Q2BSTUDIO Team

Storage challenges and solutions in AI/ML workloads

The deployment of large-scale artificial intelligence and machine learning workloads has become a strategic priority for modern enterprises. However, moving these massive, highly stateful and demanding data volumes into cloud native infrastructures creates bottlenecks that traditional architectures cannot resolve. The recent CNCF TAG Infrastructure white paper, titled "Data On Kubernetes – Data Analytics and AI/ML Workloads", delves into the storage challenges engineering teams face when trying to feed GPUs and accelerators with enormous datasets. Below, we analyze its key findings from a practical and business perspective, and show how a software development company like Q2BSTUDIO can help overcome these barriers.

One of the most recurring issues is the so-called "small file trap": datasets composed of millions of tiny files saturate storage metadata, drastically reducing performance. Added to this is the dissociation between compute and storage which, although it scales efficiently, can generate a high overhead of API calls and low GPU utilization. Furthermore, load profiles vary drastically: batch training demands a constant data flow, while production inference requires low-latency responses and intermittent traffic patterns. To address these situations, the white paper proposes a layered architecture that includes hybrid data lakehouses with open formats like Apache Parquet or Iceberg, vector databases like Milvus for high-dimensional embeddings, and distributed caching systems like Fluid, a CNCF project that optimizes data locality within Kubernetes.

Another critical aspect is the standardization of interfaces. The use of CSI (Container Storage Interface) for block or file storage, COSI (Container Object Storage Interface) for objects, and FUSE CSI drivers allows AI applications to communicate with storage uniformly and efficiently. Modern data pipelines are also evolving: the white paper recommends migrating from traditional batch processes to real-time streams using Change Data Capture (CDC) and event platforms like Apache Kafka. This is essential both for continuous training and for inference in AI agent systems, which require short-term and long-term memories, event logs, and intermediate artifact repositories.

The document also breaks down storage profiles according to the AI lifecycle phase. During model training, the priority is to maximize GPU utilization, which requires tolerating non-sequential accesses due to random data shuffling and managing massive synchronized write bursts during checkpointing. In inference, latency sensitivity forces the use of advanced memory architectures like KV Caching and Prefix Caching to avoid redundant calculations. And in the emerging field of agentic AI, systems require mutable short-term memory, an append-only event history, and consolidation of past sessions into persistent storage.

Faced with this landscape, companies need technology partners who master both cloud native infrastructure and custom application development. Q2BSTUDIO offers precisely that: a team expert in custom software that integrates artificial intelligence and AWS and Azure cloud services to build scalable and secure solutions. For example, our cybersecurity practice ensures that critical data pipelines are protected, while our business intelligence services with Power BI allow visualizing model performance metrics. We also support the implementation of AI for enterprises and AI agents that require the architectures described in the CNCF white paper.

If your organization is adopting AI workloads on Kubernetes and needs technical guidance to resolve storage bottlenecks, we invite you to learn how our artificial intelligence offering can help you design and implement the right solutions. Additionally, to ensure a robust cloud infrastructure, we recommend exploring our AWS and Azure cloud services, where we integrate storage optimized for AI workloads. Ultimately, the combination of a solid architectural approach and an experienced technology partner is the key to harnessing the full potential of artificial intelligence in cloud native environments.

A BREAK?

Play for a moment before you go

OUR SERVICES

How we can help you

Do you have a project in mind?

Tell us your vision and we'll turn it into a software solution. Whatever the scope, we make your idea real.