Batch size optimization in AI model training is a critical challenge for maximizing hardware usage and reducing computation time. Traditionally, fixed or manually tuned schedules are used, but adaptive approaches like gradient noise scale (GNS) offer a more robust alternative. However, these methods assume Euclidean geometries inherent to SGD, which conflicts with modern optimizers like signSGD (based on l8 norm) and Muon (based on spectral norm ??8). A recent work proposes extending the GNS concept to these non-Euclidean spaces, deriving noise metrics that naturally emerge from the geometry of their dual norms. To estimate them efficiently, a variance procedure is introduced that leverages local mini-batches in distributed data-parallel systems. Experimental results show that adaptive strategies based on this non-Euclidean noise can match the validation loss of constant-batch baselines, reducing training steps by up to 66% in models like Llama with 160 million parameters.
This innovation has direct implications for companies developing AI for business and seeking to optimize their machine learning pipelines. At Q2BSTUDIO, we offer custom applications that integrate these advanced techniques, combining them with AWS and Azure cloud services to scale training, and with AI agents that automate resource management. Additionally, our expertise in custom software and cybersecurity ensures that sensitive data environments are protected. For organizations that need to monitor their model performance, business intelligence services with Power BI allow visualizing key metrics such as gradient noise evolution. Implementing adaptive batch sizes with non-Euclidean optimizers is a step toward more efficient and sustainable training, and at Q2BSTUDIO we accompany our clients at every stage of the process, from conceptualization to production deployment.

.jpg)


