CoCurve: Training-Free Structured Pruning for LLMs

CoCurve: Jointly prune attention and FFN units in LLMs without training or fine-tuning. Using co-pruning curvature for efficient structured compression.

sábado, 25 de julio de 2026 • 3 min read • Q2BSTUDIO Team

Poda de LLMs sin necesidad de reentrenamiento

Compressing large language models (LLMs) is a strategic priority for companies aiming to deploy artificial intelligence efficiently. Structured pruning removes entire computational units, such as attention heads and feed-forward (FFN) channel groups, to reduce model size without compromising functionality. However, traditional training-free methods often rank these units independently, incorrectly assuming that the loss from removing a set is the sum of individual losses. In Transformers, sublayers are coupled through a shared residual stream, meaning two individually weak units can be jointly indispensable. Ignoring this interdependence leads to removing them together, degrading performance.

This is where CoCurve comes in, an innovative joint pruning approach that overcomes this limitation. CoCurve is a calibration-only, fine-tuning-free method that prunes attention and FFN units simultaneously. It uses a second-order Taylor expansion of the token-level KL divergence between the frozen model and its masked copy, yielding a single Fisher matrix. The diagonal elements represent classical node saliency, while off-diagonal entries are co-pruning curvature edges: the extra damage of removing two units together. Under a single-ablation additivity approximation, this matrix reduces to a Gram product of single-unit ablation features, so the full M x M interaction is recovered with M forward passes, with no pairwise sweeps or gradients. Pruning then becomes one budgeted quadratic program, solved in a single shot under a shared attention—FFN budget, with no labels, fine-tuning, or recovery.

From a technical perspective, the beauty of CoCurve lies in its computational efficiency. While conventional methods require evaluating unit pairs to capture interactions, CoCurve models all relationships with a linear number of passes. This is crucial for modern LLMs with hundreds of millions of parameters, where exhaustive exploration would be prohibitive. The resulting Fisher matrix captures joint curvature, allowing the algorithm to decide which unit combinations to preserve to minimize information loss. In practice, this yields smaller models that maintain near-original performance, speeding up inference and reducing hardware requirements.

For companies developing AI-based solutions, adopting techniques like CoCurve can be a competitive advantage. At Q2BSTUDIO, we understand that efficiency is not just a technical luxury but an operational necessity. That's why we offer AI and process automation services that integrate advanced compression methods, ensuring your models are lightweight, fast, and accurate. It's not just about pruning, but doing it intelligently, preserving critical knowledge and removing redundancies. This approach aligns perfectly with our philosophy of developing custom software, where each solution is tailored to the client's specific needs.

Furthermore, CoCurve fits into a broader ecosystem of model optimization. Companies operating in the cloud, whether on AWS or Azure, can benefit from smaller models that reduce resource consumption and inference costs. At Q2BSTUDIO, we offer consulting on cloud AWS/Azure to help our clients deploy these compressed models efficiently. Similarly, cybersecurity is strengthened: smaller models mean a reduced attack surface and easier auditing. And in data analysis, combining compressed models with Business Intelligence (BI) tools like Power BI enables real-time processing of large data volumes without overwhelming systems.

In short, CoCurve represents a significant advancement in LLM pruning. By accounting for cross-module dependencies, it achieves more aggressive compression without sacrificing quality. For companies developing AI agents, conversational assistants, or recommendation systems, adopting this technique can make the difference between a viable product and one that consumes too many resources. At Q2BSTUDIO, we are committed to innovation in custom software, artificial intelligence, and automation, helping our clients leverage cutting-edge technologies like CoCurve to their fullest potential.

OUR SERVICES

How we can help you

Do you have a project in mind?

Tell us your vision and we'll turn it into a software solution. Whatever the scope, we make your idea real.