Early detection of heart disease remains one of the biggest challenges in preventive medicine. Heart murmurs, indicating abnormal blood flow, can be identified through auscultation with a stethoscope. However, human interpretation is subjective and requires years of experience. This is where artificial intelligence (AI) offers a promising solution: analyzing heart sound recordings with convolutional neural networks (CNNs) to automatically classify whether a heartbeat is normal or abnormal. This article compares three audio representation methodologies—log-mel spectrogram, PCEN normalization, and multi-resolution spectrogram—on the same CNN model, evaluating their performance and applicability in real-world settings.
The study is based on the PhysioNet 2016 dataset, a benchmark in the field. The CNN architecture, optimizer, and random seed were kept constant, varying only how the audio is converted into images that the network can process. This isolates the impact of spectral representation on diagnostic accuracy.
The log-mel spectrogram is the most widespread technique: it transforms the audio signal into the frequency domain using the Short-Time Fourier Transform (STFT), applies mel filters, and logarithmic scaling. Its simplicity and low computational cost make it a common choice, but it can lose detail in low-energy regions. On the other hand, PCEN (Per-Channel Energy Normalization) stabilizes energy over time in each frequency band, reducing volume variations and background noise. This is especially useful in clinical environments where recording conditions are uncontrolled. Finally, the multi-resolution spectrogram stacks several analysis windows of different sizes, capturing both fine details and global patterns. Although it increases computational complexity, it provides a more complete view of cardiac dynamics.
The results are revealing: all three methods achieve a sensitivity (true positive rate) close to 0.95, meaning they correctly identify most abnormal cases. However, on the official PhysioNet metric (which balances precision and sensitivity), the log-mel spectrogram scores 0.910, while PCEN and multi-resolution achieve 0.915 and 0.916 respectively. This improvement, though modest, can translate to dozens of additional correct diagnoses in large populations.
To understand what the model learned, Grad-CAM was applied, a visualization technique that highlights the image regions most influencing the decision. The analysis showed that the CNN predominantly focuses on low frequencies, where the S1 and S2 heart sounds are located. This validates that the model is learning real physiological features and not spurious artifacts. From a business perspective, this robust behavior is key to deploying assisted diagnosis systems in hospitals and clinics.
Implementing these models in production requires a comprehensive technological approach. Companies developing custom software can integrate these algorithms into telemedicine platforms or IoT devices. For example, a solution based on cloud AWS/Azure allows scaling the processing of thousands of daily recordings, while AI tools such as intelligent agents can automate the workflow: from audio acquisition to clinical report generation. Furthermore, cybersecurity is critical when handling sensitive patient data, requiring encryption and access control protocols.
Q2BSTUDIO, as a software and technology development company, offers specialized services in creating custom applications for the healthcare sector. Our team combines expertise in artificial intelligence, cloud computing, and business intelligence to build robust solutions. For instance, a murmur detection system can be integrated with Power BI dashboards so medical teams can visualize trends and alerts in real time. Likewise, process automation through AI agents reduces administrative burden, allowing professionals to focus on patient care.
The choice of the right spectrogram depends on the use context. If computational resources are limited and recordings are made in controlled environments, log-mel may suffice. In contrast, for mobile or rural applications where ambient noise is variable, PCEN offers greater robustness without significantly increasing latency. The multi-resolution approach, though more costly, is ideal when precision is critical and advanced hardware is available.
Beyond numerical results, this study shows that small variations in preprocessing can have a measurable impact on performance. For companies looking to implement AI solutions in cardiac diagnosis, it is advisable to conduct a comparative evaluation of different representations before deciding on the final architecture. Collaboration with technology partners like Q2BSTUDIO accelerates this process, leveraging preconfigured machine learning frameworks and DevOps best practices.
In conclusion, heart murmur detection via CNN is an accessible technical reality. Both PCEN and multi-resolution spectrograms slightly improve accuracy over traditional log-mel, but all offer excellent sensitivity levels. The key is selecting the representation that best fits the deployment environment and business requirements. With the support of companies like Q2BSTUDIO, specialized in custom software, it is possible to turn this technology into a reliable and scalable clinical tool, helping save lives through early detection of heart disease.





