PROTECT YOUR DNA WITH QUANTUM TECHNOLOGY
Orgo-Life the new way to the future Advertising by AdpathwayBreast cancer remains one of the most common cancers affecting women worldwide, and the single most powerful weapon against it is early detection. Mammography, the X-ray imaging of the breast, has been the gold standard for screening for decades, but reading mammograms is notoriously difficult. Radiologists must spot subtle differences in tissue density, tiny calcifications, and faint masses that can distinguish a malignant tumor from benign tissue, all while working through long queues of images. A new study published in Neural Computing and Applications now offers a detailed look at how pre-trained deep learning models perform on exactly this task, and how a simple but powerful technique called data augmentation can push their accuracy to levels that could meaningfully support clinical decision-making.
The research, conducted by Elham Tahsin Yasin, Ilkay Cinar, and Murat Koklu at Selcuk University in Turkey, set out to benchmark five well-known convolutional neural network architectures on mammographic image classification: DenseNet201, InceptionResNetV2, MobileNetV2, NASNetLarge, and Xception. Rather than building networks from scratch, the team used transfer learning, a strategy in which models pre-trained on millions of everyday images are repurposed for a specialized task. The underlying idea is that features learned from natural photographs, such as edges, textures, and shapes, provide a useful starting point that can be fine-tuned for the specific visual patterns of cancerous versus non-cancerous breast tissue.
The experimental design was deliberately comparative. Each of the five architectures was trained and evaluated under two distinct conditions: once using the original dataset of 745 mammographic images, and once using an expanded dataset of 9,685 images produced through data augmentation. Augmentation artificially enlarges a training set by applying transformations such as rotations, flips, and other image manipulations that preserve the diagnostic content of each image while creating new variations for the network to learn from. This matters enormously in medical imaging, where collecting thousands of annotated patient images is expensive, slow, and constrained by privacy regulations. The question the researchers posed was straightforward: how much does augmentation actually help when the task is distinguishing cancerous from non-cancerous mammograms?
The answer, it turns out, is quite a lot. Without augmentation, DenseNet201 led the field with an accuracy of 96.24 percent, followed closely by InceptionResNetV2 and Xception, both at 95.43 percent, NASNetLarge at 95.31 percent, and MobileNetV2 at 94.64 percent. Even in this baseline condition, the models demonstrated that transfer learning can extract clinically relevant information from a relatively small dataset of just 745 images. But the numbers climbed across the board once augmentation was applied to the larger dataset. DenseNet201 again took first place, improving to 98.06 percent accuracy, while MobileNetV2 reached 97.76 percent, Xception 97.67 percent, InceptionResNetV2 97.63 percent, and NASNetLarge 97.55 percent.
Several aspects of these results deserve attention. First, the gap between the best and worst models narrowed dramatically after augmentation, from roughly 1.6 percentage points to just half a point. This suggests that with sufficient training data, the choice of architecture becomes less critical than the quality and quantity of the training examples, a finding with practical implications for research groups with limited computational budgets. Second, MobileNetV2, the lightest and most computationally efficient model in the lineup, closed most of the gap with its heavier rivals, hinting that lightweight networks suitable for deployment in clinics with modest hardware could perform nearly as well as resource-hungry giants.
DenseNet201’s consistent dominance is also technically interesting. The DenseNet family of architectures is built around dense connectivity, in which each layer receives feature maps from all preceding layers. This design encourages feature reuse, strengthens gradient flow during training, and tends to perform well with limited data, which may explain why it excelled in both experimental conditions. Xception, by contrast, relies on depthwise separable convolutions that separate the processing of spatial patterns and cross-channel relationships, while NASNetLarge was discovered through automated neural architecture search rather than human design. The fact that all three approaches converge on similar performance after augmentation underscores a broader trend in the field: modern architectures have largely saturated the achievable accuracy on well-defined binary classification tasks, leaving data quality as the main lever for improvement.
The study is also notable for its explainability-aware framing. In medical applications, a model that simply outputs a percentage score is of limited use if clinicians cannot see why it reached its decision. Techniques such as gradient-weighted class activation mapping, which highlight the image regions most influential in a network’s prediction, have become central to building trust in artificial intelligence diagnostics. By benchmarking models with explainability in mind, the researchers align their work with a growing consensus that accuracy alone is not enough for clinical adoption; a model must also offer interpretable evidence that its attention falls on the actual lesion rather than on artifacts, imaging markers, or background tissue.
The dataset underpinning the study, known as Mammogram Mastery, was made publicly available through the Mendeley Data repository, which is significant for reproducibility. Medical imaging research has long been hampered by restricted access to patient data, and open datasets allow independent teams to verify results, compare methods on equal footing, and build cumulative knowledge. The researchers also employed stratified k-fold cross-validation, a rigorous evaluation strategy that preserves the class balance of cancerous and non-cancerous images across each validation split, reducing the risk that reported accuracies are inflated by lucky partitions of the data.
What do these numbers mean for patients? An accuracy of 98 percent on a binary classification task is impressive, but the path from benchmark to bedside involves many additional hurdles. Clinical systems must handle the full complexity of real screening practice, including different imaging devices, varying breast densities, image quality issues, and the far more granular classification schemes radiologists actually use, such as the BI-RADS categories. False positives generate unnecessary anxiety and biopsies, while false negatives can delay life-saving treatment, so the balance of sensitivity and specificity matters as much as raw accuracy. The authors’ results should therefore be read as strong evidence of feasibility rather than a claim of clinical readiness.
Nevertheless, the trajectory is clear and encouraging. The study demonstrates that data augmentation can lift the performance of pre-trained deep learning models by more than a full percentage point even at very high baseline accuracy, and that the resulting systems are accurate enough to serve as reliable second readers or triage tools that flag suspicious mammograms for priority review. As explainable artificial intelligence methods mature and open datasets grow, the combination of transfer learning, augmentation, and transparent decision-making may bring automated breast cancer detection from the research laboratory into routine screening practice, helping radiologists catch cancers earlier, when treatment is most effective.
Subject of Research: Explainability-aware benchmarking of pre-trained deep learning models for breast cancer classification in mammographic images
Article Title: Explainability-aware benchmarking of pre-trained deep learning models for breast cancer classification in mammographic images
Article References: Yasin, E. T., Cinar, I., & Koklu, M. (2026). Explainability-aware benchmarking of pre-trained deep learning models for breast cancer classification in mammographic images. Neural Computing and Applications, 38(19), Article 795. https://doi.org/10.1007/s00521-026-12513-1
Image Credits: AI Generated
DOI: 10.1007/s00521-026-12513-1
Keywords: breast cancer, mammography, deep learning, transfer learning, data augmentation, DenseNet201, convolutional neural networks, medical imaging, explainable AI, cancer screening, computer-aided diagnosis, benchmarking


1 hour ago
7




















English (US) ·
French (CA) ·