EnDeC: An Enhanced Convolutional Neural Network for Brain Tumor Classification From Magnetic Resonance Imaging
Abstract:
Accurate classification of brain tumors from magnetic resonance imaging (MRI) is important for supporting timely diagnosis and subsequent clinical decision-making. In this study, EnDeC, an enhanced convolutional neural network, was developed for automated multiclass classification of brain MRI images. The Crystal Clean: Brain Tumors MRI Dataset, comprising 21,672 images categorized as glioma, meningioma, pituitary tumor, or normal, was used for model development and evaluation. Following data cleaning and class balancing, 12,264 images were retained for model training. Image preprocessing and augmentation procedures, including resizing, noise injection, and rotation, were applied. An efficient input pipeline was implemented using the TensorFlow tf.data application programming interface, with caching and prefetching incorporated to reduce input/output bottlenecks during training. The EnDeC architecture was constructed using successive convolutional and pooling layers for hierarchical feature extraction, followed by fully connected layers for multiclass classification. Model performance was evaluated using accuracy, precision, recall, F1-score, and specificity. Average values of 99.60%, 99.18%, 99.19%, 99.18%, and 99.74% were obtained for these metrics, respectively. Among the evaluated optimization configurations, the Adam optimizer yielded the highest test accuracy of 99.13%. These findings demonstrate that discriminative imaging features associated with multiple brain tumor categories can be effectively learned using the proposed convolutional neural network framework. Nevertheless, validation using independent, patient-level, multi-institutional datasets is required before clinical applicability can be established. The proposed approach provides a computational framework for automated brain MRI classification and may support the further development of computer-aided diagnostic systems for brain tumor assessment.1. Introduction
A brain tumor is an abnormal mass of cells that develops in the brain. Brain tumors have two categories: primary and secondary tumors. Primary brain tumors form in the brain tissue, while secondary brain tumors are a spread of cancer to the brain from a different part of the body like lungs or kidneys. Brain tumors can also be either benign or malignant. The benign tumors do not cause cancer, while malignant tumors do and are usually more aggressive. While benign tumors are not cancerous, they should not be ignored and early diagnosis, appropriate treatment and proper care are crucial.
The 2021 World Health Organization Classification of Tumors of the Central Nervous System provides a detailed classification of the different types of brain tumors. Adult-type diffuse gliomas are classified as astrocytoma, oligodendroglioma, and glioblastoma. Circumscribed astrocytic gliomas include pilocytic astrocytoma, pleomorphic xanthoastrocytoma, subependymal giant cell astrocytoma, chordoid glioma and astroblastoma. The glioneuronal and neuronal tumors include ganglioglioma, desmoplastic infantile ganglioglioma, dysembryoplastic neuroepithelial tumor, papillary glioneuronal tumor, rosette-forming glioneuronal tumor, myxoid glioneuronal tumor, diffuse leptomeningeal glioneuronal tumor, gangliocytoma, multinodular and vacuolating neuronal tumor, dysplastic cerebellar gangliocytoma (Lhermitte-Duclos disease), central neurocytoma, extraventricular neurocytoma, and cerebellar liponeurocytoma. There are several types of ependymal tumors, such as supratentorial ependymoma, posterior fossa ependymoma, spinal ependymoma, myxopapillary ependymoma, and subependymoma. Choroid plexus tumors can be choroid plexus papilloma, atypical choroid plexus papilloma or choroid plexus carcinoma. Medulloblastoma is a type of embryo tumor. Pineal tumors include pineocytoma, pineoblastoma and papillary tumor of pineal region. Tumors of the cranial nerves and paraspinal nerves include schwannoma, neurofibroma, perineurioma, hybrid nerve sheath tumor and paraganglioma. Meningioma is the representation of meningiomas. Non-meningothelial tumors are called mesenchymatous tumors, such as solitary fibrous tumor, hemangioblastoma (vascular) and rhabdomyosarcoma (skeletal muscle). Germ cell tumors are made up of mature teratoma, immature teratoma, teratoma with somatic-type malignancy, germinoma, embryonal carcinoma, yolk sac tumor, and choriocarcinoma. The tumors of the sellar region encompass adamantinomatous craniopharyngioma, papillary craniopharyngioma, pituitary adenoma/pituitary neuroendocrine tumor and pituitary blastoma (Louis et al., 2021). The proposed study is to concentrate on gliomas, meningiomas and pituitary tumors.
According to the statistical report published by the Central Brain Tumor Registry of the United States (CBTRUS) on primary brain and other central nervous system tumors diagnosed in the United States from 2015 to 2019 (Ostrom et al., 2022), the rise of tumor cases, and death and loss incurred due to brain and central nervous system tumors were beyond imagination. None of the age groups were immune to it. Children, adolescents, adults, and the elderly had been badly affected. A lot of lives had already been lost, and a great number were at risk. These statistics provide a clear indication of the severity of the problem and serve as an alarming signal of the need to remain vigilant. Although life and the world we live in offer many positive experiences, life is also accompanied by unexpected and adverse events. These may include natural disasters, such as cloudbursts, earthquakes, floods, and avalanches, as well as pandemics and serious diseases, including heart disease, brain tumors, and kidney failure. Humans have always striven to tackle and resolve the challenges they encounter, with brain tumors being a significant concern. Although not all brain tumors are fatal, some can be extremely dangerous and life-threatening. It's crucial to pay attention to all types of tumors to save as many lives as possible. Considerable progress has been made in combating and curing brain tumors, but the quest for better and more efficient solutions continues. Advances in the medical field have introduced various treatment methods, including surgery, chemotherapy, and radiation therapy. Ongoing research aims to enhance the effectiveness of these treatments. With the advent of machine learning, deep learning, and artificial intelligence, approaches to many aspects of life, including medical science, are evolving rapidly. Technological advancements continue to emerge rapidly in the current era of scientific and technological development. In recent years, machine learning has significantly contributed to improving tumor segmentation, classification, and detection. Deep learning, in particular, is driving substantial research efforts to simplify and enhance medical processes. This study proposes a modest contribution to this ongoing research, seeking to explore valuable insights and further the progress in the field of brain tumor analysis and treatment.
This study aims to propose a deep learning model based on the convolutional neural network architecture, a very common and acceptable technique often used for image analysis and recognition. The proposed deep learning technique is designed to provide an easy-to-understand framework for multiclass image classification. Specifically, the proposed model aims to classify brain tumors from magnetic resonance imaging (MRI) samples while achieving high classification performance.
A transdisciplinary approach is an innovative approach that transcends traditional disciplinary boundaries, integrating knowledge from various fields to address complex problems (Bharali, 2024; Ertas & Tate, 2024; Kapçiu et al., 2024). Digital engineering and science exemplify this by combining principles from computer science, engineering, mathematics, and domain-specific knowledge to solve intricate issues across diverse domains, including healthcare. The application of a deep learning convolutional neural network model for enhanced brain tumor detection and classification exemplifies transdisciplinary research. It integrates:
(i) Medical Knowledge: Expertise in brain anatomy and pathology.
(ii) Computer Science: Development of algorithms and machine learning models.
(iii) Engineering: Ensuring the robustness and scalability of diagnostic tools.
(iv) Mathematics: Improving model accuracy and reliability through statistical methods.
This study employs a transdisciplinary approach by combining advanced medical imaging techniques (MRI and computed tomography scans), deep learning algorithms, data science for large dataset analysis, and clinical collaboration to validate and implement the proposed model in real-world settings.
The key contributions of this study are as follows:
(i) Development of a deep learning model based on convolutional neural networks for multiclass classification of brain MRI scans.
(ii) Presentation of an easy-to-understand and reproducible framework that balances model accuracy with computational efficiency, making it accessible for both clinical and academic use.
(iii) Integration of a transdisciplinary approach by combining medical expertise, advanced imaging techniques, and machine learning methods to enhance diagnostic reliability.
(iv) Evaluation of the model on benchmark datasets to demonstrate its effectiveness and potential for supporting clinical decision-making in real-world healthcare environments.
2. Literature Review
With technological progress, brain tumor research has grown significantly and huge amounts of data are available via numerous online repositories. This section reviews some previous studies concerning brain tumor detection, segmentation and classification using machine learning and artificial intelligence techniques. These studies have played an important role in the advancement of automated methods for brain tumor analysis. Zhao et al. (2018) proposed a model that combined fully convolutional neural networks and conditional random fields for brain tumor segmentation using deep learning. Imaging data from the Multimodal Brain Tumor Segmentation Challenge (BraTS) 2013, BraTS 2015 and BraTS 2016 were used for the study. The proposed approach yielded good performance on the BraTS 2013 and BraTS 2015 test sets and ranked first in the multi-temporal evaluation of the BraTS 2016 test set. A three-dimensional conditional random field was used as a post-processing method to further enhance the segmentation performance.
Abd El Kader et al. (2021) modeled a deep wavelet auto-encoder for brain tumor detection and classification using MRI images. Five magnetic resonance brain image databases were used in this study: BraTS 2012, BraTS 2013, BraTS 2014, BraTS 2015, and Ischemic Stroke Lesion Segmentation (ISLES). The deep wavelet auto-encoder model was successful in analyzing the pixel patterns of magnetic resonance brain images with high accuracy, fast detection speed and low validation loss. The model achieved an average accuracy of 99.3%, a sensitivity of 95.6%, a specificity of 96.9%, a precision of 97.4%, a false positive rate of 0.0625, a false negative rate of 0.031 and a Jaccard similarity index of 93.3%. Gull et al. (2021) used convolutional neural networks to automatically detect brain tumors from MRI. To improve the features extracted from brain MRI images, deep transfer learning techniques were employed and the image classification was performed by using the convolutional neural network architecture of GoogleNet. The model was evaluated on the BraTS 2018, BraTS 2019 and BraTS 2020 datasets, yielding accuracies of 96.50%, 97.50%, and 98% for brain tumor segmentation and 96.49%, 97.31%, and 98.79% for brain tumor classification, respectively.
To identify brain tumors, Amin et al. (2022) created a brain tumor detection model using ensemble transfer learning and a quantum variational classifier. Three benchmark datasets (Kaggle, BraTS 2020, and a set of locally collected images) were selected to test the proposed approach. The model was divided into three parts: feature extraction, classification, and segmentation. Feature extraction was performed with a pretrained Inception-v3 model, and the score vector that the model produced was passed to a quantum learning mechanism for tumor classification. Later, SegNet was used to segment the tumor lesions. The model gave an accuracy rate of 99.44% for the "no tumor" class, 99.25% for the "meningioma" class, 98.03% for the "pituitary tumor" class and 99.34% for the "glioma" class on the Kaggle dataset. An accuracy of 93.33% was achieved when the locally collected images were classified as tumor or non-tumor. The model on the BraTS 2020 dataset obtained an accuracy of 90.91% in the classification of high-grade glioma and low-grade glioma slices. The modified SegNet also obtained a global accuracy of 98.20% on the Kaggle dataset, 99.9% on the privately collected images and 99.70% on the BraTS 2020 dataset.
Yazdan et al. (2022) developed an efficient multi-scale convolutional neural network for multiclass classification of brain MRI images as Software as a Medical Device (SaMD). This research was based on an open-source Kaggle dataset consisting of 3,264 MRI images classified into four groups: glioma (926), meningioma (937), pituitary tumor (500), and non-tumor (901). The proposed architecture was a multi-scale convolutional neural network (MCNN) with three versions: MCNN1, MCNN2 and MCNN3, with two, three and four parallel convolutional neural network paths, respectively. Among the three architectures, MCNN2 using denoised images achieved the highest accuracy of 94.19%, with a precision of 94.45%, recall of 93.74%, specificity of 92.62%, and F1-score of 94.06%. More complex neural network architectures have been used to improve brain tumor detection. Maqsood et al. (2022) experimented on two publicly available datasets: Figshare MRI dataset, and BraTS 2018. The proposed approach used a 17-layer convolutional neural network for brain tumor segmentation, MobileNetV2 for feature extraction and a multiclass support vector machine framework for the brain tumor detection. The model was tested on the BraTS 2018 dataset and the Figshare dataset, with 97.47% and 98.92% accuracies, respectively.
Khan et al. (2022) used two datasets: the Figshare dataset for multiclass classification of glioma, meningioma, and pituitary tumors, and the Harvard Medical Dataset for binary classification into tumor and no-tumor classes. The study utilized a 23-layer convolutional neural network and a fine-tuned convolutional neural network based on transfer learning with the 16-layer Visual Geometry Group network (VGG16) model architecture. The proposed approach yielded prediction accuracies of 100% in the Harvard Medical Dataset and 97.8% in the Figshare dataset. ZainEldin et al. (2022) proposed a deep learning-based method for brain tumor detection and classification incorporating an adaptive dynamic sine-cosine fitness grey wolf optimizer. The proposed method was evaluated using the BraTS 2021 Task 1 dataset and compared with a convolutional neural network-based brain tumor classification model. The adaptive dynamic sine-cosine fitness grey wolf optimizer was used to optimize the convolutional neural network parameters. The proposed model was able to attain an accuracy of 99.99%. The added optimization process, however, required a significant amount of processing time, which is a significant drawback in the approach.
Abdusalomov et al. (2023) conducted a study using a dataset obtained from the publicly available Kaggle MRI dataset, which included four classes, namely, glioma (2,548 images), meningioma (2,582 images), pituitary (2,658 images) and no tumor (2,500 images). The study incorporated the Convolutional Block Attention Module (CBAM) attention mechanism, Spatial Pyramid Pooling Fast+ and bi-directional feature pyramid network modules into the You Only Look Once version 7 (YOLOv7) model. A prediction accuracy of 99.5% was achieved using the enhanced model. In 2023, Ullah et al. (2023) introduced TumorDetNet, a comprehensive deep learning model designed for both brain tumor detection and classification. This model utilized six datasets for three distinct tasks: the Tumor_Detection_MRI and Brain MRI Scans for Brain Tumor Detection datasets from Kaggle for brain tumor detection; the Tumor Classification Data and BTTypes datasets from Kaggle for binary tumor classification; and the Brain Tumor Classification dataset from Kaggle and the contrast-enhanced MRI dataset from Figshare for multiclass tumor classification. The enhanced MobileNetV2 model employed in this study integrated the activation functions of the leaky rectified linear unit and the rectified linear unit and consisted of a total of 134 layers. The TumorDetNet model demonstrated remarkable performance with a 99.83% accuracy in brain tumor detection, a 100% accuracy in binary tumor classification, and a 99.27% accuracy in multiclass classification. However, its complexity is acknowledged as a notable characteristic alongside its high effectiveness.
3. Dataset
The Crystal Clean: Brain Tumors MRI Dataset, which can be easily accessed at Kaggle (Hashemi, 2023), was utilized by the proposed deep learning convolutional neural network model. The dataset contained a total of 21,672 images with four classes: glioma (6,307), meningioma (6,391), pituitary (5,908), and normal (3,066). The proposed model was experimented on the reduced balanced dataset such that each class contained 3,066 images, as demonstrated in Figure 1.

To balance the dataset, 3,066 images were randomly retained from each class, resulting in a balanced dataset of 12,264 images. The resulting dataset was divided into training, validation, and testing subsets, with 1,248 images allocated to the test set.
A sample of MRI images is shown in Figure 2.

4. Methodology
This study aims to put forward an easy-to-understand methodology which can help in the effective multiclass classification of images and tumor detection. Data cleaning and augmentation were performed for the MRI images. The data cleaning processes included removal of duplicate samples, correction of mislabeled images, and image resizing to 224 × 224. The data augmentation processes included salt-and-pepper noise, histogram equalization, rotation, horizontal and vertical flipping, and brightness adjustment (Hashemi, 2023). Data preprocessing was also performed for the experimental dataset utilized in the study, including resizing, rescaling, random rotation, and random flipping. As shown in Figure 3, the dataset was divided at the image level into training, validation, and testing subsets; patient-level separation was not performed. Augmentation was restricted to the training set.

An efficient input pipeline was implemented using the TensorFlow tf.data application programming interface. Prefetching helped to overlap data preprocessing with model execution during a training step. While the model was trained on one data batch, prefetching prepared the next batch before it was requested. While the graphics processing unit was busy, the central processing unit prefetched the data and kept them ready in a buffer for the next step. Caching allowed a dataset to be stored either in memory or on local storage. File opening and reading were performed only during the first epoch, while subsequent epochs reused the data cached in memory or local storage. Prefetching and caching optimized the performance of the proposed model (TensorFlow, 2023). The proposed deep learning model architecture was carefully designed and developed to achieve a highly optimized and efficient model capable of producing accurate results. The convolutional layers, pooling layers, and fully connected layers were labeled as , , and , respectively, where x is the layer index (LeCun et al., 1998).
The sequence and configuration of the layers are as follows:
(a) The first convolutional layer (C1) included 32 filters, a 3 × 3 convolutional kernel, a rectified linear unit activation function, and a bias vector. Each filter corresponded to a feature map; hence, the first convolutional layer contained 32 feature maps, each with a size of 222 × 222. C1 contained 896 trainable parameters.
(b) The first max pooling layer (P1), with a pool size of and a stride of 2 × 2, contained 32 pooled feature maps, each with a size of 111 × 111.
(c) The second convolutional layer (C2) included 32 filters, a 3 × 3 convolutional filter, a rectified linear unit activation function, and a bias vector. Each filter corresponded to a feature map; hence, C2 contained 32 feature maps, each with a size of 109 × 109. C2 contained 9,248 trainable parameters.
(d) The second max pooling layer (P2), with a pool size of 2 × 2 and a stride of 2, contained 32 pooled feature maps, each with a size of 54× 54.
(e) The third convolutional layer, C3, included 64 filters, a 3 × 3 convolutional kernel, a rectified linear unit activation function, and a bias vector. Each filter corresponded to a feature map; hence, C3 contained 64 feature maps, each with a size of 52 × 52. C3 contained 18,496 trainable parameters.
(f) The third max pooling layer, P3, with a pool size of 2 × 2 and a stride of 2, contained 64 pooled feature maps, each with a size of 26 × 26.
(g) The fourth convolutional layer, C4, contained 64 filters, a 3 × 3 convolutional kernel, a rectified linear unit activation function, and a bias vector. Each filter corresponded to a feature map; hence, C4 contained 64 feature maps, each with a size of 24 × 24. C4 contained 36,928 trainable parameters.
(h) The fourth max pooling layer, P4, with a pool size of 2 × 2 and a stride of 2, contained 64 pooled feature maps, each with a size of 12 × 12.
(i) The fifth convolutional layer, C5, contained 128 filters, a 3 × 3 convolutional kernel, a rectified linear unit activation function, L2 regularization with a lambda value of 0.0067, and a bias vector. Each filter corresponded to a feature map; hence, C5 contained 128 feature maps, each with a size of 10 × 10. C5 contained 73,856 trainable parameters.
(j) The fifth max pooling layer, P5, with a pool size of 2 × 2 and a stride of 2, contained 128 pooled feature maps, each with a size of 5 × 5.
(k) A flatten layer contained 3,200 flattened features.
(l) A hidden dense layer, F1, contained 256 neurons, a rectified linear unit activation function, a bias vector, and a lambda value of 0.0067 for L2 regularization. F1 contained 819,456 trainable parameters.
(m) A dense output layer, F2, contained 4 neurons, a softmax activation function, and a bias vector. F2 contained 1,028 trainable parameters.
The model contained a total of 959,908 parameters, representing the weights and biases of the neural network. The number of total trainable parameters was the same as that of total parameters. Trainable parameters can be updated during training with back propagation, while non-trainable parameters remain fixed. The details are shown in Table 1.
Layer (Type) | Output Shape | Number of Parameters |
Sequential (Sequential) | (None, 224, 224, 3) | 0 |
Sequential_1 (Sequential) | (None, 224, 224, 3) | 0 |
C1: Conv2D | (None, 222, 222, 32) | 896 |
P1: MaxPooling2D | (None, 111, 111, 32) | 0 |
C2: Conv2D | (None, 109, 109, 32) | 9,248 |
P2: MaxPooling2D | (None, 54, 54, 32) | 0 |
C3: Conv2D | (None, 52, 52, 64) | 18,496 |
P3: MaxPooling2D | (None, 26, 26, 64) | 0 |
C4: Conv2D | (None, 24, 24, 64) | 36,928 |
P4: MaxPooling2D | (None, 12, 12, 64) | 0 |
C5: Conv2D | (None, 10, 10, 128) | 73,856 |
P5: MaxPooling2D | (None, 5, 5, 128) | 0 |
Flatten (Flatten) | (None, 3,200) | 0 |
F1: Dense | (None, 256) | 819,456 |
F2: Dense (output) | (None, 4) | 1,028 |
Total parameters: 959,908 Trainable parameters: 959,908 Non-trainable parameters: 0 | – | – |
An input image with dimensions of 224 × 224 was processed using the proposed convolutional neural network architecture. Convolutional layers were used as feature extractors, employing linear convolution operations and non-linear activation functions, such as the rectified linear unit, to introduce non-linearity into the network. These layers were used to extract features from the input image, starting with simple edges and gradually progressing to more complex patterns, which were assembled by deeper layers to extract semantically meaningful object parts or specific types of objects (Yamashita et al., 2018). Each convolution layer was followed by a max pooling layer, helping to achieve spatial invariance by reducing the resolution of the feature maps. Max pooling was used to map a subregion to its maximum value within each input patch of each feature map, helping to retain important information while reducing computational complexity. After the convolutional and pooling layers, the feature maps were flattened into a one-dimensional column vector. This vector served as the input layer to the fully connected layers (Wang et al., 2020). Flattening occurred before the fully connected layers, where each input neuron was connected to every output neuron through learnable weights. The flattened feature vector was then connected to the neurons of the hidden dense layer (F1), and the dense output layer (F2) received input from all neurons of F1. Furthermore, F2 was used to encode the object class with four neurons for multiclass classification.
Finally, the output layer utilized the softmax function to normalize the output real values from the last fully connected layer (F2) into target class probabilities. These probabilities ranged between 0 and 1, with all values summing up to 1, facilitating the final prediction and multiclass classification of images (Scherer et al., 2010). Figure 4 illustrates the proposed deep learning convolutional neural network model architecture, depicting the sequential flow of operations outlined above.

L2 regularization was applied to the fifth convolutional layer (C5) and the fully connected layer (F1), while bias vectors were incorporated into all convolutional and fully connected layers to improve and optimize the model’s performance. For convolutional neural network compilation, sparse categorical cross-entropy was used as the loss function, accuracy was selected as the evaluation metric, and Adam, Adamax, and root mean squared propagation were employed as optimizers.
The model was trained using the training dataset for 120 epochs with a batch size of 32 and a verbose value of 1. During the training process, the validation dataset was used to monitor and assess model performance. Two callback functions, namely learning-rate reduction on plateau and early stopping, were incorporated. The initial learning rate was set to 0.001, while early stopping was configured with a patience of ten epochs based on the validation loss. The learning rate was reduced by a factor of 0.5 when no improvement was observed for five consecutive epochs. A random seed of 42 was maintained to ensure reproducibility. The experiments were conducted using TensorFlow 2.13.0 and Python 3.10 on an NVIDIA GPU.
5. Experimental Results and Analysis
The model was trained on the training dataset and evaluated on the validation dataset. Figure 5 displays the training logs of the proposed deep learning convolutional neural network model.

To find the best fit model for this study, the outcomes were recorded for different optimizers, including root mean squared propagation, Adam, and Adamax. For root mean squared propagation, the training, validation, and test accuracies were 97.73%, 97.91%, and 97.55%, respectively, and the training, validation, and test losses were 11.59%, 10.97%, and 11.34%, respectively. For Adam, the training, validation, and test accuracies were 99.36%, 99.10%, and 99.20%, respectively, and the training, validation, and test losses were 6.241%, 6.65%, and 6.51%, respectively. For Adamax, the training, validation, and test accuracies were 98.68%, 99.22%, and 97.97%, respectively, and the training, validation, and test losses were 12.65%, 11.49%, and 13.53%, respectively. Among all, the Adam optimizer achieved the highest training and test accuracies, and the lowest training, validation, and test losses. Its validation accuracy was higher than that of root mean squared propagation, but was slightly lower than that of Adamax. Thus, the Adam optimizer was selected for the final model based on its observed performance across the evaluated metrics. Table 2 shows the records obtained.
No. | Epochs | L2 Regularization | Output Activation | Optimizer | Training Accuracy | Training Loss | Validation Accuracy | Validation Loss | Test Accuracy | Test Loss |
1 | 120* | 0.0067 | Softmax | Root mean squared propagation | 97.73% | 11.59% | 97.91% | 10.97% | 97.55% | 11.34% |
2 | 120 | 0.0067 | Softmax | Adam | 99.36% | 6.24% | 99.10% | 6.65% | 99.20% | 6.51% |
3 | 120** | 0.0067 | Softmax | Adamax | 98.68% | 12.65% | 99.22% | 11.49% | 97.97% | 13.53% |
The models were trained for up to 120 epochs using the respective optimization algorithms. Among the evaluated optimizers, Adam achieved the highest training, validation, and test accuracies. Therefore, the Adam optimizer was selected for the final model. Figure 6 displays the training loss and validation loss curves for the model. Initially, the training and validation losses were relatively high; however, as the number of epochs increased, the losses decreased to relatively small values, indicating convergence of the model.

Figure 7 displays the training accuracy and validation accuracy curves for the model. Initially, low training and validation accuracies were observed. However, as the number of epochs increased, the curves improved, and the accuracies rose considerably close to 1.0.

Table 3 displays the classification result of the model. The table displays precision, recall, F1-score, accuracy, macro average, and weighted average scores.
Precision | Recall | F1-Score | Support | |
Glioma tumor | 0.99 | 0.98 | 0.99 | 317 |
Meningioma tumor | 0.98 | 0.99 | 0.98 | 295 |
Normal | 1.00 | 1.00 | 1.00 | 281 |
Pituitary tumor | 1.00 | 1.00 | 1.00 | 355 |
Accuracy | – | – | 0.99 | 1,248 |
Macro average | 0.99 | 0.99 | 0.99 | 1,248 |
Weighted average | 0.99 | 0.99 | 0.99 | 1,248 |
Figure 8 shows the confusion matrix of the model. For the glioma tumor, there were 311 true positive cases, 928 true negative cases, 6 false negative cases, and 3 false positive cases. For the meningioma tumor, there were 291 true positive cases, 948 true negative cases, 4 false negative cases, and 5 false positive cases. For normal images, there were 281 true positive cases, 966 true negative cases, and 1 false positive case. For the pituitary tumor, there were 355 true positive cases, 892 true negative cases, and 1 false positive case. The test dataset contained a total of 1,248 MRI samples.

Figure 9 presents the receiver operating characteristic curves and the corresponding area under the curve values for all classes. For each class, receiver operating characteristic curves are very close to 1.0, and the corresponding area under the curve values are close to 1.00.

Figure 10 displays a sample of MRI images of the unseen test dataset predicted by the proposed deep learning convolutional neural network model along with the confidence scores. All the samples in the figure were correctly predicted.

Table 4 and Figure 11 display the evaluation metrics of the model. For glioma, the accuracy, precision, recall, F1-score, specificity, and sensitivity were 99.28%, 99.04%, 98.11%, 98.57%, 99.68%, and 98.11%, respectively. For meningioma, the accuracy, precision, recall, F1-score, specificity, and sensitivity were 99.28%, 98.31%, 98.64%, 98.48%, 99.48%, and 98.64%, respectively. For normal class, the accuracy, precision, recall, F1-score, specificity, and sensitivity were 99.92%, 99.64%, 100%, 99.82%, 99.90%, and 100%, respectively. For the pituitary, the accuracy, precision, recall, F1-score, specificity, and sensitivity were 99.92%, 99.72%, 100%, 99.86%, 99.89%, and 100%, respectively. The average accuracy, precision, recall, F1-score, specificity, and sensitivity were 99.60%, 99.18%, 99.19%, 99.18%, 99.74%, and 99.19%, respectively.
No. | Class | True Positive | True Negative | False Positive | False Negative | Accuracy | Precision | Recall | F1-Score | Specificity | Sensitivity |
1 | Glioma tumor | 311 | 928 | 3 | 6 | 99.28% | 99.04% | 98.11% | 98.57% | 99.68% | 98.11% |
2 | Meningioma tumor | 291 | 948 | 5 | 4 | 99.28% | 98.31% | 98.64% | 98.48% | 99.48% | 98.64% |
3 | Normal | 281 | 966 | 1 | 0 | 99.92% | 99.64% | 100% | 99.82% | 99.90% | 100% |
4 | Pituitary tumor | 355 | 892 | 1 | 0 | 99.92% | 99.72% | 100% | 99.86% | 99.89% | 100% |
Average | 99.60% | 99.18% | 99.19% | 99.18% | 99.74% | 99.19% | |||||

6. Comparative Analysis
Table 5 compares the proposed model with some of the latest models described in the previous studies. Zhao et al. (2018) proposed a deep learning approach that used fully convolutional neural networks and conditional random fields to segment brain tumors. The method achieved the best performance in the multi-temporal evaluation of the BraTS 2016 dataset and also demonstrated strong performance on the BraTS 2013 and BraTS 2015 test datasets. Abd El Kader et al. (2021) proposed a deep wavelet auto-encoder model for brain tumor detection and classification. The method achieved an average accuracy of 99.3%, sensitivity of 95.6%, specificity of 96.9%, precision of 97.4%, a Dice similarity coefficient of 96.55%, a false positive rate of 0.0625, a false negative rate of 0.031, and a Jaccard similarity index of 93.3%. Gull et al. (2021) proposed a convolutional neural network for automated brain tumor detection. The segmentation accuracies were 96.50%, 97.50%, and 98%, and the classification accuracies were 96.49%, 97.31%, and 98.79%, respectively, on three datasets. Amin et al. (2022) proposed a model for brain tumor detection using ensemble transfer learning and a quantum variational classifier. The model achieved a global accuracy of 98.20% on the Kaggle dataset, 99.9% on the privately collected dataset, and 99.70% on the BraTS 2020 dataset. Yazdan et al. (2022) proposed an efficient multi-scale convolutional neural network based on multiclass brain MRI classification, with an accuracy of 94.19%, precision of 94.45%, recall of 93.74%, specificity of 92.62%, and an F1-score of 94.06%. Maqsood et al. (2022) used a deep neural network and a multiclass support vector machine for multimodal brain tumor detection, with accuracies of 97.47% and 98.92% for the BraTS 2018 and Figshare datasets, respectively. Khan et al. (2022) used a deep convolutional neural network for accurate brain tumor detection, which achieved accuracies of 100% on the Harvard Medical Dataset and 97.8% on the Figshare dataset. ZainEldin et al. (2022) used deep learning and sine-cosine fitness grey wolf optimization for brain tumor detection and classification, which achieved an accuracy of 99.99%. Abdusalomov et al. (2023) proposed a deep learning approach for brain tumor detection based on MRI with a 99.5% prediction accuracy. Ullah et al. (2023) proposed TumorDetNet, a unified deep learning model designed for brain tumor detection and classification. The approach obtained an accuracy of 99.83% in brain tumor detection, 100% in binary tumor classification and 99.27% in multiclass tumor classification. However, the convolutional neural network-based model proposed in the current study achieved average accuracy, precision, recall, F1-score, specificity, and sensitivity values of 99.60%, 99.18%, 99.19%, 99.18%, 99.74%, and 99.19%, respectively. The proposed model was found to be competitive in terms of various evaluation parameters compared to the various existing brain tumor detection and classification methods. Higher average accuracy, precision, sensitivity and specificity values were achieved compared to the values in the study by Abd El Kader et al. (2021). A higher classification accuracy was achieved than that proposed by Gull et al. (2021). The proposed model yielded a good accuracy, as reported by Amin et al. (2022), although its accuracy was slightly lower. The accuracy, precision, recall, F1-score, and specificity of the proposed model were higher than those of the approach proposed by Yazdan et al. (2022). Likewise, the proposed model achieved a higher accuracy than that proposed by Maqsood et al. (2022). Some results of previous studies are more accurate in some instances, but the proposed model is still competitive at the overall accuracy level. The proposed model achieved a marginally lower accuracy than that of the approach proposed by ZainEldin et al. (2022), but its accuracy was higher than that of approaches proposed by Abdusalomov et al. (2023) and Ullah et al. (2023). The competitive experimental results of the proposed model show its potential for effective brain tumor detection and classification.
Study | Dataset | Technique | Results |
Zhao et al. (2018) | Multimodal Brain Tumor Segmentation Challenge (BraTS) 2013, BraTS 2015, and BraTS 2016 | Fully convolutional neural networks and conditional random fields | The model ranked first in the multi-temporal evaluation on the BraTS 2016 dataset and demonstrated promising performance on the BraTS 2013 and BraTS 2015 test datasets. |
Abd El Kader et al. (2021) | BraTS 2012, BraTS 2013, BraTS 2014, BraTS 2015, and Ischemic Stroke Lesion Segmentation (ISLES) | Deep wavelet auto-encoder model | Average accuracy: 99.3%; sensitivity: 95.6%; specificity: 96.9%; precision: 97.4%; Dice similarity coefficient: 96.55%; false positive rate: 0.0625; false negative rate: 0.031; Jaccard similarity index: 93.3% |
Gull et al. (2021) | BraTS 2018, BraTS 2019, and BraTS 2020 | Fully convolutional neural network + conditional random field for segmentation; GoogleNet with transfer learning for classification | Segmentation accuracy: 96.50%, 97.50%, and 98.00%; classification accuracy: 96.49%, 97.31%, and 98.79% on the BraTS 2018, BraTS 2019, and BraTS 2020 datasets, respectively |
Amin et al. (2022) | Kaggle dataset, BraTS 2020, and locally gathered images | Convolutional neural network and quantum neural network | Accuracy: 98.20% on the Kaggle dataset, 99.70% on BraTS 2020, and 99.9% on locally collected images |
Yazdan et al. (2022) | Kaggle dataset | Multi-scale convolutional neural networks (MCNN1, MCNN2, and MCNN3) | MCNN2—accuracy: 94.19%; precision: 94.45%; recall: 93.74%; specificity: 92.62%; F1-score: 94.06% |
Maqsood et al. (2022) | Figshare magnetic resonance imaging (MRI) dataset and BraTS 2018 | Convolutional neural network, MobileNetV2, and multiclass support vector machine | Accuracies of 98.92% and 97.47% |
Khan et al. (2022) | Figshare dataset and Harvard Medical Dataset | Convolutional neural network and fine-tuned 16-layer Visual Geometry Group network (VGG16) | Accuracies of 97.8% and 100% |
ZainEldin et al. (2022) | BraTS 2021 Task 1 dataset | Convolutional neural network-based brain tumor classification model | Accuracy: 99.99% |
Abdusalomov et al. (2023) | Publicly available MRI dataset obtained from Kaggle | Enhanced You Only Look Once version 7 (YOLOv7) with Convolutional Block Attention Module (CBAM), Spatial Pyramid Pooling Fast+, and bi-directional feature pyramid network | Accuracy: 99.5% |
Ullah et al. (2023) | Tumor_Detection_MRI, Brain MRI Scans for Brain Tumor Detection, Tumor Classification Data, BTTypes, Brain Tumor Classification, and contrast-enhanced MRI datasets | TumorDetNet | Accuracy: 99.83% for brain tumor detection; 100% for brain tumor classification; 99.27% for multiclass classification |
Proposed model | Crystal Clean: Brain Tumors MRI Dataset | Convolutional neural network | Average accuracy: 99.60%; average precision: 99.18%; average recall/sensitivity: 99.19%; average F1-score: 99.18%; average specificity: 99.74%; average area under the curve: 1.00 |
Unlike many existing approaches that rely on complex hybrid models, heavy feature engineering, or fine-tuned pre-trained networks, the proposed convolutional neural network-based framework demonstrates competitive performance while maintaining a relatively simple and reproducible architecture, making it more practical for potential use in computer-aided medical image analysis. In addition, the model demonstrates consistently high values for accuracy, precision, recall, and specificity across different evaluation metrics, ensuring reliability in critical diagnostic scenarios. The approach also emphasizes ease of implementation, which enables seamless integration into existing healthcare systems without requiring extensive computational resources or specialized hardware.
7. Discussion
The study introduces a deep learning model based on the convolutional neural network architecture, leveraging the Crystal Clean: Brain Tumors MRI Dataset available on Kaggle. This dataset comprises four classes, namely, glioma, meningioma, pituitary, and normal, with a total of 12,264 MRI scanned samples utilized in the experimentation. Various preprocessing stages applied to the dataset include removal of duplicate samples, correction of mislabeled images, resizing, histogram equalization, rotation, flipping, and rescaling.
The proposed deep learning model comprises convolution layers, pooling layers, and fully connected layers. Convolutional layers extract features from the input data, while max pooling downsamples the feature maps and highlights significant features. Flattening converts the pooled feature maps into a one-dimensional feature vector, which acts as the input layer for the fully connected layers. The final fully connected output layer categorizes MRI images into their respective classes. Throughout training, the model shows smooth loss and accuracy curves, with decreasing loss and increasing accuracy as epochs progress, as depicted in Figure 6 and Figure 7. The receiver operating characteristic curves approach a value close to 1.0, indicating optimal performance, as shown in Figure 9. The area under the receiver operating characteristic curve is nearly 1.00, further confirming the model's high performance. In Table 4, the model demonstrates average accuracy, precision, recall, F1-score, and specificity values of 99.60%, 99.18%, 99.19%, 99.18%, and 99.74%, respectively, demonstrating high classification performance on the evaluated dataset. Figure 10 illustrates the correct predictions made by the model on the images. Furthermore, the proposed model exhibits substantial performance compared to other models, as evidenced by the comparative model analysis presented in Table 5.
The enhanced performance of the proposed convolutional neural network model can be attributed to several critical design and preprocessing choices. The rigorous dataset refinement, including duplicate removal, noise reduction, and histogram equalization, ensures high-quality and balanced input data, minimizing bias and improving feature clarity. Data augmentation strategies such as rotation and flipping expand the effective training set, enabling the model to generalize better across unseen samples. Architecturally, the careful combination of convolutional and pooling layers allows efficient hierarchical feature extraction, while the fully connected layers ensure robust classification. Additionally, the simplicity of the model design avoids overfitting and reduces computational overhead, allowing the network to achieve competitive performance while maintaining scalability and reproducibility. These factors collectively explain the superior evaluation metrics achieved in this study compared to existing approaches. However, the study is limited to a single public dataset and does not include external clinical validation or independent scanner-based testing, which may limit the generalizability of the results. Moreover, patient-level splitting was not performed, which may introduce data leakage when multiple images from the same patient are present.
8. Conclusion
In this study, a convolutional neural network model was developed for the multiclass classification of brain MRI images and detection of tumors. The proposed model achieved average accuracy, precision, recall or sensitivity, F1-score, and specificity values of 99.60%, 99.18%, 99.19%, 99.18%, and 99.74%, respectively. The area under the receiver operating characteristic curve was approximately equivalent to 1.00 for all the classes of the dataset. Thus, the model demonstrates high classification performance on the evaluated dataset, while further validation is required before clinical use. The proposed model can classify images into multiple categories using techniques like convolutional neural networks. Additionally, it can classify MRI images into different tumor and non-tumor categories. This classification capability may support computer-aided analysis of brain MRI images. The proposed model should therefore be considered a preliminary computer-aided classification approach rather than a clinically validated diagnostic tool. In future work, the model could be trained on a larger dataset. A hybrid deep learning model could also be incorporated. With the guidance of radiologists and health experts, the model could be further validated using clinical data for potential real-world application.
Conceptualization: M.Z. and M.B.; methodology: S.A.F. and I.A.; software, M.B.; validation, I.A. and S.A.F.; formal analysis, M.B.; investigation, I.A.; resources, I.A.; data curation, S.A.F.; writing—M.B.; visualization: M.B.; supervision: M.Z.; project administration: M.Z. All authors have read and agreed to the published version of the manuscript.
The data used to support the research findings are available from the corresponding author upon request.
The authors declare no conflicts of interest
