Radar-Based Drone Target Recognition Using Machine Learning Under Clutter
Abstract:
Accurate classification of drone targets using radar systems is becoming increasingly important for surveillance, airspace monitoring, and security applications. This paper presents a machine learning-based framework for radar drone detection and classification using a synthetically generated radar dataset under both ideal and realistic operating conditions. The dataset incorporates key radar features, including radar cross section (RCS), Doppler shift, range, range resolution, aspect angle, and micro-Doppler signatures generated by rotating propellers, along with realistic clutter and noise associated with ground, weather, and urban environments to emulate practical radar scenarios. Three machine learning classifiers, Support Vector Machine (SVM), Random Forest (RF), and Boosted Trees (BT), were trained and evaluated. Under ideal conditions, all three models achieved classification accuracies close to 99%. When realistic clutter and noise were introduced, performance degraded substantially, with accuracy dropping to 78.8% for both SVM and RF and 78.5% for BT. Feature selection and hyperparameter tuning via Bayesian Optimization (BO) recovered comparable overall accuracies of 86.8–87.2% across all three classifiers. However, the optimized models exhibited a consistent trade-off rather than a single dominant classifier: SVM achieved the highest recall (90.4%), RF achieved the highest Area Under the Receiver Operating Characteristic Curve (AUC) (0.953) and the lowest false-positive rate (3.4%), and BT achieved the highest overall accuracy (87.2%) and F1-score (87.15%). No single classifier dominates across all metrics, so identifying the most suitable model requires an explicit evaluation priority. This study adopts the position that, in surveillance and security contexts, missed detections typically carry a greater operational cost than false alarms; under this recall-weighted priority, SVM is identified as the preferred classifier, while RF is preferable where minimizing false alarms is instead the priority. These results demonstrate that optimized machine learning techniques, combined with radar-derived target features, can substantially improve drone classification robustness in cluttered environments, while highlighting that classifier selection should be guided by the relative cost of missed detections versus false alarms specific to the deployment context.
1. Introduction
Early work on radar-based drone classification focused on single features. Kim et al. [1] trained a convolutional neural network on merged Doppler images to separate drone types. This was one of the first deep learning approaches in this space, and it showed that Doppler imagery alone carries enough structure for classification. Small, low-altitude drones are precisely the class of target most easily masked by clutter, since their weak, slowly moving returns sit close to the same range-Doppler region as ground, weather, and urban interference; this overlap is widely regarded as the central obstacle to reliable radar-based drone recognition [2]. Two families of radar target signature dominate the literature on this problem: high-resolution range profiles (HRRP), which resolve the target’s physical extent along the radar line of sight, and micro-Doppler signatures, which capture the modulation imposed by rotating blades or flapping wings. On the micro-Doppler side, Larrat and Sales [3] trained Long Short-Term Memory (LSTM), Gated Recurrent Unit (GRU), Conv1D, and Transformer models on 60 GHz millimeter-wave amplitude and phase data drawn from the dataset of Kang et al. [4], building on earlier convolutional classifiers such as that of Kumawat et al. [5] and on the established micro-Doppler signal model of Chen [6]. Their Multimodal Transformer, which augments the raw signal with skewness and kurtosis, remained accurate under four injected noise types, but only after considerable architectural and training effort, and the underlying deep models still presuppose a dataset large enough to train them. Fine-grained micro-Doppler and range-profile discrimination has also been pursued through multi-antenna hardware: Ogawa et al. [7] generated inverse synthetic aperture radar imagery with a millimeter-wave Multiple Input Multiple Output (MIMO) radar and classified it with a convolutional neural network, achieving high accuracy for several drone models but at the cost of the additional transmit and receive channels a MIMO architecture requires. Where clutter is addressed directly, it is almost always treated as a single, specific interference source rather than as a realistic mixture. Fan et al. [8] suppressed ground clutter in micro-Doppler data using orthogonal matching pursuit, improving signal clarity for that one interference type. At the track level, Liu et al. [9], building on the simulation-based work of Mohajerin et al. [10] and the field study of Chen et al. [11], used a small set of hand-engineered kinematic descriptors and a Random Forest (RF) classifier to separate drones, birds, and dynamic precipitation clutter from S-band surveillance tracks. This is the closest precedent to a lightweight, classical machine learning pipeline evaluated against clutter, yet it too considers only one clutter source, relies on coarse track data rather than radar cross section (RCS) or micro-Doppler information, and does not address feature or hyperparameter optimization beyond feature-importance ranking. Outside the Unmanned Aerial Vehicle (UAV) literature specifically, feature-fusion architectures that combine engineered and learned descriptors have repeatedly improved robustness under degraded or cluttered conditions, for instance in multikernel convolutional fusion for synthetic aperture radar targets [12] and attention-augmented detection for small aerial objects [13]. These results support the broader premise that careful feature and model optimization, rather than architectural scale alone, is what recovers performance lost to clutter.
Taken together, these studies show that both fine-grained, signal-level methods and coarser, track-level behavioral methods can be effective, but none of them jointly (i) targets the small-radar-cross-section, low-altitude drone regime specifically, (ii) evaluates performance under a realistically clutter-rich environment combining ground, weather, and urban sources rather than a single interference type, and (iii) does so using a single-aperture radar architecture that avoids the added hardware cost of multi-antenna systems. This study is designed to close that combined gap. It presents a machine learning-based framework for radar drone classification using a synthetically generated dataset that incorporates realistic radar characteristics and a jointly modeled multi-source clutter environment, evaluating three widely used classifiers, Support Vector Machine (SVM), RF, and Boosted Trees (BT), in their ability to distinguish drone targets under both ideal and non-homogeneous operating conditions using only a standard phased-array radar. Model performance is further analyzed before and after feature optimization and hyperparameter tuning, quantifying how much of the accuracy lost to clutter can be recovered without resorting to specialized hardware or to deep learning architectures that demand substantially larger training datasets.
The main contributions of this work are as follows:
(i) A synthetic radar dataset that jointly incorporates RCS, Doppler, micro-Doppler, and three distinct clutter models within a single unified phased-array radar simulation.
(ii) A systematic comparison of SVM, RF, and BT classifiers under matched ideal vs. cluttered conditions, isolating the accuracy degradation attributable to clutter alone.
(iii) A ReliefF-based feature selection strategy integrated with Bayesian Optimization (BO) and grid-search hyperparameter tuning to improve classification robustness under cluttered radar conditions.
(iv) Posterior probability calibration of the SVM output for more reliable confidence estimates in a decision-support context.
In addition, the performance of the models is analyzed before and after feature optimization and hyperparameter tuning to assess their effectiveness in improving classification robustness in complex radar environments.
2. Methodology
This study presents a machine learning-based framework for classifying drone targets from a generic non-drone class under both ideal and cluttered radar environments. A synthetic radar dataset was generated using Matrix Laboratory (MATLAB R2021B) to simulate target detection and classification scenarios under controlled and realistic operating conditions. The proposed framework consists of five main stages: radar dataset generation, feature extraction, data preprocessing, model training and optimization, and performance evaluation. To investigate the classification performance under different operating conditions, two synthetic datasets containing 1,000 samples each were generated. Each dataset initially comprised 700 drone samples and 300 non-drone samples. The first dataset represented an ideal clutter-free radar environment, whereas the second represented a realistic environment incorporating environmental clutter and additive noise. The non-drone category was modeled as a single generic class.
For each generated sample, the target and radar parameters were randomly varied within predefined class-specific ranges to introduce sample-to-sample variability in the simulated radar returns. As summarized in Table II, drone targets were assigned radial velocities ranging from 0 to 30 ms$^{-1}$, RCS values from 0.01 to 1 m$^2$, aspect angles from 0° to 180°, and SNR values from -5 to 30 dB. For non-drone targets, the corresponding radial velocity, RCS, aspect angle, and SNR ranges were 0–20 ms$^{-1}$, 0.001–10 m$^2$, 0°–180°, and -10–25 dB, respectively. The target range was varied from 50 to 1,000 m for both classes. The Doppler shift was associated with the target radial velocity and resulted in ranges of 0–1880.5 Hz for drone targets and 0–1253.7 Hz for non-drone targets. The parameters were independently varied across samples within their predefined ranges to represent different target motion states, observation geometries, reflectivity characteristics, and signal conditions. A fixed random seed of 42 was used in the MATLAB classification procedure to ensure reproducible randomized data partitioning and model evaluation. The ideal dataset was generated without environmental clutter or additive noise to provide a baseline representation of clean radar observations. In contrast, the realistic dataset incorporated additive white Gaussian noise and statistical models for different environmental clutter sources. Urban clutter was modeled using a log-normal distribution to represent the statistical characteristics of radar backscatter from urban environments. Weather clutter was modeled using a Rayleigh distribution to represent the statistical behavior of atmospheric backscatter, while ground clutter was modeled using a Weibull distribution to represent variations in land-surface radar backscatter characteristics. The received radar signal was generated using the monostatic radar signal model described by Fan et al. [8], which accounts for transmitted power, antenna gains, operating wavelength, target RCS, target range, system losses, carrier frequency, signal phase, and additive noise. In addition, drone targets were characterized using micro-Doppler signatures associated with rotating propeller blades, providing additional information related to the periodic motion of the target's rotating components. These parameters and signal characteristics were incorporated to produce radar observations with varying motion, range, reflectivity, aspect angle, SNR, and micro-Doppler characteristics.
Following radar-signal generation, relevant features were extracted from the simulated radar returns, including maximum SNR, mean Doppler frequency, standard deviation of Doppler frequency, range corresponding to the peak return, duration above the detection threshold, estimated RCS, peak received power, and micro-Doppler-related characteristics. These features were selected to represent complementary information concerning target motion, range, reflectivity, signal strength, and structural characteristics. During preprocessing, constant-valued features were removed because they provide no discriminatory information, and median imputation was used as a safeguard for any missing values. Additional engineered features, including the power-to-SNR ratio, Doppler coefficient of variation, RCS-to-range ratio, and SNR-range interaction, were also generated to enhance the discriminative representation of the radar data. The resulting features were standardized before classification, and informative features were selected using the ReliefF algorithm. For the realistic dataset, the initial class distribution of 700 drone and 300 non-drone samples was balanced during dataset preparation using SMOTE oversampling, resulting in 795 drone samples and 795 non-drone samples, for a total of 1,590 samples. The resulting balanced dataset was then divided using stratified holdout sampling, with 70% (1,113 samples) allocated to model development and the remaining 30% (477 samples) retained as an independent test set. The development partition was subsequently divided using a further stratified 20% holdout, resulting in 890 samples for model training and 223 samples for validation. The validation subset was used exclusively for decision-threshold tuning, whereas the independent test set remained isolated until the final performance evaluation. Thus, the reported test-set performance was obtained from data that were not used for classifier training or threshold selection.
3. System Design
The experiment was conducted in MATLAB R2021B, considering various radar operating conditions and target scenarios. Figure 1 illustrates the overall framework of the proposed machine learning-based model for classifying drone targets from non-drone objects. The evaluation was performed under three scenarios: (i) an ideal environment, where noise, clutter, and other interference sources were absent, resulting in clean radar returns; (ii) a non-homogeneous environment, simulated by incorporating multiple sources of clutter and noise, including ground, urban, and weather reflections, to represent realistic radar operating conditions; and (iii) a feature-selection and hyperparameter-optimization scenario, applied to improve classification performance in the presence of clutter.

As shown in Figure 1, the process begins with the initialization of radar parameters, followed by radar signal generation for both drone and non-drone objects. Two synthetic datasets were generated: one representing ideal operating conditions and the other representing realistic cluttered environments. The simulated radar returns were used to extract target features such as RCS, Doppler shift, micro-Doppler signatures, range, range resolution, and aspect angle. These features were then used to construct the radar dataset for training and testing the machine learning models. Finally, feature optimization and hyperparameter tuning were applied, and classification performance was evaluated using standard performance metrics.
4. Design Parameters
The dataset was generated using a radar signal model that represents both ideal and realistic operating environments. To simulate practical radar scenarios, the model incorporates additive noise and environmental clutter arising from ground, urban, and weather reflections. The received radar signal from a drone target is modeled using the monostatic radar range equation, which relates the transmitted power, RCS, antenna characteristics, operating wavelength, and target range. The received power can be expressed as [14]:
where, $P_t$ is the transmitted power, $G_t$ and $G_r$ are the transmitting and receiving antenna gains, respectively, $\lambda$ is the operating wavelength, $\sigma$ is the RCS of the target, $R$ is the target range, and $L$ represents the total system losses. In this study, the radar range equation is used to determine the received signal power as a function of target distance and radar reflectivity. For drone targets, the effective RCS varies with changes in aspect angle, orientation, flight maneuver, and structural configuration. These variations directly affect the strength of the received radar echoes and provide valuable information for target classification.
The RCS of a drone depends on its size, shape, material composition, and viewing angle. To model the radar scattering characteristics of drone targets, the drone body is approximated as an equivalent scattering object. The RCS is expressed as [14]:
where, $r$ is the effective radius of the drone. This simplified model enables the simulation of drones with different sizes and reflectivity characteristics for radar-based classification.
Doppler shift is a fundamental radar parameter that results from the relative motion between the radar and a target. It is one of the most important features for detecting and classifying moving objects. In a radar system, a signal with carrier frequency $f_c$ is transmitted through an antenna. When the transmitted signal encounters a moving drone, the reflected signal returns with a frequency shift proportional to the target's radial velocity $V_r$. This frequency change, known as the Doppler shift, provides valuable information about the motion characteristics of the target and is expressed as [15]:
where, $f_d$ denotes the Doppler frequency shift (Hz), $v_r$ denotes the radial velocity of the target (ms$^{-1}$), $f_c$ denotes the carrier frequency (Hz), and $c$ denotes the speed of light. The carrier frequency is assumed to be 9.4 GHz. The velocity of a drone depends on its propulsion system, payload, and flight maneuver. In this study, the radial velocity of the drone is modeled as a variable parameter within a predefined operating range to represent different drone types and flight conditions. The corresponding Doppler shift is calculated from the target radial velocity and is used as an important feature for radar-based drone classification [15].
where, $v_r$ is the radial velocity of the drone, and $\lambda$ is the radar wavelength.
The radar echo signal $s(t)$ received from a drone target is derived from the radar range equation and depends on parameters such as transmitted power $P_t$, antenna gain $G_t$, operating wavelength $\lambda$, RCS $\sigma$, and target range $R$. The received signal contains information about the target's reflectivity and motion characteristics. The exponential term represents the phase variation of the received signal, while $n(t)$ denotes additive noise and environmental clutter effects. This signal model is used to generate realistic radar returns for drone target classification [14].
Radar is a key sensor for detecting and classifying drone targets in surveillance systems. However, radar returns are often contaminated by environmental clutter and noise, making target recognition challenging. Due to their small RCS and low-altitude operation, drones can be difficult to distinguish from surrounding clutter. In this study, realistic radar conditions are simulated by incorporating multiple clutter sources, including ground, urban, and weather clutter, to evaluate the robustness of the proposed classification framework.
Ground clutter originates from reflections from the Earth's surface, terrain, buildings, and vegetation. It typically appears as strong stationary echoes and can significantly affect the detection of low-altitude drone targets. In this study, ground clutter is modeled using the Weibull distribution, which effectively represents variations in land-surface backscatter characteristics [15]:
where, $k$ denotes the shape parameter, $\Theta$ denotes the scale parameter, and the variable $x$ represents the value of the random variable whose distribution is being modeled. Weather clutter results from atmospheric phenomena such as rain, snow, clouds, and fog. It can significantly affect radar performance by introducing unwanted echoes that may obscure or distort drone returns, particularly for small targets. In this study, weather clutter is modeled using the Rayleigh distribution to represent the statistical characteristics of atmospheric backscatter [15]:
Urban clutter originates from buildings, vehicles, and other man-made structures in densely populated areas. It produces strong reflections that can interfere with drone detection and classification. In this study, urban clutter is modeled using the log-normal distribution to represent the statistical characteristics of radar backscatter in urban environments [15]:
where, $\mu$ is the log-mean, and $\sigma$ is the log-standard deviation of the distribution.
The two-dimensional array factor models the directional response of an M × N planar antenna array, where $d_x$ and $d_y$ denote the element spacings and $\beta_{m n}$ represents the applied phase excitation. This expression characterizes the azimuth and elevation beam formation and steering of the array. In target classification, the array factor is crucial because the antenna's spatial response influences angle-dependent measurements, thereby ensuring that the extracted target features reflect true target behavior rather than array-induced pattern variations. This shows that increasing the number of elements enhances directional gain and narrows the beam. For a 2D array (e.g., M × N) [16]:
where, $(\theta, \phi)=$ elevation and azimuth angles, $d_x d_y=$ element spacing in $x$ and $y$ directions, $\beta_{m n}=$ phase shift for element $(m, n)$. where, for steering to azimuth $\varphi_0$ and elevation $\theta_0$.
In this study, three supervised machine learning algorithms, namely SVM, RF, and BT, were employed for the classification of drone targets. Among these methods, the SVM identifies an optimal hyperplane that maximizes the separation between drone and non-drone classes, thereby improving classification performance. The mathematical formulation of the SVM classifier is given by [17]:
where, $\omega$ and $b$ represent the weight vector and bias, respectively.
The RF algorithm is an ensemble learning method that constructs multiple decision trees using bootstrap sampling of the training data. Each tree independently predicts the target class, and the final classification is determined through majority voting among all trees. This ensemble approach improves classification accuracy and reduces overfitting. The probability of classification is expressed as [18]:
where, $P(c k x)$ denotes the estimated probability that the input feature vector $x$ belongs to class $c$; $T$ is the number of trees; $h_t$ denotes each classifier; and $I(\cdot)$ is the indicator function, which returns 1 if $h_t(x)=c$ and 0 otherwise. Thus, the class probability is computed as the proportion of decision trees that vote for class $c$.
The boosted tree (specifically adaptive boosting) iteratively updates the model by minimizing the exponential loss expression as [19]:
This enhances weak learners’ performance through weighted re-training.
5. Simulation Parameters
Target detection is a fundamental function of radar systems; however, accurately classifying detected objects remains a significant challenge, particularly in cluttered environments. Since acquiring real radar datasets of drone targets under diverse operating conditions can be costly and time-consuming, a synthetic radar dataset was generated in this study using MATLAB R2021B, employing the phased array system toolbox for antenna-array and radar-signal simulation and the statistics and machine learning toolbox for ReliefF feature selection, classifier training (SVM, RF, and BT), and hyperparameter optimization.

A 9.4 GHz phased array radar system was simulated, employing a 10 × 10 phased array antenna consisting of 100 radiating elements arranged in a two-dimensional configuration, enabling high-resolution target detection, beam steering, and tracking capabilities. This radar configuration provides an effective platform for evaluating the proposed drone classification framework under both ideal and cluttered environments. The color bar represents the normalized array response in decibels (dB), where 0 dB corresponds to the maximum array gain (main lobe). Negative values indicate lower relative gain, with -50 dB representing directions in which the array response is highly attenuated. Figure 2 demonstrates that the proposed 10 × 10 uniform rectangular array concentrates its radiation in the desired steering direction while effectively suppressing sidelobes and unwanted responses. With a gain exceeding 35 dBi, the phased array radar produces a narrow beam width of only a few degrees, enabling accurate target localization, effective interference suppression, and long-range detection. Two scenarios were considered for dataset generation: an ideal environment and a non-homogeneous environment. To emulate realistic radar operation, multiple clutter combinations were incorporated into the non-homogeneous scenario. Key radar features such as RCS, Doppler shift, velocity, aspect angle, operating frequency, micro-Doppler signature, and range profile were modeled using standard radar equations. Drone targets were simulated with distinct RCS, velocity, and micro-Doppler characteristics produced by rotating propellers, while non-drone objects exhibited different radar signatures. The resulting dataset consisted of labeled time-series radar observations suitable for training and evaluating machine learning models, since radar systems continuously sample reflected echoes over time and the temporal sequence of observations contains important information for target classification. Table 1 shows the system parameters for simulation. Table 2 presents the parameters that are dependent on the target.
Parameters | Symbol | Value/Expression |
|---|---|---|
Antenna | — | Phased array antenna |
Propagation speed | $c$ | $3\times10^{8}$ m s$^{-1}$ |
Operating frequency | $f$ | 9.4 GHz |
Wavelength | $\lambda$ | 31.9 mm |
Pulse repetition frequency | $PRF$ | 1 kHz |
Bandwidth | $B$ | 20 MHz |
Antenna gain | $G$ | 35 dBi |
Maximum range | $R$ | 1 km |
Data sampling rate | $f_s$ | 40 MHz |
Clutter | $f(x)$ | ground, urban, and weather |
Range resolution | $\Delta R$ | 7.5 m |
Parameter | Symbol | Drone Target | Non-Drone Target |
|---|---|---|---|
Radial velocity | $v$ | 0–30 ms$^{-1}$ | 0–20 ms$^{-1}$ |
Doppler shift | $f_d$ | 0–1880.5 Hz | 0–1253.7 Hz |
Radar cross section (RCS) | $\sigma$ | 0.01–1 m$^2$ | 0.001–10 m$^2$ |
Signal-to-noise ratio | SNR | –5 to 30 dB | –10 to 25 dB |
6. Results and Analysis
Due to the limited availability of publicly accessible radar datasets for drone targets, a synthetic radar dataset was generated for drone target classification. The required radar parameters were defined, and realistic noise and clutter effects were incorporated into the simulation environment to emulate practical operating conditions. Figure 3 shows the visualization of drone and non-drone class under ideal conditions.

Two datasets were created: one under ideal clutter-free conditions and another under realistic conditions containing environmental clutter effects. The dataset includes key radar features such as RCS, Doppler shift, target velocity, angle of arrival, operating frequency, range, and micro-Doppler signatures, used for training and evaluating the machine learning classifiers. Under ideal conditions, where clutter and noise were absent, all three machine learning models achieved near-perfect classification performance, with an accuracy of approximately 99%.
This indicates that the models were able to effectively distinguish drone targets from non-drone objects when the radar features were clearly separable. However, when realistic environmental effects, including measurement noise and clutter from ground, urban, and weather clutter, were incorporated into the dataset, classification performance decreased significantly. Under these realistic conditions, the classification accuracy dropped to 78.8% for both the SVM and the RF, and 78.5% for the BT. These results highlight the challenges associated with drone classification in complex radar environments, where clutter and noise can obscure target signatures and reduce class separability. The receiver operating characteristic comparison of the three machine learning models for the cluttered dataset is presented in Figure 4. To improve classification performance under realistic radar conditions, BO and grid search were employed to optimize the hyperparameters of all three classifiers. BO proved particularly effective for models with complex and nonlinear hyperparameter interactions, such as the SVM and the BT.

Compared with conventional grid search, BO achieved higher classification performance while requiring fewer training iterations by adaptively exploring the most promising regions of the search space. After optimization, the precision increased to 86.98% for the SVM, 88.42% for the RF, and 88.02% for the BT, accompanied by corresponding improvements in F1-score and overall classification accuracy. These results demonstrate that hyperparameter optimization can significantly enhance the generalization capability and robustness of machine learning models for radar-based drone classification in cluttered environments. The receiver operating characteristic curves shown in Figure 5 illustrate the trade-off between the true positive rate (sensitivity) and the false positive rate for the three machine learning classifiers. Among the optimized models, the RF classifier achieved the highest Area Under the Receiver Operating Characteristic Curve (AUC) of 0.953, indicating superior discrimination between drone and non-drone class. The BT classifier achieved an AUC of 0.949, demonstrating strong overall discrimination with a comparatively low false positive rate (5.46%). In contrast, the SVM model recorded the lowest AUC of 0.937, indicating comparatively weaker class separability, but with the highest recall (90.38%) among the three classifiers.

Figure 6 shows how effectively the optimized SVM classifier separates drone from non-drone class in cluttered conditions. Figure 7 shows that the confusion matrix of the optimized SVM classifier correctly identified 216 drone targets, while 198 non-drone objects were correctly classified. The classifier produced 23 false negatives, corresponding to drone targets incorrectly classified as non-drone objects, and 40 false positives, corresponding to non-drone objects incorrectly classified as drone targets. This study targets surveillance and security applications, where missed detections are treated as more costly than false alarms. Therefore, recall is designated as a priority metric alongside accuracy, AUC, and F1-score in reporting classifier performance. These results yield an overall classification accuracy of 86.79%, with a drone detection recall of 90.38% and a false positive rate of 16.81%.


The confusion matrix of the optimized SVM classifier shows that 216 drone targets were correctly identified, while 198 non-drone objects were correctly classified. The classifier produced 23 false negatives, corresponding to drone targets incorrectly classified as non-drone objects, and 40 false positives, corresponding to non-drone objects incorrectly classified as drone targets. These results yield an overall classification accuracy of 86.79%, with a drone detection recall of 90.38% and a false positive rate of 16.81%.
All three models achieved similar overall accuracy, but each showed different strengths depending on the metric considered. The SVM classifier reached an accuracy of 86.79% and the highest drone detection recall among the three models, at 90.38%, showing that it was very effective at correctly identifying drone targets. However, this came at the cost of a higher false positive rate (16.81%), meaning it more often mistook non-drone objects for drones. The RF classifier achieved a slightly higher accuracy of 87.00%, but with a lower drone recall of 77.41%. In exchange, RF was much better at rejecting non-drone objects, achieving the highest specificity (96.64%) and the lowest false positive rate (3.36%) of the three models. The BT classifier obtained the highest overall accuracy (87.21%) and the highest macro F1-score (87.15%), reflecting a more balanced performance across both drone and non-drone classes. Its recall (79.92%) and specificity (94.54%) fell between those of SVM and RF, offering a compromise between sensitivity and false-alarm control.
Selecting a single “best” classifier from Table 3 requires first specifying the metric that matters most for the intended application, since no model dominates on every metric simultaneously. If the goal is to maximize discrimination ability independent of operating threshold, RF is preferable, achieving the highest AUC (0.953), the highest specificity (96.64%), and the lowest false-positive rate (3.36%). If the goal is to maximize overall correctness across both classes at a fixed threshold, BT is preferable, achieving the highest accuracy (87.21%) and the highest macro F1-score (87.15%). If the goal is to minimize missed detections and maximize the probability that an actual drone is correctly flagged, regardless of the resulting false-alarm cost, SVM is preferable, achieving the highest recall (90.38%), at the cost of the highest false-positive rate (16.81%) among the three models.
Metric | SVM | RF | BT |
|---|---|---|---|
Accuracy | 86.79% | 87.00% | 87.21% |
Drone Recall (TPR) | 90.38% | 77.41% | 79.92% |
Drone Precision | 84.375% | 95.85% | 93.63% |
Specificity | 83.19% | 96.64% | 94.54% |
False Positive Rate | 16.81% | 3.36% | 5.46% |
Macro Precision | 86.98% | 88.42% | 88.02% |
Macro F1-score | 86.77% | 86.89% | 87.15% |
AUC | 0.937 | 0.953 | 0.949 |
In a surveillance or security context, a missed detection and a false alarm typically carry asymmetric operational costs: a missed detection can result in an undetected incursion, whereas a false alarm typically results in an unnecessary but recoverable operator response. Under this asymmetric-cost assumption, which we explicitly adopt as our evaluation priority rather than treating as self-evident, recall is the metric that should be weighted most heavily, which favors SVM. We emphasize that this preference is conditional on the stated priority, not a claim that SVM dominates RF or BT on classification performance in general; Table 4 shows the opposite is true on AUC, specificity, and false-positive rate. A system operator whose priority is minimizing false alarms for example, in a setting where operator response is costly, or false alarms erode trust in the system, would instead be better served by RF.
Study | Sensing / Approach | Reported Outcome | Relation to This Work |
|---|---|---|---|
[20] | SVM, LSTM, and SqueezeNet | 97% with SVM, 98% with SqueezeNet, and 99.3% with LSTM, with LSTM achieving the best performance in ideal condition. | Under ideal conditions, the model attained high classification accuracy for synthetic radar targets |
[9] | Surveillance radar; the RF classification of birds vs. drones using motion-based features | Demonstrated that motion-characteristic features can separate birds from drones using a RF model on real surveillance-radar data | Supports this study's use of the RF as a strong baseline; unlike [9], the present work adds a jointly modeled multi-source clutter environment (ground, weather, and urban) rather than a single bird-vs.-drone motion contrast. |
[21] | mmWave radar; deep learning classification using radar cross section (RCS) signatures | Proposed deep learning-based drone classification using radar cross section signatures at millimeter-wave frequencies | Directly relevant since RCS is a primary feature in this study's framework; [21] uses RCS with a deep model at mmWave, while this work uses RCS (among other features) with classical ML (SVM/RF/BT) at a lower 9.4 GHz band, under explicit multi-source clutter |
This work | 9.4 GHz phased-array radar; SVM/RF/BT with ReliefF feature selection and BO / grid search | SVM achieved the highest recall (90.4%), RF achieved the highest AUC (0.953) and lowest false-positive rate (3.4%), while BT achieved the highest accuracy (87.2%) and macro F1-score (87.15%). | Provides a systematic, matched ideal-vs.-cluttered comparison and quantifies the accuracy recovered specifically by feature selection and hyperparameter tuning |
7. Conclusion
This study has demonstrated the efficacy of a machine learning-based framework for radar-based drone classification, employing synthetically generated target features, including RCS, Doppler shift, range, aspect angle, and micro-Doppler signatures under both idealized and realistic operating conditions. Under ideal, noise-free conditions, all three classifiers examined (SVM, RF, and BT) attained near-perfect classification accuracy of approximately 99%, thereby confirming the strong discriminative capacity of the selected radar feature set for distinguishing drone from non-drone targets. The introduction of realistic clutter and noise, representative of ground, urban, and weather environments, resulted in a substantial degradation of classifier performance, with accuracy falling to the range of 78.5–78.8%. This degradation underscores a critical limitation often overlooked in idealized radar classification studies: models trained and validated under clean conditions may not generalize reliably to operational environments. Through systematic feature optimization and hyperparameter tuning via BO and Grid Search, classification accuracy was recovered to 86.8–87.2% across all three models, demonstrating that principled optimization can substantially mitigate the adverse effects of environmental clutter. Notably, while the three optimized classifiers converged to comparable overall accuracy, their underlying error characteristics diverged in operationally significant ways. The SVM classifier achieved the highest recall (90.4%), rendering it the most sensitive to true drone targets, whereas the RF and BT classifiers achieved markedly lower false-positive rates (3.4% and 5.5%, respectively) at the expense of increased missed detections. The BT classifier achieved the highest macro F1-score (87.2%), reflecting the most balanced performance across both classes. This divergence illustrates a fundamental precision–recall trade-off inherent to the classification task, one that carries direct operational consequences. Given that surveillance and security applications generally impose a greater penalty on missed detections than on false alarms, the SVM classifier is identified as the preferred model for deployment in this context; nonetheless, the RF classifier may be more appropriate in scenarios where the suppression of false alarms is the primary operational concern. Collectively, these findings affirm that optimized machine learning techniques, when coupled with radar-derived target features, can meaningfully enhance the robustness of drone classification under cluttered and noise-corrupted conditions. They further suggest that classifier selection in practice should not rest solely on aggregate accuracy, but should be informed by the relative operational cost of false negatives versus false positives. Building upon the findings and acknowledged limitations of this study, future research will proceed along the following five directions:
• Validation using measured radar returns from real drone flights to verify whether the classification performance observed on the synthetic radar dataset generalizes to real-world operating conditions and practical clutter characteristics.
• Evaluation of deep learning architectures (e.g., convolutional neural networks or LSTM applied to raw micro-Doppler spectrograms) through direct comparison with the handcrafted-feature SVM/RF/BT classifiers used in this study to determine whether learned features offer additional robustness under clutter.
• Expansion of the target set to include fixed-wing unmanned aerial vehicles, drone swarms, and payload-carrying drones, since the current study considers a single drone/non-drone dichotomy.
• Robustness evaluation under adversarial or spoofed radar returns, relevant to security-critical counter-unmanned aircraft system deployments.
• Assessment of real-time computational feasibility for embedded or edge deployment, including inference latency and memory footprint of the optimized classifiers on representative hardware.
The data used to support the research findings are available from the corresponding author upon request.
The author declares no conflicts of interest.
During the preparation of this manuscript, the author used Grammarly, QuillBot, and ChatGPT as AI-assisted tools for grammar checking, language editing, paraphrasing, and improving the clarity and readability of the manuscript. All AI-assisted content was reviewed, verified, and revised by the authors, who remain fully responsible for the accuracy, originality, integrity, and final content of the manuscript. These tools were not used to fabricate or manipulate data, generate or alter research results, or create false references. No generative AI or AI-assisted tool is listed as an author of this work.
