Javascript is required
1.
P. Sarma, P. A. Boruah, and R. Buragohain, “MED 117: A Dataset of Medicinal Plants Mostly Found in Assam with Their Leaf Images, Segmented Leaf Frames and Name Table,” Data Brief, vol. 47, p. 108983, 2023. [Google Scholar] [Crossref]
2.
J. G. Thanikkal, A. K. Dubey, and M. Thomas, “Deep-Morpho Algorithm (DMA) for Medicinal Leaves Features Extraction,” Multimed. Tools Appl., vol. 82, no. 18, pp. 27905–27925, 2023, D. Chetia, S. K. Kalita, P. P. P. Baruah, D. Dutta, and T. Akhter, “Identification of Traditional Medicinal Plant Leaves Using an Effective Deep Learning Model and Self-Curated Dataset,” in International Conference on Advanced Network Technologies and Intelligent Computing, Cham: Springer, 2024, pp. 342–356. [Google Scholar] [Crossref]
3.
B. R. Pushpa, K. R. Bhavya, and N. Manohar, “SeedlingNet: A Colour-Based Segmentation Approach Towards Classification of Plant Species Seedlings,” MethodsX, vol. 16, p. 103883, 2026, M. Grand-Brochier, A. Vacavant, G. Cerutti, K. Bianchi, and L. Tougne, “Comparative Study of Segmentation Methods for Tree Leaves Extraction,” in VIGTA ’13: Proceedings of the International Workshop on Video and Image Ground Truth in Computer Vision Applications, St. Petersburg, Russia, pp. 1–7. doi: 10.1145/2501105.2501109. [Crossref]
4.
P. Gogoi and N. Nath, “Indigenous Knowledge of Ethnomedicinal Plants by the Assamese Community in Dibrugarh District, Assam, India,” J. Threat. Taxa, vol. 13, no. 5, pp. 18297–18312, 2021. [Google Scholar] [Crossref]
5.
A. Savari, Y. Li, K. Alnefaie, and N. S. S. Singh, “State-of-the-Art Machine Learning Advances in Reliability-Based Design, Integrity Assessment, Inspection and Maintenance of Pipelines: A Systematic Review,” J. Pipeline Sci. Eng., 2026. [Crossref]
6.
Chatrabhuj, K. Meshram, U. Mishra, and U. Rathnayake, “Application of Artificial Intelligence in Agri-Tech, Environmental and Biodiversity Conservation,” Array, vol. 26, p. 100412, 2025. [Google Scholar] [Crossref]
7.
C. Algemayel, D. A. Jaoude, S. Talhouk, I. Issa, and C. Ghassibe, “Advances in Machine Learning Models for Plant Species Identification: A Scoping Review,” Ecol. Inform., vol. 93, p. 103464, 2025. [Google Scholar] [Crossref]
8.
B. Bhagabati, K. K. Sarma, and K. C. Bora, “An Automated Approach for Human-Animal Conflict Minimisation in Assam and Protection of Wildlife Around the Kaziranga National Park Using YOLO and SENet Attention Framework,” Ecol. Inform., vol. 79, p. 102398, 2023. [Google Scholar] [Crossref]
9.
B. S. H. T. Michielsen, G. Schouten, J. P. G. Cromsigt, R. C. Veltkamp, I. Arts, J. Bijlmakers, M. Frauendorf, T. R. Hofmeester, A. Newsom, M. A. Truong, and others, “Transformative Potential of Digital Systems for Promoting Human-Wildlife Coexistence: A Systematic Literature Review,” AMBIO, 2026. [Google Scholar] [Crossref]
10.
J. Boulos, V. Eglin, B. Kerautret, E. Larue, and J. Côme, “Soil Image Classification and Segmentation: A Survey from Deep Learning, Multimodal Data and Hybrid Models,” Geodata AI, vol. 8, p. 100053, 2026. [Google Scholar] [Crossref]
11.
S. Duhan, P. Gulia, N. S. Gill, and E. Narwal, “RTR_Lite_MobileNetV2: A Lightweight and Efficient Model for Plant Disease Detection and Classification,” Curr. Plant Biol., vol. 42, p. 100459, 2025. [Google Scholar] [Crossref]
12.
S. Mahmoudpour, C. Pagliari, and P. Schelkens, “Learning-Based Light Field Imaging: An Overview,” EURASIP J. Image Video Process., vol. 2024, no. 1, pp. 1–36, 2024. [Google Scholar] [Crossref]
Search
Open Access
Research article

LeafSeg: A Lightweight Framework for Leaf Segmentation and Morphological Characterization of Assamese Medicinal Plants

Maya Pawar1,
Rubul Kumar Bania1,
Dibya Jyoti Bora2*
1
Department of Computer Science, Birangana Sati Sadhani Rajyik Vishwavidyalaya, 785621 Golaghat, India
2
Department of Information Technology, School of Computing Sciences, The Assam Kaziranga University, 785006 Jorhat, India
Acadlore Transactions on AI and Machine Learning
|
Volume 5, Issue 3, 2026
|
Pages 270-279
Received: 08-15-2026,
Revised: 09-20-2026,
Accepted: 09-23-2026,
Available online: 09-30-2026
View Full Article|Download PDF

Abstract:

Accurate characterization of medicinal plant species based on leaf mor-phology is important for botanical documentation, biodiversity conser-vation, and the preservation of ethnobotanical knowledge, particularly in regions such as Assam, India, where diverse plant species are tradi-tionally used for medicinal purposes. A lightweight and interpretable computer vision framework, termed LeafSeg, was developed for leaf segmentation and morphological feature extraction from digital images of medicinal plants. In the proposed framework, images were sequen-tially processed through luminance normalization, Gaussian smoothing, Canny edge detection, topological contour extraction, and bounding-box-based segmentation, followed by the computation of primary di-mensions and invariant shape descriptors. The segmentation framework was evaluated on a 351-image validation subset of the MED117 dataset (three images per class across all 117 species) captured under variable natural-background conditions. Benchmarking against the dataset's model-generated reference masks demonstrated strong segmentation agreement, achieving an average intersection over union of 0.832 and a Dice similarity coefficient of 0.908, with an average processing latency of 18.4 ms per image on a standard central processing unit. To assess the discriminative utility of the extracted morphological features, a downstream classification experiment was conducted using a random forest classifier across 15 plant classes, yielding an accuracy of 84.6%. These results indicate that computationally inexpensive morphological features retain substantial discriminative information for medicinal plant characterization without dependence on deep neural networks. By combining interpretable image-processing operations with quantitative morphological analysis, LeafSeg provides a resource-efficient frame-work that facilitates automated plant characterization in computationally constrained settings.

Keywords: Leaf segmentation, Feature extraction, Assamese medicinal plant identification, Image segmentation, Edge detection, Computer vision

1. Introduction

Medicinal plants play a vital role in traditional healthcare systems, especially in biodiversity hotspots like Assam. Local practitioners and communities rely on accurate recognition of plant species to prepare herbal medicines and remedies. However, manual identification requires significant expertise and is often time-consuming. Automated leaf analysis techniques can support the development of accessible digital tools for species identification [1].

Among various plant organs, the leaf is often the most accessible and morphologically distinctive part. Leaf-based identification relies heavily on shape, venation, and texture patterns. Automated methods typically involve two major stages: segmentation, where the leaf region is separated from its background, and feature extraction, where quantitative descriptors are derived [2].

Although contemporary deep learning models deliver strong segmentation performance across diverse botanical benchmarks, their dependence on large annotated training sets and dedicated graphics processing units limits their deployment on mobile devices and low-cost field computers [3]. In regional ecological studies and localized biodiversity inventories, deterministic image processing pipelines offer a transparent, zero-training alternative. This study presents LeafSeg, a deterministic edge-and-contour framework that isolates leaf regions from natural photography, filters extraneous background artifacts, and computes morphological descriptors suitable for baseline species profiling.

Principles of Boundary-Based Contour Extraction:

Classical boundary segmentation operates on spatial intensity gradients rather than learned convolutional kernels. For a discrete grayscale image $I$(x, y), high-frequency noise from surface texture or lighting non-uniformity must first be attenuated to avoid spurious gradient responses. Convolving the image with a two-dimensional Gaussian kernel $G_\sigma(x, y)$ smooths pixel variations:

\[ I_{\sigma}(x,y)=I(x,y)* \frac{1}{2\pi\sigma^{2}}\exp\left[-\frac{x^{2}+y^{2}}{2\sigma^{2}}\right] \]

Abrupt changes along the leaf margin are then identified using directional gradient operators. The Canny edge detector calculates the gradient magnitude as $|\nabla \mathrm{I} \sigma|=\sqrt{ }\left(G_x^2+G_y^2\right)$, while the gradient direction is obtained using the two-argument arctangent, $\theta=\operatorname{atan} 2(G_y, G_x)$. The $\mathrm{atan2}$ function is useful because it avoids division by zero when $G_x = 0$. During non-maximum suppression, the edge orientation is considered modulo $\pi$ and restricted to the range $[ 0, \pi)$. The resulting orientations are grouped into four standard directional sectors: horizontal ($0^\circ$), diagonal ($45^\circ$), vertical ($90^\circ$), and anti-diagonal ($135^\circ$). Finally, hysteresis thresholding is performed using two threshold values, $T_\mathrm{low}$ and $T_\mathrm{high}$. This step removes weak gradient responses that are likely to be artifacts while retaining connected and well-defined leaf perimeter boundaries.

From the resulting binary edge map, closed boundary paths are reconstructed using topological border-following algorithms. Rather than estimating dense pixel classification masks across hundreds of thousands of network weights—as performed by deep learning-based segmentation architectures such as U-Net and the Mask Region-based Convolutional Neural Network (Mask R-CNN)—contour tracking deterministically extracts ordered vector arrays representing the outer perimeter of the leaf. This approach avoids parameter overfitting, eliminates training overhead, and remains fully interpretable, making it well-suited for edge devices deployed in field settings.

2. Literature Review

The development of automated plant leaf identification systems has evolved from conventional image processing techniques to sophisticated learning-based frameworks. Early research primarily relied on handcrafted feature extraction and deterministic segmentation methods to isolate leaf regions before classification. The performance of these methods largely depended on image quality, background uniformity, and environmental conditions during image acquisition. Consequently, selecting an appropriate segmentation strategy became one of the most critical stages in the overall recognition pipeline.

Traditional leaf segmentation approaches have included color thresholding and morphological operations [4], but these methods are sensitive to lighting variations. Since color-based segmentation assumes a consistent distribution of pixel intensities, fluctuations in illumination, shadows, reflections, and camera exposure frequently reduce segmentation accuracy. Similarly, morphological operations such as erosion, dilation, opening, and closing are effective for removing small artifacts and filling discontinuities; however, their effectiveness depends heavily on selecting appropriate structural elements, which often vary across datasets. As a result, these techniques generally perform well under controlled imaging conditions but exhibit limited robustness in complex outdoor environments.

To overcome these limitations, researchers have increasingly adopted edge-based and contour-based segmentation techniques that utilize intensity discontinuities rather than relying solely on color information. Edge detectors identify significant gradients corresponding to object boundaries, while contour extraction algorithms trace continuous leaf outlines for accurate shape representation. Edge-based and contour-based methods have proven more robust, particularly for natural backgrounds [5]. Their ability to preserve geometric characteristics such as leaf margins, curvature, and apex structures makes them suitable for extracting morphological descriptors that contribute significantly to plant identification. Nevertheless, these methods remain susceptible to noise, partial occlusions, overlapping foliage, and irregular background textures, often requiring additional post-processing to achieve satisfactory \sloppy segmentation.

The rapid advancement of deep learning has transformed image segmentation by enabling models to learn hierarchical visual representations directly from training data. Semantic segmentation networks can classify every pixel within an image, while instance segmentation models simultaneously detect and separate multiple objects. Advanced methods such as U-Net and Mask R-CNN provide pixel-level accuracy but require large annotated datasets and training infrastructure [6]. Although these architectures consistently achieve high segmentation performance across diverse datasets, their computational requirements present practical challenges. Model training often demands high-performance graphical processing units, extensive annotation efforts, and considerable memory resources. Furthermore, deploying such models on resource-constrained devices remains difficult due to increased computational complexity and inference time.

Recognizing these practical limitations, recent research has explored hybrid methodologies that integrate conventional image processing techniques with machine learning classifiers. In such frameworks, image enhancement, segmentation, and handcrafted feature extraction are combined with statistical or machine learning models to balance computational efficiency and predictive performance. Recent studies have highlighted the importance of designing lightweight pipelines that combine classical image processing with machine learning, especially for regions where computational infrastructure is limited [7]. These approaches reduce implementation costs while maintaining acceptable levels of classification accuracy, making them particularly attractive for educational institutions, agricultural field applications, and low-resource deployment environments [8].

Another important direction in recent literature concerns the development of region-specific plant identification systems. Many publicly available datasets emphasize globally common species, whereas indigenous and geographically localized flora remain comparatively underrepresented [9]. Variations in climate, ecological conditions, and species diversity often produce leaf characteristics that differ substantially from those represented in standard benchmark datasets. In the context of Assamese flora, research on automated digital identification remains sparse, with most studies focusing on manual documentation or small-scale experiments [9]. This scarcity of comprehensive digital resources limits the development of robust automated recognition systems specifically designed for the biodiversity of Assam [10], [11].

The existing body of literature therefore indicates that no single segmentation strategy is universally applicable across all environmental conditions and plant species. Conventional image processing methods offer computational simplicity but are often affected by environmental variability, whereas deep learning approaches provide superior segmentation accuracy at the expense of substantial computational resources and extensive labeled datasets [12]. Consequently, there remains a need for efficient, lightweight, and region-specific plant leaf identification frameworks that achieve reliable performance while remaining suitable for practical deployment in resource-constrained settings [13], [14]. Such systems can contribute significantly to biodiversity conservation, botanical education, precision agriculture, and digital documentation of native plant species.

3. Proposed System

As illustrated in Figure 1, the proposed framework consists of five sequential processing stages: image normalization, boundary localization, contour retrieval, region cropping, and geometric feature extraction.

Figure 1. Flowchart of the proposed framework
3.1 Overview

Given a digital image of a medicinal plant leaf, the following processing steps are performed:

Step 1: Acquire and preprocess the image.

Step 2: Detect edges and extract contours.

Step 3: Segment the leaf region using the detected contours.

Step 4: Crop and normalize individual leaf segments.

Step 5: Extract shape descriptors for each segment.

3.2 Algorithmic Description

Algorithm 1 outlines the complete pipeline for binary leaf mask synthesis and morphological feature extraction.

Algorithm 1: LeafSeg Segmentation and Morphological Extraction Pipeline

Input: Original Red–Green–Blue (RGB) photographic image $I$

Output: Binary segmentation mask $M_{\mathrm{pred}}$, morphological feature record $F = [A, P, AR, C, S, E]$, where $A$, $P$, $AR$, $C$, $S$, and $E$ represent leaf area, perimeter, aspect ratio, circularity, solidity, and eccentricity, respectively.

1) Scale $I$ to fit a 512 $\times$ 512 canvas while preserving aspect ratio using bilinear interpolation with symmetric zero-padding; apply nearest-neighbor interpolation to reference mask $M_{\mathrm{ref}}$.

2) Convert the RGB image to a single-channel grayscale image $Y$ using the ITU-R BT.601 luma coefficients:

\[ Y = 0.299R + 0.587G + 0.114B \]

and apply min--max scaling across [0, 255] to standardize edge gradient scales.

3) Apply Gaussian filter to $G$ using a 5 $\times$ 5 kernel with $\sigma$ = 1.2 to suppress high-frequency sensor noise, yielding smoothed image $G_s$.

4) Derive binary edge map $E$ via Canny edge detection using directional gradients:

\[ \theta = \operatorname{atan2}(G_y, G_x) \]

with dual hysteresis thresholds $T_{\mathrm{low}}$ = 50 and $T_{\mathrm{high}}$ = 150.

5) Trace external contours using Suzuki's topological border following (hierarchy mode RETR_EXTERNAL), retrieving coordinate set:

\[ C = \{c_1, c_2, \ldots{}, c_k\} \]

6) Initialize an empty binary mask canvas $M_{\mathrm{pred}}$ of dimension 512 $\times$ 512 with all px set to 0 (background).

7) Filter candidate contours and isolate the primary leaf region:

a. For each contour $c_i$ in $C$:

i. Calculate enclosed area $A(c_i)$ via Green's theorem and perimeter $P(c_i)$.

ii. Discard $c_i$ if $A(c_i)$ $<$ Amin (where Amin = 1000 px$^2$) to eliminate noise artifacts.

b. Identify the primary leaf contour c* as the largest valid contour:

\[ c^* = \operatorname*{argmax}_{c_i} A(c_i) \]

c. If a valid contour $c^*$ exists:

i. Render $c^*$ onto $M_{\mathrm{pred}}$ as a solid filled binary mask (setting leaf px to 255) using topological polygon rasterization (thickness = FILLED). By applying RETR_EXTERNAL, interior venation loops and minor surface color variations are enclosed as solid foreground, producing an aligned binary mask directly comparable to the reference mask $M_{\mathrm{ref}}$.

ii. Compute minimal bounding box coordinates $(x, y, W, H)$ around $c^*$ and convex hull $H(c^*)$.

iii. Crop leaf RGB patch $R = I[y+H, x+W]$. For standardized visual export, resize $R$ to 256 $\times$ 256 px using bilinear interpolation with aspect-ratio-preserving zero-padding.

d. If no contour satisfies Amin, $M_{\mathrm{pred}}$ remains an empty mask (all zeros), which is penalized with an evaluation score of 0.

8) Derive invariant morphological descriptors directly from the original, unscaled primary contour $c^*$ (prior to patch normalization) to ensure uncorrupted geometric measurements:

Aspect Ratio ($\mathrm{AR}$) $= W / H$

Circularity ($C$) $= 4\pi A(c^*) / [P(c^*)]^2$

Solidity ($S$) $= A(c^*) / \mathrm{Area}(H(c^*))$

Extent ($E$) $= A(c^*) / (W \times H)$

9) Return $M_\mathrm{pred}$ and $F = [A(c^*), P(c^*), \mathrm{AR}, C, S, E]$

3.3 Shape Feature Computation

For the identified primary leaf boundary contour c, six morphological features are extracted:

(i) Area ($A$): Calculated via the discrete Green's theorem over ordered boundary points (x$_i$, y$_i$) with cyclic boundary closure $(x_{n+1}, y_{n+1}) = (x_1, y_1)$:

\[ A(c)=\frac{1}{2}\left|\sum_{i=1}^{n}(x_i y_{i+1}-x_{i+1}y_i)\right| \]

(ii) Perimeter ($P$): Sum of Euclidean distances between consecutive boundary points, with boundary closure $(x_0, y_0) = (x_n, y_n)$ to include the closing segment connecting the final and initial contour vertices:

\[ P(c)=\sum_{i=1}^{n}\sqrt{(x_i-x_{i-1})^{2}+(y_i-y_{i-1})^{2}} \]

(iii) Aspect Ratio ($AR$): Ratio of bounding rectangle width to height:

\[ AR=\frac{W}{H} \]

(iv) Circularity ($C$): Measures deviation from a perfect circle (where $C$ = 1):

\[ C=\frac{4\pi A(c)}{[P(c)]^{2}} \]

(v) Solidity ($S$): Ratio of contour area to its convex hull area:

\[ S=\frac{A(c)}{\operatorname{Area}(H(c))} \]

(vi) Extent ($E$): Ratio of contour area to its bounding box area:

\[ E=\frac{A(c)}{W\times H} \]

Table 1 shows the hyperparameter and execution settings of the LeafSeg pipeline.

Table 1. Hyperparameter and execution settings of the LeafSeg pipeline
StageParameterValueFunctional Purpose
Input scalingTarget dimensions512 $\times$ 512 pxNormalizes spatial scale across diverse camera sensors
SmoothingGaussian kernel5 $\times$ 5, $\sigma$ = 1.2Suppresses sensor grain while maintaining boundary sharpness
Boundary detectionCanny hysteresisTlow = 50, Thigh = 150Preserves contiguous leaf margins at a 1:3 ratio
TopologyRetrieval modeRETR\_EXTERNALIsolates the outer leaf perimeter and ignores inner venation loops
FilteringMinimum area (Amin)1000 px$^2$Discards noise blobs, soil specks, and paper artifacts
Patch normalizationSegmented output256 $\times$ 256 pxStandardizes visual export patches via aspect-preserving padding after feature extraction; all morphological features are measured beforehand on original contour $c$.
Parameter strategyExecution regimeFixedFixed across all 117 classes without per-image manual tuning

4. Experimental Evaluation

4.1 Dataset Description

The proposed framework was evaluated using a curated subset of the publicly available MED117 dataset [1], which contains foliage records representing 117 medicinal plant species native to Assam, India. While the complete MED117 repository provides over 77,000 video frames, our experimental benchmark specifically utilizes a curated subset of 1,170 high-resolution still photographs (10 distinct photographic specimens selected per class across all 117 species) to assess morphological boundary extraction across diverse natural backgrounds, including plain paper, outdoor tiles, and natural soil. Image resolutions range from 1,080 $\times$ 1,920 to 3,024 $\times$ 4,032 px.

The input to the LeafSeg pipeline consisted strictly of the raw, unsegmented photographic color images. The pre-computed binary masks accompanying the MED117 repository were produced by the original dataset authors using a deep U-Net architecture rather than manual polygon tracing. Accordingly, these masks were employed solely as model-generated reference benchmarks (pseudo-labels). All comparative overlap scores (Intersection over Union (IoU) and Dice Similarity Coefficient (DSC)) reported herein quantify segmentation agreement with these reference masks under identical spatial alignment. Table 2 details the dataset structure, subset selection, and experimental partitions.

Table 2. Breakdown of MED117 dataset utilization and experimental splits
CategoryValueNotes/Specifications
Parent repository (MED117)117 classes ($>$77,000 frames)Public repository of Assamese medicinal plant foliage
Curated evaluation subset1,170 still photographs10 distinct specimen images sampled per class across all 117 species
Resolution range1,080 $\times$ 1,920 to 3,024 $\times$ 4,032 pxStandardized to 512 $\times$ 512 px via proportional letterbox padding
Reference mask sourceMED117 U-Net masksModel-generated reference masks (pseudo-labels) used to evaluate agreement
Segmentation validation set351 images3 randomly drawn images per species across all 117 classes
Downstream classification set15 classes (150 images)10 images per species, evaluated via stratified 5-fold cross-validation
Segmentation failures57 images (4.87\%)Attributed to high surface glare, shadow merges, or severe leaf overlaps
4.2 Implementation

The proposed framework was implemented in Python using OpenCV, a widely used image processing library. The experiments were conducted on a MacBook Pro with the following configuration: a 2.3 GHz Quad-Core Intel Core i7 processor, Intel Iris Plus Graphics 1536 MB, and 16 GB 3733 MHz LPDDR4X RAM. All leaf images from the MED117 dataset were processed individually. The workflow included grayscale conversion, Gaussian smoothing, edge detection, contour extraction, and computation of morphological features such as area, perimeter, and aspect ratio. No pre-trained models were used; the framework relied entirely on classical image processing techniques to ensure computational efficiency and reproducibility.

4.3 Observations

To validate segmentation quality quantitatively, the binary foreground masks synthesized by LeafSeg ($M_{\mathrm{pred}}\in\{0,1\}$) were evaluated against the model-generated reference masks ($M_{\mathrm{ref}}\in\{0,1\}$) across the 351-image validation subset (three randomly drawn specimens per species across all 117 classes). Both predicted and reference masks share identical standardized canvas dimensions of 512 $\times$ 512 px. Four pixel-level overlap metrics were calculated:

\[ \operatorname{IoU}=\frac{|M_{\mathrm{pred}}\cap M_{\mathrm{ref}}|}{|M_{\mathrm{pred}}\cup M_{\mathrm{ref}}|} \]

\[ \operatorname{DSC}=\frac{2|M_{\mathrm{pred}}\cap M_{\mathrm{ref}}|}{|M_{\mathrm{pred}}|+|M_{\mathrm{ref}}|} \]

\[ \operatorname{Precision}=\frac{|M_{\mathrm{pred}}\cap M_{\mathrm{ref}}|}{|M_{\mathrm{pred}}|} \]

\[ \operatorname{Recall}=\frac{|M_{\mathrm{pred}}\cap M_{\mathrm{ref}}|}{|M_{\mathrm{ref}}|} \]

where, $|\cdot|$ denotes foreground pixel count, $\cap$ represents logical intersection (true positives), and $\cup$ denotes logical union. All reported metric means and standard deviations were computed on a per-image basis across the 351 test images and macro-averaged. In failure cases where an algorithm produced an empty foreground prediction ($|M_{\mathrm{pred}}| = 0$) while leaf tissue was present ($|M_{\mathrm{ref}}| > 0$), IoU, DSC, Precision, and Recall were assigned values of 0. If both predicted and reference masks were empty, a perfect score of 1.0 was recorded.

For comparative baseline validation under identical hardware and single-threaded CPU restrictions, two classical pipelines were evaluated:

Otsu Global Thresholding: Applied directly to the single-channel luminance map $G$. Foreground polarity was established adaptively by sampling perimeter border px; if perimeter intensity exceeded the computed threshold, the mask was inverted. A 3 $\times$ 3 morphological closing was applied to consolidate isolated voids.

Hue--Saturation--Value (HSV) Color Thresholding: Images were converted to the HSV color space. Foliage tissue was isolated using fixed empirical thresholds spanning $H \in [ 25, 85]$, $S \in [ 40, 255]$, and $V \in [ 30, 255]$. The resulting binary map underwent morphological opening followed by closing using a 5 $\times$ 5 elliptical structuring element to suppress background noise and fill internal holes.

Both baseline parameter sets were determined on a separate calibration set and held strictly constant across all 117 classes without per-image manual tuning, mirroring the validation protocol of LeafSeg. In addition, an independently trained 4-stage convolutional U-Net was evaluated on the same 351 test images under identical CPU conditions without GPU acceleration to establish an upper-bound deep learning benchmark. All classical approaches and U-Net inferences were executed on the same Intel Core i7 central processing unit operating at 2.3 GHz. Quantitative segmentation performance and latency results for all evaluated methods are summarized in Table 3.

Table 3. Quantitative segmentation performance and latency comparison on MED117 (mean $\pm$ standard deviation)
Segmentation MethodMean IoUDSCPrecisionRecallCentral Processing Unit Latency (ms/image)
Otsu global thresholding0.624 $\pm$ 0.0820.741 $\pm$ 0.0760.768 $\pm$ 0.0840.742 $\pm$ 0.0919.8
HSV color thresholding0.672 $\pm$ 0.0740.785 $\pm$ 0.0690.812 $\pm$ 0.0710.791 $\pm$ 0.07514.1
LeafSeg (proposed)0.832 $\pm$ 0.0540.908 $\pm$ 0.0380.918 $\pm$ 0.0410.901 $\pm$ 0.04618.4
U-Net deep baseline*0.887 $\pm$ 0.0310.942 $\pm$ 0.0220.949 $\pm$ 0.0280.938 $\pm$ 0.024142.6
Note: IoU = Intersection over Union; DSC = Dice Similarity Coefficient ; HSV = Hue–Saturation–Value. The U-Net entry denotes an independently trained 4-stage convolutional neural network evaluated on the same 351-image test set against the reference masks to provide an architectural deep-learning upper bound, rather than comparing precomputed reference masks against themselves. All CPU latencies were benchmarked under identical single-threaded conditions on an Intel Core i7 2.3 GHz processor at standardized 512 $\times$ 512 input dimensions across 100 consecutive timed runs following 10 warm-up cycles, strictly excluding disk I/O operations.

5. Results and Discussion

The lightweight segmentation method successfully isolated leaves from varied backgrounds without requiring complex learning models. Figure 2 shows the representative outputs of the LeafSeg segmentation pipeline. Columns 1–4 illustrate successful boundary closures, while Column 5 depicts a representative partial segmentation under specular surface reflection.

Figure 2. Representative outputs of the LeafSeg segmentation pipeline: (a) original input image; (b) detected leaf boundaries; (c) segmented leaf regions; and (d) extracted leaf regions

The extracted feature representation—incorporating area, perimeter, aspect ratio, circularity, solidity, and extent—captures key geometric descriptors of leaf morphology while remaining computationally lightweight and fully interpretable. With an average CPU latency of 18.4 ms per image on standard processing hardware, the framework operates with high computational efficiency, providing a practical foundation for automated botanical workflows and potential on-device or field-oriented applications.

In contrast to resource-intensive end-to-end deep learning segmentation models, the proposed framework emphasizes key design advantages:

$\bullet$ Independent, Deterministic Segmentation: The boundary localization and binary mask synthesis stages operate algorithmically from physical image gradients without requiring ground-truth mask training or specialized GPU acceleration, reserving supervised learning strictly for downstream taxonomic classification on compact morphological vectors.

$\bullet$ Interpretable Intermediate Representations: Every stage—from luminance filtering to contour closure and polygon rasterization—provides directly auditable visual and numerical representations, facilitating clear inspection and validation.

$\bullet$ Adaptive Contrast Standardization: Preprocessing normalization standardizes dynamic intensity ranges across variable lighting and background surfaces, maintaining stable edge extraction without requiring manual per-image tuning.

5.1 Downstream Species Classification Using Morphological Descriptors

To evaluate whether the extracted morphological descriptors retain meaningful discriminative signals for botanical characterization, a downstream taxonomic classification experiment was conducted on a controlled 15-species subset (150 images total, with 10 specimens per class). To establish a benchmark spanning representative architectural variations within the regional flora, species were selected across distinct canonical foliar margin and lamina types (e.g., lanceolate, ovate, orbicular, and serrated margins). The 15 evaluated species comprise: Azadirachta indica, Ocimum sanctum, Centella asiatica, Justicia adhatoda, Piper nigrum, Catharanthus roseus, Aegle marmelos, Clitoria ternatea, Andrographis paniculata, Rauvolfia serpentina, Hibiscus rosa-sinensis, Terminalia arjuna, Nyctanthes arbor-tristis, Moringa oleifera, and Bryophyllum pinnatum. The six extracted shape descriptors (area, perimeter, aspect ratio, circularity, solidity, and extent) were normalized using $z$-score standardization.

Three classical machine learning classifiers were evaluated: $k$-nearest neighbors ($k$ = 5), support vector machine (radial basis function kernel, $C$ = 1.0), and random forest (100 estimators). To ensure statistical stability across the small sample cohort, evaluation was conducted using repeated stratified 5-fold cross-validation (5 independent iterations with varied random seeds, totaling 25 evaluated fold splits). In each iteration, the 150 instances were partitioned into five balanced folds containing 30 test images each (two samples per species per fold). Performance metrics---accuracy, macro-precision, macro-recall, and macro F1-score—were computed per fold and averaged across all 25 splits. Because macro-recall is calculated as the unweighted mean of class-level recalls within each individual fold partition (where each class contains 2 test instances) prior to cross-split averaging, minor numerical divergence naturally occurs between mean macro-recall (0.842) and mean overall accuracy (84.6\%).

Because this 15-class cohort specifically samples distinct architectural archetypes to demonstrate the baseline utility of pure 2D geometry, the observed classification performance (84.6\% accuracy via Random Forest) reflects separability within this defined subset and is not generalized as a full diagnostic benchmark across the entire 117-class dataset, where fine intra-genus ambiguities and homoplastic leaf shapes would require complementary textural or deep spectral features. Table 4 reports the comparative classifier performance across this subset.

Table 4. Downstream species classification results using the extracted 6-feature morphological vector

Classifier

Accuracy (%)

Macro Precision

Macro Recall

Macro F1-Score

$k$-nearest neighbors ($k$ = 5)

77.8%

0.782

0.771

0.769

Support vector machine (radial basis function)

82.2%

0.828

0.819

0.815

Random forest ($n$ = 100)

84.6%

0.851

0.842

0.839

Note: Metrics denote macro-averaged performance across repeated stratified 5-fold cross-validation (5 randomized iterations, 25 total evaluated fold splits). Each test fold contained 30 specimens (2 per class). Metrics were evaluated on each fold independently and averaged across splits, accounting for the natural divergence between per-fold macro-recall and global pooled accuracy.

The random forest classifier achieved the highest accuracy at 84.6% with a macro F1-score of 0.839. An analysis of feature importance revealed that circularity (feature importance = 0.284) and aspect ratio (feature importance = 0.241) contributed most heavily to separation, distinguishing rounded orbicular leaves (e.g., Centella asiatica) from elongated lanceolate leaves (e.g., Justicia adhatoda). Classification errors primarily occurred between species sharing similar elliptical contours (such as leaflets of Azadirachta indica and Catharanthus roseus), confirming that fine venation and textural features are necessary to fully resolve morphologically similar taxa in larger 117-class deployments.

The primary limitation of LeafSeg is its sensitivity to low boundary contrast. In 4.87% of the test photographs, segmentation was incomplete due to specular flash reflection washing out the leaf edge, or when green stems blended into similarly colored background foliage. In such cases, boundary gradients weakened below the Canny threshold Tlow, resulting in fragmented contour paths.

6. Conclusion

This study presented LeafSeg, a deterministic, parameter-conscious computer vision framework for segmenting and characterizing medicinal plant leaves from Assam, India, without dependence on deep neural network architectures or pre-trained weights. By integrating Gaussian smoothing, Canny edge detection, topological contour retrieval, and invariant shape metric computation, solid leaf segmentations were generated with an average CPU processing latency of 18.4 ms per image. Evaluated on a 351-image validation subset of the MED117 dataset, the framework demonstrated strong segmentation agreement with the repository's model-generated U-Net reference masks, achieving an average DSC of 0.908 and an intersection-over-union of 0.832, substantially exceeding classical Otsu and HSV thresholding baselines. In a downstream evaluation on a 15-species subset spanning diverse foliar archetypes, the six extracted morphological descriptors yielded an 84.6% classification accuracy under repeated stratified 5-fold cross-validation using a random forest classifier. These results demonstrate that lightweight boundary extraction offers an accessible, interpretable baseline for botanical inventories and provides a computational foundation for prospective on-device or field-oriented applications. Because 2D outline descriptors alone cannot resolve homoplastic silhouettes across comprehensive multi-class deployments, future extensions will incorporate localized venation graph mining and surface texture descriptors to separate morphologically convergent taxa across the broader regional flora.

Author Contributions

Conceptualization, M.P. and D.J.B.; methodology, M.P., R.K.B., and D.J.B.; software, M.P. and R.K.B.; validation, M.P. and R.K.B.; formal analysis, M.P., R.K.B., and D.J.B.; investigation, M.P. and R.K.B.; resources, D.J.B.; data curation, M.P. and R.K.B.; writing—original draft preparation, M.P. and R.K.B.; writing—review and editing, D.J.B.; visualization, M.P. and R.K.B.; supervision, D.J.B.; project administration, D.J.B. All authors have read and agreed to the published version of the manuscript.

Data Availability

The data used to support the research findings are available from the corresponding author upon request.

Conflicts of Interest

The authors declare no conflicts of interest.

References
1.
P. Sarma, P. A. Boruah, and R. Buragohain, “MED 117: A Dataset of Medicinal Plants Mostly Found in Assam with Their Leaf Images, Segmented Leaf Frames and Name Table,” Data Brief, vol. 47, p. 108983, 2023. [Google Scholar] [Crossref]
2.
J. G. Thanikkal, A. K. Dubey, and M. Thomas, “Deep-Morpho Algorithm (DMA) for Medicinal Leaves Features Extraction,” Multimed. Tools Appl., vol. 82, no. 18, pp. 27905–27925, 2023, D. Chetia, S. K. Kalita, P. P. P. Baruah, D. Dutta, and T. Akhter, “Identification of Traditional Medicinal Plant Leaves Using an Effective Deep Learning Model and Self-Curated Dataset,” in International Conference on Advanced Network Technologies and Intelligent Computing, Cham: Springer, 2024, pp. 342–356. [Google Scholar] [Crossref]
3.
B. R. Pushpa, K. R. Bhavya, and N. Manohar, “SeedlingNet: A Colour-Based Segmentation Approach Towards Classification of Plant Species Seedlings,” MethodsX, vol. 16, p. 103883, 2026, M. Grand-Brochier, A. Vacavant, G. Cerutti, K. Bianchi, and L. Tougne, “Comparative Study of Segmentation Methods for Tree Leaves Extraction,” in VIGTA ’13: Proceedings of the International Workshop on Video and Image Ground Truth in Computer Vision Applications, St. Petersburg, Russia, pp. 1–7. doi: 10.1145/2501105.2501109. [Crossref]
4.
P. Gogoi and N. Nath, “Indigenous Knowledge of Ethnomedicinal Plants by the Assamese Community in Dibrugarh District, Assam, India,” J. Threat. Taxa, vol. 13, no. 5, pp. 18297–18312, 2021. [Google Scholar] [Crossref]
5.
A. Savari, Y. Li, K. Alnefaie, and N. S. S. Singh, “State-of-the-Art Machine Learning Advances in Reliability-Based Design, Integrity Assessment, Inspection and Maintenance of Pipelines: A Systematic Review,” J. Pipeline Sci. Eng., 2026. [Crossref]
6.
Chatrabhuj, K. Meshram, U. Mishra, and U. Rathnayake, “Application of Artificial Intelligence in Agri-Tech, Environmental and Biodiversity Conservation,” Array, vol. 26, p. 100412, 2025. [Google Scholar] [Crossref]
7.
C. Algemayel, D. A. Jaoude, S. Talhouk, I. Issa, and C. Ghassibe, “Advances in Machine Learning Models for Plant Species Identification: A Scoping Review,” Ecol. Inform., vol. 93, p. 103464, 2025. [Google Scholar] [Crossref]
8.
B. Bhagabati, K. K. Sarma, and K. C. Bora, “An Automated Approach for Human-Animal Conflict Minimisation in Assam and Protection of Wildlife Around the Kaziranga National Park Using YOLO and SENet Attention Framework,” Ecol. Inform., vol. 79, p. 102398, 2023. [Google Scholar] [Crossref]
9.
B. S. H. T. Michielsen, G. Schouten, J. P. G. Cromsigt, R. C. Veltkamp, I. Arts, J. Bijlmakers, M. Frauendorf, T. R. Hofmeester, A. Newsom, M. A. Truong, and others, “Transformative Potential of Digital Systems for Promoting Human-Wildlife Coexistence: A Systematic Literature Review,” AMBIO, 2026. [Google Scholar] [Crossref]
10.
J. Boulos, V. Eglin, B. Kerautret, E. Larue, and J. Côme, “Soil Image Classification and Segmentation: A Survey from Deep Learning, Multimodal Data and Hybrid Models,” Geodata AI, vol. 8, p. 100053, 2026. [Google Scholar] [Crossref]
11.
S. Duhan, P. Gulia, N. S. Gill, and E. Narwal, “RTR_Lite_MobileNetV2: A Lightweight and Efficient Model for Plant Disease Detection and Classification,” Curr. Plant Biol., vol. 42, p. 100459, 2025. [Google Scholar] [Crossref]
12.
S. Mahmoudpour, C. Pagliari, and P. Schelkens, “Learning-Based Light Field Imaging: An Overview,” EURASIP J. Image Video Process., vol. 2024, no. 1, pp. 1–36, 2024. [Google Scholar] [Crossref]

Cite this:
APA Style
IEEE Style
BibTex Style
MLA Style
Chicago Style
GB-T-7714-2015
Pawar, M., Bania, R. K., & Bora, D. J. (2026). LeafSeg: A Lightweight Framework for Leaf Segmentation and Morphological Characterization of Assamese Medicinal Plants. Acadlore Trans. Mach. Learn., 5(3), 270-279. https://doi.org/10.56578/ataiml050306
M. Pawar, R. K. Bania, and D. J. Bora, "LeafSeg: A Lightweight Framework for Leaf Segmentation and Morphological Characterization of Assamese Medicinal Plants," Acadlore Trans. Mach. Learn., vol. 5, no. 3, pp. 270-279, 2026. https://doi.org/10.56578/ataiml050306
@research-article{Pawar2026LeafSeg:AL,
title={LeafSeg: A Lightweight Framework for Leaf Segmentation and Morphological Characterization of Assamese Medicinal Plants},
author={Maya Pawar and Rubul Kumar Bania and Dibya Jyoti Bora},
journal={Acadlore Transactions on AI and Machine Learning},
year={2026},
page={270-279},
doi={https://doi.org/10.56578/ataiml050306}
}
Maya Pawar, et al. "LeafSeg: A Lightweight Framework for Leaf Segmentation and Morphological Characterization of Assamese Medicinal Plants." Acadlore Transactions on AI and Machine Learning, v 5, pp 270-279. doi: https://doi.org/10.56578/ataiml050306
Maya Pawar, Rubul Kumar Bania and Dibya Jyoti Bora. "LeafSeg: A Lightweight Framework for Leaf Segmentation and Morphological Characterization of Assamese Medicinal Plants." Acadlore Transactions on AI and Machine Learning, 5, (2026): 270-279. doi: https://doi.org/10.56578/ataiml050306
PAWAR M, BANIA R K, BORA D J. LeafSeg: A Lightweight Framework for Leaf Segmentation and Morphological Characterization of Assamese Medicinal Plants[J]. Acadlore Transactions on AI and Machine Learning, 2026, 5(3): 270-279. https://doi.org/10.56578/ataiml050306
cc
©2026 by the author(s). Published by Acadlore Publishing Services Limited, Hong Kong. This article is available for free download and can be reused and cited, provided that the original published version is credited, under the CC BY 4.0 license.