Javascript is required
Search
Volume 5, Issue 2, 2026

Abstract

Full Text|PDF|XML

Multitask learning (MTL) is a machine learning paradigm in which several related tasks are learned simultaneously to improve generalization performance. Kernel-based methods provide a mathematically rigorous and flexible framework for MTL, especially when training data are limited or uncertainty estimation is important. This study examined the research landscape of MTL using the Scopus database. The findings reveal that MTL has experienced remarkable growth in recent years, with over 93% of publications produced between 2016 and 2026, demonstrating its increasing relevance in modern scientific and technological research. The subject-area restriction further confirmed the highly interdisciplinary nature of the field, particularly across computer vision, medical imaging, predictive analytics, and artificial intelligence applications. Despite the rapid expansion of deep learning-based multitask approaches, the analysis identified only a very limited number of studies specifically focused on kernel-based MTL, indicating a significant research gap within the literature. This scarcity suggests that kernel-based methods remain largely underexplored despite their strong mathematical foundations, interpretability, and effectiveness in nonlinear modeling. The study therefore concludes that kernel-based MTL presents substantial opportunities for future theoretical development and practical applications, making it a promising direction for advancing MTL research.

Abstract

Full Text|PDF|XML
As multi-vendor Data Over Cable Service Interface Specification (DOCSIS)-based implementations have rapidly progressed from piloting to rollout in national broadband networks, differences in modulation, channelization, and error correction at the chipset level have elevated the importance of architectural standardization. This article examined both the technical and operational case for a behaviorally uniform DOCSIS chipset platform. Drawing on qualitative engineering analysis grounded in published DOCSIS specifications, cable plant operational practices, and established deployment frameworks, the article investigated chipset platforming strategies in multi-vendor and phased deployment environments, as well as centralized network monitoring architectures. Behaviorally consistent chipset platforms reduce cross-vendor interoperability failures, lower the engineering training burden, and improve predictability of network operations at scale, thereby reducing mean time to resolution (MTTR) through a standardized diagnostics and monitoring interface. The convergence of core DOCSIS functions onto a single chipset reduces configuration error rates, simplifies maintenance processes, and increases throughput stability. These benefits are complemented by the industry’s ongoing migration toward DOCSIS 4.0 and high-split spectrum architectures, which transform a forward-compatible chipset platform into the most operationally and economically rational path for operators seeking to accelerate nationally-scaled deployments while preserving service quality.

Abstract

Full Text|PDF|XML

Isolated sign language recognition (ISLR) is an important video information processing task that supports accessible communication for people with hearing impairments. Existing methods rely predominantly on red-green-blue (RGB) appearance information and make limited use of the geometric cues contained in video frames. In addition, unbalanced contributions from different representations may restrict the effectiveness and generalizability of multimodal information fusion. This study investigates whether complementary geometric representations derived from RGB videos can improve signer-independent ISLR without requiring additional sensing equipment. Depth maps were estimated from RGB frames using MiDaS-family DPT-Large model, and surface-normal maps were calculated from depth gradients. A shared Swin Transformer equipped with three lightweight adapters was used to encode the RGB, depth, and normal representations within a unified feature extraction framework. Their interactions were modelled using a dynamic cross-attention gated fusion (DCAGate) module, while entropy regularization was applied to prevent persistent dominance by a single representation. A class-embedding classification head with a margin loss was used to improve fine-grained discrimination. On the Chinese Sign Language 500 (CSL-500) dataset, the resulting Shared Swin with Gated Multimodal Fusion Network (SGNNet) achieved a Top-1 accuracy of 96.2% under a signer-independent evaluation protocol. Compared with the independent-branch baseline, in which separate backbones are used for the RGB, depth, and normal inputs, it increased accuracy by 3.8 percentage points, reduced graphics processing unit (GPU) memory consumption by 34%, and increased recognition-stage throughput by 25%. The gate weights remained relatively balanced across the evaluated representation combinations. These results indicate that complementary geometric representations and regulated cross-modal interaction can improve ISLR while limiting recognition-stage resource consumption. The proposed framework provides an efficient approach to multimodal video information processing without dependence on additional depth sensors.

Open Access
Review article
Multimodal Representation Learning for Binary Code Similarity Analysis: A Systematic Review and Conceptual Framework
rupesh kohli ,
harish parshuram bhabad ,
atmeshkumar subhashbhai patel ,
vijay m. rakhade ,
nandini s. patil ,
vedant kadlag
|
Available online: 06-30-2026

Abstract

Full Text|PDF|XML

Binary code similarity analysis is essential for software reverse engineering, vulnerability discovery, and malware analysis. However, conventional unimodal representations relying exclusively on linear opcode streams, control-flow graph (CFG) topologies, or dynamic system-call traces lack robustness when confronted with compiler transformations, adversarial obfuscation (e.g., Ultimate Packer for eXecutables (UPX) packing, control-flow flattening), and anti-analysis evasion. This study addresses these limitations through a systematic literature review and a unified conceptual framework. Following Preferred Reporting Items for Systematic Reviews and Meta-Analyses (PRISMA) 2020 guidelines, we systematically search five major academic digital databases (IEEE Xplore, ACM Digital Library, ScienceDirect, Scopus, and SpringerLink) covering 2020 through 2026. From 5,650 initially identified records, 11 primary benchmark and foundational studies are retained and categorized as direct multimodal or component-level evidence under an operational eight-dimension quality rubric (Q1–Q8). Based on the evidence synthesis, a conceptual architecture termed Multi-Modal Contrastive Binary Similarity Analysis (MM-CBSA) was proposed. Opcode sequences, control-flow graphs, and system-call traces are encoded using Transformer, Graph Isomorphism Network (GIN), and Bidirectional Long Short-Term Memory (BiLSTM) encoders, respectively, and projected onto a shared unit hypersphere. Cross-modal alignment is achieved via multi-channel Information Noise-Contrastive Estimation (InfoNCE) objectives with volumetric Gram regularization, paired with reliability-aware dynamic gating to accommodate degraded or missing modalities. Technical feasibility was examined using an associated open-source reference implementation. On a stored 1,500-sample test artifact (744 malware, 756 benign), the implementation achieved preliminary classification performance of 99.73% accuracy and 0.9973 F1-score, with a mean neural forward-pass latency of 28.87 ms on Central Processing Unit (CPU). Under preliminary stress testing, accuracy decreased to 88.0% under UPX packing and to 80.0% under dead-code insertion. Crucially, the manuscript establishes that current empirical evidence supports binary classification rather than direct binary similarity retrieval, and stored temporal and family-holdout artifacts warrant independent experimental revalidation. These observations provide technical-feasibility evidence only and do not constitute direct validation of binary code similarity retrieval. Accordingly, a reproducible empirical validation protocol is formulated to support future evaluation of multimodal binary code similarity analysis under compiler variation, obfuscation, distribution shift, and modality degradation.

- no more data -