Binary code similarity analysis is essential for software reverse engineering, vulnerability discovery, and malware analysis. However, conventional unimodal representations relying exclusively on linear opcode streams, control-flow graph (CFG) topologies, or dynamic system-call traces lack robustness when confronted with compiler transformations, adversarial obfuscation (e.g., Ultimate Packer for eXecutables (UPX) packing, control-flow flattening), and anti-analysis evasion. This study addresses these limitations through a systematic literature review and a unified conceptual framework. Following Preferred Reporting Items for Systematic Reviews and Meta-Analyses (PRISMA) 2020 guidelines, we systematically search five major academic digital databases (IEEE Xplore, ACM Digital Library, ScienceDirect, Scopus, and SpringerLink) covering 2020 through 2026. From 5,650 initially identified records, 11 primary benchmark and foundational studies are retained and categorized as direct multimodal or component-level evidence under an operational eight-dimension quality rubric (Q1–Q8). Based on the evidence synthesis, a conceptual architecture termed Multi-Modal Contrastive Binary Similarity Analysis (MM-CBSA) was proposed. Opcode sequences, control-flow graphs, and system-call traces are encoded using Transformer, Graph Isomorphism Network (GIN), and Bidirectional Long Short-Term Memory (BiLSTM) encoders, respectively, and projected onto a shared unit hypersphere. Cross-modal alignment is achieved via multi-channel Information Noise-Contrastive Estimation (InfoNCE) objectives with volumetric Gram regularization, paired with reliability-aware dynamic gating to accommodate degraded or missing modalities. Technical feasibility was examined using an associated open-source reference implementation. On a stored 1,500-sample test artifact (744 malware, 756 benign), the implementation achieved preliminary classification performance of 99.73% accuracy and 0.9973 F1-score, with a mean neural forward-pass latency of 28.87 ms on Central Processing Unit (CPU). Under preliminary stress testing, accuracy decreased to 88.0% under UPX packing and to 80.0% under dead-code insertion. Crucially, the manuscript establishes that current empirical evidence supports binary classification rather than direct binary similarity retrieval, and stored temporal and family-holdout artifacts warrant independent experimental revalidation. These observations provide technical-feasibility evidence only and do not constitute direct validation of binary code similarity retrieval. Accordingly, a reproducible empirical validation protocol is formulated to support future evaluation of multimodal binary code similarity analysis under compiler variation, obfuscation, distribution shift, and modality degradation.