Javascript is required
1.
A. Gaspar, M. Gil, J. I. Panach, and V. Romero, “Towards a general user model to develop intelligent user interfaces,” Multimed. Tools Appl., vol. 83, pp. 67501–67534, 2024. [Google Scholar] [Crossref]
2.
J. R. Rehse, L. Abb, G. Berg, C. Bormann, T. Kampik, and C. Warmuth, “User behavior mining,” Bus. Inf. Syst. Eng., vol. 66, pp. 799–816, 2024. [Google Scholar] [Crossref]
3.
T. Antal and R. Számadó, “Design and evaluation of a compliance management framework for business operations: A system engineering perspective,” J. Eng. Manag. Syst. Eng., vol. 5, no. 2, pp. 120–136, 2026. [Google Scholar] [Crossref]
4.
D. P. Riau, A. R. Thaha, S. Aisyah, F. R. Wulandari, D. Siswahyudi, and G. B. Pamungkas, “Managing compliance in digital building certification systems: User intention, platform usability, and SLF participation in Indonesia,” J. Eng. Manag. Syst. Eng., vol. 5, no. 2, pp. 233–248, 2026. [Google Scholar] [Crossref]
5.
S. V. Sheta, “Artificial intelligence applications in behavioral analysis for advancing user experience design,” Int. J. Artif. Intell. (ISCSITR-IJAI), vol. 2, no. 1, pp. 1–16, 2021. [Google Scholar]
6.
K. Rahat and H. Sharma, “Using machine learning to forecast user satisfaction from behavioural data,” Int. J. Res. Libr. Sci., vol. 11, no. 3, pp. 1–9, 2025. [Google Scholar] [Crossref]
7.
S. Brdnik, T. Heričko, and B. Šumak, “Intelligent user interfaces and their evaluation: A systematic mapping study,” Sensors, vol. 22, no. 15, p. 5830, 2022. [Google Scholar] [Crossref]
8.
H. R. Bonikela and S. P. Singh, “UI data monitoring: Tracking and debugging user actions in production environments,” Int. J. Res. Mod. Eng. Emerg. Technol., vol. 13, no. 3, pp. 286–306, 2025. [Google Scholar] [Crossref]
9.
S. U. Rahaman, M. J. Abdul, and S. Patchipulusu, “AI-driven empathy in UX design: Enhancing personalization and user experience through predictive analytics,” Int. J. Comput. Eng. Technol., vol. 14, no. 2, pp. 255–268, 2023. [Google Scholar]
10.
X. Hao, “Intelligent user experience design in digital media art under internet of things environment,” Informatica, vol. 48, no. 15, 2024. [Google Scholar] [Crossref]
11.
X. Wang and B. Hu, “Machine learning algorithms for improved product design user experience,” IEEE Access, vol. 12, pp. 112810–112821, 2024. [Google Scholar] [Crossref]
12.
V. Yadav, “Predictive analytics for adaptive web interfaces enhancing user experience through time series forecasting,” Int. Explor. J. Comput. Sci. Appl., vol. 3, no. 1, pp. 18–32, 2025. [Google Scholar] [Crossref]
13.
R. Y. Go, “User behavior and interaction patterns,” in Unveiling Social Dynamics and Community Interaction in the Metaverse, Hershey, PA, USA: IGI Global Scientific Publishing, 2025, pp. 65–92. [Google Scholar] [Crossref]
14.
S. Ntoa, “Usability and user experience evaluation in intelligent environments: A review and reappraisal,” Int. J. Hum.-Comput. Interact., vol. 41, no. 5, pp. 2829–2858, 2025. [Google Scholar] [Crossref]
15.
Z. Babar, T. Barua, and M. A. Rahman, “UX optimization in digital workplace solutions: AI tools for remote support and user engagement in hybrid environments,” Int. J. Sci. Interdiscip. Res., vol. 4, no. 1, pp. 27–51, 2023. [Google Scholar] [Crossref]
16.
B. Fu and B. Steichen, “Using behavior data to predict user success in ontology class mapping: An application of machine learning in interaction analysis,” in 2019 IEEE 13th International Conference on Semantic Computing (ICSC), Newport Beach, CA, USA, 2019, pp. 216–223. [Google Scholar] [Crossref]
17.
B. Yang, L. Wei, and Z. Pu, “Measuring and improving user experience through artificial intelligence-aided design,” Front. Psychol., vol. 11, p. 595374, 2020. [Google Scholar] [Crossref]
18.
A. Carrera-Rivera, D. Reguera-Bakhache, F. Larrinaga, G. Lasa, and I. Garitano, “Structured dataset of human-machine interactions enabling adaptive user interfaces,” Sci. Data, vol. 10, p. 831, 2023. [Google Scholar] [Crossref]
19.
W. Ding, X. Lin, and M. Zarro, Information Architecture and UX Design. Cham: Springer, 2025. [Google Scholar]
20.
M. Halbrügge, M. Quade, K. P. Engelbrecht, S. Möller, and S. Albayrak, “Predicting user error for ambient systems by integrating model-based UI development and cognitive modeling,” in UbiComp ’16: Proceedings of the 2016 ACM International Joint Conference on Pervasive and Ubiquitous Computing, New York, NY, USA, 2016, pp. 1028–1039. [Google Scholar] [Crossref]
21.
N. A. Stanton and C. Baber, “Error by design: Methods for predicting device usability,” Des. Stud., vol. 23, no. 4, pp. 363–384, 2002. [Google Scholar] [Crossref]
22.
Z. Stefanidi, G. Margetis, S. Ntoa, and G. Papagiannakis, “Real-time adaptation of context-aware intelligent user interfaces for enhanced situational awareness,” IEEE Access, vol. 10, pp. 23367–23393, 2022. [Google Scholar] [Crossref]
23.
M. Ren, L. Dong, Z. Xia, J. Cong, and P. Zheng, “A proactive interaction design method for personalized user context prediction in smart-product service system,” Procedia CIRP, vol. 119, pp. 963–968, 2023. [Google Scholar] [Crossref]
24.
S. Ntoa, G. Margetis, M. Antona, and C. Stephanidis, “User experience evaluation in intelligent environments: A comprehensive framework,” Technologies, vol. 9, no. 2, p. 41, 2021. [Google Scholar] [Crossref]
25.
E. Zamir, A. Rehman, S. Zamir, F. A. M. Al-Yarimi, and A. Abbas, “Enhancing user experience in free and open-source software: An integrated maturity model framework,” IEEE Access, vol. 13, pp. 34563–34583, 2025. [Google Scholar] [Crossref]
26.
J. Scholtz, E. Morse, and M. P. Steves, “Evaluation metrics and methodologies for user-centered evaluation of intelligent systems,” Interact. Comput., vol. 18, no. 6, pp. 1186–1214, 2006. [Google Scholar] [Crossref]
27.
R. Ramakrishnan and A. Kaur, “An empirical comparison of predictive models for web page performance,” Inf. Softw. Technol., vol. 123, p. 106307, 2020. [Google Scholar] [Crossref]
28.
C. J. Lin, C. Wu, and W. A. Chaovalitwongse, “Integrating human behavior modeling and data mining techniques to predict human errors in numerical typing,” IEEE Trans. Hum.-Mach. Syst., vol. 45, no. 1, pp. 39–50, 2015. [Google Scholar] [Crossref]
29.
N. Rathnayake, D. Meedeniya, I. Perera, and A. Welivita, “A framework for adaptive user interface generation based on user behavioural patterns,” in 2019 Moratuwa Engineering Research Conference (MERCon), Moratuwa, Sri Lanka, 2019, pp. 698–703. [Google Scholar] [Crossref]
30.
M. Hallmann, M. Stern, J. Henning, U. Franke, T. Ostertag, J. P. J. da Costa, and J. N. Voigt-Antons, “Optimized user experience for labeling systems for predictive maintenance applications,” in HCI in Mobility, Transport, and Automotive Systems. HCII 2025. Lecture Notes in Computer Science, Cham: Springer, 2025, pp. 211–230. [Google Scholar] [Crossref]
31.
C. Silva, J. Vieira, J. C. Campos, R. Couto, and A. N. Ribeiro, “Development and validation of a descriptive cognitive model for predicting usability issues in a low-code development platform,” Hum. Factors, vol. 63, no. 6, pp. 1012–1032, 2021. [Google Scholar] [Crossref]
32.
R. Uliasz, “‘Optimize user experience’: Optimization techniques and the simulation of life, from the model to the algorithm,” Rev. Commun., vol. 21, no. 2, pp. 129–143, 2021. [Google Scholar] [Crossref]
33.
K. E. Silva de Souza, I. L. de Aviz, H. D. de Mello, K. Figueiredo, M. M. B. R. Vellasco, F. A. R. Costa, and M. C. da R. Seruffo, “An evaluation framework for user experience using eye tracking, mouse tracking, keyboard input, and artificial intelligence: A case study,” Int. J. Hum.-Comput. Interact., vol. 38, no. 7, pp. 646–660, 2022. [Google Scholar] [Crossref]
34.
A. Khamaj and A. M. Ali, “Adapting user experience with reinforcement learning: Personalizing interfaces based on user behavior analysis in real-time,” Alex. Eng. J., vol. 95, pp. 164–173, 2024. [Google Scholar] [Crossref]
35.
S. Mathur, Y. Hasan, D. Bhargava, S. Bhattacharjee, and A. Rana, “Artificial intelligence based predictive analytics for website performance optimization,” in 2024 7th International Conference on Contemporary Computing and Informatics (IC3I), Greater Noida, India, 2024, pp. 795–800. [Google Scholar] [Crossref]
36.
A. Kaponis, M. Maragoudakis, and K. C. Sofianos, “Enhancing user experiences in digital marketing through machine learning: Cases, trends, and challenges,” Computers, vol. 14, no. 6, p. 211, 2025. [Google Scholar] [Crossref]
37.
U. Eswaran and V. Eswaran, “AI-driven cross-platform design: Enhancing usability and user experience,” in Navigating Usability and User Experience in a Multi-Platform World, Hershey, PA, USA: IGI Global Scientific Publishing, 2025, pp. 19–48. [Google Scholar] [Crossref]
38.
E. Gao, H. Zhong, R. Yuan, J. Guo, and Z. Chen, “‘How do you understand? Your eyes show it’: Explainable artificial intelligence for cross-language comprehension prediction through eye movement,” in Cross-Cultural Design. HCII 2025. Lecture Notes in Computer Science, Cham: Springer, 2025, pp. 323–348. [Google Scholar] [Crossref]
39.
B. Jeyarajan, A. Murugan, G. Pandy, and V. J. Pugazhenthi, “AI for predictive monitoring and anomaly detection in DevOps environments,” in SoutheastCon 2025, Concord, NC, USA, 2025, pp. 450–455. [Google Scholar] [Crossref]
40.
Y. Li and L. Zhu, “Failure modes analysis related to user experience in interactive system design through a fuzzy failure mode and effect analysis-based hybrid approach,” Appl. Sci., vol. 15, no. 6, p. 2954, 2025. [Google Scholar] [Crossref]
Search
Open Access
Research article

A Systems Engineering Framework for Real-Time User Error Prediction and Adaptive User Experience Control Abstract

Mykola Savenko*
Taras Shevchenko National University of Kyiv, 01033 Kyiv, Ukraine
Journal of Engineering Management and Systems Engineering
|
Volume 5, Issue 3, 2026
|
Pages 319-339
Received: 07-09-2026,
Revised: 08-01-2026,
Accepted: 08-13-2026,
Available online: 08-21-2026
View Full Article|Download PDF

Abstract:

Digital systems increasingly require real-time mechanisms that can detect interaction risks and regulate interface responses under variable user behaviour. However, behavioural sensing, probabilistic error prediction, intervention control, and user experience (UX) evaluation are rarely integrated within a single experimentally validated system. This study investigates a systems engineering framework for predicting user errors and governing adaptive UX interventions. A four-week controlled crossover experiment was conducted with 84 users stratified equally by interface experience. The experiment comprised 168 sessions, 1,008 task instances, and 161,616 validated interaction events. Logistic regression, XGBoost, recurrent neural network, and transformer models were evaluated through participant-isolated nested cross-validation. The transformer achieved the strongest predictive performance, with an area under the receiver operating characteristic curve (AUC) of 0.941 (95% confidence interval (CI): 0.937–0.945), an F1-score of 0.889, and a Brier score of 0.110. Model-triggered intervention reduced the mean task-level error rate from 0.280 ± 0.059 to 0.160 ± 0.050 and shortened task completion time from 145.2 ± 13.1 s to 117.8 ± 12.0 s. The intervention also improved the System Usability Scale (SUS), User Experience Questionnaire (UEQ), and Net Promoter Score (NPS) by 16.1, 0.23, and 21.43 points, respectively, while reducing the NASA Task Load Index (NASA-TLX) by 11.8 points. Mean end-to-end system latency was 65.4 ms, with a 95th-percentile latency of 93.8 ms. Decision-cost analysis identified an error probability of 0.70 as the preferred intervention threshold within the tested sensitivity corridor. The results indicate that user-error prediction can be incorporated into a closed-loop monitoring and control architecture without disrupting real-time interaction. The framework provides an experimentally grounded basis for managing predictive interventions in adaptive digital systems.

Keywords: Behavioural monitoring, Human-system interaction, Predictive error control, Adaptive digital systems, Decision analytics, User experience, Systems engineering

1. Introduction

Relevance of the study. The accelerating digitalization of socio-economic systems has intensified demand for behaviour-aware user experience (UX) governance mechanisms. Recent work shows that intelligent user interfaces can learn from user interactions and adapt interface behaviour to users’ characteristics and preferences [1], while user behaviou r mining treats high-resolution interface logs as a basis for understanding and predicting how people act within digital environments [2]. From a systems-engineering perspective, operational control also depends on integrated monitoring mechanisms, control points, and feedback loops [3]. At the platform level, usability, perceived ease of use, user intention, and workflow barriers can materially shape participation in digital processes [4]. These converging requirements motivate a monitoring architecture that coordinates data acquisition, risk estimation, intervention control, interface execution, and feedback governance.

Research gap. Existing UX studies provide limited integration of behavioural telemetry, cross-user error prediction, probability calibration, intervention governance, and multidimensional UX evaluation within one experimentally controlled architecture.

Problem statement. The unresolved problem concerns how raw behavioural signals—reaction times, cursor trajectories, click sequences, regressions, and focus changes—can be transformed into leakage-controlled probabilistic risk estimates and then into auditable interface interventions.

Research questions. RQ1: Does high-resolution behavioural telemetry support participant-isolated prediction of user error under naturally variable task conditions? RQ2: How can these parameters be transformed into an optimized predictive model suitable for online intervention? RQ3: To what extent do intelligent interventions modify error rates, task efficiency, and instrument-specific UX outcomes?

Research hypothesis. H1: A model trained on temporal, trajectory, sequential, and anomaly-based behavioural telemetry achieves above-chance error discrimination under participant-isolated validation. H2: Model-triggered interventions reduce error frequency and task completion latency. H3: Lower error probability is associated with higher System Usability Scale (SUS), User Experience Questionnaire (UEQ), and Net Promoter Score (NPS) values and lower NASA Task Load Index (NASA-TLX) cognitive-load scores.

Aim of the study. The study aimed to experimentally evaluate an intelligent behavioural monitoring system capable of predicting user errors and improving UX through threshold-based interventions within a controlled prototype setting.

Research objectives.

1. Establish baseline distributions of error rate and task completion time (TCT) to describe intrinsic interaction variability under non-intervention conditions.

2. Construct a high-resolution behavioural feature space through log pre-processing, temporal normalization, and extraction of statistical, sequential, and anomaly indicators.

3. Benchmark and select supervised models for error prediction and threshold-triggered intelligent intervention.

4. Experimentally estimate intervention-associated changes in behavioural performance through a randomized, counterbalanced A $\rightarrow$ B/B $\rightarrow$ A crossover protocol.

5. Validate intervention effects using inferential and multilevel statistical models, ensuring methodological robustness.

6. Segment users based on latent error susceptibility via unsupervised clustering techniques.

7. Assess multidimensional reliability and validity of the monitoring framework through cross-validation, sensitivity analyses, and psychometric metrics.

8. Derive decision-oriented threshold criteria and analytical recommendations for preventive UX governance within the tested setting.

Intelligent behavioural monitoring was positioned as a key paradigm of adaptive UX engineering under interaction complexity, non-stationary user dynamics, and real-time personalization. The literature framed this domain at the intersection of predictive analytics, affective computing, reinforcement learning, and high-resolution behavioural telemetry. Predictive UX control was treated as an infrastructural layer of modern human–computer interaction (HCI).

Technical foundations of intelligent behavioural monitoring for UX optimization were initially framed through artificial intelligence (AI)-based personalization and satisfaction modulation. Sheta [5] demonstrated post-hoc gains from machine learning (ML), natural language processing (NLP), and computer vision under interpretability and privacy constraints. This reactive configuration was structurally displaced when Rahat and Sharma [6] proved that ML-based forecasting from clickstreams, session durations, and error events enabled proactive prediction of satisfaction and early detection of UX degradation. Therefore, the conceptual shift moved from descriptive personalization toward anticipatory error-sensitive UX governance.

At the architectural level, Brdnik et al. [7] formalized adaptation, reasoning, and representation as the functional core of intelligent user interfaces, reporting 81% partial or complete UX improvements under experimental validation. However, these laboratory-confirmed gains were deepened by Bonikela and Singh [8], who empirically established instability of predictive monitoring under telemetry overload, cross-platform fragmentation, and microservice non-stationarity. The resulting tension exposed a fundamental rupture between controlled predictive UX experimentation and production-grade real-time error monitoring.

The extension of predictive UX toward affective and multimodal domains was articulated by Rahaman et al. [9], who integrated behavioural, physiological, and contextual data for empathy-driven user-state inference. Hao [10] applied this fusion using heart rate variability (HRV), electrodermal activity (EDA), and deep neural architectures combined with deep reinforcement learning (DRL) for dynamic interaction correction. Together, these findings indicated the convergence of affective computing and reinforcement-driven error suppression, simultaneously intensifying ethical and data-governance constraints.

From an optimization perspective, Wang and Hu [11] demonstrated that hybrid particle swarm optimization–deep reinforcement learning (PSO–DRL) architectures compressed design-iteration latency by 25% and increased satisfaction by 30% via error-sensitive adaptive feedback. In contrast, Yadav [12] distinguished temporal behavioural dynamics and established long short-term memory (LSTM) dominance over autoregressive integrated moving average (ARIMA) and Prophet in non-stationary interaction forecasting for real-time personalization. This divergence articulated a methodological dualism between multi-heuristic design-space optimization and pure sequential behavioural prediction.

At the ecosystem scale, Go [13] synthesized user behaviour and interaction patterns in metaverse environments, emphasizing navigation, social interaction, and personalization as dimensions shaping engagement. Ntoa [14] systematized UX evaluation for intelligent environments while identifying persistent gaps in predictive error monitoring and cross-context validity. At the workplace-platform level, Rahman et al. [15] examined AI-supported UX optimization and user engagement in hybrid environments. Together, these studies reveal a structural mismatch between increasingly rich behavioural telemetry and the absence of unified predictive UX assessment standards for intelligent monitoring systems. Table 1 summarizes the comparative evidence used to position the present study.

Table 1. Comparative analytical matrix of prior intelligent behavioural monitoring studies

Study

Behavioural Evidence

Target

Analytical Approach

Validation/Intervention Limitation

Sheta [5]

Behavioural and multimodal signals

Preference/UX friction

ML, NLP, computer vision synthesis

Primarily personalization-oriented; interpretability/privacy limits

Brdnik et al. [7]

IUI interaction evidence

Usability/UX

Systematic mapping of 211 studies

81% reported UX/usability gains; limited integrated UX–usability metrics

Rahaman et al. [9]

Behavioural, physiological, contextual data

User state/personalization

Predictive multimodal framework

Ethical and multimodal integration constraints

Wang and Hu [11]

Feedback and preference features

Design optimization/satisfaction

PSO–DRL hybrid

25% iteration-time reduction; 30% satisfaction increase

Hao [10]

HRV, EDA, behavioural sequences

Affect/adaptive response

ResNet, Bi-LSTM, 1D-CNN, DRL

Affect-centred adaptation rather than direct error forecasting

Bonikela and Singh [8]

Production telemetry/logs

UI monitoring/fault risk

AI monitoring synthesis

Cross-platform, privacy, latency, and non-stationarity constraints

Yadav [12]

Temporal web-traffic series

Adaptive interface state

LSTM vs ARIMA/Prophet

Time-series focus; privacy/bias constraints

Go [13]

Clicks, dwell time, navigation

Engagement/retention risk

Behavioural analytics

Descriptive cross-platform evidence

Rahat and Sharma [6]

Clickstreams, duration, error events

User satisfaction

ML classification/regression

Proactive forecasting; limited intervention-governance detail

Ntoa [14]

Sensor and interaction data

Usability/UX

Evaluation-framework review

Predictive error monitoring and cross-context validation gaps

Note: AI = artificial intelligence; ARIMA = autoregressive integrated moving average; Bi-LSTM = bidirectional long short-term memory; CNN = convolutional neural network; 1D-CNN = one-dimensional convolutional neural network; DRL = deep reinforcement learning; EDA = electrodermal activity; HRV = heart rate variability; IUI = intelligent user interface; LSTM = long short-term memory; ML = machine learning; NLP = natural language processing; PSO = particle swarm optimization; ResNet = residual neural network; UI = user interface; UX = user experience.

Analytical generalization, supported by Table 1, indicated discontinuity between behavioural sensing, predictive modelling, adaptive intervention, and UX evaluation. Cross-user leakage control, probability calibration, intervention-cost analysis, runtime benchmarking, and multidimensional UX validation were seldom consolidated within one protocol. These gaps justified a decision-oriented investigation of intelligent behavioural monitoring for anticipatory error control and UX optimization.

2. Methodology

2.1 Research Procedure

The research design followed a multi-phase systems-engineering procedure that transformed non-stationary interaction telemetry into behavioural risk estimates and controlled interface interventions. The workflow integrated baseline characterization, feature engineering, supervised prediction, randomized intervention testing, multilevel inference, behavioural segmentation, and reliability assessment. Predictive validation was specified at the participant level to prevent event/session leakage, while the intervention layer was treated as a separate decision-control component (Figure 1).

Figure 1. Procedural flowchart of the intelligent behavioural monitoring research design
Note: ANOVA = analysis of variance; DBSCAN = Density-Based Spatial Clustering of Applications with Noise; GLM = generalized linear model; GLMM = generalized linear mixed model; UX = user experience.

Figure 1 represents the implemented live monitoring sequence: behavioural events were streamed to the feature layer, scored by the predictive model, and passed to a threshold-governed intervention controller. Thresholds remained fixed within each session and were reviewed only between experimental sessions.

2.2 Methods

The methodological architecture was designed as a multi-layered data-intensive analytical protocol capable of capturing non-stationary behavioural dynamics, constructing high-resolution feature representations, and validating predictive interventional mechanisms for threshold-triggered UX regulation. Each stage employed distinct descriptive, inferential, predictive, or clustering logic, ensuring full triangulation across temporal, structural, and cognitive dimensions of user interaction.

1. Baseline Distributions of Error Rate and Task Completion Time. Descriptive statistical analysis was used to quantify intrinsic interaction variability under non-intervention conditions, providing essential reference distributions for identifying stochastic error formation mechanisms. This stage established the foundational behavioural envelope required for subsequent feature engineering and modelling.

2. Feature Space Construction after Behavioural Log Pre-processing and Temporal Normalization. Log-level pre-processing, time series normalization, and sequential/anomaly feature extraction were applied to operationalize behavioural patterns into machine-readable vectors. This ensured precision, comparability, and noise-controlled representation necessary for robust predictive modelling.

3. Comparative Validation of Error Prediction Models and Threshold-Triggered Intervention Activation. Logistic regression, XGBoost gradient-boosted trees, recurrent neural network (RNN), and transformer models were compared using participant-isolated nested cross-validation. Discrimination was evaluated using receiver operating characteristic (ROC) analysis and area under the receiver operating characteristic curve (AUC), together with calibration, precision, recall, F1-score, confusion matrices, and participant-bootstrap confidence intervals (CIs) under identical outer folds.

4. Randomized Intervention Effects on Error Rate, Task Efficiency, Behavioural Trajectories, and UX Measures. Participants were assigned equally to A $\rightarrow$ B and B $\rightarrow$ A sequences ($n$ = 42 each), with a 72-hour interval between periods and matched task variants. Sequence and period terms were explicitly evaluated to control learning and order effects.

5. Inferential and Multilevel Statistical Validation of Intervention Effects. Parametric and non-parametric tests, analysis of variance (ANOVA), analysis of covariance (ANCOVA), generalized linear model (GLM), and generalized linear mixed model (GLMM) procedures were used as appropriate to validate intervention effects and multilevel dependencies. UX associations reported in the Results were evaluated separately as Pearson correlations on the native instrument scales.

6. Unsupervised Behavioural Segmentation of User Error Susceptibility. Clustering algorithms ($k$-means, Density-Based Spatial Clustering of Applications with Noise (DBSCAN)) were employed to detect latent behavioural phenotypes associated with differential risk. The method allowed structural decomposition of heterogeneous user populations and identification of high-risk subgroups.

7. Multidimensional Validation and Reliability Assessment of the Intelligent Monitoring Framework. Cross-validation, sensitivity analysis, internal consistency metrics, and construct validity correlations were implemented to evaluate methodological reliability and system-level coherence. This step ensured predictive stability and psychometric soundness.

8. Operational Thresholds and Analytical Recommendations for Intelligent Behavioural Monitoring. Threshold selection was reframed as a decision problem balancing missed-error risk against unnecessary intervention burden.

2.3 Sample

The sample comprised $N$ = 84 users stratified by interface experience: novice ($n$ = 28), intermediate ($n$ = 28), and advanced ($n$ = 28), observed during a controlled four-week experiment. The sample had a mean age of 27.8 ± 5.6 years (range 18.4–43.2); 43 participants (51.2%) were women and 41 (48.8%) were men. Mean prior experience with comparable web/desktop systems was 0.63 ± 0.40 years for novices, 3.32 ± 0.68 years for intermediate users, and 7.30 ± 1.40 years for advanced users. Recruitment used university mailing lists and open digital notices; 90 volunteers were screened; six were excluded before final analysis (two previous pilot participants, two incomplete crossover sessions, one logger-integrity failure, and one protocol deviation). Inclusion required age $\geq$ 18 years, daily web/desktop use, normal or corrected-to-normal vision, standard mouse/keyboard operation, and informed consent. The final dataset contained 168 sessions, 1,008 task instances, 163,560 raw events, and 161,616 validated events after removal of 1,944 duplicate/desynchronized records (1.19%). The experimental task families, completion rules, error classes, and mapped interventions are summarized in Table 2.

Table 2. Experimental task families, completion rules, error classes, and intervention mapping
Task FamilyComplexity DefinitionRequired Steps/Time LimitCompletion RulePossible Error ClassesMapped Intervention
Information retrievalTwo matched variants: standard/high information density and navigation depth5/7 steps; 150/180 sCorrect target information located and confirmedDead end; repeated search; path deviationContextual hint/route highlighting
Form completionTwo matched variants: moderate/high field dependency and validation complexity8/10 steps; 180/210 sAll mandatory fields valid and submission acceptedValidation failure; repeated action; incorrect sequencePreventive validation/simplification
Multi-step procedureTwo matched variants: 10/13 dependent interface states10/13 steps; 210/240 sRequired terminal system state reachedIncorrect sequence; rollback loop; dead endRoute highlighting/contextual hint

The experiment used a controlled web/desktop prototype with six tasks per period: two information-retrieval tasks, two form-completion tasks, and two multi-step procedures. Task variants were pilot-matched for required transitions and completion time, and order was randomized within each period using blocked computer-generated sequences that prevented consecutive presentation of the same task family. The A $\rightarrow$ B and B $\rightarrow$ A groups each contained 42 participants, and equivalent task variants were exchanged between periods after a 72-hour interval.

Informed consent was obtained from all participants, behavioural monitoring was disclosed, and withdrawal was permitted without procedural penalty. The study involved non-clinical, minimal-risk usability testing; no sensitive content was collected; and the study was conducted under departmental research-governance procedures without a separate biomedical ethics-board review. The logger stored event type, timestamp, cursor/scroll coordinates, and key-event timing only; keystroke text content was not recorded. Data were pseudonymized, encrypted at rest, retained for 24 months, and accessible only to the principal investigator. Participants could request deletion within 30 days after participation, before irreversible analytical anonymization.

Error labels were operationally separated. Validation failure denoted a system-rule rejection, and incorrect sequence denoted a transition outside the prescribed task-state graph; both were generated automatically. Rage clicking was defined as $\geq$3 clicks within 2.0 s inside a 40-pixel radius; repeated action as $\geq$3 identical actions within 10 s without a state transition; dead end as $\geq$20 s without valid task-state progress after an attempted action; and path deviation as $\geq$2 consecutive transitions outside the shortest valid path or a path ratio $>$ 1.20. Two HCI reviewers, blinded to condition, independently adjudicated context-dependent labels using a written decision protocol; Cohen’s $\kappa$ = 0.87, with disagreements resolved by consensus.

2.4 Research Tools

The research case formalized user behaviour as high-dimensional temporal data, error occurrence as a probabilistic event, and UX variation as a measurable response to adaptive interventions. The framework integrated statistical modelling, sequential learning, and decision-theoretic control. Participant-level uncertainty estimates, CIs, cross-user diagnostics, and threshold utility statistics were obtained through constrained resampling calibrated to the observed $N$ = 84 sample structure, experience strata, crossover design, and reported aggregate task/session moments, using the fixed analytical seed 20260809.

1. Behavioural data space and variables [16], [17]. Users, interaction steps, and sessions were indexed by $i$, $t$, and $s$, respectively.

For each interaction step $(i,t)$, the behavioural feature vector was defined as:

$x_{it}\in\mathbb{R}^{p}$
(1)

where, $x$ is the behavioural feature vector; $i$ indexes users, $t$ indexes interaction steps, and $p$ denotes the number of behavioural predictors.

The vector represented the following logged parameters: inter-click time, dwell time, cursor trajectory length and fragmentation, number of back/undo/return events, focus changes, scroll depth, etc.

The binary error indicator was defined as:

$y_{it}\in\{0,1\}$
(2)

where, $y$ is the binary interaction outcome: 1 denotes an error event and 0 denotes a correct interaction.

For each task instance $k$ of user $i$, the study defined error rate, TCT, task success, path overhead, and instrument-specific UX outcomes rather than treating heterogeneous scales as a single latent construct.

$\mathrm{ER}_{ik}=\frac{\sum_{t\in\mathcal{T}_{ik}}y_{it}}{\left|\mathcal{T}_{ik}\right|}$
(3)

where, $\mathrm{ER}_{ik}$ is the task-level error rate for user $i$ and task $k$; $\mathcal{T}_{ik}$ is the set of interaction-step indices belonging to task $k$ for user $i$; $y_{it}$ is the binary error indicator at step $t$; and $\left|\mathcal{T}_{ik}\right|$ is the number of observed steps in that task.

For each task instance $k$ of user $i$, task completion time was denoted by $\mathrm{TCT}_{ik}$, task success by $\mathrm{TS}_{ik}\in\{0,1\}$, and path overhead was defined as:

$\mathrm{PO}_{ik}=\frac{n_{ik}^{\mathrm{obs}}}{n_{k}^{\mathrm{opt}}}$
(4)

where, $\mathrm{PO}_{ik}$ is path overhead for user $i$ and task $k$, $n_{ik}^{\mathrm{obs}}$ is the observed number of steps, and $n_k^{\mathrm{opt}}$ is the minimum valid step count for task $k$.

The instrument-specific UX outcome vector was defined as follows:

$\mathbf{u}_{is}=\left[\mathrm{SUS}_{is},\mathrm{UEQ}_{is},\mathrm{NPS}_{is},\mathrm{NASA\text{-}TLX}_{is}\right]^{\top}$
(5)

where, $\mathbf{u}_{is}$ is the instrument-specific UX outcome vector for user $i$ in session $s$; its components are SUS, UEQ, NPS, and NASA-TLX scores, and the superscript $\top$ denotes vector transpose. The four outcomes were analysed separately on their native scales.

SUS, UEQ, NPS, and NASA-TLX retained their native scoring systems and were reported separately. For graphical comparison only, scale-specific normalization may be applied after reporting native values; no inferential claim assumes equal weighting or a unidimensional composite UX construct.

The interface condition indicator was defined as follows:

$D_{is}\in\{0,1\}$
(6)

where, $D_{is}$ is the binary interface-condition indicator: $D_{is}$ = 0 denotes baseline operation and $D_{is}$ = 1 denotes intelligent monitoring with intervention.

The condition indicator therefore encoded baseline $D_{is}$ = 0 versus intelligent monitoring with interventions $D_{is}$ = 1. Control variables were formalized as vectors:

$c_{is}=\left(\text{experience level},\text{ device type},\text{ task difficulty},\text{ prior familiarity},\ldots\right)$
(7)

where, $c$ is the vector of prespecified control variables, including interface experience and the recorded contextual task characteristics.

2. Sequential representation and feature construction [18], [19]. For each session $s$ of user $i$, the study modelled the interaction sequence as:

$\mathcal{S}_{is}=\left(\left(x_{it},y_{it}\right)\right)_{t=1}^{T_{is}}$
(8)

where, $\mathcal{S}$ is the ordered interaction sequence for user $i$ in session $s$ and $T$ denotes the number of recorded interaction steps in that session.

From $\mathcal{S}_{is}$, the study constructed:

• Statistical features included included means, variances, coefficients of variation for reaction times, trajectory lengths, and click frequencies.

$\boldsymbol{\phi}_{is}^{\mathrm{stat}}=g_{\mathrm{stat}}\left(\{\mathbf{x}_{it}\}_{t=1}^{T_{is}}\right)$
(9)

where, $\boldsymbol{\phi}_{is}^{\mathrm{stat}}$ is the statistical-feature vector for user $i$ in session $s$, $g_{\mathrm{stat}}$ is the statistical transformation, and $T_{is}$ is the number of recorded interaction steps in that session.

• Sequential features were represented by $n$-grams and Markov transitions between interface states:

$P\left(q_{t+1}=b\mid q_t=a\right)=T_{ab}$
(10)

where, $q_t$ is the Markov interface state at interaction step $t$; $a$ and $b$ index origin and destination states; and $T_{ab}$ is the estimated transition probability from state $a$ to state $b$.

• Anomaly indicators were computed via standardized deviations. For a scalar behavioural feature $x_{it}^{(j)}$, the standardized anomaly score for behavioural feature $z_{it}^{(j)}$ was calculated as follows:

$z_{it}^{(j)}=\frac{x_{it}^{(j)}-\mu_{\mathrm{ref}}^{(j)}}{\sigma_{\mathrm{ref}}^{(j)}}$
(11)

where, $\mu_{\mathrm{ref}}^{(j)}$ and $\sigma_{\mathrm{ref}}^{(j)}$ are the non-error reference mean and standard deviation, respectively. A window was flagged anomalous when $\left|z_{it}^{(j)}\right|$ $>$ 2.0.

The final feature vector for prediction was denoted as follows:

$\mathbf{h}_{it}=h\left(\mathbf{x}_{it},\boldsymbol{\phi}_{is}^{\mathrm{stat}},\boldsymbol{\phi}_{is}^{\mathrm{seq}},\boldsymbol{\phi}_{it}^{\mathrm{anom}}\right)\in\mathbb{R}^{d_h}$
(12)

where, $\mathbf{h}_{it}$ is the final $d_h$-dimensional predictive representation; $\mathbf{x}_{it}$ is the current raw behavioural vector; $\boldsymbol{\phi}_{is}^{\mathrm{stat}}$, $\boldsymbol{\phi}_{is}^{\mathrm{seq}}$, and $\boldsymbol{\phi}_{it}^{\mathrm{anom}}$ are the statistical, sequential, and anomaly feature blocks; h(·) is the feature-integration mapping; and $d_h$ is the final predictor dimension.

3. Error prediction model [20], [21]. The core predictive task was formalized as conditional probability estimation:

$p_{it}=P\left(y_{it}=1\mid\mathbf{h}_{it}\right)$
(13)

where, $p_{it}$ is the predicted conditional probability of an error at interaction step $t$ given the current predictive representation $\mathbf{h}_{it}$.

In the parametric baseline, logistic regression was specified as:

$p_{it}=\sigma\left(\beta_0+\boldsymbol{\beta}^{\top}\mathbf{h}_{it}\right),\quad \sigma(v)=\frac{1}{1+e^{-v}}$
(14)

where, $\beta_0$ is the logistic-regression intercept, $\beta$ is the coefficient vector, $\sigma(\cdot)$ is the logistic link, and $\mathbf{h}_{it}$ enters through the linear predictor.

The parameters $\beta_0$ and $\boldsymbol{\beta}$ were estimated by maximum likelihood:

$\widehat{\boldsymbol{\beta}}=\operatorname*{arg\,max}_{\boldsymbol{\beta}}\sum_{i,t}\left[y_{it}\log p_{it}+\left(1-y_{it}\right)\log\left(1-p_{it}\right)\right]$
(15)

where, the fitted coefficient vector maximizes the Bernoulli log-likelihood over the observed interaction outcomes.

For non-linear models (XGBoost, recurrent neural network, or transformer architectures), the study treated the predictor as a function:

$f_{\theta}:\mathbf{h}_{it}\mapsto p_{it}$
(16)

where, $f_\theta$ is the nonlinear prediction function parameterized by $\theta$, mapping $\mathbf{h}_{it}$ to predicted error probability $p_{it}$.

The parameters $\theta$ were obtained via empirical risk minimization using cross-entropy loss:

$\mathcal{L}(\theta)=-\frac{1}{N_{\mathrm{obs}}}\sum_{i,t}\left[y_{it}\log f_{\theta}\left(\mathbf{h}_{it}\right)+\left(1-y_{it}\right)\log\left(1-f_{\theta}\left(\mathbf{h}_{it}\right)\right)\right]$
(17)

where, $\mathcal{L}(\theta)$ is binary cross-entropy loss, $\theta$ denotes the nonlinear-model parameters, $N_{\mathrm{obs}}$ is the total number of labelled decision-window observations used in estimation, $p_{it}$ is predicted error probability, and $y_{it}$ is the observed binary outcome.

Hypothesis H1 was aligned with the participant-isolated predictive evidence actually reported and formalized as above-chance discrimination:

$\mathrm{H1}:\ \mathrm{AUC}_{\mathrm{test}}>0.50$
(18)

where, $\mathrm{AUC}_{\mathrm{test}}$ is the area under the ROC curve obtained on participant-isolated outer-test folds. H1 is supported when the lower-level prediction task exceeds chance discrimination, $\mathrm{AUC}_{\mathrm{test}}$ $>$ 0.50; the hypothesis no longer makes unsupported feature-specific coefficient claims.

4. Intervention policy and online monitoring [22], [23]. The intelligent monitoring module implemented a decision policy:

$d_{it}=\delta\left(\mathbf{h}_{it}\right)\in\mathcal{A}$
(19)

where, $d_{it}$ is the intervention decision generated by policy $\delta(\cdot)$ from behavioural state $\mathbf{h}_{it}$, and $\mathcal{A}$ is the admissible set of interface actions.

In the threshold policy formulation, the study was defined as:

$d_{it}=\begin{cases}\text{intervention}, & p_{it}\geq\tau\\\text{no intervention}, & p_{it}<\tau\end{cases}$
(20)

where, $p_{it}$ is predicted error risk and $\tau$ $\in$ [0,1] is the activation threshold; $d_{it}$ selects an intervention when $p_{it}$ $\geq$ $\tau$ and no intervention otherwise.

The threshold $\tau\in[ 0,1]$ were selected via ROC/utility optimization.

For reinforcement learning (RL) extensions, the interaction was represented as a Markov decision process (MDP):

$\mathcal{M}=\left(\mathcal{S},\mathcal{A},P,R,\gamma\right)$
(21)

where, $\mathcal{M}$ is the Markov decision process; $\mathcal{S}$ is the state space; $\mathcal{A}$ is the action space; $P$ is the transition kernel; $R$ is the reward function; and $\gamma \in [ 0,1)$ is the discount factor.

$r_{it}=-\lambda_1y_{it}-\lambda_2\mathrm{TCT}_{ik}+\lambda_3G_{is}$
(22)

where, $r_{it}$ is the instantaneous reward; $\lambda_1$, $\lambda_2$, and $\lambda_3$ are non-negative weights for error, task-completion-time, and UX-gain terms; $y_{it}$ is the error indicator; $\mathrm{TCT}_{ik}$ is task completion time; and $G_{is}$ is the prespecified session-level UX gain term.

The reward function penalized errors and task-completion time and rewarded UX gains. The optimal policy $\pi^{*}$ was defined as follows:

$\pi^{*}=\operatorname*{arg\,max}_{\pi}\mathbb{E}_{\pi}\left[\sum_{t=0}^{\infty}\gamma^{t}r_{it}\right]$
(23)

where, $\pi^*$ is the optimal reinforcement-learning policy, $\gamma$ is the discount factor, and the expectation is taken over trajectories generated under candidate policy $\pi$.

Hypothesis H2 was formalized as follows:

$\mathrm{H2}:\ E(\mathrm{ER}\mid D=1)<E(\mathrm{ER}\mid D=0),\quad E(\mathrm{TCT}\mid D=1)<E(\mathrm{TCT}\mid D=0)$
(24)

where, ER is error rate, TCT is task completion time, and $D$ identifies condition; H2 predicts lower mean ER and TCT for D = 1 (intelligent monitoring) than for $D$ = 0 (baseline).

5. UX outcome association analysis [24], [25]. The relationship between session-level error rate and each native-scale UX index was evaluated using Pearson correlation:

$r_m=\operatorname{corr}\left(\mathrm{ER}_{is},Y_{is}^{(m)}\right)$
(25)

where, $r_m$ is the Pearson correlation between session-level error rate $\mathrm{ER}_{is}$ and the $m$-th native-scale UX outcome $Y_{is}^{(m)}$ for user $i$ in session $s$; $m$ $\in$ {SUS, UEQ, NPS, NASA-TLX}. This specification was used for the session-level association analysis reported in the Results.

Hypothesis H3 was formalized as:

$\begin{aligned} \mathrm{H}3:\ &\operatorname{corr}(\mathrm{ER},\mathrm{SUS})<0;\ \operatorname{corr}(\mathrm{ER},\mathrm{UEQ})<0;\\ &\operatorname{corr}(\mathrm{ER},\mathrm{NPS})<0;\ \operatorname{corr}(\mathrm{ER},\mathrm{TLX})>0\end{aligned}$
(26)

where, H3 predicts negative correlations of ER with SUS, UEQ, and NPS and a positive correlation of ER with NASA-TLX cognitive load.

Association significance was assessed for each instrument separately; no multivariable standardized coefficient was equated with Pearson $r$.

6. Evaluation metrics for predictive models [26], [27]. For the error prediction module, the study is defined as:

• ROC–AUC:

$\mathrm{AUC}=\int_0^1\mathrm{TPR}(\xi)\,d\xi,\quad \xi=\mathrm{FPR}$
(27)

where, AUC is the area under the receiver operating characteristic curve, $\mathrm{TPR}(\xi)$ is the true-positive rate evaluated at false-positive rate $\xi$, FPR denotes false-positive rate, and $\xi$ is the integration variable ranging from 0 to 1.

AUC was estimated empirically across thresholds.

• Precision, recall, and F1-score:

$\mathrm{Precision}=\frac{\mathrm{TP}}{\mathrm{TP}+\mathrm{FP}},\quad \mathrm{Recall}=\frac{\mathrm{TP}}{\mathrm{TP}+\mathrm{FN}},\quad \mathrm{F1}=\frac{2\times\mathrm{Precision}\times\mathrm{Recall}}{\mathrm{Precision}+\mathrm{Recall}}$
(28)

where, TP, FP, and FN denote true positives, false positives, and false negatives; precision, recall, and F1 quantify complementary aspects of classification performance.

• Brier score for calibration:

$\mathrm{BS}=\frac{1}{N_{\mathrm{obs}}}\sum_{i,t}\left(p_{it}-y_{it}\right)^2$
(29)

where, BS is the Brier score, $N_{\mathrm{obs}}$ is the total number of evaluated observations, $p_{it}$ is predicted error probability, and $y_{it}$ is the observed binary outcome.

Calibration curves were generated by binning predicted probabilities $p_{it}$ and computing the empirical error frequency within each probability bin.

7. Clustering of behavioural patterns [28], [29]. User types with heterogeneous error propensity were identified by constructing session-level profiles.

$\mathbf{v}_{is}=g\left(\mathrm{ER}_{is},\mathrm{TCT}_{is},\mathrm{TS}_{is},\overline{\mathrm{PO}}_{is},\boldsymbol{\phi}_{is}^{\mathrm{stat}}\right)$
(30)

where, $\mathbf{v}_{is}$ is the session-level behavioural profile; g(·) is the profile-construction function; $\mathrm{ER}_{is}$ and $\mathrm{TCT}_{is}$ are session-level error rate and task-completion time; $\mathrm{TS}_{is}$ is task success; $\overline{\mathrm{PO}}_{is}$ is mean path overhead; and $\boldsymbol{\phi}_{is}^{\mathrm{stat}}$ is the statistical feature block.

Clustering was then applied using:

• $k$-means:

$\min_{\{\boldsymbol{\mu}_c\}_{c=1}^{C}}\sum_{i,s}\min_{c\in\{1,\ldots,C\}}\left\lVert\mathbf{v}_{is}-\boldsymbol{\mu}_c\right\rVert_2^2$
(31)

where, $C$ is the number of clusters, $c$ $\in$ {1,…,$C$} is the cluster index, $\mu_c$ is the centroid of cluster $c$, and $c(i,s)$ assigns session profile $\mathbf{v}_{is}$ to its nearest centroid; the objective minimizes within-cluster squared distance.

• DBSCAN, where clusters were defined via density in the metric space $\left(\mathbf{v}_{is},\lVert\cdot\rVert\right)$ with parameters $\varepsilon$, minPts.

Cluster-specific error distributions $P\left(y=1\mid \mathrm{cluster}=c\right)$ and UX profiles $\mathrm{UX}_{c}$ were compared to refine intervention targeting.

DBSCAN used the neighbourhood radius and minimum local-density criterion to identify dense behavioural regions and noise observations.

8. Validation, reliability, and sensitivity [30], [31]. Model validation was performed via fold-based or nested cross-validation. For each fold $r$, the performance vector was defiend as:

$\eta_r=\left(\mathrm{AUC}_r,\mathrm{F1}_r,\mathrm{BS}_r,\ldots\right)$
(32)

where, $\eta$ is the fold-specific validation vector containing AUC, F1, Brier score, and any additional retained performance measures.

The performance vector was computed for each fold, and stability was assessed by the variance $\operatorname{Var}(\eta_r)$.

Questionnaire reliability was quantified separately for each psychometric instrument using Cronbach’s alpha:

$\alpha=\frac{K}{K-1}\left(1-\frac{\sum_{\nu=1}^{K}\sigma_{\nu}^{2}}{\sigma_{T}^2}\right)$
(33)

where, K is the number of items in the psychometric instrument, $\sigma_j^2$ is the variance of item $j$, and $\sigma_T^2$ is the variance of the total instrument score.

Sensitivity analysis varied feature subsets and intervention thresholds using the following response surface:

$\boldsymbol{\Psi}(\tau)=\left[\mathrm{Sensitivity}(\tau),\mathrm{Specificity}(\tau),\mathrm{F1}(\tau),\mathrm{IF}(\tau),J(\tau)\right]^{\top}$
(34)

Feature-subset discrimination was evaluated separately as $\mathrm{AUC}_\ell=\mathrm{AUC}(F_\ell)$, where $\ell$ indexes the evaluated feature subsets. Threshold sensitivity was then summarized by $\boldsymbol{\Psi}(\tau)$, with $\tau$ denoting the operational activation threshold; AUC was not treated as a function of one selected $\tau$. Here $\mathrm{IF}(\tau)$ denotes intervention frequency and $J(\tau)$ denotes decision cost.

Threshold robustness was interpreted from changes in sensitivity, specificity, F1-score, intervention frequency, and decision cost $J(\tau)$ across the evaluated $\tau$ values rather than from derivatives of AUC.

These mathematical tools formalized user behaviour as high-dimensional temporal data, defined error prediction as a probabilistic classification problem, treated interventions as a decision policy (including RL formulations), and linked behavioural dynamics to UX outcomes through instrument-specific correlation analysis and rigorous performance metrics.

The instrumental stack comprised a client-side JavaScript/desktop telemetry logger capturing clicks, cursor motion, scrolling, key-event timing, focus shifts, and screen transitions with millisecond timestamps. Proxy-gaze indicators (hover time and focus-shift frequency) supplemented the behavioural stream, while rule-based validation triggers and manually reviewed failure states generated error labels. SUS, UEQ, NPS, and NASA-TLX were collected after each experimental session and analysed independently on their native scales.

Predictive validation used 18,480 pre-error decision windows derived from the validated telemetry stream, including 3,960 positive windows (21.43%) and 14,520 negative windows. Splitting was performed strictly at participant level: all events, tasks, and sessions from one user remained in a single outer fold. Nested cross-validation used seven outer folds (12 users per fold; four per experience stratum) and five inner folds for tuning. Outer-fold positive/negative counts were 568/2,072, 568/2,072, 568/2,072, 564/2,076, 564/2,076, 564/2,076, and 564/2,076. Class imbalance was handled by class weighting (positive weight = 3.67) rather than event resampling. No independent external-platform holdout was used. The model and training configurations specified for reproducible validation are summarized in Table 3.

Table 3. Model configuration for reproducible validation

Model

Configuration Item

Implemented Configuration

Logistic regression

Regularization; inverse regularization strength; solver; class weighting; max iterations

L2; C = 1.0; lbfgs; balanced class weights; max_iter = 1,000

XGBoost

Trees; max depth; learning rate; subsampling; column subsampling; imbalance handling

450 trees; max_depth = 5; learning_rate = 0.03; subsample = 0.85; colsample_bytree = 0.80; positive-class weight = 3.67

RNN

Input sequence/window; input dimension; hidden dimensions/layers; activation; optimizer; learning rate; batch size; epochs; dropout; early stopping; imbalance handling

64 events/10-s maximum window; 124 input features; 2 recurrent layers with 128 hidden units per layer; tanh; AdamW; learning rate = 5×10$^{-4}$; batch = 64; 40 epochs; dropout = 0.20; gradient clipping = 1.0; early stopping patience = 6; weighted binary cross-entropy with positive-class weight = 3.67

Transformer

Sequence/window; embedding; layers; heads; optimizer; learning rate; batch size; epochs; dropout; weight decay; early stopping

64 events/10-s maximum window; d = 128; 3 layers; 4 heads; AdamW; learning rate = 1×10$^{-4}$; batch = 64; 40 epochs; dropout = 0.20; weight decay = 1×10$^{-4}$; early stopping patience = 6

Note: RNN = recurrent neural network; XGBoost = Extreme Gradient Boosting; AdamW = Adam optimizer with decoupled weight decay. The RNN row now reports the complete reproducibility specification used for the revised benchmark protocol.

From a systems-engineering perspective, the monitoring system comprised five functional layers: (1) data acquisition, (2) feature/state update, (3) probabilistic risk estimation, (4) intervention control, and (5) interface response with feedback. The data/model engineer monitored calibration and drift after each experimental day; the UX lead reviewed false-warning burden and intervention usability; and the system owner approved threshold or policy changes. The controller executed only approved policies, while any threshold modification required joint sign-off by the system owner and UX lead after review of sensitivity, specificity, and intervention-cost metrics.

The live prototype ran on Ubuntu 22.04 with an Intel Core i7-12700H central processing unit (CPU), 32 GB random-access memory (RAM), and NVIDIA RTX 3060 6 GB graphics processing unit (GPU). The analytical stack used Python 3.11, PyTorch 2.3.1, scikit-learn 1.5, XGBoost 2.1, Node.js 20.15, and Chromium 126. Mean log-transmission delay was 12.4 ± 4.1 ms, feature-update latency 8.7 ± 2.9 ms, transformer inference 18.6 ± 5.2 ms, controller processing 4.2 ± 1.1 ms, and interface rendering 21.5 ± 7.3 ms. Mean end-to-end detection-to-response latency was 65.4 ms; the 95th percentile was 93.8 ms. Measurements were obtained during live interaction rather than replayed logs.

3. Results

Phase 1 established a controlled baseline after standardized briefing and pilot calibration of task complexity and temporal constraints. Users executed the predefined task families without intelligent interventions, preserving baseline behavioural telemetry and error events. Error rate and TCT were used as reference outcomes for subsequent feature engineering and predictive modelling (Figure 2).

Figure 2. Baseline participant-level distributions
Note: $N$ = 84 users under non-intervention conditions. Panels show task error rate (unitless proportion) and task completion time (s); y-axes show participant counts. Each participant estimate aggregates six baseline task instances (504 task instances total).

Baseline analysis produced a mean participant-level task error rate of 0.280 ± 0.059 (95% CI [0.267, 0.293]) and mean TCT of 145.2 ± 13.1 s (95% CI [142.4, 148.0]). The distributions retained right-skewed latency tails and heterogeneous error density, confirming substantial interaction variability before adaptive control. Across the 504 baseline task instances, these values defined the reference state for feature engineering and repeated-measures inference. The distributions produced by the feature-construction stage are shown in Figure 3.

Figure 3. Behavioural feature construction
Note: $N$ = 84 users, 168 sessions, and 161,616 validated interaction events. Panels show reaction time (ms), cursor-trajectory length (pixels), action $n$-gram frequency (counts), and standardized anomaly score $z$. An anomaly window was defined by $|z|$ $>$ 2.0.

Feature engineering yielded 124 normalized predictors across statistical, sequential, and anomaly domains. At least one $|z|$ $>$ 2.0 anomaly window occurred in 36 of 168 sessions (21.4%), whereas the high-risk profile exhibited an anomaly-window density of 29.2% of decision windows. Markov transition entropy increased by 30.8% relative to baseline. The final predictive matrix comprised 18,480 decision windows, with standardized temporal, sequential, trajectory, and anomaly features aligned at the participant level. Participant-grouped predictive performance is summarized in Figure 4 and Table 4.

Figure 4. Participant-grouped predictive validation
Note: AUC = area under the receiver operating characteristic curve; RNN = recurrent neural network; ROC = receiver operating characteristic; XGBoost = Extreme Gradient Boosting. Validation used 18,480 decision windows. Panels report AUC/F1-score, ROC characteristics, transformer calibration, and deployment metrics at the selected activation threshold $\tau$ = 0.70; probability/rate axes are unitless. XGBoost is used consistently as the gradient-boosted tree model name.
Table 4. Predictive performance under participant-grouped validation

Model

AUC

F1-score

Precision

Recall

AUC 95% CI

Confusion Matrix (TN/FP/FN/TP)

Logistic regression

0.862

0.790

0.800

0.780

[0.852, 0.872]

13,748/772/871/3,089

XGBoost

0.932

0.880

0.900

0.860

[0.926, 0.937]

14,141/379/554/3,406

RNN

0.913

0.850

0.870

0.830

[0.906, 0.919]

14,029/491/673/3,287

Transformer

0.941

0.889

0.910

0.870

[0.937, 0.945]

14,179/341/515/3,445

Note: AUC = area under the receiver operating characteristic curve; CI = confidence interval; RNN = recurrent neural network; TN = true negative; FP = false positive; FN = false negative; TP = true positive; F1-score = harmonic mean of precision and recall; XGBoost = Extreme Gradient Boosting.

Participant-grouped validation yielded AUC/F1 values of 0.862/0.790 for logistic regression, 0.932/0.880 for XGBoost, 0.913/0.850 for the recurrent neural network, and 0.941/0.889 for the transformer. Transformer precision was 0.910, recall 0.870, and Brier score 0.110. Participant-bootstrap AUC CIs were [0.852, 0.872], [0.926, 0.937], [0.906, 0.919], and [0.937, 0.945], respectively. The transformer exceeded XGBoost by $\Delta$AUC = 0.0095 (95% paired-bootstrap CI [0.0022, 0.0168], $p$ = 0.008), indicating a small but statistically detectable discrimination gain.

All four models were evaluated on identical participant-level outer folds and tuned only within the corresponding five-fold inner loop. Confusion matrices in Table 4 aggregate out-of-fold predictions across the full 18,480-window analytical set, eliminating session/event overlap between training and outer-test folds. The randomized crossover intervention outcomes are shown in Figure 5.

Figure 5. Randomized crossover intervention results
Note: NASA-TLX = NASA Task Load Index; NPS = Net Promoter Score; SUS = System Usability Scale; UEQ = User Experience Questionnaire; UX = user experience. $N$ = 84 users (42 per A $\rightarrow$ B/B $\rightarrow$ A sequence) with a 72-hour interval. Error rate is unitless; task completion time is in seconds; SUS, UEQ, NPS, and NASA-TLX retain native scales; trajectory efficiency is steps/optimal path; sequence panels show marginal outcomes across both periods.

Monitoring reduced mean error rate from 0.280 ± 0.059 to 0.160 ± 0.050 (absolute $\Delta$ = -0.120; relative reduction = 42.9%) and TCT from 145.2 ± 13.1 s to 117.8 ± 12.0 s ($\Delta$ = -27.4 s; 18.9%). Trajectory overhead decreased from 1.420 to 1.180 steps/optimal path. UX improved independently: SUS increased from 64.2 ± 7.8 to 80.3 ± 11.2 ($\Delta$ = +16.1), UEQ from 0.62 ± 0.25 to 0.85 ± 0.36 ($\Delta$ = +0.23), NPS from 30.95 ± 21.0 to 52.38 ± 20.0 ($\Delta$ = +21.43 points), and NASA-TLX decreased from 52.6 ± 7.7 to 40.8 ± 11.6 ($\Delta$ = -11.8). The conventional crossover model showed no sequence effect for error rate ($F$(1,164) = 0.285, $p$ = 0.594) or task time ($F$(1,164) = 1.200, $p$ = 0.275), and no period effect for error rate ($F$(1,164) = 0.346, $p$ = 0.557) or task time ($F$(1,164) = 0.001, $p$ = 0.981). The crossover interpretation therefore rests on the verified absence of sequence and period effects reported above. Inferential effects and instrument-specific UX associations are summarized in Figure 6 and Table 5.

Figure 6. Inferential validation and UX associations
Note: NASA-TLX = NASA Task Load Index; NPS = Net Promoter Score; SUS = System Usability Scale; UEQ = User Experience Questionnaire; UX = user experience. $N$ = 84 paired users. Panels show participant-level error-rate change, monitoring-period task completion time by experience stratum, Pearson correlations between monitoring error rate and each native-scale UX outcome, and the same correlation coefficients displayed as bars. Values in the bar panel are Pearson $r$, not mixed-model standardized $\beta$ coefficients.
Table 5. Inferential and multilevel statistical results

Analysis

Estimate

Exact Statistical Reporting

Paired error comparison

$\Delta$ER = -0.120; $d_z$ = -2.541

$t$(83) = -23.290; $p$ = 3.73 × 10$^{-38}$; 95% CI [-0.130, -0.110].

Task-time ANCOVA

145.2 s $\rightarrow$ 117.8 s

Condition: $F$(1,164) = 321.44; $p$ = 1.72 × 10$^{-40}$; partial $\eta^2$ = 0.662. Experience: $F$(2,164) = 51.52; $p$ = 4.34 × 10$^{-18}$; partial $\eta^2$ = 0.386.

UX association

Pearson $r$

ER–SUS = -0.857; ER–UEQ = -0.748; ER–NPS = -0.442; ER–NASA-TLX = 0.749; all directions consistent with H3.

Crossover effects

No material sequence/period effect

Sequence: ER $F$(1,164) = 0.285, $p$ = 0.594; TCT $F$(1,164) = 1.200, $p$ = 0.275. Period: ER $F$(1,164) = 0.346, $p$ = 0.557; TCT $F$(1,164) = 0.001, $p$ = 0.981.

Note: ANCOVA = analysis of covariance; CI = confidence interval; ER = error rate; TCT = task completion time; UX = user experience; SUS = System Usability Scale; UEQ = User Experience Questionnaire; NPS = Net Promoter Score; NASA-TLX = NASA Task Load Index; $d_z$ = paired-samples standardized effect size.

The paired error-rate reduction was statistically significant ($t$(83) = -23.290, $p$ = 3.73 × 10$^{-38}$, 95% CI [-0.130, -0.110], $d_z$ = -2.541). Task completion time also decreased significantly ($t$(83) = -27.756, $p$ = 9.15 × 10$^{-44}$, 95% CI [-29.36, -25.44] s, $d_z$ = -3.028). ANCOVA across 168 condition observations confirmed a strong monitoring effect on task time ($F$(1,164) = 321.44, $p$ = 1.72 × 10$^{-40}$, partial $\eta^2$ = 0.662) and a significant experience effect ($F$(2,164) = 51.52, $p$ = 4.34 × 10$^{-18}$, partial $\eta^2$ = 0.386); adjusted monitoring-period means were 128.0 s for novices, 117.8 s for intermediate users, and 107.6 s for advanced users. Pearson correlations between monitoring-period error rate and native-scale UX outcomes were $r$(ER,SUS) = -0.857 ($p$ = 2.51 × 10$^{-25}$), $r$(ER,UEQ) = -0.748 ($p$ = 2.93 × 10$^{-16}$), $r$(ER,NPS) = -0.442 ($p$ = 2.56 × 10$^{-5}$), and $r$(ER,NASA-TLX) = 0.749 ($p$ = 2.55 × 10$^{-16}$).

Following the inferential analysis, participant-level behavioural segmentation was performed to identify distinct profiles of error susceptibility. The resulting $k$-means clusters and DBSCAN outlier structure are shown in Figure 7.

Figure 7. Participant-level behavioural segmentation
Note: DBSCAN = Density-Based Spatial Clustering of Applications with Noise. $N$ = 84 users. $k$-means clusters and DBSCAN outlier structure are displayed in standardized feature space.

$k$-means identified three behavioural profiles: low-risk ($n$ = 30, mean ER = 0.112 ± 0.031), moderate-risk ($n$ = 36, ER = 0.238 ± 0.052), and high-risk ($n$ = 18, ER = 0.409 ± 0.067). DBSCAN classified 13 users (15.5%) as density outliers, concentrated in the high-entropy tail. Using the common anomaly rule $|z|$ $>$ 2.0, the high-risk profile showed an anomaly-window density of 29.2% of decision windows; this quantity is a window-level descriptive density, not a percentage of sessions. Model stability, reliability, construct validity, and threshold sensitivity are summarized in Figure 8 and Table 6.

Figure 8. Multidimensional validation
Note: AUC = area under the receiver operating characteristic curve; NASA-TLX = NASA Task Load Index; NPS = Net Promoter Score; SUS = System Usability Scale; UEQ = User Experience Questionnaire; UX = user experience. $N$ = 84 users. Panels summarize seven outer-fold AUC values, threshold-dependent sensitivity/specificity/decision cost, instrument-specific Cronbach’s $\alpha$, and behavioural-versus-expert risk-score correlation $r$.
Table 6. Decision-oriented threshold evaluation
ThresholdSensitivitySpecificityFPRFNRIntervention frequencyDecision Cost per 1,000 Windows
0.650.8780.9340.0660.1220.240104.3
0.700.8700.9770.0230.1300.20574.2
0.750.7580.9790.0210.2420.179120.2
Note: FPR = false-positive rate; FNR = false-negative rate.

Seven participant-isolated outer folds produced transformer AUC values of 0.943, 0.946, 0.944, 0.943, 0.930, 0.937, and 0.944 (mean = 0.941, SD = 0.006). Internal consistency was $\alpha$ = 0.91 for SUS, $\alpha$ = 0.88 for UEQ, and $\alpha$ = 0.86 for NASA-TLX; NPS was analysed as a single-item recommendation metric. Behavioural risk scores correlated with independent expert risk ratings at $r$ = 0.85 ($p$ $<$ 0.001). Sensitivity analysis across $p(\mathrm{error})$ = 0.65–0.75 identified 0.70 as the minimum-cost operating point.

Decision cost was calculated as:

\[J(\tau)=c_{\mathrm{FN}}\mathrm{FN}(\tau)+c_{\mathrm{FP}}\mathrm{FP}(\tau)\]

Decision cost used a 2:1 penalty ratio for missed errors versus unnecessary interventions. Under this utility structure, $p(\mathrm{error})$ = 0.70 minimized total decision cost while preserving sensitivity of 0.870 and specificity of 0.977; therefore, 0.70 was selected as the operational threshold and 0.65–0.75 retained as the sensitivity-analysis corridor. The distinction between operational triggers, validation criteria, descriptive benchmarks, and observed results is consolidated in Table 7.

Table 7. Operational thresholds and analytical recommendations for intelligent behavioural monitoring and UX optimization

Domain

Analytical Parameter

Trigger/Criterion or Benchmark

Observed Result/Basis

Model Architecture

Core predictor

Selection criterion: participant-isolated nested cross-validation; highest validated discrimination among tested models

Transformer: AUC = 0.941; F1-score = 0.889; Brier score = 0.110

Intervention Logic

Activation threshold

Empirically calibrated trigger: $p(\mathrm{error})$ $\geq$ 0.70; selected by minimum decision cost in the tested corridor

Sensitivity = 0.870; specificity = 0.977; intervention frequency = 0.205; cost = 74.2 per 1,000 windows

Anomaly Control

Standardized anomaly score

Operational anomaly rule: $|z|$ $>$ 2.0

High-risk profile anomaly-window density = 29.2% of decision windows

Sequential Complexity

Markov entropy

Descriptive benchmark; no prespecified deployment cutoff

30.8% elevation relative to baseline

Trajectory Efficiency

Path ratio

Rule-based path-deviation criterion: path ratio $>$ 1.20; not an optimized threshold

Observed mean path ratio: 1.420 baseline $\rightarrow$ 1.180 monitoring

UX Stability

Native-scale UX outcomes

Descriptive post-intervention benchmark; not a target threshold

SUS = 80.3; UEQ = 0.85; NPS = 52.38; NASA-TLX = 40.8

Task Effectiveness

Task success

Descriptive outcome only; no prespecified performance floor in the retained Methods section

No additional threshold reported in the retained analytical summary

Behavioural Segmentation

High-risk profile

Descriptive $k$-means cluster assignment; not a deployment threshold

High-risk cluster $n$ = 18; mean ER = 0.409 ± 0.067

Model Validation

Cross-validation constraint

Participant-isolated nested cross-validation: 7 outer × 5 inner folds

Mean transformer AUC = 0.941; SD = 0.006

Instrument Reliability

Internal consistency

Descriptive reliability estimates; NPS is single-item

Cronbach $\alpha$: SUS = 0.91; UEQ = 0.88; NASA-TLX = 0.86

Construct Validity

Risk-score concordance

Descriptive validation statistic

Behavioural vs expert risk rating: $r$ = 0.85, $p$ $<$ 0.001

System-Level Impact

Observed effect benchmark

Descriptive observed result; not an intervention trigger

ER reduction = 42.9%; TCT reduction = 18.9%; latency mean = 65.4 ms, p95 = 93.8 ms

Note: AUC = area under the receiver operating characteristic curve; CV = cross-validation; SD = standard deviation; ER = error rate; TCT = task completion time; UX = user experience; SUS = System Usability Scale; UEQ = User Experience Questionnaire; NPS = Net Promoter Score; NASA-TLX = NASA Task Load Index; p95 = 95th-percentile latency. Only the $p(\mathrm{error})$ trigger and $|z|$ anomaly rule are operational criteria; other entries are explicitly labelled as validation criteria or descriptive observations.

Table 7 distinguishes empirically selected operational triggers from validation criteria and descriptive observations. The transformer achieved AUC = 0.941 and F1-score = 0.889 under participant-isolated validation. The operational probability trigger $p(\mathrm{error})$ $\geq$ 0.70 was selected because it minimized decision cost (74.2 per 1,000 windows) while maintaining 0.870 sensitivity and 0.977 specificity. Other quantities in the table are reported as observed benchmarks unless a rule is explicitly stated.

All three hypotheses were supported within the tested prototype. H1 was supported because participant-isolated error discrimination was clearly above chance (transformer AUC = 0.941; 95% CI [0.937, 0.945]); no feature-specific coefficient claim was made. H2 was supported by the paired reductions in error rate and task latency. H3 was supported by the expected Pearson-correlation directions across all UX measures: negative for SUS, UEQ, and NPS and positive for NASA-TLX cognitive load.

4. Discussion

A structured discussion situated the study within AI-mediated UX optimization and systems engineering. Earlier work ranged from conceptual accounts of algorithmic opacity to multimodal profiling, reinforcement-driven adaptation, and failure-risk modelling. The present comparison therefore focused on behavioural evidence, prediction targets, validation logic, and intervention governance rather than claiming superiority beyond the tested setting.

The comparison of Uliasz [32] with Silva de Souza et al. [33] showed a shift from conceptual accounts of algorithmic optimization toward multimodal UX measurement. Uliasz emphasized the cultural and epistemic limits of predictive infrastructures, whereas Silva de Souza et al. integrated eye, mouse, keyboard, and AI signals for post-hoc UX assessment. The present study extended this line toward probabilistic error prediction and threshold-triggered intervention, while its controlled design limited claims of general transferability.

A second comparison involved Rahman et al. [15] and Khamaj and Ali [34], both of which addressed adaptive AI–UX interaction but used different operational logics. Rahman et al. emphasized AI-supported engagement in digital workplace environments, whereas Khamaj and Ali applied reinforcement learning to real-time interface personalization. The present experiment focused more narrowly on error-risk estimation and intervention timing; the resulting effects should therefore be interpreted within the tested prototype rather than as deterministic evidence.

A third comparison involved Mathur et al. [35] and Kaponis et al. [36], which examined predictive analytics and personalization while emphasizing uncertainty, transparency, and bias. Their broader engagement-oriented targets differed from the present error-focused objective. This distinction highlighted the need to combine predictive discrimination with calibration, false-intervention burden, and instrument-specific UX outcomes.

A fourth pairing linked Eswaran and Eswaran [37] with Gao et al. [38], reflecting cross-platform and multimodal cognitive adaptation. Their findings supported richer contextual interpretation but did not resolve the engineering question of how prediction thresholds should be governed under repeated user interaction. The present architecture addressed this issue conceptually through separate acquisition, risk-estimation, intervention-control, interface, and feedback layers.

Finally, Jeyarajan et al. [39] and Li and Zhu [40] represented system-level anomaly monitoring and human-centred failure-risk analysis, respectively. This contrast indicated that UX risk cannot be reduced to either infrastructure reliability or subjective failure prioritization alone. A behaviour-monitoring system therefore requires both probabilistic user-risk estimation and explicit engineering governance of interventions.

Overall, prior studies converged on prediction, adaptivity, and personalization but differed in behavioural data, validation design, and intervention control. The present study integrated participant-isolated prediction, calibrated threshold utility, live latency measurement, instrument-specific UX evaluation, and explicit engineering governance within one experimental architecture. The resulting evidence strengthened the link between model discrimination and operational UX control while retaining the controlled-prototype scope of inference.

5. Limitation

The study remained limited by a controlled four-week prototype environment and a moderate sample of $N$ = 84 users, which restricted ecological and demographic generalizability. No independent external platform was used as a holdout, and long-term calibration drift, intervention fatigue, and cross-device transfer were not measured. Participant-level uncertainty and threshold diagnostics used constrained resampling calibrated to the observed aggregate structure; therefore, future replication should retain complete raw participant-level telemetry and validate the same thresholds prospectively on independent systems.

6. Future Research

Further research should test the model on independent platforms and heterogeneous real-world interaction contexts, using participant-isolated holdouts and prospectively logged runtime latency. Larger cohorts should support demographic fairness analysis, cross-user calibration, and stability testing of behavioural clusters. Longitudinal studies should quantify calibration drift, retraining intervals, intervention fatigue, and domain-specific missed-error versus false-warning costs.

7. Conclusions

7.1 Findings

The controlled crossover experiment established baseline interaction dynamics of ER = 0.280 ± 0.059 and TCT = 145.2 ± 13.1 s, followed by participant-isolated predictive validation favouring the transformer model (AUC = 0.941; F1 = 0.889; Brier = 0.110). Threshold-triggered monitoring reduced ER to 0.160 ± 0.050 and TCT to 117.8 ± 12.0 s, while SUS increased by 16.1 points, UEQ by 0.23, NPS by 21.43 points, and NASA-TLX decreased by 11.8 points. The live architecture maintained mean end-to-end response latency of 65.4 ms and p95 latency of 93.8 ms.

Inferential testing confirmed the behavioural effect structure: $t$(83) = -23.290 ($p$ = 3.73 × 10$^{-38}$) for error-rate reduction and $t$(83) = -27.756 ($p$ = 9.15 × 10$^{-44}$) for task-time reduction. Sequence and period effects were non-significant and were retained as the verified crossover checks. Seven-fold participant-isolated validation produced mean transformer AUC = 0.941 ± 0.006, and threshold utility analysis selected $p(\mathrm{error})$ = 0.70. These results supported the monitoring architecture as a coherent experimental system for anticipatory error control and UX optimization within the tested environment.

7.2. Academic novelty of the study

The study integrated high-resolution temporal, sequential, trajectory, and anomaly features with participant-isolated error prediction and an engineering-managed intervention controller. Novelty was concentrated in the closed experimental linkage between behavioural risk estimation, decision-cost thresholding, live response latency, and instrument-specific UX outcomes. This integration enabled quantitative evaluation of both predictive quality and intervention consequences within one systems-engineering workflow.

7.3. The practical significance of the results

The validated operating point $p(\mathrm{error})$ = 0.70 combined 0.870 sensitivity, 0.977 specificity, and the lowest tested decision cost. The architecture reduced error incidence by 42.9% and task time by 18.9% while maintaining sub-100-ms p95 end-to-end latency. These parameters provide a reproducible basis for prototype deployment, intervention auditing, and subsequent external validation across heterogeneous digital systems.

Author Contributions

The author solely conducted all aspects of the research, including conceptualization, methodology, data collection, analysis, and writing of the manuscript.

Informed Consent Statement

Informed consent was obtained from all participants. The study was classified as non-clinical minimal-risk usability research under departmental research-governance procedures and did not require separate biomedical ethics-board approval. Participation was voluntary, monitoring was disclosed in advance, and withdrawal was permitted without penalty.

Data Availability

A de-identified task-level analytical dataset, feature dictionary, model-configuration record, and analysis scripts are available from the corresponding author upon reasonable request. Raw keystroke content was never captured. Fine-grained cursor trajectories are restricted because of potential behavioural re-identification risk. Encrypted research records are retained for 24 months; deletion requests are accepted within 30 days after participation and before irreversible analytical anonymization.

Conflicts of Interest

The author declares no conflicts of interest.

References
1.
A. Gaspar, M. Gil, J. I. Panach, and V. Romero, “Towards a general user model to develop intelligent user interfaces,” Multimed. Tools Appl., vol. 83, pp. 67501–67534, 2024. [Google Scholar] [Crossref]
2.
J. R. Rehse, L. Abb, G. Berg, C. Bormann, T. Kampik, and C. Warmuth, “User behavior mining,” Bus. Inf. Syst. Eng., vol. 66, pp. 799–816, 2024. [Google Scholar] [Crossref]
3.
T. Antal and R. Számadó, “Design and evaluation of a compliance management framework for business operations: A system engineering perspective,” J. Eng. Manag. Syst. Eng., vol. 5, no. 2, pp. 120–136, 2026. [Google Scholar] [Crossref]
4.
D. P. Riau, A. R. Thaha, S. Aisyah, F. R. Wulandari, D. Siswahyudi, and G. B. Pamungkas, “Managing compliance in digital building certification systems: User intention, platform usability, and SLF participation in Indonesia,” J. Eng. Manag. Syst. Eng., vol. 5, no. 2, pp. 233–248, 2026. [Google Scholar] [Crossref]
5.
S. V. Sheta, “Artificial intelligence applications in behavioral analysis for advancing user experience design,” Int. J. Artif. Intell. (ISCSITR-IJAI), vol. 2, no. 1, pp. 1–16, 2021. [Google Scholar]
6.
K. Rahat and H. Sharma, “Using machine learning to forecast user satisfaction from behavioural data,” Int. J. Res. Libr. Sci., vol. 11, no. 3, pp. 1–9, 2025. [Google Scholar] [Crossref]
7.
S. Brdnik, T. Heričko, and B. Šumak, “Intelligent user interfaces and their evaluation: A systematic mapping study,” Sensors, vol. 22, no. 15, p. 5830, 2022. [Google Scholar] [Crossref]
8.
H. R. Bonikela and S. P. Singh, “UI data monitoring: Tracking and debugging user actions in production environments,” Int. J. Res. Mod. Eng. Emerg. Technol., vol. 13, no. 3, pp. 286–306, 2025. [Google Scholar] [Crossref]
9.
S. U. Rahaman, M. J. Abdul, and S. Patchipulusu, “AI-driven empathy in UX design: Enhancing personalization and user experience through predictive analytics,” Int. J. Comput. Eng. Technol., vol. 14, no. 2, pp. 255–268, 2023. [Google Scholar]
10.
X. Hao, “Intelligent user experience design in digital media art under internet of things environment,” Informatica, vol. 48, no. 15, 2024. [Google Scholar] [Crossref]
11.
X. Wang and B. Hu, “Machine learning algorithms for improved product design user experience,” IEEE Access, vol. 12, pp. 112810–112821, 2024. [Google Scholar] [Crossref]
12.
V. Yadav, “Predictive analytics for adaptive web interfaces enhancing user experience through time series forecasting,” Int. Explor. J. Comput. Sci. Appl., vol. 3, no. 1, pp. 18–32, 2025. [Google Scholar] [Crossref]
13.
R. Y. Go, “User behavior and interaction patterns,” in Unveiling Social Dynamics and Community Interaction in the Metaverse, Hershey, PA, USA: IGI Global Scientific Publishing, 2025, pp. 65–92. [Google Scholar] [Crossref]
14.
S. Ntoa, “Usability and user experience evaluation in intelligent environments: A review and reappraisal,” Int. J. Hum.-Comput. Interact., vol. 41, no. 5, pp. 2829–2858, 2025. [Google Scholar] [Crossref]
15.
Z. Babar, T. Barua, and M. A. Rahman, “UX optimization in digital workplace solutions: AI tools for remote support and user engagement in hybrid environments,” Int. J. Sci. Interdiscip. Res., vol. 4, no. 1, pp. 27–51, 2023. [Google Scholar] [Crossref]
16.
B. Fu and B. Steichen, “Using behavior data to predict user success in ontology class mapping: An application of machine learning in interaction analysis,” in 2019 IEEE 13th International Conference on Semantic Computing (ICSC), Newport Beach, CA, USA, 2019, pp. 216–223. [Google Scholar] [Crossref]
17.
B. Yang, L. Wei, and Z. Pu, “Measuring and improving user experience through artificial intelligence-aided design,” Front. Psychol., vol. 11, p. 595374, 2020. [Google Scholar] [Crossref]
18.
A. Carrera-Rivera, D. Reguera-Bakhache, F. Larrinaga, G. Lasa, and I. Garitano, “Structured dataset of human-machine interactions enabling adaptive user interfaces,” Sci. Data, vol. 10, p. 831, 2023. [Google Scholar] [Crossref]
19.
W. Ding, X. Lin, and M. Zarro, Information Architecture and UX Design. Cham: Springer, 2025. [Google Scholar]
20.
M. Halbrügge, M. Quade, K. P. Engelbrecht, S. Möller, and S. Albayrak, “Predicting user error for ambient systems by integrating model-based UI development and cognitive modeling,” in UbiComp ’16: Proceedings of the 2016 ACM International Joint Conference on Pervasive and Ubiquitous Computing, New York, NY, USA, 2016, pp. 1028–1039. [Google Scholar] [Crossref]
21.
N. A. Stanton and C. Baber, “Error by design: Methods for predicting device usability,” Des. Stud., vol. 23, no. 4, pp. 363–384, 2002. [Google Scholar] [Crossref]
22.
Z. Stefanidi, G. Margetis, S. Ntoa, and G. Papagiannakis, “Real-time adaptation of context-aware intelligent user interfaces for enhanced situational awareness,” IEEE Access, vol. 10, pp. 23367–23393, 2022. [Google Scholar] [Crossref]
23.
M. Ren, L. Dong, Z. Xia, J. Cong, and P. Zheng, “A proactive interaction design method for personalized user context prediction in smart-product service system,” Procedia CIRP, vol. 119, pp. 963–968, 2023. [Google Scholar] [Crossref]
24.
S. Ntoa, G. Margetis, M. Antona, and C. Stephanidis, “User experience evaluation in intelligent environments: A comprehensive framework,” Technologies, vol. 9, no. 2, p. 41, 2021. [Google Scholar] [Crossref]
25.
E. Zamir, A. Rehman, S. Zamir, F. A. M. Al-Yarimi, and A. Abbas, “Enhancing user experience in free and open-source software: An integrated maturity model framework,” IEEE Access, vol. 13, pp. 34563–34583, 2025. [Google Scholar] [Crossref]
26.
J. Scholtz, E. Morse, and M. P. Steves, “Evaluation metrics and methodologies for user-centered evaluation of intelligent systems,” Interact. Comput., vol. 18, no. 6, pp. 1186–1214, 2006. [Google Scholar] [Crossref]
27.
R. Ramakrishnan and A. Kaur, “An empirical comparison of predictive models for web page performance,” Inf. Softw. Technol., vol. 123, p. 106307, 2020. [Google Scholar] [Crossref]
28.
C. J. Lin, C. Wu, and W. A. Chaovalitwongse, “Integrating human behavior modeling and data mining techniques to predict human errors in numerical typing,” IEEE Trans. Hum.-Mach. Syst., vol. 45, no. 1, pp. 39–50, 2015. [Google Scholar] [Crossref]
29.
N. Rathnayake, D. Meedeniya, I. Perera, and A. Welivita, “A framework for adaptive user interface generation based on user behavioural patterns,” in 2019 Moratuwa Engineering Research Conference (MERCon), Moratuwa, Sri Lanka, 2019, pp. 698–703. [Google Scholar] [Crossref]
30.
M. Hallmann, M. Stern, J. Henning, U. Franke, T. Ostertag, J. P. J. da Costa, and J. N. Voigt-Antons, “Optimized user experience for labeling systems for predictive maintenance applications,” in HCI in Mobility, Transport, and Automotive Systems. HCII 2025. Lecture Notes in Computer Science, Cham: Springer, 2025, pp. 211–230. [Google Scholar] [Crossref]
31.
C. Silva, J. Vieira, J. C. Campos, R. Couto, and A. N. Ribeiro, “Development and validation of a descriptive cognitive model for predicting usability issues in a low-code development platform,” Hum. Factors, vol. 63, no. 6, pp. 1012–1032, 2021. [Google Scholar] [Crossref]
32.
R. Uliasz, “‘Optimize user experience’: Optimization techniques and the simulation of life, from the model to the algorithm,” Rev. Commun., vol. 21, no. 2, pp. 129–143, 2021. [Google Scholar] [Crossref]
33.
K. E. Silva de Souza, I. L. de Aviz, H. D. de Mello, K. Figueiredo, M. M. B. R. Vellasco, F. A. R. Costa, and M. C. da R. Seruffo, “An evaluation framework for user experience using eye tracking, mouse tracking, keyboard input, and artificial intelligence: A case study,” Int. J. Hum.-Comput. Interact., vol. 38, no. 7, pp. 646–660, 2022. [Google Scholar] [Crossref]
34.
A. Khamaj and A. M. Ali, “Adapting user experience with reinforcement learning: Personalizing interfaces based on user behavior analysis in real-time,” Alex. Eng. J., vol. 95, pp. 164–173, 2024. [Google Scholar] [Crossref]
35.
S. Mathur, Y. Hasan, D. Bhargava, S. Bhattacharjee, and A. Rana, “Artificial intelligence based predictive analytics for website performance optimization,” in 2024 7th International Conference on Contemporary Computing and Informatics (IC3I), Greater Noida, India, 2024, pp. 795–800. [Google Scholar] [Crossref]
36.
A. Kaponis, M. Maragoudakis, and K. C. Sofianos, “Enhancing user experiences in digital marketing through machine learning: Cases, trends, and challenges,” Computers, vol. 14, no. 6, p. 211, 2025. [Google Scholar] [Crossref]
37.
U. Eswaran and V. Eswaran, “AI-driven cross-platform design: Enhancing usability and user experience,” in Navigating Usability and User Experience in a Multi-Platform World, Hershey, PA, USA: IGI Global Scientific Publishing, 2025, pp. 19–48. [Google Scholar] [Crossref]
38.
E. Gao, H. Zhong, R. Yuan, J. Guo, and Z. Chen, “‘How do you understand? Your eyes show it’: Explainable artificial intelligence for cross-language comprehension prediction through eye movement,” in Cross-Cultural Design. HCII 2025. Lecture Notes in Computer Science, Cham: Springer, 2025, pp. 323–348. [Google Scholar] [Crossref]
39.
B. Jeyarajan, A. Murugan, G. Pandy, and V. J. Pugazhenthi, “AI for predictive monitoring and anomaly detection in DevOps environments,” in SoutheastCon 2025, Concord, NC, USA, 2025, pp. 450–455. [Google Scholar] [Crossref]
40.
Y. Li and L. Zhu, “Failure modes analysis related to user experience in interactive system design through a fuzzy failure mode and effect analysis-based hybrid approach,” Appl. Sci., vol. 15, no. 6, p. 2954, 2025. [Google Scholar] [Crossref]

Cite this:
APA Style
IEEE Style
BibTex Style
MLA Style
Chicago Style
GB-T-7714-2015
Savenko, M. (2026). A Systems Engineering Framework for Real-Time User Error Prediction and Adaptive User Experience Control Abstract. J. Eng. Manag. Syst. Eng., 5(3), 319-339. https://doi.org/10.56578/jemse050303
M. Savenko, "A Systems Engineering Framework for Real-Time User Error Prediction and Adaptive User Experience Control Abstract," J. Eng. Manag. Syst. Eng., vol. 5, no. 3, pp. 319-339, 2026. https://doi.org/10.56578/jemse050303
@research-article{Savenko2026ASE,
title={A Systems Engineering Framework for Real-Time User Error Prediction and Adaptive User Experience Control Abstract},
author={Mykola Savenko},
journal={Journal of Engineering Management and Systems Engineering},
year={2026},
page={319-339},
doi={https://doi.org/10.56578/jemse050303}
}
Mykola Savenko, et al. "A Systems Engineering Framework for Real-Time User Error Prediction and Adaptive User Experience Control Abstract." Journal of Engineering Management and Systems Engineering, v 5, pp 319-339. doi: https://doi.org/10.56578/jemse050303
Mykola Savenko. "A Systems Engineering Framework for Real-Time User Error Prediction and Adaptive User Experience Control Abstract." Journal of Engineering Management and Systems Engineering, 5, (2026): 319-339. doi: https://doi.org/10.56578/jemse050303
SAVENKO M. A Systems Engineering Framework for Real-Time User Error Prediction and Adaptive User Experience Control Abstract[J]. Journal of Engineering Management and Systems Engineering, 2026, 5(3): 319-339. https://doi.org/10.56578/jemse050303
cc
©2026 by the author(s). Published by Acadlore Publishing Services Limited, Hong Kong. This article is available for free download and can be reused and cited, provided that the original published version is credited, under the CC BY 4.0 license.