Javascript is required
1.
P. Guarda and S. Qian, “Statistical inference of travelers’ route choice preferences with system-level data,” Transp. Res. B Methodol., vol. 179, p. 102853, 2024. [Google Scholar] [Crossref]
2.
C. Advani, A. Bhaskar, and M. M. Haque, “Bi-level clustering of vehicle trajectories for path choice set and its nested structure identification,” Transp. Res. C Emerg. Technol., vol. 144, p. 103895, 2022. [Google Scholar] [Crossref]
3.
D. McFadden, “Conditional logit analysis of qualitative choice behavior,” in Frontiers in Econometrics, Academic Press, 1974, pp. 105–142. [Google Scholar]
4.
M. Ben-Akiva and S. R. Lerman, Discrete Choice Analysis: Theory and Application to Travel Demand. MIT Press, 1985. [Google Scholar]
5.
K. E. Train, Discrete Choice Methods With Simulation (2nd ed.). Cambridge University Press, 2009. [Google Scholar]
6.
J. D. Ortúzar and L. G. Willumsen, Modelling Transport (4th ed.). John Wiley & Sons, 2011. [Google Scholar]
7.
E. Cascetta, Transportation Systems Analysis: Models and Applications (2nd ed.). Springer, 2009. [Online]. Available: [Google Scholar] [Crossref]
8.
C. G. Prato, “Route choice modeling: Past, present and future research directions,” J. Choice Model., vol. 2, no. 1, pp. 65–100, 2009. [Google Scholar] [Crossref]
9.
D. McFadden and K. Train, “Mixed MNL models for discrete response,” J. Appl. Econom., vol. 15, no. 5, pp. 447–470, 2000, <447::AID-JAE570>3.0.CO;2-1. [Google Scholar] [Crossref]
10.
D. A. Hensher, J. M. Rose, and W. H. Greene, Applied Choice Analysis. Cambridge University Press, 2015. [Online]. Available: [Google Scholar] [Crossref]
11.
L. Cazor, L. C. Duncan, D. P. Watling, O. A. Nielsen, and T. K. Rasmussen, “A closed-form bounded route choice model accounting for heteroscedasticity, overlap, and choice set formation,” Transp. Res. B Methodol., vol. 199, p. 103275, 2025. [Google Scholar] [Crossref]
12.
T. K. Rasmussen, L. C. Duncan, D. P. Watling, and O. A. Nielsen, “Local detouredness: A new phenomenon for modelling route choice and traffic assignment,” Transp. Res. B Methodol., vol. 190, p. 103052, 2024. [Google Scholar] [Crossref]
13.
S. Bekhor, M. E. Ben-Akiva, and M. S. Ramming, “Evaluation of choice set generation algorithms for route choice models,” Ann. Oper. Res., vol. 144, no. 1, pp. 235–247, 2006. [Google Scholar] [Crossref]
14.
E. Frejinger, M. Bierlaire, and M. Ben-Akiva, “Sampling of alternatives for route choice modeling,” Transp. Res. B Methodol., vol. 43, no. 10, pp. 984–994, 2009. [Google Scholar] [Crossref]
15.
D. Liu, D. Li, K. Gao, Y. Song, and T. Zhang, “Enhancing choice-set generation and route choice modeling with data- and knowledge-driven approach,” Transp. Res. C Emerg. Technol., vol. 162, p. 104618, 2024. [Google Scholar] [Crossref]
16.
L. Cazor, L. C. Duncan, D. P. Watling, O. A. Nielsen, and T. K. Rasmussen, “A smooth bounded choice model: Formulation and application in three large-scale case studies,” J. Choice Model., vol. 57, p. 100574, 2025. [Google Scholar] [Crossref]
17.
L. Breiman, “Random forests,” Mach. Learn., vol. 45, no. 1, pp. 5–32, 2001. [Google Scholar] [Crossref]
18.
J. H. Friedman, “Greedy function approximation: A gradient boosting machine,” Ann. Stat., vol. 29, no. 5, pp. 1189–1232, 2001. [Google Scholar] [Crossref]
19.
T. Chen and C. Guestrin, “XGBoost: A scalable tree boosting system,” in Proceedings of the 22nd ACM SIGKDD International Conference on Knowledge Discovery and Data Mining. San Francisco, CA, USA, pp. 785–794, 2016. [Google Scholar] [Crossref]
20.
G. Ke, Q. Meng, T. Finley, T. Wang, W. Chen, W. Ma, Q. Ye, and T. Y. Liu, “LightGBM: A highly efficient gradient boosting decision tree,” in Advances in Neural Information Processing Systems 30. pp. 3146–3154, 2017. [Google Scholar]
21.
L. Cheng, X. Chen, J. De Vos, X. Lai, and F. Witlox, “Applying a random forest method approach to model travel mode choice behavior,” Travel Behav. Soc., vol. 14, pp. 1–10, 2019. [Google Scholar] [Crossref]
22.
Y. Zhang and A. Haghani, “A gradient boosting method to improve travel time prediction,” Transp. Res. C Emerg. Technol., vol. 58, pp. 308–324, 2015. [Google Scholar] [Crossref]
23.
Z. Zhao and Y. Liang, “A deep inverse reinforcement learning approach to route choice modeling with context-dependent rewards,” Transp. Res. C Emerg. Technol., vol. 149, p. 104079, 2023. [Google Scholar] [Crossref]
24.
S. Qiu, G. Qin, M. Wong, and J. Sun, “RoutesFormer: A sequence-based route choice transformer for efficient path inference from sparse trajectories,” Transp. Res. C Emerg. Technol., vol. 162, p. 104552, 2024. [Google Scholar] [Crossref]
25.
H. Wang, E. Moylan, and D. Levinson, “Ensemble methods for route choice,” Transp. Res. C Emerg. Technol., vol. 167, p. 104803, 2024. [Google Scholar] [Crossref]
26.
J. Arriagada, C. A. Guevara, M. A. Munizaga, and S. Gao, “An experiential learning-based transit route choice model using large-scale smart-card data,” Transportation, vol. 52, no. 4, pp. 1543–1568, 2025. [Google Scholar] [Crossref]
27.
Y. Han, F. C. Pereira, M. Ben-Akiva, and C. Zegras, “A neural-embedded discrete choice model: Learning taste representation with strengthened interpretability,” Transp. Res. B Methodol., vol. 163, pp. 166–186, 2022. [Google Scholar] [Crossref]
28.
B. Sifringer, V. Lurkin, and A. Alahi, “Enhancing discrete choice models with representation learning,” Transp. Res. B Methodol., vol. 140, pp. 236–261, 2020. [Google Scholar] [Crossref]
29.
S. Wang, B. Mo, and J. Zhao, “Theory-based residual neural networks: A synergy of discrete choice models and deep neural networks,” Transp. Res. B Methodol., vol. 146, pp. 333–358, 2021. [Google Scholar] [Crossref]
30.
S. van Cranenburgh, S. Wang, A. Vij, F. C. Pereira, and J. Walker, “Choice modelling in the age of machine learning—Discussion paper,” J. Choice Model., vol. 42, p. 100340, 2022. [Google Scholar] [Crossref]
31.
S. B. Ayaz, H. Tian, S. Gao, and D. L. Fisher, “Proactive route choice with real-time information: Learning and effects of network complexity and cognitive load,” Transp. Res. C Emerg. Technol., vol. 149, p. 104035, 2023. [Google Scholar] [Crossref]
32.
B. Zhou and R. Liu, “A generalized rationally inattentive route choice model with non-uniform marginal information costs,” Transp. Res. B Methodol., vol. 189, p. 102993, 2024. [Google Scholar] [Crossref]
33.
F. Ahmad and L. Al-Fagih, “Travel behaviour and game theory: A review of route choice modeling behaviour,” J. Choice Model., vol. 50, p. 100472, 2024. [Google Scholar] [Crossref]
34.
Y. Tian, W. Zhu, and F. Song, “Route choice modelling for an urban rail transit network: Past, recent progress and future prospects,” Eur. Transp. Res. Rev., vol. 16, p. 52, 2024. [Google Scholar] [Crossref]
35.
People’s Government of Guangdong Province, “Approval on adjusting the toll charging method for toll roads: Yue Fu Han [2019] No. 416,” 2019, [Online]. Available: https://www.gd.gov.cn/zwgk/gongbao/2019/36/content/post_3366602.html [Google Scholar]
36.
T. Akiba, S. Sano, T. Yanase, T. Ohta, and M. Koyama, “Optuna: A next-generation hyperparameter optimization framework,” in Proceedings of the 25th ACM SIGKDD International Conference on Knowledge Discovery & Data Mining, Anchorage, AK, USA, 2019, pp. 2623–2631. [Google Scholar] [Crossref]
37.
K. A. Small, E. T. Verhoef, and R. Lindsey, The Economics of Urban Transportation. Routledge, 2007. [Online]. Available: [Google Scholar] [Crossref]
38.
H. Yang and H. J. H. J. Huang, Mathematical and Economic Theory of Road Pricing. Elsevier, 2005. [Google Scholar]
39.
D. Brownstone and K. A. Small, “Valuing time and reliability: Assessing the evidence from road pricing demonstrations,” Transp. Res. A Policy Pract., vol. 39, no. 4, pp. 279–293, 2005. [Google Scholar] [Crossref]
40.
D. A. Hensher, “Measurement of the valuation of travel time savings,” J. Transp. Econ. Policy, vol. 35, no. 1, pp. 71–98, 2001. [Google Scholar]
41.
G. de Jong, M. Kouwenhoven, J. Bates, P. Koster, E. Verhoef, L. Tavasszy, and P. Warffemius, “New SP-values of time and reliability for freight transport in the Netherlands,” Transp. Res. E Logist. Transp. Rev., vol. 64, pp. 71–87, 2014. [Google Scholar] [Crossref]
Search
Open Access
Research article

Data-Driven Route Choice Prediction for Expressway Operations Using Ensemble Learning: Sensitivity Analysis of Differentiated Toll Policies

Yuejiang Su1,
Liang Zhang2,
Minxian Yuan1,
Dexin Wu1,
Weiwei Qi3,
Yunlong Cao4,
Liuhua Zhang5,
Nanfeng Zhang6*
1
Guangzhou Transport Research Institute Co., Ltd., 510635 Guangzhou, China
2
School of Artificial Intelligence, Xidian University, 710068 Xi’an, China
3
School of Civil Engineering and Transportation, South China University of Technology, 510641 Guangzhou, China
4
Guangzhou Communications Investment Group Co., Ltd., 510635 Guangzhou, China
5
China Electronic Product Reliability and Environment Test Research Institute, 511300 Guangzhou, China
6
Guangdong Provincial Key Laboratory of Intelligent Port Security Inspection, 510700 Guangzhou, China
Journal of Industrial Intelligence
|
Volume 3, Issue 4, 2025
|
Pages 226-245
Received: 09-23-2025,
Revised: 11-08-2025,
Accepted: 11-23-2025,
Available online: 11-30-2025
View Full Article|Download PDF

Abstract:

Intelligent expressway management requires reliable prediction of travellers' route choices and quantitative assessment of their responses to operational interventions. This study investigates a data-driven route choice framework that links large-scale toll transaction data with ensemble learning to support differentiated tolling and route management. Actual vehicle routes were reconstructed from toll transaction records, electronic toll collection (ETC) gantry records, and expressway network topology, from which revealed-preference candidate route sets were derived. A total of 28 features were constructed to characterize route attributes, relative differences among candidate routes, choice-set composition, and trip attributes. A Light Gradient Boosting Machine (LightGBM) model was then developed using each “trip $\times$ candidate route” pair as the unit of observation, and its predictive performance was compared with random forest, Extreme Gradient Boosting (XGBoost), multinomial logit (MNL), nested logit, and binary logistic regression. The empirical analysis used 329,699 trips from the Guangdong expressway network, including 32,473 trips in the test set. The Light Gradient Boosting Machine (LightGBM) model achieved a route-level hit rate of 85.71%, exceeding the random-guessing benchmark by 43.81 percentage points and outperforming the three conventional discrete-choice and logistic models, while showing performance comparable to random forest and XGBoost. The sensitivity analysis further showed that toll responses varied substantially with a route’s competitive position, choice-set size, and vehicle type. Routes with baseline choice probabilities of 50%–80% were more responsive to toll changes than strongly dominant routes, while trucks showed greater toll sensitivity than cars. The results demonstrate that large-scale operational data and ensemble learning can provide a reproducible framework for route choice prediction and scenario-based evaluation of differentiated toll policies. The proposed framework supports data-driven decision-making for intelligent expressway operations and provides a practical basis for evaluating route-level traffic management strategies.

Keywords: Intelligent expressway operations, Route choice prediction, Ensemble learning, Light Gradient Boosting Machine, Toll transaction data, Differentiated tolling, Data-driven decision-making

1. Introduction

As expressway networks become increasingly dense and interconnected, travellers are increasingly confronted with multiple feasible routes between the same origin–destination (OD) pair. The routes selected by individual travellers determine how traffic demand is distributed across parallel corridors and, consequently, influence the effectiveness of operational measures such as differentiated tolling, congestion management, and route guidance. Reliable route choice prediction is therefore an important component of data-driven expressway operation and intelligent traffic management. Conventional route choice studies have largely relied on household travel surveys, stated-preference experiments, Global Positioning System (GPS) trajectories, or other individual-level observations. Although these data sources provide valuable information on travel behaviour, their coverage can be constrained by sample size, network extent, vehicle classes, and operating conditions. The increasing availability of network-scale operational data provides a different empirical basis for route choice analysis. System-level observations have been used to infer travellers' route choice preferences over large transportation networks [1], while data-driven clustering of vehicle trajectories has demonstrated how observed paths can be used to identify representative route alternatives and their underlying choice structures [2]. For expressway systems equipped with electronic toll collection (ETC) infrastructure, the integration of toll transactions, gantry passage sequences, and network topology makes it possible to reconstruct realised vehicle routes at large scale, creating a direct observational basis for intelligent and data-driven traffic management.

Route choice modelling has traditionally been grounded in random utility theory. Conditional logit and multinomial logit (MNL) models established the foundation for disaggregate choice analysis [3], which was subsequently developed into a systematic framework for modelling individual travel decisions [4], while simulation-based estimation methods further extended the range of choice structures that can be estimated in practice [5]. These models retain an important advantage in behavioural interpretation because estimated parameters can be related to travel-time valuation, cost sensitivity, substitution among alternatives, and other quantities with direct behavioural meaning. They can also be incorporated into traffic assignment and wider transportation-system analysis [6], [7]. Route choice, however, presents several difficulties that are less pronounced in many other discrete-choice settings. Candidate routes frequently overlap over substantial portions of a network, so their unobserved utilities cannot always be regarded as independent. This limitation has motivated a range of alternative route-choice formulations and corrections designed to represent correlation and similarity among alternatives [8], [9], [10]. Recent studies show that this problem remains methodologically important. Cazor et al. [11] developed a bounded route choice formulation that jointly addresses heteroscedasticity, route overlap, and implicit choice-set formation, while Rasmussen et al. [12] demonstrated that both local and global detours can materially affect route generation and route-choice probabilities. These developments indicate that route similarity, detour structure, and the competitive relationship among alternatives remain central to contemporary route-choice modelling.

Choice-set construction constitutes a closely related challenge. A route-choice model cannot reproduce observed behaviour reliably if the set of alternatives presented to the model does not reasonably represent the routes available or perceived by travellers. Classical approaches have therefore relied on enumeration, sampling, and dedicated path-generation algorithms to construct candidate sets [13], [14]. Such approaches remain useful, but their performance may depend strongly on network size, parameter settings, and assumptions about which alternatives are behaviourally relevant. The growth of large trajectory datasets has encouraged more data-driven treatment of this problem. Liu et al. [15] integrated data- and knowledge-driven methods for choice-set generation and route choice modelling, using a conditional variational autoencoder to represent the choice-set generation process together with a learning-based choice model. More recently, Cazor et al. [16] proposed a smooth bounded choice formulation in which consideration-set formation is incorporated implicitly within the choice process and evaluated the approach using several large-scale applications, including route-choice cases. These studies reflect an important methodological shift: the set of behaviourally relevant routes is increasingly treated as part of the modelling problem rather than merely as an external preprocessing step.

A second major research stream approaches route choice primarily as a predictive learning problem. Random forest [17], gradient boosting [18], Extreme Gradient Boosting (XGBoost) [19], and LightGBM [20] can represent nonlinear relationships, threshold effects, and higher-order interactions without requiring the complete form of a utility function to be specified in advance. Their computational scalability also makes them attractive for large operational datasets containing large numbers of trips and route alternatives. Machine-learning methods have already demonstrated substantial predictive capability in transportation applications such as travel-mode classification [21] and travel-time prediction [22]. Their use in route choice has expanded as richer trajectory and network data have become available. Zhao and Liang [23] applied deep inverse reinforcement learning to recover context-dependent routing preferences from observed trajectories, providing a framework in which route-selection behaviour can be learned from sequential decisions rather than represented entirely through a predefined utility function. Qiu et al. [24] developed RoutesFormer, a sequence-based Transformer that integrates path inference and route-choice representation and uses attention mechanisms to capture complex sequential dependencies in sparse trajectory data. These developments demonstrate that route-choice regularities can increasingly be extracted directly from high-dimensional observations.

Of particular relevance to the present study is the recent application of ensemble and deep-learning methods to route choice prediction. Wang et al. [25] systematically evaluated ensemble methods for network-wide route choice and showed that the relative performance of individual models can vary across training and testing datasets, while appropriately constructed ensembles can provide robust prediction of aggregate route flows. Their results are important for operational applications because they suggest that model usefulness should not be assessed solely by identifying a universally superior algorithm; instead, predictive stability across heterogeneous observations and network conditions is equally important. Recent studies have also examined learning-based route choice using large-scale observed travel data, showing that travellers' accumulated experience and previous choices can provide additional information for modelling route-selection behaviour [26]. Together, these recent studies indicate that route choice research is moving beyond a simple comparison between conventional discrete-choice models and machine-learning algorithms. The emerging question is how flexible predictive models can exploit large observational datasets while still producing results that remain meaningful for transport operations and behavioural analysis.

Predictive flexibility, however, does not eliminate the need for interpretation. A machine-learning model may provide high predictive accuracy without directly yielding behavioural quantities such as marginal effects, elasticities, welfare measures, or explicit substitution patterns. A parallel research stream has therefore attempted to combine flexible computational structures with interpretable choice mechanisms. Neural-embedded discrete choice models have been developed to learn heterogeneous taste representations while retaining a utility-based structure [27]. Representation-learning methods have likewise been incorporated into conventional discrete-choice models to capture nonlinear effects that are difficult to specify a priori [28]. Theory-based residual neural networks provide another approach by retaining a behavioural model as the structural component and allowing a neural network to learn residual patterns not captured by the theoretical specification [29]. Van Cranenburgh et al. [30] provide a broader discussion of this convergence between theory-driven choice modelling and data-driven machine learning, highlighting both the opportunities created by flexible predictive methods and the importance of maintaining behavioural validity. These developments are directly relevant to operational route-choice applications, where predictive performance alone is insufficient if the model cannot be related to controllable variables or interpreted in a way that supports management decisions.

Route choice is also shaped by the information available to travellers and by the broader decision environment. Real-time traffic information can alter perceived route attractiveness, affect expectations regarding downstream conditions, and change the way travellers evaluate future diversion opportunities. Ayaz et al. [31] examined proactive route choice under real-time information and showed that learning, network complexity, and cognitive load can influence route-selection behaviour. Zhou and Liu [32] approached the problem from the perspective of rational inattention, explicitly modelling heterogeneous information-processing costs and deriving route-choice probabilities when travellers do not process all available information equally. From a broader behavioural perspective, Ahmad and Al-Fagih [33] reviewed game-theoretic approaches to route choice and highlighted their relevance to congestion, tolling, transportation policy, and strategic interactions among travellers. Tian et al. [34], in a recent review of route-choice modelling, likewise showed that contemporary research increasingly relies on richer transport data, more flexible behavioural structures, and advanced computational approaches, while continuing to face challenges related to data integration, model calibration, validation, and representation of traveller heterogeneity. Taken together, this literature suggests that modern route-choice analysis increasingly lies at the intersection of behavioural modelling, machine learning, network data, and operational decision support.

Despite these advances, several gaps remain between route choice prediction and its application to intelligent expressway operations. First, conventional utility-based models require predefined functional structures and may have difficulty representing complex nonlinearities, interactions, and threshold effects among route attributes when very large operational datasets are available. Second, although recent machine-learning, deep-learning, and ensemble approaches have improved predictive flexibility, much of the literature still concentrates primarily on model fit or predictive accuracy rather than explicitly connecting fitted models to operational variables that network operators can adjust. Third, prediction errors themselves are seldom examined systematically. A high aggregate accuracy does not reveal which route-choice situations are difficult to predict, whether errors are concentrated among minority alternatives, or where the practical boundary of a fitted model lies. Fourth, large-scale operational observations are increasingly available, but relatively few studies integrate revealed route sets, absolute and relative route attributes, ensemble prediction, error diagnosis, and management-oriented scenario analysis within a single framework. These limitations are particularly important for differentiated tolling, where the operational value of a predictive model depends not only on identifying the route most likely to be chosen but also on estimating how predicted route probabilities change when toll relationships among competing alternatives are modified.

Accordingly, this study develops a data-driven route choice prediction framework for expressway operations using large-scale toll transaction data and ensemble learning. Following the analytical sequence of “toll transaction data → choice-set and feature construction → route choice probability prediction → influencing-factor analysis → policy sensitivity”, a LightGBM model is developed with each “trip $\times$ candidate route” pair treated as an observation. In addition to route-specific attributes, the feature system represents the relative position of each route within its corresponding choice set, allowing the model to distinguish between the absolute characteristics of an alternative and its competitive position relative to other feasible routes. Predictive performance is evaluated using the route-level hit rate and compared with alternative machine-learning and conventional choice models. Key influencing factors are examined through statistical association tests and gain-based feature importance, while mis-predicted trips are analysed to identify the conditions under which observed route choices are difficult to distinguish. The fitted model is subsequently used for scenario-based sensitivity analysis by varying tolls, a directly adjustable operational variable, and recalculating the affected route characteristics and choice probabilities. In this way, the study integrates large-scale operational data, ensemble learning, route-relative feature representation, model interpretation, error diagnosis, and differentiated-toll scenario evaluation within a unified analytical framework. The resulting approach provides a reproducible basis for route choice prediction and supports data-driven decision-making for differentiated tolling, route guidance, and intelligent expressway operations. The overall research framework is illustrated in Figure 1.

Figure 1. Research framework

2. Methodological Framework

2.1 Route Choice Formulation

The modelling objective is to estimate the probability that each candidate route is selected for a given trip. Rather than treating a physical route as having a fixed propensity to be chosen, the proposed formulation evaluates the route relative to the alternatives available for a specific trip. The same physical route may therefore constitute a positive observation in one trip and a negative observation in another, depending on its attributes and competitive position within the corresponding choice set.

For each trip, candidate routes were identified from revealed travel behaviour. Routes accounting for no less than 2% of the observed trips for the corresponding OD pair were retained to form the revealed choice set, with no fewer than two and no more than five candidate routes for each OD pair. Each modelling observation represents a “trip $\times$ candidate route” pair. Its feature vector is expressed as:

$x_{ij} = (a_{ij}, r_{ij}, s_i, t_i)$
(1)

where, $x_{ij}$ is the feature vector of candidate route $j$ for trip $i$; $a_{ij}$ contains the route-specific attributes, including mileage, toll, travel time, number of switches, and operating speed; $r_{ij}$ describes the relative position of the candidate route within the choice set, including ratios and differences relative to the best alternative, dominance indicators, and margins relative to the second-best alternative; $s_i$ represents choice-set characteristics, including the number of alternatives and their dispersion; and $t_i$ contains trip-level attributes such as vehicle type, departure period, and peak-hour status.

This formulation allows the model to learn not only the absolute characteristics of a route but also its competitive position among the alternatives available to the traveller. Once the model has been trained, an operational variable such as toll can be changed for a target route, the affected relative features can be recalculated, and the resulting choice probabilities can be predicted under the modified scenario. This provides the basis for the policy sensitivity analysis conducted later in the study.

2.2 Ensemble-Learning-Based Route Choice Model

Route choice probability was estimated using LightGBM, a gradient-boosting decision-tree algorithm. Gradient boosting constructs an ensemble sequentially, with each new decision tree fitted to reduce the prediction error remaining from the preceding iterations. LightGBM further employs histogram-based feature discretisation and a leaf-wise tree-growth strategy [20], making it computationally suitable for large datasets while retaining the ability to represent nonlinear relationships and interactions among predictors.

For the feature vector defined in Eq. (1), the predicted probability that candidate route $j$ is chosen for trip $i$ is expressed as:

$\hat{p}_{ij} = f(x_{ij}) = \sum_{n=1}^{N} \beta_n h_n(x_{ij})$
(2)

where, $\hat{p}_{ij}$ denotes the predicted choice probability of candidate route $j$ for trip $i$; $f(\ )$ is a fitted gradient-boosting model consisting of $N$ decision trees; $\beta_n$ denotes the contribution of the $n$-th tree; and $h_{n}(\ )$ is its output.

The use of route-specific, relative, choice-set, and trip-level features enables the model to represent nonlinear differences among competing alternatives without requiring a predefined utility-function form. This property is particularly relevant to the present application because the attractiveness of an expressway route depends not only on its absolute mileage, toll, or travel time but also on how these attributes compare with those of the other routes available for the same trip.

2.3 Model Evaluation

Predictive performance was evaluated at the trip level. For each trip in the test set, the candidate route with the highest predicted probability was identified as the model-predicted route and compared with the route actually taken. The route-level hit rate is defined as:

$H = \frac{1}{M}\sum_{i=1}^{M}{II}\left(\hat{j}_i=j_i\right)$
(3)

where, $H$ is the route-level hit rate; $M$ is the number of trips in the test set; $\hat{j}_i$ denotes the candidate route with the highest predicted probability for trip $i$; $j_i$ denotes the route actually chosen; and $II(\ )$ is the indicator function, which equals 1 when the predicted and observed routes coincide and 0 otherwise.

This trip-level measure directly evaluates whether the model identifies the route actually selected from the available alternatives and therefore provides an interpretable basis for comparing predictive performance across different model specifications.

3. Data and Feature Construction

3.1 Data Sources, Preprocessing, and Sample Construction

The empirical analysis used one day of toll-station transaction records, ETC gantry records, and network topology data from the Guangdong expressway network. Toll transaction records contained information on entry and exit stations, vehicle type, and toll amount, while ETC gantry records provided the sequence of gantries passed by each vehicle. By integrating these two data sources with the directed expressway network, the actual travel route of each vehicle was reconstructed. Route mileage, toll, peak/off-peak travel time, and the number of switches between expressway segments were subsequently derived. Toll values were converted from the standard class-1 vehicle charge according to the vehicle-type coefficients specified in the Classification of Vehicle Types and Toll Coefficients for Guangdong Expressways (Yue Fu Han [2019] No. 416) [35].

The raw dataset contained 1,282,059 transaction records covering 11,760 OD pairs. Data preprocessing included duplicate removal, field validation, spatiotemporal consistency checks, and gantry-sequence stitching to address missing or inconsistent observations. After preprocessing, 1,255,584 records covering 11,722 OD pairs were retained, corresponding to a removal rate of 2.06%.

Revealed choice sets were then constructed from observed route use. For each OD pair, routes accounting for at least 2% of observed trips were retained, and OD pairs with fewer than two candidate routes were excluded because they did not represent a route choice problem. No more than five candidate routes were retained for an OD pair. This procedure yielded 5,596 OD pairs, 329,699 trips, and 858,681 candidate-route-level observations. The substantial reduction in the number of OD pairs primarily resulted from the fact that 52.3% of OD pairs exhibited only one observed travel route and therefore contained no observed route-choice variation.

The modelling sample was randomly divided by trip into training and test sets at a ratio of 9:1. The training set contained 297,226 trips and 774,256 candidate-route-level observations, while the test set contained 32,473 trips and 84,425 candidate-route-level observations. Table 1 summarises the sample construction and data split. The distribution of choice-set size is presented in Figure 2. Among the 5,596 OD pairs retained for modelling, 3,216 had two candidate routes, 1,512 had three, 591 had four, and 277 had five. These groups accounted for 60.34%, 24.69%, 9.14%, and 5.82% of the modelling trips, respectively, with an average choice-set size of 2.60 routes.

Table 1. Sample construction and data split

Stage

Number

Description

Raw toll transaction records

1,282,059

Raw entry–exit–gantry records

Records after cleaning

1,255,584

Duplicates, missing values, and spatio-temporal conflicts removed (2.06% dropped)

Modelling trips

329,699

5,596 OD pairs with at least 2 candidates

Candidate-route-level records

858,681

Positive cases (chosen): 329,699, 38.40%

Training set

297,226 trips/774,256 records

Split by trip, 90%

Test set

32,473 trips/84,425 records

Split by trip, 10%

Figure 2. Distribution of choice-set size
Note: OD = origin–destination.
3.2 Sample Characteristics

The distribution of the 329,699 modelling trips was examined across departure period, toll level, trip mileage, travel time, vehicle type, and number of expressway-segment switches. Table 2 summarises these characteristics and provides an overview of the operating conditions represented in the dataset.

Table 2. Trip distribution characteristics of the modelling sample

Dimension

Group

Trips

Share/%

Dimension

Group

Trips

Share/%

Departure period

0:00–6:00 (small hours)

28854

8.75

Toll/CNY

$<$10

10771

3.27

Departure period

7:00–8:00 (morning peak)

42572

12.91

Toll/CNY

10–20

74863

22.71

Departure period

9:00–11:00 (morning off-peak)

65975

20.01

Toll/CNY

20–30

81094

24.60

Departure period

12:00–16:00 (afternoon off-peak)

92199

27.96

Toll/CNY

30–50

81878

24.83

Departure period

17:00–18:00 (evening peak)

44469

13.49

Toll/CNY

50–80

43347

13.15

Departure period

19:00–23:00 (night)

55630

16.87

Toll/CNY

$\geq$80

37746

11.45

Trip mileage/km

$<$10

117

0.04

Travel time/min

$<$15

11405

3.46

Trip mileage/km

10–20

22242

6.75

Travel time/min

15–25

72527

22.00

Trip mileage/km

20–30

57966

17.58

Travel time/min

25–40

125481

38.06

Trip mileage/km

30–50

119830

36.35

Travel time/min

40–60

80861

24.53

Trip mileage/km

50–80

93295

28.30

Travel time/min

60–90

33879

10.28

Trip mileage/km

$\geq$80

36249

10.99

Travel time/min

$\geq$90

5546

1.68

Car/truck

Car

247294

75.01

Number of switches

1

37473

11.37

Car/truck

Truck

82405

24.99

Number of switches

2

152414

46.23

Vehicle-type breakdown

Car class 1

243636

73.90

Number of switches

3

90500

27.45

Vehicle-type breakdown

Truck class 1

47579

14.43

Number of switches

$\geq$4

49312

14.96

Note: CNY = Chinese Yuan.

Medium- and long-distance movements constituted the majority of the modelling sample. Trips of 30–80 km accounted for 64.65% of all observations, while trips with tolls of 20–50 CNY accounted for 49.43%, and those with travel times of 15–60 min represented 84.59%. In addition, 88.64% of trips involved at least two switches between expressway segments, indicating that most observations represented journeys involving multiple expressway sections and meaningful alternatives within the network.

The dataset also covered different temporal and vehicle operating conditions. Afternoon off-peak trips accounted for the largest share (27.96%), while the morning and evening peak periods together represented 26.40%. Night and early-morning trips accounted for a further 25.62%. Cars represented 75.01% of all trips, including 73.90% classified as class-1 cars, whereas trucks accounted for 24.99%. The truck subsample contained 82,405 trips, providing a substantial observational basis for the vehicle-type comparisons reported later in the analysis.

3.3 Route Choice Feature Construction and Statistical Association

The feature system was designed to represent both the absolute characteristics of each candidate route and its competitive position within the corresponding choice set. A total of 28 features were constructed and organised into four broad components: route-specific attributes, relative characteristics of candidate routes, choice-set characteristics, and trip-level attributes. The underlying physical dimensions included mileage, toll, travel time, and switching between expressway segments. Table 3 reports the complete feature definitions.

Table 3. Route choice feature vector and chi-square test results (28 features)
CodeDimensionFeatureDefinition$p$-value
C1Route's own attributesRoute mileageMileage of the candidate route, km$<0.001$
C2Route's own attributesRoute tollToll of the candidate route after conversion for the vehicle type, Chinese Yuan (CNY)$<0.001$
C3Route's own attributesRoute travel timePeak/off-peak travel time of the candidate route, min$<0.001$
C4Route's own attributesNumber of switchesNumber of switches between expressway segments along the candidate route$<0.001$
C5Relative to the bestDetour ratioCandidate mileage $\div$ shortest mileage of this trip ($\geq1$)$<0.001$
C6Relative to the bestToll ratioCandidate toll $\div$ lowest toll of this trip ($\geq1$)$<0.001$
C7Relative to the bestTime ratioCandidate travel time $\div$ shortest travel time of this trip ($\geq1$)$<0.001$
C8Relative to the bestSwitch differenceCandidate switches $-$ minimum switches of this trip ($\geq0$)$<0.001$
C9Absolute differenceMileage differenceCandidate mileage $-$ shortest mileage of this trip, km$<0.001$
C10Absolute differenceToll differenceCandidate toll $-$ lowest toll of this trip, CNY$<0.001$
C11Absolute differenceTime differenceCandidate travel time $-$ shortest travel time of this trip, min$<0.001$
C12Dominance indicatorShortest mileage1 if the candidate has the shortest mileage in this trip$<0.001$
C13Dominance indicatorLowest toll1 if the candidate has the lowest toll in this trip$<0.001$
C14Dominance indicatorShortest travel time1 if the candidate has the shortest travel time in this trip$<0.001$
C15Dominance indicatorNo. of dominant dimensionsNumber of dimensions (mileage/toll/travel time/switching) in which the candidate is best (0–4)$<0.001$
C16Relative to the second bestMileage leadFor the best candidate, the relative margin by which it leads the second best ($\geq0$); otherwise the negative of the margin by which it trails the best ($\leq0$)$<0.001$
C17Relative to the second bestToll leadAs above, computed on the toll dimension$<0.001$
C18Relative to the second bestTravel time leadAs above, computed on the travel time dimension$<0.001$
C19Operating levelTravel speedTravel speed of the candidate route, km/h$<0.001$
C20Operating levelSpeed ratioCandidate travel speed $\div$ highest speed of this trip ($\leq1$)$<0.001$
C21Operating levelToll per kmCandidate toll $\div$ mileage, CNY/km$<0.001$
C22Choice setChoice-set sizeNumber of candidate routes in this trip (2–5)$<0.001$
C23Choice setMileage dispersionStandard deviation $\div$ mean of the candidate mileage in this trip$<0.001$
C24Choice setTravel time dispersionStandard deviation $\div$ mean of the candidate travel time in this trip$<0.001$
C25TravellerVehicle-type codeCar classes 1–4 coded 1–4; truck classes 1–6 coded 11–16$<0.001$
C26TravellerCar/truck0 for car, 1 for truck$<0.001$
C27TravellerDeparture hourHour of the entry passage timestamp (0–23)$<0.001$
C28TravellerPeak hour1 for 7:00–9:00 and 17:00–19:00$<0.001$

For relative features, each candidate route was compared with the best alternative within the same trip using ratios, absolute differences, dominance indicators, and margins relative to the second-best alternative. These features enabled the model to capture both the absolute characteristics and relative competitiveness of each route within the traveller's choice set.

The statistical association between each feature and observed route choice was examined using chi-square tests based on the 858,681 candidate-route-level observations. Continuous variables were discretised by deciles, and categories representing less than 1% of the sample were combined with adjacent or residual categories as appropriate. The null hypothesis for each test was that the corresponding feature was independent of whether the candidate route was selected. As shown in Table 3, all 28 features were statistically associated with observed route choice at ($p <\ $0.001). Given the large sample size, these significance tests are interpreted as evidence of statistical association rather than as direct measures of effect magnitude. The relative contribution of individual features to predictive performance is therefore examined separately through the model-based feature-importance analysis presented in Section 4.

4. Model Application and Results

4.1 Route Choice Prediction and Model Performance
4.1.1 Hyperparameter tuning

The hyperparameters of LightGBM were optimised using the Bayesian optimisation framework Optuna [36]. Binary log-loss was adopted as the optimisation objective, and 10% of the training trips were reserved as a validation set during hyperparameter tuning. The resulting optimal configuration is reported in Table 4, with the number of boosting iterations determined as 1,532.

Table 4. Optimal hyperparameters of the Light Gradient Boosting Machine (LightGBM)

Parameter

Value

Parameter

Value

learning_rate

0.0279

feature_fraction

0.6732

num_leaves

128

bagging_fraction

0.8115

max_depth

9

lambda_l1

0.0284

min_data_in_leaf

53

lambda_l2

0.6911

min_gain_to_split

0.3126

Number of iterations

1,532

4.1.2 Route-level prediction performance

The route-level hit rate was calculated for the 32,473 trips in the test set using Eq. (3). The LightGBM model achieved an overall hit rate of 85.71%. For comparison, random selection with equal probability assigned to each candidate route within its corresponding choice set produced an aggregate hit rate of 41.90%. The proposed model therefore exceeded this benchmark by 43.81 percentage points. Prediction performance across different trip and choice-set characteristics is reported in Table 5 and Figure 3.

Table 5. Route-level hit rate by scenario

Scenario

Group

Trips

Route-Level Hit Rate

Scenario

Group

Trips

Route-Level Hit Rate

Overall

All test trips

32473

85.71%

Choice-set size

2 routes

19648

92.24%

Car/truck

Car

24260

85.11%

Choice-set size

3 routes

8022

82.36%

Car/truck

Truck

8213

87.47%

Choice-set size

4 routes

2952

68.67%

Time period

Off-peak

23787

85.59%

Choice-set size

5 routes

1851

58.02%

Time period

Peak (7:00–9:00/17:00–19:00)

8686

86.02%

Trip mileage/km

$<$20

2140

93.79%

Trip mileage/km

20–40

12385

89.62%

Trip mileage/km

40–60

9777

82.00%

Trip mileage/km

$\geq$60

8171

82.11%

Travel time/min

$<$20

4378

92.51%

Travel time/min

20–40

16496

84.87%

Travel time/min

40–60

7859

84.15%

Travel time/min

$\geq$60

3740

84.73%

Figure 3. Route-level hit rate by scenario

The most pronounced variation in prediction performance was associated with choice-set size. The hit rate decreased from 92.24% for trips with two candidate routes to 82.36% for three routes, 68.67% for four routes, and 58.02% for five routes. Although prediction became more difficult as the number of alternatives increased, the hit rate for five-route choice sets remained substantially above the corresponding random benchmark of 20%.

Differences across vehicle types and departure periods were comparatively small. The hit rate was 87.47% for trucks and 85.11% for cars, while peak- and off-peak trips yielded hit rates of 86.02% and 85.59%, respectively. A clearer difference was observed across trip-distance groups.

Trips shorter than 20 km achieved a hit rate of 93.79%, compared with approximately 82% for trips of 40 km or more. These results suggest that the model distinguishes route choices more reliably when the set of competing alternatives is relatively small. As the choice set expands or route characteristics become less differentiated, identifying the observed route becomes more difficult.

4.1.3 Analysis of mis-predicted routes

To examine the conditions under which prediction errors occurred, the 4,641 mis-predicted trips were analysed according to the historical choice share of the route actually taken within the corresponding OD pair. The results are reported in Table 6 and Figure 4. The median historical share of the observed route was 92.59% among correctly predicted trips but only 20.00% among mis-predicted trips. The hit rate increased monotonically with historical route share, from 1.25% for routes with a share below 10% to 99.96% for routes accounting for at least 90% of observed trips. Overall, 66.90% of all mis-predictions involved routes with historical shares below 30%.

Table 6. Route-level hit rate by historical share of the observed route

Observed-Route Share Within OD Pair

Trips

Route-Level Hit Rate

Mis-Predicted Trips

Share of Mis-Predictions/%

$<$10%

1356

1.25%

1339

28.85

10%–20%

1067

8.34%

978

21.07

20%–30%

901

12.54%

788

16.98

30%–50%

2199

44.97%

1210

26.07

50%–70%

3202

92.29%

247

5.32

70%–90%

6664

98.92%

72

1.55

$\geq$90%

17084

99.96%

7

0.15

Note: OD = origin-destination.
Figure 4. Route-level hit rate by historical share of the observed route
Note: OD = origin-destination.

This pattern indicates that prediction errors were concentrated among trips in which the observed route differed from the route most commonly selected for the same OD pair. Such deviations may reflect factors not represented in the current feature set, including individual driving preferences, navigation recommendations, service-area requirements, temporary traffic controls, or other trip-specific circumstances. The results therefore identify an important boundary of the present data-driven model: route choices that deviate substantially from recurrent population-level patterns are more difficult to infer from the available route and trip attributes alone.

Table 7 provides an illustrative mis-predicted case involving two nearly equivalent alternatives. The two routes differed by less than 3.4% in mileage, toll, and travel time: the observed route measured 70.41 km, cost 76.51 CNY, and required 61.56 min, whereas the model's first-ranked route measured 70.56 km, cost 74.05 CNY, and required 62.22 min. Their predicted probabilities were correspondingly close, at 56.27% and 56.43%, a difference of only 0.16 percentage points.

Table 7. Illustrative case of a mis-predicted route

Item

Case (Nearly Equivalent)

OD

Shatian East–Guangfo Xinguangxian

Vehicle type/departure time

Truck/11:00

Choice-set size

2 routes

Actual travel route

Guanfan Expressway–Humen Second Bridge–South Second Ring–West Line Section II–West Line Section I–Guangming Foshan–Fobei North

Model's first choice

Guanfan Expressway–Humen Second Bridge–South Second Ring–Dongxin Expressway–Guangming Guangzhou–Guangming Foshan–Fobei North

Actual: mileage/toll/travel time

70.41 km/76.51 CNY/61.56 min

First choice: mileage/toll/travel time

70.56 km/74.05 CNY/62.22 min

Predicted probability, actual/first choice

56.27%/56.43%

Share of the actual travel route within the OD pair

39.13% (9/23)

Note: OD = origin-destination; CNY = Chinese Yuan.
4.1.4 Comparison with alternative models

To evaluate the predictive performance of LightGBM relative to alternative modelling approaches, the same test set was used to compare it with random forest (RF) [17], XGBoost [19], binary logistic regression, MNL [3], and nested logit (NL) [4]. MNL and NL were implemented as route-level choice models that directly estimated probabilities across the alternatives within each choice set. The results are reported in Table 8 and Figure 5.

Table 8. Comparison of route-level hit rates across models

Model

Route-Level Hit Rate

Improvement Over the Random Benchmark/Percentage Points

LightGBM (this paper)

85.71%

43.81

XGBoost

85.61%

43.71

Random forest (RF)

85.56%

43.66

Multinomial logit (MNL)

78.63%

36.73

Nested logit (NL)

78.59%

36.69

Binary logistic regression

75.97%

34.07

Random-guessing benchmark

41.90%

—

Note: LightGBM = Light Gradient Boosting Machine; XGBoost = Extreme Gradient Boosting; — means not applicable.
Figure 5. Comparison of route-level hit rates across models
Note: LightGBM = Light Gradient Boosting Machine; XGBoost = Extreme Gradient Boosting

The three ensemble-learning models produced very similar route-level hit rates, ranging from 85.56% to 85.71%. LightGBM achieved the highest value at 85.71%, followed closely by XGBoost at 85.61% and random forest at 85.56%. The differences among these three models were therefore marginal. In contrast, all three ensemble models achieved higher hit rates than the conventional choice and logistic models. LightGBM exceeded MNL and NL by approximately 7.1 percentage points and binary logistic regression by approximately 9.7 percentage points.

These results indicate that the main predictive advantage arises from the ability of ensemble-learning models to represent nonlinear relationships and interactions among route and trip characteristics rather than from a substantial performance difference among the individual ensemble algorithms. LightGBM was retained for the subsequent interpretation and sensitivity analyses because it combined competitive predictive performance with computational efficiency and direct calculation of gain-based feature importance.

4.2 Analysis of Feature Importance

For the fitted LightGBM model, feature importance was measured using gain, defined as the cumulative reduction in the model loss attributable to splits involving a given feature. Gain-based importance therefore reflects the contribution of each feature to the model's predictive discrimination rather than a causal effect on travellers' behaviour. Figure 6 presents the 12 features with the highest gain values, while Table 9 and Figure 7 aggregate the results by feature dimension.

Figure 6. Gain-based feature importance (top 12 features)
Table 9. Feature importance aggregated by dimension

Physical Dimension

No. of Features

Gain Share/%

Physical Dimension

No. of Features

Gain Share/%

Mileage

5

65.64

Switching

2

3.43

Toll

6

13.06

Operating level

2

3.25

Travel time

5

5.56

Traveller

4

3.08

Choice set

3

5.44

Multi-dimension dominance

1

0.54

Figure 7. Feature importance aggregated by dimension
Note: LightGBM = Light Gradient Boosting Machine; traveller attr. = traveller attributes; multi-dim. dominance = multi-dimension dominance.

At the individual-feature level, mileage lead accounted for the largest gain share (40.16%), followed by detour ratio (13.60%), mileage difference (9.95%), toll lead (4.36%), toll per kilometre (3.40%), and toll difference (2.30%). Because several related variables represent different aspects of the same underlying route characteristic, individual feature importance alone does not provide a direct comparison among physical dimensions. The feature-level results were therefore aggregated according to their corresponding dimensions.

Mileage-related features accounted for 65.64% of the total gain, followed by toll-related features at 13.06% and travel-time features at 5.56%. Choice-set characteristics accounted for 5.44%, switching for 3.43%, operating-level variables for 3.25%, traveller attributes for 3.08%, and multi-dimension dominance for 0.54%.

The prominence of mileage-related variables indicates that relative distance characteristics played a major role in the model's discrimination among candidate routes. Toll-related features ranked second and are particularly relevant from an operational perspective because toll is directly adjustable by the expressway operator, whereas route mileage is largely fixed by network structure in the short term. Toll was therefore selected as the principal intervention variable for the scenario-based sensitivity analysis. The following analysis examines whether predicted responses to toll changes vary with the competitive position of a route, choice-set size, and vehicle type.

4.3 Scenario-Based Sensitivity Analysis of Differentiated Toll Policies

The sensitivity analysis examined how predicted route choice probabilities changed under hypothetical toll adjustments. Network structure, choice-set composition, and trip attributes were held constant, while the toll of the target route or competing routes was modified according to the specified scenario. All toll-dependent features were then recalculated and the fitted LightGBM model was applied again to obtain the corresponding route choice probabilities.

The resulting changes should be interpreted as model-based scenario responses under otherwise unchanged conditions rather than as causal estimates of travellers’ behavioural responses or network-wide equilibrium effects.

4.3.1 Toll increase for the target route

For each of the 32,473 observed routes in the test set, the toll of the target route was increased independently by 10%, 20%, 30%, and 50%, while the attributes of the other candidate routes and all trip-level variables were held constant. Table 10 reports the median change in predicted choice probability across baseline-probability groups together with selected structural characteristics of the routes in each group.

Table 10. Median response to increases in the toll of the target route, by baseline probability

Baseline Probability

Routes

Share/%

Median Probability

Mean Dominance Count

Shortest-Route Share/%

+10% Toll

+30% Toll

+50% Toll

Probability Decrease at +50% Toll/%

$<$50%

5440

16.8

24.25%

1.30

22.2

+0.02

$-$0.77

$-$3.31

68.1

50%–70%

3204

9.9

61.09%

1.67

51.2

$-$5.99

$-$9.91

$-$16.99

87.0

70%–80%

2342

7.2

75.64%

2.13

68.1

$-$4.86

$-$9.31

$-$14.84

84.7

80%–90%

5195

16.0

86.26%

2.58

84.9

$-$1.58

$-$4.04

$-$6.05

82.8

$\geq$90%

16292

50.2

95.10%

2.92

96.5

$-$0.46

$-$1.15

$-$1.84

85.0

Actual travel routes

32473

100.0

90.06%

2.41

75.7

$-$0.71

$-$1.93

$-$3.33

82.0

Note: Probability changes are in percentage points; the number of dominant dimensions is the number of dimensions (mileage, toll, travel time, and switching, 4 in total) in which the route takes the minimum value.

The response to toll adjustment varied substantially with the initial competitive position of the target route. The largest median reductions occurred among routes with baseline probabilities between 50% and 80%. For the 5,546 routes in this range, representing 17.1% of the test sample, a 30% toll increase reduced the median predicted probability by 9.31–9.91 percentage points, while a 50% increase produced reductions of 14.84–16.99 percentage points.

Routes with baseline probabilities of at least 90% behaved differently. This group contained 16,292 routes, or 50.2% of the test sample, and had a median baseline probability of 95.10%. Among these routes, 96.5% were also the shortest-distance alternatives. A 50% toll increase reduced their median predicted probability by only 1.84 percentage points. Their relatively limited response is consistent with the strong competitive position of these routes across several attributes, which reduces the extent to which a change in toll alone alters their overall position within the choice set.

At the other end of the distribution, the 5,440 routes with baseline probabilities below 50% also exhibited comparatively small median responses. Only 22.2% of these routes were the shortest-distance alternatives, and they were dominant in an average of 1.30 dimensions. A 10% toll increase produced a median change of +0.02 percentage points, while a 50% increase reduced the median probability by 3.31 percentage points. Taken together, these results reveal a non-monotonic relationship between baseline competitive position and model-predicted toll sensitivity, with the strongest responses concentrated among routes occupying an intermediate competitive position.

Choice-set size was also associated with the magnitude of the predicted response. For choice sets containing two, three, four, and five routes, a 50% increase in the target-route toll reduced the median predicted probability by 1.96, 5.91, 8.54, and 8.25 percentage points, respectively. The response therefore generally increased as the number of available alternatives rose from two to four, before declining slightly for five-route choice sets. Within the modelled scenarios, toll adjustments consequently produced larger probability shifts where travellers had a broader set of competing alternatives.

These response patterns across baseline probability bands and choice-set sizes are further illustrated in Figure 8.

(a)
(b)
Figure 8. Median decrease in predicted choice probability following an increase in the target-route toll: (a) by baseline probability; (b) by choice-set size
4.3.2 Toll reduction for competing routes

A second scenario kept the toll of the target route unchanged while reducing the tolls of the other candidate routes within the same choice set. This design allowed the predicted response to an increase in the target-route toll to be compared with the response generated by making competing routes less expensive.

Across all 32,473 observed routes, reductions of 10%, 30%, and 50% in the tolls of competing routes decreased the median predicted probability of the target route by 0.78, 2.98, and 7.55 percentage points, respectively. The corresponding changes produced by increasing the target-route toll were 0.71, 1.93, and 3.33 percentage points.

Across the full range of adjustment magnitudes reported in Table 11, reductions in competing-route tolls produced larger median changes than equivalent increases in the target-route toll. At adjustment levels of 10%, 20%, 30%, and 50%, the respective ratios were 1.10, 1.24, 1.54, and 2.27. The differences were relatively small at 10% and 20% but became more pronounced at 30% and 50%. Within the present model, changing the costs of several competing alternatives therefore produced a larger shift in the target route's predicted probability than changing the target route alone, particularly under larger toll adjustments. The route-level responses under the two toll-adjustment strategies are compared in Figure 9.

Table 11. Median probability change under increases in the target-route toll and reductions in competing-route tolls

Scenario

Item

0

10%

20%

30%

50%

Toll increase of the target route

Median probability change/percentage points

0.00

$-0.71$

$-1.28$

$-1.93$

$-3.33$

Toll reduction of the other routes

Median probability change/percentage points

0.00

$-0.78$

$-1.59$

$-2.98$

$-7.55$

Note: The sample comprises all 32,473 actual travel routes in the test set; the values are the median change in the choice probability of each route (percentage points), measured on the same basis as Table 10; the columns are the toll adjustment magnitudes.
Figure 9. Target-route toll increase versus competing-route toll reduction: route-level comparison

The response also differed by vehicle type, as shown in Table 12.

Table 12. Median probability change by vehicle type (percentage points)

Vehicle type

Routes

Median probability

+10% toll

+30% toll

+50% toll

$-$10% toll

$-$30% toll

$-$50% toll

Decrease at +50% toll/%

Decrease at $-50$% toll/%

Car

24,260

89.64%

$-$0.65

$-$1.67

$-$2.87

$-$0.62

$-$2.55

$-$6.77

80.3

84.7

Truck

8,213

90.97%

$-$0.94

$-$3.07

$-$5.11

$-$1.28

$-$4.27

$-$10.23

87.0

90.9

For trucks, a 50% increase in the target-route toll reduced the median predicted probability by 5.11 percentage points, compared with 2.87 percentage points for cars. Under a 50% reduction in competing-route tolls, the corresponding decreases were 10.23 and 6.77 percentage points. The truck subsample therefore exhibited larger model-predicted responses to toll changes across the examined scenarios. This difference may partly reflect the higher absolute toll exposure of trucks, although the present analysis does not separately identify the behavioural mechanisms underlying the vehicle-type difference.

4.3.3 Operational implications

The scenario results provide several implications for the use of differentiated tolling as a data-driven expressway management instrument.

First, the modelled response was strongest for routes occupying an intermediate competitive position. A 50% toll increase for routes with baseline probabilities of 50%–80% reduced their median predicted probabilities by 14.84–16.99 percentage points, whereas the corresponding reduction was only 1.84 percentage points for routes with baseline probabilities of at least 90%. This suggests that differentiated tolling may produce greater route-level shifts when applied to corridors where competing alternatives already have relatively similar levels of attractiveness.

Second, vehicle type should be considered when evaluating differentiated toll scenarios. Trucks exhibited larger predicted probability changes than cars under both target-route toll increases and competing-route toll reductions. The result indicates that a uniform toll adjustment may generate heterogeneous responses across vehicle classes and that vehicle-specific effects should therefore be assessed when designing or evaluating differentiated toll schemes.

Third, the simulations indicate that reducing the tolls of competing routes can produce larger predicted shifts than increasing the toll of the target route by the same proportion, particularly at larger adjustment magnitudes. For example, at 50%, the median probability reduction was 7.55 percentage points when competing-route tolls were reduced, compared with 3.33 percentage points when the target-route toll was increased. This finding provides a quantitative basis for comparing alternative toll-differentiation strategies before implementation.

Fourth, the magnitude of the predicted response depended on both the size of the toll adjustment and the number of alternatives in the choice set. In the present scenarios, differences between the two toll-adjustment strategies were relatively limited at 10% and 20% but became more pronounced at 30% and 50%. Larger responses were also observed for OD pairs with four or more candidate routes than for those with only two. These findings suggest that adjustment magnitude and the structure of available route alternatives should be considered jointly rather than applying a uniform toll-differentiation rule across the network.

These implications are conditional on the modelled scenarios and should not be interpreted as direct causal estimates or prescriptive toll-setting rules. Their primary value lies in demonstrating how a data-driven route choice model can be used to screen alternative operational strategies and identify route groups for more detailed evaluation before implementation.

Several limitations define the scope of these findings. First, the current feature system does not incorporate dynamic variables such as real-time traffic conditions and weather, while temporal variation is represented primarily through departure time and peak/off-peak conditions. Second, the revealed choice sets are derived from historically observed routes and therefore do not include previously unobserved alternatives that may become relevant under temporary traffic restrictions, roadworks, or other network changes. Third, the sensitivity analysis represents a partial-equilibrium scenario assessment in which network structure, choice-set composition, and travel demand are held constant. The resulting probability changes should therefore not be interpreted as causal estimates of actual behavioural responses or as network-equilibrium effects. More general road-pricing considerations, including second-best pricing, also remain outside the present framework[37], [38].

Future research can address these limitations in several directions. Dynamic traffic states and weather information can be incorporated to improve the temporal representation of route conditions. Toll and travel-time responses can also be placed on a common behavioural scale by drawing on established research on the value of travel time [39], [40], [41]. Causal inference designs based on observed toll changes would provide a stronger basis for identifying behavioural responses to pricing interventions, while integration with a traffic assignment framework [7] would allow route-level probability changes to be propagated through the network and evaluated in terms of system-wide traffic effects. These extensions would move the present framework from scenario-based route choice prediction toward more comprehensive intelligent decision support for dynamic expressway management.

5. Conclusions

This study developed a data-driven framework for route choice prediction and differentiated toll analysis by integrating large-scale expressway transaction data with ensemble learning. Actual travel routes were reconstructed from toll transaction records, ETC gantry records, and network topology, and each “trip $\times$ candidate route” pair was treated as an individual modelling observation. The resulting framework connected route choice prediction, feature interpretation, error diagnosis, and scenario-based toll sensitivity analysis, providing a quantitative basis for data-driven expressway operation and traffic management.

The LightGBM model achieved a route-level hit rate of 85.71% for the 32,473 test trips, exceeding the random-guessing benchmark of 41.90% by 43.81 percentage points. It also achieved higher hit rates than MNL (78.63%), nested logit (78.59%), and binary logistic regression (75.97%), while its performance was comparable to that of random forest (85.56%) and XGBoost (85.61%). Prediction performance varied substantially with choice-set size: the hit rate decreased from 92.24% for two-route choice sets to 58.02% for five-route choice sets, although the latter remained well above the corresponding random benchmark of 20%. Analysis of the 4,641 mis-predicted trips further showed that errors were concentrated among less frequently selected routes. The median historical share of the observed route was 92.59% for correctly predicted trips but only 20.00% for mis-predicted trips, and 66.90% of all mis-predictions involved routes with historical shares below 30%. These findings indicate that choices deviating from recurrent OD-level route patterns are more difficult to infer from the available route and trip attributes.

The feature analysis showed that mileage-related variables accounted for 65.64% of the total model gain, followed by toll-related variables at 13.06% and travel-time variables at 5.56%. Because toll is directly adjustable in expressway operations, it was subsequently used as the intervention variable in the scenario-based sensitivity analysis. The predicted response to toll adjustment varied markedly with the competitive position of the target route. For routes with baseline choice probabilities of 50%–80%, which accounted for 17.1% of the test sample, a 50% toll increase reduced the median predicted probability by 14.84–16.99 percentage points. By comparison, routes with baseline probabilities of at least 90%, representing 50.2% of the sample, showed a median reduction of only 1.84 percentage points. The magnitude of the response generally increased as the choice set expanded from two to four alternatives and decreased slightly for five-route choice sets. For OD pairs with four or five candidate routes, a 50% increase reduced the median predicted probability by 8.54 and 8.25 percentage points, respectively, compared with 1.96 percentage points for two-route choice sets.

Vehicle type and the form of the toll intervention also influenced the modelled response. Under a 50% increase in the target-route toll, the median predicted probability decreased by 5.11 percentage points for trucks and 2.87 percentage points for cars. Across the examined adjustment magnitudes, reducing the tolls of competing routes produced changes 1.10–2.27 times as large as those obtained by increasing the toll of the target route by the same proportion, with the difference becoming more pronounced as the adjustment magnitude increased. Taken together, the scenario results indicate that toll-based route management is unlikely to produce uniform responses across a network. The competitive position of a route, the number of available alternatives, vehicle type, and the form and magnitude of the toll intervention all affect the predicted response and should therefore be considered jointly when alternative differentiated toll strategies are evaluated.

From an operational perspective, the proposed framework provides a means of screening toll-management scenarios before implementation rather than prescribing a universal toll-setting rule. The results suggest that routes occupying an intermediate competitive position may warrant particular attention because their predicted choice probabilities were more responsive to toll changes than those of strongly dominant routes. They also show that vehicle-specific responses and the availability of competing routes can materially alter the predicted effect of an intervention. By connecting large-scale operational data with route-level prediction and scenario analysis, the framework provides a reproducible decision-support approach for evaluating differentiated tolling and other route-management strategies in intelligent expressway operations.

Author Contributions

Conceptualization, Y.J.S., D.X.W., and N.F.Z.; methodology, Y.J.S., L.Z., and D.X.W.; software, L.Z. and M.X.Y.; validation, M.X.Y., W.W.Q., and Y.L.C.; formal analysis, Y.J.S. and D.X.W.; investigation, M.X.Y., W.W.Q., and Y.L.C.; resources, L.H.Z. and N.F.Z.; data curation, D.X.W., M.X.Y., and L.H.Z.; writing—original draft preparation, Y.J.S. and D.X.W.; writing—review and editing, Y.J.S., L.Z., D.X.W., and N.F.Z.; visualization, Y.J.S. and L.Z.; supervision, N.F.Z.; project administration, N.F.Z.; funding acquisition, L.H.Z. and N.F.Z. All authors have read and agreed to the published version of the manuscript.

Data Availability

The data used to support the research findings are available from the corresponding author upon request.

Conflicts of Interest

The authors declare no conflicts of interest.

References
1.
P. Guarda and S. Qian, “Statistical inference of travelers’ route choice preferences with system-level data,” Transp. Res. B Methodol., vol. 179, p. 102853, 2024. [Google Scholar] [Crossref]
2.
C. Advani, A. Bhaskar, and M. M. Haque, “Bi-level clustering of vehicle trajectories for path choice set and its nested structure identification,” Transp. Res. C Emerg. Technol., vol. 144, p. 103895, 2022. [Google Scholar] [Crossref]
3.
D. McFadden, “Conditional logit analysis of qualitative choice behavior,” in Frontiers in Econometrics, Academic Press, 1974, pp. 105–142. [Google Scholar]
4.
M. Ben-Akiva and S. R. Lerman, Discrete Choice Analysis: Theory and Application to Travel Demand. MIT Press, 1985. [Google Scholar]
5.
K. E. Train, Discrete Choice Methods With Simulation (2nd ed.). Cambridge University Press, 2009. [Google Scholar]
6.
J. D. Ortúzar and L. G. Willumsen, Modelling Transport (4th ed.). John Wiley & Sons, 2011. [Google Scholar]
7.
E. Cascetta, Transportation Systems Analysis: Models and Applications (2nd ed.). Springer, 2009. [Online]. Available: [Google Scholar] [Crossref]
8.
C. G. Prato, “Route choice modeling: Past, present and future research directions,” J. Choice Model., vol. 2, no. 1, pp. 65–100, 2009. [Google Scholar] [Crossref]
9.
D. McFadden and K. Train, “Mixed MNL models for discrete response,” J. Appl. Econom., vol. 15, no. 5, pp. 447–470, 2000, <447::AID-JAE570>3.0.CO;2-1. [Google Scholar] [Crossref]
10.
D. A. Hensher, J. M. Rose, and W. H. Greene, Applied Choice Analysis. Cambridge University Press, 2015. [Online]. Available: [Google Scholar] [Crossref]
11.
L. Cazor, L. C. Duncan, D. P. Watling, O. A. Nielsen, and T. K. Rasmussen, “A closed-form bounded route choice model accounting for heteroscedasticity, overlap, and choice set formation,” Transp. Res. B Methodol., vol. 199, p. 103275, 2025. [Google Scholar] [Crossref]
12.
T. K. Rasmussen, L. C. Duncan, D. P. Watling, and O. A. Nielsen, “Local detouredness: A new phenomenon for modelling route choice and traffic assignment,” Transp. Res. B Methodol., vol. 190, p. 103052, 2024. [Google Scholar] [Crossref]
13.
S. Bekhor, M. E. Ben-Akiva, and M. S. Ramming, “Evaluation of choice set generation algorithms for route choice models,” Ann. Oper. Res., vol. 144, no. 1, pp. 235–247, 2006. [Google Scholar] [Crossref]
14.
E. Frejinger, M. Bierlaire, and M. Ben-Akiva, “Sampling of alternatives for route choice modeling,” Transp. Res. B Methodol., vol. 43, no. 10, pp. 984–994, 2009. [Google Scholar] [Crossref]
15.
D. Liu, D. Li, K. Gao, Y. Song, and T. Zhang, “Enhancing choice-set generation and route choice modeling with data- and knowledge-driven approach,” Transp. Res. C Emerg. Technol., vol. 162, p. 104618, 2024. [Google Scholar] [Crossref]
16.
L. Cazor, L. C. Duncan, D. P. Watling, O. A. Nielsen, and T. K. Rasmussen, “A smooth bounded choice model: Formulation and application in three large-scale case studies,” J. Choice Model., vol. 57, p. 100574, 2025. [Google Scholar] [Crossref]
17.
L. Breiman, “Random forests,” Mach. Learn., vol. 45, no. 1, pp. 5–32, 2001. [Google Scholar] [Crossref]
18.
J. H. Friedman, “Greedy function approximation: A gradient boosting machine,” Ann. Stat., vol. 29, no. 5, pp. 1189–1232, 2001. [Google Scholar] [Crossref]
19.
T. Chen and C. Guestrin, “XGBoost: A scalable tree boosting system,” in Proceedings of the 22nd ACM SIGKDD International Conference on Knowledge Discovery and Data Mining. San Francisco, CA, USA, pp. 785–794, 2016. [Google Scholar] [Crossref]
20.
G. Ke, Q. Meng, T. Finley, T. Wang, W. Chen, W. Ma, Q. Ye, and T. Y. Liu, “LightGBM: A highly efficient gradient boosting decision tree,” in Advances in Neural Information Processing Systems 30. pp. 3146–3154, 2017. [Google Scholar]
21.
L. Cheng, X. Chen, J. De Vos, X. Lai, and F. Witlox, “Applying a random forest method approach to model travel mode choice behavior,” Travel Behav. Soc., vol. 14, pp. 1–10, 2019. [Google Scholar] [Crossref]
22.
Y. Zhang and A. Haghani, “A gradient boosting method to improve travel time prediction,” Transp. Res. C Emerg. Technol., vol. 58, pp. 308–324, 2015. [Google Scholar] [Crossref]
23.
Z. Zhao and Y. Liang, “A deep inverse reinforcement learning approach to route choice modeling with context-dependent rewards,” Transp. Res. C Emerg. Technol., vol. 149, p. 104079, 2023. [Google Scholar] [Crossref]
24.
S. Qiu, G. Qin, M. Wong, and J. Sun, “RoutesFormer: A sequence-based route choice transformer for efficient path inference from sparse trajectories,” Transp. Res. C Emerg. Technol., vol. 162, p. 104552, 2024. [Google Scholar] [Crossref]
25.
H. Wang, E. Moylan, and D. Levinson, “Ensemble methods for route choice,” Transp. Res. C Emerg. Technol., vol. 167, p. 104803, 2024. [Google Scholar] [Crossref]
26.
J. Arriagada, C. A. Guevara, M. A. Munizaga, and S. Gao, “An experiential learning-based transit route choice model using large-scale smart-card data,” Transportation, vol. 52, no. 4, pp. 1543–1568, 2025. [Google Scholar] [Crossref]
27.
Y. Han, F. C. Pereira, M. Ben-Akiva, and C. Zegras, “A neural-embedded discrete choice model: Learning taste representation with strengthened interpretability,” Transp. Res. B Methodol., vol. 163, pp. 166–186, 2022. [Google Scholar] [Crossref]
28.
B. Sifringer, V. Lurkin, and A. Alahi, “Enhancing discrete choice models with representation learning,” Transp. Res. B Methodol., vol. 140, pp. 236–261, 2020. [Google Scholar] [Crossref]
29.
S. Wang, B. Mo, and J. Zhao, “Theory-based residual neural networks: A synergy of discrete choice models and deep neural networks,” Transp. Res. B Methodol., vol. 146, pp. 333–358, 2021. [Google Scholar] [Crossref]
30.
S. van Cranenburgh, S. Wang, A. Vij, F. C. Pereira, and J. Walker, “Choice modelling in the age of machine learning—Discussion paper,” J. Choice Model., vol. 42, p. 100340, 2022. [Google Scholar] [Crossref]
31.
S. B. Ayaz, H. Tian, S. Gao, and D. L. Fisher, “Proactive route choice with real-time information: Learning and effects of network complexity and cognitive load,” Transp. Res. C Emerg. Technol., vol. 149, p. 104035, 2023. [Google Scholar] [Crossref]
32.
B. Zhou and R. Liu, “A generalized rationally inattentive route choice model with non-uniform marginal information costs,” Transp. Res. B Methodol., vol. 189, p. 102993, 2024. [Google Scholar] [Crossref]
33.
F. Ahmad and L. Al-Fagih, “Travel behaviour and game theory: A review of route choice modeling behaviour,” J. Choice Model., vol. 50, p. 100472, 2024. [Google Scholar] [Crossref]
34.
Y. Tian, W. Zhu, and F. Song, “Route choice modelling for an urban rail transit network: Past, recent progress and future prospects,” Eur. Transp. Res. Rev., vol. 16, p. 52, 2024. [Google Scholar] [Crossref]
35.
People’s Government of Guangdong Province, “Approval on adjusting the toll charging method for toll roads: Yue Fu Han [2019] No. 416,” 2019, [Online]. Available: https://www.gd.gov.cn/zwgk/gongbao/2019/36/content/post_3366602.html [Google Scholar]
36.
T. Akiba, S. Sano, T. Yanase, T. Ohta, and M. Koyama, “Optuna: A next-generation hyperparameter optimization framework,” in Proceedings of the 25th ACM SIGKDD International Conference on Knowledge Discovery & Data Mining, Anchorage, AK, USA, 2019, pp. 2623–2631. [Google Scholar] [Crossref]
37.
K. A. Small, E. T. Verhoef, and R. Lindsey, The Economics of Urban Transportation. Routledge, 2007. [Online]. Available: [Google Scholar] [Crossref]
38.
H. Yang and H. J. H. J. Huang, Mathematical and Economic Theory of Road Pricing. Elsevier, 2005. [Google Scholar]
39.
D. Brownstone and K. A. Small, “Valuing time and reliability: Assessing the evidence from road pricing demonstrations,” Transp. Res. A Policy Pract., vol. 39, no. 4, pp. 279–293, 2005. [Google Scholar] [Crossref]
40.
D. A. Hensher, “Measurement of the valuation of travel time savings,” J. Transp. Econ. Policy, vol. 35, no. 1, pp. 71–98, 2001. [Google Scholar]
41.
G. de Jong, M. Kouwenhoven, J. Bates, P. Koster, E. Verhoef, L. Tavasszy, and P. Warffemius, “New SP-values of time and reliability for freight transport in the Netherlands,” Transp. Res. E Logist. Transp. Rev., vol. 64, pp. 71–87, 2014. [Google Scholar] [Crossref]

Cite this:
APA Style
IEEE Style
BibTex Style
MLA Style
Chicago Style
GB-T-7714-2015
Su, Y. J., Zhang, L., Yuan, M. X., Wu, D. X., Qi, W. W., Cao, Y. L., Zhang, L. H., & Zhang, N. F. (2025). Data-Driven Route Choice Prediction for Expressway Operations Using Ensemble Learning: Sensitivity Analysis of Differentiated Toll Policies. J. Ind Intell., 3(4), 226-245. https://doi.org/10.56578/jii030403
Y. J. Su, L. Zhang, M. X. Yuan, D. X. Wu, W. W. Qi, Y. L. Cao, L. H. Zhang, and N. F. Zhang, "Data-Driven Route Choice Prediction for Expressway Operations Using Ensemble Learning: Sensitivity Analysis of Differentiated Toll Policies," J. Ind Intell., vol. 3, no. 4, pp. 226-245, 2025. https://doi.org/10.56578/jii030403
@research-article{Su2025Data-DrivenRC,
title={Data-Driven Route Choice Prediction for Expressway Operations Using Ensemble Learning: Sensitivity Analysis of Differentiated Toll Policies},
author={Yuejiang Su and Liang Zhang and Minxian Yuan and Dexin Wu and Weiwei Qi and Yunlong Cao and Liuhua Zhang and Nanfeng Zhang},
journal={Journal of Industrial Intelligence},
year={2025},
page={226-245},
doi={https://doi.org/10.56578/jii030403}
}
Yuejiang Su, et al. "Data-Driven Route Choice Prediction for Expressway Operations Using Ensemble Learning: Sensitivity Analysis of Differentiated Toll Policies." Journal of Industrial Intelligence, v 3, pp 226-245. doi: https://doi.org/10.56578/jii030403
Yuejiang Su, Liang Zhang, Minxian Yuan, Dexin Wu, Weiwei Qi, Yunlong Cao, Liuhua Zhang and Nanfeng Zhang. "Data-Driven Route Choice Prediction for Expressway Operations Using Ensemble Learning: Sensitivity Analysis of Differentiated Toll Policies." Journal of Industrial Intelligence, 3, (2025): 226-245. doi: https://doi.org/10.56578/jii030403
SU Y J, ZHANG L, YUAN M X, et al. Data-Driven Route Choice Prediction for Expressway Operations Using Ensemble Learning: Sensitivity Analysis of Differentiated Toll Policies[J]. Journal of Industrial Intelligence, 2025, 3(4): 226-245. https://doi.org/10.56578/jii030403
cc
©2025 by the author(s). Published by Acadlore Publishing Services Limited, Hong Kong. This article is available for free download and can be reused and cited, provided that the original published version is credited, under the CC BY 4.0 license.