Data-Driven Route Choice Prediction for Expressway Operations Using Ensemble Learning: Sensitivity Analysis of Differentiated Toll Policies
Abstract:
Intelligent expressway management requires reliable prediction of travellers' route choices and quantitative assessment of their responses to operational interventions. This study investigates a data-driven route choice framework that links large-scale toll transaction data with ensemble learning to support differentiated tolling and route management. Actual vehicle routes were reconstructed from toll transaction records, electronic toll collection (ETC) gantry records, and expressway network topology, from which revealed-preference candidate route sets were derived. A total of 28 features were constructed to characterize route attributes, relative differences among candidate routes, choice-set composition, and trip attributes. A Light Gradient Boosting Machine (LightGBM) model was then developed using each “trip $\times$ candidate route” pair as the unit of observation, and its predictive performance was compared with random forest, Extreme Gradient Boosting (XGBoost), multinomial logit (MNL), nested logit, and binary logistic regression. The empirical analysis used 329,699 trips from the Guangdong expressway network, including 32,473 trips in the test set. The Light Gradient Boosting Machine (LightGBM) model achieved a route-level hit rate of 85.71%, exceeding the random-guessing benchmark by 43.81 percentage points and outperforming the three conventional discrete-choice and logistic models, while showing performance comparable to random forest and XGBoost. The sensitivity analysis further showed that toll responses varied substantially with a route’s competitive position, choice-set size, and vehicle type. Routes with baseline choice probabilities of 50%–80% were more responsive to toll changes than strongly dominant routes, while trucks showed greater toll sensitivity than cars. The results demonstrate that large-scale operational data and ensemble learning can provide a reproducible framework for route choice prediction and scenario-based evaluation of differentiated toll policies. The proposed framework supports data-driven decision-making for intelligent expressway operations and provides a practical basis for evaluating route-level traffic management strategies.
1. Introduction
As expressway networks become increasingly dense and interconnected, travellers are increasingly confronted with multiple feasible routes between the same origin–destination (OD) pair. The routes selected by individual travellers determine how traffic demand is distributed across parallel corridors and, consequently, influence the effectiveness of operational measures such as differentiated tolling, congestion management, and route guidance. Reliable route choice prediction is therefore an important component of data-driven expressway operation and intelligent traffic management. Conventional route choice studies have largely relied on household travel surveys, stated-preference experiments, Global Positioning System (GPS) trajectories, or other individual-level observations. Although these data sources provide valuable information on travel behaviour, their coverage can be constrained by sample size, network extent, vehicle classes, and operating conditions. The increasing availability of network-scale operational data provides a different empirical basis for route choice analysis. System-level observations have been used to infer travellers' route choice preferences over large transportation networks [1], while data-driven clustering of vehicle trajectories has demonstrated how observed paths can be used to identify representative route alternatives and their underlying choice structures [2]. For expressway systems equipped with electronic toll collection (ETC) infrastructure, the integration of toll transactions, gantry passage sequences, and network topology makes it possible to reconstruct realised vehicle routes at large scale, creating a direct observational basis for intelligent and data-driven traffic management.
Route choice modelling has traditionally been grounded in random utility theory. Conditional logit and multinomial logit (MNL) models established the foundation for disaggregate choice analysis [3], which was subsequently developed into a systematic framework for modelling individual travel decisions [4], while simulation-based estimation methods further extended the range of choice structures that can be estimated in practice [5]. These models retain an important advantage in behavioural interpretation because estimated parameters can be related to travel-time valuation, cost sensitivity, substitution among alternatives, and other quantities with direct behavioural meaning. They can also be incorporated into traffic assignment and wider transportation-system analysis [6], [7]. Route choice, however, presents several difficulties that are less pronounced in many other discrete-choice settings. Candidate routes frequently overlap over substantial portions of a network, so their unobserved utilities cannot always be regarded as independent. This limitation has motivated a range of alternative route-choice formulations and corrections designed to represent correlation and similarity among alternatives [8], [9], [10]. Recent studies show that this problem remains methodologically important. Cazor et al. [11] developed a bounded route choice formulation that jointly addresses heteroscedasticity, route overlap, and implicit choice-set formation, while Rasmussen et al. [12] demonstrated that both local and global detours can materially affect route generation and route-choice probabilities. These developments indicate that route similarity, detour structure, and the competitive relationship among alternatives remain central to contemporary route-choice modelling.
Choice-set construction constitutes a closely related challenge. A route-choice model cannot reproduce observed behaviour reliably if the set of alternatives presented to the model does not reasonably represent the routes available or perceived by travellers. Classical approaches have therefore relied on enumeration, sampling, and dedicated path-generation algorithms to construct candidate sets [13], [14]. Such approaches remain useful, but their performance may depend strongly on network size, parameter settings, and assumptions about which alternatives are behaviourally relevant. The growth of large trajectory datasets has encouraged more data-driven treatment of this problem. Liu et al. [15] integrated data- and knowledge-driven methods for choice-set generation and route choice modelling, using a conditional variational autoencoder to represent the choice-set generation process together with a learning-based choice model. More recently, Cazor et al. [16] proposed a smooth bounded choice formulation in which consideration-set formation is incorporated implicitly within the choice process and evaluated the approach using several large-scale applications, including route-choice cases. These studies reflect an important methodological shift: the set of behaviourally relevant routes is increasingly treated as part of the modelling problem rather than merely as an external preprocessing step.
A second major research stream approaches route choice primarily as a predictive learning problem. Random forest [17], gradient boosting [18], Extreme Gradient Boosting (XGBoost) [19], and LightGBM [20] can represent nonlinear relationships, threshold effects, and higher-order interactions without requiring the complete form of a utility function to be specified in advance. Their computational scalability also makes them attractive for large operational datasets containing large numbers of trips and route alternatives. Machine-learning methods have already demonstrated substantial predictive capability in transportation applications such as travel-mode classification [21] and travel-time prediction [22]. Their use in route choice has expanded as richer trajectory and network data have become available. Zhao and Liang [23] applied deep inverse reinforcement learning to recover context-dependent routing preferences from observed trajectories, providing a framework in which route-selection behaviour can be learned from sequential decisions rather than represented entirely through a predefined utility function. Qiu et al. [24] developed RoutesFormer, a sequence-based Transformer that integrates path inference and route-choice representation and uses attention mechanisms to capture complex sequential dependencies in sparse trajectory data. These developments demonstrate that route-choice regularities can increasingly be extracted directly from high-dimensional observations.
Of particular relevance to the present study is the recent application of ensemble and deep-learning methods to route choice prediction. Wang et al. [25] systematically evaluated ensemble methods for network-wide route choice and showed that the relative performance of individual models can vary across training and testing datasets, while appropriately constructed ensembles can provide robust prediction of aggregate route flows. Their results are important for operational applications because they suggest that model usefulness should not be assessed solely by identifying a universally superior algorithm; instead, predictive stability across heterogeneous observations and network conditions is equally important. Recent studies have also examined learning-based route choice using large-scale observed travel data, showing that travellers' accumulated experience and previous choices can provide additional information for modelling route-selection behaviour [26]. Together, these recent studies indicate that route choice research is moving beyond a simple comparison between conventional discrete-choice models and machine-learning algorithms. The emerging question is how flexible predictive models can exploit large observational datasets while still producing results that remain meaningful for transport operations and behavioural analysis.
Predictive flexibility, however, does not eliminate the need for interpretation. A machine-learning model may provide high predictive accuracy without directly yielding behavioural quantities such as marginal effects, elasticities, welfare measures, or explicit substitution patterns. A parallel research stream has therefore attempted to combine flexible computational structures with interpretable choice mechanisms. Neural-embedded discrete choice models have been developed to learn heterogeneous taste representations while retaining a utility-based structure [27]. Representation-learning methods have likewise been incorporated into conventional discrete-choice models to capture nonlinear effects that are difficult to specify a priori [28]. Theory-based residual neural networks provide another approach by retaining a behavioural model as the structural component and allowing a neural network to learn residual patterns not captured by the theoretical specification [29]. Van Cranenburgh et al. [30] provide a broader discussion of this convergence between theory-driven choice modelling and data-driven machine learning, highlighting both the opportunities created by flexible predictive methods and the importance of maintaining behavioural validity. These developments are directly relevant to operational route-choice applications, where predictive performance alone is insufficient if the model cannot be related to controllable variables or interpreted in a way that supports management decisions.
Route choice is also shaped by the information available to travellers and by the broader decision environment. Real-time traffic information can alter perceived route attractiveness, affect expectations regarding downstream conditions, and change the way travellers evaluate future diversion opportunities. Ayaz et al. [31] examined proactive route choice under real-time information and showed that learning, network complexity, and cognitive load can influence route-selection behaviour. Zhou and Liu [32] approached the problem from the perspective of rational inattention, explicitly modelling heterogeneous information-processing costs and deriving route-choice probabilities when travellers do not process all available information equally. From a broader behavioural perspective, Ahmad and Al-Fagih [33] reviewed game-theoretic approaches to route choice and highlighted their relevance to congestion, tolling, transportation policy, and strategic interactions among travellers. Tian et al. [34], in a recent review of route-choice modelling, likewise showed that contemporary research increasingly relies on richer transport data, more flexible behavioural structures, and advanced computational approaches, while continuing to face challenges related to data integration, model calibration, validation, and representation of traveller heterogeneity. Taken together, this literature suggests that modern route-choice analysis increasingly lies at the intersection of behavioural modelling, machine learning, network data, and operational decision support.
Despite these advances, several gaps remain between route choice prediction and its application to intelligent expressway operations. First, conventional utility-based models require predefined functional structures and may have difficulty representing complex nonlinearities, interactions, and threshold effects among route attributes when very large operational datasets are available. Second, although recent machine-learning, deep-learning, and ensemble approaches have improved predictive flexibility, much of the literature still concentrates primarily on model fit or predictive accuracy rather than explicitly connecting fitted models to operational variables that network operators can adjust. Third, prediction errors themselves are seldom examined systematically. A high aggregate accuracy does not reveal which route-choice situations are difficult to predict, whether errors are concentrated among minority alternatives, or where the practical boundary of a fitted model lies. Fourth, large-scale operational observations are increasingly available, but relatively few studies integrate revealed route sets, absolute and relative route attributes, ensemble prediction, error diagnosis, and management-oriented scenario analysis within a single framework. These limitations are particularly important for differentiated tolling, where the operational value of a predictive model depends not only on identifying the route most likely to be chosen but also on estimating how predicted route probabilities change when toll relationships among competing alternatives are modified.
Accordingly, this study develops a data-driven route choice prediction framework for expressway operations using large-scale toll transaction data and ensemble learning. Following the analytical sequence of “toll transaction data → choice-set and feature construction → route choice probability prediction → influencing-factor analysis → policy sensitivity”, a LightGBM model is developed with each “trip $\times$ candidate route” pair treated as an observation. In addition to route-specific attributes, the feature system represents the relative position of each route within its corresponding choice set, allowing the model to distinguish between the absolute characteristics of an alternative and its competitive position relative to other feasible routes. Predictive performance is evaluated using the route-level hit rate and compared with alternative machine-learning and conventional choice models. Key influencing factors are examined through statistical association tests and gain-based feature importance, while mis-predicted trips are analysed to identify the conditions under which observed route choices are difficult to distinguish. The fitted model is subsequently used for scenario-based sensitivity analysis by varying tolls, a directly adjustable operational variable, and recalculating the affected route characteristics and choice probabilities. In this way, the study integrates large-scale operational data, ensemble learning, route-relative feature representation, model interpretation, error diagnosis, and differentiated-toll scenario evaluation within a unified analytical framework. The resulting approach provides a reproducible basis for route choice prediction and supports data-driven decision-making for differentiated tolling, route guidance, and intelligent expressway operations. The overall research framework is illustrated in Figure 1.

2. Methodological Framework
The modelling objective is to estimate the probability that each candidate route is selected for a given trip. Rather than treating a physical route as having a fixed propensity to be chosen, the proposed formulation evaluates the route relative to the alternatives available for a specific trip. The same physical route may therefore constitute a positive observation in one trip and a negative observation in another, depending on its attributes and competitive position within the corresponding choice set.
For each trip, candidate routes were identified from revealed travel behaviour. Routes accounting for no less than 2% of the observed trips for the corresponding OD pair were retained to form the revealed choice set, with no fewer than two and no more than five candidate routes for each OD pair. Each modelling observation represents a “trip $\times$ candidate route” pair. Its feature vector is expressed as:
where, $x_{ij}$ is the feature vector of candidate route $j$ for trip $i$; $a_{ij}$ contains the route-specific attributes, including mileage, toll, travel time, number of switches, and operating speed; $r_{ij}$ describes the relative position of the candidate route within the choice set, including ratios and differences relative to the best alternative, dominance indicators, and margins relative to the second-best alternative; $s_i$ represents choice-set characteristics, including the number of alternatives and their dispersion; and $t_i$ contains trip-level attributes such as vehicle type, departure period, and peak-hour status.
This formulation allows the model to learn not only the absolute characteristics of a route but also its competitive position among the alternatives available to the traveller. Once the model has been trained, an operational variable such as toll can be changed for a target route, the affected relative features can be recalculated, and the resulting choice probabilities can be predicted under the modified scenario. This provides the basis for the policy sensitivity analysis conducted later in the study.
Route choice probability was estimated using LightGBM, a gradient-boosting decision-tree algorithm. Gradient boosting constructs an ensemble sequentially, with each new decision tree fitted to reduce the prediction error remaining from the preceding iterations. LightGBM further employs histogram-based feature discretisation and a leaf-wise tree-growth strategy [20], making it computationally suitable for large datasets while retaining the ability to represent nonlinear relationships and interactions among predictors.
For the feature vector defined in Eq. (1), the predicted probability that candidate route $j$ is chosen for trip $i$ is expressed as:
where, $\hat{p}_{ij}$ denotes the predicted choice probability of candidate route $j$ for trip $i$; $f(\ )$ is a fitted gradient-boosting model consisting of $N$ decision trees; $\beta_n$ denotes the contribution of the $n$-th tree; and $h_{n}(\ )$ is its output.
The use of route-specific, relative, choice-set, and trip-level features enables the model to represent nonlinear differences among competing alternatives without requiring a predefined utility-function form. This property is particularly relevant to the present application because the attractiveness of an expressway route depends not only on its absolute mileage, toll, or travel time but also on how these attributes compare with those of the other routes available for the same trip.
Predictive performance was evaluated at the trip level. For each trip in the test set, the candidate route with the highest predicted probability was identified as the model-predicted route and compared with the route actually taken. The route-level hit rate is defined as:
where, $H$ is the route-level hit rate; $M$ is the number of trips in the test set; $\hat{j}_i$ denotes the candidate route with the highest predicted probability for trip $i$; $j_i$ denotes the route actually chosen; and $II(\ )$ is the indicator function, which equals 1 when the predicted and observed routes coincide and 0 otherwise.
This trip-level measure directly evaluates whether the model identifies the route actually selected from the available alternatives and therefore provides an interpretable basis for comparing predictive performance across different model specifications.
3. Data and Feature Construction
The empirical analysis used one day of toll-station transaction records, ETC gantry records, and network topology data from the Guangdong expressway network. Toll transaction records contained information on entry and exit stations, vehicle type, and toll amount, while ETC gantry records provided the sequence of gantries passed by each vehicle. By integrating these two data sources with the directed expressway network, the actual travel route of each vehicle was reconstructed. Route mileage, toll, peak/off-peak travel time, and the number of switches between expressway segments were subsequently derived. Toll values were converted from the standard class-1 vehicle charge according to the vehicle-type coefficients specified in the Classification of Vehicle Types and Toll Coefficients for Guangdong Expressways (Yue Fu Han [2019] No. 416) [35].
The raw dataset contained 1,282,059 transaction records covering 11,760 OD pairs. Data preprocessing included duplicate removal, field validation, spatiotemporal consistency checks, and gantry-sequence stitching to address missing or inconsistent observations. After preprocessing, 1,255,584 records covering 11,722 OD pairs were retained, corresponding to a removal rate of 2.06%.
Revealed choice sets were then constructed from observed route use. For each OD pair, routes accounting for at least 2% of observed trips were retained, and OD pairs with fewer than two candidate routes were excluded because they did not represent a route choice problem. No more than five candidate routes were retained for an OD pair. This procedure yielded 5,596 OD pairs, 329,699 trips, and 858,681 candidate-route-level observations. The substantial reduction in the number of OD pairs primarily resulted from the fact that 52.3% of OD pairs exhibited only one observed travel route and therefore contained no observed route-choice variation.
The modelling sample was randomly divided by trip into training and test sets at a ratio of 9:1. The training set contained 297,226 trips and 774,256 candidate-route-level observations, while the test set contained 32,473 trips and 84,425 candidate-route-level observations. Table 1 summarises the sample construction and data split. The distribution of choice-set size is presented in Figure 2. Among the 5,596 OD pairs retained for modelling, 3,216 had two candidate routes, 1,512 had three, 591 had four, and 277 had five. These groups accounted for 60.34%, 24.69%, 9.14%, and 5.82% of the modelling trips, respectively, with an average choice-set size of 2.60 routes.
Stage | Number | Description |
|---|---|---|
Raw toll transaction records | 1,282,059 | Raw entry–exit–gantry records |
Records after cleaning | 1,255,584 | Duplicates, missing values, and spatio-temporal conflicts removed (2.06% dropped) |
Modelling trips | 329,699 | 5,596 OD pairs with at least 2 candidates |
Candidate-route-level records | 858,681 | Positive cases (chosen): 329,699, 38.40% |
Training set | 297,226 trips/774,256 records | Split by trip, 90% |
Test set | 32,473 trips/84,425 records | Split by trip, 10% |

The distribution of the 329,699 modelling trips was examined across departure period, toll level, trip mileage, travel time, vehicle type, and number of expressway-segment switches. Table 2 summarises these characteristics and provides an overview of the operating conditions represented in the dataset.
Dimension | Group | Trips | Share/% | Dimension | Group | Trips | Share/% |
|---|---|---|---|---|---|---|---|
Departure period | 0:00–6:00 (small hours) | 28854 | 8.75 | Toll/CNY | $<$10 | 10771 | 3.27 |
Departure period | 7:00–8:00 (morning peak) | 42572 | 12.91 | Toll/CNY | 10–20 | 74863 | 22.71 |
Departure period | 9:00–11:00 (morning off-peak) | 65975 | 20.01 | Toll/CNY | 20–30 | 81094 | 24.60 |
Departure period | 12:00–16:00 (afternoon off-peak) | 92199 | 27.96 | Toll/CNY | 30–50 | 81878 | 24.83 |
Departure period | 17:00–18:00 (evening peak) | 44469 | 13.49 | Toll/CNY | 50–80 | 43347 | 13.15 |
Departure period | 19:00–23:00 (night) | 55630 | 16.87 | Toll/CNY | $\geq$80 | 37746 | 11.45 |
Trip mileage/km | $<$10 | 117 | 0.04 | Travel time/min | $<$15 | 11405 | 3.46 |
Trip mileage/km | 10–20 | 22242 | 6.75 | Travel time/min | 15–25 | 72527 | 22.00 |
Trip mileage/km | 20–30 | 57966 | 17.58 | Travel time/min | 25–40 | 125481 | 38.06 |
Trip mileage/km | 30–50 | 119830 | 36.35 | Travel time/min | 40–60 | 80861 | 24.53 |
Trip mileage/km | 50–80 | 93295 | 28.30 | Travel time/min | 60–90 | 33879 | 10.28 |
Trip mileage/km | $\geq$80 | 36249 | 10.99 | Travel time/min | $\geq$90 | 5546 | 1.68 |
Car/truck | Car | 247294 | 75.01 | Number of switches | 1 | 37473 | 11.37 |
Car/truck | Truck | 82405 | 24.99 | Number of switches | 2 | 152414 | 46.23 |
Vehicle-type breakdown | Car class 1 | 243636 | 73.90 | Number of switches | 3 | 90500 | 27.45 |
Vehicle-type breakdown | Truck class 1 | 47579 | 14.43 | Number of switches | $\geq$4 | 49312 | 14.96 |
Medium- and long-distance movements constituted the majority of the modelling sample. Trips of 30–80 km accounted for 64.65% of all observations, while trips with tolls of 20–50 CNY accounted for 49.43%, and those with travel times of 15–60 min represented 84.59%. In addition, 88.64% of trips involved at least two switches between expressway segments, indicating that most observations represented journeys involving multiple expressway sections and meaningful alternatives within the network.
The dataset also covered different temporal and vehicle operating conditions. Afternoon off-peak trips accounted for the largest share (27.96%), while the morning and evening peak periods together represented 26.40%. Night and early-morning trips accounted for a further 25.62%. Cars represented 75.01% of all trips, including 73.90% classified as class-1 cars, whereas trucks accounted for 24.99%. The truck subsample contained 82,405 trips, providing a substantial observational basis for the vehicle-type comparisons reported later in the analysis.
The feature system was designed to represent both the absolute characteristics of each candidate route and its competitive position within the corresponding choice set. A total of 28 features were constructed and organised into four broad components: route-specific attributes, relative characteristics of candidate routes, choice-set characteristics, and trip-level attributes. The underlying physical dimensions included mileage, toll, travel time, and switching between expressway segments. Table 3 reports the complete feature definitions.
| Code | Dimension | Feature | Definition | $p$-value |
|---|---|---|---|---|
| C1 | Route's own attributes | Route mileage | Mileage of the candidate route, km | $<0.001$ |
| C2 | Route's own attributes | Route toll | Toll of the candidate route after conversion for the vehicle type, Chinese Yuan (CNY) | $<0.001$ |
| C3 | Route's own attributes | Route travel time | Peak/off-peak travel time of the candidate route, min | $<0.001$ |
| C4 | Route's own attributes | Number of switches | Number of switches between expressway segments along the candidate route | $<0.001$ |
| C5 | Relative to the best | Detour ratio | Candidate mileage $\div$ shortest mileage of this trip ($\geq1$) | $<0.001$ |
| C6 | Relative to the best | Toll ratio | Candidate toll $\div$ lowest toll of this trip ($\geq1$) | $<0.001$ |
| C7 | Relative to the best | Time ratio | Candidate travel time $\div$ shortest travel time of this trip ($\geq1$) | $<0.001$ |
| C8 | Relative to the best | Switch difference | Candidate switches $-$ minimum switches of this trip ($\geq0$) | $<0.001$ |
| C9 | Absolute difference | Mileage difference | Candidate mileage $-$ shortest mileage of this trip, km | $<0.001$ |
| C10 | Absolute difference | Toll difference | Candidate toll $-$ lowest toll of this trip, CNY | $<0.001$ |
| C11 | Absolute difference | Time difference | Candidate travel time $-$ shortest travel time of this trip, min | $<0.001$ |
| C12 | Dominance indicator | Shortest mileage | 1 if the candidate has the shortest mileage in this trip | $<0.001$ |
| C13 | Dominance indicator | Lowest toll | 1 if the candidate has the lowest toll in this trip | $<0.001$ |
| C14 | Dominance indicator | Shortest travel time | 1 if the candidate has the shortest travel time in this trip | $<0.001$ |
| C15 | Dominance indicator | No. of dominant dimensions | Number of dimensions (mileage/toll/travel time/switching) in which the candidate is best (0–4) | $<0.001$ |
| C16 | Relative to the second best | Mileage lead | For the best candidate, the relative margin by which it leads the second best ($\geq0$); otherwise the negative of the margin by which it trails the best ($\leq0$) | $<0.001$ |
| C17 | Relative to the second best | Toll lead | As above, computed on the toll dimension | $<0.001$ |
| C18 | Relative to the second best | Travel time lead | As above, computed on the travel time dimension | $<0.001$ |
| C19 | Operating level | Travel speed | Travel speed of the candidate route, km/h | $<0.001$ |
| C20 | Operating level | Speed ratio | Candidate travel speed $\div$ highest speed of this trip ($\leq1$) | $<0.001$ |
| C21 | Operating level | Toll per km | Candidate toll $\div$ mileage, CNY/km | $<0.001$ |
| C22 | Choice set | Choice-set size | Number of candidate routes in this trip (2–5) | $<0.001$ |
| C23 | Choice set | Mileage dispersion | Standard deviation $\div$ mean of the candidate mileage in this trip | $<0.001$ |
| C24 | Choice set | Travel time dispersion | Standard deviation $\div$ mean of the candidate travel time in this trip | $<0.001$ |
| C25 | Traveller | Vehicle-type code | Car classes 1–4 coded 1–4; truck classes 1–6 coded 11–16 | $<0.001$ |
| C26 | Traveller | Car/truck | 0 for car, 1 for truck | $<0.001$ |
| C27 | Traveller | Departure hour | Hour of the entry passage timestamp (0–23) | $<0.001$ |
| C28 | Traveller | Peak hour | 1 for 7:00–9:00 and 17:00–19:00 | $<0.001$ |
For relative features, each candidate route was compared with the best alternative within the same trip using ratios, absolute differences, dominance indicators, and margins relative to the second-best alternative. These features enabled the model to capture both the absolute characteristics and relative competitiveness of each route within the traveller's choice set.
The statistical association between each feature and observed route choice was examined using chi-square tests based on the 858,681 candidate-route-level observations. Continuous variables were discretised by deciles, and categories representing less than 1% of the sample were combined with adjacent or residual categories as appropriate. The null hypothesis for each test was that the corresponding feature was independent of whether the candidate route was selected. As shown in Table 3, all 28 features were statistically associated with observed route choice at ($p <\ $0.001). Given the large sample size, these significance tests are interpreted as evidence of statistical association rather than as direct measures of effect magnitude. The relative contribution of individual features to predictive performance is therefore examined separately through the model-based feature-importance analysis presented in Section 4.
4. Model Application and Results
The hyperparameters of LightGBM were optimised using the Bayesian optimisation framework Optuna [36]. Binary log-loss was adopted as the optimisation objective, and 10% of the training trips were reserved as a validation set during hyperparameter tuning. The resulting optimal configuration is reported in Table 4, with the number of boosting iterations determined as 1,532.
Parameter | Value | Parameter | Value |
|---|---|---|---|
learning_rate | 0.0279 | feature_fraction | 0.6732 |
num_leaves | 128 | bagging_fraction | 0.8115 |
max_depth | 9 | lambda_l1 | 0.0284 |
min_data_in_leaf | 53 | lambda_l2 | 0.6911 |
min_gain_to_split | 0.3126 | Number of iterations | 1,532 |
The route-level hit rate was calculated for the 32,473 trips in the test set using Eq. (3). The LightGBM model achieved an overall hit rate of 85.71%. For comparison, random selection with equal probability assigned to each candidate route within its corresponding choice set produced an aggregate hit rate of 41.90%. The proposed model therefore exceeded this benchmark by 43.81 percentage points. Prediction performance across different trip and choice-set characteristics is reported in Table 5 and Figure 3.
Scenario | Group | Trips | Route-Level Hit Rate | Scenario | Group | Trips | Route-Level Hit Rate |
|---|---|---|---|---|---|---|---|
Overall | All test trips | 32473 | 85.71% | Choice-set size | 2 routes | 19648 | 92.24% |
Car/truck | Car | 24260 | 85.11% | Choice-set size | 3 routes | 8022 | 82.36% |
Car/truck | Truck | 8213 | 87.47% | Choice-set size | 4 routes | 2952 | 68.67% |
Time period | Off-peak | 23787 | 85.59% | Choice-set size | 5 routes | 1851 | 58.02% |
Time period | Peak (7:00–9:00/17:00–19:00) | 8686 | 86.02% | Trip mileage/km | $<$20 | 2140 | 93.79% |
Trip mileage/km | 20–40 | 12385 | 89.62% | Trip mileage/km | 40–60 | 9777 | 82.00% |
Trip mileage/km | $\geq$60 | 8171 | 82.11% | Travel time/min | $<$20 | 4378 | 92.51% |
Travel time/min | 20–40 | 16496 | 84.87% | Travel time/min | 40–60 | 7859 | 84.15% |
Travel time/min | $\geq$60 | 3740 | 84.73% |

The most pronounced variation in prediction performance was associated with choice-set size. The hit rate decreased from 92.24% for trips with two candidate routes to 82.36% for three routes, 68.67% for four routes, and 58.02% for five routes. Although prediction became more difficult as the number of alternatives increased, the hit rate for five-route choice sets remained substantially above the corresponding random benchmark of 20%.
Differences across vehicle types and departure periods were comparatively small. The hit rate was 87.47% for trucks and 85.11% for cars, while peak- and off-peak trips yielded hit rates of 86.02% and 85.59%, respectively. A clearer difference was observed across trip-distance groups.
Trips shorter than 20 km achieved a hit rate of 93.79%, compared with approximately 82% for trips of 40 km or more. These results suggest that the model distinguishes route choices more reliably when the set of competing alternatives is relatively small. As the choice set expands or route characteristics become less differentiated, identifying the observed route becomes more difficult.
To examine the conditions under which prediction errors occurred, the 4,641 mis-predicted trips were analysed according to the historical choice share of the route actually taken within the corresponding OD pair. The results are reported in Table 6 and Figure 4. The median historical share of the observed route was 92.59% among correctly predicted trips but only 20.00% among mis-predicted trips. The hit rate increased monotonically with historical route share, from 1.25% for routes with a share below 10% to 99.96% for routes accounting for at least 90% of observed trips. Overall, 66.90% of all mis-predictions involved routes with historical shares below 30%.
Observed-Route Share Within OD Pair | Trips | Route-Level Hit Rate | Mis-Predicted Trips | Share of Mis-Predictions/% |
|---|---|---|---|---|
$<$10% | 1356 | 1.25% | 1339 | 28.85 |
10%–20% | 1067 | 8.34% | 978 | 21.07 |
20%–30% | 901 | 12.54% | 788 | 16.98 |
30%–50% | 2199 | 44.97% | 1210 | 26.07 |
50%–70% | 3202 | 92.29% | 247 | 5.32 |
70%–90% | 6664 | 98.92% | 72 | 1.55 |
$\geq$90% | 17084 | 99.96% | 7 | 0.15 |

This pattern indicates that prediction errors were concentrated among trips in which the observed route differed from the route most commonly selected for the same OD pair. Such deviations may reflect factors not represented in the current feature set, including individual driving preferences, navigation recommendations, service-area requirements, temporary traffic controls, or other trip-specific circumstances. The results therefore identify an important boundary of the present data-driven model: route choices that deviate substantially from recurrent population-level patterns are more difficult to infer from the available route and trip attributes alone.
Table 7 provides an illustrative mis-predicted case involving two nearly equivalent alternatives. The two routes differed by less than 3.4% in mileage, toll, and travel time: the observed route measured 70.41 km, cost 76.51 CNY, and required 61.56 min, whereas the model's first-ranked route measured 70.56 km, cost 74.05 CNY, and required 62.22 min. Their predicted probabilities were correspondingly close, at 56.27% and 56.43%, a difference of only 0.16 percentage points.
Item | Case (Nearly Equivalent) |
|---|---|
OD | Shatian East–Guangfo Xinguangxian |
Vehicle type/departure time | Truck/11:00 |
Choice-set size | 2 routes |
Actual travel route | Guanfan Expressway–Humen Second Bridge–South Second Ring–West Line Section II–West Line Section I–Guangming Foshan–Fobei North |
Model's first choice | Guanfan Expressway–Humen Second Bridge–South Second Ring–Dongxin Expressway–Guangming Guangzhou–Guangming Foshan–Fobei North |
Actual: mileage/toll/travel time | 70.41 km/76.51 CNY/61.56 min |
First choice: mileage/toll/travel time | 70.56 km/74.05 CNY/62.22 min |
Predicted probability, actual/first choice | 56.27%/56.43% |
Share of the actual travel route within the OD pair | 39.13% (9/23) |
To evaluate the predictive performance of LightGBM relative to alternative modelling approaches, the same test set was used to compare it with random forest (RF) [17], XGBoost [19], binary logistic regression, MNL [3], and nested logit (NL) [4]. MNL and NL were implemented as route-level choice models that directly estimated probabilities across the alternatives within each choice set. The results are reported in Table 8 and Figure 5.
Model | Route-Level Hit Rate | Improvement Over the Random Benchmark/Percentage Points |
|---|---|---|
LightGBM (this paper) | 85.71% | 43.81 |
XGBoost | 85.61% | 43.71 |
Random forest (RF) | 85.56% | 43.66 |
Multinomial logit (MNL) | 78.63% | 36.73 |
Nested logit (NL) | 78.59% | 36.69 |
Binary logistic regression | 75.97% | 34.07 |
Random-guessing benchmark | 41.90% | — |

The three ensemble-learning models produced very similar route-level hit rates, ranging from 85.56% to 85.71%. LightGBM achieved the highest value at 85.71%, followed closely by XGBoost at 85.61% and random forest at 85.56%. The differences among these three models were therefore marginal. In contrast, all three ensemble models achieved higher hit rates than the conventional choice and logistic models. LightGBM exceeded MNL and NL by approximately 7.1 percentage points and binary logistic regression by approximately 9.7 percentage points.
These results indicate that the main predictive advantage arises from the ability of ensemble-learning models to represent nonlinear relationships and interactions among route and trip characteristics rather than from a substantial performance difference among the individual ensemble algorithms. LightGBM was retained for the subsequent interpretation and sensitivity analyses because it combined competitive predictive performance with computational efficiency and direct calculation of gain-based feature importance.
For the fitted LightGBM model, feature importance was measured using gain, defined as the cumulative reduction in the model loss attributable to splits involving a given feature. Gain-based importance therefore reflects the contribution of each feature to the model's predictive discrimination rather than a causal effect on travellers' behaviour. Figure 6 presents the 12 features with the highest gain values, while Table 9 and Figure 7 aggregate the results by feature dimension.

Physical Dimension | No. of Features | Gain Share/% | Physical Dimension | No. of Features | Gain Share/% |
|---|---|---|---|---|---|
Mileage | 5 | 65.64 | Switching | 2 | 3.43 |
Toll | 6 | 13.06 | Operating level | 2 | 3.25 |
Travel time | 5 | 5.56 | Traveller | 4 | 3.08 |
Choice set | 3 | 5.44 | Multi-dimension dominance | 1 | 0.54 |

At the individual-feature level, mileage lead accounted for the largest gain share (40.16%), followed by detour ratio (13.60%), mileage difference (9.95%), toll lead (4.36%), toll per kilometre (3.40%), and toll difference (2.30%). Because several related variables represent different aspects of the same underlying route characteristic, individual feature importance alone does not provide a direct comparison among physical dimensions. The feature-level results were therefore aggregated according to their corresponding dimensions.
Mileage-related features accounted for 65.64% of the total gain, followed by toll-related features at 13.06% and travel-time features at 5.56%. Choice-set characteristics accounted for 5.44%, switching for 3.43%, operating-level variables for 3.25%, traveller attributes for 3.08%, and multi-dimension dominance for 0.54%.
The prominence of mileage-related variables indicates that relative distance characteristics played a major role in the model's discrimination among candidate routes. Toll-related features ranked second and are particularly relevant from an operational perspective because toll is directly adjustable by the expressway operator, whereas route mileage is largely fixed by network structure in the short term. Toll was therefore selected as the principal intervention variable for the scenario-based sensitivity analysis. The following analysis examines whether predicted responses to toll changes vary with the competitive position of a route, choice-set size, and vehicle type.
The sensitivity analysis examined how predicted route choice probabilities changed under hypothetical toll adjustments. Network structure, choice-set composition, and trip attributes were held constant, while the toll of the target route or competing routes was modified according to the specified scenario. All toll-dependent features were then recalculated and the fitted LightGBM model was applied again to obtain the corresponding route choice probabilities.
The resulting changes should be interpreted as model-based scenario responses under otherwise unchanged conditions rather than as causal estimates of travellers’ behavioural responses or network-wide equilibrium effects.
For each of the 32,473 observed routes in the test set, the toll of the target route was increased independently by 10%, 20%, 30%, and 50%, while the attributes of the other candidate routes and all trip-level variables were held constant. Table 10 reports the median change in predicted choice probability across baseline-probability groups together with selected structural characteristics of the routes in each group.
Baseline Probability | Routes | Share/% | Median Probability | Mean Dominance Count | Shortest-Route Share/% | +10% Toll | +30% Toll | +50% Toll | Probability Decrease at +50% Toll/% |
|---|---|---|---|---|---|---|---|---|---|
$<$50% | 5440 | 16.8 | 24.25% | 1.30 | 22.2 | +0.02 | $-$0.77 | $-$3.31 | 68.1 |
50%–70% | 3204 | 9.9 | 61.09% | 1.67 | 51.2 | $-$5.99 | $-$9.91 | $-$16.99 | 87.0 |
70%–80% | 2342 | 7.2 | 75.64% | 2.13 | 68.1 | $-$4.86 | $-$9.31 | $-$14.84 | 84.7 |
80%–90% | 5195 | 16.0 | 86.26% | 2.58 | 84.9 | $-$1.58 | $-$4.04 | $-$6.05 | 82.8 |
$\geq$90% | 16292 | 50.2 | 95.10% | 2.92 | 96.5 | $-$0.46 | $-$1.15 | $-$1.84 | 85.0 |
Actual travel routes | 32473 | 100.0 | 90.06% | 2.41 | 75.7 | $-$0.71 | $-$1.93 | $-$3.33 | 82.0 |
The response to toll adjustment varied substantially with the initial competitive position of the target route. The largest median reductions occurred among routes with baseline probabilities between 50% and 80%. For the 5,546 routes in this range, representing 17.1% of the test sample, a 30% toll increase reduced the median predicted probability by 9.31–9.91 percentage points, while a 50% increase produced reductions of 14.84–16.99 percentage points.
Routes with baseline probabilities of at least 90% behaved differently. This group contained 16,292 routes, or 50.2% of the test sample, and had a median baseline probability of 95.10%. Among these routes, 96.5% were also the shortest-distance alternatives. A 50% toll increase reduced their median predicted probability by only 1.84 percentage points. Their relatively limited response is consistent with the strong competitive position of these routes across several attributes, which reduces the extent to which a change in toll alone alters their overall position within the choice set.
At the other end of the distribution, the 5,440 routes with baseline probabilities below 50% also exhibited comparatively small median responses. Only 22.2% of these routes were the shortest-distance alternatives, and they were dominant in an average of 1.30 dimensions. A 10% toll increase produced a median change of +0.02 percentage points, while a 50% increase reduced the median probability by 3.31 percentage points. Taken together, these results reveal a non-monotonic relationship between baseline competitive position and model-predicted toll sensitivity, with the strongest responses concentrated among routes occupying an intermediate competitive position.
Choice-set size was also associated with the magnitude of the predicted response. For choice sets containing two, three, four, and five routes, a 50% increase in the target-route toll reduced the median predicted probability by 1.96, 5.91, 8.54, and 8.25 percentage points, respectively. The response therefore generally increased as the number of available alternatives rose from two to four, before declining slightly for five-route choice sets. Within the modelled scenarios, toll adjustments consequently produced larger probability shifts where travellers had a broader set of competing alternatives.
These response patterns across baseline probability bands and choice-set sizes are further illustrated in Figure 8.


A second scenario kept the toll of the target route unchanged while reducing the tolls of the other candidate routes within the same choice set. This design allowed the predicted response to an increase in the target-route toll to be compared with the response generated by making competing routes less expensive.
Across all 32,473 observed routes, reductions of 10%, 30%, and 50% in the tolls of competing routes decreased the median predicted probability of the target route by 0.78, 2.98, and 7.55 percentage points, respectively. The corresponding changes produced by increasing the target-route toll were 0.71, 1.93, and 3.33 percentage points.
Across the full range of adjustment magnitudes reported in Table 11, reductions in competing-route tolls produced larger median changes than equivalent increases in the target-route toll. At adjustment levels of 10%, 20%, 30%, and 50%, the respective ratios were 1.10, 1.24, 1.54, and 2.27. The differences were relatively small at 10% and 20% but became more pronounced at 30% and 50%. Within the present model, changing the costs of several competing alternatives therefore produced a larger shift in the target route's predicted probability than changing the target route alone, particularly under larger toll adjustments. The route-level responses under the two toll-adjustment strategies are compared in Figure 9.
Scenario | Item | 0 | 10% | 20% | 30% | 50% |
|---|---|---|---|---|---|---|
Toll increase of the target route | Median probability change/percentage points | 0.00 | $-0.71$ | $-1.28$ | $-1.93$ | $-3.33$ |
Toll reduction of the other routes | Median probability change/percentage points | 0.00 | $-0.78$ | $-1.59$ | $-2.98$ | $-7.55$ |

The response also differed by vehicle type, as shown in Table 12.
Vehicle type | Routes | Median probability | +10% toll | +30% toll | +50% toll | $-$10% toll | $-$30% toll | $-$50% toll | Decrease at +50% toll/% | Decrease at $-50$% toll/% |
|---|---|---|---|---|---|---|---|---|---|---|
Car | 24,260 | 89.64% | $-$0.65 | $-$1.67 | $-$2.87 | $-$0.62 | $-$2.55 | $-$6.77 | 80.3 | 84.7 |
Truck | 8,213 | 90.97% | $-$0.94 | $-$3.07 | $-$5.11 | $-$1.28 | $-$4.27 | $-$10.23 | 87.0 | 90.9 |
For trucks, a 50% increase in the target-route toll reduced the median predicted probability by 5.11 percentage points, compared with 2.87 percentage points for cars. Under a 50% reduction in competing-route tolls, the corresponding decreases were 10.23 and 6.77 percentage points. The truck subsample therefore exhibited larger model-predicted responses to toll changes across the examined scenarios. This difference may partly reflect the higher absolute toll exposure of trucks, although the present analysis does not separately identify the behavioural mechanisms underlying the vehicle-type difference.
The scenario results provide several implications for the use of differentiated tolling as a data-driven expressway management instrument.
First, the modelled response was strongest for routes occupying an intermediate competitive position. A 50% toll increase for routes with baseline probabilities of 50%–80% reduced their median predicted probabilities by 14.84–16.99 percentage points, whereas the corresponding reduction was only 1.84 percentage points for routes with baseline probabilities of at least 90%. This suggests that differentiated tolling may produce greater route-level shifts when applied to corridors where competing alternatives already have relatively similar levels of attractiveness.
Second, vehicle type should be considered when evaluating differentiated toll scenarios. Trucks exhibited larger predicted probability changes than cars under both target-route toll increases and competing-route toll reductions. The result indicates that a uniform toll adjustment may generate heterogeneous responses across vehicle classes and that vehicle-specific effects should therefore be assessed when designing or evaluating differentiated toll schemes.
Third, the simulations indicate that reducing the tolls of competing routes can produce larger predicted shifts than increasing the toll of the target route by the same proportion, particularly at larger adjustment magnitudes. For example, at 50%, the median probability reduction was 7.55 percentage points when competing-route tolls were reduced, compared with 3.33 percentage points when the target-route toll was increased. This finding provides a quantitative basis for comparing alternative toll-differentiation strategies before implementation.
Fourth, the magnitude of the predicted response depended on both the size of the toll adjustment and the number of alternatives in the choice set. In the present scenarios, differences between the two toll-adjustment strategies were relatively limited at 10% and 20% but became more pronounced at 30% and 50%. Larger responses were also observed for OD pairs with four or more candidate routes than for those with only two. These findings suggest that adjustment magnitude and the structure of available route alternatives should be considered jointly rather than applying a uniform toll-differentiation rule across the network.
These implications are conditional on the modelled scenarios and should not be interpreted as direct causal estimates or prescriptive toll-setting rules. Their primary value lies in demonstrating how a data-driven route choice model can be used to screen alternative operational strategies and identify route groups for more detailed evaluation before implementation.
Several limitations define the scope of these findings. First, the current feature system does not incorporate dynamic variables such as real-time traffic conditions and weather, while temporal variation is represented primarily through departure time and peak/off-peak conditions. Second, the revealed choice sets are derived from historically observed routes and therefore do not include previously unobserved alternatives that may become relevant under temporary traffic restrictions, roadworks, or other network changes. Third, the sensitivity analysis represents a partial-equilibrium scenario assessment in which network structure, choice-set composition, and travel demand are held constant. The resulting probability changes should therefore not be interpreted as causal estimates of actual behavioural responses or as network-equilibrium effects. More general road-pricing considerations, including second-best pricing, also remain outside the present framework[37], [38].
Future research can address these limitations in several directions. Dynamic traffic states and weather information can be incorporated to improve the temporal representation of route conditions. Toll and travel-time responses can also be placed on a common behavioural scale by drawing on established research on the value of travel time [39], [40], [41]. Causal inference designs based on observed toll changes would provide a stronger basis for identifying behavioural responses to pricing interventions, while integration with a traffic assignment framework [7] would allow route-level probability changes to be propagated through the network and evaluated in terms of system-wide traffic effects. These extensions would move the present framework from scenario-based route choice prediction toward more comprehensive intelligent decision support for dynamic expressway management.
5. Conclusions
This study developed a data-driven framework for route choice prediction and differentiated toll analysis by integrating large-scale expressway transaction data with ensemble learning. Actual travel routes were reconstructed from toll transaction records, ETC gantry records, and network topology, and each “trip $\times$ candidate route” pair was treated as an individual modelling observation. The resulting framework connected route choice prediction, feature interpretation, error diagnosis, and scenario-based toll sensitivity analysis, providing a quantitative basis for data-driven expressway operation and traffic management.
The LightGBM model achieved a route-level hit rate of 85.71% for the 32,473 test trips, exceeding the random-guessing benchmark of 41.90% by 43.81 percentage points. It also achieved higher hit rates than MNL (78.63%), nested logit (78.59%), and binary logistic regression (75.97%), while its performance was comparable to that of random forest (85.56%) and XGBoost (85.61%). Prediction performance varied substantially with choice-set size: the hit rate decreased from 92.24% for two-route choice sets to 58.02% for five-route choice sets, although the latter remained well above the corresponding random benchmark of 20%. Analysis of the 4,641 mis-predicted trips further showed that errors were concentrated among less frequently selected routes. The median historical share of the observed route was 92.59% for correctly predicted trips but only 20.00% for mis-predicted trips, and 66.90% of all mis-predictions involved routes with historical shares below 30%. These findings indicate that choices deviating from recurrent OD-level route patterns are more difficult to infer from the available route and trip attributes.
The feature analysis showed that mileage-related variables accounted for 65.64% of the total model gain, followed by toll-related variables at 13.06% and travel-time variables at 5.56%. Because toll is directly adjustable in expressway operations, it was subsequently used as the intervention variable in the scenario-based sensitivity analysis. The predicted response to toll adjustment varied markedly with the competitive position of the target route. For routes with baseline choice probabilities of 50%–80%, which accounted for 17.1% of the test sample, a 50% toll increase reduced the median predicted probability by 14.84–16.99 percentage points. By comparison, routes with baseline probabilities of at least 90%, representing 50.2% of the sample, showed a median reduction of only 1.84 percentage points. The magnitude of the response generally increased as the choice set expanded from two to four alternatives and decreased slightly for five-route choice sets. For OD pairs with four or five candidate routes, a 50% increase reduced the median predicted probability by 8.54 and 8.25 percentage points, respectively, compared with 1.96 percentage points for two-route choice sets.
Vehicle type and the form of the toll intervention also influenced the modelled response. Under a 50% increase in the target-route toll, the median predicted probability decreased by 5.11 percentage points for trucks and 2.87 percentage points for cars. Across the examined adjustment magnitudes, reducing the tolls of competing routes produced changes 1.10–2.27 times as large as those obtained by increasing the toll of the target route by the same proportion, with the difference becoming more pronounced as the adjustment magnitude increased. Taken together, the scenario results indicate that toll-based route management is unlikely to produce uniform responses across a network. The competitive position of a route, the number of available alternatives, vehicle type, and the form and magnitude of the toll intervention all affect the predicted response and should therefore be considered jointly when alternative differentiated toll strategies are evaluated.
From an operational perspective, the proposed framework provides a means of screening toll-management scenarios before implementation rather than prescribing a universal toll-setting rule. The results suggest that routes occupying an intermediate competitive position may warrant particular attention because their predicted choice probabilities were more responsive to toll changes than those of strongly dominant routes. They also show that vehicle-specific responses and the availability of competing routes can materially alter the predicted effect of an intervention. By connecting large-scale operational data with route-level prediction and scenario analysis, the framework provides a reproducible decision-support approach for evaluating differentiated tolling and other route-management strategies in intelligent expressway operations.
Conceptualization, Y.J.S., D.X.W., and N.F.Z.; methodology, Y.J.S., L.Z., and D.X.W.; software, L.Z. and M.X.Y.; validation, M.X.Y., W.W.Q., and Y.L.C.; formal analysis, Y.J.S. and D.X.W.; investigation, M.X.Y., W.W.Q., and Y.L.C.; resources, L.H.Z. and N.F.Z.; data curation, D.X.W., M.X.Y., and L.H.Z.; writing—original draft preparation, Y.J.S. and D.X.W.; writing—review and editing, Y.J.S., L.Z., D.X.W., and N.F.Z.; visualization, Y.J.S. and L.Z.; supervision, N.F.Z.; project administration, N.F.Z.; funding acquisition, L.H.Z. and N.F.Z. All authors have read and agreed to the published version of the manuscript.
The data used to support the research findings are available from the corresponding author upon request.
The authors declare no conflicts of interest.
