Dodo6/Topic_Modelling_using_LDA
1
1logistics2Article3 4Travel Time Prediction in a Multimodal Freight5Transport Relation Using Machine6Learning Algorithms7Nikolaos Servos 1, *, Xiaodi Liu 1 , Michael Teucke 2 and Michael Freitag 2,3819210311 12*13 14Bosch Connected Industry, Robert Bosch Manufacturing Solutions GmbH, Leitzstrasse 47, 70469 Stuttgart,15Germany; taylorkkxiaodiliu@gmail.com16BIBA—Bremer Institut für Produktion und Logistik GmbH, University of Bremen, Hochschulring 20,1728359 Bremen, Germany; tck@biba.uni-bremen.de (M.T.); fre@biba.uni-bremen.de (M.F.)18Faculty of Production Engineering, University of Bremen, Badgasteiner Straße 1, 28359 Bremen, Germany19Correspondence: nikolaos.servos@de.bosch.com20 21Received: 18 November 2019; Accepted: 23 December 2019; Published: 25 December 201922 232425 26Abstract: Accurate travel time prediction is of high value for freight transports, as it allows supply27chain participants to increase their logistics quality and efficiency. It requires both sufficient input data,28which can be generated, e.g., by mobile sensors, and adequate prediction methods. Machine Learning29(ML) algorithms are well suited to solve non-linear and complex relationships in the collected30tracking data. Despite that, only a minority of recent publications use ML for travel time prediction31in multimodal transports. We apply the ML algorithms extremely randomized trees (ExtraTrees),32adaptive boosting (AdaBoost), and support vector regression (SVR) to this problem because of their33ability to deal with low data volumes and their low processing times. Using different combinations of34features derived from the data, we have built several models for travel time prediction. Tracking data35from a real-world multimodal container transport relation from Germany to the USA are used for36evaluation of the established models. We show that SVR provides the best prediction accuracy, with a37mean absolute error of 17 h for a transport time of up to 30 days. We also show that our model38performs better than average-based approaches.39Keywords: logistics; supply chain management; multimodal freight transports; travel time prediction;40machine learning41 421. Introduction43An accurate travel time prediction is of high value for freight transports, as it allows supply44chain participants to increase their logistics quality [1–3]. A material planner at the receiving plant45can identify delayed deliveries in advance, and forecasts of material stocks can be optimized through46this. Furthermore, a plant can adjust its capacities in time, e.g., staff or machinery, and increase47its efficiency. Logistic service providers benefit in the same way. At warehouses, ports, or other48hubs, capacities concerning staff, ramps, forklifts, etc. can be scheduled accordingly. Consequently,49manufacturers and logistic service providers can enhance their efficiency, optimize their processes,50and increase planning accuracy [1,3,4].51Generally speaking, a supply chain involves all collaborative operations by all involved companies52to transform raw materials into the final product. This includes activities such as sourcing raw materials,53manufacturing, assembly, and distribution to the end customer. To do so, logistics for the handling,54transport, and storage of materials is required. While logistics deals with the operational level,55supply chain management deals with the planning and management of supply chains. Logistics also56 57Logistics 2020, 4, 1; doi:10.3390/logistics401000158 59www.mdpi.com/journal/logistics60 61Logistics 2020, 4, 162 632 of 2264 65includes the task to provide information between the various stages in the supply chain. Usually,66a number of companies or organizations are responsible for the transport. Especially in multimodal67transports with various transshipment points, these conditions make the supply chain significantly68more complex. [5] This work focuses on estimating travel time accurately in multimodal transports.69Accurate travel time estimates are important for efficient management of transport operations and70logistics in supply chains.71For travel time estimation, a continuous monitoring of freight transports is required, e.g., by using72mobile sensors attached to transported goods [6]. At the same time, achieving the required transparency73is one of the most challenging tasks in logistics [7]. Travel time prediction is also difficult, as many74factors, such as weather, traffic, vehicle, route, or transport relation influence it [8].75Currently, the majority of the published literature deals with passenger transportation, such as76bus arrival times or highway travel times, rather than freight transports. Hereby, only rides with77a duration of up to two hours are considered, which is not applicable to freight transports [9,10].78For travel time estimation, the authors use average-based approaches [11,12], Kalman filters [13–17],79or machine learning (ML) algorithms (see Section 2). Furthermore, in their research, the route and80stops are known in advance.81As per [9], only ML algorithms can adequately deal with complex and dynamic behavior during82transports. ML has the ability to deal better with complex and non-linear relationships between83predictors and can process complex and noisy data [10]. The literature also confirms this statement,84as ML approaches usually perform better than average-based approaches [18–20]. Despite this situation,85ML has only been applied in a minority of recent publications dealing with travel time prediction in86freight transports [1,2,21,22].87In the forecasting of multimodal freight transports based on real-time tracking data, only very88limited research has been conducted, which mainly uses average-based approaches [23]. Consequently,89this paper focuses on the evaluation of ML for long term forecasting to predict travel times of90multimodal freight transports. Tracking data is generated by sensors, which are attached to the91transported goods. The sensor data is transmitted with a low frequency in longer periods of 30 min,92due to a small battery coverage. Except for the origin and destination, no further information concerning93the exact route and trans-loading points are available beforehand. Finally, we can show that machine94learning algorithms can adequately predict travel time under the given constraints. In our use-case,95support vector regression (SVR) provides the best prediction accuracy, with a mean absolute error96of 17 h for a transport time of up to 30 days. We also show that our model performs better than97average-based approaches.98We organize the rest of the paper as follows: In Section 2, we provide an overview of literature99dealing with travel time estimation using machine learning to identify the research gap. Section 3100describes the selected learning algorithms and the framework of this work. Section 4 presents the101use-cases and the collected tracking data, which will be processed and transformed in Section 5.102Section 6 describes the process of creating the prediction model. The results are evaluated and103compared to average-based models in Section 7. In Section 8, we conclude the paper and suggest104further work in this field.1052. Literature Review106In this paper, we evaluate the capability of machine learning algorithms to predict the travel time107in multimodal transports. Most research in this realm has been conducted on travel time prediction for108segments of streets or highways and bus rides. Only a minor part of the research is related to freight109transports or similar applications.110The authors of [24] use data from loop detectors and a global positioning system (GPS) to111determine speed, occupancy, and volume of cars on highways, and include previous transport times to112predict travel times on a four-mile stretch of highway. They use artificial neural networks (ANNs)113as a learning algorithm. Similarly, the authors of [25] only include loop detectors and use support114 115Logistics 2020, 4, 1116 1173 of 22118 119vector machines to estimate travel time on highway segments of up to 350 km. Ref. [26] identifies120random forests (RF) as the best approach for predicting travel times on urban street segments using121GPS data from 300 probe vehicles. In a similar use-case to ref. [24], the authors of [27] evaluate a122gradient boosting method and RF for travel time estimation on highway segments and show that the123gradient boosting method performs best. The authors of [8] also show that gradient boosting provides124a good prediction. The authors also evaluate different features based on GPS data. They establish that125including the speed between two transmissions increases prediction accuracy by 5%. The authors126of [28] evaluate deep neuronal networks using 12 months of GPS data and the previous 12 travel times127as features. Even for a 131 km stretch, their approach results in a mean absolute error (MAE) of 3.25 min.128The above approaches do not consider any end-to-end transport relations as with freight transports.129In addition, exact advance knowledge of the route and a high frequency of GPS measurements are130required, or side-based approaches such as loop detectors are used.131Considering bus travel time prediction, nearly all authors, such as [18,29–33], use ANNs in their132research. In all cases, the route is known in advance and is separated based on the bus stops. Most of133the approaches only differ when considering the evaluated use-case and the used features. Usually,134features such as the time of departure, public holidays, dwelling time at bus stops, current travel time,135distance to destination, and average speed are used. Similarly, the authors of [20] also apply ANNs to136tramways. The authors of [31] identify ANNs as a better approach than k-nearest neighbors (kNN),137while the authors of [34] show that support vector regression (SVR) performs better. The authors of [35]138use a multilinear regression for prediction. To consider all previous stops in their model, the authors139add the cumulated number of stops and dwelling time. In contrast to the previous approaches, the bus140ride is related to an origin and destination considering an end-to-end transport relation including141stops, similar to freight logistics. However, the authors have not considered that many stops within142supply chains are often unknown to supply chain participants and can differ between transports using143the same transport relation. Furthermore, their approaches require a high number of rides and assume144a high frequency of GPS measurements.145The authors of [2] apply SVR on milk run freight transports. They comment that SVR can better146generalize information than neuronal networks. Hereby, the stops of the milk run are known in advance.147Their prediction model achieves an accuracy of three minutes. However, it requires 1800 rides for148training. The authors of [36] consider random forests (RF) as a reasonable algorithm to predict aircraft149arrival times. Flight related data such as delays, origin, destination, position, weather data, and data150from the air traffic control have been taken to consideration. The authors can prove that their estimation151is better than the one provided by air traffic control. The authors of [1] apply neuronal networks and152SVR to predict the time of arrival of container transports. The timestamp, distance to destination,153and geo-coordinates have been chosen as features, as well as weather information. The authors show154that SVR performs better than ANNs and that weather data does not have a significant influence on155the transport time. Finally, the authors of [21] evaluate kNN, SVR, and RF for arrival time prediction of156open-pit trucks. Hereby, a site-based approach is used. The position is only measured at a few discrete157nodes of the route network, and no GPS data is used. RF provides the best prediction results.158The literature review shows that none of the identified research has yet been conducted on159multimodal transports with the transported goods themselves being equipped with a sensor instead160of the transport vehicle. Furthermore, the research has also not yet considered the scenario where161transported goods are equipped with sensors with a small battery coverage that are consequently162transmitting position data with a low frequency, in this case every 30 min, in comparable distances.163The concluding algorithm has to be capable of deriving a good prediction with a lower amount of data.164In addition, the case of the route being unknown has not been considered in the research.165 166Logistics 2020, 4, 1167 1684 of 22169 1703. Methodology171The following chapter provides a description of applied machine learning algorithms and the172corresponding parameters of models based on the knowledge from the literature review. A framework173of the proposed approach is also presented.1743.1. Choice of Learning Algorithms175This research will select extremely randomized trees (ExtraTrees), adaptive boosting (AdaBoost)176and support vector regression (SVR) as representative algorithms for modelling, according to the177results of the literature review. All three algorithms are capable of adapting to complex systems and178are robust in dealing with complex and small data sets. They have shown superior performance in179previous research with low processing time [1,2,25,37].180Support vector machine (SVM) is a classification technique based on the concept of supervised181learning. It aims to find the decision boundary between data points and separate them by an optimally182derived hyperplane maximizing margin [25] (pp. 277–278). SVM is applied to regression problems as183SVR. Briefly speaking, SVR defines a weight for each variable of the model and learns the behavior184of the data through a training and testing process. The kernel trick is applied to find a nonlinear185decision boundary by creating a much higher dimension of new features transformed from existing186features [38] (pp. 282–284). A common and well performing kernel function for travel time prediction187is the radial basis function (RBF) [39] (p. 216). The other two hyper-parameters, ε and C, are important188for structuring a decision boundary. While the parameter, ε, describes an insensitive loss function that189penalizes prediction residuals larger than the value specified by it, C describes the tolerance towards190prediction errors as a regularization parameter [40] (pp. 68–73).191SVR models have shown plausible performance in recent research due to their good ability to192generalize data and guarantee global minima compared to other methods. The processing time can193also be low, given the constraints of a small data size and high problem complexity [25,39]. For these194reasons, SVR is chosen as one of the algorithms in our study.195Extremely randomized trees (ExtraTrees) is a tree-based ensemble method for both classification196and regression problems. The basic principle behind ExtraTrees is to randomly create a number of197different trees (n_estimators) with randomly chosen features. This process minimizes the variance198of the prediction results [41] (pp. 3–4). In this study, regularization parameter, random_state, is also199considered, which initiates the random number generator to randomize characteristics in trees [42]200(p. 619).201Adaptive boosting (AdaBoost) is one of the most widely used boosting approaches, based on202the concept of combining weak and inaccurate learners, mostly decision trees, to a strong and203accurate learner. In regression problems, AdaBoost acts as a meta-estimator, which uses a weighted204sum of prediction errors for a number of regressors to increase performance. These weights on205the dataset are subsequently updated based on the prediction errors of the previous regressor.206Important hyper-parameters of this method are the number of estimators (n_estimators) after which the207algorithm will stop and the learning_rate that determines the degree of the weight adjustment [43].208According to [21,44], both ExtraTrees and AdaBoost are used for travel time prediction and have209shown reasonable results with low errors. With the support of previous practical success, we also210chose these two methods to build our prediction model.2113.2. Framework212In this research, the framework has been derived from the theory of knowledge discovery in213databases (KDD) data mining process flow [45] (pp. 72–73) and is shown in Figure 1.214 215Logistics 2020, 4, x FOR PEER REVIEW216 2175 of 22218 219on the results, the best model is selected. (8) Finally, the best machine learning model is compared to220Logistics 2020, 4, 12215 of 22222commonly223used average-based approaches.224 225Figure2261. 1.227Framework228used229forfor230travel231time232prediction.233Figure234Framework235used236travel237time238prediction.239 240The241beenisapplied242using243Python with244Jupyter245notebook.246The used247Thedescribed248proposedapproach249approach,has250which251based on252the historical253tracking254data255of transports256in the257clustering258anduse-case259prediction260algorithms261have been262implemented263library264scikit-learn.265considered266described267in Section2684.1, involves269eightusing270steps:the271(1)machine272Relevant273data for274modelling275is selected and cleansed to eliminate outliers. (2) Then, a train and test set splitting is performed.2764.The277Use-Case278Data Generation279test setand280represents281unknown data that is used to evaluate the performance of the developed models,282while283the284train285set286represents287known288for training289models.290(3) To include291information292In this chapter, our use-case293anddata294theused295system296used to the297track298the transported299goods300is brieflyon301the302route,303a304clustering305approach306is307used308to309identify310stops311and312trans-loading313points.314Only315the316train set is317introduced. The generated data is also described.318used to identify the clusters. (4) Finally, features, which are considered as relevant, are derived from the319tracking320data321and the clustering results based on domain knowledge and literature. (5) In the modelling3224.1.323Use-Case324Description325phase, different methods have been used to select features and to derive different experiments as326In this use-case, real-life data derived from a multimodal supply chain extending from Bremen327feature subsets. (6) Then, for each experiment and learning algorithm, grid search and cross validation328in Germany to Vance in the United States has been used. Trucks, trains, and ships carry out the329with five train and validation sets have been used for parameter tuning. (7) The models are finally330included transports. The process is shown in Figure 2. First, the required goods are packed into a331tested and evaluated, considering the experiments and learning algorithms. Based on the results,332container at a container freight station of a logistics service provider in Bremen. Hereafter, it will be333the best model is selected. (8) Finally, the best machine learning model is compared to commonly used334picked up by truck and transported to the port of Bremerhaven. After storing the container at the335average-based approaches.336port’s container yard for several days, the container is transported by vessel to the port of Charleston337The described approach has been applied using Python with Jupyter notebook. The used clustering338via the port of Norfolk. In Charleston, the container is loaded onto a train driving to a train yard in339and prediction algorithms have been implemented using the machine library scikit-learn.340Bessemer, with several unknown intermediate stops. There, it is picked up by truck and transported341to4.itsUse-Case342final destination343inGeneration344Vance. Hereby, it should be noted that the actual stops during transport are345and Data346not limited to the stops previously described. More stops can occur during the transport that have347In this chapter, our use-case and the system used to track the transported goods is briefly348not been known by most of the supply chain participants. The total transport takes between 22 and349introduced. The generated data is also described.35030 days.3514.1. Use-Case Description352In this use-case, real-life data derived from a multimodal supply chain extending from Bremen in353Germany to Vance in the United States has been used. Trucks, trains, and ships carry out the included354transports. The process is shown in Figure 2. First, the required goods are packed into a container at a355container freight station of a logistics service provider in Bremen. Hereafter, it will be picked up by356truck and transported to the port of Bremerhaven. After storing the container at the port’s container357yard for several days, the358container359is transported360by considered361vessel to the362port of Charleston via the port363Figure3642. Transport365process of the366use-case.367of Norfolk. In Charleston, the container is loaded onto a train driving to a train yard in Bessemer,368with369several370unknown371intermediate stops. There, it is picked up by truck and transported to its final3724.2.373Data374Generation375and Description376destination in Vance. Hereby, it should be noted that the actual stops during transport are not limited377Fourty-three pallets have been distributed among seven container shipments. Those pallets have378to the stops previously described. More stops can occur during the transport that have not been known379been equipped with a hybrid sensor system, as shown in Figure 3, for a period of seven weeks for the380by most of the supply chain participants. The total transport takes between 22 and 30 days.381 382port’s container yard for several days, the container is transported by vessel to the port of Charleston383via the port of Norfolk. In Charleston, the container is loaded onto a train driving to a train yard in384Bessemer, with several unknown intermediate stops. There, it is picked up by truck and transported385to its final destination in Vance. Hereby, it should be noted that the actual stops during transport are386not limited to the stops previously described. More stops can occur during the transport that have387Logistics3882020,3894, 1known by most of the supply chain participants. The total transport takes between 22 and6 of 22390not391been39230 days.393 394Figure395Transportprocess396process of397of the398Figure3992. 2.400Transport401the considered402considereduse-case.403use-case.404 4054.2. Data406Generation407andand408Description4094.2. Data410Generation411Description412Fourty-three413pallets414have415been416distributedamong417among seven418shipments.419Those420pallets421have have422Fourty-three423pallets424have425been426distributed427sevencontainer428container429shipments.430Those431pallets432been433equipped434with435a436hybrid437sensor438system,439as440shown441in442Figure4433,444for445a446period447of448seven449weeks450for451the for452been equipped with a hybrid sensor system, as shown in Figure 3, for a period of seven weeks453Logistics 2020, 4, x FOR PEER REVIEW4546 of 22455the purpose of tracking. Therefore, each pallet within the container has been equipped with a sensor.456The sensor457is associated458with transport459related460information461from462label463attached464to athe465palette,466purpose467of tracking. Therefore,468each pallet469within470the container471hasabeen472equipped473with474sensor.475The e.g.,476the serial477number.478This ‘pairing479process’480is performed481The sensor482is measuring483sensor484is associated485with transport486related487informationwith488fromaascan.489label attached490to the491palette, e.g., quality492the493serial494number.495This496‘pairing and497process’498is performed499with500a scan. The501is measuring502related503data,504such as505humidity506temperature,507and508transmits509this sensor510data together511withquality512its ID via513related514such515as humidity516and517temperature,518and transmits519data called520together521with its ID522via has523Bluetooth524lowdata,525energy526(BLE).527Then, the528sensor529data is transmitted530to athis531device532a gateway,533which534Bluetooth535low536energy537(BLE).538Then,539the540sensor541data542is543transmitted544to545a546device547called548a549gateway,550which551been attached on the container and is receiving the sensor data at a predefined interval. Due to the552has been attached on the container and is receiving the sensor data at a predefined interval. Due to553long transport554and a restricted battery life, sensor measurements are received and transmitted only555the long transport and a restricted battery life, sensor measurements are received and transmitted556every 30 min. The gateway then adds geo-positioning data to the received sensor data using GPS and557only every 30 min. The gateway then adds geo-positioning data to the received sensor data using558transmits559the data to the Bosch IoT Cloud, a central data cloud system, through the global system560GPS and transmits the data to the Bosch IoT Cloud, a central data cloud system, through the global561for mobile562(GSM) network.563If no GSM564is available,565the sensor566data is567systemcommunications568for mobile communications569(GSM) network.570If noconnection571GSM connection572is available,573the sensor574buffered575gateway576together577with578the GPS579untildata580a connection581for transmission582is available583dataon584is the585buffered586on the587gateway588together589withdata590the GPS591until a connection592for transmission593is594again.available595All transmitted596is storeddata597in the598cloud for599later600retrieval601andretrieval602analysis.603The604gateways605again. Alldata606transmitted607is stored608in the609cloud610for later611and612analysis.613The can614gateways615can also bewith616usedfixed617stationary618with fixed geo-coordinates.619endtransport,620of each transport,621also be622used stationary623geo-coordinates.624During the During625end of the626each627the link to628the629link630to631the632transported633unit634load635is636removed637automatically638using639BLE640via641this642type643of stationary644the transported unit load is removed automatically using BLE via this type of stationary645gateway.646gateway.647processas648is ‘un-pairing’.649described as ‘un-pairing’.650This process651is This652described653 654Figure6553. Architecture656system.657Figure6583. Architectureofofthe659theused660used tracking661tracking system.662 663BasedBased664on the665described666dataattributes667attributes668shown669in Table670on previously671the previously672describedsystem,673system,the674the collected675collected data676areare677shown678in Table6791. 1.680The system681provides682table683of tracking684for each685between686the pairing687un-pairing688The system689provides690a tablea of691tracking692tracestraces693for each694unit unit695loadload696between697the pairing698andand699un-pairing700events.701events.702Table 1. Collected tracking data from the tracking system.703Table 1. Collected tracking data from the tracking system.704 705Column706Column707UnitUnit708loadload709ID ID710Start/End of tracking711Start/End of tracking712Gateway ID713Gateway ID714Last Update715Last716Update717Latitude,718longitude719Latitude,720longitude721Temperature,722humidity723Temperature, humidity724 725Description726Description727Uniqueidentifier728identifierofofpaired729pairedtransport730transportunit731unitload732load733Unique734Pairing and un-pairing Unix time in seconds735Pairing and un-pairing Unix time in seconds736Unique identifier of transmitting gateway737Unique identifier of transmitting gateway738Unix time of transmission in milliseconds739Unix time740of transmission741in milliseconds742Geo-location743of the last744update745Geo-location746ofof747the748last update749Quality750related data751sensor752of last update753Quality related data of sensor of last update754 755Apart from the sensor data, information considering the origin and destination location in the756form of a geo-fence is provided. The geo-fence consists of a geo-coordinate and a radius.7575. Data Pre-Processing758 759Logistics 2020, 4, 1760 7617 of 22762 763Apart from the sensor data, information considering the origin and destination location in the764form of a geo-fence is provided. The geo-fence consists of a geo-coordinate and a radius.7655. Data Pre-Processing766In this chapter, two major tasks in the research will be presented: Data cleansing and feature767engineering. Since GPS data can often be quite noisy, data pre-processing is required to increase the768quality of the used data. Further, relevant features are derived from the cleansed data in order to build769an accurate prediction model.770Logistics 2020, 4, x FOR PEER REVIEW771 7727 of 22773 7745.1. Data Cleansing7755.1. Data Cleansing776After a qualitative analysis of the sensor data, the data is cleansed based on the following rules:777After a qualitative analysis of the sensor data, the data is cleansed based on the following rules:778•779If the test pairings have no GPS transmissions falling within its destination geo-fence, they will be780•781If the test pairings782have783nothe784GPS785transmissions786within787its destination geo-fence, they will788removed.789This means790that791goods792have neverfalling793reached794the destination.795◦796be797removed.798This799means800that801the802goods803have804never805reached806the807destination.808•809If transmissions have a latitude and longitude set to 0 , the transmission810will be removed and811•812If813transmissions814have815a816latitude817and818longitude819set820to8210°,822the823transmission824will be removed and825ignored as it implies the absence of a GPS signal.826ignored827as828it829implies830the831absence832of833a834GPS835signal.836•837While storing the containers at the ports and at other intermediate stops, outliers have been detected.838•839While840storing841the containers842the ports843and844at other intermediate845stops,846outliers847been848To849identify850outliers,851the speedatbetween852two853consecutive854stops has been855used,856whichhave857exceeded858detected.859To860identify861outliers,862the863speed864between865two866consecutive867stops868has869been870used,871which872120 kmph in this case. Case 1 in Figure 4 shows one outlier where the speed exceeds more than873exceeded874kmph875in thisand876case.877Case 1 in878Figure 4 shows879onesame880outlier881where882the speed883exceeds884120885kmph 120886to the887previous888following889transmission.890At the891time,892the speed893between894the895more896than897120898kmph899to900the901previous902and903following904transmission.905At906the907same908time,909the910speed911transmissions before and after the outlier is significantly smaller, with a speed below 5 kmph.912between913transmissions914before andoutliers915after the916outlier917significantly918with aclose919speed920In921Case 2 the922of Figure9234, two consecutive924have925been is926detected,927whichsmaller,928are relatively929to930below9315932kmph.933In934Case9352936of937Figure9384,939two940consecutive941outliers942have943been944detected,945which946are947each other. The speed between the outliers is below 5 kmph, while the speed to the previous and948relatively transmission949close to each again950other.exceeds951The speed952outliers953below9545 kmph,955whilethe956theoutliers957speed958following959120 between960kmph. Asthe961before,962the is963speed964before965and after966to967the968previous969and970following971transmission972again973exceeds974120975kmph.976As977before,978the979speed980before981is below 5 kmph. The transmissions identified as outliers based on the described rules have been982and after as983theshown984outliers985is below986removed,987in Figure9884. 5 kmph. The transmissions identified as outliers based on the989described rules have been removed, as shown in Figure 4.990•991As only the transport itself should be considered, data points other than the last transmission in the992•993As only the transport itself should be considered, data points other than the last transmission in994origin geo-fence and the first transmission in the destination geo-fence are removed. This situation995the origin geo-fence and the first transmission in the destination geo-fence are removed. This996is shown in Figure 4, Case 3.997situation is shown in Figure 4, Case 3.998 999Figure 4. Data cleansing of the raw data.1000 1001In addition, one has to consider that we have received tracking data from a number of pallets1002per container. To avoid using the data from the same ride several times in the training1003training and1004and test1005test set,1006set,1007only the tracking data of one pallet per1008per transport1009transport is1010is considered.1011considered.10125.2.1013Clustering for1014for Route1015Route Identification1016Identification10175.2. Clustering1018As1019our1020approach1021is only1022basedbased1023on real-time1024tracking1025data from1026thefrom1027material1028As previously1029previouslymentioned,1030mentioned,1031our1032approach1033is only1034on real-time1035tracking1036data1037the1038flow.1039This1040means1041that1042no1043information1044about1045the1046route1047and1048possible1049stops1050or1051trans-loading1052points1053is1054material flow. This means that no information about the route and possible stops or trans-loading1055available.1056Thus, toThus,1057include1058route information1059in the prediction1060model,model,1061those locations1062have to1063be1064points is available.1065to include1066route information1067in the prediction1068those locations1069have1070identified1071based1072on1073the1074geo1075data1076provided.1077To1078do1079so,1080a1081clustering1082algorithm1083is1084required1085that1086can1087deal1088to be identified based on the geo data provided. To do so, a clustering algorithm is required that can1089 1090deal with geo data and noise without knowing the number of clusters. A suitable and proven1091algorithm for this purpose is density based spatial clustering of applications with noise (DBSCAN),1092[46] (p. 39) and [47].1093Similar to [46] (p. 39), before applying the clustering algorithm on all tracking data, potential1094stop points are determined to decrease noise in the identified clusters. In a previously performed1095 1096Logistics 2020, 4, 11097 10988 of 221099 1100with geo data and noise without knowing the number of clusters. A suitable and proven algorithm1101for this purpose is density based spatial clustering of applications with noise (DBSCAN), [46] (p. 39)1102and [47].1103Similar to [46] (p. 39), before applying the clustering algorithm on all tracking data, potential stop1104points are determined to decrease noise in the identified clusters. In a previously performed experiment,1105a gateway had been attached to a container, which was standing at the same position for two days.1106Hereby, noise of up to 300 m has been identified in the GPS data. Similarly, only tracking points with1107a distance smaller than 300 m either to the previous or next transmission are chosen for clustering.1108This rule applies to all points, P1 to P7 . Further, to identify only significant stops, at least four1109consecutive points have to be identified, with a distance below 300 m between each other on a ride.1110This1111rule1112applies1113only1114toREVIEW1115points P1 to P5 . This circumstance is shown in Figure 5.1116Logistics11172020,11184, x FOR1119PEER11208 of 221121 1122Figure 5. Identification of potential stop points.1123 1124Before1125and1126test1127setset1128splitting1129is required,1130as the1131testtest1132set1133Before applying1134applyingthe1135theDBSCAN1136DBSCANalgorithm,1137algorithm,a atraining1138training1139and1140test1141splitting1142is required,1143as the1144represents1145datadata1146unknown1147to the1148andand1149should1150not be1151to identify1152the clusters.1153The The1154test1155set represents1156unknown1157to model1158the model1159should1160notconsidered1161be considered1162to identify1163the clusters.1164set1165is1166also1167relevant1168for1169evaluation1170purposes,1171which1172are1173performed1174in1175Section11767.1177As1178some1179features1180test set is also relevant for evaluation purposes, which are performed in Section 7. As some features1181depend1182on values1183valuesof1184ofprevious1185previousdata1186datapoints1187pointsinin1188a transport,1189data1190is split1191by unit1192andby1193not1194by1195depend on1196a transport,1197thethe1198data1199is split1200by unit