A data perspective on trip planners

A data perspective on trip planners

Stating the obvious: trip planners would not exist without the data that feeds their algorithms! As a direct consequence, data integration is a key factor that influences the scope of the trip planning service, its update, its comprehensiveness and its ability to support multimodal travel for all.

There are several approaches to data integration depending on the type of data, that we are detailling below. One main takeaway: different types of data require different integration pipelines.

The in-house geographical database

The geographical (or topological) database is not used solely for the application’s display. It also powers the search autocomplete function along as being the first feature that any user see. It is based on geographical and topographical resources that provide addresses, points of interest, base maps, etc. To build this database, journey planners generally rely on an existing base, whether collected in-house or provided by a third party. In most cases, planners prefer to manage this aspect directly as it has a significant impact on the user experience.

For stakeholders in the multimodal mobility domain, it is beneficial to enrich this map database with:

  • pedestrian routes, particularly accessible ones,
  • points of interest.

Whenever they have it, producers of multimodal mobility data should share this rich information directly with trip planners and/or upload it to open databases (e.g., OpenStreetMap) using standardised data formats.

The integration of standardised datasets

Direct integration is the most common way of feeding data into trip planner’s databases, provided the datasets are shared in standardised exchange formats. Most often, these datasets describe public transport services, whether expressed in NeTEx / GTFS Schedule for schedule data or in SIRI / GTFS Realtime for real-time data.

Each dataset is then processed individually before being integrated into the database that feeds the routing engine. This database is built on key concepts of public transport service and defines the functional scope of the trip planner, based on a data model specific to the trip planner. It may be based on a recognised conceptual model such as Transmodel. This database, which forms the foundation of the public transport component of the routing engine, typically contains:

  • stops and/or stations for shared vehicles,
  • routes,
  • timetables,
  • service schedules,
  • fares,
  • vehicle locations,
  • delays,
  • service cancellations,
  • disruptions,
  • occupancy levels,
  • etc.

Data integration is facilitated by the adoption of standardised data formats across the multimodal mobility data chain, particularly by software developers and solution providers for scheduling, offer management, branding management, CAD/AVL, passenger information management, etc. The ability to integrate standardised data significantly improves the accuracy and relevance of the trip planning results provided to travellers. It also greatly reduces the effort required to adapt the dataset to the inner data model of the routing engine before its integration.

As for the integration of shared mobility services, on-demand transport, car-sharing and private hire services, it is carried out in the same way as for conventional public transport. The only difference is that the information is often collected via an API call to the mobility operator rather than through a dataset that may have been collected or puoblished via open data portals.

What about open data in all this?

Driven by the MMTIS Delegated Regulation, open data has been increasingly used for the provision of information describing public transport (i.e., multimodal mobility data). For journey planners, open data offers two major advantages:

  • The immediate identification of available resources for integration into the journey planner’s databases,
  • An overall assessment of available public transport services within a single geographical area, to gauge the effort required to integrate the corresponding datasets.

Furthermore, open data ensures that every journey planner has access to the same information base which meets the needs of the majority of travellers. It has also helped to increase the appeal of standardised data formats for the provision of mutlimodal mobility data by public transport authorities and operators.

However, open data is far from sufficient, as journey planners still need to:

  • Transform the data to integrate it into their databases,
  • Enrich it according to their own criteria to power their routing algorithms.

Direct integration of data repositories

However, integration carried out on a dataset-by-dataset basis has drawbacks such as:

  • The resources required to collect each dataset,
  • The individual work involved in transforming each dataset before it is integrated into the trip planner’s database,
  • The failure to consolidate multimodal transport services when they fall within the same geographical area (for example, managing duplicates).

Such drawbacks can be addressed by the use of a centralised data repository (or platform) as a single source of data for a given geographical area. Such data platforms harmonise, validate, aggregate and distribute data from multiple operators before sharing it to trip planner under an aggregated form. Generally, the deployment of a multi-operator, multimodal journey planner adopts this model.

This simplifies data integration through a single point of access. It improves data governance, quality and standardisation. It also enables high scalability when adding new operators and mobility services into the supervision of the same public transport authority (geographical area).

Data integration methods are not mutually exclusive and are often combined in hybrid approaches to feed journey planners.

➡️ Why it is crucial to manage data ahead of integration in trip planners?

More articles

Lettre dans une enveloppe

Subscribe to our newsletter

Receive the latest articles and news directly in your inbox.