July 2014

ORNL Report

Connected Vehicle Data Privacy Assessment – Personally Identifiable Information Analysis of Research Data Exchange Data

By:
Carter, Jason M; Lamb, Logan M
Publication Date:
July 21, 2014

Abstract

The goal of this portion of the Connected Vehicle Data Privacy Assessment was to ensure the existing RDE is as free as possible of data artifacts that could expose the identity of an individual driver or that driver’s vehicle; it also aided the development of more effective de-identification algorithms. Nine data environments were reviewed and several very specific ways to improve the privacy preserving features of the RDE have been provided in separate reports to data technicians. This report describes the technical how and why of our process. Visual inspection of plotted trips remains the most effective way of finding trip artifacts that could expose an individual’s identity or travel behavior; however, several automated methods are emerging that help focus this effort. Automatic detection helps with automatic removal. The goal of the data manager is to work toward providing an “evidence of absence” guarantee in terms of identifying characteristics; this involves proving the absence of any person’s ability to re-identify a single driver and vehicle using any auxiliary information available today and in the future. The re-identifier wins by re-identifying a single person. The tasks are truly disproportionate. With that in mind, we have found two important ways to improve de-identification: remove certain vehicle characteristics and their connections to BSM-based trips and integrate map data into the de-identification process. The former removes a surprisingly valuable personal link between drivers and the cars they choose to drive; the later facilitates a means to introduce uncertainty into de-identified trips thereby obfuscating them from re-identifiers. For the Safety Pilot One-Day Sample (SPODS), seven analysis techniques were used to find trip segments that could lead to identifying a specific driver and vehicle. Trip beginning/end segments, travel regions, and vehicle dynamics are the most useful data characteristics to the re-identifier. One of the most difficult and interesting aspects of our analysis was understanding the connections between the data generated by different sensor types (e.g., an RSEs and VADs); the de-identification of one sensor’s data could be exposed by another sensor’s data. De-identification should remain consistent across all file types in a data environment, but file schemas and purpose vary from experiment to experiment and device to device, and this heterogeneity makes normalization interesting. Using the beginnings of trips that had been “hidden” from de-identification algorithms by normal but unexpected GPS device behavior, two possible individuals were re-identified during our analysis of the SPODS; sensors and collection devices are error-prone and unreliable – this can become advantageous to the re-identifier. We were unable to re-identify with any certainty a driver or vehicle using the San Diego data, but there were trips that seemed to escape normal de-identification. The trips generated by a single driver in the San Diego data set were not connected; connecting these trips makes re-identification easier. We developed several rules to help connect trips, but ultimately how the data was aggregated made connecting trips possible; it is essential that the archival process not introduce features that help the re-identifier. In the NCAR dataset, two driver identities were possibly discovered because they deviated from a prescribed route. The Pasadena dataset contained image data; the most important data characteristic that protected privacy was poor image resolution. Some distinct vehicle types can be identified, and license plate silhouettes seen, but language and facial features are obscured due to resolution. Image data can be linked to trip data using geoposition and time. Many of the files in the RDE provide a way to measure traffic throughput and road characteristics; these files do not present a threat to individual privacy. Level of re-identification effort remains difficult to quantify; however, we estimate the 5990 trips in the SPODS would take 200 hours to visibly inspect to determine re-identification candidates. This is after an experienced person has spent approximately 24 hours pre-processing the data. With a pool of candidates, the work required to gather auxiliary information and build a re-identification case varies widely depending on geolocation and the quality of the auxiliary information sources. Outside of the context of the RDE, it is worth examining how this type of data may present a risk to group privacy. Although individual privacy may be protected by precise de-identification methods, group behaviors might remain inferable. In a somewhat tangentially related way, intra-trip drop off points may not present a direct privacy problem for the driver, but they could present a mis-identification problem. To the extent possible, we think all trips to residences should be removed from an open dataset. This report concludes with a number of recommendations based on our research of related areas and working with RDE data. Keeping as much of the data collected during connected vehicle pilots and experiments as possible is important as researchers find new ways to use connected vehicle technology to improve the safety and efficiency of our roads.


Related Researchers