July 2014

ORNL Report

Connected Vehicle Data Privacy Assessment – Leesburg Research Data Exchange Dataset

By:
Carter, Jason M; Lamb, Logan M
Publication Date:
July 21, 2014

Abstract

This document outlines an attempt to re-identify the driver, vehicle, and vehicle owner that generated the Basic Safety Messages (BSM) contained in the Research Data Exchange Leesburg dataset. Seven different analysis techniques were explored, and publicly available web-based search tools were used to gather the auxiliary information necessary for the re-identification task. Using the available documents and data, the auxiliary information, and open source tools, the driver, driver’s residence, and vehicle’s identification number (VIN) were re-identified. Much of this analysis was labor-intensive, manual investigation, but time and motivation are something adversaries have in abundance. Finding data relationships, or data linkages, is the foundation of re-identification; almost any piece of information can link two ostensibly de-identified data sources. Quasi-identifiers, those data attributes that create linkages, establish an information chain between disparate data sources; the composition of the attributes in these sources can create a resource that makes re-identification possible. Although singular, unique identifiers exist, almost anything can be classified as “personally identifying information” when the right quasi-identifier and auxiliary dataset is found. Some of our success could have been hindered if a more consistent de-identification method were used on the individual trips. More specifically, trip truncation methods that fail to introduce future geopoint uncertainty are potentially ineffective. The methods used on the Leesburg dataset were effective in some cases and ineffective in others. The challenge of perfect de-identification cannot be understated. The re-identification and de-identification efforts are imbalanced: The re-identifier only requires one incorrectly obfuscated trip and the right auxiliary information to successfully identify a driver; the de-identifier must obfuscate all trips while accounting for all relevant auxiliary information. This work initiated examination of more comprehensive methods of analyzing geotrack data. In particular, BSM data contains a very rich set of attributes; there may be details that have not been explored sufficiently to determine how helpful they would be to the re-identifier in a multi-vehicle, multi-driver environment. This work did not reveal a silver-bullet re-identification algorithm for an arbitrary collection of BSM data; this is good news. The successes achieved, however, do reveal some significant privacy concerns. The Internet has become our own worst enemy. Most people continue to contribute to their web footprint, and this growing footprint persists; this makes privacy assurance a significant challenge. The lessons learned during re-identification of the Leesburg driver and vehicle provide insight into several de-identification challenges; some of these problems can be solved, but the most significant problem is the growing quantity of auxiliary information that can be coupled with geotrack data to discern driver behavior. When we approach the task of data de-identification, we must understand what tasks we would like researchers to accomplish; these tasks determine the method and degree to which the data must be de-identified and whether it is warranted at all.


Related Researchers