New mobile apps routinely request access to your location information, address book, or other applications, and of course, websites like Amazon or Netflix track your browsing history to make personalized recommendations.
At the same time, several recent studies have shown that it is surprisingly easy to identify unidentified individuals in supposedly anonymous databases, even those containing millions of records. So, if we want to reap the benefits of data mining—such as personalized recommendations or location-based services—how can we protect our privacy?
In the latest issue of PLoS ONE, MIT researchers offer a possible answer. Their prototype system, openPDS—for storing personal data—stores data from your digital devices in a single location you specify: It could be an encrypted server in the cloud, but it could also be a computer in a locked box under your desk.
Any mobile application, online service, or big data research team that wants to use your data has to query the data warehouse, which returns only the amount of information that is required.
Shared code, not the data.
"The example I like to use is personalized music," says Yves-Alexandre de Montjoye, a graduate student in arts and sciences and the first author of the new paper. "Pandora, for example, boils down to what they call the music genome, which contains a summary of your musical tastes. To recommend a song, all it needs is the last 10 songs you listened to—just to make sure it doesn't keep recommending the same one again—and this music genome doesn't need the list of every song I've been listening to."
With openPDS, Montjoye explains, "You share code; you don't share data. Instead of sending data to Pandora to define your music preferences, Pandora sends a piece of code to you to define your music preferences, and then sends it back."
Montjoye was joined on the paper by his doctoral advisor, Alex "Sandy" Pentland, the Toshiba Professor of Media Arts and Sciences; Erez Shmueli, a postdoctoral fellow in Pentland's group; and Samuel Wang, a Foursquare software engineer who was a graduate student in the Department of Electrical Engineering and Computer Science when the research was conducted.
After an initial rollout involving 21 people who used openPDS to regulate access to their medical records, researchers are testing the system with several telecommunications companies in Italy and Denmark. Although openPDS can, in principle, run on any machine of the user's choosing, in the trials, the data is stored in the cloud.
Meaningful Permissions
One of the benefits of openPDS, Montjoye explains, is that it requires applications to specify what information they need and how it will be used. Today, he says, "when you install an application, it tells you 'this application has access to your GPS location' or 'has access to the SD card.' You as the user have absolutely no way of knowing what that means. The permissions tell you nothing."
In fact, apps often collect far more data than they actually need. Service providers and app developers don't always know in advance which data will be most useful, so they store as much as they can in case they might need it later. It could turn out, for example, that for some music listeners, the album cover is a better indicator of whether they'll like the songs, rather than anything captured by Pandora's music genome.
OpenPDS stores all potentially useful data, but in a repository controlled by the end user, not the application developer or service provider. A developer who discovers that a previously underutilized piece of information is useful must request access to it from the user. If the request seems unnecessarily intrusive, the user can simply deny it.
Of course, a malicious developer could try to manipulate the system, constructing requests that elicit more information than the user intends to reveal. A navigation app might, for example, be authorized to identify the nearest subway stop or parking garage to the user. But it shouldn't need both pieces of information simultaneously, and by requesting them, it could infer more detailed location information than the user wants to disclose.
Creating safeguards against such data leaks will have to be done on a case-by-case, application-by-application basis, Montjoye acknowledges, and at least initially, the full implications of some query combinations may not be obvious. But “even if it’s not 100 percent secure, it’s still a huge improvement over the current situation,” he says. “If we can get people access to most of their data, and if we can have the cutting-edge technology to anonymously move into interactive systems, that would be a huge victory.”
Written by Larry Hardesty, MIT News Office
