As more personal data moves online, a more pressing concern may be its accidental misuse by individuals authorized to access it. Every month seems to bring another story of private information accidentally leaked by government agencies or providers of digital products or services.

At the same time, stricter access restrictions could undermine data sharing. Coordination between agencies and providers could be key to quality care; you might want your family to be able to share the photos you post on a social media site.
Researchers at the Decentralized Information Group (DIG) at MIT’s Computer Science and Artificial Intelligence Laboratory (CSAIL) believe the solution may be transparency rather than secrecy. To that end, they are developing a protocol they call “HTTP with Responsibility,” or HTTPA, that will automatically track the transmission of private data and allow the data owner to examine how it is being used.

At the IEEE Conference on Privacy, Security and Trust in July, Oshani Seneviratne, an MIT graduate student in electrical and computer engineering, and Lalana Kagal, a senior research scientist at CSAIL, will present a paper that provides an overview of HTTPA and presents a sample application, involving a health care records system that Seneviratne implemented on the experimental PlanetLab network.

The DIG is led by Tim Berners-Lee, the inventor of the Web and Founding Professor of Engineering at MIT, and shares office space with the World Wide Web Consortium (W3C), the organization, also led by Berners-Lee, that oversees the development of internet protocols such as HTTP, XML, and CSS. The DIG's role is to develop new technologies that leverage these protocols.

With HTTPA, each piece of private data is assigned its own Uniform Resource Identifier (URI), a key component of the Semantic Web, a new set of technologies championed by the W3C that would transform the Web from essentially a collection of searchable text files into a giant database.
Remote access to a web server would be much more tightly controlled than it is now, through passwords and encryption. Every time the server transmits a piece of sensitive data, it would also send a description of the restrictions on the data's use. It would initiate the transaction, using only the URI, somewhere in a network of encrypted, special-purpose servers.
HTTPA would be voluntary: It would be up to software developers to adhere to its specifications when designing their systems. But HTTPA compliance could become a selling point for companies that offer services handling private data.
"It's not that difficult to transform an existing website into an HTTPA-aware website," says Seneviratne. "In each HTTP request, the server must say, 'OK, here are the usage restrictions for this resource,' and log transactions on the special purpose server network."

An HTTPA-compliant program also incurs certain responsibilities if data supplied by another HTTPA-compliant source is reused. Suppose, for example, that a consultant in a physician network wants to access data created by a patient's primary care physician and augment that data with their own notes. Their system would then create its own record, with its own URI. But using standard Semantic Web techniques, it would mark that record as "derivatives" of the PCP record and label it with the same usage limits.
The network of servers is where the heavy lifting happens. When the data owner requests an audit, the servers work through the chain of derivations, identifying everyone who has accessed the data and what they have done with it.

Seneviratne uses a technology known as distributed hash tables—the technology at the heart of peer-to-peer networks like BitTorrent—to distribute transaction logs across servers. Redundantly storing the same data on multiple servers serves two purposes: First, it ensures that if some servers go down, the data will still be accessible. Second, it provides a way to determine if someone has tried to manipulate the transaction logs for a particular piece of data—such as deleting a record of illicit use. A server whose logs differ from those of its peers would be easy to find.
To test the system, Seneviratne built a rudimentary healthcare records system from scratch and populated it with data supplied by 25 volunteers. He then simulated a set of transactions—pharmacy visits, referrals to specialists, use of anonymized data for research purposes, and the like—that the volunteers reported had occurred over the course of a year.
Seneviratne uses 300 servers at PlanetLab to store transaction logs; in experiments, the system efficiently tracked data stored across the network and handled the inference chains needed to audit the propagation of data across multiple providers. In practice, audit servers could be maintained by a network base, much like the servers that host BitTorrent files or record Bitcoin transactions.
# # #
Written by Larry Hardesty, MIT News Office