Warning: Undefined array key "mm" in /www/wwwroot/www.ai-bt.com/si.php on line 10 Deprecated: trim(): Passing null to parameter #1 ($string) of type string is deprecated in /www/wwwroot/www.ai-bt.com/si.php on line 10 Anonymizing datasets with demographics and diagnosis codes in the presence of utility constraints.

Literature DB >> 27832965

Anonymizing datasets with demographics and diagnosis codes in the presence of utility constraints.

Giorgos Poulis¹, Grigorios Loukides², Spiros Skiadopoulos³, Aris Gkoulalas-Divanis⁴.

Abstract

Publishing data about patients that contain both demographics and diagnosis codes is essential to perform large-scale, low-cost medical studies. However, preserving the privacy and utility of such data is challenging, because it requires: (i) guarding against identity disclosure (re-identification) attacks based on both demographics and diagnosis codes, (ii) ensuring that the anonymized data remain useful in intended analysis tasks, and (iii) minimizing the information loss, incurred by anonymization, to preserve the utility of general analysis tasks that are difficult to determine before data publishing. Existing anonymization approaches are not suitable for being used in this setting, because they cannot satisfy all three requirements. Therefore, in this work, we propose a new approach to deal with this problem. We enforce the requirement (i) by applying (k,km)-anonymity, a privacy principle that prevents re-identification from attackers who know the demographics of a patient and up to m of their diagnosis codes, where k and m are tunable parameters. To capture the requirement (ii), we propose the concept of utility constraint for both demographics and diagnosis codes. Utility constraints limit the amount of generalization and are specified by data owners (e.g., the healthcare institution that performs anonymization). We also capture requirement (iii), by employing well-established information loss measures for demographics and for diagnosis codes. To realize our approach, we develop an algorithm that enforces (k,km)-anonymity on a dataset containing both demographics and diagnosis codes, in a way that satisfies the specified utility constraints and with minimal information loss, according to the measures. Our experiments with a large dataset containing more than 200,000 electronic health records show the effectiveness and efficiency of our algorithm.

Entities: Species

Keywords: Demographics; Diagnosis codes; Generalization; Privacy; Suppression; Utility constraints

Mesh：

Year: 2016 PMID： 27832965 DOI： 10.1016/j.jbi.2016.11.001

Source DB: PubMed Journal: J Biomed Inform ISSN： 1532-0464 Impact factor: 6.317

Keyword Cloud
Cited

3 in total

Anonymizing datasets with demographics and diagnosis codes in the presence of utility constraints.

1. Returning to our roots: The use of geospatial data for nurse-led community research.

2. Privacy Policy and Technology in Biomedical Data Science.

3. Probabilistic record linkage of de-identified research datasets with discrepancies using diagnosis codes.