详细信息
DPMasker: A Context-Aware Reversible Differential Privacy Perturbation Algorithm for Medical Text ( EI收录)
文献类型:期刊文献
英文题名:DPMasker: A Context-Aware Reversible Differential Privacy Perturbation Algorithm for Medical Text
作者:Sun, Yifan[1]; Guo, Weibin[1]
机构:[1] East China University of Science and Technology, Department of Computer Science, Shanghai, China
年份:2026
起止页码:2374
外文期刊名:2026 9th International Conference on Advanced Algorithms and Control Engineering, ICAACE 2026
收录:EI(收录号:20262420909797)
语种:英文
外文关键词:Bioinformatics - Digital storage - Electronic health record - Medical computing - Privacy-preserving techniques - Sampling - Semantics - Text processing
摘要:Automatic clinical text anonymization is crucial for unlocking the secondary application of Electronic Health Records (EHRs) in medical research, yet existing methods often struggle to strike a balance among privacy protection strength, semantic preservation, and data reversibility. Traditional methods based on Named Entity Recognition (NER) and rewriting are prone to poor generalization or the disruption of medical contexts, while differential privacy (DP) perturbation based on vector distance is often constrained by rigid candidate word distributions. To address these issues, this paper proposes DPMasker, a reversible privacy perturbation algorithm for medical texts. Rather than relying on vector distances, this algorithm utilizes Bio-ClinicalBERT to extract the contextual feature distribution of target entities. By introducing a dual-stage dynamic truncation mechanism (head risk isolation and nucleus truncation semantic filtering), DPMasker strictly removes high-risk privacy words and long-tail noise that disrupts medical logic. Subsequently, controlled sampling is performed on a secure candidate set combining the DP exponential mechanism, achieving mathematically strict, quantifiable privacy constraints and maximizing semantic utility. Furthermore, the algorithm designs a DP embedding hash mapping storage mechanism based on HMACSHA256 to meet the zero-trust secure traceback requirements of high-privilege hospital environments. Extensive experiments on the HealthCareMagic real-world doctor-patient dialogue dataset demonstrate that DPMasker successfully achieves absolute zero leakage against privacy extraction attacks. Meanwhile, in the generative utility evaluation of large language models (LLMs) such as Qwen3, GLM-4.6, and Kimi-K2, its BLEU and ROUGE-L scores significantly surpass those of existing mainstream baseline methods. ? 2026 IEEE.
参考文献:
正在载入数据...
