Attacking Neural Networks with Neural Networks: Towards Deep Synchronization for Backdoor Attacks

Document Type

Conference Proceeding

Publication Date

10-21-2023

Abstract

Backdoor attacks inject poisoned samples into training data, where backdoor triggers are embedded into the model trained on the mixture of poisoned and clean samples. An interesting phenomenon can be observed in the training process: the loss of poisoned samples tends to drop significantly faster than that of clean samples, which we call the early-fitting phenomenon. Early-fitting provides a simple but effective evidence to defend against backdoor attacks, where the poisoned samples can be detected by selecting the samples with the lowest loss values in the early training epochs. Then, two questions naturally arise: (1) What characteristics of poisoned samples cause early-fitting? (2) Does a stronger attack exist which could circumvent the defense methods? To answer the first question, we find that early-fitting could be attributed to a unique property among poisoned samples called synchronization, which depicts the similarity between two samples at different layers of a model. Meanwhile, the degree of synchronization could be controlled based on whether it is captured by shallow or deep layers of the model. Then, we give an affirmative answer to the second question by proposing a new backdoor attack method, Deep Backdoor Attack (DBA), which utilizes deep synchronization to reverse engineer trigger patterns by activating neurons in the deep layer of a base neural network. Experimental results validate our propositions and the effectiveness of DBA. Our code is available at https://github.com/GuanZihan/Deep-Backdoor-Attack.

Identifier

85178103602 (Scopus)

ISBN

[9798400701245]

Publication Title

International Conference on Information and Knowledge Management Proceedings

External Full Text Location

https://doi.org/10.1145/3583780.3614784

First Page

608

Last Page

618

Grant

2223768

Fund Ref

National Science Foundation

This document is currently not available here.

Share

COinS