Revisiting Character-level Adversarial Attacks for Language Models

Adversarial attacks in Natural Language Processing apply perturbations in the character or token levels. Token-level attacks, gaining prominence for their use of gradient-based methods, are susceptible to altering sentence semantics, leading to invalid adversarial examples. While characterlevel attacks easily maintain semantics, they have received less attention as they cannot easily adopt popular gradient-based methods, and are thought to be easy to defend. Challenging these beliefs, we introduce Charmer, an efficient query-based adversarial attack capable of achieving high attack success rate (ASR) while generating highly similar adversarial examples. Our method successfully targets both small (BERT) and large (Llama 2) models. Specifically, on BERT with SST-2, Charmer improves the ASR in 4.84% points and the USE similarity in 8% points with respect to the previous art. Our implementation is available in github.com/LIONS-EPFL Charmer.

Chat with Graph Search

Ask any question about EPFL courses, lectures, exercises, research, news, etc. or try the example questions below.

DISCLAIMER: The Graph Chatbot is not programmed to provide explicit or categorical answers to your questions. Rather, it transforms your questions into API requests that are distributed across the various IT services officially administered by EPFL. Its purpose is solely to collect and recommend relevant references to content that you can explore to help you answer your questions.

Revisiting Character-level Adversarial Attacks for Language Models

Graph Chatbot

Chat with Graph Search

Understanding generalization and robustness in modern deep learning

Infusing structured knowledge priors in neural models for sample-efficient symbolic reasoning

Explainable Face Verification via Feature-Guided Gradient Backpropagation

Understanding generalization and robustness in modern deep learning

Infusing structured knowledge priors in neural models for sample-efficient symbolic reasoning

Explainable Face Verification via Feature-Guided Gradient Backpropagation