Bergeron: Combating Adversarial Attacks through a Conscience-Based Alignment Framework

Pisano, Matthew; Ly, Peter; Sanders, Abraham; Yao, Bingsheng; Wang, Dakuo; Strzalkowski, Tomek; Si, Mei

Computer Science > Cryptography and Security

arXiv:2312.00029v1 (cs)

[Submitted on 16 Nov 2023 (this version), latest version 15 Mar 2024 (v2)]

Title:Bergeron: Combating Adversarial Attacks through a Conscience-Based Alignment Framework

Authors:Matthew Pisano, Peter Ly, Abraham Sanders, Bingsheng Yao, Dakuo Wang, Tomek Strzalkowski, Mei Si

View PDF

Abstract:Modern Large language models (LLMs) can still generate responses that may not be aligned with human expectations or values. While many weight-based alignment methods have been proposed, many of them still leave models vulnerable to attacks when used on their own. To help mitigate this issue, we introduce Bergeron, a framework designed to improve the robustness of LLMs against adversarial attacks. Bergeron employs a two-tiered architecture. Here, a secondary LLM serves as a simulated conscience that safeguards a primary LLM. We do this by monitoring for and correcting potentially harmful text within both the prompt inputs and the generated outputs of the primary LLM. Empirical evaluation shows that Bergeron can improve the alignment and robustness of several popular LLMs without costly fine-tuning. It aids both open-source and black-box LLMs by complementing and reinforcing their existing alignment training.

Subjects:	Cryptography and Security (cs.CR); Artificial Intelligence (cs.AI); Computation and Language (cs.CL)
Cite as:	arXiv:2312.00029 [cs.CR]
	(or arXiv:2312.00029v1 [cs.CR] for this version)
	https://doi.org/10.48550/arXiv.2312.00029

Submission history

From: Matthew Pisano [view email]
[v1] Thu, 16 Nov 2023 07:31:18 UTC (370 KB)
[v2] Fri, 15 Mar 2024 20:13:06 UTC (196 KB)

Computer Science > Cryptography and Security

Title:Bergeron: Combating Adversarial Attacks through a Conscience-Based Alignment Framework

Submission history

Access Paper:

References & Citations

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Computer Science > Cryptography and Security

Title:Bergeron: Combating Adversarial Attacks through a Conscience-Based Alignment Framework

Submission history

Access Paper:

References & Citations

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators