RobustNLP/CipherChat
A framework for evaluating the generalizability of safety alignment in large language models using cipher-encoded prompts.

CipherChat is a systematic evaluation framework that tests whether safety alignments in LLMs, which are trained on natural language human feedback, can be bypassed using non-natural ciphers. The framework teaches the model to comprehend cipher language by designating it as a cipher expert, then probes for safety vulnerabilities across different cipher methods and instruction domains. Results are provided as query-response pairs that can be loaded and analyzed.
Frequently asked
- What is RobustNLP/CipherChat?
- A framework for evaluating the generalizability of safety alignment in large language models using cipher-encoded prompts.
- Is CipherChat open source?
- Yes — RobustNLP/CipherChat is open source, released under the MIT license.
- What language is CipherChat written in?
- RobustNLP/CipherChat is primarily written in Python.
- How popular is CipherChat?
- RobustNLP/CipherChat has 628 stars on GitHub.
- Where can I find CipherChat?
- RobustNLP/CipherChat is on GitHub at https://github.com/RobustNLP/CipherChat.