Zephyr
HuggingFace H4
Structures found in this family (1)
By model (1)
Zephyr-7B-Beta · 7B
Observations (2)
Papers
The WMDP Benchmark: Measuring and Reducing Malicious Use With Unlearning (2024), Refusal in Language Models Is Mediated by a Single Direction (2024), Representation Engineering: A Top-Down Approach to AI Transparency (2023), Programming Refusal with Conditional Activation Steering (2024), Refusal Direction is Universal Across Safety-Aligned Languages (2025)