๐Ÿ›ก๏ธ Prompt Safety Classifier

Fine-tuned Llama 3.2 1B Instruct using QLoRA + Unsloth

This model classifies prompts into one of three categories:

  • ๐ŸŸข Benign
  • ๐Ÿ”ด Harmful
  • ๐ŸŸ  Jailbreak

Examples

Label Definitions

๐ŸŸข Benign

Safe and harmless prompts.

๐Ÿ”ด Harmful

Requests involving dangerous, illegal, or malicious activities.

๐ŸŸ  Jailbreak

Prompts attempting to manipulate or bypass the model's safety behavior.


Base Model: Llama 3.2 1B Instruct

Fine-tuning: QLoRA + Unsloth

Task: Prompt Safety Classification