Update README.md
Browse files
README.md
CHANGED
|
@@ -17,7 +17,7 @@ pipeline_tag: text-generation
|
|
| 17 |
|
| 18 |
AMALIAGuard is a content safety guard model for LLM pipelines, designed specifically for **European Portuguese (pt-PT)**. It classifies user prompts and assistant responses as safe or unsafe across a 12-category taxonomy that combines standard universal harm categories with **six GDPR-specific risk categories** — addressing a gap left by existing guard models, which are predominantly English-centric and lack explicit coverage of European data protection regulation.
|
| 19 |
|
| 20 |
-
AMALIAGuard-4B is fine-tuned from [Qwen/Qwen3Guard-Gen-4B](https://huggingface.co/Qwen/Qwen3Guard-Gen-4B) on a three-layer synthetic AART pipeline covering both pillars in pt-PT and English, augmented with translated subsets of WildGuardMix and ToxicChat for broader generalization.
|
| 21 |
|
| 22 |
---
|
| 23 |
|
|
@@ -160,7 +160,7 @@ AMALIAGuard was trained with **absent-category augmentation**: when a violated c
|
|
| 160 |
|
| 161 |
## Evaluation
|
| 162 |
|
| 163 |
-
|
| 164 |
|
| 165 |
- **In-domain (pt-PT held-out test set):** 99.65% overall F1, substantially outperforming zero-shot Qwen3Guard-Gen baselines (78–91% F1) at all three scales (0.6B, 4B, 8B), confirming that the AMALIAGuard taxonomy needs task-specific fine-tuning.
|
| 166 |
- **External benchmarks:** augmenting training with translated WildGuardMix/ToxicChat (the *ext* condition) closes most of the synthetic-to-real gap seen in models trained on in-domain data alone. On ToxicChat, fine-tuned models clearly beat the zero-shot baseline (76.4% vs. 63.7% F1); on WildGuardMix, the best fine-tuned model comes within ~1 point of the zero-shot baseline. HarmBench recall reaches 93.5% (EN).
|
|
|
|
| 17 |
|
| 18 |
AMALIAGuard is a content safety guard model for LLM pipelines, designed specifically for **European Portuguese (pt-PT)**. It classifies user prompts and assistant responses as safe or unsafe across a 12-category taxonomy that combines standard universal harm categories with **six GDPR-specific risk categories** — addressing a gap left by existing guard models, which are predominantly English-centric and lack explicit coverage of European data protection regulation.
|
| 19 |
|
| 20 |
+
AMALIAGuard-4B is fine-tuned from [Qwen/Qwen3Guard-Gen-4B](https://huggingface.co/Qwen/Qwen3Guard-Gen-4B) on a three-layer synthetic AART pipeline covering both pillars in pt-PT and English, augmented with translated subsets of WildGuardMix and ToxicChat for broader generalization.
|
| 21 |
|
| 22 |
---
|
| 23 |
|
|
|
|
| 160 |
|
| 161 |
## Evaluation
|
| 162 |
|
| 163 |
+
Key findings:
|
| 164 |
|
| 165 |
- **In-domain (pt-PT held-out test set):** 99.65% overall F1, substantially outperforming zero-shot Qwen3Guard-Gen baselines (78–91% F1) at all three scales (0.6B, 4B, 8B), confirming that the AMALIAGuard taxonomy needs task-specific fine-tuning.
|
| 166 |
- **External benchmarks:** augmenting training with translated WildGuardMix/ToxicChat (the *ext* condition) closes most of the synthetic-to-real gap seen in models trained on in-domain data alone. On ToxicChat, fine-tuned models clearly beat the zero-shot baseline (76.4% vs. 63.7% F1); on WildGuardMix, the best fine-tuned model comes within ~1 point of the zero-shot baseline. HarmBench recall reaches 93.5% (EN).
|