Skip to content

feat(eval): add lexical_slur to toxicity eval, standardize dataset - #161

Draft
PritamSGB wants to merge 6 commits into
ProjectTech4DevAI:mainfrom
PritamSGB:feat/toxicity-lexical-slur-eval
Draft

PritamSGB wants to merge 6 commits into
ProjectTech4DevAI:mainfrom
PritamSGB:feat/toxicity-lexical-slur-eval

Conversation

@PritamSGB

Copy link
Copy Markdown

Summary

  • Adds LexicalSlur as a fourth validator in the toxicity eval
  • Adds --validators / --sources flags to run a subset of validators or source datasets
  • Adds a combined_pred column and a combined metrics entry — a logical OR across whichever validators ran, derived from the already-produced *_pred columns (nothing is re-run) — plus a source_metrics breakdown per origin dataset for every validator
  • Consolidates the three toxicity datasets into one standardized toxicity_test_combined.csv (text, label, language, dataset), with one predictions.csv / metrics.json output pair instead of a file per dataset
  • Imports nsfw_text / profanity_free lazily from their guardrails_ai.* packages and llamaguard_7b lazily from guardrails.hub, so selecting a subset of validators no longer requires llamaguard_7b's still-unmigrated hub path to be importable (a top-level import currently fails outright, since .guardrails/hub_registry.json still points at the defunct guardrails_grhub_llamaguard_7b package)

Results from a local run

--validators lexical_slur profanity_free across all 1801 rows:

Scope Precision Recall F1
Overall (combined) 0.85 0.72 0.78
hasoc 0.88 0.49 0.63
sharechat 0.77 0.47 0.58
lexical 0.90 1.00 0.95

Note the lexical source is the lexical-slur eval's own curated set, so its near-perfect recall is expected and inflates the overall figure; hasoc + sharechat alone give precision 0.78 / recall 0.48 / F1 0.59.

Stacked on #159 — please review that one first. Until it merges, the diff here also includes its commits.

Test plan

  • python3 -m app.evaluation.toxicity.run --validators lexical_slur profanity_free
  • Confirm metrics.json has a per-validator entry plus combined, each with a source_metrics breakdown
  • Confirm predictions.csv has combined_pred equal to the row-wise OR of the individual *_pred columns

🤖 Generated with Claude Code

PritamSGB and others added 5 commits September 23, 2026 19:07
…val output

Surfaces overall precision/recall/f1/accuracy (not just per-entity) and the
PIIRemover config (entity types, threshold, NLP engine, model, on_fail) used
for the run, so metrics.json is self-describing for external readers.
Clarifies that this is PII's overall detection-level score alongside
entity_metrics, distinct from the shared compute_binary_metrics() helper
used as the sole metric in other eval scripts.
…s output

Mirrors the pii_remover eval script, which already records the config used
(entity types, threshold, etc.) alongside the metrics for reproducibility.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
…valuations

Introduces build_validator_config() in common/helper.py, which reads on_fail
from validator.on_fail_descriptor (set by every guardrails Validator base
class) and normalizes any extra constructor params passed in (enums -> their
.value). PII and gender_assumption_bias now use it instead of hand-rolled
dicts, and ban_list, lexical_slur, topic_relevance, and toxicity gain a
config block in their metrics.json for the first time.
Adds LexicalSlur alongside the existing toxicity validators with
--validators/--sources filtering and a combined (OR) metric per source
(computed from already-produced predictions, no re-running), now against
a single standardized dataset (text/label/language/dataset) instead of
three separate files. Also lazily imports nsfw_text/profanity_free from
their new guardrails_ai.* packages and llamaguard_7b from guardrails.hub,
so selecting a subset of validators doesn't require llamaguard_7b's
still-unmigrated hub path to be importable at all.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
@coderabbitai

coderabbitai Bot commented Sep 23, 2026 •

Copy link
Copy Markdown

Important

Review skipped

Auto reviews are limited based on label configuration.

🏷️ Required labels (at least one) (1)
  • ready-for-review

Please check the settings in the CodeRabbit UI or the .coderabbit.yaml file in this repository. To trigger a single review, invoke the @coderabbitai review command.

⚙️ Run configuration

Configuration used: Repository: ProjectTech4DevAI/kaapi-guardrails/.coderabbit.yaml

Review profile: CHILL

Plan: Advanced

Run ID: 5c06098f-a2c6-4d5c-917c-25cb1e139f97

You can disable this status message by setting the reviews.review_status to false in the CodeRabbit configuration file.

Use the checkbox below for a quick retry:

  • 🔍 Trigger review

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

…icity eval

The toxicity eval runs uli_slur_match over a superset of the rows the
standalone script used, producing identical metrics and config for that
subset (--validators lexical_slur --sources lexical reproduces it exactly),
so the separate script and its dataset are redundant.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant