Skip to content

feat(eval): add --backend flag and cross-domain combined metrics to topic relevance eval - #160

Draft
PritamSGB wants to merge 5 commits into
ProjectTech4DevAI:mainfrom
PritamSGB:feat/topic-relevance-backend-flag
Draft

PritamSGB wants to merge 5 commits into
ProjectTech4DevAI:mainfrom
PritamSGB:feat/topic-relevance-backend-flag

Conversation

@PritamSGB

Copy link
Copy Markdown

Summary

  • Adds a --backend flag to the topic relevance eval so a single backend (topic_relevance or topic_relevance_llm) can be run instead of both
  • topic_relevance (the LLMCritic-based validator) is now imported lazily, so --backend topic_relevance_llm works without the llm_critic hub validator installed
  • Adds combine_binary_metrics() to the shared helper and writes a combined-metrics.json aggregating the education + healthcare domain results by summing confusion-matrix counts

Stacked on #159 — please review that one first. Until it merges, the diff here also includes its commits.

Test plan

  • python3 -m app.evaluation.topic_relevance.run --backend topic_relevance_llm
  • Confirm combined-metrics.json totals match the sum of the two per-domain metrics files

🤖 Generated with Claude Code

PritamSGB and others added 5 commits September 23, 2026 19:07
…val output

Surfaces overall precision/recall/f1/accuracy (not just per-entity) and the
PIIRemover config (entity types, threshold, NLP engine, model, on_fail) used
for the run, so metrics.json is self-describing for external readers.
Clarifies that this is PII's overall detection-level score alongside
entity_metrics, distinct from the shared compute_binary_metrics() helper
used as the sole metric in other eval scripts.
…s output

Mirrors the pii_remover eval script, which already records the config used
(entity types, threshold, etc.) alongside the metrics for reproducibility.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
…valuations

Introduces build_validator_config() in common/helper.py, which reads on_fail
from validator.on_fail_descriptor (set by every guardrails Validator base
class) and normalizes any extra constructor params passed in (enums -> their
.value). PII and gender_assumption_bias now use it instead of hand-rolled
dicts, and ban_list, lexical_slur, topic_relevance, and toxicity gain a
config block in their metrics.json for the first time.
…s to topic relevance eval

Lets topic_relevance/run.py target a single validator backend without
installing llm_critic, and adds a shared helper to aggregate per-domain
binary metrics into one combined report.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
@coderabbitai

coderabbitai Bot commented Sep 23, 2026

Copy link
Copy Markdown

Important

Review skipped

Auto reviews are limited based on label configuration.

🏷️ Required labels (at least one) (1)
  • ready-for-review

Please check the settings in the CodeRabbit UI or the .coderabbit.yaml file in this repository. To trigger a single review, invoke the @coderabbitai review command.

⚙️ Run configuration

Configuration used: Repository: ProjectTech4DevAI/kaapi-guardrails/.coderabbit.yaml

Review profile: CHILL

Plan: Advanced

Run ID: bd05c977-2ea5-4e87-9909-23c2fa9bbb13

You can disable this status message by setting the reviews.review_status to false in the CodeRabbit configuration file.

Use the checkbox below for a quick retry:

  • 🔍 Trigger review

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant