Skip to content

feat(llm): allow setting query transformation behavior in BaseRAGQuestionAnswerer / BaseRAGQA (#67) - #280

Open
omm-prakash18 wants to merge 1 commit into
pathwaycom:mainfrom
omm-prakash18:feat/rag-query-transform
Open

omm-prakash18 wants to merge 1 commit into
pathwaycom:mainfrom
omm-prakash18:feat/rag-query-transform

Conversation

@omm-prakash18

Copy link
Copy Markdown

Summary

Fixes #67

This pull request introduces configurable query transformation behavior to BaseRAGQuestionAnswerer (and exposes the BaseRAGQA class alias).

Users can now configure whether and how user queries are rewritten or expanded before vector index retrieval (supporting standard query rewriting, HyDE - Hypothetical Document Embeddings, or custom prompt callables/UDFs), defaulting to None for backward compatibility.


What Changed

  • Configurable query_transform:
    • Added query_transform: str | Callable | pw.UDF | None = None parameter to BaseRAGQuestionAnswerer.__init__ and AdaptiveRAGQuestionAnswerer.__init__.
    • Added robust resolution:
      • None (default): Skips transformation and queries the vector index with the raw user prompt.
      • "rewrite": Uses prompts.prompt_query_rewrite to generate keyword/entity-optimized search queries.
      • "hyde": Uses prompts.prompt_query_rewrite_hyde for Hypothetical Document Embeddings.
      • Callable / pw.UDF: Allows arbitrary custom transformation logic.
  • Query Pipeline Execution:
    • In answer_query, when query_transform is active, the query is transformed with the LLM before calling self.indexer.retrieve_query. The raw user query is preserved for final answer generation.
  • Ergonomic Alias:
    • Exported BaseRAGQA = BaseRAGQuestionAnswerer.
  • Unit & Integration Tests:
    • Added test cases in python/pathway/xpacks/llm/tests/test_rag.py covering default None, "rewrite", "hyde", custom functions, invalid arguments, and the BaseRAGQA alias.

Example Usage

from pathway.xpacks.llm.question_answering import BaseRAGQA

# 1. Standard Query Rewriting
rag = BaseRAGQA(
    llm=chat_model,
    indexer=vector_server,
    query_transform="rewrite",
)

# 2. HyDE (Hypothetical Document Embeddings)
rag_hyde = BaseRAGQA(
    llm=chat_model,
    indexer=vector_server,
    query_transform="hyde",
)

# 3. Raw Query / Direct Search (Default)
rag_default = BaseRAGQA(
    llm=chat_model,
    indexer=vector_server,
    query_transform=None,
)

@CLAassistant

CLAassistant commented Oct 3, 2026 •

Copy link
Copy Markdown

CLA assistant check
All committers have signed the CLA.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Allow setting query transformers in the BaseRAGQA

2 participants