Skip to content

v0.9.4: db contention fixes, databricks genie, snowflake cortex, additional search connectors - #8356

Open
waleedlatif1 wants to merge 22 commits into
mainfrom
staging
Open

waleedlatif1 wants to merge 22 commits into
mainfrom
staging

Conversation

@waleedlatif1

@waleedlatif1 waleedlatif1 commented Sep 28, 2026 •

Copy link
Copy Markdown
Collaborator

waleedlatif1 and others added 21 commits September 26, 2026 13:11
…une finished outbox events (#8330)

- Processing recovery discovered candidates by walking the global `(uploaded_at, id)` recovery index, joining each document to its knowledge base and connector, and rejecting it only afterwards. Retained inputs of paused or dormant sources sit at the front of that order, so every call read through all of them first. Follow-up batches (`NOT IN` attempted connectors) walked the same range again
- Discovery now starts from the eligible connectors and reads each one's oldest candidates through a new partial index `doc_connector_processing_recovery_idx (connector_id, uploaded_at, id)` with a `CROSS JOIN LATERAL`, then takes the oldest overall. Paused sources cost nothing. Batch size, the follow-up batches for refused connectors, blocked knowledge bases, the lock-time recheck, and liveness are unchanged, so recovery picks the same candidates as before
- The old `doc_processing_recovery_idx` stays for now and is dropped in a later contract migration
- Outbox retention: `outbox_event` never deleted completed rows. The outbox processor now prunes `completed` rows older than 7 days in bounded, `SKIP LOCKED` batches, only for `knowledge.document.processing.recover` and `knowledge.document.storage.cleanup`. Both have random ids, and nothing reads their completed rows. Every other event type is kept, including idempotency-keyed ones and the checkpoint expiry events, whose completed rows are read back by id, as are pending, processing, and dead-letter rows
…nd event (#8328)

- Each chat turn loaded the chat's full transcript through `resolveOrCreateChat` just to read the MCP server ids earlier user messages were tagged with. `resolveOrCreateChat` now never loads the transcript, and a new `loadChatMcpServerIds` reads only the MCP context ids with one jsonb query, in the same first-tagged order
- Split the chat loaders: `getAccessibleCopilotChatDetail` (row plus authorization, no transcript) backs both `resolveOrCreateChat` and `getAccessibleCopilotChatWithMessages`. Dropped the `includeTranscript` option and the `conversationHistory` field, which nothing read
- Client: a `completed` event for the viewer's own live stream no longer refetches the chat detail, because the client's own stream finalization already refetches it. `renamed` marks the detail stale without refetching; the lists that show titles still refetch
…indexes (#8333)

- Since #8314, workspace knowledge search ranks on `embedding`/`document`. The only live reader of `embedding_keyword_search` and `embedding_keyword_tin` is dormant indexed search, and only for bases that are search indexes. Yet every chunk write still upserted a keyword row, ran a Tin `DELETE` on insert, and every document insert fanned an ACL/connector sync out to the projections
- Script migration `0025_scope_keyword_projections`, all `CREATE OR REPLACE` in one transaction with the 0024 lock-timeout retry:
  - `sync_embedding_keyword_search` takes the shared membership lock and writes keyword rows only for `is_search_index` bases. On an update that moves a chunk out of a search index, it deletes the chunk's row
  - A new flip trigger backfills a base's keyword rows when it becomes a search index and deletes them when it stops being one, under the exclusive membership lock. It is separate from the Tin flip trigger, which only exists where Tin is installed. Its upsert skips unchanged rows
  - The Tin chunk trigger skips its `DELETE` on insert, since a new chunk can't have a Tin row
  - `projection_source_acl_sync` becomes `AFTER UPDATE OF connector_id, acl ... WHEN` the value actually changed. A new document has no chunks yet (FK), so the insert arm never matched anything
- The projector's keyword page writes only for search-index bases, keyed on the base flag
- Fresh installs, `db:push`, and late Tin adoption end with the same trigger definitions: 0019 installs the final guarded, key-share-locked Tin triggers, and 0025 re-runs after 0016. A `db:push` re-run of the legacy 0016 backfill may refill keyword rows for non-search bases; they're unread and left in place like the existing ones
- Both adoption backfills read their chunks `FOR KEY SHARE`, so a chunk delete racing an adoption is waited out instead of failing it
- Rollback floor: v0.9.1 and earlier ranked workspace keyword search through `embedding_keyword_search`. Once 0025 has run, don't roll back below v0.9.2, and self-hosters should upgrade through v0.9.2 or later
- Existing keyword rows for non-search bases are left in place (unread) and removed separately
- Every transaction-scoped advisory lock in `apps/sim` ran the same statement, `SELECT pg_advisory_xact_lock(hashtextextended($1, 0))`, so query insights showed all lock waits as one fingerprint with no way to tell which caller was waiting
- New `acquireAdvisoryXactLock(tx, tag, key)` and `tryAcquireAdvisoryXactLock(tx, tag, key)` in `lib/db/advisory-locks.ts` issue the same call with a trailing SQLCommenter tag (`/*lock='<family>'*/`), which PlanetScale query tags can filter on. The tag must be a static snake_case identifier, so it can't break out of the comment
- About 45 call sites now go through the helper, each tagged with its lock family (`user_table_rows_pos`, `organization_membership`, `usage_log_event`, …). Keys, lock types, timeouts, and transaction boundaries are unchanged
- Left alone: statements whose SQL text is already unique (`packages/db` locks, the file-search dispatch `generate_series` claim) and one untyped repair script
…s processes (#8325)

- Enterprise reporting-window usage sums (a year-long ledger scan) are now shared across processes through Redis with a 30s TTL. Trigger.dev runs every task in a fresh process, so the existing in-process LRU was always cold there, and the scan ran on nearly every document-processing and execution admission check
- The in-process LRU stays in front and still coalesces concurrent misses. The Redis GET has a 250ms deadline, and any Redis error, timeout, or unreadable value falls through to the exact sum. The write is a fire-and-forget `SET … NX` with a jittered TTL, so a slower, older sum never overwrites a fresher one or extends its life. There is no lock or lease
- The usage threshold email is now level-triggered, instead of being edge-triggered off an exact before/after org sum on every workflow completion. A claim keyed on (billing period, limit) (`claimCreditsThreshold`) sends each threshold at most once. A new period or a changed limit re-arms it with no reset write, and concurrent completions can't both send
- Recipients are resolved before claiming, so a period's email isn't used up when nobody can receive it. Zero-cost completions don't claim. The personal usage baseline and the claimed period come from the same billing context
- Invoicing, cycle close, overage, and threshold billing still read the ledger exactly
- Rollout note: accounts already at or above 80% this billing period get one threshold email after deploy; there is no backfill
…e unique index (#8327)

- The File block's URL fetch no longer saves every fetched URL into workspace Files. Nothing ever read the saved copy: the parser keeps its own execution file, and the reuse the save once served was removed earlier. Repeated fetches were piling up `name (N).html` copies in the Files root
- Renamed `fetchExternalUrlToWorkspace` to `fetchExternalUrl` and removed its save, permission, and upload branch
- Name existence checks (`fileNameExistsInWorkspaceFolder`, `getWorkspaceFileByName`) now spell the folder predicate as `coalesce(folder_id, '') = $folder`, matching the unique `(workspace_id, coalesce(folder_id, ''), original_name)` index, so each check is a point lookup. The old `folder_id IS NULL` form made root lookups scan the folder-id index
- `allocateUniqueWorkspaceFileName` probes the base name plus `(1)`…`(20)`, then falls back to a short-id suffix instead of probing up to 1,000 candidates and then failing. The unique index and the existing conflict retry remain the authority
…where counts are shown (#8329)

- v1 knowledge base list, detail, and update returned `docCount`/`tokenCount` over every document in the base, including connector documents the caller cannot read. They now count through the caller's access (`resolveV1KnowledgeReadAccess`), the same way the v1 documents routes already do. v1 detail also now returns its real `connectorTypes`
- The internal KB list joined `document` through the full access predicate on every call, but only the Knowledge page shows the totals. The list contract gains `includeCounts` (default false), and without it the query reads only `knowledge_base` columns. The service option is `countsFor: access`, so a count without an access filter can't be written
- React Query: counted lists get their own key beside `list()`, used only by the Knowledge page and its server prefetch. Document, upload, and connector mutations refresh only the counted lists; KB create, rename, delete, restore, and move refresh both
- Removed `getKnowledgeBaseById`, whose unfiltered count join ran on every context resolution. Callers use `getActiveKnowledgeBaseReference`, and single-base totals come only from `attachKnowledgeBaseConnectors(kb, access)`
- The v2 list keeps returning totals, since its public contract requires them
* fix(copilot): bound agent CLI grep matching

* fix(copilot): bound grep context expansion
* fix(api): constrain workflow response headers

* fix(api): reserve additional browser policy headers
…#8332)

- Drops 8 indexes on hot write tables in one `DROP INDEX CONCURRENTLY` migration (0386), following the 0239 pattern: `COMMIT` breakpoint, `lock_timeout 0`, idempotent replay
- `{emb,doc}_date1_idx` / `date2_idx`: tag filters compare `col::date`, which a plain timestamp btree can never match, so these are written on every chunk insert and document update and never serve a query
- Four left-prefix duplicates of an existing non-partial superset, which keeps serving equality lookups and FK cascades on the leading column:
  - `emb_kb_id_idx`: covered by `emb_kb_enabled_idx` / `emb_kb_model_idx`
  - `emb_doc_id_idx`: covered by `emb_doc_chunk_idx` / `emb_doc_enabled_idx`
  - `usage_log_workspace_id_idx`: covered by `usage_log_workspace_created_at_idx`
  - `copilot_runs_execution_id_idx`: covered by `copilot_runs_execution_started_at_idx`
- Kept on purpose:
  - Every `number` and `boolean` tag-slot index. The vector leg's short-result rescue probe checks tag filters with an `EXISTS` over `embedding` that relies on them, and without them a selective number/boolean filter turns that probe into a table-wide scan
  - The small `copilot_runs` chat/workspace indexes
…8343)

* feat(databricks): add Genie agent operations to the Databricks block

* fix(databricks): fail Genie asks on nested errors and document per-tool outputs
* fix(execution): bound reference and context processing

* fix(execution): preserve compatibility and tighten input budgets
…8346)

* fix(desktop): wait for native update staging and verify installation

* fix(desktop): distinguish refresh failures from native update errors
…#8347)

* feat(snowflake): add Cortex Analyst operations to the Snowflake block

* fix(snowflake): keep Cortex Analyst SQL context for quoted and multi-source views and replay suggestions

* fix(snowflake): run Cortex Analyst SQL with the resolved token, the chosen source's schema, and document access requirements
* feat(search): expand live providers and discussion reads

* fix(search): preserve transcript details and account search

* fix(search): require complete review and date evidence

* fix(search): align mixed-query acceptance limits

* fix(search): bind acceptance evidence to matched comments
…tamps stay off the index (#8334)

- Connector sync stamps `document.source_seen_at` on every listed document. `source_seen_at` is a key column of `doc_connector_reconciliation_idx`, so every stamp was a non-HOT update that wrote every index on `document`. This is release A of two: move every reader off that key so release B can drop it and the stamps become HOT
- New `doc_connector_reconciliation_v2_idx (connector_id, id)` with the same partial predicate, built concurrently. The old index stays until release B
- The absence walks (ACL revoke, soft delete, hard delete) and the member resurrection walk page by id instead of `(seen, id)`, with absence as a plain filter. Each window is built by a recursive keyset walk that fetches one row per step (`id > previous ORDER BY id LIMIT 1`) up to 5,000 ids, so every statement reads at most one window whatever plan the database picks. A single `ORDER BY id LIMIT` could be planned as a bitmap read of the whole connector plus a sort. Matches are filtered within the window, so cost is bounded by ids scanned, not matches found. A walk whose absence count is zero is skipped
- Both seen stamps skip rows already stamped at or after the run start (`staleSeen`), so current rows aren't rewritten and a later stamp is never overwritten by an earlier one
* fix(chat): use hostname brand icons for inline sources

* fix(chat): recognize www aliases in source branding
* feat(chat): add shared find menu

* fix(chat): preserve find focus and scroll ownership

* fix(chat): resume pending send scroll after find closes
@vercel

vercel Bot commented Sep 28, 2026 •

Copy link
Copy Markdown

The latest updates on your projects. Learn more about Vercel for GitHub.

1 Skipped Deployment
Project Deployment Actions Updated
docs Skipped Skipped Sep 28, 2026 12:48am UTC

Request Review

@github-actions github-actions Bot added the requires-mothership-merge Has a companion PR on the mothership/copilot side — merge in lockstep label Sep 28, 2026
@github-actions

Copy link
Copy Markdown

⚠️ Cross-repo companion check

One or more companion PRs aren't merged into main yet (aggregated across the feature PRs in this release). Merging this without them will leave copilot and sim out of sync — merge them in lockstep.

  • ⚠️ simstudioai/mothership#539 — merged into staging (this PR targets main) — feat(search): expand live provider contracts and query guidance

Comment thread apps/sim/scripts/test-search-discussions-live.ts Dismissed
Comment thread apps/sim/scripts/test-search-discussions-live.ts Dismissed

@cubic-dev-ai cubic-dev-ai Bot left a comment •

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

19 issues found across 278 files

Confidence score: 2/5

  • apps/sim/tools/databricks/utils.ts: databricksUrl can send a Databricks bearer token to an attacker-controlled host, exposing credentials through Genie operations; validate the host before making requests.
  • apps/sim/lib/db/advisory-locks.ts: Passing the pool-level db runs the lock statement in autocommit, releasing the lock before the protected work and risking concurrent updates; require DbTransaction in both helpers.
  • apps/sim/lib/knowledge/connectors/member-observations.ts: A scan that exceeds its lifecycle budget restarts from the first document on every run, potentially leaving later documents unprocessed; persist a resume cursor.
  • apps/sim/lib/billing/core/reporting-usage-cache.ts: A slow ledger sum can overwrite a newer cached total with an older value and fresh TTL, making usage reporting stale; prevent writes from sums that exceed the cache’s age limit.
Prompt for AI agents (unresolved issues)

Check if these issues are valid — if so, understand the root cause of each and fix them. If appropriate, use sub-agents to investigate and fix each issue separately.


<file name="apps/sim/lib/db/advisory-locks.ts">

<violation number="1" location="apps/sim/lib/db/advisory-locks.ts:22">
P2: `DbOrTx` also admits the pool-level `db`; passing it directly runs this statement in autocommit and releases the lock before the caller's work. Require `DbTransaction` in both lock helpers so this misuse fails at compile time.</violation>
</file>

<file name="apps/sim/lib/knowledge/connectors/member-observations.ts">

<violation number="1" location="apps/sim/lib/knowledge/connectors/member-observations.ts:916">
P2: This walk restarts from the first document ID after a deadline, so a connector whose scan exceeds the remaining lifecycle budget repeats the same prefix on every run and may never resurrect later documents. Persist and resume the walk cursor so each run makes durable progress.</violation>
</file>

<file name="apps/sim/lib/sim-search/live/notion-mcp.ts">

<violation number="1" location="apps/sim/lib/sim-search/live/notion-mcp.ts:158">
P2: The selected tool already identifies an AI search, but this only marks results partial when the MCP payload includes `type: 'ai_search'`. Derive this flag from `tool` so a response without that discriminator is not presented as complete.</violation>
</file>

<file name="apps/sim/lib/billing/core/usage.ts">

<violation number="1" location="apps/sim/lib/billing/core/usage.ts:726">
P2: For paid personal plans, `usageBefore` already excludes weekly-refresh credits, but `costDelta` is the raw ledger increment. Adding the full delta counts refresh-covered spend toward the threshold and can send warning or limit-reached emails early; calculate the post-record effective usage or pass a delta in the same effective-usage basis.</violation>
</file>

<file name="apps/sim/lib/credential-groups/managed-mcp-service.ts">

<violation number="1" location="apps/sim/lib/credential-groups/managed-mcp-service.ts:328">
P2: `validateServerUrl` performs DNS/egress I/O while the supplied executor is the caller's open transaction, holding the accounts lock and DB connection unnecessarily. Validate the URL before opening that transaction, or pass a prevalidated result into this function.</violation>
</file>

<file name="apps/sim/app/workspace/[workspaceId]/home/components/user-message-content/utils.ts">

<violation number="1" location="apps/sim/app/workspace/[workspaceId]/home/components/user-message-content/utils.ts:31">
P2: Slash-command contexts use `/` tokens, but this branch assigns them `@`, so persisted `/command` mentions never match their context. Include `slash_command` among the slash-prefixed kinds.</violation>

<violation number="2" location="apps/sim/app/workspace/[workspaceId]/home/components/user-message-content/utils.ts:33">
P2: Stored-context mentions ending in punctuation are not recognized, so `@Workflow,` renders without a chip and remains prefixed in chat search. Match a non-word boundary after the token so punctuation-delimited mentions behave consistently.</violation>
</file>

<file name="apps/sim/lib/sim-search/live/discussion.ts">

<violation number="1" location="apps/sim/lib/sim-search/live/discussion.ts:48">
P2: This catch turns `AbortError` and `TimeoutError` into a coverage warning, so a cancelled or timed-out live read can return successfully. Re-throw cancellation and timeout errors, and only downgrade provider retrieval failures to partial coverage.</violation>
</file>

<file name="apps/sim/lib/knowledge/__integration__/scale.integration.ts">

<violation number="1" location="apps/sim/lib/knowledge/__integration__/scale.integration.ts:371">
P2: This assertion makes the scale test depend on one planner index choice even though its one-connector fixture is the case where PostgreSQL may legitimately choose `document_pkey`. Accept either the v2 index or the primary-key walk, or seed enough unrelated rows to make the v2 choice deterministic.</violation>
</file>

<file name="apps/sim/app/workspace/[workspaceId]/home/components/mothership-chat/mothership-chat.tsx">

<violation number="1" location="apps/sim/app/workspace/[workspaceId]/home/components/mothership-chat/mothership-chat.tsx:694">
P3: Passing the transient `chatId` as the find scope closes the find bar and clears its query when a pending chat receives its persisted id, even though the conversation did not change. Preserve state across the undefined-to-id transition or provide a stable conversation scope.</violation>
</file>

<file name="apps/sim/lib/billing/core/reporting-usage-cache.ts">

<violation number="1" location="apps/sim/lib/billing/core/reporting-usage-cache.ts:108">
P2: The readiness check does not prevent a disconnect before `set`; with offline queuing enabled, ioredis can replay this fire-and-forget write after reconnect and extend a stale usage value. Use a cache-specific Redis path with offline queuing disabled and treat reconnecting as a write miss.</violation>

<violation number="2" location="apps/sim/lib/billing/core/reporting-usage-cache.ts:108">
P2: `NX` only protects the key while it exists; a slow ledger sum can finish after a newer entry expires and reintroduce its older total with a fresh TTL. Track the sum age or skip writes from sums that exceed the cache TTL so stale producers cannot overwrite newer usage.</violation>
</file>

<file name="apps/sim/lib/sim-search/live/linear.ts">

<violation number="1" location="apps/sim/lib/sim-search/live/linear.ts:141">
P2: Oldest-sorted searches silently claim completion when Linear omits `hasPreviousPage`, because the response guard validates only `hasNextPage`. Validate `hasPreviousPage` whenever `backwards` is true so incomplete reverse pages fail closed or are reported partial.</violation>
</file>

<file name="apps/sim/tools/databricks/utils.ts">

<violation number="1" location="apps/sim/tools/databricks/utils.ts:75">
P1: `databricksUrl` sends the Databricks bearer token to any host supplied in the `host` parameter, enabling SSRF and credential disclosure through the new Genie operations. Validate the host with `validateDatabricksWorkspaceHost` before constructing the URL and reject hosts outside Databricks-owned domains.</violation>
</file>

<file name="apps/sim/scripts/test-search-discussions-live.ts">

<violation number="1" location="apps/sim/scripts/test-search-discussions-live.ts:254">
P2: Validate the GitHub free-text length before this loop. The schema allows 2,000 characters, so an overlong case makes real oracle `/search/issues` calls before `searchGitHub` returns its 256-character error, wasting quota and obscuring the intended failure.

(Based on your team's feedback about GitHub free-text query limits.)</violation>
</file>

<file name="apps/sim/lib/sim-search/live/meeting-content.ts">

<violation number="1" location="apps/sim/lib/sim-search/live/meeting-content.ts:7">
P3: `truncate` slices by UTF-16 code unit, so when the cut lands inside an astral character (emoji, math symbols) the result ends with a lone surrogate half, which this codebase's own string module documents as a fidelity hazard (see `truncateAtCodePoint`, "never cuts inside a surrogate pair"). Meeting transcripts routinely contain such characters, and the 199,922-unit cut in the default 200,000 limit will eventually split one. Use `truncateAtCodePoint(content, limit - notice.length, '')` instead, and update the import on line 1 accordingly.</violation>
</file>

<file name="apps/sim/lib/sim-search/live/google.ts">

<violation number="1" location="apps/sim/lib/sim-search/live/google.ts:128">
P2: An aborted comments request is converted into a coverage warning, so `readDrive` continues and returns a document after cancellation. Preserve abort errors instead of treating them as recoverable discussion failures.</violation>
</file>

<file name="apps/sim/lib/billing/core/limit-notifications.ts">

<violation number="1" location="apps/sim/lib/billing/core/limit-notifications.ts:101">
P2: This collapses all period starts on one UTC date into the same claim key. A billing-period reset later that day therefore leaves both thresholds claimed and suppresses the new-period email; store the exact period-start timestamp instead.</violation>
</file>

<file name="apps/sim/app/workspace/[workspaceId]/home/components/mothership-chat/chat-find-text.ts">

<violation number="1" location="apps/sim/app/workspace/[workspaceId]/home/components/mothership-chat/chat-find-text.ts:25">
P2: GFM footnote references are childless nodes, so this branch removes their visible marker from the search projection. Searching text after `[^1]` then uses offsets shorter than the rendered DOM, causing highlights to land on the wrong characters; preserve the reference's rendered footprint and keep it aligned with `findRanges`.</violation>
</file>

Note: This PR contains a large number of files. cubic selects up to 200 of the highest-priority eligible files for this review, so some files may not have been reviewed.
Tip: instead of fixing issues one by one fix them all with cubic

Re-trigger cubic

Comment thread apps/sim/tools/snowflake/cortex_analyst_ask.ts
Comment thread apps/sim/lib/api/contracts/v2/knowledge.ts
.trim()
.replace(/^https?:\/\//, '')
.replace(/\/$/, '')
return `https://${normalizedHost}${path}`

@cubic-dev-ai cubic-dev-ai Bot Sep 28, 2026 •

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P1: databricksUrl sends the Databricks bearer token to any host supplied in the host parameter, enabling SSRF and credential disclosure through the new Genie operations. Validate the host with validateDatabricksWorkspaceHost before constructing the URL and reject hosts outside Databricks-owned domains.

Prompt for AI agents
Check if this issue is valid — if so, understand the root cause and fix it. At apps/sim/tools/databricks/utils.ts, line 75:

<comment>`databricksUrl` sends the Databricks bearer token to any host supplied in the `host` parameter, enabling SSRF and credential disclosure through the new Genie operations. Validate the host with `validateDatabricksWorkspaceHost` before constructing the URL and reject hosts outside Databricks-owned domains.</comment>

<file context>
@@ -0,0 +1,421 @@
+    .trim()
+    .replace(/^https?:\/\//, '')
+    .replace(/\/$/, '')
+  return `https://${normalizedHost}${path}`
+}
+
</file context>
Fix with cubic

Comment thread apps/sim/lib/logs/execution/logger.ts
* Blocks until the transaction-scoped advisory lock for `key` is held. The lock
* releases on commit or rollback.
*/
export async function acquireAdvisoryXactLock(tx: DbOrTx, tag: string, key: string): Promise<void> {

@cubic-dev-ai cubic-dev-ai Bot Sep 28, 2026 •

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2: DbOrTx also admits the pool-level db; passing it directly runs this statement in autocommit and releases the lock before the caller's work. Require DbTransaction in both lock helpers so this misuse fails at compile time.

Prompt for AI agents
Check if this issue is valid — if so, understand the root cause and fix it. At apps/sim/lib/db/advisory-locks.ts, line 22:

<comment>`DbOrTx` also admits the pool-level `db`; passing it directly runs this statement in autocommit and releases the lock before the caller's work. Require `DbTransaction` in both lock helpers so this misuse fails at compile time.</comment>

<file context>
@@ -0,0 +1,39 @@
+ * Blocks until the transaction-scoped advisory lock for `key` is held. The lock
+ * releases on commit or rollback.
+ */
+export async function acquireAdvisoryXactLock(tx: DbOrTx, tag: string, key: string): Promise<void> {
+  await tx.execute(sql`SELECT pg_advisory_xact_lock(hashtextextended(${key}, 0)) ${lockTag(tag)}`)
+}
</file context>
Fix with cubic

Comment thread packages/db/knowledge-projection.ts
)
return typed ? [query] : [`${query} is:issue`, `${query} is:pr`]
})
for (const query of oracleQueries) {

@cubic-dev-ai cubic-dev-ai Bot Sep 28, 2026 •

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2: Validate the GitHub free-text length before this loop. The schema allows 2,000 characters, so an overlong case makes real oracle /search/issues calls before searchGitHub returns its 256-character error, wasting quota and obscuring the intended failure.

(Based on your team's feedback about GitHub free-text query limits.)

Prompt for AI agents
Check if this issue is valid — if so, understand the root cause and fix it. At apps/sim/scripts/test-search-discussions-live.ts, line 254:

<comment>Validate the GitHub free-text length before this loop. The schema allows 2,000 characters, so an overlong case makes real oracle `/search/issues` calls before `searchGitHub` returns its 256-character error, wasting quota and obscuring the intended failure.

(Based on your team's feedback about GitHub free-text query limits.) </comment>

<file context>
@@ -0,0 +1,586 @@
+      )
+      return typed ? [query] : [`${query} is:issue`, `${query} is:pr`]
+    })
+    for (const query of oracleQueries) {
+      const nativeDateRange = (query.match(/"[^"]*"|\S+/g) ?? []).some((token) =>
+        /^updated:/i.test(token)
</file context>
Fix with cubic

export function boundedMeetingContent(content: string, limit = 200_000): string {
if (content.length <= limit) return content
const notice = '[Meeting content truncated. Open the original meeting for the remainder.]\n\n'
return notice + truncate(content, limit - notice.length, '')

@cubic-dev-ai cubic-dev-ai Bot Sep 28, 2026 •

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P3: truncate slices by UTF-16 code unit, so when the cut lands inside an astral character (emoji, math symbols) the result ends with a lone surrogate half, which this codebase's own string module documents as a fidelity hazard (see truncateAtCodePoint, "never cuts inside a surrogate pair"). Meeting transcripts routinely contain such characters, and the 199,922-unit cut in the default 200,000 limit will eventually split one. Use truncateAtCodePoint(content, limit - notice.length, '') instead, and update the import on line 1 accordingly.

Prompt for AI agents
Check if this issue is valid — if so, understand the root cause and fix it. At apps/sim/lib/sim-search/live/meeting-content.ts, line 7:

<comment>`truncate` slices by UTF-16 code unit, so when the cut lands inside an astral character (emoji, math symbols) the result ends with a lone surrogate half, which this codebase's own string module documents as a fidelity hazard (see `truncateAtCodePoint`, "never cuts inside a surrogate pair"). Meeting transcripts routinely contain such characters, and the 199,922-unit cut in the default 200,000 limit will eventually split one. Use `truncateAtCodePoint(content, limit - notice.length, '')` instead, and update the import on line 1 accordingly.</comment>

<file context>
@@ -0,0 +1,8 @@
+export function boundedMeetingContent(content: string, limit = 200_000): string {
+  if (content.length <= limit) return content
+  const notice = '[Meeting content truncated. Open the original meeting for the remainder.]\n\n'
+  return notice + truncate(content, limit - notice.length, '')
+}
</file context>
Fix with cubic

Comment thread apps/desktop/e2e/fixtures/updater.ts
})

const find = useChatFind({
chatId,

@cubic-dev-ai cubic-dev-ai Bot Sep 28, 2026 •

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P3: Passing the transient chatId as the find scope closes the find bar and clears its query when a pending chat receives its persisted id, even though the conversation did not change. Preserve state across the undefined-to-id transition or provide a stable conversation scope.

Prompt for AI agents
Check if this issue is valid — if so, understand the root cause and fix it. At apps/sim/app/workspace/[workspaceId]/home/components/mothership-chat/mothership-chat.tsx, line 694:

<comment>Passing the transient `chatId` as the find scope closes the find bar and clears its query when a pending chat receives its persisted id, even though the conversation did not change. Preserve state across the undefined-to-id transition or provide a stable conversation scope.</comment>

<file context>
@@ -695,6 +690,23 @@ export function MothershipChat({
   })
 
+  const find = useChatFind({
+    chatId,
+    messages,
+    hiddenUserByIndex: interactionPairing.hiddenUserByIndex,
</file context>
Fix with cubic

@greptile-apps

greptile-apps Bot commented Sep 28, 2026 •

Copy link
Copy Markdown
Contributor

RetriggerConfidence Score: 4/5

[High risk] Database contention fixes and new search connectors with API integration changes.

The PR should not merge until chat find counts and highlights the same rendered matches; the updater checkpoint and legacy-index concerns are follow-up improvements.

Findings

  1. P1 Citation spacing breaks chat find ▶
  2. P2 Native errors leave stale checkpoints ▶
  3. P2 Legacy index preserves write cost ▶

Summary

This release combines database and knowledge-index contention work with new live-search and warehouse integrations, chat improvements, and desktop updater staging verification.

  • Knowledge recovery, reconciliation, projections, counts, and outbox retention receive new bounded or scoped paths.
  • Chat gains a find menu and avoids repeated transcript loads; Databricks Genie, Snowflake Cortex Analyst, and additional search providers gain operations.
  • Desktop updates now wait for native staging and record installation outcomes.
Diagram
%%{init: {'theme': 'neutral'}}%%
flowchart LR
  A[Connector documents] --> B[Recovery and id-window reconciliation]
  B --> C[Knowledge projections]
  C --> D[Indexed and live search]
  E[Chat transcript] --> F[Chat find index]
  F --> G[Rendered-match highlighting]
  H[Desktop download] --> I[Native staging]
  I --> J[Install checkpoint]
  J --> K[Next-launch result]
Loading

Reviews (1) · Last reviewed commit: "feat(chat): add shared find menu (#8351)"

Comment thread apps/desktop/src/main/updater.ts
Comment thread packages/db/schema.ts

This branch was previously deployed

1 inactive deployment
Preview — 5c446f1b Deployed Sep 28, 2026 by vercel[bot]
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

requires-mothership-merge Has a companion PR on the mothership/copilot side — merge in lockstep

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants