Repository navigation
fix(lake-formation-tag-sync): option to keep connection key in schema name - #2802
Conversation
… name The schema segment of every qualifiedName is the Lake Formation DatabaseName with its first "_"-separated segment (the connection_map key) stripped off. Where the crawled schema name keeps that prefix, every schema, table and column lookup misses and is skipped as not found, and no connection_map.json value can restore the prefix. Add an opt-in keep_database_prefix input (default false). When set, the first segment is still used to resolve the connection, but the full DatabaseName is used as the schema name. Existing configurations are unaffected. Also skip and log a DatabaseName with no "_" instead of throwing IndexOutOfBoundsException when destructuring the split. Refs CSA-628 Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Fix the "Drom" typo in the remove_schema help text and give both schema-name options a concrete example. The keep_database_prefix help now states the default behaviour (the connection key is stripped from the schema name) so it is clear when to turn it on. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
End-to-end verificationTested on an internal Atlan tenant with a branch image built through the manual
After run 2 I confirmed in the catalog that all 33 assets carry the value written by the run. Backward compatibility: a workflow saved before this input existed ran unchanged. Argo applied the template default, and Unrelated issue found while testing: That stops every package that loads the cache. Not addressed here; I'll raise it separately. |
|
Run 1 (backwards compatibility): https://fs3.atlan.com/workflows/profile/csa-lake-formation-tag-sync-1768497496/runs?name=csa-lake-formation-tag-sync-1768497496-pdllf Run 2 (showing skip due to new logic): https://fs3.atlan.com/workflows/profile/csa-lake-formation-tag-sync-1768497496/runs?name=csa-lake-formation-tag-sync-1768497496-68ghc |
|
Nice work on this, Tin. The fix is clean and keeping it opt-in protects existing configs. One thing before merge: the green That way the PR has a record that the tests passed. Thanks! |
|
Thanks John, good catch on These tests don't need a tenant: they call |
jtgatlan
left a comment
There was a problem hiding this comment.
Thanks Tin, the local run plus the end-to-end runs on the test tenant (including the backward-compat check) give us everything we need. Approving.
Two small things for after merge: the marketplace-packages bump so the new option shows up for users, and a Linear ticket for the CustomMetadataCache primitiveType issue you found, since that one can stop every package on an affected tenant. Great work on this.
What
Lake Formation Tag Sync builds each schema qualifiedName from the Lake Formation
DatabaseNamewith its first_-separated segment (theconnection_map.jsonkey) stripped off:prod_sales_analytics→ keyprod, schemasales_analytics.That's correct when the crawled schema name excludes the prefix. When the crawled schema keeps it (
prod_sales_analytics), every schema, table and column lookup misses and is skipped as not found in update-only mode. Noconnection_map.jsonvalue can work around it: the code always puts/between the mapped connection and the stripped remainder, and an empty-string key can never match.Separately, a
DatabaseNamewith no_throwsIndexOutOfBoundsExceptionwhen the split result is destructured.How
keep_database_prefix("Keep connection key in schema name"), defaultfalse. Whentrue, the first segment still resolves the connection, but the fullDatabaseNameis used as the schema name.remove_schemastrips whichever schema name is in use.DatabaseNamewith no_is logged and skipped instead of throwing.remove_schemaand gave both schema-name options a concrete example; thekeep_database_prefixhelp states the default (theconnection_map.jsonkey is stripped from the schema name). Labels and keys are unchanged, but theremove_schemahelp text changes for every tenant.LakeFormationTagSyncCfg.ktregenerated frompackage.pklviagenCustomPkg.Default behaviour is unchanged, so existing tenant configurations that rely on the stripping keep working.
Testing
New
CSVProducerSchemaNameTest(no tenant needed), 4/4 passing locally on JDK 17:DatabaseNamefor schema, table and column qualifiedNamesremove_schemastrips the full name from the table nameDatabaseNameis skipped, not thrownExisting tenant-backed tests (
CSVProducerTest,LakeTagSynchronizerTest,EnumCreatorTest) were not run locally.Follow-up
After release, bump the
csa-lake-formation-tag-syncimage and add the new field inmarketplace-packages/packages/csa/lake-formation-tag-sync.Refs CSA-628
🤖 Generated with Claude Code