diff --git a/.gitignore b/.gitignore
index 872223992..edcddfd0b 100644
--- a/.gitignore
+++ b/.gitignore
@@ -9,6 +9,7 @@ core/target
utils/target
utils/dependency-reduced-pom.xml
examples/target
+azure/target
compat/target
core/node
deploy*.bat
diff --git a/README.md b/README.md
index e945c7770..ebdd2e145 100644
--- a/README.md
+++ b/README.md
@@ -485,6 +485,101 @@ The token is stored as **plaintext JSON protected by file permissions** — `060
`FileTokenStore` is safe to share between processes that sign in as the same identity: each update is written atomically (so a concurrent reader never sees a half-written credential), and when the identity provider rotates the refresh token on each refresh, the read-refresh-write is serialized across processes with a lock file so they do not race each other into an unnecessary re-prompt. The lock file's staleness is judged by its modification time, so this coordination assumes the processes share a clock — a single machine, or machines with synchronized clocks; under significant clock skew (for example a store directory on NFS shared across hosts) a live lock can be mis-judged stale or a dead one never expire. `clearCache()` removes the persisted entry under the same lock, but across processes it is best-effort: a peer that still holds a live in-memory token may legitimately re-persist afterwards (it always forces a fresh sign-in for the calling process).
+### Rotating Bearer Tokens (Microsoft Entra ID and Other Identity Platforms)
+
+Service identities - an Azure managed identity, a service principal, an AKS workload identity - authenticate with
+bearer tokens that expire every hour or every day. Select a **token provider** and the client fetches tokens itself,
+keeps them in a shared cache, and refreshes them in the background well before they expire, so long-lived senders and
+query clients keep reconnecting with a valid token.
+
+**Microsoft Entra ID from a connect string.** Add the optional `questdb-client-azure` artifact (same version as
+`questdb-client`); it brings Azure Identity and registers `token_provider=azure`:
+
+```xml
+
+ org.questdb
+ questdb-client-azure
+ 1.0.0
+
+```
+
+```java
+try (QuestDB db = QuestDB.connect(
+ "wss::addr=qdb1:9000,qdb2:9000;token_provider=azure;azure_resource=api://;")) {
+ // ... use db ...
+}
+```
+
+- `azure_resource` (required) is the application ID URI (`api://`) or the client ID of the QuestDB app
+ registration; the client requests the scope `/.default`.
+- `azure_client_id` (optional) selects a user-assigned managed identity or a workload identity.
+- The credential comes from Azure Identity's `DefaultAzureCredential`: an environment service principal
+ (`AZURE_TENANT_ID`, `AZURE_CLIENT_ID`, `AZURE_CLIENT_SECRET` or a certificate), a workload identity, a managed
+ identity, or developer tools. No secret ever goes into the connect string, so it stays safe to log and to put in
+ `QDB_CLIENT_CONF`.
+- `token_provider` requires `wss::`, and cannot be combined with `token`, `username`/`password` or an
+ application-supplied provider.
+- Every sender, query client and pooled connection built from equivalent connect strings shares one provider per
+ process: one refresher thread, one call to Entra at a time.
+
+On the server, validate tokens locally (`acl.oidc.groups.encoded.in.token=true`, `acl.oidc.groups.claim=roles`,
+`acl.oidc.sub.claim=oid`), set the QuestDB app registration to issue v2 tokens, and map its app roles to groups with
+`CREATE GROUP ... WITH EXTERNAL ALIAS ''`.
+
+**Any other identity platform.** Wrap a `TokenSource` in a `RefreshingTokenProvider` and pass the provider to
+`QuestDB.connect(config, provider)`, `httpTokenProvider(...)` or `withBearerTokenProvider(...)`. The application owns
+the provider: close it after the clients that use it.
+
+```java
+import io.questdb.client.cutlass.auth.ExpiringToken;
+import io.questdb.client.cutlass.auth.RefreshingTokenProvider;
+import io.questdb.client.cutlass.auth.TokenUnavailableException;
+
+RefreshingTokenProvider tokens = RefreshingTokenProvider.builder(() -> {
+ MyToken t = myIdentityPlatform.requestToken(); // runs on the provider's own thread
+ return new ExpiringToken(t.value(), t.expiresAtEpochMillis());
+}).build();
+try (QuestDB db = QuestDB.connect("wss::addr=qdb1:9000;", tokens)) {
+ // ... use db ...
+} finally {
+ tokens.close();
+}
+```
+
+A source reports a failure by throwing `TokenUnavailableException.retryable(...)` (network trouble, throttling,
+5xx) or `TokenUnavailableException.permanent(...)` (missing or wrong configuration); anything else counts as
+retryable. Never put a token or a raw response body in the message. The provider keeps retrying with backoff and
+keeps serving the current token while it is still valid.
+
+How failures are handled:
+
+- A token is refreshed at about half its lifetime, so an idle client always holds a valid one.
+- When the server rejects a token with `401`, the client asks the provider for a new one and retries the same server
+ once, immediately.
+- At startup, `initial_connect_retry=off` (the default) fails if no token can be obtained; `on` keeps retrying a
+ retryable provider failure within `reconnect_max_duration_millis`; `async` retries in the background.
+- Once connected, a store-and-forward sender rides out credential outages and `401`/`403` indefinitely, buffering
+ rows and reporting each failure to the error handler. Set `auth_failure_max_duration_millis` to make the sender
+ fail instead once such an outage lasts that long (unacknowledged rows stay on disk).
+
+### Connection Health
+
+A sender that rides out an outage keeps accepting rows, so check its connection health to see that it is not
+reaching the server. Reading health never blocks and never contains a credential, so it can back a health endpoint.
+
+```java
+ConnectionHealth.Aggregate health = db.health(); // every pooled sender and query client
+if (health.count(ConnectionHealth.State.RECONNECTING) > 0) {
+ long since = health.getOldestOutageSinceEpochMillis();
+ ConnectionHealth.Failure failure = health.getLastFailure(); // e.g. AUTH_REJECTED 401, CREDENTIAL_UNAVAILABLE
+ // ...
+}
+```
+
+`Sender.health()` and `QwpQueryClient.health()` return the same information for a single client: its state
+(`CONNECTING`, `CONNECTED`, `RECONNECTING`, `FAILED`, `CLOSED`), the last successful connect, the start of the current
+outage, the number of failed connect rounds since, and the last failure with its class.
+
### Explicit Timestamps
```java
@@ -529,6 +624,10 @@ schema::key1=value1;key2=value2;
| `tls_roots_password` | | Optional JKS/PKCS#12 password; omit when `tls_roots` is PEM |
| `connect_timeout` | _(OS)_ | TCP connect + TLS handshake timeout, in milliseconds |
| `auth_timeout_ms` | `15000` | Authentication/upgrade request timeout, in milliseconds |
+| `token_provider` | | Refreshing bearer-token provider, `wss` only: `azure` (needs `questdb-client-azure`) |
+| `azure_resource` | | `token_provider=azure`: application ID URI or client ID of the QuestDB app registration |
+| `azure_client_id` | | `token_provider=azure`: client ID of a user-assigned managed identity or workload identity |
+| `auth_failure_max_duration_millis` | _(none)_ | Ingest: fail the sender once a credential outage or `401`/`403` lasts this long |
### Pool keys (facade only)
diff --git a/azure/pom.xml b/azure/pom.xml
new file mode 100644
index 000000000..754ad21cb
--- /dev/null
+++ b/azure/pom.xml
@@ -0,0 +1,307 @@
+
+
+
+
+ 4.0.0
+
+ org.questdb
+ questdb-client-azure
+ 1.3.10-SNAPSHOT
+ jar
+ QuestDB client - Microsoft Entra ID token provider
+ token_provider=azure for the QuestDB Java client: Microsoft Entra ID bearer tokens through Azure Identity's DefaultAzureCredential
+ https://questdb.io/
+
+
+
+ Apache 2.0
+ https://www.apache.org/licenses/LICENSE-2.0.txt
+ repo
+
+
+
+
+
+ QuestDB Team
+ hello@questdb.io
+
+
+
+
+ https://github.com/questdb/java-questdb-client
+ scm:git:https://github.com/questdb/java-questdb-client.git
+ scm:git:https://github.com/questdb/java-questdb-client.git
+ HEAD
+
+
+
+ UTF-8
+
+ 1.3.8
+ 1.5.25
+
+
+
+
+
+ com.azure
+ azure-sdk-bom
+ ${azure-sdk-bom.version}
+ pom
+ import
+
+
+
+
+
+
+ org.questdb
+ questdb-client
+ ${project.version}
+
+
+ com.azure
+ azure-identity
+
+
+
+
+ junit
+ junit
+ 4.13.2
+ test
+
+
+ ch.qos.logback
+ logback-classic
+ ${logback.version}
+ test
+
+
+
+
+
+
+ org.apache.maven.plugins
+ maven-compiler-plugin
+ 3.11.0
+
+ ${javac.compile.source}
+ ${javac.compile.target}
+
+
+
+ org.apache.maven.plugins
+ maven-surefire-plugin
+ 3.5.3
+
+
+ false
+
+
+
+ org.apache.maven.plugins
+ maven-jar-plugin
+ 3.0.1
+
+
+
+
+ io.questdb.client.azure
+ ${project.version}
+
+
+
+
+
+
+ org.apache.maven.plugins
+ maven-deploy-plugin
+ 2.8.2
+
+ true
+
+
+
+
+
+
+
+ java11+
+
+ [11,)
+
+
+
+ 11
+ 11
+
+
+
+ java8
+
+ 1.8
+
+
+ 1.8
+ 1.8
+
+ 1.3.15
+
+
+
+ javadoc
+
+
+
+ org.apache.maven.plugins
+ maven-javadoc-plugin
+ 3.5.0
+
+
+ attach-javadocs
+
+ jar
+
+
+
+
+ none
+ ${javac.compile.source}
+ false
+
+
+
+
+
+
+ release-artifacts
+
+
+
+ org.apache.maven.plugins
+ maven-javadoc-plugin
+ 3.5.0
+
+
+ attach-javadocs
+
+ jar
+
+
+
+
+ none
+ ${javac.compile.source}
+ false
+
+
+
+ org.apache.maven.plugins
+ maven-source-plugin
+ 3.0.1
+
+
+ attach-sources
+
+ jar
+
+
+
+
+
+ org.apache.maven.plugins
+ maven-gpg-plugin
+ 3.2.7
+
+
+ sign-artifacts
+ verify
+
+ sign
+
+
+
+
+ gpg
+
+ --pinentry-mode
+ loopback
+
+
+
+
+
+
+
+ maven-central-publish
+
+
+
+
+ org.apache.maven.plugins
+ maven-enforcer-plugin
+ 3.0.0-M3
+
+
+ enforce-publish-from-jdk8
+
+ enforce
+
+
+
+
+ [1.8,1.9)
+ questdb-client-azure is published with questdb-client, from JDK 8.
+
+
+
+
+
+
+
+ org.sonatype.central
+ central-publishing-maven-plugin
+ 0.9.0
+ true
+
+ central
+ false
+ validated
+
+
+
+
+
+
+
diff --git a/azure/src/main/java/io/questdb/client/azure/AzureTokenProviderFactory.java b/azure/src/main/java/io/questdb/client/azure/AzureTokenProviderFactory.java
new file mode 100644
index 000000000..950f1bbcb
--- /dev/null
+++ b/azure/src/main/java/io/questdb/client/azure/AzureTokenProviderFactory.java
@@ -0,0 +1,91 @@
+/*+*****************************************************************************
+ * ___ _ ____ ____
+ * / _ \ _ _ ___ ___| |_| _ \| __ )
+ * | | | | | | |/ _ \/ __| __| | | | _ \
+ * | |_| | |_| | __/\__ \ |_| |_| | |_) |
+ * \__\_\\__,_|\___||___/\__|____/|____/
+ *
+ * Copyright (c) 2014-2019 Appsicle
+ * Copyright (c) 2019-2026 QuestDB
+ *
+ * Licensed under the Apache License, Version 2.0 (the "License");
+ * you may not use this file except in compliance with the License.
+ * You may obtain a copy of the License at
+ *
+ * http://www.apache.org/licenses/LICENSE-2.0
+ *
+ * Unless required by applicable law or agreed to in writing, software
+ * distributed under the License is distributed on an "AS IS" BASIS,
+ * WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied.
+ * See the License for the specific language governing permissions and
+ * limitations under the License.
+ *
+ ******************************************************************************/
+
+package io.questdb.client.azure;
+
+import com.azure.identity.DefaultAzureCredentialBuilder;
+import io.questdb.client.cutlass.auth.TokenProviderFactory;
+import io.questdb.client.cutlass.auth.TokenProviderSpec;
+import io.questdb.client.cutlass.auth.TokenSource;
+
+import java.util.Map;
+
+/**
+ * The {@code azure} token provider (design/qwp-token-provider-spec.md, section 7.5): Microsoft Entra ID tokens
+ * from Azure Identity's {@code DefaultAzureCredential}. Discovered through {@link java.util.ServiceLoader}; put
+ * this artifact on the class path or module path and connect with
+ *
+ * or let a connect string select {@code token_provider=azure}, which uses {@code DefaultAzureCredential}.
+ *
+ * Expiry and the refresh hint come from the library's {@link AccessToken}. Failures are classified as section
+ * 7.5 prescribes:
+ *
+ *
permanent - "no credential available in the chain" ({@link CredentialUnavailableException}), and
+ * authentication failures Entra reports: an HTTP 400 or 401 from the token endpoint, or a known AADSTS
+ * configuration error such as an invalid client or an application that does not exist;
+ *
retryable - network errors, timeouts, throttling, 5xx responses, and anything else. A
+ * {@code Retry-After} in seconds is passed on.
+ *
+ * Each attempt is bounded (30 s by default). Library exceptions never travel as a cause - their response
+ * objects reference the raw HTTP request, which can carry a client secret - only their sanitized message does.
+ */
+public final class AzureTokenSource implements TokenSource {
+ /**
+ * The bound on one token request.
+ */
+ public static final Duration DEFAULT_TIMEOUT = Duration.ofSeconds(30);
+ private static final String DEFAULT_SCOPE_SUFFIX = "/.default";
+ private static final int MAX_CAUSE_DEPTH = 16;
+ // Entra (AADSTS) errors that a retry cannot fix: the configuration is wrong, an operator must act.
+ private static final String[] PERMANENT_AADSTS = {
+ "AADSTS700016", // application not found in the directory
+ "AADSTS7000215", // invalid client secret
+ "AADSTS7000222", // client secret expired
+ "AADSTS700027", // client assertion signature invalid
+ "AADSTS70011", // invalid scope
+ "AADSTS500011", // resource principal not found: azure_resource is wrong
+ "AADSTS90002", // tenant not found
+ "AADSTS700213", // no matching federated identity record
+ "AADSTS70021", // no matching federated identity record
+ "AADSTS7000112", // application disabled
+ "AADSTS50049", // unknown or invalid instance
+ };
+ private final TokenRequestContext context;
+ private final TokenCredential credential;
+ private final String scope;
+ private final Duration timeout;
+
+ /**
+ * @param credential the Azure Identity credential
+ * @param resource the application ID URI ({@code api://}) or client ID of the QuestDB app
+ * registration; a trailing {@code /.default} is accepted
+ */
+ public AzureTokenSource(TokenCredential credential, String resource) {
+ this(credential, resource, DEFAULT_TIMEOUT);
+ }
+
+ /**
+ * @param credential the Azure Identity credential
+ * @param resource the application ID URI ({@code api://}) or client ID of the QuestDB app
+ * registration; a trailing {@code /.default} is accepted
+ * @param timeout the bound on one token request
+ */
+ public AzureTokenSource(TokenCredential credential, String resource, Duration timeout) {
+ if (credential == null) {
+ throw new IllegalArgumentException("credential must not be null");
+ }
+ if (resource == null || resource.isEmpty()) {
+ throw new IllegalArgumentException("resource must not be empty");
+ }
+ if (timeout == null || timeout.isNegative() || timeout.isZero()) {
+ throw new IllegalArgumentException("timeout must be positive");
+ }
+ this.credential = credential;
+ this.scope = resource.endsWith(DEFAULT_SCOPE_SUFFIX) ? resource : resource + DEFAULT_SCOPE_SUFFIX;
+ this.context = new TokenRequestContext().addScopes(scope);
+ this.timeout = timeout;
+ }
+
+ /**
+ * Classifies a failure from Azure Identity per section 7.5. Never attaches the library exception.
+ */
+ static TokenUnavailableException classify(Throwable failure, String scope) {
+ final String description = describe(failure);
+ Throwable t = failure;
+ for (int depth = 0; t != null && depth < MAX_CAUSE_DEPTH; depth++, t = t.getCause()) {
+ if (t instanceof CredentialUnavailableException) {
+ return TokenUnavailableException.permanent(
+ "no credential available in the Azure Identity chain for " + scope + ": " + description);
+ }
+ if (t instanceof HttpResponseException) {
+ final HttpResponse response = ((HttpResponseException) t).getResponse();
+ if (response != null) {
+ final int status = response.getStatusCode();
+ if (status == 400 || status == 401) {
+ return TokenUnavailableException.permanent(
+ "Entra rejected the token request for " + scope + " with HTTP " + status + ": " + description);
+ }
+ return TokenUnavailableException.retryable(
+ "the token request for " + scope + " failed with HTTP " + status + ": " + description,
+ retryAfterMillis(response));
+ }
+ }
+ final String message = t.getMessage();
+ if (message != null) {
+ for (String code : PERMANENT_AADSTS) {
+ if (message.contains(code)) {
+ return TokenUnavailableException.permanent(
+ "Entra rejected the token request for " + scope + " (" + code + "): " + description);
+ }
+ }
+ }
+ }
+ return TokenUnavailableException.retryable("the token request for " + scope + " failed: " + description);
+ }
+
+ private static String describe(Throwable t) {
+ final String message = t.getMessage();
+ return message == null ? t.getClass().getName() : t.getClass().getSimpleName() + ": " + message;
+ }
+
+ private static long retryAfterMillis(HttpResponse response) {
+ final String value;
+ try {
+ value = response.getHeaderValue("Retry-After");
+ } catch (RuntimeException e) {
+ return TokenUnavailableException.NO_RETRY_AFTER;
+ }
+ if (value == null) {
+ return TokenUnavailableException.NO_RETRY_AFTER;
+ }
+ try {
+ // delta-seconds only; an HTTP-date is ignored and the provider's own backoff applies
+ return TimeUnit.SECONDS.toMillis(Long.parseLong(value.trim()));
+ } catch (NumberFormatException e) {
+ return TokenUnavailableException.NO_RETRY_AFTER;
+ }
+ }
+
+ private static long refreshAtMillis(AccessToken token) {
+ try {
+ final OffsetDateTime refreshAt = token.getRefreshAt();
+ return refreshAt == null ? ExpiringToken.NO_REFRESH_AT : refreshAt.toInstant().toEpochMilli();
+ } catch (NoSuchMethodError e) {
+ // an azure-core older than the refresh hint
+ return ExpiringToken.NO_REFRESH_AT;
+ }
+ }
+
+ @Override
+ public ExpiringToken fetchToken() {
+ final AccessToken token;
+ try {
+ token = credential.getToken(context).block(timeout);
+ } catch (RuntimeException e) {
+ throw classify(e, scope);
+ }
+ if (token == null) {
+ throw TokenUnavailableException.retryable("Azure Identity returned no token for " + scope);
+ }
+ final OffsetDateTime expiresAt = token.getExpiresAt();
+ if (token.getToken() == null || expiresAt == null) {
+ // the shape of the problem only, never the response
+ throw TokenUnavailableException.retryable("Azure Identity returned a token for " + scope
+ + " without " + (token.getToken() == null ? "a value" : "an expiry"));
+ }
+ try {
+ return new ExpiringToken(token.getToken(), expiresAt.toInstant().toEpochMilli(), refreshAtMillis(token));
+ } catch (IllegalArgumentException e) {
+ throw TokenUnavailableException.retryable("Azure Identity returned an unusable token for " + scope
+ + ": " + e.getMessage());
+ }
+ }
+
+ /**
+ * @return the requested scope, {@code /.default}
+ */
+ public String getScope() {
+ return scope;
+ }
+
+ @Override
+ public String toString() {
+ return "AzureTokenSource{scope=" + scope + ", credential=" + credential.getClass().getSimpleName() + '}';
+ }
+}
diff --git a/azure/src/main/resources/META-INF/services/io.questdb.client.cutlass.auth.TokenProviderFactory b/azure/src/main/resources/META-INF/services/io.questdb.client.cutlass.auth.TokenProviderFactory
new file mode 100644
index 000000000..97f279b6a
--- /dev/null
+++ b/azure/src/main/resources/META-INF/services/io.questdb.client.cutlass.auth.TokenProviderFactory
@@ -0,0 +1 @@
+io.questdb.client.azure.AzureTokenProviderFactory
diff --git a/azure/src/test/java/io/questdb/client/azure/test/AzureTokenSourceTest.java b/azure/src/test/java/io/questdb/client/azure/test/AzureTokenSourceTest.java
new file mode 100644
index 000000000..d8d6b470b
--- /dev/null
+++ b/azure/src/test/java/io/questdb/client/azure/test/AzureTokenSourceTest.java
@@ -0,0 +1,292 @@
+/*+*****************************************************************************
+ * ___ _ ____ ____
+ * / _ \ _ _ ___ ___| |_| _ \| __ )
+ * | | | | | | |/ _ \/ __| __| | | | _ \
+ * | |_| | |_| | __/\__ \ |_| |_| | |_) |
+ * \__\_\\__,_|\___||___/\__|____/|____/
+ *
+ * Copyright (c) 2014-2019 Appsicle
+ * Copyright (c) 2019-2026 QuestDB
+ *
+ * Licensed under the Apache License, Version 2.0 (the "License");
+ * you may not use this file except in compliance with the License.
+ * You may obtain a copy of the License at
+ *
+ * http://www.apache.org/licenses/LICENSE-2.0
+ *
+ * Unless required by applicable law or agreed to in writing, software
+ * distributed under the License is distributed on an "AS IS" BASIS,
+ * WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied.
+ * See the License for the specific language governing permissions and
+ * limitations under the License.
+ *
+ ******************************************************************************/
+
+package io.questdb.client.azure.test;
+
+import com.azure.core.credential.AccessToken;
+import com.azure.core.credential.TokenCredential;
+import com.azure.core.credential.TokenRequestContext;
+import com.azure.core.exception.ClientAuthenticationException;
+import com.azure.core.exception.HttpResponseException;
+import com.azure.core.http.HttpHeaders;
+import com.azure.core.http.HttpResponse;
+import com.azure.identity.AuthenticationRequiredException;
+import com.azure.identity.CredentialUnavailableException;
+import io.questdb.client.azure.AzureTokenProviderFactory;
+import io.questdb.client.azure.AzureTokenSource;
+import io.questdb.client.cutlass.auth.ExpiringToken;
+import io.questdb.client.cutlass.auth.RefreshingTokenProvider;
+import io.questdb.client.cutlass.auth.TokenProviderRegistry;
+import io.questdb.client.cutlass.auth.TokenProviderSpec;
+import io.questdb.client.cutlass.auth.TokenSource;
+import io.questdb.client.cutlass.auth.TokenUnavailableException;
+import io.questdb.client.cutlass.qwp.client.QwpQueryClient;
+import io.questdb.client.impl.ConfigString;
+import io.questdb.client.impl.ConfigView;
+import org.junit.Assert;
+import org.junit.Test;
+import reactor.core.publisher.Flux;
+import reactor.core.publisher.Mono;
+
+import java.nio.ByteBuffer;
+import java.nio.charset.Charset;
+import java.time.Duration;
+import java.time.Instant;
+import java.time.OffsetDateTime;
+import java.time.ZoneOffset;
+import java.util.HashMap;
+import java.util.List;
+import java.util.Map;
+import java.util.concurrent.atomic.AtomicReference;
+
+/**
+ * The {@code azure} provider (design/qwp-token-provider-spec.md, section 7.5) against fake Azure Identity
+ * credentials - no network: expiry and refresh-hint mapping, the requested scope, the classification of library
+ * errors, and discovery through {@link java.util.ServiceLoader}.
+ */
+public class AzureTokenSourceTest {
+ private static final OffsetDateTime EXPIRES = OffsetDateTime.of(2027, 1, 15, 12, 0, 0, 0, ZoneOffset.UTC);
+ private static final OffsetDateTime REFRESH = EXPIRES.minusMinutes(30);
+
+ @Test
+ public void testAuthenticationFailuresEntraReportsArePermanent() {
+ assertPermanent(new ClientAuthenticationException("bad request", new FakeResponse(400, null)), "HTTP 400");
+ assertPermanent(new ClientAuthenticationException("unauthorized", new FakeResponse(401, null)), "HTTP 401");
+ // DefaultAzureCredential often wraps MSAL without a response: recognize the Entra error code
+ assertPermanent(new ClientAuthenticationException(
+ "AADSTS700016: Application with identifier 'x' was not found in the directory", null, (Object) null),
+ "AADSTS700016");
+ assertPermanent(new RuntimeException("wrapper", new IllegalStateException(
+ "AADSTS7000215: Invalid client secret provided.")), "AADSTS7000215");
+ }
+
+ @Test
+ public void testConnectStringResolvesTheAzureFactory() {
+ // C18 with the module present: the azure keys parse on both clients and nothing is fetched
+ String cfg = "wss::addr=localhost:9000;token_provider=azure;azure_resource=api://qdb/.default;"
+ + "azure_client_id=AAAAAAAA-BBBB-CCCC-DDDD-EEEEEEEEEEEE;";
+ TokenProviderSpec spec = TokenProviderSpec.parse(new ConfigView(ConfigString.parse(cfg)), true);
+ Assert.assertNotNull(spec);
+ Assert.assertTrue(spec.factory() instanceof AzureTokenProviderFactory);
+ Assert.assertEquals("api://qdb", spec.params().get(TokenProviderSpec.KEY_AZURE_RESOURCE));
+ Assert.assertEquals("azure[resource=api://qdb, client_id=aaaaaaaa-bbbb-cccc-dddd-eeeeeeeeeeee]", spec.describe());
+ QwpQueryClient.validateConfig(new ConfigView(ConfigString.parse(cfg)), true);
+ try {
+ TokenProviderSpec.parse(new ConfigView(ConfigString.parse("wss::addr=localhost:9000;token_provider=azure;")), true);
+ Assert.fail();
+ } catch (IllegalArgumentException e) {
+ Assert.assertTrue(e.getMessage(), e.getMessage().contains("requires azure_resource"));
+ }
+ try {
+ TokenProviderSpec.parse(new ConfigView(ConfigString.parse("wss::addr=localhost:9000;token_provider=nope;")), true);
+ Assert.fail();
+ } catch (IllegalArgumentException e) {
+ Assert.assertTrue(e.getMessage(), e.getMessage().contains("supported values: [azure]"));
+ }
+ }
+
+ @Test
+ public void testCredentialUnavailableIsPermanent() {
+ assertPermanent(new CredentialUnavailableException("EnvironmentCredential authentication unavailable"),
+ "no credential available");
+ assertPermanent(new AuthenticationRequiredException("interactive sign-in required", new TokenRequestContext()),
+ "no credential available");
+ assertPermanent(new RuntimeException("chain failed", new CredentialUnavailableException("none")),
+ "no credential available");
+ }
+
+ @Test
+ public void testExpiryAndRefreshHintAreMapped() {
+ AzureTokenSource source = new AzureTokenSource(
+ ctx -> Mono.just(new AccessToken("eyJ.token", EXPIRES, REFRESH)), "api://qdb");
+ ExpiringToken token = source.fetchToken();
+ Assert.assertEquals("eyJ.token", token.getToken());
+ Assert.assertEquals(EXPIRES.toInstant().toEpochMilli(), token.getExpiresAtEpochMillis());
+ Assert.assertEquals(REFRESH.toInstant().toEpochMilli(), token.getRefreshAtEpochMillis());
+
+ AzureTokenSource noHint = new AzureTokenSource(ctx -> Mono.just(new AccessToken("t", EXPIRES)), "api://qdb");
+ Assert.assertFalse(noHint.fetchToken().hasRefreshAt());
+ }
+
+ @Test
+ public void testFactoryBuildsADefaultAzureCredentialSourceWithoutNetworkIo() {
+ AzureTokenProviderFactory factory = new AzureTokenProviderFactory();
+ Map params = new HashMap<>();
+ params.put(TokenProviderSpec.KEY_AZURE_RESOURCE, "api://qdb");
+ params.put(TokenProviderSpec.KEY_AZURE_CLIENT_ID, "aaaaaaaa-bbbb-cccc-dddd-eeeeeeeeeeee");
+ TokenSource source = factory.createSource(params);
+ Assert.assertTrue(source instanceof AzureTokenSource);
+ Assert.assertEquals("api://qdb/.default", ((AzureTokenSource) source).getScope());
+ Assert.assertTrue(source.toString(), source.toString().contains("DefaultAzureCredential"));
+ }
+
+ @Test
+ public void testFactoryIsDiscoveredThroughServiceLoader() {
+ Assert.assertTrue(TokenProviderRegistry.findFactory("azure") instanceof AzureTokenProviderFactory);
+ List supported = TokenProviderRegistry.supportedProviders();
+ Assert.assertTrue(supported.toString(), supported.contains("azure"));
+ }
+
+ @Test
+ public void testLibraryExceptionsAreNeverAttached() {
+ // a library exception's response references the raw HTTP request, which can carry a client secret
+ TokenUnavailableException e = fetchFailure(new HttpResponseException("throttled", new FakeResponse(429, "3")));
+ Assert.assertNull(e.getCause());
+ }
+
+ @Test
+ public void testOtherFailuresAreRetryable() {
+ assertRetryable(new HttpResponseException("throttled", new FakeResponse(429, "7")), "HTTP 429", 7_000);
+ assertRetryable(new HttpResponseException("unavailable", new FakeResponse(503, null)), "HTTP 503", -1);
+ assertRetryable(new HttpResponseException("gone", new FakeResponse(410, "Wed, 21 Oct 2015 07:28:00 GMT")), "HTTP 410", -1);
+ assertRetryable(new IllegalStateException("connection reset"), "connection reset", -1);
+ assertRetryable(new ClientAuthenticationException("ManagedIdentityCredential: IMDS endpoint timed out", null,
+ (Object) null), "timed out", -1);
+ }
+
+ @Test
+ public void testResultsWithoutAValueOrExpiryAreRetryable() {
+ TokenUnavailableException e = expectFailure(new AzureTokenSource(
+ ctx -> Mono.just(new AccessToken("t", null)), "api://qdb"));
+ Assert.assertTrue(e.isRetryable());
+ Assert.assertTrue(e.getMessage(), e.getMessage().contains("without an expiry"));
+ e = expectFailure(new AzureTokenSource(ctx -> Mono.empty(), "api://qdb"));
+ Assert.assertTrue(e.isRetryable());
+ e = expectFailure(new AzureTokenSource(ctx -> Mono.just(new AccessToken("bad\r\ntoken", EXPIRES)), "api://qdb"));
+ Assert.assertTrue(e.isRetryable());
+ Assert.assertFalse("the token must not be echoed", e.getMessage().contains("bad"));
+ }
+
+ @Test
+ public void testScopeIsResourceDefault() {
+ AtomicReference> scopes = new AtomicReference<>();
+ TokenCredential credential = ctx -> {
+ scopes.set(ctx.getScopes());
+ return Mono.just(new AccessToken("t", EXPIRES));
+ };
+ new AzureTokenSource(credential, "api://qdb").fetchToken();
+ Assert.assertEquals("[api://qdb/.default]", scopes.get().toString());
+ new AzureTokenSource(credential, "11111111-2222-3333-4444-555555555555/.default").fetchToken();
+ Assert.assertEquals("[11111111-2222-3333-4444-555555555555/.default]", scopes.get().toString());
+ }
+
+ @Test(timeout = 30_000)
+ public void testSlowCredentialIsBoundedAndRetryable() {
+ AzureTokenSource source = new AzureTokenSource(ctx -> Mono.never(), "api://qdb", Duration.ofMillis(200));
+ long start = System.nanoTime();
+ TokenUnavailableException e = expectFailure(source);
+ Assert.assertTrue(e.isRetryable());
+ Assert.assertTrue("the attempt must be bounded", System.nanoTime() - start < 10_000_000_000L);
+ }
+
+ @Test(timeout = 30_000)
+ public void testWorksBehindTheRefreshingProvider() {
+ OffsetDateTime expires = OffsetDateTime.now(ZoneOffset.UTC).plusHours(1);
+ AzureTokenSource source = new AzureTokenSource(
+ ctx -> Mono.just(new AccessToken("eyJ.cached", expires, expires.minusMinutes(50))), "api://qdb");
+ try (RefreshingTokenProvider provider = RefreshingTokenProvider.builder(source).build()) {
+ Assert.assertTrue(provider.awaitReady(10_000));
+ Assert.assertEquals("eyJ.cached", provider.getToken().toString());
+ Assert.assertEquals(expires.toInstant().toEpochMilli(), provider.getTokenExpiresAtEpochMillis());
+ Assert.assertTrue(provider.getLastSuccessEpochMillis() <= Instant.now().toEpochMilli());
+ }
+ }
+
+ private static void assertPermanent(Throwable libraryFailure, String fragment) {
+ TokenUnavailableException e = fetchFailure(libraryFailure);
+ Assert.assertFalse("expected permanent: " + e.getMessage(), e.isRetryable());
+ Assert.assertTrue(e.getMessage(), e.getMessage().contains(fragment));
+ }
+
+ private static void assertRetryable(Throwable libraryFailure, String fragment, long retryAfterMillis) {
+ TokenUnavailableException e = fetchFailure(libraryFailure);
+ Assert.assertTrue("expected retryable: " + e.getMessage(), e.isRetryable());
+ Assert.assertTrue(e.getMessage(), e.getMessage().contains(fragment));
+ Assert.assertEquals(retryAfterMillis, e.getRetryAfterMillis());
+ }
+
+ private static TokenUnavailableException expectFailure(AzureTokenSource source) {
+ try {
+ source.fetchToken();
+ Assert.fail("expected the fetch to fail");
+ return null;
+ } catch (TokenUnavailableException e) {
+ return e;
+ }
+ }
+
+ private static TokenUnavailableException fetchFailure(Throwable libraryFailure) {
+ return expectFailure(new AzureTokenSource(ctx -> Mono.error(libraryFailure), "api://qdb"));
+ }
+
+ private static final class FakeResponse extends HttpResponse {
+ private final String retryAfter;
+ private final int status;
+
+ FakeResponse(int status, String retryAfter) {
+ super(null);
+ this.status = status;
+ this.retryAfter = retryAfter;
+ }
+
+ @Override
+ public Flux getBody() {
+ return Flux.empty();
+ }
+
+ @Override
+ public Mono getBodyAsByteArray() {
+ return Mono.empty();
+ }
+
+ @Override
+ public Mono getBodyAsString() {
+ return Mono.empty();
+ }
+
+ @Override
+ public Mono getBodyAsString(Charset charset) {
+ return Mono.empty();
+ }
+
+ @Override
+ public String getHeaderValue(String name) {
+ return "Retry-After".equalsIgnoreCase(name) ? retryAfter : null;
+ }
+
+ @Override
+ public HttpHeaders getHeaders() {
+ HttpHeaders headers = new HttpHeaders();
+ if (retryAfter != null) {
+ headers.set("Retry-After", retryAfter);
+ }
+ return headers;
+ }
+
+ @Override
+ public int getStatusCode() {
+ return status;
+ }
+ }
+}
diff --git a/core/src/main/java/io/questdb/client/ConnectionHealth.java b/core/src/main/java/io/questdb/client/ConnectionHealth.java
new file mode 100644
index 000000000..3ea731523
--- /dev/null
+++ b/core/src/main/java/io/questdb/client/ConnectionHealth.java
@@ -0,0 +1,290 @@
+/*+*****************************************************************************
+ * ___ _ ____ ____
+ * / _ \ _ _ ___ ___| |_| _ \| __ )
+ * | | | | | | |/ _ \/ __| __| | | | _ \
+ * | |_| | |_| | __/\__ \ |_| |_| | |_) |
+ * \__\_\\__,_|\___||___/\__|____/|____/
+ *
+ * Copyright (c) 2014-2019 Appsicle
+ * Copyright (c) 2019-2026 QuestDB
+ *
+ * Licensed under the Apache License, Version 2.0 (the "License");
+ * you may not use this file except in compliance with the License.
+ * You may obtain a copy of the License at
+ *
+ * http://www.apache.org/licenses/LICENSE-2.0
+ *
+ * Unless required by applicable law or agreed to in writing, software
+ * distributed under the License is distributed on an "AS IS" BASIS,
+ * WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied.
+ * See the License for the specific language governing permissions and
+ * limitations under the License.
+ *
+ ******************************************************************************/
+
+package io.questdb.client;
+
+import java.time.Instant;
+
+/**
+ * An immutable snapshot of a QWP client's connection health (design/qwp-token-provider-spec.md, section 8.4).
+ * A sender in store-and-forward mode retries an outage indefinitely - a revoked credential, an unreachable
+ * cluster - while the application keeps writing; this snapshot is how an application, or a health endpoint,
+ * finds out that it has not reached the server for an hour even when no error handler is registered.
+ *
+ * Reading it never waits on I/O or on a connection thread: {@link Sender#health()},
+ * {@code QwpQueryClient.health()} and {@link QuestDB#health()} return a published snapshot, safe to call from any
+ * thread at any rate. It never contains a credential.
+ */
+public final class ConnectionHealth {
+ /**
+ * Value of the timestamp getters when there is nothing to report.
+ */
+ public static final long NONE = -1L;
+ private final long failedRounds;
+ private final Failure lastFailure;
+ private final long lastConnectedAtEpochMillis;
+ private final long outageSinceEpochMillis;
+ private final State state;
+
+ public ConnectionHealth(
+ State state,
+ long lastConnectedAtEpochMillis,
+ long outageSinceEpochMillis,
+ long failedRounds,
+ Failure lastFailure
+ ) {
+ this.state = state;
+ this.lastConnectedAtEpochMillis = lastConnectedAtEpochMillis;
+ this.outageSinceEpochMillis = outageSinceEpochMillis;
+ this.failedRounds = failedRounds;
+ this.lastFailure = lastFailure;
+ }
+
+ private static String instant(long epochMillis) {
+ return epochMillis == NONE ? "none" : Instant.ofEpochMilli(epochMillis).toString();
+ }
+
+ /**
+ * Number of failed connect rounds - walks over the configured endpoints that ended without a connection -
+ * since {@link #getOutageSinceEpochMillis()}. Reset when a connection is established.
+ */
+ public long getFailedRounds() {
+ return failedRounds;
+ }
+
+ /**
+ * The most recent failed connect round: its class, a sanitized message and its time. Kept after the
+ * connection recovers. Null when no round has failed.
+ */
+ public Failure getLastFailure() {
+ return lastFailure;
+ }
+
+ /**
+ * Wall-clock time of the most recent successful upgrade, or {@link #NONE}.
+ */
+ public long getLastConnectedAtEpochMillis() {
+ return lastConnectedAtEpochMillis;
+ }
+
+ /**
+ * Start of the current period without a connection, or {@link #NONE} while connected.
+ */
+ public long getOutageSinceEpochMillis() {
+ return outageSinceEpochMillis;
+ }
+
+ public State getState() {
+ return state;
+ }
+
+ @Override
+ public String toString() {
+ return "ConnectionHealth{state=" + state
+ + ", lastConnectedAt=" + instant(lastConnectedAtEpochMillis)
+ + ", outageSince=" + instant(outageSinceEpochMillis)
+ + ", failedRounds=" + failedRounds
+ + ", lastFailure=" + lastFailure + '}';
+ }
+
+ /**
+ * The class of a failed connect round.
+ */
+ public enum FailureClass {
+ /**
+ * The token provider could not supply a credential; no endpoint was contacted.
+ */
+ CREDENTIAL_UNAVAILABLE,
+ /**
+ * An endpoint rejected the credential with {@code 401} or {@code 403}; see {@link Failure#getStatusCode()}.
+ */
+ AUTH_REJECTED,
+ /**
+ * Every reachable endpoint has the wrong role, such as a replica during a failover.
+ */
+ ROLE_REJECTED,
+ /**
+ * No endpoint could be reached: connection refused, timed out, or dropped.
+ */
+ TRANSPORT,
+ /**
+ * Anything else: another HTTP status at the upgrade, a protocol or capability mismatch.
+ */
+ OTHER
+ }
+
+ /**
+ * The connection state.
+ */
+ public enum State {
+ /**
+ * No successful upgrade yet.
+ */
+ CONNECTING,
+ /**
+ * Connected.
+ */
+ CONNECTED,
+ /**
+ * The connection was lost and is being re-established. For a query client, which connects on demand: the
+ * last connect or failover reconnect failed and the next operation will try again.
+ */
+ RECONNECTING,
+ /**
+ * Terminal: the client will not connect again.
+ */
+ FAILED,
+ /**
+ * Closed by the application.
+ */
+ CLOSED
+ }
+
+ /**
+ * Health across many connections - every pooled connection of a {@link QuestDB} handle: the number of
+ * connections in each state, the oldest outage and the most recent failure.
+ */
+ public static final class Aggregate {
+ private final int[] counts;
+ private final Failure lastFailure;
+ private final long oldestOutageSinceEpochMillis;
+
+ private Aggregate(int[] counts, long oldestOutageSinceEpochMillis, Failure lastFailure) {
+ this.counts = counts;
+ this.oldestOutageSinceEpochMillis = oldestOutageSinceEpochMillis;
+ this.lastFailure = lastFailure;
+ }
+
+ /**
+ * Aggregates the given snapshots.
+ */
+ public static Aggregate of(Iterable healths) {
+ int[] counts = new int[State.values().length];
+ long oldest = NONE;
+ Failure last = null;
+ for (ConnectionHealth h : healths) {
+ counts[h.state.ordinal()]++;
+ if (h.outageSinceEpochMillis != NONE && (oldest == NONE || h.outageSinceEpochMillis < oldest)) {
+ oldest = h.outageSinceEpochMillis;
+ }
+ if (h.lastFailure != null && (last == null || h.lastFailure.epochMillis > last.epochMillis)) {
+ last = h.lastFailure;
+ }
+ }
+ return new Aggregate(counts, oldest, last);
+ }
+
+ /**
+ * Number of connections in {@code state}.
+ */
+ public int count(State state) {
+ return counts[state.ordinal()];
+ }
+
+ /**
+ * The most recent failed connect round across all connections, or null.
+ */
+ public Failure getLastFailure() {
+ return lastFailure;
+ }
+
+ /**
+ * The oldest {@code outage_since} across all connections, or {@link #NONE} when all are connected.
+ */
+ public long getOldestOutageSinceEpochMillis() {
+ return oldestOutageSinceEpochMillis;
+ }
+
+ /**
+ * Number of connections aggregated.
+ */
+ public int total() {
+ int n = 0;
+ for (int c : counts) {
+ n += c;
+ }
+ return n;
+ }
+
+ @Override
+ public String toString() {
+ StringBuilder sb = new StringBuilder("ConnectionHealth.Aggregate{");
+ State[] states = State.values();
+ for (int i = 0; i < states.length; i++) {
+ sb.append(states[i].name().toLowerCase(java.util.Locale.ROOT)).append('=').append(counts[i]).append(", ");
+ }
+ return sb.append("oldestOutageSince=").append(instant(oldestOutageSinceEpochMillis))
+ .append(", lastFailure=").append(lastFailure).append('}').toString();
+ }
+ }
+
+ /**
+ * A failed connect round. Never carries a credential.
+ */
+ public static final class Failure {
+ private final long epochMillis;
+ private final FailureClass failureClass;
+ private final String message;
+ private final int statusCode;
+
+ public Failure(FailureClass failureClass, int statusCode, String message, long epochMillis) {
+ this.failureClass = failureClass;
+ this.statusCode = statusCode;
+ this.message = message;
+ this.epochMillis = epochMillis;
+ }
+
+ /**
+ * Wall-clock time of the failure.
+ */
+ public long getEpochMillis() {
+ return epochMillis;
+ }
+
+ public FailureClass getFailureClass() {
+ return failureClass;
+ }
+
+ /**
+ * A sanitized description of the failure, at most 256 characters.
+ */
+ public String getMessage() {
+ return message;
+ }
+
+ /**
+ * The HTTP status of the upgrade rejection - {@code 401} or {@code 403} for
+ * {@link FailureClass#AUTH_REJECTED} - or {@code 0} when none applies.
+ */
+ public int getStatusCode() {
+ return statusCode;
+ }
+
+ @Override
+ public String toString() {
+ return "Failure{class=" + failureClass + (statusCode > 0 ? ", status=" + statusCode : "")
+ + ", at=" + instant(epochMillis) + ", message=" + message + '}';
+ }
+ }
+}
diff --git a/core/src/main/java/io/questdb/client/HttpTokenProvider.java b/core/src/main/java/io/questdb/client/HttpTokenProvider.java
index cf28bbf50..02b318e68 100644
--- a/core/src/main/java/io/questdb/client/HttpTokenProvider.java
+++ b/core/src/main/java/io/questdb/client/HttpTokenProvider.java
@@ -47,6 +47,18 @@
* resolution to the OS. A provider that builds its own HTTP client should bound its connect and TLS
* handshake likewise, or a black-holed token endpoint stalls a flush for the OS connect timeout. An exception from {@link #getToken()} fails the
* in-flight flush (HTTP) or the connection attempt (WebSocket).
+ *
+ * Failure classification (WebSocket). A {@link io.questdb.client.cutlass.auth.TokenUnavailableException} marked
+ * retryable is a transient credential outage: a {@code Sender} whose initial connect retries
+ * ({@code initial_connect_retry=on}) keeps retrying it within {@code reconnect_max_duration_millis}. Any other
+ * exception counts as a permanent credential failure, so startup fails fast, as it does for an OIDC device-flow
+ * provider that is not signed in yet. Once a sender has connected, every credential failure is retried while
+ * store-and-forward keeps the rows.
+ *
+ * For a token that rotates on a schedule - a Microsoft Entra ID managed identity or service principal, an OAuth
+ * client-credentials grant - wrap a {@link io.questdb.client.cutlass.auth.TokenSource} in a
+ * {@link io.questdb.client.cutlass.auth.RefreshingTokenProvider}, which caches the token and refreshes it in the
+ * background before it expires.
*
* @see QuestDB#connect(CharSequence, HttpTokenProvider)
* @see QuestDBBuilder#httpTokenProvider(HttpTokenProvider)
@@ -106,4 +118,23 @@ static void validateToken(CharSequence token) {
* @return the current HTTP authentication token
*/
CharSequence getToken();
+
+ /**
+ * Tells the provider that a server rejected a token it handed out, so that a caching provider can refresh
+ * early. WebSocket ingest and query clients call this when an upgrade that presented {@code token} is
+ * answered with {@code 401} - and, when the server sent a {@code WWW-Authenticate: Bearer} challenge that
+ * carries an {@code error}, only when that error is {@code invalid_token}. They then call {@link #getToken()}
+ * again and, if the token changed, retry the same endpoint once. A {@code 403} is never reported: it is an
+ * authorization decision that a new token does not change.
+ *
+ * The call may block briefly while the provider refreshes - a {@code RefreshingTokenProvider} waits up to its
+ * {@code forced_wait}, 5 s by default - and must honour an interrupt by returning promptly with the
+ * interrupt flag still set. It should not throw; the clients ignore anything it throws. The default does
+ * nothing, which suits a provider that always returns its freshest token anyway.
+ *
+ * @param token the rejected token, without the {@code "Bearer "} prefix
+ * @param httpStatus the HTTP status of the rejection
+ */
+ default void onTokenRejected(CharSequence token, int httpStatus) {
+ }
}
diff --git a/core/src/main/java/io/questdb/client/QuestDB.java b/core/src/main/java/io/questdb/client/QuestDB.java
index cf669a166..8cfc437b7 100644
--- a/core/src/main/java/io/questdb/client/QuestDB.java
+++ b/core/src/main/java/io/questdb/client/QuestDB.java
@@ -160,6 +160,18 @@ static QuestDB connect(CharSequence configurationString, HttpTokenProvider token
*/
Sender borrowSender();
+ /**
+ * Aggregated connection health of every pooled connection - ingest senders and query clients: the number of
+ * connections in each state, the oldest outage and the most recent failed connect round
+ * (design/qwp-token-provider-spec.md, section 8.4). Cheap and safe to call from any thread, so it can back a
+ * health endpoint; it never waits on I/O and never contains a credential.
+ *
+ * @return the aggregate
+ */
+ default ConnectionHealth.Aggregate health() {
+ throw new UnsupportedOperationException("connection health is not available from this QuestDB implementation");
+ }
+
/**
* Shuts down the pools and their published clients. Idempotent. Threads
* currently blocked in {@link #borrowSender()} or {@link Query#submit()}
diff --git a/core/src/main/java/io/questdb/client/QuestDBBuilder.java b/core/src/main/java/io/questdb/client/QuestDBBuilder.java
index e846ad129..b37cdcf42 100644
--- a/core/src/main/java/io/questdb/client/QuestDBBuilder.java
+++ b/core/src/main/java/io/questdb/client/QuestDBBuilder.java
@@ -192,6 +192,10 @@ public QuestDB build() {
throw new IllegalArgumentException(
"httpTokenProvider cannot be combined with token, username, or password in the configuration");
}
+ if (httpTokenProvider != null && view.has("token_provider")) {
+ throw new IllegalArgumentException(
+ "httpTokenProvider cannot be combined with token_provider in the configuration");
+ }
// Validate the single cluster config exactly as both pools will, but
// without connecting: the full Sender parse plus validateParameters
// (ingress value keys are registry-STRING, so only the real parse
diff --git a/core/src/main/java/io/questdb/client/Sender.java b/core/src/main/java/io/questdb/client/Sender.java
index 645d7b254..57ff7bf59 100644
--- a/core/src/main/java/io/questdb/client/Sender.java
+++ b/core/src/main/java/io/questdb/client/Sender.java
@@ -25,6 +25,8 @@
package io.questdb.client;
import io.questdb.client.cutlass.auth.AuthUtils;
+import io.questdb.client.cutlass.auth.TokenProviderRegistry;
+import io.questdb.client.cutlass.auth.TokenProviderSpec;
import io.questdb.client.cutlass.line.AbstractLineTcpSender;
import io.questdb.client.cutlass.line.LineChannel;
import io.questdb.client.cutlass.line.LineSenderException;
@@ -77,6 +79,7 @@
import java.util.Base64;
import java.util.List;
import java.util.concurrent.TimeUnit;
+import java.util.concurrent.atomic.AtomicBoolean;
import java.util.function.Supplier;
/**
@@ -623,6 +626,22 @@ default long getAckedFsn() {
return -1L;
}
+ /**
+ * A snapshot of this sender's connection health: whether it is connected, since when it has been without a
+ * connection, how many connect rounds have failed since, and the most recent failure with its class
+ * (design/qwp-token-provider-spec.md, section 8.4). A store-and-forward sender retries an outage - a revoked
+ * credential, an unreachable cluster - indefinitely while the application keeps writing; this is how to see
+ * that it has not reached the server. Cheap and safe to call from any thread: it never waits on I/O or on the
+ * sender's I/O thread, so it can back a health endpoint. It never contains a credential.
+ *
+ * @return the current health
+ * @throws UnsupportedOperationException for transports that do not track connection health: only the
+ * WebSocket (QWP) sender does
+ */
+ default ConnectionHealth health() {
+ throw new UnsupportedOperationException("connection health is only available for WebSocket (QWP) senders");
+ }
+
/**
* Add a column with a 32-bit signed integer value.
*
@@ -947,6 +966,8 @@ enum Transport {
*/
final class LineSenderBuilder {
private static final int AUTO_FLUSH_DISABLED = 0;
+ // Warn once per process when an application-supplied token provider is used over ws:: (spec section 9).
+ private static final AtomicBoolean CLEARTEXT_PROVIDER_WARNED = new AtomicBoolean();
private static final String TLS_ROOTS_INSECURE_CONFIG_ERROR = "tls_roots cannot be combined with tls_verify=unsafe_off; remove tls_verify to use custom roots, or remove tls_roots to disable certificate validation";
// close() drain timeout. Default applied at build() time. 0 or -1
// means "fast close" (skip the drain entirely); any positive value
@@ -1041,6 +1062,8 @@ final class LineSenderBuilder {
OrphanScanner.QUARANTINE_SLOT_INFIX;
private final ObjList hosts = new ObjList<>();
private final IntList ports = new IntList();
+ // Optional authentication-outage deadline (spec section 8.5); 0 = not set, the default.
+ private long authFailureMaxDurationMillis;
private long authTimeoutMillis = QwpWebSocketSender.DEFAULT_AUTH_TIMEOUT_MS;
private int autoFlushBytes = PARAMETER_NOT_SET_EXPLICITLY;
private int autoFlushIntervalMillis = PARAMETER_NOT_SET_EXPLICITLY;
@@ -1093,6 +1116,10 @@ final class LineSenderBuilder {
private int httpTimeout = PARAMETER_NOT_SET_EXPLICITLY;
private String httpToken;
private HttpTokenProvider httpTokenProvider;
+ // A token_provider selected by the connect string (wss:: only). Resolved through the process-wide
+ // TokenProviderRegistry when build() connects - never at parse time, so validating a configuration
+ // fetches no token - and released when the sender closes.
+ private TokenProviderSpec tokenProviderSpec;
// Drives the initial-connect strategy. null means "not set
// explicitly", which build() resolves to SYNC when any reconnect_*
// knob was tuned by the user, otherwise OFF. SYNC retries on the
@@ -1274,6 +1301,33 @@ public LineSenderBuilder connectTimeoutMillis(int millis) {
return this;
}
+ /**
+ * Optional authentication-outage deadline (design/qwp-token-provider-spec.md, section 8.5): the sender
+ * becomes terminal when a connect round fails with an authentication-class failure - the token provider
+ * could not supply a credential, or an endpoint answered {@code 401}/{@code 403} - once that outage has
+ * lasted this long. Not set by default: such outages are retried indefinitely while store-and-forward
+ * keeps the rows. The clock starts at the first authentication-class failure of an outage and only a
+ * successful upgrade resets it; failures of other classes neither reset nor fire it. Applies while the
+ * sender is established and during an {@code async} initial connect; orphan drains are unaffected. When
+ * it fires, the error names the failure class and the elapsed time, goes to the error handler as
+ * terminal, and is thrown from later producer calls; unacknowledged rows stay in on-disk
+ * store-and-forward for a later sender or an orphan drain. WebSocket transport only. Connect-string key:
+ * {@code auth_failure_max_duration_millis}.
+ *
+ * @param millis the deadline, {@code > 0}
+ * @return this instance for method chaining
+ */
+ public LineSenderBuilder authFailureMaxDurationMillis(long millis) {
+ if (protocol != PARAMETER_NOT_SET_EXPLICITLY && protocol != PROTOCOL_WEBSOCKET) {
+ throw new LineSenderException("auth_failure_max_duration_millis is only supported for WebSocket transport");
+ }
+ if (millis <= 0) {
+ throw new LineSenderException("auth_failure_max_duration_millis must be > 0: ").put(millis);
+ }
+ this.authFailureMaxDurationMillis = millis;
+ return this;
+ }
+
/**
* Per-endpoint timeout on the WebSocket upgrade response read. Default
* {@value QwpWebSocketSender#DEFAULT_AUTH_TIMEOUT_MS} ms.
@@ -1455,355 +1509,21 @@ public Sender build() {
}
if (protocol == PROTOCOL_WEBSOCKET) {
- if (hosts.size() < 1) {
- throw new LineSenderException("WebSocket transport requires at least one host:port pair");
- }
-
- int actualAutoFlushRows = autoFlushRows == PARAMETER_NOT_SET_EXPLICITLY ? DEFAULT_WS_AUTO_FLUSH_ROWS : autoFlushRows;
- int actualAutoFlushBytes = autoFlushBytes == PARAMETER_NOT_SET_EXPLICITLY ? DEFAULT_WS_AUTO_FLUSH_BYTES : autoFlushBytes;
- long actualAutoFlushIntervalNanos = autoFlushIntervalMillis == PARAMETER_NOT_SET_EXPLICITLY
- ? DEFAULT_WS_AUTO_FLUSH_INTERVAL_NANOS
- : TimeUnit.MILLISECONDS.toNanos(autoFlushIntervalMillis);
-
- Supplier wsAuthHeader = buildWebSocketAuthHeader();
-
- ClientTlsConfiguration wsTlsConfig = null;
- if (tlsEnabled) {
- assert trustStorePassword == null || trustStorePath != null;
- wsTlsConfig = new ClientTlsConfiguration(
- trustStorePath,
- trustStorePassword,
- tlsValidationMode == TlsValidationMode.DEFAULT
- ? ClientTlsConfiguration.TLS_VALIDATION_MODE_FULL
- : ClientTlsConfiguration.TLS_VALIDATION_MODE_NONE
- );
+ if (tokenProviderSpec == null) {
+ return buildWebSocket(httpTokenProvider);
}
-
- // Setting sfDir enables store-and-forward (mmap'd, recoverable
- // across sender restarts); omitting it gives memory-only mode
- // (same lock-free architecture, no disk involvement).
- // Durability-combination validation lives in validateParameters
- // so build() and no-connect validation apply the same rules.
- long actualSfMaxSegmentBytes = sfMaxSegmentBytes == PARAMETER_NOT_SET_EXPLICITLY
- ? DEFAULT_SEGMENT_BYTES
- : sfMaxSegmentBytes;
- // Default cap depends on backing: RAM (memory mode) is tight
- // by default; disk (SF mode) is cheap so the default is
- // generous enough that normal traffic never hits it.
- long defaultMaxTotal = sfDir == null
- ? DEFAULT_MAX_BYTES_MEMORY
- : DEFAULT_MAX_BYTES_SF;
- long actualSfMaxTotalBytes = sfMaxTotalBytes == PARAMETER_NOT_SET_EXPLICITLY
- ? Math.max(defaultMaxTotal, actualSfMaxSegmentBytes * 2)
- : sfMaxTotalBytes;
- long actualCloseFlushTimeoutMillis = closeFlushTimeoutMillis == CLOSE_FLUSH_TIMEOUT_NOT_SET
- ? DEFAULT_CLOSE_FLUSH_TIMEOUT_MILLIS
- : closeFlushTimeoutMillis;
- long actualReconnectMaxDurationMillis =
- reconnectMaxDurationMillis == PARAMETER_NOT_SET_EXPLICITLY
- ? CursorWebSocketSendLoop.DEFAULT_RECONNECT_MAX_DURATION_MILLIS
- : reconnectMaxDurationMillis;
- long actualReconnectInitialBackoffMillis =
- reconnectInitialBackoffMillis == PARAMETER_NOT_SET_EXPLICITLY
- ? CursorWebSocketSendLoop.DEFAULT_RECONNECT_INITIAL_BACKOFF_MILLIS
- : reconnectInitialBackoffMillis;
- long actualReconnectMaxBackoffMillis =
- reconnectMaxBackoffMillis == PARAMETER_NOT_SET_EXPLICITLY
- ? CursorWebSocketSendLoop.DEFAULT_RECONNECT_MAX_BACKOFF_MILLIS
- : reconnectMaxBackoffMillis;
- // Resolve the initial-connect mode. An explicit user choice
- // (via initialConnectMode/initialConnectRetry, or the
- // initial_connect_retry conf key) wins unconditionally --
- // including initial_connect_retry=off paired with a tuned
- // reconnect budget. When the user left it unset and tuned
- // any reconnect_* knob, promote to SYNC so the budget they
- // wrote actually applies to the first connect: the knob
- // name reads as a generic retry budget but the underlying
- // path only governs reconnects from an established
- // connection, and silently ignoring the budget on the
- // initial connect is the canonical footgun this implicit
- // upgrade removes.
- InitialConnectMode actualInitialConnectMode;
- if (initialConnectMode != null) {
- actualInitialConnectMode = initialConnectMode;
- } else if (reconnectMaxDurationMillis != PARAMETER_NOT_SET_EXPLICITLY
- || reconnectInitialBackoffMillis != PARAMETER_NOT_SET_EXPLICITLY
- || reconnectMaxBackoffMillis != PARAMETER_NOT_SET_EXPLICITLY) {
- actualInitialConnectMode = InitialConnectMode.SYNC;
- } else {
- actualInitialConnectMode = InitialConnectMode.OFF;
- }
- long actualDurableAckKeepaliveIntervalMillis =
- durableAckKeepaliveIntervalMillis == DURABLE_ACK_KEEPALIVE_NOT_SET
- ? CursorWebSocketSendLoop.DEFAULT_DURABLE_ACK_KEEPALIVE_INTERVAL_MILLIS
- : durableAckKeepaliveIntervalMillis;
- int actualMaxFrameRejections = maxFrameRejections != PARAMETER_NOT_SET_EXPLICITLY
- ? maxFrameRejections
- : CursorWebSocketSendLoop.DEFAULT_MAX_HEAD_FRAME_REJECTIONS;
- long actualPoisonMinEscalationWindowMillis = poisonMinEscalationWindowMillis != PARAMETER_NOT_SET_EXPLICITLY
- ? poisonMinEscalationWindowMillis
- : CursorWebSocketSendLoop.DEFAULT_POISON_MIN_ESCALATION_WINDOW_MILLIS;
- long actualCatchUpCapGapMinEscalationWindowMillis =
- catchUpCapGapMinEscalationWindowMillis != PARAMETER_NOT_SET_EXPLICITLY
- ? catchUpCapGapMinEscalationWindowMillis
- : CursorWebSocketSendLoop.DEFAULT_CATCHUP_CAP_GAP_MIN_ESCALATION_WINDOW_MILLIS;
-
- // sfDir is the parent (group root); the actual slot lives
- // under sfDir/senderId. This is what the engine sees — the
- // slot lock and segment files all live one level deeper than
- // the user-supplied path. Memory mode skips this composition
- // (slotPath stays null).
- //
- // The slot ctor inside CursorSendEngine creates the slot
- // directory itself, but Files.mkdir is non-recursive — so we
- // must ensure the parent group root exists first.
- String slotPath;
- if (sfDir == null) {
- slotPath = null;
- } else {
- if (!Files.exists(sfDir)) {
- int rc = Files.mkdir(sfDir, Files.DIR_MODE_DEFAULT);
- // mkdir is non-zero on failure, but "already exists"
- // is one such failure. Multiple SF senders sharing one
- // sf_dir can be built concurrently (the pool calls
- // build() outside its lock), so two threads can both
- // pass the exists() check and race into mkdir; the
- // loser gets EEXIST. Treat a benign creation race --
- // the dir now exists -- as success and only fail when
- // the directory is genuinely absent afterwards.
- if (rc != 0 && !Files.exists(sfDir)) {
- throw new LineSenderException(
- "could not create sf_dir: " + sfDir + " rc=" + rc);
- }
- }
- if (sfDurability == SfDurability.PERIODIC
- && Files.fsyncParentDir(sfDir) != 0) {
- throw new LineSenderException(
- "could not sync parent directory for sf_dir: " + sfDir);
- }
- slotPath = sfDir + "/" + senderId;
- }
- long actualSfAppendDeadlineNanos =
- sfAppendDeadlineMillis == PARAMETER_NOT_SET_EXPLICITLY
- ? CursorSendEngine.DEFAULT_APPEND_DEADLINE_NANOS
- : sfAppendDeadlineMillis * 1_000_000L;
- long actualSfSyncIntervalNanos = sfDurability == SfDurability.PERIODIC
- ? (sfSyncIntervalMillis == PARAMETER_NOT_SET_EXPLICITLY
- ? DEFAULT_SF_SYNC_INTERVAL_MILLIS : sfSyncIntervalMillis) * 1_000_000L
- : 0L;
- QwpWebSocketSender connected = null;
- // The parent-anchored logical lock is stable across a slot rename. Keep it
- // from before the directory-local lock is acquired until connect() has either
- // adopted that engine or quarantine has closed, renamed and recreated it.
- // This closes the inode-swap window in which an already-queued orphan drainer
- // could otherwise acquire the renamed directory's old .lock and later operate
- // on the fresh slot through the original pathname.
- try (SlotLock logicalSlotLock = slotPath == null
- ? null
- : SlotLock.acquireLogical(slotPath)) {
- // The constructor's own recovery seed can also fail terminally, and
- // not only as UnreplayableSlotException: when SegmentRing.openExisting
- // had to skip an unreadable segment it throws SfRecoveryException (it
- // constructs UnreplayableSlotException nowhere), and where it cannot
- // even prove the chain's identity -- no manifest -- it quarantines the
- // corrupt files and returns an EMPTY recovery rather than refusing.
- // Either way the frame range cannot be shown already-acked, so recovery
- // sets the slot aside rather than risk seeding the ack cursor past
- // frames that were never delivered. All three types below are load
- // bearing; narrowing this catch to UnreplayableSlotException would
- // restore the permanent build() brick for the segment-skip case. That verdict gets
- // the exact same quarantine-and-continue treatment as the connect()-time
- // verdict below -- constructing cursorEngine is not inside the loop below,
- // so a throw here would otherwise escape build() entirely, uncaught.
- // quarantineTornSlot(null, ...) renames the WHOLE slot directory aside
- // (not just the unreadable segment file) before building the replacement
- // at the original slotPath, so the replacement starts on a genuinely empty
- // directory with nothing left to skip -- it cannot throw the same way
- // twice, which is what makes looping unnecessary here.
- boolean quarantined = false;
- CursorSendEngine cursorEngine;
- try {
- try {
- cursorEngine = new CursorSendEngine(
- slotPath, actualSfMaxSegmentBytes,
- actualSfMaxTotalBytes, actualSfAppendDeadlineNanos,
- actualSfSyncIntervalNanos);
- } catch (SfSanitizedResidueException first) {
- // NOT terminal, and it must be intercepted ahead of its
- // SfRecoveryException parent below. Recovery durably zeroed
- // proven-dead sealed residue BEFORE failing closed, so the
- // chain on disk is already healed: quarantining here would
- // set aside a slot whose backlog replays perfectly. Retry
- // once over the healed chain; a repeat is genuine and takes
- // the terminal arm.
- LOG.info("sf slot {}: sealed residue sanitized during recovery ({}); "
- + "retrying over the healed chain",
- slotPath, first.getMessage());
- cursorEngine = new CursorSendEngine(
- slotPath, actualSfMaxSegmentBytes,
- actualSfMaxTotalBytes, actualSfAppendDeadlineNanos,
- actualSfSyncIntervalNanos);
- }
- } catch (UnreplayableSlotException | SfRecoveryException
- | MmapSegmentCorruptionException e) {
- // The terminal recovery verdicts, and the only ones build()
- // sets a slot aside for. UnreplayableSlotException says the
- // symbol dictionary cannot be rebuilt from any source;
- // SfRecoveryException and MmapSegmentCorruptionException say
- // the durable chain itself is proven corrupt or incomplete.
- // None of the three clears on a retry, and senderId is stable
- // with a not-fully-drained slot retained on close -- so
- // without this arm every restart re-recovers the same slot and
- // throws again, and the application cannot construct a Sender
- // at all, not even to BUFFER new rows.
- //
- // Deliberately NOT catching plain MmapSegmentException or
- // SfOperationalException: those are operational (EMFILE,
- // ENOMEM, an unreadable-but-possibly-intact file). Aborting
- // startup on them is correct; quarantining on them would
- // convert a transient into the permanent loss of a healthy
- // slot's durable frames.
- if (slotPath == null) {
- throw e;
- }
- quarantined = true;
- cursorEngine = quarantineTornSlot(
- null, e, sfDir, senderId, slotPath, actualSfMaxSegmentBytes,
- actualSfMaxTotalBytes, actualSfAppendDeadlineNanos,
- actualSfSyncIntervalNanos, errorHandler);
- }
- int actualErrorInboxCapacity = errorInboxCapacity != PARAMETER_NOT_SET_EXPLICITLY
- ? errorInboxCapacity
- : io.questdb.client.cutlass.qwp.client.sf.cursor.SenderErrorDispatcher.DEFAULT_CAPACITY;
- int actualConnectionListenerInboxCapacity = connectionListenerInboxCapacity != PARAMETER_NOT_SET_EXPLICITLY
- ? connectionListenerInboxCapacity
- : io.questdb.client.cutlass.qwp.client.sf.cursor.SenderConnectionDispatcher.DEFAULT_CAPACITY;
- List wsEndpoints =
- new ArrayList<>(hosts.size());
- for (int i = 0, n = hosts.size(); i < n; i++) {
- wsEndpoints.add(new QwpWebSocketSender.Endpoint(hosts.getQuick(i), ports.getQuick(i)));
- }
- // The recovery seed inside connect() is the authority on whether a recovered
- // slot can be replayed: it rebuilds the dictionary from its intact prefix and
- // then from the surviving frames' own delta sections, and throws
- // UnreplayableSlotException only once neither source holds the missing ids.
- // Quarantining on anything weaker would set aside slots that recovery can
- // still rescue, so build() waits for that verdict rather than pre-judging it.
- while (connected == null) {
- try {
- connected = QwpWebSocketSender.connectWithCredentialSupplier(
- wsEndpoints,
- wsTlsConfig,
- actualAutoFlushRows,
- actualAutoFlushBytes,
- actualAutoFlushIntervalNanos,
- wsAuthHeader,
- requestDurableAck,
- cursorEngine,
- actualCloseFlushTimeoutMillis,
- actualReconnectMaxDurationMillis,
- actualReconnectInitialBackoffMillis,
- actualReconnectMaxBackoffMillis,
- actualInitialConnectMode,
- errorHandler,
- actualErrorInboxCapacity,
- actualDurableAckKeepaliveIntervalMillis,
- authTimeoutMillis,
- connectTimeoutMillis == PARAMETER_NOT_SET_EXPLICITLY ? 0 : connectTimeoutMillis,
- connectionListener,
- actualConnectionListenerInboxCapacity,
- actualMaxFrameRejections,
- actualPoisonMinEscalationWindowMillis,
- actualCatchUpCapGapMinEscalationWindowMillis
- );
- } catch (UnreplayableSlotException e) {
- // The one failure build() recovers from. The slot's frames reference ids
- // that nothing still holds, so they can never go on the wire -- but that is
- // no reason to take the producer down with them. Before this, the throw
- // escaped build() and, because senderId is stable and a not-fully-drained
- // slot is retained on close, every retry re-recovered the same slot and
- // threw again: the application could not construct a Sender at all, so it
- // could not even BUFFER new rows. An already-lost batch became an unbounded
- // outage of everything after it.
- //
- // Set the slot aside instead, keep its bytes for forensics and resend, and
- // start the producer on a clean one. Once only: a second such failure would
- // mean the FRESH slot is unreplayable, which cannot happen, so let it out
- // rather than loop.
- if (quarantined || slotPath == null) {
- try {
- // close(false): we still hold the logical slot lock.
- cursorEngine.close(false);
- } catch (Throwable ignored) {
- // best-effort
- }
- throw e;
- }
- quarantined = true;
- cursorEngine = quarantineTornSlot(
- cursorEngine, e, sfDir, senderId, slotPath, actualSfMaxSegmentBytes,
- actualSfMaxTotalBytes, actualSfAppendDeadlineNanos,
- actualSfSyncIntervalNanos, errorHandler);
- } catch (Throwable t) {
- // connect() failed before ownership of cursorEngine
- // transferred — close it ourselves. close(false)
- // because logicalSlotLock is still held here: a fresh
- // slot is fully drained, so the default close would
- // unlink the very lock file this scope holds.
- try {
- cursorEngine.close(false);
- } catch (Throwable ignored) {
- // best-effort
- }
- throw t;
- }
- }
- }
- // connect() succeeded — `connected` now owns cursorEngine
- // via setCursorEngine(engine, true). From here on, ANY
- // failure must close `connected` (which closes the engine
- // through ownsCursorEngine), not cursorEngine directly:
- // closing the engine alone would leak the I/O thread,
- // dispatcher daemon, drainer pool, microbatch buffers and
- // WebSocketClient inside the abandoned `connected`.
- connected.setTransactional(transactional);
+ // token_provider: share the process-wide provider for this configuration. The lease is the
+ // sender's from here on: released when the sender closes, or right here if build() fails.
+ TokenProviderRegistry.Lease lease = TokenProviderRegistry.global().acquire(tokenProviderSpec);
try {
- // Install the drainer listener BEFORE startOrphanDrainers
- // below: drainers must see the listener at submit time so
- // no early drainer event is lost to a late installation.
- if (drainerListener != null) {
- connected.setDrainerListener(drainerListener);
- }
- // Once the foreground sender is up, dispatch drainers
- // for any sibling orphan slots. Scan AFTER we acquire
- // our own slot lock so we never accidentally try to
- // adopt our own data; the OrphanScanner.scan filter
- // also excludes our sender_id.
- if (drainOrphans && sfDir != null) {
- io.questdb.client.std.ObjList orphans =
- io.questdb.client.cutlass.qwp.client.sf.cursor.OrphanScanner
- .scan(sfDir, senderId, orphanDrainBase, orphanDrainSlotCount);
- if (orphans.size() > 0) {
- org.slf4j.LoggerFactory.getLogger(LineSenderBuilder.class)
- .info("dispatching drainers for {} orphan slot(s) under {} "
- + "(max_background_drainers={})",
- orphans.size(), sfDir, maxBackgroundDrainers);
- connected.startOrphanDrainers(
- orphans,
- maxBackgroundDrainers,
- actualSfMaxSegmentBytes,
- actualSfMaxTotalBytes,
- actualSfSyncIntervalNanos);
- }
- }
- return connected;
- } catch (Throwable t) {
- try {
- connected.close();
- } catch (Throwable ignored) {
- // best-effort
+ QwpWebSocketSender sender = buildWebSocket(lease.provider());
+ sender.setCredentialLease(lease);
+ lease = null;
+ return sender;
+ } finally {
+ if (lease != null) {
+ lease.close();
}
- throw t;
}
}
@@ -2282,6 +2002,9 @@ public LineSenderBuilder httpTimeoutMillis(int httpTimeoutMillis) {
* @return this instance for method chaining
*/
public LineSenderBuilder httpToken(String token) {
+ if (this.tokenProviderSpec != null) {
+ throw new LineSenderException("token cannot be combined with token_provider");
+ }
if (this.username != null) {
throw new LineSenderException("authentication username was already configured ")
.put("[username=").put(this.username).put("]");
@@ -2341,6 +2064,10 @@ public LineSenderBuilder httpToken(String token) {
* @return this instance for method chaining
*/
public LineSenderBuilder httpTokenProvider(HttpTokenProvider httpTokenProvider) {
+ if (this.tokenProviderSpec != null) {
+ throw new LineSenderException("an application-supplied token provider cannot be combined with "
+ + "token_provider in the configuration");
+ }
if (this.username != null) {
throw new LineSenderException("authentication username was already configured ")
.put("[username=").put(this.username).put("]");
@@ -2370,6 +2097,9 @@ public LineSenderBuilder httpTokenProvider(HttpTokenProvider httpTokenProvider)
* @see #httpToken(String)
*/
public LineSenderBuilder httpUsernamePassword(String username, String password) {
+ if (this.tokenProviderSpec != null) {
+ throw new LineSenderException("username/password cannot be combined with token_provider");
+ }
if (this.username != null) {
throw new LineSenderException("authentication username was already configured ")
.put("[username=").put(this.username).put("]");
@@ -3424,7 +3154,371 @@ private void appendAddress(String host, int port) {
ports.add(port);
}
- private Supplier buildWebSocketAuthHeader() {
+ private QwpWebSocketSender buildWebSocket(HttpTokenProvider tokenProvider) {
+ if (hosts.size() < 1) {
+ throw new LineSenderException("WebSocket transport requires at least one host:port pair");
+ }
+ if (tokenProvider != null && !tlsEnabled && CLEARTEXT_PROVIDER_WARNED.compareAndSet(false, true)) {
+ // design/qwp-token-provider-spec.md, section 9: a configured token_provider is rejected on ws::,
+ // but an application-supplied provider is the application's call - warn once instead.
+ LOG.warn("a token provider is used over ws:: (no TLS): bearer tokens cross the network in "
+ + "cleartext; use wss:: in production");
+ }
+
+ int actualAutoFlushRows = autoFlushRows == PARAMETER_NOT_SET_EXPLICITLY ? DEFAULT_WS_AUTO_FLUSH_ROWS : autoFlushRows;
+ int actualAutoFlushBytes = autoFlushBytes == PARAMETER_NOT_SET_EXPLICITLY ? DEFAULT_WS_AUTO_FLUSH_BYTES : autoFlushBytes;
+ long actualAutoFlushIntervalNanos = autoFlushIntervalMillis == PARAMETER_NOT_SET_EXPLICITLY
+ ? DEFAULT_WS_AUTO_FLUSH_INTERVAL_NANOS
+ : TimeUnit.MILLISECONDS.toNanos(autoFlushIntervalMillis);
+
+ Supplier wsAuthHeader = buildWebSocketAuthHeader(tokenProvider);
+
+ ClientTlsConfiguration wsTlsConfig = null;
+ if (tlsEnabled) {
+ assert trustStorePassword == null || trustStorePath != null;
+ wsTlsConfig = new ClientTlsConfiguration(
+ trustStorePath,
+ trustStorePassword,
+ tlsValidationMode == TlsValidationMode.DEFAULT
+ ? ClientTlsConfiguration.TLS_VALIDATION_MODE_FULL
+ : ClientTlsConfiguration.TLS_VALIDATION_MODE_NONE
+ );
+ }
+
+ // Setting sfDir enables store-and-forward (mmap'd, recoverable
+ // across sender restarts); omitting it gives memory-only mode
+ // (same lock-free architecture, no disk involvement).
+ // Durability-combination validation lives in validateParameters
+ // so build() and no-connect validation apply the same rules.
+ long actualSfMaxSegmentBytes = sfMaxSegmentBytes == PARAMETER_NOT_SET_EXPLICITLY
+ ? DEFAULT_SEGMENT_BYTES
+ : sfMaxSegmentBytes;
+ // Default cap depends on backing: RAM (memory mode) is tight
+ // by default; disk (SF mode) is cheap so the default is
+ // generous enough that normal traffic never hits it.
+ long defaultMaxTotal = sfDir == null
+ ? DEFAULT_MAX_BYTES_MEMORY
+ : DEFAULT_MAX_BYTES_SF;
+ long actualSfMaxTotalBytes = sfMaxTotalBytes == PARAMETER_NOT_SET_EXPLICITLY
+ ? Math.max(defaultMaxTotal, actualSfMaxSegmentBytes * 2)
+ : sfMaxTotalBytes;
+ long actualCloseFlushTimeoutMillis = closeFlushTimeoutMillis == CLOSE_FLUSH_TIMEOUT_NOT_SET
+ ? DEFAULT_CLOSE_FLUSH_TIMEOUT_MILLIS
+ : closeFlushTimeoutMillis;
+ long actualReconnectMaxDurationMillis =
+ reconnectMaxDurationMillis == PARAMETER_NOT_SET_EXPLICITLY
+ ? CursorWebSocketSendLoop.DEFAULT_RECONNECT_MAX_DURATION_MILLIS
+ : reconnectMaxDurationMillis;
+ long actualReconnectInitialBackoffMillis =
+ reconnectInitialBackoffMillis == PARAMETER_NOT_SET_EXPLICITLY
+ ? CursorWebSocketSendLoop.DEFAULT_RECONNECT_INITIAL_BACKOFF_MILLIS
+ : reconnectInitialBackoffMillis;
+ long actualReconnectMaxBackoffMillis =
+ reconnectMaxBackoffMillis == PARAMETER_NOT_SET_EXPLICITLY
+ ? CursorWebSocketSendLoop.DEFAULT_RECONNECT_MAX_BACKOFF_MILLIS
+ : reconnectMaxBackoffMillis;
+ // Resolve the initial-connect mode. An explicit user choice
+ // (via initialConnectMode/initialConnectRetry, or the
+ // initial_connect_retry conf key) wins unconditionally --
+ // including initial_connect_retry=off paired with a tuned
+ // reconnect budget. When the user left it unset and tuned
+ // any reconnect_* knob, promote to SYNC so the budget they
+ // wrote actually applies to the first connect: the knob
+ // name reads as a generic retry budget but the underlying
+ // path only governs reconnects from an established
+ // connection, and silently ignoring the budget on the
+ // initial connect is the canonical footgun this implicit
+ // upgrade removes.
+ InitialConnectMode actualInitialConnectMode;
+ if (initialConnectMode != null) {
+ actualInitialConnectMode = initialConnectMode;
+ } else if (reconnectMaxDurationMillis != PARAMETER_NOT_SET_EXPLICITLY
+ || reconnectInitialBackoffMillis != PARAMETER_NOT_SET_EXPLICITLY
+ || reconnectMaxBackoffMillis != PARAMETER_NOT_SET_EXPLICITLY) {
+ actualInitialConnectMode = InitialConnectMode.SYNC;
+ } else {
+ actualInitialConnectMode = InitialConnectMode.OFF;
+ }
+ long actualDurableAckKeepaliveIntervalMillis =
+ durableAckKeepaliveIntervalMillis == DURABLE_ACK_KEEPALIVE_NOT_SET
+ ? CursorWebSocketSendLoop.DEFAULT_DURABLE_ACK_KEEPALIVE_INTERVAL_MILLIS
+ : durableAckKeepaliveIntervalMillis;
+ int actualMaxFrameRejections = maxFrameRejections != PARAMETER_NOT_SET_EXPLICITLY
+ ? maxFrameRejections
+ : CursorWebSocketSendLoop.DEFAULT_MAX_HEAD_FRAME_REJECTIONS;
+ long actualPoisonMinEscalationWindowMillis = poisonMinEscalationWindowMillis != PARAMETER_NOT_SET_EXPLICITLY
+ ? poisonMinEscalationWindowMillis
+ : CursorWebSocketSendLoop.DEFAULT_POISON_MIN_ESCALATION_WINDOW_MILLIS;
+ long actualCatchUpCapGapMinEscalationWindowMillis =
+ catchUpCapGapMinEscalationWindowMillis != PARAMETER_NOT_SET_EXPLICITLY
+ ? catchUpCapGapMinEscalationWindowMillis
+ : CursorWebSocketSendLoop.DEFAULT_CATCHUP_CAP_GAP_MIN_ESCALATION_WINDOW_MILLIS;
+
+ // sfDir is the parent (group root); the actual slot lives
+ // under sfDir/senderId. This is what the engine sees — the
+ // slot lock and segment files all live one level deeper than
+ // the user-supplied path. Memory mode skips this composition
+ // (slotPath stays null).
+ //
+ // The slot ctor inside CursorSendEngine creates the slot
+ // directory itself, but Files.mkdir is non-recursive — so we
+ // must ensure the parent group root exists first.
+ String slotPath;
+ if (sfDir == null) {
+ slotPath = null;
+ } else {
+ if (!Files.exists(sfDir)) {
+ int rc = Files.mkdir(sfDir, Files.DIR_MODE_DEFAULT);
+ // mkdir is non-zero on failure, but "already exists"
+ // is one such failure. Multiple SF senders sharing one
+ // sf_dir can be built concurrently (the pool calls
+ // build() outside its lock), so two threads can both
+ // pass the exists() check and race into mkdir; the
+ // loser gets EEXIST. Treat a benign creation race --
+ // the dir now exists -- as success and only fail when
+ // the directory is genuinely absent afterwards.
+ if (rc != 0 && !Files.exists(sfDir)) {
+ throw new LineSenderException(
+ "could not create sf_dir: " + sfDir + " rc=" + rc);
+ }
+ }
+ if (sfDurability == SfDurability.PERIODIC
+ && Files.fsyncParentDir(sfDir) != 0) {
+ throw new LineSenderException(
+ "could not sync parent directory for sf_dir: " + sfDir);
+ }
+ slotPath = sfDir + "/" + senderId;
+ }
+ long actualSfAppendDeadlineNanos =
+ sfAppendDeadlineMillis == PARAMETER_NOT_SET_EXPLICITLY
+ ? CursorSendEngine.DEFAULT_APPEND_DEADLINE_NANOS
+ : sfAppendDeadlineMillis * 1_000_000L;
+ long actualSfSyncIntervalNanos = sfDurability == SfDurability.PERIODIC
+ ? (sfSyncIntervalMillis == PARAMETER_NOT_SET_EXPLICITLY
+ ? DEFAULT_SF_SYNC_INTERVAL_MILLIS : sfSyncIntervalMillis) * 1_000_000L
+ : 0L;
+ QwpWebSocketSender connected = null;
+ // The parent-anchored logical lock is stable across a slot rename. Keep it
+ // from before the directory-local lock is acquired until connect() has either
+ // adopted that engine or quarantine has closed, renamed and recreated it.
+ // This closes the inode-swap window in which an already-queued orphan drainer
+ // could otherwise acquire the renamed directory's old .lock and later operate
+ // on the fresh slot through the original pathname.
+ try (SlotLock logicalSlotLock = slotPath == null
+ ? null
+ : SlotLock.acquireLogical(slotPath)) {
+ // The constructor's own recovery seed can also fail terminally, and
+ // not only as UnreplayableSlotException: when SegmentRing.openExisting
+ // had to skip an unreadable segment it throws SfRecoveryException (it
+ // constructs UnreplayableSlotException nowhere), and where it cannot
+ // even prove the chain's identity -- no manifest -- it quarantines the
+ // corrupt files and returns an EMPTY recovery rather than refusing.
+ // Either way the frame range cannot be shown already-acked, so recovery
+ // sets the slot aside rather than risk seeding the ack cursor past
+ // frames that were never delivered. All three types below are load
+ // bearing; narrowing this catch to UnreplayableSlotException would
+ // restore the permanent build() brick for the segment-skip case. That verdict gets
+ // the exact same quarantine-and-continue treatment as the connect()-time
+ // verdict below -- constructing cursorEngine is not inside the loop below,
+ // so a throw here would otherwise escape build() entirely, uncaught.
+ // quarantineTornSlot(null, ...) renames the WHOLE slot directory aside
+ // (not just the unreadable segment file) before building the replacement
+ // at the original slotPath, so the replacement starts on a genuinely empty
+ // directory with nothing left to skip -- it cannot throw the same way
+ // twice, which is what makes looping unnecessary here.
+ boolean quarantined = false;
+ CursorSendEngine cursorEngine;
+ try {
+ try {
+ cursorEngine = new CursorSendEngine(
+ slotPath, actualSfMaxSegmentBytes,
+ actualSfMaxTotalBytes, actualSfAppendDeadlineNanos,
+ actualSfSyncIntervalNanos);
+ } catch (SfSanitizedResidueException first) {
+ // NOT terminal, and it must be intercepted ahead of its
+ // SfRecoveryException parent below. Recovery durably zeroed
+ // proven-dead sealed residue BEFORE failing closed, so the
+ // chain on disk is already healed: quarantining here would
+ // set aside a slot whose backlog replays perfectly. Retry
+ // once over the healed chain; a repeat is genuine and takes
+ // the terminal arm.
+ LOG.info("sf slot {}: sealed residue sanitized during recovery ({}); "
+ + "retrying over the healed chain",
+ slotPath, first.getMessage());
+ cursorEngine = new CursorSendEngine(
+ slotPath, actualSfMaxSegmentBytes,
+ actualSfMaxTotalBytes, actualSfAppendDeadlineNanos,
+ actualSfSyncIntervalNanos);
+ }
+ } catch (UnreplayableSlotException | SfRecoveryException
+ | MmapSegmentCorruptionException e) {
+ // The terminal recovery verdicts, and the only ones build()
+ // sets a slot aside for. UnreplayableSlotException says the
+ // symbol dictionary cannot be rebuilt from any source;
+ // SfRecoveryException and MmapSegmentCorruptionException say
+ // the durable chain itself is proven corrupt or incomplete.
+ // None of the three clears on a retry, and senderId is stable
+ // with a not-fully-drained slot retained on close -- so
+ // without this arm every restart re-recovers the same slot and
+ // throws again, and the application cannot construct a Sender
+ // at all, not even to BUFFER new rows.
+ //
+ // Deliberately NOT catching plain MmapSegmentException or
+ // SfOperationalException: those are operational (EMFILE,
+ // ENOMEM, an unreadable-but-possibly-intact file). Aborting
+ // startup on them is correct; quarantining on them would
+ // convert a transient into the permanent loss of a healthy
+ // slot's durable frames.
+ if (slotPath == null) {
+ throw e;
+ }
+ quarantined = true;
+ cursorEngine = quarantineTornSlot(
+ null, e, sfDir, senderId, slotPath, actualSfMaxSegmentBytes,
+ actualSfMaxTotalBytes, actualSfAppendDeadlineNanos,
+ actualSfSyncIntervalNanos, errorHandler);
+ }
+ int actualErrorInboxCapacity = errorInboxCapacity != PARAMETER_NOT_SET_EXPLICITLY
+ ? errorInboxCapacity
+ : io.questdb.client.cutlass.qwp.client.sf.cursor.SenderErrorDispatcher.DEFAULT_CAPACITY;
+ int actualConnectionListenerInboxCapacity = connectionListenerInboxCapacity != PARAMETER_NOT_SET_EXPLICITLY
+ ? connectionListenerInboxCapacity
+ : io.questdb.client.cutlass.qwp.client.sf.cursor.SenderConnectionDispatcher.DEFAULT_CAPACITY;
+ List wsEndpoints =
+ new ArrayList<>(hosts.size());
+ for (int i = 0, n = hosts.size(); i < n; i++) {
+ wsEndpoints.add(new QwpWebSocketSender.Endpoint(hosts.getQuick(i), ports.getQuick(i)));
+ }
+ // The recovery seed inside connect() is the authority on whether a recovered
+ // slot can be replayed: it rebuilds the dictionary from its intact prefix and
+ // then from the surviving frames' own delta sections, and throws
+ // UnreplayableSlotException only once neither source holds the missing ids.
+ // Quarantining on anything weaker would set aside slots that recovery can
+ // still rescue, so build() waits for that verdict rather than pre-judging it.
+ while (connected == null) {
+ try {
+ connected = QwpWebSocketSender.connectWithCredentialSupplier(
+ wsEndpoints,
+ wsTlsConfig,
+ actualAutoFlushRows,
+ actualAutoFlushBytes,
+ actualAutoFlushIntervalNanos,
+ wsAuthHeader,
+ requestDurableAck,
+ cursorEngine,
+ actualCloseFlushTimeoutMillis,
+ actualReconnectMaxDurationMillis,
+ actualReconnectInitialBackoffMillis,
+ actualReconnectMaxBackoffMillis,
+ actualInitialConnectMode,
+ errorHandler,
+ actualErrorInboxCapacity,
+ actualDurableAckKeepaliveIntervalMillis,
+ authTimeoutMillis,
+ connectTimeoutMillis == PARAMETER_NOT_SET_EXPLICITLY ? 0 : connectTimeoutMillis,
+ connectionListener,
+ actualConnectionListenerInboxCapacity,
+ actualMaxFrameRejections,
+ actualPoisonMinEscalationWindowMillis,
+ actualCatchUpCapGapMinEscalationWindowMillis
+ );
+ } catch (UnreplayableSlotException e) {
+ // The one failure build() recovers from. The slot's frames reference ids
+ // that nothing still holds, so they can never go on the wire -- but that is
+ // no reason to take the producer down with them. Before this, the throw
+ // escaped build() and, because senderId is stable and a not-fully-drained
+ // slot is retained on close, every retry re-recovered the same slot and
+ // threw again: the application could not construct a Sender at all, so it
+ // could not even BUFFER new rows. An already-lost batch became an unbounded
+ // outage of everything after it.
+ //
+ // Set the slot aside instead, keep its bytes for forensics and resend, and
+ // start the producer on a clean one. Once only: a second such failure would
+ // mean the FRESH slot is unreplayable, which cannot happen, so let it out
+ // rather than loop.
+ if (quarantined || slotPath == null) {
+ try {
+ // close(false): we still hold the logical slot lock.
+ cursorEngine.close(false);
+ } catch (Throwable ignored) {
+ // best-effort
+ }
+ throw e;
+ }
+ quarantined = true;
+ cursorEngine = quarantineTornSlot(
+ cursorEngine, e, sfDir, senderId, slotPath, actualSfMaxSegmentBytes,
+ actualSfMaxTotalBytes, actualSfAppendDeadlineNanos,
+ actualSfSyncIntervalNanos, errorHandler);
+ } catch (Throwable t) {
+ // connect() failed before ownership of cursorEngine
+ // transferred — close it ourselves. close(false)
+ // because logicalSlotLock is still held here: a fresh
+ // slot is fully drained, so the default close would
+ // unlink the very lock file this scope holds.
+ try {
+ cursorEngine.close(false);
+ } catch (Throwable ignored) {
+ // best-effort
+ }
+ throw t;
+ }
+ }
+ }
+ // connect() succeeded — `connected` now owns cursorEngine
+ // via setCursorEngine(engine, true). From here on, ANY
+ // failure must close `connected` (which closes the engine
+ // through ownsCursorEngine), not cursorEngine directly:
+ // closing the engine alone would leak the I/O thread,
+ // dispatcher daemon, drainer pool, microbatch buffers and
+ // WebSocketClient inside the abandoned `connected`.
+ connected.setTransactional(transactional);
+ if (authFailureMaxDurationMillis > 0) {
+ // The I/O loop may already be running (async initial connect), but it tracks the outage clock
+ // from the first authentication-class failure regardless; this only arms the deadline.
+ connected.setAuthFailureMaxDurationMillis(authFailureMaxDurationMillis);
+ }
+ try {
+ // Install the drainer listener BEFORE startOrphanDrainers
+ // below: drainers must see the listener at submit time so
+ // no early drainer event is lost to a late installation.
+ if (drainerListener != null) {
+ connected.setDrainerListener(drainerListener);
+ }
+ // Once the foreground sender is up, dispatch drainers
+ // for any sibling orphan slots. Scan AFTER we acquire
+ // our own slot lock so we never accidentally try to
+ // adopt our own data; the OrphanScanner.scan filter
+ // also excludes our sender_id.
+ if (drainOrphans && sfDir != null) {
+ io.questdb.client.std.ObjList orphans =
+ io.questdb.client.cutlass.qwp.client.sf.cursor.OrphanScanner
+ .scan(sfDir, senderId, orphanDrainBase, orphanDrainSlotCount);
+ if (orphans.size() > 0) {
+ org.slf4j.LoggerFactory.getLogger(LineSenderBuilder.class)
+ .info("dispatching drainers for {} orphan slot(s) under {} "
+ + "(max_background_drainers={})",
+ orphans.size(), sfDir, maxBackgroundDrainers);
+ connected.startOrphanDrainers(
+ orphans,
+ maxBackgroundDrainers,
+ actualSfMaxSegmentBytes,
+ actualSfMaxTotalBytes,
+ actualSfSyncIntervalNanos);
+ }
+ }
+ return connected;
+ } catch (Throwable t) {
+ try {
+ connected.close();
+ } catch (Throwable ignored) {
+ // best-effort
+ }
+ throw t;
+ }
+ }
+
+ private Supplier buildWebSocketAuthHeader(HttpTokenProvider provider) {
// A constant credential goes through fixedAuthHeader, not a bare lambda: the tag is what lets
// the store-and-forward drainer tell a permanently-wrong password from a rotating token that a
// fresh pull can repair, and so decide whether a 401 may quarantine an orphan slot for good.
@@ -3437,21 +3531,13 @@ private Supplier buildWebSocketAuthHeader() {
String header = "Bearer " + httpToken;
return QwpWebSocketSender.fixedAuthHeader(header);
}
- if (httpTokenProvider != null) {
- // pull a fresh token at each (re)handshake so a long-lived WebSocket follows token
- // refreshes; validateToken rejects a null/empty/blank return, or a token carrying a
- // control or non-ASCII char (both forbidden by the HttpTokenProvider contract), rather
- // than send a malformed or CR/LF-injected "Bearer " header
- final HttpTokenProvider provider = httpTokenProvider;
- return () -> {
- // snapshot before validating: the concatenation below re-reads the sequence, and a
- // provider is free to reuse a mutable buffer, so validating the live sequence checks
- // bytes the header need not carry. See HttpTokenProvider.validateToken.
- CharSequence pulled = provider.getToken();
- CharSequence token = pulled == null ? null : pulled.toString();
- HttpTokenProvider.validateToken(token);
- return "Bearer " + token;
- };
+ if (provider != null) {
+ // Pull a fresh token at each (re)handshake so a long-lived WebSocket follows token refreshes.
+ // The supplier snapshots and validates every pull (validateToken rejects a null/empty/blank
+ // return, or a token carrying a control or non-ASCII char, rather than send a malformed or
+ // CR/LF-injected "Bearer " header), and it carries the provider's onTokenRejected back-channel
+ // that the connect walk uses for its one retry after a 401.
+ return QwpWebSocketSender.tokenProviderAuthHeader(provider);
}
return null;
}
@@ -3973,6 +4059,14 @@ private LineSenderBuilder fromConfig(CharSequence configurationString) {
// genuine value-parse error names the offending key.
String reservedKey = Chars.toString(sink);
pos = getValue(configurationString, pos, sink, reservedKey);
+ } else if (Chars.equals("auth_failure_max_duration_millis", sink)) {
+ throw new LineSenderException("auth_failure_max_duration_millis is only supported for WebSocket transport");
+ } else if (Chars.equals("token_provider", sink)
+ || Chars.equals("azure_resource", sink)
+ || Chars.equals("azure_client_id", sink)) {
+ // Dynamic bearer credentials are defined for QWP over wss:: only (decision D9).
+ throw new LineSenderException(Chars.toString(sink)
+ + " is only supported with the wss:: schema (QWP over WebSocket)");
} else {
// sf-client.md §4.6: parser must reject unknown keys.
// Forward-compat is via the spec, not silent ignore — silent
@@ -4020,6 +4114,16 @@ private LineSenderBuilder fromConfigWebSocket(CharSequence configurationString)
ConfigString cs = ConfigString.parse(configurationString);
ConfigView view = new ConfigView(cs);
validateWsConfig(view, tlsEnabled);
+ // Validates token_provider and its keys (wss:: only, exclusive with static credentials, a
+ // supported provider) without fetching anything; build() acquires the provider.
+ TokenProviderSpec spec = TokenProviderSpec.parse(view, tlsEnabled);
+ if (spec != null) {
+ if (httpTokenProvider != null) {
+ throw new LineSenderException("token_provider cannot be combined with an "
+ + "application-supplied token provider");
+ }
+ tokenProviderSpec = spec;
+ }
view.getHostPorts("addr", DEFAULT_WEBSOCKET_PORT, this::appendAddress);
@@ -4055,6 +4159,9 @@ private LineSenderBuilder fromConfigWebSocket(CharSequence configurationString)
// int, and an over-int value must reject, not wrap.
connectTimeoutMillis(view.getInt("connect_timeout", 0));
}
+ if (view.has("auth_failure_max_duration_millis")) {
+ authFailureMaxDurationMillis(view.getLong("auth_failure_max_duration_millis", 0));
+ }
s = view.getStr("auto_flush_rows");
if (s != null) {
@@ -4335,6 +4442,12 @@ public java.util.Map wsConfigSnapshotForTest() {
m.put("tls_verify", tlsValidationMode == null ? null : tlsValidationMode.name());
m.put("tls_roots", trustStorePath);
m.put("tls_roots_password", trustStorePassword == null ? null : new String(trustStorePassword));
+ m.put("auth_failure_max_duration_millis", authFailureMaxDurationMillis);
+ m.put("token_provider", tokenProviderSpec == null ? null : tokenProviderSpec.name());
+ m.put("azure_resource", tokenProviderSpec == null ? null
+ : tokenProviderSpec.params().get(TokenProviderSpec.KEY_AZURE_RESOURCE));
+ m.put("azure_client_id", tokenProviderSpec == null ? null
+ : tokenProviderSpec.params().get(TokenProviderSpec.KEY_AZURE_CLIENT_ID));
return m;
}
diff --git a/core/src/main/java/io/questdb/client/cutlass/auth/CredentialRedaction.java b/core/src/main/java/io/questdb/client/cutlass/auth/CredentialRedaction.java
new file mode 100644
index 000000000..a84be9fe4
--- /dev/null
+++ b/core/src/main/java/io/questdb/client/cutlass/auth/CredentialRedaction.java
@@ -0,0 +1,145 @@
+/*+*****************************************************************************
+ * ___ _ ____ ____
+ * / _ \ _ _ ___ ___| |_| _ \| __ )
+ * | | | | | | |/ _ \/ __| __| | | | _ \
+ * | |_| | |_| | __/\__ \ |_| |_| | |_) |
+ * \__\_\\__,_|\___||___/\__|____/|____/
+ *
+ * Copyright (c) 2014-2019 Appsicle
+ * Copyright (c) 2019-2026 QuestDB
+ *
+ * Licensed under the Apache License, Version 2.0 (the "License");
+ * you may not use this file except in compliance with the License.
+ * You may obtain a copy of the License at
+ *
+ * http://www.apache.org/licenses/LICENSE-2.0
+ *
+ * Unless required by applicable law or agreed to in writing, software
+ * distributed under the License is distributed on an "AS IS" BASIS,
+ * WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied.
+ * See the License for the specific language governing permissions and
+ * limitations under the License.
+ *
+ ******************************************************************************/
+
+package io.questdb.client.cutlass.auth;
+
+import io.questdb.client.std.str.DisplaySafe;
+
+import java.nio.charset.StandardCharsets;
+import java.security.MessageDigest;
+import java.security.NoSuchAlgorithmException;
+
+/**
+ * Redaction rules for bearer credentials and for error text that crosses the client boundary, as required by
+ * the dynamic-credential specification (design/qwp-token-provider-spec.md, section 9).
+ *
+ *
A token is never rendered. {@link #describeToken(CharSequence)} gives its length and at most the first
+ * 8 hex digits of its SHA-256 - enough to correlate a rotation across log lines, useless to replay.
+ *
Error text that came from a library or an identity endpoint goes through
+ * {@link #sanitizeErrorText(CharSequence)} before it reaches a client error or a log line: control,
+ * bidirectional and other non-displayable characters are removed, and the result is capped at
+ * {@value #MAX_ERROR_TEXT_LENGTH} characters.
+ *
+ */
+public final class CredentialRedaction {
+ /**
+ * Upper bound, in characters, on error text from a library or endpoint once sanitized.
+ */
+ public static final int MAX_ERROR_TEXT_LENGTH = 256;
+ private static final int FINGERPRINT_HEX_DIGITS = 8;
+ private static final char[] HEX = "0123456789abcdef".toCharArray();
+ private static final String TRUNCATION_MARKER = "...";
+
+ private CredentialRedaction() {
+ }
+
+ /**
+ * Renders a token as {@code }: its length plus the first 8 hex digits of
+ * its SHA-256. Never the token itself.
+ *
+ * @param token the token, may be null
+ * @return a description that is safe to log
+ */
+ public static String describeToken(CharSequence token) {
+ if (token == null) {
+ return "";
+ }
+ return "';
+ }
+
+ /**
+ * The first 8 hex digits of the SHA-256 of {@code token}, encoded as UTF-8.
+ *
+ * @param token the token, never null
+ * @return 8 lowercase hex digits
+ */
+ public static String fingerprint(CharSequence token) {
+ final byte[] digest;
+ try {
+ digest = MessageDigest.getInstance("SHA-256").digest(token.toString().getBytes(StandardCharsets.UTF_8));
+ } catch (NoSuchAlgorithmException e) {
+ // every Java platform is required to provide SHA-256
+ throw new IllegalStateException("SHA-256 is not available", e);
+ }
+ final char[] out = new char[FINGERPRINT_HEX_DIGITS];
+ for (int i = 0; i < FINGERPRINT_HEX_DIGITS / 2; i++) {
+ out[2 * i] = HEX[(digest[i] >> 4) & 0xf];
+ out[2 * i + 1] = HEX[digest[i] & 0xf];
+ }
+ return new String(out);
+ }
+
+ /**
+ * Prepares untrusted error text - from a token source, an identity library or an endpoint - for a client
+ * error or a log line: removes every code point {@link DisplaySafe} rejects (controls including CR/LF,
+ * bidirectional overrides, zero-width and other format characters, lone surrogates) and caps the result at
+ * {@value #MAX_ERROR_TEXT_LENGTH} characters, ending in {@code ...} when it had to cut.
+ *
+ * This does not, and cannot, remove a secret the text already carries. Keeping tokens and response bodies
+ * out of error text is the job of whoever produces it.
+ *
+ * @param text untrusted text, may be null
+ * @return the sanitized text, or {@code null} when {@code text} is null
+ */
+ public static String sanitizeErrorText(CharSequence text) {
+ if (text == null) {
+ return null;
+ }
+ final StringBuilder sb = new StringBuilder(Math.min(text.length(), MAX_ERROR_TEXT_LENGTH));
+ final int limit = MAX_ERROR_TEXT_LENGTH - TRUNCATION_MARKER.length();
+ for (int i = 0, n = text.length(); i < n; ) {
+ final int cp = Character.codePointAt(text, i);
+ final int count = Character.charCount(cp);
+ if (DisplaySafe.isDisplaySafe(cp)) {
+ if (sb.length() + count > limit) {
+ // Only cut when something displayable is left to drop: text that fits exactly is kept whole.
+ if (hasDisplayableFrom(text, i, MAX_ERROR_TEXT_LENGTH - sb.length())) {
+ sb.append(TRUNCATION_MARKER);
+ return sb.toString();
+ }
+ }
+ sb.appendCodePoint(cp);
+ }
+ i += count;
+ }
+ return sb.toString();
+ }
+
+ // True when the displayable remainder of text, starting at from, does not fit in budget characters.
+ private static boolean hasDisplayableFrom(CharSequence text, int from, int budget) {
+ int needed = 0;
+ for (int i = from, n = text.length(); i < n; ) {
+ final int cp = Character.codePointAt(text, i);
+ final int count = Character.charCount(cp);
+ if (DisplaySafe.isDisplaySafe(cp)) {
+ needed += count;
+ if (needed > budget) {
+ return true;
+ }
+ }
+ i += count;
+ }
+ return false;
+ }
+}
diff --git a/core/src/main/java/io/questdb/client/cutlass/auth/ExpiringToken.java b/core/src/main/java/io/questdb/client/cutlass/auth/ExpiringToken.java
new file mode 100644
index 000000000..95713352b
--- /dev/null
+++ b/core/src/main/java/io/questdb/client/cutlass/auth/ExpiringToken.java
@@ -0,0 +1,148 @@
+/*+*****************************************************************************
+ * ___ _ ____ ____
+ * / _ \ _ _ ___ ___| |_| _ \| __ )
+ * | | | | | | |/ _ \/ __| __| | | | _ \
+ * | |_| | |_| | __/\__ \ |_| |_| | |_) |
+ * \__\_\\__,_|\___||___/\__|____/|____/
+ *
+ * Copyright (c) 2014-2019 Appsicle
+ * Copyright (c) 2019-2026 QuestDB
+ *
+ * Licensed under the Apache License, Version 2.0 (the "License");
+ * you may not use this file except in compliance with the License.
+ * You may obtain a copy of the License at
+ *
+ * http://www.apache.org/licenses/LICENSE-2.0
+ *
+ * Unless required by applicable law or agreed to in writing, software
+ * distributed under the License is distributed on an "AS IS" BASIS,
+ * WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied.
+ * See the License for the specific language governing permissions and
+ * limitations under the License.
+ *
+ ******************************************************************************/
+
+package io.questdb.client.cutlass.auth;
+
+import java.time.Instant;
+
+/**
+ * A bearer token together with the instant it expires, as returned by a {@link TokenSource}. This is the
+ * {@code TokenResult} of the dynamic-credential specification (design/qwp-token-provider-spec.md, section 4).
+ *
+ * The token is opaque and is stored without the {@code "Bearer "} prefix. It must be non-blank printable ASCII
+ * ({@code 0x20}-{@code 0x7e}); the constructors reject anything else, never echoing the token in the message.
+ *
+ * {@code expiresAtEpochMillis} is an absolute wall-clock instant. A source that receives a relative lifetime
+ * (such as {@code expires_in}) should add it to its own local time of receipt, which makes the cache robust to
+ * clock skew between the client and the identity platform; a source that cannot determine the expiry must
+ * supply a conservative value. {@code refreshAtEpochMillis}, when present, is the earliest instant the platform
+ * suggests refreshing; {@link #NO_REFRESH_AT} means none.
+ *
+ * {@link #toString()} never renders the token: it shows the length, an 8-hex-digit SHA-256 fingerprint and the
+ * timestamps.
+ */
+public final class ExpiringToken {
+ /**
+ * Value of {@code refreshAtEpochMillis} meaning the platform gave no refresh hint.
+ */
+ public static final long NO_REFRESH_AT = 0;
+ private final long expiresAtEpochMillis;
+ private final long refreshAtEpochMillis;
+ private final String token;
+
+ /**
+ * @param token the token, without the {@code "Bearer "} prefix
+ * @param expiresAtEpochMillis the absolute expiry, in milliseconds since the Unix epoch
+ * @throws IllegalArgumentException if the token is null, blank or not printable ASCII
+ */
+ public ExpiringToken(String token, long expiresAtEpochMillis) {
+ this(token, expiresAtEpochMillis, NO_REFRESH_AT);
+ }
+
+ /**
+ * @param token the token, without the {@code "Bearer "} prefix
+ * @param expiresAtEpochMillis the absolute expiry, in milliseconds since the Unix epoch
+ * @param refreshAtEpochMillis the earliest instant the platform suggests refreshing, or
+ * {@link #NO_REFRESH_AT} (any value {@code <= 0}) for none
+ * @throws IllegalArgumentException if the token is null, blank or not printable ASCII
+ */
+ public ExpiringToken(String token, long expiresAtEpochMillis, long refreshAtEpochMillis) {
+ String problem = describeInvalidToken(token);
+ if (problem != null) {
+ throw new IllegalArgumentException(problem);
+ }
+ this.token = token;
+ this.expiresAtEpochMillis = expiresAtEpochMillis;
+ this.refreshAtEpochMillis = refreshAtEpochMillis > 0 ? refreshAtEpochMillis : NO_REFRESH_AT;
+ }
+
+ /**
+ * Applies the token rules of the specification's section 3: non-null, not blank, and every character within
+ * printable ASCII. Returns a token-free description of the first violation, or {@code null} when the token
+ * is acceptable.
+ */
+ static String describeInvalidToken(CharSequence token) {
+ if (token == null || token.length() == 0) {
+ return "token is null or empty";
+ }
+ boolean blank = true;
+ for (int i = 0, n = token.length(); i < n; i++) {
+ char c = token.charAt(i);
+ if (c < 0x20 || c > 0x7e) {
+ return "token contains a control or non-ASCII character";
+ }
+ if (c != ' ') {
+ blank = false;
+ }
+ }
+ return blank ? "token is blank" : null;
+ }
+
+ private static String instant(long epochMillis) {
+ try {
+ return Instant.ofEpochMilli(epochMillis).toString();
+ } catch (RuntimeException e) {
+ return Long.toString(epochMillis);
+ }
+ }
+
+ /**
+ * @return the absolute expiry, in milliseconds since the Unix epoch
+ */
+ public long getExpiresAtEpochMillis() {
+ return expiresAtEpochMillis;
+ }
+
+ /**
+ * @return the platform's earliest suggested refresh instant, or {@link #NO_REFRESH_AT} when none was given
+ */
+ public long getRefreshAtEpochMillis() {
+ return refreshAtEpochMillis;
+ }
+
+ /**
+ * @return the token, without the {@code "Bearer "} prefix
+ */
+ public String getToken() {
+ return token;
+ }
+
+ /**
+ * @return true when the platform gave a refresh hint
+ */
+ public boolean hasRefreshAt() {
+ return refreshAtEpochMillis != NO_REFRESH_AT;
+ }
+
+ @Override
+ public String toString() {
+ StringBuilder sb = new StringBuilder("ExpiringToken{token=")
+ .append(CredentialRedaction.describeToken(token))
+ .append(", expiresAt=").append(instant(expiresAtEpochMillis));
+ if (hasRefreshAt()) {
+ sb.append(", refreshAt=").append(instant(refreshAtEpochMillis));
+ }
+ return sb.append('}').toString();
+ }
+}
diff --git a/core/src/main/java/io/questdb/client/cutlass/auth/RefreshingTokenProvider.java b/core/src/main/java/io/questdb/client/cutlass/auth/RefreshingTokenProvider.java
new file mode 100644
index 000000000..796eb3aea
--- /dev/null
+++ b/core/src/main/java/io/questdb/client/cutlass/auth/RefreshingTokenProvider.java
@@ -0,0 +1,1027 @@
+/*+*****************************************************************************
+ * ___ _ ____ ____
+ * / _ \ _ _ ___ ___| |_| _ \| __ )
+ * | | | | | | |/ _ \/ __| __| | | | _ \
+ * | |_| | |_| | __/\__ \ |_| |_| | |_) |
+ * \__\_\\__,_|\___||___/\__|____/|____/
+ *
+ * Copyright (c) 2014-2019 Appsicle
+ * Copyright (c) 2019-2026 QuestDB
+ *
+ * Licensed under the Apache License, Version 2.0 (the "License");
+ * you may not use this file except in compliance with the License.
+ * You may obtain a copy of the License at
+ *
+ * http://www.apache.org/licenses/LICENSE-2.0
+ *
+ * Unless required by applicable law or agreed to in writing, software
+ * distributed under the License is distributed on an "AS IS" BASIS,
+ * WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied.
+ * See the License for the specific language governing permissions and
+ * limitations under the License.
+ *
+ ******************************************************************************/
+
+package io.questdb.client.cutlass.auth;
+
+import io.questdb.client.HttpTokenProvider;
+import io.questdb.client.std.Chars;
+import io.questdb.client.std.QuietCloseable;
+import org.jetbrains.annotations.TestOnly;
+import org.slf4j.Logger;
+import org.slf4j.LoggerFactory;
+
+import java.time.Instant;
+import java.util.concurrent.ScheduledFuture;
+import java.util.concurrent.ScheduledThreadPoolExecutor;
+import java.util.concurrent.TimeUnit;
+import java.util.concurrent.ThreadLocalRandom;
+import java.util.concurrent.atomic.AtomicInteger;
+import java.util.concurrent.locks.Condition;
+import java.util.concurrent.locks.ReentrantLock;
+import java.util.function.DoubleSupplier;
+
+/**
+ * A shared, proactively refreshing cache in front of a {@link TokenSource}: the refreshing provider of the
+ * dynamic-credential specification (design/qwp-token-provider-spec.md, section 5). Hand one instance to
+ * {@code Sender.builder(...).httpTokenProvider(...)}, {@code QwpQueryClient.withBearerTokenProvider(...)} or
+ * {@code QuestDB.builder().httpTokenProvider(...)}; every connection then reads the cached token, and the
+ * provider fetches a new one in the background well before the current one expires.
+ *
Prefetch. The first fetch starts as soon as the provider is built, on the provider's own daemon
+ * thread. {@link #awaitReady(long)} lets an application gate startup on it.
+ *
Proactive refresh. After a fetch received at {@code f} with expiry {@code e}, the next fetch is
+ * scheduled at a jittered half-life, {@code f + (e - f) * U(0.45, 0.55)}, no later than
+ * {@code e - min(refresh_margin, (e - f) / 2)}, no earlier than {@code f + min(min_refresh_interval,
+ * (e - f) / 2)}, and no later than the platform's refresh hint when it gave one. While the source is healthy,
+ * even an idle client therefore always holds a usable token.
+ *
Hand-out. A token is usable only while {@code now < e - handout_floor}. {@link #getToken()}
+ * returns a usable token at once - a volatile read, no I/O, no lock. With no usable token it starts a fetch
+ * (unless one is running or the provider is backing off after failures) and waits up to {@code cold_wait},
+ * then throws a {@link TokenUnavailableException} carrying the classification of the most recent failed
+ * fetch, or a retryable one when no fetch failed. It never hands out an unusable token. Any number of
+ * concurrent callers cause at most one fetch.
+ *
Failures. A failed fetch is retried with jittered exponential backoff, from
+ * {@code backoff_initial} up to {@code backoff_max}, honouring the source's {@code retry_after} up to
+ * {@code retry_after_max}. Retries continue whatever the classification for as long as the provider is open,
+ * the current token keeps being served while it remains usable, and failures are logged as warnings at most
+ * once a minute, with only the classification and a sanitized message.
+ *
Forced refresh. {@link #onTokenRejected(CharSequence, int)} with {@code 401} for the current
+ * token starts a fetch, at most once per {@code forced_min_interval}, and waits up to {@code forced_wait} for
+ * it. A stale token (one that has already rotated) and any other status are ignored.
+ *
Interrupts. A caller interrupted while waiting gets a retryable {@link TokenUnavailableException}
+ * promptly, with its interrupt flag still set.
+ *
Close. {@link #close()} stops the refresher and wakes every waiter; afterwards
+ * {@link #getToken()} fails with a permanent error. Clients never close a provider the application supplied:
+ * close it yourself once every client using it is closed.
+ *
Clocks. Expiry comparisons use the wall clock; delays and waits use a monotonic clock.
+ *
+ * Nothing here ever renders the token: {@link #toString()}, log lines and error messages show at most its
+ * length and an 8-hex-digit SHA-256 fingerprint.
+ */
+public final class RefreshingTokenProvider implements HttpTokenProvider, QuietCloseable {
+ public static final long DEFAULT_BACKOFF_INITIAL_MILLIS = 500;
+ public static final long DEFAULT_BACKOFF_MAX_MILLIS = 60_000;
+ public static final long DEFAULT_COLD_WAIT_MILLIS = 30_000;
+ public static final long DEFAULT_FORCED_MIN_INTERVAL_MILLIS = 30_000;
+ public static final long DEFAULT_FORCED_WAIT_MILLIS = 5_000;
+ public static final long DEFAULT_HANDOUT_FLOOR_MILLIS = 60_000;
+ public static final long DEFAULT_MIN_REFRESH_INTERVAL_MILLIS = 30_000;
+ public static final long DEFAULT_REFRESH_MARGIN_MILLIS = 300_000;
+ public static final long DEFAULT_RETRY_AFTER_MAX_MILLIS = 300_000;
+ /**
+ * Value returned by the observability getters when there is nothing to report.
+ */
+ public static final long NONE = -1;
+ private static final long CLOSE_AWAIT_MILLIS = 1_000;
+ private static final long FAILURE_LOG_INTERVAL_NANOS = TimeUnit.MINUTES.toNanos(1);
+ private static final Logger LOG = LoggerFactory.getLogger(RefreshingTokenProvider.class);
+ // Delays are bookkept as monotonic deadlines compared by subtraction, which is only sound while every
+ // difference stays below 2^63. Clamping each delay to 2^61 ns (~73 years) keeps that true for any token
+ // lifetime a source can report, including Long.MAX_VALUE "never expires".
+ private static final long MAX_DELAY_NANOS = Long.MAX_VALUE >> 2;
+ private static final AtomicInteger THREAD_IDS = new AtomicInteger();
+ // A waiter re-checks its deadline at least this often when the clock is not the system clock, so a test
+ // that advances a fake clock is noticed. With the system clock the waits are untimed beyond their deadline.
+ private static final long WAIT_SLICE_NANOS = TimeUnit.MILLISECONDS.toNanos(20);
+ private final long backoffInitialMillis;
+ private final long backoffMaxMillis;
+ private final Clock clock;
+ private final long coldWaitMillis;
+ private final long forcedMinIntervalNanos;
+ private final long forcedWaitNanos;
+ private final long handoutFloorMillis;
+ private final ReentrantLock lock = new ReentrantLock();
+ private final long maxWaitSliceNanos;
+ private final long minRefreshIntervalMillis;
+ private final String name;
+ private final DoubleSupplier random;
+ private final long refreshMarginMillis;
+ private final long retryAfterMaxMillis;
+ private final Scheduler scheduler;
+ private final TokenSource source;
+ // Signalled whenever a fetch completes and on close.
+ private final Condition stateChanged = lock.newCondition();
+ // ---- guarded by lock ----
+ private boolean backoffActive;
+ private long backoffUntilNanos;
+ private volatile boolean closed;
+ private long completedFetches;
+ private volatile int consecutiveFailures;
+ private boolean failureLogged;
+ private volatile boolean fetchRunning;
+ private boolean forcedRefreshed;
+ // The current token, or null. Immutable snapshot, published with a volatile write so the warm path in
+ // getToken() needs no lock.
+ private volatile Held held;
+ private volatile FetchFailure lastFailure;
+ private long lastFailureLogNanos;
+ private long lastForcedNanos;
+ private volatile long lastSuccessEpochMillis = NONE;
+ private long scheduledDueNanos;
+ private long scheduledSeq;
+ private ScheduledTask scheduledTask;
+
+ private RefreshingTokenProvider(Builder b) {
+ this.source = b.source;
+ this.name = b.name;
+ this.clock = b.clock;
+ this.maxWaitSliceNanos = b.clock == SystemClock.INSTANCE ? Long.MAX_VALUE : WAIT_SLICE_NANOS;
+ this.scheduler = b.scheduler != null ? b.scheduler : new DefaultScheduler();
+ this.random = b.random;
+ this.refreshMarginMillis = b.refreshMarginMillis;
+ this.minRefreshIntervalMillis = b.minRefreshIntervalMillis;
+ this.handoutFloorMillis = b.handoutFloorMillis;
+ this.coldWaitMillis = b.coldWaitMillis;
+ this.backoffInitialMillis = b.backoffInitialMillis;
+ this.backoffMaxMillis = b.backoffMaxMillis;
+ this.retryAfterMaxMillis = b.retryAfterMaxMillis;
+ this.forcedMinIntervalNanos = TimeUnit.MILLISECONDS.toNanos(b.forcedMinIntervalMillis);
+ this.forcedWaitNanos = TimeUnit.MILLISECONDS.toNanos(b.forcedWaitMillis);
+ }
+
+ /**
+ * Starts building a provider around {@code source}.
+ *
+ * @param source obtains new tokens; called only from the provider's refresher thread
+ * @return a builder whose defaults are the specification's
+ */
+ public static Builder builder(TokenSource source) {
+ if (source == null) {
+ throw new IllegalArgumentException("token source must not be null");
+ }
+ return new Builder(source);
+ }
+
+ /**
+ * The refresh instant of the specification's section 5.2 for a token received at {@code f} that expires at
+ * {@code e}, given a uniform sample {@code u} in {@code [0, 1)}. Pure; exposed for the conformance tests.
+ */
+ @TestOnly
+ public static long computeRefreshAtMillis(
+ long f,
+ long e,
+ long refreshAt,
+ double u,
+ long refreshMarginMillis,
+ long minRefreshIntervalMillis
+ ) {
+ final long life = e - f;
+ final long half = life / 2;
+ final long margin = Math.min(refreshMarginMillis, half);
+ final long floor = Math.min(minRefreshIntervalMillis, half);
+ final double fraction = 0.45 + 0.10 * Math.max(0.0, Math.min(1.0, u));
+ long r = f + (long) (life * fraction);
+ r = Math.min(r, e - margin);
+ r = Math.max(r, f + floor);
+ if (refreshAt > 0) {
+ r = Math.min(r, Math.max(refreshAt, f + floor));
+ }
+ return r;
+ }
+
+ private static String instant(long epochMillis) {
+ try {
+ return Instant.ofEpochMilli(epochMillis).toString();
+ } catch (RuntimeException e) {
+ return Long.toString(epochMillis);
+ }
+ }
+
+ private static long millisToNanos(long millis) {
+ return Math.min(TimeUnit.MILLISECONDS.toNanos(Math.max(0, millis)), MAX_DELAY_NANOS);
+ }
+
+ /**
+ * Waits until the provider holds a usable token, starting a fetch if none is running and the provider is
+ * not backing off. Use it to gate application startup on the first fetch.
+ *
+ * @param timeoutMillis the longest to wait
+ * @return true once a usable token is held; false on timeout, when the provider is closed, or when the
+ * calling thread is interrupted (its interrupt flag is then left set)
+ */
+ public boolean awaitReady(long timeoutMillis) {
+ lock.lock();
+ try {
+ final long start = clock.monotonicNanos();
+ final long deadline = start + millisToNanos(timeoutMillis);
+ if (!usable(held, clock.wallClockMillis())) {
+ requestFetchNowLocked(start);
+ }
+ while (true) {
+ if (closed) {
+ return false;
+ }
+ if (usable(held, clock.wallClockMillis())) {
+ return true;
+ }
+ final long remaining = deadline - clock.monotonicNanos();
+ if (remaining <= 0) {
+ return false;
+ }
+ try {
+ stateChanged.awaitNanos(Math.min(remaining, maxWaitSliceNanos));
+ } catch (InterruptedException e) {
+ Thread.currentThread().interrupt();
+ return false;
+ }
+ }
+ } finally {
+ lock.unlock();
+ }
+ }
+
+ /**
+ * Stops scheduled fetches, interrupts a fetch in progress and wakes every waiter. Afterwards
+ * {@link #getToken()} fails with a permanent {@link TokenUnavailableException}. Idempotent. Waits at most
+ * one second for the refresher thread to exit.
+ */
+ @Override
+ public void close() {
+ lock.lock();
+ try {
+ if (closed) {
+ return;
+ }
+ closed = true;
+ cancelScheduledLocked();
+ held = null;
+ stateChanged.signalAll();
+ } finally {
+ lock.unlock();
+ }
+ scheduler.shutdown();
+ LOG.debug("token provider {} closed", name);
+ }
+
+ /**
+ * @return the number of consecutive failed fetches since the last successful one
+ */
+ public int getConsecutiveFailures() {
+ return consecutiveFailures;
+ }
+
+ /**
+ * The most recent failed fetch - its classification, sanitized message and time - or null when no fetch
+ * has failed. Kept after the source recovers.
+ */
+ public FetchFailure getLastFailure() {
+ return lastFailure;
+ }
+
+ /**
+ * @return the wall-clock time of the last successful fetch, or {@link #NONE}
+ */
+ public long getLastSuccessEpochMillis() {
+ return lastSuccessEpochMillis;
+ }
+
+ /**
+ * @return the non-secret label this provider uses in log lines and errors
+ */
+ public String getName() {
+ return name;
+ }
+
+ /**
+ * Returns a usable token: at once when one is held, otherwise after waiting up to {@code cold_wait} for a
+ * fetch.
+ *
+ * @return the current token, without the {@code "Bearer "} prefix
+ * @throws TokenUnavailableException when no usable token arrives in time (classified as the most recent
+ * failed fetch, or retryable when none failed), when the calling thread
+ * is interrupted (retryable; the interrupt flag stays set), or when the
+ * provider is closed (permanent)
+ */
+ @Override
+ public CharSequence getToken() {
+ if (closed) {
+ throw closedException();
+ }
+ final Held h = held;
+ if (h != null) {
+ final long wallNow = clock.wallClockMillis();
+ if (usable(h, wallNow)) {
+ // Warm path: a volatile read, no I/O and no lock. Only when the refresh is overdue - its
+ // scheduled fetch has not run, say after the host slept through it - does a caller take the
+ // lock, to start that fetch in the background.
+ if (!fetchRunning && consecutiveFailures == 0
+ && (wallNow >= h.refreshAtEpochMillis || clock.monotonicNanos() - h.refreshAtNanos >= 0)) {
+ startOverdueRefresh();
+ }
+ return h.token;
+ }
+ }
+ return awaitUsableToken();
+ }
+
+ /**
+ * @return the current token's expiry, or {@link #NONE} when no token is held
+ */
+ public long getTokenExpiresAtEpochMillis() {
+ final Held h = held;
+ return h == null ? NONE : h.expiresAtEpochMillis;
+ }
+
+ /**
+ * @return true once {@link #close()} has been called
+ */
+ public boolean isClosed() {
+ return closed;
+ }
+
+ /**
+ * Forced refresh. When a server answered {@code 401} to {@code token} and that token is still the current
+ * one, starts a fetch now - joining one already running, and respecting failure backoff - and waits up to
+ * {@code forced_wait} for it. Ignored for any other status, for a token that has already rotated, and when
+ * a forced refresh ran less than {@code forced_min_interval} ago. A forced fetch that returns the same token
+ * counts as a normal success. Never throws; returns promptly when interrupted, with the flag still set.
+ */
+ @Override
+ public void onTokenRejected(CharSequence token, int httpStatus) {
+ if (httpStatus != 401 || token == null) {
+ return;
+ }
+ lock.lock();
+ try {
+ if (closed) {
+ return;
+ }
+ final Held h = held;
+ if (h == null || !Chars.equals(h.token, token)) {
+ return; // stale: the token has already rotated, so the caller's next pull gets the new one
+ }
+ final long now = clock.monotonicNanos();
+ if (forcedRefreshed && now - lastForcedNanos < forcedMinIntervalNanos) {
+ return;
+ }
+ forcedRefreshed = true;
+ lastForcedNanos = now;
+ LOG.info("token provider {}: the server rejected the current token with 401, refreshing early", name);
+ final long target = completedFetches + 1;
+ requestFetchNowLocked(now);
+ final long deadline = now + forcedWaitNanos;
+ while (!closed && completedFetches < target) {
+ final long remaining = deadline - clock.monotonicNanos();
+ if (remaining <= 0) {
+ break;
+ }
+ stateChanged.awaitNanos(Math.min(remaining, maxWaitSliceNanos));
+ }
+ } catch (InterruptedException e) {
+ Thread.currentThread().interrupt();
+ } finally {
+ lock.unlock();
+ }
+ }
+
+ @Override
+ public String toString() {
+ final Held h = held;
+ final FetchFailure f = lastFailure;
+ final StringBuilder sb = new StringBuilder("RefreshingTokenProvider{name=").append(name);
+ if (h == null) {
+ sb.append(", token=");
+ } else {
+ sb.append(", token=").append(CredentialRedaction.describeToken(h.token))
+ .append(", expiresAt=").append(instant(h.expiresAtEpochMillis))
+ .append(", refreshAt=").append(instant(h.refreshAtEpochMillis));
+ }
+ final long success = lastSuccessEpochMillis;
+ if (success != NONE) {
+ sb.append(", lastSuccess=").append(instant(success));
+ }
+ sb.append(", consecutiveFailures=").append(consecutiveFailures);
+ if (f != null) {
+ sb.append(", lastFailure=").append(f);
+ }
+ if (closed) {
+ sb.append(", closed");
+ }
+ return sb.append('}').toString();
+ }
+
+ private CharSequence awaitUsableToken() {
+ lock.lock();
+ try {
+ final long start = clock.monotonicNanos();
+ final long deadline = start + millisToNanos(coldWaitMillis);
+ // A caller that arrives without a usable token starts one fetch - unless one is running or the
+ // provider is backing off - and then waits for it or for the fetches already scheduled. It does not
+ // start another each time a fetch completes: a source that keeps returning a token inside the
+ // hand-out floor would otherwise be fetched in a tight loop for the whole wait.
+ if (!usable(held, clock.wallClockMillis())) {
+ requestFetchNowLocked(start);
+ }
+ while (true) {
+ if (closed) {
+ throw closedException();
+ }
+ final Held h = held;
+ if (usable(h, clock.wallClockMillis())) {
+ return h.token;
+ }
+ final long now = clock.monotonicNanos();
+ // The earliest permitted fetch lies beyond the caller's deadline: nothing can arrive in time,
+ // so say so now rather than parking the caller for nothing.
+ if (!fetchRunning && inBackoffLocked(now) && backoffUntilNanos - deadline > 0) {
+ throw unavailableLocked(now, start);
+ }
+ final long remaining = deadline - now;
+ if (remaining <= 0) {
+ throw unavailableLocked(now, start);
+ }
+ try {
+ stateChanged.awaitNanos(Math.min(remaining, maxWaitSliceNanos));
+ } catch (InterruptedException e) {
+ Thread.currentThread().interrupt();
+ throw new TokenUnavailableException("interrupted while waiting for a token from " + name, true);
+ }
+ }
+ } finally {
+ lock.unlock();
+ }
+ }
+
+ private long backoffMillis(int failures, long retryAfterMillis) {
+ final int shift = Math.min(failures - 1, 40);
+ long base = backoffInitialMillis << shift;
+ if (base < 0 || (base >> shift) != backoffInitialMillis) {
+ base = backoffMaxMillis; // overflowed
+ }
+ base = Math.min(base, backoffMaxMillis);
+ final long half = base / 2;
+ long delay = half + (long) (half * uniform());
+ if (retryAfterMillis >= 0) {
+ delay = Math.max(delay, Math.min(retryAfterMillis, retryAfterMaxMillis));
+ }
+ return delay;
+ }
+
+ private void cancelScheduledLocked() {
+ scheduledSeq++;
+ final ScheduledTask t = scheduledTask;
+ scheduledTask = null;
+ if (t != null) {
+ t.cancel();
+ }
+ }
+
+ private FetchFailure checkResult(ExpiringToken result, long receivedAt) {
+ if (result == null) {
+ return new FetchFailure(true, TokenUnavailableException.NO_RETRY_AFTER,
+ "token source returned null", receivedAt);
+ }
+ final String problem = ExpiringToken.describeInvalidToken(result.getToken());
+ if (problem != null) {
+ return new FetchFailure(true, TokenUnavailableException.NO_RETRY_AFTER,
+ "token source returned an unusable token: " + problem, receivedAt);
+ }
+ if (result.getExpiresAtEpochMillis() <= receivedAt) {
+ return new FetchFailure(true, TokenUnavailableException.NO_RETRY_AFTER,
+ "token source returned a token that expired at " + instant(result.getExpiresAtEpochMillis())
+ + ", not after it was received at " + instant(receivedAt),
+ receivedAt);
+ }
+ return null;
+ }
+
+ private FetchFailure classify(Throwable t, long receivedAt) {
+ if (t instanceof TokenUnavailableException) {
+ final TokenUnavailableException e = (TokenUnavailableException) t;
+ return new FetchFailure(e.isRetryable(), e.getRetryAfterMillis(), e.getMessage(), receivedAt);
+ }
+ // An unclassified failure is retryable (the specification's rule when unsure). Name its type: the
+ // message alone ("connect timed out") often does not say what failed.
+ final String message = t.getMessage();
+ return new FetchFailure(true, TokenUnavailableException.NO_RETRY_AFTER,
+ message == null ? t.getClass().getName() : t.getClass().getName() + ": " + message, receivedAt);
+ }
+
+ private TokenUnavailableException closedException() {
+ return new TokenUnavailableException("token provider " + name + " is closed", false);
+ }
+
+ private boolean inBackoffLocked(long now) {
+ return backoffActive && now - backoffUntilNanos < 0;
+ }
+
+ private String onFailureLocked(FetchFailure failure, long receivedNanos) {
+ final int failures = consecutiveFailures + 1;
+ consecutiveFailures = failures;
+ lastFailure = failure;
+ final long delayMillis = backoffMillis(failures, failure.retryAfterMillis);
+ final long delayNanos = millisToNanos(delayMillis);
+ backoffActive = true;
+ backoffUntilNanos = receivedNanos + delayNanos;
+ scheduleFetchLocked(delayNanos, receivedNanos);
+ if (failureLogged && receivedNanos - lastFailureLogNanos < FAILURE_LOG_INTERVAL_NANOS) {
+ return null;
+ }
+ failureLogged = true;
+ lastFailureLogNanos = receivedNanos;
+ return "token provider " + name + ": fetch failed [retryable=" + failure.retryable
+ + ", consecutiveFailures=" + failures + ", retryInMillis=" + delayMillis + "]: " + failure.message;
+ }
+
+ private void onSuccessLocked(ExpiringToken result, long receivedAt, long receivedNanos) {
+ final long refreshAt = computeRefreshAtMillis(
+ receivedAt,
+ result.getExpiresAtEpochMillis(),
+ result.getRefreshAtEpochMillis(),
+ uniform(),
+ refreshMarginMillis,
+ minRefreshIntervalMillis
+ );
+ final long delayNanos = millisToNanos(refreshAt - receivedAt);
+ final int failures = consecutiveFailures;
+ held = new Held(result.getToken(), result.getExpiresAtEpochMillis(), refreshAt, receivedNanos + delayNanos);
+ lastSuccessEpochMillis = receivedAt;
+ consecutiveFailures = 0;
+ backoffActive = false;
+ scheduleFetchLocked(delayNanos, receivedNanos);
+ if (failures > 0) {
+ LOG.info("token provider {}: fetch succeeded after {} consecutive failure(s) [token={}, expiresAt={}]",
+ name, failures, CredentialRedaction.describeToken(result.getToken()),
+ instant(result.getExpiresAtEpochMillis()));
+ } else if (LOG.isDebugEnabled()) {
+ LOG.debug("token provider {}: fetched [token={}, expiresAt={}, refreshAt={}]",
+ name, CredentialRedaction.describeToken(result.getToken()),
+ instant(result.getExpiresAtEpochMillis()), instant(refreshAt));
+ }
+ }
+
+ // Starts a fetch now unless one is running, one is already due, or the provider is in failure backoff.
+ private void requestFetchNowLocked(long now) {
+ if (closed || fetchRunning || inBackoffLocked(now)) {
+ return;
+ }
+ if (scheduledTask != null && now - scheduledDueNanos >= 0) {
+ return; // already due: the scheduler runs it as soon as it can
+ }
+ scheduleFetchLocked(0, now);
+ }
+
+ private void runFetch(long seq) {
+ lock.lock();
+ try {
+ if (closed || seq != scheduledSeq || fetchRunning) {
+ return; // cancelled, superseded, or a fetch is already in flight
+ }
+ scheduledTask = null;
+ fetchRunning = true;
+ } finally {
+ lock.unlock();
+ }
+ ExpiringToken result = null;
+ Throwable error = null;
+ try {
+ result = source.fetchToken();
+ } catch (Throwable t) {
+ // Every failure, an Error included, is a failed fetch to retry with backoff: rethrowing would only
+ // be swallowed by the scheduler and leave the provider with no fetch scheduled, ever again.
+ error = t;
+ }
+ final long receivedAt = clock.wallClockMillis();
+ final long receivedNanos = clock.monotonicNanos();
+ FetchFailure failure = error != null ? classify(error, receivedAt) : checkResult(result, receivedAt);
+ if (failure != null) {
+ failure = failure.sanitized();
+ }
+ String warning = null;
+ lock.lock();
+ try {
+ fetchRunning = false;
+ completedFetches++;
+ if (!closed) {
+ if (failure == null) {
+ onSuccessLocked(result, receivedAt, receivedNanos);
+ } else {
+ warning = onFailureLocked(failure, receivedNanos);
+ }
+ }
+ stateChanged.signalAll();
+ } finally {
+ lock.unlock();
+ }
+ if (warning != null) {
+ LOG.warn(warning);
+ }
+ }
+
+ private void scheduleFetchLocked(long delayNanos, long now) {
+ if (closed) {
+ return;
+ }
+ cancelScheduledLocked();
+ final long seq = scheduledSeq;
+ scheduledDueNanos = now + delayNanos;
+ scheduledTask = scheduler.schedule(() -> runFetch(seq), delayNanos);
+ }
+
+ private void startOverdueRefresh() {
+ lock.lock();
+ try {
+ requestFetchNowLocked(clock.monotonicNanos());
+ } finally {
+ lock.unlock();
+ }
+ }
+
+ private double uniform() {
+ return random.getAsDouble();
+ }
+
+ private TokenUnavailableException unavailableLocked(long now, long start) {
+ final FetchFailure f = consecutiveFailures > 0 ? lastFailure : null;
+ final long waitedMillis = TimeUnit.NANOSECONDS.toMillis(Math.max(0, now - start));
+ if (f == null) {
+ return new TokenUnavailableException("no usable token from " + name + " after waiting "
+ + waitedMillis + " ms", true);
+ }
+ final long retryAfter = backoffActive && backoffUntilNanos - now > 0
+ ? TimeUnit.NANOSECONDS.toMillis(backoffUntilNanos - now)
+ : TokenUnavailableException.NO_RETRY_AFTER;
+ return new TokenUnavailableException("no usable token from " + name + " after waiting " + waitedMillis
+ + " ms [retryable=" + f.retryable + ", consecutiveFailures=" + consecutiveFailures + "]: "
+ + f.message, f.retryable, retryAfter);
+ }
+
+ private boolean usable(Held h, long wallNow) {
+ return h != null && wallNow < h.expiresAtEpochMillis - handoutFloorMillis;
+ }
+
+ /**
+ * Time source. Expiry comparisons use {@link #wallClockMillis()}; delays and waits use
+ * {@link #monotonicNanos()}.
+ */
+ public interface Clock {
+ long monotonicNanos();
+
+ long wallClockMillis();
+ }
+
+ /**
+ * Runs fetches on the provider's background context. The default owns one daemon thread. A test may supply
+ * its own to run fetches deterministically.
+ */
+ public interface Scheduler {
+ /**
+ * Runs {@code task} once, after {@code delayNanos} on the monotonic clock.
+ */
+ ScheduledTask schedule(Runnable task, long delayNanos);
+
+ /**
+ * Called once from {@link RefreshingTokenProvider#close()}: stop running tasks and interrupt a task in
+ * progress.
+ */
+ void shutdown();
+ }
+
+ /**
+ * Handle to a task submitted to a {@link Scheduler}.
+ */
+ public interface ScheduledTask {
+ /**
+ * Best-effort: a task that has already started may still run, and the provider ignores it.
+ */
+ void cancel();
+ }
+
+ /**
+ * Builder for {@link RefreshingTokenProvider}. The defaults are the specification's (section 5.1).
+ */
+ public static final class Builder {
+ private final TokenSource source;
+ private long backoffInitialMillis = DEFAULT_BACKOFF_INITIAL_MILLIS;
+ private long backoffMaxMillis = DEFAULT_BACKOFF_MAX_MILLIS;
+ private Clock clock = SystemClock.INSTANCE;
+ private long coldWaitMillis = DEFAULT_COLD_WAIT_MILLIS;
+ private long forcedMinIntervalMillis = DEFAULT_FORCED_MIN_INTERVAL_MILLIS;
+ private long forcedWaitMillis = DEFAULT_FORCED_WAIT_MILLIS;
+ private long handoutFloorMillis = DEFAULT_HANDOUT_FLOOR_MILLIS;
+ private long minRefreshIntervalMillis = DEFAULT_MIN_REFRESH_INTERVAL_MILLIS;
+ private String name = "token-provider";
+ private DoubleSupplier random = () -> ThreadLocalRandom.current().nextDouble();
+ private long refreshMarginMillis = DEFAULT_REFRESH_MARGIN_MILLIS;
+ private long retryAfterMaxMillis = DEFAULT_RETRY_AFTER_MAX_MILLIS;
+ private Scheduler scheduler;
+
+ private Builder(TokenSource source) {
+ this.source = source;
+ }
+
+ private static long nonNegative(String name, long value) {
+ if (value < 0) {
+ throw new IllegalArgumentException(name + " must be >= 0: " + value);
+ }
+ return value;
+ }
+
+ /**
+ * First retry delay after a failed fetch. Default 500 ms.
+ */
+ public Builder backoffInitialMillis(long millis) {
+ if (millis <= 0) {
+ throw new IllegalArgumentException("backoff_initial must be > 0: " + millis);
+ }
+ this.backoffInitialMillis = millis;
+ return this;
+ }
+
+ /**
+ * Largest retry delay after failed fetches. Default 60 s.
+ */
+ public Builder backoffMaxMillis(long millis) {
+ if (millis <= 0) {
+ throw new IllegalArgumentException("backoff_max must be > 0: " + millis);
+ }
+ this.backoffMaxMillis = millis;
+ return this;
+ }
+
+ /**
+ * Builds the provider and starts its first fetch in the background.
+ */
+ public RefreshingTokenProvider build() {
+ if (backoffMaxMillis < backoffInitialMillis) {
+ throw new IllegalArgumentException("backoff_max must be >= backoff_initial [backoff_max="
+ + backoffMaxMillis + ", backoff_initial=" + backoffInitialMillis + ']');
+ }
+ RefreshingTokenProvider provider = new RefreshingTokenProvider(this);
+ provider.lock.lock();
+ try {
+ provider.scheduleFetchLocked(0, provider.clock.monotonicNanos());
+ } finally {
+ provider.lock.unlock();
+ }
+ return provider;
+ }
+
+ /**
+ * Test seam: the clock for expiry comparisons, delays and waits.
+ */
+ @TestOnly
+ public Builder clock(Clock clock) {
+ if (clock == null) {
+ throw new IllegalArgumentException("clock must not be null");
+ }
+ this.clock = clock;
+ return this;
+ }
+
+ /**
+ * The longest {@code getToken()} waits when no usable token is held. Default 30 s.
+ */
+ public Builder coldWaitMillis(long millis) {
+ this.coldWaitMillis = nonNegative("cold_wait", millis);
+ return this;
+ }
+
+ /**
+ * Minimum spacing between forced refreshes. Default 30 s.
+ */
+ public Builder forcedMinIntervalMillis(long millis) {
+ this.forcedMinIntervalMillis = nonNegative("forced_min_interval", millis);
+ return this;
+ }
+
+ /**
+ * The longest {@code onTokenRejected} waits for a forced refresh. Default 5 s.
+ */
+ public Builder forcedWaitMillis(long millis) {
+ this.forcedWaitMillis = nonNegative("forced_wait", millis);
+ return this;
+ }
+
+ /**
+ * A token is usable only while {@code now < expires_at - handout_floor}. Default 60 s.
+ */
+ public Builder handoutFloorMillis(long millis) {
+ this.handoutFloorMillis = nonNegative("handout_floor", millis);
+ return this;
+ }
+
+ /**
+ * Lower bound on the time between a fetch and the next scheduled refresh. Default 30 s.
+ */
+ public Builder minRefreshIntervalMillis(long millis) {
+ this.minRefreshIntervalMillis = nonNegative("min_refresh_interval", millis);
+ return this;
+ }
+
+ /**
+ * A non-secret label for log lines, errors and the refresher thread, e.g.
+ * {@code azure[resource=api://...]}. Default {@code token-provider}.
+ */
+ public Builder name(String name) {
+ if (name == null || name.isEmpty()) {
+ throw new IllegalArgumentException("name must not be empty");
+ }
+ this.name = name;
+ return this;
+ }
+
+ /**
+ * Test seam: the source of uniform samples in {@code [0, 1)} for the refresh and backoff jitter.
+ */
+ @TestOnly
+ public Builder random(DoubleSupplier uniform) {
+ if (uniform == null) {
+ throw new IllegalArgumentException("random must not be null");
+ }
+ this.random = uniform;
+ return this;
+ }
+
+ /**
+ * Upper bound on how close to expiry a scheduled refresh may run. Default 5 min.
+ */
+ public Builder refreshMarginMillis(long millis) {
+ this.refreshMarginMillis = nonNegative("refresh_margin", millis);
+ return this;
+ }
+
+ /**
+ * Cap applied to a source's {@code retry_after}. Default 5 min.
+ */
+ public Builder retryAfterMaxMillis(long millis) {
+ this.retryAfterMaxMillis = nonNegative("retry_after_max", millis);
+ return this;
+ }
+
+ /**
+ * Test seam: where fetches run. The provider calls {@link Scheduler#shutdown()} from {@code close()}.
+ */
+ @TestOnly
+ public Builder scheduler(Scheduler scheduler) {
+ if (scheduler == null) {
+ throw new IllegalArgumentException("scheduler must not be null");
+ }
+ this.scheduler = scheduler;
+ return this;
+ }
+ }
+
+ /**
+ * A failed fetch: its classification, sanitized message and time. Never carries the token.
+ */
+ public static final class FetchFailure {
+ private final long epochMillis;
+ private final String message;
+ private final long retryAfterMillis;
+ private final boolean retryable;
+
+ FetchFailure(boolean retryable, long retryAfterMillis, String message, long epochMillis) {
+ this.retryable = retryable;
+ this.retryAfterMillis = retryAfterMillis < 0 ? TokenUnavailableException.NO_RETRY_AFTER : retryAfterMillis;
+ this.message = message;
+ this.epochMillis = epochMillis;
+ }
+
+ /**
+ * @return the wall-clock time of the failure
+ */
+ public long getEpochMillis() {
+ return epochMillis;
+ }
+
+ /**
+ * @return the failure message, sanitized and at most 256 characters
+ */
+ public String getMessage() {
+ return message;
+ }
+
+ /**
+ * @return the source's suggested wait before retrying, or {@link TokenUnavailableException#NO_RETRY_AFTER}
+ */
+ public long getRetryAfterMillis() {
+ return retryAfterMillis;
+ }
+
+ /**
+ * @return whether the source classified the failure as retryable
+ */
+ public boolean isRetryable() {
+ return retryable;
+ }
+
+ @Override
+ public String toString() {
+ return "FetchFailure{retryable=" + retryable + ", at=" + instant(epochMillis) + ", message=" + message + '}';
+ }
+
+ FetchFailure sanitized() {
+ final String clean = CredentialRedaction.sanitizeErrorText(message);
+ return new FetchFailure(retryable, retryAfterMillis, clean == null || clean.isEmpty()
+ ? "token source failed without a message" : clean, epochMillis);
+ }
+ }
+
+ private static final class DefaultScheduler implements Scheduler {
+ private final ScheduledThreadPoolExecutor executor;
+ private volatile Thread thread;
+
+ DefaultScheduler() {
+ final String threadName = "qdb-token-refresh-" + THREAD_IDS.incrementAndGet();
+ this.executor = new ScheduledThreadPoolExecutor(1, r -> {
+ Thread t = new Thread(r, threadName);
+ t.setDaemon(true);
+ thread = t;
+ return t;
+ });
+ executor.setRemoveOnCancelPolicy(true);
+ executor.setExecuteExistingDelayedTasksAfterShutdownPolicy(false);
+ executor.setContinueExistingPeriodicTasksAfterShutdownPolicy(false);
+ }
+
+ @Override
+ public ScheduledTask schedule(Runnable task, long delayNanos) {
+ final ScheduledFuture> future = executor.schedule(task, delayNanos, TimeUnit.NANOSECONDS);
+ return () -> future.cancel(false);
+ }
+
+ @Override
+ public void shutdown() {
+ executor.shutdownNow();
+ if (Thread.currentThread() == thread) {
+ return; // closed from inside a fetch: the thread exits once this task unwinds
+ }
+ // Interrupt-neutral, like the other close paths in this library: a carried interrupt flag would
+ // turn the bounded wait into an immediate return.
+ final boolean interrupted = Thread.interrupted();
+ try {
+ executor.awaitTermination(CLOSE_AWAIT_MILLIS, TimeUnit.MILLISECONDS);
+ } catch (InterruptedException e) {
+ Thread.currentThread().interrupt();
+ } finally {
+ if (interrupted) {
+ Thread.currentThread().interrupt();
+ }
+ }
+ }
+ }
+
+ private static final class Held {
+ final long expiresAtEpochMillis;
+ final long refreshAtEpochMillis;
+ final long refreshAtNanos;
+ final String token;
+
+ Held(String token, long expiresAtEpochMillis, long refreshAtEpochMillis, long refreshAtNanos) {
+ this.token = token;
+ this.expiresAtEpochMillis = expiresAtEpochMillis;
+ this.refreshAtEpochMillis = refreshAtEpochMillis;
+ this.refreshAtNanos = refreshAtNanos;
+ }
+ }
+
+ private static final class SystemClock implements Clock {
+ static final SystemClock INSTANCE = new SystemClock();
+
+ @Override
+ public long monotonicNanos() {
+ return System.nanoTime();
+ }
+
+ @Override
+ public long wallClockMillis() {
+ return System.currentTimeMillis();
+ }
+ }
+}
diff --git a/core/src/main/java/io/questdb/client/cutlass/auth/TokenProviderFactory.java b/core/src/main/java/io/questdb/client/cutlass/auth/TokenProviderFactory.java
new file mode 100644
index 000000000..2e96e1fd0
--- /dev/null
+++ b/core/src/main/java/io/questdb/client/cutlass/auth/TokenProviderFactory.java
@@ -0,0 +1,83 @@
+/*+*****************************************************************************
+ * ___ _ ____ ____
+ * / _ \ _ _ ___ ___| |_| _ \| __ )
+ * | | | | | | |/ _ \/ __| __| | | | _ \
+ * | |_| | |_| | __/\__ \ |_| |_| | |_) |
+ * \__\_\\__,_|\___||___/\__|____/|____/
+ *
+ * Copyright (c) 2014-2019 Appsicle
+ * Copyright (c) 2019-2026 QuestDB
+ *
+ * Licensed under the Apache License, Version 2.0 (the "License");
+ * you may not use this file except in compliance with the License.
+ * You may obtain a copy of the License at
+ *
+ * http://www.apache.org/licenses/LICENSE-2.0
+ *
+ * Unless required by applicable law or agreed to in writing, software
+ * distributed under the License is distributed on an "AS IS" BASIS,
+ * WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied.
+ * See the License for the specific language governing permissions and
+ * limitations under the License.
+ *
+ ******************************************************************************/
+
+package io.questdb.client.cutlass.auth;
+
+import java.util.Map;
+
+/**
+ * Service-provider interface behind the {@code token_provider} connect-string key (design/qwp-token-provider-spec.md,
+ * section 7). An implementation is found through {@link java.util.ServiceLoader}: list it in
+ * {@code META-INF/services/io.questdb.client.cutlass.auth.TokenProviderFactory}, or declare it with
+ * {@code provides} in a module descriptor. The optional {@code questdb-client-azure} artifact supplies the
+ * {@code azure} factory this way.
+ *
+ * A factory turns the provider-specific connect-string keys into a {@link TokenSource}. The client wraps that
+ * source in a {@link RefreshingTokenProvider} that the process-wide {@link TokenProviderRegistry} shares between
+ * every client built from an equivalent connect string, so a process runs one refresher per configuration.
+ *
+ * The connect-string vocabulary is fixed - unknown keys are rejected before a factory ever sees them - so a
+ * factory receives only the keys the client defines for it, already normalized: for {@code azure}, the
+ * {@code azure_resource} (with any trailing {@code /.default} removed) and the optional, lower-cased
+ * {@code azure_client_id}. Values in these keys are never secret: secrets come from the platform's standard
+ * configuration, never from the connect string.
+ */
+public interface TokenProviderFactory {
+
+ /**
+ * Creates the token source for one configuration. Called once per distinct configuration, when the first
+ * client using it connects. It must not perform network I/O: the provider's refresher thread calls
+ * {@link TokenSource#fetchToken()} later.
+ *
+ * @param params the provider-specific keys, normalized; unmodifiable
+ * @return the token source
+ * @throws IllegalArgumentException when the parameters cannot be used
+ */
+ TokenSource createSource(Map params);
+
+ /**
+ * A non-secret label for log lines and errors, e.g. {@code azure[resource=api://...]}.
+ *
+ * @param params the provider-specific keys, normalized
+ * @return the label
+ */
+ default String describe(Map params) {
+ return params.isEmpty() ? name() : name() + params;
+ }
+
+ /**
+ * The value of {@code token_provider} this factory serves, such as {@code azure}.
+ */
+ String name();
+
+ /**
+ * Validates the provider-specific keys when a connect string is parsed, before any client connects. Must be
+ * pure: it runs during configuration validation that must not fetch a token (section 7.2).
+ *
+ * @param params the provider-specific keys, normalized; unmodifiable
+ * @throws IllegalArgumentException naming the offending key
+ */
+ default void validate(Map params) {
+ }
+}
diff --git a/core/src/main/java/io/questdb/client/cutlass/auth/TokenProviderRegistry.java b/core/src/main/java/io/questdb/client/cutlass/auth/TokenProviderRegistry.java
new file mode 100644
index 000000000..7c55a6f74
--- /dev/null
+++ b/core/src/main/java/io/questdb/client/cutlass/auth/TokenProviderRegistry.java
@@ -0,0 +1,302 @@
+/*+*****************************************************************************
+ * ___ _ ____ ____
+ * / _ \ _ _ ___ ___| |_| _ \| __ )
+ * | | | | | | |/ _ \/ __| __| | | | _ \
+ * | |_| | |_| | __/\__ \ |_| |_| | |_) |
+ * \__\_\\__,_|\___||___/\__|____/|____/
+ *
+ * Copyright (c) 2014-2019 Appsicle
+ * Copyright (c) 2019-2026 QuestDB
+ *
+ * Licensed under the Apache License, Version 2.0 (the "License");
+ * you may not use this file except in compliance with the License.
+ * You may obtain a copy of the License at
+ *
+ * http://www.apache.org/licenses/LICENSE-2.0
+ *
+ * Unless required by applicable law or agreed to in writing, software
+ * distributed under the License is distributed on an "AS IS" BASIS,
+ * WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied.
+ * See the License for the specific language governing permissions and
+ * limitations under the License.
+ *
+ ******************************************************************************/
+
+package io.questdb.client.cutlass.auth;
+
+import io.questdb.client.std.QuietCloseable;
+import org.jetbrains.annotations.TestOnly;
+import org.slf4j.Logger;
+import org.slf4j.LoggerFactory;
+
+import java.util.ArrayList;
+import java.util.HashMap;
+import java.util.Iterator;
+import java.util.List;
+import java.util.Map;
+import java.util.ServiceConfigurationError;
+import java.util.ServiceLoader;
+import java.util.TreeSet;
+import java.util.concurrent.ScheduledFuture;
+import java.util.concurrent.ScheduledThreadPoolExecutor;
+import java.util.concurrent.TimeUnit;
+import java.util.concurrent.atomic.AtomicBoolean;
+
+/**
+ * Process-wide registry of the token providers that connect strings select with {@code token_provider}
+ * (design/qwp-token-provider-spec.md, section 7.4). Every client instance built from a connect string - each
+ * sender, each query client, each pooled connection - acquires a {@link Lease} when it connects and releases it
+ * when it closes. All leases for the same configuration - the provider name plus its normalized parameters -
+ * share one {@link RefreshingTokenProvider}, so the process runs one refresher, and makes at most one call to
+ * the identity platform at a time, per configuration.
+ *
+ * When the last lease is released the provider stays alive for the registry's linger period (60 s for the
+ * global registry) and is closed only if no new lease arrives in the meantime: a pool that recycles its
+ * connections, or an application that closes one sender and opens another, keeps the warm cache.
+ *
+ * Providers that an application supplies itself never pass through here; sharing those is the application's
+ * responsibility.
+ */
+public final class TokenProviderRegistry {
+ public static final long DEFAULT_LINGER_MILLIS = 60_000;
+ private static final TokenProviderRegistry GLOBAL = new TokenProviderRegistry(DEFAULT_LINGER_MILLIS);
+ private static final Logger LOG = LoggerFactory.getLogger(TokenProviderRegistry.class);
+ @TestOnly
+ private static volatile List factoriesForTesting;
+ private final Map entries = new HashMap<>();
+ private final long lingerMillis;
+ private final Object lock = new Object();
+ private ScheduledThreadPoolExecutor lingerTimer;
+
+ /**
+ * Creates a private registry. Production code uses {@link #global()}; a separate instance is for tests and
+ * for applications that want to scope provider sharing themselves.
+ *
+ * @param lingerMillis how long a provider outlives its last lease; {@code <= 0} closes it at once
+ */
+ public TokenProviderRegistry(long lingerMillis) {
+ this.lingerMillis = lingerMillis;
+ }
+
+ /**
+ * Finds the factory serving {@code name} through {@link ServiceLoader}, trying the thread context class
+ * loader first and then the loader of this library.
+ *
+ * @return the factory, or null when none is installed
+ */
+ public static TokenProviderFactory findFactory(String name) {
+ for (TokenProviderFactory factory : factories()) {
+ if (name.equals(factory.name())) {
+ return factory;
+ }
+ }
+ return null;
+ }
+
+ /**
+ * The registry every client built from a connect string uses.
+ */
+ public static TokenProviderRegistry global() {
+ return GLOBAL;
+ }
+
+ /**
+ * Test seam: replaces {@link ServiceLoader} discovery with a fixed list of factories. Pass null to restore
+ * discovery.
+ */
+ @TestOnly
+ public static void setFactoriesForTesting(List factories) {
+ factoriesForTesting = factories;
+ }
+
+ /**
+ * The names of every installed factory, sorted: the values {@code token_provider} accepts.
+ */
+ public static List supportedProviders() {
+ TreeSet names = new TreeSet<>();
+ for (TokenProviderFactory factory : factories()) {
+ names.add(factory.name());
+ }
+ return new ArrayList<>(names);
+ }
+
+ private static List factories() {
+ final List override = factoriesForTesting;
+ if (override != null) {
+ return override;
+ }
+ final List found = new ArrayList<>();
+ final ClassLoader context = Thread.currentThread().getContextClassLoader();
+ load(context == null ? ServiceLoader.load(TokenProviderFactory.class)
+ : ServiceLoader.load(TokenProviderFactory.class, context), found);
+ final ClassLoader own = TokenProviderFactory.class.getClassLoader();
+ if (own != context) {
+ load(ServiceLoader.load(TokenProviderFactory.class, own), found);
+ }
+ return found;
+ }
+
+ private static void load(ServiceLoader loader, List into) {
+ final Iterator it = loader.iterator();
+ while (true) {
+ final TokenProviderFactory factory;
+ try {
+ if (!it.hasNext()) {
+ return;
+ }
+ factory = it.next();
+ } catch (ServiceConfigurationError e) {
+ // one broken provider must not hide the others
+ LOG.warn("could not load a token provider factory: {}",
+ CredentialRedaction.sanitizeErrorText(String.valueOf(e.getMessage())));
+ continue;
+ }
+ boolean duplicate = false;
+ for (int i = 0, n = into.size(); i < n; i++) {
+ if (into.get(i).getClass() == factory.getClass()) {
+ duplicate = true;
+ break;
+ }
+ }
+ if (!duplicate) {
+ into.add(factory);
+ }
+ }
+ }
+
+ /**
+ * Acquires a lease on the shared provider for {@code spec}, creating the provider - and starting its first
+ * fetch - when none is alive. Cancels a pending linger close.
+ *
+ * @throws IllegalArgumentException when the factory rejects the parameters
+ */
+ public Lease acquire(TokenProviderSpec spec) {
+ synchronized (lock) {
+ Entry entry = entries.get(spec.registryKey());
+ if (entry == null || entry.provider.isClosed()) {
+ TokenSource source = spec.factory().createSource(spec.params());
+ if (source == null) {
+ throw new IllegalStateException("token provider factory " + spec.name() + " returned no token source");
+ }
+ RefreshingTokenProvider provider = RefreshingTokenProvider.builder(source).name(spec.describe()).build();
+ entry = new Entry(spec.registryKey(), provider);
+ entries.put(entry.key, entry);
+ LOG.info("started token provider {}", entry.provider.getName());
+ }
+ entry.refs++;
+ if (entry.lingerTask != null) {
+ entry.lingerTask.cancel(false);
+ entry.lingerTask = null;
+ }
+ return new Lease(entry);
+ }
+ }
+
+ /**
+ * Whether a provider for {@code spec} is alive - leased, or lingering after its last lease.
+ */
+ public boolean isActive(TokenProviderSpec spec) {
+ synchronized (lock) {
+ Entry entry = entries.get(spec.registryKey());
+ return entry != null && !entry.provider.isClosed();
+ }
+ }
+
+ /**
+ * Number of leases currently held on the provider for {@code spec}.
+ */
+ public int leaseCount(TokenProviderSpec spec) {
+ synchronized (lock) {
+ Entry entry = entries.get(spec.registryKey());
+ return entry == null ? 0 : entry.refs;
+ }
+ }
+
+ private void expire(Entry entry) {
+ synchronized (lock) {
+ if (entry.refs > 0 || entries.get(entry.key) != entry) {
+ return; // re-leased during the linger, or already replaced
+ }
+ entries.remove(entry.key);
+ entry.lingerTask = null;
+ }
+ LOG.info("closing token provider {}: no client has used it for {} ms", entry.provider.getName(), lingerMillis);
+ entry.provider.close();
+ }
+
+ private ScheduledThreadPoolExecutor lingerTimer() {
+ if (lingerTimer == null) {
+ lingerTimer = new ScheduledThreadPoolExecutor(1, r -> {
+ Thread t = new Thread(r, "qdb-token-provider-registry");
+ t.setDaemon(true);
+ return t;
+ });
+ lingerTimer.setRemoveOnCancelPolicy(true);
+ lingerTimer.setKeepAliveTime(5, TimeUnit.SECONDS);
+ lingerTimer.allowCoreThreadTimeOut(true);
+ }
+ return lingerTimer;
+ }
+
+ private void release(Entry entry) {
+ RefreshingTokenProvider toClose = null;
+ synchronized (lock) {
+ if (--entry.refs > 0) {
+ return;
+ }
+ if (lingerMillis <= 0) {
+ entries.remove(entry.key, entry);
+ toClose = entry.provider;
+ } else {
+ entry.lingerTask = lingerTimer().schedule(() -> expire(entry), lingerMillis, TimeUnit.MILLISECONDS);
+ }
+ }
+ if (toClose != null) {
+ toClose.close();
+ }
+ }
+
+ private static final class Entry {
+ final String key;
+ final RefreshingTokenProvider provider;
+ ScheduledFuture> lingerTask;
+ int refs;
+
+ Entry(String key, RefreshingTokenProvider provider) {
+ this.key = key;
+ this.provider = provider;
+ }
+ }
+
+ /**
+ * One client's hold on a shared provider. Release it exactly when the client closes; releasing twice is
+ * harmless.
+ */
+ public final class Lease implements QuietCloseable {
+ private final Entry entry;
+ private final AtomicBoolean released = new AtomicBoolean();
+
+ private Lease(Entry entry) {
+ this.entry = entry;
+ }
+
+ @Override
+ public void close() {
+ if (released.compareAndSet(false, true)) {
+ release(entry);
+ }
+ }
+
+ /**
+ * The shared provider. Do not close it: release the lease instead.
+ */
+ public RefreshingTokenProvider provider() {
+ return entry.provider;
+ }
+
+ @Override
+ public String toString() {
+ return "TokenProviderRegistry.Lease{" + entry.provider.getName() + (released.get() ? ", released}" : "}");
+ }
+ }
+}
diff --git a/core/src/main/java/io/questdb/client/cutlass/auth/TokenProviderSpec.java b/core/src/main/java/io/questdb/client/cutlass/auth/TokenProviderSpec.java
new file mode 100644
index 000000000..82e09019e
--- /dev/null
+++ b/core/src/main/java/io/questdb/client/cutlass/auth/TokenProviderSpec.java
@@ -0,0 +1,248 @@
+/*+*****************************************************************************
+ * ___ _ ____ ____
+ * / _ \ _ _ ___ ___| |_| _ \| __ )
+ * | | | | | | |/ _ \/ __| __| | | | _ \
+ * | |_| | |_| | __/\__ \ |_| |_| | |_) |
+ * \__\_\\__,_|\___||___/\__|____/|____/
+ *
+ * Copyright (c) 2014-2019 Appsicle
+ * Copyright (c) 2019-2026 QuestDB
+ *
+ * Licensed under the Apache License, Version 2.0 (the "License");
+ * you may not use this file except in compliance with the License.
+ * You may obtain a copy of the License at
+ *
+ * http://www.apache.org/licenses/LICENSE-2.0
+ *
+ * Unless required by applicable law or agreed to in writing, software
+ * distributed under the License is distributed on an "AS IS" BASIS,
+ * WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied.
+ * See the License for the specific language governing permissions and
+ * limitations under the License.
+ *
+ ******************************************************************************/
+
+package io.questdb.client.cutlass.auth;
+
+import io.questdb.client.impl.ConfigView;
+
+import java.util.Collections;
+import java.util.List;
+import java.util.Locale;
+import java.util.Map;
+import java.util.TreeMap;
+
+/**
+ * A validated {@code token_provider} selection from a {@code ws}/{@code wss} connect string: the provider name,
+ * its normalized parameters, and the factory that serves it (design/qwp-token-provider-spec.md, section 7).
+ *
+ * {@link #parse(ConfigView, boolean)} enforces every rule of section 7.2 and never fetches a token: it only
+ * looks the factory up. A client turns the spec into a provider when it connects, through
+ * {@link TokenProviderRegistry#acquire(TokenProviderSpec)}.
+ *
+ */
+public final class TokenProviderSpec {
+ /**
+ * The {@code token_provider} value of the Microsoft Entra ID provider, served by the optional
+ * {@code questdb-client-azure} module.
+ */
+ public static final String AZURE = "azure";
+ /**
+ * Reserved for a zero-dependency Azure IMDS provider (decision D5); not implemented.
+ */
+ public static final String AZURE_IMDS = "azure_imds";
+ public static final String AZURE_MODULE = "org.questdb:questdb-client-azure";
+ public static final String KEY_AZURE_CLIENT_ID = "azure_client_id";
+ public static final String KEY_AZURE_RESOURCE = "azure_resource";
+ public static final String KEY_TOKEN_PROVIDER = "token_provider";
+ private static final String AZURE_DEFAULT_SCOPE_SUFFIX = "/.default";
+ private static final String[] STATIC_CREDENTIAL_KEYS = {"token", "username", "password"};
+ private final TokenProviderFactory factory;
+ private final String name;
+ private final Map params;
+ private final String registryKey;
+
+ private TokenProviderSpec(TokenProviderFactory factory, String name, Map params) {
+ this.factory = factory;
+ this.name = name;
+ this.params = Collections.unmodifiableMap(params);
+ StringBuilder key = new StringBuilder(name);
+ for (Map.Entry e : params.entrySet()) {
+ key.append('|').append(e.getKey()).append('=').append(e.getValue());
+ }
+ this.registryKey = key.toString();
+ }
+
+ /**
+ * Validates and resolves the {@code token_provider} keys of a {@code ws}/{@code wss} connect string. Never
+ * fetches a token and never starts a provider. Rejects, naming the offending key:
+ *
+ *
{@code token_provider} combined with {@code token}, {@code username} or {@code password};
+ *
an empty, unknown, reserved or uninstalled {@code token_provider}, listing the supported values;
+ *
a provider-specific key the selected provider does not accept, such as {@code azure_resource}
+ * without {@code token_provider=azure};
+ *
a missing required provider key;
+ *
{@code token_provider} on {@code ws::}: the bearer token would cross the network in cleartext
+ * (decision D4).
+ *
+ * Combining {@code token_provider} with a provider the application supplies is rejected by the builders.
+ *
+ * @param view the parsed connect string
+ * @param tls true for the {@code wss} schema
+ * @return the spec, or null when the string selects no provider
+ * @throws IllegalArgumentException on any violation
+ */
+ public static TokenProviderSpec parse(ConfigView view, boolean tls) {
+ final String name = view.getStr(KEY_TOKEN_PROVIDER);
+ final String resource = view.getStr(KEY_AZURE_RESOURCE);
+ final String clientId = view.getStr(KEY_AZURE_CLIENT_ID);
+ if (name == null) {
+ if (resource != null) {
+ throw new IllegalArgumentException(KEY_AZURE_RESOURCE + " requires " + KEY_TOKEN_PROVIDER + '=' + AZURE);
+ }
+ if (clientId != null) {
+ throw new IllegalArgumentException(KEY_AZURE_CLIENT_ID + " requires " + KEY_TOKEN_PROVIDER + '=' + AZURE);
+ }
+ return null;
+ }
+ for (String key : STATIC_CREDENTIAL_KEYS) {
+ if (view.has(key)) {
+ throw new IllegalArgumentException(KEY_TOKEN_PROVIDER + " cannot be combined with " + key);
+ }
+ }
+ if (name.trim().isEmpty()) {
+ throw new IllegalArgumentException(KEY_TOKEN_PROVIDER + " must not be empty; " + supportedValues());
+ }
+ if (!tls) {
+ throw new IllegalArgumentException(KEY_TOKEN_PROVIDER + " requires the wss:: schema: over ws:: the bearer "
+ + "token would cross the network in cleartext");
+ }
+ if (AZURE_IMDS.equals(name)) {
+ throw new IllegalArgumentException(KEY_TOKEN_PROVIDER + '=' + AZURE_IMDS
+ + " is reserved and not implemented by this client; " + supportedValues());
+ }
+ final TokenProviderFactory factory = TokenProviderRegistry.findFactory(name);
+ if (factory == null) {
+ if (AZURE.equals(name)) {
+ throw new IllegalArgumentException(KEY_TOKEN_PROVIDER + '=' + AZURE + " requires the " + AZURE_MODULE
+ + " module on the class path or module path; " + supportedValues());
+ }
+ throw new IllegalArgumentException("unsupported " + KEY_TOKEN_PROVIDER + ": " + safe(name) + "; "
+ + supportedValues());
+ }
+ if (!AZURE.equals(name)) {
+ if (resource != null) {
+ throw new IllegalArgumentException(KEY_AZURE_RESOURCE + " is only valid with " + KEY_TOKEN_PROVIDER + '=' + AZURE);
+ }
+ if (clientId != null) {
+ throw new IllegalArgumentException(KEY_AZURE_CLIENT_ID + " is only valid with " + KEY_TOKEN_PROVIDER + '=' + AZURE);
+ }
+ }
+ final Map params = new TreeMap<>();
+ if (AZURE.equals(name)) {
+ if (resource == null) {
+ throw new IllegalArgumentException(KEY_TOKEN_PROVIDER + '=' + AZURE + " requires " + KEY_AZURE_RESOURCE);
+ }
+ params.put(KEY_AZURE_RESOURCE, normalizeAzureResource(resource));
+ if (clientId != null) {
+ params.put(KEY_AZURE_CLIENT_ID, normalizeGuid(clientId));
+ }
+ }
+ factory.validate(Collections.unmodifiableMap(params));
+ return new TokenProviderSpec(factory, name, params);
+ }
+
+ /**
+ * Lower-cases a GUID after checking its {@code 8-4-4-4-12} hex shape.
+ */
+ static String normalizeGuid(String value) {
+ if (value.length() != 36) {
+ throw invalidClientId(value);
+ }
+ for (int i = 0; i < 36; i++) {
+ char c = value.charAt(i);
+ boolean dash = i == 8 || i == 13 || i == 18 || i == 23;
+ if (dash ? c != '-' : Character.digit(c, 16) < 0) {
+ throw invalidClientId(value);
+ }
+ }
+ return value.toLowerCase(Locale.ROOT);
+ }
+
+ private static IllegalArgumentException invalidClientId(String value) {
+ return new IllegalArgumentException("invalid " + KEY_AZURE_CLIENT_ID
+ + ": expected a GUID such as 00000000-0000-0000-0000-000000000000, got " + safe(value));
+ }
+
+ // The application ID URI (api://) or the client ID of the QuestDB app registration. A trailing
+ // "/.default" is stripped: the client always requests /.default.
+ private static String normalizeAzureResource(String value) {
+ String resource = value;
+ if (resource.endsWith(AZURE_DEFAULT_SCOPE_SUFFIX)) {
+ resource = resource.substring(0, resource.length() - AZURE_DEFAULT_SCOPE_SUFFIX.length());
+ }
+ if (resource.isEmpty()) {
+ throw new IllegalArgumentException(KEY_AZURE_RESOURCE + " must not be empty");
+ }
+ for (int i = 0, n = resource.length(); i < n; i++) {
+ char c = resource.charAt(i);
+ if (c <= 0x20 || c > 0x7e) {
+ throw new IllegalArgumentException(KEY_AZURE_RESOURCE
+ + " must be an application ID URI (api://) or a client ID without spaces or "
+ + "non-ASCII characters, got " + safe(resource));
+ }
+ }
+ return resource;
+ }
+
+ private static String safe(String value) {
+ return CredentialRedaction.sanitizeErrorText(value);
+ }
+
+ private static String supportedValues() {
+ List names = TokenProviderRegistry.supportedProviders();
+ if (names.isEmpty()) {
+ return "supported values: none installed (token_provider=azure needs " + AZURE_MODULE + ')';
+ }
+ return "supported values: " + names;
+ }
+
+ /**
+ * A non-secret label for log lines and errors.
+ */
+ public String describe() {
+ return factory.describe(params);
+ }
+
+ public TokenProviderFactory factory() {
+ return factory;
+ }
+
+ /**
+ * The provider name, the value of {@code token_provider}.
+ */
+ public String name() {
+ return name;
+ }
+
+ /**
+ * The normalized provider-specific keys; unmodifiable.
+ */
+ public Map params() {
+ return params;
+ }
+
+ /**
+ * The registry key: the provider name plus the normalized parameters. Equivalent connect strings share it.
+ */
+ public String registryKey() {
+ return registryKey;
+ }
+
+ @Override
+ public String toString() {
+ return "TokenProviderSpec{" + registryKey + '}';
+ }
+}
diff --git a/core/src/main/java/io/questdb/client/cutlass/auth/TokenSource.java b/core/src/main/java/io/questdb/client/cutlass/auth/TokenSource.java
new file mode 100644
index 000000000..31575460d
--- /dev/null
+++ b/core/src/main/java/io/questdb/client/cutlass/auth/TokenSource.java
@@ -0,0 +1,63 @@
+/*+*****************************************************************************
+ * ___ _ ____ ____
+ * / _ \ _ _ ___ ___| |_| _ \| __ )
+ * | | | | | | |/ _ \/ __| __| | | | _ \
+ * | |_| | |_| | __/\__ \ |_| |_| | |_) |
+ * \__\_\\__,_|\___||___/\__|____/|____/
+ *
+ * Copyright (c) 2014-2019 Appsicle
+ * Copyright (c) 2019-2026 QuestDB
+ *
+ * Licensed under the Apache License, Version 2.0 (the "License");
+ * you may not use this file except in compliance with the License.
+ * You may obtain a copy of the License at
+ *
+ * http://www.apache.org/licenses/LICENSE-2.0
+ *
+ * Unless required by applicable law or agreed to in writing, software
+ * distributed under the License is distributed on an "AS IS" BASIS,
+ * WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied.
+ * See the License for the specific language governing permissions and
+ * limitations under the License.
+ *
+ ******************************************************************************/
+
+package io.questdb.client.cutlass.auth;
+
+/**
+ * Obtains a new bearer token from an identity platform - Microsoft Entra ID, an OAuth client-credentials
+ * endpoint, a secrets manager - together with its expiry. A {@link RefreshingTokenProvider} wraps a source and
+ * caches its tokens, refreshing them in the background before they expire. This is the token-source contract of
+ * the dynamic-credential specification (design/qwp-token-provider-spec.md, section 4).
+ *
+ * Threading. The provider calls {@link #fetchToken()} only from its own background refresher thread, never from
+ * a connection thread, and never concurrently with itself. The call may block, but it should bound its own
+ * network operations; 30 seconds per attempt is recommended. A provider being closed interrupts its refresher,
+ * so a source blocked in an interruptible wait should let the interrupt end it.
+ *
+ * Failures. Throw {@link TokenUnavailableException} to classify a failure:
+ *
+ *
retryable - network failures, timeouts, HTTP 429, HTTP 5xx, and IMDS 404 or 410. Pass the platform's
+ * {@code Retry-After}, when it gave one, as {@code retryAfterMillis};
+ *
permanent - configuration that is missing or wrong: no credential configured, identity not found,
+ * invalid client.
+ *
+ * When unsure, report the failure as retryable. Any other exception is treated as retryable. The provider keeps
+ * retrying either way - an operator can repair a "permanent" failure, such as a missing role assignment,
+ * without restarting the application - but the classification decides how a caller waiting for a token reacts.
+ *
+ * Secrets. A failure message must not contain a token or a raw response body: report only the shape of a
+ * problem, such as a missing or invalid field. Never attach an object that serializes a raw HTTP request, such
+ * as a client-credentials body carrying a {@code client_secret}, to an exception.
+ */
+@FunctionalInterface
+public interface TokenSource {
+
+ /**
+ * Obtains a new token from the identity platform.
+ *
+ * @return the token and its expiry; never null
+ * @throws TokenUnavailableException to report a classified failure; any other exception counts as retryable
+ */
+ ExpiringToken fetchToken();
+}
diff --git a/core/src/main/java/io/questdb/client/cutlass/auth/TokenUnavailableException.java b/core/src/main/java/io/questdb/client/cutlass/auth/TokenUnavailableException.java
new file mode 100644
index 000000000..6e08754b8
--- /dev/null
+++ b/core/src/main/java/io/questdb/client/cutlass/auth/TokenUnavailableException.java
@@ -0,0 +1,128 @@
+/*+*****************************************************************************
+ * ___ _ ____ ____
+ * / _ \ _ _ ___ ___| |_| _ \| __ )
+ * | | | | | | |/ _ \/ __| __| | | | _ \
+ * | |_| | |_| | __/\__ \ |_| |_| | |_) |
+ * \__\_\\__,_|\___||___/\__|____/|____/
+ *
+ * Copyright (c) 2014-2019 Appsicle
+ * Copyright (c) 2019-2026 QuestDB
+ *
+ * Licensed under the Apache License, Version 2.0 (the "License");
+ * you may not use this file except in compliance with the License.
+ * You may obtain a copy of the License at
+ *
+ * http://www.apache.org/licenses/LICENSE-2.0
+ *
+ * Unless required by applicable law or agreed to in writing, software
+ * distributed under the License is distributed on an "AS IS" BASIS,
+ * WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied.
+ * See the License for the specific language governing permissions and
+ * limitations under the License.
+ *
+ ******************************************************************************/
+
+package io.questdb.client.cutlass.auth;
+
+import io.questdb.client.cutlass.line.LineSenderException;
+
+/**
+ * No token could be obtained. This is the {@code TokenError} of the dynamic-credential specification
+ * (design/qwp-token-provider-spec.md, section 4) and also its {@code token-unavailable} error.
+ *
+ *
A {@link TokenSource} throws it to classify a failed fetch.
+ *
A {@link RefreshingTokenProvider} throws it from {@code getToken()} when it holds no usable token and
+ * cannot get one in time. It then carries the classification of the provider's most recent failed fetch, or
+ * is retryable when no fetch has failed (the wait timed out, or the caller was interrupted).
+ *
+ * {@link #isRetryable()} tells a caller whether trying again can succeed. A permanent failure is configuration
+ * that is missing or wrong; a retryable one is a network failure, a timeout, throttling or a server error.
+ * {@link #getRetryAfterMillis()} is the platform's suggested wait before the next attempt, or
+ * {@link #NO_RETRY_AFTER}.
+ *
+ * Connection code treats an exception from an application-supplied token provider by type: only a
+ * {@code TokenUnavailableException} marked retryable is retried while a {@code Sender} with
+ * {@code initial_connect_retry=on} starts up; anything else fails startup fast (decision D8).
+ *
+ * The message must never contain a token or a raw response body.
+ */
+public class TokenUnavailableException extends LineSenderException {
+ /**
+ * Value of {@link #getRetryAfterMillis()} meaning the platform suggested no wait.
+ */
+ public static final long NO_RETRY_AFTER = -1;
+ private final long retryAfterMillis;
+ private final boolean retryable;
+
+ /**
+ * @param message a description of the failure, never containing a token or a raw response body
+ * @param retryable whether retrying can succeed
+ */
+ public TokenUnavailableException(CharSequence message, boolean retryable) {
+ this(message, retryable, NO_RETRY_AFTER);
+ }
+
+ /**
+ * @param message a description of the failure, never containing a token or a raw response body
+ * @param retryable whether retrying can succeed
+ * @param retryAfterMillis the platform's suggested wait before retrying, or {@link #NO_RETRY_AFTER}
+ */
+ public TokenUnavailableException(CharSequence message, boolean retryable, long retryAfterMillis) {
+ super(message, retryable);
+ this.retryable = retryable;
+ this.retryAfterMillis = retryAfterMillis < 0 ? NO_RETRY_AFTER : retryAfterMillis;
+ }
+
+ /**
+ * Carries a cause. Attach only a cause that is safe to log: never one whose message holds a token or a raw
+ * response body, and never an object that serializes a raw HTTP request.
+ *
+ * @param message a description of the failure, never containing a token or a raw response body
+ * @param retryable whether retrying can succeed
+ * @param retryAfterMillis the platform's suggested wait before retrying, or {@link #NO_RETRY_AFTER}
+ * @param cause the underlying failure, may be null
+ */
+ public TokenUnavailableException(String message, boolean retryable, long retryAfterMillis, Throwable cause) {
+ super(message, cause);
+ this.retryable = retryable;
+ this.retryAfterMillis = retryAfterMillis < 0 ? NO_RETRY_AFTER : retryAfterMillis;
+ }
+
+ /**
+ * A permanent failure: configuration that is missing or wrong.
+ */
+ public static TokenUnavailableException permanent(CharSequence message) {
+ return new TokenUnavailableException(message, false);
+ }
+
+ /**
+ * A retryable failure without a suggested wait.
+ */
+ public static TokenUnavailableException retryable(CharSequence message) {
+ return new TokenUnavailableException(message, true);
+ }
+
+ /**
+ * A retryable failure with the platform's suggested wait, such as an HTTP {@code Retry-After}.
+ */
+ public static TokenUnavailableException retryable(CharSequence message, long retryAfterMillis) {
+ return new TokenUnavailableException(message, true, retryAfterMillis);
+ }
+
+ /**
+ * @return the platform's suggested wait before the next attempt, in milliseconds, or
+ * {@link #NO_RETRY_AFTER} when it gave none
+ */
+ public long getRetryAfterMillis() {
+ return retryAfterMillis;
+ }
+
+ /**
+ * @return true when retrying can succeed (a network failure, timeout, throttling or server error), false for
+ * configuration that is missing or wrong
+ */
+ @Override
+ public boolean isRetryable() {
+ return retryable;
+ }
+}
diff --git a/core/src/main/java/io/questdb/client/cutlass/http/BearerChallenge.java b/core/src/main/java/io/questdb/client/cutlass/http/BearerChallenge.java
new file mode 100644
index 000000000..914f97f00
--- /dev/null
+++ b/core/src/main/java/io/questdb/client/cutlass/http/BearerChallenge.java
@@ -0,0 +1,158 @@
+/*+*****************************************************************************
+ * ___ _ ____ ____
+ * / _ \ _ _ ___ ___| |_| _ \| __ )
+ * | | | | | | |/ _ \/ __| __| | | | _ \
+ * | |_| | |_| | __/\__ \ |_| |_| | |_) |
+ * \__\_\\__,_|\___||___/\__|____/|____/
+ *
+ * Copyright (c) 2014-2019 Appsicle
+ * Copyright (c) 2019-2026 QuestDB
+ *
+ * Licensed under the Apache License, Version 2.0 (the "License");
+ * you may not use this file except in compliance with the License.
+ * You may obtain a copy of the License at
+ *
+ * http://www.apache.org/licenses/LICENSE-2.0
+ *
+ * Unless required by applicable law or agreed to in writing, software
+ * distributed under the License is distributed on an "AS IS" BASIS,
+ * WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied.
+ * See the License for the specific language governing permissions and
+ * limitations under the License.
+ *
+ ******************************************************************************/
+
+package io.questdb.client.cutlass.http;
+
+/**
+ * Reads the {@code error} parameter of a {@code Bearer} challenge in a {@code WWW-Authenticate} header
+ * (RFC 6750, section 3; RFC 7235, section 4.1). Clients use it to decide whether a {@code 401} is worth a token
+ * refresh: {@code invalid_token} is, {@code insufficient_scope} is not.
+ *
+ * The parser is tolerant. It walks a comma-separated list of challenges - {@code scheme [token68 | auth-param,
+ * ...]} - accepts quoted and unquoted parameter values, and ignores what it cannot read. Only the first
+ * {@code error} parameter of a {@code Bearer} challenge counts.
+ */
+public final class BearerChallenge {
+ /**
+ * The {@code error} value that means the presented token itself is bad, so a new one can help.
+ */
+ public static final String INVALID_TOKEN = "invalid_token";
+
+ private BearerChallenge() {
+ }
+
+ /**
+ * @param wwwAuthenticate the value of one or more {@code WWW-Authenticate} headers (joined with commas), or
+ * null
+ * @return the {@code error} parameter of the first {@code Bearer} challenge that carries one, or null when
+ * there is no such challenge
+ */
+ public static String bearerError(CharSequence wwwAuthenticate) {
+ if (wwwAuthenticate == null) {
+ return null;
+ }
+ final int n = wwwAuthenticate.length();
+ boolean inBearer = false;
+ int i = 0;
+ while (i < n) {
+ char c = wwwAuthenticate.charAt(i);
+ if (c == ',' || isWhitespace(c)) {
+ i++;
+ continue;
+ }
+ final int start = i;
+ while (i < n && isTokenChar(wwwAuthenticate.charAt(i))) {
+ i++;
+ }
+ if (i == start) {
+ i++; // not a token character: skip it
+ continue;
+ }
+ final String token = wwwAuthenticate.subSequence(start, i).toString();
+ int j = skipWhitespace(wwwAuthenticate, i);
+ if (j < n && wwwAuthenticate.charAt(j) == '=') {
+ // auth-param: name = ( token / quoted-string )
+ j = skipWhitespace(wwwAuthenticate, j + 1);
+ final String value;
+ if (j < n && wwwAuthenticate.charAt(j) == '"') {
+ final StringBuilder sb = new StringBuilder();
+ j++;
+ while (j < n) {
+ char q = wwwAuthenticate.charAt(j++);
+ if (q == '\\' && j < n) {
+ sb.append(wwwAuthenticate.charAt(j++));
+ } else if (q == '"') {
+ break;
+ } else {
+ sb.append(q);
+ }
+ }
+ value = sb.toString();
+ } else {
+ final int valueStart = j;
+ while (j < n && wwwAuthenticate.charAt(j) != ',' && !isWhitespace(wwwAuthenticate.charAt(j))) {
+ j++;
+ }
+ value = wwwAuthenticate.subSequence(valueStart, j).toString();
+ }
+ if (inBearer && "error".equalsIgnoreCase(token)) {
+ return value;
+ }
+ i = j;
+ } else {
+ // a new challenge's auth-scheme (or a token68, which carries no parameters)
+ inBearer = "Bearer".equalsIgnoreCase(token);
+ i = j;
+ }
+ }
+ return null;
+ }
+
+ /**
+ * Whether a {@code 401} carrying {@code wwwAuthenticate} warrants refreshing the token: true when the header
+ * has no {@code Bearer} challenge with an {@code error} parameter (the status code alone decides), or when
+ * that error is {@code invalid_token}.
+ */
+ public static boolean allowsTokenRefresh(CharSequence wwwAuthenticate) {
+ final String error = bearerError(wwwAuthenticate);
+ return error == null || INVALID_TOKEN.equals(error);
+ }
+
+ private static boolean isTokenChar(char c) {
+ if ((c >= 'a' && c <= 'z') || (c >= 'A' && c <= 'Z') || (c >= '0' && c <= '9')) {
+ return true;
+ }
+ switch (c) {
+ case '!':
+ case '#':
+ case '$':
+ case '%':
+ case '&':
+ case '\'':
+ case '*':
+ case '+':
+ case '-':
+ case '.':
+ case '^':
+ case '_':
+ case '`':
+ case '|':
+ case '~':
+ return true;
+ default:
+ return false;
+ }
+ }
+
+ private static boolean isWhitespace(char c) {
+ return c == ' ' || c == '\t';
+ }
+
+ private static int skipWhitespace(CharSequence s, int i) {
+ while (i < s.length() && isWhitespace(s.charAt(i))) {
+ i++;
+ }
+ return i;
+ }
+}
diff --git a/core/src/main/java/io/questdb/client/cutlass/http/client/WebSocketClient.java b/core/src/main/java/io/questdb/client/cutlass/http/client/WebSocketClient.java
index 1290fcab5..8c8d07bd8 100644
--- a/core/src/main/java/io/questdb/client/cutlass/http/client/WebSocketClient.java
+++ b/core/src/main/java/io/questdb/client/cutlass/http/client/WebSocketClient.java
@@ -163,6 +163,11 @@ public abstract class WebSocketClient implements QuietCloseable {
private int serverNegotiatedZstdLevel;
private int serverQwpVersion = 1;
private String upgradeRejectRole;
+ // WWW-Authenticate value(s) from the most recent rejected upgrade, joined with ", " when the server sent
+ // several, or null when absent. Lets a 401 carrying "Bearer error=insufficient_scope" be told apart from an
+ // expired token, so a dynamic credential is refreshed only when a new token can help. Reset to null on every
+ // upgrade() invocation.
+ private String upgradeRejectWwwAuthenticate;
// Server-advertised zone identifier from the most recent rejected upgrade,
// captured from the X-QuestDB-Zone response header on a 421. Null when the
// header was absent or empty. Per failover.md §5 servers SHOULD emit this
@@ -374,6 +379,14 @@ public String getUpgradeRejectRole() {
return upgradeRejectRole;
}
+ /**
+ * {@code WWW-Authenticate} value(s) on the most recent rejected upgrade, joined with {@code ", "} when the
+ * server sent several, or null when absent.
+ */
+ public String getUpgradeRejectWwwAuthenticate() {
+ return upgradeRejectWwwAuthenticate;
+ }
+
/**
* Zone identifier from {@code X-QuestDB-Zone} on the most recent rejected
* upgrade, or null when the header was absent or empty (after trimming).
@@ -648,6 +661,7 @@ public void upgrade(CharSequence path, int timeout, CharSequence authorizationHe
return; // Already upgraded
}
upgradeRejectRole = null;
+ upgradeRejectWwwAuthenticate = null;
upgradeRejectZone = null;
upgradeStatusCode = 0;
@@ -878,6 +892,34 @@ private static String extractRoleHeader(String response) {
return null;
}
+ // Every value of the named header, joined with ", " (RFC 7230 section 3.2.2), or null when absent.
+ private static String extractHeaderValues(String response, String headerName) {
+ int headerLen = headerName.length();
+ int responseLen = response.length();
+ StringBuilder values = null;
+ int lineStart = response.indexOf("\r\n");
+ while (lineStart >= 0 && lineStart + 2 + headerLen <= responseLen) {
+ int hStart = lineStart + 2;
+ if (response.regionMatches(true, hStart, headerName, 0, headerLen)) {
+ int valueStart = hStart + headerLen;
+ int lineEnd = response.indexOf('\r', valueStart);
+ if (lineEnd < 0) {
+ lineEnd = responseLen;
+ }
+ String value = response.substring(valueStart, lineEnd).trim();
+ if (!value.isEmpty()) {
+ if (values == null) {
+ values = new StringBuilder(value);
+ } else {
+ values.append(", ").append(value);
+ }
+ }
+ }
+ lineStart = response.indexOf("\r\n", hStart);
+ }
+ return values == null ? null : values.toString();
+ }
+
private static String extractZoneHeader(String response) {
int headerLen = QUESTDB_ZONE_HEADER_NAME.length();
int responseLen = response.length();
@@ -1346,6 +1388,9 @@ private void validateUpgradeResponse(int headerEnd) {
if (!response.startsWith("HTTP/1.1 101")) {
String statusLine = response.split("\r\n")[0];
upgradeStatusCode = parseStatusCode(statusLine);
+ if (upgradeStatusCode == 401) {
+ upgradeRejectWwwAuthenticate = extractHeaderValues(response, "WWW-Authenticate:");
+ }
if (upgradeStatusCode == 421) {
upgradeRejectRole = extractRoleHeader(response);
upgradeRejectZone = extractZoneHeader(response);
diff --git a/core/src/main/java/io/questdb/client/cutlass/qwp/client/QwpAuthFailedException.java b/core/src/main/java/io/questdb/client/cutlass/qwp/client/QwpAuthFailedException.java
index d2a5714a2..e14c3a7da 100644
--- a/core/src/main/java/io/questdb/client/cutlass/qwp/client/QwpAuthFailedException.java
+++ b/core/src/main/java/io/questdb/client/cutlass/qwp/client/QwpAuthFailedException.java
@@ -24,26 +24,61 @@
package io.questdb.client.cutlass.qwp.client;
+import io.questdb.client.cutlass.auth.CredentialRedaction;
+import io.questdb.client.cutlass.http.BearerChallenge;
import io.questdb.client.cutlass.http.client.HttpClientException;
/**
- * WebSocket upgrade rejected with {@code 401} or {@code 403}. Terminal across all
- * configured endpoints: a rejected credential is uniformly rejected across the
- * cluster, so failing fast surfaces the configuration error immediately. Path
- * mismatches ({@code 404}) are NOT routed through this exception because a single
- * misconfigured node mid-deploy can return 404 while peers are healthy.
+ * WebSocket upgrade rejected with {@code 401} or {@code 403}: the {@code auth-rejected} failure class of the
+ * dynamic-credential specification (design/qwp-token-provider-spec.md, section 8.1). The message starts with
+ * {@code auth-rejected} so every report of it names the class.
+ *
+ * A credential is uniformly accepted or rejected across a cluster, so the connect walk stops at the first such
+ * rejection instead of trying the remaining endpoints. What happens next depends on the phase (section 8.3):
+ * startup fails, an established store-and-forward sender keeps retrying, an orphan drainer rides a rotating
+ * credential out before quarantining, and a query fails. Before any of that, a {@code 401} against a dynamic
+ * credential earns one same-endpoint retry with a refreshed token (section 8.2) - see
+ * {@link #isTokenRefreshable()}. Path mismatches ({@code 404}) are NOT routed through this exception because a
+ * single misconfigured node mid-deploy can return 404 while peers are healthy.
*/
public final class QwpAuthFailedException extends HttpClientException {
+ private static final int MAX_BEARER_ERROR_LENGTH = 64;
+ private final String bearerError;
private final String host;
private final int port;
private final int statusCode;
public QwpAuthFailedException(int statusCode, String host, int port) {
- super("WebSocket upgrade rejected with HTTP ");
- put(statusCode).put(" for ").put(host).put(':').put(port);
+ this(statusCode, host, port, null);
+ }
+
+ /**
+ * @param wwwAuthenticate the {@code WWW-Authenticate} header value(s) of the rejection, or null
+ */
+ public QwpAuthFailedException(int statusCode, String host, int port, String wwwAuthenticate) {
+ super("auth-rejected: WebSocket upgrade rejected with HTTP ");
+ put(statusCode);
+ String error = BearerChallenge.bearerError(wwwAuthenticate);
+ if (error != null) {
+ error = CredentialRedaction.sanitizeErrorText(error);
+ if (error.length() > MAX_BEARER_ERROR_LENGTH) {
+ error = error.substring(0, MAX_BEARER_ERROR_LENGTH);
+ }
+ put(" [error=").put(error).put(']');
+ }
+ put(" for ").put(host).put(':').put(port);
this.statusCode = statusCode;
this.host = host;
this.port = port;
+ this.bearerError = error;
+ }
+
+ /**
+ * The {@code error} parameter of the rejection's {@code WWW-Authenticate: Bearer} challenge (for example
+ * {@code invalid_token} or {@code insufficient_scope}), sanitized, or null when the server sent none.
+ */
+ public String getBearerError() {
+ return bearerError;
}
public String getHost() {
@@ -57,4 +92,13 @@ public int getPort() {
public int getStatusCode() {
return statusCode;
}
+
+ /**
+ * Whether a new token can help: the status is {@code 401}, and the server either sent no {@code Bearer}
+ * challenge with an {@code error} or sent {@code error="invalid_token"} (RFC 6750, section 3.1). A
+ * {@code 403} never qualifies - it is an authorization decision a new token does not change (decision D2).
+ */
+ public boolean isTokenRefreshable() {
+ return statusCode == 401 && (bearerError == null || BearerChallenge.INVALID_TOKEN.equals(bearerError));
+ }
}
diff --git a/core/src/main/java/io/questdb/client/cutlass/qwp/client/QwpConnectionHealthTracker.java b/core/src/main/java/io/questdb/client/cutlass/qwp/client/QwpConnectionHealthTracker.java
new file mode 100644
index 000000000..a26b681b4
--- /dev/null
+++ b/core/src/main/java/io/questdb/client/cutlass/qwp/client/QwpConnectionHealthTracker.java
@@ -0,0 +1,182 @@
+/*+*****************************************************************************
+ * ___ _ ____ ____
+ * / _ \ _ _ ___ ___| |_| _ \| __ )
+ * | | | | | | |/ _ \/ __| __| | | | _ \
+ * | |_| | |_| | __/\__ \ |_| |_| | |_) |
+ * \__\_\\__,_|\___||___/\__|____/|____/
+ *
+ * Copyright (c) 2014-2019 Appsicle
+ * Copyright (c) 2019-2026 QuestDB
+ *
+ * Licensed under the Apache License, Version 2.0 (the "License");
+ * you may not use this file except in compliance with the License.
+ * You may obtain a copy of the License at
+ *
+ * http://www.apache.org/licenses/LICENSE-2.0
+ *
+ * Unless required by applicable law or agreed to in writing, software
+ * distributed under the License is distributed on an "AS IS" BASIS,
+ * WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied.
+ * See the License for the specific language governing permissions and
+ * limitations under the License.
+ *
+ ******************************************************************************/
+
+package io.questdb.client.cutlass.qwp.client;
+
+import io.questdb.client.ConnectionHealth;
+import io.questdb.client.cutlass.auth.CredentialRedaction;
+import io.questdb.client.cutlass.http.client.HttpClientException;
+import io.questdb.client.cutlass.http.client.WebSocketUpgradeException;
+import io.questdb.client.cutlass.line.LineSenderException;
+
+/**
+ * Tracks one QWP client's connection health (design/qwp-token-provider-spec.md, section 8.4) and publishes it as
+ * an immutable {@link ConnectionHealth} snapshot. Connection code reports transitions - a successful upgrade, a
+ * lost connection, a failed connect round, a terminal failure, close - from whichever thread observes them;
+ * readers get the last published snapshot with a single volatile read and never wait on connection threads.
+ * Writers serialize on a private monitor that is never held across I/O.
+ */
+public final class QwpConnectionHealthTracker {
+ private static final int MAX_CAUSE_DEPTH = 16;
+ private final Object lock = new Object();
+ private long failedRounds;
+ private ConnectionHealth.Failure lastFailure;
+ private long lastConnectedAt = ConnectionHealth.NONE;
+ private long outageSince;
+ private volatile ConnectionHealth snapshot;
+ private ConnectionHealth.State state = ConnectionHealth.State.CONNECTING;
+
+ public QwpConnectionHealthTracker() {
+ // A client that has never connected has been without a connection since it was created.
+ this.outageSince = System.currentTimeMillis();
+ publish();
+ }
+
+ /**
+ * Classifies a failed connect round's exception into the failure classes of section 8.4.
+ */
+ public static ConnectionHealth.Failure classify(Throwable failure, long nowMillis) {
+ Throwable t = failure;
+ for (int depth = 0; t != null && depth < MAX_CAUSE_DEPTH; depth++, t = t.getCause()) {
+ if (t instanceof QwpCredentialUnavailableException) {
+ return failure(ConnectionHealth.FailureClass.CREDENTIAL_UNAVAILABLE, 0, t, nowMillis);
+ }
+ if (t instanceof QwpAuthFailedException) {
+ return failure(ConnectionHealth.FailureClass.AUTH_REJECTED,
+ ((QwpAuthFailedException) t).getStatusCode(), t, nowMillis);
+ }
+ if (t instanceof QwpRoleMismatchException || t instanceof QwpIngressRoleRejectedException) {
+ return failure(ConnectionHealth.FailureClass.ROLE_REJECTED, 0, t, nowMillis);
+ }
+ if (t instanceof WebSocketUpgradeException) {
+ WebSocketUpgradeException e = (WebSocketUpgradeException) t;
+ return e.isRoleMismatch()
+ ? failure(ConnectionHealth.FailureClass.ROLE_REJECTED, 0, t, nowMillis)
+ : failure(ConnectionHealth.FailureClass.OTHER, Math.max(0, e.getStatusCode()), t, nowMillis);
+ }
+ if (t instanceof QwpVersionMismatchException || t instanceof QwpDurableAckMismatchException) {
+ return failure(ConnectionHealth.FailureClass.OTHER, 0, t, nowMillis);
+ }
+ }
+ return failure(failure instanceof HttpClientException || failure instanceof LineSenderException
+ ? ConnectionHealth.FailureClass.TRANSPORT
+ : ConnectionHealth.FailureClass.OTHER, 0, failure, nowMillis);
+ }
+
+ private static ConnectionHealth.Failure failure(
+ ConnectionHealth.FailureClass failureClass,
+ int statusCode,
+ Throwable t,
+ long nowMillis
+ ) {
+ String message = CredentialRedaction.sanitizeErrorText(t.getMessage());
+ return new ConnectionHealth.Failure(failureClass, statusCode,
+ message == null || message.isEmpty() ? t.getClass().getSimpleName() : message, nowMillis);
+ }
+
+ /**
+ * The client was closed. Sticky.
+ */
+ public void closed() {
+ synchronized (lock) {
+ state = ConnectionHealth.State.CLOSED;
+ publish();
+ }
+ }
+
+ /**
+ * An established connection was lost and is being re-established.
+ */
+ public void connectionLost() {
+ synchronized (lock) {
+ if (state == ConnectionHealth.State.CONNECTED) {
+ state = ConnectionHealth.State.RECONNECTING;
+ outageSince = System.currentTimeMillis();
+ failedRounds = 0;
+ publish();
+ }
+ }
+ }
+
+ /**
+ * The client will not connect again. Sticky; keeps the last connect-round failure.
+ */
+ public void failed() {
+ synchronized (lock) {
+ if (state != ConnectionHealth.State.CLOSED) {
+ state = ConnectionHealth.State.FAILED;
+ publish();
+ }
+ }
+ }
+
+ /**
+ * A connect round - one walk over the configured endpoints - ended without a connection.
+ */
+ public void roundFailed(Throwable failure) {
+ final long now = System.currentTimeMillis();
+ final ConnectionHealth.Failure classified = classify(failure, now);
+ synchronized (lock) {
+ if (state == ConnectionHealth.State.FAILED || state == ConnectionHealth.State.CLOSED) {
+ return;
+ }
+ if (state == ConnectionHealth.State.CONNECTED) {
+ state = ConnectionHealth.State.RECONNECTING;
+ outageSince = now;
+ failedRounds = 0;
+ }
+ failedRounds++;
+ lastFailure = classified;
+ publish();
+ }
+ }
+
+ /**
+ * @return the current snapshot; never null
+ */
+ public ConnectionHealth snapshot() {
+ return snapshot;
+ }
+
+ /**
+ * A connect round ended with a successful upgrade.
+ */
+ public void upgraded() {
+ synchronized (lock) {
+ if (state == ConnectionHealth.State.FAILED || state == ConnectionHealth.State.CLOSED) {
+ return;
+ }
+ state = ConnectionHealth.State.CONNECTED;
+ lastConnectedAt = System.currentTimeMillis();
+ outageSince = ConnectionHealth.NONE;
+ failedRounds = 0;
+ publish();
+ }
+ }
+
+ // caller holds lock
+ private void publish() {
+ snapshot = new ConnectionHealth(state, lastConnectedAt, outageSince, failedRounds, lastFailure);
+ }
+}
diff --git a/core/src/main/java/io/questdb/client/cutlass/qwp/client/QwpCredentialUnavailableException.java b/core/src/main/java/io/questdb/client/cutlass/qwp/client/QwpCredentialUnavailableException.java
index 2652a7f79..963119527 100644
--- a/core/src/main/java/io/questdb/client/cutlass/qwp/client/QwpCredentialUnavailableException.java
+++ b/core/src/main/java/io/questdb/client/cutlass/qwp/client/QwpCredentialUnavailableException.java
@@ -24,23 +24,28 @@
package io.questdb.client.cutlass.qwp.client;
+import io.questdb.client.cutlass.auth.TokenUnavailableException;
import io.questdb.client.cutlass.line.LineSenderException;
/**
* Signals that the client could not OBTAIN an Authorization credential for a
* (re)connect handshake: the configured {@code httpTokenProvider} threw instead of
- * returning a token -- a failed silent refresh, or no sign-in yet.
+ * returning a token -- a failed silent refresh, or no sign-in yet. This is the
+ * {@code credential-unavailable} failure class of the dynamic-credential specification
+ * (design/qwp-token-provider-spec.md, section 8.1).
*
* Distinct from {@link QwpAuthFailedException}, which means the server rejected a
- * credential the client did present (a terminal auth failure). A credential the client
- * cannot ACQUIRE is instead handled by connection phase, exactly like a transport outage:
+ * credential the client did present. A credential the client cannot ACQUIRE is instead
+ * handled by connection phase, exactly like a transport outage:
* the RUNNING store-and-forward drainer retries it indefinitely with capped backoff under
* Invariant B -- the IdP becomes reachable again, or the user completes an interactive
* sign-in -- holding the un-acked rows in SF meanwhile, and NEVER bounds it by
* {@code reconnectMaxDurationMillis} nor latches a terminal (either would drop a producer
- * store-and-forward promised to keep alive). Only the foreground/SYNC initial connect
- * fails fast, because a connectivity error is the caller's to see during initialization,
- * not after the drainer is running.
+ * store-and-forward promised to keep alive). During initialization the classification
+ * decides: {@link #isRetryable()} is true only when the provider threw a
+ * {@link TokenUnavailableException} marked retryable, which a SYNC initial connect keeps
+ * retrying within its budget; any other provider failure is permanent and fails startup
+ * fast (decision D8), as does any failure on an OFF initial connect.
*
* It exists so the send loop can tell "the provider failed" apart from "the network
* failed", and it carries the provider's own exception so a handler can surface that
@@ -52,7 +57,9 @@
* connect in {@code QwpWebSocketSender} - both catch it and rethrow
* {@link #providerFailure()}, so a token-provider failure reaches the caller as the
* provider's own exception; the running background drainer catches it and retries under
- * the invariant above. It is public because both of those packages handle it, and
+ * the invariant above. The query client ({@code QwpQueryClient}) does throw it from
+ * {@code connect()}, with a message that starts with {@code credential-unavailable}. It is
+ * public because both of those packages handle it, and
* because {@code QwpWebSocketSender.newReconnectFactory()} is public: a caller that
* drives {@code ReconnectFactory.reconnect()} itself runs the endpoint walk directly and
* so can receive this type unwrapped. Such a caller should treat it as the provider
@@ -63,12 +70,30 @@ public class QwpCredentialUnavailableException extends LineSenderException {
private final RuntimeException providerFailure;
public QwpCredentialUnavailableException(RuntimeException providerFailure) {
- super(providerFailure.getMessage() == null
+ this(providerFailure.getMessage() == null
? "token provider failed to supply a credential"
: providerFailure.getMessage(), providerFailure);
+ }
+
+ /**
+ * @param message the message, which must not contain a token
+ * @param providerFailure the exception the token provider threw
+ */
+ public QwpCredentialUnavailableException(String message, RuntimeException providerFailure) {
+ super(message, providerFailure);
this.providerFailure = providerFailure;
}
+ /**
+ * True only when the provider threw a {@link TokenUnavailableException} marked retryable. Any other
+ * exception from a provider counts as a permanent credential failure (decision D8), so that startup fails
+ * fast, as it does for an OIDC device-flow provider that is not signed in yet.
+ */
+ @Override
+ public boolean isRetryable() {
+ return providerFailure instanceof TokenUnavailableException && ((TokenUnavailableException) providerFailure).isRetryable();
+ }
+
/**
* The exception the token provider threw, for a caller that must surface the
* provider's own error rather than this wrapper. Never null: the wrapper is only
diff --git a/core/src/main/java/io/questdb/client/cutlass/qwp/client/QwpQueryClient.java b/core/src/main/java/io/questdb/client/cutlass/qwp/client/QwpQueryClient.java
index bfc924e52..380fc0eae 100644
--- a/core/src/main/java/io/questdb/client/cutlass/qwp/client/QwpQueryClient.java
+++ b/core/src/main/java/io/questdb/client/cutlass/qwp/client/QwpQueryClient.java
@@ -25,12 +25,14 @@
package io.questdb.client.cutlass.qwp.client;
import io.questdb.client.ClientTlsConfiguration;
+import io.questdb.client.ConnectionHealth;
import io.questdb.client.HttpTokenProvider;
import io.questdb.client.cutlass.http.client.HttpClientException;
import io.questdb.client.cutlass.http.client.WebSocketClient;
import io.questdb.client.cutlass.http.client.WebSocketClientFactory;
import io.questdb.client.cutlass.http.client.WebSocketFrameHandler;
-import io.questdb.client.cutlass.line.LineSenderException;
+import io.questdb.client.cutlass.auth.TokenProviderRegistry;
+import io.questdb.client.cutlass.auth.TokenProviderSpec;
import io.questdb.client.cutlass.qwp.protocol.QwpConstants;
import io.questdb.client.impl.ConfigString;
import io.questdb.client.impl.ConfigView;
@@ -151,6 +153,8 @@ public class QwpQueryClient implements QuietCloseable {
* is already in the client's kernel recv buffer by the time this wait starts.
*/
private static final int DEFAULT_SERVER_INFO_TIMEOUT_MS = 5_000;
+ // Warn once per process when an application-supplied token provider is used over ws:: (spec section 9).
+ private static final AtomicBoolean CLEARTEXT_PROVIDER_WARNED = new AtomicBoolean();
private static final Logger LOG = LoggerFactory.getLogger(QwpQueryClient.class);
// Reusable typed bind-value sink. Populated on the user thread by the
// {@link QwpBindSetter} passed to execute(); the pre-encoded bytes are
@@ -166,6 +170,8 @@ public class QwpQueryClient implements QuietCloseable {
private final List endpoints = new ArrayList<>();
private final AtomicBoolean executing = new AtomicBoolean();
private final Random failoverRandom = new Random();
+ // Connection health (design/qwp-token-provider-spec.md, section 8.4).
+ private final QwpConnectionHealthTracker healthTracker = new QwpConnectionHealthTracker();
private long authTimeoutMs = DEFAULT_AUTH_TIMEOUT_MS;
private String authorizationHeader;
// Deterministic lifecycle barrier used by facade shutdown tests. Null in
@@ -270,6 +276,9 @@ public class QwpQueryClient implements QuietCloseable {
// currentRequestId. Volatile so a cancel from any thread is visible to the
// worker thread's post-requestId read.
private volatile boolean pendingCancel;
+ // Set when a failover reconnect inside execute() failed, leaving the client disconnected: the next
+ // execute() reconnects instead of throwing "not connected". Cleared by every successful connect.
+ private volatile boolean reconnectOnNextExecute;
// Decoded SERVER_INFO from the current connection's handshake. Null before
// connect() has succeeded; non-null on every established connection (the
// server always emits the frame). Volatile so getServerInfo(), callable
@@ -300,6 +309,12 @@ public class QwpQueryClient implements QuietCloseable {
// token rotation. Mutually exclusive with the fixed authorizationHeader
// synthesized by withBearerToken/withBasicAuth; null when unset.
private HttpTokenProvider tokenProvider;
+ // The registry lease behind tokenProvider when the connect string selected a token_provider. Acquired on
+ // the first connect() - never by fromConfig(), so building or validating a client fetches no token - and
+ // released by close().
+ private volatile TokenProviderRegistry.Lease tokenProviderLease;
+ // A token_provider selected by the connect string (wss:: only); null otherwise.
+ private TokenProviderSpec tokenProviderSpec;
private char[] trustStorePassword;
private String trustStorePath;
private volatile WebSocketClient webSocketClient;
@@ -385,6 +400,7 @@ public static QwpQueryClient fromConfig(CharSequence configurationString) {
}
ConfigView view = new ConfigView(cs);
validateConfig(view, tls);
+ TokenProviderSpec tokenProviderSpec = TokenProviderSpec.parse(view, tls);
List parsedEndpoints = new ArrayList<>();
view.getHostPorts("addr", DEFAULT_WS_PORT, (h, p) -> parsedEndpoints.add(new Endpoint(h, p)));
@@ -484,6 +500,7 @@ public static QwpQueryClient fromConfig(CharSequence configurationString) {
}
if (hasBasic) client.withBasicAuth(username, password);
if (token != null) client.withBearerToken(token);
+ client.tokenProviderSpec = tokenProviderSpec;
if (cid != null) client.withClientId(cid);
if (maxBatchRows > 0) client.withMaxBatchRows(maxBatchRows);
if (zone != null) client.withZone(zone);
@@ -561,6 +578,9 @@ public static void validateConfig(ConfigView view, boolean tls) {
if (tlsRoots != null && "unsafe_off".equals(tlsVerify)) {
throw new IllegalArgumentException(TLS_ROOTS_INSECURE_CONFIG_ERROR);
}
+ // token_provider and its keys (design/qwp-token-provider-spec.md, section 7.2): resolves the factory, never
+ // fetches a token.
+ TokenProviderSpec.parse(view, tls);
// Mirror fromConfig's effective values: a missing bound takes its
// default, so the ordering is enforced even when only one key is set
// (e.g. failover_backoff_max_ms alone, below the default initial backoff).
@@ -640,6 +660,7 @@ public void close() {
// scratch, double-freeing it.
return;
}
+ healthTracker.closed();
Runnable hook = beforeCloseHook;
beforeCloseHook = null;
if (hook != null) {
@@ -720,6 +741,13 @@ public void close() {
// (submitQuery copies its bytes into sendScratch), so it is safe to free
// even when we otherwise leak the I/O thread and buffer pool.
bindValues.close();
+ // Release the token_provider lease: the registry keeps the shared provider alive for its linger
+ // period, which also covers an I/O thread that failed to join above.
+ TokenProviderRegistry.Lease lease = tokenProviderLease;
+ tokenProviderLease = null;
+ if (lease != null) {
+ lease.close();
+ }
if (wasInterrupted) {
// Hand the caller's cancellation back exactly as it arrived. Restoring it here rather
// than earlier keeps it out of the joins above, which is the whole point.
@@ -758,6 +786,36 @@ public synchronized void connect() {
if (connected) {
return;
}
+ try {
+ connectWalk();
+ healthTracker.upgraded();
+ } catch (RuntimeException e) {
+ healthTracker.roundFailed(e);
+ throw e;
+ }
+ }
+
+ /**
+ * Connection health (design/qwp-token-provider-spec.md, section 8.4). A query client connects on demand, so
+ * {@link ConnectionHealth.State#RECONNECTING} means its last connect or failover reconnect failed and the next
+ * operation will try again. Cheap and safe from any thread; never contains a credential.
+ */
+ public ConnectionHealth health() {
+ return healthTracker.snapshot();
+ }
+
+ // One connect round: resolve the credential, walk the endpoints. See connect().
+ private void connectWalk() {
+ if (tokenProviderSpec != null && tokenProviderLease == null) {
+ // The connect string selected a token_provider: share the process-wide provider for it (spec section
+ // 7.4). The lease is this client's until close().
+ tokenProviderLease = TokenProviderRegistry.global().acquire(tokenProviderSpec);
+ tokenProvider = tokenProviderLease.provider();
+ } else if (tokenProvider != null && !tlsEnabled && CLEARTEXT_PROVIDER_WARNED.compareAndSet(false, true)) {
+ // spec section 9: an application-supplied provider over ws:: is the application's call - warn once
+ LOG.warn("a token provider is used over ws:: (no TLS): bearer tokens cross the network in cleartext; "
+ + "use wss:: in production");
+ }
lastCloseTimedOut = false;
if (hostTracker == null) {
hostTracker = new QwpHostHealthTracker(
@@ -774,8 +832,18 @@ public synchronized void connect() {
// instead of being folded into "all endpoints unreachable", and avoids re-querying the provider
// once per endpoint.
String authHeader = resolveAuthorizationHeader();
+ // One same-endpoint retry after a refreshable 401 per connect, for a token provider only
+ // (design/qwp-token-provider-spec.md, section 8.2). The retried attempt records no health penalty.
+ boolean authRetryAvailable = tokenProvider != null;
+ int retryIdx = -1;
while (true) {
- int i = hostTracker.pickNext();
+ int i;
+ if (retryIdx >= 0) {
+ i = retryIdx;
+ retryIdx = -1;
+ } else {
+ i = hostTracker.pickNext();
+ }
if (i < 0) {
break;
}
@@ -784,6 +852,17 @@ public synchronized void connect() {
connectToEndpoint(ep, authHeader);
} catch (QwpAuthFailedException ae) {
cleanupFailedConnect();
+ if (authRetryAvailable && ae.isTokenRefreshable()) {
+ authRetryAvailable = false;
+ String refreshed = refreshAfterRejection(authHeader, ae);
+ if (refreshed != null && !refreshed.equals(authHeader)) {
+ LOG.info("QwpQueryClient {}:{} rejected the token with {}; retrying once with a refreshed token",
+ ep.host, ep.port, ae.getStatusCode());
+ authHeader = refreshed;
+ retryIdx = i;
+ continue;
+ }
+ }
throw ae;
} catch (QwpIngressRoleRejectedException re) {
lastTransportError = re;
@@ -819,6 +898,7 @@ public synchronized void connect() {
spawnIoThread();
hostTracker.recordSuccess(i);
currentEndpointIndex = i;
+ reconnectOnNextExecute = false;
connected = true;
return;
}
@@ -973,6 +1053,11 @@ public java.util.Map configSnapshotForTest() {
m.put("tls_verify", tlsValidationMode);
m.put("tls_roots", trustStorePath);
m.put("tls_roots_password", trustStorePassword == null ? null : new String(trustStorePassword));
+ m.put("token_provider", tokenProviderSpec == null ? null : tokenProviderSpec.name());
+ m.put("azure_resource", tokenProviderSpec == null ? null
+ : tokenProviderSpec.params().get(TokenProviderSpec.KEY_AZURE_RESOURCE));
+ m.put("azure_client_id", tokenProviderSpec == null ? null
+ : tokenProviderSpec.params().get(TokenProviderSpec.KEY_AZURE_CLIENT_ID));
return m;
}
@@ -1126,6 +1211,9 @@ public QwpQueryClient withConnectTimeout(int connectTimeoutMs) {
*/
public QwpQueryClient withBasicAuth(String username, String password) {
checkPreConnect("withBasicAuth");
+ if (tokenProviderSpec != null) {
+ throw new IllegalStateException("withBasicAuth cannot be combined with token_provider in the configuration");
+ }
if (tokenProvider != null) {
throw new IllegalStateException("withBasicAuth cannot be combined with withBearerTokenProvider");
}
@@ -1146,6 +1234,9 @@ public QwpQueryClient withBasicAuth(String username, String password) {
*/
public QwpQueryClient withBearerToken(String token) {
checkPreConnect("withBearerToken");
+ if (tokenProviderSpec != null) {
+ throw new IllegalStateException("withBearerToken cannot be combined with token_provider in the configuration");
+ }
if (tokenProvider != null) {
throw new IllegalStateException("withBearerToken cannot be combined with withBearerTokenProvider");
}
@@ -1179,6 +1270,10 @@ public QwpQueryClient withBearerTokenProvider(HttpTokenProvider provider) {
if (provider == null) {
throw new IllegalArgumentException("provider must not be null");
}
+ if (tokenProviderSpec != null) {
+ throw new IllegalStateException(
+ "withBearerTokenProvider cannot be combined with token_provider in the configuration");
+ }
if (authorizationHeader != null) {
throw new IllegalStateException("withBearerTokenProvider cannot be combined with withBearerToken or withBasicAuth");
}
@@ -1578,7 +1673,20 @@ private void executeImpl(CharSequence sql, QwpBindSetter binds, QwpColumnBatchHa
throw new IllegalStateException("QwpQueryClient is closed");
}
if (!connected) {
- throw new IllegalStateException("QwpQueryClient not connected; call connect() first");
+ if (!reconnectOnNextExecute) {
+ throw new IllegalStateException("QwpQueryClient not connected; call connect() first");
+ }
+ // A previous failover reconnect failed. Reconnect on this operation rather than leave the client
+ // permanently unusable (design/qwp-token-provider-spec.md, section 8.3, "Egress recovery") -- a
+ // pooled client is never discarded by its pool on its own, so without this a single failed
+ // failover (say, a token outage) would poison the slot for the life of the pool.
+ try {
+ connect();
+ } catch (RuntimeException e) {
+ handler.onError(-1L, WebSocketResponse.STATUS_INTERNAL_ERROR,
+ "reconnect failed: " + e.getMessage());
+ return;
+ }
}
hostTracker.beginRound(false);
long failoverDeadlineNanos;
@@ -1602,6 +1710,9 @@ private void executeImpl(CharSequence sql, QwpBindSetter binds, QwpColumnBatchHa
return;
}
if (!failoverEnabled) {
+ // With failover off the transport failure stays latched: every later execute() reports it.
+ healthTracker.connectionLost();
+ healthTracker.failed();
handler.onError(probe.interceptedRequestId, probe.interceptedStatus, probe.interceptedMessage);
return;
}
@@ -1621,6 +1732,10 @@ private void executeImpl(CharSequence sql, QwpBindSetter binds, QwpColumnBatchHa
}
cleanupFailedConnect();
connected = false;
+ healthTracker.connectionLost();
+ // Every exit from here on - an exhausted deadline, an interrupted backoff, a failed reconnect - leaves
+ // the client disconnected; the next execute() reconnects rather than throw "not connected".
+ reconnectOnNextExecute = true;
if (failoverInitialBackoffMs > 0L) {
long base = failoverInitialBackoffMs << Math.min(attempt - 1, 30);
if (base < 0L) base = failoverMaxBackoffMs;
@@ -1661,18 +1776,31 @@ private void executeImpl(CharSequence sql, QwpBindSetter binds, QwpColumnBatchHa
}
}
try {
- reconnectViaTracker();
+ try {
+ reconnectViaTracker();
+ healthTracker.upgraded();
+ } catch (RuntimeException e) {
+ healthTracker.roundFailed(e);
+ throw e;
+ }
} catch (QwpAuthFailedException authErr) {
// failover.md S6: AuthError is terminal across all hosts.
// Credentials are cluster-wide, so retrying floods server logs
// without recovery. Surface a distinct message so monitoring
- // can pull auth incidents apart from generic transport failures.
+ // can pull auth incidents apart from generic transport failures;
+ // it names the failure class (design/qwp-token-provider-spec.md, 8.3).
handler.onError(probe.interceptedRequestId, probe.interceptedStatus,
- "auth failure during failover reconnect [host="
+ "auth-rejected during failover reconnect [host="
+ authErr.getHost() + ':' + authErr.getPort()
+ ", status=" + authErr.getStatusCode()
+ ", last error: " + probe.interceptedMessage + ']');
return;
+ } catch (QwpCredentialUnavailableException credentialErr) {
+ // The token provider could not supply a credential; no endpoint was contacted.
+ handler.onError(probe.interceptedRequestId, probe.interceptedStatus,
+ "failover reconnect failed: " + credentialErr.getMessage()
+ + " [last error: " + probe.interceptedMessage + ']');
+ return;
} catch (RuntimeException reconnectErr) {
handler.onError(probe.interceptedRequestId, probe.interceptedStatus,
"failover reconnect failed after " + attempt + " attempt"
@@ -1873,8 +2001,16 @@ private void reconnectViaTracker() {
// reason as connect(): a provider failure is cluster-wide, so surface it directly rather than
// as a per-endpoint transport error retried across every host.
String authHeader = resolveAuthorizationHeader();
+ boolean authRetryAvailable = tokenProvider != null;
+ int retryIdx = -1;
while (true) {
- int i = hostTracker.pickNext();
+ int i;
+ if (retryIdx >= 0) {
+ i = retryIdx;
+ retryIdx = -1;
+ } else {
+ i = hostTracker.pickNext();
+ }
if (i < 0) {
if (!retriedAfterReset) {
hostTracker.beginRound(true);
@@ -1888,6 +2024,17 @@ private void reconnectViaTracker() {
connectToEndpoint(ep, authHeader);
} catch (QwpAuthFailedException ae) {
cleanupFailedConnect();
+ if (authRetryAvailable && ae.isTokenRefreshable()) {
+ authRetryAvailable = false;
+ String refreshed = refreshAfterRejection(authHeader, ae);
+ if (refreshed != null && !refreshed.equals(authHeader)) {
+ LOG.info("QwpQueryClient {}:{} rejected the token with {} on failover; retrying once with "
+ + "a refreshed token", ep.host, ep.port, ae.getStatusCode());
+ authHeader = refreshed;
+ retryIdx = i;
+ continue;
+ }
+ }
throw ae;
} catch (QwpIngressRoleRejectedException re) {
lastError = re;
@@ -1917,6 +2064,7 @@ private void reconnectViaTracker() {
spawnIoThread();
hostTracker.recordSuccess(i);
currentEndpointIndex = i;
+ reconnectOnNextExecute = false;
connected = true;
return;
}
@@ -1934,30 +2082,54 @@ private String resolveAuthorizationHeader() {
// With a token provider, query it once per connect()/reconnect (the caller resolves before the
// endpoint walk) so a reconnect presents a freshly refreshed token; validateToken rejects a
// null/empty/blank return, or one carrying a control or non-ASCII character, before it reaches
- // the "Bearer " header. A provider that throws (a failed silent refresh, or not signed in yet)
- // fails connect()/reconnect as a LineSenderException, preserving the provider failure as its cause.
+ // the "Bearer " header. A provider that throws (a failed silent refresh, or not signed in yet), or
+ // returns a token that fails validation, fails connect()/reconnect with the credential-unavailable
+ // failure class (design/qwp-token-provider-spec.md, section 8.1): a QwpCredentialUnavailableException
+ // whose message names the class and whose cause is the provider failure.
if (tokenProvider != null) {
- CharSequence pulled;
+ String token;
try {
- pulled = tokenProvider.getToken();
- } catch (LineSenderException e) {
- throw e;
+ CharSequence pulled = tokenProvider.getToken();
+ // snapshot before validating, for the reason HttpTokenProvider.validateToken gives: the
+ // concatenation below re-reads the sequence, and the provider may be reusing its buffer
+ token = pulled == null ? null : pulled.toString();
+ HttpTokenProvider.validateToken(token);
} catch (RuntimeException e) {
- throw new LineSenderException(
- e.getMessage() == null
+ throw new QwpCredentialUnavailableException(
+ "credential-unavailable: " + (e.getMessage() == null
? "token provider failed to supply a credential"
- : e.getMessage(),
+ : e.getMessage()),
e);
}
- // snapshot before validating, for the reason HttpTokenProvider.validateToken gives: the
- // concatenation below re-reads the sequence, and the provider may be reusing its buffer
- CharSequence token = pulled == null ? null : pulled.toString();
- HttpTokenProvider.validateToken(token);
return "Bearer " + token;
}
return authorizationHeader;
}
+ /**
+ * Steps 1-2 of the one-retry-after-401 rule (design/qwp-token-provider-spec.md, section 8.2): tells the
+ * token provider that the token in {@code presentedHeader} was rejected, then resolves the header again.
+ * Returns the new header, or null when no credential could be obtained - the connect then ends with the
+ * rejection, carrying the provider's failure as a suppressed diagnostic.
+ */
+ private String refreshAfterRejection(String presentedHeader, QwpAuthFailedException rejection) {
+ if (presentedHeader != null && presentedHeader.startsWith("Bearer ")) {
+ try {
+ tokenProvider.onTokenRejected(presentedHeader.substring("Bearer ".length()),
+ rejection.getStatusCode());
+ } catch (RuntimeException e) {
+ // The contract says it must not throw; a provider that does still gets its token re-pulled.
+ LOG.debug("token provider onTokenRejected threw {}", e.getClass().getName());
+ }
+ }
+ try {
+ return resolveAuthorizationHeader();
+ } catch (RuntimeException e) {
+ rejection.addSuppressed(e);
+ return null;
+ }
+ }
+
private long resolveQueryFlags(boolean resetSymbolDict) {
if (!resetSymbolDict) {
return 0L;
diff --git a/core/src/main/java/io/questdb/client/cutlass/qwp/client/QwpUpgradeFailures.java b/core/src/main/java/io/questdb/client/cutlass/qwp/client/QwpUpgradeFailures.java
index c0709f53d..388c0cad3 100644
--- a/core/src/main/java/io/questdb/client/cutlass/qwp/client/QwpUpgradeFailures.java
+++ b/core/src/main/java/io/questdb/client/cutlass/qwp/client/QwpUpgradeFailures.java
@@ -48,7 +48,8 @@ static HttpClientException classify(WebSocketClient client, String host, int por
}
int status = client.getUpgradeStatusCode();
if (status == 401 || status == 403) {
- QwpAuthFailedException ae = new QwpAuthFailedException(status, host, port);
+ QwpAuthFailedException ae = new QwpAuthFailedException(
+ status, host, port, client.getUpgradeRejectWwwAuthenticate());
ae.initCause(ex);
return ae;
}
diff --git a/core/src/main/java/io/questdb/client/cutlass/qwp/client/QwpWebSocketSender.java b/core/src/main/java/io/questdb/client/cutlass/qwp/client/QwpWebSocketSender.java
index 2f9c5a122..588a37e07 100644
--- a/core/src/main/java/io/questdb/client/cutlass/qwp/client/QwpWebSocketSender.java
+++ b/core/src/main/java/io/questdb/client/cutlass/qwp/client/QwpWebSocketSender.java
@@ -25,6 +25,8 @@
package io.questdb.client.cutlass.qwp.client;
import io.questdb.client.ClientTlsConfiguration;
+import io.questdb.client.ConnectionHealth;
+import io.questdb.client.HttpTokenProvider;
import io.questdb.client.Sender;
import io.questdb.client.SenderConnectionEvent;
import io.questdb.client.SenderConnectionListener;
@@ -65,6 +67,7 @@
import io.questdb.client.std.Numbers;
import io.questdb.client.std.NumericException;
import io.questdb.client.std.ObjList;
+import io.questdb.client.std.QuietCloseable;
import io.questdb.client.std.bytes.DirectByteSlice;
import org.jetbrains.annotations.NotNull;
import org.jetbrains.annotations.TestOnly;
@@ -185,6 +188,10 @@ public class QwpWebSocketSender implements Sender {
// work, and neither the foreground's reconnect nor close() can queue
// behind a drainer's endpoint walk.
private final ReentrantLock connectWalkLock = new ReentrantLock();
+ // Connection health of the FOREGROUND connection (design/qwp-token-provider-spec.md, section 8.4): rounds and
+ // upgrades from buildAndConnect, connection loss and terminal failure from the I/O loop, close from close().
+ // Background drainer walks never report here.
+ private final QwpConnectionHealthTracker healthTracker = new QwpConnectionHealthTracker();
private final QwpHostHealthTracker hostTracker;
// Per-table encoded body byte counts captured during flushPendingRows' combined
// encode. flushPendingRowsSplit uses them both for preflight sizing and to walk
@@ -233,8 +240,13 @@ public class QwpWebSocketSender implements Sender {
// Test-only lifecycle witness. close() invokes and clears it strictly after
// publishing closed=true and before starting any drain or teardown work.
private volatile Runnable closeStartedHook;
+ // Optional authentication-outage deadline (spec section 8.5) handed to the I/O loop; 0 = none.
+ private volatile long authFailureMaxDurationMillis;
private boolean connected;
private SenderConnectionDispatcher connectionDispatcher;
+ // A connect-string token_provider's registry lease (TokenProviderRegistry), released when this sender
+ // closes so the shared provider can linger and then stop. Null for any other credential.
+ private volatile QuietCloseable credentialLease;
// Async-delivery sink for SenderConnectionEvent notifications. Default
// installed at construction; the builder hook can swap before connect()
// runs, and post-connect setConnectionListener() propagates to the live
@@ -1005,6 +1017,22 @@ public static Supplier fixedAuthHeader(String header) {
return header == null ? null : new FixedAuthHeader(header);
}
+ /**
+ * Wraps a token provider as a DYNAMIC {@code Authorization} header supplier: each {@code get()} pulls the
+ * provider's current token, snapshots it, validates it ({@link HttpTokenProvider#validateToken}) and returns
+ * {@code "Bearer " + token}. Unlike a bare lambda, the wrapper also carries the provider's
+ * {@link HttpTokenProvider#onTokenRejected} back-channel, so a {@code 401} on an upgrade that presented this
+ * credential tells the provider - a caching provider then refreshes early - before the connect walk pulls
+ * again and retries the same endpoint once with the new token (design/qwp-token-provider-spec.md, section
+ * 8.2).
+ *
+ * @param provider the token provider, or null when no credential is configured
+ * @return a dynamic supplier, or null when {@code provider} is null
+ */
+ public static Supplier tokenProviderAuthHeader(HttpTokenProvider provider) {
+ return provider == null ? null : new TokenProviderAuthHeader(provider);
+ }
+
@Override
public void at(long timestamp, ChronoUnit unit) {
checkNotClosed();
@@ -1302,6 +1330,7 @@ public void close() {
}
private void close0(boolean[] restoreInterrupt) {
+ healthTracker.closed();
Runnable hook = closeStartedHook;
closeStartedHook = null;
if (hook != null) {
@@ -1453,6 +1482,9 @@ private void close0(boolean[] restoreInterrupt) {
terminalError = captureCloseError(terminalError, e);
}
}
+ // Nothing pulls a credential any more (or, after a failed stop, the registry's linger outlasts the
+ // straggler), so the shared token provider may go.
+ releaseCredentialLease();
// Always free resources the I/O thread never touches:
// encoder and table buffers are user-thread-only.
@@ -1507,6 +1539,54 @@ public boolean isCloseCleanupComplete() {
return closeCleanupComplete;
}
+ /**
+ * {@inheritDoc}
+ *
+ * Tracks the foreground connection only; orphan-slot drainers report through
+ * {@link BackgroundDrainerListener} and the error handler.
+ */
+ @Override
+ public ConnectionHealth health() {
+ return healthTracker.snapshot();
+ }
+
+ /**
+ * Arms the optional authentication-outage deadline (design/qwp-token-provider-spec.md, section 8.5); see
+ * {@code Sender.LineSenderBuilder.authFailureMaxDurationMillis(long)}. {@code <= 0} disarms it. May be called
+ * while the I/O loop runs.
+ */
+ public void setAuthFailureMaxDurationMillis(long millis) {
+ this.authFailureMaxDurationMillis = millis;
+ CursorWebSocketSendLoop loop = cursorSendLoop;
+ if (loop != null) {
+ loop.setAuthFailureMaxDurationMillis(millis);
+ }
+ }
+
+ /**
+ * Hands this sender the registry lease of the {@code token_provider} its credential comes from; the sender
+ * releases it when it closes. {@code Sender.LineSenderBuilder.build()} calls this right after connecting. A
+ * lease handed to a sender that is already closed is released at once.
+ */
+ public void setCredentialLease(QuietCloseable lease) {
+ this.credentialLease = lease;
+ if (closed) {
+ releaseCredentialLease();
+ }
+ }
+
+ private void releaseCredentialLease() {
+ QuietCloseable lease = credentialLease;
+ credentialLease = null;
+ if (lease != null) {
+ try {
+ lease.close();
+ } catch (Throwable e) {
+ LOG.error("Error releasing the token provider lease: {}", String.valueOf(e));
+ }
+ }
+ }
+
/**
* True once the store-and-forward slot flock has been released. False
* means an I/O or manager worker did not stop and close() retained the
@@ -3211,7 +3291,15 @@ private WebSocketClient buildAndConnect(ReconnectSupplier ctx, CursorWebSocketSe
}
connectWalkLock.lock();
try {
- return connectWalk(ctx, cancellation);
+ WebSocketClient client = connectWalk(ctx, cancellation);
+ healthTracker.upgraded();
+ return client;
+ } catch (RuntimeException e) {
+ // One failed connect round. A walk that a close() aborted is not a failure of the connection.
+ if (!ctx.isAborted()) {
+ healthTracker.roundFailed(e);
+ }
+ throw e;
} finally {
connectWalkLock.unlock();
}
@@ -3232,6 +3320,47 @@ private static void clearInFlight(CursorWebSocketSendLoop.ConnectCancellation ca
}
}
+ /**
+ * Steps 1-2 of the one-retry-after-401 rule (design/qwp-token-provider-spec.md, section 8.2): tells a
+ * provider-backed credential that {@code presentedHeader} was rejected - a caching provider refreshes early,
+ * waiting briefly for the fetch - and pulls the credential again. Returns the new header, or null when no
+ * credential could be obtained, in which case the round's outcome stays the rejection.
+ *
+ * Both calls run caller-supplied provider code that may block, so on a cancellable walk this thread is
+ * published as being inside a credential pull, exactly like the pull before the walk: {@code close()} then
+ * breaks it with an interrupt.
+ */
+ private String refreshCredentialAfterRejection(
+ String presentedHeader,
+ QwpAuthFailedException rejection,
+ ReconnectSupplier ctx,
+ CursorWebSocketSendLoop.ConnectCancellation cancellation
+ ) {
+ final Supplier supplier = authorizationHeaderSupplier;
+ if (cancellation != null) {
+ cancellation.publishCredentialPull(Thread.currentThread());
+ if (cancellation.isCancelled()) {
+ cancellation.clearCredentialPull();
+ throw new LineSenderException(ctx.abortMessage());
+ }
+ }
+ try {
+ if (supplier instanceof TokenProviderAuthHeader) {
+ ((TokenProviderAuthHeader) supplier).onRejected(presentedHeader, rejection.getStatusCode());
+ }
+ return supplier.get();
+ } catch (RuntimeException e) {
+ // No fresh credential to retry with. Keep the provider's failure as a diagnostic of the rejection
+ // the round ends with.
+ rejection.addSuppressed(e);
+ return null;
+ } finally {
+ if (cancellation != null) {
+ cancellation.clearCredentialPull();
+ }
+ }
+ }
+
private WebSocketClient connectWalk(ReconnectSupplier ctx, CursorWebSocketSendLoop.ConnectCancellation cancellation) {
// Background (drainer) factories share this connect walk -- endpoint
// list and hostTracker HEALTH state (never the shared round: a
@@ -3330,7 +3459,7 @@ private WebSocketClient connectWalk(ReconnectSupplier ctx, CursorWebSocketSendLo
throw new LineSenderException(ctx.abortMessage());
}
}
- final String authHeader;
+ String authHeader;
try {
authHeader = authorizationHeaderSupplier == null ? null : authorizationHeaderSupplier.get();
} catch (RuntimeException e) {
@@ -3349,11 +3478,24 @@ private WebSocketClient connectWalk(ReconnectSupplier ctx, CursorWebSocketSendLo
cancellation.clearCredentialPull();
}
}
+ // One retry after a 401 per round (design/qwp-token-provider-spec.md, section 8.2): when an upgrade that
+ // presented a DYNAMIC credential is rejected with a refreshable 401, the provider is told, the credential
+ // is pulled again and, if it changed, the SAME endpoint is retried at once. The retried attempt is not an
+ // endpoint failure: it records no health penalty, fires no event, and does not consume the round's pick.
+ // A static credential never gets it - re-presenting the same bytes cannot change the answer.
+ boolean authRetryAvailable = hasDynamicCredential();
+ int retryIdx = -1;
while (true) {
if (ctx.isAborted()) {
throw new LineSenderException(ctx.abortMessage());
}
- int idx = background ? cursor.next() : hostTracker.pickNext();
+ int idx;
+ if (retryIdx >= 0) {
+ idx = retryIdx;
+ retryIdx = -1;
+ } else {
+ idx = background ? cursor.next() : hostTracker.pickNext();
+ }
if (idx < 0) break;
Endpoint ep = endpoints.get(idx);
lastEndpoint = ep;
@@ -3425,12 +3567,26 @@ private WebSocketClient connectWalk(ReconnectSupplier ctx, CursorWebSocketSendLo
continue;
}
if (classified instanceof QwpAuthFailedException) {
+ QwpAuthFailedException rejection = (QwpAuthFailedException) classified;
+ if (authRetryAvailable && rejection.isTokenRefreshable()) {
+ authRetryAvailable = false;
+ String refreshed = refreshCredentialAfterRejection(
+ authHeader, rejection, ctx, cancellation);
+ if (refreshed != null && !refreshed.equals(authHeader)) {
+ LOG.info("{}:{} rejected the token with {}; retrying the same endpoint once with a "
+ + "refreshed token", ep.host, ep.port, rejection.getStatusCode());
+ authHeader = refreshed;
+ retryIdx = idx;
+ continue;
+ }
+ }
// Auth is uniform across the cluster; we won't keep walking
// endpoints. Fire AUTH_FAILED before throwing so the user
- // listener observes the terminal classification at the
- // moment the I/O thread gives up, ahead of the producer
+ // listener observes the classification of the round's final
+ // outcome at the moment the walk gives up, ahead of the producer
// thread learning via LineSenderException on the next
- // API call.
+ // API call. A 401 that earned the same-endpoint retry above
+ // fires nothing: only the round's final outcome is reported.
if (!background) {
dispatchConnectionEvent(
SenderConnectionEvent.Kind.AUTH_FAILED,
@@ -4090,6 +4246,8 @@ private void ensureConnected() {
// the loop no longer fires a terminal budget-exhaustion event -- it
// retries indefinitely.)
cursorSendLoop.setConnectionDispatcher(connectionDispatcher);
+ cursorSendLoop.setConnectionHealthTracker(healthTracker);
+ cursorSendLoop.setAuthFailureMaxDurationMillis(authFailureMaxDurationMillis);
cursorSendLoop.start();
} catch (Throwable t) {
// start() (or dispatcher construction) failed after cursorSendLoop was
@@ -5338,6 +5496,42 @@ public String get() {
}
}
+ /**
+ * A dynamic {@code Authorization} header backed by an {@link HttpTokenProvider}. See
+ * {@link #tokenProviderAuthHeader(HttpTokenProvider)}.
+ */
+ private static final class TokenProviderAuthHeader implements Supplier {
+ private static final String BEARER_PREFIX = "Bearer ";
+ private final HttpTokenProvider provider;
+
+ private TokenProviderAuthHeader(HttpTokenProvider provider) {
+ this.provider = provider;
+ }
+
+ @Override
+ public String get() {
+ // Snapshot before validating: the concatenation below re-reads the sequence, and a provider is free
+ // to reuse a mutable buffer, so validating the live sequence checks bytes the header need not carry.
+ // See HttpTokenProvider.validateToken.
+ CharSequence pulled = provider.getToken();
+ String token = pulled == null ? null : pulled.toString();
+ HttpTokenProvider.validateToken(token);
+ return BEARER_PREFIX + token;
+ }
+
+ void onRejected(String presentedHeader, int httpStatus) {
+ if (presentedHeader == null || !presentedHeader.startsWith(BEARER_PREFIX)) {
+ return;
+ }
+ try {
+ provider.onTokenRejected(presentedHeader.substring(BEARER_PREFIX.length()), httpStatus);
+ } catch (RuntimeException e) {
+ // The contract says it must not throw; a provider that does still gets its token re-pulled.
+ LOG.debug("token provider onTokenRejected threw {}", e.getClass().getName());
+ }
+ }
+ }
+
private final class ReconnectSupplier implements CursorWebSocketSendLoop.ReconnectFactory {
/**
* Optional caller-owned liveness gate. {@code null} means this factory
diff --git a/core/src/main/java/io/questdb/client/cutlass/qwp/client/sf/cursor/CursorWebSocketSendLoop.java b/core/src/main/java/io/questdb/client/cutlass/qwp/client/sf/cursor/CursorWebSocketSendLoop.java
index 6643bf036..a64409f56 100644
--- a/core/src/main/java/io/questdb/client/cutlass/qwp/client/sf/cursor/CursorWebSocketSendLoop.java
+++ b/core/src/main/java/io/questdb/client/cutlass/qwp/client/sf/cursor/CursorWebSocketSendLoop.java
@@ -403,6 +403,18 @@ public final class CursorWebSocketSendLoop implements QuietCloseable {
// indefinitely and never gives up on a wall-clock budget).
private volatile SenderConnectionDispatcher connectionDispatcher;
private volatile SenderErrorDispatcher errorDispatcher;
+ // Optional authentication-outage deadline (design/qwp-token-provider-spec.md, section 8.5), FOREGROUND only;
+ // 0 = none, the default. Written by the owner thread, read by the I/O thread.
+ private volatile long authFailureMaxDurationNanos;
+ // The outage clock of that deadline: started by the first authentication-class failure (credential-unavailable
+ // or a 401/403 upgrade rejection) of the current outage, reset only by a successful upgrade. Failures of other
+ // classes neither reset nor fire it. Tracked whether or not a deadline is armed, so arming it late - the
+ // builder does so right after an async connect has started - still measures the whole outage. I/O thread only.
+ private boolean authOutageActive;
+ private long authOutageStartNanos;
+ // Foreground connection health (spec section 8.4): the loop reports connection loss and its terminal failure.
+ // Null for orphan drainers.
+ private volatile io.questdb.client.cutlass.qwp.client.QwpConnectionHealthTracker healthTracker;
// The send cursor has two coordinate systems:
//
// FSN: durable frame sequence number in the local cursor engine. This is
@@ -1051,18 +1063,32 @@ public static WebSocketClient connectWithRetry(
contextLabel, e.getMessage());
throw e;
} catch (QwpCredentialUnavailableException e) {
- // A credential the client cannot ACQUIRE (the configured token provider threw) is NOT a
- // transport outage: retrying the connect cannot conjure a token the provider will not hand
- // over, so fail fast with the provider's own exception rather than burn the whole connect
- // budget treating it as a reachable-server problem (which would block build() for up to
- // maxDurationMillis, default 5 min, and surface a transport-shaped wrapper). Mirrors the
- // foreground OFF-mode connect (QwpWebSocketSender) and the background reconnect loop above,
- // which both give credential acquisition its own terminal handling; only this SYNC
- // initial-connect path lacked it. QwpCredentialUnavailableException is a LineSenderException,
- // disjoint from the HttpClientException-based terminal set above, so it reaches here.
- LOG.error("{} could not acquire a credential, won't retry: {}",
- contextLabel, e.getMessage());
- throw e.providerFailure();
+ // A credential the client cannot ACQUIRE (the configured token provider threw) is classified by
+ // the provider (design/qwp-token-provider-spec.md, section 8.3 and decisions D6/D8):
+ // - a TokenUnavailableException marked retryable - the IdP is unreachable, timing out or
+ // throttling - is a transient outage like any other, so it consumes the connect budget and is
+ // retried within it (D6);
+ // - anything else is permanent - no sign-in yet, a missing or wrong configuration - and retrying
+ // cannot conjure a token the provider will not hand over, so fail fast with the provider's own
+ // exception rather than burn the whole budget (up to maxDurationMillis, default 5 min) and
+ // surface a transport-shaped wrapper (D8). This is how an OIDC device-flow provider that is
+ // not signed in behaves, and how every provider failure behaved before D6.
+ // An interrupt on the calling thread also ends the retries: a provider wait it cut short reports
+ // retryable, and re-entering it with the flag still set would only spin through the budget.
+ // QwpCredentialUnavailableException is a LineSenderException, disjoint from the
+ // HttpClientException-based terminal set above, so it reaches here.
+ if (!e.isRetryable() || Thread.currentThread().isInterrupted()) {
+ LOG.error("{} could not acquire a credential, won't retry: {}",
+ contextLabel, e.getMessage());
+ throw e.providerFailure();
+ }
+ lastError = e;
+ long now = System.nanoTime();
+ if (now - lastLogNanos >= RECONNECT_LOG_THROTTLE_NANOS) {
+ LOG.warn("{} attempt {}: credential-unavailable (retryable): {}; retrying within connect budget",
+ contextLabel, attempts, e.getMessage());
+ lastLogNanos = now;
+ }
} catch (Throwable e) {
if (e instanceof Error) {
// JVM/programming failure (OOM, LinkageError): not a
@@ -1105,7 +1131,11 @@ public static WebSocketClient connectWithRetry(
backoffMillis = Math.min(backoffMillis * 2, maxBackoffMillis);
}
long elapsedMs = (System.nanoTime() - startNanos) / 1_000_000L;
- String lastMsg = lastError == null ? "no attempts made" : lastError.getMessage();
+ String lastMsg = lastError == null
+ ? "no attempts made"
+ : lastError instanceof QwpCredentialUnavailableException
+ ? "credential-unavailable: " + lastError.getMessage()
+ : lastError.getMessage();
throw new LineSenderException(
contextLabel + " failed after " + elapsedMs + "ms / "
+ attempts + " attempts: " + lastMsg,
@@ -1550,6 +1580,25 @@ public void setConnectionDispatcher(SenderConnectionDispatcher dispatcher) {
this.connectionDispatcher = dispatcher;
}
+ /**
+ * Arms the optional authentication-outage deadline (design/qwp-token-provider-spec.md, section 8.5). A
+ * FOREGROUND loop then latches a terminal when a connect round fails with an authentication-class failure -
+ * credential-unavailable, or a 401/403 upgrade rejection - and the current outage's first such failure lies
+ * at least this far back. {@code <= 0} disarms it. Ignored by orphan drainers, which have their own policy.
+ * May be called while the loop runs.
+ */
+ public void setAuthFailureMaxDurationMillis(long millis) {
+ this.authFailureMaxDurationNanos = millis <= 0 ? 0 : TimeUnit.MILLISECONDS.toNanos(millis);
+ }
+
+ /**
+ * Plugs the owning sender's connection-health tracker: the loop reports connection loss and its terminal
+ * failure there. Set before {@link #start()}; foreground loops only.
+ */
+ public void setConnectionHealthTracker(io.questdb.client.cutlass.qwp.client.QwpConnectionHealthTracker tracker) {
+ this.healthTracker = tracker;
+ }
+
/**
* Plug an async-delivery sink for {@link SenderError} notifications.
* Idempotent — set once before {@link #start()}; later reassignment is
@@ -1746,6 +1795,12 @@ private void connectLoop(Throwable initial, String phase, long paceFirstAttemptM
snapshotReplayTarget();
LOG.warn("cursor I/O loop entering {} loop: {}",
phase, initial.getMessage());
+ if (hasEverConnected) {
+ io.questdb.client.cutlass.qwp.client.QwpConnectionHealthTracker tracker = healthTracker;
+ if (tracker != null) {
+ tracker.connectionLost();
+ }
+ }
long outageStartNanos = System.nanoTime();
// INVARIANT B: a store-and-forward loop must NEVER terminate on a
// wall-clock reconnect budget. A replica-only / all-endpoints-replica
@@ -1790,6 +1845,8 @@ private void connectLoop(Throwable initial, String phase, long paceFirstAttemptM
try {
WebSocketClient newClient = reconnectFactory.reconnect(connectCancellation);
if (newClient != null) {
+ // a successful upgrade ends the authentication outage (spec section 8.5)
+ authOutageActive = false;
if (!running) {
// close() ran while this connect attempt was in
// flight. Its latch await may have been interrupted
@@ -1875,6 +1932,9 @@ private void connectLoop(Throwable initial, String phase, long paceFirstAttemptM
}
resetCatchUpCapGapEpisode();
lastReconnectError = e;
+ if (e instanceof QwpAuthFailedException && authOutageDeadlineFired(e)) {
+ return;
+ }
dispatchRetriedEndpointPolicyFailure(
SenderError.Category.SECURITY_ERROR, "ws-upgrade-failed: " + e.getMessage());
long now = System.nanoTime();
@@ -1942,6 +2002,9 @@ private void connectLoop(Throwable initial, String phase, long paceFirstAttemptM
// cap-gap dwell (see MAX_CATCHUP_CAP_GAP_ATTEMPTS).
resetCatchUpCapGapEpisode();
lastReconnectError = e;
+ if (authOutageDeadlineFired(e)) {
+ return;
+ }
// Retrying must not be programmatically INVISIBLE, exactly as for the auth/upgrade and
// durable-ack policy failures above: a revoked refresh token or a permanently unreachable IdP
// is not self-healing, yet flush() keeps returning success while SF absorbs the rows. Without
@@ -2048,6 +2111,53 @@ private void connectLoop(Throwable initial, String phase, long paceFirstAttemptM
phase, elapsedMs, attempts, lastMsg);
}
+ /**
+ * The optional authentication-outage deadline (design/qwp-token-provider-spec.md, section 8.5), evaluated
+ * when a connect round of a FOREGROUND loop fails with an authentication-class failure ({@code failure} is a
+ * credential-unavailable or a 401/403 rejection). Starts the outage clock on the first such failure; when a
+ * deadline is armed and the clock has reached it, latches a terminal that names the failure class and the
+ * elapsed time, reports it to the error handler, and returns true. Unacknowledged rows stay in
+ * store-and-forward.
+ */
+ private boolean authOutageDeadlineFired(Throwable failure) {
+ if (reconnectPolicy != ReconnectPolicy.FOREGROUND) {
+ return false;
+ }
+ final long now = System.nanoTime();
+ if (!authOutageActive) {
+ authOutageActive = true;
+ authOutageStartNanos = now;
+ }
+ final long deadline = authFailureMaxDurationNanos;
+ if (deadline <= 0 || now - authOutageStartNanos < deadline) {
+ return false;
+ }
+ final long elapsedMillis = TimeUnit.NANOSECONDS.toMillis(now - authOutageStartNanos);
+ final String failureClass = failure instanceof QwpCredentialUnavailableException
+ ? "credential-unavailable" : "auth-rejected";
+ final String message = "authentication outage deadline exceeded: " + failureClass + " persisted for "
+ + elapsedMillis + "ms (auth_failure_max_duration_millis="
+ + TimeUnit.NANOSECONDS.toMillis(deadline) + "); last failure: " + failure.getMessage();
+ LOG.error("{} -- the sender stops; unacknowledged rows stay in store-and-forward", message);
+ long fromFsn = engine.ackedFsn() + 1L;
+ long toFsn = Math.max(fromFsn, engine.publishedFsn());
+ SenderError err = new SenderError(
+ SenderError.Category.SECURITY_ERROR,
+ SenderError.Policy.TERMINAL,
+ SenderError.NO_STATUS_BYTE,
+ message,
+ SenderError.NO_MESSAGE_SEQUENCE,
+ fromFsn,
+ toFsn,
+ null,
+ System.nanoTime()
+ );
+ totalServerErrors.incrementAndGet();
+ recordFatal(new LineSenderServerException(err));
+ dispatchError(err);
+ return true;
+ }
+
/**
* Reports an endpoint-policy rejection a FOREGROUND sender is riding out.
*
@@ -2549,6 +2659,10 @@ private void positionCursorInSegment(MmapSegment seg, long targetFsn) {
* every rethrow delivers the same instance.
*/
private void recordFatal(Throwable t) {
+ io.questdb.client.cutlass.qwp.client.QwpConnectionHealthTracker tracker = healthTracker;
+ if (tracker != null) {
+ tracker.failed();
+ }
if (terminalError == null) {
terminalError = t instanceof LineSenderException
? (LineSenderException) t
diff --git a/core/src/main/java/io/questdb/client/impl/ConfigSchema.java b/core/src/main/java/io/questdb/client/impl/ConfigSchema.java
index c9529a13e..38eebb6fc 100644
--- a/core/src/main/java/io/questdb/client/impl/ConfigSchema.java
+++ b/core/src/main/java/io/questdb/client/impl/ConfigSchema.java
@@ -57,10 +57,19 @@ public final class ConfigSchema {
str("tls_roots_password", Side.COMMON);
longRange("auth_timeout_ms", Side.COMMON, 0, OPEN_MAX, true, false); // > 0
longRange("connect_timeout", Side.COMMON, 0, OPEN_MAX, true, false); // > 0
+ // Dynamic bearer credentials (design/qwp-token-provider-spec.md, section 7.1): a refreshing token
+ // provider selected by name, wss:: only, mutually exclusive with token/username/password. The values are
+ // never secret. Validated by TokenProviderSpec on both clients.
+ str("token_provider", Side.COMMON);
+ str("azure_resource", Side.COMMON);
+ str("azure_client_id", Side.COMMON);
// INGRESS -- the WebSocket Sender applies. STRING in the registry; the
// Sender parses suffix/mode values (off/on, 64k, durability) with its
// own helpers, byte-for-byte.
+ // Optional authentication-outage deadline (design/qwp-token-provider-spec.md, section 8.5). Not set by
+ // default; > 0 when set.
+ longRange("auth_failure_max_duration_millis", Side.INGRESS, 0, OPEN_MAX, true, false);
str("auto_flush", Side.INGRESS);
str("auto_flush_bytes", Side.INGRESS);
str("auto_flush_interval", Side.INGRESS);
diff --git a/core/src/main/java/io/questdb/client/impl/PooledSender.java b/core/src/main/java/io/questdb/client/impl/PooledSender.java
index 7b4e5f802..06f3940f9 100644
--- a/core/src/main/java/io/questdb/client/impl/PooledSender.java
+++ b/core/src/main/java/io/questdb/client/impl/PooledSender.java
@@ -24,6 +24,7 @@
package io.questdb.client.impl;
+import io.questdb.client.ConnectionHealth;
import io.questdb.client.Sender;
import io.questdb.client.cutlass.line.array.DoubleArray;
import io.questdb.client.cutlass.line.array.LongArray;
@@ -276,6 +277,11 @@ public long getAckedFsn() {
return slot.live(generation).getAckedFsn();
}
+ @Override
+ public ConnectionHealth health() {
+ return slot.live(generation).health();
+ }
+
@Override
public Sender intColumn(CharSequence name, int value) {
slot.live(generation).intColumn(name, value);
diff --git a/core/src/main/java/io/questdb/client/impl/QueryClientPool.java b/core/src/main/java/io/questdb/client/impl/QueryClientPool.java
index 8e858c019..4f3c7b566 100644
--- a/core/src/main/java/io/questdb/client/impl/QueryClientPool.java
+++ b/core/src/main/java/io/questdb/client/impl/QueryClientPool.java
@@ -24,6 +24,7 @@
package io.questdb.client.impl;
+import io.questdb.client.ConnectionHealth;
import io.questdb.client.HttpTokenProvider;
import io.questdb.client.QueryException;
import io.questdb.client.cutlass.qwp.client.QwpQueryClient;
@@ -485,6 +486,21 @@ void discard(QueryWorker w, long gen) {
}
}
+ /**
+ * Adds the connection health of every pooled query client to {@code into}. Reading a client's health never
+ * waits on its I/O, so this is safe under the pool lock.
+ */
+ void collectHealth(java.util.List into) {
+ lock.lock();
+ try {
+ for (int i = 0, n = all.size(); i < n; i++) {
+ into.add(all.get(i).client().health());
+ }
+ } finally {
+ lock.unlock();
+ }
+ }
+
void reapIdle() {
if (closed) {
return;
diff --git a/core/src/main/java/io/questdb/client/impl/QuestDBImpl.java b/core/src/main/java/io/questdb/client/impl/QuestDBImpl.java
index 574d6b59e..0f0c4c64f 100644
--- a/core/src/main/java/io/questdb/client/impl/QuestDBImpl.java
+++ b/core/src/main/java/io/questdb/client/impl/QuestDBImpl.java
@@ -24,6 +24,7 @@
package io.questdb.client.impl;
+import io.questdb.client.ConnectionHealth;
import io.questdb.client.HttpTokenProvider;
import io.questdb.client.QuestDB;
import io.questdb.client.Query;
@@ -206,6 +207,14 @@ public Sender borrowSender() {
return senderPool.borrow();
}
+ @Override
+ public ConnectionHealth.Aggregate health() {
+ java.util.List healths = new java.util.ArrayList<>();
+ senderPool.collectHealth(healths);
+ queryPool.collectHealth(healths);
+ return ConnectionHealth.Aggregate.of(healths);
+ }
+
// synchronized so concurrent close() callers serialize THROUGH the bounded
// shutdown sequence, not merely through the `closed` flip. `closed` is set
// before the teardown chain runs, so a plain volatile guard (or a bare CAS)
diff --git a/core/src/main/java/io/questdb/client/impl/SenderPool.java b/core/src/main/java/io/questdb/client/impl/SenderPool.java
index 1a4390269..fd899158c 100644
--- a/core/src/main/java/io/questdb/client/impl/SenderPool.java
+++ b/core/src/main/java/io/questdb/client/impl/SenderPool.java
@@ -24,6 +24,7 @@
package io.questdb.client.impl;
+import io.questdb.client.ConnectionHealth;
import io.questdb.client.HttpTokenProvider;
import io.questdb.client.Sender;
import io.questdb.client.SenderConnectionListener;
@@ -1923,6 +1924,25 @@ public int availableSize() {
}
}
+ /**
+ * Adds the connection health of every live pooled sender to {@code into}. Reading a sender's health never
+ * waits on its I/O, so this is safe under the pool lock.
+ */
+ public void collectHealth(java.util.List into) {
+ lock.lock();
+ try {
+ for (int i = 0, n = all.size(); i < n; i++) {
+ try {
+ into.add(all.get(i).delegate().health());
+ } catch (UnsupportedOperationException ignored) {
+ // a test double without connection health
+ }
+ }
+ } finally {
+ lock.unlock();
+ }
+ }
+
/** Snapshot of the total number of live slots (idle + in-use). For tests and introspection. */
public int totalSize() {
lock.lock();
diff --git a/core/src/main/java/module-info.java b/core/src/main/java/module-info.java
index 8383221e4..4f4bc9852 100644
--- a/core/src/main/java/module-info.java
+++ b/core/src/main/java/module-info.java
@@ -71,4 +71,8 @@
exports io.questdb.client.cutlass.qwp.client.sf.cursor;
exports io.questdb.client.cutlass.qwp.protocol;
exports io.questdb.client.cutlass.qwp.websocket;
+
+ // token_provider= in a connect string resolves through this SPI; the optional questdb-client-azure
+ // artifact provides the "azure" factory (design/qwp-token-provider-spec.md, section 7).
+ uses io.questdb.client.cutlass.auth.TokenProviderFactory;
}
diff --git a/core/src/test/java/io/questdb/client/test/cutlass/auth/RefreshingTokenProviderTest.java b/core/src/test/java/io/questdb/client/test/cutlass/auth/RefreshingTokenProviderTest.java
new file mode 100644
index 000000000..629377106
--- /dev/null
+++ b/core/src/test/java/io/questdb/client/test/cutlass/auth/RefreshingTokenProviderTest.java
@@ -0,0 +1,841 @@
+/*+*****************************************************************************
+ * ___ _ ____ ____
+ * / _ \ _ _ ___ ___| |_| _ \| __ )
+ * | | | | | | |/ _ \/ __| __| | | | _ \
+ * | |_| | |_| | __/\__ \ |_| |_| | |_) |
+ * \__\_\\__,_|\___||___/\__|____/|____/
+ *
+ * Copyright (c) 2014-2019 Appsicle
+ * Copyright (c) 2019-2026 QuestDB
+ *
+ * Licensed under the Apache License, Version 2.0 (the "License");
+ * you may not use this file except in compliance with the License.
+ * You may obtain a copy of the License at
+ *
+ * http://www.apache.org/licenses/LICENSE-2.0
+ *
+ * Unless required by applicable law or agreed to in writing, software
+ * distributed under the License is distributed on an "AS IS" BASIS,
+ * WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied.
+ * See the License for the specific language governing permissions and
+ * limitations under the License.
+ *
+ ******************************************************************************/
+
+package io.questdb.client.test.cutlass.auth;
+
+import io.questdb.client.cutlass.auth.CredentialRedaction;
+import io.questdb.client.cutlass.auth.ExpiringToken;
+import io.questdb.client.cutlass.auth.RefreshingTokenProvider;
+import io.questdb.client.cutlass.auth.TokenUnavailableException;
+import io.questdb.client.test.cutlass.auth.TokenTestKit.FakeClock;
+import io.questdb.client.test.cutlass.auth.TokenTestKit.ManualScheduler;
+import io.questdb.client.test.cutlass.auth.TokenTestKit.ScriptedSource;
+import org.junit.Assert;
+import org.junit.Test;
+
+import java.util.ArrayList;
+import java.util.List;
+import java.util.UUID;
+import java.util.concurrent.CountDownLatch;
+import java.util.concurrent.TimeUnit;
+import java.util.concurrent.atomic.AtomicBoolean;
+import java.util.concurrent.atomic.AtomicReference;
+
+import static io.questdb.client.test.cutlass.auth.TokenTestKit.await;
+
+/**
+ * Conformance tests C1-C8 and the cache half of C20 of the dynamic-credential specification
+ * (design/qwp-token-provider-spec.md, section 10) for {@link RefreshingTokenProvider}.
+ *
+ * Schedule, backoff, hand-out and forced-refresh tests run on a fake clock and a scheduler the test drives, so
+ * they assert exact delays without waiting. The concurrency, interrupt and lifecycle tests run on the real
+ * refresher thread.
+ */
+public class RefreshingTokenProviderTest {
+ private static final long MIN = 60_000L;
+ private static final long HOUR = 60 * MIN;
+ private static final long DAY = 24 * HOUR;
+ // A fixed wall-clock origin keeps the expected instants literal.
+ private static final long T0 = 1_800_000_000_000L;
+
+ @Test
+ public void testAlreadyExpiredResultIsARetryableFailure() {
+ FakeClock clock = new FakeClock(T0);
+ ManualScheduler scheduler = new ManualScheduler(clock);
+ ScriptedSource source = new ScriptedSource()
+ .thenToken("EXPIRED", T0 - 1)
+ .thenToken("FRESH", T0 + HOUR);
+ try (RefreshingTokenProvider provider = provider(source, clock, scheduler, 0.0).build()) {
+ scheduler.runNext();
+ Assert.assertEquals(1, provider.getConsecutiveFailures());
+ RefreshingTokenProvider.FetchFailure failure = provider.getLastFailure();
+ Assert.assertNotNull(failure);
+ Assert.assertTrue("a result that has already expired is a retryable failure", failure.isRetryable());
+ Assert.assertTrue(failure.getMessage(), failure.getMessage().contains("expired"));
+ Assert.assertEquals(RefreshingTokenProvider.NONE, provider.getTokenExpiresAtEpochMillis());
+ // retried after backoff, not dropped
+ scheduler.advanceAndRunNext();
+ Assert.assertEquals("FRESH", provider.getToken().toString());
+ Assert.assertEquals(0, provider.getConsecutiveFailures());
+ }
+ }
+
+ @Test(timeout = 30_000)
+ public void testAwaitReadyGatesOnTheFirstFetch() {
+ ScriptedSource source = new ScriptedSource().thenToken("READY", System.currentTimeMillis() + HOUR);
+ source.setGate();
+ try (RefreshingTokenProvider provider = RefreshingTokenProvider.builder(source).build()) {
+ Assert.assertFalse("the first fetch has not finished", provider.awaitReady(50));
+ source.openGate();
+ Assert.assertTrue(provider.awaitReady(10_000));
+ Assert.assertEquals("READY", provider.getToken().toString());
+ }
+ }
+
+ @Test
+ public void testBackoffFollowsTheSpec() {
+ // section 5.4: base = min(backoff_max, backoff_initial * 2^(n-1)); delay = base/2 + U(0, base/2).
+ // u = 0 pins the delay to the lower bound base/2, u = 0.5 to 3/4 of base.
+ assertBackoffDelays(0.0, new long[]{250, 500, 1_000, 2_000, 4_000, 8_000, 16_000, 30_000, 30_000});
+ assertBackoffDelays(0.5, new long[]{375, 750, 1_500, 3_000, 6_000, 12_000, 24_000, 45_000, 45_000});
+ }
+
+ @Test
+ public void testCloseFailsGetTokenPermanentlyAndStopsTheRefresher() {
+ ScriptedSource source = new ScriptedSource().thenToken("TOKEN", System.currentTimeMillis() + HOUR);
+ List before = refresherThreads();
+ RefreshingTokenProvider provider = RefreshingTokenProvider.builder(source).build();
+ Assert.assertTrue(provider.awaitReady(10_000));
+ List started = refresherThreads();
+ started.removeAll(before);
+ Assert.assertEquals("one refresher thread per provider", 1, started.size());
+
+ provider.close();
+ provider.close(); // idempotent
+ Assert.assertTrue(provider.isClosed());
+ await(() -> !started.get(0).isAlive(), 5_000, "the refresher thread to exit");
+ try {
+ provider.getToken();
+ Assert.fail("getToken() after close() must fail");
+ } catch (TokenUnavailableException e) {
+ Assert.assertFalse("a closed provider is a permanent failure", e.isRetryable());
+ }
+ Assert.assertFalse(provider.awaitReady(10));
+ provider.onTokenRejected("TOKEN", 401); // no-op, no throw
+ }
+
+ @Test(timeout = 30_000)
+ public void testColdBurstAfterTheWallClockJumpsCausesOneFetch() throws Exception {
+ // C3, second shape: a token is held, but the wall clock jumped past its expiry (a host that slept
+ // through its scheduled refresh). 64 callers find no usable token at once; exactly one fetch runs.
+ SkewedClock clock = new SkewedClock();
+ ScriptedSource source = new ScriptedSource()
+ .thenToken("OLD", System.currentTimeMillis() + HOUR)
+ .then(() -> new ExpiringToken("NEW", clock.wallClockMillis() + HOUR));
+ try (RefreshingTokenProvider provider = RefreshingTokenProvider.builder(source).clock(clock).build()) {
+ Assert.assertTrue(provider.awaitReady(10_000));
+ Assert.assertEquals("OLD", provider.getToken().toString());
+ source.setGate();
+ clock.skewMillis = 2 * HOUR;
+ Assert.assertEquals(listOf("NEW"), burst(provider, 64, source));
+ Assert.assertEquals("the prefetch plus exactly one fetch for the whole burst", 2, source.calls());
+ }
+ }
+
+ @Test(timeout = 30_000)
+ public void testColdBurstCausesExactlyOneFetch() throws Exception {
+ // C3: 64 concurrent callers with no usable token cause exactly one fetch.
+ ScriptedSource source = new ScriptedSource().thenToken("TOKEN-1", System.currentTimeMillis() + HOUR);
+ source.setGate(); // the prefetch blocks until every caller is waiting
+ try (RefreshingTokenProvider provider = RefreshingTokenProvider.builder(source).build()) {
+ Assert.assertEquals(listOf("TOKEN-1"), burst(provider, 64, source));
+ Assert.assertEquals(1, source.calls());
+ }
+ }
+
+ @Test(timeout = 30_000)
+ public void testColdFailureFailsImmediatelyWhenBackoffOutlastsTheWait() {
+ // section 5.3: "If the earliest permitted fetch is later than now + cold_wait, fail immediately."
+ ScriptedSource source = new ScriptedSource()
+ .thenThrow(TokenUnavailableException.retryable("HTTP 429 from the token endpoint", 60_000));
+ try (RefreshingTokenProvider provider = RefreshingTokenProvider.builder(source)
+ .coldWaitMillis(5_000)
+ .build()) {
+ await(() -> provider.getConsecutiveFailures() == 1, 10_000, "the prefetch to fail");
+ long start = System.nanoTime();
+ try {
+ provider.getToken();
+ Assert.fail("expected the cold wait to fail");
+ } catch (TokenUnavailableException e) {
+ long elapsedMillis = TimeUnit.NANOSECONDS.toMillis(System.nanoTime() - start);
+ Assert.assertTrue("must fail at once, not wait out cold_wait; took " + elapsedMillis + " ms",
+ elapsedMillis < 2_000);
+ Assert.assertTrue(e.isRetryable());
+ Assert.assertTrue("the remaining backoff is the hint: " + e.getRetryAfterMillis(),
+ e.getRetryAfterMillis() > 30_000);
+ Assert.assertTrue(e.getMessage(), e.getMessage().contains("HTTP 429"));
+ }
+ Assert.assertEquals("the caller must not bypass the backoff with a fetch of its own", 1, source.calls());
+ }
+ }
+
+ @Test(timeout = 30_000)
+ public void testColdFailurePassesAPermanentClassificationThrough() {
+ // C4: a permanent classification is passed through to the caller.
+ ScriptedSource source = new ScriptedSource()
+ .thenThrow(TokenUnavailableException.permanent("no credential is configured"));
+ try (RefreshingTokenProvider provider = RefreshingTokenProvider.builder(source)
+ .coldWaitMillis(300)
+ .backoffInitialMillis(20)
+ .backoffMaxMillis(40)
+ .build()) {
+ try {
+ provider.getToken();
+ Assert.fail("expected a token-unavailable error");
+ } catch (TokenUnavailableException e) {
+ Assert.assertFalse("the source's permanent classification must reach the caller", e.isRetryable());
+ Assert.assertTrue(e.getMessage(), e.getMessage().contains("no credential is configured"));
+ }
+ // the provider keeps retrying a "permanent" failure: an operator can repair it without a restart
+ int calls = source.calls();
+ await(() -> source.calls() > calls, 5_000, "another retry of the permanent failure");
+ }
+ }
+
+ @Test(timeout = 30_000)
+ public void testColdFailureRaisesARetryableErrorWithinColdWait() {
+ // C4: a retryable error is raised within cold_wait.
+ ScriptedSource source = new ScriptedSource()
+ .thenThrow(TokenUnavailableException.retryable("IMDS answered HTTP 503"));
+ long coldWaitMillis = 400;
+ try (RefreshingTokenProvider provider = RefreshingTokenProvider.builder(source)
+ .coldWaitMillis(coldWaitMillis)
+ .backoffInitialMillis(20)
+ .backoffMaxMillis(40)
+ .build()) {
+ long start = System.nanoTime();
+ try {
+ provider.getToken();
+ Assert.fail("expected a token-unavailable error");
+ } catch (TokenUnavailableException e) {
+ long elapsedMillis = TimeUnit.NANOSECONDS.toMillis(System.nanoTime() - start);
+ Assert.assertTrue(e.isRetryable());
+ Assert.assertTrue(e.getMessage(), e.getMessage().contains("HTTP 503"));
+ Assert.assertTrue("the wait must be bounded by cold_wait; took " + elapsedMillis + " ms",
+ elapsedMillis < coldWaitMillis + 2_000);
+ }
+ Assert.assertTrue("short failures inside cold_wait are retried, not surfaced on the first one",
+ source.calls() > 1);
+ }
+ }
+
+ @Test
+ public void testCredentialRedactionHelpers() {
+ String token = "abc.def.ghi";
+ String fingerprint = CredentialRedaction.fingerprint(token);
+ Assert.assertTrue(fingerprint, fingerprint.matches("[0-9a-f]{8}"));
+ String described = CredentialRedaction.describeToken(token);
+ Assert.assertEquals("', described);
+ Assert.assertEquals("", CredentialRedaction.describeToken(null));
+
+ // controls, CR/LF and bidi overrides are removed; the result is capped at 256 characters
+ String hostile = "line1\r\nFORGED: yes\u202Eevil\u0007" + repeat('x', 1_000);
+ String clean = CredentialRedaction.sanitizeErrorText(hostile);
+ Assert.assertTrue(clean.length() <= CredentialRedaction.MAX_ERROR_TEXT_LENGTH);
+ Assert.assertTrue(clean.endsWith("..."));
+ Assert.assertFalse(clean.contains("\r") || clean.contains("\n") || clean.contains("\u202E")
+ || clean.contains("\u0007"));
+ Assert.assertTrue(clean.startsWith("line1FORGED: yesevil"));
+ // text that fits is kept whole
+ String exact = repeat('y', CredentialRedaction.MAX_ERROR_TEXT_LENGTH);
+ Assert.assertEquals(exact, CredentialRedaction.sanitizeErrorText(exact));
+ Assert.assertNull(CredentialRedaction.sanitizeErrorText(null));
+ }
+
+ @Test
+ public void testExpiringTokenRejectsInvalidTokensWithoutEchoingThem() {
+ String secret = "SECRET-" + UUID.randomUUID();
+ assertRejected(null);
+ assertRejected("");
+ assertRejected(" ");
+ assertRejected(secret + "\r\nX-Injected: 1");
+ assertRejected(secret + "\u00e9");
+ try {
+ new ExpiringToken(secret + '\n', T0);
+ Assert.fail();
+ } catch (IllegalArgumentException e) {
+ Assert.assertFalse("the token must never be echoed", e.getMessage().contains(secret));
+ }
+ ExpiringToken ok = new ExpiringToken(secret, T0, -5);
+ Assert.assertFalse("a non-positive refresh hint means none", ok.hasRefreshAt());
+ Assert.assertEquals(ExpiringToken.NO_REFRESH_AT, ok.getRefreshAtEpochMillis());
+ Assert.assertTrue(new ExpiringToken(secret, T0 + HOUR, T0 + MIN).hasRefreshAt());
+ }
+
+ @Test(timeout = 30_000)
+ public void testForcedRefreshIsRateLimitedIgnoresStaleTokensAndAcceptsTheSameToken() throws Exception {
+ // C8: rate limited; a stale token is ignored; a same-token result is accepted.
+ FakeClock clock = new FakeClock(T0);
+ ManualScheduler scheduler = new ManualScheduler(clock);
+ ScriptedSource source = new ScriptedSource()
+ .thenToken("T1", T0 + HOUR)
+ .thenToken("T2", T0 + HOUR)
+ .thenToken("T2", T0 + HOUR);
+ try (RefreshingTokenProvider provider = provider(source, clock, scheduler, 0.5)
+ .forcedMinIntervalMillis(30_000)
+ .forcedWaitMillis(5_000)
+ .build()) {
+ scheduler.runNext();
+ Assert.assertEquals("T1", provider.getToken().toString());
+
+ // 403 never triggers a forced refresh (decision D2)
+ assertReturnsPromptly(() -> provider.onTokenRejected("T1", 403));
+ Assert.assertNotEquals("a 403 must not schedule a fetch", 0, scheduler.peek().delayNanos);
+
+ // 401 for the current token: a fetch starts now and the caller waits for it
+ Thread rejecter = start(() -> provider.onTokenRejected("T1", 401));
+ await(() -> scheduler.peek() != null && scheduler.peek().delayNanos == 0, 5_000,
+ "the forced fetch to be scheduled");
+ scheduler.runNext();
+ rejecter.join(5_000);
+ Assert.assertFalse("onTokenRejected must return once the forced fetch completes", rejecter.isAlive());
+ Assert.assertEquals("T2", provider.getToken().toString());
+ Assert.assertEquals(2, source.calls());
+
+ // a stale token - it has already rotated - is ignored
+ assertReturnsPromptly(() -> provider.onTokenRejected("T1", 401));
+ // rate limited: a second forced refresh within forced_min_interval is ignored
+ clock.advanceMillis(30_000 - 1);
+ assertReturnsPromptly(() -> provider.onTokenRejected("T2", 401));
+ Assert.assertNotEquals("no forced fetch may be scheduled", 0, scheduler.peek().delayNanos);
+ Assert.assertEquals(2, source.calls());
+
+ // once the interval has passed, a forced fetch that returns the same token is a normal success
+ clock.advanceMillis(1);
+ rejecter = start(() -> provider.onTokenRejected("T2", 401));
+ await(() -> scheduler.peek() != null && scheduler.peek().delayNanos == 0, 5_000,
+ "the second forced fetch to be scheduled");
+ scheduler.runNext();
+ rejecter.join(5_000);
+ Assert.assertFalse(rejecter.isAlive());
+ Assert.assertEquals(3, source.calls());
+ Assert.assertEquals("T2", provider.getToken().toString());
+ Assert.assertEquals(0, provider.getConsecutiveFailures());
+ Assert.assertNull("a same-token result is not a failure", provider.getLastFailure());
+ Assert.assertEquals(clock.wallClockMillis(), provider.getLastSuccessEpochMillis());
+ }
+ }
+
+ @Test(timeout = 30_000)
+ public void testHandoutFloorTokenIsNeverHandedOut() throws Exception {
+ // C6: a token inside handout_floor is never handed out.
+ FakeClock clock = new FakeClock(T0);
+ ManualScheduler scheduler = new ManualScheduler(clock);
+ ScriptedSource source = new ScriptedSource()
+ .thenToken("T1", T0 + 10 * MIN)
+ .thenToken("T2", T0 + 70 * MIN);
+ try (RefreshingTokenProvider provider = provider(source, clock, scheduler, 0.5).build()) {
+ scheduler.runNext();
+ clock.jumpWallMillis(10 * MIN - RefreshingTokenProvider.DEFAULT_HANDOUT_FLOOR_MILLIS - 1);
+ Assert.assertEquals("one millisecond outside the floor the token is still usable",
+ "T1", provider.getToken().toString());
+ clock.jumpWallMillis(1);
+
+ // exactly at expires_at - handout_floor the token is unusable: the caller must get a new one
+ AtomicReference