Repository navigation
Too large queries produce MaxRetryError #413
Description
Activity
@rth Can you please enable debug logging and share the log?
import logging from databricks import sql logging.basicConfig(level=logging.DEBUG) # other your codeAlso, can you try to access that failed link using wget/curl/browser? I'm curious if there are some SSL issues on server, or if that's something on our side
+ additional question: do you use any kind of proxy, firewall, VPN, or something that may affect ssl cert validation?
+ if you are able to path
databricks/sqlon your machine - can you try this? locate this file and line - https://github.com/databricks/databricks-sql-python/blob/main/src/databricks/sql/cloudfetch/downloader.py#L96 - and add averify=FalseparameterThanks for your feedback @kravets-levko !
Yes, I'm behind a corporate proxy that does SSL cerificate rewrites, so by itself SSLError s are expected.
If I modifydownloader.py#L96to addverify=Falseactually it works even for a big query that previously failed. It's just confusing because I though I disabled SSL verification since smaller queries worked fine.Any chance you could allows users to disable ssl verification in that section without editing the code? For instance either via the
_tls_no_verify(orssl_verify) parameter passed toconnector even via monkeypatching some object indatabricks.sql.cloudfetch.downloaderif option a) would be difficult.If still relevant some of the debug logs in the first case where it failed are below,
Details
DEBUG:databricks.sql.thrift_backend:retry parameter: _retry_delay_min given_or_default 1.0 DEBUG:databricks.sql.thrift_backend:retry parameter: _retry_delay_max given_or_default 60.0 DEBUG:databricks.sql.thrift_backend:retry parameter: _retry_stop_after_attempts_count given_or_default 30 DEBUG:databricks.sql.thrift_backend:retry parameter: _retry_stop_after_attempts_duration given_or_default 900.0 DEBUG:databricks.sql.thrift_backend:retry parameter: _retry_delay_default given_or_default 5.0 DEBUG:databricks.sql.thrift_backend:Sending request: OpenSession(<REDACTED>) DEBUG:urllib3.connectionpool:Starting new HTTPS connection (1): xxx.14.azuredatabricks.net:443 DEBUG:urllib3.connectionpool:https://xxxx.azuredatabricks.net:443 "POST /sql/protocolv1/o/xxxx/0605-141634-2g07xge9 HTTP/11" 200 171 DEBUG:databricks.sql.thrift_backend:Received response: TOpenSessionResp(<REDACTED>) INFO:databricks.sql.client:Successfully opened session 37f9a404-d4f2-4f21-aafc-464e03cf22e0 DEBUG:databricks.sql.thrift_backend:Sending request: ExecuteStatement(<REDACTED>) DEBUG:urllib3.connectionpool:https://xxxx.azuredatabricks.net:443 "POST /sql/protocolv1/o/xxxxx/0605-141634-2g07xge9 HTTP/11" 200 14371 DEBUG:databricks.sql.thrift_backend:Received response: TExecuteStatementResp(<REDACTED>) DEBUG:databricks.sql.utils:Initialize CloudFetch loader, row set start offset: 0, file list: DEBUG:databricks.sql.utils:- start row offset: 0, row count: 49152 DEBUG:databricks.sql.utils:- start row offset: 49152, row count: 49152 DEBUG:databricks.sql.utils:- start row offset: 98304, row count: 22480 DEBUG:databricks.sql.utils:- start row offset: 120784, row count: 49152 DEBUG:databricks.sql.utils:- start row offset: 169936, row count: 49152 DEBUG:databricks.sql.cloudfetch.download_manager:ResultFileDownloadManager: adding file link, start offset 0, row count: 49152 DEBUG:databricks.sql.cloudfetch.download_manager:ResultFileDownloadManager: adding file link, start offset 49152, row count: 49152 DEBUG:databricks.sql.cloudfetch.download_manager:ResultFileDownloadManager: adding file link, start offset 98304, row count: 22480 DEBUG:databricks.sql.cloudfetch.download_manager:ResultFileDownloadManager: adding file link, start offset 120784, row count: 49152 DEBUG:databricks.sql.cloudfetch.download_manager:ResultFileDownloadManager: adding file link, start offset 169936, row count: 49152 DEBUG:databricks.sql.utils:CloudFetchQueue: Trying to get downloaded file for row 0 DEBUG:databricks.sql.cloudfetch.download_manager:ResultFileDownloadManager: schedule downloads DEBUG:databricks.sql.cloudfetch.download_manager:- start: 0, row count: 49152 DEBUG:databricks.sql.cloudfetch.downloader:ResultSetDownloadHandler: starting file download, offset 0, row count 49152 DEBUG:databricks.sql.cloudfetch.download_manager:- start: 49152, row count: 49152 DEBUG:databricks.sql.cloudfetch.downloader:ResultSetDownloadHandler: starting file download, offset 49152, row count 49152 DEBUG:databricks.sql.cloudfetch.download_manager:- start: 98304, row count: 22480 DEBUG:databricks.sql.cloudfetch.downloader:ResultSetDownloadHandler: starting file download, offset 98304, row count 22480 DEBUG:databricks.sql.cloudfetch.download_manager:- start: 120784, row count: 49152 DEBUG:databricks.sql.cloudfetch.downloader:ResultSetDownloadHandler: starting file download, offset 120784, row count 49152 DEBUG:databricks.sql.cloudfetch.download_manager:- start: 169936, row count: 49152 DEBUG:databricks.sql.cloudfetch.downloader:ResultSetDownloadHandler: starting file download, offset 169936, row count 49152 DEBUG:urllib3.connectionpool:Starting new HTTPS connection (1): xxx.blob.core.windows.net:443 DEBUG:urllib3.connectionpool:Starting new HTTPS connection (1): xxx.blob.core.windows.net:443 DEBUG:urllib3.connectionpool:Starting new HTTPS connection (1): xxx.blob.core.windows.net:443 DEBUG:urllib3.connectionpool:Starting new HTTPS connection (1): xxx.blob.core.windows.net:443 DEBUG:urllib3.connectionpool:Starting new HTTPS connection (1): xxx.blob.core.windows.net:443 DEBUG:urllib3.util.retry:Incremented Retry for (url='/jobs/xxxx/sql/2024-07-11/15/results_2024-07-11T15:40:05Z_05e1b6a6-3439-440c-8522-b428f40d3b6f?sig=xxx%2FxxtrcDCT8U%3D&se=2024-07-11T15%3A55%3A07Z&sv=2019-02-02&spr=https&sp=r&sr=b'): Retry(total=4, connect=None, read=None, redirect=None, status=None)Reacted by Levko KravetsThe thing is that for smaller results CloudFetch is not used, that's why you were able to get result. Thank you for your help and all the feedback you provide, and also I really appreciate that you were able to help me with debugging. PR will come in a minute
Reacted by Roman Yurchak
Previously when a too big query was made #383 we got 0 rows as output (as discussed in that issue) . With the changes in #405 for me it now produces an MaxRetryError which is better but the error message is misleading (and also retying so many times is slow).
The minimal code I'm using is,
If the query is small, it works with no warnings.
If the query is too big produces the following MaxRetryError with a nested SSLError. There is no way to detect a too big query in the HTTP response status without retying N times and reaching MaxRetryError? And also I have the impression that _tls_no_verify is not passed somewhere in this case, which produces those SSLError. cc @kravets-levko
Details
Versions