Visitar URL original
dbt errors since latest update · Issue #140 · databricks/databricks-sql-python · GitHub
Skip to content

dbt errors since latest update #140

Description

@Frunkh

We're getting weird errors in DBT:

image

The errors occur in random models, different every run.

Downgrading to 2.5.2 resolves this issue.

Logs aren't very clear: Databricks adapter: <class 'databricks.sql.exc.RequestError'>: Error during request to server: Remote end closed connection without response

Activity

  1. susodapop commented on Jun 8, 2023

    @susodapop
    Contributor

    Thanks for the report. Do you connect to Databricks through a proxy?

  2. Frunkh commented on Jun 8, 2023

    @Frunkh
    Author

    We do not, just a regular connection to databricks (Azure)

  3. susodapop commented on Jun 8, 2023

    @susodapop
    Contributor

    Please talk to Databricks support so that we can investigate this non-publicly. I'll update this issue with relevant info once we chase down the root cause.

    Version 2.6 of the connector uses urllib3 to and connection pools / http 1.1 which is a pretty big change under the hood. It's not clear yet what causes the exception you found. So we need more information about your environment to debug this properly.

  4. susodapop commented on Jun 8, 2023

    @susodapop
    Contributor

    I've chased this down. It's hard to do with dbt directly because right now the error messages from databricks-sql-connector are not piped to stdout by dbt (we are going to open a PR against dbt-databricks to fix that -- it's a pesky issue that has blocked several debugging attempts for customers).

    The actual error that caused dbt to fall over was:

    INFO     databricks.sql.thrift_backend:thrift_backend.py:260 Error during request to server: {"method": "ExecuteStatement", "session-id": "b'\\x01\\xee\\x06 \\xa8\\xa5\\x14@\\x8c\\xe8\\xf3\\xfe\\xc7\\xff\\xa4\\xa4'", "query-id": null, "http-code": 200, "error-message": "", "original-exception": "('Connection aborted.', BadStatusLine('0\\r\\n'))", "no-retry-reason": "non-retryable error", "bounded-retry-delay": null, "attempt": "1/30", "elapsed-seconds": "0.00203704833984375/900.0"}
    

    After some research, I found that the BadStatusLine error is raised by Python's http.client if it reads an unexpected status code from an HTTP response. The specific bytes it found 0\r\n are the trailing bytes of a response sent with transfer-encoding: chunked. The 0\r\n trailing bytes are meant to indicate to the client that "this is the last chunk". It seems that in some cases these trailing bytes are not consumed by thrift in pysql. And since we release our HTTP connection back to the connection pool after each thrift RPC, it could happen that the next RPC would actually read-in the trailing bytes of the previous RPC while expecting to read the status code and headers of the current RPC. This raised an error.

    The solution is to specifically drain the connection buffer before releasing the connection back to the pool. I've opened #141 to fix this. We will release an update to dbt-databricks that pins this version.

  5. susodapop commented on Jun 9, 2023

    @susodapop
    Contributor

    @Frunkh you can upgrade to dbt-databricks==1.5.4 to capture this update to the databricks-sql-connector.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions