Skip to content

Bind typed ARRAY/MAP values and LargeBinary correctly - #80

Open
aminghadersohi wants to merge 2 commits into
databricks:mainfrom
aminghadersohi:fix-typed-collection-and-binary-binds
Open

aminghadersohi wants to merge 2 commits into
databricks:mainfrom
aminghadersohi:fix-typed-collection-and-binary-binds

Conversation

@aminghadersohi

@aminghadersohi aminghadersohi commented Sep 26, 2026 •

Copy link
Copy Markdown

What type of PR is this?

  • Bug Fix

Description

Problem

  • ARRAY/MAP. DatabricksArray / DatabricksMap binds hand a Python list/dict to the connector, which sends an ARRAY/MAP parameter. On the SQL warehouse I tested (databricks-sql-connector 4.6.0), the elements of such a parameter were ignored under every Thrift encoding I tried (protocol versions V7–V9), so the row was written with an empty collection and no error (as reported in Unable to insert Python lists into ARRAY<STRING> columns using pandas to_sql #73). This may depend on the connector and warehouse version: the repository's integration run on main (June 2026, locked connector) passed. A NULL collection column cannot be bound at all ('NoneType' object is not iterable).
  • LargeBinary. The bind processor calls the DB-API Binary() constructor, which databricks.sql does not define (AttributeError). Raw bytes are a Sequence, so the connector sends them as an ARRAY, and the warehouse rejects BINARY as a parameter type (INVALID_PARAMETER_MARKER_VALUE.INVALID_DATA_TYPE).

Change

  • A typed ARRAY/MAP value is sent as one JSON STRING and rebuilt with from_json(:p, '<schema>', map('mode', 'FAILFAST')). JSON object keys are strings, so maps parse as MAP<STRING, V> and are CAST to the declared type. FAILFAST makes malformed JSON fail the statement instead of becoming NULL. For maps, the CAST also rejects keys that can't be converted when ANSI mode is on. The SQL does not depend on the element count and uses no lambdas (the warehouse rejects transform inside INSERT ... VALUES), so multi-row VALUES and executemany work. Each element and key first goes through its own type's bind processor (so Uuid, Time and TypeDecorator elements behave as before). Decimals then serialize as plain digits, dates/datetimes as ISO 8601 strings, and None as SQL NULL.
  • SQLAlchemy binary types map (via colspecs) to DatabricksBinary, which binds a hex string decoded with unhex().

How is this tested?

  • Unit tests

  • E2E Tests

  • Manually

  • N/A

  • New tests/test_local/test_collection_binds.py: 4 tests that fail on main and pass here, plus a test that Uuid/Time elements and Uuid map keys still go through their bind processors. The local unit suite (tests/test_local, excluding e2e) passes (304), and black --check src and mypy are clean.

  • Live check on a SQL warehouse (SQLAlchemy 2.0.52, databricks-sql-connector 4.6.0): a Core executemany of ARRAY<BIGINT> [9007199254740993, NULL], MAP<INT,STRING> {7: '雪'}, ARRAY<DECIMAL(38,18)>, all 256 byte values in a BINARY column, plus NULL and empty collections, read back exactly. Date/timestamp elements are covered by unit tests only.

Related Tickets & Documents

Fixes #73.

The warehouse ignores the elements of every Thrift ARRAY/MAP parameter
encoding, so DatabricksArray/DatabricksMap values were persisted as empty
collections without an error, and a NULL collection could not be bound.
Send a typed value as one JSON string rebuilt with
from_json(..., FAILFAST) (maps parsed with string keys and CAST to the
declared type).

LargeBinary binds raised AttributeError (the DB-API module has no
Binary()) and raw bytes were sent as an ARRAY parameter; bind them as a
hex string decoded with unhex().
Serializing collection values to JSON skipped each element type's own
bind processor, so DatabricksArray(Uuid), DatabricksArray(Time) and
DatabricksMap(Uuid, ...) raised TypeError. Run the element's and key's
dialect bind processor before serializing. Numeric has no bind processor
here, so decimals still serialize exactly.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Unable to insert Python lists into ARRAY<STRING> columns using pandas to_sql

1 participant