Reflect nested ARRAY/MAP types and tolerate unknown column types - #78
Open
aminghadersohi wants to merge 2 commits into
Open
aminghadersohi wants to merge 2 commits into
aminghadersohi wants to merge 2 commits into
Conversation
get_columns mapped only the first word of TYPE_NAME: ARRAY and MAP columns reflected as String, and any type missing from GET_COLUMNS_TYPE_MAP (e.g. TIME, INTERVAL, VOID) raised KeyError and failed reflection of the whole table or view. Parse TYPE_NAME recursively into DatabricksArray/DatabricksMap and reflect unrecognised types as NullType with a warning, as other SQLAlchemy dialects do.
aminghadersohi
requested review from
deeksha-db,
gopalldb,
jackyhu-db,
jayantsing-db,
jprakash-db,
samikshya-db,
shivam2680 and
vikrantpuppala
as code owners
September 26, 2026 09:53
A nested DECIMAL element spelled with a space or in lower case, e.g. ARRAY<DECIMAL(10, 2)>, raised AttributeError and failed reflection of the whole table. Match case-insensitively, tolerate whitespace, and fall back to Numeric() for a bare DECIMAL.
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
What type of PR is this?
Description
Problem
parse_column_info_from_tgetcolumnsresponsemaps only the first word ofTYPE_NAME:ARRAY<...>andMAP<...>columns reflect asString, so reflected/autoloaded tables lose their element types (and tools built on reflection report e.g. anARRAY<STRING>column asSTRING).GET_COLUMNS_TYPE_MAPraisesKeyError, which fails reflection of the whole table or view. Observed on a SQL warehouse (DBSQL 2026.36): a table with aTIME(6)column →KeyError: 'time'; a view selecting anINTERVALliteral →KeyError: 'interval'; aNULLcolumn in a view (VOID) →KeyError: 'void'.Change
parse_type_name()parsesTYPE_NAMErecursively:ARRAY<T>→DatabricksArray(T),MAP<K, V>→DatabricksMap(K, V)(nested at any depth, includingDECIMAL(p,s)elements).DECIMAL(p,s)is parsed in any case and with any spacing (DECIMAL(10, 2),decimal(10,2)), and a bareDECIMALbecomesNumeric(). Before, those spellings raisedAttributeError, which nested parsing would have turned into a whole-table reflection failure.NullTypewith anSAWarning, the convention SQLAlchemy dialects use, instead of raising.Stringmapping (no STRUCT type exists in the dialect yet).How is this tested?
New cases in
tests/test_local/test_parsing.py(nested parsing, DECIMAL spellings, and TIME/INTERVAL/VOID/GEOMETRY no longer raising). They fail onmainand pass with this change. The local unit suite (tests/test_local, excluding e2e) passes (314), andblack --check srcand mypy are clean.Related Tickets & Documents
None.