Skip to content
Open
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
31 changes: 10 additions & 21 deletions DEVELOPING.md
Original file line number Diff line number Diff line change
Expand Up @@ -39,27 +39,16 @@ env \
pytest --tb=auto -v
```

## Testing with Magic Support

To run tests with magic functionality, install the required dependencies manually:

```sh
pip install .
pip install IPython sparksql-magic
```

Then run tests as normal. Any magic-related tests will automatically detect and use the available dependencies.

## Testing without Magic Support

To run tests without the magic dependencies, simply install the base package:

```sh
pip install .
pytest
```

Tests that require magic functionality will be automatically skipped if the dependencies are not available.
## Testing the Interactive Extras

The dependencies of the interactive extras (`explore_dataframe`, `%dpip`,
`%%sparksql`) live in the `interactive` extra
(`pip install '.[interactive]'`), and `requirements-dev.txt` already includes
them. The tests in
`tests/unit/test_init.py` assume `google-colabsqlviz`, `sparksql-magic`,
`ipython` and `traitlets` are importable and will fail, not skip, if they are
not. Missing dependencies are simulated within the tests themselves, by
setting `sys.modules[name] = None`.

The integration tests in particular can take a while to run. To speed up the
testing cycle, you can run them in parallel. You can do so using the `xdist`
Expand Down
81 changes: 50 additions & 31 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -121,53 +121,72 @@ To create or connect to a named session:

5. A session with a given ID that is in a TERMINATED state cannot be reused. It must be deleted before a new session with the same ID can be created.

### Using Spark SQL Magic Commands (Jupyter Notebooks)
### Interactive Extras

The package supports the [sparksql-magic](https://github.com/cryeo/sparksql-magic) library for executing Spark SQL queries directly in Jupyter notebooks.
For notebooks and other interactive use, install the `interactive` extra:

**Installation**: To use magic commands, install the required dependencies manually:
```bash
pip install google-cloud-spark-connect
pip install IPython sparksql-magic
```sh
pip install 'google-cloud-spark-connect[interactive]'
```

1. Load the magic extension:
```python
%load_ext sparksql_magic
```
When you import the package inside an IPython kernel, it then automatically sets
up a few interactive conveniences:

- `explore_dataframe()` from
[google-colabsqlviz](https://pypi.org/project/google-colabsqlviz/) is injected
into your notebook globals.
- The `%dpip` line magic is loaded.
- The `%%sparksql` cell magic from
[sparksql-magic](https://github.com/cryeo/sparksql-magic) is loaded.

```python
import google.cloud.managed_spark_connect # extras load here
```

Each extra is only set up if its dependency is importable, so without the
`interactive` extra you get whichever of them your environment already happens
to provide. On import, a single line lists the extras that were loaded.

The extras won't override anything you've already set up, for example if `explore_dataframe` is already present, or `%%sparksql` magic is loaded from somewhere else, these are left alone.

#### Opting out

Set the environment variable before starting the kernel:

```sh
export MANAGED_SPARK_CONNECT_ENABLE_EXTRAS=false
```

Alternatively, configure it through IPython, either persistently in
`~/.ipython/profile_default/ipython_config.py`:

2. Configure default settings (optional):
```python
c.ManagedSparkConnect.enable_extras = False
```

Or at runtime:

```python
%config ManagedSparkConnect.enable_extras = False
```

An explicit IPython setting takes precedence over the environment variable.

#### Using `%%sparksql`

1. Configure default settings (optional):
```python
%config SparkSql.limit=20
```

3. Execute SQL queries:
2. Execute SQL queries:
```python
%%sparksql
SELECT * FROM your_table
```

4. Advanced usage with options:
```python
# Cache results and create a view
%%sparksql --cache --view result_view df
SELECT * FROM your_table WHERE condition = true
```

Available options:
- `--cache` / `-c`: Cache the DataFrame
- `--eager` / `-e`: Cache with eager loading
- `--view VIEW` / `-v VIEW`: Create a temporary view
- `--limit N` / `-l N`: Override default row display limit
- `variable_name`: Store result in a variable

See [sparksql-magic](https://github.com/cryeo/sparksql-magic) for more examples.

**Note**: Magic commands are optional. If you only need basic ManagedSparkSession functionality without Jupyter magic support, install only the base package:
```bash
pip install google-cloud-spark-connect
```

## Migrating from dataproc-spark-connect

The `dataproc-spark-connect` package has been renamed to `google-cloud-spark-connect`. This is a breaking change with no compatibility shims — you need to update your code in the following places when you switch to the new package.
Expand Down
12 changes: 12 additions & 0 deletions google/cloud/managed_spark_connect/__init__.py
Original file line number Diff line number Diff line change
Expand Up @@ -12,6 +12,7 @@
# See the License for the specific language governing permissions and
# limitations under the License.
import importlib.metadata
import importlib.util
import warnings

from .session import ManagedSparkSession
Expand All @@ -28,3 +29,14 @@
)
except Exception:
pass

# traitlets, like every other dependency of the interactive extras, only comes
# with the [interactive] extra. IPython depends on it, so without it there can
# be no shell and nothing to initialize.
if importlib.util.find_spec("traitlets") is not None:
from ._ipython import (
ManagedSparkConnect,
_init_extras,
)

_init_extras()
Loading
Loading