Skip to content

feat: Automatically load notebook extras on import with opt-out controls - #199

Open
dborowitz wants to merge 1 commit into
GoogleCloudDataproc:mainfrom
dborowitz:colabsqlviz
Open

dborowitz wants to merge 1 commit into
GoogleCloudDataproc:mainfrom
dborowitz:colabsqlviz

Conversation

@dborowitz

Copy link
Copy Markdown

Automatically initialize interactive notebook extras when importing google.cloud.managed_spark_connect inside an IPython kernel:

  • Inject colabsqlviz's explore_dataframe() into IPython user_ns
  • Load the %dpip line magic extension (google.cloud.managed_spark_magics)
  • Load the %%sparksql cell magic extension (sparksql_magic)
  • Add google-colabsqlviz>=0.3.0 and sparksql-magic>=0.0.3 to dependencies

Add opt-out and runtime configuration controls:

  • Environment variable: MANAGED_SPARK_CONNECT_ENABLE_EXTRAS=false
  • IPython traitlet: ManagedSparkConnect.enable_extras = False (supports both persistent file-based config and runtime toggling via %config, tracking and only undoing changes that managed_spark_connect itself performed).

Dependency & Footprint Justification

The new hard dependencies added by this PR (google-colabsqlviz, sparksql-magic) add just a few MiB of wheel downloads. In more detail:

  1. Zero Version Conflicts or Package Mutations:
    All core data dependencies (pyspark, pandas>=2.0.0, pyarrow>=10.0.1, protobuf>=4.24.0, packaging>=20.0) are already satisfied by pyspark[connect] and google-api-core. Installing both libraries causes 0 upgrades or downgrades to existing packages.
  2. sparksql-magic Has Zero Extra Transitive Cost:
    sparksql-magic (4.2 KiB wheel, 6.9 KiB extracted, 14.4 KiB with .pyc) only depends on pyspark and ipython—a strict subset of google-colabsqlviz and ipykernel. Adding it alongside google-colabsqlviz adds 0 additional transitive dependencies.
  3. Negligible Footprint Relative to the ~840–930 MiB Base:
    In local notebook kernels (e.g., VS Code or Jupyter, which require ipykernel), the entire ipython stack is already present. Adding both libraries requires downloading just 3.59 MiB of wheels (~1.7% of the existing install size dominated by pyspark @ ~435 MiB wheel and pyarrow @ ~130 MiB):
Environment Scenario Base Installed Size Added Wheel Download Added Extracted (no .pyc, e.g. uv) Added Extracted + .pyc (pip) Relative Increase (pip) New Packages Added (Both Libraries)
Local Kernel + Widgets (ipykernel + ipywidgets) ~945 MiB 1.08 MiB 3.88 MiB 5.36 MiB +0.6% 4 (google-colabsqlviz, sparksql-magic, anywidget, psygnal)
Minimal VS Code Kernel (ipykernel only) ~933 MiB 3.59 MiB 13.79 MiB 15.75 MiB +1.7% 7 (above 4 + ipywidgets, jupyterlab-widgets, widgetsnbextension)
Bare Headless venv (no kernel installed) ~840 MiB 11.50 MiB 38.85 MiB 53.50 MiB +6.4% 24 (above 7 + ipython stack)

(Note: Pure Python code across all 7 packages added in the minimal VS Code kernel scenario is < 1 MiB. The remaining extracted space is static frontend JS bundles/source maps in widgetsnbextension and anywidget [~11.5 MiB], psygnal's compiled mypyc .so binary [~1.3 MiB], and pip's .pyc bytecode cache duplicating embedded JS strings in colabsqlviz [~1.0 MiB].)

@dborowitz
dborowitz requested a review from medb September 21, 2026 18:18

@gemini-code-assist gemini-code-assist Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Code Review

This pull request introduces Jupyter Notebook Extras to automatically load interactive conveniences like explore_dataframe(), %dpip, and %%sparksql when importing the package inside an IPython kernel. It includes configuration options to opt out of these extras, updates dependencies, and adds comprehensive tests. The review feedback highlights a potential memory leak in _ipython.py due to strong references to the IPython shell in _SHELL_STATES, recommending the use of weakref.WeakKeyDictionary. Additionally, a minor grammatical redundancy was pointed out in the README.md documentation.

Comment thread google/cloud/managed_spark_connect/_ipython.py
Comment thread README.md Outdated
Automatically initialize interactive notebook extras when importing
google.cloud.managed_spark_connect inside an IPython kernel:

- Inject colabsqlviz's explore_dataframe() into IPython user_ns
- Load the %dpip line magic extension (google.cloud.managed_spark_magics)
- Load the %%sparksql cell magic extension (sparksql_magic)
- Add google-colabsqlviz>=0.3.0 and sparksql-magic>=0.0.3 to dependencies

Add opt-out and runtime configuration controls:

- Environment variable: MANAGED_SPARK_CONNECT_ENABLE_EXTRAS=false
- IPython traitlet: ManagedSparkConnect.enable_extras = False (supports both
  persistent file-based config and runtime toggling via %config, tracking and
  only undoing changes that managed_spark_connect itself performed).
Comment thread setup.py
Comment on lines +35 to +38
# Imported directly by managed_spark_connect._ipython and
# managed_spark_magics; previously these only arrived transitively via
# google-colabsqlviz and sparksql-magic.
"ipython>=8.0",

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I feel like we may want to make all of these optional - we have customers that do not use notebooks w/ Spark Connect, for them these dependencies are not necessary.

@dborowitz dborowitz Sep 25, 2026 •

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The goal is to make the out-of-the box experience as smooth as possible for notebook users.

I agree that we shouldn't introduce performance regressions for non-notebook users. The impact of these changes for non-notebook users are:

  • 11M of extra wheel downloads
  • 0 runtime cost (short-circuits when get_ipython is not importable or returns None)

I'm sympathetic to wanting to avoid the extra wheels, but there is no perfect option:

  1. pip install google-cloud-spark-connect[notebook] is extra wordy, and agents/humans might fail to add the extra
  2. It's not technically possible to make pip install google-cloud-spark-connect[no-notebook] remove the dependencies
  3. Optionally importing them only if available means you need to do pip install google-cloud-spark-connect sparksql-magic google-colabsqlviz [...this list will grow over time...] which is like (1) but worse

On balance, we're saying that the extra wheel download cost is better any of the alternatives 1-3

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I think that it's beyond wheel download, it can cause actual incompatibilities in env, and agents does not need this as well?

I think that to make it practically useful we would need to split it out in the and even exclude pyspark by default:

pip install google-cloud-spark-connect[notebook] pyspark-client 

This is still one line step that users/agents will copy from docs, but it allow us to support all the use cases.

Comment thread setup.py
except Exception:
pass

_init_extras()

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Did we measure latency of this call?

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I can, but it's just loading a couple python modules so I don't imagine it will be more than a few ms.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

We had 20s regression with some AI client libs - I think worth to check this ahead of time.

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Importing is slower than I thought, but still only ~35ms on my laptop. ~27ms of that is importing anywidget. Assuming they are going to run any Spark operations I think this is negligible.

$ t() { MANAGED_SPARK_CONNECT_ENABLE_EXTRAS=0 ipython -c 'import os, time, google.cloud.managed_spark_connect._ipython as m
m._SHELL_STATES.clear(); os.environ["MANAGED_SPARK_CONNECT_ENABLE_EXTRAS"] = "1"
t0 = time.perf_counter(); m._init_extras(); print(f"{(time.perf_counter() - t0) * 1e3:.0f}ms")' 2>/dev/null | tail -1; }
evict() { if [ "$(uname)" = Darwin ]; then sudo purge; else python -c 'import os, site; [os.posix_fadvise(fd := os.open(os.path.join(d, f), os.O_RDONLY), 0, 0, os.POSIX_FADV_DONTNEED) or os.close(fd) for p in site.getsitepackages() for d, _, fs in os.walk(p) for f in fs]'; fi; }
for i in 1 2 3; do echo "hot: $(t)  cold page cache: $(evict; t)  no .pyc: $(PYTHONPYCACHEPREFIX=$(mktemp -d) t)"; done
MANAGED_SPARK_CONNECT_ENABLE_EXTRAS=0 python -X importtime -m IPython -c 'import os, sys, google.cloud.managed_spark_connect._ipython as m
m._SHELL_STATES.clear(); os.environ["MANAGED_SPARK_CONNECT_ENABLE_EXTRAS"] = "1"; print("MARK", file=sys.stderr, flush=True); m._init_extras()' 2>&1 >/dev/null | awk -F'|' 'f && $2 >= 5000 {l[n++] = sprintf("%6.1fms %s", $2 / 1000, $3)} /^MARK$/ {f = 1} END {print "cumulative import time during _init_extras (>=5ms):"; while (n) print l[--n]}'
hot: 47ms  cold page cache: 56ms  no .pyc: 98ms
hot: 36ms  cold page cache: 45ms  no .pyc: 102ms
hot: 35ms  cold page cache: 47ms  no .pyc: 102ms
cumulative import time during _init_extras (>=5ms):
  33.6ms  google.colabsqlviz.explore_dataframe
  31.6ms    google.colabsqlviz.interactive_viz
  27.6ms      anywidget
  20.2ms        anywidget.widget
  19.8ms          ipywidgets
  19.4ms            ipywidgets.widgets
   6.8ms        anywidget._traits
   6.7ms          anywidget._descriptor
   6.0ms            anywidget._file_contents
   5.8ms              psygnal

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants