Skip to content

Python 3.13 support 05 release - #781

Draft
kmontemayor2-sc wants to merge 49 commits into
mainfrom
python-3.13-support-05-release
Draft

kmontemayor2-sc wants to merge 49 commits into
mainfrom
python-3.13-support-05-release

Conversation

@kmontemayor2-sc

Copy link
Copy Markdown
Collaborator

Scope of work done

Where is the documentation for this feature?: N/A

Did you add automated tests or write a test plan?

Updated Changelog.md? NO

Ready for code review?: NO

kmonte and others added 30 commits September 2, 2026 22:23
…ilures fatal

Python 3.12 removed `distutils` and the deprecated `unittest` aliases, and GiGL uses
both. The `distutils` imports work today only because `setuptools` is installed
transitively and ships a compatibility shim; a library must not depend on that.

- Add `gigl.common.utils.parse.str_to_bool` and use it in place of
  `distutils.util.strtobool` at all 15 call sites. Accepted and rejected spellings
  match `strtobool`. Every call site already coerced the result with `bool()` or used
  it as a condition, so the `int` to `bool` return change is not observable.
- Rename the 26 `assertEquals` / `assertNotEquals` uses to `assertEqual` /
  `assertNotEqual`, and select ruff `UP005` so they cannot return.
- Make a failed `install_glt.sh` fatal. `main()` returned the child's status but the
  `__main__` block discarded it, so `requirements/install_py_deps.sh` saw exit 0 under
  `set -e` and every base image build continued after a failed GLT install. Measured
  against a stub that exits 7: the old script exits 0, the new one exits 7. A
  successful install still exits 0.
- Run `ty` twice in `make type_check`, at the 3.11 floor and at 3.13. `ty` resolves the
  standard library against one version per invocation, so the floor pass accepts
  modules 3.13 removed and only the ceiling pass rejects them.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
tensorflow-data-validation 1.21.0 publishes only manylinux_2_39_x86_64 wheels, so no
image below glibc 2.39 can take it. The CUDA base was Ubuntu 22.04 (glibc 2.35) and the
Dataflow base was Debian bookworm (glibc 2.36). Both move here, on today's lockfile and
still on Python 3.11, so that when the dependency stack moves the failure mode is
attributable to the stack rather than to the OS or interpreter underneath it.

- CUDA base: `nvidia/cuda:12.8.1-cudnn-devel-ubuntu24.04`, digest pinned. Drops
  `UV_SYSTEM_PYTHON` and the conda interpreter, so the image now builds its own
  `/gigl_deps/.venv` from `.python-version` exactly as the CPU base does. Verified in the
  built image: Python 3.11.14, torch 2.8.0+cu128, CUDA 12.8, glibc 2.39.
- Dataflow base: adopts Beam's custom-container shape — own base OS, the boot harness
  copied from the SDK image, and an explicit ENTRYPOINT, which was previously inherited.
  `RUN_PYTHON_SDK_IN_DEFAULT_ENVIRONMENT=1` is required: without it boot builds a nested
  venv that cannot see this one, and every worker dies reporting apache-beam missing.
- Dataflow images are now built for linux/amd64 only. tensorflow-data-validation ships no
  aarch64 wheel, so a working arm64 image was never possible; the published manifest's
  arm64 entry shares layer digests with amd64.
- Deletes the `--inexact` branch in the installers. It existed only for the two images
  that set `UV_SYSTEM_PYTHON`, and neither does now.
- Hardens `has_cuda_driver()` in both copies. Callers use it as an `if` condition, which
  suspends `set -e`, so a missing `whereis` reported "no CUDA" while the build still
  exited 0. That was survivable while the CUDA base shipped torch preinstalled and
  `--inexact` kept it; without either, it would silently install CPU torch and a
  WITH_CUDA=OFF GLT into a GPU image.
- Adds `scripts/smoke_test_image.py`, which asserts interpreter, ABI tag, active venv,
  glibc floor, a caller-declared import set, and optionally CUDA, Beam version and the
  boot environment variable. The import set is required rather than defaulted: base
  images carry a metadata-only gigl-core, and the Dataflow image has no GLT by design, so
  a fixed list cannot describe every image.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
`libcuda.so.1` and `libnvidia-ml.so.1` are the host NVIDIA driver and ship in no image.
Vertex AI and GKE bind-mount them into /usr/local/nvidia/lib{,64} at container start, so
both directories are empty while the image builds. `nvidia/cuda` sets
LD_LIBRARY_PATH=/usr/local/cuda/lib64 and drops them, which broke GPU two ways at once:

- `fbgemm_gpu` names libcuda.so.1 in DT_NEEDED across 19 of its extensions, so
  `import torchrec` raised outright. Every trainer reaching `gigl.nn` died, because
  `gigl/nn/__init__.py` imports `gigl/nn/models.py` eagerly.
- Everything else imported torch fine, reported no CUDA device, and trained on CPU while
  reporting success. Five of nine e2e pipelines passed that way on two T4s each.

Move to `pytorch/pytorch:2.10.0-cuda12.8-cudnn9-devel`, the oldest torch tag on Ubuntu
24.04, which exports the driver paths as the previous 2.8.0 base did. Its own torch and
Python 3.12 go unused: GiGL pins `requires-python = "==3.11.*"`, so install_py_deps.sh
still builds /gigl_deps/.venv from .python-version, and the venv leads PATH. Set both
variables here anyway rather than inherit them, because upstream shuffles them between
releases: 2.11.0 already puts /usr/local/cuda/lib64 first.

Verified in the built image: Ubuntu 24.04, glibc 2.39, nvcc 12.8, Python 3.11.14 from the
venv, torch 2.8.0+cu128. With the driver directory left empty, `import torchrec` fails as
it did in production; with the host driver mounted into it, the same image imports
torchrec and torch 2.8.0+cu128.

`smoke_test_image.py` gains `torchrec` as a checkable import, so the broken chain can be
asserted directly, and `--require-nvidia-driver-path`, which compares LD_LIBRARY_PATH
entries whole. That check needs no GPU, so it can gate an image build on any machine.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
# Conflicts:
#	gigl/scripts/post_install.py
main added tests/unit/scripts_test/post_install_test.py, which runs the shipped
post_install.py with bash shimmed on PATH and asserts that a failing and a succeeding
install_glt.sh set the process exit code. That is the behaviour this branch's
tests/unit/scripts/post_install_test.py checked. With main's post_install.py taken in the
merge, the second file only duplicates the first.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Moves the TFX family (TFDV, TFMD, TFT, tfx-bsl) from 1.16 to 1.21, TensorFlow from 2.16
to 2.21 and Beam from 2.56 to 2.76, and the Dataflow base's SDK image with it. The lock
resolves pyarrow 25, NumPy 2.4 (via `uv lock --upgrade-package numpy`; a plain re-lock
keeps 1.26), protobuf 6, pandas 2 and kfp 2.17. torch, torchrec, fbgemm-gpu and PyG do
not move. requires-python stays ==3.11.*.

tensorflow-data-validation 1.21 publishes manylinux_2_39 wheels only, which the runtime
images already satisfy.

The bump broke six things, each fixed here:

- TensorFlow 2.21 dropped its TensorBoard dependency, so `tf.summary` raised
  TBNotInstalledError and the v1 node-classification trainer failed on its first epoch.
  Declare `tensorboard~=2.21`. No unit test reaches that path; the full unit suite
  passed without it.
- Beam's Dataflow job object is a proto-plus message with snake_case fields, so reading
  `_job.projectId` raised after every successful Dataflow preprocessing job.
- ApproximateQuantiles.Globally widened its input hint to Union[T, Sequence[T], ...],
  so Beam inferred List[Union[float, List[float]]] for a flat list[float] and rejected
  the next step. Declare the output type the transform really produces.
- NumPy 2.4 refuses to cast the object array that pyarrow returns for one-value list
  columns, which TFT emits for scalar features. Unwrap list, large_list and
  fixed_size_list columns, rejecting nulls and rows without exactly one value.
- DirectRunner can hand pipelines to Prism, which returns from run() while the pipeline
  still executes in-process. Overlapping TFT pipelines then trip TFT's process-global
  object tracker. Wait for every non-Dataflow pipeline inside the existing lock.
- Beam's default pickler is now cloudpickle, which cannot pickle the weakref inside
  jaxtyping's import-hook wrapper, so TFT tests failed under the test-only shape-check
  hook. The hook now leaves functions without shape annotations undecorated; annotated
  functions are checked as before.

Adds an integration test for `pretrained_tft_model_uri`, which no test covered, plus unit
tests for the Dataflow console URI and for the feature-matrix list-column handling.

Validated in a glibc 2.39 image: the full unit suite (1204 tests), and the data
preprocessor, feature quantization, pretrained transform and node-classification
trainer integration tests.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Declare `requires-python = ">=3.11,<3.14"` in both pyproject.toml and
gigl-core/pyproject.toml, with 3.11, 3.12 and 3.13 classifiers on both. The upper bound
holds because TensorFlow publishes no cp314 wheel. The glibc 2.39 floor comes from
tensorflow-data-validation, not the interpreter.

3.13 is the primary: `.python-version` moves to 3.13.15, so the runtime images, the
builder image and the full suites run on it. The Dataflow base takes its boot harness from
apache/beam_python3.13_sdk:2.76.0. Dockerfile.builder now copies `.python-version` to where
uv looks for it; before, it landed in tmp/ and uv picked the interpreter from
requires-python alone.

Both locks are regenerated. Only backports-tarfile and google-apitools differ by minor.
Every locked package has a cp311, cp312 and cp313 wheel, or a pure-Python sdist, on
manylinux_2_39 for both torch variants. torch, torchrec, fbgemm-gpu and PyG do not move.

`setuptools>=70.1` joins gigl-core-build-backend: GLT's `python setup.py bdist_wheel` needs
it, and a fresh uv venv has none.

3.11 and 3.12 get one install-and-import job each, on every PR comment trigger and in the
merge queue. Each runs the production install path under `UV_PYTHON` with
`UV_PYTHON_PREFERENCE=only-managed`, because hosted runners ship a system Python 3.12 that
uv would otherwise pick. `.github/scripts/smoke_test_python_minor.sh` then asserts the
interpreter, ABI tag and venv, imports the native per-minor wheels including torchrec, and
runs three hermetic unit files. The jobs have stable names so they can be made required.

Verified locally: the 3.13.15 stack installs and imports in the CPU base on glibc 2.39,
and the smoke script passes end to end on 3.11.16 and 3.12.14.

This depends on the uv 0.12.9 upgrade: uv 0.9.5 cannot fetch Python 3.13.15. It must land
together with the base-image rebuild and bot commit. The published images move to 3.13 with
it, and the builder image must match `.python-version`.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
pytorch/pytorch:2.10.0-cuda12.8-cudnn9-devel ships uv 0.9.26 at /usr/local/bin/uv.
install_py_deps.sh installs the pinned uv only when no uv is on PATH, so this image
built its venv with the base's uv instead. uv 0.9.26 knows CPython up to 3.13.11, so
with .python-version at 3.13.15 the build failed with "No interpreter found for Python
3.13.15". Earlier images hid this because the base's uv knew 3.11.14.

With it removed, the image installs the pinned uv 0.12.9 and builds a 3.13.15 venv.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
gigl-core is a pybind11 extension that links torch, so each supported interpreter needs
its own wheel per torch variant. release.yml now builds six (cpu and cu128 for each of
3.11, 3.12 and 3.13) and publishes only after all of them pass.

Each build job asserts its interpreter minor, its torch variant (and, for cu128, that the
extension links libtorch_cuda.so), the wheel's cpXY tag and its Requires-Python, and then
imports the wheel in a clean venv. The gigl wheel is built once. Publishing runs one job
per registry, behind the release environment. It is the only job that holds cloud
credentials; the build jobs need none.

A publish that stops partway can be finished with "Re-run failed jobs". That works
because:

- builds are byte-reproducible (SOURCE_DATE_EPOCH is the tagged commit's time), so
  --check-url sees an identical file and skips it;
- the check URL carries the username, so uv authenticates the check against the private
  cu128 index instead of treating its 401 as "absent";
- the two registries publish independently (fail-fast: false).

The workflow refuses any ref that is not the release tag matching pyproject.toml.

create_release.yml re-locks after the version bump. It fails if either lockfile changes
anything beyond the gigl and gigl-core version lines, which catches a moved pin in either
torch variant or in gigl-core's own lock.

RELEASING.md describes the tag dispatch and the new job layout. The CHANGELOG and the
installation docs describe Python 3.11 to 3.13 support, the TFX-library 1.21 / TensorFlow
2.21 / Beam 2.76 stack, the glibc 2.39 image floor, and credentials for the cu128 index.

Validated locally by running every build step in glibc 2.39 containers, with no gcloud
or gsutil present: all six wheels build, pass their checks and import, and repeated
builds are byte-identical. The re-lock gate passes on minor and nightly bumps and fails
on direct, transitive, cu128-only and gigl-core lock changes.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The Dataflow image moved to a single-arch build, and `docker build`
without a platform targets the host's architecture, so an image built on
an arm64 host would not run on the amd64 Dataflow and Vertex AI workers.
Every image built through this path is amd64-only already.

The Beam SDK stage is now pinned by its index digest, as the CUDA base
pins its base, so a re-push of the tag cannot change the copied `boot`.

The CUDA base also removes the uv that pytorch/pytorch ships at
/usr/local/bin, because install_py_deps.sh skips installing the pinned
uv whenever one is already on PATH.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…pport-03-stack-bump

# Conflicts:
#	containers/Dockerfile.dataflow.base
The data preprocessor comment now names Beam's PrismRunner in full so
the reference reads as the Apache Beam runner, and the Dataflow base
smoke-test example checks for the Beam version this branch pins.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…eclare-range

Brings in the Beam SDK digest pin, resolved to the digest of the
beam_python3.13_sdk:2.76.0 index this branch uses.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
tests.unit.main exits 0 when a pattern matches no tests, so a renamed or
mistyped file in smoke_test_python_minor.sh passed silently. Each file
must now report at least one test run.

The smoke jobs run only in the merge queue, and on-pr-comment.yml runs
from main, so neither path can run them on this PR before it merges.
A workflow_dispatch trigger on CI Tests, enabled only for the two smoke
jobs, lets them run once on this branch. Every other job still gates on
merge_group.

The smoke_test_image.py examples for the CUDA and Dataflow bases now use
Python 3.13, which is what both bases install from .python-version.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…5-release

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Re-running failed jobs reuses the run's artifacts, so it is safe without
reproducible builds. The docs now say so, and describe SOURCE_DATE_EPOCH
as a best-effort backstop for a fresh dispatch, where --check-url fails
on a differing file rather than overwriting it.

The build jobs, including the cu128 jobs on the self-hosted GPU pool,
run before the release environment's approval. That is intended: they
hold no cloud credentials, and since workflow_dispatch runs the workflow
file from the dispatched ref, a gate on them would not constrain anyone
who can push a tag. The workflow and RELEASING.md now state this.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
type_check ran ty twice with hard-coded minors (3.11 from pyproject,
3.13 from the Makefile). A single PYTHON_VERSION variable now picks the
interpreter for every uv command and the ty target version, so CI can
run the same targets once per minor (make ... PYTHON_VERSION=3.13).

PYTHON_VERSION defaults to the contents of .python-version and is
exported as UV_PYTHON; uv resolves that to the same interpreter it
picks from .python-version today. type_check passes the major.minor to
ty, which makes [tool.ty.environment] python-version redundant, so it
is removed.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
This branch now only widens the supported range to 3.11-3.13; the
default interpreter stays 3.11. Images for 3.12 and 3.13 are built per
minor in a follow-up, so .python-version, the Dataflow base's Beam SDK
image, the workbench PYTHON_VERSION and the smoke_test_image examples
go back to their 3.11 values, and comments that called 3.13 the primary
now call .python-version the default.

The 3.11/3.12 install-and-import smoke jobs, their script and the
workflow_dispatch trigger that existed only for them are removed. A
unit-test matrix over all three minors replaces them in a follow-up.

The range itself (requires-python, classifiers, locks), setuptools>=70.1,
and the builder's .python-version COPY fix are unchanged.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
type_check no longer has a separate 3.13 pass; the 3.13 check is
make type_check PYTHON_VERSION=3.13.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
kmonte and others added 19 commits September 29, 2026 17:11
The previous branch keeps 3.11 as the default interpreter and drops the
3.11/3.12 smoke jobs, so the CHANGELOG entry and the installation guide
no longer call 3.13 the primary or mention install-and-import tests.
3.12 and 3.13 are supported, gigl-core wheels are published for all
three minors, and consumers on 3.12 or 3.13 build their own images
until published per-minor images follow.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The workbench Dockerfile declared ARG PYTHON_VERSION="3.9". Docker exposes
a build ARG to RUN steps as an environment variable, and the Makefile now
reads PYTHON_VERSION to choose the interpreter, so `make install_deps` in
that image would ask uv for Python 3.9, which GiGL does not support. The
conda env and GiGL now use the same supported minor.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The comment explaining why non-Dataflow pipelines are waited on inside
the lock nested a relative clause mid-sentence; it now states each fact
on its own.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Describe the tfx-bsl pin by the constraint it lives under rather than
relative to "the latest" release, which goes stale, and state the Beam
SDK reference's coupling plainly.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…pport-03-stack-bump

# Conflicts:
#	containers/Dockerfile.dataflow.base
#	pyproject.toml
…port-04-declare-range

# Conflicts:
#	containers/Dockerfile.dataflow.base
The CHANGELOG entry packed the dependency stack into one paragraph; a
nested list makes each piece scannable. The workflow comments and the
Docker-image docs say the same things in fewer words, and the docs no
longer promise future images.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@semgrep-code-snapchat

Copy link
Copy Markdown

Semgrep found 6 github-actions-mutable-action-tag findings:

🟡 Medium severity issue identified in your code:

GitHub Actions step uses a mutable tag or branch reference. Tags and branch names can be silently repointed by the action owner, enabling supply-chain attacks — as seen in the trivy-action and kics-github-action compromises. Pin the reference to a full 40-character commit SHA instead, e.g. uses: actions/checkout@8ade135a41bc03ea155e62e844d188df1ea18608.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants