Python 3.13 support 05 release - #781
Draft
kmontemayor2-sc wants to merge 49 commits into
Draft
kmontemayor2-sc wants to merge 49 commits into
kmontemayor2-sc wants to merge 49 commits into
Conversation
…ilures fatal Python 3.12 removed `distutils` and the deprecated `unittest` aliases, and GiGL uses both. The `distutils` imports work today only because `setuptools` is installed transitively and ships a compatibility shim; a library must not depend on that. - Add `gigl.common.utils.parse.str_to_bool` and use it in place of `distutils.util.strtobool` at all 15 call sites. Accepted and rejected spellings match `strtobool`. Every call site already coerced the result with `bool()` or used it as a condition, so the `int` to `bool` return change is not observable. - Rename the 26 `assertEquals` / `assertNotEquals` uses to `assertEqual` / `assertNotEqual`, and select ruff `UP005` so they cannot return. - Make a failed `install_glt.sh` fatal. `main()` returned the child's status but the `__main__` block discarded it, so `requirements/install_py_deps.sh` saw exit 0 under `set -e` and every base image build continued after a failed GLT install. Measured against a stub that exits 7: the old script exits 0, the new one exits 7. A successful install still exits 0. - Run `ty` twice in `make type_check`, at the 3.11 floor and at 3.13. `ty` resolves the standard library against one version per invocation, so the floor pass accepts modules 3.13 removed and only the ceiling pass rejects them. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
tensorflow-data-validation 1.21.0 publishes only manylinux_2_39_x86_64 wheels, so no image below glibc 2.39 can take it. The CUDA base was Ubuntu 22.04 (glibc 2.35) and the Dataflow base was Debian bookworm (glibc 2.36). Both move here, on today's lockfile and still on Python 3.11, so that when the dependency stack moves the failure mode is attributable to the stack rather than to the OS or interpreter underneath it. - CUDA base: `nvidia/cuda:12.8.1-cudnn-devel-ubuntu24.04`, digest pinned. Drops `UV_SYSTEM_PYTHON` and the conda interpreter, so the image now builds its own `/gigl_deps/.venv` from `.python-version` exactly as the CPU base does. Verified in the built image: Python 3.11.14, torch 2.8.0+cu128, CUDA 12.8, glibc 2.39. - Dataflow base: adopts Beam's custom-container shape — own base OS, the boot harness copied from the SDK image, and an explicit ENTRYPOINT, which was previously inherited. `RUN_PYTHON_SDK_IN_DEFAULT_ENVIRONMENT=1` is required: without it boot builds a nested venv that cannot see this one, and every worker dies reporting apache-beam missing. - Dataflow images are now built for linux/amd64 only. tensorflow-data-validation ships no aarch64 wheel, so a working arm64 image was never possible; the published manifest's arm64 entry shares layer digests with amd64. - Deletes the `--inexact` branch in the installers. It existed only for the two images that set `UV_SYSTEM_PYTHON`, and neither does now. - Hardens `has_cuda_driver()` in both copies. Callers use it as an `if` condition, which suspends `set -e`, so a missing `whereis` reported "no CUDA" while the build still exited 0. That was survivable while the CUDA base shipped torch preinstalled and `--inexact` kept it; without either, it would silently install CPU torch and a WITH_CUDA=OFF GLT into a GPU image. - Adds `scripts/smoke_test_image.py`, which asserts interpreter, ABI tag, active venv, glibc floor, a caller-declared import set, and optionally CUDA, Beam version and the boot environment variable. The import set is required rather than defaulted: base images carry a metadata-only gigl-core, and the Dataflow image has no GLT by design, so a fixed list cannot describe every image. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
`libcuda.so.1` and `libnvidia-ml.so.1` are the host NVIDIA driver and ship in no image.
Vertex AI and GKE bind-mount them into /usr/local/nvidia/lib{,64} at container start, so
both directories are empty while the image builds. `nvidia/cuda` sets
LD_LIBRARY_PATH=/usr/local/cuda/lib64 and drops them, which broke GPU two ways at once:
- `fbgemm_gpu` names libcuda.so.1 in DT_NEEDED across 19 of its extensions, so
`import torchrec` raised outright. Every trainer reaching `gigl.nn` died, because
`gigl/nn/__init__.py` imports `gigl/nn/models.py` eagerly.
- Everything else imported torch fine, reported no CUDA device, and trained on CPU while
reporting success. Five of nine e2e pipelines passed that way on two T4s each.
Move to `pytorch/pytorch:2.10.0-cuda12.8-cudnn9-devel`, the oldest torch tag on Ubuntu
24.04, which exports the driver paths as the previous 2.8.0 base did. Its own torch and
Python 3.12 go unused: GiGL pins `requires-python = "==3.11.*"`, so install_py_deps.sh
still builds /gigl_deps/.venv from .python-version, and the venv leads PATH. Set both
variables here anyway rather than inherit them, because upstream shuffles them between
releases: 2.11.0 already puts /usr/local/cuda/lib64 first.
Verified in the built image: Ubuntu 24.04, glibc 2.39, nvcc 12.8, Python 3.11.14 from the
venv, torch 2.8.0+cu128. With the driver directory left empty, `import torchrec` fails as
it did in production; with the host driver mounted into it, the same image imports
torchrec and torch 2.8.0+cu128.
`smoke_test_image.py` gains `torchrec` as a checkable import, so the broken chain can be
asserted directly, and `--require-nvidia-driver-path`, which compares LD_LIBRARY_PATH
entries whole. That check needs no GPU, so it can gate an image build on any machine.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
# Conflicts: # gigl/scripts/post_install.py
main added tests/unit/scripts_test/post_install_test.py, which runs the shipped post_install.py with bash shimmed on PATH and asserts that a failing and a succeeding install_glt.sh set the process exit code. That is the behaviour this branch's tests/unit/scripts/post_install_test.py checked. With main's post_install.py taken in the merge, the second file only duplicates the first. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Moves the TFX family (TFDV, TFMD, TFT, tfx-bsl) from 1.16 to 1.21, TensorFlow from 2.16 to 2.21 and Beam from 2.56 to 2.76, and the Dataflow base's SDK image with it. The lock resolves pyarrow 25, NumPy 2.4 (via `uv lock --upgrade-package numpy`; a plain re-lock keeps 1.26), protobuf 6, pandas 2 and kfp 2.17. torch, torchrec, fbgemm-gpu and PyG do not move. requires-python stays ==3.11.*. tensorflow-data-validation 1.21 publishes manylinux_2_39 wheels only, which the runtime images already satisfy. The bump broke six things, each fixed here: - TensorFlow 2.21 dropped its TensorBoard dependency, so `tf.summary` raised TBNotInstalledError and the v1 node-classification trainer failed on its first epoch. Declare `tensorboard~=2.21`. No unit test reaches that path; the full unit suite passed without it. - Beam's Dataflow job object is a proto-plus message with snake_case fields, so reading `_job.projectId` raised after every successful Dataflow preprocessing job. - ApproximateQuantiles.Globally widened its input hint to Union[T, Sequence[T], ...], so Beam inferred List[Union[float, List[float]]] for a flat list[float] and rejected the next step. Declare the output type the transform really produces. - NumPy 2.4 refuses to cast the object array that pyarrow returns for one-value list columns, which TFT emits for scalar features. Unwrap list, large_list and fixed_size_list columns, rejecting nulls and rows without exactly one value. - DirectRunner can hand pipelines to Prism, which returns from run() while the pipeline still executes in-process. Overlapping TFT pipelines then trip TFT's process-global object tracker. Wait for every non-Dataflow pipeline inside the existing lock. - Beam's default pickler is now cloudpickle, which cannot pickle the weakref inside jaxtyping's import-hook wrapper, so TFT tests failed under the test-only shape-check hook. The hook now leaves functions without shape annotations undecorated; annotated functions are checked as before. Adds an integration test for `pretrained_tft_model_uri`, which no test covered, plus unit tests for the Dataflow console URI and for the feature-matrix list-column handling. Validated in a glibc 2.39 image: the full unit suite (1204 tests), and the data preprocessor, feature quantization, pretrained transform and node-classification trainer integration tests. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Declare `requires-python = ">=3.11,<3.14"` in both pyproject.toml and gigl-core/pyproject.toml, with 3.11, 3.12 and 3.13 classifiers on both. The upper bound holds because TensorFlow publishes no cp314 wheel. The glibc 2.39 floor comes from tensorflow-data-validation, not the interpreter. 3.13 is the primary: `.python-version` moves to 3.13.15, so the runtime images, the builder image and the full suites run on it. The Dataflow base takes its boot harness from apache/beam_python3.13_sdk:2.76.0. Dockerfile.builder now copies `.python-version` to where uv looks for it; before, it landed in tmp/ and uv picked the interpreter from requires-python alone. Both locks are regenerated. Only backports-tarfile and google-apitools differ by minor. Every locked package has a cp311, cp312 and cp313 wheel, or a pure-Python sdist, on manylinux_2_39 for both torch variants. torch, torchrec, fbgemm-gpu and PyG do not move. `setuptools>=70.1` joins gigl-core-build-backend: GLT's `python setup.py bdist_wheel` needs it, and a fresh uv venv has none. 3.11 and 3.12 get one install-and-import job each, on every PR comment trigger and in the merge queue. Each runs the production install path under `UV_PYTHON` with `UV_PYTHON_PREFERENCE=only-managed`, because hosted runners ship a system Python 3.12 that uv would otherwise pick. `.github/scripts/smoke_test_python_minor.sh` then asserts the interpreter, ABI tag and venv, imports the native per-minor wheels including torchrec, and runs three hermetic unit files. The jobs have stable names so they can be made required. Verified locally: the 3.13.15 stack installs and imports in the CPU base on glibc 2.39, and the smoke script passes end to end on 3.11.16 and 3.12.14. This depends on the uv 0.12.9 upgrade: uv 0.9.5 cannot fetch Python 3.13.15. It must land together with the base-image rebuild and bot commit. The published images move to 3.13 with it, and the builder image must match `.python-version`. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
pytorch/pytorch:2.10.0-cuda12.8-cudnn9-devel ships uv 0.9.26 at /usr/local/bin/uv. install_py_deps.sh installs the pinned uv only when no uv is on PATH, so this image built its venv with the base's uv instead. uv 0.9.26 knows CPython up to 3.13.11, so with .python-version at 3.13.15 the build failed with "No interpreter found for Python 3.13.15". Earlier images hid this because the base's uv knew 3.11.14. With it removed, the image installs the pinned uv 0.12.9 and builds a 3.13.15 venv. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
gigl-core is a pybind11 extension that links torch, so each supported interpreter needs its own wheel per torch variant. release.yml now builds six (cpu and cu128 for each of 3.11, 3.12 and 3.13) and publishes only after all of them pass. Each build job asserts its interpreter minor, its torch variant (and, for cu128, that the extension links libtorch_cuda.so), the wheel's cpXY tag and its Requires-Python, and then imports the wheel in a clean venv. The gigl wheel is built once. Publishing runs one job per registry, behind the release environment. It is the only job that holds cloud credentials; the build jobs need none. A publish that stops partway can be finished with "Re-run failed jobs". That works because: - builds are byte-reproducible (SOURCE_DATE_EPOCH is the tagged commit's time), so --check-url sees an identical file and skips it; - the check URL carries the username, so uv authenticates the check against the private cu128 index instead of treating its 401 as "absent"; - the two registries publish independently (fail-fast: false). The workflow refuses any ref that is not the release tag matching pyproject.toml. create_release.yml re-locks after the version bump. It fails if either lockfile changes anything beyond the gigl and gigl-core version lines, which catches a moved pin in either torch variant or in gigl-core's own lock. RELEASING.md describes the tag dispatch and the new job layout. The CHANGELOG and the installation docs describe Python 3.11 to 3.13 support, the TFX-library 1.21 / TensorFlow 2.21 / Beam 2.76 stack, the glibc 2.39 image floor, and credentials for the cu128 index. Validated locally by running every build step in glibc 2.39 containers, with no gcloud or gsutil present: all six wheels build, pass their checks and import, and repeated builds are byte-identical. The re-lock gate passes on minor and nightly bumps and fails on direct, transitive, cu128-only and gigl-core lock changes. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The Dataflow image moved to a single-arch build, and `docker build` without a platform targets the host's architecture, so an image built on an arm64 host would not run on the amd64 Dataflow and Vertex AI workers. Every image built through this path is amd64-only already. The Beam SDK stage is now pinned by its index digest, as the CUDA base pins its base, so a re-push of the tag cannot change the copied `boot`. The CUDA base also removes the uv that pytorch/pytorch ships at /usr/local/bin, because install_py_deps.sh skips installing the pinned uv whenever one is already on PATH. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…pport-03-stack-bump # Conflicts: # containers/Dockerfile.dataflow.base
The data preprocessor comment now names Beam's PrismRunner in full so the reference reads as the Apache Beam runner, and the Dataflow base smoke-test example checks for the Beam version this branch pins. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…eclare-range Brings in the Beam SDK digest pin, resolved to the digest of the beam_python3.13_sdk:2.76.0 index this branch uses. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
tests.unit.main exits 0 when a pattern matches no tests, so a renamed or mistyped file in smoke_test_python_minor.sh passed silently. Each file must now report at least one test run. The smoke jobs run only in the merge queue, and on-pr-comment.yml runs from main, so neither path can run them on this PR before it merges. A workflow_dispatch trigger on CI Tests, enabled only for the two smoke jobs, lets them run once on this branch. Every other job still gates on merge_group. The smoke_test_image.py examples for the CUDA and Dataflow bases now use Python 3.13, which is what both bases install from .python-version. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…5-release Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Re-running failed jobs reuses the run's artifacts, so it is safe without reproducible builds. The docs now say so, and describe SOURCE_DATE_EPOCH as a best-effort backstop for a fresh dispatch, where --check-url fails on a differing file rather than overwriting it. The build jobs, including the cu128 jobs on the self-hosted GPU pool, run before the release environment's approval. That is intended: they hold no cloud credentials, and since workflow_dispatch runs the workflow file from the dispatched ref, a gate on them would not constrain anyone who can push a tag. The workflow and RELEASING.md now state this. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
type_check ran ty twice with hard-coded minors (3.11 from pyproject, 3.13 from the Makefile). A single PYTHON_VERSION variable now picks the interpreter for every uv command and the ty target version, so CI can run the same targets once per minor (make ... PYTHON_VERSION=3.13). PYTHON_VERSION defaults to the contents of .python-version and is exported as UV_PYTHON; uv resolves that to the same interpreter it picks from .python-version today. type_check passes the major.minor to ty, which makes [tool.ty.environment] python-version redundant, so it is removed. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…port-04-declare-range
…pport-03-stack-bump
This branch now only widens the supported range to 3.11-3.13; the default interpreter stays 3.11. Images for 3.12 and 3.13 are built per minor in a follow-up, so .python-version, the Dataflow base's Beam SDK image, the workbench PYTHON_VERSION and the smoke_test_image examples go back to their 3.11 values, and comments that called 3.13 the primary now call .python-version the default. The 3.11/3.12 install-and-import smoke jobs, their script and the workflow_dispatch trigger that existed only for them are removed. A unit-test matrix over all three minors replaces them in a follow-up. The range itself (requires-python, classifiers, locks), setuptools>=70.1, and the builder's .python-version COPY fix are unchanged. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
type_check no longer has a separate 3.13 pass; the 3.13 check is make type_check PYTHON_VERSION=3.13. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…support-05-release
…pport-03-stack-bump
…port-04-declare-range
The previous branch keeps 3.11 as the default interpreter and drops the 3.11/3.12 smoke jobs, so the CHANGELOG entry and the installation guide no longer call 3.13 the primary or mention install-and-import tests. 3.12 and 3.13 are supported, gigl-core wheels are published for all three minors, and consumers on 3.12 or 3.13 build their own images until published per-minor images follow. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The workbench Dockerfile declared ARG PYTHON_VERSION="3.9". Docker exposes a build ARG to RUN steps as an environment variable, and the Makefile now reads PYTHON_VERSION to choose the interpreter, so `make install_deps` in that image would ask uv for Python 3.9, which GiGL does not support. The conda env and GiGL now use the same supported minor. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…pport-03-stack-bump
…port-04-declare-range
…support-05-release
The comment explaining why non-Dataflow pipelines are waited on inside the lock nested a relative clause mid-sentence; it now states each fact on its own. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Describe the tfx-bsl pin by the constraint it lives under rather than relative to "the latest" release, which goes stale, and state the Beam SDK reference's coupling plainly. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…pport-03-stack-bump # Conflicts: # containers/Dockerfile.dataflow.base # pyproject.toml
…port-04-declare-range # Conflicts: # containers/Dockerfile.dataflow.base
The CHANGELOG entry packed the dependency stack into one paragraph; a nested list makes each piece scannable. The workflow comments and the Docker-image docs say the same things in fewer words, and the docs no longer promise future images. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…support-05-release
…pport-03-stack-bump
…port-04-declare-range
…support-05-release
|
Semgrep found 6
🟡 Medium severity issue identified in your code: GitHub Actions step uses a mutable tag or branch reference. Tags and branch names can be silently repointed by the action owner, enabling supply-chain attacks — as seen in the trivy-action and kics-github-action compromises. Pin the reference to a full 40-character commit SHA instead, e.g. |
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Scope of work done
Where is the documentation for this feature?: N/A
Did you add automated tests or write a test plan?
Updated Changelog.md? NO
Ready for code review?: NO