You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
{{ message }}
Repository navigation
DOCS-3001: Add an upgrade notes page for 3.33 - #3044
Adds operations/upgrading/upgrade-notes.mdx and moves the upgrade notes off the release notes.
Fifteen notes, action-required first, grouped by the release that introduced them. Each says who it affects, named as something the reader can check, what happens if they do nothing, then numbered steps in upgrade order.
The three upgrade guides name the page as the first thing to read.
The release notes keep a short pointer in both places an upgrade comes up.
The guides are scoped to the two previous releases. They advertised v3.15 on Kubernetes and v3.0 on OpenStack. The supported range is now stated through a shared partial, with no release numbers to maintain.
Removed as below that floor: the Upgrade OwnerReferences section, which starts at v3.28 and was duplicated in the Kubernetes and OpenShift guides, and three steps about cleaning up a temporary allow-all-upgrade policy from before v3.14. Those three were already broken, since they told the reader to undo pre-upgrade steps that are not on the page and the component behind them is imported nowhere. That component is deleted too.
The OpenStack policy data migration stays, because it applies to upgrades from earlier than v3.32 and 3.31 is a supported source.
Collect the changes that need action at upgrade time onto one page, rather
than leaving them in the release notes where a reader upgrading across
several releases would have to find them one page at a time. Grouped by
the release that introduced them, so skipping releases means reading every
group in between.
Every note follows the same shape: who it affects, named as something the
reader can check; what changes and what happens if they do nothing; then
numbered steps in upgrade order. Action-required notes come first.
Nine notes for 3.33. Four are new from vetting the build PR discussion:
the single calico image, the new policy rule and selector limits, the
eBPF kernel floor, and the Felix metrics client auth default. Three move
off the release notes. FIPS removal and the Envoy Gateway Kubernetes
floor are drawn from the changelog.
The three upgrade guides now name the page as the first thing to read,
and the release notes point at it in both places an upgrade comes up.
Document only the upgrade paths we support. For 3.33 that is 3.31 and
3.32, so raise the floors the guides advertise, from v3.15 on Kubernetes
and v3.0 on OpenStack, and say the supported range on every guide through
a shared partial, so a release cut updates it in one place.
Remove what only applied below that floor:
- Upgrade OwnerReferences, which starts at v3.28, and was duplicated
verbatim in the Kubernetes and OpenShift guides.
- Three steps about cleaning up a temporary allow-all-upgrade policy after
upgrading from earlier than v3.14. These were already broken: each told
the reader to undo pre-upgrade steps "above" that no longer exist,
because the component that creates that policy is imported by no page in
3.33 or 3.32. Delete the orphaned component with them.
The OpenStack policy data migration stays. It applies to upgrades from
earlier than v3.32, and 3.31 is a supported source.
The required pin update is not actionable with the accompanying 3.33 documentation: operations/image-options/imageset.mdx:87-96 still shows the removed component images, and its generator's allowlist at line 164 omits calico/calico. An operator following that documented workflow will not pin the consolidated image and can fail in a digest-only/private-registry setup. Update those 3.33 ImageSet examples and the script for the consolidated image, and link the corrected procedure here.
FIPS installations must disable fipsMode before upgrading
Moving off the image variant is not sufficient for operator-managed FIPS installations. The 3.33 Installation API states that leaving fipsMode: Enabled marks the installation degraded (reference/installation/_api.mdx:2016), so this required-action step must also tell users to disable that field before upgrading.
Name no release numbers in the support statement. Hardcoding 3.31 and
3.32 puts a maintenance burden on every cut, and the numbers go stale
silently. The partial and the guides now say the range relative to the
release, and the surrounding sentences no longer pin a floor either.
Drop the claim that upgrading from an earlier release means going to the
previous one first. That is not true; $[prodname] has always documented
upgrades across several releases, and N-2 scopes what we document rather
than what is possible.
Plain text rather than an admonition, merged with each guide's opening so
the page does not introduce itself twice.
The OpenStack policy data migration still names v3.32, which is correct.
That is a behavior change in a specific release, not a support boundary.
This replacement list is incomplete. The consolidated binary also absorbs the previously published calico/dikastes, calico/csi, calico/node-driver-registrar, calico/pod2daemon-flexvol, calico/key-cert-provisioner, calico/whisker-backend, and flannel-migration-controller components; none appears in the 3.33 release image set. Readers who use this list to update mirrors, pins, or scanners can therefore leave stale references that will not pull.
Changing only the image reference is insufficient for custom workloads that invoked an old component image directly. calico/calico has /usr/bin/calico as its entrypoint and requires component <name> before the component's old arguments; for example, the documented Dikastes sidecar currently uses args: ["server", ...], which becomes an unknown top-level command after swapping the image. The upgrade steps need to call out this command migration or pinned/custom manifests will still fail.
Missing Gateway API CRD v1.6.1 upgrade prerequisite
The Envoy Gateway 1.9 upgrade also requires Gateway API CRDs v1.6.1 to be installed first; upstream explicitly marks this as an upgrade prerequisite because 1.9 reconciles TCPRoute and UDPRoute through gateway.networking.k8s.io/v1. The operator's default PreferExisting mode does not update existing CRDs, so these steps can leave Envoy Gateway unable to start or reconcile routes even after the Kubernetes version is upgraded.
Minimum Kubernetes version conflicts with native CRD prerequisites
This minimum conflicts with the 3.33 native-v3-CRD prerequisites and installation flows, which require Kubernetes 1.34 or later; they also require the MutatingAdmissionPolicy feature gate on 1.34 and 1.35. As written, readers may attempt a default native-CRD install on 1.32/1.33 even though the documented manifests do not support that path.
Fixes found by reviewing the page against the pages and CRDs it describes,
rather than against the release notes it was drafted from.
Wrong, and would have cost a reader real time:
- FIPS. Setting fipsMode to Enabled marks the installation degraded in
3.33, and that field is how an operator-managed cluster runs the -fips
images. The note reassured the reader the field was staying and told
them only to move off image tags, which is not the action that helps.
- Native v3 CRDs need Kubernetes 1.34, not 1.32, and on 1.34 and 1.35
they also need the MutatingAdmissionPolicy feature gate, which is not
on by default.
- The single image note named 6 of the 14 images that stopped being
published. The missing 8 include csi, node-driver-registrar and
key-cert-provisioner, which are in the calico-node pod by default, so
a mirrored registry would still have failed to pull.
- Policy limits are ratcheted by Kubernetes, so an update that does not
touch an over-limit field is accepted. Saying the policy is rejected by
any update sends people into unnecessary policy splitting. The limits
also do not apply on etcdv3, which has no CRDs.
- The eBPF note contradicted both eBPF pages, which support Red Hat 8.4
on kernel 4.18.0-305 by backport.
- The crd-migration pointer called the migration briefly locked. That
page asks for a maintenance window.
- The Ingress Gateway note asserted the Kubernetes versions $[prodname]
supports. The 3.33 and 3.32 requirements pages disagree on that, so the
note no longer makes the claim at all.
Added, all of them changes that need action and had no coverage:
- BPFAttachType defaults to Netkit, which must be changed before rolling
back to a release without netkit support.
- eBPF overlay traffic now uses the node's main IP, so rules matching a
tunnel address stop matching.
- The default CNI configuration requires containerd v1.6 or CRI-O v1.24.
- Overlapping IP pools are rejected on write.
Also: an Upgrading to 3.32 group, because the page tells readers to read
every group between their version and the target, N-2 makes 3.31 a
supported source, and the OpenStack policy-name migration lives there.
Steps that named no command now name one. The OpenShift description no
longer advertises the OwnerReferences section this branch removed.
Candidates for upgrade notes that this PR does not cover, recorded so they are not lost. Each is a 3.33 change that alters behavior someone may depend on. They can be added here or carried as a follow-up.
The first seven are quoted from the 3.33 release notes in this tree, so they are as verified as the notes that did make it in.
Felix, not BIRD, programs IPIP cluster routes by default. BIRD programming is deprecated and intended for removal in v3.35. Route programming changes hands during the upgrade, by default, on every IPIP cluster. calico 13470.
New write-time validation on the projectcalico.org/v3 CRDs. Invalid BGP peer IPs, network set entries and policy rule protocols are now rejected on write, and the libcalico-go write defaults and numeric ranges are applied in the CRD schemas. GitOps pipelines can start failing on values kubectl previously accepted. This is the same shape as the policy limits note already on the page, and could fold into it. calico 13967.
The cni.projectcalico.org/ipAddrs annotation now honours the pool's allowedUses. Requesting an IP from a pool that does not permit workload use fails instead of allocating. Pods that scheduled before the upgrade stop scheduling after it. calico 13301.
Default BIRD logging widens from states to states, routes, filters and events. A log-volume step change for large BGP deployments, with disk-fill risk. The opt-out is LogSeverityScreen None in BGPConfiguration. calico 12578.
Dikastes normalizes HTTP request-targets before evaluating Application Layer Policy path rules, and prefix matches are anchored to segment boundaries. Existing ALP path rules can change meaning. The Kubernetes upgrade guide already has an ALP section, which is the natural home. calico 12531.
Annotations set directly on a generated host endpoint are removed on the next reconcile. They have to move to the KubeControllersConfiguration auto host endpoint template. calico 13205.
Helm deprecates defaultFelixConfiguration in the tigera-operator chart. Helm upgraders with that value set get no warning. calico 13953.
The last one is from Envoy Gateway's own v1.9 release notes rather than ours, so it needs checking against what Calico actually bundles before it is written up.
Envoy Gateway v1.9 beyond the Kubernetes floor and the Gateway API CRD bump this PR already covers: Lua EnvoyExtensionPolicy is disabled by default and needs a new enableLua field, tracing client sampling defaults to 0 per cent instead of 100, and EndpointSliceIndex is on by default and can raise controller memory in clusters with many EndpointSlices.
Worth recording why the page was fixed rather than retired: Envoy-based application layer policy is deprecated in Calico Enterprise and Calico Cloud from 3.23, but data/feature-status.yaml records no status for it in Open Source, and the page carries no deprecation notice, so it is a supported path in 3.33 and the manifests have to work.
I checked the other twelve removed images across the 3.33 tree for the same problem. The only other hit is getting-started/kubernetes/hardway/install-typha.mdx, which pins calico/typha:v3.8.0. That tag still resolves on quay, so it is stale rather than broken, and it is outside this change.
- 3.32 dropped enforcement of AdminNetworkPolicy and
BaselineAdminNetworkPolicy in favor of ClusterNetworkPolicy. Those
resources stay in the cluster and stop taking effect, with no error, so
a 3.31 source loses policy enforcement silently. Added to the 3.32
group, which previously held only the OpenStack migration.
- The Ingress Gateway note gave a version floor without saying what
happens below it.
- The OpenStack note told readers to run calico-resync without saying it
ships in its own package.
- The yum and apt repository instructions told readers to substitute
$[version], which renders v3.33 while the repository is calico-3.33.
Following it literally gives calico-v3.33 and a failed update. Mine,
from the N-2 commit.
flexvol is not a separate image; it is the component key/alias for calico/pod2daemon-flexvol (see version-3.32/variables.js:40), which is already in this list. Listing it twice makes the count fourteen and may send mirror operators looking for a nonexistent repository. Remove the duplicate and change the count to thirteen.
Correct overlap validation behavior for native v3 CRDs
This behavior is only true in aggregation API-server mode. With native projectcalico.org/v3 CRDs, overlap validation is asynchronous: the pool is created, receives a Disabled status condition, and IPAM stops allocating from it (operations/native-v3-crds.mdx:172-174). Because upgrades preserve CRD mode, the current wording and conditional precheck give native-mode users the wrong outcome and can leave an overlapping pool unusable after upgrade.
3.33 tests against Kubernetes 1.35 to 1.37, so two notes warned about
versions no supported cluster can be on.
The Ingress Gateway note led with Envoy Gateway's 1.33 floor, which every
supported cluster clears by two minors. What still bites is the Gateway
API CRD bump to v1.6, so the note is now about that, and keeps the two
fields it rejects.
The native v3 CRDs note gave a 1.34 floor and a feature gate needed on
1.34 and 1.35. Only 1.35 is in range, so only 1.35 needs the gate, and
from 1.36 the feature is generally available.
The group headings read "Upgrading to 3.33" and "Upgrading to 3.32", but
nobody reading the 3.33 docs is upgrading to 3.32. The second group holds
changes that arrived in 3.32 and still need action if you are coming from
3.31 and jumping over it, which is what its own subtitle said while the
heading above it said the opposite.
Headings now name the release the change came from, and the page
introduction tells you to read every group newer than the version you are
upgrading from.
The groupings were scaffolding for us, not something a reader needs. The
page is now a flat list of notes, each at the same level, and each one
already says who it affects.
Provenance moves into an MDX comment above each note, giving the release
it came from and the upstream PR where there is one, so we can still tell
where a note originated without putting it on the page.
The two notes that arrived in 3.32 now carry their own scoping, since the
group heading that used to say so is gone. They state that they apply
when upgrading from 3.31, and that a 3.32 cluster already has them.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Adds operations/upgrading/upgrade-notes.mdx and moves the upgrade notes off the release notes.
Fifteen notes, action-required first, grouped by the release that introduced them. Each says who it affects, named as something the reader can check, what happens if they do nothing, then numbered steps in upgrade order.
Removed as below that floor: the Upgrade OwnerReferences section, which starts at v3.28 and was duplicated in the Kubernetes and OpenShift guides, and three steps about cleaning up a temporary allow-all-upgrade policy from before v3.14. Those three were already broken, since they told the reader to undo pre-upgrade steps that are not on the page and the component behind them is imported nowhere. That component is deleted too.
The OpenStack policy data migration stays, because it applies to upgrades from earlier than v3.32 and 3.31 is a supported source.
The page on the preview