Best way to go 1.5.4 → 1.6.21 and move tunnelers without endpoints dropping

Hi all,

Thanks for all the work on OpenZiti — it's been great to run. I've got an upgrade coming up and I'd really value a second opinion before I touch production, because I'm worried about endpoints dropping offline partway through the rollout.

Here's where I'm at. It's a Helm/GitOps setup, so the control-plane bump is really just an image-tag change. Current state and where I want to get to:

Component Now Target
Controller + routers 1.5.4 1.6.21
Linux ziti-edge-tunnel 1.7.4 1.18.7
Windows Desktop Edge 2.7.4.0 2.11.7.0
ziti-controller Helm chart 1.3.4 (appVersion 1.5.4) —

One detail that might matter: our controller's client API is served on its own hostname on :443 with a separate server certificate — i.e. a different cert chain than the controller's internal ziti identity (the identity.server_cert issued by the network's own OpenZiti CA). As I understand it - this is the default behaviour of the official ziti-controller Helm chart. We run chart version 1.3.4 where webBindingPki.enabled defaults to `true`, which the chart describes as generate a separate PKI root of trust.. — so the edge client API's server cert chains to a different root CA than the controller's own identity cert, out of the box. So I suspect anyone running the chart with defaults is in a similar situation.

What I'm trying to avoid is any window where devices lose connectivity, and there are two things in particular that worry me — and they seem to be triggered by different components, which is the whole reason I think there might be a safe ordering:

1. The API-session-refresh problem (gated by the router version)

Older tunnelers against 1.6–1.8 routers hit the issue where everything works for about half an hour (default session time) and then new dials start failing with 401 NOT_AUTHORIZED until you restart the tunneler.

2. The OIDC auth problem (gated by the controller version)

A newer tunneler talking to an older controller that serves its client API from a separate cert chain.

I actually ran into that second one when I tried to be clever and test the new 1.18.7 tunneler against the still-1.5.4 controller before upgrading anything. It authenticates over OIDC, gets a token, and then the controller turns around and rejects that same token on current-identity and current-api-session, over and over:

INFO  Ziti C SDK version 1.18.7 starting
INFO  ctrl\[https://client-api.example.com:443\] controller initialized
INFO  version_pre_auth_cb() connected to controller https://client-api.example.com:443 version v1.5.4
INFO  version_pre_auth_cb() using OIDC authentication method
INFO  oidc\[internal\] initializing with provider\[https://client-api.example.com:443/oidc\]
INFO  request_token() requesting token path\[.../oidc/oauth/token\] auth\[<redacted>\]
INFO  on_ziti_event() ... authorization success
ERROR ctrl_body_cb() API request\[/current-identity\] failed code\[UNAUTHORIZED\]
       message\[The request could not be completed. The session is not authorized or the credentials are invalid\]
WARN  update_identity_data() api session is no longer valid. Trying to re-auth
ERROR ctrl_body_cb() API request\[/current-api-session\] failed code\[UNAUTHORIZED\]
WARN  oidc_refresh_cb() OIDC token refresh failed: 400 Bad Request { "error": "invalid_grant" }
# ...then it loops: request_token → authorization success → UNAUTHORIZED → invalid_grant

What I think is going on (happy to be corrected)

My read — and I could well be wrong here, so please correct me — is that the newer SDK picks OIDC because the 1.5.4 controller advertises the capability, and that 1.5.4 then mishandles the token when the client API has a separate cert chain. At least to my eyes that lines up with [#3231](OIDC authentication fails if the client api has a separate cert chain · Issue #3231 · openziti/ziti · GitHub). As far as I can tell the thing that decides whether a client is affected is simply the auth method its SDK selects:

  • Old tunnelers (our 1.7.4) seem to use legacy auth — I don't think their older C SDK does OIDC at all — and they connect to the 1.5.4 controller without any trouble.
  • Newer tunnelers (1.18.7) appear to auto-select OIDC when the controller advertises it, and those are the ones I saw fall over against 1.5.4.

So my tentative interpretation is that this is an old-controller + new-client (OIDC) combination rather than anything wrong with the new tunneler itself — though I'd really like confirmation. I also couldn't find a way to force legacy auth from the client side, and disableOidcAutoBinding looks like it only arrived in 2.0, which is why I've landed on "upgrade the controller first" — but I may be missing something.

The plan I'm considering

The two problems seem to be gated by different components — the OIDC one by the controller, the session-refresh one by the routers — so my idea is to split the control-plane upgrade and do the routers last:

  1. Controller → 1.6.21 first (routers stay on 1.5.4).
    Should fix the OIDC issue, and my old 1.7.4 tunnelers should keep working throughout — they use legacy auth (which a 1.6.21 controller should still accept), and with the routers still on 1.5.4 the session-refresh bug can't bite either.

  2. Tunnelers → 1.18.7, at my own pace.
    Realistically, rolling the fleet out takes me about a week — the devices aren't all reachable at once and I have to stage it — so a 30-minute window after a router upgrade just isn't something I can hit. With the routers still on 1.5.4, the session-refresh bug can't trigger, so I'm not racing that clock across the fleet.

  3. Routers → 1.6.21 last
    Once every client is already on a fixed tunneler.

The idea is that no device ever sees the broken combination, so nothing "falls off the grid" during the rollout.

My questions (whenever someone has a moment)

  • Transition/bridge version (my main question):
    Is there a middle step that avoids a hard cutover — either a controller version that plays nicely with my current legacy-auth 1.7.4 tunnelers and the newer OIDC ones, or a tunneler version that works against both a 1.5.4 and a 1.6.21 controller?

  • Routers-last ordering: is doing the routers last like this a sensible and supported way to deal with the session-refresh window? It matters because, as above, I realistically need ~a week to get every tunneler updated, so I can't rely on hitting a 30-minute window. Anything to watch out for running a 1.6.21 controller with 1.5.4 routers for a week or so (I'll keep that gap as short as I can)?

  • Fleet rollout:
    Is there an officially recommended way to roll a larger tunneler fleet forward without the rush — a documented order, or bumping edge.api.sessionTimeout temporarily to widen the window?

  • Root cause:
    Does the OIDC loop above really match [#3231](OIDC authentication fails if the client api has a separate cert chain · Issue #3231 · openziti/ziti · GitHub), i.e. is it purely a controller-side thing that clears once I'm on 1.6.8+, with no supported way to run a modern tunneler against 1.5.4?

  • Target version:
    For a 1.6.21 control plane, is the latest tunneler (linux 1.18.7 / windows 2.11.7.0) the right target, or is there any reason to pin a bit lower?

Happy to share more of the logs or our controller web/identity config if that helps. Thanks a lot for reading, and for all the work on the project!