One client with wrong clock caused whole network down

Short summary before the long read:

A windows client with a bad clock, trying to recconnect too often( about 40 times per second ) which causing api sessions number to raise and cleanup locking bbolt database, which in the end causing error 429 for new clients, and in case of controller restart - dead network.

Hello, our company is using OpenZiti for connections and one day a newbie installed client on his windows and contacted our support with a problem that he wasn't able to access company's resources. Checking his logs revealed that he's having a bad clock:

[2026-10-01T18:36:58.804Z]    WARN ziti-sdk:legacy_auth.c:397 refresh_delay() legacy_auth[ziti-ctl.nekki.com]: api session is already expired according to local clock
[2026-10-01T18:36:58.804Z]    INFO ziti-sdk:legacy_auth.c:322 legacy_session_cb() legacy_auth[ziti-ctl.nekki.com]: refresh in 0 s
[2026-10-01T18:36:58.804Z]    WARN ziti-sdk:legacy_auth.c:333 legacy_timer_cb() legacy_auth[ziti-ctl.nekki.com]: session expired according to local clock
[2026-10-01T18:36:58.804Z]    WARN ziti-sdk:ziti.c:242 ztx_set_unauthenticated() ztx[1] auth error: session token has expired

So, I have adviced him to fix this, which was done.

Soon after this, he reported that he's having another issue, how with error 429:

[2026-10-01T09:16:59.616Z]    INFO ziti-sdk:legacy_auth.c:322 legacy_session_cb() legacy_auth[ziti-ctl.nekki.com]: backoff in 18 s
[2026-10-01T09:17:11.119Z]    INFO ziti-edge-tunnel:ziti-edge-tunnel.c:3509 endpoint_status_change() Received session unlocked event
[2026-10-01T09:17:11.119Z]    WARN ziti-sdk:posture.c:1220 ziti_endpoint_state_change() ztx[1] endpoint is disabled
[2026-10-01T09:17:18.346Z]    WARN ziti-sdk:legacy_auth.c:315 legacy_session_cb() legacy_auth[ziti-ctl.nekki.com]: failed[3] to acquire response: 429 Too Many Requests

We started looking into this, and soon started getting same reports from other people who were trying to connect.

Whoever was connected prior to this - were fine.

We were unable get a responce from an API on web\cli, were getting the same error code - 429.

We tried restarting the controller, and we were able to run some commands after, but only for a few moments, then it again restponded with 429 errors.

Starting from this restart, whole network was dead.

Controller log was filled with such messages:

Oct 01 10:12:32 controller ziti[586128]: {"circuitId":"4tLohFDRuYTSNwEPDKTeGf","file":"``github.com/openziti/ziti/controller/network/fault.go:49","func":"github.com/openziti/ziti/controller/network.(*Network).fault","level":"info","msg":"sent`` unroute for circuit to router in response to forwarding fault","routerId":"r1-JRG5d.","time":"2026-10-01T10:12:32.300Z"}

We have noticed that bbolt.db got about 3x bigger that it was previos day, and stated digging there, found that database was locked with a process which seem was trying to cleanup sessions:

goroutine 79 [runnable]:
github.com/openziti/storage/boltz.FieldToString(0x5?, {0x7f6ce03940ab, 0x19, 0x19?})
github.com/openziti/storage@v0.4.28/boltz/typed_bucket.go:268 +0x428
github.com/openziti/storage/boltz.(*rowCursorImpl).EvalString(0x1e8a0aef5dc0, {0x1e8a0e3407f0, 0xa})
github.com/openziti/storage@v0.4.28/boltz/query_cursor.go:132 +0x67
github.com/openziti/storage/ast.(*StringSymbolNode).EvalString(0x37?, {0x482d910?, 0x1e8a0aef5dc0?})
github.com/openziti/storage@v0.4.28/ast/node_symbol.go:250 +0x2b
github.com/openziti/storage/ast.(*BinaryStringExprNode).EvalBool(0x1e89f78358c0, {0x482d910, 0x1e8a0aef5dc0})
github.com/openziti/storage@v0.4.28/ast/node_expr.go:381 +0x3f
github.com/openziti/storage/ast.(*queryNode).EvalBool(0x1e89fa3f4200?, {0x482d910?, 0x1e8a0aef5dc0?})
github.com/openziti/storage@v0.4.28/ast/node_query.go:108 +0x25
github.com/openziti/storage@v0.4.28/boltz/query_scanners.go:158 +0x199
github.com/openziti/storage/boltz.newFilteredCursor(0x1e8a05e3dc00, {0x4861c80, 0x1e89f71ed7a0}, {0x480ad30, 0x1e89fa3f4200}, {0x4816460, 0x1e89f78358f0})
github.com/openziti/storage@v0.4.28/boltz/query_scanners.go:89 +0x1be
github.com/openziti/storage/boltz.(*BaseStore[...]).IterateIds(0x48d8680, 0x1e8a05e3dc00, {0x4816460, 0x1e89f78358f0})
github.com/openziti/storage@v0.4.28/boltz/store_query.go:386 +0x6b
github.com/openziti/storage/boltz.(*BaseStore[...]).IterateValidIds(0x48d8680, 0x1e8a05e3dc00, {0x4816460, 0x1e89f78358f0?})
github.com/openziti/storage@v0.4.28/boltz/store_query.go:391 +0x37
github.com/openziti/storage/boltz.(*fkDeleteCascadeConstraint).ProcessBeforeDelete(0x1e89f70f14b8, 0x1e89ff392e40)
github.com/openziti/storage@v0.4.28/boltz/indexes.go:1032 +0x4be
github.com/openziti/storage@v0.4.28/boltz/indexes.go:235 +0x74
github.com/openziti/storage/boltz.(*BaseStore[...]).processDeleteConstraints(0x48e0a80, {0x4824a20, 0x1e89fd663130}, {0x1e8a10c671a0, 0x19})
github.com/openziti/storage@v0.4.28/boltz/store_crud.go:378 +0x154
github.com/openziti/storage/boltz.(*BaseStore[...]).DeleteById(0x48e0a80, {0x4824a20, 0x1e89fd663130}, {0x1e8a10c671a0, 0x19})
github.com/openziti/storage@v0.4.28/boltz/store_crud.go:412 +0x38b
github.com/openziti/ziti/controller/model/base_manager.go:324 +0x99
github.com/openziti/storage@v0.4.28/boltz/db.go:154 +0x63
go.etcd.io/bbolt.(*DB).Update(0x1889425?, 0x1e89fa56dc70)
go.etcd.io/bbolt@v1.4.3/db.go:908 +0x6f
github.com/openziti/storage/boltz.(*DbImpl).Update(0x1e89f6f7e1e0, {0x4824a20?, 0x1e89fd663130?}, 0x1e8a081cdd10)
github.com/openziti/storage@v0.4.28/boltz/db.go:152 +0x170
github.com/openziti/ziti/controller/model.(*baseEntityManager[...]).deleteEntityBatch(0x48a43a0, {0x1e89f8a3a008, 0x1f4, 0x1f4}, 0x1e8a03272490)
github.com/openziti/ziti/controller/model/base_manager.go:322 +0xeb
github.com/openziti/ziti/controller/internal/policy/api_session_enforcer.go:102 +0x55a
github.com/openziti/ziti/common/runner/runner.go:121 +0xcf
created by github.com/openziti/ziti/common/runner.(*LimitedRunner).Start in goroutine 1
github.com/openziti/ziti/common/runner/runner.go:114 +0xda

So, we have deleted sessions with "ziti controller delete-sessions-from-db". Which resolved the problem.

Later we found in the db, that one client had 50+ thousand sessions, this was that user with the wrong clock.

Cheking his log, we found that he had 18k messages 'api session is already expired according to local clock' in 8 minutes, which seems to be the root cause of the problem.

Packages versions from our controller if needed:

openziti-1.6.19-1.x86_64
openziti-controller-1.6.19-1.noarch
openziti-console-4.5.3-1.x86_64

Hi @AlexZ, welcome to the community and to OpenZiti!

Yikes... That doesn't sound good... Would you be willing to send us a 'feedback' zip file from that Windows client so we can have a look at those logs? I don't quite see where the 'hot loop' was from the snippets you sent. Also I don't see the ZDEW version/ziti-edge-tunnel version? (I might be missing it there's a lot of good detail you sent)

client_logs.zip (1013.2 KB)

Here are logs from client, user is saying he was using Ziti.Desktop.Edge.Client-2.11.7.0.msi file for client install.

Hope this helps, please let me know if there's anything else I can provide.

It shows where the hot loop is. Yes very helpful thanks. I/we will check to see if that hot loop is already fixed or not. Appreciate those logs!

I've filed legacy_auth: hot loop when the controller returns an already-expired API session · Issue #1163 · openziti/ziti-sdk-c · GitHub

thanks again

OpenZiti has moved away from durably stored API Sessions for similar scaling reasons. Even in normally operating networks, a scaling bottleneck naturally occurs once there are enough users. Additionally, herding scenarios could occur where partial underlay network outages could cause mass reconnect/re-auth events.

I'll talk internally about whether we want to defend the controller against this. This would fall under 1.6.x LTS releases. In 2.0.x LTS, OIDC auth has replaced "legacy auth".