What rewrites and loads endpoints.yml when ctrl.endpointsFile is commented out?

Hi,

We are running OpenZiti Router v2.0.4. Our controller is exposed through an nginx TCP/SNI proxy on public port 443; its internal listener uses port 1280.

The router configuration contains:

ctrl:
  endpoint: tls:controller.example.com:443
  #endpointsFile: /var/lib/ziti-router/endpoints.yml

The public 443 endpoint works. Connecting directly to public port 1280 does not, as that port is only used internally behind nginx.

Despite endpointsFile being commented out, the router logs:

loading controller endpoints from [endpoints.yml]

The router successfully bootstraps through 443, but then receives this update from the controller:

update ctrl endpoints message received
endpoints=["tls:controller.example.com:1280"]

Afterward, endpoints.yml contains:

controllers:
  - id: controller
    endpoints:
      - address: tls:controller.example.com:1280

endpoints:
  - tls:controller.example.com:1280

Controller configuration:

v: 3

cluster:
  dataDir: "/var/lib/ziti-controller/raft"

ctrl:
  listener: tls:0.0.0.0:1280
  options:
    advertiseAddress: tls:controller.example.com:443

edge:
  api:
    sessionTimeout: 30m
    address: controller.example.com:443

web:
  - name: client
    bindPoints:
      - interface: 0.0.0.0:1280
        address: controller.example.com:443
    apis:
      - binding: edge-client
        options: {}
      - binding: edge-oidc
        options: {}
      - binding: health-checks
        options: {}

Questions:

  1. Why does the router default to endpoints.yml when ctrl.endpointsFile is not configured?
  2. Why does the generated endpoint use the controller’s internal listener port 1280 instead of the externally reachable 443 address?
  3. What takes precedence between ctrl.endpoint and the generated endpoints.yml?
  4. What is the supported configuration for preventing the router from learning or loading the unreachable internal endpoint?

Hi @montwepa, thank you for reporting this.

Let's look at how things work, what's broken and then short term and longer term fixes.

This is tied into HA. Pre-HA, there was no advertise option because routers had to already know how to reach the one controller.

In the HA world, however, controllers need to be able to reach each other and tell new members how to reach each node in the cluster. Tied into this is that when you add or remove a node from a cluster, we want to update the routers automatically.

Finally, routers and controller peers all connect on the same endpoint, b/c most organizations only want a single port open, usually 443.

So the endpoints.yml isn't optional for HA because routers need to be able to react to cluster changes. The endpointsFile just lets you customize the location away from the default location. The file is written when the controller updates the router with the latest endpoints.

The problem you're seeing is that the advertise address configuration is only referenced when adding a node to the cluster (or for the initial node, at bootstrap time). This is because the addresses are stored inside raft as part of raft configuration. If you want to change it, you need to either remove the node from the cluster and then add it back or just re-add it (it will see the ID is the same and just update the address to what the added controller tells it is the advertise addr).

When you've got a single node running in HA mode, you can't remove the node from the cluster, because it's a single node. You have to back up the node, including a db snapshot, clear the raft and db and re-bootstrap from the snapshotted db, with an updated config file.

That's quite cumbersome and not intuitive. I had started work on a CLI agent command to update the advertise address, see Can't change the advertise address of a single node HA cluster controller · Issue #3588 · openziti/ziti · GitHub. I ran into some issues because updating things through raft doesn't reflect in the config file. The config file is still needed for non-HA setups as the source of truth when creating enrollment JWTs.

Given that this is clearly causing some pain, I'm going to take another stab at fixing this, likely by treating the config file as the source of truth and updating raft to match on startup, if it sees changes. We'll see if I run into any blockers.

If you want to verify that a mismatch between raft and the config is your problem, you can run:

ziti agent cluster list

on the controller, and you can see what raft thinks the advertise address is.

Let me know if that's helpful,
Paul

Hi @plorenz, thanks for the reply. Glad to hear that the issue was known and not something I did wrong on my end.

I have used the old Ziti HA when it was in beta, think back then you could point to multiple endpoints in the controller yaml it self? Still getting used to the 2.0 version of Ziti :grin:

To confirm my understanding, the likely reason 1280 is being published is that it was stored as this controller member’s advertised address in Raft during the original bootstrap, and later changing ctrl.options.advertiseAddress to 443 did not update that stored value?

I’ll verify this with ziti agent cluster list. If Raft already reports 443, then something else appears to be deriving or publishing 1280.

But anyways, glad to hear you are working on it. Hope you can solve it.