# Proxmox Backup over Ziti fails on large datasets (HTTP/2.0 connection failed)

**URL:** <https://openziti.discourse.group/t/proxmox-backup-over-ziti-fails-on-large-datasets-http-2-0-connection-failed/5975>\
**Category:** Ziti Overlay\
**Created:** [August 1, 2026, 1:42pm UTC](https://openziti.discourse.group/t/proxmox-backup-over-ziti-fails-on-large-datasets-http-2-0-connection-failed/5975 "2026-08-01T13:42:31Z")\
**Posts on this page:** 11\
**Page:** 1

<div class="post-metadata">

**Author:** ![golden-monkey](https://avatars.discourse-cdn.com/v4/letter/g/e5b9ba/32.png) [@golden-monkey](https://openziti.discourse.group/u/golden-monkey)\
**Post date:** [August 1, 2026, 1:42pm UTC](https://openziti.discourse.group/t/proxmox-backup-over-ziti-fails-on-large-datasets-http-2-0-connection-failed/5975/1 "2026-08-01T13:42:31Z")

</div>

Hi community,

we are experiencing an issue regarding our remote backups which uses a ziti overlay network to connect the source to our remote backup target. Currently, every scheduled backup we run fails.

We are using the `promox-backup-client` cli tool to backup zfs-snapshots to a virtualized Promox Backup Server in a separate location. The locations are connected via public internet (Germany) using private household internet connections.

This setup used to work, when the amount of data to be backuped was small (\<20 GiB). After the dataset size increased to approx. 560 GiB we are continuously experiencing the following issue (`promox-backup-client` logs)

```bash
HTTP/2.0 connection failed
unfinished encoder state dropped
finished encoder state with errors
unclosed encoder dropped
closed encoder dropped with state
unfinished encoder state dropped
catalog upload error - channel closed    
Error: stream closed because of a broken pipe

```

In the relevant logs on the source machine (pve\_hydrogen.log) there is a ziti-edge-tunnel installed. We observe lines like these:

```bash
Jul 12 03:35:33 hydrogen ziti-edge-tunnel[2020]: (2020)[244105.375] ERROR ziti-sdk:channel.c:688 latency_timeout() ch[0] no read/write traffic on channel since before latency probe was sent, closing channel
Jul 12 03:35:33 hydrogen ziti-edge-tunnel[2020]: (2020)[244105.375] WARN ziti-sdk:channel.c:852 on_channel_close() ch[0] disconnected from edge router[private0] -110(connection timed out)
Jul 12 03:35:33 hydrogen ziti-edge-tunnel[2020]: (2020)[244105.375] ERROR ziti-sdk:connect.c:1049 connect_reply_cb() conn[1.118394/A2lfGNaY/Connecting](pbs-hydrogen.svc) failed to connect [-20/operation did not complete in time]
Jul 12 03:35:44 hydrogen ziti-edge-tunnel[2020]: (2020)[244115.568] ERROR tunnel-sdk:tunnel_tcp.c:191 on_tcp_client_err() client=tcp:100.64.0.1:48310 err=-14, terminating connection

```

On the target machine (pbs\_helium.log) there is a ziti-edge-tunnel installed. We observe:

```bash
Jul 12 05:35:36 pbs-helium ziti-edge-tunnel[195]: (195)[244211.073] WARN ziti-sdk:channel.c:403 on_channel_send() ch[0] write delay = 5.293 q=29 qs=526400
Jul 12 05:35:31 pbs-helium ziti-edge-tunnel[195]: (195)[244205.671] ERROR ziti-sdk:connect.c:1477 process_edge_message() conn[1.49358/Pf23Und1/Accepting](pbs-helium.svc) failed to connect, reason=close called 

```

On the “first” router (private0\_hydrogen.log) to which the source machine connects:

```bash
Jul 12 05:35:39 private0 ziti[123]: {"_context":"ch{edge}->u{classic}->i{ziti-sdk-c[0]@hydrogen/Zylw}","error":"write tcp 10.x.y.z:3022->192.a.b.c:42566: write: connection reset by peer",..."msg":"write error",...}
Jul 12 05:35:20 private0 ziti[123]: {... "error":"error creating route for [c/43sDS3qFenLAxpoIa90kTa]: timeout waiting for message reply: context deadline exceeded",..."msg":"failure while handling route update",...}
Jul 12 05:34:36 private0 ziti[123]: {... "level":"warning","msg":"closing while buffer contains unacked payloads","payloadCount":2,...}

```

On the “last” router (private1\_helium.log) which connects with the target machine:

```bash
Jul 12 03:34:31 private1 ziti[129]: {"circuitCount":5,"ctrlId":"atlas-ctrl-client","file":"github.com/openziti/ziti/router/forwarder/faulter.go:107","func":"github.com/openziti/ziti/router/forwarder.(*Faulter).run","level":"warning","msg":"reported forwarding faults","time":"2026-07-12T03:34:31.323Z"}
Jul 12 03:34:31 private1 ziti[129]: {"circuitId":"4KRh8X9sjo0t1ryh94FK19","file":"github.com/openziti/ziti/router/forwarder/forwarder.go:155","func":"github.com/openziti/ziti/router/forwarder.(*Forwarder).Unroute","level":"info","msg":"circuit unrouted","time":"2026-07-12T03:34:31.346Z"}
Jul 12 03:34:35 private1 ziti[129]: {"_context":"{c/HnyMaDyXq4UU8AciIX0C8|@/6q65xsfekE08jWiJ2ll67J}\u003cInitiator\u003e","circuitId":"HnyMaDyXq4UU8AciIX0C8","error":"cannot forward payload, no destination for circuit=HnyMaDyXq4UU8AciIX0C8 src=6q65xsfekE08jWiJ2ll67J dst=5WvGiniQ0JG3TsRVoSf7D9","file":"github.com/openziti/ziti/router/handler_xgress/data_plane.go:58","func":"github.com/openziti/ziti/router/handler_xgress.(*dataPlaneAdapter).ForwardPayload","level":"error","msg":"unable to forward payload","origin":0,"seq":9,"time":"2026-07-12T03:34:35.284Z"}

```

On “in-between” routers (public0\_hydrogen.log & public1\_helium.log) on the overlay network we see:

```bash
Jul 12 03:34:36 public0 ziti[685]: {"circuitId":"2VSvrd7ZMFcATcHhzcICdD","ctrlId":"atlas-ctrl-client","file":"github.com/openziti/ziti/router/forwarder/scanner.go:85","func":"github.com/openziti/ziti/router/forwarder.(*Scanner).scan","idleThreshold":60000000000,"idleTime":69532000000,"level":"warning","msg":"circuit exceeds idle threshold","time":"2026-07-12T03:34:36.773Z"}

```

We are sharing the logs of all “stations” in the ziti overlay network. We anonymized them to the best of our knowledge, if you see any sensitive data we forgot to redact, please send us a private message so we can take the info down.

As this issue is occurring every night, we were able to visualize some data. The “HTTP/2.0 connection failed” error usually occurs right before 20 GiB transferred are reached just before the one hour mark.

 ![image](https://global.discourse-cdn.com/free1/uploads/netfoundry/original/2X/9/9f016ce3601a3189e51343dfbe455f16a7d6730c.png)

 ![image](https://global.discourse-cdn.com/free1/uploads/netfoundry/original/2X/4/47f9f164cfaa96dbb7aa1159aeff3548795f33ba.png)

One of things we have tried (not part of the visualization), is to rate limit the transfer/upload speed of the `promox-backup-client` to `--rate 2875KB` and `--burst 128KB`. With these settings we still observe the "HTTP/2.0 connection failed” error at about 175 GiB transferred at the 16h mark. Note that the `--rate` parameter is set to approx. 40% uplink speed of the source servers internet connection. (The uplink is the bottleneck in our scenario.)

**Question:**

We cannot get the backups to complete without running into the error and we suspect that it has to do with the ziti overlay (or at least the interplay between `promox-backup-client` and ziti).

As we are at a loss as how to further troubleshoot this issue, we are happy for any ideas and suggestions. Let us know if you require any further information like network architecture, etc. Maybe someone had a similar issue or just knows a solution to our problem.

**Than you very much!**

[public1\_helium.log](https://openziti.discourse.group/uploads/short-url/u5gEFkuSkcRNf56FaUuNb102iOJ.log) (2.1 KB)

[pbs\_helium.log](https://openziti.discourse.group/uploads/short-url/z33ORu6FBf19mL6OGHz8ogAgice.log) (16.8 KB)

[pve\_hydrogen.log](https://openziti.discourse.group/uploads/short-url/6KHRoIdOHqflFvMiP2HN7NqPOG.log) (28.2 KB)

[public0\_hydrogen.log](https://openziti.discourse.group/uploads/short-url/ncZl2zAc9ibvlQYdHnjrTb08nCx.log) (2.6 KB)

[private1\_helium.log](https://openziti.discourse.group/uploads/short-url/qMNRPEuxdk6oG7m2RIhybJqINkV.log) (122.7 KB)

[private0\_hydrogen.log](https://openziti.discourse.group/uploads/short-url/33R48EfvdyvsCwr1Y4VUrSwpy9Y.log) (66.3 KB)

---

<div class="post-metadata">

**Author:** ![golden-monkey](https://avatars.discourse-cdn.com/v4/letter/g/e5b9ba/32.png) [@golden-monkey](https://openziti.discourse.group/u/golden-monkey)\
**Post date:** [August 3, 2026, 8:57am UTC](https://openziti.discourse.group/t/proxmox-backup-over-ziti-fails-on-large-datasets-http-2-0-connection-failed/5975/2 "2026-08-03T08:57:51Z")

</div>

Ideas anyone? Or a similar observations? 🥲

---

<div class="post-metadata">

**Author:** ![ekoby](https://yyz2.discourse-cdn.com/free1/user_avatar/openziti.discourse.group/ekoby/32/14_2.png) [@ekoby](https://openziti.discourse.group/u/ekoby)\
**Post date:** [August 3, 2026, 3:00pm UTC](https://openziti.discourse.group/t/proxmox-backup-over-ziti-fails-on-large-datasets-http-2-0-connection-failed/5975/3 "2026-08-03T15:00:44Z")

</div>

what versions of ziti network and tunnelers are you using?

these warnings indicate that hosting SDK has trouble sending data to the ER (similar errors are in the log from the other side)

```auto
WARN ziti-sdk:channel.c:403 on_channel_send() ch[0] write delay = 5.293 q=29 qs=526400

```

would it be possible to capture network packets (tcpdump/wireshark) between SDKs and ERs?

---

<div class="post-metadata">

**Author:** ![golden-monkey](https://avatars.discourse-cdn.com/v4/letter/g/e5b9ba/32.png) [@golden-monkey](https://openziti.discourse.group/u/golden-monkey)\
**Post date:** [August 3, 2026, 3:30pm UTC](https://openziti.discourse.group/t/proxmox-backup-over-ziti-fails-on-large-datasets-http-2-0-connection-failed/5975/4 "2026-08-03T15:30:25Z")

</div>

@ekoby  
I am looking into the packet capture with wireshark. i will post once i have that done. meanwhile here are the versions

#### Location 1 - “Hydrogen”:

Proxmox VE:

- `ziti-edge-tunnel version` v1.18.1

Private Router:

- `ziti --version` v1.6.12

Public Router:

- `ziti --version` v1.6.12

Controller:

- `ziti --version` v1.6.14

* * *

#### Location 2 “Helium”:

Private Router:

- `ziti --version` v1.6.14

Public Router:

- `ziti --version` v1.6.14

Edit: cleaned confusing services that have no influence on the product

---

<div class="post-metadata">

**Author:** ![ekoby](https://yyz2.discourse-cdn.com/free1/user_avatar/openziti.discourse.group/ekoby/32/14_2.png) [@ekoby](https://openziti.discourse.group/u/ekoby)\
**Post date:** [August 3, 2026, 4:16pm UTC](https://openziti.discourse.group/t/proxmox-backup-over-ziti-fails-on-large-datasets-http-2-0-connection-failed/5975/5 "2026-08-03T16:16:42Z")

</div>

is it the same tunneler version on both (Hydrogen and Helium) sites?

---

<div class="post-metadata">

**Author:** ![golden-monkey](https://avatars.discourse-cdn.com/v4/letter/g/e5b9ba/32.png) [@golden-monkey](https://openziti.discourse.group/u/golden-monkey)\
**Post date:** [August 3, 2026, 4:25pm UTC](https://openziti.discourse.group/t/proxmox-backup-over-ziti-fails-on-large-datasets-http-2-0-connection-failed/5975/6 "2026-08-03T16:25:20Z")

</div>

No, the versions are different

Proxmox VE “Hydrogen”:

- `ziti-edge-tunnel version` v1.18.1

Proxmox VE "Helium":

- `ziti-edge-tunnel version` v1.12.0

I can update all of them to the latest version if required

---

<div class="post-metadata">

**Author:** ![ekoby](https://yyz2.discourse-cdn.com/free1/user_avatar/openziti.discourse.group/ekoby/32/14_2.png) [@ekoby](https://openziti.discourse.group/u/ekoby)\
**Post date:** [August 4, 2026, 12:49pm UTC](https://openziti.discourse.group/t/proxmox-backup-over-ziti-fails-on-large-datasets-http-2-0-connection-failed/5975/7 "2026-08-04T12:49:44Z")

</div>

I would definitely recommend updating ZET to latest.

---

<div class="post-metadata">

**Author:** ![golden-monkey](https://avatars.discourse-cdn.com/v4/letter/g/e5b9ba/32.png) [@golden-monkey](https://openziti.discourse.group/u/golden-monkey)\
**Post date:** [August 5, 2026, 8:25am UTC](https://openziti.discourse.group/t/proxmox-backup-over-ziti-fails-on-large-datasets-http-2-0-connection-failed/5975/8 "2026-08-05T08:25:20Z")

</div>

I have not updated (yet)... was working on the packet capture over night. 1.5 GB packets later i have the following:

[backup\_failed\_relevant\_time\_frame\_20260804.pcapng.gz](https://openziti.discourse.group/uploads/short-url/671k0NdHmmqb18ecGqa4P2PlixT.gz) (4.4 KB)

---

<div class="post-metadata">

**Author:** ![golden-monkey](https://avatars.discourse-cdn.com/v4/letter/g/e5b9ba/32.png) [@golden-monkey](https://openziti.discourse.group/u/golden-monkey)\
**Post date:** [August 15, 2026, 6:15pm UTC](https://openziti.discourse.group/t/proxmox-backup-over-ziti-fails-on-large-datasets-http-2-0-connection-failed/5975/9 "2026-08-15T18:15:33Z")

</div>

@ekoby

Hi again,

sorry for the long radio silence. We were able to conduct further investigation. We used `iperf3` to test the connectivity.

We ran the test on the ziti-overlay and using a “non-ziti”/direct connection.

Our hypothesis is that the issue is caused by virtualizing the ziti-router on the same machine the ziti-edge-tunnel connects to. As the bandwidth of the first “hop“ is much greater than the bandwidth of the internet connection, the causes the buffer fill-up leading to tcp zero-windows and ultimately the connection dropping.

### Context:

In our source location we run a Proxmox server. On this server, we are running two virtualized ziti-routers. One we call `private0` , the other one we call `public0`. `private0` connects to all of the services we host. This includes the Proxmox host itself, which is the source of of the backup. `private0` can not connect to the internet and has to route all non-local traffic through `public1`. The idea behind this architecture was to be able to rate limit traffic so internet for the local users of the network remains stable, even if someone uploads/or downloads a large file from the server.

On the target side, we run the same setup.

### Observations:

1. Our MTU seems to be larger than the network connection can support. This worsens the observed issue, but does not seem to be the main culprit.

2. The “direct, non-ziti” internet connection is stable and runs at its advertised bandwidth.

3. When running the same tests on the ziti-overlay with the default circuit (source ZET → `private0` → `public0` → `public1` → `private1` → target ZET) we observe a volatile bandwidth with around 25% of datagrams lost (in udp mode). We observe the rmem allocated on `private0` to jump up to the buffer size of the socket.

4. When shortening the circuit to (source ZET → public0 → target ZET) the issue persists.

5. When changing the circuit to (source ZET → public1 → tartget ZET) the issue significantly improves. We can almost reach advertised bandwidth and the datagrams lost drop to below 5%.

---

<div class="post-metadata">

**Author:** ![golden-monkey](https://avatars.discourse-cdn.com/v4/letter/g/e5b9ba/32.png) [@golden-monkey](https://openziti.discourse.group/u/golden-monkey)\
**Post date:** [August 25, 2026, 12:52pm UTC](https://openziti.discourse.group/t/proxmox-backup-over-ziti-fails-on-large-datasets-http-2-0-connection-failed/5975/10 "2026-08-25T12:52:59Z")

</div>

@ekoby We are still experiencing the issue but we may have found the cause. However, it is unclear to us why that happens. Do you have any insight on why this might occur? It seems weird that the edge tunnel would need to have a direct route to a router facing the internet for TCP flow and congestion control to work.

---

<div class="post-metadata">

**Author:** ![gormami](https://yyz2.discourse-cdn.com/free1/user_avatar/openziti.discourse.group/gormami/32/1010_2.png) [@gormami](https://openziti.discourse.group/u/gormami)\
**Post date:** [August 26, 2026, 11:41am UTC](https://openziti.discourse.group/t/proxmox-backup-over-ziti-fails-on-large-datasets-http-2-0-connection-failed/5975/11 "2026-08-26T11:41:04Z")

</div>

Are the packet losses at any particular point or spread randomly throughout the transfer?

OpenZiti has an internal flow control, much like a TCP window, and of course the mTLS links usually run over TCP. It may be that things like latency differences are allowing some piling up somewhere, I would guess the side going into the tunnel, unless there are connections being made and broken during the transfers that you aren't seeing. Once UDP enters the overlay, it is transported over TCP.

First, I would check the circuit events, and see if there is one connection, or multiple ending and restarting quietly during the test.

Then, I'd make sure the ER's TCP stack settings are modified to make things easier and more performant in general. We use the following settings in sysctl.

`net.core.rmem_max = 16777216`  
`net.core.wmem_max = 16777216`  
`net.core.rmem_default = 16777216`  
`net.core.wmem_default = 16777216`  
`net.ipv4.tcp_rmem = 4096 87380 16777216`  
`net.ipv4.tcp_wmem = 4096 65536 16777216`  
`net.ipv4.tcp_mem = 8388608 8388608 16777216`  
`net.ipv4.udp_mem = 8388608 8388608 16777216`  
`net.ipv4.tcp_retries2 = 8`

In the OpenZiti setup, each router can take options on the link dialer and listener configs as well, to modify xgress, which handles the flow control. Depending on where your loss is, I would suggest upping the start size. If you are sending a strong flow of UDP into the tunneler, and there is back pressure from the system as the various windows "warm up", it could cause some significant loss, particularly early. The settings and basic descriptions are listed here.

> **[Conventions | NetFoundry Documentation](https://netfoundry.io/docs/openziti/reference/configuration/conventions/)**
>
> The following conventions apply to multiple areas of the configuration files for routers and

Lastly, we've seen issues with iperf3 before, particularly in UDP testing. Are you using the latest one? Especially, the "latest" using a lot of package managers vs. the latest downloading from the iperf site. The site had a newer version for a long time that did much better in UDP testing.
