This appears to be a BGP EVPN (VXLAN)-like architecture, with the VTEPs pushed down to the hypervisors instead of living on the ToR switches. (Also OpenStack's network architecture.)
Oh, my Hetzner VM just had a 13 hour outage from yesterday evening to today morning. They say there was some incident with their DNS servers or something.
I don't recall that ever happening during my 12 years at Digital Ocean. But yes, the Hetzner VM is both cheaper and beefier than the DO one was.
This is very, very suboptimal. I like the VXLAN support a lot, and of course this will have a lot of features while requiring minimal actual network knowledge (good luck getting everything to work correctly when combined, but you can certainly configure them ... and so you're making it the customer's problem)
Oh and this is going to cause out-of-order delivery and slow down user applications by a lot (because of channel bonding that looks like it's just left at the defaults), it is horribly inefficient. People don't put the host network either next to the VMs or on a separate network card for nothing.
8) host outgoing bonding virtual port -> physical port
Each of these steps requires at the very least a memory allocation, inserting a step on a work queue, waiting on that work queue. Also very likely 5 of these steps require a context switch (at minimum waiting for the process scheduler to reschedule a task, on linux still usually requires 1ms minimum wait, more under load). So this inserts 5ms of latency minimum (and under load it's going to balloon) before the packet even arrives on the ring buffer of the first physical network card. And, as stated before, it's also going to cause out-of-order delivery.
And that's, of course, before application developers put a multiplication factor before this cost by using something like nginx or even multiple layers of nginx. I get the flexibility gain, and of course application developers get to do whatever they want, but ... why?
This is also eating a lot of processing power of the machines (and everything that comes with that, power use, even co2). And a further issue with that is that this is kernel networking path, which doesn't show in top, and doesn't show in most kernel metrics, you have to really know what you're looking for. And the cost that is incurred on the application side by due to the delay and the out of order packets doesn't show up anywhere except on the customer's bill, but good luck finding that it's wasted capacity.
If you do something like ML training from an NFS or S3 mount or NVMEoE or RoCE you will clearly notice the flaws in this design. The gains you can make there approach the gains you can make by switching from ethernet to fibre channel/infiniband.
What is possible with a great design: 0/1 context switch from guest userspace to network card ring buffer (zero context switches is possible by either using io/uring in guest userspace, or by binding the physical hardware to the guest VM and then into the application). Ping times to same-building VMs that consistently stay below 0.1ms, even with machine loads over 98%. Wish someone would pay me to do that.
Zero context switches while maintaining all features is possible. A lot of work, but possible. I don't believe anyone has yet done it, but it is possible.
And, please, move the linux host into the OVS ... just that little step will save about half the cost and it only requires being a bit more careful in operations (or having actual OOB, like serial or an extra hardware network card, the cheapest thing you're throwing away will easily do for that purpose)
And yes, I worked on the networking stack of one of the hyperscalers. They are at 2 to 3 context switches, more if you use any kind of tunneling (it's a crime that VXLAN is not supported ...). A lot better than this design, but not really close to perfectly optimal. It would be great to work on getting that closer to optimal in a large hoster. VPP + DPDK right into guest VMs. Sigh. Back to AI networking.
why are you mentioning roce and ml training? These are cpu-only machines with 10 Gbit uplinks shared between every virtual machine. I'm not super familiar with the helmet offering, but last time I checked their GPU servers were baremetal
Maybe calling the setup "suboptimal" is incorrect and it is optimal given their fleet and customers
I don't recall that ever happening during my 12 years at Digital Ocean. But yes, the Hetzner VM is both cheaper and beefier than the DO one was.
We triple the price from one day to another and don’t answer support tickets on weekends.
Oh and this is going to cause out-of-order delivery and slow down user applications by a lot (because of channel bonding that looks like it's just left at the defaults), it is horribly inefficient. People don't put the host network either next to the VMs or on a separate network card for nothing.
In their case a packet walk would show:
1) guest userspace -> guest kernel
2) guest kernel virtio_net (hopefully) -> host kernel virtio_net
3) host kernel virtio_net -> host kernel
4) host kernel -> OVS vswitch data path
5) OVS vswitch data path -> host kernel bridge port
6) host kernel bridge port -> host kernel switch/networking stack
7) host kernel switch/networking stack -> host outgoing bonding virtual port
8) host outgoing bonding virtual port -> physical port
Each of these steps requires at the very least a memory allocation, inserting a step on a work queue, waiting on that work queue. Also very likely 5 of these steps require a context switch (at minimum waiting for the process scheduler to reschedule a task, on linux still usually requires 1ms minimum wait, more under load). So this inserts 5ms of latency minimum (and under load it's going to balloon) before the packet even arrives on the ring buffer of the first physical network card. And, as stated before, it's also going to cause out-of-order delivery.
And that's, of course, before application developers put a multiplication factor before this cost by using something like nginx or even multiple layers of nginx. I get the flexibility gain, and of course application developers get to do whatever they want, but ... why?
This is also eating a lot of processing power of the machines (and everything that comes with that, power use, even co2). And a further issue with that is that this is kernel networking path, which doesn't show in top, and doesn't show in most kernel metrics, you have to really know what you're looking for. And the cost that is incurred on the application side by due to the delay and the out of order packets doesn't show up anywhere except on the customer's bill, but good luck finding that it's wasted capacity.
If you do something like ML training from an NFS or S3 mount or NVMEoE or RoCE you will clearly notice the flaws in this design. The gains you can make there approach the gains you can make by switching from ethernet to fibre channel/infiniband.
What is possible with a great design: 0/1 context switch from guest userspace to network card ring buffer (zero context switches is possible by either using io/uring in guest userspace, or by binding the physical hardware to the guest VM and then into the application). Ping times to same-building VMs that consistently stay below 0.1ms, even with machine loads over 98%. Wish someone would pay me to do that.
Zero context switches while maintaining all features is possible. A lot of work, but possible. I don't believe anyone has yet done it, but it is possible.
And, please, move the linux host into the OVS ... just that little step will save about half the cost and it only requires being a bit more careful in operations (or having actual OOB, like serial or an extra hardware network card, the cheapest thing you're throwing away will easily do for that purpose)
And yes, I worked on the networking stack of one of the hyperscalers. They are at 2 to 3 context switches, more if you use any kind of tunneling (it's a crime that VXLAN is not supported ...). A lot better than this design, but not really close to perfectly optimal. It would be great to work on getting that closer to optimal in a large hoster. VPP + DPDK right into guest VMs. Sigh. Back to AI networking.
Maybe calling the setup "suboptimal" is incorrect and it is optimal given their fleet and customers