This is very, very suboptimal. I like the VXLAN support a lot, and of course this will have a lot of features while requiring minimal actual network knowledge (good luck getting everything to work correctly when combined, but you can certainly configure them ... and so you're making it the customer's problem)
Oh and this is going to cause out-of-order delivery and slow down user applications by a lot (because of channel bonding that looks like it's just left at the defaults), it is horribly inefficient. People don't put the host network either next to the VMs or on a separate network card for nothing.
In their case a packet walk would show:
1) guest userspace -> guest kernel
2) guest kernel virtio_net (hopefully) -> host kernel virtio_net
3) host kernel virtio_net -> host kernel
4) host kernel -> OVS vswitch data path
5) OVS vswitch data path -> host kernel bridge port
6) host kernel bridge port -> host kernel switch/networking stack
7) host kernel switch/networking stack -> host outgoing bonding virtual port
8) host outgoing bonding virtual port -> physical port
Each of these steps requires at the very least a memory allocation, inserting a step on a work queue, waiting on that work queue. Also very likely 5 of these steps require a context switch (at minimum waiting for the process scheduler to reschedule a task, on linux still usually requires 1ms minimum wait, more under load). So this inserts 5ms of latency minimum (and under load it's going to balloon) before the packet even arrives on the ring buffer of the first physical network card. And, as stated before, it's also going to cause out-of-order delivery.
And that's, of course, before application developers put a multiplication factor before this cost by using something like nginx or even multiple layers of nginx. I get the flexibility gain, and of course application developers get to do whatever they want, but ... why?
This is also eating a lot of processing power of the machines (and everything that comes with that, power use, even co2). And a further issue with that is that this is kernel networking path, which doesn't show in top, and doesn't show in most kernel metrics, you have to really know what you're looking for. And the cost that is incurred on the application side by due to the delay and the out of order packets doesn't show up anywhere except on the customer's bill, but good luck finding that it's wasted capacity.
If you do something like ML training from an NFS or S3 mount or NVMEoE or RoCE you will clearly notice the flaws in this design. The gains you can make there approach the gains you can make by switching from ethernet to fibre channel/infiniband.
What is possible with a great design: 0/1 context switch from guest userspace to network card ring buffer (zero context switches is possible by either using io/uring in guest userspace, or by binding the physical hardware to the guest VM and then into the application). Ping times to same-building VMs that consistently stay below 0.1ms, even with machine loads over 98%. Wish someone would pay me to do that.
Zero context switches while maintaining all features is possible. A lot of work, but possible. I don't believe anyone has yet done it, but it is possible.
And, please, move the linux host into the OVS ... just that little step will save about half the cost and it only requires being a bit more careful in operations (or having actual OOB, like serial or an extra hardware network card, the cheapest thing you're throwing away will easily do for that purpose)
And yes, I worked on the networking stack of one of the hyperscalers. They are at 2 to 3 context switches, more if you use any kind of tunneling (it's a crime that VXLAN is not supported ...). A lot better than this design, but not really close to perfectly optimal. It would be great to work on getting that closer to optimal in a large hoster. VPP + DPDK right into guest VMs. Sigh. Back to AI networking.
why are you mentioning roce and ml training? These are cpu-only machines with 10 Gbit uplinks shared between every virtual machine. I'm not super familiar with the helmet offering, but last time I checked their GPU servers were baremetal
Maybe calling the setup "suboptimal" is incorrect and it is optimal given their fleet and customers
What do you think of Oxide.computer networking?
If you want Nitro you know where to find it... and what it costs.
The people using hetzner cloud are probably hosting CRUD apps
This is like a nice advert for SQLite. Or for putting your app and database on the same server.
> Oh and this is going to cause out-of-order delivery and slow down user applications by a lot (because of channel bonding that looks like it's just left at the defaults),
Doesn't lag default to hashing by something? I expect it to use the rss hash, or src/dest ip and hopefully port.
On the number of layers, I totally agree. You can't fix the problems that come from having too many layers with more layers. And, you'll never get those delays back. If this really adds 5 ms (in each direction!), that's wild, but it would be clearly visible in pings... 10 ms round trip to get onto the network is like moving your servers 300 miles away from everyone.
> Wish someone would pay me to do that.
I feel you. Everyonce in a while, I get to do some really neat networking stuff, but I don't know how to make that my whole job.