SDN & Routing
The networking side of the Rocknet Lab: a private VXLAN overlay built with Proxmox SDN and provisioned by Terraform, and a four-router BGP, BFD and IS-IS lab running across it, with a live dashboard of every router and link.
The VXLAN Overlay
On top of the physical cluster network sits a private virtual network, built with Proxmox's SDN (software-defined networking) stack and provisioned entirely through Terraform rather than clicked together by hand in the UI.
It's a VXLAN overlay: each node wraps overlay traffic in UDP and tunnels it to the other two over the physical LAN, so all three nodes share one private layer-2 segment — no physical switch port, no VLAN tagging, no shared broadcast domain with the production LAN. Every container in the lab has a second network interface on it, which gives them a private east-west path to each other no matter which node they run on. Their primary LAN interface stays exactly as it was, so nothing about how they reach the internet or how the site is served changed. The whole thing can be created, changed, or torn down with a single Terraform command. As with the rest of this page, no addresses are shown here by design.
Routing Lab — BGP, BFD & IS-IS
Four routers running FRR, each in its own container and spread across all three nodes, make up a small service-provider-style network I use to practise the protocols carrier backbones run on. IS-IS is the interior gateway protocol, BFD sits on every link for sub-second failure detection, iBGP carries customer routes through a route reflector, and an eBGP session connects to a second autonomous system with prefix-list policy on both sides.
Every router-to-router link is its own point-to-point VXLAN segment built with the same Proxmox SDN and Terraform tooling as the overlay above, so each IS-IS adjacency and BFD session really does cross the physical LAN between two different machines. Addresses in this section show only their first three octets — the last one is masked as .xxx.
Live Routing Lab Dashboard
The drawings above show how the lab is designed; this panel shows what it's doing right now. A small dashboard container on the cluster polls all four routers every five seconds and checks every SDN link from both ends: is the IS-IS adjacency up, is the BFD session up, does a ping make it across, how much traffic is on it, and has a fault been injected on purpose. Each router reports its BGP, IS-IS and BFD sessions, and every change of state lands in the event log with a timestamp.
The point is practice. When I break a link deliberately, as in the failover drill below, this is where I watch BFD catch the fault, IS-IS route around it and everything recover, with a time on each step, instead of piecing it together from four router consoles. It's also plain proof that the lab is real and running, not just a diagram. The routers are read through an SSH key that can only run one read-only collection script and only from the dashboard container, and this web server republishes a curated snapshot with route tables left out and addresses masked, the same approach as the cluster telemetry on the data center page.
| Time | Source | Event |
|---|---|---|
| Waiting for events… | ||
| Router | Guest | Node | AS | Loopback | Advertises | Role |
|---|---|---|---|---|---|---|
| R1 | LXC 111 | PVEServer | 65001 | 10.255.0.xxx/32 | 172.16.1.xxx/24 | iBGP route reflector |
| R2 | LXC 112 | PVE2 | 65001 | 10.255.0.xxx/32 | 172.16.2.xxx/24 | Edge router — eBGP to AS 65002, next-hop-self |
| R3 | LXC 113 | PVE3 | 65001 | 10.255.0.xxx/32 | 172.16.3.xxx/24 | Internal router, RR client |
| R4 | LXC 114 | PVE3 | 65002 | 10.255.0.xxx/32 | 198.51.100.xxx/24 | External peer |
IS-IS
The interior routing protocol for AS 65001. It only has to do one job: make every router's loopback reachable over the best path. All three links are point-to-point circuits, Level-2 only, with wide metrics — the way most service-provider backbones run it.
Level-2 · point-to-point · loopbacks passiveBFD
A lightweight hello protocol that runs underneath IS-IS and the eBGP session. It sends a packet every 300 ms and declares the neighbour dead after three misses, so a failure that doesn't take the link down — a silent drop somewhere in the path — is caught in under a second instead of waiting out IS-IS's 30-second hold time.
300 ms tx/rx · multiplier 3 · on every linkiBGP & the Route Reflector
Customer prefixes are carried in BGP, not IS-IS. Instead of a full mesh of iBGP sessions, R2 and R3 each peer only with R1, which reflects routes between them. Sessions run loopback to loopback, so they survive any single link failure as long as IS-IS can find another path.
R1 reflector · R2, R3 clients · update-source loopbackeBGP & Policy
R2 is the edge: it peers with R4 in a separate AS and rewrites next-hop to itself before passing routes inward. Both sides filter with prefix-lists and route-maps — AS 65001 only accepts the external customer prefix and only announces its own customer routes.
prefix-list + route-map in and out · BFD on the sessionCustomer Prefixes
Each router has a dummy interface standing in for a customer network. They're deliberately kept out of IS-IS so the only way the rest of the network learns them is BGP — the same split between infrastructure and customer routes a real provider keeps.
172.16.x.xxx/24 · 198.51.100.xxx/24 · BGP onlyFailover Test: BFD vs. the IS-IS Hold Timer
To prove BFD was doing its job, I made the link between R1 and R2 silently drop every packet using Linux traffic control, without taking the interface down — the kind of failure a router can't see from its own link state.
-
Silent Fault Injected
Every packet leaving R1 toward R2 is dropped. The interface stays up, so nothing at the link layer tells either router something is wrong.
R1 lk12 · 100% packet loss · interface still up -
BFD Detects It
Three BFD hellos go missing and the session goes down, which immediately tears down the IS-IS adjacency on that link.
IS-IS: "Adjacency to r2 (lk12) … changed from Up to Down, bfd session went down" · ≈1 s -
IS-IS Reroutes
R1 recalculates and reaches R2's loopback the long way round, through R3. The iBGP session between them never drops because it rides the loopbacks, and pings between customer prefixes keep working throughout.
R1 → R2 now via R3 · iBGP sessions stay up -
Recovery
With the fault removed, BFD and IS-IS come back within a couple of seconds, but the routers take around 30 seconds longer to re-advertise the restored link and shift traffic back — failure is fast, recovery is deliberately cautious. Tuning that is the next experiment.
adjacency up in ~2 s · shortest path restored ~30 s later
One real-world lesson from the build: IS-IS pads its hello packets to the full interface MTU to prove large frames make it across. The VXLAN links can only carry 1450 bytes, but the containers were resetting their interfaces to 1500, so every hello was silently too big and no adjacency ever formed. Pinning the MTU to 1450 fixed it — the same class of MTU mismatch that bites on real provider networks.