Skip to content
Go back

VLAN Trunking Gotchas Nobody Warns You

By KingPin 10 min read
VLAN Trunking Gotchas Nobody Warns You
Contents

Your VLANs Are Configured. Nothing Reaches Anything.

You followed the basics. Access ports assigned. Trunk port tagged. pfSense has its VLAN interfaces, DHCP is handing out leases, firewall rules look clean. Then you add a second switch, or spin up a Proxmox VM, or drop a container onto a macvlan network, and suddenly a device that should reach the NAS gets nothing. No error. No log line pointing at the problem. Just silence.

This is where VLANs get interesting, past the point every basics guide stops. Two devices can both be configured correctly and still not talk, because correctly meant something slightly different on each end. Below are the ones I keep tripping over, each with the command or setting that proves it. The short version: most “broken VLAN” tickets come down to an untagged frame landing in a VLAN nobody expected.

The Native VLAN Mismatch: Two Switches Disagreeing

Every trunk port has a native VLAN (Cisco’s term) or a PVID (everyone else’s term) for the one VLAN it sends and receives untagged. If switch A’s trunk treats untagged frames as VLAN 1 and switch B’s trunk treats the same untagged frames as VLAN 99, the two switches disagree about what an untagged frame means, and neither is wrong by its own config.

Cisco gear at least tells you. Cisco Discovery Protocol carries the native VLAN in its advertisements, and a mismatch throws a console warning like %CDP-4-NATIVE_VLAN_MISMATCH: Native VLAN mismatch discovered on GigabitEthernet0/1 (1), with Switch2 GigabitEthernet0/2 (99). Cisco switches running PVST+ or Rapid PVST+ go further: those BPDUs carry a VLAN ID, so spanning tree spots the mismatch and blocks the port as PVID inconsistent until you fix it. Plain 802.1D or RSTP on a non-Cisco switch has no such field and never notices.

Cheap home lab switches don’t run CDP and won’t warn you at all. The traffic lands in the wrong place, silently, while you spend an hour convinced your firewall rules are broken. Fix it by matching the native VLAN or PVID on both ends of every trunk:

interface GigabitEthernet0/1
switchport trunk native vlan 99

On a TP-Link or Netgear switch that’s the PVID setting on the trunk port, on both switches, and it has to be the same number.

Membership and PVID Live on Different Pages

On TP-Link’s Easy Smart line (the TL-SG108E from the basics post included), 802.1Q VLAN membership, which ports are tagged or untagged for a given VLAN, lives under VLAN > 802.1Q VLAN. PVID lives on a separate page, VLAN > 802.1Q VLAN PVID Setting, and every port defaults to PVID 1.

Set a port as an untagged member of VLAN 20 and stop there, and that port’s PVID is still 1. Untagged frames coming in get tagged VLAN 1 internally before the switch looks at your VLAN 20 membership table at all. Symptom: the device on that port gets a lease from your VLAN 1 scope, or no lease if VLAN 1 has no DHCP server listening on it. Two pages, one setting each, and both have to match or the port does nothing useful.

Don’t Configure Your Own Way Out of the Web UI

Removing your own management port from VLAN 1, or repointing its PVID, cuts your browser off from the switch mid change. The switch is fine. You’re just no longer in the VLAN that can reach it. Change your PC’s port last, and leave one access port permanently pinned to the switch’s actual management VLAN so you keep a way back in. Ask anyone who’s driven back to a rack at midnight for a factory reset how they know this.

Proxmox: The Bridge Is VLAN Aware, the Switch Port Might Not Be

A VLAN aware Linux bridge tags VM traffic in software:

auto vmbr0
iface vmbr0 inet static
address 192.168.1.5/24
gateway 192.168.1.1
bridge-ports eno1
bridge-stp off
bridge-fd 0
bridge-vlan-aware yes
bridge-vids 2-4094

Set a VM’s NIC “VLAN Tag” field to 20 in the Proxmox GUI, and Proxmox tags its frames going out eno1. None of that matters if the physical switch port feeding eno1 is still an access port. An access port expects untagged frames. Depending on the switch, a frame arriving with a VLAN 20 tag gets dropped or handled in ways you didn’t configure, and either way your VM’s VLAN 20 traffic never gets where it’s going. The switch port carrying a VLAN aware bridge’s uplink has to be a trunk, every time, no exceptions.

Outside Proxmox, a plain Linux box gets its own tagged subinterface:

Terminal window
ip link add link eth0 name eth0.10 type vlan id 10
ip link set eth0.10 up
ip addr add 192.168.10.5/24 dev eth0.10

Netplan wants the same information as YAML:

network:
version: 2
ethernets:
eth0:
dhcp4: false
vlans:
eth0.10:
id: 10
link: eth0
addresses: [192.168.10.5/24]

eth0 itself needs no address if it only carries tagged VLANs. Give it one anyway and the server also lives on whatever VLAN the switch port’s PVID assigns to untagged frames. If that PVID is still 1, you’ve got the same trap as the switch section above, just on a server now.

Seeing the Tag, or Not, Because the NIC Ate It

tcpdump -e -i eth0 vlan prints the 802.1Q header for tagged frames on that interface. Before you trust an empty capture as proof tagging is broken, check NIC offload:

Terminal window
ethtool -k eth0 | grep vlan

If rx-vlan-offload is on, the NIC strips the tag in hardware and hands the kernel a plain frame with the VLAN ID stored as metadata. Recent libpcap reads that metadata back and shows the tag, but older builds and some drivers don’t, and that’s how you get a capture that swears the wire is untagged. Turn offload off while you debug and capture again:

Terminal window
ethtool -K eth0 rxvlan off

Turn it back on with ethtool -K eth0 rxvlan on when you’re done. If tags show up with offload off, your switch was fine all along.

Router on a Stick Shares One Cable, Both Directions

If your router has one trunk link back to the switch and routes between VLANs over it, the classic pfSense setup from the basics post, every inter-VLAN packet travels up that trunk to the router and back down the same trunk to reach the other VLAN. A 1 Gbps trunk is full duplex: 1 Gbps from switch to router and 1 Gbps from router to switch, independently. A single transfer from your NAS on VLAN 40 to a laptop on VLAN 10 uses both directions at once, up to the router on the VLAN 40 tag and back down on the VLAN 10 tag. That one backup job can fill the trunk in both directions, and every other routed flow, internet traffic included, queues behind it. If iperf3 between two VLANs tops out well under the same test inside one VLAN, this cable is why. The fixes are a faster trunk (2.5 or 10 Gbps), a LAG, or putting the two chatty devices in the same VLAN.

Broadcasting separate SSIDs onto separate VLANs means the AP’s uplink to the switch carries tagged traffic for every one of those VLANs: a trunk port, not access. The AP’s own management interface, the one you log into to configure it, typically sits on the native or untagged VLAN of that trunk. Get the native VLAN wrong and either the AP vanishes from your management network, or its management traffic leaks onto whatever VLAN you left untagged instead.

Docker Macvlan: The Host Can’t Talk to Its Own Container

Terminal window
docker network create -d macvlan \
--subnet 192.168.20.0/24 --gateway 192.168.20.1 \
-o parent=eth0.20 vlan20net

Point parent at a subinterface that doesn’t exist yet and Docker creates eth0.20 automatically. Every other device on VLAN 20 reaches the container fine. The Proxmox host or bare metal box running Docker cannot, because the Linux kernel’s macvlan driver won’t hairpin traffic between a parent interface and its own macvlan children. That’s the driver working as designed, not a firewall rule and not a routing bug. Fix it with a second macvlan shim interface on the host bound to the same parent and given its own IP in that subnet, or attach the container to a second bridge network for host access and keep macvlan for LAN facing traffic only.

Debug Checklist

Work through these before you touch a single firewall rule:

Terminal window
# Is the tag actually on the wire
tcpdump -e -i eth0 vlan
# Is the NIC hiding it from you
ethtool -k eth0 | grep vlan
# What VLAN subinterfaces exist and their parent
cat /proc/net/vlan/config
ip -d link show eth0.10
# What's the VLAN aware bridge actually doing per port (Proxmox)
bridge vlan show

On the switch side: confirm PVID and 802.1Q membership independently, on both ends of every trunk, and confirm the native VLAN matches on both ends too. Skip one of those checks and a “fixed” VLAN breaks again a week later, usually right after you add a device.

Common Questions

Do I need to lower MTU for VLAN tagging?

No. The 802.1Q tag adds 4 bytes, extending the maximum Ethernet frame from 1518 to 1522 bytes. Switches and NICs have supported this frame size since the IEEE 802.3ac amendment in 1998. Leave your MTU at the standard 1500 unless you have an unrelated jumbo frame reason to change it.

Can a Proxmox host with one NIC carry multiple VLANs?

Yes. A single NIC carries every VLAN once the Proxmox bridge on it has bridge-vlan-aware yes and the switch port it plugs into is a trunk. Each VM gets its VLAN from the “VLAN Tag” field on its network device. The Proxmox host’s own management IP stays on the trunk’s untagged native VLAN, or on a vmbr0.X interface.

What’s the difference between native VLAN and PVID?

Native VLAN is Cisco’s name for the VLAN a trunk port uses for untagged frames. PVID (Port VLAN ID) is the name TP-Link, Netgear, and most Linux tooling use, and PVID applies to every port, access ports included. On a trunk, the two settings do the same job, so both ends of the trunk must agree on the number.

Should I move my trunks off VLAN 1 in a home lab?

Yes, when the switch allows it. Using VLAN 1 as the native VLAN on trunks is the default that VLAN hopping attacks (double tagging) rely on, so Cisco hardening guides move the native VLAN to an unused ID. On a home lab switch, the bigger win is not leaving user devices on VLAN 1 by accident.


Share this post on:

Send a Webmention

Written about this post on your own site? Send a webmention and it'll show up above once verified.


Next Post
journald Structured Logging, No ELK

Discussion

Powered by Garrul . Sign in with GitHub or Google, or post anonymously.

Related Posts