Skip to content
Go back

Network Namespaces by Hand, No Docker

By KingPin 15 min read
Network Namespaces by Hand, No Docker
Contents

You’ve Typed docker network create a Hundred Times

You run docker network create, then docker run -p 8080:80, and a web server appears on your host’s port. It works, so you never look underneath. Then one night a container can’t reach the internet, or a published port answers from the wrong place, and you discover you have no mental model of what Docker built.

The model is short. A Docker bridge network is a network namespace per container, a veth pair per container, one Linux bridge, one sysctl, and a few NAT rules. That is about six iproute2 commands plus three nftables rules. Build it by hand once and your 2 AM self will thank you, because every Docker networking bug turns out to be one of those pieces misbehaving.

If you want the other seven namespace types, that’s a different post. This one covers networking only.

Full example: Clone the working files at github.com/KingPin/sumguy-examples/linux/network-namespaces-by-hand/

The Parts List

Five pieces, one sentence each.

I tested every command below on iproute2 7.2.0 and nftables 1.1.7 on a recent kernel. The exceptions are called out where they come up.

Step 1: Two Namespaces

Terminal window
sudo ip netns add red
sudo ip netns add blue
sudo ip -n red link set lo up
sudo ip -n blue link set lo up

ip -n red ... is shorthand for ip netns exec red ip .... A fresh namespace has only a loopback interface, and it starts down. Bring it up or even ping 127.0.0.1 fails inside.

Step 2: The Bridge (Your Own docker0)

Terminal window
sudo ip link add br0 type bridge
sudo ip addr add 10.200.0.1/24 dev br0
sudo ip link set br0 up

Docker’s default bridge is docker0 on 172.17.0.0/16. I picked 10.200.0.0/24 on purpose. If you reuse 172.17.0.0/16 on a machine that runs Docker, you get two routes to the same subnet and a very confusing evening. Pick a range nothing else on the box uses.

The bridge gets an IP because it doubles as the gateway. Containers send their off-subnet traffic to it, and the host routes it onward.

Step 3: Cables and Plugs

Create a veth pair with one end already inside the namespace, then plug the host end into the bridge:

Terminal window
sudo ip link add veth-red type veth peer name eth0 netns red
sudo ip link set veth-red master br0 up
sudo ip -n red addr add 10.200.0.2/24 dev eth0
sudo ip -n red link set eth0 up
sudo ip -n red route add default via 10.200.0.1

Same again for blue, with 10.200.0.3:

Terminal window
sudo ip link add veth-blue type veth peer name eth0 netns blue
sudo ip link set veth-blue master br0 up
sudo ip -n blue addr add 10.200.0.3/24 dev eth0
sudo ip -n blue link set eth0 up
sudo ip -n blue route add default via 10.200.0.1

Test it:

Terminal window
sudo ip netns exec red ping -c1 10.200.0.3

If that works, you have a working L2 network. This is exactly what docker run does for each container: create the pair, name the inside end eth0, attach the outside end to the bridge. Run ip -br link on any Docker host and the vethXXXX@if.. entries are the plugs.

The default route matters. Without it, red can reach blue and the bridge but nothing beyond. Docker sets the same default route to the bridge’s IP (172.17.0.1 on docker0).

Step 4: Let the Host Route

Terminal window
sudo sysctl -w net.ipv4.ip_forward=1

Make it permanent with a file in /etc/sysctl.d/:

net.ipv4.ip_forward = 1

Docker flips this on at startup. Its documentation says it also sets the forward policy to drop, so the host doesn’t become an open router for everything. Keep that in mind for Step 6.

Step 5: Masquerade (Reaching the Internet)

The namespace has the private address 10.200.0.2. The internet can’t route replies to that, so the host rewrites the source address on the way out. Find your uplink interface (the one with the default route), call it eth0 on the host here, and add the rule:

Terminal window
sudo nft add table ip netns_lab
sudo nft add chain ip netns_lab postrouting '{ type nat hook postrouting priority srcnat; }'
sudo nft add rule ip netns_lab postrouting ip saddr 10.200.0.0/24 oifname "eth0" masquerade

That is the first of the three rules. Masquerade is the same job as SNAT, except it looks up the outgoing interface’s address at packet time. That makes it the right choice for a laptop or a DHCP box where the address changes.

Then check the result:

Terminal window
sudo ip netns exec red ping -c1 1.1.1.1

Step 6: Port Forwarding (The -p 8080:80 Part)

Add a prerouting chain that rewrites the destination of anything arriving on port 8080:

Terminal window
sudo nft add chain ip netns_lab prerouting '{ type nat hook prerouting priority dstnat; }'
sudo nft add rule ip netns_lab prerouting tcp dport 8080 dnat to 10.200.0.2:80
sudo nft add chain ip netns_lab output '{ type nat hook output priority dstnat; }'
sudo nft add rule ip netns_lab output ip daddr 10.200.0.1 tcp dport 8080 dnat to 10.200.0.2:80

That is the second rule, plus its output twin. Prerouting only sees packets that arrive on an interface; the host’s own connections go through the output hook instead, so a test from the host needs both. The DNAT rewrites the destination before routing, so the kernel now sees a packet for 10.200.0.2 and forwards it through the bridge.

Forwarded packets then hit the forward hook. On a host with a default-drop forward policy (Docker sets one, and so do many firewall setups), they die there unless something accepts them. A self-contained forward chain that accepts them (the third rule, plus its helpers). Do this on a lab box: on a live Docker host a second forward chain with policy drop blocks new connections to Docker’s published ports, so use policy accept there, as the companion script does.

Terminal window
sudo nft add chain ip netns_lab forward '{ type filter hook forward priority filter; policy drop; }'
sudo nft add rule ip netns_lab forward ct state established,related accept
sudo nft add rule ip netns_lab forward iifname br0 accept
sudo nft add rule ip netns_lab forward ip daddr 10.200.0.2 tcp dport 80 accept

The gotcha: in nftables, an accept in one base chain ends only that chain. If another table hooked on forward says drop, the packet is still dropped. So adding accept rules in your own table does not override Docker’s drop policy. You have to put the accept in the place that is dropping, which for Docker’s iptables backend is the DOCKER-USER chain. Docker’s iptables documentation describes it as a placeholder for user rules that are processed before DOCKER-FORWARD and DOCKER.

Step 7: DNS Inside the Namespace

A namespace shares the host’s filesystem, so it reads the host’s /etc/resolv.conf. On a systemd-resolved box that holds 127.0.0.53, which is a loopback address in the host’s namespace. Inside red, 127.0.0.53 is red’s own loopback, where nothing listens. DNS fails and you stare at it.

man ip-netns documents the fix. Files under /etc/netns/NAME/ are bind-mounted over /etc/ for processes started with ip netns exec:

Terminal window
sudo mkdir -p /etc/netns/red
echo "nameserver 1.1.1.1" | sudo tee /etc/netns/red/resolv.conf
sudo ip netns exec red cat /etc/resolv.conf

Docker solves the same problem differently. It writes its own resolv.conf and bind-mounts it into each container. On a user-defined network the nameserver in that file is 127.0.0.11, Docker’s embedded DNS server, which resolves container names. On the default bridge it copies the host’s resolvers.

Step 8: Test the Whole Thing

Terminal window
sudo ip netns exec red python3 -m http.server 80 &
sudo ip netns exec blue curl -s -o /dev/null -w '%{http_code}\n' http://10.200.0.2/
curl -s -o /dev/null -w '%{http_code}\n' http://10.200.0.1:8080/

The first curl goes blue to red across the bridge. The second goes from the host to the bridge IP on port 8080, and the output DNAT sends it to red. Traffic from another machine takes the prerouting rule instead.

Notice I used 10.200.0.1, not 127.0.0.1. DNAT on loopback traffic needs net.ipv4.conf.all.route_localnet=1 and extra masquerading, or the kernel treats 127.0.0.1 as a bad source address on the way out of the loopback. Docker’s answer to that is a small userland process, which brings us to the next section.

Teardown

Terminal window
sudo nft delete table ip netns_lab
sudo ip netns del red
sudo ip netns del blue
sudo ip link del br0

Deleting a namespace destroys the interfaces inside it, and a veth pair dies when either end is gone. I checked: after ip netns del blue, veth-blue on the host no longer exists. You don’t delete the veths yourself.

Never name your table ip nat or ip filter on a Docker host. The iptables-nft shim stores Docker’s rules in tables with exactly those names, so nft add table ip nat silently joins Docker’s table and nft delete table ip filter wipes Docker’s firewall. That is why every command above uses netns_lab: teardown removes only its own rules.

What Docker Actually Does Differently

You now know 90 percent of it. The remaining differences:

The default bridge. On this machine docker0 is 172.17.0.1/16, matching the docs and defaults you’ve seen everywhere.

docker-proxy. When you publish a port, Docker adds DNAT rules like Step 6. By default it also starts a docker-proxy process that listens on the host port. The daemon reference lists userland-proxy as “Use userland proxy for loopback traffic”, default true. That process exists for the loopback case above: DNAT can’t cleanly handle 127.0.0.1, so a proxy accepts the connection and opens a new one to the container. Setting it to false drops the proxy and relies on rules alone:

{
"userland-proxy": false
}

Restart the daemon after editing /etc/docker/daemon.json. Check the binary Docker would use with docker info; mine reports /usr/bin/docker-proxy.

The firewall backend. Docker’s documentation says it creates rules with iptables by default and also supports nftables, marked experimental, through the firewall-backend option. Your hand-built rules above are plain nftables. Docker’s rules, on a default install, are iptables rules. The next section covers how the two coexist.

DOCKER-USER. Docker’s rule chains run before anything you append to FORWARD. Its docs say rules appended to FORWARD are processed after Docker’s, so put your policy in DOCKER-USER instead.

Where the namespaces live. Docker doesn’t register its namespaces in /var/run/netns, so ip netns list shows nothing for containers. Ask Docker for the path instead:

Terminal window
docker inspect -f '{{.NetworkSettings.SandboxKey}}' mycontainer

On my host that prints a path like /var/run/docker/netns/ac927f1858ad. Then step in with nsenter:

Terminal window
sudo nsenter --net=/var/run/docker/netns/ac927f1858ad ip addr

You run your host’s tools (ss, tcpdump, ip) inside the container’s network, which is useful when the image has no tools of its own. You can also symlink that file into /var/run/netns/ under a friendly name and use plain ip netns exec.

nftables vs iptables for These Rules

Use nftables. It is the kernel’s current packet filtering framework, rules live in tables you create and delete by name, and one nft list ruleset shows everything. Old tutorials use iptables, so the translation, once:

Jobnftablesiptables
Masqueradenft add rule ip netns_lab postrouting ip saddr 10.200.0.0/24 oifname "eth0" masqueradeiptables -t nat -A POSTROUTING -s 10.200.0.0/24 -o eth0 -j MASQUERADE
Port forwardnft add rule ip netns_lab prerouting tcp dport 8080 dnat to 10.200.0.2:80iptables -t nat -A PREROUTING -p tcp --dport 8080 -j DNAT --to-destination 10.200.0.2:80
Allow forwardnft add rule ip netns_lab forward iifname br0 acceptiptables -A FORWARD -i br0 -j ACCEPT
Remove your stuffnft delete table ip netns_labdelete each rule by number or -D

The confusing part is that iptables on a current distro is usually a shim over nftables:

Terminal window
iptables -V
iptables v1.8.13 (nf_tables)

That (nf_tables) means the iptables command translates your rule into nftables objects. Both tools write to the same kernel layer, so the two styles can coexist and even trip over each other. If a rule seems to vanish, check nft list ruleset and look for a table you didn’t write.

When This Blows Up

Check these in order. They cover most of the 2 AM calls.

  1. ip_forward is 0. Pings reach the bridge IP but nothing goes further. Run sysctl net.ipv4.ip_forward.
  2. Forward chain drops it. Counters on a drop rule climb while your test runs. Add the accept where the drop lives (DOCKER-USER on a Docker host).
  3. Firewalld or ufw rewrote the ruleset. They own the firewall on distros that ship them and can flush or shadow your tables on reload. Put your rules in their zone or policy system instead.
  4. rp_filter throws out the reply. Reverse path filtering drops packets whose return route looks wrong. Check sysctl net.ipv4.conf.all.rp_filter, and look at ip route get <source> to find the mismatch.
  5. Bridged traffic hits the forward chain. If the br_netfilter module is loaded and net.bridge.bridge-nf-call-iptables is 1, frames crossing the bridge between two namespaces are shown to the IP firewall too. Same-subnet traffic that worked suddenly drops under a default-drop policy. That sysctl does not exist until the module is loaded, so sysctl: cannot stat means it isn’t.
  6. DNS fails, ping to an IP works. Back to Step 7. Check what ip netns exec red cat /etc/resolv.conf says.
  7. Masquerade matches the wrong interface. If you changed uplinks (Wi-Fi to Ethernet, VPN up), oifname "eth0" no longer matches. Drop the oifname match or fix it.

One honesty note. I built this walkthrough inside an unprivileged user namespace (unshare -rmn), which has no route to the internet and can’t load the nft_nat module. So Steps 1 through 4, the forward chain, the DNS file handling, and teardown are tested end to end, while the masquerade and the DNAT port forward are verified for syntax only. The companion script runs as root on a real box; run it and tell me if the forward misbehaves.

The SumGuy Take

Docker isn’t doing anything you can’t do with a screwdriver set. It’s a forklift, and for ten containers a forklift is what you want. But when a container’s network goes sideways, knowing it’s a bridge, a cable, a route, and a few rules means you read ip -br link, ip route, and nft list ruleset and fix it, instead of restarting the daemon and hoping.

Common Questions

Can I see Docker’s network namespaces with ip netns?

No. ip netns list reads only /var/run/netns, and Docker keeps its namespace files under /var/run/docker/netns instead. Find a container’s path with docker inspect -f '{{.NetworkSettings.SandboxKey}}', then use nsenter --net=<path>, or symlink that file into /var/run/netns to make ip netns see it.

Do network namespaces survive a reboot?

No. Namespaces, veth pairs, bridges, and nftables rules created at runtime all vanish on reboot. To persist them, wrap the up commands in a systemd oneshot unit that runs after the network is up, or describe the bridge and veths in systemd-networkd files. The /etc/netns/*/resolv.conf files persist, since they are ordinary files.

Do I need nftables, or can I use iptables for this?

You can use either, since both write to the same kernel packet filtering layer. Nftables is the better choice for new setups because you can create and delete a whole named table in one command. Every nft rule in this post has an iptables equivalent in the comparison table above.

Does this work with rootless Docker, Podman, or WSL2?

Not directly. Rootless Docker and rootless Podman run their networking inside a user namespace with a userspace network stack, so there is no host bridge to inspect the same way. WSL2 runs a real Linux kernel, so the commands work there if the kernel includes bridge and nftables support. Test with ip link add br0 type bridge first.


Share this post on:

Send a Webmention

Written about this post on your own site? Send a webmention and it'll show up above once verified.


Previous Post
gVisor vs Firecracker vs Kata for Agents
Next Post
Kill-A-Watt Tour of Every Box in the Rack

Discussion

Powered by Garrul . Sign in with GitHub or Google, or post anonymously.

Related Posts