Inside the Linux Networking Stack: From NIC to App
The core purpose of a network is to establish communication between multiple nodes. That is already a difficult problem in almost any system, and it becomes even more demanding in critical real-time systems such as the power grid. But some of the hardest networking problems do not happen between different devices at all. They happen inside one single machine.
It is easy to imagine that once an Ethernet frame reaches the network interface, most of the work is already done and everything after that is trivial. It is not. Behind the black box that we casually call the Linux networking stack is another complex system, with its own objects, queues, memory, control paths and state transitions. This is the system I want to open in this part.
Why should we care about what happens inside one machine? Because some of the most difficult, bug-prone and timing-sensitive problems I have faced were not hidden in the coordination of a huge number of servers or in enormous traffic volumes. They were hidden inside one particular computer that had to process network traffic with extremely strict timing requirements, where a lot could go wrong at very different levels.
So now I want to open one of the most underestimated black boxes we constantly depend on when working with networked systems: the Linux networking stack. The route I tried several ways to explain the whole concept, and the approach I want to use now is the one I wish I had when I first faced this topic. Instead of showing isolated diagrams, we will keep the state of the whole system visible at the same time. We will see userspace, kernel state, memory structures and the communication between them while one simple story evolves. We will start before our tracked network interface exists in the kernel. Then we will watch Linux discover and initialize the NIC, create the interface representation, bring it administratively and operationally up, prepare the receive path, configure the interface, and finally trace one single Ethernet frame from the cable to a userspace application. Once that receive journey is complete, we will briefly turn the direction around and see how the same machinery participates when data is sent back out.
In this post we will discuss the Linux networking stack by following one machine as its networking state evolves. Instead of describing every observable value twice, we will use two languages in parallel: the text will explain why a transition happens and what it means, while the dashboard will show what exists now, what changed, which command was executed, and what the system returned. Long-lived state such as net_device, routing state and the RX machinery will stay in fixed places, while transient objects such as sk_buff will appear only when they actually exist.
There is one deliberate simplification. We follow one physical NIC and the Linux interface that represents it. State 0 therefore shows our tracked net_device, RX queue and routes as absent even though a real Linux system can already contain other interfaces, especially loopback. The goal is to make the birth and evolution of this particular interface visible without pretending that the rest of the operating system does not exist.
Main concepts and characters
Before the actual journey starts, we need a small vocabulary. The important point is that not everything in the diagram belongs to the same category. Some things are regions, some are long-lived objects, some are transient packet objects, and some are mechanisms that connect them.
Regions
- Hardware is where the physical NIC, PHY, firmware and hardware buses live. Hardware can read and write selected areas of RAM through DMA and can notify the CPU through interrupts.
- Kernel is the privileged part of the operating system. For our story it contains the network driver, interrupt handling, NAPI receive processing, the networking stack, sockets, routing state and kernel objects such as
net_deviceandsk_buff. - Userspace is where normal processes run. This includes applications and tools such as NetworkManager,
ip,nmcliandethtool.
Actors and control interfaces
- NIC is the physical network interface controller. It receives frames from the link and cooperates with the driver through queues, descriptors, DMA and interrupts.
- CPU executes the kernel and userspace code involved in our story.
- Driver is kernel code that knows how to control a specific NIC. It creates and configures much of the receive machinery that we will watch appear in the dashboard.
- Netlink is an important message-based interface between userspace and kernel subsystems. Networking tools can use it to request configuration changes or query state, and the kernel can also send notifications back to userspace. In the visualization, Netlink is better represented as a message channel or event log than as a permanent state object.
There is not one single Linux networking tool because there is not one single level of networking state. ip is the most direct tool in our story for inspecting and changing current kernel networking state. It works with objects such as links, addresses, routes and neighbors, mainly through rtnetlink. When we run ip link or ip route, we are asking what the kernel believes right now, or asking it to change that current state.
NetworkManager is a long-running userspace daemon that sits one level higher. It keeps connection profiles and policy, decides which configuration should be active, and then applies that configuration to the kernel. It can manage addresses, routes, DNS, autoconnect behavior and many device settings. This means NetworkManager has its own idea of what the machine should look like, while the kernel contains what is actually active at this moment.
nmcli is the command-line client for NetworkManager. It does not configure the driver directly. It talks to the NetworkManager daemon over D-Bus, and NetworkManager then uses Netlink and other kernel interfaces to apply the requested configuration.
ethtool looks lower in the stack. It focuses mostly on NIC and driver properties such as link modes, channels, ring sizes, offloads, interrupt coalescing and hardware statistics. Modern ethtool operations use an ethtool Generic Netlink family, while some older operations still use legacy ioctls.
So these tools do not control separate networks. They give us different views and control points over the same running system. ip is close to current kernel network state, nmcli works through NetworkManager's managed configuration and policy, and ethtool focuses on the device and driver layer. Their responsibilities can overlap, which is exactly why it is useful to keep all of them visible in the same story.
Persistent kernel and memory state
net_deviceis the central kernel representation of a network interface. It is long-lived relative to individual packets and contains or references interface state such as the interface identity, flags, MTU, queue configuration and driver operations.- RX queue is a logical receive path associated with the NIC and driver. For this visualization, each RX queue will own or reference the receive resources used to process incoming traffic.
- RX descriptor ring is a circular array of descriptors shared conceptually between the driver and the NIC. A descriptor does not store the complete Ethernet frame itself. Instead it describes a receive buffer and carries metadata and ownership or completion information that lets the NIC and driver coordinate packet reception.
- Receive buffers are areas of RAM prepared for incoming packet data. RX descriptors point to, or otherwise identify, these buffers so that the NIC can DMA frame data into them.
- Routing and forwarding state is separate kernel networking state. It is not a field inside
net_device. Linux organizes routes in routing tables and uses this information through its FIB, Forwarding Information Base, when making forwarding decisions. Routes describe how destinations should be reached and may reference a network interface, next hop, scope and other routing information. In the dashboard we will show the routes relevant to our tracked interface so commands such asip routehave visible kernel state to inspect.
Transient packet state
sk_buff is Linux's central packet representation while network data moves through much of the kernel networking stack. It contains metadata and references to packet data. It should be treated very differently from net_device: net_device belongs to the machine, while sk_buff belongs to a particular packet processing episode.
Mechanisms and events
- DMA, Direct Memory Access, allows the NIC to access DMA-mapped memory without the CPU copying each byte itself. During receive processing, the NIC can write the incoming frame into a prepared receive buffer in RAM.
- IRQ, Interrupt Request, is a hardware notification mechanism. For our story, the NIC can signal that receive work is available. The CPU then runs the interrupt handling path registered by the driver. With NAPI, the interrupt handler typically does not perform the whole receive path itself; it arranges for receive polling to continue in the appropriate kernel context.
Before anything moves: State 0
This is the deliberately empty state from which the story begins. The physical NIC exists in the machine, but Linux has not yet created the kernel representation and receive resources that we want to watch appear.
State 0 dashboard. The physical NIC exists, while the tracked net_device, RX queue, receive resources, routing entries and sk_buff are deliberately absent. The Terminal and Netlink panels are idle.
From this point on, the dashboard itself is the state description. Empty panels remain visible as NOT CREATED, NONE or NO ENTRIES, because seeing an object appear is part of the explanation. When a field changes, we will change it there instead of repeating the same information in a table below the image.
Now we can finally change something. The first transition happens when Linux discovers the NIC and the driver introduces it to the networking subsystem.
State 1: the NIC becomes a Linux network interface
The next transition starts when the system boots or the NIC is hot-plugged. Linux discovers the PCI (Peripheral Component Interconnect) device, matches it with the appropriate driver, and invokes the driver's probe() callback. probe() is the driver's initialization function for that device. It is where the driver starts setting up the hardware and creating the kernel-side structures needed to use it.
Inside that initialization path, the driver prepares access to the hardware, including mapping the device registers it needs to communicate with the NIC. It also creates its own internal state. The important moment for us comes when the driver creates and registers a net_device. This is the point where our NIC stops being merely a piece of detected hardware and becomes a network interface that the Linux
Comments
No comments yet. Start the discussion.