Ethtool: Tuning Network Interface Buffers, Auditing Hardware Offloads, and Diagnosing Dropped Frames in Production
Yet when you log into the servers, your monitoring dashboards present a maddening contradiction. CPU utilisation is sitting comfortably at twenty-five percent, memory usage is well within safe margins, and standard Linux diagnostic tools insist that everything is running perfectly. There are no crashed processes, no saturated disk queues, and not a single error reported in the system logs.
The reason for this phantom failure is a blind spot built into standard operating system utilities. Tools like netstat, ss, and ip only see network traffic that has successfully made it past the hardware barrier and into the Linux kernelβs memory buffers. When network packets are dropped at the physical wire before the operating system even knows they exist, your server remains completely oblivious, reporting zero errors while your application quietly starves.
To uncover what is actually happening, you need a diagnostic tool that bypasses software abstractions and interrogates the physical silicon of your network interface card directly: ethtool.
Before you dive into complex network topologies, there is one essential triage command that every engineer should run the moment silent network issues strike:
sudo ethtool -S eth0 | grep -E "(drop|miss|overrun|fifo|buff|error)"
This single command pulls raw hardware counters directly from the network interface controller (NIC) registers. If counters like rx_missed_errors or rx_no_buffer_count are ticking upward while your software metrics report zero dropped packets, you have instantly located your culprit: incoming data is overwhelming the physical hardware queue and vanishing before Linux can even touch it.
1. What It Does in Plain English
The ethtool utility is the administrative bridge between the Linux operating system and physical network hardware. While traditional networking commands inspect traffic from within the operating system's software layers, ethtool speaks directly to the driver hooks (ethtool_ops) controlling the network cardβs underlying silicon.
Think of it as a low-level diagnostic probe and tuning console for your physical network port. With ethtool, you can query raw hardware error counters, expand or shrink direct memory access (DMA) queue capacities, toggle silicon-level offloading engines, negotiate physical link speeds, and even read real-time temperature and optical power levels directly from connected fiber transceivers.
2. Core Flags & Quick Start
The ethtool suite provides dedicated flags for inspecting and configuring different facets of the network adapter:
| Flag / Option | Functional Domain | Operational Purpose |
|---|---|---|
-S |
Device Statistics | Queries raw hardware and MAC/PHY driver counters directly from the NIC silicon. |
-g / -G |
Ring Buffer Descriptors | Inspects (-g) and alters (-G) the size of RX and TX DMA descriptor queues. |
-k / -K |
Hardware Offloads | Displays (-k) and modifies (-K) silicon-level acceleration features (TSO, GRO, LRO, Checksums). |
-m / --module-info |
Optical Telemetry | Queries Digital Optical Monitoring (DOM/DDM) EEPROM registers on SFP+/QSFP transceivers. |
-c / -C |
Interrupt Coalescing | Inspects (-c) and tunes (-C) hardware interrupt throttling timers and packet thresholds. |
-i / --driver |
Driver Information | Displays the running kernel driver module name, firmware version, and PCIe bus address. |
-p / --identify |
Physical Identification | Flashes the physical LED indicators on the NIC chassis port for easy identification in a rack. |
The Essential Baseline Invocation
To construct a comprehensive baseline of any network interface, query its physical link state, autonegotiation parameters, and supported media types:
sudo ethtool eth0
Settings for eth0:
Supported ports: [ FIBRE ]
Supported link modes: 10000baseT/Full
25000baseCR/Full
25000baseSR/Full
Supported pause frame use: Symmetric Receive-only
Supports auto-negotiation: Yes
Supported FEC modes: None BaseR RS
Speed: 25000Mb/s
Duplex: Full
Auto-negotiation: on
Port: Direct Attach Copper
PHYAD: 0
Transceiver: internal
Link detected: yes
3. The Architecture of the Linux Network Boundary
To understand why packets can vanish without a trace, you need to follow the physical path an Ethernet frame takes when it strikes your network card.
- DMA Ring Allocation: When the network driver loads, it reserves an array of memory buffers in host RAM called the RX Descriptor Ring and shares these memory addresses with the NIC.
- Hardware Ingress: When an electrical or optical signal arrives at the PHY layer, the NIC's Direct Memory Access (DMA) engine claims an available descriptor and writes the raw packet directly into host RAM.
- Interrupt Generation: Once the data is deposited into RAM, the NIC fires a PCIe MSI-X hardware interrupt to alert the CPU.
- NAPI Scheduling: Linux responds using the New API (NAPI) framework. It temporarily disables hardware interrupts for that queue to avoid overwhelming the CPU, scheduling a software interrupt (
NET_RX_SOFTIRQ) handled by kernel worker threads (ksoftirqd). - Driver Polling & Allocation: The driver extracts the frames from the ring buffer, wraps them into kernel socket buffer objects (
sk_buff), and passes them up to the TCP/IP stack. - Hardware Drop on Depletion: If packets arrive faster than the CPU can drain the RX descriptor ring, the ring runs completely out of slots. The NIC hardware cannot write the incoming frames to RAM and immediately drops them directly on the wire.
Because these dropped frames never reach host memory, no sk_buff structure is ever created. The Linux kernel remains entirely unaware of their existence, leaving standard monitoring software completely blind. The only record of these dropped frames lives inside the hardware registers on the network card itself.
4. Five Real-World Production Use Cases
Case 1: Diagnosing Silent Hardware Packet Drops and Overruns
Operational Scenario
A high-frequency trading node experiences sporadic latency spikes during the market opening bell. High-level socket monitors (netstat -s, ss -s) report zero packet drops and healthy TCP connections. You must determine whether packets are being dropped by the physical hardware queues before the kernel can construct a socket buffer.
Diagnostic Invocation
sudo ethtool -S eth0 | grep -E "(drop|miss|overrun|fifo|buff|error)"
Raw Command Output
rx_dropped: 0
tx_dropped: 0
rx_missed_errors: 1849204
rx_fifo_errors: 0
rx_no_buffer_count: 1849204
rx_over_errors: 0
rx_crc_errors: 0
rx_frame_errors: 0
alloc_rx_buff_failed: 0
rx_discards_phy: 0
Line-by-Line Telemetry Analysis
rx_dropped: 0: The driver-level software abstraction reports zero dropped frames. This counter is frequently misleading because many network drivers only increment this metric when a drop occurs after the driver has successfully claimed the descriptor.rx_missed_errors: 1849204: A critical hardware-level indicator. The MAC sublayer received valid Ethernet frames, but the controller could not transfer them to host RAM because no RX descriptors were available.rx_no_buffer_count: 1849204: Confirms the root cause. This vendor-specific counter (frequently found on Intelixgbeandi40echipsets) increments synchronously withrx_missed_errors, proving that the hardware descriptor ring was exhausted when the DMA transfer was attempted.rx_fifo_errors: 0&rx_crc_errors: 0: Confirms that the physical signal, transceiver integrity, and internal PCIe FIFO queues are operating properly without link degradation or electrical noise.
What the Admin Does Next
The telemetry confirms that incoming traffic bursts are exhausting the hardware ring queues. The engineer should:
1. Inspect current ring buffer capacities with ethtool -g eth0 and expand them to the hardware maximum.
2. Check /proc/net/softnet_stat to verify whether CPU cores are exhausting their NAPI polling budgets.
3. Pin the network cardβs MSI-X interrupt vectors to dedicated, non-isolated CPU cores using smp_affinity.
Case 2: Mitigating Microburst Drops via RX/TX Ring Buffer Sizing
Operational Scenario
A distributed storage cluster (such as Ceph or NVMe-over-Fabrics) suffers severe throughput drops during write operations. High-intensity I/O microbursts lasting under 50 milliseconds are overwhelming default hardware queues, triggering TCP window collapses and severe tail latencies.
Diagnostic and Tuning Invocations
Step A: Inspect Ring Sizing Constraints
sudo ethtool -g eth0
Ring parameters for eth0:
Pre-set maximums:
RX: 4096
RX Mini: n/a
RX Jumbo: n/a
TX: 4096
Current hardware settings:
RX: 512
RX Mini: n/a
RX Jumbo: n/a
TX: 512
Step B: Expand Descriptors to Hardware Maximums
sudo ethtool -G eth0 rx 4096 tx 4096
Step C: Verify Reconfigured Ring Allocation
sudo ethtool -g eth0
Ring parameters for eth0:
Pre-set maximums:
RX: 4096
RX Mini: n/a
RX Jumbo: n/a
TX: 4096
Current hardware settings:
RX: 4096
RX Mini: n/a
RX Jumbo: n/a
TX: 4096
Line-by-Line Engineering Analysis
Pre-set maximums: RX: 4096, TX: 4096: Identifies the hardware boundary enforced by the NIC ASIC. The controller can track up to 4,096 descriptors per ring queue.Current hardware settings: RX: 512, TX: 512: The operating system vendor configured a conservative default (512 descriptors) to minimize base memory overhead and reduce CPU cache pollution.sudo ethtool -G eth0 rx 4096 tx 4096: Reallocates the circular DMA descriptor ring in host RAM. At 4,096 entries, the interface can absorb an eightfold surge in simultaneous ingress frames during microbursts before the silicon exhausts its descriptors.
| Queue Profile | Configured Capacity | Microburst Absorption Capacity | Operational Trade-off |
|---|---|---|---|
| Factory Default | 512 Descriptors | ~512 Packets | Low RAM/cache usage, but drops sudden bursts |
| Tuned Maximum | 4096 Descriptors | ~4096 Packets | Absorbs 8x traffic volume, minor RAM overhead |
What the Admin Does Next
Monitor hardware registers with ethtool -S eth0 to confirm that rx_missed_errors stops incrementing during heavy I/O workloads. If microburst drops continue despite maximum queue sizing, tune interrupt coalescing (ethtool -C eth0 rx-usecs ...) or distribute queues across more CPU cores using Receive Side Scaling (RSS).
Case 3: Auditing and Toggling Hardware Acceleration & Offload Engines
Operational Scenario
After deploying an overlay network utilizing VXLAN/Geneve encapsulation or enabling jumbo frames (MTU 9000), throughput degrades drastically. Packets are intermittently corrupted or dropped due to incompatible hardware offloading logic in the NICβs physical ASIC.
Diagnostic and Tuning Invocations
Step A: Audit All Offload Features
sudo ethtool -k eth0 | grep -E "(segmentation|checksum|receive-offload)"
rx-checksumming: on
tx-checksumming: on
tx-checksum-ipv4: on
tx-checksum-ip-generic: off [fixed]
tx-checksum-ipv6: on
tx-checksum-fcoe-crc: off [fixed]
tx-checksum-sctp: on
scatter-gather: on
tx-scatter-gather: on
tx-scatter-gather-fraglist: off [fixed]
tcp-segmentation-offload: on
tx-tcp-segmentation: on
tx-tcp-ecn-segmentation: on
tx-tcp-mangleid-segmentation: on
tx-tcp6-segmentation: on
generic-segmentation-offload: on
generic-receive-offload: on
large-receive-offload: off [fixed]
rx-gro-hw: off [fixed]
Step B: Disable Problematic Offload Engines
sudo ethtool -K eth0 tso off gso on rx off tx off
Step C: Verify State Mutation
sudo ethtool -k eth0 | grep -E "(tcp-segmentation-offload|checksumming)"
rx-checksumming: off
tx-checksumming: off
tx-checksum-ipv4: off
tx-checksum-ipv6: off
tx-checksum-sctp: off
tcp-segmentation-offload: off
tx-tcp-segmentation: off
tx-tcp-ecn-segmentation: off
tx-tcp-mangleid-segmentation: off
tx-tcp6-segmentation: off
Line-by-Line Engineering Analysis
tcp-segmentation-offload (TSO): Allows the kernel's network stack to construct large TCP payloads (up to 64KB) and offload the packet slicing into MTU-sized frames down to the NIC silicon. If the NIC firmware contains microcode bugs with encapsulated packets, it can generate malformed frame headers or invalid outer IP checksums.generic-segmentation-offload (GSO): The software fallback mechanism within the kernel. When TSO is disabled (tso off), the kernel performs segmentation in software before the packet reaches the driver, bypassing buggy NIC hardware.generic-receive-offload (GRO): The software aggregation engine that reassembles contiguous incoming packets into single large structures before handing them to the IP stack. Unlike Large Receive Offload (LRO), GRO preserves end-to-end transport layer headers and is safe for routing and forwarding nodes.
64KB Super-Packet"] --> H1["NIC ASIC Hardware:
Slices into 1500B MTU Frames"] --> W1["Physical Wire"] end subgraph GSO_Path["Software GSO Fallback Path (Kernel Emulation)"] direction LR K2["Kernel Stack:
64KB Super-Packet"] --> S2["Kernel GSO Engine:
Slices into 1500B MTU Frames"] --> D2["Driver & NIC Passthrough"] --> W2["Physical Wire"] end
What the Admin Does Next
Test payload delivery across the overlay network. If disabling TSO resolves packet corruption and drops, keep hardware segmentation disabled while upgrading the physical NIC's firmware and kernel driver to the latest stable vendor release.
Case 4: Triaging Physical Link Layer, Autonegotiation, and Speed Constraints
Operational Scenario
Following an emergency top-of-rack switch maintenance window in a remote data center, a dual-port 100GbE interface operates at only a fraction of its intended bandwidth. You must verify link negotiation parameters and illuminate the physical port LED so an on-site technician can locate the exact cable.
Diagnostic and Control Invocations
Step A: Interrogate Physical Link Constraints
sudo ethtool eth1
Settings for eth1:
Supported ports: [ FIBRE ]
Supported link modes: 10000baseSR4/Full
40000baseSR4/Full
100000baseSR4/Full
Supported pause frame use: Symmetric
Supports auto-negotiation: Yes
Supported FEC modes: None BaseR RS
Advertised link modes: 10000baseSR4/Full
40000baseSR4/Full
100000baseSR4/Full
Advertised pause frame use: Symmetric
Advertised auto-negotiation: Yes
Advertised FEC modes: RS
Speed: 10000Mb/s
Duplex: Full
Auto-negotiation: on
Port: FIBRE
PHYAD: 1
Transceiver: internal
Link detected: yes
Step B: Enforce Correct Speed, Duplex, and FEC Profiles
sudo ethtool -s eth1 speed 100000 duplex full autoneg off
Step C: Trigger Physical Port Identification Beacon
sudo ethtool -p eth1 30
Line-by-Line Telemetry Analysis
Supported link modes: 10000baseSR4/Full, 40000baseSR4/Full, 100000baseSR4/Full: Confirms that both the controller silicon and the installed optical transceiver support up to 100Gbps operation.Speed: 10000Mb/s: Identifies the fault. Despite being capable of 100Gbps, the link autonegotiated down to 10Gbps due to an autonegotiation mismatch or link training degradation with the top-of-rack switch.sudo ethtool -s eth1 speed 100000 duplex full autoneg off: Overrides autonegotiation and forces the local MAC/PHY to establish the link strictly at 100Gbps full duplex.sudo ethtool -p eth1 30: Enters visual identification mode, commanding the physical chassis LED on theeth1RJ45 or SFP cage to blink rapidly for 30 seconds, allowing data center personnel to pinpoint the physical cable.
What the Admin Does Next
If manually setting 100Gbps causes the link to drop completely (Link detected: no), the problem is physical: an improperly seated optic, an incompatible Forward Error Correction (FEC) mode (RS versus BaseR), or dirty optical end-faces.
Case 5: Inspecting Optical Transceiver Telemetry and SFP+ DDM/DOM Health
Operational Scenario
An edge hypervisor connected over a 10km single-mode fiber span suffers intermittent frame discards and CRC errors. You must evaluate the optical module's physical health metrics without requiring an on-site optical power meter.
Diagnostic Invocation
sudo ethtool -m eth0
Raw Command Output
Identifier : 0x03 (SFP)
Extended identifier : 0x04 (GBIC/SFP defined AC-coupled)
Connector : 0x07 (LC)
Transceiver codes : 0x20 0x00 0x00 0x00 0x00 0x00 0x00 0x00
Transceiver type : 10G Ethernet: 10G Base-LR
Encoding : 0x06 (64B/66B)
BR, Nominal : 10300MBd
Rate identifier : 0x00 (unspecified)
Link length supported for 9/125um fiber : 10km
Laser wavelength : 1310nm
Module Temperature : 41.24 degrees C / 106.23 degrees F
Module Voltage : 3.2840 V
Laser Bias Current : 32.140 mA
Laser Output Power : 0.7240 mW / -1.40 dBm
Receiver Signal Average Optical Power : 0.0381 mW / -14.19 dBm
Module temperature alarm high flag : Off
Module temperature warning high flag : Off
Laser bias current alarm high flag : Off
Laser output power alarm low flag : Off
Rx power alarm low flag : On
Rx power warning low flag : On
Line-by-Line Telemetry Analysis
Transceiver type: 10G Base-LR/Laser wavelength: 1310nm: Confirms the physical medium is a long-range 1310nm single-mode optical module designed for up to 10 kilometers.Laser Output Power: 0.7240 mW / -1.40 dBm: The local transmitting laser is operating normally, pushing -1.40 dBm into the transmit fiber (standard 10G Base-LR launch power is between -8.2 dBm and +0.5 dBm).Receiver Signal Average Optical Power: 0.0381 mW / -14.19 dBm: The received optical power has dropped to -14.19 dBm.Rx power alarm low flag: On: The transceiver's internal microcontroller has tripped an alarm threshold. While the receiver sensitivity limit is typically -14.4 dBm, operating at -14.19 dBm leaves almost no link margin, leading to intermittent bit errors, CRC check failures, and dropped frames.
Launch Power: -1.40 dBm (Healthy)"] -->|"Optical Fiber Run (12.79 dB Attenuation)"| B["Receiving Node (Remote Transceiver)
Received Power: -14.19 dBm (Critical: Alarm Low ON)"]
What the Admin Does Next
The telemetry confirms that the local host's driver and transmit laser are functioning properly, but the incoming optical signal is severely degraded. The engineer should: 1. Dispatch data center technicians to clean and inspect the LC fiber connectors using an optical scope. 2. Inspect the fiber run for tight bend radiuses or damaged patch cables. 3. Check the remote switch optic to ensure its transmission power has not degraded.
5. What Can Go Wrong: Pitfalls, Latency Spikes, and Recovery
While ethtool is invaluable for production diagnostics, misapplying hardware settings can trigger unexpected outages.
| Pitfall Action | Immediate Consequence | Remediation / Prevention |
|---|---|---|
Resizing Ring Descriptors (-G) |
Link reset, momentary frame drop, link flap to switch. | Apply changes in maintenance windows or across bonded interfaces. |
| Excessive Ring Sizing (e.g. 8192) | Bufferbloat, CPU cache thrashing, increased tail latency. | Balance ring capacity against strict application latency budgets. |
| Disabling GRO/TSO at 40GbE+ Speeds | CPU core saturation at high throughput (>10Gbps). | Only disable for active triage or if driver firmware is proven buggy. |
1. Link Flaps During Ring Buffer Mutation
When you execute ethtool -G <interface> rx <size>, most network drivers must reinitialize the hardware ring descriptors. To complete this operation, the driver stops the interface, drains existing queues, reallocates memory, and restarts the MAC engine.
2. The Bufferbloat and L3 Cache Penalty of Oversized Queues
It is tempting to expand ring buffers to their absolute maximum (such as 4,096 or 8,192 descriptors) across all hosts. However, oversized queues introduce two performance penalties: * Bufferbloat: During downstream congestion, an excessively large descriptor ring holds thousands of packets in memory rather than dropping them early. This prevents TCP congestion algorithms (like CUBIC or BBR) from recognizing bottlenecks, increasing round-trip latencies from sub-milliseconds to hundreds of milliseconds. * Cache Thrashing: Very large queues continuously evict application data from CPU L3 caches to accommodate incoming DMA buffers, degrading performance for latency-sensitive applications.
3. CPU Saturation from Indiscriminate Offload Disabling
Disabling hardware offloads like TSO, GSO, and GRO is a great diagnostic technique to isolate driver bugs. However, operating a 25Gbps, 40Gbps, or 100Gbps interface without hardware segmentation forces the host CPU to process every 1,500-byte frame individually.
A single 100Gbps interface running at line rate without GRO/TSO can flood a host with up to 8.1 million packets per second, pinning multiple CPU cores entirely on software interrupt handling (ksoftirqd).
6. Persisting ethtool Parameters Across Reboots
Adjustments made with ethtool update running hardware registers immediately, but they do not persist across reboots. To ensure your tuning survives a system restart, use your distribution's native declarative configuration framework.
Method A: Declarative Configuration via systemd-networkd .link Files
On modern systemd-based Linux distributions, configure persistent hardware properties via a declarative .link file in /etc/systemd/network/.
Create /etc/systemd/network/10-eth0-hw-tuning.link:
[Match]
OriginalName=eth0
[Link]
Description=Hardware tuning for eth0 high-throughput interface
BitsPerSecond=25G
Duplex=full
AutoNegotiation=yes
# Hardware Ring Buffer Configuration
RxBufferSize=4096
TxBufferSize=4096
# Offload Engines
GenericReceiveOffload=yes
GenericSegmentationOffload=yes
TCPSegmentationOffload=yes
LargeReceiveOffload=no
For a comprehensive list of configuration directives, see the systemd.link man page.
Method B: Persistent Udev Rules for Specific MAC Addresses
For standalone systems or environments without systemd-networkd, apply rules the moment the kernel identifies the hardware using udev.
Create /etc/udev/rules.d/99-network-tuning.rules:
ACTION=="add", SUBSYSTEM=="net", ATTR{address}=="00:1b:21:bb:cc:dd", RUN+="/usr/sbin/ethtool -G %k rx 4096 tx 4096", RUN+="/usr/sbin/ethtool -K %k tso on gso on gro on"
Method C: NetworkManager Dispatcher Scripts
On enterprise distributions managed by NetworkManager (such as RHEL, Rocky Linux, or AlmaLinux), apply parameters using a dispatcher script in /etc/NetworkManager/dispatcher.d/.
Create /etc/NetworkManager/dispatcher.d/90-ethtool-tuning.sh:
#!/bin/bash
INTERFACE=$1
ACTION=$2
if [ "$INTERFACE" = "eth0" ] && [ "$ACTION" = "up" ]; then
/usr/sbin/ethtool -G eth0 rx 4096 tx 4096
/usr/sbin/ethtool -K eth0 tso on gro on
fi
Set execution permissions on the script:
sudo chmod +x /etc/NetworkManager/dispatcher.d/90-ethtool-tuning.sh
7. Comparative Reference Matrix
When troubleshooting network performance anomalies, use the matrix below to match symptoms with the correct diagnostic flags and counters:
| Symptom / Anomaly | Suspected Subsystem | Primary ethtool Flag | Critical Metrics / Parameters to Evaluate |
|---|---|---|---|
| Application packet loss with zero socket errors | DMA Ring Depletion | ethtool -S <iface> |
rx_missed_errors, rx_no_buffer_count, rx_fifo_errors |
| Intermittent throughput collapse during I/O bursts | Hardware Ring Sizing | ethtool -g <iface> |
Pre-set maximums vs. Current hardware settings |
| Corrupted payloads, MTU blackholes, overlay faults | Silicon Acceleration | ethtool -k <iface> |
tcp-segmentation-offload, generic-receive-offload, rx-checksumming |
| Link operating at lower speeds than rated hardware | PHY Autoneg / FEC | ethtool <iface> |
Speed:, Duplex:, Supported link modes:, Advertised link modes: |
| Intermittent CRC/FCS errors on optical fiber | SFP+/QSFP Optic | ethtool -m <iface> |
Receiver Signal Average Optical Power, Laser Bias Current, Rx power alarm low flag |
For further technical reading on Linux kernel networking architecture and tuning strategies, refer to the Linux Kernel Networking Scaling Documentation, the Ethtool Official Linux Manual, the Arch Linux Network Configuration Guide, and the Red Hat Enterprise Linux Network Performance Tuning Guide.
8. Today's Takeaway
The true health of a Linux network interface cannot be assessed from high-level operating system socket tools alone. Open a terminal right now and execute sudo ethtool -S <primary_interface> | grep -Ei "(drop|miss|error|fail|loss)" on your most critical production node. Review the output for any incrementing hardware-level counters, such as rx_missed_errors or rx_no_buffer_count. If those registers are greater than zero, your system is silently dropping traffic in hardware before the Linux kernel ever has a chance to process it. Pair that check with sudo ethtool -g <primary_interface> to verify whether your ring buffers are still operating at small factory defaults, and optimize them to keep pace with your infrastructure's real-world traffic.