Debugging packet loss without tcpdump
Reaching for tcpdump first is habit, not strategy. A capture answers
"what was on the wire", but most loss reports are really one of three questions:
is the kernel dropping, is the path dropping, or is the receiver overwhelmed. Those
are all visible in counters you already have.
Start with the socket
ss -tin on the listening or established socket gives retransmit and
congestion state per connection. Look at retrans, rtt and the
congestion window before anything else:
ss -tin state established '( dport = :443 or sport = :443 )'
A single connection with a high retransmit count points at the path. Every connection showing retransmits at once points at the host — usually a saturated queue or an interface dropping under load.
Then the interface
ip -s link show separates receive-side drops from transmit-side ones.
Errors on RX with clean TX is a driver or ring-buffer problem; drops on TX is
almost always queueing.
ip -s link show dev eth0
nstat -az | grep -E 'TcpRetrans|TcpExtTCPLostRetransmit|UdpInErrors'
The three counters that explain most cases
- TcpRetransSegs — normal at a fraction of a percent; a sustained percentage in the high single digits is a real path problem.
- listen queue overflow — a burst of SYNs that outran
accept(). Usually an application problem wearing a network costume. - receive buffer errors — the NIC ran out of descriptors or the socket buffer was too small for the bandwidth-delay product.
When a capture is actually worth it
Take the capture when the counters disagree with each other, or when you need to prove which hop drops. At that point, filter narrowly — a full capture of a busy interface mostly produces a large file and a slow conclusion:
tcpdump -ni eth0 -s 96 -c 2000 'host 203.0.113.10 and tcp[tcpflags] & tcp-syn != 0'
The habit worth building is to read the counters first. They are free, they are always on, and they are usually enough.