Wire Level libpcap Tests for Trading Desk Keypad Latency
Share
Yes: measure the time from the moment a trader’s finger actuates the keypad button to the moment the venue sends back an execution acknowledgment, and capture both ends at the wire level with hardware timestamps rather than application logs. The biggest accuracy gains come from fixing your clock source and eliminating software timestamping jitter. The biggest latency gains come from addressing device firmware, the OS input path, the network relay, and the exchange itself, roughly in that order of what you control.
TL;DR:
- Wire-level capture at the gateway or NIC, with hardware timestamping, provides the most accurate measurement of keypad-to-exchange latency compared to application logs.
- Clocks must be synchronized to UTC(NIST) traceability, with PTP or GNSS clocks achieving microsecond to nanosecond precision for HFT-grade testing.
- The 99th and 99.9th percentile latencies reveal tail events that can significantly impact trading costs, unlike average latency which hides worst-case delays.
- Fixing device firmware and USB polling, OS-level tuning, and removing unnecessary relays typically yields the biggest improvements before considering colocation.
- Collecting artifacts such as raw pcap files and detailed logs is essential for compliance, forensic review, and accurate reporting of latency performance.
Table of Contents
- What to measure: start and end events for keypad-to-exchange latency
- Measurement methods and tooling: wire-level capture vs application logs
- Clock synchronization and expected accuracy: NTP, PTP, and UTC(NIST) traceability
- Step-by-step test procedure: from button press to percentile histogram
- Interpreting results and optimization priorities
- Validation, logging, and reporting for trading desks and compliance
- Vendor perspective: how keypad design choices affect measurable latency
- Key-Trade Professional Trading Keyboard: built for the measurement window above
- FAQ
- Sources
What to measure: start and end events for keypad-to-exchange latency
A latency number only means something when everyone agrees on where the clock starts and stops. The start event can be the hardware actuation signal inside the keypad, the moment the HID/USB transmit leaves the device, or the moment your trading application generates a send event. Each has tradeoffs: actuation signals are the truest start but need a GPIO tap or oscilloscope, while an application send event is easy to log but already carries OS scheduling delay baked in.
The end event is usually the venue’s ExecutionReport or order acknowledgment, not the raw exchange timestamp on the fill, because the ACK is what tells the desk the order was actually received and processed. Turnaround time is then a simple subtraction: Order-Processed Timestamp minus Order-Initiated Timestamp. Keeping that formula fixed across every test run is what makes results comparable across keypads, platforms, and network paths.
Measurement methods and tooling: wire-level capture vs application logs
Wire-level packet capture is the standard for anyone serious about this measurement. Tools built on libpcap, including tcpdump and Wireshark, let you tap traffic at the trading gateway or directly on the workstation NIC, capturing both the outbound order packet and the inbound acknowledgment without touching the application’s own clock. Correlating those captures with FIX message sequencing, as practitioners recommend, is what lets you reconstruct true end-to-end timing rather than relying on what the application believes happened.
Application logs are easier to generate but mislead in one specific way: operating system scheduling and buffering routinely add tens to over a hundred microseconds of jitter to a software timestamp, so two runs with identical network paths can show very different numbers for reasons that have nothing to do with the order itself. They are acceptable for rough sanity checks, not for benchmarking or compliance reporting.
For desks pushing toward microsecond-level confidence, hardware timestamping on a PTP-capable NIC removes the OS from the measurement path entirely.
A practical capture setup includes:
- A mirrored port or tap at the gateway, not just the workstation, to see real wire timing.
- Hardware or PTP-capable NICs where sub-microsecond comparisons matter.
- Synchronized clocks across every capture point, with offsets logged, not assumed.
- FIX SendingTime (tag 52) correlation rather than a single timestamp field.
Pro Tip: Timestamp the moment your device driver hands the packet to the network stack, then correlate that against the NIC-level capture. That isolates input-path delay from everything that happens after the packet leaves the machine.
Clock synchronization and expected accuracy: NTP, PTP, and UTC(NIST) traceability

None of this matters if your clocks disagree with each other. NIST guidance on time synchronization for financial markets describes UTC(NIST) traceability as the reference chain exchanges and trading systems use to keep timestamps meaningful across different firms and systems, and notes that exchanges commonly hold synchronization to within tens of microseconds in practice. If your test rig’s clock cannot be tied back to that same reference, your latency numbers are not comparable to anyone else’s.
Accuracy expectations scale with what you are trying to prove:
- NTP typically holds offsets in the low microsecond to low millisecond range, adequate for manual or semi-automated desk monitoring.
- PTP or GNSS-disciplined clocks with hardware timestamping reach microsecond to nanosecond accuracy, the tier needed for automated systems and anything approaching HFT-grade claims.
- Well-configured NTP on a LAN can show server offsets around 20 microseconds with round-trip delays near 300 microseconds in NIST timing examples, a useful sanity baseline before you invest in PTP.
Monitor clock offset continuously rather than once at setup, since drift accumulates and a capture point that was accurate last week may not be accurate today.
Step-by-step test procedure: from button press to percentile histogram
A reproducible test run follows the same five stages every time, which is what lets you trust a comparison between two keypads or two network configurations.
- Prep: isolate a test network segment, choose your capture point at the gateway or NIC, and record the clock offset for every device in the chain before you press a single button.
- Instrument: timestamp the exact key actuation, either through a GPIO trigger on the keypad or an application-level tap, and start your libpcap capture on both the outbound order path and the inbound ExecutionReport path.
- Correlate: match each outbound packet to its inbound acknowledgment using FIX sequence IDs and SendingTime (tag 52) rather than TransactTime (tag 60), which reflects the venue’s clock, not yours.
- Execute: run single-press trials first to catch obvious faults, then steady bursts, then a sample of at least 1,000 messages, since percentile statistics below that size are noisy.
- Analyze: build a percentile histogram at 50th, 90th, 99th, and 99.9th percentiles, then compute turnaround time across the full sample.
A libpcap-based matching-engine experiment correlating outbound order packets with inbound acknowledgments produced a latency histogram running from a 23-microsecond minimum to 538 microseconds at the 99.99th percentile, showing how far tail behavior can diverge from the median even in an optimized system. That gap between median and tail is exactly what a keypad-to-exchange test is built to expose.
Interpreting results and optimization priorities
Mean latency flatters almost every setup. What actually costs money on a trading desk is the tail: the 99th and 99.9th percentile events, because those are the fills that arrive late enough to matter, often during volatile moves when speed matters most. A desk that only reports an average is hiding its worst outcomes.
Captures usually point to one of four sources: the keypad device itself (firmware send timing, USB polling interval, debounce logic), the OS input path (driver buffering, scheduler delay), the network (an unnecessary relay or VPN hop), or the exchange side (queue depth at the matching engine). Each leaves a distinct fingerprint in the percentile shape, a device issue tends to show up as consistent added latency across every percentile, while a network issue often shows as a fat tail with a clean median.
Fix priorities generally run:
- Firmware and USB polling adjustments on the keypad, the cheapest fix with the most direct effect on input-path delay.
- OS-level tuning, such as disabling unnecessary background processes competing for the scheduler.
- Removing cold-start bridges or unnecessary relays between the keypad and the trading platform.
- NIC hardware timestamping and, for the most latency-sensitive desks, colocation near the exchange.
Pro Tip: Escalate to infrastructure or colocation spending only after device and OS fixes stop moving the 99th percentile. Spending on colocation before fixing a slow keypad driver wastes the investment.
Validation, logging, and reporting for trading desks and compliance
Every test run should leave behind an artifact trail, not just a headline number. Keep the raw pcap extracts, a timestamp correlation table mapping outbound orders to inbound acknowledgments, the percentile histogram itself, and a log of clock offsets at the time of capture.
Retention matters for two reasons: forensic review after an execution dispute, and traceability back to UTC(NIST) if a regulator or counterparty asks how your timestamps were generated. FINRA’s own guidance notes that the vast majority of held market orders in NMS stocks execute within 500 milliseconds, a useful external benchmark when a desk lead wants to know whether an internal latency number is reasonable or an outlier.
A short report template for desk leads should include:
- Test date, clock offsets, and capture point used.
- Sample size and percentile breakdown (50th, 90th, 99th, 99.9th).
- Identified bottleneck and the fix applied, if any.
Vendor perspective: how keypad design choices affect measurable latency
Specialized trading keypads have design choices that move latency numbers, which are rarely visible from a spec sheet: USB polling interval, debounce timing, and how quickly firmware hands a packet to the host. Before buying, ask about polling rate, debounce implementation, and any capture-based test data.
— key-trade
Key-Trade Professional Trading Keyboard: built for the measurement window above
A dedicated keypad shortens the input path this whole guide is built around, which is why the Key-Trade Professional Trading Keyboard integrates with TradingView, MetaTrader, cTrader, NinjaTrader, and other platforms without a software relay in between.

Check specs or request test data on the Key-Trade trading keyboard product page.
FAQ
What is the correct measurement window for keypad latency?
The window runs from the physical button actuation on the keypad to the venue’s execution acknowledgment, not the raw exchange fill timestamp. Using the ACK as the end event, matched to your local SendingTime rather than the venue’s TransactTime, keeps comparisons across setups consistent.
Why use wire-level capture instead of application timestamps?
Wire-level capture with tools built on libpcap avoids the operating system jitter that application logs inherit from scheduling and buffering delays. Hardware or PTP-capable NIC timestamping removes that jitter entirely when sub-microsecond accuracy is the goal.
How accurate do my clocks need to be?
Accuracy needs depend on use case: NTP typically suits manual or semi-automated monitoring, while PTP or GNSS-disciplined clocks with hardware timestamping are needed for automated or HFT-grade testing. NIST examples show well-configured NTP holding offsets around 20 microseconds on a LAN, a reasonable baseline before considering PTP.
What sample size do I need for reliable percentile statistics?
A run of at least 1,000 messages gives stable 99th and 99.9th percentile readings, since smaller samples produce noisy tail estimates. Running steady bursts after initial single-press trials catches both one-off faults and sustained jitter.
How does keypad latency compare to typical execution benchmarks?
FINRA statistics show that the vast majority of held market orders in NMS stocks execute within 500 milliseconds, a far larger timescale than the microsecond figures seen in matching-engine benchmarks. Desk-level keypad tests should be judged against the 500-millisecond retail benchmark unless the setup specifically targets matching-engine-grade latency.