"Connection reset by peer."
Three words that have ended many a trading system's weekend. At 3:47 AM on a Saturday, while the engineer who deployed the system sleeps, the WebSocket connection silently dies. The system never recovers. When markets open Monday morning, there's no data. There's no alert. There's just silence where there should be ticks.
This is not a hypothetical. This is a recurring incident across financial data infrastructure, and it almost always traces back to the same root cause: a WebSocket implementation without proper heartbeat handling.
The Problem No One Talks About Until It's Too Late
WebSocket connections are not permanent. They are TCP connections that appear permanent, but under the hood, they are subject to forces that silently terminate them:
Network infrastructure timeouts. Load balancers, NAT gateways, and firewalls all maintain connection tables. When a connection goes idle, these devices drop the mapping. Common timeout values range from 30 seconds to 5 minutes. The connection appears alive to both endpoints, but the network path has been severed.
Cloud provider idle limits. AWS Application Load Balancer drops connections after 60 seconds of inactivity by default. Google Cloud Load Balancer uses 10 minutes. Azure uses 380 seconds. If your WebSocket client hasn't sent anything in that window, the connection is gone — and neither endpoint is notified.
Peer process restarts. The server restarts. The connection terminates. The client doesn't know.
NAT table expiration. Carrier-grade NAT on mobile networks or shared WiFi can reassign the external IP mapping to another user, orphaning your connection.
Without heartbeat mechanisms, your application cannot distinguish between "no data because the market is closed" and "no data because the connection died." You are flying blind.
RFC 6455 and the ping/pong Protocol
The WebSocket specification (RFC 6455) defines a built-in heartbeat mechanism using control frames:
| Frame Type | Opcode | Direction | Purpose |
|---|---|---|---|
ping |
0x9 | Server → Client | "Are you alive?" |
pong |
0xA | Client → Server | "Yes, I'm alive" |
pong |
0xA | Server → Client | Server heartbeat response |
The specification is elegant: the server sends a ping frame, and the client responds with a pong. If no pong arrives within a reasonable timeout, the connection is dead. This is a clean, standards-compliant approach that requires zero application-level code.
However, RFC 6455 places one critical constraint on clients: "A Pong frame MAY be sent unsolicited. This serves as a unidirectional heartbeat." The word "MAY" means servers are not obligated to send ping frames. Many market data vendors — including some major ones — do not implement server-side ping. They leave it to clients to implement their own heartbeat on top of the application layer.
This is where engineering complexity explodes.
The Hidden Cost: When Heartbeat Becomes Your Problem
Let's examine what happens when a WebSocket client must implement its own heartbeat.
The Naive Approach (What Most People Write First)
import websocket
import threading
import time
class NaiveWebSocketClient:
def __init__(self, url):
self.url = url
self.ws = None
self.running = False
self.last_message_time = time.time()
self.heartbeat_interval = 30 # seconds
def on_message(self, ws, message):
self.last_message_time = time.time()
# Process message...
def on_ping(self, ws, data):
# websocket-client library handles this automatically
# But if server doesn't send pings, this never fires
pass
def start(self):
self.ws = websocket.WebSocketApp(
self.url,
on_message=self.on_message,
on_ping=self.on_ping
)
self.running = True
self.ws.run_forever()
This implementation has a fatal flaw: it relies on the server sending ping frames. If the server doesn't, last_message_time only updates when actual market data arrives. During quiet periods (pre-market, weekends, thin trading), the connection can die silently, and the client never knows.
The Application-Layer Heartbeat Approach (What Polygon Requires)
When the server doesn't send ping frames, clients must send their own application-layer heartbeat — typically by emitting a special JSON message that the server acknowledges.
import websocket
import threading
import time
import json
class PolygonHeartbeatClient:
"""
Polygon requires application-layer heartbeat.
Client must send a ping message every 20 seconds,
and the server responds with a pong.
"""
def __init__(self, url, api_key):
self.url = url
self.api_key = api_key
self.ws = None
self.running = False
self.last_pong_time = time.time()
self.ping_interval = 20 # Polygon requires pings every 20 seconds
self.pong_timeout = 30 # Disconnect if no pong within 30 seconds
self.heartbeat_thread = None
def _heartbeat_loop(self):
"""Dedicated thread for sending heartbeats."""
while self.running:
if self.ws and self.ws.sock and self.ws.sock.connected:
try:
# Application-layer ping — not RFC 6455 ping
# This is a JSON message, not a WebSocket control frame
ping_message = json.dumps({"action": "ping"})
self.ws.send(ping_message)
# Wait for pong with timeout
time.sleep(self.pong_timeout)
if time.time() - self.last_pong_time > self.pong_timeout:
print(f"[HEARTBEAT] No pong received for {self.pong_timeout}s. Reconnecting...")
self._reconnect()
except Exception as e:
print(f"[HEARTBEAT] Send failed: {e}")
self._reconnect()
else:
time.sleep(1)
def _reconnect(self):
"""Reconnection with exponential backoff."""
self.running = False
if self.heartbeat_thread:
self.heartbeat_thread.join(timeout=5)
delay = 1
max_delay = 60
while True:
try:
print(f"[RECONNECT] Attempting connection in {delay}s...")
self.running = True
self.heartbeat_thread = threading.Thread(target=self._heartbeat_loop)
self.heartbeat_thread.daemon = True
self.heartbeat_thread.start()
self.ws = websocket.WebSocketApp(
self.url,
on_message=self._on_message,
on_pong=self._on_pong,
on_close=self._on_close
)
self.ws.run_forever(ping_interval=None) # We handle ping ourselves
break
except Exception as e:
print(f"[RECONNECT] Failed: {e}")
time.sleep(delay)
delay = min(delay * 2, max_delay)
def _on_pong(self, ws, payload):
self.last_pong_time = time.time()
def _on_message(self, ws, message):
try:
data = json.loads(message)
if data.get("action") == "pong":
self.last_pong_time = time.time()
else:
# Process market data...
pass
except json.JSONDecodeError:
pass
def _on_close(self, ws, close_status_code, close_msg):
print(f"[CLOSE] Connection closed: {close_status_code} {close_msg}")
self._reconnect()
def start(self):
self._reconnect()
What This Code Reveals
Look at the engineering surface area of this application-layer heartbeat implementation:
| Concern | Complexity | Risk if Mishandled |
|---|---|---|
| Separate heartbeat thread | Must be managed, started, stopped, joined | Thread leak, zombie threads |
| Application-layer ping message format | Must match server's expected format | Silent heartbeat failure |
| Pong timeout tracking | Must maintain state, check on interval | False positives, spurious reconnects |
| Reconnection coordination | Heartbeat thread and WebSocket thread must coordinate | Race conditions on reconnect |
| Ping interval synchronization | Must match server's expected interval | Server may disconnect client |
| Thread safety | last_pong_time shared between threads |
Data races, incorrect timeout detection |
This is not a trivial amount of code. This is production infrastructure that requires careful testing, monitoring, and ongoing maintenance. Every bug in this code translates to missed data at the worst possible moment.
TickDB's Native Approach: What RFC 6455 Actually Provides
TickDB implements server-side ping frames as specified in RFC 6455. The WebSocket protocol handles heartbeats at the transport layer, without any application-layer involvement.
import os
import websocket
import time
import threading
class TickDBWebSocketClient:
"""
TickDB WebSocket client using native RFC 6455 ping/pong.
Key advantages:
1. No application-layer heartbeat thread required
2. WebSocket library handles ping/pong automatically
3. Connection state is managed at the protocol layer
4. Cleaner code, fewer failure modes
"""
def __init__(self, url):
# ⚠️ For production HFT workloads, use aiohttp/asyncio for non-blocking I/O
self.url = url
self.ws = None
self.running = False
self.reconnect_delay = 1
self.max_reconnect_delay = 60
self.last_ping_time = time.time()
self.ping_timeout = 60 # If no activity for 60s, connection is suspect
def _on_ping(self, ws, payload):
"""RFC 6455 ping received — websocket-client auto-responds with pong."""
self.last_ping_time = time.time()
print(f"[PING] Server heartbeat received at {self.last_ping_time}")
def _on_pong(self, ws, payload):
"""Our outbound ping got a response."""
self.last_ping_time = time.time()
print(f"[PONG] Connection confirmed alive")
def _on_message(self, ws, message):
"""Process incoming market data."""
self.last_ping_time = time.time()
# Parse and process message...
pass
def _on_close(self, ws, close_status_code, close_msg):
"""Connection closed — trigger reconnect."""
print(f"[CLOSE] Status {close_status_code}: {close_msg}")
self.running = False
self._reconnect_with_backoff()
def _on_error(self, ws, error):
"""Log errors for monitoring."""
print(f"[ERROR] {error}")
def _reconnect_with_backoff(self):
"""Exponential backoff with jitter to prevent thundering herd."""
delay = self.reconnect_delay
while not self.running:
print(f"[RECONNECT] Next attempt in {delay:.1f}s")
time.sleep(delay)
try:
api_key = os.environ.get("TICKDB_API_KEY")
if not api_key:
raise ValueError("TICKDB_API_KEY not set")
# WebSocket auth via URL parameter (not header)
ws_url = f"{self.url}?api_key={api_key}"
self.ws = websocket.WebSocketApp(
ws_url,
on_message=self._on_message,
on_ping=self._on_ping,
on_pong=self._on_pong,
on_close=self._on_close,
on_error=self._on_error
)
# run_forever with ping_interval enables automatic RFC 6455 ping/pong
# TickDB server sends pings; library handles pong responses
self.running = True
self.ws.run_forever(
ping_interval=30, # Send our own ping every 30s (belt and suspenders)
ping_timeout=10, # Wait 10s for pong response
ping_payload="heartbeat"
)
# If we exit run_forever, connection was lost — reset delay
self.reconnect_delay = 1
except Exception as e:
print(f"[RECONNECT] Failed: {e}")
# Exponential backoff with jitter
self.reconnect_delay = min(self.reconnect_delay * 2, self.max_reconnect_delay)
import random
jitter = random.uniform(0, self.reconnect_delay * 0.1)
self.reconnect_delay += jitter
def start(self):
"""Start the WebSocket connection."""
self._reconnect_with_backoff()
# Usage
if __name__ == "__main__":
# Subscribe to depth channel for AAPL
url = "wss://api.tickdb.ai/ws/v1/depth"
client = TickDBWebSocketClient(url)
client.start()
The Difference in Practice
| Aspect | Polygon (Application-Layer) | TickDB (RFC 6455 Native) |
|---|---|---|
| Heartbeat mechanism | Custom JSON ping/pong | RFC 6455 control frames |
| Code complexity | 100+ lines for heartbeat alone | ~20 lines total reconnect logic |
| Failure modes | Thread bugs, format mismatches, race conditions | Standard WebSocket failures only |
| Server requirements | Client must ping every 20s | Server sends pings automatically |
| Protocol overhead | JSON serialization on every heartbeat | 2-byte control frame |
| Standard compliance | Proprietary extension | RFC 6455 compliant |
| Debugging | Is it the format? The interval? The thread? | Check ping_interval/ping_timeout params |
Production Considerations Beyond the Code
The Monitoring Gap
When you implement application-layer heartbeat, you must also implement monitoring for the heartbeat itself. What happens when your heartbeat thread crashes? What happens when your last_pong_time variable becomes corrupted? What happens when your ping message format drifts from what the server expects?
With native RFC 6455 ping/pong, the WebSocket library handles these concerns internally. You monitor the connection state; the library monitors liveness.
The Compliance Testing Burden
If you're building a system that interfaces with multiple WebSocket APIs (market data feeds, exchange connections, prime brokerage streams), each vendor may have different application-layer heartbeat requirements:
- Vendor A: JSON
{"type": "ping"}every 15 seconds - Vendor B: Custom binary format every 30 seconds
- Vendor C: No heartbeat required, but must send
{"action": "subscribe"}every 60 seconds or disconnect - Vendor D: RFC 6455 ping/pong
Every deviation is engineering surface area. Every deviation is a potential bug. Every deviation requires testing.
TickDB uses RFC 6455 natively. Your monitoring infrastructure, your testing harness, and your operational runbook need only understand one heartbeat mechanism.
The Latency Cost
Application-layer heartbeat messages consume bandwidth and processing cycles. A JSON {"action": "ping"} message is approximately 20 bytes, serialized and deserialized on both ends, every 20 seconds. Over a 24-hour period, that's 86,400 bytes of heartbeat traffic — negligible in isolation, but multiplied across thousands of connections and integrated into your data pipeline, it adds latency jitter and processing overhead.
RFC 6455 ping frames are 2 bytes of overhead, handled by the WebSocket library without application-layer involvement. No serialization. No deserialization. No processing latency.
Architecture Comparison
┌─────────────────────────────────────────────────────────────────────┐
│ APPLICATION-LAYER HEARTBEAT │
│ (Polygon Pattern) │
├─────────────────────────────────────────────────────────────────────┤
│ │
│ ┌──────────────┐ ┌──────────────┐ ┌──────────────┐ │
│ │ Market │ │ Heartbeat │ │ Reconnect │ │
│ │ Data │ │ Thread │ │ Logic │ │
│ │ Handler │ │ │ │ │ │
│ └──────┬───────┘ └──────┬───────┘ └──────┬───────┘ │
│ │ │ │ │
│ │ JSON ping │ JSON pong │ │
│ │◄────────────────────┼────────────────────► │
│ │ │ │
│ ┌──────▼─────────────────────▼───────────────────────────────┐ │
│ │ Application Layer │ │
│ │ (JSON serialization, thread safety) │ │
│ └─────────────────────────────────────────────────────────────┘ │
│ │
└─────────────────────────────────────────────────────────────────────┘
┌─────────────────────────────────────────────────────────────────────┐
│ NATIVE PING/PONG │
│ (TickDB Pattern) │
├─────────────────────────────────────────────────────────────────────┤
│ │
│ ┌──────────────┐ ┌──────────────┐ ┌──────────────┐ │
│ │ Market │ │ Ping │ │ Reconnect │ │
│ │ Data │ │ Handler │ │ Logic │ │
│ │ Handler │ │ (Library) │ │ │ │
│ └──────┬───────┘ └──────┬───────┘ └──────┬───────┘ │
│ │ │ │ │
│ │ ping (0x9) │ pong (0xA) │ │
│ │◄────────────────────┼────────────────────► │
│ │ │ │
│ ┌──────▼─────────────────────▼───────────────────────────────┐ │
│ │ WebSocket Library │ │
│ │ (Internal, no application code) │ │
│ └─────────────────────────────────────────────────────────────┘ │
│ │
└─────────────────────────────────────────────────────────────────────┘
The TickDB architecture is simpler, fewer lines of code to maintain, and leverages the WebSocket library's battle-tested implementation rather than reinventing heartbeat logic in application code.
When This Matters Most
The native ping/pong advantage is not theoretical. It becomes critical in specific production scenarios:
High-frequency trading systems. Every millisecond of unnecessary processing is lost alpha. Removing JSON serialization from the heartbeat path is a measurable improvement.
Systems with multiple WebSocket connections. Managing N heartbeat threads for N connections is complexity that scales poorly. Native ping/pong requires zero per-connection threads.
Weekend and overnight operations. Markets close, data stops flowing, and connections are most likely to be silently terminated by infrastructure. The heartbeat mechanism must work flawlessly during these quiet periods.
Systems deployed across cloud regions. Cross-AZ connections, VPN tunnels, and shared infrastructure all introduce variable network path reliability. Robust heartbeat handling is not optional.
Conclusion
The WebSocket heartbeat problem is deceptively simple. It seems like a solved problem — connections should stay alive, right? But in production financial systems, "seems like" is not good enough. Silent connection failures during market hours are not acceptable.
TickDB's native RFC 6455 ping/pong support eliminates an entire class of engineering problems:
- No application-layer heartbeat threads to manage
- No JSON ping/pong format to synchronize with the server
- No thread-safety bugs in heartbeat state tracking
- No race conditions between heartbeat and reconnection logic
- No custom monitoring for application-layer heartbeat failures
- Standards-compliant implementation tested by browser vendors and network equipment manufacturers
When you're debugging a production incident at 4 AM, the last thing you want is to wonder whether your heartbeat implementation is the culprit. TickDB's native heartbeat lets you focus on your strategy, not your infrastructure.
Next Steps
If you're evaluating WebSocket market data providers, ask each vendor two questions: "Do you send RFC 6455 ping frames?" and "What happens if I don't send application-layer heartbeats?" The answers will tell you how much engineering you'll be doing for the vendor's infrastructure.
If you want to connect to TickDB's WebSocket API:
- Sign up at tickdb.ai (free tier available, no credit card required)
- Generate an API key in the dashboard
- Set the
TICKDB_API_KEYenvironment variable - Copy the production-ready code from this article
If you're building a multi-vendor data infrastructure, consider the total cost of heartbeat management across all your feeds. The vendor that handles heartbeat natively may be cheaper to operate than the one that passes the complexity to you.
This article does not constitute investment advice. Markets involve risk; past performance does not guarantee future results.