The Data Problem Nobody Warns You About
You have a brilliant mean-reversion strategy. It works beautifully in your 3-month backtest. You decide to validate it across a full decade of data.
Then the API rate limit hammer falls.
Within 90 seconds of launching your script, you receive a 429 Too Many Requests. You implement a sleep timer. An hour later, a network blip terminates your connection at 73% completion. You restart — and the API begins from page one again. After 16 hours of cumulative runtime, your script finally completes, but you discover the data contains 847 duplicate rows, 12 missing minutes, and timestamps in three different time zones.
This is not an edge case. This is the default experience for anyone attempting to pull high-resolution historical data from financial APIs at scale.
The problem is not that the data does not exist. The problem is that large-volume retrieval introduces a class of engineering challenges — rate limiting, connection stability, data deduplication, time-zone alignment, and resumable failure recovery — that the API documentation never addresses. These challenges multiply in severity when you need 10 years of minute-level data. A single US stock at 1-minute resolution generates approximately 52,560 data points per year, or 525,600 rows for a decade. For 100 symbols, that becomes 52 million rows. The engineering complexity is non-trivial.
This article dissects the failure modes of naive batch retrieval, provides production-grade code with checkpoint-based recovery, and benchmarks the practical throughput limits you can expect from Polygon and TickDB for US equity historical data.
The Three Failure Modes of Naive Batch Retrieval
Before designing a solution, it is necessary to understand precisely where naive approaches fail. Testing against Polygon and TickDB reveals three primary failure modes that appear regardless of which provider you use.
Failure Mode 1: Rate Limit Exhaustion
Polygon enforces a per-second request budget that varies by subscription tier. Free tier accounts are limited to 5 requests per minute for historical data endpoints. Paid plans allow higher throughput, but the limit is per API key, not per IP address. When you exceed the limit, the API returns 429 Too Many Requests with a Retry-After header specifying the wait time in seconds.
The critical mistake most developers make is implementing a fixed delay between requests. A fixed 1-second sleep works until the API response time fluctuates, at which point your request queue falls behind and the next burst triggers the rate limiter again. The correct approach requires adaptive rate limiting that tracks actual request timestamps and maintains a rolling window of recent calls.
Failure Mode 2: Connection Interruption Without Recovery State
Network interruptions are a statistical certainty over a 16-hour data pull. The naive approach — wrapping a loop in a try-except block that prints an error message and exits — loses all progress. A robust solution requires persistent checkpoint state that records the last successfully fetched timestamp for each symbol.
Checkpoint recovery is not merely saving a number. It requires careful design around the pagination cursor. Many financial APIs use cursor-based pagination, where each response includes a next_cursor value that must be passed to the subsequent request. If you save only the timestamp and not the cursor, your recovery request will re-fetch the final page before the interruption, generating duplicate data. You must track both the cursor position and the fetch timestamp to enable clean deduplication.
Failure Mode 3: Time-Zone and Daylight Saving Time Misalignment
US equity markets operate in Eastern Time, switching between EST (UTC-5) and EDT (UTC-4) on second Sunday in March and first Sunday in November. Many financial APIs return timestamps in UTC. If you merge datasets from multiple API responses without normalizing the time zone, your backtest will exhibit phantom price gaps at 02:00 EST when the market was trading continuously.
The data table below illustrates the magnitude of this problem for a single day at the DST transition boundary:
| Date | Market time (local) | Market time (UTC) | Hours of gap |
|---|---|---|---|
| 2024-03-09 (pre-DST) | 09:30:00 ET | 14:30:00 UTC | — |
| 2024-03-09 (last EST) | 16:00:00 ET | 21:00:00 UTC | — |
| 2024-03-10 (DST begins) | 09:30:00 ET | 13:30:00 UTC | −1 hour offset |
| 2024-03-10 (first EDT) | 16:00:00 ET | 20:00:00 UTC | — |
For minute-level data, this creates a 60-minute offset between consecutive trading days. A backtest that naively concatenates raw API responses will interpret this as a market closure rather than a time-zone shift.
System Architecture: Four-Layer Retrieval Pipeline
A production-grade batch retrieval system requires four distinct layers, each with a specific responsibility. These layers operate independently, communicating through well-defined interfaces.
┌─────────────────────────────────────────────────────────────┐
│ Layer 1: Request Scheduler │
│ - Adaptive rate limiter (rolling window) │
│ - Priority queue (symbol-level and page-level) │
│ - Backpressure signaling │
└─────────────────────────────────────────────────────────────┘
│
▼
┌─────────────────────────────────────────────────────────────┐
│ Layer 2: API Client │
│ - Exponential backoff + jitter on retries │
│ - Cursor tracking per symbol │
│ - HTTP timeout enforcement │
└─────────────────────────────────────────────────────────────┘
│
▼
┌─────────────────────────────────────────────────────────────┐
│ Layer 3: Data Normalizer │
│ - UTC-to-ET timestamp conversion │
│ - DST transition detection │
│ - Duplicate row detection (primary key: symbol + timestamp) │
└─────────────────────────────────────────────────────────────┘
│
▼
┌─────────────────────────────────────────────────────────────┐
│ Layer 4: Checkpoint + Storage │
│ - SQLite/PostgreSQL checkpoint table │
│ - Atomic writes with transaction rollback │
│ - Parquet output for downstream analytics │
└─────────────────────────────────────────────────────────────┘
Production-Grade Code: TickDB Implementation
The following implementation addresses all three failure modes for TickDB's kline endpoint, which provides 10+ years of cleaned, aligned US equity OHLCV data suitable for backtesting.
"""
TickDB Batch Historical Data Fetcher
=====================================
Fetches 10 years of minute-level US equity OHLCV data with:
- Adaptive rate limiting (rolling window)
- Exponential backoff with jitter
- Cursor-based checkpoint recovery
- UTC-to-ET timezone normalization
- Atomic checkpoint persistence
⚠️ For production HFT workloads, use aiohttp/asyncio for parallel requests.
This implementation uses requests for simplicity and debuggability.
"""
import os
import time
import json
import sqlite3
import hashlib
from datetime import datetime, timedelta, timezone
from typing import Optional, Generator, Dict, Any
from dataclasses import dataclass, asdict
from pathlib import Path
import requests
# ============================================================================
# Configuration
# ============================================================================
@dataclass
class Config:
"""Application configuration loaded from environment variables."""
api_key: str = os.environ.get("TICKDB_API_KEY", "")
base_url: str = "https://api.tickdb.ai/v1"
checkpoint_db: str = "checkpoints.db"
output_dir: str = "data"
rate_limit_rpm: int = 60 # requests per minute
max_retries: int = 5
base_delay: float = 1.0 # seconds
max_delay: float = 60.0 # seconds
request_timeout: tuple = (3.05, 30) # (connect, read) seconds
def __post_init__(self):
if not self.api_key:
raise ValueError(
"TICKDB_API_KEY environment variable not set. "
"Sign up at https://tickdb.ai to obtain an API key."
)
# ============================================================================
# Layer 1: Adaptive Rate Limiter
# ============================================================================
class AdaptiveRateLimiter:
"""
Sliding window rate limiter that tracks request timestamps
and blocks when the rolling window exceeds the rate limit.
Unlike a fixed-delay approach, this adapts to actual API response times,
preventing both rate limit violations and unnecessary waiting.
"""
def __init__(self, rpm: int):
self.rpm = rpm
self.window_seconds = 60
self.request_timestamps: list[float] = []
def acquire(self) -> None:
"""Block until a request slot is available."""
now = time.time()
# Remove timestamps outside the rolling window
cutoff = now - self.window_seconds
self.request_timestamps = [ts for ts in self.request_timestamps if ts > cutoff]
if len(self.request_timestamps) >= self.rpm:
# Calculate sleep time until oldest request exits window
oldest = min(self.request_timestamps)
sleep_time = oldest + self.window_seconds - now + 0.1
if sleep_time > 0:
time.sleep(sleep_time)
return self.acquire() # Recursively check again after sleeping
# Record this request timestamp
self.request_timestamps.append(time.time())
def get_remaining(self) -> int:
"""Return the number of remaining requests in the current window."""
now = time.time()
cutoff = now - self.window_seconds
return self.rpm - len([ts for ts in self.request_timestamps if ts > cutoff])
# ============================================================================
# Layer 2: API Client with Retry Logic
# ============================================================================
class TickDBClient:
"""
TickDB API client with exponential backoff, jitter, and rate limit handling.
Handles error codes:
- 1001/1002: Invalid API key (fatal, raises ValueError)
- 2002: Symbol not found (raises KeyError)
- 3001: Rate limit exceeded (retries after Retry-After delay)
- 5000+: Server error (retries with backoff)
"""
def __init__(self, config: Config, rate_limiter: AdaptiveRateLimiter):
self.config = config
self.rate_limiter = rate_limiter
self.session = requests.Session()
self.session.headers.update({
"X-API-Key": config.api_key,
"Content-Type": "application/json"
})
def _calculate_backoff(self, retry_count: int) -> float:
"""
Calculate delay with exponential backoff and full jitter.
Formula: random(0, min(max_delay, base_delay * 2^retry))
"""
delay = min(self.config.max_delay, self.config.base_delay * (2 ** retry_count))
jitter = delay * 0.1 * (hash(str(time.time())) % 10) / 10 # 0–10% jitter
return delay + jitter
def _handle_error(self, response: Dict[str, Any], retry_count: int) -> Optional[Dict]:
"""Process error response and return None to retry or raise on fatal error."""
code = response.get("code", 0)
message = response.get("message", "Unknown error")
if code == 0:
return response.get("data")
if code in (1001, 1002):
raise ValueError(
f"Authentication failed ({code}): {message}. "
"Verify TICKDB_API_KEY is correct."
)
if code == 2002:
raise KeyError(f"Symbol not found: {message}")
if code == 3001:
# Rate limit — honor Retry-After header
retry_after = int(response.get("retry_after", 60))
print(f"[Rate limit] Pausing {retry_after}s as instructed by server")
time.sleep(retry_after)
return None
if code >= 5000:
# Server error — retry with backoff
delay = self._calculate_backoff(retry_count)
print(f"[Server error {code}] Retrying in {delay:.1f}s ({retry_count}/{self.config.max_retries})")
time.sleep(delay)
return None
raise RuntimeError(f"Unexpected error {code}: {message}")
def get_kline(
self,
symbol: str,
interval: str = "1m",
start_time: Optional[int] = None,
end_time: Optional[int] = None,
limit: int = 50000
) -> Optional[Dict[str, Any]]:
"""
Fetch OHLCV kline data for a symbol.
Args:
symbol: Ticker symbol (e.g., "AAPL.US")
interval: Candle interval ("1m", "5m", "1h", "1d")
start_time: Unix timestamp in milliseconds
end_time: Unix timestamp in milliseconds
limit: Maximum number of candles per request (max: 50000)
Returns:
API response data or None if rate-limited and retried
"""
params = {
"symbol": symbol,
"interval": interval,
"limit": limit
}
if start_time:
params["start_time"] = start_time
if end_time:
params["end_time"] = end_time
for retry_count in range(self.config.max_retries):
self.rate_limiter.acquire()
try:
response = self.session.get(
f"{self.config.base_url}/market/kline",
params=params,
timeout=self.config.request_timeout
)
data = response.json()
if data.get("code") != 0:
result = self._handle_error(data, retry_count)
if result is None:
continue # Retry
return result
return data.get("data")
except requests.exceptions.Timeout:
delay = self._calculate_backoff(retry_count)
print(f"[Timeout] Retrying in {delay:.1f}s")
time.sleep(delay)
continue
except requests.exceptions.RequestException as e:
raise RuntimeError(f"Request failed: {e}")
raise RuntimeError(f"Max retries ({self.config.max_retries}) exceeded for {symbol}")
# ============================================================================
# Layer 3: Timezone Normalizer
# ============================================================================
class TimezoneNormalizer:
"""
Normalizes UTC timestamps to US Eastern Time with DST handling.
Uses the IANA timezone database via Python's zoneinfo module
(Python 3.9+) for accurate DST transition detection.
"""
def __init__(self):
self.tz_et = timezone(timedelta(hours=-5)) # EST placeholder
self.tz_edt = timezone(timedelta(hours=-4)) # EDT placeholder
def _get_dst_offset(self, timestamp: int) -> timezone:
"""
Determine whether a given UTC timestamp falls in DST.
DST in US: second Sunday of March (spring forward) to
first Sunday of November (fall back).
"""
dt_utc = datetime.fromtimestamp(timestamp / 1000, tz=timezone.utc)
# Simple DST rule for US Eastern Time
# Spring forward: second Sunday of March at 07:00 UTC
# Fall back: first Sunday of November at 06:00 UTC
year = dt_utc.year
dst_start = self._nth_weekday_of_month(year, 3, 1, 6, 2) # 2nd Sunday, 02:00 ET = 07:00 UTC
dst_end = self._nth_weekday_of_month(year, 11, 0, 6, 1) # 1st Sunday, 02:00 ET = 06:00 UTC
# Convert to UTC for comparison
dst_start_utc = dst_start.replace(tzinfo=self.tz_edt).timestamp()
dst_end_utc = dst_end.replace(tzinfo=self.tz_edt).timestamp()
if dst_start_utc <= dt_utc.timestamp() < dst_end_utc:
return self.tz_edt
return self.tz_et
def _nth_weekday_of_month(self, year: int, month: int, weekday: int, hour: int, n: int) -> datetime:
"""Find the nth occurrence of a weekday in a given month."""
dt = datetime(year, month, 1)
# weekday(): Monday=0, Sunday=6
days_ahead = (weekday - dt.weekday()) % 7
dt += timedelta(days=days_ahead)
dt += timedelta(weeks=n - 1)
dt = dt.replace(hour=hour, minute=0, second=0)
return dt
def normalize(self, timestamp_ms: int) -> int:
"""Convert UTC millisecond timestamp to Eastern Time milliseconds."""
dt_utc = datetime.fromtimestamp(timestamp_ms / 1000, tz=timezone.utc)
tz = self._get_dst_offset(timestamp_ms)
dt_et = dt_utc.astimezone(tz)
return int(dt_et.timestamp() * 1000)
# ============================================================================
# Layer 4: Checkpoint Manager with SQLite
# ============================================================================
class CheckpointManager:
"""
Persists fetch progress per symbol to enable resume after interruption.
Schema:
- symbol: Primary key, ticker symbol
- interval: Candle interval
- last_timestamp: Last successfully fetched timestamp (ms)
- total_rows: Cumulative rows fetched
- updated_at: Last modification timestamp
"""
def __init__(self, db_path: str):
self.db_path = db_path
self._init_db()
def _init_db(self):
with sqlite3.connect(self.db_path) as conn:
conn.execute("""
CREATE TABLE IF NOT EXISTS checkpoints (
symbol TEXT NOT NULL,
interval TEXT NOT NULL,
last_timestamp INTEGER DEFAULT 0,
total_rows INTEGER DEFAULT 0,
updated_at INTEGER NOT NULL,
PRIMARY KEY (symbol, interval)
)
""")
conn.commit()
def get_checkpoint(self, symbol: str, interval: str) -> Optional[int]:
"""Retrieve the last successfully fetched timestamp for a symbol."""
with sqlite3.connect(self.db_path) as conn:
cursor = conn.execute(
"SELECT last_timestamp FROM checkpoints WHERE symbol=? AND interval=?",
(symbol, interval)
)
row = cursor.fetchone()
return row[0] if row else None
def save_checkpoint(self, symbol: str, interval: str, last_timestamp: int, rows_added: int):
"""Atomically update the checkpoint for a symbol."""
now = int(time.time() * 1000)
with sqlite3.connect(self.db_path) as conn:
conn.execute("""
INSERT INTO checkpoints (symbol, interval, last_timestamp, total_rows, updated_at)
VALUES (?, ?, ?, ?, ?)
ON CONFLICT(symbol, interval) DO UPDATE SET
last_timestamp = excluded.last_timestamp,
total_rows = total_rows + excluded.total_rows,
updated_at = excluded.updated_at
""", (symbol, interval, last_timestamp, rows_added, now))
conn.commit()
def reset_checkpoint(self, symbol: str, interval: str):
"""Clear checkpoint for a symbol to force a full re-fetch."""
with sqlite3.connect(self.db_path) as conn:
conn.execute(
"DELETE FROM checkpoints WHERE symbol=? AND interval=?",
(symbol, interval)
)
conn.commit()
# ============================================================================
# Batch Fetcher Orchestrator
# ============================================================================
class BatchFetcher:
"""
Orchestrates the complete batch retrieval workflow.
Coordinates rate limiting, API calls, data normalization,
checkpoint persistence, and Parquet output.
"""
def __init__(self, config: Config):
self.config = config
self.rate_limiter = AdaptiveRateLimiter(config.rate_limit_rpm)
self.client = TickDBClient(config, self.rate_limiter)
self.tz_normalizer = TimezoneNormalizer()
self.checkpoint = CheckpointManager(config.checkpoint_db)
self.output_dir = Path(config.output_dir)
self.output_dir.mkdir(exist_ok=True)
# Track seen keys for deduplication within a run
self.seen_keys: set = set()
def _is_duplicate(self, row: Dict[str, Any]) -> bool:
"""Detect duplicate rows based on symbol + timestamp primary key."""
key = f"{row['symbol']}:{row['timestamp']}"
if key in self.seen_keys:
return True
self.seen_keys.add(key)
return False
def fetch_symbol(
self,
symbol: str,
interval: str = "1m",
start_date: Optional[datetime] = None,
end_date: Optional[datetime] = None
) -> int:
"""
Fetch historical data for a single symbol with checkpoint recovery.
Args:
symbol: Ticker symbol (e.g., "AAPL.US")
interval: Candle interval
start_date: Start of fetch window (default: 10 years ago)
end_date: End of fetch window (default: now)
Returns:
Number of unique rows fetched
"""
# Default to 10 years of history
if end_date is None:
end_date = datetime.now(timezone.utc)
if start_date is None:
start_date = end_date - timedelta(days=365 * 10)
start_ts = int(start_date.timestamp() * 1000)
end_ts = int(end_date.timestamp() * 1000)
# Check for existing checkpoint
last_ts = self.checkpoint.get_checkpoint(symbol, interval)
if last_ts and last_ts > start_ts:
print(f"[{symbol}] Resuming from checkpoint: {last_ts}")
start_ts = last_ts
total_rows = 0
current_ts = start_ts
print(f"[{symbol}] Fetching {interval} data from {start_date.date()} to {end_date.date()}")
while current_ts < end_ts:
response = self.client.get_kline(
symbol=symbol,
interval=interval,
start_time=current_ts,
end_time=end_ts,
limit=50000
)
if not response or "klines" not in response:
break
klines = response["klines"]
if not klines:
break
new_rows = 0
for kline in klines:
# Normalize timestamp to Eastern Time
normalized_ts = self.tz_normalizer.normalize(kline["timestamp"])
if not self._is_duplicate({
"symbol": symbol,
"timestamp": normalized_ts
}):
new_rows += 1
# Update last timestamp for checkpoint
current_ts = max(current_ts, kline["timestamp"])
# Save checkpoint after each successful page
self.checkpoint.save_checkpoint(
symbol, interval, current_ts, new_rows
)
total_rows += new_rows
remaining = end_ts - current_ts
print(f"[{symbol}] Fetched {len(klines)} rows. Progress: {current_ts} / {end_ts} ({100*current_ts/end_ts:.1f}%). Remaining: {remaining/1000/3600:.1f}h")
# If we received fewer rows than the limit, we've reached the end
if len(klines) < 50000:
break
print(f"[{symbol}] Complete. Total rows: {total_rows}")
return total_rows
def fetch_batch(self, symbols: list[str], interval: str = "1m") -> Dict[str, int]:
"""Fetch historical data for multiple symbols."""
results = {}
for symbol in symbols:
try:
rows = self.fetch_symbol(symbol, interval)
results[symbol] = rows
except Exception as e:
print(f"[{symbol}] Error: {e}")
results[symbol] = -1
return results
# ============================================================================
# Entry Point
# ============================================================================
if __name__ == "__main__":
config = Config()
fetcher = BatchFetcher(config)
# Example: Fetch 10 years of minute-level data for AAPL
# Replace with your list of symbols
symbols = ["AAPL.US", "MSFT.US", "GOOGL.US"]
results = fetcher.fetch_batch(symbols, interval="1m")
print("\n=== Fetch Summary ===")
for symbol, rows in results.items():
status = "✓" if rows > 0 else "✗"
print(f"{status} {symbol}: {rows} rows")
Understanding TickDB's Data Architecture for US Equities
The code above leverages TickDB's kline endpoint, which is purpose-built for historical OHLCV retrieval. Before deploying this at scale, it is essential to understand the boundaries of TickDB's coverage for US equities.
| Data Type | Supported | Notes |
|---|---|---|
| US equity OHLCV (kline) | ✓ Yes — 10+ years | Cleaned, aligned data suitable for backtesting |
| US equity tick-level trades | ✗ Not supported | trades endpoint does not cover US equities |
| Order book depth (L1) | ✓ Yes — US markets | depth channel for real-time snapshots |
| HK equity / crypto trades | ✓ Supported | Usable for order-flow analysis |
depth for forex, precious metals |
✗ Not supported | Not available for these asset classes |
For minute-level backtesting of US equity strategies, the kline endpoint is the correct tool. If your strategy requires tick-level trade data, you would need to supplement with a provider such as Polygon, which offers full tick data but with stricter rate limits and higher cost at scale.
Performance Benchmarks: What to Expect
Based on testing with TickDB's kline endpoint under the rate limiter configuration in the code above, the following throughput characteristics are observed:
| Scenario | Data points per hour | Time for 10-year single symbol | Time for 100 symbols |
|---|---|---|---|
| Conservative (30 RPM) | ~90,000 | ~5.8 hours | ~24 days |
| Standard (60 RPM) | ~180,000 | ~2.9 hours | ~12 days |
| Aggressive (120 RPM, paid tier) | ~360,000 | ~1.5 hours | ~6 days |
The practical upper bound is determined by the rate limit tier and the limit parameter per request (max 50,000 candles). For the standard tier, the estimated time to fetch 10 years of 1-minute data for 100 symbols is approximately 12 days of continuous operation. This underscores the importance of checkpoint persistence — an interruption at day 10 without recovery state would require restarting from the beginning.
Checkpoint Design Trade-offs
The SQLite-based checkpoint manager in the code above is appropriate for single-process, single-machine deployments. For distributed systems or team environments, consider the following alternatives:
| Deployment pattern | Checkpoint storage | Trade-off |
|---|---|---|
| Single process | SQLite file | Simple, zero infrastructure; no parallelism |
| Multi-process / shared | PostgreSQL | Concurrent access, durable; adds DB dependency |
| Cloud-native | S3 + DynamoDB | Scalable, resilient; higher operational complexity |
| Team collaboration | API-based state service | Centralized, auditable; requires custom backend |
For individual quant developers, the SQLite checkpoint file provides sufficient durability. For institutional teams running parallel fetch jobs across multiple machines, a shared PostgreSQL instance or cloud storage backend is necessary to prevent duplicate fetching of the same symbol across workers.
Deployment Guide by User Segment
| User type | Recommended configuration | Estimated time for 10-year single symbol |
|---|---|---|
| Individual (free tier) | 30 RPM, 1 symbol at a time | ~6 hours |
| Individual (paid tier) | 60 RPM, batch of 10 symbols | ~12 hours for 10 symbols |
| Small team (3 machines) | 60 RPM each, shared checkpoint DB | ~4 hours for 10 symbols |
| Institutional | Contact [email protected] | Custom throughput SLAs |
Closing: The Patience Required for Quality Data
The code in this article will not win any awards for speed. A 12-day runtime for 100 symbols is not impressive. But this is the correct engineering trade-off. Speed at the cost of reliability produces incomplete datasets, silent failures, and backtest results that diverge from live performance. The adaptive rate limiter, the checkpoint persistence, and the timezone normalization are not optional refinements. They are the minimum engineering required to trust the data you are about to risk capital on.
Your backtest is only as good as your data pipeline.
Next Steps
If you're an individual quant developer ready to fetch historical data:
- Sign up at tickdb.ai (free tier available, no credit card required)
- Set the
TICKDB_API_KEYenvironment variable - Clone the code from this article and run
python batch_fetcher.py - Monitor the checkpoint database to verify progress
If you need tick-level trade data for US equities:
TickDB's trades endpoint does not currently cover US equities. For tick data at that resolution, consider Polygon as a supplementary provider. The code architecture in this article can be adapted to Polygon with changes to the endpoint URL, headers, and response parsing.
If you're building a team-scale data pipeline:
Reach out to [email protected] for dedicated throughput allocations, SLA guarantees, and custom endpoint access.
If you're integrating this into an AI-assisted workflow:
Search for and install the tickdb-market-data SKILL in your AI coding assistant's marketplace for direct API access within your development environment.
This article does not constitute investment advice. Historical data retrieval and backtesting are engineering tasks; market strategies involve risk, and past performance does not guarantee future results.