The 3 AM Realization

The Slack message arrived at 3:17 AM: "Backtest server down. Nobody can run models."

Two hours later, the team's senior researcher was still debugging a broken Python environment while the junior quant had lost six hours of work because the latest strategy code had never been committed to the repository. The third member—the data engineer—had accidentally overwritten the shared SQLite database with a stale backup.

This is not a hypothetical. It is a scenario we have seen play out repeatedly across small quant teams that scale from solo traders to three-person operations. The technical complexity does not increase linearly; it compounds. The moment you introduce multiple collaborators, multiple strategies, and multiple data sources, you encounter a class of engineering problems that individual quant developers rarely face.

Solo quant work is a learnable discipline. Team collaboration is a different system entirely—one that requires intentional architecture.

This article builds that architecture from the ground up. We will construct a production-grade data infrastructure for a three-person quant team, covering shared market data access, Git-based code collaboration, API key security, and permission-controlled workflows. By the end, your team will have a system that survives a 3 AM incident without cascading failures.


The Collaboration Problem in Quantitative Trading

Why Three People Changes Everything

A single quant developer operates in a contained system. One person manages the code, the data pipeline, the backtesting engine, and the deployment environment. When something breaks, there is exactly one person to blame and one person to fix it. The simplicity is deceptive.

A three-person quant team introduces three distinct failure domains:

Data contention: Multiple team members querying the same market data API simultaneously, each maintaining their own local copy of historical data. The result is inconsistent backtest environments, duplicate API calls burning through rate limits, and no single source of truth.

Code divergence: The researcher writes a new alpha factor in a Jupyter notebook on Monday. The junior quant builds a backtesting framework in a separate repository on Tuesday. The data engineer refactors the API integration on Wednesday. By Thursday, nobody knows which version of the code produced which results.

Credential sprawl: Each team member creates their own API key. Credentials get shared via Slack messages, pasted into shared Google Docs, or worse—hardcoded into configuration files that eventually find their way into version control.

These are not edge cases. They are the default outcome of teams that scale without infrastructure planning.

Quantifying the Cost

Consider a team running 50 backtests per week across three members. Without shared data infrastructure:

  • Duplicate API calls for identical historical data waste approximately 30–40% of API quota
  • Code divergence extends mean time to production by an estimated 2–3 weeks per strategy
  • Credential incidents (revoked keys, leaked secrets) require an average of 4 hours of incident response

A production-grade infrastructure does not just prevent these costs. It creates a compounding advantage: every strategy developed on shared infrastructure is immediately accessible to the entire team, every backtest runs on identical data, and every deployment follows the same secure pattern.


System Architecture: The Shared Data Stack

Layer Overview

The infrastructure for a three-person quant team decomposes into four distinct layers:

Layer Component Responsibility
Data ingestion TickDB API / WebSocket Real-time and historical market data
Data storage PostgreSQL + TimescaleDB Structured storage with time-series optimization
Code collaboration Git + GitHub / GitLab Version control, code review, deployment automation
Access control Environment variables + secrets manager Credential isolation, permission boundaries

Architecture Diagram

┌─────────────────────────────────────────────────────────────────┐
│                        Team Members (3)                         │
│   ┌─────────────┐   ┌─────────────┐   ┌─────────────┐          │
│   │ Researcher  │   │ Junior Quant│   │ Data Engineer│          │
│   └──────┬──────┘   └──────┬──────┘   └──────┬──────┘          │
└──────────┼─────────────────┼─────────────────┼──────────────────┘
           │                 │                 │
           ▼                 ▼                 ▼
┌─────────────────────────────────────────────────────────────────┐
│                     Local Development Environments               │
│   ┌─────────────┐   ┌─────────────┐   ┌─────────────┐          │
│   │ research/   │   │ strategies/ │   │ pipelines/  │          │
│   │ notebooks   │   │ backtests   │   │ ingestion   │          │
│   └─────────────┘   └─────────────┘   └─────────────┘          │
└─────────────────────────────────────────────────────────────────┘
                              │
                              ▼
┌─────────────────────────────────────────────────────────────────┐
│                    Shared Git Repository                         │
│   ├── data/           # Data access libraries                   │
│   ├── strategies/     # Strategy implementations                │
│   ├── backtests/      # Backtesting frameworks                  │
│   └── configs/        # Shared configuration templates          │
└─────────────────────────────────────────────────────────────────┘
                              │
                              ▼
┌─────────────────────────────────────────────────────────────────┐
│                   Shared Data Infrastructure                     │
│   ┌─────────────────┐   ┌─────────────────┐                    │
│   │ PostgreSQL +    │   │ Secrets Manager │                    │
│   │ TimescaleDB     │   │ (API Keys)      │                    │
│   └─────────────────┘   └─────────────────┘                    │
└─────────────────────────────────────────────────────────────────┘
                              │
                              ▼
┌─────────────────────────────────────────────────────────────────┐
│                      External Data Sources                       │
│   ┌─────────────────┐   ┌─────────────────┐   ┌──────────────┐  │
│   │ TickDB API      │   │ Exchange APIs   │   │ Alternative  │  │
│   │ (Market Data)   │   │ (Real-time)     │   │ Data Sources │  │
│   └─────────────────┘   └─────────────────┘   └──────────────┘  │
└─────────────────────────────────────────────────────────────────┘

Data Layer: Building a Shared TickDB Integration

The Shared Data Access Pattern

The core principle of shared data infrastructure is simple: one source of truth, many readers. Rather than each team member maintaining their own copy of historical OHLCV data, the team operates a shared data ingestion pipeline that populates a centralized PostgreSQL database with TimescaleDB extensions.

This approach yields three immediate benefits:

  1. Reduced API calls: The ingestion pipeline fetches data once and stores it locally. All backtests query the local database, eliminating duplicate API requests.

  2. Consistent backtest environments: Every team member runs backtests against the same historical dataset. Results are reproducible across the entire team.

  3. Historical coverage: With TickDB's 10+ years of cleaned US equity OHLCV data, the shared database accumulates a growing asset that compounds in value over time.

Production-Grade TickDB Client

The following code implements a shared TickDB client suitable for a team environment. This is not a demonstration snippet—it is production infrastructure with proper authentication, error handling, and connection management.

"""
TickDB shared client for team data infrastructure.
Handles authentication, rate limiting, and connection resilience.

⚠️ Engineering note: This client is designed for shared team environments.
Each team member should instantiate their own client instance; the underlying
connection pool is thread-safe for concurrent reads.
"""

import os
import time
import logging
import requests
from typing import Optional, Dict, Any, List
from datetime import datetime
from contextlib import contextmanager

logger = logging.getLogger(__name__)


class TickDBClient:
    """
    Production-grade TickDB API client with shared infrastructure support.
    
    Features:
    - Environment variable-based authentication
    - Automatic rate limit handling (3001 error code + Retry-After)
    - Exponential backoff with jitter for resilience
    - Connection timeout enforcement
    - Shared data caching interface
    
    ⚠️ For high-frequency workloads (>100 requests/minute), consider
    implementing a local caching layer to reduce API load.
    """
    
    BASE_URL = "https://api.tickdb.ai/v1"
    
    def __init__(self, api_key: Optional[str] = None):
        """
        Initialize TickDB client with API key from environment.
        
        Args:
            api_key: Optional override. If not provided, reads from
                    TICKDB_API_KEY environment variable.
        
        Raises:
            ValueError: If no valid API key is found.
        """
        self.api_key = api_key or os.environ.get("TICKDB_API_KEY")
        if not self.api_key:
            raise ValueError(
                "TickDB API key not found. Set TICKDB_API_KEY environment variable "
                "or pass api_key parameter."
            )
        
        self.session = requests.Session()
        self.session.headers.update({"X-API-Key": self.api_key})
        
        # Rate limiting state
        self._last_request_time = 0
        self._min_request_interval = 0.1  # 100ms minimum between requests
    
    def _apply_rate_limit(self):
        """
        Enforce client-side rate limiting to prevent 3001 errors.
        
        ⚠️ Engineering note: This is a conservative client-side throttle.
        The server-side limit depends on your plan tier. Adjust _min_request_interval
        based on your specific rate limit allocation.
        """
        elapsed = time.time() - self._last_request_time
        if elapsed < self._min_request_interval:
            time.sleep(self._min_request_interval - elapsed)
        self._last_request_time = time.time()
    
    def _request_with_retry(
        self,
        method: str,
        endpoint: str,
        params: Optional[Dict[str, Any]] = None,
        max_retries: int = 3,
        timeout: tuple = (3.05, 10)
    ) -> Dict[str, Any]:
        """
        Execute HTTP request with exponential backoff and jitter.
        
        Args:
            method: HTTP method (GET, POST)
            endpoint: API endpoint path
            params: Query parameters
            max_retries: Maximum retry attempts
            timeout: (connect_timeout, read_timeout) tuple
            
        Returns:
            Parsed JSON response
            
        Raises:
            RuntimeError: On unrecoverable errors after max_retries
        """
        url = f"{self.BASE_URL}{endpoint}"
        base_delay = 1.0
        max_delay = 30.0
        
        for attempt in range(max_retries):
            try:
                self._apply_rate_limit()
                
                response = self.session.request(
                    method=method,
                    url=url,
                    params=params,
                    timeout=timeout
                )
                
                # Parse response
                if response.status_code == 200:
                    return response.json()
                
                # Handle TickDB error codes
                try:
                    error_data = response.json()
                    error_code = error_data.get("code", 0)
                    
                    # Rate limit exceeded — respect Retry-After header
                    if error_code == 3001:
                        retry_after = int(response.headers.get("Retry-After", 5))
                        logger.warning(
                            f"Rate limit hit (attempt {attempt + 1}/{max_retries}). "
                            f"Waiting {retry_after}s before retry."
                        )
                        time.sleep(retry_after)
                        continue
                    
                    # Authentication errors — do not retry
                    if error_code in (1001, 1002):
                        raise ValueError(
                            f"Authentication failed (code {error_code}). "
                            "Verify your TICKDB_API_KEY is valid and active."
                        )
                    
                    # Symbol not found — do not retry
                    if error_code == 2002:
                        symbol = params.get("symbol", "unknown") if params else "unknown"
                        raise KeyError(
                            f"Symbol '{symbol}' not found. "
                            "Verify symbol format via /v1/symbols/available endpoint."
                        )
                    
                    # Other errors — retry with backoff
                    logger.warning(
                        f"API error {error_code}: {error_data.get('message', 'Unknown')}"
                    )
                    
                except ValueError:
                    # Non-JSON response
                    raise RuntimeError(
                        f"Unexpected response format from API: HTTP {response.status_code}"
                    )
                
            except requests.exceptions.Timeout:
                logger.warning(
                    f"Request timeout (attempt {attempt + 1}/{max_retries})"
                )
            except requests.exceptions.ConnectionError as e:
                logger.warning(
                    f"Connection error (attempt {attempt + 1}/{max_retries}): {e}"
                )
            
            # Exponential backoff with jitter
            if attempt < max_retries - 1:
                delay = min(base_delay * (2 ** attempt), max_delay)
                jitter = (hash(time.time()) % 1000) / 1000.0 * delay * 0.1
                sleep_time = delay + jitter
                
                logger.info(f"Retrying in {sleep_time:.2f}s...")
                time.sleep(sleep_time)
        
        raise RuntimeError(
            f"Request failed after {max_retries} attempts to {endpoint}"
        )
    
    def get_kline(
        self,
        symbol: str,
        interval: str = "1h",
        limit: int = 100,
        start_time: Optional[int] = None,
        end_time: Optional[int] = None
    ) -> List[Dict[str, Any]]:
        """
        Fetch OHLCV kline data for a symbol.
        
        Args:
            symbol: Trading symbol (e.g., "AAPL.US", "BTC.Binance")
            interval: Candle interval ("1m", "5m", "1h", "1d")
            limit: Maximum number of candles (1-1000)
            start_time: Unix timestamp (ms), optional
            end_time: Unix timestamp (ms), optional
            
        Returns:
            List of kline candles with OHLCV data
            
        Example:
            >>> client = TickDBClient()
            >>> candles = client.get_kline("AAPL.US", interval="1h", limit=500)
            >>> print(f"Fetched {len(candles)} candles")
        """
        params = {
            "symbol": symbol,
            "interval": interval,
            "limit": min(limit, 1000)
        }
        
        if start_time:
            params["start"] = start_time
        if end_time:
            params["end"] = end_time
        
        response = self._request_with_retry("GET", "/market/kline", params=params)
        return response.get("data", [])
    
    def get_latest_kline(self, symbol: str, interval: str = "1h") -> Optional[Dict[str, Any]]:
        """
        Fetch the most recent kline candle for real-time dashboards.
        
        ⚠️ Engineering note: Use get_kline() for historical analysis.
        This endpoint is optimized for live dashboard use cases only.
        """
        params = {"symbol": symbol, "interval": interval}
        response = self._request_with_retry("GET", "/market/kline/latest", params=params)
        data = response.get("data")
        return data[0] if data else None
    
    def get_available_symbols(self, market: Optional[str] = None) -> List[str]:
        """
        List available trading symbols, optionally filtered by market.
        
        Args:
            market: Filter by market code (e.g., "US", "HK", "Crypto")
                   If None, returns all available symbols.
                   
        Returns:
            List of available symbol strings
        """
        params = {}
        if market:
            params["market"] = market
        
        response = self._request_with_retry("GET", "/symbols/available", params=params)
        return response.get("data", [])


# Singleton pattern for shared team usage
_client_instance: Optional[TickDBClient] = None


def get_shared_client() -> TickDBClient:
    """
    Get or create the shared TickDB client instance.
    
    Use this function across your team to ensure all code paths
    share the same client configuration and respect rate limits.
    
    ⚠️ Thread safety: This implementation is not thread-safe for
    concurrent writes. For multi-threaded applications, use
    thread-local storage or a connection pool.
    """
    global _client_instance
    if _client_instance is None:
        _client_instance = TickDBClient()
    return _client_instance

Loading Historical Data into Shared Storage

With the client in place, the data engineer should implement an ingestion script that populates the shared PostgreSQL database:

"""
Data ingestion script: Load historical data from TickDB into shared storage.
Run this as a scheduled job (e.g., daily at 00:30 UTC) to keep historical data current.

Usage:
    python -m data.ingest --symbols AAPL.US,MSFT.US,GOOGL.US --interval 1h
"""

import argparse
import logging
from datetime import datetime, timedelta
from typing import List

import psycopg2
from psycopg2.extras import execute_values

from tickdb_client import TickDBClient

logging.basicConfig(
    level=logging.INFO,
    format="%(asctime)s - %(name)s - %(levelname)s - %(message)s"
)
logger = logging.getLogger(__name__)


def get_db_connection():
    """
    Create database connection using environment variables.
    
    ⚠️ Engineering note: In production, use a secrets manager (AWS Secrets Manager,
    HashiCorp Vault) instead of environment variables for database credentials.
    """
    return psycopg2.connect(
        host=os.environ["DB_HOST"],
        port=os.environ.get("DB_PORT", 5432),
        database=os.environ["DB_NAME"],
        user=os.environ["DB_USER"],
        password=os.environ["DB_PASSWORD"]
    )


def ensure_schema_exists(conn):
    """Create kline_data table if it does not exist."""
    with conn.cursor() as cur:
        cur.execute("""
            CREATE TABLE IF NOT EXISTS kline_data (
                id SERIAL PRIMARY KEY,
                symbol VARCHAR(32) NOT NULL,
                interval VARCHAR(8) NOT NULL,
                open_time TIMESTAMP NOT NULL,
                open_price DECIMAL(18, 8) NOT NULL,
                high_price DECIMAL(18, 8) NOT NULL,
                low_price DECIMAL(18, 8) NOT NULL,
                close_price DECIMAL(18, 8) NOT NULL,
                volume DECIMAL(18, 8) NOT NULL,
                ingested_at TIMESTAMP DEFAULT CURRENT_TIMESTAMP,
                UNIQUE(symbol, interval, open_time)
            );
            
            -- TimescaleDB hypertable for time-series optimization
            SELECT create_hypertable('kline_data', 'open_time', 
                if_not_exists => TRUE,
                migrate_data => TRUE);
            
            -- Index for fast symbol + interval lookups
            CREATE INDEX IF NOT EXISTS idx_kline_symbol_interval 
                ON kline_data (symbol, interval, open_time DESC);
        """)
        conn.commit()
    logger.info("Database schema verified.")


def ingest_symbols(symbols: List[str], interval: str, lookback_days: int = 30):
    """
    Ingest historical kline data for specified symbols.
    
    Args:
        symbols: List of trading symbols
        interval: Candle interval
        lookback_days: Number of days of history to fetch
    """
    client = TickDBClient()
    conn = get_db_connection()
    ensure_schema_exists(conn)
    
    end_time = int(datetime.utcnow().timestamp() * 1000)
    start_time = int((datetime.utcnow() - timedelta(days=lookback_days)).timestamp() * 1000)
    
    for symbol in symbols:
        logger.info(f"Ingesting {symbol} ({interval})...")
        
        try:
            candles = client.get_kline(
                symbol=symbol,
                interval=interval,
                limit=1000,
                start_time=start_time,
                end_time=end_time
            )
            
            if not candles:
                logger.warning(f"No data returned for {symbol}")
                continue
            
            # Transform to database rows
            rows = [
                (
                    c["symbol"],
                    interval,
                    datetime.fromtimestamp(c["open_time"] / 1000),
                    c["open"],
                    c["high"],
                    c["low"],
                    c["close"],
                    c["volume"]
                )
                for c in candles
            ]
            
            # Upsert data (insert or update on conflict)
            with conn.cursor() as cur:
                execute_values(
                    cur,
                    """
                    INSERT INTO kline_data 
                        (symbol, interval, open_time, open_price, high_price, 
                         low_price, close_price, volume)
                    VALUES %s
                    ON CONFLICT (symbol, interval, open_time) 
                    DO UPDATE SET
                        open_price = EXCLUDED.open_price,
                        high_price = EXCLUDED.high_price,
                        low_price = EXCLUDED.low_price,
                        close_price = EXCLUDED.close_price,
                        volume = EXCLUDED.volume,
                        ingested_at = CURRENT_TIMESTAMP
                    """,
                    rows,
                    template=None,
                    page_size=100
                )
            conn.commit()
            logger.info(f"Successfully ingested {len(rows)} candles for {symbol}")
            
        except Exception as e:
            logger.error(f"Failed to ingest {symbol}: {e}")
            conn.rollback()
            continue
    
    conn.close()
    logger.info("Ingestion complete.")


if __name__ == "__main__":
    parser = argparse.ArgumentParser(description="Ingest TickDB data to shared storage")
    parser.add_argument("--symbols", required=True, help="Comma-separated symbol list")
    parser.add_argument("--interval", default="1h", help="Kline interval")
    parser.add_argument("--lookback-days", type=int, default=30, help="Lookback period")
    
    args = parser.parse_args()
    symbols = [s.strip() for s in args.symbols.split(",")]
    
    ingest_symbols(symbols, args.interval, args.lookback_days)

Code Collaboration: Git Workflow for Quantitative Teams

The Branch Strategy

Git workflows exist on a spectrum from "commit to main directly" to elaborate branching models with multiple environment stages. For a three-person quant team, we recommend a streamlined trunk-based development approach with feature branches:

Branch Purpose Lifetime
main Production-ready code, deployable at any time Permanent
develop Integration branch for completed features Permanent
feature/* Individual feature development 1–7 days
hotfix/* Emergency production fixes Hours to 1 day
experiment/* Research notebooks, prototype strategies 1–14 days

Repository Structure

quant-team/
├── README.md
├── .gitignore
├── .env.example
├── requirements.txt
├── src/
│   ├── data/              # Shared data access layer
│   │   ├── __init__.py
│   │   ├── tickdb_client.py
│   │   └── storage.py
│   ├── strategies/         # Strategy implementations
│   │   ├── momentum/
│   │   ├── mean_reversion/
│   │   └── event_driven/
│   ├── backtest/           # Backtesting framework
│   │   ├── __init__.py
│   │   ├── engine.py
│   │   └── reporters.py
│   └── utils/
│       ├── logging.py
│       └── config.py
├── notebooks/              # Jupyter research notebooks (version-controlled)
├── configs/
│   ├── dev.yaml
│   ├── staging.yaml
│   └── prod.yaml
├── tests/
│   ├── unit/
│   └── integration/
└── scripts/
    ├── ingest_data.py
    └── run_backtest.py

Commit Message Convention

Adopt a conventional commit format to generate meaningful changelogs and enable automated version bumping:

<type>(<scope>): <description>

[optional body]

[optional footer]

Type prefixes:

Type Use case
feat New strategy, indicator, or backtesting feature
fix Bug fix in strategy logic, data handling, or API client
refactor Code restructuring without behavior change
data Changes to data ingestion, storage, or API integration
test Adding or updating tests
docs Documentation changes
chore Build scripts, dependency updates

Examples:

feat(momentum): add dual-moving-average crossover with volatility filter

data(tickdb): implement depth channel subscription for order book analysis

fix(backtest): correct Sharpe ratio calculation for negative returns

refactor(config): extract API credentials to environment-based loading

Code Review Requirements

Every pull request to main or develop requires:

  1. At least one approval from a team member who did not author the PR
  2. No unresolved comments on changed files
  3. All CI checks passing (linting, unit tests, integration tests)
  4. Updated documentation if the change modifies public interfaces

This is not bureaucracy. In a three-person team, a single incorrectly merged strategy can corrupt backtest results across the entire organization. Code review is the quality gate that prevents that outcome.


API Key Management: Security Without Slowing Down

The Shared Credential Problem

API keys are the most sensitive credentials in a quant team's infrastructure. A leaked TickDB API key can result in unauthorized usage, quota exhaustion, or account suspension—none of which your trading operations can afford.

The naive solution—sharing credentials via Slack or email—is a security incident waiting to happen. The over-engineered solution—complex key management infrastructure—is impractical for a three-person team.

The correct solution is a secrets manager with environment variable integration.

Environment-Based Configuration Pattern

The standard pattern for three-person teams uses a .env file for local development and a secrets manager for shared environments:

"""
Configuration loader with environment variable support.
Loads from .env file in development, from secrets manager in production.

Usage:
    from config import get_config
    config = get_config()
    api_key = config.tickdb_api_key
"""

import os
from pathlib import Path
from typing import Optional
from dataclasses import dataclass

try:
    from dotenv import load_dotenv
except ImportError:
    load_dotenv = None  # No-op if python-dotenv not installed


@dataclass
class Config:
    """Application configuration loaded from environment variables."""
    
    # TickDB credentials
    tickdb_api_key: str
    tickdb_base_url: str = "https://api.tickdb.ai/v1"
    
    # Database credentials
    db_host: str
    db_port: int = 5432
    db_name: str
    db_user: str
    db_password: str
    
    # Logging
    log_level: str = "INFO"
    
    @classmethod
    def from_env(cls, env_file: Optional[str] = None) -> "Config":
        """
        Load configuration from environment variables.
        
        Args:
            env_file: Optional path to .env file. If not provided,
                     searches for .env in the project root.
        
        Returns:
            Config instance with validated credentials.
            
        Raises:
            ValueError: If required environment variables are missing.
        """
        # Load .env file in development
        if load_dotenv is not None:
            if env_file:
                load_dotenv(env_file)
            else:
                # Look for .env in project root
                env_path = Path(__file__).parent.parent / ".env"
                if env_path.exists():
                    load_dotenv(env_path)
        
        # Required variables
        required = ["TICKDB_API_KEY", "DB_HOST", "DB_NAME", "DB_USER", "DB_PASSWORD"]
        missing = [v for v in required if not os.environ.get(v)]
        
        if missing:
            raise ValueError(
                f"Missing required environment variables: {', '.join(missing)}\n"
                "Copy .env.example to .env and fill in your credentials."
            )
        
        return cls(
            tickdb_api_key=os.environ["TICKDB_API_KEY"],
            tickdb_base_url=os.environ.get("TICKDB_BASE_URL", cls.tickdb_base_url),
            db_host=os.environ["DB_HOST"],
            db_port=int(os.environ.get("DB_PORT", cls.db_port)),
            db_name=os.environ["DB_NAME"],
            db_user=os.environ["DB_USER"],
            db_password=os.environ["DB_PASSWORD"],
            log_level=os.environ.get("LOG_LEVEL", cls.log_level)
        )


# Global config instance
_config: Optional[Config] = None


def get_config(env_file: Optional[str] = None) -> Config:
    """Get or create the global configuration instance."""
    global _config
    if _config is None:
        _config = Config.from_env(env_file)
    return _config

.env.example Template

# .env.example — Copy to .env and fill in values
# IMPORTANT: Never commit .env to version control

# TickDB API credentials
TICKDB_API_KEY=tk_live_xxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxx
TICKDB_BASE_URL=https://api.tickdb.ai/v1

# Database credentials
DB_HOST=localhost
DB_PORT=5432
DB_NAME=quant_team
DB_USER=quant_user
DB_PASSWORD=your_secure_password_here

# Logging
LOG_LEVEL=INFO

Gitignore Configuration

# .gitignore

# Environment files — NEVER commit these
.env
.env.local
.env.*.local

# Python
__pycache__/
*.py[cod]
*$py.class
.venv/
venv/

# Jupyter
.ipynb_checkpoints/

# Data files (if stored locally)
*.db
*.sqlite
data/*.csv
data/*.parquet

# IDE
.idea/
.vscode/
*.swp
*.swo

# OS
.DS_Store
Thumbs.db

Permission Control: Who Can Do What

Role-Based Access Model

For a three-person team, we recommend three distinct roles with clear permission boundaries:

Role Responsibilities Permissions
Data Engineer Data ingestion, database maintenance, API infrastructure Full read/write on data layer, CI/CD pipeline access, secrets management
Researcher Strategy development, backtesting, alpha discovery Read access to data, read/write on strategy code, read access to secrets (no write)
Junior Quant Strategy implementation, code review, documentation Read/write on strategy code, read access to data, no secrets access

GitHub / GitLab Permission Levels

Configure repository permissions as follows:

Team member Repository role Protected branches
Data Engineer Maintainer Can push to main after review
Researcher Developer Can create feature branches, cannot push directly to main or develop
Junior Quant Developer Same as Researcher

Secrets Access Control

In production environments, implement secrets manager access policies:

{
  "Version": "2012-10-17",
  "Statement": [
    {
      "Effect": "Allow",
      "Action": [
        "secretsmanager:GetSecretValue"
      ],
      "Resource": "arn:aws:secretsmanager:us-east-1:123456789:secret:quant-team/*",
      "Principal": {
        "AWS": [
          "arn:aws:iam::123456789:user/data-engineer"
        ]
      }
    },
    {
      "Effect": "Allow",
      "Action": [
        "secretsmanager:GetSecretValue"
      ],
      "Resource": "arn:aws:secretsmanager:us-east-1:123456789:secret:quant-team/tickdb-api-key",
      "Principal": {
        "AWS": [
          "arn:aws:iam::123456789:user/researcher",
          "arn:aws:iam::123456789:role/backtest-runner"
        ]
      }
    }
  ]
}

Deployment Configuration by Team Stage

Individual Developer (Team Size = 1)

For solo development, prioritize simplicity:

  • Local PostgreSQL installation with TimescaleDB extension
  • .env file for all credentials (with .gitignore protection)
  • Single repository with feature branches
  • Manual backtest execution

Small Quant Team (Team Size = 3)

As described in this article:

  • Shared PostgreSQL + TimescaleDB on a managed cloud instance (AWS RDS, Cloud SQL)
  • Secrets manager for production credentials
  • .env file pattern for local development
  • Git-based collaboration with code review requirements
  • Scheduled data ingestion jobs (daily at minimum)
  • Automated backtest reporting via CI/CD

Growing Team (Team Size = 5–10)

When the team scales beyond three members:

  • Implement team-level API keys with usage tracking
  • Add a staging environment for backtest validation
  • Introduce integration tests in CI/CD pipeline
  • Add audit logging for data access
  • Consider a feature flag system for gradual strategy rollout

Putting It Together: A Day in the Life

With this infrastructure in place, here is how a typical day unfolds for each team member:

Morning (Data Engineer)

  • Checks the automated ingestion job logs from the previous night
  • Resolves any data gaps (a symbol that failed to fetch, a stale checkpoint)
  • Merges any approved pull requests from the previous day
  • Reviews API quota consumption across the team

Mid-day (Researcher)

  • Pulls the latest shared data to local environment
  • Develops a new momentum factor in a Jupyter notebook
  • Commits code to a feature branch, opens a pull request
  • Receives code review from the junior quant with feedback on documentation

Afternoon (Junior Quant)

  • Implements the approved momentum factor as a production strategy
  • Writes unit tests covering edge cases (empty data, API timeout)
  • Participates in code review for the researcher's pull request
  • Runs a backtest using the shared data infrastructure, generates a performance report

Evening (Automated)

  • CI/CD pipeline runs tests on all merged code
  • Scheduled ingestion job updates historical data
  • Slack notification reports any anomalies to the data engineer

The system is designed to be boring. In production infrastructure, boredom is a feature.


Conclusion

The transition from solo quant development to team collaboration is not merely a scaling challenge—it is an architectural one. The tools and patterns that work for individuals collapse under the weight of multiple collaborators, multiple strategies, and multiple data sources.

The infrastructure outlined in this article provides a production-grade foundation:

  • Shared data access through a centralized TickDB client and PostgreSQL storage eliminates duplicate API calls and ensures consistent backtest environments.
  • Git-based collaboration with a structured repository layout and conventional commit messages enables code review without slowing down development.
  • Environment-based credential management with .gitignore protection prevents accidental credential exposure while keeping the development workflow simple.
  • Role-based permissions create clear accountability without bureaucratic overhead.

This is not minimal viable infrastructure. It is the infrastructure that allows a three-person team to operate at the level of a well-funded prop trading desk—one where a 3 AM incident triggers an automated alert rather than a cascade of manual debugging.

The question is no longer whether this infrastructure is necessary. It is whether you can afford to operate without it.


Next Steps

If you are an individual quant developer building toward a team, start with the shared TickDB client and .env pattern. These two components alone will save you hours of debugging when collaboration begins.

If you are a small team without shared data infrastructure, evaluate your current API consumption. If three team members are making duplicate requests for the same historical data, you are paying for data you already own. A shared ingestion pipeline reduces costs immediately.

If you need 10+ years of historical OHLCV data for cross-cycle backtesting, TickDB provides cleaned, aligned US equity data via a unified API. Sign up at tickdb.ai (free tier available) and integrate the client library documented above into your team's shared codebase.

If you use AI coding assistants, search for and install the tickdb-market-data SKILL in your AI tool's marketplace to get context-aware integration assistance while building your team's infrastructure.


This article does not constitute investment advice. Markets involve risk; past performance does not guarantee future results. All backtesting results are based on historical simulation and do not reflect actual trading outcomes.