dropbox

Inside the feature store powering real-time AI in Dropbox Dash (opens in new tab)

Dropbox Dash’s ranking system depends on a hybrid feature store that can combine real-time user behavior with large-scale historical data. Because Dropbox operates across on-premises and cloud environments, and because each query can trigger thousands of feature lookups, off-the-shelf systems could not meet its latency, scale, and integration requirements. The resulting architecture uses Feast for orchestration, Spark for computation, Dynovault for low-latency storage, and a custom Go serving layer, achieving roughly 25–35 ms p95 latency while keeping features fresh.

Goals and Requirements

  • Dash ranks documents, images, and conversations using behavioral, contextual, and real-time signals.
  • A single query can fan out into thousands of feature lookups across many candidate files.
  • The feature store needed to:
    • Support sub-100 ms search latency.
    • Reflect user actions within seconds or minutes.
    • Bridge Dropbox’s on-premises services and Spark-based cloud infrastructure.
    • Handle both streaming-style updates and batch computations.
    • Let engineers develop features without managing serving and orchestration details.

Choosing a Hybrid Architecture

  • Dropbox evaluated Feast, Hopsworks, Featureform, Feathr, Databricks, and Tecton.
  • Feast was selected because:
    • It separates feature definitions from infrastructure concerns.
    • Engineers can focus on PySpark transformations.
    • Its modular adapter system supports existing Dropbox infrastructure.
  • Feast’s DynamoDB adapter enabled integration with Dynovault, Dropbox’s DynamoDB-compatible storage system.
  • The architecture combines:
    • Feast for orchestration and serving APIs.
    • Spark jobs for feature computation and ingestion.
    • Cloud storage for offline indexing and data management.
    • Dynovault for online, low-latency lookups.
    • A custom Go service replacing Feast’s Python online serving path.
  • Dynovault is colocated with inference workloads and provides approximately 20 ms client-side latency.
  • Monitoring covers job failures, feature freshness, and data lineage.

Replacing Python with Go for Low Latency

  • The initial Feast-based Python service struggled under heavy concurrency.
  • CPU-bound JSON parsing and Python’s Global Interpreter Lock became bottlenecks.
  • A multi-process design helped temporarily but introduced coordination overhead.
  • The serving layer was rewritten in Go using:
    • Lightweight goroutines.
    • Shared memory.
    • Faster JSON parsing.
  • The Go service now handles thousands of requests per second.
  • It adds only about 5–10 ms beyond Dynovault latency and achieves roughly 25–35 ms p95 latency.

Keeping Features Fresh

  • Fresh signals are essential for ranking quality; actions such as opening a document should influence subsequent searches quickly.
  • Fully real-time computation is impractical for features requiring large joins, aggregations, and historical context.
  • Dropbox therefore built a three-part ingestion strategy.
  • Batch ingestion handles complex, high-volume transformations using a medallion architecture.
  • Intelligent change detection updates only modified records rather than rewriting every feature.
  • This reduced online-store writes from hundreds of millions to fewer than one million per run and significantly shortened update time.

Practical Takeaway

The system demonstrates that a feature store does not need to be entirely off-the-shelf or entirely real-time. Combining a modular framework with custom serving, colocated storage, batch optimization, and freshness monitoring allowed Dropbox to meet demanding latency and scale requirements while keeping feature development manageable.