Inside the feature store powering real-time AI in Dropbox Dash (opens in new tab)
Dropbox Dash’s ranking system depends on a hybrid feature store that can combine real-time user behavior with large-scale historical data. Because Dropbox operates across on-premises and cloud environments, and because each query can trigger thousands of feature lookups, off-the-shelf systems could not meet its latency, scale, and integration requirements. The resulting architecture uses Feast for orchestration, Spark for computation, Dynovault for low-latency storage, and a custom Go serving layer, achieving roughly 25–35 ms p95 latency while keeping features fresh.
Goals and Requirements
- Dash ranks documents, images, and conversations using behavioral, contextual, and real-time signals.
- A single query can fan out into thousands of feature lookups across many candidate files.
- The feature store needed to:
- Support sub-100 ms search latency.
- Reflect user actions within seconds or minutes.
- Bridge Dropbox’s on-premises services and Spark-based cloud infrastructure.
- Handle both streaming-style updates and batch computations.
- Let engineers develop features without managing serving and orchestration details.
Choosing a Hybrid Architecture
- Dropbox evaluated Feast, Hopsworks, Featureform, Feathr, Databricks, and Tecton.
- Feast was selected because:
- It separates feature definitions from infrastructure concerns.
- Engineers can focus on PySpark transformations.
- Its modular adapter system supports existing Dropbox infrastructure.
- Feast’s DynamoDB adapter enabled integration with Dynovault, Dropbox’s DynamoDB-compatible storage system.
- The architecture combines:
- Feast for orchestration and serving APIs.
- Spark jobs for feature computation and ingestion.
- Cloud storage for offline indexing and data management.
- Dynovault for online, low-latency lookups.
- A custom Go service replacing Feast’s Python online serving path.
- Dynovault is colocated with inference workloads and provides approximately 20 ms client-side latency.
- Monitoring covers job failures, feature freshness, and data lineage.
Replacing Python with Go for Low Latency
- The initial Feast-based Python service struggled under heavy concurrency.
- CPU-bound JSON parsing and Python’s Global Interpreter Lock became bottlenecks.
- A multi-process design helped temporarily but introduced coordination overhead.
- The serving layer was rewritten in Go using:
- Lightweight goroutines.
- Shared memory.
- Faster JSON parsing.
- The Go service now handles thousands of requests per second.
- It adds only about 5–10 ms beyond Dynovault latency and achieves roughly 25–35 ms p95 latency.
Keeping Features Fresh
- Fresh signals are essential for ranking quality; actions such as opening a document should influence subsequent searches quickly.
- Fully real-time computation is impractical for features requiring large joins, aggregations, and historical context.
- Dropbox therefore built a three-part ingestion strategy.
- Batch ingestion handles complex, high-volume transformations using a medallion architecture.
- Intelligent change detection updates only modified records rather than rewriting every feature.
- This reduced online-store writes from hundreds of millions to fewer than one million per run and significantly shortened update time.
Practical Takeaway
The system demonstrates that a feature store does not need to be entirely off-the-shelf or entirely real-time. Combining a modular framework with custom serving, colocated storage, batch optimization, and freshness monitoring allowed Dropbox to meet demanding latency and scale requirements while keeping feature development manageable.