Data Modeling

7 posts

netflix2 min readCurated summary

Modeling Device Capabilities for Analytics

Netflix models device capabilities to determine which features can be safely supported across its diverse hardware ecosystem. By tracking hardware, software, and platform limitations in scalable analytical datasets, Netflix can measure feature reach and identify adoption bottlenecks. This enables more precise feature management for capabilities such as 4K, spatial audio, cloud gaming, and new UI experiences. ## Building a Device Capability Model - Devices vary significantly in RAM, CPU cores, display resolution, audio support, and platform capabilities. - Netflix maintains detailed capability data for each device model, including: - Screen dimensions and resolution - Supported video profiles and codecs - Surround sound support - RAM capacity - Software version and platform information - Internal feature flags are integrated into the model to connect device capabilities with feature availability. ## Cumulative Tables for Current Device State - Netflix uses a cumulative table to track the latest known capabilities for each device. - Capabilities are stored in a structured format, such as supported screen sizes and video profiles. - This design supports large-scale analytics and reporting by providing an up-to-date view of device functionality. ## Histogram Tables for Feature Distribution - A histogram table measures active devices over the previous 28 days. - Results are broken down by device model and software version. - The table also counts how many devices support particular capabilities. - For example, Netflix can analyze external display support on streaming sticks: - 100% of devices may support the HD PlayReady profile. - Only 20% may support the UHD HEVC profile. ## Using Analytics for Feature Management - Netflix uses these datasets to evaluate feature penetration for products such as: - 4K Ultra HD - Netflix Spatial Audio - Cloud Gaming - Updated user interfaces - Capability data helps teams identify hardware or software bottlenecks. - Feature decisions can therefore be made at a more granular level, improving performance, reliability, and user experience. Netflix’s approach demonstrates that a detailed, analytics-focused capability model is essential for managing features across a global and highly varied device ecosystem.

Read original(opens in new tab)
netflix3 min readCurated summary

Scaling Global Storytelling: Modernizing Localization Analytics at Netflix

Netflix is modernizing its localization analytics to support more than 300 million members across 190+ countries and 50+ languages. Rapid growth created duplicated pipelines, inconsistent business logic, and siloed dashboards, making basic questions such as who produced a dub difficult to answer reliably. The company’s solution is to consolidate data foundations, improve usability, and centralize reusable business logic. ## The Challenge of Fragmented Localization Data - Localization metrics were historically built independently across different teams and workflows. - Determining who created a dub or subtitle required combining multiple sources with complex, frequently changing rules. - Duplicated logic led to: - Inconsistent reporting across tools - High maintenance costs when upstream systems changed - Siloed analytics and dashboards ## Auditing and Consolidating Analytics - Netflix audited more than 40 dashboards and tools for usage, quality, and code health. - The focus shifted from repeatedly fixing frontend visualizations to consolidating backend data pipelines. - Three legacy dashboards covering dubbing-partner operations, capacity, and finances are being unified around a shared data and backend layer. - This foundation can support multiple future frontend experiences instead of forcing each dashboard to maintain separate logic. ## Reducing User Experience Debt - Netflix defines “Not-So-Tech Debt” as stakeholder friction caused by confusing tools or weak analytical storytelling. - The Language Asset Consumption tool was redesigned to combine audio and text languages into a single consumption-language view. - This distinguishes: - Original-language viewing from localized consumption - Subtitle, dubbing, or combined preferences - Recurring member preferences for a given language - The result is more intuitive analysis aligned with real stakeholder questions. ## Centralizing Reusable Business Logic - Netflix is adopting a “write once, read many” architecture. - Shared tables, including a Language Asset Producer table, solve common questions in one centralized location. - The same trusted data can feed downstream domains such as Dub Quality and Translation Quality. - Updates to business rules propagate across the analytics ecosystem instead of requiring changes in multiple pipelines. ## Moving Toward Event-Level Analytics - Future work will analyze individual timed-text events rather than only complete language assets. - A generic model will capture details such as individual subtitle lines and reading speed. - Netflix plans to connect subtitle characteristics with member engagement. - These findings can improve style guidelines for subtitle linguists and ultimately enhance the localized viewing experience. Netflix’s recommendation is to treat analytics modernization as both a technical and product-quality effort: consolidate data foundations, centralize business logic, and design tools around how stakeholders actually make decisions. This creates more trustworthy reporting while enabling deeper analysis of how localization affects member enjoyment.

Read original(opens in new tab)
figma3 min readCurated summary

Redefining Impact as a Data Scientist | Figma Blog

Data science impact is not limited to experiments, forecasting, or optimization. In complex, high-stakes systems such as billing, data scientists can create value by making workflows understandable, validating correctness, and improving operational safety. Figma’s experience shows that effective data science may require domain modeling, cross-functional collaboration, instrumentation, and production-quality tools. ## Data Science as a Full-Stack Discipline - The role of data science varies by team: it may involve experimentation, product analysis, data modeling, instrumentation, or operational tooling. - Billing combines a user-facing product with backend infrastructure, so accuracy directly affects customer trust. - Supporting Billing required: - Building deep domain expertise - Partnering with engineers and other functions - Creating tools that explain and verify system behavior - Experimentation and opportunity analysis remained useful, but represented a smaller portion of the actual work. - Figma’s full-stack model encouraged the team to define the right data science support collaboratively rather than follow a fixed playbook. ## Explaining Complex Systems Beyond Charts and Models - Some of the most valuable data science work explains existing or historical outcomes rather than predicting future ones. - A single invoice seat charge may depend on: - Seat assignments and removals - Permission changes - Contract terms - Workspace state - Billing rules - The timing of state transitions - Figma built the **Invoice Seat Report** to reconstruct the complete reasoning behind each charge. - The application combines product events, contract metadata, billing rules, and historical state transitions, presenting the result in plain language. - Building it required: - Reconciling fragmented schemas and inconsistent historical data - Validating assumptions with engineers - Adding instrumentation where logs recorded what happened but not why - Translating billing rules into traceable and debuggable SQL transformations - The team also had to account for legacy multiyear contracts, sparse seat histories, early upgrades, and other cases that could create gaps in the data. ## Shaping Technical Direction Through Data - Data scientists can turn business rules into measurable checks that define expected system behavior. - These validations can detect drift, regressions, and anomalies in both development and production. - For Billing, automated verification is especially important because small errors in seat states or invoice calculations can affect customer charges and trust. - During Figma’s billing-model re-architecture, data science helped verify that: - Data moved correctly through pipelines - New pricing and billing logic produced intended outcomes - Customers did not enter unexpected billing states - The system could be monitored consistently across environments The practical lesson is to look beyond conventional analytics when assessing data science impact. In complex domains, building reliable data foundations, explanatory tools, and correctness checks may be more valuable than running another experiment.

Read original(opens in new tab)
airbnb3 min readCurated summary

My Journey to Airbnb: Peter Coles

Peter Coles’s career connects mathematical training, academic economics, and practical data science. After studying game theory and market design, he moved from Harvard Business School to eBay and then Airbnb, where he could apply economic models to real-world marketplaces. At Airbnb, he helped build economics and data science teams, guide policy decisions, investigate pandemic-driven changes, and measure the company’s broader impact. ## From Mathematics to Economics - Coles grew up in Milwaukee and developed an early interest in marketplaces by trying to run a neighborhood rock stand. - He studied math at Princeton after briefly pursuing ancient history. - He earned a PhD in economics at Stanford, focusing on game theory—the study of strategic decision-making. - His mentor, Jon Levin, taught him to simplify complex research problems. - While studying in Germany, Coles traveled around Europe and stayed with strangers connected to classmates, unintentionally experimenting with a model similar to Airbnb. ## Studying Markets and Market Design - At Harvard Business School, Coles researched market design and taught with Al Roth, who later won the Nobel Prize in Economics. - His work focused on “matching,” or designing systems that pair participants from two groups when prices cannot directly balance supply and demand. - He studied participant strategy, signaling, and market mechanisms, including improvements to the market for PhD economists. - He also wrote business cases about companies such as Zillow, Microsoft, and Craigslist. - Although he valued academia, he found the long research and peer-review cycle was not a good long-term fit. ## Applying Economics at eBay - In 2013, Coles joined eBay as technology and the sharing economy were rapidly expanding. - He led an economics team created by Steve Tadelis and helped combine it with another group to form eBay’s Data Labs. - One notable project, “What’s It Worth,” developed a method for estimating the fair market value of items sold on eBay. - The work combined economic reasoning, practical marketplace knowledge, and statistical modeling. ## Building Airbnb’s Economics and Data Science Functions - In 2015, Coles joined Airbnb to help address the company’s growing regulatory challenges. - He built a global team of economists and data scientists to study short-term rentals and their relationship with cities. - The team used data to inform policy discussions and evaluate Airbnb’s effects on guests, hosts, and communities. - This role allowed Coles to connect economic theory with decisions affecting a rapidly expanding platform. ## Central Strategy & Insights - As Airbnb grew, executives needed analysis that crossed organizational boundaries. - Coles and Jackson Wang founded Central Strategy & Insights, known as CSI. - The team acted as “forensic investigators,” assembling evidence and narratives from company-wide data. - During the pandemic, CSI analyzed major changes in guest travel patterns and determined what kinds of supply Airbnb would need. - The team also led business reviews and prepared analyses for shareholders before Airbnb’s IPO. ## Measuring Airbnb’s Broader Impact - Coles later returned to policy-focused work with a larger economics organization. - The team developed models to guide Airbnb’s response to governments as travel recovered after the pandemic. - Economists and analysts evaluated Airbnb’s impact on hosts, guests, and society. - Their work included the US Economic Impact Report and expanded collaboration with academic researchers using Airbnb data. Coles’s experience suggests that marketplace companies benefit from combining rigorous economic research with hands-on data science. Moving between academia and industry enabled him to turn theories about market design into practical tools for product strategy, policy, and impact measurement.

Read original(opens in new tab)
daangnOriginal article

Why Karrot made User (opens in new tab)

Daangn transitioned from manually calculating user activation metrics to a centralized "Activation Layer" built on DBT to solve inconsistencies and high operational overhead. By standardizing the definitions of user states and transitions, the team provides a reliable foundation for analyzing why active user counts fluctuate rather than just reporting the final numbers. This common data layer improves data reliability and cost-efficiency while allowing various teams to reuse the same logic for different core user behaviors. ### The Role of User Activation Analysis * While Active User counts show "what" happened, User Activation explains "why" by breaking users down into specific categories. * The system tracks **Activation States**, classifying users as New, Retained, Reactivated, or Inactive at any given time. * It monitors **State Transitions** to identify how users move between categories, such as "New to Retained" or "Reactivated to Inactive." * The layer provides granular behavioral metadata, including continuous activity streaks, the interval between visits, and the duration of churned periods. ### Ensuring Reliability via Fact Models * Raw event logs are often tied to specific UI elements and contain "noise" that makes them unreliable for direct activation analysis. * To ensure consistency, the Activation Layer uses **Fact Models** as its primary input, which are refined datasets where business logic and core behaviors are already defined. * A strict naming convention (`fact_name_activation_time_grain`) is enforced so that users can immediately identify which specific behavior is being analyzed. * This structure ensures that "Active" status is interpreted identically across the entire organization, regardless of which team is performing the analysis. ### Incremental Processing for Cost Efficiency * Calculating the entire history of user activity every day is computationally expensive and leads to high cloud infrastructure costs. * The architecture utilizes a **FirstLast model** to store only the essential metadata for each user: the date of their very first activity and their most recent activity. * By joining daily activity logs with this lightweight FirstLast table, the system can calculate new states and transitions incrementally. * This approach maintains data idempotency and ensures high performance even as the volume of user interaction data grows. ### Scaling with DBT Macros * To support various metrics—such as app visits, item sales, or community posts—the team encapsulated the complex transition logic into **DBT Macros**. * This abstraction allows data engineers to generate a new activation model by simply specifying the source Fact model and the desired time grain (daily, weekly, or monthly). * Centralizing the logic in macros ensures that any bug fixes or improvements to the activation calculation are automatically reflected across all related data models. * The standardized output format allows for the creation of universal dashboards and analysis templates that work for any tracked behavior. Centralizing User Activation logic into a common data layer allows organizations to move beyond surface-level vanity metrics and gain deep, actionable behavioral insights. By combining DBT’s macro capabilities with incremental modeling, teams can maintain high data quality and operational efficiency even as the variety of tracked user behaviors expands.

tossOriginal article

Toss People: Designing a structure (opens in new tab)

Data architecture is evolving from a reactive "cleanup" task into a proactive, end-to-end design process that ensures high data quality from the moment of creation. In fast-paced platform environments, the role of a Data Architect is to bridge the gap between rapid product development and reliable data structures, ultimately creating a foundation that both humans and AI can interpret accurately. By shifting from mere post-processing to foundational governance, organizations can maintain technical agility without sacrificing the integrity of their data assets. **From Post-Processing to End-to-End Governance** * Traditional data management often involves "fixing" or "matching puzzles" at the end of the pipeline after a service has already changed, leading to perpetual technical debt. * Effective data architecture requires a culture where data is treated as a primary design object from its inception, rather than a byproduct of application development. * The transition to an end-to-end governance model ensures that data quality is maintained throughout its entire lifecycle—from initial generation in production systems to final analysis and consumption. **Machine-Understandable Data and Ontologies** * Modern data design must move beyond human-readable metadata to structures that AI can autonomously process and understand. * The implementation of semantic-based standard dictionaries and ontologies reduces the need for "inference" or guessing by either humans or machines. * By explicitly defining the relationships and conceptual meanings of columns and tables, organizations create a high-fidelity environment where AI can provide accurate, context-aware responses without interpretive errors. **Balancing Development Speed with Data Quality** * In high-growth environments, insisting on "perfect" design can hinder competitive speed; therefore, architects must find a middle ground that allows for future extensibility. * Practical strategies include designing for current needs while leaving "logical room" for anticipated changes, ensuring that future cleanup is minimally disruptive. * Instead of enforcing rigid rules, architects should design systems where following the standard is the "path of least resistance," making high-quality data entry easier for developers than the alternative. **The Role of the Modern Data Architect** * The role has shifted from a fixed, corporate function to a dynamic problem-solver who uses structural design to solve business bottlenecks. * A successful architect must act as a mediator, convincing stakeholders that investing in a 5% quality improvement (e.g., moving from 90 to 95 points) provides significant long-term ROI in decision-making and AI reliability. * Aspiring architects should focus on incremental structural improvements, as any data professional who cares about how data functions is already operating on the path to data architecture.

datadog3 min readCurated summary

Piecewise regression: When one line simply isn’t enough

Piecewise regression models a timeseries with multiple linear segments when one line is insufficient. Datadog’s approach automatically detects both breakpoints and the number of segments, while avoiding a brute-force search of all possible partitions. It starts with an intentionally overfit model and greedily merges neighboring segments until the increase in error indicates that further merging would lose important structure. ## Objectives - **Automated breakpoint detection** - The algorithm identifies where one linear trend changes into another. - This is necessary for running hundreds of regressions per second without manual input. - **Automated segment-count selection** - The number of segments is not specified in advance. - The method must distinguish between data best represented by one line and data requiring several. - **No continuity requirement** - Adjacent regression lines do not need to meet at their shared breakpoint. - This allows the model to represent discontinuous changes in the data. ## Challenges - **Large search space** - A timeseries can be partitioned in exponentially many ways. - Although dynamic programming is more efficient than brute force, it remains too slow for Datadog’s performance requirements. - A greedy heuristic is used to eliminate large portions of the search space quickly. - **Balancing fit and simplicity** - More segments generally reduce the sum of squared errors. - Using one segment per point could produce nearly zero error but would provide little useful information for interpolation or extrapolation. - The goal is therefore to find the fewest segments that model the data accurately. ## Greedy Merging Algorithm - Begin with approximately **n/2 segments** for a timeseries containing *n* observations. - Fit each segment using ordinary least squares regression. - Repeatedly examine every pair of neighboring segments: - Calculate the increase in total squared error if the pair were merged. - Merge the pair producing the smallest error increase. - Continue merging until only one segment remains. - Record the segmentation state immediately before a merge appears to go too far. - If no merge triggers the stopping rule, select one large segment; otherwise, return the last recorded segmentation. ## Stopping Criteria - A merge becomes a potential stopping point when its increase in total squared error exceeds that of every earlier merge. - To avoid stopping prematurely on data that is fundamentally linear, the increase must also be less than **3% of the total error from a single-line regression**. - The 3% threshold is heuristic but was found to work well in practice. - For data generated from one noisy linear trend, error increases gradually as segments are merged, so no merge qualifies as an adequate stopping point and the algorithm ultimately selects one segment. The method provides a practical compromise between exhaustive optimization and model quality: greedy merging makes automated regression fast, while the error-based stopping rule limits overfitting and preserves meaningful changes in trend.

Read original(opens in new tab)