Ai Ethics

1 posts

figma3 min readCurated summary

Ovetta Sampson on Inputs and Outputs | Figma Blog

Minimum viable data is the idea that AI projects should begin with representative, high-quality data—not with the most powerful model or feature. Ovetta Sampson argues that model outputs are overwhelmingly determined by their inputs, which reflect human choices and historical biases. Product builders should therefore question whether AI is necessary, who it serves, and whether the data is equitable enough to avoid harming overlooked groups. ## Data Quality Shapes AI Outcomes - The quality of an AI system depends primarily on the data used to train and operate it. - Decisions about what data to collect, exclude, label, and measure determine who the system recognizes and how it behaves. - Data is never purely objective: it is generated, engineered, and transformed by people. - Treating data as disconnected from human lives can produce “traumatized data sets,” embedding social, cultural, and economic harms into models. ## The Consequences of Omission - Historical datasets often exclude entire groups: - U.S. credit and mortgage models were developed before women could independently obtain mortgages or credit cards. - The U.S. Census did not recognize LGBTQ individuals until 2021, despite those people existing in earlier populations. - When people are absent from the data, models may fail to serve them or may expose them to harmful decisions. - The central question is not simply whether data exists, but whether it represents the people affected by the system. ## Define the Problem Before Choosing AI - Teams should first identify the problem they are trying to solve and determine whether ML or AI is appropriate. - The fact that a problem can be addressed with AI does not mean it should be. - Builders should ask: - Who is the product for? - Is the data equitable and sufficiently high quality? - What is the minimum data and technology needed? - Could the proposed solution increase human risks or reduce people to data points? - Minimum viable data means collecting what is necessary for a useful, responsible solution rather than indiscriminately gathering more data. ## Putting People Back in Control - Product builders and the public need to participate in decisions about how AI systems are designed and governed. - Important questions include who defines “good” data, who decides what enters a training set, and how much data is truly required. - Sampson recommends learning from work such as *Weapons of Math Destruction*, *Ghost Work*, and research on the lack of attention given to data work in AI development. AI development should start with the people affected by a system and the data needed to represent them fairly. Choosing the smallest appropriate dataset and validating its quality can be more responsible—and more effective—than pursuing larger models or unnecessary AI features.

Read original(opens in new tab)