Techlist.io - Korean Tech Blog Curator

googleOriginal article

MLE-STAR: A state-of-the-art machine learning engineering agent (opens in new tab)

MLE-STAR is a state-of-the-art machine learning engineering agent designed to automate complex ML tasks by treating them as iterative code optimization challenges. Unlike previous agents that rely solely on an LLM’s internal knowledge, MLE-STAR integrates external web searches and targeted ablation studies to pinpoint and refine specific pipeline components. This approach allows the agent to achieve high-performance results, evidenced by its ability to win medals in 63% of Kaggle competitions within the MLE-Bench-Lite benchmark. ## External Knowledge and Targeted Ablation The core of MLE-STAR’s effectiveness lies in its ability to move beyond generic machine learning libraries by incorporating external research and specific performance testing. * The agent uses web search to retrieve task-specific, state-of-the-art models and approaches rather than defaulting to familiar libraries like scikit-learn. * Instead of modifying an entire script at once, the system conducts an ablation study to evaluate the impact of individual pipeline components, such as feature engineering or model selection. * By identifying which code blocks have the most significant impact on performance, the agent can focus its reasoning and optimization efforts where they are most needed. ## Iterative Refinement and Intelligent Ensembling Once the critical components are identified, MLE-STAR employs a specialized refinement process to maximize the effectiveness of the generated solution. * Targeted code blocks undergo iterative refinement based on LLM-suggested plans that incorporate feedback from prior experimental failures and successes. * The agent features a unique ensembling strategy where it proposes multiple candidate solutions and then designs its own method to merge them. * Rather than using simple validation-score voting, the agent iteratively improves the ensemble strategy itself, treating the combination of models as a distinct optimization task. ## Robustness and Safety Verification To ensure the generated code is both functional and reliable for real-world deployment, MLE-STAR incorporates three specialized diagnostic modules. * **Debugging Agent:** Automatically analyzes tracebacks and execution errors in Python scripts to provide iterative corrections. * **Data Leakage Checker:** Reviews the solution script prior to execution to ensure the model does not improperly access test dataset information during the training phase. * **Data Usage Checker:** Analyzes whether the script is utilizing all available data sources, preventing the agent from overlooking complex data formats in favor of simpler files like CSVs. By combining external grounding with a granular, component-based optimization strategy, MLE-STAR represents a significant shift in automated machine learning. For organizations looking to scale their ML workflows, such an agent suggests a future where the role of the engineer shifts from manual coding to high-level supervision of autonomous agents that can navigate the vast landscape of research and data engineering.

figma3 min readCurated summary

Figma’s IPO: Design Is Everyone’s Business | Figma Blog

Figma’s IPO marks a new phase for the company, but not a change in its founding mission: narrowing the gap between imagination and reality. CEO Dylan Field argues that AI will make design more accessible and important, while emphasizing that Figma will prioritize decades-long growth over short-term efficiency or share-price performance. He sees Figma as a collaborative platform where more people can shape products and ideas. ## Figma’s IPO and Long-Term Mission - Figma went public to improve corporate governance, increase brand awareness, provide liquidity, strengthen its acquisition currency, and access capital markets. - Field especially values public ownership because it allows the broader Figma community to share in the company’s success. - He cautions investors that public-market performance is unpredictable and does not promise share-price growth. - Figma will prioritize supporting designers’ evolving needs and pursuing long-term growth over maximizing quarterly efficiency. - The company expects to take significant risks, including large platform investments and mergers and acquisitions. ## AI as a Strategic Investment - Figma is investing heavily in AI and plans to increase that investment further. - This spending may reduce efficiency for several years, but Field considers AI central to the future of design workflows. - Existing capabilities, including Figma Make and other AI features, are presented as only the beginning. - AI could help designers work more effectively and bring more people into the design process. ## Design as a Competitive Advantage - Creating a minimum viable product is easier than ever, making design, craft, and distinctive perspective more important differentiators. - Design is no longer an afterthought focused only on form and function; it can determine whether a product succeeds or fails. - As design becomes more central, companies will involve more types of contributors and encourage greater experimentation and creativity. - Figma must balance accessibility with professional power while supporting collaboration and decision-making across large organizations. ## The Evolution of AI Interfaces - Field compares current AI interfaces, which rely heavily on prompts, to the MS-DOS era of computing. - He expects new, domain-specific design patterns to make AI capabilities easier and more intuitive to use. - Just as graphical user interfaces expanded access to computers, well-designed AI interfaces could make advanced capabilities available to everyday users. ## Figma’s Role in the Future - Field describes Figma as a “peaceful garden” where individuals and teams can develop ideas together. - He believes tools alone do not change the world; people use them to create meaningful change. - Figma’s long-term ambition is to help more people participate in design and turn ideas into reality. Figma’s direction is therefore one of patient, ambitious investment: accept near-term inefficiency, expand access to design, and build AI-powered tools for a much broader creative community.

Read original(opens in new tab)
figma2 min readCurated summary

Figma Announces Pricing of Initial Public Offering | Figma Blog

Figma announced pricing for its initial public offering at **$33 per share**, valuing an offering of 36,937,080 Class A shares. Trading is expected to begin on the New York Stock Exchange under **“FIG”** on July 31, 2025, with closing scheduled for August 1. The IPO marks Figma’s transition from a design tool into a broader, AI-powered product development platform. ## IPO Details - Figma is offering **12,472,657 shares** of Class A common stock. - Existing stockholders are selling **24,464,423 shares**. - The offering’s public price is **$33.00 per share**. - Certain selling stockholders granted underwriters a 30-day option to purchase up to **5,540,561 additional shares** for over-allotments. - Figma will not receive proceeds from shares sold by existing stockholders. ## Trading and Underwriting - Shares are expected to trade on the **New York Stock Exchange** under ticker symbol **FIG**. - The offering is expected to close on **August 1, 2025**, subject to customary conditions. - Morgan Stanley, Goldman Sachs, Allen & Company, and J.P. Morgan are joint lead book-running managers. - BofA Securities, Wells Fargo Securities, and RBC Capital Markets are additional book-running managers. - William Blair and Wolfe | Nomura Alliance are serving as co-managers. - The SEC declared the related registration statement effective on July 30, 2025. ## Figma’s Platform and Evolution - Founded in 2012, Figma describes itself as a collaborative platform for digital product development. - The company has expanded beyond interface design into an AI-powered system supporting: - Ideation - Design - Building - Product shipping - Its central value proposition is helping teams collaborate more efficiently while maintaining alignment throughout the product lifecycle. Figma’s IPO makes its public-market debut with a $33-per-share offering and positions the company as a broader collaborative platform for designing and delivering digital products.

Read original(opens in new tab)
figma3 min readCurated summary

A Tale of Two Parameter Architectures—and How We Unified Them | Figma Blog

Figma’s component properties and variables both let users define a value once and apply it across a design, but they were originally built on separate architectures. This led to conflicting bindings, inconsistent rendering, and duplicated engineering effort. Figma unified the systems into one parameter architecture, improving consistency and creating a scalable foundation for future products. ## Two Different Parameter Systems - **Component properties**, introduced in 2022, addressed the gap between design and code. - They enabled **scoped parametrization**: parameters could be defined by a component and applied only to its internal layers. - This gave design-system authors a clear customization contract and mirrored how component properties work in code. - The model supported features such as: - Boolean properties - Text and instance customization - Component variants and states - Component properties helped power products including Figma Sites, Figma Make, and Code Connect. ## Variables and Global Parametrization - Variables provided a broader, more globally reusable parameter system. - They supported values such as colors and could define different values for contextual modes, such as light and dark themes. - Unlike component properties, variables were designed to be reused across many unrelated layers and components. - This made variables well suited to design tokens and system-wide configuration. ## Problems Caused by Separate Architectures - The two systems evolved independently despite having similar goals. - A variable and a component property could bind to the same layer property, producing inconsistent results in the editor. - Users had to learn different behaviors and rules for concepts that appeared similar. - Maintaining two implementations increased engineering complexity and slowed feature development. - Extending parametrization to other Figma products risked reproducing the same technical limitations. ## Unifying the Architecture - Figma built a single underlying parameter architecture to support both component properties and variables. - The unified system preserves the distinction between: - Parameters scoped to a component - Parameters reused globally across a project - It eliminates competing bindings and makes parameter behavior more predictable. - The change also improves developer velocity by allowing new capabilities to build on shared infrastructure. ## Broader Impact - Designers receive a more consistent mental model for configuring components and design systems. - Rendering behavior is more reliable when multiple parameter types interact. - Figma can extend parametrization across existing and future products without maintaining separate technical foundations. - The unified architecture supports Figma’s broader goal of connecting design-system behavior with code and interactive products. Figma’s recommendation in practice is to treat component properties and variables as complementary uses of one parameter model: use scoped properties for component contracts and variables for reusable, contextual design tokens.

Read original(opens in new tab)
googleOriginal article

Simulating large systems with Regression Language Models (opens in new tab)

Researchers from Google have introduced Regression Language Models (RLMs) as a universal solution for numeric prediction tasks by framing regression as a text-to-text problem. By converting complex, unstructured system data into strings, RLMs can predict performance metrics without the need for manual feature engineering or data normalization. This approach allows large language models to move beyond subjective human feedback and directly model raw operational data for large-scale software and industrial infrastructures. ## Conceptualizing Text-to-Text Regression * Traditional regression methods rely on tabular data—fixed-length numeric vectors—which are difficult and laborious to maintain for evolving systems like software logs or hardware patterns. * RLMs represent the input state ($x$) as a structured text string (such as JSON or YAML) and the numerical output ($y$) as a text string. * The model is trained using standard next-token prediction and cross-entropy loss, allowing it to function as a universal approximator for complex data types. * This paradigm eliminates the need for manual feature engineering, as the model learns directly from the raw textual representation of the system state. ## Architecture and Training for Large Systems * The research utilizes a compact RLM consisting of a two-layer encoder-decoder architecture with 60 million parameters. * To manage large inputs that can reach up to 1 million tokens, the system reorders features by importance at the beginning of the string so that critical data is preserved when truncated to the model's 8k token limit. * Pre-training the RLM on diverse regression tasks enables few-shot adaptation, allowing the model to adjust to new data types with minimal gradient updates. * Numerical values are processed as-is within the text, removing the requirement for traditional scaling or normalization common in standard machine learning pipelines. ## Optimizing Google's Borg Infrastructure * The method was specifically applied to Google’s Borg system to predict MIPS per GCU (Millions of Instructions Per Second per Google Compute Unit), a vital efficiency metric. * The RLM simulates the outcomes of complex bin-packing algorithms within a "digital twin" framework to optimize resource allocation across CPUs and TPUs. * By analyzing execution traces and textual metadata, the model provides high-accuracy forecasting for diverse workloads including Gmail, YouTube, and Maps. ## Density Capture and Uncertainty Modeling * Unlike traditional regressors that provide a single point estimate, RLMs can capture full probability distributions by sampling the decoded output multiple times. * This density estimation is critical for modeling aleatoric uncertainty, which represents the inherent randomness and stochastic load demands of large-scale compute environments. * The ability to visualize these distributions helps engineers identify the range of possible outcomes and the inherent variability of the system's performance over time. This research demonstrates that small, specialized language models can effectively replace traditional regression methods in highly dynamic environments. For practitioners looking to implement these capabilities, the open-source `regress-lm` library provides a framework for simulating large systems and predicting performance across varied industrial and scientific use cases.

lineOriginal article

Introducing a case of utilizing DDD in (opens in new tab)

LY Corporation’s ABC Studio developed a specialized retail Merchant system by leveraging Domain-Driven Design (DDD) to overcome the functional limitations of a legacy food-delivery infrastructure. The project demonstrates that the primary value of DDD lies not just in technical implementation, but in aligning organizational structures and team responsibilities with domain boundaries. By focusing on the roles and responsibilities of the system rather than just the code, the team created a scalable platform capable of supporting diverse consumer interfaces. ### Redefining the Retail Domain * The legacy system treated retail items like restaurant entries, creating friction for specialized retail services; the new system was built to be a standalone platform. * The team narrowed the domain focus to five core areas: Shop, Item, Category, Inventory, and Order. * Sales-specific logic, such as coupons and promotions, was delegated to external "Consumer Platforms," allowing the Merchant system to serve as a high-performance information provider. ### Clean Architecture and Modular Composition * The system utilizes Clean Architecture to ensure domain entities remain independent of external frameworks, which also provided a manageable learning curve for new team members. * Services are split into two distinct modules: "API" modules for receiving external requests and "Engine" modules for processing business logic. * Communication between these modules is handled asynchronously via gRPC and Apache Kafka, using the Decaton library to increase throughput while maintaining a low partition count. * The architecture prioritizes eventual consistency, allowing for high responsiveness and scalability across the platform. ### Global Collaboration and Conway’s Law * Development was split between teams in Korea (Core Domain) and Japan (System Integration and BFF), requiring a shared understanding of domain boundaries. * Architectural Decision Records (ADR) were implemented to document critical decisions and prevent "knowledge drift" during long-term collaboration. * The organizational structure was intentionally designed to mirror the system architecture, with specific teams (Core, Link, BFF, and Merchant Link) assigned to distinct domain layers. * This alignment, reflecting Conway’s Law, ensures that changes to external consumer platforms have minimal impact on the stable core domain logic. Successful DDD adoption requires moving beyond technical patterns like hexagonal architecture and focusing on establishing a shared understanding of roles across the organization. By structuring teams to match domain boundaries, companies can build resilient systems where the core business logic remains protected even as the external service ecosystem evolves.

googleOriginal article

SensorLM: Learning the language of wearable sensors (opens in new tab)

SensorLM is a new family of foundation models designed to bridge the gap between high-dimensional wearable sensor data and natural language descriptions. By training on a massive dataset of nearly 60 million hours of de-identified health data, the models learn to interpret complex physiological signals to provide meaningful context for human activities. This research demonstrates that integrating multimodal sensor signals with language models enables sophisticated health insights, such as zero-shot activity recognition and automated health captioning, that significantly outperform general-purpose large language models. ## Dataset Scale and Automated Annotation * The models were pre-trained on an unprecedented 59.7 million hours of multimodal sensor data collected from over 103,000 individuals across 127 countries. * To overcome the high cost of manual annotation, researchers developed a hierarchical pipeline that automatically generates text descriptions by calculating statistics and identifying trends within the raw sensor streams. * Data was sourced from Fitbit and Pixel Watch devices, representing nearly 2.5 million person-days of activity and health information. ## Hybrid Training Architecture * SensorLM unifies two primary multimodal strategies: contrastive learning and generative pre-training. * Through contrastive learning, the model learns to discriminate between different states—such as a "light swim" versus a "strength workout"—by matching sensor segments to corresponding text descriptions. * The generative component allows the model to "speak" for the sensors, producing nuanced, context-aware natural language captions directly from high-dimensional biometric signals. ## Activity Recognition and Cross-Modal Capabilities * The model demonstrates state-of-the-art performance in zero-shot human activity recognition, accurately classifying 20 different activities without any specific fine-tuning. * Its few-shot learning capabilities allow the model to adapt to new tasks or individual user patterns with only a handful of examples. * SensorLM facilitates cross-modal retrieval, enabling users or experts to find specific sensor patterns using natural language queries or to generate descriptions based on specific sensor inputs. ## Generative Health Captioning * Beyond simple classification, the model can generate hierarchical captions that describe the statistical, structural, and semantic dimensions of a user’s data. * Experimental results using metrics like BERTScore show that SensorLM produces captions that are more factually correct and coherent than those created by powerful non-specialist LLMs. * This capability allows for the translation of abstract data points, such as heart rate variability or step counts, into readable summaries that explain the "why" behind physiological changes. By providing a framework where wearable data can be understood through the lens of human language, SensorLM paves the way for more intuitive and personalized health monitoring. This technology holds the potential to transform raw biometric streams into actionable insights, helping users better understand the relationship between their activities and their overall physical well-being.

figma2 min readCurated summary

Figma Announces Increase in IPO Price Range | Figma Blog

Figma announced an increased expected price range of **$30–$32 per share** for its proposed initial public offering. The company has amended its Form S-1 registration statement and plans to list Class A common stock on the New York Stock Exchange under **“FIG.”** The announcement follows the launch of Figma’s IPO roadshow. ## Updated IPO Terms - Figma plans to offer **12,472,657 shares** of Class A common stock. - Existing stockholders plan to sell an additional **24,464,423 shares**. - Selling stockholders may grant underwriters a 30-day option to purchase up to **5,540,561 additional shares** to cover over-allotments. - Figma will not receive proceeds from shares sold by existing stockholders. ## Underwriters - Joint lead book-running managers: - Morgan Stanley - Goldman Sachs - Allen & Company - J.P. Morgan - Additional book-running managers include BofA Securities, Wells Fargo Securities, and RBC Capital Markets. - William Blair and Wolfe | Nomura Alliance will serve as co-managers. ## Regulatory Status - The offering is being made through a preliminary prospectus. - Figma’s registration statement has been filed with the SEC but is not yet effective. - Shares cannot be sold, and purchase offers cannot be accepted, until the registration statement becomes effective and applicable securities-law requirements are satisfied. ## About Figma - Founded in 2012, Figma describes itself as an AI-powered platform for collaborative product development. - Its product supports the full workflow from ideation and design through development and shipping. - The company emphasizes collaboration, efficiency, and keeping product teams aligned. Figma’s revised price range signals increased expectations for its upcoming IPO, though the final offering price and completion of the listing remain subject to market conditions and regulatory approval.

Read original(opens in new tab)
figma2 min readCurated summary

The making of a product icon | Figma Blog

Figma product icons are the result of extensive exploration rather than a single inspired sketch. Designer Tim Van Damme combines consistent visual rules with product-specific research, then develops dozens or hundreds of variations before refining the final symbol. The process balances recognizability, scalability, and a distinct identity across Figma’s expanding product suite. ## Building a Consistent Icon System - Figma previously had no formal icon-design system, so Tim created unofficial guidelines to ensure the icons worked as a cohesive family. - Core principles include: - One-pixel-wide strokes - One-pixel-thick cutouts - Rounded caps - Balanced compositions that use the available space - Three standard sizes for different surfaces - Icons must work everywhere from toolbars to large marketing displays and billboards. - Designs are rejected when lines blur or visual elements become indistinguishable at smaller sizes. ## Finding the Product’s Core Idea - Each icon must follow Figma’s broader visual language while representing the identity of its specific product. - Tim collaborates with design and product teams to list concepts, themes, and associations connected to the product. - For Figma Community, ideas such as shared learning, connection, and knowledge exchange inspired experiments with people, trees, and books. - For Figma Buzz, themes including magic, creation, and AI image generation led to explorations involving glass orbs, stars, and bees. ## Exploring Through Extensive Iteration - Tim uses Figma’s variable-width strokes and simultaneous vector-layer editing to speed up experimentation. - He develops large families of related symbols, gradually adjusting shapes, proportions, movement, and composition. - Creative exploration can transform one idea into another—for example, trees into faces, faces into mandalas, or abstract shapes into interlocking forms. - For Figma Make, he explored motion and transformation through wheels, butterflies, compasses, and fidget spinners. - For Figma Buzz, he created hundreds of bee variations, changing wings, antennas, and line treatments before narrowing the options. The practical lesson is that effective product icons emerge from a structured but highly exploratory process: establish shared rules, identify the product’s essential meaning, and iterate freely until the symbol is both distinctive and immediately legible.

Read original(opens in new tab)
googleOriginal article

Synthetic and federated: Privacy-preserving domain adaptation with LLMs for mobile applications (opens in new tab)

Researchers at Google have developed a framework for improving both small and large language models (LMs) in mobile applications like Gboard by utilizing privacy-preserving synthetic data and federated learning. This approach combines differential privacy (DP) with large language model (LLM) generation to minimize data memorization risks while achieving significant gains in production metrics like next-word prediction and proofreading. The result is a robust pipeline that allows models to adapt to specific user domains without compromising individual privacy or requiring centralized data storage. ### Strengthening Privacy with DP-FL * Gboard has transitioned all production LMs trained on user data to a Federated Learning with Differential Privacy (DP-FL) framework, ensuring data remains on-device and is never memorized. * The deployment utilizes the **BLT-DP-FTRL** algorithm, which offers an optimized trade-off between privacy guarantees and model utility while being easier to deploy in production. * Engineers adopted the **SI-CIFG** model architecture to facilitate efficient on-device training, ensuring the hardware can handle local updates while maintaining compatibility with DP constraints. ### Synthetic Data Generation via Public LLMs * Powerful LLMs trained on public web data are prompted to synthesize high-quality text that mimics mobile user interactions without ever accessing actual private user data. * The process involves a two-step prompting strategy: first, filtering public datasets to identify topics common in mobile communication, and second, generating new, domain-specific text based on those patterns. * This synthetic data serves as a bridge for pre-training small LMs, which are then refined through private post-training on-device to capture the nuances of user behavior. ### Adapting LLMs for Mobile Proofreading * To support advanced features like Gboard's "Proofread," researchers developed a "Synthesize-then-Adapt" pipeline specifically for error correction. * LLMs generate synthetic "corrupted" text to simulate common mobile typing errors, providing the necessary training pairs (error/correction) that are difficult to find in public datasets. * Federated learning is then used to adapt these error-correction models to specific app domains (such as messaging or email) using on-device signals, ensuring the model understands the specific context of the user's typing. The success of these techniques in Gboard demonstrates that synthetic data can effectively replace or augment private data throughout the machine learning lifecycle. For developers working with sensitive user information, adopting a "synthetic-first" approach combined with federated learning provides a scalable path to model improvement that adheres to the core principles of data minimization and anonymization.

figma3 min readCurated summary

Figma Make Is Now Available to All Users | Figma Blog

Figma has moved Figma Make and other AI features out of beta and made them available to all users. Figma Make is positioned as a prompt-to-app tool that lets people create interactive, high-fidelity prototypes without extensive technical skills. The company argues that this speeds up exploration and alignment while enabling designers, engineers, researchers, and product managers to participate more directly in product development. ## Figma Make Becomes Generally Available - Users can create prototypes and web apps through natural-language prompts. - New capabilities include: - Importing an existing Figma library, including colors, typography, styling, and usage guidelines. - Connecting prototypes to backend data through a Supabase integration. - Figma’s other AI features, including Make and Edit Image and Boost Resolution, are also leaving beta. - Publishing Figma Make files remains in beta. ## Availability, Seats, and AI Credits - **Full seat users** receive the full Figma Make and AI experience, including private sharing and publishing capabilities. - **View, Collab, and Dev seat users** can create unlimited Make files in drafts and use the AI features available to their seats. - **Starter plan users** can create unlimited draft files and share up to three Make files with their team. - AI credits are included with all seats. - Full-seat monthly allocations are: - Professional: 3,000 credits, estimated at 50–70 Make prompts. - Organization: 3,500 credits, estimated at 60–80 prompts. - Enterprise: 4,250 credits, estimated at 80–100 prompts. - Figma plans to offer additional credits for purchase later. Full-seat limits will not initially be strictly enforced, while lower limits for Starter, View, Collab, and Dev users are enforced daily and monthly. ## Faster, More Collaborative Product Development - Figma says Make supports the broader development process, not just prototyping. - Designers can test complex interactions. - Engineers can produce high-fidelity design artifacts. - Product managers can build functional mockups to validate ideas. - The tool encourages teams to experiment, take risks, and discuss concrete artifacts rather than abstract descriptions. ## Making Prototyping More Accessible - Figma argues that language alone often fails to communicate how a product should look and feel. - Traditional prototypes can be time-consuming or inaccessible to people without design or development expertise. - Natural-language prompting allows more team members to turn ideas into realistic prototypes quickly. - Figma researcher Rie McGwier used Make to build a Survey Traffic Calculator for managing roughly 30 surveys across eight products. - The prototype modeled weekly active users, audience overlap, response rates, recontact windows, sampling requirements, and other constraints, producing a dashboard to help avoid oversampling. Figma’s recommendation is to use Make as a collaborative exploration tool: teams can rapidly turn ideas into working artifacts, test assumptions, and refine solutions before committing to full implementation.

Read original(opens in new tab)
lineOriginal article

Milvus: Building a Large-Scale (opens in new tab)

LINE VOOM transitioned its recommendation system from a batch-based offline process to a real-time infrastructure to solve critical content freshness issues. By adopting Milvus, an open-source vector database, the team enabled the immediate indexing and searching of new video content as soon as it is uploaded. This implementation ensures that time-sensitive posts are recommended to users without the previous 24-hour delay, significantly enhancing user engagement. ### Limitations of the Legacy Recommendation System * The original system relied on daily offline batch processing for embedding generation and similarity searches. * New content, such as holiday greetings or trending sports clips, suffered from a "lack of immediacy," often taking up to a full day to appear in user feeds. * To improve user experience, the team needed to shift from offline candidate pools to an online system capable of real-time Approximate Nearest Neighbor (ANN) searches. ### Selecting Milvus as the Vector Database * The team evaluated Milvus and Qdrant based on performance, open-source status, and on-premise compatibility. * Milvus was selected due to its superior performance, handling 2,406 requests per second compared to Qdrant's 326, with lower query latency (1ms vs 4ms). * Key architectural advantages of Milvus included the separation of storage and computing, support for both stream and batch inserts, and a diverse range of supported in-memory index types. ### Reliability Verification via Chaos Testing * Given the complexity of Milvus clusters, the team performed chaos testing by intentionally injecting failures like pod kills and scaling events. * Tests revealed critical vulnerabilities: killing the `Querycoord` led to collection release and search failure, while losing the `Etcd` quorum caused total metadata loss. * These findings highlighted the need for robust high-availability (HA) configurations to prevent service interruptions during component failures. ### High Availability (HA) Implementation Strategies * **Collection-Level HA:** To prevent search failures during coordinator issues, the team implemented a dual-writing system where embeddings are recorded in two separate collections simultaneously. * **Alias Switching:** Client applications use an "alias" to reference collections; if the primary collection becomes unavailable, the system instantly switches the alias to the backup collection to minimize downtime. * **Coordinator-Level HA:** To eliminate single points of failure, coordinators (such as `Indexcoord`) were configured in an Active-Standby mode, ensuring a backup is always ready to take over management tasks. To successfully deploy a large-scale real-time recommendation engine, it is critical to select a vector database that decouples storage from compute and to implement multi-layered high-availability strategies, such as dual-collection writing and active-standby coordinators, to ensure production stability.

googleOriginal article

LSM-2: Learning from incomplete wearable sensor data (opens in new tab)

LSM-2 introduces a paradigm shift in processing wearable sensor data by treating naturally occurring data gaps as inherent features rather than errors to be corrected. By utilizing the Adaptive and Inherited Masking (AIM) framework, the model learns directly from fragmented, real-world data streams without the need for biased imputation or data-discarding filters. This approach allows LSM-2 to achieve state-of-the-art performance in health-related classification and regression tasks, maintaining robustness even when sensors fail or data is highly interrupted. ## The Challenge of Pervasive Missingness * Real-world wearable data is almost never continuous; factors such as device charging, motion artifacts, and battery-saving modes create frequent "missingness." * Traditional self-supervised learning models require complete data, forcing researchers to use imputation—which can introduce artificial bias—or aggressive filtering that discards over 90% of potentially useful samples. * In a dataset of 1.6 million day-long windows, research found that not a single sample had 0% missingness, highlighting the impracticality of training only on complete datasets. ## Adaptive and Inherited Masking (AIM) * AIM extends the Masked Autoencoder (MAE) framework by treating "inherited" masks (naturally occurring gaps) and "artificial" masks (training objectives) as equivalent. * The framework utilizes a dual masking strategy: it employs token dropout on a fixed ratio of tokens to ensure computational efficiency during encoding. * To handle the unpredictable and variable nature of real-world gaps, AIM uses attention masking within the transformer blocks for any remaining masked tokens. * During evaluation and fine-tuning, the model relies solely on attention masking to navigate naturally occurring gaps, allowing for accurate physiological modeling without filling in missing values. ## Scale and Training Architecture * LSM-2 was trained on a massive dataset comprising 40 million hours of de-identified wearable data from more than 60,000 participants using Fitbit and Google Pixel devices. * The model learns to understand underlying physiological structures by reconstructing masked segments across multimodal inputs, including heart signals, sleep patterns, and activity levels. * Because it is trained on fragmented data, the resulting foundation model is significantly more resilient to sensor dropouts in downstream tasks like hypertension prediction or stress monitoring. LSM-2 demonstrates that foundation models for health should be built to embrace the messiness of real-world environments. By integrating missingness directly into the self-supervised learning objective, developers can bypass the computational and statistical overhead of imputation while building more reliable diagnostic and monitoring tools.

figma2 min readCurated summary

Launching the roadshow for Figma’s proposed IPO | Figma Blog

Figma announced the launch of its roadshow for a proposed initial public offering of Class A common stock. The offering includes new shares from Figma and shares sold by existing stockholders, with an expected price of $25–$28 per share. Figma plans to list on the New York Stock Exchange under “FIG,” pending regulatory effectiveness. ## Proposed Share Offering - Figma will offer 12,472,657 shares of Class A common stock. - Existing stockholders will offer 24,646,423 additional shares. - Selling stockholders may grant underwriters a 30-day option to purchase up to 5,540,561 more shares for over-allotments. - Figma will not receive proceeds from shares sold by existing stockholders. - The expected IPO price is $25–$28 per share. ## Listing and Underwriters - Figma has applied to list its Class A common stock on the New York Stock Exchange under the ticker symbol **FIG**. - Morgan Stanley, Goldman Sachs, Allen & Company, and J.P. Morgan are joint lead book-running managers. - BofA Securities, Wells Fargo Securities, and RBC Capital Markets are book-running managers. - William Blair and Wolfe | Nomura Alliance are serving as co-managers. ## Regulatory Status - Figma’s registration statement has been filed with the SEC but is not yet effective. - The shares cannot be sold, and purchase offers cannot be accepted, until the registration process is complete. - The offering will be made only through a prospectus. ## Figma’s Business - Founded in 2012, Figma describes itself as a collaborative platform for digital product development. - It has expanded from a design tool into a connected, AI-powered platform supporting ideation, design, development, and product delivery. Figma’s announcement marks a formal step toward becoming a public company, but the IPO remains subject to SEC approval and final offering conditions.

Read original(opens in new tab)
figma2 min readCurated summary

Figma Deepens Roots in Australia with Local Data Hosting | Figma Blog

Figma is expanding its investment in Australia by introducing enterprise governance features and local hosting for Figma file data. Starting in Q4 2025, Australian customers will be able to store data locally, supporting organizations with strict security and compliance requirements. The move strengthens Figma’s position among regulated industries and marks its first data-residency offering in Asia Pacific. ## Local Data Hosting in Australia - Figma will host file data locally in Australia, including content from: - Figma - FigJam - Make - Sites - Buzz - Slides - Local hosting is intended for industries such as: - Government and the public sector - Healthcare - Financial services - The option provides greater control over data location while preserving Figma’s platform capabilities and scalability. - Australia is Figma’s first local data-hosting market in Asia Pacific, extending similar enterprise offerings already available in Europe and the United States. - Figma opened its Sydney office in November 2024 and serves customers including NAB, Safety Culture, and Atlassian. ## Governance+ for Enterprise Customers Governance+ gives enterprises more control over how employees access and use Figma. - **Centralized controls** - Enforce use of approved Figma instances and networks. - Use IP Allowlisting and Network Access Restrictions to prevent data from moving into unauthorized spaces. - **Account security** - Require two-factor authentication. - Extend idle session timeouts. - Support for multiple SSO configurations is planned. - **Data governance** - Monitor Figma activity through tools such as the Discovery Pipeline. - Support electronic communications retention and legal discovery requirements. ## Existing Enterprise Security Features Governance+ builds on existing enterprise capabilities, including: - Action logs - SAML single sign-on - Role assignments connected to identity-management systems - Restrictions on external collaborators joining an organization Governance+ is available now to customers on Figma’s Enterprise plan. Figma’s Australian data residency option will be particularly useful for organizations that must meet local storage, privacy, and regulatory obligations. Enterprise customers can adopt Governance+ immediately and register interest in local hosting ahead of its planned Q4 2025 launch.

Read original(opens in new tab)