Techlist.io - Korean Tech Blog Curator

googleOriginal article

Teaching Gemini to spot exploding stars with just a few examples (opens in new tab)

Researchers have demonstrated that Google’s Gemini model can classify cosmic events with 93% accuracy, rivaling specialized machine learning models while providing human-readable explanations. By utilizing few-shot learning with only 15 examples per survey, the model addresses the "black box" limitation of traditional convolutional neural networks used in astronomy. This approach enables scientists to efficiently process the millions of alerts generated by modern telescopes while maintaining a transparent and interactive reasoning process. ## Bottlenecks in Modern Transient Astronomy * Telescopes like the Vera C. Rubin Observatory are expected to generate up to 10 million alerts per night, making manual verification impossible. * The vast majority of these alerts are "bogus" signals caused by satellite trails, cosmic rays, or instrumental artifacts rather than real supernovae. * Existing specialized models often provide binary "real" or "bogus" labels without context, forcing astronomers to either blindly trust the output or spend hours on manual verification. ## Multimodal Few-Shot Learning for Classification * The research utilized few-shot learning, providing Gemini with only 15 annotated examples for three major surveys: Pan-STARRS, MeerLICHT, and ATLAS. * Input data consisted of image triplets—a "new" alert image, a "reference" image of the same sky patch, and a "difference" image—each 100x100 pixels in size. * The model successfully generalized across different telescopes with varying pixel scales, ranging from 0.25" per pixel for Pan-STARRS to 1.8" per pixel for ATLAS. * Beyond simple labels, Gemini generates a textual description of observed features and an interest score to help astronomers prioritize follow-up observations. ## Expert Validation and Self-Assessment * A panel of 12 professional astronomers evaluated the model using a 0–5 coherence rubric, confirming that Gemini’s logic aligned with expert reasoning. * The study found that Gemini can effectively assess its own uncertainty; low self-assigned "coherence scores" were strong indicators of likely classification errors. * This ability to flag its own potential mistakes allows the model to act as a reliable partner, alerting scientists when a specific case requires human intervention. The transition from "black box" classifiers to interpretable AI assistants allows the astronomical community to scale with the data flood of next-generation telescopes. By combining high-accuracy classification with transparent reasoning, researchers can maintain scientific rigor while processing millions of cosmic events in real time.

lineOriginal article

Essential Element for App Success: Error Monitoring (opens in new tab)

Effective mobile app management requires proactive outage monitoring to prevent user churn caused by failures in critical flows like registration or payment. Relying on user reports is often too late, so developers must implement systematic event collection and real-time dashboards to identify issues the moment they arise. By integrating tools like Sentry or Firebase, teams can maintain high quality through immediate response and detailed performance analysis. ### Implementing Sentry in Flutter * **Dependency and Initialization**: Integration begins by adding `sentry_flutter` and `sentry_dio` to the project. The initialization process involves setting the Data Source Name (DSN), environment tags (e.g., production vs. staging), and release versions to ensure logs are correctly categorized. * **Performance and Privacy**: Developers should configure `tracesSampleRate` and `profilesSampleRate` to balance monitoring depth with costs. Additionally, the `beforeSend` callback allows for masking sensitive user data like authorization headers or IP addresses before they are transmitted. * **Contextual Tracking**: To aid debugging, the system captures user IDs via `Sentry.configureScope` and tracks user movement using `SentryNavigatorObserver`. Utilizing `SentryInterceptor` with the Dio library allows for automatic tracking of HTTP request performance and API bottlenecks. ### Strategic Log Level Design * **Debug and Info**: Debug logs remain local to the terminal to save resources. Info logs are reserved for significant user actions that change data, such as successful sign-ups or purchases, while high-frequency read actions like "viewing a product list" are excluded to reduce noise and costs. * **Warning**: This level tracks external system failures, such as failed API calls or push notification losses. To prevent "alert fatigue," client-side network issues (e.g., timeouts or offline status) are ignored, and alerts are triggered only when specific thresholds are met, such as 100 failures within 10 minutes. * **Error**: Error logs represent internal logic failures that bypass defensive coding, such as null object errors, parsing failures, or unreachable code branches. These require immediate notification to the development team to facilitate rapid hotfixes. * **Fatal**: This level is dedicated to application crashes and unhandled exceptions. When configured at the app's entry point, the system automatically captures these critical failures to provide a comprehensive "crash-free users" metric. ### Creating Effective Dashboards * **Naming Conventions**: Logs should follow a strict structure, using tags for modules and event names (e.g., `[API] [postLogin] success`). This consistency allows for granular querying and clearer visualization on monitoring dashboards. * **Data Enrichment**: Using the `extra` field in log events provides vital context for troubleshooting, such as including the specific endpoint, request body, and response status code for a failed transaction. * **Actionable Metrics**: Effective monitoring focuses on key performance indicators like API error rates and the failure percentage of core business events (login, registration, payment) rather than just raw crash counts. A robust monitoring strategy shifts the focus from simple crash reporting to comprehensive service health. By standardizing log levels and automating event collection, development teams can distinguish between transient network blips and critical logic errors, ensuring they spend their time fixing high-impact issues.

googleOriginal article

Solving virtual machine puzzles: How AI is optimizing cloud computing (opens in new tab)

Google researchers have developed LAVA, a scheduling framework designed to optimize virtual machine (VM) allocation in large-scale data centers by accurately predicting and adapting to VM lifespans. By moving beyond static, one-time predictions toward a "continuous re-prediction" model based on survival analysis, the system significantly improves resource efficiency and reduces fragmentation. This approach allows cloud providers to solve the complex "bin packing" problem more effectively, leading to better capacity utilization and easier system maintenance. ### The Challenge of Long-Tailed VM Distributions * Cloud workloads exhibit a extreme long-tailed distribution: while 88% of VMs live for less than an hour, these short-lived jobs consume only 2% of total resources. * The rare VMs that run for 30 days or longer account for a massive fraction of compute resources, meaning their placement has a disproportionate impact on host availability. * Poor allocation leads to "resource stranding," where a server's remaining capacity is too small or unbalanced to host new VMs, effectively wasting expensive hardware. * Traditional machine learning models that provide only a single prediction at VM creation are often fragile, as a single misprediction can block a physical host from being cleared for maintenance or new tasks. ### Continuous Re-prediction via Survival Analysis * Instead of predicting a single average lifetime, LAVA uses an ML model to generate a probability distribution of a VM's expected duration. * The system employs "continuous re-prediction," asking how much longer a VM is expected to run given how long it has already survived (e.g., a VM that has run for five days is assigned a different remaining lifespan than a brand-new one). * This adaptive approach allows the scheduling logic to automatically correct for initial mispredictions as more data about the VM's actual behavior becomes available over time. ### Novel Scheduling and Rescheduling Algorithms * **Non-Invasive Lifetime Aware Scheduling (NILAS):** Currently deployed on Google’s Borg cluster manager, this algorithm ranks potential hosts by grouping VMs with similar expected exit times to increase the frequency of "empty hosts" available for maintenance. * **Lifetime-Aware VM Allocation (LAVA):** This algorithm fills resource gaps on hosts containing long-lived VMs with jobs that are at least an order of magnitude shorter. This ensures the short-lived VMs exit quickly without extending the host's overall occupation time. * **Lifetime-Aware Rescheduling (LARS):** To minimize disruptions during defragmentation, LARS identifies and migrates the longest-lived VMs first while allowing short-lived VMs to finish their tasks naturally on the original host. By integrating survival-analysis-based predictions into the core logic of data center management, cloud providers can transition from reactive scheduling to a proactive model. This system not only maximizes resource density but also ensures that the physical infrastructure remains flexible enough to handle large, resource-intensive provisioning requests and essential system updates.

googleOriginal article

Using AI to identify genetic variants in tumors with DeepSomatic (opens in new tab)

DeepSomatic is an AI-powered tool developed by Google Research to identify cancer-related mutations by analyzing a tumor's genetic sequence with higher accuracy than current methods. By leveraging convolutional neural networks (CNNs), the model distinguishes between inherited genetic traits and acquired somatic variants that drive cancer progression. This flexible tool supports multiple sequencing platforms and sample types, offering a critical resource for clinicians and researchers aiming to personalize cancer treatment through precision medicine. ## Challenges in Somatic Variant Detection * Somatic variants are genetic mutations acquired after birth through environmental exposure or DNA replication errors, making them distinct from the germline variants found in every cell of a person's body. * Detecting these mutations is technically difficult because tumor samples are often heterogeneous, containing a diverse set of variants at varying frequencies. * Sequencing technologies often introduce small errors that can be difficult to distinguish from actual somatic mutations, especially when the mutation is only present in a small fraction of the sampled cells. ## CNN-Based Variant Calling Architecture * DeepSomatic employs a method pioneered by DeepVariant, which involves transforming raw genetic sequencing data into a set of multi-channel images. * These images represent various data points, including alignment along the chromosome, the quality of the sequence output, and other technical variables. * The convolutional neural network processes these images to differentiate between three categories: the human reference genome, non-cancerous germline variants, and the somatic mutations driving tumor growth. * By analyzing tumor and non-cancerous cells side-by-side, the model effectively filters out sequencing artifacts that might otherwise be misidentified as mutations. ## System Versatility and Application * The model is designed to function in multiple modes, including "tumor-normal" (comparing a biopsy to a healthy sample) and "tumor-only" mode, which is vital for blood cancers like leukemia where isolating healthy cells is difficult. * DeepSomatic is platform-agnostic, meaning it can process data from all major sequencing technologies and adapt to different types of sample processing. * The tool has demonstrated the ability to generalize its learning to various cancer types, even those not specifically included in its initial training sets. ## Open-Source Contributions to Precision Medicine * Google has made the DeepSomatic tool and the CASTLE dataset—a high-quality training and evaluation set—openly available to the global research community. * This initiative is part of a broader effort to use AI for early detection and advanced research in various cancers, including breast, lung, and gynecological cancers. * The release aims to accelerate the development of personalized treatment plans by providing a more reliable way to identify the specific genetic drivers of an individual's disease. By providing a more accurate and adaptable method for variant calling, DeepSomatic helps researchers pinpoint the specific drivers of a patient's cancer. This tool represents a significant advancement in deep learning for genomics, potentially shortening the path from biopsy to targeted therapeutic intervention.

datadogOriginal article

Failure is inevitable: Learning from a large outage, and building for reliability in depth at Datadog | Datadog (opens in new tab)

Following a major 2023 incident that caused a near-total platform outage despite partial infrastructure availability, Datadog shifted its engineering philosophy from "never-fail" architectures to a model of graceful degradation. The company identified that prioritizing absolute data correctness during systemic stress created "square-wave" failures, where the entire platform appeared down if even a portion of data was missing. By moving toward a "fail better" mindset, Datadog now focuses on maintaining core functionality and data persistence even when underlying infrastructure is compromised. ## Limitations of the Never-Fail Approach * Classical root-cause analysis focused on a legacy, unsupervised global update mechanism that disconnected 50–60% of production Kubernetes nodes. * While the "precipitating event" was easily identified and disabled, the engineering team realized that fixing the trigger did not address the systemic fragility that caused a binary (up/down) failure pattern. * Prioritizing absolute accuracy meant that systems would wait for all data tags to process before displaying results; under stress, this caused the UI to show no data at all rather than "almost correct" data. * Sequential queuing, aggressive retry logic, and node-specific processing requirements exacerbated the bottleneck, preventing real-time recovery. ## Prioritizing Graceful Degradation * The incident prompted a shift away from relying solely on redundancy to prevent outages, acknowledging that some level of failure is eventually inevitable at scale. * Engineering priorities were redefined to ensure that data is never lost (even if delayed) and that real-time data is processed before stale backlogs. * The platform now aims to serve partial-but-accurate results to customers during an incident, providing visibility rather than a complete blackout. * Implementation is handled as a company-wide program where individual product teams adapt these principles to their specific architectural needs. ## Strengthening Data Persistence at Intake * Analysis revealed that data was lost during the outage because it was stored in memory or on local disks before being replicated to persistent stores. * The original design favored low-latency responses by acknowledging receipt of data before it was fully replicated, making that data unrecoverable if the node failed. * Downstream failures caused intake nodes to overflow their local buffers, leading to data loss even on nodes that remained online. * New architectural changes focus on implementing disk-based persistence at the very beginning of the processing pipeline to ensure data survives node restarts and downstream congestion. To build truly resilient systems, engineering teams must move beyond trying to prevent every possible failure trigger. Instead, focus on designing services that can survive partial infrastructure loss by prioritizing data persistence and allowing for degraded states that still provide value to the end user.

datadog3 min readCurated summary

Failure is inevitable: Learning from a large outage, and building for reliability in depth at Datadog

Datadog’s March 2023 outage exposed a fundamental weakness in its reliability strategy: although 40–50% of production Kubernetes nodes remained operational, customers experienced the platform as entirely unavailable. The incident showed that preventing every failure is impossible and that systems must instead continue delivering useful, accurate service when components fail. Datadog consequently began redesigning products around graceful degradation, prioritizing data preservation, fresh information, and partial results. ## Lessons from the March 2023 Incident - An unsupervised global update triggered a restart interaction that disconnected roughly 50–60% of production Kubernetes nodes. - The web interface recovered quickly, but logs, metrics, alerts, traces, and other core features became unavailable. - Pages loaded without displaying customer data, creating a nearly complete outage from the user’s perspective. ## Limits of Traditional Root-Cause Analysis - Datadog identified the legacy global security-update mechanism as the immediate trigger and disabled it. - Fixing that mechanism alone could not address the broader class of failures caused by certificates, configuration changes, overloads, date-handling bugs, or other unexpected events. - The company concluded that resilience requires reducing the impact of failures, not merely preventing one specific failure mode. ## Why Partial Infrastructure Became a Total User-Facing Failure - Datadog’s systems historically favored complete correctness over partial visibility. - For example, metric queries could wait until all relevant tags were processed to avoid showing misleading values or triggering false alerts. - During a large outage, this behavior created a “square-wave” failure: missing some data caused the system to show no data. - Ordered queues could stall fresh results behind stuck work, retries could overload already-strained services, and node-specific processing could make surviving capacity ineffective. - The underlying design assumption was that systems should either function fully or stop, rather than degrade while continuing to provide value. ## Prioritizing Graceful Degradation Datadog shifted from relying primarily on redundancy and “never-fail” architectures to explicitly designing for inevitable failures. - Customer data should never be lost, even if delivery is delayed. - Fresh, real-time data should take priority over stale backlog processing. - Systems should provide partial but accurate results whenever possible instead of returning nothing. ## Persistent Storage at the Start of Processing Pipelines - The outage caused a limited but non-zero amount of irreversible customer data loss. - Some pipelines acknowledged data before writing it to replicated storage, leaving unreplicated data only in memory or on a local disk. - When a node failed, that data disappeared and could not be recovered through agent retries. - After the node loss, surviving intake nodes also struggled to write to downstream replicated stores. - Their memory and local-disk buffers eventually filled, causing additional data loss as the outage continued. - Datadog therefore identified persistent intake storage as a key requirement for preserving data during large-scale failures. The broader recommendation is to design systems not only to prevent outages, but also to remain useful during them: preserve every accepted event, prioritize current information, and expose accurate partial results instead of failing completely.

Read original(opens in new tab)
googleOriginal article

Coral NPU: A full-stack platform for Edge AI (opens in new tab)

Coral NPU is a new full-stack, open-source platform designed to bring advanced AI directly to power-constrained edge devices and wearables. By prioritizing a matrix-first hardware architecture and a unified software stack, Google aims to overcome traditional bottlenecks in performance, ecosystem fragmentation, and data privacy. The platform enables always-on, low-power ambient sensing while providing developers with a flexible, RISC-V-based environment for deploying modern machine learning models. ## Overcoming Edge AI Constraints * The platform addresses the "performance gap" where complex ML models typically exceed the power, thermal, and memory budgets of battery-operated devices. * It eliminates the "fragmentation tax" by providing a unified architecture, moving away from proprietary processors that require costly, device-specific optimizations. * On-device processing ensures a high standard of privacy and security by keeping personal context and data off the cloud. ## AI-First Hardware Architecture * Unlike traditional chips, this architecture prioritizes the ML matrix engine over scalar compute to optimize for efficient on-device inference. * The design is built on RISC-V ISA compliant architectural IP blocks, offering an open and extensible reference for system-on-chip (SoC) designers. * The base design delivers performance in the 512 giga operations per second (GOPS) range while consuming only a few milliwatts of power. * The architecture is tailored for "always-on" use cases, making it ideal for hearables, AR glasses, and smartwatches. ## Core Architectural Components * **Scalar Core:** A lightweight, C-programmable RISC-V frontend that manages data flow using an ultra-low-power "run-to-completion" model. * **Vector Execution Unit:** A SIMD co-processor compliant with the RISC-V Vector instruction set (RVV) v1.0 for simultaneous operations on large datasets. * **Matrix Execution Unit:** A specialized engine using quantized outer product multiply-accumulate (MAC) operations to accelerate fundamental neural network tasks. ## Unified Developer Ecosystem * The platform is a C-programmable target that integrates with modern compilers such as IREE and TFLM (TensorFlow Lite Micro). * It supports a wide range of popular ML frameworks, including TensorFlow, JAX, and PyTorch. * The software toolchain utilizes MLIR and the StableHLO dialect to facilitate the transition from high-level models to hardware-executable code. * Developers have access to a complete suite of tools, including a simulator, custom kernels, and a general-purpose MLIR compiler. SoC designers and ML developers looking to build the next generation of wearables should leverage the Coral NPU reference architecture to balance high-performance AI with extreme power efficiency. By utilizing the open-source documentation and RISC-V-based tools, teams can significantly reduce the complexity of deploying private, always-on ambient sensing.

figma2 min readCurated summary

15+ Ways We're Improving Accessibility in Figma | Figma Blog

Figma is rolling out more than 15 accessibility improvements to make its products easier to use with keyboards, screen readers, and enhanced contrast settings. The updates expand keyboard control across Figma Design, FigJam, Slides, Buzz, commenting, and Dev Mode, while improving screen-reader navigation and object descriptions. Together, they aim to make collaboration, canvas editing, and handoff more reliable for people with different access needs. ## Expanded Keyboard Controls - Users can navigate and manipulate more canvas objects without a mouse. - **Figma Design:** Add and edit lines, adjust ruler guides, and create or edit arcs from ellipses. - **FigJam:** Manage table rows and columns, adjust stamps, votes, washi tape, marker lines, and highlighter strokes, and navigate embedded content and other canvas objects. - **Figma Slides:** Resize presenter notes and adjust writing tone with AI. - **Across products:** Open and move between links in edit or view-only mode. ## Keyboard Support for Comments and Collaboration - Add, move, and navigate comments across Figma products. - Navigate and manage Dev Mode annotations with shortcuts. - Move between discussions without losing keyboard focus. - New personalization toggles let users disable Figma-exclusive shortcuts while typing and choose whether Spotlight automatically follows other users. ## Improved Screen Reader Support - Tab navigation through buttons, menus, panels, and other actions now follows a more logical order. - Users can jump directly to specific actions, such as opening menus or activating toolbar controls. - Object announcements include details such as type, name, and state. - More consistent announcements help users detect new comments and file updates. - Screen readers preserve rich-text meaning, including bold, italics, lists, and links. - Canvas objects in Buzz and Slides can now be recognized and announced. ## Enhanced Color Contrast - A new setting increases contrast between text, interface elements, and backgrounds in both light and dark modes. - The option can be enabled through Accessibility settings, the Actions menu, or General settings. - Stronger contrast improves text and icon legibility, clarifies interface structure, and makes buttons and outlines easier to identify. - It can also improve visibility in glare, sunlight, and prolonged or multitasked screen use. Figma’s updates make accessibility a broader part of everyday editing and collaboration rather than a separate workflow. Users who rely on keyboards or screen readers should explore the new controls, while all users may benefit from enabling enhanced contrast when working in difficult lighting or for extended periods.

Read original(opens in new tab)
discord3 min readCurated summary

Staff Picks, September 2025: Welcome to Our Video Game Museum

The Discord staff’s September 2025 feature celebrates National Video Games Day by asking employees which games deserve a place in their personal “video game history museum.” Their answers emphasize nostalgia, formative experiences, influential design, and the communities built around gaming. Recent favorites include *Hollow Knight*, *Balatro*, *Final Fantasy XIV*, *Oblivion*, and *The Legend of Heroes: Trails of Cold Steel III*. ## Games That Shaped Personal Gaming Histories - **Super Mario 64** - Became a formative experience through childhood play with a neighbor. - Its 3D graphics, controls, music, level structure, and iconic painting-based worlds made a lasting impression. - It helped inspire a later favorite, *Super Mario Sunshine*. - **The Legend of Zelda: Ocarina of Time** - Described as a lifelong comfort game and the contributor’s definitive Zelda experience. - It has been completed 100% multiple times and remains deeply tied to personal memories. - The proposed exhibit would include the original cartridge and a heavily worn official strategy guide. - **Kingdom Hearts** - Represents powerful childhood nostalgia through its combination of Disney and *Final Fantasy*. - Obtaining the PlayStation 2 game and its strategy guide led to hours of play. - Although *Kingdom Hearts II* is the contributor’s favorite entry, the original is valued as the groundbreaking starting point of the series. - **The Legend of Heroes: Trails of Cold Steel III** - Highlights a sprawling JRPG series with interconnected arcs, countries, and characters. - Its strengths include extensive world-building, large character casts, and constantly updated dialogue for even minor NPCs. - The series’ eventual crossovers are compared to Avengers-style ensemble moments. - **Professor Layton and the Curious Village** - A last-minute purchase made because a strict birthday budget ruled out a new Pokémon game. - Its challenging puzzles and mystery narrative introduced the contributor to the appeal of story-driven games. - The experience encouraged a broader willingness to discover unexpected favorites. ## Games Staff Are Playing Recently - *Hollow Knight* is being replayed ahead of *Silksong*, with one contributor nearing 100% completion and another finally progressing after initially abandoning it twice. - *Balatro* has become highly addictive, with one player only one joker away from completing the collection. - *Final Fantasy XIV* has regained momentum after a partner began playing. - The *Elder Scrolls Oblivion* remake and *Lost Soul Aside* are also mentioned as current or anticipated games. - The contributors’ comments repeatedly connect current playtime to anticipation for *Silksong*. ## Gaming as Nostalgia and Community - The feature presents games as personal cultural artifacts rather than merely entertainment. - Physical items such as cartridges and strategy guides carry emotional value alongside the games themselves. - Several responses stress how games introduce players to new genres, stories, friendships, and communities. - The playful editorial notes about disappearing staff members after *Silksong*’s release reinforce the article’s informal, gaming-community tone. The article’s central recommendation is simple: revisit the games that shaped you, but also stay open to unexpected discoveries—whether through a discounted purchase, a friend’s recommendation, or a long-awaited sequel.

Read original(opens in new tab)
airbnb3 min readCurated summary

From Static Rate Limiting to Adaptive Traffic Management in Airbnb’s Key-Value Store

Airbnb evolved Mussel’s QoS system from static, per-client QPS limits into adaptive traffic management designed to maximize goodput. The newer approach accounts for the actual cost of requests, prioritizes critical workloads under stress, and detects hot keys or attack traffic before they overwhelm storage. Together, resource-aware quotas and real-time load shedding provide stronger protection against traffic spikes, uneven workloads, and DDoS-like bursts. ## Why Static QPS Limits Fell Short - Mussel is a multi-tenant key-value store serving millions of point and range reads across Airbnb. - Its original Redis-backed limiter assigned each client a fixed requests-per-second quota. - Requests exceeding the quota received HTTP 429 responses. - This model worked when backend effort roughly matched request count. - As usage grew, it could not account for: - The difference between a cheap one-row lookup and a 100,000-row scan. - Hot keys accessed by many clients simultaneously. - Localized storage-shard overload that affected unrelated traffic. - Sudden events such as bot floods, DDoS attacks, or large uploads. ## Resource-Aware Rate Control - Mussel replaced raw request counting with request units (RU), which represent estimated backend work. - RU calculations incorporate: - Fixed per-request overhead. - Rows and payload bytes processed. - Request latency, which distinguishes cached operations from disk-heavy ones. - The system uses calibrated linear formulas for reads and writes, with weights based on compute, network, and disk-I/O measurements. - Dispatchers debit a local token bucket according to each request’s RU cost rather than charging every request equally. - Periodic RU refills preserve simple, static quotas while making them more proportional to actual resource consumption. - Requests are rejected with HTTP 419 when the RU bucket is exhausted. - Load shedding remains separate, allowing latency-based protection to react dynamically without changing the underlying quota-refill mechanism. ## Load Shedding Under Sudden Stress - RU rate limiting smooths normal traffic but may react too slowly to rapidly changing workloads. - Mussel adds a load-shedding layer based on: - Traffic criticality. - A real-time latency ratio. - A CoDel-inspired queue-management policy. - Each dispatcher compares long-term p95 latency with short-term p95 latency. - A ratio near 1.0 indicates stable performance; a drop toward 0.3 signals rapidly increasing latency. - When stress crosses the threshold: - The system raises the effective RU cost for a designated lower-priority client class. - That class’s token bucket drains faster, causing its traffic to back off. - If conditions worsen, the penalty expands to additional classes. - Critical workloads, such as customer support and trust-and-safety traffic, can remain responsive while less important traffic is reduced. - The latency estimate uses the constant-memory P² algorithm, avoiding raw sample storage and cross-node coordination. ## Hot-Key Detection and DDoS Protection - Client-level quotas cannot prevent overload when many clients request the same popular key. - Mussel therefore detects skewed access patterns in real time. - When duplicate requests target a hot key, the system can protect storage by: - Serving responses from cache. - Coalescing identical requests before they reach the backend. - This approach protects the underlying shard whether the traffic comes from legitimate popularity, automation, or a DDoS burst. Mussel’s experience suggests that mature multi-tenant services should move beyond fixed QPS limits. Combining resource-based accounting, priority-aware load shedding, and hot-key mitigation provides a more effective way to preserve reliability while maximizing useful work during unpredictable traffic conditions.

Read original(opens in new tab)
discord3 min readCurated summary

Discord Patch Notes: October 7, 2025

Discord’s October 7, 2025 patch focuses on reducing update interruptions, improving search and personalization, and resolving numerous platform-specific bugs. Notable changes include less frequent mandatory desktop updates on Windows, online indicators in Quick Switcher results, improved iOS notification clearing, and per-server Nameplates. The release also contains extensive fixes across general navigation, the Shop, Overlay, mobile apps, and chat. ## Desktop Updates and Personalization - Windows desktop updates are now delayed unless a release is mandatory or the client is several versions behind. - Users can still update immediately through the green arrow in the top-right corner. - Nameplates can now be configured individually for each server. - The Shop now supports search and filtering, making collectibles easier to find. - Quick Switcher search results display users’ online status. - Quick Switcher reliability was improved for users with per-server nicknames. ## Notifications and Platform Reliability - iOS notifications read on another device should now clear more reliably. - Remaining stale notifications are cleared when Discord launches. - Fixed blank screens when navigating backward on iOS. - Improved Android navigation between Channel Details tabs. - Fixed improperly rendered invite links on iOS and Android. - macOS title-bar traffic lights now render more consistently. - Desktop update windows should no longer occasionally display a blank screen. ## General Interface and Overlay Fixes - Dismissing the in-game overlay now correctly returns focus to the game. - Fixed drifting “New” badges, inconsistent Nitro badge borders, and various alignment issues. - Corrected problems with appearance settings, including the disabled “Reset to Default” button at 87.5% size. - Improved keyboard navigation in Server Settings. - Fixed unnecessary text truncation in the Multi-Factor Authentication modal. - Resolved issues with pasting content in several dialogs and screens. - Overlay context menus, tooltips, animations, voice widgets, toast notifications, and transparency received multiple visual and behavior fixes. - Fixed issues involving onboarding question deletion, server description limits, channel topics, forum sorting, and server invite behavior. - Updated the outdated Nitro logo and corrected capitalization in the Shop search tooltip. ## Chat and Messaging - Sending a DM from a profile modal now reliably loads the message into the chat view. - Fixed overlapping emoji in iOS reply previews. - Corrected the placement of the “Slowmode is enabled” message. - Android users can once again tap spoiler tags to reveal them. - Fixed scroll-position resets when navigating away from and back to the same Android channel. - Corrected typing indicators appearing over message text in Desktop chat history. - Fixed the iOS composer moving behind the keyboard after pasting a link. - Improved handling of character limits for the first message in a thread. - Fixed missing unread-message banners and incorrect positioning of iOS’s “Jump to Latest Message” button. Discord users should receive these improvements as the changes roll out across platforms, though availability may vary. Users encountering additional problems can report them through Discord’s community bug megathread, while iOS users can test upcoming changes through TestFlight.

Read original(opens in new tab)
discord3 min readCurated summary

From Single-Node to Multi-GPU Clusters: How Discord Made Distributed Compute Easy for ML Engineers

Discord argues that distributed machine learning becomes practical when developer experience is treated as a first-class engineering problem. Ray provided the distributed-computing foundation, while Discord built a platform around it with a CLI, Dagster and KubeRay orchestration, and the X-Ray observability interface. This transformed GPU-intensive ML from manual experimentation into reproducible production pipelines, enabling Ads Ranking to move to multi-GPU neural networks and produce major business gains. ## Scaling Beyond Single-Node ML - Discord’s ML systems grew from simple classifiers to complex models serving hundreds of millions of users. - Teams needed: - Multiple GPUs for training - Datasets larger than a single machine - More compute than existing infrastructure could provide - Ray addressed the distributed-computing challenge, but Discord still needed a standardized internal platform to make it easy to use. ## Problems with Ad-Hoc Ray Clusters - Early ML engineers manually created Ray clusters using open-source documentation. - This led to: - Inconsistent cluster configurations - Uneven resource management - No centralized scheduling - Limited monitoring - Multiple teams independently rebuilding infrastructure solutions - Discord concluded that Ray needed an internal platform layer rather than direct, manual use. ## A Parameterized CLI for Cluster Creation - Discord replaced numerous GPU-specific YAML templates with one parameterized template. - Engineers specify requirements such as: - GPU type - Worker count - Memory - The CLI generates Kubernetes configuration, security settings, and hardware-specific resource requests. - It manages the full cluster lifecycle, including creation and deletion. - This made multi-GPU environments available through a single command and standardized deployments across teams. ## Automated Orchestration with Dagster, KubeRay, and Ray - Discord combined three systems: - **Dagster** defines workflows, dependencies, schedules, and validated configuration. - **KubeRay** dynamically provisions Ray clusters on Kubernetes with the appropriate namespace, service account, and GPU node pool. - **Ray** executes distributed training, evaluation, and batch inference. - The workflow is: 1. An engineer launches or schedules a Dagster pipeline. 2. Dagster submits the job specification. 3. KubeRay creates the required Ray cluster. 4. Ray distributes the workload across GPUs. 5. Logs and metrics flow back to Dagster and monitoring systems. - The approach provides predictable, reproducible jobs with centralized visibility. - Discord’s ad relevance model now trains daily without engineers manually editing cluster configurations. ## Centralized Observability with X-Ray - Discord built X-Ray as a web UI for monitoring Ray infrastructure. - It displays: - Active clusters - Cluster ownership - Machine types - Current status - Engineers can inspect dashboards and launch interactive notebooks for experimentation from one place. ## Ads Ranking as a Production Test - Ads Ranking determines which Quest advertisements are most relevant to individual users. - Before Ray, the system relied on XGBoost and lacked: - Model sharding - Multi-GPU support - Scalable, frequent retraining - Ray enabled sharded neural networks trained on multi-GPU clusters. - Reported results included: - Twice as many players joining Quests - Ad coverage increasing from roughly 40% to nearly 100% - A production pipeline that retrains daily and continuously delivers new model versions Discord’s experience suggests that distributed ML succeeds when powerful infrastructure is paired with simple interfaces, automated orchestration, and strong observability. Organizations adopting Ray should build comparable platform tooling around it rather than expecting ML engineers to manage clusters, scheduling, and monitoring themselves.

Read original(opens in new tab)
googleOriginal article

XR Blocks: Accelerating AI + XR innovation (opens in new tab)

XR Blocks is an open-source, cross-platform framework designed to bridge the technical gap between mature AI development ecosystems and high-friction extended reality (XR) prototyping. By providing a modular architecture and high-level abstractions, the toolkit enables creators to rapidly build and deploy intelligent, immersive web applications without managing low-level system integration. Ultimately, the framework empowers developers to move from concept to interactive prototype across both desktop simulators and mobile XR devices using a unified codebase. ### Core Design Principles * **Simplicity and Readability:** Drawing inspiration from the "Zen of Python," the framework prioritizes human-readable abstractions where a developer’s script reflects a high-level description of the experience rather than complex boilerplate code. * **Creator-Centric Workflow:** The architecture is designed to handle the "plumbing" of XR—such as sensor fusion, AI model integration, and cross-platform logic—allowing creators to focus entirely on user interaction and experience. * **Pragmatic Modularity:** Rather than attempting to be a perfect, all-encompassing system, XR Blocks favors an adaptable and simple architecture that can evolve alongside the rapidly changing fields of AI and spatial computing. ### The Reality Model Abstractions * **The Script Primitive:** Acts as the logical center of an application, separating the "what" of an interaction from the "how" of its underlying technical implementation. * **User and World:** Provides built-in support for tracking hands, gaze, and avatars while allowing the system to query the physical environment for depth, estimated lighting conditions, and object recognition. * **AI and Agents:** Facilitates the integration of intelligent assistants, such as the "Sensible Agent," which can provide proactive, context-aware suggestions within the XR environment. * **Virtual Interfaces:** Offers tools to augment blended reality with virtual UI elements that respond to the user's physical context. ### Technical Implementation and Integration * **Web-Based Foundation:** The framework is built upon accessible, standard technologies including WebXR, three.js, and LiteRT (formerly TFLite) to ensure a low barrier to entry for web developers. * **Advanced AI Support:** It features native integration with Gemini for high-level reasoning and context-aware applications. * **Cross-Platform Deployment:** Developers can prototype depth-aware, physics-based interactions in a desktop simulator and deploy the exact same code to Android XR devices. * **Open-Source Resources:** The project includes a comprehensive suite of templates and live demos covering specific use cases like depth mapping, gesture modeling, and lighting estimation. By lowering the barrier to entry for intelligent XR development, XR Blocks serves as a practical starting point for researchers and developers aiming to explore the next generation of human-centered computing. Interested creators can access the source code on GitHub to begin building immersive, AI-driven applications that function seamlessly across the web and specialized XR hardware.

figma2 min readCurated summary

4 Ways for Design Teams to Chart New Territory With Figma Make | Figma Blog

Figma Make helps design teams turn early ideas into interactive, high-fidelity prototypes without starting from scratch. The article argues that this accelerates buy-in, exploration, collaboration, and design-system consistency, allowing designers to focus more on strategy, vision, and refinement. Maven Clinic’s experience shows how a prototype can revive a shelved feature and shorten development dramatically. ## Getting Ideas Back on the Roadmap - Maven Clinic had postponed a map-based fertility clinic finder because of launch deadlines. - Product Design Manager Loric Avanessian used initial designs with Figma Make to create an interactive prototype. - The prototype looked and felt like part of Maven’s existing product, generating enthusiasm across the company and even attracting CEO attention. - Its realism helped restore the feature to the roadmap despite competing priorities. - Designers could iterate between Figma Make and Figma Design, refining the concept and testing details. - The prototype exposed important micro-interactions early and enabled Maven to design, develop, test, and launch an MVP in fewer than four sprints—after the idea had remained in the backlog for two years. ## Exploring Unfamiliar Directions - Figma Make helps teams quickly generate and compare multiple ideas. - By providing something tangible to react to, it reduces the blank-canvas problem and supports divergent exploration before requirements are finalized. ## Collaborating on New Interfaces - Interactive prototypes make ideas easier for cross-functional partners to understand and critique. - Teams can share prototypes, gather focused feedback, and identify issues earlier than they might with static mockups. ## Incorporating Design Systems Early - Figma Make can produce explorations that remain visually consistent with established product patterns. - This reduces redundant layout work while allowing designers to concentrate on taste, strategic decisions, and product vision. Figma Make is most valuable when used as an early design partner: create a credible first version quickly, gather feedback, refine it in Figma Design, and use the result to align stakeholders before development begins.

Read original(opens in new tab)
slack4 min readCurated summary

Deploy Safety: Reducing customer impact from change

Slack’s Deploy Safety Program reduced customer-impact hours by 90% from its peak by focusing on safer change across all deployment systems, rather than optimizing individual services in isolation. The program combines measurable reliability goals, automated detection and rollback, blast-radius reduction, and cultural change. Its core lesson is to invest broadly, measure results, and expand approaches that demonstrably reduce customer impact without slowing development. ## Defining the Reliability Problem - Slack became increasingly mission-critical, raising customer expectations for reliability. - In analysis of customer-facing incidents, 73% were triggered by Slack-induced change, especially code deployments. - Incidents occurred across hundreds of services and multiple deployment systems, producing inconsistent levels of customer impact. - Customers reported that interruptions became significantly more disruptive after roughly 10 minutes. - Earlier reliability efforts often focused on individual deployment systems or services, leading to manual processes that slowed innovation and reduced engineering morale. ## North Star Goals and the Deploy Safety Manifesto The initial program goals applied to Slack’s highest-importance services: - Detect and automatically remediate deployment problems within 10 minutes. - Detect and manually remediate problems within 20 minutes. - Identify problematic deployments before they reach 10% of the fleet. - Preserve Slack’s engineering and development velocity. These goals later evolved into a Deploy Safety Manifesto covering all deployment systems and processes, including: - Automated safety improvements. - Deployment guardrails. - Changes to engineering practices and safety culture. ## Measuring Customer Impact Slack defined its primary program metric as: - **Hours of customer impact from high-severity and selected medium-severity change-triggered incidents.** The metric is an imperfect proxy for customer sentiment because: - Incident severity reflects current or anticipated impact, not always the final customer experience. - Medium-severity incidents require additional filtering to determine whether their actual impact is relevant. - It can be difficult to connect an individual engineering project directly to changes in customer sentiment. Slack evaluates the metric using four principles: - Measure outcomes rather than activity. - Distinguish real measurements from proxy metrics. - Apply subjective criteria consistently. - Regularly validate the metric against feedback from leaders who speak directly with customers. ## Choosing Where to Invest At the beginning of the program, Slack did not know which projects would produce the greatest benefit or when results would appear. Incident data is inherently delayed, while customers are experiencing reliability problems immediately. The investment strategy therefore emphasized: - Broad initial investment and a bias toward action. - Addressing known customer pain first. - Expanding successful projects and repeatable patterns. - Reducing investment in areas with limited impact. - Maintaining a flexible roadmap that could change as results emerged. Projects were prioritized according to whether they could: - Detect deployment problems earlier. - Improve automatic remediation time. - Improve manual rollback and remediation time. - Reduce severity by limiting deployment blast radius. ## Improving Webapp Backend Deployments Slack identified Webapp backend deployments as the largest source of change-triggered incidents and iteratively improved their safety: - Built automated metric monitoring. - Added automatic alerts and manual rollback procedures to validate alignment with customer impact. - Introduced automatic deployments and rollback. - Demonstrated that repeated automatic rollbacks could keep customer impact below 10 minutes. - Expanded monitoring to additional metrics. - Optimized manual rollback processes. - Added manual rollback capability for the frontend. - Began consolidating deployment practices through a centralized orchestration system inspired by ReleaseBot and AWS Pipelines. - Extended metrics-based deployment and automatic remediation beyond Bedrock and Kubernetes. These improvements made Webapp backend, frontend, and some infrastructure deployments significantly safer, with continued quarter-over-quarter improvement. ## Iterative Expansion Slack applied the same pattern across other areas: - Try an intervention. - Measure whether customer impact improves. - Invest further when the approach succeeds. - Reuse successful patterns in other systems. - Reduce or redirect investment when results are limited. The article notes that some efforts, such as faster mobile-app issue detection, were successful, while others produced less noticeable improvements. Slack’s experience suggests that deployment safety works best as an ongoing program: establish measurable customer-focused goals, automate detection and recovery, control blast radius, and continuously replicate proven practices without sacrificing delivery speed.

Read original(opens in new tab)