Techlist.io - Korean Tech Blog Curator

discord3 min readCurated summary

Osprey: Open Sourcing our Rule Engine

Discord is open-sourcing Osprey, a rule engine designed to help platforms detect and respond to emerging safety threats in real time. Built with ROOST and internet.dev, it processes platform events, evaluates configurable rules, and produces actionable verdicts with minimal engineering effort. Osprey emphasizes scale, rapid rule deployment, transparency, extensibility, and continuous improvement. ## Goals for a Modern Rule Engine Osprey was designed around several requirements: - Process thousands of events per second in real time. - Let teams create and deploy expressive rules within minutes. - Return clear verdicts indicating whether activity is safe, suspicious, or malicious. - Explain how rules were executed and expose errors for investigation and debugging. - Support feedback loops that improve future detection rules. - Remain extensible enough to address new attack patterns. ## Osprey’s Processing Model Osprey accepts platform events called **Actions** through either: - Synchronous gRPC requests. - Asynchronous message queues. The engine evaluates these actions using rules written in SML, a Python-based rule language. Rules can use Python UDFs, Features, and Effects, while synchronous requests can return Verdict effects directly to callers. Outputs are sent to Apache Druid, which powers investigation and analysis tools. ## Actions Actions are JSON-like events submitted to Osprey. - Each action type has a unique name and schema. - Callers can customize the payload with relevant platform data. - Example data includes login attempts, user IDs, usernames, email addresses, and IP addresses. - Rules extract and evaluate values from these action payloads. ## Rules and SML Rules are the central mechanism for detecting suspicious behavior. - SML uses a Python-inspired syntax intended to be accessible to less-technical rule authors. - Rules can reference other rules and extracted data. - Static validation enforces consistent rule-writing practices. - Validation can be extended with Python, from naming conventions to more complex domain-specific checks. - Example rules identify a known spammer by email and apply a `spammer` label to the associated user entity. ## User-Defined Functions UDFs are regular Python functions that extend Osprey’s rule language and standard library. - Built-in capabilities such as `Rule`, `WhenRules`, and `JsonData` are implemented as UDFs. - Teams can add their own UDFs when integrating Osprey into other products. - UDFs can retrieve information from external services, including machine-learning models. - They can be configured for asynchronous execution and access external-service providers through the execution context. - A sample UDF obtains a link-spam score from an external prediction service. ## Features and Entities Features are globally named variables produced during Osprey executions. - Features are exported to Apache Druid for later querying and investigation. - Prefixing a variable name with `_` keeps it local instead of exporting it. - Examples include `UserId` and `UserEmail`, extracted from JSON action data. - Entities are a specialized type of Feature representing persistent objects such as users, servers, or email addresses. - Entities can receive effects such as labels, classifications, and signals. - Entity types determine which effects are valid through static validation. - The Osprey interface provides dedicated Entity Views for examining an entity’s history. ## Effects Effects are outcomes triggered when rules evaluate as true. - They are validated and processed in aggregate after execution. - Effects can modify or annotate entities with labels, classifications, or signals. - Verdict effects can be returned synchronously to inform the requesting service of a safety determination. Osprey’s open-source release gives platforms a reusable foundation for real-time trust and safety enforcement. Teams interested in adopting it can explore the repository at [github.com/roostorg/osprey](https://github.com/roostorg/osprey).

Read original(opens in new tab)
discord3 min readCurated summary

Discord Patch Notes: February 4, 2026

Discord’s February 4, 2026 patch focuses on performance, streaming, permissions, mobile consistency, and bug fixes across desktop, iOS, and Android. The largest improvements include faster desktop rendering, zoomable screenshares, a new Slowmode bypass permission, and continued migration of Discord’s audio/video backend to Rust. The update also adds usability refinements and resolves numerous interface, profile, notification, and stability issues. ## Performance and Infrastructure - Desktop navigation and interaction are significantly faster, especially on systems with limited processing capacity. - The improvements were primarily achieved by replacing slow CSS selectors rather than optimizing APIs or components. - More than 80% of Discord’s audio/video traffic now runs through the Rust-based backend. - Stream previews load substantially faster during the start-stream process. ## Screensharing and Media - Users can zoom and pan screenshares and game streams with a mouse wheel or laptop trackpad. - Android now correctly plays videos containing multiple audio tracks, such as microphone and game audio. - Currently playing voice messages stop correctly when users log out. - Android Group DM notifications now show the group name instead of the sender’s name, matching iOS behavior. ## Server Permissions and Time Formatting - A new **Bypass Slowmode** permission lets server administrators trust selected members to bypass slowmode restrictions. - The permission can be integrated into existing roles through Role Settings until February 23. - Desktop now includes an `@time` command for generating localized Linux timestamps that automatically adapt to each viewer’s time zone. ## General Usability Improvements - Vibing Wumpus has been updated to version 2.0 and can be summoned with: - `Ctrl+Alt+Shift+W` on Windows/Linux - `Cmd+Option+Shift+W` on macOS - Pressing Escape while editing a desktop profile no longer closes the entire profile modal. - Copying with nothing selected no longer clears the clipboard. - Browse Channels on mobile has officially left beta. - Clicking role mentions now shows users assigned to that role for all stable users. - Invite Friends to Server supports both usernames and display names. - Ignoring a friend request correctly changes the action from “Accept Friend Request” to “Add Friend.” ## Profile, Themes, and Server Interface Fixes - Fixed cursor-jumping while editing the About Me section. - Avatar Decorations now render correctly in per-server profile editing. - Profile Theme gradients fill their modal correctly. - Server tags now appear accurately in desktop profile previews. - Fixed alignment and styling issues in Members, Emoji, Connections, and Display Name Styles interfaces. - iOS Client Themes no longer scroll at extremely high speed. - Channel categories on iOS no longer resemble unread channels. - Android no longer displays an extra bottom bar that blocks interaction. ## Stability and Bug Fixes - Fixed crashes involving related stickers in servers without text channels. - Prevented duplicate Checkpoint and gift-reception modals. - Emoji editing no longer creates duplicates when Finish is clicked twice quickly. - Desktop Inbox filtering works correctly again. - Fixed a desktop lockup that could occur after adding a game to a profile without saving. - Mobile Community setup now renders properly. - The Shop displays multiple columns correctly on affected Android devices. - Keybinds involving keys above F10 now persist correctly. - Event previews properly render Markdown. - On iOS, GIF avatars selected through the GIF picker remain animated instead of becoming static images. - Additional fixes address onboarding overlays, dangerous-download modal spacing, role selection transparency, event buttons, server uploads, and localized status text. Discord’s update is primarily a quality and performance release. Users should see the greatest immediate benefits from faster desktop interactions, improved streaming controls, better mobile behavior, and more precise server administration tools, although individual fixes may continue rolling out by platform.

Read original(opens in new tab)
gitlabOriginal article

GitLab backs 99.9% availability SLA with service credits (opens in new tab)

GitLab has introduced a 99.9% availability service-level agreement (SLA) specifically for Ultimate customers on GitLab.com and GitLab Dedicated. This commitment is backed by service credits to ensure that mission-critical DevSecOps workflows remain uninterrupted and to align GitLab's interests with customer business outcomes. By formalizing this uptime guarantee, GitLab aims to provide a reliable foundation for high-velocity teams that depend on continuous code pushes and automated deployments. ## Scope of Covered Services The SLA covers the core platform experiences essential to daily software delivery workflows: * Issues and merge requests management. * Git operations, including push, pull, and clone actions via both HTTPS and SSH protocols. * Operations within the Container Registry and Package Registry. * API requests associated with the aforementioned core services. ## Defining and Measuring Downtime Service availability is tracked via automated monitoring across multiple geographic locations to reflect actual user experience. * A "downtime minute" is triggered when 5% or more of valid customer requests result in server errors. * Server errors are strictly defined as HTTP 5xx status codes or connection timeouts exceeding 30 seconds. * While monitoring focuses on server-side failures, GitLab will also holistically review claims for issues that might not trigger 5xx errors, such as Sidekiq job processing outages or specific application bugs. ## Service Credit Claim Procedure To maintain accountability, GitLab has established a formal process for Ultimate customers to recoup costs during outages: * Customers must submit a support request at support.gitlab.com within 30 days of the end of the month in which the downtime occurred. * The GitLab team validates the claim against internal and external monitoring data. * Validated service credits are applied directly to the customer's next issued invoice, with the credit amount scaled based on the severity of the availability shortfall. Ultimate customers should familiarize their operations teams with these specific performance thresholds and the 30-day claim window to ensure they are adequately compensated during significant service disruptions.

datadog1 min readCurated summary

How we reduced the size of our Agent Go binaries by up to 77% | Datadog

The supplied text does not include the tech blog post itself. It contains Datadog navigation links and a promotional banner announcing its recognition as a Leader in the 2026 Gartner Magic Quadrant for Observability Platforms, but no article body or technical sections. ## Available content - Datadog promotes observability products covering: - Infrastructure and Kubernetes monitoring - Application performance monitoring - Logs and database monitoring - Security - Digital experience monitoring - Software delivery and CI visibility - Service management - AI-powered investigation and monitoring - The page links to an engineering article at: - `/blog/engineering/agent-go-binaries/` - No technical explanation, examples, conclusions, or section content from that article is included. Please provide the blog post’s full text or relevant excerpt for a substantive summary.

Read original(opens in new tab)
datadog3 min readCurated summary

How we reduced the size of our Agent Go binaries by up to 77%

The Datadog Agent’s Linux artifact grew from 428 MiB in version 7.16.0 to 1.22 GiB in 7.60.0, creating problems for serverless, IoT, and containerized environments. Rather than remove features, Datadog reduced Go binary sizes by up to 77% between versions 7.60.0 and 7.68.0. The effort combined dependency analysis, targeted code refactoring, and renewed use of Go linker optimizations. ## Why the Agent Became So Large - The Agent supports many operating systems, architectures, distributions, and deployment environments. - Its codebase contains hundreds of dependencies, including cloud SDKs, container runtimes, and security tools. - Build tags and dependency injection determine which features are included in each binary. - The compressed Linux amd64 Debian package grew from 126 MiB to 265 MiB. - Its uncompressed size increased from 428 MiB to 1,248 MiB—a 192% increase over five years. - Go binaries represented a substantial portion of that growth and became the primary optimization target. ## How Go Selects Dependencies - Go compiles required packages individually before the linker combines them into a binary. - Files are included only when they: - Are not test files ending in `_test.go` - Match the current operating system, architecture, and build tags - Satisfy other constraints such as CGO settings, compiler version, or architecture features - Starting from the main package, Go transitively includes imported packages and the runtime required by every Go binary. - Unnecessary dependencies can be excluded by: - Adding a build tag to the file that imports them - Moving dependency-using symbols into a separate package imported only by relevant binaries ## Analyzing Imports and Dependencies - `go list` reveals all packages used for a specific OS, architecture, and set of build tags. - `goda` generates dependency graphs, including indirect imports. - `goda` can also show only the paths leading to a particular target package using its `reach` function. - These tools account for `GOOS`, `GOARCH`, and build constraints, making them useful for examining platform-specific builds. ## Why Package Lists Are Not Enough - A package’s presence does not directly indicate its binary size impact. - The linker removes symbols that are not reachable from the program’s entry points. - The same package can therefore contribute different amounts of code depending on how it is used. - Importing a package can still have significant side effects: - `init` functions execute. - Global variables are initialized. - These behaviors may force otherwise unnecessary symbols to remain in the binary. - Certain uses of reflection can also limit linker optimizations. - Datadog used `go-size-analyzer` to measure the contribution of individual dependencies more accurately than import graphs alone. ## Overall Optimization Strategy - Datadog systematically audited dependencies rather than removing product capabilities. - The work focused on restructuring imports, isolating optional functionality, and restoring linker optimizations that had been disabled or undermined over time. - The resulting improvements brought artifact sizes close to levels from roughly five years earlier. - Some compiler and linker behaviors uncovered during the effort led to improvements benefiting other large Go projects, including Kubernetes. The practical lesson is to treat binary size as an ongoing dependency and architecture concern: analyze actual symbol reachability, isolate optional features behind build constraints or packages, and verify each build variant independently.

Read original(opens in new tab)
pinterest4 min readCurated summary

Drastically Reducing Out-of-Memory Errors in Apache Spark at Pinterest

Pinterest developed **Auto Memory Retries** to reduce Spark out-of-memory failures without permanently assigning oversized executors to every task. The system detects OOM failures and retries affected tasks with progressively larger resource profiles, reducing both on-call incidents and wasted compute. Instead of tuning every job for its peak memory demand, Pinterest can size jobs around typical usage while handling exceptional tasks elastically. ## Pinterest’s Spark Environment - Pinterest processes more than **90,000 Spark jobs daily** across tens of thousands of nodes. - Its infrastructure includes: - Kubernetes clusters - Spark 3.2, with Spark 3.5 adoption underway - Apache Celeborn for shuffle - Apache YuniKorn for scheduling - Apache Gluten and Meta’s Velox for acceleration - Archer, Pinterest’s internal submission service - More than **4.6% of job failures** were caused by OOM errors. ## Why Manual Memory Tuning Was Insufficient - Pinterest’s clusters are memory-bound, so simply increasing executor sizes is expensive and difficult. - Automatic tuning generally reduces executor memory to match historical usage and improve resource efficiency. - Manual tuning can work, but requires substantial expertise because: - Different stages perform different operations. - Individual tasks may have very different memory needs because of data skew. - Configurations that work for most tasks may fail for a small number of high-memory tasks. - Auto Memory Retries allow jobs to target approximately their **P90 memory usage**, while automatically giving unusually demanding tasks more capacity. ## How Spark Executor Memory Works - An executor’s memory and CPU capacity determine how many tasks can run concurrently. - By default, each CPU core provides a task slot. - For example, with `spark.task.cpus=2`, an executor with two usable task slots and 8 GB of memory provides roughly 4 GB per task on average. - Memory is shared, so one task may temporarily use more than its average allocation if another uses less. - An OOM occurs when the combined memory usage of concurrent tasks exceeds the executor’s available memory. ## Auto Memory Retries Design Pinterest modified Spark’s scheduling loop so individual tasks can use resource profiles different from their parent `TaskSet`. - Each task can store an optional `taskRpId` identifying its retry resource profile. - Pinterest creates immutable retry profiles at **2x, 3x, and 4x** the base profile. - If off-heap memory is enabled, it is scaled as well. - Retries use a hybrid strategy: - **First retry:** Double `cpus per task`, allowing the task to run on an existing executor with fewer concurrent tasks. - **Later retry:** Launch a physically larger executor if the task still fails or already requires the entire executor. - The approach prioritizes reusing existing executors before provisioning larger ones. ## Changes to Spark Internals Pinterest extended core Spark components through Pinterest-specific subclasses rather than using a listener-only implementation. - **Task** - Stores the optional task resource profile ID. - **TaskSetManager** - Tracks tasks with non-default profiles. - Assigns the next larger retry profile after an OOM. - **TaskSchedulerImpl** - Allows tasks with increased CPU requirements to run on standard executors. - **ExecutorAllocationManager** - Tracks pending tasks by retry profile. - Requests larger executors when physical memory is required. - The feature-specific classes are loaded only when Auto Memory Retries is enabled. - The Spark UI was updated to display each task’s resource profile ID. ## Handling Tasks After an OOM - When a task fails on an executor with more than one core, its first retry doubles `spark.task.cpus`. - Other tasks in the same stage or future stages are unaffected. - Spark cannot reliably determine which concurrent task caused the executor-level OOM. - As a result, Pinterest treats **all tasks running on the terminated executor** as having failed due to OOM and routes them to retries that do not share the executor with other tasks. ## Practical Conclusion Pinterest’s approach makes executor sizing elastic at the task level: configure jobs for normal memory usage, then progressively increase resources only for tasks that need them. This can reduce OOM-related failures and operational load while avoiding the cost of running every task on oversized executors.

Read original(opens in new tab)
figma3 min readCurated summary

Redefining Impact as a Data Scientist | Figma Blog

Data science impact is not limited to experiments, forecasting, or optimization. In complex, high-stakes systems such as billing, data scientists can create value by making workflows understandable, validating correctness, and improving operational safety. Figma’s experience shows that effective data science may require domain modeling, cross-functional collaboration, instrumentation, and production-quality tools. ## Data Science as a Full-Stack Discipline - The role of data science varies by team: it may involve experimentation, product analysis, data modeling, instrumentation, or operational tooling. - Billing combines a user-facing product with backend infrastructure, so accuracy directly affects customer trust. - Supporting Billing required: - Building deep domain expertise - Partnering with engineers and other functions - Creating tools that explain and verify system behavior - Experimentation and opportunity analysis remained useful, but represented a smaller portion of the actual work. - Figma’s full-stack model encouraged the team to define the right data science support collaboratively rather than follow a fixed playbook. ## Explaining Complex Systems Beyond Charts and Models - Some of the most valuable data science work explains existing or historical outcomes rather than predicting future ones. - A single invoice seat charge may depend on: - Seat assignments and removals - Permission changes - Contract terms - Workspace state - Billing rules - The timing of state transitions - Figma built the **Invoice Seat Report** to reconstruct the complete reasoning behind each charge. - The application combines product events, contract metadata, billing rules, and historical state transitions, presenting the result in plain language. - Building it required: - Reconciling fragmented schemas and inconsistent historical data - Validating assumptions with engineers - Adding instrumentation where logs recorded what happened but not why - Translating billing rules into traceable and debuggable SQL transformations - The team also had to account for legacy multiyear contracts, sparse seat histories, early upgrades, and other cases that could create gaps in the data. ## Shaping Technical Direction Through Data - Data scientists can turn business rules into measurable checks that define expected system behavior. - These validations can detect drift, regressions, and anomalies in both development and production. - For Billing, automated verification is especially important because small errors in seat states or invoice calculations can affect customer charges and trust. - During Figma’s billing-model re-architecture, data science helped verify that: - Data moved correctly through pipelines - New pricing and billing logic produced intended outcomes - Customers did not enter unexpected billing states - The system could be monitored consistently across environments The practical lesson is to look beyond conventional analytics when assessing data science impact. In complex domains, building reliable data foundations, explanatory tools, and correctness checks may be more valuable than running another experiment.

Read original(opens in new tab)
gitlabOriginal article

Claude Opus 4.6 now available in GitLab Duo Agent Platform (opens in new tab)

GitLab has integrated Anthropic’s Claude Opus 4.6 into its Duo Agent Platform, providing developers with a high-intelligence frontier model designed for complex agentic workflows. By combining a 1-million-token context window with native access to DevSecOps data, the update enables more autonomous task execution and deeper reasoning within the software development lifecycle. This integration allows teams to delegate multi-step tasks to AI agents that can now process entire codebases and project histories in a single interaction. ## Advanced Agentic Capabilities and Reasoning * Claude Opus 4.6 features enhanced "agentic" behavior, meaning it can proactively take actions and drive tasks forward with minimal human intervention. * The model supports multi-agent orchestration, allowing it to spin up subagents and coordinate parallel workstreams to solve complex, multi-step problems. * Adaptive thinking capabilities allow the model to calibrate its reasoning depth based on the query, using extended thinking for difficult tasks while maintaining speed for simpler ones. * Deep reasoning via test-time compute helps the model navigate challenging development bottlenecks and architectural decisions. ## Full-Context DevSecOps Integration * The model boasts a 1-million-token context window—a fivefold increase over Opus 4.5—enabling the processing of massive codebases and extensive documentation. * Integration with the GitLab Duo Agent Platform provides the model with direct access to repositories, merge requests, pipelines, and security findings. * Enterprise-grade security features, including human-in-the-loop controls and group-based access, ensure that agentic actions remain transparent and governed. * Native integration ensures developers can utilize these frontier capabilities without leaving their established GitLab workflows. ## Availability and Resource Consumption * Opus 4.6 is currently available for GitLab.com users via the Duo Agent Platform and Agentic Chat, though it is not supported for GitLab Duo Classic features. * Support for the model within various Integrated Development Environments (IDEs) is expected to be released in the near future. * Usage is managed via GitLab credits, with multipliers determined by the size of the prompt. * Prompts containing 200k tokens or fewer are charged at 1.2 requests per credit, while larger prompts exceeding 200k tokens are charged at 0.7 requests per credit. Organizations aiming to automate complex development workstreams should migrate their specialized agents to Claude Opus 4.6 to take advantage of its superior orchestration and context handling. By leveraging the model's ability to coordinate parallel subagents, teams can significantly reduce the manual effort required for codebase-wide refactors and security remediation.

aws2 min readCurated summary

Amazon EC2 Hpc8a Instances powered by 5th Gen AMD EPYC processors are now available | Amazon Web Services

Amazon EC2 Hpc8a instances are now generally available for tightly coupled, compute-intensive HPC workloads. Powered by 5th Gen AMD EPYC processors reaching 4.5 GHz, they provide up to 40% more performance, 42% higher memory bandwidth, and 25% better price-performance than Hpc7a instances. AWS targets applications such as fluid dynamics, weather modeling, design simulations, and crash analysis. ## Instance Specifications - Available in a single `96xlarge` configuration: - 192 CPU cores - 768 GiB memory - 300 Gbps Elastic Fabric Adapter (EFA) networking - Uses a 1:4 core-to-memory ratio. - Customers can customize the number of cores at launch to better match workload requirements. - Simultaneous Multithreading (SMT) is disabled to maximize HPC performance. - Sixth-generation AWS Nitro cards handle virtualization, storage, and networking tasks separately from the CPUs. ## Supported HPC Services - Integrates with AWS ParallelCluster and AWS Parallel Computing Service (AWS PCS) for cluster creation and job submission. - Supports Amazon FSx for Lustre, offering sub-millisecond latency and throughput of up to hundreds of gigabytes per second. - High-bandwidth, low-latency networking is designed for workloads requiring extensive communication between compute nodes. ## Availability and Purchasing - Initially available in: - US East (Ohio) - Europe (Stockholm) - Offered through On-Demand Instances and Savings Plans. - Regional availability and future expansion can be checked through AWS Capabilities by Region. Hpc8a instances are best suited for organizations needing faster simulation results and efficient scaling across tightly coupled HPC workloads. Teams can launch them through the Amazon EC2 console and combine them with AWS cluster and storage services for a complete HPC environment.

Read original(opens in new tab)
aws3 min readCurated summary

Announcing Amazon SageMaker Inference for custom Amazon Nova models | Amazon Web Services

Amazon SageMaker Inference now generally supports deploying and scaling full-rank customized Amazon Nova models. The feature gives production workloads more control over instance types, autoscaling, context length, concurrency, and batch settings while improving cost efficiency through optimized GPU utilization. Customers can train Nova Micro, Nova Lite, and Nova 2 Lite models with SageMaker Training Jobs or HyperPod, then deploy them as managed real-time or asynchronous endpoints. ## Custom Nova Model Support - Supports customized Nova Micro, Nova Lite, and Nova 2 Lite models. - Models can use: - Continued pre-training - Supervised fine-tuning - Reinforcement fine-tuning - Custom models can be trained through Amazon SageMaker Training Jobs or Amazon HyperPod. - SageMaker Inference provides managed deployment, scaling, and HTTPS access for production workloads. - GPU utilization and inference costs can be optimized with Amazon EC2 G5 and G6 instances instead of relying exclusively on P5 instances. - Autoscaling can respond to five-minute usage patterns. - Configurable context length, concurrency, and batch size help balance latency, cost, and accuracy. ## Deploying Through SageMaker Studio - In SageMaker Studio, users select a trained Nova model from the Models menu. - Choosing **Deploy**, **SageMaker AI**, and **Create new endpoint** starts deployment. - Deployment settings include: - Endpoint name - Instance type - Initial and maximum instance counts - Permissions - Networking configuration - Supported launch instance types vary by model: - Nova Micro: G5, G6, and P5 options, including `g5.12xlarge` through `g6.48xlarge` and `p5.48xlarge` - Nova Lite: `g5.48xlarge`, `g6.48xlarge`, and `p5.48xlarge` - Nova 2 Lite: `p5.48xlarge` - Provisioning takes time because SageMaker must create infrastructure, download model artifacts, and initialize the inference container. - Once the endpoint is `InService`, users can test it in the Studio Playground using chat prompts. ## Deploying with the SageMaker SDK - Deployment requires two SageMaker resources: - A model object referencing the Nova artifacts and inference container - An endpoint configuration specifying the instance type and count - Model artifacts can be stored in Amazon S3 and referenced with an S3 prefix. - Environment variables configure inference behavior, including: - `CONTEXT_LENGTH` - `MAX_CONCURRENCY` - `DEFAULT_TEMPERATURE` - `DEFAULT_TOP_P` - The endpoint configuration creates a real-time endpoint, such as one using an `ml.g5.12xlarge` instance. - SageMaker supports network isolation and execution roles for secure deployment. ## Inference and Request Configuration - Endpoints support synchronous real-time inference in streaming or non-streaming modes. - Asynchronous endpoints are available for batch-style processing. - Requests can configure: - Maximum output tokens - Temperature - Top-p and top-k sampling - Log probabilities - Streaming usage statistics - Reasoning effort, with `low` and `high` options - The example request asks the model to compare quarterly spending against budget and identify variances above 10 percent. SageMaker Inference provides a complete path from Nova customization to production deployment. Teams should select instance types and tune context length, concurrency, batching, and sampling parameters based on their workload’s latency, cost, and accuracy requirements.

Read original(opens in new tab)
aws3 min readCurated summary

AWS Weekly Roundup: Amazon EC2 M8azn instances, new open weights models in Amazon Bedrock, and more (February 16, 2026) | Amazon Web Services

AWS’s February 16, 2026 roundup highlights the launch of Amazon EC2 M8azn instances, which deliver substantial performance gains for compute-intensive workloads. It also covers expanded Amazon Bedrock model and networking support, improved observability in EKS Auto Mode, more efficient OpenSearch Serverless capacity management, and configurable RDS backup settings during snapshot restoration. The post concludes with upcoming AWS conferences, summits, and community events. ## Amazon EC2 M8azn Instances - Powered by fifth-generation AMD EPYC processors with a maximum frequency of 5 GHz. - Compared with M5zn instances, they provide: - Up to 2× compute performance - 4.3× higher memory bandwidth - 10× larger L3 cache - Up to 2× networking throughput - Up to 3× EBS throughput - Built on the AWS Nitro System with sixth-generation Nitro Cards. - Available in nine sizes, from 2 to 96 vCPUs and up to 384 GiB of memory, including two bare-metal options. - Designed for high-performance workloads such as financial analytics, high-frequency trading, CI/CD, gaming, simulations, and HPC. ## New Open-Weight Models in Amazon Bedrock - Bedrock now supports six fully managed models: - DeepSeek V3.2 - MiniMax M2.1 - GLM 4.7 - GLM 4.7 Flash - Kimi K2.5 - Qwen3 Coder Next - The models target reasoning, agentic intelligence, autonomous coding, and cost-efficient production deployments. - They use Project Mantle and support OpenAI-compatible APIs. - DeepSeek V3.2, MiniMax 2.1, and Qwen3 Coder Next are also available in Kiro. ## Amazon Bedrock PrivateLink Support - AWS PrivateLink now supports the `bedrock-mantle` endpoint in addition to `bedrock-runtime`. - Project Mantle provides serverless inference, quality-of-service controls, automated capacity management, and OpenAI API compatibility. - PrivateLink support for OpenAI-compatible endpoints is available in 14 AWS Regions. ## EKS Auto Mode Logging - EKS Auto Mode now supports CloudWatch Vended Logs for managed capabilities such as: - Compute autoscaling - Block storage - Load balancing - Pod networking - Logs can be delivered to CloudWatch Logs, Amazon S3, or Amazon Data Firehose. - The feature includes AWS authentication and authorization and is offered at a lower price than standard CloudWatch Logs. ## OpenSearch Serverless Collection Groups - Collection Groups allow multiple collections to share OpenSearch Compute Units while retaining separate KMS keys and access controls. - Shared capacity can reduce OCU costs. - Administrators can define both minimum and maximum OCU limits, ensuring baseline capacity for latency-sensitive applications. ## RDS Snapshot Restore Improvements - RDS now lets users view and configure backup retention periods and preferred backup windows before or during snapshot restoration. - Restored databases no longer need post-restore backup configuration changes. - The feature supports all major RDS engines, Aurora editions, commercial AWS Regions, and GovCloud at no additional cost. ## Upcoming AWS Events - AWS Summits in Paris, London, and Bengaluru during April 2026. - AWS AI and Data Conference in Ireland on March 12, focusing on Bedrock, SageMaker, QuickSight, agent deployment, data integration, and governance. - AWS Community Days in Ahmedabad, Slovakia, and Pune. Overall, the announcements emphasize faster specialized compute, broader managed AI model access, stronger private connectivity, and improved operational controls across AWS services.

Read original(opens in new tab)
google3 min readCurated summary

Teaching AI to read a map

MapTrace addresses a major weakness in multimodal language models: recognizing objects on maps is easier for them than understanding connectivity, obstacles, and valid routes. The authors propose a synthetic-data pipeline that generates maps, identifies walkable areas, constructs navigation graphs, and verifies computed paths with AI critics. They report releasing 2 million map question-answer pairs and show that fine-tuning on a much smaller subset improves route tracing on unseen real-world maps. ## The Challenge: Weak Spatial Grounding - MLLMs may recognize locations and objects in an image but still draw routes through walls, buildings, enclosures, or shops. - Effective navigation requires understanding: - Which regions are traversable - How paths connect - That routes are ordered sequences of connected points - The geometric and topological relationships between map features - Existing image-text training rarely teaches this “spatial grammar.” - Manual pixel-level route annotation would be expensive and difficult to scale. - Many useful maps of malls, museums, and theme parks are proprietary, limiting access to real-world training data. ## A Scalable Synthetic-Data Pipeline MapTrace uses generative AI to create diverse maps and automatically produce valid route annotations. ### Generating Diverse Maps - An LLM creates detailed prompts for environments such as: - Zoos with interconnected habitats - Shopping malls with food courts - Fantasy theme parks with themed areas - A text-to-image model renders the prompts as map images. - This approach provides control over map diversity and complexity. ### Identifying Walkable Areas with a Mask Critic - Pixels are clustered by color to produce candidate masks representing possible walkways. - An MLLM reviews each mask alongside the original map. - The “Mask Critic” rejects masks that do not represent realistic, connected traversable regions. - Accepted areas may include sidewalks, crosswalks, and pedestrian paths. ### Converting Maps into Navigation Graphs - The selected traversable mask is converted into a pixel-based graph. - Walkway intersections become nodes, while connected stretches become edges. - This graph captures the map’s connectivity and enables computational route planning. ### Generating and Validating Routes - Thousands of random start and end points are sampled for each map. - Dijkstra’s algorithm computes the shortest path between each pair. - A “Path Critic” checks the overlaid route to ensure it: - Stays within traversable regions - Avoids obstacles - Follows a logical human route - Routes approved by the critic become training examples. ## Dataset and Evaluation - The pipeline generated 2 million annotated map question-answer pairs. - The authors note that generated maps sometimes contain incorrect text, but the study focuses primarily on path fidelity. - They fine-tuned models including Gemma 3 27B and Gemini 2.5 Flash on 23,000 generated paths. - Performance was evaluated on MapBench, which contains unseen real-world maps. - Route accuracy was measured using normalized dynamic time warping (NDTW), which compares predicted and reference coordinate sequences while accounting for differences in sampling and travel speed. - Lower NDTW scores indicate closer agreement with the reference route. ## Conclusion The work suggests that targeted synthetic training data can teach MLLMs map-based spatial reasoning that is largely missing from general pretraining. The released dataset and pipeline provide a foundation for improving visual navigation, while better image-generation models should reduce remaining typography and rendering artifacts.

Read original(opens in new tab)
figma2 min readCurated summary

The Future of Design Is Code and Canvas | Figma Blog

The post argues that the future of product creation combines code with visual design canvases rather than treating them as separate, linear stages. Figma’s integration with Claude Code lets developers send rendered browser work into Figma as editable layers, enabling teams to explore alternatives visually and move changes back into code. The broader goal is to help builders avoid tunnel vision and choose better solutions before committing to implementation. ## Code and Canvas as Complementary Tools - Code is powerful for building and expressing ideas, while the canvas is better for comparing and navigating many possibilities. - Figma supports: - Divergent exploration of multiple approaches - Side-by-side comparison of designs - Direct manipulation of visual details - Big-picture evaluation before implementation ## Claude Code to Figma - Users can install the Figma MCP and type “Send this to Figma” in Claude Code. - The browser’s rendered state is translated into fully editable Figma layers. - After refining the design in Figma, Figma MCP can transfer design changes back into the codebase. - This creates a bidirectional workflow between production code and visual design. ## Moving Beyond Linear Workflows - Traditional product development often followed a sequence: brainstorm, design, then code. - AI and connected tools allow work to begin in a terminal, prompt box, visual interface, or sketch and move between formats. - Teams can now reconsider direction during development instead of simply advancing the first workable concept. ## Design as the Main Differentiator - As AI makes it easier to generate almost any articulated possibility, the difficult work becomes identifying the best solution. - Design judgment, craft, and point of view remain essential. - Figma positions the canvas as a space for stepping back, examining alternatives, and escaping the momentum of building the first version. The practical recommendation is to combine code-driven speed with canvas-based exploration: use code to create, Figma to compare and refine, and MCP integrations to keep both workflows connected.

Read original(opens in new tab)
figma2 min readCurated summary

From Claude Code to Figma: Turning Production Code into Editable Figma Designs | Figma Blog

Claude Code to Figma lets users capture working interfaces from production, staging, or localhost and convert them into editable Figma frames. The workflow combines code’s speed for building functional prototypes with Figma’s strengths in collaboration, comparison, and exploration. Its central argument is that teams can move faster without stopping at the first working implementation. ## From Code to an Editable Canvas - Developers can capture real UI screens from Claude Code workflows. - Captured screens can be pasted into any Figma file as editable frames. - The workflow supports interfaces running in production, staging, or locally. - Multiple screens can be captured in one session, preserving flow sequence and context. ## Start Anywhere, Then Collaborate - Code-first exploration is fast but often isolated: one person manages the branch, server, and context. - Sharing screenshots, recordings, or local builds creates friction when feedback is needed. - Once imported into Figma, screens can be organized, duplicated, refined, annotated, and shared. - Teams can discuss and explore the interface without switching environments or modifying code for every idea. ## Build the Best Idea, Not Just the First One - AI makes it easier to produce an initial prototype quickly, shifting attention toward evaluating alternatives. - Figma Make supports a similar workflow by bringing generated prototypes onto the design canvas. - Claude Code to Figma extends this approach to code-created interfaces. - Both workflows aim to turn an initial tangible result into deeper design exploration. ## Explore Systems and Variations Visually - Side-by-side frames make patterns, inconsistencies, gaps, and trade-offs easier to identify. - Teams can duplicate frames, rearrange steps, and test structural changes without reimplementing code. - Keeping alternatives visible supports continued exploration, including previously rejected ideas. - Designers, engineers, and product managers can make decisions using the same high-fidelity artifact. - Shared context helps surface questions and resolve direction earlier. Claude Code to Figma is intended as a bridge between functional prototyping and collaborative design. Teams can use code to quickly discover what works, then move the result into Figma to compare options, gather feedback, and establish shared direction.

Read original(opens in new tab)
cloudflare3 min readCurated summary

Shedding old code with ecdysis: graceful restarts for Rust services at Cloudflare

Cloudflare’s open-source Rust library **ecdysis** enables zero-downtime restarts for high-volume network services. It preserves listening sockets and existing connections while a new process initializes, avoiding refused connections and dropped requests. After five years of production use, Cloudflare uses it to safely deploy fixes, security patches, and new features across its global infrastructure. ## Why Conventional Restarts Fail - Stopping the old process before starting the new one creates a period when no process is listening. - New clients receive `ECONNREFUSED`; even a 100 ms gap can drop hundreds of connections at a busy location. - Existing connections—including file uploads, video streams, WebSockets, and gRPC streams—are terminated when the old process exits. - `SO_REUSEPORT` allows multiple processes to bind the same port, but can orphan connections: - The kernel assigns an incoming `SYN` to one listening socket. - If that process exits before calling `accept()`, the queued connection is terminated. - This makes simply overlapping two independently bound processes unsafe for graceful upgrades. ## The ecdysis Restart Model ecdysis uses a process-forking approach pioneered by NGINX: - The parent calls `fork()` to create a child. - The child replaces itself with the new executable using `execve()`. - The child inherits the listening socket file descriptors through a named pipe shared with the parent. - The parent continues serving traffic while the child initializes. - Once the child signals readiness, the parent closes its copy of the listening socket and drains existing connections. - Both processes may briefly accept connections during the transition, but this is intentional and avoids coverage gaps. ## Crash Safety and Upgrade Requirements - The old process can fully shut down after the replacement is ready. - The new process receives time to initialize before taking over. - If initialization fails—for example, because of invalid configuration—the child exits while the parent continues serving normally. - Upgrades are serialized so that only one runs at a time, preventing cascading failures. - The unchanged listening socket ensures that new connections are not refused during the handoff. ## Rust and System Integration - ecdysis provides native Tokio stream wrappers for asynchronous Rust services. - Synchronous services can use it without an async runtime. - With the `systemd_notify` feature enabled, it integrates with systemd lifecycle notifications. - Configuring a service with `Type=notify-reload` allows systemd to track graceful upgrades correctly. Cloudflare’s approach demonstrates that graceful restarts require coordination between processes rather than simply starting a second server. Services needing reliable zero-downtime upgrades can use ecdysis to preserve connections, tolerate failed deployments, and safely roll out new Rust binaries.

Read original(opens in new tab)