Techlist.io - Korean Tech Blog Curator

meta4 min readCurated summary

How Meta Used AI to Map Tribal Knowledge in Large-Scale Data Pipelines

AI coding assistants struggle when they lack a map of a large, proprietary codebase. To address this, the team built a pre-compute system using 50+ specialized agents that analyzed over 4,100 files across four repositories and three languages, producing 59 concise context files. The approach gave agents complete module coverage, captured previously undocumented tribal knowledge, reduced tool calls by about 40%, and made complex development tasks much faster. ## The Problem: Powerful Tools Without Codebase Context - The pipeline combines Python configuration, C++ services, and Hack automation across multiple repositories. - A seemingly simple change, such as adding a data field, can affect: - Configuration registries - Routing logic - DAG composition - Validation rules - C++ code generation - Automation scripts - AI agents often explored repeatedly, guessed at conventions, and produced code that compiled but was subtly incorrect. - Important examples of missing context included: - Different field names for the same operation in separate configuration modes - “Deprecated” enum values that must remain for serialization compatibility - Hidden intermediate field names used between pipeline stages ## The Pre-Compute Approach The team used a large-context model and orchestrated specialized agents in several phases: - Two agents explored and mapped the codebase. - Eleven analysts read every file and answered five questions: - What does the module configure? - How is it commonly modified? - What non-obvious patterns can cause failures? - What are its cross-module dependencies? - What tribal knowledge is hidden in comments? - Writers generated context files. - More than ten critic passes reviewed quality across three rounds. - Fixers, upgraders, gap-fillers, prompt testers, and final critics corrected and validated the results. - In total, more than 50 specialized tasks were coordinated in one session. This process uncovered over 50 non-obvious design patterns, including naming conventions and append-only identifier rules that were not documented elsewhere. ## Context Files: “A Compass, Not an Encyclopedia” Each of the 59 context files is intentionally short—about 25–35 lines or roughly 1,000 tokens—and contains: - Quick Commands for common operations - Key Files limited to the most relevant three to five files - Non-Obvious Patterns - See Also references to related modules Together, the files use less than 0.1% of a modern model’s context window. They are designed for targeted, opt-in use rather than being loaded into every task. ## Routing and Dependency Navigation - An orchestration layer routes natural-language requests to the appropriate tool. - Operational questions can trigger dashboard scans and matching against more than 85 historical incident patterns. - Development requests can launch configuration generation and multi-phase validation. - A cross-repository dependency index and data-flow maps show how changes propagate. - Dependency questions that previously required about 6,000 tokens of exploration can be answered through a graph lookup using roughly 200 tokens. ## Results and Quality Controls - Preliminary tests across six tasks showed approximately 40% fewer tool calls and tokens. - Work that previously required around two days of research and engineer consultation took about 30 minutes. - Critic reviews raised quality scores from 3.65 to 4.20 out of 5. - Every referenced file path was verified, with no hallucinated paths. - Coverage expanded from navigation guidance for roughly 5% of modules to all 4,100+ files across three repositories. ## Why This Differs from Generic Context Files Research has found that AI-generated context files can reduce agent performance on familiar open-source projects. The team argues that this result does not directly apply to proprietary systems whose conventions and tribal knowledge are absent from model training data. Their approach addresses common problems by making context: - Concise rather than encyclopedic - Opt-in rather than always loaded - Quality-gated through independent critics - Continuously refreshed to prevent stale information Without this context, agents typically spend 15–25 tool calls exploring and remain vulnerable to subtle domain-specific errors. ## Keeping the Knowledge Fresh Automated jobs refresh the system every few weeks by: - Validating file paths - Detecting coverage gaps - Re-running critic reviews - Finding and repairing stale references - Updating routing and dependency information The system treats AI not merely as a consumer of documentation, but as the engine that creates and maintains it. ## Applying the Method Elsewhere Teams can adapt the approach by: - Identifying where agents most often fail due to undocumented conventions or dependencies - Applying the five-question analysis framework to each module - Keeping context files short and action-oriented - Using independent quality critics before publishing generated guidance - Automating freshness checks and self-repair The practical recommendation is to build a small, targeted, continuously maintained knowledge layer for proprietary codebases. Concise navigation and dependency context can reduce exploration costs while preventing the subtle errors that arise when agents lack domain-specific understanding.

Read original(opens in new tab)
spotify3 min readCurated summary

Background Coding Agents: Supercharging Downstream Consumer Dataset Migrations (Honk, Part 4) | Spotify Engineering

Spotify used its Honk background coding agent with Backstage and Fleet Management to automate migrations from two deprecated datasets to new versions. The effort targeted roughly 1,800 downstream pipelines and produced 240 automated pull requests, potentially saving about 10 engineering weeks. The experience showed that agents perform best when repositories follow standardized patterns, prompts contain precise technical context, and automated testing is available. ## The Challenge of Large-Scale Dataset Migrations - Two heavily used datasets needed replacement to support new dimensions and features. - The datasets had approximately 1,800 direct downstream pipelines and affected thousands more indirectly. - Migrations spanned three frameworks: - BigQuery Runner - dbt - Scala-based Scio - Manual migration was estimated to require around 10 engineering weeks within a six-month deadline. ## Using Backstage to Identify Consumers - Backstage’s endpoint lineage pages revealed downstream dataset consumers. - Its Codesearch plugin located relevant repositories across Spotify’s GitHub Enterprise environment. - The Fleetshift plugin used those results to organize and orchestrate repository migrations. - Backstage also provided a centralized view for tracking progress and opening generated pull requests. ## Context Engineering for Honk - Honk needed detailed, self-contained prompts because it could not access external documentation, dataset schemas, MCPs, or custom Claude skills during execution. - Scio was excluded because its flexible, inconsistent implementations made it difficult to describe all migration cases in one reliable prompt. - BigQuery Runner and dbt were more standardized, making them better candidates for automation. - An initial prompt based on a human migration guide was insufficient and caused incorrect assumptions about field mappings. - Explicit mapping tables in the context file significantly improved results. - Prompts also specified cases where fields should not be migrated automatically. - Honk left those fields unchanged. - It added comments linking to human migration guidance for later review. ## Testing and Automated Pull Requests - BigQuery Runner and dbt repositories generally lacked build-time unit tests. - As a result, Honk could not automatically verify and correct its changes, one of its key capabilities. - Downstream teams had to manually test the generated pull requests before merging. - Despite this limitation, the team successfully created 240 automated migration PRs. - Fleetshift’s Backstage interface simplified monitoring, troubleshooting, repository navigation, and communication with owning teams. ## Lessons for Future Agent-Driven Maintenance - Large-scale automation depends on standardizing frameworks and data practices across repositories. - Consistent testing and validation requirements are essential so agents can verify their own changes. - Future Honk functionality will allow agents to gather context from sources such as JIRA tickets and documentation before editing code. - Better context gathering should reduce the need for exhaustive prompt files and improve migration quality. Spotify’s experience suggests that background coding agents can substantially reduce migration toil, but their effectiveness depends on disciplined standardization, explicit migration rules, and strong automated testing.

Read original(opens in new tab)
spotify4 min readCurated summary

Indexing the Data Lake for Online Point Queries | Spotify Engineering

Random Access Parquet (RAP) enables low-latency point queries directly against massive data lakes. It addresses the mismatch between fast cloud storage and query engines such as Trino or BigQuery, whose planning and scheduling overhead can make single-row lookups take seconds. By indexing keys to exact Parquet files and rows, RAP avoids broad scans and retrieves only the required data while preserving a shared source of truth for analytics, ML, and online services. ## Why Conventional Query Engines Struggle - Online applications and AI agents need fast access to per-user histories, often stored across billions of records. - Keeping all data in systems such as Bigtable or DynamoDB is prohibitively expensive when the lake contains exabytes. - Object storage latency is increasingly suitable for online workloads: - GCS requests typically take 30–100 ms. - S3 Express One Zone and GCS Rapid Storage can offer single-digit millisecond latency. - Distributed SQL engines introduce seconds of scheduling and planning overhead, making them better suited to analytical throughput than point lookups. ## Narrowing the Search Still Leaves a Problem - A 90-day query over daily listening data could involve approximately 90,000 Parquet files. - Key-based partitioning can reduce candidates substantially; with 1,000 buckets per day, the set may fall to 90 files. - Bloom filters can eliminate files that do not contain the requested user, potentially reducing the set to around 12 files. - The remaining files still require multiple dependent reads: - Fetch the Parquet footer. - Parse row-group metadata. - Scan the key column. - Locate relevant pages through column and page indexes. - Read the corresponding value pages. - These sequential, dependent requests consume both latency and cloud-storage bandwidth. ## The RAP Approach - RAP replaces scanning with direct lookup. - An external index maps each key to: - The relevant Parquet file. - The row numbers containing the key. - Optionally, the number of values for pagination. - Cached metadata maps row numbers to page locations. - The reader then issues precise ranged reads for the required pages. - Reads can be performed in parallel because they no longer depend on a chain of discovery operations. - The approach benefits cloud storage, SSDs, and memory because it removes dependent reads at every tier. ## The External Index - RAP works with existing, unmodified Parquet files. - An index builder reads file footers and page locations, scans key columns, and writes key-to-location mappings. - New pipeline output adds index fragments rather than modifying existing index data. - The index is a multimap, allowing one key to occur in multiple files and partitions. - Index size is typically much smaller than the data: - Terabytes of data may produce gigabytes of index. - Petabytes of data may produce terabytes of index. - Large indexes can be distributed using hash bucketing. - Unlike Parquet Bloom filters and PageIndex structures, the external index is definitive: it identifies exact files and rows instead of merely narrowing a scan. - With unmodified files, RAP may still need to read an entire page—for example, several megabytes to retrieve a small value—so write-time preparation can further reduce read cost. ## Preparing Parquet for Faster Point Reads - Once the reader knows the exact key, row, and columns, file layout can prioritize smaller and fewer final reads rather than in-file discovery. - Relevant optimizations fall into three broad categories: - Concentrating a key’s data. - Reducing bytes per read. - Reducing the number of reads. - Sorting by key places a key’s rows together, minimizing the number of pages required. - Hash bucketing deterministically places each key in one file per partition. - Co-grouping can store one row per key with values in repeated or nested structures, such as an array of timestamp, track URI, and duration fields. - Coarser partitioning can reduce how many files a key spans. RAP therefore provides a way to serve interactive point queries from the same Parquet data used for batch analytics, reducing duplication and avoiding specialized serving copies. For latency-sensitive workloads, indexing should be combined with write-time layout decisions that cluster keys and minimize page reads.

Read original(opens in new tab)
discord3 min readCurated summary

Discord Patch Notes: April 6, 2026

Discord’s April 6, 2026 patch focuses on performance, accessibility, media sharing, and a broad set of usability fixes across desktop, iOS, Android, and Linux. The most notable improvements reduce desktop voice-channel deadlocks by about 30% and reduce iOS image-upload sizes by 17% and latency by 12%. Discord also continues a major accessibility audit while addressing numerous navigation, layout, search, and platform-specific bugs. ## Performance and Media Sharing - Desktop changes reduced deadlocked Voice threads by approximately 30%, making users less likely to remain stuck on “Connecting.” - iOS image uploads now use files roughly 17% smaller and complete about 12% faster. - Mobile landscape mode was improved by calculating padding on a screen-by-screen basis rather than relying on global padding rules. - Android server reordering now scrolls more smoothly. ## Accessibility Improvements - Discord is continuing a large accessibility audit across its clients. - Fixes addressed keyboard focus rings, button alignment, text overflow, and controls becoming inaccessible in longer languages. - Desktop profile buttons, settings controls, and the Server Invite modal received layout and focus improvements. - Discord encourages users to report remaining accessibility problems through its bug-reporting channels. ## Search, Navigation, and Account Behavior - Search negation now works correctly with filters such as `has:-image`. - The desktop `CMD/CTRL+F` shortcut no longer opens server search while a modal is active. - Browser-style back and forward navigation now behaves correctly after switching accounts. - Opening a Discord link from a browser no longer replaces the current channel without preserving a usable way to return. - Fixed several Settings and modal-navigation problems, including unexpected jumps to the bottom of the page. ## Desktop and Settings Fixes - Corrected keybind-button alignment and keyboard focus positioning. - Fixed profile buttons extending beyond the visible area in languages with longer text. - Shop item modals can now be closed by clicking outside them at minimum window height. - Profile bio changes are properly cleared by the Reset button. - Removed a persistent “NEW” badge from the `@time` command. - Corrected outdated role designs in Server Template previews. - Fixed visual issues involving nameplates, profile banners, Nitro perk badges, and Shop error-message spacing. - The Desktop update indicator no longer appears clickable while an update is still downloading. ## Mobile and Platform-Specific Fixes - Wayland now properly detects when Linux users become active or go AFK. - Android now provides visual confirmation after sending a friend request through a QR code. - Android theme changes update the entire Settings interface immediately. - Android QR-code login buttons now appear active when usable. - Android’s In-App Browser setting correctly opens links inside Discord. - iOS users can now switch out of Invisible status reliably. - iOS server lists no longer jump when switching servers. - iOS Server Guide progress bars and welcome messages now display and dismiss correctly. - Fixed an iOS crash that could occur when canceling a Nitro Classic subscription. - iOS search results now display images in bot-message containers at the correct size. ## Messaging, Profiles, and Social Features - Removed duplicate friend suggestions for users who had already received a request. - Fixed incorrect profile connection icons for external links. - Corrected duplicate usernames shown in pending friend-request tooltips. - Channel-name inputs no longer incorrectly offer custom emoji, which channel names do not support. - Student Hub join-method filtering now works properly. - Server Tags and badges received alignment fixes on Android. - The Nitro gift emoji picker’s “Add Emoji” button now functions correctly. Discord recommends updating as fixes reach each platform, with additional early testing available through the iOS TestFlight release.

Read original(opens in new tab)
netflix3 min readCurated summary

Powering Multimodal Intelligence for Video Search

Video search is difficult because it must combine many kinds of information—characters, scenes, dialogue, labels, and embeddings—across enormous volumes of footage. The post argues that solving this problem requires a distributed pipeline that separates reliable ingestion, computationally intensive data fusion, and low-latency search indexing. Temporal bucketing, hybrid ranking, and deduplication turn billions of model outputs into searchable moments for editors. ## Why Video Search Is Complex - Video contains multiple overlapping modalities, each analyzed by specialized models. - Models produce different outputs, including: - Text labels such as characters or objects - Scene classifications - High-dimensional embedding vectors - Time ranges with varying boundaries - Overlapping model timelines must be synchronized into a chronological representation. - A 2,000-hour archive may contain more than 216 million frames, expanding to billions of records after multimodal processing. - Search must avoid returning thousands of redundant clips from continuous shots. - Ranking therefore combines: - Symbolic text matching for precision and interpretability - Semantic vector similarity for contextual relevance - Clustering and deduplication to identify the best moments - Sub-second response times are essential because delays interrupt editors’ creative workflows. ## Three-Stage Ingestion and Fusion Pipeline ### Transactional Persistence - Raw model annotations are ingested through highly available pipelines. - Apache Cassandra stores the annotations with an emphasis on: - Data integrity - Distributed availability - High write throughput - An annotation can include a type, nanosecond time range, embedding vector, label, and confidence score. ### Offline Data Fusion - After persistence, Apache Kafka publishes an event that starts asynchronous processing. - The offline pipeline performs expensive temporal intersections without slowing ingestion or search. - Model outputs are normalized into fixed one-second time buckets. - The fusion process: - Maps continuous detections into discrete intervals - Intersects annotations sharing a bucket - Combines them into unified records - Writes the enriched records back to Cassandra - For example, a “Joey” character detection from seconds 2–8 can be combined with a “kitchen” scene detection from seconds 4–9 to create a fused record for the 4–5 second interval. - Each fused record retains links to the original annotations and source asset. ### Real-Time Search Indexing - Enriched buckets are later sent from Cassandra to Elasticsearch. - Upserts use a composite key consisting of the asset ID and time bucket. - If a bucket already exists, it is updated rather than duplicated. - This creates one consistent record for each second of footage while allowing new model results to be incorporated. The overall recommendation is to treat multimodal video search as a distributed data-fusion problem rather than a single-model retrieval task. Decoupling ingestion, offline processing, and indexing allows the system to handle massive archives while preserving reliable data capture and fast, context-rich search.

Read original(opens in new tab)
aws2 min readCurated summary

Amazon Bedrock Guardrails supports cross-account safeguards with centralized control and management | Amazon Web Services

Amazon Bedrock Guardrails now supports cross-account safeguards, allowing organizations to centrally enforce safety controls across AWS accounts and organizational units. Administrators can apply immutable, versioned guardrails to all Bedrock model invocations while still allowing account- or application-specific policies. The capability is generally available across commercial and GovCloud Regions where Bedrock Guardrails is supported. ## Centralized Organization- and Account-Level Enforcement - **Organization-level enforcement** uses an Amazon Bedrock policy created in the AWS Organizations management account. - Policies can attach a specified guardrail and version to: - The organization root - Organizational units - Individual AWS accounts - The selected guardrail is automatically applied to Bedrock inference requests across targeted member entities. - Different policies and guardrails can be assigned to different accounts or organizational units. - **Account-level enforcement** applies a configured guardrail to all Bedrock inference API calls within one account and Region. ## Configuring Guardrail Coverage - Guardrails must use a specific version so their configuration remains immutable and cannot be changed by member accounts. - Administrators can choose whether enforcement: - Includes or excludes specific Bedrock models - Covers all or only selected system and user prompt content - **Comprehensive** mode guards all content, regardless of caller-provided tags. - **Selective** mode relies on callers to identify content requiring protection, reducing processing for pre-validated inputs. ## Testing and Verification - Account-level enforcement can be configured in the Amazon Bedrock Guardrails console. - Enforcement can be tested with: - `InvokeModel` - `InvokeModelWithResponseStream` - `Converse` - `ConverseStream` - Responses include guardrail assessment details and identify the enforced guardrail. - Member accounts can verify organization-level enforcement in the Bedrock console. ## Important Considerations - Organizations must meet prerequisites such as configuring resource-based policies for guardrails. - Incorrect or invalid guardrail ARNs can cause policy violations, prevent safeguards from being enforced, and block model inference. - Automated Reasoning checks are not supported. - Charges apply for each enforced guardrail based on its configured safeguards. ## Availability Cross-account safeguards are generally available in all commercial and GovCloud AWS Regions where Amazon Bedrock Guardrails is available. Organizations can enable the feature through the Amazon Bedrock and AWS Organizations consoles. Overall, the capability gives security teams a centralized way to enforce responsible AI requirements while reducing the need to audit guardrail settings independently in every account and application.

Read original(opens in new tab)
github3 min readCurated summary

The uphill climb of making diff lines performant

GitHub rebuilt the pull request **Files changed** experience to keep diff reviews responsive across everything from tiny fixes to massive changes. The core conclusion is that no single optimization solves performance at scale; instead, targeted rendering improvements, virtualization, and simpler components must work together. Early results show that even small reductions in DOM size can have substantial effects on memory usage and interaction latency in large pull requests. ## Performance Challenges at GitHub’s Scale - Pull requests may contain thousands of files and millions of lines. - In extreme cases, the old experience reached: - More than **1 GB of JavaScript heap usage** - Over **400,000 DOM nodes** - Unacceptably high Interaction to Next Paint (INP) scores - Large reviews became sluggish or nearly unusable, despite the experience remaining fast for most smaller pull requests. ## A Strategy Based on Pull Request Size GitHub concluded that different pull request sizes require different performance strategies: - **Optimize diff-line components** so medium and large reviews remain fast without losing expected browser behavior, such as native find-in-page. - **Use virtualization for the largest reviews**, rendering only the content currently needed to preserve responsiveness and stability. - **Improve foundational components and rendering**, allowing performance gains to benefit every pull request size. ## Problems with the Original Diff Architecture The first React implementation made each diff line unnecessarily expensive: - Unified view used roughly **10 DOM elements per line**; split view used about **15**, before syntax highlighting added more `<span>` elements. - Each unified diff line typically involved at least **eight React components**, while split view involved at least **13**. - Additional states—such as comments, hover, and focus—could add still more components. - Small components often registered five or six React event handlers each, resulting in **20 or more handlers per line**. - These costs multiplied across thousands of lines, increasing JavaScript heap usage and worsening INP. - The component-heavy design was initially reasonable when React was introduced, but proved unsustainable for unbounded data sets. ## Incremental Improvements in the New Design GitHub’s second version focused on simplification and removing unnecessary structure: - Reduced state, JavaScript, React components, and DOM elements. - Removed redundant `<code>` tags from line-number cells. - Eliminating just two nodes per line saves approximately **20,000 DOM nodes across 10,000 lines**. - The example demonstrates how seemingly minor changes compound into meaningful improvements at large scale. The practical lesson is that performant large-scale interfaces require layered optimizations: simplify every repeated element, reduce per-item overhead, and use virtualization when rendering everything at once is no longer viable.

Read original(opens in new tab)
line4 min readCurated summary

From Hive to Iceberg: The Secret to 12x Faster Data Reflection

LINE Plus replaced a full-dump ETL pipeline for product data with incremental processing using Apache Iceberg and Apache Flink. The previous HBase/Hive workflow rewrote hundreds of millions of rows for every update, causing high compute costs and delays that left data up to an hour out of date. With the new architecture, update intervals were reduced from 60 minutes to 5 minutes—roughly a 12× improvement—while preserving consistency and fault tolerance. ## Limitations of Full-Data ETL - The existing HBase and Hive pipeline continuously collected CDC data in HDFS but had to merge it with existing data and rewrite the entire table before changes became queryable. - This caused: - High compute and storage costs - Dependence on limited shared Hadoop resources - Delayed updates and stale data - Snapshot-based extraction provides consistency, but large snapshots can take hours and retain old versions through MVCC, increasing system overhead. - Processing only the changed rows would reduce the workload from hundreds of millions of records to tens of thousands, separating update cost from total dataset size. ## Introducing Apache Iceberg - Iceberg manages data through metadata and table snapshots rather than relying solely on directory structures like traditional Hive tables. - It supports row-level `upsert` and `delete` operations. - This allows incremental changes to be written without rewriting the entire table, making much shorter ETL intervals possible. ## Requirements for the Streaming Pipeline The team evaluated Spark and Flink against three essential requirements: - **Data freshness:** Late-arriving compensation or replay data must not overwrite newer records. - **End-to-end exactly-once processing:** Iceberg updates and Kafka status messages must not partially succeed. - **Fault tolerance and state management:** Processing state must survive failures and restarts. A Kafka message indicating that all CDC data through a specific timestamp—such as 13:03—has been applied serves as the signal that a bulk extraction can safely begin. This requires complete confidence that the message accurately represents the Iceberg table’s committed state. ## Why Two-Phase Commit Was Necessary - Iceberg and Kafka are independent systems, so writing to one while failing to write to the other could create inconsistent state. - Two-phase commit (2PC) prevents partial success: - Both systems prepare their writes. - They commit only when all required operations succeed. - Any failure causes the operation to roll back. - Exactly-once processing also prevents duplicate or missing records during retries, network failures, or node restarts. - Together, these guarantees make Kafka status messages a reliable representation of the Iceberg table’s state. ## Choosing Flink over Spark - Spark Structured Streaming uses a micro-batch model, which makes fine-grained event-time and state control more difficult. - Flink provides native event-by-event streaming and better support for the required consistency model. - The team used Flink state to track each record’s `updatedate`: - Older late-arriving events are ignored. - Replayed historical data cannot overwrite newer values. - Flink checkpoints: - Persist streaming state externally. - Enable recovery from the latest consistent point. - Integrate with the Kafka sink’s 2PC mechanism. - Kafka messages remain in a pre-commit state until the Iceberg write and checkpoint both succeed. ## Kubernetes Deployment Options - The team compared: - **Native Kubernetes:** Requires manually configuring roles, service accounts, services, routing, deployments, slots, and jobs. - **Flink Kubernetes Operator:** Represents Flink infrastructure and jobs as custom resources, automating configuration such as routing and the web UI through Helm values. - Although Flink has greater operational complexity and a steeper learning curve than Spark, it was selected because it was the only option that satisfied all three core requirements at the engine level. The recommended architecture is an incremental Iceberg pipeline powered by Flink, with stateful processing, checkpoints, and two-phase commit between Iceberg and Kafka. This approach keeps data current, avoids expensive full-table rewrites, and provides reliable recovery and consistency at a five-minute update interval.

Read original(opens in new tab)
toss3 min readCurated summary

Why Toss reduced its design roles to two

On April 1, Toss’s Design Chapter consolidated six design roles into two: Product Designer and Visual Designer. The change reflects how role boundaries had already blurred as designers crossed disciplines and technology reduced the importance of tool-specific expertise. Toss’s central argument is that designers should be organized around judgment and user problems—not the tools, media, or screens they work with. ## Why Role Boundaries Became a Problem - The previous structure separated designers by tools and outputs rather than by the decisions they made. - Ambiguity emerged in areas such as: - Whether interaction in a design system belonged to Platform or Interaction Designers - Whether interactive graphics should be handled through Lottie, code, or UI design - Whether expanding a PC product to mobile belonged to a Tools Product Designer or Product Designer - These divisions sometimes determined ownership based on medium instead of capability or context. ## Designers Were Already Crossing Disciplines - Tools Product Designers began designing mobile products. - Interaction Designers worked on parts of internal design tools. - Graphic Designers created semantic icon systems. - Platform Designers built interactive web pages. - Brand Designers with visual-design backgrounds worked on lighting products. - AI and other tools have shortened the time needed to learn formerly specialized skills, including: - Video and Lottie production - Figma prototyping - Coding interactive experiences - As tool proficiency becomes less differentiating, the ability to judge what creates a good experience becomes more important. ## Product Designer - Product Designer and Tools Product Designer were merged into one role. - The distinction between mobile and PC disappeared. - The role now focuses on: - Understanding the user’s context and problems - Deciding how those problems should be solved - Designing across screen sizes and product environments ## Visual Designer - Platform, Interaction, Graphic, and Brand Designers were combined into Visual Designer. - Visual Designers are expected to work across media and produce what the experience requires, such as: - Building interactions within systems - Creating icons for prototypes - Designing interactive web experiences - The defining capability is visual judgment: deciding what is beautiful, appropriate, and correct. - The title was chosen to emphasize visual decision-making rather than a specific medium or technique. ## Lessons from Other Industries - Disney animation reduced many physical and intermediate production steps through software while preserving stages requiring important creative judgment. - Digital audio workstations allow artists such as Billie Eilish and Finneas to compose, perform, record, and mix with a laptop, but human judgment about what sounds good remains essential. - Digital cinema and streaming weakened the historical distinction between film and television production. - Across these industries, tools converged while the value of creative judgment increased. ## What Comes Next - The new job structure will not immediately change how people work. - Toss still needs to redesign hiring standards, onboarding, and career-development paths. - The consolidation is intended to give designers broader ownership and more room to make decisions across disciplines, ultimately improving the experiences delivered to users.

Read original(opens in new tab)
google3 min readCurated summary

Evaluating alignment of behavioral dispositions in LLMs

The post introduces a framework for evaluating whether LLM behavior aligns with human behavioral tendencies in realistic social and workplace situations. Instead of relying on self-report questionnaires, it converts validated psychological traits into situational judgment tests and compares model responses with judgments from human annotators. Across 25 models, larger systems align better when humans strongly agree, but models remain overconfident and often fail to represent legitimate human disagreement. ## From Psychological Self-Reports to Situational Tests - The researchers adapt statements from established instruments measuring traits such as empathy, emotion regulation, and assertiveness. - Because LLM self-reports can vary with prompt wording and may not predict real behavior, the statements are transformed into realistic user-assistant scenarios. - Each scenario presents two possible actions: - One expressing or supporting a behavioral trait. - One opposing or suppressing it. - Three annotators review each generated test to ensure the scenario and actions accurately represent the intended trait. - Models respond naturally, and an LLM judge maps each response to one of the two actions. - Human preferences are collected from 10 annotators per scenario, drawn from a pool of 550 participants. ## Measuring Directional Alignment - Directional alignment measures whether a model gives greater probability to the action favored by the human majority. - The analysis focuses on scenarios with strong human consensus: - Unanimous agreement: 10 of 10 annotators. - Very high agreement: 9 or 10. - High agreement: 8 or 9. - Smaller models, particularly those under 25 billion parameters, often perform near chance and struggle to distinguish when a trait should be expressed or restrained. - Larger models over 120 billion parameters and frontier closed-weight models perform substantially better. - These models approach near-perfect alignment when human agreement is unanimous, but performance generally plateaus in the low-to-mid 80% range when consensus is weaker. - Qualitative deviations included: - Encouraging emotional openness in professional situations where humans preferred composure. - Favoring harmony in disputes instead of standing up for one’s position. - Recommending immediate action in time-sensitive situations without sufficient logistical verification. ## Representing Human Disagreement - The study also evaluates distributional alignment: whether model confidence reflects the diversity of human opinions. - When human annotators disagree, a well-aligned model should distribute its probability more evenly between the available actions. - The results show systematic model overconfidence across all 25 evaluated systems. - Models tend to favor one action too strongly even when human preferences are divided, indicating that they often fail to preserve pluralism in human judgment. ## Broader Implications - The framework distinguishes two types of alignment gaps: - Directional gaps, where models choose differently from a clear human majority. - Distributional gaps, where models fail to reflect uncertainty or disagreement among people. - The findings suggest that scale improves behavioral alignment but does not fully solve nuanced social judgment. - Evaluating behavior in realistic scenarios may reveal limitations that conventional personality questionnaires or direct model self-reports miss. Future alignment work should assess not only whether models choose the human-majority response, but also whether their confidence and range of responses appropriately reflect genuine variation in human perspectives.

Read original(opens in new tab)
netflix3 min readCurated summary

Smarter Live Streaming at Scale: Rolling Out VBR for All Netflix Live Events

Netflix switched all Live events from Constant Bitrate (CBR) to capped Variable Bitrate (VBR), using AWS Elemental MediaLive’s QVBR setting. VBR allocates bits according to scene complexity, reducing delivery costs and improving playback quality, but its unpredictable spikes and dips invalidate traditional capacity-planning assumptions. Netflix addressed this by reserving delivery capacity according to each stream’s nominal bitrate rather than its current traffic level. ## Why Netflix Moved Live Streaming from CBR to VBR - CBR delivers streams near a fixed target, making server capacity and traffic patterns easy to predict. - However, CBR wastes bits on simple scenes and may provide insufficient bits for complex action. - VBR targets consistent visual quality instead: - Simple scenes use substantially fewer bits. - Complex scenes receive higher bitrate to prevent artifacts. - Netflix’s tests found: - Approximately 15% fewer bytes transferred on average. - Around 10% less traffic during the peak minute. - About 5% fewer rebuffers per hour. - Lower average traffic improves Open Connect scalability and can reduce startup delays and playback interruptions. ## Why VBR Creates Stability Risks - VBR bitrate can remain well below its nominal target during simple scenes, sometimes using only 2 Mbps for a 5 Mbps stream. - Delivery systems may interpret these low-traffic periods as spare server capacity and route additional sessions to the server. - When complex content appears—such as fights, confetti, rapid camera movement, or detailed crowds—bitrate can quickly rise to 6–8 Mbps or more. - If too many sessions were admitted during the low-bitrate period, aggregate traffic can exceed link or NIC capacity, causing: - Higher latency - Packet loss - Playback stalls - Quality downshifts ## Making Capacity Planning Aware of VBR - Netflix changed traffic-steering decisions so they no longer rely solely on current throughput. - Each stream reserves capacity based on its nominal bitrate, even when its current bitrate is much lower. - This treats every stream as capable of quickly returning to its expected capacity level. - The approach prevents servers from being overfilled during low-complexity scenes and keeps delivery behavior consistent across CBR and VBR. ## Matching VBR Bitrates to CBR Quality - Identical nominal bitrates do not produce identical behavior: - CBR remains clustered around its target with frequent small variations. - VBR spends far less on simple scenes and increases bitrate only when complexity demands it. - Netflix therefore needed to revisit which nominal VBR bitrates correspond to the quality previously delivered by CBR, rather than assuming the same configured bitrate would provide equivalent results. Netflix’s rollout shows that VBR is more than an encoder setting: it requires coordinated changes to bitrate ladders, capacity reservations, and traffic steering. With those safeguards, VBR can deliver comparable or better quality while using significantly less network capacity.

Read original(opens in new tab)
meta3 min readCurated summary

KernelEvolve: How Meta’s Ranking Engineer Agent Optimizes AI Infrastructure

KernelEvolve is Meta’s agentic system for automating the creation and optimization of low-level AI kernels across diverse hardware. It treats kernel tuning as a search problem rather than one-shot code generation, evaluating hundreds of alternatives with profiling and diagnostics. The system reduces optimization work from weeks to hours and has delivered over 60% higher inference throughput for an Ads model on NVIDIA GPUs and over 25% higher training throughput on Meta’s MTIA chips. ## Kernel Optimization at Meta - AI models rely on optimized kernels that translate high-level operations into hardware-specific instructions. - Meta runs models across NVIDIA GPUs, AMD GPUs, custom MTIA accelerators, and CPUs. - Production workloads require many custom operators beyond standard GEMMs and convolutions available in vendor libraries. - Kernels must be developed and tuned for each combination of: - Hardware type and generation - Model architecture - Operator type ## The Challenge of Hardware Heterogeneity - NVIDIA, AMD, MTIA, and CPU platforms differ in: - Memory architectures and hierarchies - Instruction sets - Execution models - Supported numeric data types - A kernel optimized for one platform may perform poorly or fail on another. - Hardware generations also require new optimization strategies. Meta’s MTIA roadmap includes four generations, from MTIA 300 through MTIA 500, in two years. - Manual tuning by kernel specialists cannot keep pace with these changes. ## Increasing Model and Operator Complexity - Meta’s recommendation systems have evolved from embedding-based models to sequence models with attention, GEM, and LLM-scale models such as Meta Adaptive Ranking Model. - Each new model generation introduces operators that earlier systems did not require. - Multiple model families may be involved in a single ads-serving request. - As model architectures and operator inventories grow, the number of kernel configurations expands rapidly into the thousands. ## How KernelEvolve Works - KernelEvolve generates candidate implementations in languages and DSLs including: - Triton, Cute DSL, and FlyDSL - CUDA, HIP, and MTIA C++ - A dedicated job harness compiles, runs, profiles, and evaluates each candidate. - Performance results, correctness checks, and diagnostic information are fed back to the LLM. - The system continuously searches through hundreds of alternatives instead of stopping at the first plausible implementation. - Its automated workflow includes profiling, optimization, testing, and cross-hardware debugging. ## Results and Broader Impact - KernelEvolve improved Andromeda Ads inference throughput by more than 60% on NVIDIA GPUs. - It improved training throughput for an ads model by more than 25% on Meta’s MTIA silicon. - The system operates across both public and proprietary hardware. - In production, it optimizes code supporting trillions of daily inference requests. - By automating kernel development, Meta can enable new hardware and adapt to changing model architectures with substantially less engineering effort. KernelEvolve turns kernel development from a manual, expert-driven bottleneck into a continuous automated process. Its search-based approach is particularly valuable as Meta’s hardware portfolio and model architectures continue to diversify.

Read original(opens in new tab)
dropbox3 min readCurated summary

Improving storage efficiency in Magic Pocket, our immutable blob store

Magic Pocket’s immutable design protects data integrity but makes storage efficiency dependent on continuous reclamation. A new Live Coder service reduced write amplification while unintentionally creating severely under-filled volumes, driving fragmentation and storage overhead sharply upward. Dropbox responded by rethinking compaction, since its existing steady-state strategy was too slow to recover space from the resulting long tail of sparse volumes. ## The Cost of Immutability - Magic Pocket stores user files as immutable blobs distributed across its storage fleet. - Updates and deletions never modify data in place; obsolete blobs remain until compaction. - Garbage collection identifies unreferenced blobs, while compaction physically moves live blobs into new volumes and retires old ones. - Because closed volumes cannot be reopened, deleted data creates unused space unless it is actively consolidated. - Durability also increases storage requirements: - Replication stores multiple complete copies. - Erasure coding splits data into fragments and adds parity, providing fault tolerance with less overhead. - Fragmentation determines how efficiently that redundant capacity is used: - A volume with 50% live data effectively doubles required storage. - A volume with 10% live data uses roughly ten times the necessary space. ## The Live Coder Incident - A new on-the-fly erasure-coding service created severely under-filled volumes as it rolled out to new regions. - In the worst cases, less than 5% of a volume’s capacity contained live data. - Since volumes have fixed allocations, many mostly empty volumes consumed nearly as much raw capacity as full volumes. - Dropbox detected rising effective replication-factor signals, indicating more raw storage was being used per live byte. - The existing compaction system continued reclaiming space but was not designed for a long tail of extremely sparse volumes. - The incident demonstrated that compaction must adapt when the distribution of live data changes substantially. ## Steady-State L1 Compaction - Dropbox’s baseline strategy, L1, treats compaction as a packing problem. - It selects: - A highly filled host volume with available space. - Donor volumes whose live data fits into that space. - Live blobs from the donors are written into a new volume, eventually leaving the donors empty and removable. - L1 is simple, fast, and limits placement risk and metadata changes. - However, each run can read tens of GiB while typically producing only one densely packed volume. - Fewer than one complete volume is reclaimed on average because only donor volumes are fully drained. - This works well when volumes are already near full, but performs poorly when storage overhead is concentrated in many severely under-filled volumes.

Read original(opens in new tab)
cloudflare3 min readCurated summary

Why we-re rethinking cache for the AI era

AI traffic is fundamentally changing how CDNs should think about caching. Unlike human visitors, AI crawlers make broad, high-volume, often sequential requests for long-tail content, creating low reuse and substantial cache churn. Cloudflare argues that traditional LRU-based caching and techniques such as prefetching are increasingly poorly suited to this traffic, forcing operators to rethink cache design if they want to support AI access without harming human performance. ## Why AI Traffic Is Different - Automated traffic accounts for 32% of Cloudflare’s network traffic, including crawlers, scrapers, and AI assistants. - AI agents often: - Send many requests in parallel. - Scan large portions of a website sequentially. - Request rarely visited or loosely related pages. - Fetch documentation, images, and articles from many sources. - AI crawlers represent approximately 80% of self-identified AI bot traffic. - Most single-purpose AI bot traffic is associated with model training, with search-related crawling a distant second. ## The Three Defining Characteristics of AI Crawlers - **High unique URL ratio:** More than 90% of pages observed in large-scale Common Crawl datasets are unique by content. - **Content diversity:** Different crawlers target different materials, including source code, technical documentation, media, and blog posts. - **Crawling inefficiency:** Many requests lead to 404 errors or redirects because of poor URL handling. - AI crawlers generally lack browser-side caching and shared session behavior, so independent crawler instances may repeatedly appear as new visitors. - They can also repeatedly revisit content while iteratively refining search results, but each iteration still tends to fetch mostly new pages. ## How AI Crawling Disrupts Traditional Caches - Conventional CDN caching keeps frequently requested content available near users and evicts less recently used objects when storage fills. - Cloudflare uses an LRU (least recently used) policy, but broad AI scans introduce large numbers of low-reuse objects. - These objects can evict content that human visitors are more likely to request. - AI-driven long-tail access increases cache misses and sends more requests back to origin servers. - Cache speculation and prefetching become less effective because crawler access patterns are difficult to predict. - Higher miss rates can cause: - Slower responses. - Increased origin-server load. - Greater egress costs. - Reduced cache hit rates for human traffic. ## Implications for Website Operators - Operators face a tradeoff between optimizing infrastructure for human visitors and accommodating AI crawlers. - Some organizations may want to encourage AI access: - Developers may want documentation represented in AI models. - E-commerce companies may want product information included in LLM search results. - Publishers may seek compensation through systems such as pay-per-crawl. - The challenge is supporting useful AI traffic without allowing it to degrade the cache performance experienced by human users. Cloudflare’s analysis, conducted with ETH Zurich researchers, suggests that CDN caching strategies need to evolve beyond traditional assumptions about popularity and reuse. Cache systems designed specifically for AI-era traffic may need to isolate crawler workloads or otherwise prevent broad, low-reuse scans from displacing content valuable to human users.

Read original(opens in new tab)
stripe3 min readCurated summary

Insights from Shoptalk 2026: How agents are changing retail

Agentic commerce is already reshaping retail, especially product discovery, embedded checkout, and customer engagement across AI-powered surfaces. However, retailers still lack a common strategy for managing product data, choosing channels, and deciding between first-party and third-party agent experiences. The article argues that success will depend not only on agent-compatible infrastructure, but also on strong brands, unified customer data, and frictionless checkout. ## Agentic Commerce Needs a Standard Framework - Retailers are experimenting with where to begin, which partners to use, and how to syndicate accurate product data across AI platforms. - Search and discovery are changing quickly: - Sephora is using loyalty data in its ChatGPT app to personalize recommendations and highlight benefits such as samples and free shipping. - OpenAI reported that more than half of its searches are discovery-oriented, with 70% containing detailed constraints or context. - AI agents increasingly function as storefronts. Brands that are not discoverable through these systems risk losing visibility. - Direct product feeds are becoming important because they provide agents with more structured and current information than web crawling. - Stripe’s Agentic Commerce Suite allows retailers to connect catalogs and syndicate them across supported agents without building separate integrations. - Many companies are using test-and-learn programs to measure how products are discovered, recommended, and purchased through AI surfaces. ## Commerce Is Expanding Beyond Chat Interfaces - Agentic commerce is appearing across: - Embedded checkout - Product discovery - Customer service - Catalog enrichment - Post-purchase systems - Meta demonstrated a Facebook checkout flow using the Agentic Commerce Protocol, allowing shoppers to move from an ad to product information, AI-generated review summaries, and in-app purchase. - New consumer brands may increasingly be built on agentic infrastructure, reducing customer acquisition costs and dependence on standalone websites. - Retailers must decide how much to invest in first-party experiences versus third-party agents across categories such as fashion, beauty, and home goods. - The market is unlikely to be controlled by one large language model or channel; instead, commerce will spread across many specialized applications and surfaces. ## Brand Trust Becomes More Important - As AI simplifies comparison shopping, trust, consistency, and emotional connection will play a larger role in brand selection. - New Balance is emphasizing consistent quality, store improvements, and better-trained associates rather than relying primarily on discounts. - Tapestry is studying Gen Z to maintain Coach’s relevance, while Victoria’s Secret is focusing on comforting, confidence-building store experiences. - Stitch Fix is using first-party customer data to power Stitch Fix Vision, an AI tool for personalized outfit visualization. - Retailers will need unified customer data and systems that preserve identity and context across websites, stores, apps, and AI agents. ## Checkout and Commerce Infrastructure Remain Fundamental - Customers arriving through agent-driven journeys may be ready to buy and less tolerant of checkout friction. - Stripe says its Optimized Checkout Suite selects payment methods using more than 100 signals and typically increases conversion by 2%–3%. - Core requirements remain unchanged: - Fast, branded checkout - Relevant payment methods - Effective fraud prevention - Connected online, in-store, and in-app commerce data - Stripe’s Agentic Commerce Suite is designed to let businesses connect their catalog and commerce systems once, then expand into compatible agents and channels. Retailers should begin with structured product data, measurable experiments, unified customer systems, and a frictionless checkout experience. Agentic channels are developing rapidly, but durable brand value and strong commerce fundamentals will remain essential as those channels multiply.

Read original(opens in new tab)