Techlist.io - Korean Tech Blog Curator

cloudflare3 min readCurated summary

The post-quantum EO is an important milestone. Now it’s time to get to work

The post welcomes Executive Order 14409 as a major step toward post-quantum security, setting federal deadlines of 2030 for encryption and 2031 for authentication. It argues that the threat timeline has accelerated and that organizations must begin migration now, especially to prevent “harvest-now-decrypt-later” attacks. Cloudflare views the order as a strong foundation but believes agencies need clearer guidance and a coordinated migration roadmap. ## Federal Post-Quantum Requirements - The order primarily covers: - **High Value Assets (HVAs)**, such as systems containing sensitive employee records, classified intelligence, or federal financial data. - **High impact systems** rated “high” under FIPS 199, where compromise could cause severe harm. - Key deadlines include: - **July 2026:** Agencies name a post-quantum migration lead. - **September 2026:** Agencies inventory HVAs and high-impact systems, create migration plans, and submit them to OMB and the National Cyber Director. - **December 2030:** Key establishment must use post-quantum cryptography. - **December 2031:** Digital signatures and certificates must use post-quantum cryptography. - National Security Systems are excluded from these deadlines and remain on a separate NSA-managed schedule. - The order directly binds federal agencies, not state and local governments, critical infrastructure, academia, or civil society. ## Encryption and Authentication Are Separate Migrations - **Post-quantum encryption** protects key establishment and should begin immediately. - It prevents attackers from collecting encrypted data now and decrypting it after quantum computers become capable of breaking RSA and elliptic-curve cryptography. - This is especially important for government, financial, healthcare, defense, and telecommunications data with long-term value. - **Post-quantum authentication** protects digital signatures, certificates, software signatures, and system access. - It prevents future quantum computers from impersonating servers or forging trusted signatures. - Its primary threat emerges once a cryptographically relevant quantum computer exists. - The order’s 2031 authentication deadline suggests the U.S. government considers an operational quantum computer around that period a meaningful possibility. ## Standardized Cryptography Over Quantum Key Distribution - The order emphasizes NIST-standardized post-quantum algorithms. - The authors support this focus because Quantum Key Distribution requires specialized hardware and dedicated physical links, making it unsuitable for Internet-scale deployment. - Cloudflare reports that more than two-thirds of browser traffic reaching its network already uses post-quantum encryption. - Its Cloudflare One platform supports post-quantum protection across TLS, MASQUE, and IPsec, while broader post-quantum authentication deployment is still beginning. ## Why Authentication Is More Difficult - Post-quantum ML-DSA signatures are larger than traditional signatures, potentially reducing performance in systems such as short-lived TLS connections. - Cloudflare is working with Google Chrome on Merkle Tree Certificates to reduce this TLS overhead. - Authentication requires coordinated upgrades across a larger ecosystem: - Clients and servers - Certificate authorities - Certificate transparency logs - Root stores - Web browsers - By comparison, post-quantum key establishment is already more widely available and easier to deploy incrementally. Organizations should begin with asset inventories, risk assessments, and post-quantum key-establishment upgrades now rather than waiting for the federal deadlines. Authentication migration should also start early because its broader dependency chain and larger signatures make it the more complex transition.

Read original(opens in new tab)
meta3 min readCurated summary

How Meta Engineered Ultra-Narrow Batteries for AI Glasses

Smart glasses require batteries that fit inside extremely narrow temple arms while powering cameras, speakers, AI processing, and displays. Meta addressed this limitation by developing ultra-narrow steel-can cells, including batteries as thin as 7 mm, with redesigned electrode structures and tighter manufacturing tolerances. The approach increased capacity and peak-power performance while enabling different configurations across successive generations of Meta’s wearables. ## Why Traditional Pouch Cells Fall Short - Common in phones and laptops, pouch cells are difficult to shrink and reshape. - Folding, manufacturing tolerances, and wasted internal volume are especially costly in glasses. - Small pouch cells may also struggle to deliver peak power when several features operate simultaneously, such as recording video while an AI model processes a request. - Smart glasses instead need rigid, precisely shaped batteries that use nearly every available micron. ## Designing Ultra-Narrow Steel-Can Cells - Steel-can batteries are established in products such as watches and power tools, but Meta needed unprecedented widths down to 7 mm. - Engineers replaced the conventional wound “jelly roll” electrode with die-cut, stacked layers. - This architecture reduces impedance, helping prevent power drops or brownouts during simultaneous high-demand tasks. - Steel cans maintain their shape to approximately 100 microns, preserving usable space and improving energy density in narrow cells. ## Increasing Capacity Through System Design - The second-generation Ray-Ban Meta battery increased from 160 mAh to 210 mAh, about a 30% capacity increase. - The glasses nevertheless claimed roughly twice the runtime because of broader hardware and software improvements. - Gains came from better power management, tighter firmware control, and a form factor that accommodated a larger cell. - This demonstrates that battery life depends on the entire system, not chemistry alone. ## Managing Multiple Batteries and Higher Power Demands - Oakley Meta Vanguards use one battery in each temple arm. - Although the cells are symmetrical, the electrical loads are not evenly distributed. - Engineers had to address cross-charging risks and coordinate battery sequencing during startup and shutdown. - Meta Ray-Ban Display glasses created a sustained power demand because the display draws energy continuously rather than in short bursts. - They use a 248 mAh steel-can cell, the largest in Meta’s lineup. ## Scaling the Technology - Meta’s narrow steel-can design could support other wearable form factors beyond smart glasses. - The company is working to scale production across multiple vendors and build a more resilient supply chain. - Developing these cells required coordination among electrical, mechanical, firmware, manufacturing, and global collaboration teams. Meta’s steel-can battery technology shows how wearable battery improvements come from rethinking both cell construction and overall device engineering. For future compact wearables, precisely shaped, low-impedance cells combined with system-level power optimization offer a practical path to longer runtime and more demanding features.

Read original(opens in new tab)
github1 min readCurated summary

I automated my job (and it made me a better leader)

Ashley Willis is GitHub’s Senior Director of Developer Relations, where she focuses on open source, community, and developer advocacy. Her work combines leadership, accessibility, and inclusion, with an emphasis on making technology more human and building resilient teams. ### Leadership and Advocacy - Leads developer relations at GitHub. - Advocates for developers and open-source contributors. - Amplifies underrepresented voices in technology. ### Community and Accessibility - Builds supportive, inclusive spaces for contributors. - Focuses on creating tools that genuinely serve their users. - Works at the intersection of leadership, advocacy, and accessibility. Overall, Willis’s career centers on strengthening developer communities and making technology more inclusive, accessible, and human.

Read original(opens in new tab)
naver1 min readCurated summary

SNOW’s Journey to Adopting Automatic Sharding

The content is a minimal NAVER D2 landing-page outline rather than a substantive tech blog post. It lists links to D2 News, About D2, NAVER Developers, DEVIEW, OpenSource, and D2 STARTUP FACTORY, followed by a NAVER copyright notice. No technical argument, explanation, or conclusion is provided. ## NAVER D2 Sections - **Main navigation** - Hello world - D2 News - About D2 - NAVER Developers - DEVIEW - OpenSource - D2 STARTUP FACTORY - **Copyright** - Copyright © NAVER Corp. All Rights Reserved. There is not enough technical content to derive a practical recommendation or detailed summary.

Read original(opens in new tab)
kakao3 min readCurated summary

Solving Social Problems with AI Beyond Development

The first “SSAFY X Kakao Tech Bootcamp AI Hackathon” brought together 90 trainees from 12 teams to use AI for solving real social problems. Rather than focusing only on coding competition, the event emphasized public value, practical service prototypes, expert feedback, and collaboration across different training programs. It demonstrated that future developers need both technical ability and the capacity to work with others on meaningful problems. ## Connecting Kakao and Samsung’s Developer Programs - Held June 13–14 at Kakao’s AI Campus in Yongin. - Organized jointly by Kakao Tech Bootcamp and Samsung’s SSAFY program. - Participants came from two major digital-training initiatives supported by Korea’s K-Digital Training program. - The event aimed to create opportunities for collaboration and growth among future AI developers. ## Applying AI to Everyday Social Problems - Teams selected challenges from the government’s “Top 10 AI Projects for People’s Livelihoods.” - Topics included: - Small-business support - Voice-phishing prevention - Child and youth protection - Maritime safety - Over two intensive, sleepless days, teams: - Defined a specific social problem - Designed solutions from the user’s perspective - Built AI-powered service prototypes - The hackathon stressed that AI’s value depends not only on technical advancement, but also on how effectively it improves society. ## Practical Mentoring from Government and Industry - Officials from agencies including the National Police Agency, Ministry of Justice, and Ministry of Gender Equality and Family provided policy and field expertise. - Kakao developers delivered lectures and technical mentoring based on real-world service development. - Teams refined their ideas through questions, feedback, and discussions with experts. - This allowed trainees to connect classroom learning with actual policy and operational challenges. ## Collaboration Across Different Backgrounds - Kakao Tech Bootcamp and SSAFY use different educational approaches, giving participants varied experiences and strengths. - Teams worked with people they had not previously met and actively discussed how to incorporate AI into their products. - Participants discovered new perspectives and solutions by sharing their knowledge. - Many came to recognize communication and teamwork as essential skills alongside technical competence. ## Projects and the Future Developer Ecosystem - Five teams received awards after the final presentations. - The Ministry of Employment and Labor award went to “Golden Time” for **DRIFT**, an AI service supporting maritime rescue when communications are unavailable. - Kakao’s CEO award went to “SSAIKA” for **Mindam**, an AI-based civil complaint intake and processing service. - Other awards were presented by Samsung Electronics, the Korea Chamber of Commerce and Industry, and the Korea Radio Promotion Association. - Although the total prize money was 15 million won, the article identifies hands-on experience solving social problems as the participants’ more important achievement. - Kakao has trained more than 660 digital professionals since joining the K-Digital Training initiative in 2022. The hackathon suggests that AI education should combine technical training with real-world projects, expert guidance, and cross-organizational collaboration. Kakao plans to expand these practical opportunities to support developers who can turn technology into social value.

Read original(opens in new tab)
stripe3 min readCurated summary

Four travel and hospitality trends from HITEC 2026

Hospitality’s AI opportunity is growing, but most operators lack the data, infrastructure, and operational systems needed to turn investment into measurable returns. AI is reshaping how travelers discover and book hotels, while fragmented data and outdated payment systems create lost revenue and guest frustration. The strongest strategy is to connect accurate data, intelligent workflows, and seamless payments so technology improves the experience without becoming visible to guests. ## AI Is Changing the Direct-Booking Battle - Hotels historically relied on SEO to compete with OTAs such as Expedia and Booking.com. - AI-generated search answers are reducing traditional website traffic: - 65% of Google searches with AI Overviews end without a click. - The figure rises to 78% on mobile. - Traditional search traffic is declining by about 25%. - AI systems prioritize accurate, structured, machine-readable information rather than keyword density and backlinks. - More than 90% of accommodation websites are reportedly undetected by AI models. - Hotels should audit whether AI tools can correctly describe: - Room categories - Amenities - Policies and cancellation terms - Local context - Real-time availability - Winning direct bookings will require both AI discoverability and a modern checkout experience supporting local currencies, payment methods, and fraud protection. ## Hospitality AI Is Held Back by Fragmented Data - Only about 25% of hospitality businesses are actively scaling AI, and fewer than 10% are considered “AI future-built.” - Property management, CRM, loyalty, food and beverage, and payment systems often operate in silos. - Incomplete data weakens: - Personalization - Guest profiles - Financial reconciliation - Operational decision-making - The main challenge is not building AI features but operationalizing them reliably in real workflows. - Successful examples connect live data to timely actions: - Delta’s AI concierge uses customer and operational data to provide context-aware support. - Wynn’s revenue managers receive predictive alerts and recommended actions. - For most operators, better data connectivity matters more than using a more advanced AI model. ## Payment Friction Directly Affects Revenue - Payments are increasingly viewed as a competitive capability rather than a back-office commodity. - Survey findings cited in the article include: - 90% of executives consider payments important to growth. - 37% say limited payment options most harm the guest experience. - 58% report that fraud tools block legitimate transactions. - 74% say fragmented systems create excessive reconciliation work. - Guests may abandon a hotel when their preferred payment method is unavailable, shifting the booking to an OTA that supports it. - Modern payment infrastructure allows smaller operators to offer international payment methods and currencies without building large in-house teams. ## Invisible Technology Creates the Best Guest Experience - Guests have little tolerance for technology failures and may simply avoid returning rather than complain. - Effective hospitality technology should anticipate needs without drawing attention to itself. - The desired experience includes details such as: - A room set to the guest’s preferred temperature - Familiar television channels - Preferred pillow firmness - Hospitality is moving from remembering information guests explicitly provided to predicting preferences based on connected guest data. Operators should prioritize clean, connected data, AI systems tied to real operational actions, and flexible payment infrastructure. The goal is not to add AI for its own sake, but to make booking and stays more seamless while quietly improving revenue, efficiency, and guest loyalty.

Read original(opens in new tab)
netflix3 min readCurated summary

Toward More Controllable AI Video Editing: An Early Research Exploration at Netflix

Netflix explores AI video-editing tools designed to preserve artists’ creative control rather than regenerate entire clips indiscriminately. The research addresses two major problems: unintended changes to untouched footage and physically implausible results when objects are removed. Its proposed systems, Vera and VOID, generate targeted edits while preserving scene identity, performance, and continuity. ## Challenges in Generative Video Editing - Full-video regeneration can unintentionally change: - Actors’ identities and performances - Backgrounds and objects - Important scene details - Object removal often produces unnatural results because models erase the target without reconstructing realistic motion and physical interactions. - Professional editors need precise control over what changes and what remains untouched. ## Vera: Layered Video Diffusion - Vera generates: - An edit layer containing the requested visual change - An alpha matte defining where that change should appear - These layers are composited with the original footage, leaving pixels outside the edited region intact. - The approach supports tasks such as: - Adding objects - Changing backgrounds - This layered design helps preserve original identities, performances, and details. ## Training Dataset - Netflix created a custom dataset because existing public datasets lacked high-quality layered video data. - The dataset contains 486,000 frames at 832×480 resolution. - It includes: - **Synthetic composites:** Foreground objects with alpha mattes placed over generated backgrounds. - **Realistic single-object videos:** Real footage processed with segmentation, matting, background generation, and human review. - **Realistic multi-object videos with effects:** Objects isolated along with shadows, reflections, and other scene effects. ## Vera’s Model Architecture - Vera uses a Mixture-of-Transformers design with three specialized DiTs for: - The edit layer - The alpha matte - The composite video - Each branch has its own attention projections and feed-forward weights, allowing specialization while joint attention enables communication between layers. - The model is initialized from a pretrained text-to-video model. - Additional embeddings and input layers help distinguish source-video, mask, alpha, and composite information. ## Evaluation and Results - Netflix tested Vera on: - 72 object-addition video-prompt pairs - 69 background-change pairs - The benchmark included varied motion speeds, camera movements, object counts, and scene complexity. - Evaluation measured: - Preservation of untouched content - Compliance with text instructions - Temporal and per-frame video quality - Vera-1.3B and Vera-14B substantially outperformed existing methods on content preservation while achieving comparable instruction-following and visual quality. Netflix’s research favors localized, layered editing over unrestricted video regeneration. Vera demonstrates how separating edits from original footage can make generative tools safer and more controllable for professional workflows; the accompanying VOID research aims to apply similar principles to physically plausible object and interaction removal.

Read original(opens in new tab)
line4 min readCurated summary

From Manual to AI Prompt Tuning: Genetic Algorithm–Based Automated Optimization and Acceleration

LY Corporation automated LLM prompt tuning with the GEPA genetic algorithm, reducing a process that previously took days or weeks to roughly one hour. GEPA evolves prompt candidates using evaluation scores and natural-language feedback, allowing it to improve prompts without manually inspecting every output. The approach was applied to Yahoo! JAPAN Search’s AI responses for health and medical queries, balancing policy compliance with improved readability. ## Challenges of Manual Prompt Tuning - Each prompt change requires repeated output generation and human review. - Practical tuning knowledge often remains with individual engineers and is difficult to document or explain. - The cycle of editing, generating, and evaluating responses can take days or weeks. - Model changes and version updates can alter output quality, requiring repeated retuning. - Manual effort leaves less time for defining evaluation criteria, judging quality, and verifying policy compliance. ## Automated Prompt Optimization Approaches - **Reinforcement learning:** Learns prompt-generation policies from scalar rewards, such as with GRPO. - **Bayesian optimization:** Efficiently searches candidate instructions and few-shot examples, as in MIPROv2. - **Genetic algorithms:** Iteratively evolve a population of prompt candidates, as in GEPA. - Genetic methods are well suited to discrete, natural-language prompts because they can use natural-language reflection to identify problems and propose improvements rather than relying only on numerical rewards. ## How GEPA Works - Generates and evaluates multiple prompt candidates. - Uses **Reflective Prompt Mutation** to analyze execution results and create improved instructions. - Applies Pareto-frontier selection to preserve candidates that perform well across multiple evaluation dimensions. - Repeats the process over several to dozens of generations until prompts converge toward the evaluation objectives. - The article notes that GEPA has reportedly outperformed previous optimization methods, including results presented at ICLR 2026. ## Implementation with DSPy - DSPy allows prompt optimization to be controlled programmatically. - A task is defined as a DSPy module with a signature containing input and output fields. - The signature’s docstring becomes an instruction for the LLM. - GEPA rewrites this instruction during optimization. - Separate models can be assigned for: - Task inference - Output evaluation - Reflection and prompt improvement ## Designing the Evaluation Function - GEPA requires an overall scalar score, even when quality is judged across multiple criteria. - Individual scores can be assigned to dimensions such as accuracy, completeness, and style, then normalized and averaged. - The evaluator can also return natural-language feedback through `dspy.Prediction(score=..., feedback=...)`. - Feedback explains why a candidate was penalized, giving GEPA a clearer direction for improvement than a score alone. - Evaluation can use: - LLM-as-a-Judge - Gold answers or labels - Rule-based correctness checks - In the example, an evaluator scores three criteria from 0 to 10, averages them into a single score, and passes the explanation to GEPA for reflection. ## Yahoo! JAPAN Search Health and Medical Queries - Health-related answers must follow stricter policies than general search responses. - Requirements include: - Avoiding definitive medical diagnoses or severity judgments - Matching wording to the strength of available evidence - Recommending medical consultation appropriately - Limiting responses to general explanations where necessary - The project pursued two goals simultaneously: - Satisfy medical and health-policy requirements. - Apply readable Markdown formatting, including headings, lists, and emphasis. - Improving one goal manually could easily damage the other, making automated optimization attractive. ## Applying GEPA to the Production Task - The system takes a search query as input and generates an AI answer. - The initial prompt combined an existing general-purpose prompt with additional health and medical policy instructions. - GEPA rewrote and optimized the instruction section rather than requiring engineers to manually redesign the entire prompt. - The optimization aimed to preserve policy compliance while improving structure and readability. Overall, GEPA with DSPy provides a practical way to shorten prompt-tuning cycles and make the improvement process more reproducible. Its effectiveness depends heavily on carefully designed evaluation criteria and meaningful natural-language feedback, especially for high-risk domains such as medical information.

Read original(opens in new tab)
line4 min readCurated summary

Total Capacity Exceeds 1 EB! How Do You Connect Two HDFS Systems with Different Histories? Challenges and Design Decisions in Data Platform Integration

LY Corporation’s Tech-Verse 2026 article examines how its former LINE and Yahoo Japan organizations operated HDFS platforms exceeding one exabyte in total capacity. Although both platforms used Hadoop at scale, their access models, namespace architectures, permission systems, and operational practices differed substantially. The article argues that large-scale data platforms must be designed around actual usage patterns, not just storage capacity, and previews how the two environments were later connected after organizational integration. ## Different Operating Models - Former LINE built a unified analytics environment for broad, cross-departmental data use. - Users accessed data through a web portal that managed catalogs, permissions, and role-based approval workflows rather than interacting directly with HDFS or Apache Ranger. - BI tools, reporting systems, and ETL pipelines supported diverse use cases, but integrating multiple existing clusters made operations complex. - Former Yahoo Japan evolved from a limited-purpose Hadoop deployment into a company-wide platform. - Its user interfaces and access methods were intentionally restricted, making the system easier to stabilize and support. - Yahoo Japan retained HDFS-style POSIX permissions, which limited flexibility compared with newer data-governance models. ## Different HDFS Architectures - Both platforms split their storage across multiple namespaces to overcome NameNode scaling limitations. - Each namespace used two to four NameNodes for redundancy, but the way namespaces were exposed differed: - LINE used **ViewFS**, requiring clients to maintain mount-table configurations. - Yahoo Japan used **Router-Based Federation (RBF)**, allowing routers to direct requests to the correct NameNode. - LINE shared DataNodes across namespaces, improving resource efficiency but increasing operational complexity. - LINE also had mixed NameNode and DataNode versions because several legacy platforms had been consolidated. - Yahoo Japan’s RBF design included Observer NameNodes to distribute read load. - These architectural differences affected later federation work, including connection endpoints, client configuration, network reachability, and permission management. ## Capacity and Network Challenges at LINE - Data growth exceeded forecasts, causing HDFS capacity shortages before new servers could be delivered. - Older servers were temporarily reused, leading to frequent node additions and removals. - Large changes in node count triggered HDFS Balancer activity and block redistribution, generating substantial network traffic. - Network engineers therefore had to coordinate closely with the Hadoop operations team during infrastructure changes. ## NameNode Metadata and Small-File Problems - As file and block counts increased, NameNode heap usage and processing load grew. - Larger heaps also increased garbage-collection times, making NameNodes slower and less stable. - The team analyzed regularly dumped FSImage data stored in Hive tables to identify users, paths, file counts, block counts, and data volumes. - They prioritized tables containing many small files where file compaction could significantly reduce block counts without requiring data deletion or schema changes. - File merging reduced both NameNode metadata pressure and the number of HDFS operations, improving response times for jobs. ## Namespace-Specific Load Patterns - Different namespaces experienced different types of pressure. - Temporary-file namespaces saw frequent Spark staging-file creation and deletion, producing repeated metadata updates requiring NameNode write locks. - When HDFS Balancer moved blocks, read-lock activity increased and could delay file creation and deletion. - Increasing Balancer parallelism initially worsened contention. - The team reduced parallelism to a level compatible with available DataNode disk capacity, balancing migration speed against cluster impact. ## Connecting the Two Platforms - Organizational integration introduced additional challenges beyond storage: - Determining which platform and entry point users should access - Reconciling different permission-management models - Establishing data-transfer paths between platforms - LINE’s ViewFS model depends on correctly distributed client mount tables. - Yahoo Japan’s RBF model depends on reliable, scalable, and reachable router infrastructure. - These differences directly influence cross-platform data movement, including transfers using DistCP. Large HDFS environments should be managed according to real workload behavior, namespace characteristics, and operational dependencies. Capacity planning alone is insufficient; teams should monitor metadata growth, small-file patterns, lock contention, balancing traffic, network effects, and the distinct access models of each platform.

Read original(opens in new tab)
toss4 min readCurated summary

5. Technical Writer, A Decision to Disappear

Toss’s technical writing team argues that documentation is essential context for AI, but manually maintaining thousands of documents is impossible with only three technical writers serving roughly 4,000 people. Their solution is to automate the technical writer’s work by teaching AI the team’s implicit standards and embedding those standards into reusable Skills. The initial system supported document creation and review, but adoption remained low because users still had to install, invoke, and supply information to the AI manually. ## Why Toss Wanted to Automate Technical Writing - Documentation gives AI the organizational context it needs to work effectively. - Toss has approximately 4,000 employees but only three technical writers. - Reviewing documents individually does not scale, especially in a fast-moving organization where features change or disappear before documentation is complete. - The team’s goal to “eliminate technical writers” means transferring routine writing and editing work to AI, not abandoning documentation quality. ## Teaching AI Technical Writing Principles - The team analyzed existing technical writing review comments to identify how writers evaluate documents. - Existing writing guidelines were converted into explicit principles, such as: - Focus each page on one subject. - Present value before implementation details. - Each principle was supplemented with incorrect and correct examples so AI would understand the intent rather than apply rules mechanically. - Common document types were converted into templates. - Templates include: - Instructions explaining what each section should contain. - `(required)` markers for information that must not be omitted. - For example, an ADR template requires an overview, context, considered alternatives, decision, and rationale, while also allowing optional sections such as expected outcomes and related references. ## Skill for Writing New Documents The document-writing Skill reproduces the four stages a technical writer typically follows: - **Clarify the purpose:** Ask about the project, document goal, audience, level of detail, source materials, and expected structure. - **Design the structure:** Use a standard structure or select a relevant template, such as onboarding guides, meeting notes, or PRDs. - **Write the content:** Apply technical writing and MDX rules while using templates as structural guidance. - **Review the draft:** Check for awkward wording, missing information, and other quality issues. The Skill also distinguishes between required and optional template sections: - Required sections remain in the draft even when source information is incomplete. - Missing information is represented with questions or comments rather than guesses. - Optional sections are omitted when there is not enough source material to complete them. ## Skill for Reviewing and Improving Documents - The team initially converted past review comments into a checklist. - This produced poor results: AI overlooked important issues while generating unnecessary comments. - The problem was that good writing follows relatively stable principles, whereas bad writing can fail in many different ways. - The revised workflow lets AI independently: - Read the technical writing principles. - Analyze the document. - Identify violations. - Explain the issue and suggest revised wording. - Perform a final checklist-based review. - Previous review comments are now used as examples of how principles apply, rather than as a rigid list of required findings. - One example principle requires descriptions of parameters or properties to include their meaning, accepted format, and usage example—not merely a type such as `date: string`. ## Low Adoption Revealed a Usability Problem - Despite creating both Skills, the team found that few employees used them. - Users still had to: - Download and install the Skill manually. - Understand CLI-based setup, which was unfamiliar to non-developers. - Remember to invoke the Skill whenever they began writing documentation. - Find and provide all relevant source materials themselves. - The team concluded that improving the AI’s capabilities was not enough; the workflow also had to reduce the effort required from users. The main lesson is that AI-based documentation succeeds only when organizational knowledge, writing principles, and templates are encoded clearly—and when the system is integrated into everyday work so employees do not have to remember to use it or prepare everything manually.

Read original(opens in new tab)
toss4 min readCurated summary

6. Beyond Tools: Standards and Responsibility

Toss’s commerce domain found that reliable organizational knowledge cannot be created by writing more documents or adding automation alone. Sustainable knowledge management requires clear standards for what should be documented, who owns it, how it is maintained, and which sources can be trusted. The proposed solution combines AI-assisted documentation with domain-level responsibility and company-wide governance. ## The Limits of Writing Alone - A commerce wiki consolidated terminology, onboarding material, code references, and policy documents. - This reduced confusion over terms such as “seller” and “store” and gave teams a shared starting point. - However, product and policy changes happened faster than one Technical Writer could document them. - Important knowledge also appeared in policy changes, temporary experiments, and chat discussions that were difficult to track manually. ## Why Culture and Participation Were Not Enough - The team promoted documentation through: - A weekly “Commerce Wiki News” newsletter - A policy-question channel and bot - AI documentation workshops - A documentation guild - These efforts increased requests, wiki usage, and adoption of official terminology. - Participation rarely continued beyond an individual’s first document because documentation was not part of normal work priorities. - Writers lacked guidance on: - What information to preserve - How much detail to include - Which audience to target - How to verify whether a document was correct - Documentation became sustainable only when it was treated as a team responsibility embedded in existing workflows. ## AI Automation Reveals the Governance Problem - AI now creates draft documents nightly from two signals: - Product deployment and policy-change announcements - Questions that the commerce Q&A bot cannot answer - AI gathers supporting context and produces drafts, while humans verify the evidence and approve them. - This removes the burden of starting documents from a blank page. - Automation also exposed new problems: - Duplicate or overlapping documents - Unclear authoritative sources - Outdated policies being used in bot answers - Difficulty distinguishing current policies from completed experiments - Automation can collect and draft information, but it cannot decide who owns a policy or whether a document should still be trusted. ## Knowledge Standards and Governance - The focus shifted from “How do we create more documents?” to “How do we create knowledge people can trust?” - Toss’s knowledge-management standards state that teams should: - Preserve recurring questions, important decisions, and information needed by newcomers. - Organize knowledge so both people and AI can find it. - Connect documents to work tools such as Q&A bots and GitHub. - Assign owners and review cycles to keep information accurate and current. - Possible classification systems include: - **Technical layers** for teams with clear data or system flows - **Service domains** for teams responsible for multiple service areas - **Functional units** for systems with distinct feature boundaries - Information becomes organizational knowledge only when it helps people understand situations and make better decisions, with sufficient context and verification. ## The Role of the Knowledge Committee - The Knowledge Committee defines and maintains company-wide documentation standards and resolves conflicts between organizational rules. - Unlike a voluntary guild, it has designated members with decision-making authority. - Governance operates at two levels: - The Technical Writing Chapter manages shared standards for sources, ownership, document status, and lifecycle. - Individual domains decide how those standards apply locally, including ownership, update schedules, and retirement rules. - This balance prevents both inconsistent practices across teams and overly centralized rules that ignore local realities. - For example, commerce teams may need separate handling for permanent deployments and temporary experiments so expired policies do not remain authoritative. The practical recommendation is to treat knowledge management as an operating system for the organization, not a documentation project. AI can reduce the effort of capturing knowledge, but clear ownership, review processes, lifecycle rules, and governance are necessary to keep that knowledge reliable and useful.

Read original(opens in new tab)
datadog3 min readCurated summary

How we migrated a live routing system using AI-assisted refactoring

Stream Router evolved from a small configuration file into a critical control-plane service routing Datadog’s massive metrics workload. Its original FoundationDB key-value model eventually hit transaction-size and performance limits because relational relationships were reconstructed in application code. Datadog redesigned the system around PostgreSQL and DuckDB, using AI-assisted, test-driven refactoring to accelerate the migration without disrupting production traffic. ## Stream Router’s Role in Datadog’s Metrics Pipeline - Datadog processes more than a hundred trillion events per day. - Stream Router determines which Kafka cluster, topic, partitions, and sharding strategy should handle each datapoint. - It serves both producers and queriers but does not process Kafka messages itself. - Routing decisions change frequently as infrastructure evolves, making correctness and historical tracking essential. ## From Configuration File to Control Plane - In 2016, routing was managed through a small configuration file distributed to services. - As the platform grew, the file expanded to thousands of lines and required manual edits and rollouts. - Stream Router replaced this workflow with: - A centralized gRPC service - API-managed routes - Automated, gradual rollouts - The write path used FoundationDB, while the read path served static RocksDB snapshots restored into memory. - This eventually became a bottleneck as routing tables and operational changes grew larger. ## Why the Key-Value Model Stopped Scaling - Routes reference streams and sharding strategies, while rules reference routes. - These relationships are inherently relational and require cross-entity validation. - The KV implementation loaded tens of thousands of records into application processes and reconstructed database-like relationships in code. - Some operations exceeded FoundationDB transaction-size limits. - Moving to PostgreSQL without changing the access patterns would not solve the issue; certain operations were estimated to require 45 minutes because of thousands of sequential database round trips. - The fundamental problem was the data model and application logic, not simply the choice of database. ## Designing the New Storage Architecture - The team redesigned the schema manually before using AI tools. - The relational model introduced explicit foreign keys between: - Streams - Sharding strategies - Routes - Rules - PostgreSQL was selected for the write path because it provided the required relational semantics and transaction model. - DuckDB was selected for the read path because: - It is embeddable and suitable for snapshot-based serving - It supports array columns - Its SQL dialect is closely compatible with PostgreSQL - Shared query logic could therefore work across both storage engines. ## AI-Assisted Refactoring - Claude and Cursor were used to accelerate a systematic, test-driven migration. - For each method, developers supplied: - The old implementation - The new schema - A failing test - AI generated an initial implementation, while tests determined whether it was correct. - The models assisted with method-level refactoring rather than autonomously designing the architecture. - Human expertise remained central to schema design, migration strategy, and evaluating system-level risks. ## Foundations for a Safe Migration - The migration benefited from infrastructure already present at Datadog. - Stream Router’s storage layer was isolated behind an internal `Controller` interface. - This modularity helped contain storage changes and enabled incremental refactoring. - Existing tests and clear boundaries provided confidence in generated implementations while production traffic continued. The central lesson is that AI was most effective as an accelerator inside a disciplined engineering process. A well-designed relational schema, modular storage abstraction, and failing tests provided the safety mechanisms; AI helped implement the resulting changes faster, but did not replace human architectural judgment.

Read original(opens in new tab)
aws3 min readCurated summary

Run isolated sandboxes with full lifecycle control: AWS Lambda introduces MicroVMs | Amazon Web Services

AWS Lambda MicroVMs provide isolated, stateful execution environments for running untrusted user- or AI-generated code without managing virtual machine infrastructure. Built on Firecracker, they combine VM-level isolation, near-instant startup and resume, and persistent memory and disk state. The post concludes that MicroVMs fill the gap between slow, isolated VMs, less-secure containers, and stateless event-driven Lambda functions. ## The Need for Isolated, Stateful Execution - AI coding assistants, online development environments, analytics tools, vulnerability scanners, and game servers increasingly need a dedicated environment for each user or session. - Traditional options involve tradeoffs: - VMs provide strong isolation but often take minutes to start. - Containers launch quickly but share a kernel and require extensive hardening for untrusted workloads. - Standard serverless functions are designed for short, request-response workloads rather than long-running interactive sessions. - Building custom virtualization infrastructure requires significant security, operations, and virtualization expertise. ## What Lambda MicroVMs Provide - Each user or session receives its own Firecracker-powered MicroVM. - MicroVMs offer: - Dedicated VM-level isolation with no shared kernel between users. - Rapid launch and resume from a pre-initialized snapshot. - Persistent memory, disk state, and running processes during a session. - Automatic suspension during inactivity to reduce idle costs. - Automatic resume when new traffic arrives. - Firecracker already powers AWS Lambda at large scale, providing an established virtualization foundation. ## Creating a MicroVM Image - The example packages a Flask application and Dockerfile into a ZIP archive and uploads it to Amazon S3. - The Dockerfile uses: ```dockerfile FROM public.ecr.aws/lambda/microvms:al2023-minimal ``` - It installs Python and dependencies, copies the Flask application, and starts it with Gunicorn on port 5000. - An image is created with the `aws lambda-microvms create-microvm-image` command, specifying: - The S3 code artifact - An image name - An AWS-provided base image ARN - An IAM build role - Lambda builds the image, initializes the application, and captures its memory and disk state in a Firecracker snapshot. - Build logs are available in CloudWatch under `/aws/lambda/microvms/<image-name>`. ## Launching and Managing a MicroVM - A MicroVM is launched from the image ARN with `run-microvm`. - The example configures an idle policy that: - Suspends the MicroVM after 15 minutes of inactivity. - Keeps it suspended for up to 5 minutes. - Automatically resumes it when traffic returns. - Lambda assigns a unique MicroVM ID and provides a dedicated HTTPS endpoint. - No separate networking setup is required. - The application is already running when the MicroVM becomes available because it resumes from the image snapshot. ## Request Handling and State Preservation - Clients authenticate requests using a short-lived token in the `X-aws-proxy-auth` header. - The Flask API responds immediately after launch. - When the MicroVM becomes idle, Lambda snapshots and stores its memory and disk state. - A later request resumes the environment with the application state intact, making suspension effectively invisible to the client. ## Underlying Execution Model - Lambda MicroVMs use an image-then-launch workflow: - Build and initialize an environment once. - Snapshot the initialized state. - Launch future MicroVMs by resuming that snapshot. - This avoids repeating operating-system and application startup work. - The combination of Firecracker isolation, snapshot-based startup, and suspend/resume lifecycle control makes MicroVMs suitable for secure, interactive, multi-tenant workloads. For applications that must safely execute untrusted code while preserving session state and responsive startup times, Lambda MicroVMs offer a managed alternative to building custom VM infrastructure.

Read original(opens in new tab)
netflix3 min readCurated summary

How Netflix Simplified Batch Compute with Kueue

Netflix replaced much of its custom Compute Managed Batch (CMB) queuing and scheduling logic with Kubernetes-native Kueue. The migration preserved the existing user experience while enabling features such as preemption, fair sharing, all-or-nothing scheduling, and topology-aware placement. Kueue now manages millions of batch workloads across Netflix’s Titus-based infrastructure. ## CMB and Titus Architecture - CMB manages workloads that run to completion using: - Hierarchical tenants - Priority-based ordering - Per-tenant capacity management - Workloads ultimately run on Titus, Netflix’s container platform. - Titus provides federation across multiple Kubernetes cells and shared capacity reservations, allowing CMB to interact with a unified endpoint. - CMB tenants are either: - **Internal tenants**, which organize child tenants but do not accept jobs - **Leaf tenants**, which accept jobs through associated queues - Capacity includes: - **Reserved capacity**, providing predictable resources within a tenant hierarchy - **Shared capacity**, a global pool that tenants can burst into - CMB enforced fair sharing only at admission time because it lacked preemption; admitted jobs ran to completion even when demand changed. ## Why Netflix Chose Kueue - CMB was developed before many Kubernetes batch features became available in open source. - Kueue provided capabilities Netflix had previously built or wanted to build, including: - Fair sharing - Hierarchical tenancy - Capacity management - Priority queues - Preemption - Unlike schedulers such as YuniKorn and Volcano, Kueue works with the existing Kubernetes scheduler rather than replacing it. - This allowed Netflix to retain Titus scheduling profiles and avoid inefficient job placement. - Kueue also supports: - Multi-tenant quotas across heterogeneous hardware - Native Kubernetes objects such as `Pod` and `Job` - Higher-level workloads such as `RayJob` and `RayCluster` - All-or-nothing admission and topology-aware scheduling ## Migrating CMB Workloads - The migration, called **Netflix Batch**, was designed to: - Require no changes from CMB users - Avoid regressions in launch rates and maximum throughput - Move queuing and scheduling responsibilities to Kueue - Kueue runs in enabled Titus cells, while a custom router and Titus federation direct workloads to the appropriate cell. - Tenant enrollment was exposed as a simple operator action in Netflix’s UI, making rollout and rollback straightforward. - Internally, the migration mapped: - CMB internal tenants to Kueue **Cohorts** - Leaf tenants to **ClusterQueues** and **LocalQueues** - Capacity configurations to Kueue **resource flavors** and **nominal quotas** ## Lessons from the Rollout - Maintaining API compatibility reduced customer disruption and allowed Netflix to replace backend components incrementally. - Migrating the largest and most complex customer early exposed problems sooner and increased confidence in the broader rollout. - The production migration took approximately four weeks. - Kueue required substantially higher QPS, burst, and `groupKindConcurrency` settings than its defaults. - Netflix validated these settings early through load tests in an environment modeled on Titus. ## Kueue in Production - Kueue is fully deployed at Netflix and manages millions of batch workloads. - Netflix is extending its use to additional Titus batch workloads. - Fair sharing and preemption are being expanded to improve utilization of reserved capacity. - Netflix’s experience is also informing other internal Kubernetes-native systems, including training infrastructure. Netflix’s migration demonstrates that a batch platform can adopt Kubernetes-native scheduling incrementally without forcing users to change APIs or abandoning existing placement infrastructure. For organizations with mature custom systems, preserving the external contract while delegating queueing and admission to Kueue offers a lower-risk path to modern features and simpler long-term operations.

Read original(opens in new tab)
cloudflare3 min readCurated summary

How we found a bug in the hyper HTTP library

The Images binding’s migration to a local Unix-socket architecture exposed a rare race condition in Rust’s `hyper` HTTP library. Under slow-reader conditions, large image responses were truncated even though they returned `200 OK` and a full `Content-Length`, causing downstream processing or image decoding to fail. After six weeks of investigation, the issue was traced to premature socket shutdown and fixed with four lines of code. ## Images Bindings and the Request Path - Cloudflare’s Images service runs on Workers and uses `hyper` to manage HTTP connections. - The Images binding lets Workers send image data directly to the service, chain transformations, and receive the processed result as a stream. - The response path involved: - The Images service generating the complete encoded image. - `hyper` buffering the response. - Data moving through socket buffers managed by the kernel. - A client or intermediary reading the response. - If the reader was fast, `hyper` could flush the entire response and safely shut down the socket. - If the reader was slower, the socket’s outbound buffer filled, requiring `hyper` to pause and resume writing. ## Moving from FL to Local Unix Sockets - Initially, binding traffic passed through Cloudflare’s FL intermediary service. - In December 2025, the Images team replaced FL with an internal binding running on the same machine. - Unix sockets removed network and FL-processing overhead, including routing and DNS work. - The redesign improved performance and allowed the Images team to release binding changes independently. - The bug appeared within days of the rollout. ## Successful Responses with Truncated Bodies - The first report involved nested image-processing pipelines: - An inner Images binding composited large JPEG and PNG inputs from R2. - An outer URL-based pipeline resized, compressed, and transcoded the result. - The inner pipeline returned `200 OK` and a `Content-Length` for several megabytes, but delivered only a fraction of the body. - One response contained roughly 200 KB instead of the expected 3.3 MB. - The outer pipeline reported an end-of-file error because the body ended before the declared message length. - Depending on the image format, clients saw partially rendered images or completely broken images. ## Reproducing and Isolating the Race - Engineers recreated the nested setup, then removed layers until the failure occurred with the binding alone. - Batch testing produced failures reliably—for example, 19 of 25 requests in one run. - The amount of data received, approximately 200 KB, closely matched the production socket-buffer size. - This indicated that the failure was related to backpressure and socket-buffer exhaustion rather than the customer’s specific configuration. - Investigation eventually identified a race in `hyper` where the connection could be shut down before buffered response data had finished flushing. The incident demonstrates that HTTP success status codes do not guarantee complete response bodies when connection handling is incorrect. Systems streaming large payloads over sockets should test slow-reader and backpressure scenarios, and libraries should only close connections after all buffered data has been written.

Read original(opens in new tab)