Techlist.io - Korean Tech Blog Curator

github3 min readCurated summary

Automate repository tasks with GitHub Agentic Workflows

GitHub Agentic Workflows bring coding agents into GitHub Actions, allowing developers to describe repository tasks in Markdown instead of complex YAML. They can automate issue triage, documentation, testing, code cleanup, CI investigation, and reporting while preserving human oversight through permissions, sandboxing, logging, and review. The post presents the technology, now in technical preview, as an extension of CI/CD rather than a replacement for deterministic build and release pipelines. ## Markdown-Defined Repository Automation - Developers describe desired outcomes in plain Markdown and add the workflow to a repository. - The workflow runs in GitHub Actions using configurable coding agents such as: - GitHub Copilot CLI - Claude Code - OpenAI Codex - Because workflows operate within GitHub Actions, they benefit from repository context, audit logs, permission controls, and sandboxed execution. ## Examples of Continuous AI GitHub describes these workflows as “Continuous AI”: AI-powered automation integrated throughout the software development lifecycle. - **Issue triage:** Summarize, label, and route new issues. - **Documentation maintenance:** Update READMEs and documentation after code changes. - **Code simplification:** Find opportunities for improvement and open pull requests. - **Test improvement:** Evaluate coverage and add valuable tests. - **Quality hygiene:** Investigate CI failures and suggest targeted fixes. - **Reporting:** Produce recurring reports on repository health, activity, and trends. These tasks are difficult to implement with traditional deterministic YAML workflows because they require interpretation, judgment, and code changes. ## Relationship to CI/CD - Agentic workflows are intended to augment, not replace, existing CI/CD systems. - Traditional pipelines remain responsible for deterministic builds, tests, and releases. - Agentic workflows handle higher-level tasks involving analysis, recommendations, and repository maintenance. - GitHub Actions provides the infrastructure needed for controlled execution and observability. ## Guardrails and Human Control - Security is presented as a core design requirement, particularly against unintended behavior and prompt injection. - Workflows run with read-only permissions by default. - Write operations require explicit approval through “safe outputs,” which are designed to make changes pre-approved and reviewable. - The overall approach emphasizes inspectability, defined boundaries, and human review rather than unrestricted autonomous changes. ## Adoption Across Teams - GitHub Next reports using workflows to replace repetitive chores and assemble useful information for developers. - Home Assistant uses them to analyze large numbers of issues and identify important trends. - The Cloud Native Computing Foundation applies them to documentation automation and organizational reporting. - Carvana uses them for engineering work spanning multiple repositories. GitHub Agentic Workflows are best viewed as a controlled way to add AI judgment to repository operations. Teams should begin with focused, reviewable maintenance tasks and expand usage as they gain confidence in the workflows’ behavior and safeguards.

Read original(opens in new tab)
grammarlyOriginal article

What Is an AI Assistant? Definition, Types, and Examples (opens in new tab)

AI assistants have evolved from simple command-driven tools into sophisticated digital partners that leverage natural language processing to streamline workplace productivity. By integrating large language models with real-time data and contextual awareness, these tools enable users to automate repetitive tasks and manage information more effectively. Ultimately, their value lies in their ability to bridge the gap between open-ended human intent and actionable digital output across diverse software environments. ### The Technical Framework of AI Interaction * **Natural Language Processing (NLP):** This technology allows assistants to interpret the nuance of everyday language, distinguishing between literal questions and requests for tonal adjustments or stylistic changes. * **Large Language Models (LLMs):** These models use machine learning patterns to predict and generate helpful responses rather than relying on a pre-written script. * **Context Windows:** Modern assistants maintain a "memory" of the current conversation or document, allowing them to refer back to earlier sections and maintain consistency across long-form projects. * **Tool Integration:** Many assistants function by connecting to external APIs, enabling them to check calendars, pull data from the web, or manage task lists within other applications. ### Functional Applications in Daily Workflows * **Content Synthesis:** Assistants can ingest lengthy documents or meeting recordings to produce condensed summaries, outlines, and key takeaways. * **Drafting and Revision:** Beyond simple generation, these tools help refine existing text for clarity, length, and professional tone. * **Ideation and Brainstorming:** Users can utilize AI to overcome the "blank page" problem by generating initial project structures or exploring different angles for a specific topic. * **Technical Support:** For developers, AI assistants can interpret error messages, generate code snippets, and explain complex technical concepts in plain language. To maximize the impact of these tools, users should focus on providing detailed prompts that provide clear context and intent. As AI assistants become more deeply embedded in browsers and operating systems, understanding the balance between their generative capabilities and their contextual limitations is essential for maintaining an efficient digital workflow.

grammarlyOriginal article

How to Create an AI Assistant Step by Step: A Beginner’s Guide (opens in new tab)

Creating a custom AI assistant is no longer restricted to engineers, as modern no-code tools and APIs allow users to build specialized agents for specific personal or professional workflows. By focusing on a narrow scope and selecting the right platform, individuals can gain greater control over data, behavior, and task efficiency than generic tools provide. Ultimately, the shift toward custom assistants reflects a move away from one-size-fits-all software toward personalized AI teammates integrated directly into daily work. ## The Anatomy of an AI Assistant * Digital assistants utilize Natural Language Processing (NLP) to interpret user intent and tone through conversational prompts. * Large Language Models (LLMs) serve as the underlying engine, recognizing language patterns to generate contextually relevant responses. * Advanced implementations, such as the "Go" assistant, operate within existing apps like email and documents to eliminate context switching and manual data entry. ## Strategic Drivers for Customization * **Personalization:** Tailoring the assistant’s tone and behavior ensures it supports specific tasks exactly as the user expects. * **Data Control:** Building a custom solution offers transparency into how data is used, which is critical for teams handling sensitive internal information. * **Efficiency and Innovation:** Customizing an assistant for a niche problem—like summarizing specific document types or automating recurring questions—reduces manual effort more effectively than general tools. * **Independence:** Creating a proprietary tool reduces reliance on third-party platforms that may change their pricing or feature sets. ## Defining the Core Mission * The most successful assistants focus on one primary responsibility rather than trying to handle every possible task. * Effective planning requires answering who the user is and what specific problem the assistant is meant to solve consistently. * Starting with a narrow scope, such as a dedicated writing assistant or a customer service bot, simplifies the testing and refinement process during the initial launch. ## Development Paths and Lifecycles * Users can choose between no-code platforms for rapid deployment or API-based configurations for higher flexibility and integration. * The development process follows a standard lifecycle: strategic planning, technical configuration, launch, and continuous improvement. * Ongoing monitoring is essential to ensure the assistant remains responsible, accurate, and aligned with evolving user needs. To build a successful AI assistant, start by identifying a single high-impact task and selecting a tool that matches your technical comfort level. Prioritizing a narrow focus during the initial build will allow for more effective monitoring and easier scaling as your requirements grow.

netflix4 min readCurated summary

Scaling LLM Post-Training at Netflix

Netflix argues that LLM post-training at production scale is as much an infrastructure challenge as a modeling challenge. Its internal framework abstracts distributed data processing, model sharding, GPU orchestration, checkpointing, and complex training workflows so developers can focus on experimentation. The result is a flexible system supporting SFT, DPO, reinforcement learning, and knowledge distillation across hundreds of GPUs. ## Why Post-Training Becomes an Engineering Problem - Pre-training provides general language ability, but post-training adapts models to Netflix’s catalog, member histories, recommendation tasks, personalization, and search. - Production-scale training introduces challenges involving: - Large proprietary datasets - Multi-node GPU coordination - Distributed model state - Workflows that combine training and inference - Failure recovery and experiment tracking - A simple Hugging Face fine-tuning script is insufficient for reliable, large-scale jobs. ## Preparing Data Correctly - Chat templates serialize conversations but do not determine which tokens should contribute to the loss. - Netflix applies explicit loss masking so training focuses on assistant responses rather than prompts or other non-target text. - Variable-length examples can waste GPU memory through padding and create synchronization overhead across FSDP workers. - Sequence packing combines multiple samples into fixed-length sequences. - A document mask prevents attention across separately packed samples while improving GPU utilization. ## Loading and Optimizing Large Models - Models that do not fit on one GPU require sharding strategies such as FSDP or tensor parallelism. - Partial weights should be loaded directly onto the device mesh rather than materializing the entire checkpoint on a single device. - Developers can choose full fine-tuning or LoRA and use: - Activation checkpointing - Compilation - Appropriate precision settings - Reinforcement learning requires compatible precision between rollout generation and policy training. - Large vocabularies create memory pressure because logits have dimensions `[batch, seq_len, vocab]`. - The framework reduces peak memory by removing ignored tokens before projection and computing logits and loss in sequence chunks. ## Distributed Training and Workflow Management - The framework supports standard forward/backward training for SFT as well as workflows that interleave: - Rollout generation - Reward-model and reference-model inference - Policy updates - Ray actors orchestrate distributed jobs while keeping hardware concerns separate from modeling code. - Experiment tracking covers both quality metrics, such as loss, and efficiency metrics, such as Model FLOPS Utilization (MFU). - Standardized checkpointing allows jobs to resume after failures. ## Netflix’s Post-Training Framework - The stack is built on: - Mako for AWS GPU provisioning - PyTorch, Ray, and vLLM - Netflix’s framework library for reusable utilities and training recipes - Jobs are generally defined through configuration files that select a recipe and provide task-specific components. - Unlike narrower fine-tuning systems, the framework supports: - Custom output heads - Expanded vocabularies and semantic IDs - Special tokens - Transformer models trained on non-natural-language sequences - This flexibility is important for Netflix-specific recommendation and personalization use cases. ## Four Core Abstractions ### Data - Dataset abstractions cover SFT, reward modeling, and RL. - Streaming supports datasets larger than local disk capacity. - Asynchronous sequence packing overlaps CPU preprocessing with GPU execution to reduce idle time. ### Model - The framework supports architectures such as Qwen3 and Gemma3, including Mixture-of-Experts variants. - LoRA is integrated into model definitions. - High-level sharding APIs distribute models across device meshes without requiring developers to write low-level distributed code. ### Compute - A unified job interface scales from one node to hundreds of GPUs. - MFU measurement remains accurate for custom architectures and LoRA configurations. - Checkpoints include parameters, optimizer state, dataloader state, and data-mixer state, enabling exact resumption. ### Workflow - The system supports SFT, DPO, RL, and knowledge distillation. - Online RL uses a hybrid architecture combining a single controller with Single Program, Multiple Data (SPMD) workers. - This extends conventional SPMD training to multi-stage workflows that cannot be represented as a simple training loop. Netflix’s approach is to standardize the difficult operational parts of post-training while preserving enough flexibility for unconventional models and objectives. A framework built around reusable data, model, compute, and workflow abstractions can help teams iterate faster and scale experiments without repeatedly rebuilding distributed infrastructure.

Read original(opens in new tab)
line3 min readCurated summary

Claude Code Action: Platformizing AI Code

LINE NEXT transformed Claude Code from an individual productivity tool into an organization-wide code review platform integrated with GitHub Actions. The goal was to reduce review-quality variation, standardize policies, and make AI feedback part of the existing pull request workflow. Its central design separates simple repository-level invocation from centrally managed execution, prompts, permissions, and infrastructure. ## Why AI Code Review Needed to Be Platformized - As LINE NEXT’s services and repositories grew, human code review quality varied according to each reviewer’s experience and preferences. - Developers were already using Claude Code locally, but individual usage created several problems: - Inconsistent review criteria and perspectives - No organization-wide quality process - AI feedback disconnected from pull request workflows - Difficulty providing new employees with a consistent review experience - DevOps therefore treated the issue as a decentralized quality-process problem rather than merely a tooling problem. ## Why GitHub Actions and Claude Code - GitHub Actions was already the foundation for CI/CD and automation across LINE NEXT repositories. - It allowed the team to: - Apply a common workflow repository by repository - Centrally manage execution environments and permissions - Avoid requiring each service team to build additional infrastructure - Claude Code Action integrated directly with pull requests: - Developers could trigger reviews with an `@claude` mention. - Results appeared as GitHub comments or PR reviews. - Developers did not need to learn a separate interface. - A shared GitHub App Runner environment provided consistent execution and centralized security controls. ## Centralized Caller–Executor Architecture - Service repositories act as **callers**: - They invoke the standard workflow. - They provide only basic parameters such as service name and review type. - A centrally managed DevOps repository acts as the **executor**: - Stores prompts and review personas - Defines review policies and priorities - Manages permissions and authentication - Contains the actual execution logic - This design makes AI review an organization-wide platform capability rather than a separate configuration maintained by every project. ### Benefits of Central Control - **Consistent quality:** Central prompts and personas ensure common review depth, tone, security checks, stability checks, and priorities. - **Faster adoption:** New repositories need only add the standard workflow and specify a few parameters. - **Improved governance:** GitHub Apps, centrally managed secrets, and shared runners make it possible to track who accessed which code and with what permissions. - **Lower operational overhead:** Service teams use the platform without managing AI infrastructure themselves. ## Handling Fork-Based Pull Requests - The official Claude Code Action initially assumed that a PR branch existed in the base repository’s `origin`. - For pull requests created from forks, this caused failures such as: ```text couldn't find remote ref ``` - The original implementation fetched and checked out the branch by name: ```text git fetch origin <branch> git checkout <branch> ``` - This failed because fork branches exist in the external repository, not necessarily in the base repository. - From a platform perspective, this was a structural limitation because it blocked external contributors and collaboration repositories. - The proposed direction was to redesign the execution flow rather than simply add an exception, using GitHub’s special pull-request reference: ```text refs/pull/<PR number>/head ``` This approach allows the workflow to retrieve the actual pull request head commit regardless of whether the PR originated from the main repository or a fork.

Read original(opens in new tab)
line3 min readCurated summary

Slow Query Resolution: Optimizing Bit

LINE VOOM’s post server experienced intermittent timeouts when loading profiles belonging to users with hundreds of thousands of posts. The root cause was bitwise filtering on `category_flag` and `access_flag`, which prevented MySQL from efficiently using indexes and forced scans of all posts for a user. The team resolved the issue with MySQL 8.0.13 functional indexes and by changing the query predicates to exact decimal comparisons, reducing scanned rows from 805 to 31 in testing. ## The Slow Query and Its Root Cause - Post metadata was distributed across shards and partitioned tables. - `category_flag` and `access_flag` were stored as `bit(64)` values containing multiple status flags. - The problematic query filtered by: - `user_id` - `category_flag & 0x0100` - `access_flag & 0x0001` - For heavy users, the query scanned hundreds of thousands of posts and ran for more than 30 seconds. - Bitwise expressions operated on computed results rather than raw column values, preventing normal indexes from filtering efficiently. ## Choosing Functional Indexes - The team considered hardware upgrades, caching, and additional partitioning, but none addressed the root cause adequately. - MySQL 8.0.13 functional indexes could index expression results without changing the table schema. - The proposed composite index was: ```sql ALTER TABLE post_metadata ADD INDEX idx_user_premium_searchable ( user_id, (category_flag & 0x0100), (access_flag & 0x0001) ); ``` - Functional indexes rely on the query expression matching the index definition precisely. ## Discovering the Required Query Form - Initial attempts failed to use the index: - Truthy checks such as `category_flag & 0x0100` - Comparisons using `> 0` - Equality against hexadecimal values such as `= 0x0100` - The successful form used decimal equality: ```sql WHERE user_id = '{user_id}' AND (category_flag & 0x0100) = 256 AND (access_flag & 0x0001) = 1 ``` - In testing, scanned rows dropped from 805 to 31. - Index storage increased by approximately 24%, but the DBA team determined that production capacity was sufficient. ## Rolling Out the Indexes in Production - Indexes were created before changing the application queries. - The team used online schema changes to avoid service downtime and support pausing or rollback during replication problems. - Because dozens of tables across multiple shards were affected: - One shard was handled first for validation. - Only one or two tables were processed per day. - Work was avoided during periods when emergency DBA support was unavailable. - Index creation increased replication lag, causing newly created posts to temporarily disappear from read replicas. - The team reduced the cache expiration time for the affected post lists and accepted the remaining replication delay before resuming the rollout. ## Gradual Query Deployment and a Bitwise Logic Bug - Query changes were deployed gradually through a dynamic configuration system. - Each query pattern was tested on one shard before being expanded to the remaining shards. - This allowed changes to be rolled back immediately through configuration. - During rollout, a serious visibility bug was found. - The original condition: ```sql category_flag & 0x0110 ``` matched when either `0x0100` or `0x0010` was present, effectively representing an OR condition. - Rewriting it as: ```sql (category_flag & 0x0110) = 272 ``` required both bits to be set, creating an AND condition. - Because production data stored only the premium bit, some profiles returned no content. - The incident highlighted the need to verify the semantic meaning of bit flags before converting bitwise predicates into equality comparisons. ## Practical Recommendation For slow queries involving bit flags, consider functional indexes when using MySQL 8.0.13 or later. Ensure the query expression exactly matches the index definition, validate bitwise logic carefully, and use staged schema and query rollouts with monitoring and fast rollback mechanisms.

Read original(opens in new tab)
pinterest3 min readCurated summary

GPU-Serving Two-Tower Models for Lightweight Ads Engagement Prediction

Pinterest replaced its CPU-served two-tower model for ads lightweight ranking with a GPU-serving architecture based on MMOE and DCN. The more expressive model maintained latency comparable to the CPU baseline while reducing offline CTR loss by 5–10%. Separating standard and shopping ad models produced another 5–10% loss reduction and doubled offline iteration speed, with online improvements in CPC and CTR. ## Role of Lightweight Ranking - Lightweight ranking serves as an intermediate stage in Pinterest’s ads recommendation pipeline. - It filters a large pool of candidate ads before more complex downstream ranking models process them. - The two-tower design balances quality and latency: - The Pin tower generates ad embeddings offline through batch updates. - The query tower generates real-time user embeddings. - The prediction score is the sigmoid of the embeddings’ dot product. ## MMOE-DCN Model Architecture - The new system replaces the previous Multi-Task Multi-Domain (MTMD) model. - It combines: - Multi-gate Mixture-of-Experts (MMOE) with MLP-based gating. - Deep & Cross Network (DCN) layers for modeling feature interactions. - Each expert uses both full-rank and low-rank DCN layers. - Unlike MTMD, MMOE handles multi-task and multi-domain learning without relying on separate domain-specific modules. - GPU serving makes it practical to deploy this larger and more computationally demanding model while preserving CPU-baseline latency. ## Scenario-Specific Modeling - Standard and shopping ad scenarios are served as separate models. - Each model is trained only on data relevant to its scenario. - This specialization delivered an additional 5–10% reduction in offline loss. - Separating the models also doubled the speed of offline model iteration. ## Training Efficiency Improvements - **Dataloader optimization** - GPU prefetching prepares the next batch while the current batch is processed. - Additional worker threads take advantage of the 1 TB of CPU memory available on p4d instances. - **Model code optimization** - Operations that previously allocated zero-filled tensors on the CPU were moved to the GPU. - Fused kernels replaced multiple individual kernels to reduce execution overhead. - **Training configuration** - BF16 precision improved processing speed. - Larger batch sizes increased GPU memory utilization. ## Evaluation Results - The model uses downstream ranking scores as labels and optimizes KL divergence between those labels and its predictions. - Evaluation covers both: - Auction winners—ads ultimately inserted and shown to users. - Auction candidates—ads passed to downstream ranking. - Offline loss decreased significantly across all evaluated slices. - Online experiments showed: - Lower cost per click (CPC), which is favorable. - Higher click-through rate (CTR). GPU-serving a more complex MMOE-DCN two-tower model allowed Pinterest to improve ad engagement prediction without sacrificing serving latency. The results support using GPU infrastructure, scenario-specific models, and targeted training optimizations to scale lightweight ranking systems.

Read original(opens in new tab)
netflix3 min readCurated summary

Automating RDS Postgres to Aurora Postgres Migration

Netflix standardized on Amazon Aurora PostgreSQL after finding that PostgreSQL already supported most relational workloads and that Aurora offered stronger scalability, availability, elasticity, and ecosystem alignment. To migrate nearly 400 RDS PostgreSQL clusters efficiently, Netflix built a self-service workflow that automates replication, traffic quiescence, validation, and cutover while minimizing downtime and eliminating data loss. The Aurora read-replica method is preferred over snapshot migration because it keeps the target nearly synchronized while production continues running. ## Why Netflix Chose Aurora PostgreSQL - PostgreSQL already supported the majority of Netflix’s relational workloads. - Internal evaluations found Aurora PostgreSQL could support more than 95% of workloads running on other relational database systems. - PostgreSQL benefits from: - A broad open-source ecosystem - Strong community adoption - Compatibility with modern data platforms - Aurora’s distributed, cloud-native architecture provides: - Better scalability and elasticity - High availability - Support for globally distributed applications - The migration effort began with RDS PostgreSQL and is intended to expand to other relational systems. ## Database Migration Requires More Than Data Copying A safe migration must move both data and database functionality while preserving correctness, availability, and performance. - **Data replication:** Copy existing data and continuously apply source changes to the destination. - **Quiescence:** Stop writes to the source so the destination can catch up completely. - **Validation:** Confirm that source and destination data are synchronized. - **Cutover:** Redirect applications to the new Aurora database as the system of record. ## Operational and Technical Challenges - Manually migrating almost 400 PostgreSQL clusters would be slow, error-prone, and operationally expensive. - Coordinating downtime across dependent services is difficult. - Netflix therefore created a self-service workflow that handles orchestration, safety checks, and correctness guarantees automatically. - The system must guarantee: - Zero data loss - Extremely short downtime, especially for critical services - No performance degradation during or after migration - Migration of related resources such as parameter groups, read replicas, and replication slots - Application teams control database clients, so the platform cannot depend on them manually pausing writes. - The migration system must provide control-plane mechanisms to halt traffic safely during validation and cutover. - The workflow must operate without obtaining RDS credentials from users, since databases may be tightly secured and the migration platform may lack direct database access. - Because non-experts operate the process, the experience must be self-guided and require minimal user effort. ## Snapshot-Based Migration The snapshot approach is straightforward but requires stopping writes before migration. - Halt write traffic to the RDS PostgreSQL source. - Create a manual snapshot. - Convert the snapshot into an Aurora-compatible format. - Create an Aurora PostgreSQL cluster from the converted snapshot. - Validate the new cluster. - Redirect applications to the Aurora endpoint. This method is simple but can involve a longer interruption because the target is not continuously updated while the snapshot is created and converted. ## Aurora Read-Replica Migration The read-replica approach reduces downtime by continuously replicating the RDS database into Aurora. - Create an Aurora PostgreSQL read replica from the RDS source. - Stream changes asynchronously from RDS to Aurora while applications continue using the source. - Provision and validate Aurora configuration, connectivity, and performance in advance. - When replication lag is sufficiently low, briefly pause writes. - Allow the replica to catch up fully. - Promote it to a standalone Aurora PostgreSQL cluster. - Redirect application traffic to the Aurora endpoint. This approach keeps the destination nearly synchronized before cutover, making it substantially less disruptive than snapshot-based migration. Netflix’s automation focuses on making the read-replica migration process safe, repeatable, and self-service, with the platform handling replication, traffic control, validation, and cutover rather than relying on manual application-team coordination.

Read original(opens in new tab)
dropbox3 min readCurated summary

How low-bit inference enables efficient AI

Low-bit inference reduces the memory, compute, and energy required to serve modern AI models by representing values with fewer bits. Quantization can substantially increase GPU throughput, but its benefits depend on model accuracy, hardware support, and whether workloads prioritize latency or throughput. The article presents low-bit inference as a production trade-off rather than a universally optimal technique. ## The Rising Cost of Modern Models - Models are growing rapidly, increasing demand for: - Memory capacity - Compute power - Energy - Low-latency serving infrastructure - Dropbox uses attention-based models for Dash and other capabilities involving: - Text, image, video, and audio understanding - Search and summarization - Reasoning over large collections of content - Production deployment requires balancing model capability with hardware utilization, cost, and responsiveness. ## Where Inference Compute Is Spent - Most computation comes from repeated matrix multiplications in two areas: - **Linear layers**, including attention projections, MLP layers, and final output layers. - **Attention mechanisms**, which calculate relationships between input tokens and become increasingly expensive with longer contexts. - GPUs accelerate these operations using specialized hardware: - NVIDIA Tensor Cores - AMD Matrix Cores - These cores execute matrix multiply-accumulate operations much faster than general-purpose CUDA cores. ## How Lower Precision Improves Efficiency - Quantization reduces the number of bits used to represent model values. - Converting values from 16-bit to 8-bit or 4-bit formats: - Reduces memory usage - Lowers memory-transfer costs - Can increase matrix-operation throughput - Reduces energy consumption - GPU throughput generally improves as precision decreases; halving precision can approximately double the number of operations performed per second in suitable workloads. - Eight-bit quantization maps values into 256 discrete levels. Formats below 8 bits typically require **bitpacking**, combining multiple values into types such as `uint8` or `int32` because 4-bit values are not normally stored as native hardware types. - Newer hardware, such as Blackwell GPUs with FP4 support, can provide major energy savings compared with higher-precision systems like the H100. ## Limits of Extremely Low-Bit Formats - Binary and ternary quantization restricts weights to two or three possible values, offering greater theoretical savings. - These formats are not well matched to today’s GPUs because they cannot fully use Tensor or Matrix Cores. - Specialized accelerators could make them more practical, but adoption remains limited by: - Weak ecosystem support - Hardware availability - Concerns about model quality - Practical gains therefore depend not only on bit width, but also on how well the format is supported by existing hardware and software. ## Quantization Formats and Deployment Trade-offs - Quantization is a family of techniques with different choices for: - Numerical representation - Scaling - Execution strategy - These choices affect: - Model accuracy - Inference speed - Memory consumption - Hardware utilization - Different workloads have different priorities: - Latency-sensitive applications need fast individual requests. - Throughput-oriented workloads prioritize processing large volumes efficiently. - Depending on the workload, inference may be limited by software overhead, memory bandwidth, or specialized GPU compute units. ## Pre-MXFP and MXFP Approaches - The article divides modern low-bit formats into two broad groups following the introduction of **MXFP microscaling**: - **Pre-MXFP formats** rely on software-managed scaling and explicit dequantization. - **MXFP formats** move scaling and related operations into Tensor Core hardware. - MXFP aims to standardize low-bit data types while making them more directly usable by modern GPUs. - The choice between these approaches depends on the hardware generation and the specific performance requirements of each production workload. Low-bit inference is most effective when quantization formats, model quality, and hardware capabilities are considered together. Teams should select formats based on the actual bottleneck—memory, bandwidth, latency, or compute—rather than assuming that the fewest possible bits will always deliver the best result.

Read original(opens in new tab)
cloudflare2 min readCurated summary

Introducing Markdown for Agents

AI agents increasingly need structured, efficient access to web content, making traditional HTML a costly format for machine consumption. Cloudflare’s Markdown for Agents lets enabled websites serve HTML pages as Markdown when clients request `text/markdown`, reducing token usage and parsing overhead. The post argues that websites should treat AI agents as first-class visitors alongside humans and search engines. ## Why Markdown Matters for AI - Markdown conveys document structure with far less surrounding markup than HTML. - A Markdown heading such as `## About Us` uses roughly 3 tokens, compared with 12–15 tokens for an equivalent HTML heading. - The post’s HTML uses about 16,180 tokens, while its Markdown version uses approximately 3,150—a reduction of about 80%. - Converting HTML to Markdown inside an AI pipeline adds computation, cost, and complexity, and may not preserve the publisher’s intended structure. ## How Markdown for Agents Works - Cloudflare-enabled zones can respond to content negotiation requests containing: ```http Accept: text/markdown ``` - Cloudflare fetches the original HTML from the origin, converts it to Markdown at the network edge, and returns the converted response. - Clients can request Markdown with `curl`, while Workers-based agents can use a `fetch()` request with `Accept: "text/markdown, text/html"`. - Responses use `Content-Type: text/markdown` and include `Vary: accept`. - Existing coding agents, including Claude Code and OpenCode, already send compatible `Accept` headers. ## Token Estimates and Agent Workflows - Converted responses include an `x-markdown-tokens` header. - Agents can use this estimate to: - Determine whether content fits within a context window - Plan chunking strategies - Manage processing costs and limits ## Content Signals - Markdown responses include: ```http Content-Signal: ai-train=yes, search=yes, ai-input=yes ``` - These signals indicate that the content may be used for AI training, search results, and AI input, including agentic applications. - Cloudflare says future versions will support custom Content Signal policies. ## Availability - Cloudflare enabled Markdown for Agents on its Developer Documentation and Blog. - AI crawlers and agents can test the feature by requesting those pages with `Accept: text/markdown`. Web publishers can make their content more accessible to AI systems by supporting Markdown negotiation, while agents should request `text/markdown` whenever available to reduce tokens, parsing work, and processing cost.

Read original(opens in new tab)
microsoft4 min readCurated summary

How we built the Microsoft Learn MCP Server

Microsoft Learn MCP Server gives AI agents direct, standardized access to current Microsoft documentation through the Model Context Protocol (MCP). Rather than requiring custom APIs, scraping, or embeddings, agents can dynamically discover and use tools for searching documentation, fetching full articles, and finding code samples. Microsoft’s experience shows that successful MCP systems depend not only on retrieval quality, but also on agent-oriented tool design, operational resilience, clear descriptions, and defensive compatibility practices. ## Purpose of Learn MCP Server - Provides trusted, up-to-date Microsoft Learn content to GitHub Copilot and other AI agents. - Uses Streamable HTTP Transport so MCP-compatible clients can connect to a remote server. - Supports three tools: - `microsoft_docs_search` for titles, relevant content sections, and source URLs. - `microsoft_docs_fetch` for retrieving complete article content. - `microsoft_code_sample_search` for locating language-specific code examples. - Grounds agent responses in official Microsoft documentation rather than relying solely on model memory. ## Why MCP Instead of a Traditional API - Conventional APIs require each client to implement: - Authentication and request formatting. - Documentation and integration logic. - Error handling and compatibility maintenance. - MCP allows clients to discover available tools and schemas at runtime. - The same server can support many agents without custom integrations. - Runtime discovery helps clients adapt to evolving tool contracts and reduces hardcoded assumptions. ## Architecture - The remote MCP server sits in front of the Microsoft Learn knowledge service. - It uses the official C# MCP SDK and runs on Azure App Service. - Clients communicate through Streamable HTTP Transport. - The server uses the same content vector store as Ask Learn, providing shared: - Freshness guarantees. - Relevance ranking. - Index coverage. - Ask Learn delivers retrieval directly to users, while Learn MCP Server exposes that capability through a protocol usable by external agents. ## Designing Tools Around Agent Workflows - Internal retrieval APIs expose many low-level options, such as `topK`, index selection, thresholds, filters, and search modes. - Learn MCP Server hides that complexity behind intuitive search-and-fetch operations. - Tool contracts should reflect how agents work rather than mirror backend APIs. - Keeping retrieval details internal prevents implementation choices from leaking into the agent-facing interface. ## Operating a Remote MCP Service - A public MCP server has distributed-systems concerns despite using JSON-RPC: - Cross-region deployment. - Dynamic scaling. - CORS. - Session affinity. - Statelessness. - Data protection. - Operational design and SDK collaboration are as important as implementing the tools themselves. ## Tool Descriptions Shape Agent Behavior - Tool and parameter descriptions act as instructions for language models. - Small wording changes can significantly affect whether agents select a tool and how successfully they use it. - Microsoft created automated evaluation tooling to test descriptions against observed agent behavior and success metrics. - Updated descriptions can be delivered when clients refresh their MCP sessions. ## Combining Search and Fetch - Search and fetch are more effective together than independently. - A typical workflow is: - Search for the most relevant Learn article or section. - Fetch the full Markdown page for additional context. - Use that content to produce a better-grounded answer with stronger citations. - Explicitly describing this follow-up pattern improved downstream results. ## Handling Hardcoded Clients - Some MCP clients treat discovered tools like fixed APIs and hardcode schemas. - Renaming the `question` parameter to `query` caused 2–5% of requests to fail. - Supporting both names during a deprecation period reduced disruption. - Public MCP services must evolve defensively, even though the protocol supports dynamic discovery. - Tools such as MCP Interviewer can help identify schema and behavioral problems before deployment. ## Using Data to Guide Improvements - Usage data showed that most requests involve: - Coding tasks. - Explanations. - Troubleshooting. - The team prioritized retrieval and description changes around these intents. - Documentation-level agent instructions also encourage use of Learn tools when Microsoft technologies are involved. Microsoft Learn MCP Server replaces the manual process of searching, opening, and copying documentation into a development environment. The practical recommendation is to connect compatible agents to the server so they can retrieve official Learn content directly, while MCP tool authors should design simple contracts, measure real agent behavior, and preserve compatibility as their services evolve.

Read original(opens in new tab)
meta2 min readCurated summary

The Death of Traditional Testing: Agentic Development Broke a 50-Year-Old Field, JiTTesting Can Revive It

Just-in-Time Tests (JiTTests) are an LLM-driven testing approach designed for fast, agentic software development. Instead of maintaining static test suites, the system generates tests for each code change, simulates likely faults, and reports only meaningful regressions. The goal is to reduce test maintenance and false positives while catching serious bugs before production. ## Limitations of Traditional Testing - Tests are manually written as code changes enter the system. - They must account for both current behavior and unknown future changes. - This often leads to: - Tests that fail to detect relevant bugs. - False positives when intended changes break outdated assumptions. - Ongoing maintenance and review costs. - Agentic development increases the volume and speed of changes, making these problems harder and more expensive to manage. ## How Catching JiTTests Work - A new code change or pull request is submitted. - An LLM infers the likely intent of the change. - The system creates mutants—versions of the code containing deliberately introduced faults. - It generates and runs tests designed to expose those faults. - Rule-based and LLM-based assessors evaluate failures and filter out likely false positives. - Engineers receive a focused report when the system identifies an unexpected behavior change. Because the tests are tailored to a specific change, they can reason about intended behavior and distinguish legitimate updates from regressions. ## Benefits for Agentic Development - Tests are generated on demand and do not remain in the codebase. - There is no ongoing test maintenance or test-code review. - Each test is specific to the change being evaluated. - Tests automatically adapt as the code evolves. - Human attention is required mainly when an actual bug is detected. - Testing shifts from measuring generic code quality to determining whether a specific change introduces a real fault. Catching JiTTests are presented as a way to make testing scale with AI-assisted development by moving routine test creation and maintenance from engineers to automated systems.

Read original(opens in new tab)
dropbox2 min readCurated summary

Insights from our executive roundtable on AI and engineering productivity

Dropbox argues that AI improves engineering productivity only when tied to measurable business outcomes rather than adopted for its own sake. The company has expanded AI use across the software development lifecycle, while recognizing trade-offs involving quality, maintenance, and organizational change. Its executive roundtable concluded that leadership, formal AI competency, and stronger outcome measurement will be central to realizing AI’s potential. ## Dropbox’s AI Adoption Strategy - Dropbox made AI adoption a company-wide priority with leadership sponsorship, enabling teams to experiment more easily and reducing delays in approving new tools. - Engineers use AI across code review, documentation, debugging, testing, and other stages of development. - Because Dropbox operates a large, multilingual monorepo, it combines commercial tools such as Claude Code and Cursor with internally built systems. - One internal tool detects failed pull-request builds and uses Dropbox’s AI platform to suggest fixes. - Most developers now use at least one AI tool. - Dropbox tracks monthly pull-request throughput per engineer and has observed higher output among developers who use AI coding tools more actively. - The company also monitors engineer sentiment, reporting increased positive sentiment and reduced negative sentiment as adoption improves. ## Focus of the Executive Roundtable Leaders from multiple companies discussed engineering productivity and AI in rotating peer groups organized around three themes: - **Measuring impact** - Identifying ways to measure AI-driven productivity gains. - Connecting engineering improvements to broader business results. - **Leadership alignment** - Establishing how executives should communicate AI deployment progress. - Determining the appropriate pace and scope of adoption. - **The human element** - Recruiting, evaluating, and developing AI-capable employees. - Applying lessons from developer productivity to help non-engineering teams work more effectively. ## Lessons About AI and Productivity - **Balance is essential:** Faster development must not come at the expense of software quality or increased long-term maintenance costs. - **Leadership sets standards:** Technical managers play a key role in defining responsible and effective AI usage norms. - **AI skills should be formalized:** Including AI competency in career frameworks demonstrates that it is a lasting strategic capability rather than a temporary trend. - **Extra capacity needs direction:** Dropbox is currently using productivity gains to address technical debt, complete migrations, and improve reliability. ## Priorities for 2026 Dropbox’s main unresolved challenge is linking engineering productivity metrics to tangible business outcomes. Its next phase will focus on mapping AI-driven gains to specific results, extending operational discipline beyond engineering, and improving end-to-end product velocity.

Read original(opens in new tab)
figma2 min readCurated summary

State of the Designer 2026: Designers Are Leaning Into the Messy Middle | Figma Blog

Designers are navigating rapid change by using AI as a complement to—not a replacement for—human craft. Figma’s 2026 survey of 906 designers finds that AI is helping many work faster, collaborate better, and improve quality, while strong craft remains central to satisfaction and business performance. The report’s overall conclusion is optimistic: designers are embracing uncertainty and turning new pressures into creative momentum. ## AI Improves Speed, Collaboration, and Quality - The survey was conducted by NewtonX across North America, APAC, Europe, LATAM, and the Middle East. - Respondents answered in English, Spanish, French, Italian, Portuguese, Japanese, and Korean. - 89% of designers say AI helps them work faster. - 80% say it improves collaboration. - 91% believe AI tools improve their designs, countering concerns that AI-generated work will reduce quality. - Designers who actively use AI are 25% more likely to report job satisfaction. - AI users are also more likely to say they drive business impact and contribute to company growth. - By automating or accelerating workflow tasks, AI gives designers more time for high-impact ideas. ## Craft Remains a Human Differentiator - As AI makes prototyping more accessible, craft becomes a key way for products to stand out. - Designers define craft in several ways: - Visual polish: 58% - Thoughtful problem-solving: 47% - Clear, intuitive UX: 36% - Emotion and delight: 35% - Consistency across products: 15% - Craft can mean technical skill, careful execution, intentional decisions, artistry, or solving difficult product problems. - Designers who associate craft with visible emotional and creative outcomes often receive more recognition than those whose craft involves less visible tactical work. ## Design Excellence Supports Morale and Growth - Designers are twice as likely to feel positive about their work when leaders prioritize design excellence. - Teams that value craft report stronger morale, faster business growth, and a clearer sense of momentum. - Leadership support, recognition, and opportunities for development help designers maintain quality while adapting to new tools. - The report links investment in craft with better outcomes for both designers and their organizations. Organizations should treat AI as a way to extend designers’ capabilities while continuing to invest in human judgment, creativity, quality, and recognition.

Read original(opens in new tab)
kakao2 min readCurated summary

Introducing the new Kanana-o

Kanana-o is Kakao’s new Korean-focused omni-modal AI model, designed to understand and generate text, images, and audio naturally. Kakao is opening a closed beta for the Kanana-1.5-o-9.8b-2602 model to gather feedback from developers and partners before commercial release. The service emphasizes practical experimentation rather than large-scale traffic handling. ## Model Capabilities - Supports simultaneous processing of multiple modalities, including text, images, and audio. - Specializes in: - Deep understanding of Korean language, culture, and user intent. - Natural Korean speech with expressive intonation, pacing, and emotion. - Flexible applications such as podcast narration, multi-turn conversations, and multi-speaker text-to-speech. - Balances text-generation speed with audio-processing speed to produce more natural spoken responses. ## API Beta Service - **Service:** Kanana-o API Beta - **Model:** Kanana-1.5-o-9.8b-2602 - **Beta period:** February 27–May 27, 2026 - **Access:** Selected testers receive a fixed number of daily API uses during the beta. - The closed beta is intended for meaningful developer testing and feedback, not high-volume production workloads. ## Application and Selection - Applicants should visit [omni.kanana.ai](https://omni.kanana.ai/), sign in with a Kakao account, and submit information about: - Their organization or affiliation - Intended purpose - Expected technical scenarios - Selected applicants will receive invitations and API documentation through KakaoTalk notifications starting February 27. - Kakao is seeking developers, students, startups, and researchers with concrete implementation plans. - Specific proposals—such as building a visual shopping assistant for people with visual impairments—are favored over general interest in trying AI. Developers interested in exploring Korean-language, audio, and vision applications can apply for the beta with a clearly defined use case and prototype plan.

Read original(opens in new tab)