Techlist.io - Korean Tech Blog Curator

gitlab3 min readCurated summary

Prepare your pipeline for AI-discovered zero-days

AI is accelerating both vulnerability discovery and insecure code production, shrinking the time defenders have to respond from months to hours. The post argues that security teams cannot remain the final defense layer; security controls, automated triage, and remediation must operate directly within development pipelines. AI-generated fixes can help close the gap, but they must follow the same policies, approvals, testing, and audit requirements as human-authored code. ## The Remediation Backlog Is Already Too Large - Most exploited vulnerabilities are already known and have patches available, but organizations cannot remediate them quickly enough. - Sixty percent of breaches in the 2025 Verizon DBIR involved known vulnerabilities. - Developers spend roughly 11 hours per month fixing vulnerabilities after release. - The median time to close half of internet-facing vulnerabilities is 361 days, while exploitation can begin within hours. - AI-assisted development is increasing the volume of insecure code: - Fortune 50 repositories reportedly gained more than 10,000 security findings per month by mid-2025. - AI coding tools may introduce outdated patterns, hallucinated packages, insecure examples, and excessive dependencies. - Security AI should therefore operate within existing development policies and audit trails rather than as a disconnected tool. ## Security Enforcement Must Move Into the Pipeline - Every change should pass security controls at the merge request, which becomes the central enforcement point. - Policies should be defined once and applied consistently across teams and projects. - Exceptions should be explicitly approved and logged. - IDE checks can catch straightforward problems—such as hardcoded secrets, vulnerable imports, and deprecated APIs—before code reaches review. - This allows human reviewers to focus on complex issues such as reachability, exploitability, and architectural risk. ## Automated Triage and Governed Remediation - AI should reduce the volume of findings developers must investigate by assessing: - False positives - Reachability - Exploitability - Severity - AI-generated fixes should not bypass normal governance. - Remediation proposals should be submitted as merge requests, with: - Required scans - Policy enforcement - Human approvals - Confidence scores - Complete audit records - Human and AI-authored changes should follow the same review and compliance process. ## Example: Responding to an Emerging Vulnerability - A proof-of-concept exploit may appear before a CVE, NVD entry, or scanner signature exists. - A security agent can inspect dependency graphs across projects, identify affected versions and call paths, and rank production exposure. - Teams can then launch a coordinated remediation campaign: - Upgrade dependencies where patches exist. - Apply targeted code changes where they do not. - Block merge requests that retain the vulnerable dependency. - Require security approval for fixes. - Pipeline tests can reject faulty AI-generated patches, allowing the agent to revise them before developers approve the corrected version. - Automatically collected scan results, policies, approvals, and merge timestamps provide audit evidence without manual reconstruction. ## Strengthen the Pipeline Before Attackers Catch Up - Organizations should verify that security scans run on every merge request, not only in selected projects. - Pipelines should detect compromised or vulnerable dependencies before build time. - Critical findings should move quickly from detection to the responsible developer without unnecessary tool boundaries. - The central recommendation is to make pipeline enforcement, AI-assisted triage, and governed remediation standard parts of the software supply chain before comparable offensive AI capabilities become widely available.

Read original(opens in new tab)
gitlab3 min readCurated summary

GitHub Copilot's policy for AI training: A governance wake-up call

GitHub’s April 2026 policy change will make Copilot Free, Pro, and Pro+ interaction data—including code, prompts, outputs, and context—available for AI training by default unless users opt out. The change highlights governance risks for regulated organizations, especially when protections vary by subscription tier or can be altered through policy updates. The post presents GitLab’s no-training commitment, contractual safeguards, and transparency documentation as a stronger model for enterprise AI governance. ## What the GitHub Policy Change Means - Beginning April 24, 2026, GitHub may use Copilot Free, Pro, and Pro+ data for model training by default. - Covered data includes: - User inputs and outputs - Code snippets - Associated context - Interaction data - Users must actively opt out. - Copilot Business and Enterprise customers remain exempt under existing contracts. - Data may also be shared with GitHub affiliates, including Microsoft, for AI development. - Organizations must review license tiers, settings, contracts, and internal AI governance controls. ## Why This Matters in Regulated Industries - Source code can expose: - Proprietary business logic - Internal system architecture - Sensitive data flows - Financial algorithms and risk models - Financial institutions may face intellectual-property and model-risk concerns involving trading strategies, underwriting rules, fraud detection, and credit models. - Frameworks such as Federal Reserve SR 11-7 and DORA require documented oversight of third-party technology and material changes in vendor practices. - Public-sector environments governed by NIST 800-53 and FISMA may require sensitive code to remain within controlled boundaries. - Healthcare organizations must consider HIPAA obligations when development tools interact with clinical or patient-adjacent systems. - Default opt-in training, individual opt-out requirements, and tier-dependent protections create compliance risks. ## Requirements for Enterprise AI Vendors - **Contractual certainty:** Vendors should clearly and unconditionally define how customer data is handled. - **Auditability:** Organizations need documentation about models, training data, subprocessors, retention, and compliance status. - **Independence from vendor incentives:** Customer code should not become training data for systems that may benefit competitors. - **Operational flexibility:** Regulated customers may require self-hosting, controlled processing boundaries, or clear procedures for vendor changes. ## GitLab’s AI Governance Position - GitLab states that it does not train AI models on customer code at any pricing tier. - Its AI vendors are contractually prohibited from using GitLab customer inputs or outputs for their own purposes. - The GitLab AI Transparency Center documents: - Models powering its features - Data handling practices - Subprocessors - Retention periods - Feature compliance status - GitLab emphasizes cloud and model neutrality, supports self-hosted deployments, and addresses vendor changes through its AI Continuity Plan. - The post argues that these policies reduce vendor-concentration, compliance, and intellectual-property risks. ## Closing the Governance Gap Organizations should ask every AI vendor: - Is customer data used for model training? - Who are the model subprocessors? - What happens if data practices change? - Can AI processing remain inside the organization’s infrastructure? - What indemnification applies to AI-generated output? The post’s recommendation is to favor vendors that provide durable, contractual, and auditable answers rather than relying on defaults, temporary opt-outs, or policies that can change with short notice.

Read original(opens in new tab)
gitlab2 min readCurated summary

What’s new in Git 2.54.0?

Git 2.54.0 introduces foundational changes to Git’s storage and history-editing capabilities. Its object database is now pluggable, making alternative storage formats more feasible, while the new `git history` command simplifies common commit-history edits that previously required interactive rebases. These changes are early milestones in longer-term efforts to improve repository performance and support stacked-diff workflows. ## Pluggable Object Databases - Git already supports interchangeable reference backends, including `files` and `reftable`. - Git 2.54 extends this abstraction to object databases, which store loose objects and packfiles under `.git/objects`. - The work began in Git 2.48 and involved nearly 400 upstream commits over almost two years. - Alternative object-storage implementations can now support meaningful local workflows, including: - Creating commits - Displaying commit graphs - Performing merges - Remote operations such as fetching and pushing are not yet supported through alternate backends. - Future storage formats could: - Handle large binary files more efficiently than packfiles - Be optimized for GitLab’s repository-serving infrastructure - The project was led by Patrick Steinhardt. ## Easier Commit-History Editing - Developers often rewrite history to produce small, atomic commits with clear messages, but interactive rebases can be difficult to learn. - Interactive rebases require users to choose a base commit, edit an instruction sheet, and understand Git’s stateful rebase process. - Git 2.54 introduces `git history`, inspired partly by Jujutsu’s simpler history-editing commands. - Initial subcommands include: - `git history reword`: change a commit message - `git history split`: divide one commit into two by selecting which changes belong in each - Planned commands include: - `git history fixup` - `git history drop` - `git history reorder` - `git history squash` - The command can automatically rebase local branches that contain the edited commit, including branches other than the current one. - This behavior supports Git’s broader effort to improve stacked-diff workflows, where dependent branches are reviewed independently. - The project was led by Patrick Steinhardt with support from Elijah Newren. The release points toward a more extensible Git: repository storage can eventually be optimized for different workloads, while history editing becomes more approachable than traditional interactive rebases. Since both features are still developing, users should expect broader backend support and additional history commands in future releases.

Read original(opens in new tab)
github2 min readCurated summary

Building an emoji list generator with the GitHub Copilot CLI

Cassidy Williams describes building an AI-powered emoji list generator during GitHub’s Rubber Duck Thursdays livestream. The terminal application converts bullet points into relevant emojis, then copies the formatted Markdown list to the clipboard with a keyboard shortcut. The project demonstrates how GitHub Copilot CLI and SDK can quickly turn a small idea into a functional open-source tool. ## The Emoji List Generator - Runs directly in the terminal. - Accepts pasted or typed bullet points. - Uses AI to replace each bullet with a relevant emoji. - Generates the result when the user presses `Ctrl + S`. - Copies the completed list to the clipboard. - Exits with `Ctrl + C`. ## Technologies Used - `@opentui/core` provides the terminal user interface. - `@github/copilot-sdk` supplies the AI functionality. - `clipboardy` handles clipboard access. ## Building the Project with Copilot CLI - Development began in Copilot CLI’s plan mode using Claude Sonnet 4.6. - A natural-language prompt described the desired Markdown emoji generator and requested integration with the Copilot SDK. - Copilot asked clarifying questions about the technology stack and libraries. - It then produced a `plan.md` file for review. - The implementation was completed with Claude Opus 4.7 only a few minutes later. ## Copilot CLI Features Demonstrated The livestream project combined several Copilot CLI capabilities: - Plan mode for outlining the implementation. - Autopilot mode for carrying out development tasks. - A multi-model workflow using different Claude models. - The `allow-all` tools flag for permissive tool access. - The GitHub MCP server for GitHub-related integrations. The finished Emoji List Generator is available as a free, open-source project, alongside documentation for the GitHub Copilot CLI and SDK.

Read original(opens in new tab)
netflix3 min readCurated summary

The Human Infrastructure: How Netflix Built the Operations Layer Behind Live at Scale

Netflix’s live-streaming growth required more than resilient technology—it demanded a dedicated human and physical operations layer. In three years, Netflix expanded from one live show per month to roughly 70 events in March 2026, including a World Baseball Classic game watched concurrently by more than 9.6 million accounts. The company evolved from engineers operating improvised setups to specialized teams, permanent facilities, and standardized broadcast procedures designed for continuous global scale. ## From Improvised Launches to Global Scale - Netflix’s first live events in March 2023 were operated by the engineers who built the streaming pipeline. - There was no dedicated operations team, formal command center, or live-specific incident response process. - Engineers monitored dashboards on laptops, coordinated through Slack, and troubleshot while millions of members watched. - Temporary control rooms were assembled in conference rooms, while larger events used rented broadcast facilities and equipment. - By March 2026, Netflix was operating 24/7 from facilities in Los Gatos and Los Angeles, with international coverage from Tokyo. - The company streamed approximately 70 events in that month—nearly as many as it had streamed throughout all of 2024. ## The Broadcast Operations Center - The Broadcast Operations Center (BOC) is Netflix’s physical command center for live events. - It receives the fully produced feed from a venue and hands it off to Netflix’s streaming infrastructure. - BOC responsibilities include: - Signal ingest and inspection - Audio and video conditioning - Closed-caption validation - Graphics insertion - Advertising management - A hub-and-spoke design, dual internet circuits, and SMPTE 2022-7 seamless switching reduce dependence on venue-specific infrastructure. - Centralizing these functions makes live events more repeatable and resilient. ## Protecting the Venue Signal - Netflix requires three completely independent transmission paths for every show-critical feed. - Approved contribution methods are prioritized as follows: - Dedicated video fiber - Single-feed satellite links - Dedicated enterprise-grade internet - SRT contribution systems - Production trucks must use redundant routers and transmission hardware, including separate router line cards. - Transmission equipment requires two independent power sources, UPS battery protection, and surge conditioning. - Before each event, operators conduct FACS/FAX facilities checks, including: - Audio/video synchronization tests - Latency and quality testing - Closed-caption verification - Backup switcher validation ## The Evolution of Netflix’s Operations Teams ### Phase 1: All-Hands Engineering - Core software engineers configured, launched, monitored, and dismantled every live event. - This approach worked for early broadcasts but could not scale as event volume increased. - Requiring developers to manually operate each show limited their ability to build new platform capabilities. ### Phase 2: Specialized Engineering Teams - Streaming Operations Engineers (SOEs) took responsibility for configuring and supporting events on the live-streaming pipeline. - SOEs became the first escalation point, allowing core developers to focus on platform development. - Broadcast Operations Engineers (BOEs) were later added to manage physical broadcast facilities and hardware. - BOEs oversee facility-related issues and support all shows running during a shift. ### Phase 3: The Co-Pilot Control Room - Dedicated Broadcast Control Operators (BCOs) assumed responsibility for operating the audio and video feeds. - Two BCOs worked together in a “first captain/second captain” model similar to a pilot and co-pilot. - This arrangement provided strong focus and execution quality for one or two events per day. - It became too space- and labor-intensive when Netflix began targeting up to ten simultaneous events. Netflix’s experience shows that live streaming at global scale depends on integrating broadcast discipline with software-engineering expertise. The key recommendation is to treat operations, redundancy, facilities, and specialized human roles as core parts of the product—not as temporary support added after the technology is built.

Read original(opens in new tab)
cloudflare3 min readCurated summary

Introducing the Agent Readiness score. Check to see if your site is agent-ready

Cloudflare argues that websites must evolve beyond browser and search-engine compatibility to become usable by AI agents. Its new Agent Readiness score evaluates whether sites support standards for discovery, content access, bot control, and agent capabilities. Early data shows adoption is extremely low, creating both a challenge and an opportunity for sites that adopt these standards early. ## Agent readiness across the web - Cloudflare analyzed the 200,000 most visited domains, excluding categories unlikely to need agent interaction. - The resulting Cloudflare Radar dataset tracks adoption of AI-agent standards and will be updated weekly. - robots.txt exists on 78% of sites, but most files target traditional search crawlers rather than AI agents. - Only 4% of sites declare AI usage preferences through Content Signals. - Just 3.9% support Markdown content negotiation via `Accept: text/markdown`. - MCP Server Cards and API Catalogs based on RFC 9727 appear on fewer than 15 sites, showing how early these standards remain. ## The Agent Readiness score Site owners can test their websites at **isitagentready.com**. Cloudflare scans the site and scores it across four dimensions: - **Discoverability:** robots.txt, sitemap.xml, and Link Headers under RFC 8288. - **Content:** Markdown for Agents. - **Bot Access Control:** Content Signals, AI-specific robots.txt rules, and Web Bot Auth. - **Capabilities:** Agent Skills, API Catalogs, OAuth discovery standards, MCP Server Cards, and WebMCP. - The tool also checks commerce standards such as x402, Universal Commerce Protocol, and Agentic Commerce Protocol, though these do not yet affect the score. - Each failed check includes a prompt that can be handed to a coding agent for implementation. The service itself supports agents through a stateless MCP server with a `scan_site` tool and publishes Agent Skills documents explaining how to implement each supported standard. ## Discoverability for AI agents - robots.txt helps agents understand crawl permissions and locate sitemaps. - Sitemaps provide a structured list of site paths, reducing the need to discover content by following every HTML link. - HTTP Link headers, defined by RFC 8288, expose important resources directly in responses without requiring agents to parse page markup. - Sites can use headers such as `rel="api-catalog"` to point agents toward machine-readable capabilities. ## Making content easier to read - `llms.txt` provides an LLM-oriented reading list at the site root, describing the site and linking to important content in a format designed for model context windows. - Markdown content negotiation lets agents request a clean Markdown version of a page with `Accept: text/markdown`. - Cloudflare measured token reductions of up to 80% compared with HTML, improving speed, cost, and the likelihood that agents can consume an entire document within their context limits. Cloudflare’s recommendation is to evaluate sites with the Agent Readiness tool and adopt the relevant standards incrementally. With current adoption so low, early support can make a site significantly easier for AI agents to discover, understand, authenticate with, and use.

Read original(opens in new tab)
cloudflare4 min readCurated summary

Shared Dictionaries: compression that keeps up with the agentic web

Shared dictionaries address a growing web-performance problem: pages are getting heavier, being rebuilt more frequently, and fetched repeatedly by agents. Instead of retransmitting entire assets after every deployment, servers can compress new versions against files already cached by the browser and send only the differences. The approach could dramatically reduce bandwidth and CPU use, though adoption depends on browser support, security safeguards, and complex server-side implementation. ## The Problem: More Shipping Means Less Caching - Web pages have become 6–9% heavier annually due to frameworks, interactivity, and media. - Agentic crawlers and other automated tools increasingly request full pages; they accounted for nearly 10% of Cloudflare requests in March 2026, up about 60% year over year. - AI-assisted development leads to more frequent deployments and experiments. - Small code changes can cause bundlers to re-chunk assets and generate new filenames, forcing clients to download entire bundles again. - Conventional compression reduces the size of each response but cannot exploit the fact that the client already has most of the previous version. - Frequent deployments therefore create substantial redundant bandwidth and CPU usage. ## How Shared Dictionaries Work - A compression dictionary is shared knowledge between the client and server. - The server compresses a new response using content the client already possesses as a reference. - The client uses that same reference to reconstruct the complete file. - Brotli includes a built-in dictionary of common web patterns, while Zstandard can generate custom dictionaries from representative content. - Gzip lacks a prebuilt or custom dictionary and discovers patterns only during compression. ## Delta Compression for Versioned Assets - Shared dictionaries use the previously cached resource as the compression dictionary. - The initial response includes a `Use-As-Dictionary` header, telling the browser to retain the resource for future compression. - On a later request, the browser sends an `Available-Dictionary` header identifying what it has cached. - The server sends only the differences between the old and new versions. - A 500 KB JavaScript bundle with a one-line change could become only a few kilobytes on the wire. - The technique is especially useful for versioned JavaScript bundles, CSS, framework updates, and other incrementally changing assets. - Each release can use the immediately preceding version as its dictionary, allowing savings to continue across many deployments. - Custom and dynamic dictionaries for non-static content remain an area for future development. ## Lessons from SDCH - Google introduced Shared Dictionary Compression for HTTP (SDCH) in Chrome in 2008. - Although early adopters reported significant performance improvements, SDCH had serious security and architectural issues. - Compression side-channel attacks such as CRIME and BREACH demonstrated that attackers could infer secrets by injecting content and observing compressed response sizes. - SDCH also conflicted with the Same-Origin Policy and CORS because of its cross-origin dictionary model. - Its specification did not adequately define interactions with APIs such as the Cache API. - Chrome removed SDCH in 2017 after adoption failed to materialize. ## The Modern Standard and Remaining Challenges - RFC 9842, Compression Dictionary Transport, addresses major SDCH shortcomings. - Dictionaries are restricted to responses from the same origin, reducing conditions that enabled earlier side-channel attacks. - Chrome and Edge support the standard, while Firefox is working toward support. - Implementing the system requires servers to: - Generate or select dictionaries. - Advertise them with correct headers. - Detect `Available-Dictionary` requests. - Delta-compress responses dynamically. - Fall back cleanly for clients without dictionary support. - Cache behavior becomes more complicated because responses vary by both content encoding and dictionary availability. Cloudflare plans to offer a beta of its shared compression dictionary support on April 30, 2026. The technology is promising for frequently deployed applications and agent-heavy traffic, but broad benefits will depend on cross-browser adoption and careful handling of security and caching complexity.

Read original(opens in new tab)
cloudflare3 min readCurated summary

Introducing Flagship: feature flags built for the age of AI

AI-generated code is moving toward autonomous production deployment, making safety and controlled rollout essential. The post argues that feature flags provide the guardrails: agents can deploy disabled code, test it with limited cohorts, monitor results, and roll back automatically. Cloudflare’s new Flagship service is designed for this workflow, evaluating flags at the edge through Workers, KV, and Durable Objects. ## Feature Flags for Autonomous Deployment - Agents can ship code behind an off flag without affecting users. - They can enable features for themselves or small test cohorts, observe metrics, and expand or disable rollouts. - Humans define boundaries while flags limit the blast radius. - This separates not only deployment from release, but also routine shipping decisions from constant human attention. ## Problems with Feature Flags on Workers - Hardcoded flags are initially convenient because Workers deploy quickly. - Over time, flags become fragmented across teams, with no central visibility or audit trail. - Troubleshooting may require searching version history with tools such as `git blame`. - Calling an external flag service adds a network request to every user request, potentially introducing significant latency. - This undermines the advantage of running applications close to users at the edge. ## Why Local Evaluation Is Difficult on Workers - Traditional local-evaluation SDKs download rules into a long-lived process. - Worker isolates may be created and evicted between requests, requiring repeated initialization. - Serverless environments therefore need a distribution system with edge-local reads and managed synchronization. - Flagship uses Cloudflare KV to provide this distribution without persistent connections or per-request external calls. ## How Flagship Works - Flagship is built on Workers, Durable Objects, and KV, without external databases or centralized evaluation servers. - Durable Objects provide a globally unique, SQLite-backed source of truth for flag configuration and changelogs. - Changes are synchronized to KV within seconds and replicated throughout Cloudflare’s network. - Evaluations read configuration from KV at the edge and execute targeting and rollout logic inside the Worker isolate. - Both flag data and evaluation logic remain close to the request. ## Worker Binding and Typed Evaluation - Workers connect Flagship through a `wrangler.jsonc` binding containing a binding name and `app_id`. - The binding supports typed methods including: - `getBooleanValue()` - `getStringValue()` - `getNumberValue()` - `getObjectValue()` - `*Details()` methods return the value, matched variant, and selection reason. - Evaluation errors return the supplied default value. - Type mismatches throw exceptions because they indicate application bugs rather than temporary service failures. ## OpenFeature Integration - Flagship is built on OpenFeature, the CNCF standard for feature-flag evaluation. - It supports Workers as well as Node.js, Bun, Deno, and browser environments. - The service is currently available in closed beta. Flagship is positioned as an edge-native feature-flag system for safely automating deployment and rollout. For Cloudflare Workers, its direct binding avoids network round-trips while providing centralized configuration, targeting, auditability, and controlled release mechanisms.

Read original(opens in new tab)
cloudflare3 min readCurated summary

Agents that remember: introducing Agent Memory

Cloudflare’s Agent Memory is a managed, retrieval-based service designed to give AI agents persistent memory without continuously expanding their context windows. It addresses context rot by extracting useful information from conversations, retaining it across sessions, and retrieving synthesized answers when needed. The service is intended for production agents that run for weeks or months, where fast ingestion, affordable retrieval, and durable knowledge matter more than benchmark performance alone. ## The Challenge of Agent Memory - Larger context windows—even beyond 1 million tokens—do not eliminate context rot; excessive context can reduce model quality. - Aggressive pruning creates the opposite risk: removing information the agent may need later. - Existing memory systems vary widely: - Managed services versus self-hosted frameworks - Raw database or filesystem access versus purpose-built APIs - Full-context approaches versus retrieval-based systems - Benchmarks such as LongMemEval, LoCoMo, and BEAM help compare systems but may encourage overfitting to clean datasets that do not reflect long-running production workloads. ## Cloudflare’s Retrieval-Based Approach - Agent Memory is a managed service with an opinionated API. - It extracts and retrieves relevant information instead of exposing agents directly to a filesystem or database. - This approach is intended to: - Reduce token usage and cost - Improve retrieval performance - Support temporal reasoning, supersession, and instruction following - Keep memory operations out of the agent’s main reasoning context - Cloudflare expects programmatic querying to be useful for specialized edge cases, but not as the default interaction model. ## Memory Profiles and Core Operations Memory is organized into named profiles that can be shared across sessions, agents, and users. - `ingest`: Processes a conversation and extracts memories, typically during context compaction. - `remember`: Stores one important fact explicitly, often through direct model tool use. - `recall`: Runs the full retrieval pipeline and returns a synthesized response. - `list`: Lists stored memories. - `forget`: Removes a specific memory. For example, an agent can ingest a conversation containing a user’s preference for pnpm and dark mode, explicitly remember an operational fact such as an increased API rate limit, and later recall that the user prefers pnpm over npm. ## Integration and Supported Architectures - Agent Memory is available through a binding in Cloudflare Workers. - Agents running outside Workers can use the REST API. - The Cloudflare Agents SDK integrates it with session compaction, memory creation, and retrieval. - It can support: - Individual coding or personal agents - Self-hosted frameworks and managed agent services - Autonomous background agents that must survive restarts - Custom agent harnesses - Shared knowledge between engineers, agents, and tools - Shared profiles can preserve coding conventions, architectural decisions, and other organizational knowledge that might otherwise be lost during context pruning. ## Practical Recommendation Agent Memory is positioned as a default persistent-memory layer for production agents: use ingestion during compaction, explicit remembering for critical facts, and retrieval when the agent needs historical context. Its private beta is particularly aimed at long-running, multi-session workloads where maintaining useful memory is more important than simply fitting more text into the context window.

Read original(opens in new tab)
cloudflare3 min readCurated summary

Redirects for AI Training enforces canonical content

Cloudflare argues that AI training crawlers often ignore deprecation notices, `noindex`, and canonical tags, causing outdated documentation to enter training data. Its new Redirects for AI Training feature converts qualifying canonical tags into HTTP 301 redirects for verified AI training bots, sending them to current content instead. The feature is intended to enforce content freshness without affecting human visitors, search crawlers, or AI agents. ## The Problem with Deprecated Content - Cloudflare’s older Wrangler documentation includes deprecation banners, `noindex`, and canonical tags. - AI training crawlers consumed deprecated and current documentation at roughly the same rate. - Unlike humans, crawlers may ingest the entire page, treating warnings as ordinary text. - Blocking crawlers with `robots.txt` creates a dead end and does not tell them where the replacement content is. - Outdated content can persist in trained models even after the original page has been updated. ## Canonical Tags as Redirect Instructions - The HTML `<link rel="canonical">` tag identifies the authoritative version of a page. - Canonical tags appear on an estimated 65–69% of websites and are often generated automatically by CMS platforms. - Cloudflare’s feature uses this existing content hierarchy rather than requiring separate redirect rules. - For verified AI training crawlers, a non-self-referencing canonical becomes an HTTP `301 Moved Permanently` redirect. ## How Redirects for AI Training Works - Cloudflare identifies eligible requests using `cf.verified_bot_category`. - The AI Crawler category includes bots such as GPTBot, ClaudeBot, and Bytespider. - Cloudflare inspects the requested page’s HTML: - If it has a canonical URL on the same domain, the crawler is redirected there. - Self-referencing canonicals do not trigger redirects. - Cross-origin canonicals are excluded. - Human traffic, search crawlers, and AI Assistant or AI Search bots are unaffected. ## Limitations and Alternatives - The feature cannot remove outdated material already ingested into training datasets. - It does not cover unverified crawlers. - AI Agents visiting deprecated pages are not redirected. - Traditional redirect rules can work for a small number of paths, but they require manual maintenance, user-agent tracking, and rule capacity. - Canonical-based enforcement stays aligned with content changes automatically. ## Cloudflare’s Documentation Test - In March 2026, legacy Workers documentation received thousands of crawls from OpenAI, Anthropic, and Meta. - An AI assistant provided the deprecated Wrangler syntax `kv:key put` instead of the current `wrangler kv key put`. - After enabling the feature, 100% of AI training crawler requests for pages with non-self-referencing canonicals were redirected during the first seven days. - Cloudflare expects this to improve future AI answers, though the impact depends on training pipelines and recrawl timing. ## Enabling the Feature - In the Cloudflare dashboard, go to **AI Crawl Control → Quick Actions → Redirects for AI training** and enable the toggle. - Path-specific controls are available through Configuration Rules and Cloudflare for SaaS. - Cloudflare also added AI crawler response-status analysis to Radar’s AI Insights page, covering 2xx, 3xx, 4xx, and 5xx responses. Website owners with canonical tags can use the feature to automatically steer verified AI training crawlers toward current content. It is most useful for sites with frequently changing documentation and many deprecated URLs, but it should complement—not replace—ordinary redirects, access controls, and content maintenance.

Read original(opens in new tab)
cloudflare3 min readCurated summary

Unweight: how we compressed an LLM 22% without sacrificing quality

Unweight is Cloudflare’s lossless compression system for LLM weights, reducing model size by 15–22% while preserving bit-exact outputs. It targets the memory-bandwidth bottleneck in GPU inference by compressing weights in HBM and decompressing them directly into fast on-chip memory before tensor-core computation. On Llama-3.1-8B, the approach saves roughly 3 GB of VRAM and enables more models to run per GPU. ## The GPU Memory Bottleneck - LLM inference is often limited by memory bandwidth rather than computation. - Each generated token requires reading the model’s weights from GPU high-bandwidth memory (HBM). - NVIDIA H100 tensor cores can process data far faster than HBM can supply it. - Smaller weights reduce the amount of data transferred across the memory bus. - Decompression must be carefully integrated: if it adds latency that cannot overlap with matrix multiplication, token generation becomes slower. ## Why Lossless Compression Matters - Quantization commonly converts 16-bit values into 8- or 4-bit integers. - Because quantization is lossy, it can change model behavior and response quality unpredictably. - Unweight instead preserves exact outputs and does not require specialized hardware. - Existing systems were unsuitable because they focused on CPU decompression, custom FPGA hardware, or consumer GPUs rather than Hopper-generation GPUs and production inference. ## Compressing BF16 Weights - BF16 values contain: - A sign bit - An 8-bit exponent - A 7-bit mantissa - Sign and mantissa values appear largely random and are difficult to compress. - Exponents are highly predictable: the 16 most common exponent values account for more than 99% of weights in a typical layer. - Unweight applies Huffman coding to exponent bytes while leaving sign and mantissa bits unchanged. - Rare exponents are handled by storing an entire row of 64 weights verbatim, avoiding per-element branching during decoding. ## Selective Compression of Model Layers - Unweight compresses the MLP gate, up, and down projection matrices. - These matrices represent roughly two-thirds of model parameters and generate substantial memory traffic during decoding. - Attention weights, embeddings, and layer norms remain uncompressed. - The exponent compression produces about 30% savings in the targeted streams and approximately 20% reduction in total MLP weight size. - Overall model-size reductions reach 15–22%. ## Direct GPU Decompression - Model weights normally reside in large but slower HBM and are staged into small, fast shared memory before computation. - Conventional approaches decompress full matrices back into HBM and then run standard matrix multiplication, creating additional memory traffic. - Unweight decompresses weights in shared memory and feeds them directly to tensor cores. - Different execution strategies are used depending on the weight matrix and batch size. - An autotuner selects the fastest strategy for each workload. ## Results and Availability - Tests on Llama-3.1-8B achieved: - Around 30% compression for MLP weights - 15–22% reduction in total model size - Approximately 3 GB of VRAM savings - The savings allow more models to fit on each GPU, potentially reducing inference cost and improving global deployment coverage. - Cloudflare is publishing a technical paper and open-sourcing the GPU kernels. Unweight demonstrates that lossless, inference-time compression can improve GPU utilization without changing model behavior. The practical recommendation is to compress the portions of a model that dominate memory traffic while integrating decoding directly into the GPU execution path.

Read original(opens in new tab)
cloudflare2 min readCurated summary

Agents Week: network performance update

Cloudflare reports that it became the fastest network in 60% of the world’s 1,000 largest networks by December 2025, up from 40% during Birthday Week 2025. The improvement came from both expanding its global points of presence and optimizing connection-handling software. Cloudflare says it is continuing to target the remaining networks where competitors still lead. ### Measuring Network Performance - Cloudflare analyzes the 1,000 largest networks by estimated population, using APNIC data. - It measures TCP connection time—the time required to complete a TCP handshake—as a practical indicator of users’ perceived Internet speed. - Rankings use the **trimean**, a weighted average of the 25th percentile, median, and 75th percentile, reducing the influence of outliers. - Data comes from Real User Measurements: browsers loading Cloudflare error pages silently download small files from Cloudflare, Amazon CloudFront, Google, Fastly, and Akamai under real network conditions. ### Expanding Points of Presence - New locations in Constantine, Algeria; Malang, Indonesia; and Wroclaw, Poland brought Cloudflare physically closer to users. - In Wroclaw, free-user average round-trip time fell from 19 ms to 12 ms, a 40% improvement. - In Malang, Enterprise traffic improved from 39 ms to 37 ms, a 5% reduction. - However, new locations alone did not account for the increase from 40% to 60% of networks. ### Improving Connection Handling - Cloudflare optimized the software responsible for connection establishment, SSL/TLS termination, traffic management, and proxying. - HTTP/3 adoption and changes to congestion-window management reduced processing time. - Improvements in CPU and memory efficiency allow the global network to handle connections more effectively. - Cloudflare compares this to improving both the efficiency of highway toll booths and the routing of traffic between them. ### Results by December 2025 - Cloudflare was the fastest provider in 60% of the largest networks. - Between September and December 2025, it became fastest in: - 40 additional countries - 261 additional networks - 54 additional U.S. autonomous systems (ASNs) - During December, Cloudflare was on average 6 ms faster than the next-fastest provider. Cloudflare’s conclusion is that continued gains in network reach and software efficiency can produce measurable improvements for users. It plans to focus on the remaining networks where it is narrowly behind competitors, with the long-term goal of being fastest worldwide.

Read original(opens in new tab)
meta3 min readCurated summary

Capacity Efficiency at Meta: How Unified AI Agents Optimize Performance at Hyperscale

Meta’s Capacity Efficiency Program uses AI agents to automate both the discovery and resolution of infrastructure performance issues. By combining standardized tools with encoded expertise from senior efficiency engineers, the platform turns investigations that once took hours into minutes and has recovered hundreds of megawatts of power. The approach aims to let Meta scale efficiency improvements across more product areas without proportionally increasing engineering headcount. ## Capacity Efficiency at Hyperscale - At Meta’s scale, even a 0.1% performance regression can significantly increase power consumption across systems serving more than 3 billion people. - The program has two complementary functions: - **Offense:** Proactively identify and implement optimizations. - **Defense:** Detect production regressions, identify their causes, and deploy mitigations. - Human investigation is often the bottleneck, requiring engineers to analyze profiling data, review documentation and prior fixes, inspect deployments, and search internal discussions. - AI automation can reduce roughly 10 hours of manual diagnosis to about 30 minutes. ## A Unified Platform for AI Efficiency Agents - Meta built one platform for both offensive and defensive workflows because they share the same basic structure: - Gather relevant technical context. - Apply domain-specific reasoning. - Produce a code change for review. - **MCP tools** provide standardized interfaces for querying profiling data, retrieving experiment results, examining configuration history, searching code, and accessing documentation. - **Skills** encode expert reasoning, including which tools to use and how to interpret their results. - The same tools support both use cases, while specialized skills handle different optimization and regression scenarios. ## Defense: Automated Regression Resolution - FBDetect monitors noisy production time series and can identify regressions as small as 0.005%. - Traditional root-cause analysis correlates the regression with recent pull requests or configuration changes. - Previously, teams often rolled back problematic changes—reducing engineering velocity—or left them unresolved, allowing resource waste to accumulate. - The AI Regression Solver: - Identifies affected functions and regression symptoms. - Locates the responsible pull request, files, and changed lines. - Applies mitigation expertise appropriate to the codebase, language, or regression type. - Generates a corrective pull request and sends it to the original author for review. - Faster resolution prevents small regressions from compounding across Meta’s infrastructure. ## Offense: Converting Opportunities into Code - Efficiency opportunities describe potential improvements to existing code, but implementing them traditionally required substantial investigation and engineering time. - Meta’s AI workflow gathers: - Opportunity metadata. - Optimization documentation. - Examples of similar fixes. - Relevant files and functions. - Validation criteria. - Skills then apply specialized knowledge, such as memoizing a function to reduce CPU usage. - The agent generates a guarded candidate fix, checks syntax and style, validates that it addresses the intended issue, and presents the change in an engineer’s editor for review or one-click application. - This expands the number of optimization opportunities engineers can pursue manually. ## Scaling Efficiency with AI - The platform has already recovered hundreds of megawatts of power—enough to supply hundreds of thousands of U.S. homes for a year. - Automated regression handling reduces ongoing waste, while automated opportunity resolution increases the volume of proactive improvements. - The long-term goal is a self-sustaining efficiency engine in which AI handles the long tail of investigations and fixes, allowing engineers to focus on new products and higher-value work. Meta’s approach recommends treating performance expertise as reusable, composable software: standardize access to engineering data, encode proven reasoning into skills, and let agents carry issues from detection through ready-to-review code changes.

Read original(opens in new tab)
github3 min readCurated summary

How GitHub uses eBPF to improve deployment safety

GitHub uses eBPF to prevent deployment scripts from depending on GitHub or other services that may be unavailable during an outage. The approach applies network restrictions only to deployment processes, preserving normal traffic for stateful production hosts. By combining Linux cgroups with eBPF’s `BPF_PROG_TYPE_CGROUP_SKB` hooks, GitHub can detect or block unsafe outbound calls before they create circular deployment dependencies. ## The Deployment Circular Dependency - GitHub hosts its own source code on `github.com`, creating a basic dependency: GitHub may need GitHub to deploy a fix. - GitHub mitigates this with: - A code mirror used for “fix-forward” deployments. - Prebuilt assets used to roll back changes. - Additional circular dependencies can still be introduced by deployment scripts, internal services, or tools that download binaries dynamically. ## Three Types of Circular Dependencies The post uses a hypothetical MySQL outage to illustrate how deployment recovery can fail: - **Direct dependencies** - A deployment script downloads the latest release of an open-source tool from GitHub. - If GitHub cannot serve release data, the deployment cannot complete. - **Hidden dependencies** - A required tool is already installed locally but checks GitHub for updates when it runs. - The tool may fail or hang if that update check cannot connect. - **Transient dependencies** - The deployment calls another internal service, such as a migrations service. - That service then attempts to download a binary from GitHub, causing the failure to propagate back to the deployment. ## Why Manual Review Is Insufficient - Teams responsible for stateful hosts traditionally review deployment scripts for circular dependencies. - Many indirect or unexpected dependencies are discovered only during incidents, when they can delay recovery. - Blocking `github.com` at the host level would be too disruptive because stateful hosts continue serving customer traffic during deploys, drains, and restarts. ## Per-Process Network Filtering with eBPF - eBPF allows custom programs to run inside the Linux kernel and attach to low-level operations such as networking. - GitHub focused on `BPF_PROG_TYPE_CGROUP_SKB`, which can inspect network egress for a specific cgroup. - Linux cgroups provide process grouping, isolation, and resource controls without requiring Docker. - The proposed design: - Create a dedicated cgroup. - Place only the deployment script and its processes inside it. - Restrict or monitor outbound network access for that group. - Leave the host’s ordinary production traffic unaffected. ## Proof of Concept with Go and eBPF - The proof of concept uses Go and the `cilium/ebpf` library. - The library simplifies: - Compiling and loading eBPF programs. - Attaching programs to kernel hooks. - Reading and updating eBPF maps. - The example attaches an egress program to `/sys/fs/cgroup/system.slice`. - An eBPF array map tracks the number of egress packets, while the Go program periodically reads and reports the counter. - The same mechanism can be extended from packet counting to selectively allowing or blocking network traffic. GitHub’s approach moves dependency validation from manual inspection into enforcement at runtime. Restricting only deployment processes with cgroups and eBPF provides a practical way to make emergency deployments independent of services that may be down, without blocking the production workloads sharing the same host.

Read original(opens in new tab)
meta3 min readCurated summary

Post-Quantum Cryptography Migration at Meta: Framework, Lessons, and Takeaways

Meta argues that organizations should begin migrating to post-quantum cryptography (PQC) before quantum computers become practical. The “store now, decrypt later” threat means attackers may already be collecting encrypted data for future decryption, making long-lived sensitive information vulnerable today. Meta’s experience suggests a phased strategy based on risk prioritization, cryptographic inventories, technical readiness, deployment, and operational guardrails. ## Why PQC Migration Is Urgent - Quantum computers are expected to eventually break conventional public-key cryptography, potentially within 10–15 years. - Attackers can use “store now, decrypt later” (SNDL) attacks by collecting encrypted data today and decrypting it once quantum capabilities mature. - NIST and the UK NCSC have issued migration guidance, including target timeframes such as 2030 for protecting critical systems. - NIST has standardized algorithms including: - **ML-KEM (Kyber)** for key encapsulation - **ML-DSA (Dilithium)** for digital signatures - **HQC**, which includes contributions from Meta cryptographers ## Meta’s Migration Goals Meta’s multi-year migration is guided by four objectives: - **Effectiveness:** Protect systems against quantum-enabled adversaries. - **Timeliness:** Deploy protections as standards and technologies evolve. - **Performance:** Minimize latency, resource use, and user impact. - **Cost efficiency:** Balance investment against the risk and sensitivity of each use case. ## PQC Migration Levels Meta proposes a maturity ladder that measures how quickly an organization can respond to a relevant quantum event, such as a major technical breakthrough, new standards, or changing industry practices. - **PQ-Unaware:** The organization has not recognized the quantum threat. - **PQ-Aware:** The threat and eventual requirements have been assessed, but design work has not begun. - **PQ-Ready:** A suitable PQC solution has been identified or prepared, but deployment is deferred because of cost, prioritization, or other constraints. - **PQ-Hardened:** All currently available protections have been implemented, but complete mitigation is impossible because required primitives—such as efficient post-quantum OPRFs—do not yet exist. - **PQ-Enabled:** A post-quantum-secure solution is deployed for the use case. This is the desired end state for every application. Even reaching PQ-Ready can reduce future reaction time and create useful technical and organizational foundations, although it does not itself protect systems from quantum attacks. ## Meta’s PQC Migration Strategy Meta describes migration as several potentially overlapping workstreams: - **Define prioritization:** Classify applications by high, moderate, or low risk so the most exposed use cases move first. - **Build a cryptographic inventory:** Identify where cryptography is used and which applications rely on quantum-vulnerable algorithms. - **Address external dependencies:** Track standards, PQC-capable hardware security modules, and the maturity of available implementations. - **Implement PQC components:** Build reusable post-quantum cryptographic capabilities for later integration. - **Deploy guardrails:** Update cryptographic standards, prevent creation of new vulnerable keys, and restrict affected APIs. - **Integrate protections:** Apply PQC components to prioritized use cases and internal traffic. ## Prioritizing Applications The first prioritization category focuses on applications vulnerable to attacks that can begin now and be completed later using quantum algorithms such as Shor’s algorithm. - Applications using quantum-vulnerable public-key encryption or key-exchange mechanisms are considered high priority. - Systems handling sensitive data with long confidentiality requirements are especially exposed to SNDL attacks. - Risk-based prioritization helps organizations avoid attempting a costly, simultaneous migration of every application. Organizations should begin by identifying high-value and long-lived data, inventorying vulnerable cryptography, and moving each use case progressively toward PQ-Enabled status.

Read original(opens in new tab)