Techlist.io - Korean Tech Blog Curator

cloudflare3 min readCurated summary

Give any website a WebMCP interface

Cloudflare is launching a developer preview of WebMCP that lets browser-based AI agents use websites through structured tools instead of scraping pages or navigating human-oriented interfaces. Cloudflare injects a browser-side bridge at the edge, requiring no origin-code changes or redeployment. The system currently supports tool packs such as Content Credentials and proxying an existing MCP server, with all preview tools executing in the visitor’s browser. ## Why WebMCP Matters - Traditional websites assume a human will read pages, click controls, and submit forms. - AI agents increasingly visit the web but often rely on crawlers that copy content away from the original site. - WebMCP provides a browser-native interface through `document.modelContext`. - Sites can expose tools that agents can call directly, reducing navigation overhead and token usage. - The standard is experimental in Chrome 146 and normally requires site-level implementation. ## Cloudflare’s No-Code Integration - Cloudflare adds WebMCP support through a Dashboard setting. - Enabled sites receive groups of related tools called tool packs. - New packs can be activated later without redeploying the site. - The preview includes: - A Content Credentials pack for reading C2PA metadata. - A Site MCP Server pack for exposing tools from an existing MCP server. ## Edge Injection and Browser Bridge - Cloudflare uses `HTMLRewriter` to inject a same-origin bridge script into HTML responses. - The injection leaves the site’s original HTML and application code otherwise unchanged. - The script includes: - `data-packs`, identifying enabled tool packs. - `data-mcp-url`, identifying the site’s MCP endpoint, defaulting to `/mcp`. - The bridge exits harmlessly when the browser lacks WebMCP support. - It registers tools with `document.modelContext.registerTool`. - Static packs define tools in advance, while dynamic packs discover available tools during startup. ## MCP Tools and Site Sessions - Tools use standard MCP `Tool` and `CallToolResult` types. - Existing MCP clients can interact with these browser tools without special integration. - For a site’s MCP server, the bridge: - Retrieves the server’s tool definitions through `tools/list`. - Registers browser-side proxy tools. - Sends calls to the site’s MCP endpoint using same-origin requests. - Preserves the visitor’s existing session through `credentials: "same-origin"`. - Preview tools run locally in the visitor’s browser, without requests to Cloudflare-owned services. - The edge worker architecture leaves room for future packs that use Workers AI or AI Search. ## Reading Content Credentials - The Content Credentials pack analyzes C2PA metadata embedded in images. - `scan_images_c2pa` scans images on the page and reports: - Image counts and formats. - Whether C2PA metadata exists. - Manifest counts. - Claim generators, titles, and signing organizations. - `inspect_image_c2pa` retrieves more detailed manifest data, including edit history, authorship, and certificates. - The reader examines only the metadata near the beginning of the image rather than downloading or processing the entire image. - In the current preview, credentials are reported but not cryptographically verified; results therefore indicate `signatureVerified: false`. Cloudflare’s approach makes WebMCP adoption largely configuration-driven: sites can expose agent-friendly capabilities without changing their origin code, while retaining browser execution and the visitor’s authentication context. Developers should treat it as an experimental preview, especially because browser support and credential verification are still evolving.

Read original(opens in new tab)
discord3 min readCurated summary

Discord Patch Notes: August 4, 2026

Discord’s August 4, 2026 patch focuses on usability, reliability, performance, and interface cleanup across desktop and mobile. The update introduces another phase of the User Settings redesign and upgrades the desktop client to Electron 42, producing modest CPU improvements. Most remaining changes are bug fixes for overlays, profiles, themes, input fields, responsive layouts, and mobile rendering. ## User Settings Redesign and Performance - “Activity” has been renamed **Games & Apps**. - “Content & Social” is now **Messaging Permissions**. - Related options, including Authorized Apps, Connections, operating-system settings, and keybinds, have been consolidated. - Wording and styling were updated across Data & Privacy and Activity Privacy settings. - Discord’s desktop client now runs on **Electron 42**, improving maintenance and slightly reducing CPU usage. ## Desktop Fixes - Fixed email changes failing after CAPTCHA completion. - Resolved black screens caused by the game overlay when certain games were tabbed out. - Stabilized profile popouts that resized or jittered when game data loaded. - Fixed theme previews permanently changing the default theme. - Restored the ability to reposition the overlay chat widget and prevented overlay actions from grabbing images or opening dialogs in the wrong window. - Corrected activity-status issues when activity detection was disabled. - Fixed custom theme switching after using the theme editor. - Improved server selectors and long dropdowns, preventing selections from reverting or scrolling back to the top. - Corrected invite validation after users left a server to make room under the server limit. - Restored access to main profiles when users also had private per-server profiles. - Fixed numerous layout and interaction issues involving coachmarks, modal positioning, focus rings, close buttons, profile cards, and resized windows. - Improved truncation for long server names, connection names, custom statuses, and other profile text. - Fixed visual defects involving badges, gradients, phone verification fields, wishlist cards, and game-shop modals. - On Linux, Stable and Canary clients now use distinct window classes so taskbars and window managers group them correctly. ## Mobile Fixes - Improved Android game-profile scrolling performance and reduced dropped frames. - Fixed clipped profile controls on iPad. - Resolved an Android gesture issue that could leave the You Bar on a blank screen. - Fixed server icons failing to render after returning from picture-in-picture mode. - Corrected iOS name-style effects such as “Toon” and “Pop” when status text was present. - Prevented multiline pastes from expanding single-line fields such as nicknames, server names, and role names. - Fixed clipped game titles in the Active Now card, especially on iPad and with larger text sizes. - Improved display of long connection links on Android and iOS profiles. - Corrected mismatched colors on Android’s Server Boosts page in Light theme. - The provided patch notes end during an additional Android fix concerning the last friend in a list. The changes are merged but may still be rolling out across individual platforms. Users experiencing remaining issues are directed to Discord’s community-run bimonthly bug megathread.

Read original(opens in new tab)
gitlab2 min readCurated summary

GitLab Secrets Manager adds ESO, Terraform, API support

GitLab Secrets Manager expands beyond CI/CD by supporting Kubernetes, Terraform/OpenTofu, CLI tools, and external automation. Built on OpenBao and compatible with Vault APIs, it provides one centrally managed secret store with consistent access controls and auditing. The result is fewer duplicated credential stores and safer secret retrieval across the software delivery lifecycle. ## Kubernetes with External Secrets Operator - ESO uses its Vault provider to retrieve secrets from GitLab Secrets Manager. - A Kubernetes workload uses a short-lived GitLab-minted JWT to authenticate with OpenBao. - A `SecretStore` configures: - The Vault-compatible server and KV v2 mount - The GitLab organization, group, and project namespace - JWT authentication and the Kubernetes secret containing the token - An `ExternalSecret` maps remote secrets to a Kubernetes `Secret`. - ESO refreshes values according to `refreshInterval`, allowing rotated credentials to reach workloads without redeployment. - `remoteRef.key`, `property`, and `secretKey` define the source path, field, and destination key. ## Terraform and OpenTofu Integration - Terraform can retrieve secrets at plan or apply time instead of storing them in `.tfvars` files or CI/CD variables. - A script obtains a minted JWT and connection metadata through Terraform’s `external` data source. - The Vault provider uses that JWT to authenticate against GitLab Secrets Manager. - The `vault_kv_secret_v2` data source reads the required secret. - Outputs containing secrets should be marked `sensitive`, though downstream Terraform state handling still requires care. ## OpenBao and Vault CLI - Existing Vault-compatible scripts can access GitLab Secrets Manager without using the API directly. - Users configure `VAULT_ADDR` and `VAULT_NAMESPACE`. - A minted JWT is exchanged for an OpenBao client token through the configured JWT authentication path. - The `vault kv get` command then retrieves secrets from the KV mount. ## Secrets Manager API - The API supports automation outside GitLab CI/CD, Kubernetes, and Terraform. - A service account requests an access token through GitLab’s project API. - The response supplies the Vault server, namespace, mount, secrets path, JWT authentication path, and role. - External systems can use this information to authenticate and fetch secrets without hardcoded credentials or separate variable files. GitLab Secrets Manager is most useful when multiple deployment tools need the same credentials. Centralizing secrets in the OpenBao-backed store, using short-lived JWT authentication, and integrating through ESO, Terraform, CLI, or the API can reduce duplication and improve rotation and auditing.

Read original(opens in new tab)
gitlab3 min readCurated summary

Confidential AI for GitLab Self-Hosted

Privatemode AI enables GitLab Duo Self-Hosted to provide modern coding agents without exposing source code to GitLab, a cloud provider, or an AI operator. It uses confidential computing and remote attestation to keep prompts, code, and completions encrypted even during inference, avoiding both the compliance risks of public AI services and the operational burden of running private GPU infrastructure. ## The productivity gap for regulated teams - GitLab Duo supports more than autocomplete, including: - Merge request reviews - Cross-file refactoring - Test generation and execution - Agentic workflows running in CI - These features normally require sending source code and prompts to an external model provider. - For organizations handling regulated software or proprietary intellectual property, that data transfer may violate contracts, regulations, or internal policy. ## Why self-hosting is difficult - Affected sectors include finance, healthcare, defense, government, and critical infrastructure. - Requirements may come from NIS2, DORA, GDPR, BaFin, BSI C5, healthcare rules, and broader data-sovereignty expectations. - Public AI SaaS is often unacceptable because code leaves the organization. - Private-cloud or VPC services reduce exposure but still require trusting the cloud and service operators with plaintext. - Running models internally preserves privacy but requires expensive GPUs, specialized staff, and ongoing model operations, while often lagging behind frontier models. ## Confidential computing as the solution - Confidential computing uses hardware-based trusted execution environments (TEEs) to encrypt data while it is being processed. - The architecture relies on technologies such as: - AMD SEV or Intel TDX for CPU protection - NVIDIA Confidential Computing for GPU protection - AES-256 encryption for data in transit and at rest - Remote attestation verifies that approved code is running inside the TEE before any data is sent. - Prompts, source code, context, and completions are decrypted only inside the protected environment. - The operator and underlying cloud provider cannot inspect the data through normal system or infrastructure access. ## Privatemode AI - Privatemode, developed by Germany-based Edgeless Systems, provides confidential inference through an OpenAI-compatible API. - Its client-side proxy manages encryption and remote attestation transparently. - Existing tools and SDKs using the standard `/v1` API can work without major changes. - The current highlighted coding model is Kimi K2.6 with a 256K context window; Kimi K3 and GLM are expected to follow. - The service is presented as production-ready and already used by public-sector, financial, defense, and regulated-industry organizations. - Its post-quantum-safe cryptography is intended to protect against “harvest now, decrypt later” attacks. ## GitLab Duo integration - GitLab Duo Self-Hosted connects to a self-hosted AI Gateway. - The Gateway forwards requests to the Privatemode proxy as an OpenAI-compatible endpoint. - The proxy: - Encrypts requests before they leave the organization’s network - Verifies the remote TEE through attestation - Forwards only after verification succeeds - Developers continue using Duo features such as Code Suggestions, Chat, Code Review, and agentic workflows without changing their experience. The recommended approach for regulated organizations is to combine GitLab Duo Self-Hosted with a confidential-computing provider such as Privatemode. This provides modern AI coding capabilities while replacing contractual privacy promises with hardware-enforced protection, without requiring the organization to operate its own LLM infrastructure.

Read original(opens in new tab)
discord3 min readCurated summary

General Availability of Mobile Platform Support Arrives in Discord Social SDK Version 1.10

Discord Social SDK 1.10 makes mobile support generally available for iOS and Android, extending Discord’s social features into mobile games. The update focuses on reducing friction through improved Account Linking, richer Rich Presence options, and mobile commerce. Early partners including Tencent, Scopely, and Ninja Kiwi report stronger social engagement, easier squad formation, and increased community activity. ## Mobile Platform Support - The SDK supports: - Android 7.0 and later - iOS 15.1 and later - C++, Unreal Engine, and Unity projects - The goal is to bring desktop-like Discord social experiences to smaller screens and on-the-go play. - Discord says the SDK was refined with launch partners and tested against real mobile player behavior. ## Rich Presence on Android - Android games can now share live Rich Presence through multiple methods: - Account Linking - Remote Procedure Call (RPC) - These integrations improve game visibility and help players share what they are playing with their Discord communities. ## Faster Mobile Account Linking - Deeplink support simplifies the account-linking process. - Players can connect their Discord accounts directly and with fewer steps. - The reduced friction is intended to improve retention, engagement, onboarding, and campaign activation. ## Discord Social Commerce on Mobile - Discord is expanding Social Commerce to mobile games. - Developers can sell in-game items through Discord to existing players, friends, and potential new customers. - Early Marvel Rivals pilot results showed: - 41% of purchases were gifts. - 25% of gift buyers were lapsed or new players. - These results suggest Discord’s social graph can support both monetization and player acquisition. ## Tencent: Arena Breakout - Arena Breakout used the SDK to address fragmented cross-platform friend networks and slow squad formation. - Its integration included: - A Unified Friends List - Game Invites - Improved mobile Account Linking - Reported outcomes included faster squad formation, more cross-platform messaging, and increased activity in the game’s official Discord server. - Tencent plans to combine Discord social features with limited-time in-game events. ## Scopely: Marvel Strike Force - Scopely integrated Discord into its turn-based RPG to support: - Player communication - Alliance management - Community interaction - Customer support and moderation - The integration produced immediate account linking among new players and rapid creation of alliance channels. - Community feedback was strongly positive, with players mainly wishing the integration had arrived earlier. ## Tencent: Delta Force - Delta Force used Discord to connect in-game play with its broader player community. - The initial integration focused on Account Linking and the Unified Friends List. - Single-tap mobile linking replaced browser redirects and manual logins, improving onboarding and campaign activation. - The integration supported squad formation and community-created content, with future plans for Discord discoverability and Rich Presence invites. ## Ninja Kiwi: Bloons TD 6 - Bloons TD 6 began integrating the SDK with Account Linking and the Unified Friends List. - The provided article text ends before describing the full integration or its results. The update positions Discord’s Social SDK as a way for mobile developers to strengthen retention, social discovery, community participation, and monetization. Developers targeting mobile games should consider the SDK’s deeplink Account Linking, Android Rich Presence, Unified Friends List, invites, and emerging commerce capabilities.

Read original(opens in new tab)
cloudflare3 min readCurated summary

Cloudflare is the only vendor named a Visionary in 2026 SASE and SSE reports

Cloudflare argues that SASE and SSE are entering a major transition driven by AI agents, shadow applications, post-quantum threats, and increasingly distributed workforces. It presents Cloudflare One as a unified, programmable platform designed to address these pressures without the fragmented architectures, complex deployments, and hidden costs associated with legacy vendors. The company cites its recognition as a Visionary in both Gartner’s 2026 SASE and SSE Magic Quadrants as validation of this approach. ## The SASE Market’s Architectural Gap - Many SASE platforms are assembled through mergers and acquisitions, creating disconnected products and difficult deployments. - Cloudflare’s “connectivity cloud” uses one global network to connect and protect employees, AI agents, and infrastructure. - AI security has focused primarily on human interactions with generative AI, leaving autonomous agents and MCP server sprawl insufficiently governed. - Cloudflare claims its SASE platform provides shared visibility and policy controls for both humans and AI agents, including limits on AI inference costs. - Post-quantum protection is presented as an immediate requirement against “harvest-now, decrypt-later” attacks, rather than a future concept. - Cloudflare emphasizes predictable SASE bundles instead of charging separately for advanced capabilities or remote and office use cases. ## Technological Pressures Reshaping SASE - **AI-generated applications:** Employees can rapidly create internal “vibe-coded” tools without IT oversight. SASE platforms will need to automatically apply zero trust access, WAF, API protection, and DLP. - **Autonomous AI agents:** Future systems must issue narrowly scoped credentials for individual tasks, evaluate agent intent, and detect abnormal tool-call activity. - **Post-quantum agility:** Organizations need adaptable post-quantum encryption now, while standards continue to evolve. Cloudflare says it aims to deliver a fully quantum-secure SASE platform by 2028. - **Architectural consolidation:** Genuine platform consolidation requires shared code, control, data, and infrastructure planes—not merely multiple products marketed as a single platform. - These changes are described as current customer requirements rather than distant predictions. ## Cloudflare’s Unified Architecture - Cloudflare says it built its SASE platform from the ground up on a single global network rather than combining unrelated security products. - A composable architecture allows new security capabilities to be introduced without waiting for lengthy integration cycles. - Administrators can use familiar SASE policies to secure human AI prompts, AI-agent connections, and MCP servers. - New AI applications can inherit existing zero trust controls instead of requiring security to be retrofitted later. ## Easier SASE Deployment - Legacy platforms often route traffic through multiple inspection points, producing “tromboning,” capacity-planning challenges, and complicated operations. - Cloudflare claims every service runs across its network, eliminating specialized appliance silos and reducing deployment complexity. - Common tasks—such as extending zero trust to an application, adding DLP to Gateway traffic, or connecting an office—are intended to take days or weeks rather than months or years. - The platform is positioned as operating more like a modern SaaS service than a collection of separately managed security engines. ## Programmable SASE - Cloudflare distinguishes true programmability from basic GUI automation and APIs layered over inflexible products. - Its SASE platform runs alongside the company’s edge developer platform, allowing customers to integrate custom code directly into the security fabric. - This design is intended to let organizations enrich access decisions with real-time signals and adapt policies to their own requirements. Cloudflare’s recommendation is effectively to choose SASE platforms built on unified, composable infrastructure that can govern people, applications, and autonomous agents together. Organizations should prioritize integrated policy enforcement, native post-quantum readiness, predictable pricing, and genuine programmability over loosely bundled legacy products.

Read original(opens in new tab)
meta3 min readCurated summary

From User Sequences to Scaling Laws: A Multi-Stage Architecture for Meta’s Ads Ranking

Meta’s new sequence-learning platform improves ads recommendations by separating deep offline user modeling from fast online ranking. Combined with dense tokenization and target-aware attention, it enables richer behavioral representations, predictable compute-to-performance scaling, and major gains: 6% more Instagram conversions, 3% more Facebook conversions, and 3.5% more Facebook ad clicks. The system is also a core part of Meta’s Generative Ads Recommendation Model (GEM). ## Challenges of Earlier Sequence Models - Ads systems must rank thousands of candidates within milliseconds and process millions of candidates per second. - Hybrid architectures typically use: - One model for user event sequences. - Another for sparse feature interactions. - This design can cause: - Lossy knowledge transfer between components. - Continued dependence on manually engineered features. - Scaling limits caused by interference between sequence modeling and ranking. - Increasing sequence lengths and transformer capacity can therefore raise serving costs without delivering proportional improvements. ## Multi-Stage Sequence Modeling Meta separates sequence learning into two complementary stages: - **Offline user modeling** - Processes long user histories asynchronously. - Uses deep transformer models with thousands of events and multiple layers. - Produces cached, user-level embeddings that represent long-term behavioral patterns. - Keeps user features separate from ad and context features so embeddings remain independent of individual candidates. - **Online ranking** - Combines cached user embeddings with fresh user signals, ad features, and context. - Performs final ranking under strict latency requirements. - Uses a lightweight architecture optimized for real-time serving. This separation allows the offline model to grow in depth, width, and sequence length without proportionally increasing online serving costs. ## Dense Tokenization and Target-Aware Attention - **Dense tokenization** - Converts sparse features and sequential behavioral data into a shared dense vocabulary. - Allows the model to learn feature interactions directly instead of relying on manually engineered cross-features. - **Target-aware multi-head attention** - Combines user behavior sequences with the specific ad candidate being scored. - Lets each attention layer determine which past behaviors matter for that candidate. - Stacked attention blocks capture increasingly complex interactions and compress long histories into compact representations. - The approach is designed to be memory-efficient while preserving candidate-specific information. ## Predictable Scaling Laws - On real-world ads traffic, the architecture shows an LLM-like log-linear relationship between compute and recommendation performance. - Improvements were measured using normalized entropy across: - Model depth. - Model width. - Sequence length. - Content and semantic enrichment. - The scaling behavior suggests the architecture is well suited to continued investment in sequence learning, despite recommendation systems combining sparse IDs with temporal data rather than dense text. ## Scaling Strategies - **Balanced model shape** - Depth, width, and sequence length should grow together. - Scaling only one dimension can create bottlenecks and diminishing returns. - Meta calls this the “scaling synergy principle.” - **Multi-stage tunability** - Online models offer strong improvements per unit of compute but are constrained by request latency. - Offline models improve more gradually but can scale aggressively because inference is asynchronous. - **Sequence composition** - Longer sequences generally improve performance. - Diversity of actions is more valuable than simply adding more homogeneous events. ## Practical Conclusion Meta’s approach makes sequence learning more scalable and operationally practical by moving expensive user-history processing offline while retaining fast, target-specific ranking online. Dense tokenization and target-aware attention reduce manual feature engineering, while the observed scaling laws provide a framework for deciding where additional model capacity and compute will produce the greatest gains.

Read original(opens in new tab)
aws3 min readCurated summary

Amazon DynamoDB now supports real-time vector search at any scale | Amazon Web Services

Amazon DynamoDB now offers native vector search, allowing applications to store embeddings beside operational data and query them without a separate vector database. The serverless service provides single-digit millisecond latency, 99%+ recall, horizontal scaling, and support for trillions of vectors. This removes synchronization pipelines, data movement, and additional infrastructure for applications already built on DynamoDB. ## Native Vector Search in DynamoDB - Embeddings are stored directly in DynamoDB as lists of floating-point numbers. - Similarity searches use the `SearchVectors` API and return up to 100 ranked results. - Vector indexes scale horizontally without storage limits or servers to manage. - Pricing follows DynamoDB’s pay-per-request model. - Common use cases include: - Agent memory - Retrieval-augmented generation - Recommendations - Personalized experiences - Anomaly detection ## Supported Search Capabilities - Supports vectors with up to 4,096 dimensions. - Offers three distance functions: - **Cosine**: Useful for semantic text similarity. - **Euclidean**: Useful when vector magnitude is meaningful. - **Dot product**: Useful when both direction and magnitude affect relevance. - Supports optional partition keys to distribute data and scope searches. - Supports inline exact-match filters, but not range operators such as `BETWEEN` or `BEGINS_WITH`. - Search results can include operational attributes through index projections. ## Adding Embeddings to an Existing Table - Generate embeddings with a model such as Amazon Bedrock Titan Text Embeddings, Cohere Embed, or OpenAI embeddings. - Store them in a new attribute, such as `descriptionEmbedding`, using `UpdateItem` or other AWS tooling. - No new DynamoDB data type or schema migration is required because vectors use the existing `List` and `Number` types. ## Creating and Using a Vector Index - Create a vector index on the embedding attribute. - Configure: - Index name - Vector attribute - Embedding dimensions - Distance function - Optional partition key - Filter attributes - Generate a query embedding with the same model used for stored data. - Call `SearchVectors` with the query vector, result count, partition key, and filters. - Scores depend on the distance function: - Lower scores indicate greater similarity for Cosine and Euclidean distance. - Higher scores indicate greater similarity for Dot product. ## Example: Product Catalog Search - A `ProductCatalog` table stores product details such as `productId`, `name`, `description`, `category`, `marketplace`, and `price`. - Product descriptions receive embeddings stored in `descriptionEmbedding`. - A `ProductDescriptionIndex` can use: - `marketplace` as the partition key - `category` as an inline filter - Cosine distance for semantic matching - A query such as “lightweight running shoes for summer” can return the five most relevant footwear products in the US marketplace, along with attributes such as name and price. DynamoDB vector search is best suited to applications whose operational data already resides in DynamoDB and need semantic retrieval without operating a second database or synchronization system.

Read original(opens in new tab)
cloudflare4 min readCurated summary

The Agent Access Model

BeyondCorp established that access should depend on identity and device health rather than network location. The post argues that this human-centered model is inadequate for ephemeral, fast-moving, and highly composable software agents. It proposes the Agent Access Model (AAM), which limits an agent’s capability to a specific task, evaluates every action against evolving task state, and enforces controls in the execution harness and network rather than relying on prompts. ## The Shift from Humans to Agents - Traditional Zero Trust assumes a legible human principal: - A person uses a small number of devices. - Activity occurs at human speed. - Access decisions can be evaluated over time using SSO, device posture, and risk scoring. - Agents have a different operating model: - A task-scoped run is ephemeral and ends when its work is complete. - A long-lived service may execute many independent tasks. - Agents can access databases, source control, logs, ticketing systems, documents, and other systems in rapid succession. - Least privilege must therefore become real-time and task-specific rather than a periodic policy review. ## Why Human-Oriented Controls Fail - **Durable credentials outlive ephemeral work** - Service-account keys and broad scopes may remain available after a task ends. - Credentials can persist in memory, logs, or environment variables. - Agent credentials should expire with the task and typically live only for minutes. - **Machine-speed activity bypasses slow detection** - An agent can read sensitive data and transmit it externally before human-tuned anomaly or DLP systems react. - Preventive controls must operate inline at tool-call and network boundaries. - **Prompts cannot enforce security boundaries** - Instructions such as “do not access production” can be overridden by malicious content or unsafe model behavior. - Intent may inform risk decisions, but enforcement must occur in the harness and network layer. - **Authority can disappear across delegation chains** - Agents may call tools that invoke other agents and APIs. - Existing identity and delegation mechanisms struggle to preserve the original human, task, and permissions across multiple hops. ## The Agent Access Model - AAM’s central rule is: **do not trust the task execution graph; authorize every action.** - Each action is evaluated against: - The agent’s identity. - The human or system principal it acts for. - The authorized task. - Resources already accessed by the task execution graph. - Accumulated task state can only reduce remaining capabilities; authorization for one action does not automatically authorize later actions. - AAM complements systems such as Beyond Zero by shrinking the capability set that authorization engines must evaluate and recording the agent, principal, and task behind each decision. ## AAM’s Five Principles - **Short-lived, bound credentials** - Credentials are minted for a specific task, expire with it, and are sender-constrained. - A stolen token cannot be replayed without the harness-held proof key. - **Enforcement outside the prompt** - The harness mediates tool calls. - The network mediates packets. - Prompts communicate intent but are not security boundaries. - **Exceptional human oversight** - Human approval is reserved for genuinely consequential decisions. - Requiring approval for every step causes fatigue and habitual clicking. - **Evidence-based grant review** - Captured activity reveals whether task templates are too broad or too narrow. - Approved policy changes apply only to future tasks and never expand the permissions of an active task. - **One-way capability reduction** - A declared protected event triggers the Trust Ratchet. - Capabilities are removed across the task execution graph according to policy. - Removed authority can return only through a newly authorized task. ## Reference Architecture - AAM describes a reference architecture with: - Four active controls governing the task. - An Agent Activity Log that records evidence. - A Grant Review Loop that uses that evidence to improve future grants. - The architecture is intended to define security guarantees and component responsibilities rather than prescribe a specific wire-level implementation. - At dispatch, the Agent Identity Broker issues a verifiable, short-lived credential scoped to the task. - The credential identifies: - The agent. - The principal on whose behalf it acts. - The authorized task. - It must expire no later than the task itself and be sender-constrained to prevent token-only replay. A practical implementation should treat each agent run as a bounded, independently authorized execution graph. Enforce permissions inline at the harness and network layers, use short-lived task-bound credentials, continuously reduce capability when risk changes, and use audit evidence to refine future grants without widening permissions during an active task.

Read original(opens in new tab)
cloudflare3 min readCurated summary

How we’re rethinking work at Cloudflare with Cloudflare OS

Cloudflare built Cloudflare OS to let employees use AI agents safely after a sudden increase in demand for production access and automation capabilities. The company’s approach combines AI enablement with strict controls around data access, human accountability, organizational context, and engineering quality. Its experience suggests that successful AI adoption requires meeting both technical and non-technical users where they work. ## Why Cloudflare Built Cloudflare OS - Employees rapidly began using improved AI models and agent-building tools to create internal applications. - One sales employee requested production access to roughly a dozen systems and administrative deployment permissions for an AI-built “SuperApp.” - Cloudflare needed to enable experimentation without exposing internal systems, company data, or customer data. - The resulting platform combines existing products such as Workers and Access with custom services developed for internal AI workflows. ## Principles for AI Adoption - **Start with jobs to be done:** Teams should identify customer-related pain points, bottlenecks, or missed opportunities before selecting an AI tool. - **Give everyone access to AI capabilities:** AI interfaces should not be limited to developers using terminals, code editors, and repositories. - **Keep humans accountable:** Employees remain responsible for defining quality, testing outputs, and owning the workflows and agents they deploy. - **Prioritize organizational context:** Cloudflare-specific knowledge and canonical internal guidance matter more than simply choosing the most powerful model. - **Never expand permissions through AI:** AI tools and agents must inherit users’ existing access restrictions and receive only the permissions required for their tasks. Shared agents must respect each recipient’s permissions rather than the deployer’s. ## Engineering Guardrails with the Cloudflare Engineering Codex - Cloudflare created the Engineering Codex as an authoritative, opinionated guide to engineering practices. - Unlike policies, which define what engineers cannot do, the Codex describes what they should do. - Domain owners are responsible for defining quality standards across the codebase. - AI agents use the Codex throughout the software development lifecycle: - Planning work - Reviewing merge requests - Evaluating technical designs before implementation - Reviewing incident reports - Over four months, these agents identified nearly 250,000 potential issues, blocked 16,000 merges, and caught architectural problems in almost 600 designs. - Cloudflare is now focusing on helping engineers create evaluation loops for assessing the work produced by their agents. ## Rethinking AI Tools for Non-Engineers - Cloudflare initially gave non-engineering employees developer-oriented tools with more approachable interfaces. - This approach worked poorly for knowledge workers who create one-off deliverables and interact with many systems of record. - Code-focused harnesses encouraged excessive “vibe-coded” applications, often without a clear problem to solve. - Cloudflare then began working backward from users’ actual needs and introduced the idea of a “magic AI email bot” to which employees could delegate unwanted work. The supplied excerpt ends before describing how that system worked. Cloudflare’s experience recommends pairing broad AI access with strong identity, permission, context, and accountability systems. Organizations should design tools around real jobs to be done—not simply distribute coding agents—and provide interfaces suited to both engineers and non-engineers.

Read original(opens in new tab)
cloudflare4 min readCurated summary

Cloudflare OS: an open platform for agents, apps, and work

Cloudflare OS is an open-source platform that gives every employee an agent workspace grounded in their organization’s terminology, procedures, systems, and best practices. It combines conversational agents, code execution, connected apps, workflows, and governed access to internal data. Cloudflare’s experience showed that security and resource-level authorization must be built into the platform rather than left to individual users or app developers. ## Why Organizations Need More Than Coding Agents - Code provides a clear feedback loop: it either works or fails. - Other organizational work—documents, research, processes, relationships, and physical-world outcomes—is harder for agents to support. - Agents need both: - Context about how the company operates. - Access to the systems employees use. - Cloudflare OS was created to apply agent leverage across the entire organization, not only engineering. ## Lessons from the First Version - Cloudflare’s initial system gave employees private agent workspaces. - Early limitations included: - Static apps that were not connected to live internal systems. - Repeatedly rerunning agent skills for mostly deterministic tasks, consuming additional model tokens. - Collaboration risks when users shared workspaces, apps, and outputs. - MCP servers could define which tools an agent could call, but not which underlying resources the agent had seen. - The platform therefore needed security that tracked data access and possible downstream exposure. - The new version makes security, governance, customization, and organizational context core platform features. ## Cloudflare OS Platform Components Cloudflare OS combines: - **Agent workspaces:** Browser-based environments with sessions, persistent state, files, resource access, and isolated code runtimes. - **Security and governance:** Controlled access to internal services and data. - **Personal and collaborative apps:** Modifiable applications that users can build, share, and continue evolving. - Conversations can become documents, applications, or workflows that continue operating after the initial interaction. ## Agent Workspaces for Everyone - Employees can use workspaces through a browser without being developers or using a terminal. - Company-curated skills and context prevent users from repeatedly explaining terminology, processes, and best practices to an AI model. - Shared skills allow improvements discovered by one person to benefit the wider organization. ### Research and Analysis - Agents can research using approved company context and resources. - They can write code to search, filter, join, and analyze data without loading entire datasets into the model’s context window. ### Documents, Slides, and Spreadsheets - Agents can convert research into editable documents, presentations, and spreadsheets. - Outputs can remain connected to live data, update when sources change, and be exported to services such as Google Drive. ### Connected Team Applications - When static documents are insufficient, agents can create applications with interfaces, logic, and persistent state. - These apps can use connected company resources and support collaboration among multiple users. ### Deterministic Workflows - Repetitive jobs can be implemented as workflows rather than full agent sessions. - Code handles predictable steps, while models are used only where judgment is needed. - Workflows can run manually, on schedules, or in response to events. - Access to systems of record is provided through Gatekeepers, while existing MCP servers can be connected through MCP Server Portals. ## Security and Governance - Directly distributing API keys to employees or agents creates broad, long-lived access that is difficult to constrain and audit. - MCP improves credential handling by keeping keys in servers and exposing defined tools. - Tool-level control is not sufficient: agents may combine data from multiple systems, move it to less restricted locations, or expose it through apps and generated outputs. - Authorization must therefore consider not only which tools an agent can use, but also which resources it has observed and where that information can go. ### Default-Deny Access - Cloudflare Access controls entry into Cloudflare OS. - Within the platform, every agent and app begins with no permissions. - An agent must request access to a specific resource, which can be approved or denied. - Approved resources are exposed to generated code through typed bindings such as `env.PROJECT`. - These bindings represent narrowly scoped capabilities under a specific policy. - Credentials remain isolated from both the agent and the generated code. Cloudflare OS is intended as a customizable organizational platform: companies can deploy it, connect internal systems, encode their operating knowledge as skills, and give employees governed tools for building useful apps and workflows. Its default-deny, resource-aware security model is essential for safely sharing agent-generated work across an organization.

Read original(opens in new tab)
cloudflare4 min readCurated summary

Catching rogue AI behavior with identity-aware analytics

AI usage is difficult to govern without knowing both who made each request and what normal usage looks like for that person or agent. Cloudflare’s new Identity-aware AI Gateway and User Insights address this by attaching verified identities to requests and detecting behavior that significantly deviates from historical patterns. Together, they provide centralized visibility, per-user cost controls, and anomaly detection without requiring additional setup for traffic already routed through AI Gateway. ## AI Gateway as a Central Control Plane - AI Gateway routes requests from applications, developer tools, and agent harnesses—including Claude Code, Codex, and GitHub Copilot—through one platform. - It provides centralized observability, security, governance, and spend management across providers such as OpenAI, Anthropic, Google, and Workers AI. - This centralization makes it possible to analyze usage consistently across both human users and automated agents. ## Identity-Aware Requests with Cloudflare Access - The Cloudflare Access integration places a custom domain, such as `ai.example.com`, in front of the gateway. - Organizations can: - Authenticate users through SAML-compatible providers such as Okta or Microsoft Entra. - Apply access policies to specific users. - Avoid distributing Cloudflare API keys. - Each authenticated request includes the Access user ID as `cf.user_id`. - Administrators can filter logs, analytics, and spending by the actual requester rather than by a shared API key. - Per-user spend limits can assign each person a separate budget and either block requests or route them to cheaper models after the limit is reached. - Planned improvements will use identity-provider groups to control model access and spending—for example, granting frontier-model access to machine learning teams while limiting support teams. ## User Insights and Behavioral Baselines - User Insights is available to all AI Gateway customers at no extra cost. - It analyzes existing gateway traffic without requiring additional configuration. - The feature builds behavioral profiles for every account, including both people and agents. - It tracks cost inefficiencies such as poor cache-hit rates and oversized context windows, but focuses primarily on whether usage is normal for that particular account. - Human users and automated agents are evaluated according to their own patterns: - Agents may have regular, predictable sessions. - Humans typically have more irregular prompts, timing, and session lengths. ## Session-Based Anomaly Detection - User Insights evaluates sessions rather than individual requests, reducing noise from isolated events. - Each session is compared with the account’s rolling 95th-percentile session cost over the previous 30 days. - A session becomes a strong anomaly candidate when it exceeds twice that personal p95 baseline. - This relative comparison avoids misleading fixed thresholds: - A $500 session may be normal for a consistently heavy user. - A $50 session may be highly unusual for an agent that normally spends $5. - Baselines adjust over time as an account’s usage changes. ## Combining Personal and Organization-Wide Thresholds - User Insights also applies an organization-wide p99 cost ceiling. - In the example analysis: - Most sessions cost less than $10. - The organizational p95 is $20. - The p99 is $200, meaning only 1% of sessions reach that amount. - Alerts are triggered only when a session is both: - More than twice the account’s personal p95. - Above the organization’s p99 ceiling. - This prevents alerts for: - Small-dollar spikes that are statistically unusual but not worth investigating. - Expensive sessions that are routine for a particular user. - A dollar floor also prevents tiny accounts from triggering alerts because of insignificant percentage increases. ## Filtering for Rogue Behavior - The resulting interface presents a feed of accounts that have broken their established usage patterns. - This focuses administrators on potentially meaningful incidents instead of showing every unusual request. - The approach is designed to detect trusted users or agents that suddenly perform more of an already-authorized activity—behavior that traditional controls may not block because no new tool or forbidden action is involved. Cloudflare’s recommendation is to route AI traffic through AI Gateway, authenticate it with Cloudflare Access, and use identity-based budgets alongside behavioral baselines. This combination helps organizations connect spending and activity to specific people or agents while concentrating investigations on statistically significant, high-impact deviations.

Read original(opens in new tab)
cloudflare3 min readCurated summary

WriteGuard: Fine-grained controls for MCP Servers

Cloudflare built WriteGuard to safely expand AI agents’ write access to internal MCP servers. The system centralizes authorization, risk classification, agent attribution, and auditing, addressing failures that client-side prompts or individual user vigilance cannot reliably prevent. It preserves the human user’s permissions while making each agent session identifiable and its actions queryable. ## The Risk of Uncontrolled Agent Actions - A broadly instructed cleanup agent accidentally closed thousands of tickets. - Human and agent actions were recorded under the same employee identity, making the incident difficult to investigate and repair. - Network logs could not distinguish between multiple agent sessions. - More serious failures could involve: - Amending contracts - Sending mass customer replies - Deleting database tables - Triggering destructive production actions ## MCP Fundamentals - The Model Context Protocol connects AI applications to external tools and data. - An MCP server exposes tools with: - A name - A description - An input schema - A handler that performs the operation - When an agent selects a tool, the MCP client sends the call to the server, which interacts with the downstream application. ## Cloudflare’s MCP Expansion - Cloudflare uses MCP with local clients such as OpenCode and Cloudflare OS, as well as long-running agent services. - Its internal MCP portal grew from 13 servers to 27. - Servers initially provided read-only access to systems such as Jira, GitLab, internal documentation, and operational tools. - As agents became more capable, teams requested write actions across engineering, product, design, sales, and customer success. - Cloudflare decided centralized controls were necessary because client-side skills and elicitation prompts vary across agent harnesses and can be disabled. ## WriteGuard’s Policy and Attribution Layer - WriteGuard evaluates tool configuration together with request context. - It can: - Pass a call through unchanged - Add agent attribution to supported writes - Create a scrubbed audit event - Block a call before the tool handler executes - Policies are defined per tool and include: - Risk tier - Enabled or disabled status - Labeling configuration - Risk tiers include: - **Read Only:** Search issues or inspect merge requests - **Minimal Impact:** Add reactions or mark notifications read - **Contained Write:** Add comments, create merge requests, or update issue fields - **Critical:** Merge code, deploy to production, or bulk-delete records - Labeling allows agent context to be inserted into downstream applications in formats such as plain text or HTML without modifying the MCP server. ## Preserving Human Permissions While Identifying Agents - Agents operate through the employee’s Cloudflare Access and OAuth identity. - An agent cannot perform an action its user is not authorized to perform. - Cloudflare avoided standalone agent accounts because they would create additional permissions to manage and weaken accountability. - WriteGuard supplements the human identity with MCP client and session information. - Each write can therefore be tied to both the responsible person and the specific agent session. ## Centralized, Queryable Auditing - WriteGuard classifies every invocation as successful, failed, or blocked. - It asynchronously sends scrubbed events to an internal audit Worker. - Audit records include: - MCP server and tool - Risk tier - Outcome - User and client - Request duration - Secret and sensitive input values are omitted. - Asynchronous logging avoids adding latency to the agent’s response. - MCP portal logs show raw tool invocations, while WriteGuard adds semantic classifications, agent context, and backing-service outcomes. - Central auditing makes unusually fast or widespread agent activity easier to detect and investigate. ## Recommendation Organizations expanding MCP agents beyond read-only access should use centralized, server-side policy enforcement, preserve human authorization boundaries, attach per-session agent attribution, and maintain scrubbed audit logs. Relying solely on prompts, client configuration, or undifferentiated user identities makes destructive automation difficult to prevent and even harder to understand afterward.

Read original(opens in new tab)
figma3 min readCurated summary

Better Code, Fewer Tokens: The Benefits of Code Connect in MCP | Figma Blog

Code Connect improves how coding agents translate Figma designs into production code by supplying real design-system components, imports, and prop values. Figma’s evaluations found that Code Connect reduced median task duration by 19.6%, lowered token usage by 29.5%, and increased code quality by one point on a 1–4 scale. The central conclusion is that better design-to-code context helps agents work faster while producing code that fits existing codebases. ## The Problem: Visually Correct but Technically Wrong Code - Without production context, agents often: - Rebuild interfaces from basic primitives. - Invent components that already exist. - Choose the wrong design-system component. - Spend extra tokens searching, debugging, and rewriting. - Figma’s MCP server normally provides a React representation of the design through `get_design_context`. - This output may match the visual design but does not explain how the design maps to a company’s actual component library. ## How Code Connect Enriches MCP Responses - Code Connect links Figma components to their real implementations in a codebase. - With Code Connect templates configured, MCP responses replace generic React markup with production-relevant snippets. - Agents receive: - Correct component imports. - Accurate component names. - Appropriate property values. - Code that reflects the company’s design system. - For example, generic markup for a tab control can be replaced with an existing component such as: ```tsx <SegmentedControl value="design" options={["Design", "Code"]} /> ``` ## Coinbase Case Study - Coinbase’s Design Systems team adopted Code Connect as engineers increasingly used coding agents. - Without Code Connect, agents sometimes fabricated alternatives, such as constructing a stepper from progress bars. - With Code Connect, agents received literal imports and accurate code representations for Coinbase Design System components. - Coinbase reported improved output quality and reduced token usage. ## Evaluation Results - Figma created an evaluation harness that ran identical design-to-code tasks: - Once without Code Connect. - Once with Code Connect templates. - The evaluation covered 27 test cases. - It measured: - Code quality. - Token consumption. - Task duration. - The tests used two React-based design systems: - Simple Design System (SDS), Figma’s example system. - Figma Pattern Library (FPL), a larger internal system. - Across the tests, Code Connect produced: - **19.6% lower median task duration** - **29.5% lower median token usage** - **A one-point increase in code quality on a 1–4 Likert scale** Teams using coding agents for design-to-code work should connect their Figma components to production implementations through Code Connect. Providing exact component context reduces guesswork and tokens while helping agents produce maintainable, design-system-compliant code.

Read original(opens in new tab)
github3 min readCurated summary

How the GitHub legal team used Copilot CLI to streamline their workflows

GitHub’s legal team used Copilot CLI to turn repetitive legal work into customizable internal tools without relying on traditional software engineering. By expressing workflows, standards, and policies in plain language and Markdown, lawyers built systems that improved consistency, reduced drafting time, and preserved human oversight. The post argues that domain expertise can be operationalized into useful AI tools by anyone who can clearly define a process. ## Building a Contract Drafting Style Guide - Principal Product Counsel Ngandu Kasuku created **terms-ai** to manage varied commercial agreements involving data, infrastructure, and product integrations. - The tool stores instructions, drafting resources, workflows, and reference documents in a version-controlled repository. - An internal style guide enforces plain-language drafting and replaces repetitive prompt copying with consistent guidance. - A library of approved agreements lets the tool draw on prior work for addenda and new contracts. - Sensitive agreements remain in a controlled internal environment rather than the open-source repository. - Kasuku reports cutting drafting and review time roughly in half while producing more consistent provisions. - The main insight was that AI could support a lawyer’s own judgment and working style, not merely perform isolated tasks. ## Turning Legal Workflows into Plain-Language Instructions - Online Safety Counsel Jesse Geraci began with a workflow for analyzing source code in **DMCA** notices. - Copilot instructions covered triage, code comparison, license checks, circumvention review, policy references, and report templates. - Instead of traditional programming, the workflow encoded legal reasoning through structured instruction files. - Different modes were created for clients and lawyers, including faster client analysis, escalation recommendations, deeper legal review, and arguments for both sides. - The system later grew into a desktop application supporting contract review, NDA triage, risk assessment, compliance checks, and response drafting. - Reusable skills and agents handle tasks such as intake, playbook alignment, risk scoring, evidence verification, escalation, and report assembly. - Legal teams can still customize the system through readable Markdown, while human review remains essential. ## Broader Lessons for Nontechnical Teams - Repetitive work in almost any profession can be a starting point for automation. - Clear definitions of methodology, standards, and desired outputs can substitute for extensive programming knowledge. - Teams should begin with one bottleneck, use Copilot CLI to prototype a solution, and expand based on real usage. - These tools are decision-support systems—not replacements for professional judgment. Teams can use Copilot CLI to turn their existing expertise into repeatable, transparent workflows while retaining control over sensitive data and final decisions.

Read original(opens in new tab)