Techlist.io - Korean Tech Blog Curator

gitlab3 min readCurated summary

Automating detection gap analysis with GitLab Duo Agent Platform

GitLab’s Signals Engineering team uses GitLab Duo Agent Platform to automate detection gap analysis after security incidents. The approach replaces inconsistent manual reviews with AI agents that examine incident issues, map attacker behavior to MITRE ATT&CK, and recommend actionable detection improvements. GitLab recommends starting with the built-in Security Analyst Agent, then creating a custom agent when organization-specific context is required. ## The Detection Gap Problem - A detection gap occurs when an attacker performs an action that existing detections fail to identify. - Reviewing gaps requires analysts to: - Read incident timelines, comments, and related artifacts. - Map attacker actions to detection opportunities. - Identify missing or insufficient alerts. - Recommend concrete detection improvements. - Manual analysis is time-consuming, inconsistent across reviewers, and easy to postpone. - GitLab embeds this process in the workflow where incidents already reside: GitLab issues. ## GitLab Duo Agent Platform - Duo Agent Platform supports agents that can reason, take actions, and interact with GitLab resources such as issues, merge requests, and code. - Teams can either: - Use pre-built agents with existing domain knowledge. - Build custom agents using a name, description, and system prompt. - The system prompt defines the agent’s role, knowledge, tools, and expected behavior. ## Security Analyst Agent - The built-in Security Analyst Agent can be invoked directly from a closed incident issue. - It reviews: - Incident descriptions and timelines. - Tasks and comments. - Linked artifacts and other issue content. - It can identify missed attacker tactics, techniques, and procedures and map them to MITRE ATT&CK. - It is useful for quick, low-configuration assessments, particularly when incident documentation is thorough. - Its limitation is a lack of knowledge about an organization’s specific SIEM, log sources, detection stack, and engineering standards. ## Detection Engineering Assistant - GitLab created a custom agent to provide recommendations tailored to its environment. - Building the agent requires only: - A name. - A description. - A system prompt. - The system prompt is central to the agent’s usefulness; detailed instructions produce more consistent and relevant results. ### Defining the Agent’s Role - The prompt explicitly identifies the agent as a detection engineering assistant responsible for analyzing incidents and finding coverage gaps. - Clear framing helps anchor the agent’s responses to the team’s actual responsibilities. ### Encoding Detection Principles - GitLab describes its preferred detection characteristics: - Low false-positive rates. - High signal fidelity. - Actionable alerts with useful response context. - The prompt favors behavioral detections over indicator-of-compromise approaches when practical. - It also addresses the tradeoff between broad coverage and alert fatigue. ### Providing Environment and Telemetry Context - The agent is told which log sources are available, what SIEM is used, and what telemetry is missing. - This prevents it from recommending detections that depend on data the team cannot access. ### Structuring Findings with MITRE ATT&CK - Gap findings are organized around ATT&CK tactics and techniques. - This provides consistent reporting and supports internal coverage tracking and prioritization. ### Standardizing Output - Each finding should include: - The relevant ATT&CK technique. - What attacker behavior was missed. - The log source or data needed for detection. - A recommended detection approach. - Consistent formatting makes findings easier to triage and convert into engineering work. - GitLab’s full system prompt contains 1,870 words and 337 lines, illustrating the level of detail used to tailor the agent. ## Practical Recommendation Use the Security Analyst Agent for an immediate first pass, but build a custom detection engineering agent when recommendations need to reflect your own telemetry, tooling, standards, and detection philosophy. A detailed system prompt is the key to turning general AI analysis into repeatable, actionable security engineering work.

Read original(opens in new tab)
figma2 min readCurated summary

5 Design Skills To Sharpen in the AI Era | Figma Blog

AI is changing product creation by accelerating experimentation and expanding who can participate in design. Figma argues that designers should strengthen adaptable, technology-oriented skills rather than rely only on traditional craft. The first priority is becoming fluent with AI tools and learning to prompt them effectively, while maintaining human judgment and design fundamentals. ## AI Fluency and Prompting - AI skills are becoming essential for designers and increasingly important in non-design roles such as product management, development, and marketing. - More than half of designers and hiring managers consider AI design capabilities—such as rapid prototyping and “vibe coding”—important hiring skills. - Among designers who adopted AI during the past year: - 91% say it helps them create better designs. - 89% say it helps them work faster. - AI can support many activities, including: - Editing images directly within a workflow. - Building prototypes instead of writing traditional product requirements documents. - Testing assumptions and creating tangible artifacts for team alignment. ## Writing Better Prompts - Clear, structured prompts produce more reliable AI-generated results. - Figma recommends organizing prompts around: - The task - Context - Required elements - Behavior - Constraints - Prompting is presented as a repeatable design practice, not merely a way to get a one-off output. - Strong prompts help turn AI into a consistent design partner rather than an unpredictable experimentation tool. ## Broader Changes to Design Work - AI is lowering barriers to participation and blurring boundaries between product roles. - Designers are increasingly expected to work across disciplines and use AI to extend their capabilities. - Prototyping is becoming a faster way to communicate ideas, validate assumptions, and build momentum than relying solely on written documentation. Designers should build practical fluency with AI tools, practice structured prompting, and use prototypes to make ideas concrete—while applying their own judgment to guide and evaluate the results.

Read original(opens in new tab)
figma2 min readCurated summary

Vishal Kapoor’s 10 Rules for Building Honest Products with AI | Figma Blog

AI product development is ultimately a trust challenge, not merely a technical one. Vishal Kapoor argues that AI should accelerate exploration and execution without replacing human judgment, empathy, or accountability. His approach centers on building products that remain transparent, secure, emotionally aware, and honest—especially in sensitive areas such as personal finance. ## Start with First-Principles Thinking - Break complex problems into their fundamental components before reaching for an AI solution. - AI can accelerate ideation and iteration, but it cannot replace human intuition, taste, or a distinctive product perspective. - Question basic assumptions to uncover better alternatives. For example, Affirm challenges why customers receive three payment-plan options rather than one, five, or a customizable plan. - Thoughtful disagreement among people remains essential for generating meaningful insights; AI is best used to explore possibilities more quickly. ## Stay Close to Human Emotions - Product teams should regularly observe customers, conduct UX research, read app-store reviews, monitor social media, and speak directly with users. - Metrics and dashboards identify patterns, but they do not fully explain the emotions behind customer behavior. - Financial products especially require sensitivity to anxiety, frustration, trust, and relief—not just transactional outcomes. - Affirm uses an internal AI tool called Pluto to investigate recent customer disappointments, while still relying on human observation and empathy to interpret those experiences. ## Treat AI as a Teammate - AI is neither a guaranteed productivity multiplier nor an inevitable replacement for employees; it is another participant in a collaborative product-development process. - Tools such as Figma Make help teams convert customer insights into prototypes and test ideas faster. - AI can audit large numbers of screens and interaction patterns across web, mobile, and desktop experiences, identifying outdated or inconsistent designs. - Moving repetitive auditing and prototyping work from engineers to designers and product managers increases iteration speed and creates more room for creativity. ## Test the Edge Cases - Trustworthy products cannot be designed only around the happy path. - Teams should deliberately explore unusual inputs, failure modes, and unexpected customer situations rather than assuming normal usage. - The article begins this rule by emphasizing that authentic product quality depends on examining the difficult and overlooked scenarios where users are most likely to encounter confusion or harm. The overall recommendation is to use AI aggressively for exploration, prototyping, and repetitive analysis—but keep humans responsible for defining the problem, understanding customers, challenging assumptions, and ensuring the final product is honest.

Read original(opens in new tab)
aws3 min readCurated summary

AWS Weekly Roundup: Amazon Connect Health, Bedrock AgentCore Policy, GameDay Europe, and more (March 9, 2026) | Amazon Web Services

The March 9, 2026 AWS Weekly Roundup highlights AWS’s growing focus on agentic AI, healthcare automation, security, and developer productivity. Major updates include Amazon Connect Health, centralized policies for Bedrock agents, private AI assistants on Lightsail, and new tools for troubleshooting and durable Lambda workflows. The roundup also previews community events, including GameDay Europe, NVIDIA GTC, AWS Summits, and regional Community Days. ## Major AWS Product Launches - **Amazon Connect Health** is generally available with five healthcare-focused AI agents: - Patient verification - Appointment management - Patient insights - Ambient documentation - Medical coding - These capabilities are HIPAA-eligible and designed to integrate with existing clinical workflows within days. - **Bedrock AgentCore Policy** provides centralized, fine-grained controls for agent-to-tool interactions. - Policies can be written in natural language. - AWS converts them into Cedar, its open-source policy language. - Controls operate outside application code, supporting security and compliance teams. - **OpenClaw on Amazon Lightsail** enables deployment of private autonomous AI assistants. - Includes sandboxed sessions, security controls, HTTPS, and device-pairing authentication. - Uses Amazon Bedrock by default and supports Slack, Telegram, WhatsApp, and Discord integrations. ## Pricing, Cost Management, and Security - **VPC Encryption Controls** became a paid feature on March 1, 2026. - Monitor mode detects unencrypted traffic. - Enforce mode blocks traffic that does not meet encryption requirements. - Controls apply to traffic within and across VPCs in a region. - **Database Savings Plans** now cover Amazon OpenSearch Service and Amazon Neptune Analytics. - Customers can save up to 35% with a one-year commitment. - Savings apply across engine, instance family, size, and AWS Region. - **Amazon GameLift Servers DDoS Protection** adds a co-located relay network. - Client traffic is authenticated with access tokens. - Per-player traffic limits help mitigate attacks. - The feature adds no cost for GameLift Servers customers. ## Developer and Operations Improvements - **Elastic Beanstalk AI-powered environment analysis** sends events, health data, and logs to Amazon Bedrock when environments degrade. - It returns troubleshooting recommendations tailored to the affected environment. - AWS now allows **IAM roles to be created directly inside service workflows**, reducing the need to switch to the IAM console. Supported services include EC2, Lambda, EKS, ECS, Glue, and CloudFormation. - **Kiro’s new Lambda durable functions power** assists developers with long-running, multi-step applications and AI workflows. - It provides guidance on replay models, waits, concurrency, error handling, and deployment. ## AWS Community Projects - One community project demonstrates a persistent AI memory layer using **MCP, Amazon Bedrock, and a Chrome extension**, allowing agents to retain context across sessions and applications. - Another experimental application treats the AI model as the runtime, generating a complete interactive web application from a single prompt without a conventional codebase, framework, or persistent state. ## Community Events and AWS Activities - **AWS Community GameDay Europe** takes place March 17, offering team-based challenges using real AWS services. - AWS will participate in **NVIDIA GTC 2026** in San Jose from March 16–19, with sessions, demos, booths, and discounted passes. - Upcoming **AWS Summits** include Paris, London, and Bengaluru. - Upcoming **AWS Community Days** include events in Slovakia, Pune, and Mexico City. AWS’s latest announcements show a clear emphasis on practical AI agents, stronger governance, and automation across infrastructure and application development. Developers and cloud teams should review the new security and pricing changes while exploring the AI tools and upcoming hands-on community events.

Read original(opens in new tab)
meta3 min readCurated summary

How Advanced Browsing Protection Works in Messenger

Advanced Browsing Protection (ABP) extends Messenger’s Safe Browsing beyond on-device detection by checking links against a frequently updated database of millions of potentially malicious websites. Its central challenge is balancing effective URL matching with privacy: Messenger must identify unsafe links without revealing users’ exact queries or distributing the entire blocklist. ABP combines private information retrieval, cryptographic techniques, sharding, and client-side preprocessing to achieve this balance. ## Safe Browsing Within End-to-End Encryption - Messenger’s end-to-end encryption protects messages and calls, but it does not by itself protect users from malicious links. - Safe Browsing warns users when a link may lead to phishing, credential theft, or other harmful activity. - The standard feature uses on-device models. - Advanced Browsing Protection adds access to a continually updated watchlist containing millions of potentially malicious websites. ## Private Information Retrieval as the Foundation - Private information retrieval (PIR) allows a client to ask whether an item exists in a server-held database while revealing as little as possible about the query. - Sending the full database to each device is impractical because: - The database is large and frequently updated. - Exposing the complete list could help attackers evade detection. - Existing PIR approaches use oblivious pseudorandom functions (OPRFs) and divide the database into buckets or shards. - ABP had to address two limitations: - OPRFs are designed for exact matches, whereas URLs require prefix matching. - The client generally must identify which bucket to query, creating a privacy-versus-efficiency tradeoff. - More advanced lattice-based constructions may reduce the need for sharding, but they were not yet practical at ABP’s scale. ## Privacy-Preserving Prefix Matching for URLs - A database entry such as `example.com` should match a longer URL such as `example.com/a/b/index.html`. - Querying every prefix separately would work functionally: - `example.com` - `example.com/a` - `example.com/a/b` - `example.com/a/b/index.html` - However, each query can leak information about the original URL. If one query leaks `B` bits and there are `P` prefixes, the total leakage may reach `P × B` bits. - ABP instead groups URLs by domain so the client makes one bucket request and checks path prefixes within that bucket. - This reduces query leakage but creates uneven bucket sizes. - Domains such as link-shortening services may contain huge numbers of URLs, producing oversized buckets and potentially large padded responses. ## Preprocessing Rulesets to Balance Buckets - The server addresses bucket imbalance by generating a ruleset that tells clients how to process URLs before selecting a bucket. - Each rule maps an 8-byte hash prefix to a number of path segments that should be appended to the current URL before hashing again. - For example: - The client hashes `example.com`. - If the hash matches a ruleset entry, it appends specified path segments, such as `/a/b`. - It hashes the resulting URL again and repeats the process. - When no ruleset entry matches, the client uses the first two bytes of the final hash as the bucket identifier. - The server builds the ruleset iteratively: - It initially hashes URLs by domain. - It identifies the largest bucket. - It finds the most common domain in that bucket. - It adds rules that incorporate additional URL path segments to split the oversized bucket. - Clients receive the ruleset in advance and perform the same deterministic processing during lookups. ABP’s design demonstrates how privacy-preserving lookup can support real-world URL semantics without exposing users’ links. The combination of PIR, controlled sharding, prefix-aware processing, and adaptive rulesets allows Messenger to warn about malicious sites while limiting what the server learns about each user’s browsing query.

Read original(opens in new tab)
github3 min readCurated summary

Under the hood: Security architecture of GitHub Agentic Workflows

GitHub Agentic Workflows are designed to bring autonomous agents into CI/CD without giving them unrestricted access to repositories, secrets, or the internet. Because agents can be prompt-injected and behave unpredictably, GitHub treats them as untrusted components and compiles workflows into constrained GitHub Actions. The architecture relies on layered isolation, controlled communication, staged writes, and comprehensive auditing. ## Threat Model - Agents reason over repository state and act autonomously, so they cannot be trusted by default. - GitHub Actions normally place components in one permissive trust domain with broad access to: - Repository contents - Authentication secrets - MCP servers - Arbitrary network destinations - A malicious webpage, issue, or repository file could prompt an agent to: - Read credentials from files, environment variables, logs, or `/proc` - Upload secrets externally - Embed secrets in issues, pull requests, or comments - Make unwanted repository changes - Strict mode follows four principles: - Defense in depth - Never trust agents with secrets - Stage and vet writes - Log everything ## Layered Security Architecture GitHub Agentic Workflows use three complementary layers: - **Substrate layer** - Runs on a GitHub Actions runner VM. - Uses trusted containers, Docker isolation, network controls, and kernel-enforced boundaries. - Separates components and mediates privileged operations and system calls. - Is intended to contain damage even if an untrusted component is compromised. - **Configuration layer** - Defines which components run and how they connect. - Controls communication channels, privileges, firewall policies, Docker images, and MCP configuration. - Determines which tokens are loaded into which containers. - Converts declarative workflow configuration into a secure runtime structure. - **Planning layer** - Controls which components are active and how data moves between them over time. - Creates staged workflows with explicit data exchanges. - Uses the Safe Outputs subsystem to govern potentially dangerous operations. ## Keeping Secrets Away from Agents - In ordinary GitHub Actions, secrets may be visible through environment variables and configuration files across the shared runner trust domain. - This creates a major prompt-injection risk: an agent with shell access could discover credentials and exfiltrate them. - Agentic Workflows instead place the agent in a dedicated container with: - Firewalled internet access - MCP access through a trusted gateway - LLM communication through an API proxy - A private network connects the agent only to approved services. - The trusted MCP gateway launches MCP servers and exclusively handles MCP authentication material. - LLM authentication tokens are kept in the isolated API proxy rather than exposed directly inside the agent container. ## Controlled Execution and Writes - Open-ended workflow authoring is separated from governed execution. - Workflows are compiled into GitHub Actions with explicit constraints covering: - Permissions - Outputs - Network access - Auditability - The planning and Safe Outputs systems are intended to mediate GitHub write operations and apply controls such as call filtering, volume limits, secret removal, and moderation. GitHub’s approach is to treat agents as untrusted CI/CD components rather than granting them normal workflow privileges. Organizations adopting agentic automation should isolate agents, broker access to tools and credentials, restrict network connectivity, stage all writes for review, and maintain detailed logs.

Read original(opens in new tab)
cloudflare3 min readCurated summary

Fixing request smuggling vulnerabilities in Pingora OSS deployments

Pingora 0.8.0 fixes three HTTP/1.x request-smuggling vulnerabilities affecting standalone deployments used as Internet-facing ingress proxies. The flaws could let attackers bypass proxy security controls, desynchronize connections with backends, hijack other users’ sessions, or poison shared caches. Cloudflare’s own CDN was not affected, but Pingora users are urged to upgrade immediately. ## Scope and Impact - Vulnerabilities: - CVE-2026-2833 - CVE-2026-2835 - CVE-2026-2836 - Reported through Cloudflare’s bug bounty program in December 2025. - Affected deployments are standalone Pingora proxies exposed directly to the Internet. - Potential consequences included: - Bypassing ACL and WAF checks at the proxy layer. - Desynchronizing Pingora and backend HTTP connections. - Cross-user session or credential theft. - Cache poisoning when shared backends are used. - Cloudflare’s CDN was not vulnerable because Pingora is not used as its ingress proxy, and internal clients did not send pipelineable HTTP/1 requests to affected services. ## Premature Upgrade Without a `101` Response - Pingora treated a request containing an `Upgrade` header as an upgraded, pass-through connection immediately. - Under RFC 9110, the connection should switch protocols only after the backend returns `101 Switching Protocols`. - If the backend instead returned `200 OK`, Pingora could still forward subsequent bytes directly to the backend. - An attacker could pipeline a second, partial request—such as `/admin`—after the initial upgrade request. - This bypassed Pingora’s normal ACL or WAF processing and left Pingora and the backend disagreeing about request boundaries. - A later request from another user could complete the attacker’s partial request, causing the backend to return the attacker’s response to the wrong user. - Pingora 0.8.0 now enables pass-through mode only after receiving a valid `101` response. ## HTTP/1.0, Close-Delimited Bodies, and Transfer-Encoding - Another attack resembled a classic CL.TE desynchronization: - Pingora used `Content-Length` to determine the request body length. - The backend interpreted `Transfer-Encoding: chunked` and ended the body at the zero-length chunk. - The example combined: - HTTP/1.0 - `Connection: keep-alive` - Multiple transfer encodings - Both `Transfer-Encoding` and `Content-Length` - Pingora’s earlier transfer-encoding detection was too simplistic: - It only checked whether `Transfer-Encoding` contained “chunked.” - It assumed a single encoding or header. - HTTP specifications require the final transfer encoding to determine whether chunked framing applies, creating disagreement between Pingora and backend servers such as Node.js. - These differing interpretations of body boundaries enabled attackers to smuggle a second request through the proxy. ## Hardening and Recommendation Pingora 0.8.0 corrects the HTTP/1 framing and upgrade handling issues and adds defensive hardening. Operators running Pingora as an ingress proxy should upgrade as soon as possible and review whether their deployments expose HTTP/1 connections directly to untrusted clients.

Read original(opens in new tab)
cloudflare3 min readCurated summary

Active defense: introducing a stateful vulnerability scanner for APIs

Cloudflare is launching a beta Web and API Vulnerability Scanner to actively detect API logic flaws that defensive tools often miss. Its first target is Broken Object Level Authorization (BOLA), where authenticated users can access or modify another user’s resources through valid requests. The scanner combines Cloudflare’s existing API visibility with stateful DAST to test authorization across sequences of API calls. ## Why Defensive Security Misses API Logic Flaws - Traditional WAFs are effective against recognizable attacks such as SQL injection, XSS, malformed requests, and suspicious payloads. - API vulnerabilities often involve valid requests that violate business rules rather than protocol or schema requirements. - In the food delivery example: - User A sends a valid `PATCH` request for User B’s order. - User A’s token, headers, and request schema are all legitimate. - The vulnerability exists because the API fails to verify that User A owns the order. - A basic authorization check could prevent the issue: ```js if (order.userID != user.ID) throw Unauthorized; ``` ## Passive Detection and the Limits of Traditional DAST - Cloudflare’s existing API Shield BOLA detection passively analyzes customer traffic for abnormal usage patterns. - Effective passive detection requires context about: - Valid API calls - Variable parameters - Normal user behavior - API responses when parameters are manipulated - Passive analysis may be insufficient in development environments with little traffic or production systems without observed attacks. - DAST creates new traffic specifically for security testing and can operate in environments without relevant user activity. - Traditional DAST tools often: - Require extensive configuration - Depend on manually maintained Swagger/OpenAPI files - Struggle with modern authentication flows - Lack API-specific tests such as BOLA detection ## Cloudflare’s API Scanning Advantage - Scan results will appear in Security Insights alongside other Cloudflare security findings. - API Shield customers already benefit from Cloudflare’s API Discovery and Schema Learning, which catalog endpoints and learn traffic patterns. - The initial release requires an uploaded OpenAPI specification, though future versions are expected to work without one. - Cloudflare can use passive traffic inspection to identify possible BOLA issues and actively verify them with new HTTP requests. - Customers provide API credentials, while Cloudflare uses API schemas to construct a scan plan automatically. ## Stateful API Scanning - Conventional scanners often evaluate requests independently, making it difficult to test vulnerabilities that require a sequence of related actions. - BOLA testing may require: - Creating a resource as one user - Attempting to access or modify it as another user - Comparing the resulting behavior - Cloudflare’s scanner builds an API call graph from the OpenAPI document. - It walks that graph using separate owner and attacker contexts: - Owners create resources. - Attackers use their own valid credentials to attempt access. - This stateful approach is designed to test authorization across realistic API workflows rather than isolated requests. Cloudflare’s scanner is intended to complement—not replace—passive monitoring and edge defenses. Organizations should use the beta to actively test APIs, especially authorization controls, in environments where normal traffic provides insufficient security context.

Read original(opens in new tab)
cloudflare3 min readCurated summary

Complexity is a choice. SASE migrations shouldn’t take years.

Cloudflare argues that SASE and zero trust migrations do not need to take years. Its partners, TachTech and Adapture, reportedly reduced deployments from around 18 months to four–six weeks by using Cloudflare One’s unified, cloud-native architecture. The post concludes that programmable security infrastructure can accelerate zero trust adoption while also enabling safer use of AI. ## Faster Zero Trust Deployments - Traditional Secure Web Gateway (SWG) and Zero Trust Network Access (ZTNA) migrations can take up to 18 months for large organizations. - TachTech reduced comparable Cloudflare One deployments to four–six weeks. - Cloudflare Access is presented as lightweight and largely “no-touch” after deployment, reducing ongoing operational effort. ## Why Legacy Migrations Stall - Legacy architectures often treat migration as hardware replacement rather than software transformation. - Complex service chaining creates a “trombone effect,” increasing latency and making troubleshooting difficult. - Cloudflare’s partners accelerate migrations through: - **Identity-first on-ramps:** Existing identity-provider groups define access policies instead of rebuilding network segments. - **Consolidated policy engines:** SWG and ZTNA policies are handled together, avoiding synchronization between separate products. - **Cloud-native connectors:** Tools such as `cloudflared` provide connectivity without opening inbound firewall ports. ## Scaling Quickly - Adapture expanded one Cloudflare Access deployment from 600 contractors to 5,000 users. - The company describes the expansion as seamless compared with the lengthy implementation cycles associated with legacy SASE platforms. - Cloudflare positions rapid elasticity as important for organizations whose workforce and security needs change quickly. ## A Programmable, Extensible Edge - Cloudflare One is described as software-defined and composable, allowing partners to adapt it to specialized environments. - TachTech supported Arch Linux developer workstations by extracting binaries from an Ubuntu `.deb` package and creating a custom `PKGBUILD`. - This approach preserved device-posture checks, including disk-encryption and firewall-status verification, without creating a security exception. ## Supporting Safe AI Adoption - Cloudflare says the Secure Web Gateway is evolving from simple URL filtering toward controlling data flows to large language models. - Its AI security capabilities include: - **Shadow AI visibility:** Identifying unauthorized AI tools in use across the organization. - **AI confidence scores:** Evaluating models based on standards such as SOC 2 and ISO 42001, as well as data-handling practices. - **DLP prompt protection:** Blocking sensitive source code, personally identifiable information, and financial data from being submitted to public AI services. - **LLM discovery:** Finding and labeling internet-exposed LLM endpoints to reveal the organization’s AI attack surface. - **Request validation:** Intended to defend AI applications against prompt injection and related attacks. Cloudflare’s central recommendation is to replace fragmented, hardware-oriented security deployments with a unified, programmable platform. Doing so can shorten zero trust migrations, simplify operations, preserve consistent security controls across unusual environments, and establish a faster foundation for responsible AI adoption.

Read original(opens in new tab)
line4 min readCurated summary

Building an Enterprise LLM Service Part

FAA achieves a 96.1% response rate by favoring simple, maintainable techniques over complex AI architectures. Its design choices were to use RAG instead of knowledge-focused fine-tuning, retrieve complete documents before cutting them into question-relevant sections, and rely on a basic ReAct agent loop rather than elaborate workflows or multiple agents. The article concludes that improving documentation is more valuable than adding complexity when unanswered questions mainly result from missing source material. ## RAG Instead of Fine-Tuning - Fine-tuning was rejected as the primary method for injecting enterprise knowledge. - Research cited in the article found that fine-tuning was highly effective for changing a model’s style—about 97% success—but achieved only about 11% accuracy when teaching new factual knowledge. - FAA’s experiment with approximately 40 examples showed that the model answered the exact training question correctly but failed when the wording changed slightly. - Maintaining larger fine-tuning datasets would require experts to create, verify, and continuously update training examples whenever product documentation changes. - RAG is better suited to frequently changing product information because only the source documents need to be updated. - Fine-tuning may still be useful for domain-specific terminology or reasoning patterns, but not for keeping FAA’s product knowledge current. ## Retrieving Whole Documents Instead of Pre-Chunking - Conventional RAG systems split documents into small chunks before embedding them, improving semantic search precision. - Pre-chunking can remove essential context, especially when references such as “this case” or “the following settings” are separated from the text they depend on. - FAA’s documents are generally short, well-structured, focused on one product and topic, making whole-document retrieval practical. - Instead of chunking before search, FAA embeds and retrieves complete documents, then splits them after the relevant document is known. - The post-split process has two stages: - Split the document by Markdown headers into meaningful sections. - Use a lightweight LLM to select only the sections relevant to the user’s question. - For a question about creating and deleting a VM, the main model might receive only the “VM creation” and “VM deletion” sections. - This extra filtering call remains inexpensive because the lightweight model outputs only section indexes rather than generating a full response. - The key advantage is that splitting happens after the system understands the question, preserving context while delivering only the necessary information. ## ReAct Instead of Complex Agent Workflows - FAA tested plan-and-execute workflows, in which the model first creates a multi-step plan and then carries it out. - Planning and replanning increased system complexity without producing a noticeable improvement in answer quality. - With well-designed tools and carefully filtered context, the model was able to determine tool order on its own. - FAA therefore uses ReAct: the model reasons, takes an action, observes the result, and decides what to do next. - This approach allowed the agent to handle troubleshooting questions without a separate planning layer. ## Rejecting Multi-Agent Architectures - The team also tested specialized agents, such as separate VM and Kubernetes experts. - Delegating questions and assembling the results required additional LLM calls, increasing response time from roughly 9 seconds to 14 seconds in one test. - Multi-agent routing performed poorly for cross-domain questions, such as moving data from a VM to object storage. - Specialists could miss information outside their assigned domain, whereas a single agent could maintain the complete context. - FAA therefore kept one agent with access to progressively disclosed tools and relevant documentation. ## Documentation as the Main Bottleneck - Analysis of unanswered questions showed that about 50% were caused by a documentation gap: no reference document existed. - Other failures were mostly temporary API issues or questions outside FAA’s intended scope. - This suggests the core retrieval and agent system performs well when documentation is available. - The team shares missing questions with product teams, whose updated documents are then re-embedded and incorporated into future evaluations. The practical recommendation is to start with the simplest architecture that fits the data: use RAG for changing knowledge, preserve document context during retrieval, and let a capable model operate through a ReAct loop. In enterprise systems, improving the underlying documentation may produce greater gains than adopting more sophisticated AI frameworks.

Read original(opens in new tab)
discord3 min readCurated summary

Building on the Social Layer of Games: What’s New from GDC 2026

Discord’s GDC 2026 announcements focus on making social interaction a more seamless part of gaming. The company argues that friends significantly increase player engagement, citing longer sessions and more active game days when players connect through Discord. Its expanded Social SDK, profile features, and commerce tools aim to reduce the friction between discovering, playing, and supporting games. ## Discord as the Social Layer of Gaming - Steam players played fewer titles on average between 2021 and 2024, while Discord users on Steam increased the number of games they played. - Players spend a median of six times longer in a game when one Discord friend is present, and eight times longer when three friends are present. - Players in voice channels play games on approximately 66% more days. - Discord says more than 90 million people use the platform daily, including before, during, and after gaming sessions. ## Account Linking Connects Discord Friends to In-Game Friends - Discord’s Social SDK links a player’s Discord and game accounts, enabling persistent social presence and activity sharing. - Friends can see what someone is playing and join with less coordination, eliminating the need to exchange codes or arrange sessions manually. - Across initial integration partners, linked players showed: - A median 25% increase in active game days - 16% longer session durations - Partners include *Marvel Rivals*, *Rust*, *Pax Dei*, *Delta Force*, *Arena Breakout*, and *Predecessor*. - New Account Linking improvements include: - In-Discord linking prompts that appear contextually - Publisher authentication across multiple games - Native linking on iOS and Android - The SDK also adds: - Persistent message history - More advanced audio post-processing - Server-side moderation APIs - Custom Rich Presence displays, buttons, and activity types such as listening, watching, and competing ## Game Stats on Discord Profiles - The Game Stats Widget lets players display achievements and favorite in-game content on their Discord profiles. - In *Wuthering Waves*, users can show a favorite Resonator and recent achievements. - Discord profiles receive approximately 133 million monthly views, giving games a prominent space for player expression and discovery. - Discord is inviting developers to become early partners for custom Game Stats Widgets. ## Gifting and Game Commerce Inside Discord - Discord is expanding Social Commerce to let players buy and gift in-game items without leaving conversations. - A Game Shop can be embedded in a game’s official Discord server. - Items can also appear in direct messages, voice channels, and player profile Wishlists. - Players can gift items such as skins during a conversation, with recipients able to claim them later in-game. Discord’s overall strategy is to make social discovery, communication, achievement sharing, and in-game purchases part of the same connected experience. For developers, integrating the Social SDK and related tools can help players find games through friends, return more often, and engage with game communities beyond the game client.

Read original(opens in new tab)
gitlab2 min readCurated summary

Navigate repositories faster with the file tree browser

GitLab 18.9 introduces a collapsible file tree browser that makes repository navigation more like using an IDE. The panel keeps files and directories visible while reading code, reducing back-and-forth navigation and preserving context. It is available on GitLab.com, Self-Managed, and Dedicated, with support for accessibility, responsive layouts, and large repositories. ## Persistent Repository Context - A resizable, collapsible panel appears alongside file lists and code. - Users can expand or collapse directories and switch files without losing their place. - When opening a nested file directly, parent directories expand and the current file is highlighted. - The tree stays synchronized with the selected file or directory in the main content area. ## Filename Search - Press `F` to open the global file search dialog. - Search results can match part of a filename or extension. - Each result includes its parent directories, making the destination clear before navigation. - Press `Enter` to open the selected file. ## Keyboard and Accessibility Support - The browser follows the W3C ARIA treeview pattern. - Users can navigate with arrow keys, `Enter`, `Space`, `Home`, `End`, and character keys. - The design supports screen readers and keyboard-first workflows. ## Responsive Design and Performance - On desktop, the tree appears beside the file list and code. - On smaller screens, it becomes a toggleable left-side drawer. - On mobile, it is hidden to maximize the code-view area. - Pagination prevents large repositories from overwhelming the page and keeps the interface responsive. ## Availability and Usage - Open a repository at `/<project>/-/tree/<branch>`. - Select the file tree icon or press `Shift+F` to toggle the browser. - The feature is available on GitLab.com and was released in version 18.9 for GitLab Self-Managed and GitLab Dedicated. The file tree browser is recommended for anyone navigating large repositories, especially users who want IDE-like structure, faster file discovery, and better keyboard accessibility.

Read original(opens in new tab)
datadog2 min readCurated summary

When an AI agent came knocking: Catching malicious contributions in Datadog’s open source repos

Datadog announces that Gartner has named it a Leader in the 2026 Magic Quadrant for Observability Platforms. The surrounding product catalog presents Datadog as a broad platform spanning infrastructure, applications, data, logs, security, digital experience, software delivery, service management, and AI. However, the provided content does not include the blog post’s detailed analysis or Gartner’s specific evaluation criteria. ## Gartner Recognition - Datadog highlights its position as a **Leader** in the **Gartner Magic Quadrant for Observability Platforms 2026**. - The announcement links to a Gartner resource but provides no further details about the ranking, strengths, or limitations. ## Broad Observability Platform - **Infrastructure:** Infrastructure and container monitoring, metrics, Kubernetes autoscaling, network monitoring, serverless, cloud cost, storage, GPU monitoring, and Cloudcraft. - **Applications and data:** Application Performance Monitoring, service monitoring, profiling, dynamic instrumentation, database monitoring, data streams, data quality, and jobs monitoring. - **Logs and security:** Log management, sensitive-data scanning, audit trails, observability pipelines, cloud security, SIEM, code security, vulnerability management, and workload protection. - **Digital experience:** Browser and mobile RUM, product analytics, session replay, synthetic monitoring, mobile testing, error tracking, and experiments. - **Software delivery and service management:** CI visibility, test optimization, continuous testing, feature flags, code coverage, event management, SLOs, incident response, workflow automation, and service catalogs. - **AI capabilities:** Agent observability, GPU monitoring, AI integrations, AI agents, investigation tools, security analysis, MCP Server, and agent-building features. Datadog’s positioning is based on consolidating telemetry, security, developer, operations, and AI capabilities into one observability platform. To assess the Gartner recognition fully, readers would need the linked report or the complete article, which is not included here.

Read original(opens in new tab)
datadog3 min readCurated summary

When an AI agent came knocking: Catching malicious contributions in Datadog’s open source repos

Datadog describes how AI-powered attackers targeted its open-source repositories through malicious issues, pull requests, and comments. The campaign, attributed to the “hackerbot-claw” agent, focused on weaknesses in GitHub Actions and LLM-powered workflows. Datadog’s LLM-based review system and layered CI security controls detected the activity and helped limit its impact, while prompting further hardening. ## Why Open-Source Repositories Attract Attackers - Public repositories are attractive targets because automated CI/CD pipelines often build and execute code from external contributions. - Common attack techniques include: - Injecting user-controlled values, such as PR titles, into workflow scripts. - Using indirect poisoned pipeline execution to introduce malicious dependencies or build instructions. - Abusing `pull_request_target` workflows, which may run untrusted code with elevated permissions. - Prompt-injecting LLM-powered GitHub Actions used for issue triage, labeling, or code assistance. - Attackers may also disguise malicious changes through: - Large or obfuscated diffs. - Invisible Unicode characters. - Malicious libraries. - Imposter commits that resemble legitimate dependency references. ## Datadog’s LLM-Based Contribution Detection - Datadog receives dozens of external PRs each week across projects such as the Agent, tracers, SDKs, Vector, chaos-controller, and Stratus Red Team. - Its BewAIre system monitors GitHub events and selects security-relevant activity, including PRs and pushes. - BewAIre: - Extracts, normalizes, and enriches code diffs. - Sends them through a two-stage LLM pipeline. - Classifies changes as benign or malicious. - Produces a structured explanation for each verdict. - Malicious verdicts are forwarded to Datadog Cloud SIEM, where detection rules create enriched signals for the Security Incident Response Team to investigate. ## Hardening CI and Development Workflows - Datadog reduces the potential impact of successful attacks through multiple preventive controls: - Its `dd-octo-sts-action` generates minimally scoped, short-lived GitHub credentials using OIDC. - Long-lived and overly broad personal access tokens and GitHub Apps are being replaced. - Unused GitHub Actions secrets are identified and removed across thousands of repositories. - Organization-wide controls enforce branch protection, mandatory PR approval, commit signing, and lower-privilege default `GITHUB_TOKEN` permissions. - Engineers are provided with documented best practices and secure “golden paths” for CI development. ## The Hackerbot-Claw Campaign - Modern AI models are increasingly capable of offensive security tasks, especially when given tools, feedback loops, and autonomy. - StepSecurity reported an AI agent attacking open-source CI systems on March 1. - Between February 27 and March 2, the actor: - Opened 16 pull requests. - Created two issues and eight comments. - Targeted nine repositories across six organizations. - The activity was later linked to the hackerbot-claw agent, whose GitHub account was removed. - Datadog’s investigation began after BewAIre alerted the team to a suspicious contribution in the newly public `datadog-iac-scanner` repository on February 27. ## Practical Takeaway Organizations that accept public contributions should combine automated, AI-assisted review with least-privilege credentials, strict workflow permissions, secret management, mandatory approvals, and human incident response. Detection alone is insufficient; CI pipelines should be designed so that a malicious contribution has limited access and minimal opportunity to compromise secrets or production systems.

Read original(opens in new tab)
github3 min readCurated summary

How to scan for vulnerabilities with GitHub Security Lab’s open source AI-powered framework

GitHub Security Lab’s open-source Taskflow Agent uses AI-driven, multi-step auditing workflows to find high-impact vulnerabilities in web applications and open-source projects. The authors report more than 80 vulnerabilities, including authorization bypasses and private-data disclosures, with about 20 already disclosed. They argue that carefully designed taskflows and prompts can give LLMs enough freedom to discover vulnerabilities while reducing hallucinations and false positives. ## Running the Audits - The taskflows are available in the [`seclab-taskflows`](https://github.com/GitHubSecurityLab/seclab-taskflows) repository. - To run an audit: 1. Start a Codespace for the repository. 2. Wait for initialization. 3. Run `./scripts/audit/run_audit.sh myorg/myrepo`. - Audits may take one or two hours on a medium-sized repository. - Results are stored in SQLite and can be inspected in the `audit_results` table. - Rows marked with a check in `has_vulnerability` indicate potential findings. - A GitHub Copilot license and premium model requests are required. - The same repository should be audited multiple times because LLM results are nondeterministic; using different models may reveal different vulnerabilities. - Private repositories require changes to the Codespace configuration to grant access. ## How Taskflows Work - Taskflows are YAML files defining ordered tasks and dependencies for an LLM. - The `seclab-taskflow-agent` runs tasks sequentially and passes their results between stages. - Repository audits begin by dividing the codebase into functional components. - For each component, context is gathered, including: - Untrusted-input entry points - Intended privilege levels - Component purposes and behavior - This context is stored in a database for later auditing tasks. - Separate tasks can: - Suggest generic security issues - Carefully verify each suggested issue - Focus on specific vulnerability classes - Tasks can be reused across many components asynchronously through templated prompts and component-specific substitutions. ## Why Use Multiple Tasks - A single large prompt is less reliable because LLMs may omit steps in complex, multi-stage investigations. - Taskflows help control, debug, and structure the process even when models provide large context windows. - Breaking work into stages allows each result to be reviewed and reused as context for subsequent analysis. - Repeated task execution across components makes the approach scalable for large repositories. ## General Security Auditing - The team initially used the framework to triage CodeQL alerts, where strict instructions and predefined criteria helped limit false positives. - General auditing is more difficult because the LLM must search broadly for vulnerabilities rather than evaluate known alerts. - Greater freedom increases the risk of hallucinations and unexploitable findings. - The authors’ approach uses taskflow design and prompt engineering to preserve a high true-positive rate while allowing the model to investigate diverse security issues. ## Reported Vulnerabilities - The taskflows have found more than 80 vulnerabilities in open-source projects. - Many reported issues are high-impact, including: - Authorization bypasses - Information disclosure - Logging in as another user - Accessing private user data - Examples include exposing personally identifiable information in ecommerce shopping carts and authenticating to a chat application with arbitrary passwords. - The authors manually verify findings before reporting them and maintain an advisories page as disclosures become public. The practical recommendation is to run the open-source taskflows on your own projects, repeat audits with different models, and manually validate every result. The framework is intended to improve through shared taskflows, prompts, and findings across the security community.

Read original(opens in new tab)