Techlist.io - Korean Tech Blog Curator

line4 min readCurated summary

How We Built an SRE Bot That Reduced Our Team’s Repetitive Work by 90%

LINE Home DevOps created an SRE bot to reduce the repetitive work caused by growing services, Flava cloud migration, and increasing developer requests. By making Slack the central interface and automating Jira, Confluence, and workflow updates, the team reduced deployment-request handling from roughly 30 minutes to under one minute. The bot also improved tracking, consistency, and response speed, helping SREs move away from constant firefighting. ## Repetitive SRE Work and Its Costs - Developers frequently asked how to inspect Flava pod logs, request permissions, interpret errors, and access staging environments. - Deployment requests required manual movement between Slack, Confluence, and Jira: - Finding release checklists - Copying information into Jira - Creating missing Fix Versions - Linking Epics and active sprints - Sharing ticket links and deployment documentation - Each deployment request previously took about 30 minutes to an hour. - Manual processing caused omissions and mistakes, especially during urgent releases. - General requests were buried in Slack mentions, making ownership and completion status difficult to track. - Measurement showed that each SRE spent nearly half a day per week on repetitive work. ## Slack-Centered Automation The team adopted the principle that developers should only need Slack, while SREs should be able to manage work with a few clicks. - **Slack as the single source of truth:** Requests begin and remain trackable in Slack. - **Zero manual work:** Rule-based Jira and documentation tasks are automated. - **Immediate visibility:** Status changes and results are posted to Slack in real time. - **Permission control:** Only authorized SRE members can claim or complete requests. ## Key Technical Decisions ### Slack Workflows Instead of Slash Commands - Slash commands are easy to implement but depend on users entering correctly formatted text. - Slack Workflows provide structured forms with required-field validation. - Because Workflows are native Slack functionality, the team avoided building a separate user interface. - The lower usage barrier made adoption more likely. ### Asynchronous Processing - Slack requires event responses within three seconds. - Sequential calls to Jira, Confluence, and other APIs could exceed that limit. - The bot immediately acknowledges the request, then performs external work in the background. - Successes and failures are reported in the Slack thread, keeping processing transparent. ### Redis-Based State Management - In-memory state would be lost whenever the bot restarted. - Slack metadata APIs were considered too slow for real-time interactions such as emoji clicks. - Redis was selected for sub-100-millisecond lookups and persistent state. - A 30-day TTL limits stale data. - Redis transactions using `WATCH/MULTI/EXEC` ensure consistent updates when multiple SREs interact simultaneously. ### Hexagonal Architecture - The bot uses ports and adapters to isolate business logic from external systems. - The architecture separates: - Inbound Slack event adapters - Application use cases and business logic - Outbound Jira, Confluence, and Redis adapters - External API or SDK changes can be handled without modifying core business logic. - This structure also makes testing and future feature development easier. ## Automated Request Scenarios ### Deployment Requests - Developers submit required project, release-version, checklist, and other details through a Slack Workflow. - The bot automatically: - Creates a missing Jira Fix Version - Creates and configures the Jira ticket - Links the Epic - Adds the ticket to the active sprint - Finds the relevant deployment manual - Posts the result to the Slack thread - An SRE can click 👀 to claim the work. - Clicking ✅ completes the Jira ticket and posts a completion notification. - SRE effort falls from about 30 minutes to under one minute, with minimal risk of missing required fields. ### Emergency Deployments - Selecting an urgent request automatically sets Jira Priority to `Highest`. - The bot immediately announces the request in Slack. - An SRE can claim it with 👀, perform the deployment, and complete it with ✅. - The process reduces delays from roughly 30–40 minutes to about one minute. ### General SRE Requests - Requests such as production-access permissions are submitted through a structured Slack Workflow. - The bot creates a Jira ticket, links the Epic, assigns the active sprint, and sets an appropriate priority. - Slack retains the ticket link and status, eliminating the need to search through message history later. - SREs claim and complete the request using the same emoji-based workflow. The main recommendation is to automate repetitive, rule-based operations at the point where requests already occur. A Slack-centered, asynchronous bot with durable state and clean system boundaries can reduce manual effort while making ownership, progress, and completion visible to everyone.

Read original(opens in new tab)
discord2 min readCurated summary

Discord Update: March 24, 2026 Changelog

Discord’s March 24, 2026 changelog focuses on making desktop gaming and navigation faster and more convenient. New features include improved screen sharing, easier voice invitations, expanded game profiles, and gifting for Marvel Rivals items. Discord also introduced browser-style navigation controls, faster performance, role-member lists, and refreshed settings pages. ### Improvements for Game Nights - Screen shares can now be zoomed and panned with a mouse wheel or trackpad. - Single-window Go Live streams should start faster. - A new **Invite to Voice** option recommends server members and nearby friends for voice chats. - The Game Stats Profile Widget now supports **Wuthering Waves**, displaying information such as achievements and favorite Resonators. - Users can wishlist and gift Marvel Rivals items through the game’s Discord server and Game Shop. ### Faster Desktop Navigation - Behind-the-scenes performance improvements reduce lag when moving around the desktop app. - New **Back** and **Forward** buttons work similarly to browser navigation, including support for compatible mouse buttons. - Clicking an `@Role` mention now shows up to 100 users assigned to that role. - The Desktop Settings redesign continues with updated Notifications, Voice and Video, Clips, and Streamer Mode pages. ### Additional Developer News - Discord also highlighted new opportunities for game developers announced at this year’s Game Developers Conference, directing developers to a separate blog post for details. Overall, the update is aimed at smoother desktop performance, easier navigation, and more features for connecting around games.

Read original(opens in new tab)
google3 min readCurated summary

TurboQuant: Redefining AI efficiency with extreme compression

TurboQuant is a quantization framework designed to dramatically reduce memory use in large language models and vector search without sacrificing accuracy. It combines PolarQuant’s efficient vector compression with QJL’s one-bit residual correction to eliminate the overhead found in traditional quantization. Experiments show that it can compress KV caches to 3 bits, reduce memory by at least 6×, and accelerate attention-logit computation by up to 8×. ## The Memory Challenge in AI - High-dimensional vectors power language understanding, image features, vector search, and model attention. - These vectors consume substantial memory, particularly in the key-value (KV) cache used to store frequently accessed attention information. - Traditional vector quantization reduces vector size but often requires full-precision scaling or normalization constants for each block. - This metadata can add one or two bits per value, undermining the benefits of compression. ## TurboQuant’s Two-Stage Approach - TurboQuant first applies a random rotation to simplify the geometry of the data. - PolarQuant then compresses the transformed vectors using a standard quantizer, dedicating most bits to the vector’s primary information. - A remaining single bit is used by QJL to encode residual error. - QJL removes bias from the initial compression, improving the accuracy of attention-score calculations. - The approach requires no model training or fine-tuning. ## QJL: One-Bit Error Correction - QJL builds on the Johnson-Lindenstrauss Transform, which preserves important distances and relationships in high-dimensional data. - It represents each transformed value using only its sign: +1 or −1. - A specialized estimator combines low-precision stored data with a high-precision query. - This preserves accurate attention scores while introducing effectively zero memory overhead. ## PolarQuant: Compression Without Metadata Overhead - PolarQuant converts vectors from Cartesian coordinates into polar coordinates. - Instead of separately storing coordinate values, it represents vectors through: - A radius, capturing magnitude or signal strength - Angles, capturing direction and semantic structure - Because angular values follow a predictable, concentrated distribution, PolarQuant avoids expensive per-block normalization constants. - It recursively groups coordinate pairs and transforms their radii until the vector becomes one final radius plus a collection of angles. - This produces a compact representation with fixed, predictable boundaries. ## Experimental Results - The methods were tested on LongBench, Needle In A Haystack, ZeroSCROLLS, RULER, and L-Eval using Gemma and Mistral models. - TurboQuant achieved strong dot-product distortion and recall results while minimizing KV-cache memory. - On needle-in-a-haystack tasks, TurboQuant maintained perfect downstream performance while reducing KV memory by at least 6×. - PolarQuant was also nearly lossless on these tasks. - TurboQuant compressed KV caches to 3 bits without accuracy degradation. - Quantized models ran faster than the original uncompressed models. - On H100 GPUs, 4-bit TurboQuant delivered up to an 8× speedup for attention-logit computation compared with 32-bit keys. - The method has negligible runtime overhead and is relatively simple to implement. TurboQuant is presented as a practical way to make long-context LLMs and large-scale vector search more memory-efficient. Its combination of metadata-free PolarQuant compression and one-bit QJL correction is especially promising for deployments constrained by KV-cache capacity, latency, or GPU memory.

Read original(opens in new tab)
figma3 min readCurated summary

Agents, Meet the Figma Canvas | Figma Blog

Figma is opening its canvas to AI agents, allowing tools such as Claude Code and Codex to create and modify designs directly in Figma files. Through the `use_figma` tool and customizable skills, agents can use a team’s components, variables, design decisions, and workflows instead of producing generic designs. The feature is free during beta but is expected to become usage-based and paid. ## Agents Work Directly on the Figma Canvas - Figma’s MCP integration lets agents read and write Figma files through the `use_figma` tool. - Agents can create or update: - Design assets - Components - Files based on existing design systems - Designs linked to established variables and conventions - Teams can move between code, the command line, and Figma while keeping design context shared. - Figma positions the canvas as the place where product decisions become visible and refined. ## Working Across Code and Canvas - The existing `generate_figma_design` tool converts HTML from live apps and websites into editable Figma layers. - The new `use_figma` tool operates directly on the canvas, using existing components and variables. - The tools are intended to work together: - `generate_figma_design` brings current implementation details into Figma. - `use_figma` edits those designs or creates new system-aligned assets. ## Skills Encode Design Intent - Skills are Markdown-based instructions that tell agents: - Which workflow steps to follow - What sequence to use - Which team conventions to respect - What quality standards and specialized knowledge to apply - Anyone can author a skill without building a plugin or writing traditional code. - The foundational `/figma-use` skill teaches agents Figma’s structure and core principles. - Teams can customize that foundation to reflect their own design systems and working methods. ## Example Skills and Workflows Figma highlights skills for tasks such as: - Generating component libraries from code - Creating designs from existing components and variables - Producing accessibility specifications for VoiceOver, TalkBack, and ARIA - Creating components from structured JSON contracts - Applying design systems to existing designs - Managing spacing through variables and fallbacks - Synchronizing design tokens between code and Figma - Running parallel, multi-agent design workflows ## More Predictable and Self-Correcting Output - Skills make AI behavior more consistent by encoding repeatable instructions and implementation rules. - Agents can use screenshots to identify mismatches and iteratively refine generated screens. - Because agents work with real Figma structure—components, variables, and auto layout—corrections affect the underlying design system rather than only the visual appearance. - Team conventions become active rules that agents apply during creation, rather than static documentation they merely reference. Figma’s agent workflow is most useful when teams invest in well-defined components, variables, and skills. During the beta, teams can experiment with `use_figma` and community skills to automate design work while preserving their existing design intent and system standards.

Read original(opens in new tab)
google3 min readCurated summary

Mapping the modern world: How S2Vec learns the language of our cities

S2Vec is a self-supervised framework that converts buildings, roads, businesses, and infrastructure into general-purpose geospatial embeddings. By rasterizing these features into S2 Geometry cells and training a masked autoencoder to reconstruct missing areas, it learns the spatial “character” of neighborhoods without manually labeled data. It performs especially well for socioeconomic predictions in geographically unseen regions, while environmental tasks benefit from combining it with satellite imagery. ## Turning Geospatial Data into Images - Geospatial data is multimodal and unevenly distributed: urban blocks may contain hundreds of features, while rural areas contain few. - S2Vec uses hierarchical S2 Geometry cells to divide the Earth into regions at different resolutions. - It counts feature types within each cell—such as buildings, parks, roads, and businesses—and organizes them into multilayered raster images. - This makes complex geographic information compatible with computer vision methods developed for ordinary images. ## Learning with Masked Autoencoding - S2Vec masks portions of the rasterized map and trains a model to reconstruct the missing features from surrounding context. - Repeated training across global locations teaches relationships among urban elements, such as the likelihood of shops near residential buildings and transit stations. - The resulting embeddings are compact numerical representations of each location’s built environment. - Because training is self-supervised, S2Vec does not require worldwide labels for income, air quality, population, or other metrics. - The model can identify similar neighborhood types without being explicitly told concepts such as “financial district” or “suburban residential area.” ## Evaluation and Socioeconomic Performance - S2Vec was compared with models including SATCLIP, GEOCLIP, RS-MaMMUT, Hex2vec, and GeoVeX. - Tests covered population density, median income, carbon emissions, tree cover, and elevation. - Models were evaluated using mean squared error and both: - Interpolation, using random train/test splits - Extrapolation, predicting conditions in geographically unseen regions - S2Vec was generally the strongest individual model for zero-shot socioeconomic prediction, including population density and median income. - It performed competitively with established image-based approaches and exceeded GEOCLIP in the reported comparisons. ## Benefits of Multimodal Fusion - Combining S2Vec with satellite-image embeddings generally produced better results than either modality alone. - Built-environment data captures structures and infrastructure, while satellite imagery adds information about vegetation, terrain, and transportation patterns. - Fusion was particularly valuable for environmental prediction tasks. ## Limitations on Environmental Tasks - Built-environment features alone do not fully explain factors such as tree cover and elevation. - S2Vec was competitive for carbon-emissions prediction but weaker on some environmental metrics. - Satellite imagery embeddings improved performance by supplying information unavailable from counts of buildings, roads, and businesses. S2Vec points toward scalable geographic foundation models that replace task-specific feature engineering with reusable representations. In practice, it is most effective when its built-environment embeddings are combined with complementary imagery, especially for environmental analysis.

Read original(opens in new tab)
aws3 min readCurated summary

AWS Weekly Roundup: NVIDIA Nemotron 3 Super on Amazon Bedrock, Nova Forge SDK, Amazon Corretto 26, and more (March 23, 2026) | Amazon Web Services

This week’s AWS roundup highlights major updates across generative AI, data analytics, Java, serverless, logging, and Kubernetes. Notable announcements include NVIDIA Nemotron 3 Super on Amazon Bedrock, the Nova Forge SDK for customizing models, faster Redshift queries, and expanded EKS scaling and availability guarantees. The roundup also points readers to community initiatives, developer resources, and upcoming AWS events. ## Generative AI and Developer Tools - **NVIDIA Nemotron 3 Super** is now available through Amazon Bedrock. - Supports text generation, reasoning, summarization, and code generation. - Can be invoked through Bedrock’s unified API without managing infrastructure. - **Nova Forge SDK** simplifies fine-tuning and customizing Amazon Nova models. - Enables domain-specific adaptations for enterprise use cases. - Handles much of the underlying customization and deployment complexity. - **Kiro for students** provides free access to AI-powered development tools. - **Strands Steering Hooks** reportedly achieved 100% agent accuracy, outperforming prompt engineering and rigid workflows for controlling agent behavior. ## Data, Java, and Serverless Updates - **Amazon Redshift** now delivers up to 7x faster execution for new, uncached queries in dashboards and ETL workloads. - The improvement is especially useful for workloads with high query variability. - **Amazon Corretto 26** is generally available. - Includes current Java features, performance improvements, and security updates. - Supports Amazon Linux, Windows, macOS, and Docker environments. - **AWS Lambda** now exposes Availability Zone metadata for function invocations. - Helps with observability, troubleshooting, latency analysis, and multi-AZ architecture decisions. - **CloudWatch Logs** supports log ingestion through an HTTP-based protocol, reducing the need for custom agents or SDK integrations. ## Amazon EKS Enhancements - Provisioned Control Plane clusters now receive a **99.99% SLA**, compared with 99.95% for the standard control plane. - A new **8XL scaling tier** doubles Kubernetes API server request-processing capacity compared with the 4XL tier. - The larger tier targets demanding workloads such as AI/ML training, HPC, and large-scale data processing. ## AWS Community and Events - **AWS Builder Center badges** recognize contributions, challenges, and community participation. - AWS promotes community-driven learning through the “Keep Building Together” initiative. - Upcoming events include AWS Summits in cities such as Paris, London, Bengaluru, Singapore, Tel Aviv, and Stockholm; AWS Community Days in San Francisco and Romania; and the AWSome Women Summit LATAM in Mexico City. Overall, the announcements emphasize AWS’s continued investment in enterprise AI customization, higher-performance infrastructure, improved observability, and developer communities. Teams should evaluate the new Bedrock, Redshift, Lambda, and EKS capabilities according to their workload scale, reliability, and customization needs.

Read original(opens in new tab)
github2 min readCurated summary

GitHub expands application security coverage with AI‑powered detections

GitHub is expanding application security coverage with AI-powered detections that complement CodeQL’s traditional static analysis. The approach targets languages and frameworks that are difficult to support through semantic analysis alone, including Bash, Dockerfiles, Terraform, and PHP. Planned for public preview in early Q2, the system brings detection, automated remediation, and enforcement directly into pull requests. ## Hybrid Static Analysis and AI Detection - CodeQL remains the primary tool for deep analysis of supported languages. - AI-powered detections extend coverage to scripts, infrastructure definitions, and less-supported ecosystems. - The system can identify vulnerabilities and suggest fixes within the pull request workflow. - Internal testing analyzed more than 170,000 findings in 30 days, receiving positive feedback from over 80% of developers. - Early supported areas include: - Shell/Bash - Dockerfiles - Terraform/HCL - PHP - The capability is part of GitHub’s broader agentic detection platform, which also supports code quality and code review. ## Security Findings in Pull Requests - GitHub automatically analyzes changes when a pull request is opened. - It selects CodeQL or AI-powered detection based on the code being reviewed. - Findings appear alongside existing code-scanning results, without requiring developers to switch tools. - Example risks include: - Unsafe string-built SQL queries or commands - Weak cryptographic algorithms - Infrastructure configurations exposing sensitive resources - Detecting issues during review allows teams to address vulnerabilities before code is merged or deployed. ## Copilot Autofix for Remediation - GitHub connects detection with Copilot Autofix, which proposes fixes developers can review, test, and apply. - Autofix resolved more than 460,000 security alerts in 2025. - Alerts were resolved in an average of 0.66 hours with Autofix, compared with 1.29 hours without it. - This reduces the gap between discovering a vulnerability and correcting it. ## Security Enforcement at Merge - GitHub positions pull requests as the point where security policies can be enforced. - Detection, remediation, and governance operate within the same workflow. - Teams can reduce risk without adding separate post-deployment review steps. - GitHub plans to demonstrate the technology at RSAC, highlighting hybrid detection and developer-native remediation. GitHub’s recommendation is effectively to combine CodeQL’s precision with AI-based coverage for modern, diverse repositories, while using Copilot Autofix and merge policies to turn findings into timely, enforceable fixes.

Read original(opens in new tab)
cloudflare3 min readCurated summary

Launching Cloudflare’s Gen 13 servers- trading cache for cores for 2x edge compute performance

Cloudflare’s Gen 13 servers use AMD EPYC 5th Gen Turin processors to provide up to twice as many cores as Gen 12. However, Turin’s much smaller per-core cache caused the legacy FL1 request-handling layer to suffer severe latency increases, despite higher throughput. Cloudflare found that tuning alone could not fully solve the problem, reinforcing the need for FL2, a Rust-based rewrite designed to scale with cores rather than depend heavily on cache. ## Turin’s Core-Heavy Architecture - Gen 13 Turin processors offer: - Up to 192 cores and 384 SMT threads, compared with Gen 12’s 96 cores. - Improved instructions per cycle through the Zen 5 architecture. - Up to 32% lower power consumption per core. - DDR5-6400 support for greater memory bandwidth. - The tradeoff is substantially less cache: - Gen 12 Genoa-X provides 12 MB of L3 cache per core through 3D V-Cache. - The 192-core Turin 9965 provides only 2 MB per core. - This architecture favors aggregate throughput but challenges workloads dependent on cache locality. ## FL1’s Cache and Latency Problems - FL1, based on NGINX and LuaJIT, was optimized for Gen 12’s large cache. - AMD uProf measurements showed: - Dramatically higher L3 cache miss rates on Turin. - More requests requiring slow DRAM access. - Increasing latency as CPU utilization and cache contention rose. - An L3 hit takes roughly 50 CPU cycles, while a DRAM fetch can take more than 350 cycles. - As a result, Gen 13’s additional cores delivered throughput gains but introduced unacceptable latency penalties. ## Throughput Gains at an Unacceptable Cost - With FL1, Gen 13 produced: - 10% more throughput on the 128-core Turin 9755. - 31% more on the 160-core Turin 9845. - 62% more on the 192-core Turin 9965. - The Turin 9965 offered the strongest total-cost-of-ownership benefits. - However, latency increased by more than 50% at high CPU utilization, which would negatively affect customer experience and violate performance requirements. ## Hardware and Resource Tuning - Cloudflare tested several mitigations with AMD: - Hardware prefetcher and Data Fabric Probe Filter adjustments produced only marginal improvements. - Adding FL1 workers increased throughput but took resources away from other services. - CPU pinning and isolation provided limited benefits. - AMD’s Platform Quality of Service (PQOS) was used to control cache and memory-bandwidth sharing across Turin’s Core Complex Dies. ## Cache Isolation with PQOS - Reserving part of a single CCD’s cache for FL1 produced less than 5% additional throughput. - Configurations assigning FL1 50–75% of each CCD’s cache also delivered less than 5% improvement and caused minor degradation elsewhere. - A socket-level approach was more successful: - Six of twelve CCDs, aligned with a NUMA domain, were dedicated to FL1. - This provided more than 15% incremental throughput while keeping latency acceptable. - These results showed that workload placement and cache locality could help, but they were not a complete substitute for software designed around Turin’s cache profile. Cloudflare’s broader solution was FL2, a Rust-based rewrite of its core request-handling layer. By reducing dependence on large per-core caches, FL2 enabled Gen 13’s higher core count to translate into scalable edge-compute performance without the latency penalties seen with FL1.

Read original(opens in new tab)
cloudflare3 min readCurated summary

Inside Gen 13- how we built our most powerful server yet

Cloudflare’s Gen 13 server is a hardware redesign aligned with its Rust-based FL2 request-processing stack. By choosing a 192-core AMD EPYC Turin 9965, doubling memory, expanding storage and networking, and adding stronger security and accelerator support, Cloudflare targets up to twice the throughput of Gen 12 while remaining within latency limits. The design also improves efficiency, rack density, and operational simplicity. ## Gen 13 at a Glance - Uses a 2U, single-socket design. - Key specifications: - 192-core AMD EPYC 9965 processor - 768 GB of DDR5-6400 memory - Three 7.68 TB E1.S PCIe 5.0 NVMe drives - Dual 100 GbE OCP 3.0 networking - 1,300W Titanium-grade power supply - ASPEED AST2600 BMC and AST1060 hardware root of trust - Compared with Gen 12, Gen 13 provides: - Up to 2× throughput - Up to 50% better performance per watt - Up to 60% more throughput per rack at the same power budget - Twice the memory capacity, 1.5× the storage, and 4× the network bandwidth - PCIe encryption in addition to memory encryption - Better support for high-heat PCIe accelerators ## Choosing the CPU - Gen 12 used a 96-core AMD EPYC Genoa-X 9684X with: - 400W TDP - 1,152 MB of L3 cache - Cloudflare evaluated three Turin processors: - 9755: strongest per-core performance - 9845: lower socket power and fewer cores - 9965: highest core count and better efficiency - The Turin 9965 was selected with 192 cores and 384 threads, doubling Gen 12’s hardware threads. - Its L3 cache is much smaller—384 MB total, or 2 MB per core versus Gen 12’s 12 MB per core—but FL2 workloads depend less on large L3 caches than the previous FL1 stack. - FL2 scales nearly linearly with additional cores, allowing the 9965 to deliver up to 100% higher throughput. - Production testing showed the 9965 achieved the best aggregate requests per second and favorable performance per watt at its 500W TDP. - Higher compute density also means fewer servers to provision, patch, monitor, and operate. - Turin’s support for DDR5-6400, PCIe 5.0, and CXL 2.0 provides a longer upgrade and security-support runway. ## Memory Bandwidth and Capacity - Gen 13 doubles memory from 384 GB to 768 GB while retaining 4 GB per core. - All twelve memory channels are populated using one 64 GB DDR5-6400 ECC RDIMM per channel. - This configuration delivers approximately 614 GB/s of peak memory bandwidth per socket, a 33.3% increase over Gen 12. - Using identical DIMMs across all channels enables balanced interleaving, distributing memory accesses across the full memory subsystem. - The design is intended to prevent the 192-core processor from being starved of data during highly parallel workloads. Cloudflare’s central design choice was to match hardware to the characteristics of FL2 rather than preserve Gen 12’s cache-heavy strategy. For workloads that scale well across cores, the Turin 9965 and fully populated memory system offer higher throughput, better rack economics, and simpler fleet operations.

Read original(opens in new tab)
gitlab2 min readCurated summary

Agile planning gets a boost from new features in GitLab 18.10

GitLab 18.10 introduces a unified work items list and saved views to improve Agile planning. Teams can now manage epics, issues, and other work items in one place while saving filters, sorting, and display settings for repeated workflows. These changes reduce context switching and establish the foundation for broader planning views and customizable work item types. ## Unified Work Items List - Combines epics, issues, and other work item types into a single list. - Replaces the need to navigate between separate epic and issue pages. - Provides a foundation for future hierarchy and table views that show relationships between work items. - GitLab plans to bring additional planning workflows, including Boards, into the same unified experience. - The term “work items” supports future customization of item types and names instead of limiting the system to traditional “issues.” ## Why GitLab Is Making the Change - GitLab’s 2024 Agile planning vision identified separate epic and issue experiences as a source of friction. - The work items framework provides a shared architecture for consistent functionality across planning objects. - The new list and saved views represent progress toward that unified planning model. ## Saved Views - Save customized list configurations for future use. - Configurations can include: - Filters - Sort order - Display options - Help teams avoid repeatedly setting up common workflows. - Support consistent reporting, status checks, iteration planning, backlog refinement, and portfolio planning. - Saved views can be shared with teammates to standardize how work is reviewed. ## Future Planning Experience - Users will be able to move between list, board, table, and other views while retaining the same filter scope. - Planned hierarchy and nested table views will improve portfolio-level planning. - Future board capabilities may include swimlanes based on any work item attribute. - The unified framework is intended to make planning views more flexible and interconnected. ## Transition and Feedback - Existing workflows based on separate epic and issue pages may require adjustment. - GitLab acknowledges the change could be disruptive but says it reflects years of feedback and substantial investment in the work items framework. - Users are encouraged to try the new capabilities and provide feedback through GitLab’s feedback issue. Teams should begin testing the unified work items list and saved views in their everyday planning workflows, while offering feedback to help GitLab refine the transition.

Read original(opens in new tab)
datadog1 min readCurated summary

When upserts don't update but still write: Debugging Postgres performance at scale

The provided content does not include the tech blog post itself. It consists primarily of Datadog’s website navigation and a promotional link about its Gartner recognition, so there is not enough article content to summarize reliably. ## Available Information - Datadog was named a Leader in the 2026 Gartner® Magic Quadrant™ for Observability Platforms. - The page navigation lists Datadog products across: - Infrastructure and application monitoring - Database and log management - Security - Digital experience monitoring - Software delivery - Incident and service management - AI-powered observability tools - The URL path references `debugging-postgres-performance`, suggesting the intended article may concern PostgreSQL performance debugging, but the article text is not included. Please provide the blog post’s body or a working text extraction for a substantive summary.

Read original(opens in new tab)
datadog3 min readCurated summary

When upserts don't update but still write: Debugging Postgres performance at scale

Datadog needed to track when ephemeral hosts were last seen so inactive hosts could be deleted after seven days. A seemingly inexpensive PostgreSQL upsert caused disk writes to double and WAL syncs to quadruple, despite most operations not changing any data. Investigating the WAL revealed that conflict-handling upserts still lock conflicting rows and generate WAL activity, consuming the database’s limited write capacity. ## Tracking Host Activity Efficiently - Hosts stop reporting telemetry when they terminate, but Datadog has no direct termination signal. - Hosts inactive for seven days can be safely removed from the metadata store. - Updating the main host table on every observation would be too expensive because: - Large data centers generate more than 25,000 observations per second. - PostgreSQL MVCC creates a new row version for every update. - Updating the main table would rewrite all host metadata. - Datadog created a separate `host_last_ingested` table containing: - `host_id` as the primary key - `last_ingested` with a default timestamp - The table used `fillfactor=80` to leave page space for future updates. - No index was created on `last_ingested`, allowing updates to use Heap-Only Tuples (HOT) and avoid additional index writes. - Because only daily freshness was required, the timestamp needed to change at most once per day. ## The Conditional Upsert The initial query inserted a host if it did not exist and otherwise updated `last_ingested` only when the previous value was more than a day old: ```sql INSERT INTO host_last_ingested AS t VALUES ($1, now()) ON CONFLICT (host_id) DO UPDATE SET last_ingested = EXCLUDED.last_ingested WHERE t.last_ingested < EXCLUDED.last_ingested - '1 day'::interval; ``` - New hosts produced an insert. - Recently seen hosts matched the conflict but were expected to be no-ops because of the `WHERE` clause. - The team therefore expected most queries to avoid meaningful writes. ## Unexpected Disk and WAL Activity - During a gradual rollout at roughly 500 upserts per second: - Insertions initially increased as expected. - Actual updates remained mostly flat. - Write IOPS more than doubled. - WAL syncs increased by approximately the same amount. - This showed that the absence of an applied update did not mean the query was free. - Since PostgreSQL must flush WAL records at transaction commit, additional WAL activity directly increased disk pressure. - A PostgreSQL cluster’s single-writer design makes this write budget particularly important. ## Inspecting PostgreSQL WAL - PostgreSQL records database changes in its Write-Ahead Log, including table changes, index modifications, and related transaction activity. - The team used the `pg_walinspect` extension, available starting in PostgreSQL 15: ```sql CREATE EXTENSION pg_walinspect; ``` - Its `pg_get_wal_records_info` function allows inspection of WAL records between two Log Sequence Numbers (LSNs). - Examining the WAL helped explain why the conditional upsert generated writes even when the `WHERE` condition prevented the row update. - The underlying issue was that `ON CONFLICT DO UPDATE` still locks the conflicting row, and that locking activity is recorded in the WAL. The key lesson is that a PostgreSQL upsert that reports zero processed rows is not necessarily a true no-op. Conditional conflict updates can still create substantial WAL and locking overhead, so WAL inspection is essential when database write metrics do not match apparent update volume.

Read original(opens in new tab)
stripe2 min readCurated summary

Three of the biggest fraud trends from MRC Vegas 2026

Fraud is becoming more automated, adaptive, and difficult to detect with traditional rules-based systems. At MRC Vegas 2026, leading fraud teams emphasized dynamic authentication, fraud controls embedded directly into agentic payments, and layered identity verification to address deepfakes and synthetic identities. The common goal is to reduce friction for trusted customers while applying stronger defenses where risk is highest. ## Dynamic Authentication Based on User Intent - Universal authentication creates unnecessary friction, increases false positives, and can cause businesses to lose legitimate customers and their long-term value. - Airbnb advocates building behavioral profiles over time to measure “high-trust velocity”—the likelihood that a user’s activity reflects legitimate intent. - Trusted users can proceed without additional challenges, while authentication is reserved for the small percentage of traffic proven to be risky. - Stripe Radar’s adaptive 3DS uses AI to trigger authentication only when transaction behavior appears unusual. - Stripe reports that eligible businesses have seen fraud reductions of more than 30% with this approach. ## Fraud Detection for Agentic Commerce - Ashley Furniture’s existing rules-based system handled different authorization needs for quick-ship products and custom orders. - That model became insufficient when AI agents began making purchases across channels. - Fraud detection must be part of the payment infrastructure and evaluate transactions in real time, rather than analyzing them only after purchase. - Stripe Shared Payment Tokens let agents use a customer’s saved payment method without exposing payment credentials. - Combined with Stripe Radar, these tokens transmit risk signals such as potential disputes, card testing, stolen-card usage, and issuer declines. - These signals help distinguish legitimate, high-intent agents from low-trust automated bots. ## Deepfakes and Synthetic Identity Fraud - Fake identities are easier to create because criminals can access document templates and generative AI impersonation tools. - Fraudsters may produce convincing fake IDs, images, voices, and videos with limited resources. - Effective verification depends on identifying inconsistencies that forgeries fail to reproduce, such as incorrect signatures, mirrored photos, or mismatched expiration dates. - No single verification check is reliable enough; multiple independent checks are necessary. - Stripe Identity uses AI to detect fake documents and spoofed photos, compare ID images with selfies, and validate Social Security numbers and addresses against databases. Businesses should replace blanket controls with risk-sensitive interventions: minimize friction for trusted users, integrate fraud detection into agent-driven payment flows, and use layered identity verification to catch increasingly convincing forgeries.

Read original(opens in new tab)
line4 min readCurated summary

Internalizing without specifications: Proving equivalence through validation logic

The article describes how LINE Plus safely internalized black-box e-commerce systems without specifications or source code. The team built an automated equivalence-testing loop using Kafka, CDC, OpenSearch, and ksqlDB to compare legacy and new behavior at massive scale. By repeatedly identifying differences, fixing logic, and rechecking results, they could reduce discrepancies toward zero while also measuring performance and protecting production stability. ## Domain: Products, Catalogs, and Data Ingestion - **Products** are individual seller-listed items, potentially with different prices and shipping conditions. - **Catalogs** group products representing the same model or product type and provide derived value such as: - Real-time lowest prices - Unit-price metrics such as price per 100 ml - **Ingestion** receives large product files from sellers, validates and transforms them into internal formats, and updates product and catalog data. - Because the platform contains tens of millions of catalogs and hundreds of millions of products, small logic differences can affect the entire service. ## The Verification Loop - The goal was not merely to find errors, but to help developers understand and correct them quickly. - Inputs had to be identical for both systems, such as: - The same IDs - The same time-based snapshot - The same product files - Outputs were compared according to system type: - API response objects - Database update values - Final registered product data - The general loop consisted of: - **Trigger:** Database changes, developer requests, or file arrivals - **Execution:** Send identical inputs to legacy and new systems - **Comparison:** Apply logic suited to reads, updates, or end-to-end flows - **Processing:** Store detailed differences and produce real-time statistics - **Action:** Developers inspect dashboards or Slack alerts, fix the implementation, and repeat ## Query Logic Verification - The catalog API was difficult to reproduce because it had over 100 response fields, complex filters, undocumented defaults, and unknown sorting behavior. - CDC streamed database binary-log changes into Kafka, allowing verification to begin from many real catalog states. - The verifier made dual API calls and compared legacy and new responses field by field. - Responses were converted into `Map<String, Object>` structures and compared recursively, avoiding the need to model every response class. - If values differed only because list ordering varied, the verifier sorted serialized values and performed a second comparison. - This helped distinguish real implementation defects from harmless ordering differences. - Kafka isolated verification traffic from production services while handling large event volumes. - Difference events were written to Kafka topics and indexed in OpenSearch for detailed investigation. - ksqlDB aggregated streaming discrepancies and sent Slack notifications when abnormal patterns appeared. - Rate limiting restricted repeated errors, such as those from the same field, to a manageable sample per minute. - Because both APIs were called in parallel, the same pipeline also measured and compared their response times. ## Update Logic Verification - The second case involved recalculating catalog statistics whenever product or catalog data changed. - Unlike read verification, this process tested state transitions and asynchronous updates. - When CDC detected a relevant change: - The new statistics logic calculated an expected result. - The verifier compared it with the result actually written by the legacy logic. - Recursive Map-based comparison checked deeply nested statistics fields. - To avoid wasting resources, verification was triggered only for updates related to the catalog-statistics module. ## Handling Asynchronous Lag - Kafka-based processing caused timing gaps: the verifier could read the database before the legacy update had completed. - The team introduced an **N-attempt retry queue**: - Temporarily inconsistent events were requeued. - Only differences that remained after several retries were treated as genuine defects. - The verifier remained a separate process rather than being embedded in the production statistics stream. - This avoided adding load or latency to the existing processing pipeline while preserving independent verification. ## ETL Batch Verification for Missing Triggers - Real-time comparison could detect incorrect results, but not cases where an update should have happened and never occurred. - During refactoring, a complex combination of product and catalog field changes contained a missing trigger condition. - As a result, some statistics remained stale without generating any comparison event. - To detect these silent omissions, the team designed a separate batch-verification process using ETL data alongside the real-time stream checks. The practical recommendation is to treat system internalization as an evidence-building process: define identical inputs and observable outputs, compare legacy and replacement systems continuously, isolate verification through event streams, and supplement real-time checks with batch validation for silent or missing updates.

Read original(opens in new tab)
toss3 min readCurated summary

Automating Service Vulnerability Analysis using LLM #2

The post explains how Toss Security Research improved AI-driven vulnerability analysis in a research network. Its main challenges were efficiently providing large codebases to an AI and making analysis results consistent and complete. The solution combined a custom code-browsing MCP server with SAST tools used not to identify vulnerabilities directly, but to enumerate all input-to-function paths that the AI must review. ## Efficiently Providing Large Codebases - Tools such as Cursor and Claude Code can search large projects, but primarily rely on pattern matching with tools like ripgrep. - Without prebuilt indexes, they may miss relevant code or waste tokens exploring unnecessary files. - The team built an MCP server that: - Uses **ctags** to index symbol definitions. - Uses **tree-sitter** to parse function boundaries. - Allows AI to access code remotely, similar to IDE features such as “Go to Definition” and “Find References.” ### SourceCode Browse MCP The MCP server provides four main tools: - **`find_references()`** - Searches for symbols or patterns using ripgrep. - Returns file paths, line numbers, snippets, total matches, and whether results were truncated. - **`read_definition()`** - Looks up definitions through the ctags index. - Returns metadata such as file, line, symbol type, language, signature, and scope. - Uses tree-sitter to include the complete function body when requested. - **`read_source()`** - Reads a configurable number of lines before and after a target line. - Lets the AI retrieve only the relevant local context instead of entire files. - **`get_project_structure()`** - Returns the indexed project’s directory structure. - Provides the AI with a project “blueprint,” which is especially important in remote environments where it cannot inspect the repository locally. The MCP workflow is to locate relevant symbols with `find_references()` and `read_definition()`, inspect nearby code with `read_source()`, and use `get_project_structure()` to understand the overall project. ## Improving Consistency and Accuracy - AI analysis produced inconsistent results: for example, it might find all 10 XSS vulnerabilities in one run but only 8 in another. - This variability made the results difficult to trust. - The team combined AI analysis with SAST tooling to ensure complete coverage. ## Using SAST to Enumerate Review Candidates - Rather than passing SAST-detected vulnerabilities directly to the AI, the team used SAST as a candidate-generation tool. - This avoids limiting the AI to vulnerabilities that the SAST engine itself knows how to detect. - SAST extracts every location where untrusted input enters the application and tracks its possible flow to function calls. - Custom Semgrep taint rules identify sources such as: - Spring `@RequestParam` - `@PathVariable` - `@RequestHeader` - Fields read from `@RequestBody` DTOs - `@RequestPart` - `@ModelAttribute` - `@RequestAttribute` - Potential sinks include generic function calls and object method calls. - The AI then reviews every extracted source-to-sink path, combining the completeness of static analysis with the broader reasoning ability of an LLM. The overall approach is to use deterministic indexing and SAST for coverage, while relying on AI for deeper vulnerability interpretation.

Read original(opens in new tab)