Techlist.io - Korean Tech Blog Curator

toss3 min readCurated summary

From Intern to Solo Designer: Growth

As a Toss Bank product design intern, Jeon Nuri designed experiments to improve non-member sign-up conversion. She prioritized the funnel using speed and impact, studied previous experiments, and learned that clear, narrowly defined hypotheses were more valuable than constantly generating new ideas. The experience showed that failed experiments can still guide better decisions when they produce actionable learning. ## Prioritizing the Right Funnel Stage - The largest drop-offs occurred in the intro, consent, and identity-verification screens. - Consent and identity verification were shared modules requiring legal and compliance review, making rapid iteration difficult. - The intro screen could be changed more quickly and had the potential to affect the greatest number of users. - Based on this speed-versus-impact assessment, she chose the intro screen as the starting point. ## Learning from Previous Experiments - Instead of immediately designing new concepts, she reviewed existing experiments, including both winners and unsuccessful variations. - She examined: - The problem each experiment addressed - The reasoning behind its hypothesis - How the test variation was designed - Experiments from unrelated screens were also useful because their problem definitions and hypothesis structures could be adapted. - The main lesson was that inexperienced experimenters benefit more from systematically analyzing existing learning than from rushing to create new ideas. ## First Experiment: A Counselor Concept - The first variation presented benefits as if they were being recommended by a counselor and offered a small number of choices. - The hypothesis was vague: fewer choices would increase conversion. - The result was negative: - Click-through rate fell by more than 10%. - Conversion rate fell by more than 3%. - The design actually introduced more choices than the original, which had only one CTA button. - The experiment also failed to consider why users had entered the screen and whether they needed recommendations. - This led her to analyze the existing screen and user context before creating a hypothesis. ## Identifying and Solving Concrete Problems - Rather than inventing an entirely new design, she identified two specific weaknesses in the existing version: - The copy did not clearly communicate benefits users cared about. - Images loaded slowly, taking two to three seconds on low-end devices. - Previous experiments showed that users responded well to messages about high interest rates and receiving interest daily. - She incorporated those themes into the copy and optimized the visuals with newer graphics and lower-weight image formats. - Both click-through rate and conversion rate increased, demonstrating that a hypothesis grounded in clear problems can provide a stable direction for design. ## Making Benefits Easier to Imagine - Building on the earlier results, she changed functional wording into language that helped users imagine a concrete situation and immediate benefit. - Instead of simply explaining that interest could be earned after depositing money for one day, the revised copy foregrounded the moment when users would experience the benefit. - Copy alone increased CTR by 5% and also produced a meaningful improvement in CVR. - The result reinforced that different expressions of the same information can create significantly different first impressions. ## Principles for Designing Experiments - Break the funnel into stages and prioritize opportunities by speed and potential impact. - Understand the existing context before defining the core problem. - Study previous experiments through their hypotheses and problem definitions, not just their numerical outcomes. - Establish a clear hypothesis and success metric before designing the variation. - Make sure the experiment visibly tests the stated hypothesis. - Treat failure as input for the next decision rather than as wasted effort. A practical starting point for new designers is to begin with a small, focused experiment—but make the hypothesis precise enough to guide both the design and the next iteration.

Read original(opens in new tab)
gitlab2 min readCurated summary

AI can detect vulnerabilities, but who governs risk?

AI can increasingly detect vulnerabilities and suggest fixes, but detection alone does not make software secure. The post argues that enterprises also need governance, context, continuous assurance, and supply-chain oversight to determine which risks are acceptable and what can ship. GitLab presents its platform as the orchestration layer for enforcing these controls across AI-assisted development. ## Trust Requires Governance - AI analysis is not the same as accountability. - Humans must define acceptable risk, policies, guardrails, separation of duties, and audit requirements. - As autonomous agents gain more control over development, stronger governance becomes essential rather than optional. - Governance enables organizations to trust AI at scale without relying on unchecked autonomy. ## Context Matters Beyond Code Scanning - LLMs typically assess code in isolation, while enterprise platforms can evaluate its broader context. - Important factors include: - Who authored the change - The application’s business criticality - Its dependencies and infrastructure interactions - Whether vulnerable code is reachable in production - Whether the vulnerability is exploitable in the actual runtime environment - Context reduces noisy alerts and supports faster, more effective risk triage. ## Risk Changes Continuously - Dependencies, environments, and system interactions evolve after an initial scan. - A clean static scan does not guarantee that software remains safe at release time. - Organizations need continuous assurance embedded throughout development, testing, and deployment. - Detection identifies risk, while ongoing governance determines how that risk is managed. ## Governing AI-Generated Software - Modern software combines AI-generated code, open-source libraries, and third-party dependencies across many projects. - Governing this entire supply chain is more difficult than detecting flaws in individual code changes. - The post argues that developer-side AI tools alone are not designed to provide organization-wide enforcement and auditability. - GitLab Ultimate is positioned as a platform combining policy enforcement, security scanning, governance, and auditing within software delivery workflows. Organizations adopting AI most successfully will pair capable coding assistants with strong, continuous governance. The practical recommendation is to treat AI security as a platform and lifecycle-management problem—not merely a vulnerability-detection problem.

Read original(opens in new tab)
github2 min readCurated summary

What’s new with GitHub Copilot coding agent

GitHub Copilot coding agent is becoming more capable at handling delegated development work from issue to pull request. Recent updates let users choose models, receive self-reviewed and security-checked changes, apply team-specific workflows through custom agents, and move tasks between the cloud and local CLI without losing context. Together, these features aim to reduce cleanup and make background coding tasks more reliable. ## Model selection for different tasks - The Agents panel now includes a model picker. - Users can choose faster models for routine work, stronger models for complex refactoring or integration tests, or let GitHub select automatically. - Model selection is currently available to Copilot Pro and Pro+ users; Business and Enterprise support is planned. ## Self-review before pull requests - Copilot coding agent now runs Copilot code review on its own changes before opening a pull request. - It incorporates feedback and improves the patch, such as simplifying overly complex code. - Users can inspect the review and iteration steps in the task logs before reviewing the resulting pull request. ## Integrated security checks - The agent performs code scanning, secret scanning, and dependency vulnerability checks during its workflow. - Vulnerable dependencies, exposed API keys, and other risky patterns can be identified before a pull request is created. - These code-scanning capabilities are provided without requiring a separate GitHub Advanced Security subscription for this workflow. ## Custom agents for team processes - Teams can define specialized agents in `.github/agents/`. - Custom agents can enforce repeatable procedures, such as benchmarking code before and after a performance change. - Agents can be shared across an organization or enterprise to standardize development practices. - The article describes a custom performance agent that achieved a 99% improvement on a targeted lookup function. ## Cloud and local CLI handoff - Cloud coding-agent sessions can be continued locally with their branch, logs, and context intact. - Users can select “Continue in Copilot CLI” and run the provided command in a terminal. - Pressing `&` in the CLI delegates work back to the cloud without restarting the task. GitHub recommends using these features to match models and workflows to each task, while reviewing the agent’s logs and pull requests. Planned capabilities include private mode, planning before coding, and tasks that produce summaries or reports instead of pull requests.

Read original(opens in new tab)
aws2 min readCurated summary

AWS Security Hub Extended offers full-stack enterprise security with curated partner solutions | Amazon Web Services

AWS Security Hub Extended expands Security Hub from an AWS-focused service into a broader enterprise security platform. It combines AWS services such as GuardDuty and Inspector with curated partner solutions covering endpoints, identity, email, networks, data, cloud, AI, and security operations. The plan simplifies procurement and operations through AWS billing, normalized findings, and a unified console. ## Curated Partner Security Solutions - Includes offerings from partners such as CrowdStrike, Okta, Proofpoint, SailPoint, Splunk, Zscaler, and others. - Covers security needs across endpoint, identity, email, network, data, browser, cloud, AI, and security operations. - Lets organizations combine AWS and partner tools to detect risks spanning multiple parts of their technology stack. ## Simplified Procurement and Billing - AWS acts as the seller of record. - Customers receive pre-negotiated pay-as-you-go pricing, one monthly bill, and no long-term commitments. - Consumption-based metering is handled automatically after onboarding. - AWS Enterprise Support customers receive unified Level 1 support. ## Unified Findings and Operations - Findings from participating solutions are emitted in the Open Cybersecurity Schema Framework (OCSF). - Security Hub automatically aggregates and normalizes findings in one location. - The unified view helps teams prioritize and respond to critical risks more quickly. ## Access and Availability - Customers can find the offerings in the Security Hub console under **Management → Extended plan**. - Partner details, subscriptions, and onboarding are available directly through the console. - The plan is generally available in all commercial AWS Regions where Security Hub operates. - Pricing supports either flexible pay-as-you-go or flat-rate options. Organizations seeking broader security coverage can use Security Hub Extended to consolidate partner procurement, billing, findings, and operations through a single AWS-managed experience.

Read original(opens in new tab)
dropbox3 min readCurated summary

Using LLMs to amplify human labeling and improve Dash search relevance

Dropbox Dash improves AI answers through retrieval-augmented generation (RAG): enterprise search retrieves relevant company documents, and an LLM uses a small subset of them to generate grounded responses. Because ranking determines which documents reach the LLM, search relevance depends heavily on high-quality query–document labels. Dash combines a small set of human judgments with large-scale LLM-generated labels to produce training data efficiently while retaining human oversight. ## How Dash search ranking works - Dash uses a trained ranking model, such as XGBoost, rather than manually configured rules. - The model learns from query–document pairs labeled on a 1–5 relevance scale: - **5:** Closely matches the user’s intent. - **1:** Not useful enough to display. - Relevance depends on the query, user context, and timing; it is not an intrinsic property of a document. - Ranking quality is especially important because enterprises may have millions or billions of indexed documents, while only a small selection can be sent to the answer-generating LLM. ## Sources of relevance labels - Labels can come from: - User behavior, such as clicks or skipped results. - Human evaluators assigning relevance scores. - LLMs directly judging query–document relevance. - Behavioral signals are useful but often sparse, biased by existing rankings, and unevenly distributed, so they work best as a supplement. - Human evaluators can provide comprehensive judgments across result sets, but labeling is expensive, difficult to scale, and vulnerable to inconsistency. - Humans also cannot directly review sensitive or proprietary customer data in this process, and different content types—such as Slack messages, Jira tickets, and Salesforce records—require different contextual expertise. ## LLM-assisted relevance evaluation - LLMs can evaluate far larger candidate sets at lower cost and with greater consistency than human annotators. - They can operate across languages and analyze customer content within established compliance boundaries. - Their judgments still depend on the model’s quality and the clarity of the evaluation prompt. - LLM-generated labels therefore require calibration and validation before being used for model training. ## Combining human review with LLM scale - Dropbox first creates a relatively small, high-quality dataset using human evaluators and limited, non-sensitive internal data. - These human labels are used to tune LLM prompts and model parameters. - Once the LLM meets quality thresholds, it generates hundreds of thousands or millions of relevance labels. - This approach multiplies human labeling effort by roughly 100 times, enabling broader and more representative training data. - LLMs are used offline rather than directly at query time because production-time use would introduce excessive latency and context-window limitations. - The LLM acts as a teacher for smaller, faster ranking models that can serve searches at scale. ## Evaluation as the foundation - Dash follows an iterative process: measure performance, change the model or instructions, and measure again. - The article compares this to chess engines, where the quality of the evaluation function determines which possible moves are preserved or discarded. - The same principle applies to ranking: poor relevance judgments can cause useful search-result patterns to be eliminated, while accurate judgments guide the model toward better rankings. Dash’s approach uses humans for quality control and contextual grounding, then uses LLMs to expand that expertise into large-scale training data. This hybrid strategy offers a practical way to improve enterprise search relevance without exposing customer data to human reviewers or imposing LLM latency on every search.

Read original(opens in new tab)
toss3 min readCurated summary

Easy-to-use Toss Front SDK

The post argues that an SDK’s stability depends not only on its internal implementation but also on how safely users can interact with it. Low-level APIs may expose every operation clearly, yet still allow human errors such as missing event handlers or cleanup. The recommended solution is an intent-driven Facade interface that simplifies common workflows, prevents misuse, and still provides low-level escape hatches for advanced cases. ## Designing an SDK That Is Easy to Use - Toss Place develops an external SDK for Toss Front payment terminals. - The SDK allows third-party developers to build plugin apps that integrate with Toss services and run on the terminal. - A simple-looking server API might require users to: - Open a server. - Register connection, message, and error handlers. - Remove handlers. - Close the server. - This approach exposes implicit responsibilities to SDK users: - A message callback might never be registered after a connection. - Handlers might not be removed before shutdown. - Improper cleanup can cause memory leaks and operational issues. - Therefore, third-party implementation mistakes can directly affect platform reliability. - A safer interface hides unnecessary internal steps: ```ts const server = await sdk.start({ onConnection, onMessage }); await server.stop(); ``` ## Facade as an Intent-Driven Interface - The Facade pattern is commonly described as wrapping a complex subsystem with a simpler interface. - In SDK design, its deeper purpose is to reorganize complexity around user intent rather than merely hide functionality. - Users should express goals such as: - “Start a server” - “Upload a file” - “Request a payment” - Internal concerns—including authentication, retries, state management, listener registration, and cleanup—should be handled by the SDK. - AWS CDK illustrates this distinction: - **L1 constructs** closely represent raw CloudFormation resources and provide fine-grained control. - **L2 constructs** provide intent-based APIs, such as creating a versioned S3 bucket with `versioned: true`, while handling the underlying configuration automatically. - The goal of a Facade is to reduce cognitive load and coupling, not simply to conceal every lower-level capability. ## Combining High-Level and Low-Level APIs - A well-designed SDK should provide both abstraction levels: - **High-level Facade:** Handles the roughly 80% of common use cases through complete workflows. - **Low-level APIs:** Serve as escape hatches for the roughly 20% of specialized cases requiring precise control. - In the example: - The Facade’s `start()` method opens the server, registers listeners, coordinates connections, and returns a unified server handle. - Low-level APIs separately expose operations such as `open`, `close`, `send`, `disconnect`, and event listeners. - This layered design improves immediate developer experience while preserving long-term compatibility and extensibility. ## Trade-offs and Escape Hatches - Higher-level abstractions inevitably reduce some flexibility. - Specialized requirements—such as keeping one connection while closing others—may not fit the Facade workflow. - As orchestration becomes more sophisticated, the SDK maintainers inherit additional implementation and maintenance costs. - Low-level escape hatches are therefore essential: users should be able to bypass the Facade when they need detailed control. ## Practical Recommendation Design SDK APIs around user intent and automate error-prone lifecycle management wherever possible. Offer a concise Facade for common workflows, but retain well-defined low-level interfaces so advanced users are not blocked by the abstraction.

Read original(opens in new tab)
gitlab3 min readCurated summary

Introducing the GitLab Managed Service Provider (MSP) Partner Program

GitLab has launched a global Managed Service Provider (MSP) Partner Program for qualified providers to deliver GitLab as a fully managed DevSecOps service. The program combines financial incentives, technical enablement, marketing support, and a structured onboarding process. It aims to help customers adopt and operate GitLab while enabling MSPs to build recurring, services-led revenue. ## Why the Program Matters - Many organizations lack the resources to deploy, administer, migrate, and continuously optimize a DevSecOps platform. - MSP partners can manage GitLab’s operational needs while development teams focus on building and delivering software. - The program provides formal requirements, enablement, dedicated support, and financial benefits for MSPs worldwide. - It addresses common customer challenges such as complex migrations, fragmented toolchains, and expanding security requirements. ## Benefits for MSP Partners - Partners receive standard GitLab margins plus an additional MSP premium on transactions, new business, and renewals. - MSPs retain all service fees from deployment, migration, training, enablement, and consulting. - Quarterly technical bootcamps cover releases, new features, best practices, roadmap updates, and peer experiences. - Recommended certifications include AWS Solutions Architect Associate and GCP Associate Cloud Engineer. - Go-to-market resources include: - A GitLab Certified MSP Partner badge - Co-brandable marketing assets - Eligibility for joint customer case studies - Partner Locator placement - Marketing Development Funds for qualified campaigns ## Customer Experience Customers receive a structured and repeatable managed DevSecOps service, including: - Documented implementation and migration methodologies - Platform deployment, administration, and ongoing support - Regular business reviews - Defined response and escalation procedures - Continuous platform optimization handled by the MSP ## Supporting AI Adoption GitLab MSPs can help organizations introduce AI-assisted development through the GitLab Duo Agent Platform. - MSPs can provide governance and operational guidance for AI adoption. - Customers can test AI workflows in controlled environments. - Managed services can help address data residency, compliance, and scaling requirements. - This approach reduces the burden on internal teams while enabling broader adoption of agentic AI. ## Who Should Apply The program is designed for MSPs that: - Already manage cloud, infrastructure, or application operations - Want to expand into DevSecOps services - Have, or plan to develop, relevant technical expertise - Prefer long-term customer relationships and recurring services revenue - Are existing GitLab Select or Professional Services Partners seeking a repeatable managed offering ## How to Get Started - Confirm business and technical eligibility using the program handbook. - Apply through the GitLab Partner Portal with supporting documentation. - Complete a structured 90-day onboarding process covering contracts, technical training, sales enablement, and an initial customer engagement. - Package the managed service, define SLAs, and launch the offering. - Applications are reviewed in approximately three business days. For MSPs with the necessary operational capabilities, the program offers a formal path to build a recurring GitLab managed services practice while helping customers adopt DevSecOps and AI more effectively.

Read original(opens in new tab)
gitlab2 min readCurated summary

Secure and fast deployments to Google Agent Engine with GitLab

Google Agent Engine provides a managed, scalable runtime for AI agents built with Google’s Agent Development Kit (ADK). The post shows how to deploy an ADK agent through GitLab using Workload Identity Federation, avoiding service-account keys while integrating security scanning into CI/CD. A GitLab pipeline can automatically test and deploy the agent to Agent Engine when changes reach the main branch. ## Agent Engine and GitLab - Agent Engine manages infrastructure, scaling, sessions, memory storage, logging, monitoring, and IAM. - GitLab simplifies deployment through: - Dependency scanning, SAST, and secret detection. - Native Google Cloud integration. - Keyless authentication with Workload Identity Federation. - CI/CD templates and the ADK deployment CLI. ## Prerequisites - A Google Cloud project with the Cloud Storage and Vertex AI APIs enabled. - A GitLab project containing the agent source code. - A Google Cloud Storage bucket for deployment staging. - GitLab’s Google Cloud IAM integration configured. ## Configure IAM with Workload Identity Federation - In GitLab, configure the Google Cloud IAM integration with: - Project ID - Project number - Workload Identity Pool ID - Provider ID - Run GitLab’s generated setup script in Google Cloud Shell. - Grant the federated service principal: - `roles/aiplatform.user` - `roles/storage.objectAdmin` - This setup lets GitLab authenticate to Google Cloud without storing long-lived service-account keys. ## Build the GitLab CI/CD Pipeline - Add a `.gitlab-ci.yml` file with `test` and `deploy` stages. - Use the `google/cloud-sdk:slim` image and define variables for: - Google Cloud project and region - Staging bucket - Agent name - Agent entry point - Include GitLab templates for: - Dependency scanning - Static application security testing - Secret detection - Enable keyless authentication with: ```yaml identity: google_cloud ``` - Install the ADK and required Google Cloud libraries during the job. - Deploy with: ```bash adk deploy agent_engine \ --project=$GCP_PROJECT_ID \ --region=$GCP_REGION \ --staging_bucket=gs://$STORAGE_BUCKET \ --display_name="$AGENT_NAME" \ $AGENT_ENTRY ``` - Restrict deployment to the `main` branch. - Cache Python dependencies to speed up later pipeline runs. ## Deploy and Verify - Commit the agent code and `.gitlab-ci.yml` to GitLab. - Monitor the pipeline under **Build > Pipelines**. - Confirm that security scans complete successfully before deployment. - The deployment stage packages the agent, places it in the staging bucket, and publishes it to Agent Engine. The recommended approach is to combine GitLab’s built-in security checks and Workload Identity Federation with the ADK CLI. This provides a secure, keyless, and repeatable deployment process for Google AI agents.

Read original(opens in new tab)
gitlab2 min readCurated summary

GitLab Duo Agent Platform with Claude accelerates development

GitLab Duo Agent Platform integrates external AI models such as Anthropic’s Claude and OpenAI’s Codex directly into GitLab workflows. Instead of operating as isolated coding assistants, these agents use project context and organizational standards to handle multi-step development tasks. The result is faster delivery, more consistent quality, and less manual work across the software development lifecycle. ## From an Idea to a Working Application - An agent can use an issue’s title and detailed requirements as the foundation for a complete application. - It analyzes project context and related assets, then generates: - Backend Java classes - Frontend HTML, CSS, and JavaScript - Business logic and UI components - Build configuration - The agent creates a merge request containing the implementation for developers to test and refine through natural-language interaction. ## Automated Code Review - Developers can mention the external agent in a merge request to request a review. - The review can cover: - Code strengths and critical issues - Medium- and low-priority improvements - Security risks - Testing gaps and code metrics - Recommendations and an approval status - This provides consistent review coverage while allowing senior developers to focus on architecture and complex decisions. ## Pipeline and Container Image Creation - When a project lacks CI/CD configuration, the agent can generate the required pipeline. - It creates a Dockerfile with a suitable base image for the project’s Java version. - The pipeline can: - Build the application - Build a Docker image - Push the image to GitLab’s container registry - The resulting workflow runs automatically through build, image creation, and deployment stages. ## Broader Impact on Development - External agents remain within GitLab, reducing context switching between development tools. - They can follow project-specific coding standards and understand broader repository context. - Teams can automate work from initial requirements through implementation, review, and deployment. - Developers spend less time on repetitive tasks while maintaining stronger consistency and quality. GitLab presents Duo Agent Platform as a way to turn external AI models into integrated development collaborators. Teams can use it to accelerate coding, automate reviews, and create deployment pipelines while keeping humans focused on validation, architecture, and innovation.

Read original(opens in new tab)
figma2 min readCurated summary

Building Frontend UIs with Codex and Figma | Figma Blog

Figma’s Codex integration creates a two-way workflow between coding and visual design. Using the Figma MCP server, developers can turn Figma designs into implementation context for Codex, then bring running interfaces back into editable Figma files. The result is a faster cycle for building, comparing, refining, and collaborating on frontend experiences. ## Starting an Application from a Design - Developers can select frames or nodes in Figma Design, Figma Make, or FigJam. - They copy a direct selection link by right-clicking a frame and choosing **Copy as → Copy link to selection**. - The link is provided to Codex with an implementation prompt, such as using existing design-system components. - Codex calls the MCP server’s `get_design_context` tool to retrieve: - Layout information - Styles and visual properties - Component details - Other design context needed for code generation - The MCP server supports additional tools and prompts for extracting information from Figma files. ## Bringing Code Back to the Canvas After iterating on the implementation, developers can import the live interface into Figma rather than recreating it manually. - The application must be rendered locally or on a publicly accessible web server. - Codex uses the `generate_figma_design` tool to convert the running UI into editable Figma frames. - Codex guides users through: 1. Creating or selecting a Figma file 2. Choosing a workspace 3. Setting up the application for capture 4. Opening the application in a browser session - The capture toolbar supports: - **Entire screen:** Captures the currently displayed screen - **Select element:** Captures a specific UI component - **Open file:** Opens the resulting Figma design for inspection ## Iterating Between Code and Design Once the interface is in Figma, teams can use the canvas to explore and refine the product. - Add design-system components. - Convert styles, fonts, and colors into variables. - Adjust layouts and add annotations. - Design interactions, empty states, and alternative flows. - Collaborate on multiple visual directions. - Send the refined design back to Codex through the same MCP workflow. The article presents this round trip as a continuous loop: design informs code, code produces a working interface, and the interface returns to Figma for further exploration. This lets teams begin from either a design or an implementation while preserving context and reducing the friction between developers and designers.

Read original(opens in new tab)
toss4 min readCurated summary

The Software 3.0

The post argues that teams using the same LLM can achieve very different results because individual knowledge of context engineering varies widely. Claude Code’s plugins and marketplace could help turn personal LLM techniques into shared, executable team workflows, raising the organization’s productivity floor. The author presents this as a forward-looking hypothesis rather than a proven success story. ## The Frictionless Harness - LLM adoption loses effectiveness when developers must switch between terminals, browsers, and chat tools. - Claude Code’s terminal-based TUI reduces context switching by combining natural-language instructions and code in the developer’s existing environment. - This low-friction experience makes it easier to distribute standardized workflows across a team. ## Executable Single Source of Truth - Wikis and Notion pages become outdated because they are designed primarily for human reading. - Claude Code plugins can serve as “executable SSOT”: - Humans can read them as guidelines and manuals. - LLMs can interpret them as precise system instructions. - Updating a plugin can immediately change how team agents behave, keeping operational knowledge aligned with current practices. ## Raising the Team’s Productivity Floor - Teams have significant differences in LLM literacy, independent of coding ability. - Generic open-source plugins can provide shared best practices, but they lack company- and domain-specific context. - Each domain needs its own rules for: - Tasks the AI can perform autonomously. - Tasks requiring human approval through HITL processes. - The goal is to minimize human intervention while preserving approval at critical points. ## Extending Platform Engineering into Software 3.0 - AI workflows resemble traditional internal platform components such as authentication, logging, and payment libraries. - The analogy is: - Common software modules → AI workflow plugins - Library distribution → Marketplace publishing - The implementation changes from traditional code to prompts and agent logic. - AI workflows should receive the same quality practices as software modules, including review, optimization, and feedback on token usage and failure cases. - Marketplace-based collaboration could turn individual prompting techniques into shared organizational intelligence. ## Why Use a Marketplace Instead of Only RAG? - RAG systems can make it difficult to predict which context will be retrieved due to search, reranking, and indexing behavior. - Plugins provide more explicit and controllable instructions and code. - Developers can modify and test workflows locally in the TUI without deploying a server. - With the Claude Agent SDK, workflows validated locally could also run in server environments, improving development-production parity. - The marketplace could become the shared source of truth between experimentation and production. ## Marketplace as a Workflow Distribution Platform - Teams could package coding conventions, Git strategies, lint rules, and testing policies into private plugins or registries. - Hooks could actively correct behavior rather than merely reject violations—for example, preventing commits on `main` and creating a `feature/` branch instead. - Slash commands could distribute the best engineer’s workflow to everyone: - `/new-feature` gathers requirements. - Creates a Jira issue and branch. - Produces an implementation plan for approval. - Implements the feature and opens a pull request. - This allows less experienced users to follow a reliable, high-quality process without reproducing it manually. ## Layered Context Architecture The author proposes separating plugin knowledge into three layers: - **Global layer:** Organization-wide security rules and coding standards. - **Domain layer:** Business-specific knowledge for areas such as payments, settlement, or membership. - **Local layer:** Repository-specific implementation details and conventions. This structure avoids overwhelming the LLM with irrelevant information and creates a “living knowledge base” made of maintainable prompts and code rather than static documents. ## The Data Flywheel Hypothesis - Standardized plugins could generate high-quality instruction-tuning data. - Accumulated workflow data might eventually support domain-specific model fine-tuning. - Existing workflows could also provide evaluation criteria for those models. - Success would require sustained data collection, quality controls, and long-term organizational investment. - The proposed flywheel is: more usage creates more data, better data improves models, and better models encourage further usage. The practical recommendation is to treat LLM expertise as an organizational system rather than an individual skill. Teams should begin packaging their implicit knowledge, approval rules, and proven workflows into versioned, domain-aware plugins that can be tested, reviewed, and distributed through a marketplace or private registry.

Read original(opens in new tab)
gitlab2 min readCurated summary

Passkeys now available for passwordless sign-in and 2FA on GitLab

GitLab now supports passkeys for passwordless sign-in and phishing-resistant two-factor authentication. Built on WebAuthn and public-key cryptography, passkeys let users authenticate with a fingerprint, face recognition, or device PIN while keeping the private key on their device. Users can register multiple passkeys across browsers, mobile devices, and FIDO2 security keys, improving both security and convenience. ## Passkeys for Sign-In and 2FA - Passkeys can be used: - As a passwordless login method. - As a phishing-resistant 2FA method. - For accounts with 2FA enabled, passkeys automatically become the default 2FA option. - Authentication uses a device fingerprint, facial recognition, or PIN. ## Registration and Compatibility - Users can register passkeys under **Profile settings > Account > Manage authentication**. - Supported platforms include: - Chrome, Firefox, Safari, and Edge. - iOS 16 and later. - Android 9 and later. - FIDO2 hardware security keys. - Multiple passkeys can be registered for access across different devices. ## WebAuthn Security Model - Passkeys rely on WebAuthn and public-key cryptography. - The private key remains securely stored on the user’s device and is never sent to GitLab. - GitLab stores only the public key. - A breach of GitLab’s stored credentials would not give attackers usable private keys for account access. ## GitLab’s Security Goals - Passkeys support GitLab’s commitment under the CISA Secure by Design Pledge. - They help increase MFA adoption while providing a smoother, phishing-resistant authentication experience. - GitLab invites users to provide feedback through its community and feedback channels. Users should register passkeys in their GitLab authentication settings, ideally across multiple trusted devices or security keys for both stronger protection and account recovery.

Read original(opens in new tab)
gitlab2 min readCurated summary

GitLab metrics and registry features help reduce CI/CD bottlenecks

GitLab’s two new beta features target common CI/CD bottlenecks without requiring additional third-party tools. CI/CD Job Performance Metrics provides job-level visibility into duration and failures, while Container Virtual Registry centralizes pulls from multiple registries through a cached GitLab endpoint. Together, they help platform teams identify pipeline problems faster and simplify container management. ## CI/CD Job Performance Metrics - Available in GitLab Premium and Ultimate. - Limited beta on GitLab.com; available on Self-Managed and Dedicated with ClickHouse configured. - Adds a job-focused panel to **Analyze > CI/CD analytics**. - Shows, for the previous 30 days by default: - Median (P50) and worst-case (P95) job duration - Failure rate - Job name and pipeline stage - Supports sorting, searching, and pagination to identify slow or unreliable jobs. - GitLab plans to add stage-level aggregation for build, test, and deploy bottlenecks. ## Container Virtual Registry - Available in GitLab Premium and Ultimate; API-ready in GitLab 18.9. - Provides one GitLab endpoint for pulling images from multiple upstream registries. - Supports registries such as Docker Hub, Harbor, Quay, and other sources using long-lived token authentication. - Uses pull-through caching to: - Reduce repeated downloads and bandwidth costs - Improve availability and reliability - Centralize authentication and registry configuration - Currently configured through the API, with UI management in development. - Cloud registries requiring IAM authentication, including Amazon ECR, Google Artifact Registry, and Azure Container Registry, may be supported later. ## Beta Access and Feedback - GitLab.com users can request access through their customer success manager or the feature’s feedback issue. - Self-managed users can enable the feature flag and configure the virtual registry through the API. - GitLab is seeking feedback to guide future improvements to both features. These betas are worth evaluating if your team needs better visibility into pipeline performance or manages images across several registries. The metrics feature can replace custom dashboards, while the virtual registry can reduce registry-related configuration and operational overhead.

Read original(opens in new tab)
meta3 min readCurated summary

RCCLX: Innovating GPU communications on AMD platforms

RCCLX is Meta’s open-source enhancement of RCCL for AMD GPUs, integrated with Torchcomms to support portable distributed AI workloads. It introduces Direct Data Access (DDA) and low-precision collectives, targeting communication bottlenecks in inference and training. On AMD MI300X systems, these optimizations deliver lower latency and higher throughput while maintaining acceptable accuracy. ## RCCLX and Torchcomms Integration - RCCLX is based on RCCL and tested on Meta’s internal workloads. - It integrates CTran transport technology for AMD platforms. - CTran enables features such as `AllToAllvDynamic`, a GPU-resident collective; additional CTran capabilities are planned for future releases. - Through Torchcomms, applications can use a common communication API across AMD, NVIDIA, and other backends without major code changes. - RCCLX is intended to achieve feature parity with Meta’s NCCLX backend for NVIDIA systems. ## Direct Data Access for Intra-Node Collectives - LLM inference has two distinct phases: - **Prefill** is compute-bound and generates the model’s key-value cache. - **Decoding** is memory-bound and generates tokens incrementally. - Tensor parallelism can make AllReduce responsible for up to 30% of end-to-end latency. - RCCLX introduces two DDA algorithms: - **DDA flat** lets each rank directly read other ranks’ memory and perform local reductions. It reduces latency from O(N) to O(1) for small messages by increasing data exchange from O(n) to O(n²). - **DDA tree** divides AllReduce into reduce-scatter and all-gather phases, retaining ring-like data movement while reducing latency for somewhat larger messages. - On AMD MI300X GPUs, DDA improves over RCCL by: - 10–50% for decode workloads. - 10–30% for prefill workloads. - Approximately 10% lower time-to-incremental-token. ## Low-Precision Collectives - RCCLX provides optimized low-precision versions of AllReduce, AllGather, AlltoAll, and ReduceScatter. - These target AMD Instinct MI300 and MI350 GPUs and support FP32 and BF16 inputs. - FP8 quantization provides up to 4:1 compression, reducing communication overhead for messages of at least 16 MB. - Parallel peer-to-peer mesh communication uses AMD Infinity Fabric for bandwidth and low latency. - Computation remains in FP32 to improve numerical stability. - Users can enable the feature with: ```bash RCCL_LOW_PRECISION_ENABLE=1 ``` - Internal evaluations showed: - About a 0.3% change on GSM8K accuracy evaluations. - 9–10% lower latency. - Approximately 7% higher throughput. - The current implementation is tuned for single-node deployments. ## Getting Started - Install Torchcomms with the RCCLX backend. - Create an RCCLX communicator through Torchcomms using the `"rcclx"` backend and a HIP device. - Existing Torchcomms operations such as `allreduce` can then run without backend-specific API changes. - Distributed initialization uses standard `torchrun` environment variables such as `MASTER_ADDR`, `MASTER_PORT`, `RANK`, and `WORLD_SIZE`. RCCLX is positioned as a practical way to improve AMD-based AI training and inference without requiring applications to adopt a new communication API. Teams can use DDA for lower inference latency and selectively enable low-precision collectives for higher throughput, while evaluating numerical accuracy for their own workloads.

Read original(opens in new tab)
cloudflare4 min readCurated summary

How we rebuilt Next.js with AI in one week

Vinext is an experimental, Vite-based reimplementation of Next.js built in one week by one engineer and an AI model. It preserves much of Next.js’s API and project structure while avoiding the fragile process of adapting Next.js/Turbopack output for serverless platforms. Early results suggest builds can be up to 4× faster, client bundles up to 57% smaller, and Cloudflare Workers deployment can be handled with a single command. ## The Deployment Challenges of Next.js - Next.js provides an excellent developer experience but relies on a bespoke build and deployment toolchain. - Deploying to platforms such as Cloudflare, Netlify, or AWS Lambda requires reshaping Next.js output. - OpenNext addresses this problem but must reverse-engineer build artifacts, making it vulnerable to changes between Next.js versions. - Next.js’s planned adapters API improves deployment support but does not solve the underlying Turbopack dependency. - `next dev` runs only in Node.js, making it difficult to develop against platform-specific APIs such as Durable Objects, KV, and AI bindings. ## Vinext’s Vite-Based Architecture - Vinext reimplements the Next.js API surface directly on Vite rather than wrapping or adapting Next.js output. - Existing `app/`, `pages/`, and `next.config.js` files can be reused. - Developers install it with `npm install vinext` and replace `next` scripts with `vinext`. - It supports: - Routing - Server-side rendering - React Server Components - Server actions - Caching - Middleware - Hot module replacement - Vite’s Environment API allows the output to run across different platforms. ## Early Performance Results - Benchmarks compared vinext with Next.js 16 using the same 33-route App Router application. - Type checking and ESLint were disabled for Next.js to focus on compilation and bundling. - Static pre-rendering was disabled with `force-dynamic` for a fairer comparison. - Early results showed: - Production builds up to 4× faster - Gzipped client bundles up to 57% smaller - The results measure build performance, not serving performance, and come from a single test application. - The authors describe the figures as directional because both vinext and its supporting tools are still evolving. - Vite’s architecture and the upcoming Rust-based Rolldown bundler are identified as major sources of potential performance gains. ## Cloudflare Workers Deployment - `vinext deploy` builds the application, generates Worker configuration, and deploys it automatically. - Both the App Router and Pages Router are supported. - Applications retain client-side hydration, interactive components, navigation, and React state. - A Cloudflare KV cache handler provides Incremental Static Regeneration: ```ts import { KVCacheHandler } from "vinext/cloudflare"; import { setCacheHandler } from "next/cache"; setCacheHandler(new KVCacheHandler(env.MY_KV_NAMESPACE)); ``` - The cache layer is pluggable, allowing alternatives such as R2 or future Cache API improvements. - Because development and deployment can both run in `workerd`, applications can use Durable Objects, AI bindings, and other Cloudflare services without Node.js compatibility workarounds. ## Broader Ecosystem Potential - Although Cloudflare Workers is the initial target, roughly 95% of vinext is platform-independent Vite code. - Its routing, SSR pipeline, module shims, and React Server Components integration are not Cloudflare-specific. - A proof of concept reportedly ran on Vercel in under 30 minutes. - The project is open source and invites other hosting providers to contribute deployment targets. ## Experimental Status - Vinext is less than a week old and has not been tested under meaningful production-scale traffic. - The authors recommend caution before adopting it for critical applications. - Its test suite already includes more than 1,700 Vitest tests and 380 Playwright end-to-end tests, including tests ported from Next.js and OpenNext. - The project reportedly cost approximately $1,100 in AI-token usage to build. Vinext is best viewed as a promising experimental alternative rather than a drop-in replacement ready for every production workload. Teams interested in platform-native development and faster Vite-based builds can evaluate it carefully, while waiting for broader compatibility and real-world validation.

Read original(opens in new tab)