Unit Testing

3 posts

cloudflare3 min readCurated summary

How we built a software factory to drive Astro’s GitHub issue count to zero

AI-powered software factories can address a pressing open-source problem: maintainers are overwhelmed by the flood of AI-generated issues, pull requests, and security reports. The Astro team built an automated triage pipeline that reproduces bugs, diagnoses causes, creates fixes, and ships preview releases for verification. After several months, it reduced Astro’s open issues from more than 200 to roughly 30 without mass-closing or ignoring reports. ## Building an Issue-Triage Skill - The team began by automating issue triage, one of the most time-consuming parts of open-source maintenance. - The workflow mirrors manual debugging: - **Reproduce:** Clone the reporter’s reproduction repository and confirm the problem. - **Diagnose:** Instrument the code and add logging to identify the root cause. - **Verify:** Check tests, documentation, and comments to determine whether the behavior is actually a bug. - **Fix:** Turn the reproduction into failing tests, implement a solution, and deploy it. - Each phase runs in an isolated AI subagent to reduce the tendency to force a solution. - Subagents communicate through a sequential `report.md` file containing their findings. ## Running the Pipeline in GitHub Actions - The workflow is driven by GitHub issue labels rather than a separate internal database. - New issues begin with `triage needed`; verified fixes eventually move to `fix verified`. - The pipeline reconstructs its state from labels and existing issue comments. - When a fix is ready, it: - Creates a preview release using `pkg.pr.new`. - Posts the diagnosis, logs, and installation instructions to the issue. - Lets the original reporter test the patch. - Opens a linked pull request after confirmation. ## From a Repository Workflow to Flue - The team recognized that the process was not inherently tied to GitHub. - Its core structure consists of: - An external event. - A sequence of isolated subagents. - Separate reasoning and execution permissions. - Durable workflow state. - This generalization became **Flue**, an open, platform-agnostic framework for agent workflows that can respond to GitHub events, Slack messages, cron jobs, or webhooks. ## Effects on Maintainer and Community Work - Automation did not make the Astro team less connected to users. - Instead, it freed maintainers to spend more time: - Engaging with the community in Discord. - Participating in RFCs and feature discussions. - Collaborating with contributors. - The system is designed to resolve most incoming issues, while failures are treated as signals that the codebase needs improvement. ## Using Agent Failures to Improve the Codebase Agent mistakes often reveal problems that would also challenge human developers: - **Opaque abstractions:** Component boundaries are unclear. - **Missing documentation:** Important implementation decisions are unexplained. - **Insufficient testing:** Critical behavior lacks adequate unit tests. - For example, the bot repeatedly changed an HMR-related condition and caused regressions because the logic was poorly documented and under-tested. - Adding a precise comment clarified the intended behavior, after which the bot stopped making the same incorrect change. - Fixing these weaknesses improves both future automation and human maintainability. ## Extracting the Workflow into a GitHub Action - Initially, the triage system was embedded in the Astro monorepo, making changes risky and difficult to test. - The team separated it into the standalone `triagebot-action` repository. - This enabled independent testing and safer updates to Flue and the workflow. - The action now supports Astro and has been adopted or forked by other teams building their own automated development pipelines. The practical lesson is to start with a narrow, repeatable maintenance task, isolate agent responsibilities, make all reasoning auditable, and use failures to improve documentation, architecture, and tests.

Read original(opens in new tab)
spotify3 min readCurated summary

Congratulations to the recipients of the 2025 Spotify FOSS Fund | Spotify Engineering

Spotify’s 2025 FOSS Fund provided €30,000 to FFmpeg, €15,000 to Mock Service Worker (MSW), and support to the Xiph.Org Foundation. The grants recognize both large, foundational projects and smaller efforts maintained by individuals, emphasizing that Spotify’s technology depends on volunteer-driven open source. Maintainers describe the funds as valuable for development, infrastructure, contributor support, and long-term sustainability. ## Spotify’s FOSS Fund - Established in 2022 to support open source projects that Spotify relies on and its developers value. - The 2025 recipients were: - FFmpeg - Mock Service Worker (MSW) - Xiph.Org Foundation - FFmpeg and Xiph.Org support technologies fundamental to streaming, while MSW is primarily maintained by its creator. - Despite differences in scale, all three projects depend heavily on volunteer contributions. ## FFmpeg: €30,000 - FFmpeg has been actively developed for more than 25 years and supports: - Multimedia encoding and decoding - Transcoding - Muxing and demuxing - Filtering - Streaming and playback - Its goal is to play every multimedia file ever created, making it widely used by hobbyists and companies such as YouTube, Netflix, X, and Spotify. - Maintainer Kieran Kunhya said the project is written almost entirely by volunteers. - The funding will support: - Hardware purchases - Conference travel - Development projects - FFmpeg’s maintainers argue that company donations also raise awareness of the project’s importance and demonstrate responsibility toward critical dependencies. - They identify reliable, long-term funding and dedicated company developers as major unmet needs in open source sustainability. ## Mock Service Worker: €15,000 - MSW is a JavaScript library for API mocking used by Spotify for unit testing and across its web stack. - Creator Artem Zakharchenko is the project’s sole full-time maintainer, supported by a small group of active contributors. - Major 2025 improvements included: - A new Interceptors architecture - Custom mock agents operating at the Node.js `node:net` level - First-class Server-Sent Events support - Planned work for 2026 includes: - A redesigned internal architecture - A new approach to request interception - Remote request interception, a technically difficult feature under development for several years - The previous Spotify grant primarily helped Artem support himself while working full time on open source. - The new funding will directly support continued MSW development and efforts to make external contributions more financially sustainable. - Artem notes that many contributors have been reluctant to accept payment, creating an ongoing challenge in building a sustainable maintainer and contributor community. ## Xiph.Org Foundation - The post identifies the Xiph.Org Foundation as a 2025 recipient and groups it with FFmpeg as a project representing technology fundamental to streaming. - The provided text does not include further details about the foundation’s grant amount, projects, or maintainer comments. Spotify’s grants illustrate how companies can support both major infrastructure projects and individual maintainers. Recurring funding, paid development time, contributor compensation, and practical resources such as hardware and conference travel remain important to the long-term health of open source.

Read original(opens in new tab)
woowahanOriginal article

Test Automation with AI: Plugin Development Story (opens in new tab)

This blog post explores how a development team at Woowahan Tech successfully automated the creation of 100 unit tests in just 30 minutes by combining a custom IntelliJ plugin with Amazon Q. The author argues that while full AI automation often fails in complex multi-module environments, a hybrid approach using "compile-guaranteed templates" ensures high success rates and maintains operational stability. This strategy allows developers to bypass repetitive setup tasks while leveraging AI for logic implementation within a strictly defined, valid structure. ### Evaluating AI Assistants for Testing * The team compared various AI tools including GitHub Copilot, Cursor, and Amazon Q to determine which best fit their existing IntelliJ-based workflow. * Amazon Q was selected for its superior understanding of the entire project context and its ability to integrate seamlessly as a plugin without requiring a switch to a new IDE. * Initial manual use of AI assistants highlighted repetitive patterns: developers had to constantly specify team conventions (Kotest FunSpec, MockK) and manually fix build errors in 15% of the generated code. * On average, it took 10 minutes per class to generate and refine tests manually, prompting the team to seek a more automated solution via a custom plugin. ### The Pitfalls of Full Automation * The first version of the custom plugin attempted to generate complete test files by gathering class metadata through PSI (Program Structure Interface) and sending it to the Gemini API. * Pilot tests revealed a 90% compilation failure rate, as the AI frequently generated incorrect imports, hallucinated non-existent fields, or used mismatched data types. * A critical issue was the "loss of existing tests," where the AI-generated output would completely overwrite previous work rather than appending to it. * In complex multi-module projects, the AI struggled to identify the correct classes when multiple modules contained identical class names, leading to significant manual correction time. ### Shifting to Compile-Guaranteed Templates * To overcome the limitations of full automation, the team pivoted to a "template first" approach where the plugin generates a valid, compilable shell for the test. * The plugin handles the complex infrastructure of the test file, including correct imports, MockK setups, and empty test stubs for every method in the target class. * This approach reduces the AI's "hallucination surface" by providing it with a predefined structure, allowing tools like Amazon Q to focus solely on filling in the implementation details. * By automating the 1-minute setup and letting the AI handle the 2-minute implementation phase, the team achieved a 97% success rate across 100 test cases. ### Practical Conclusion For teams looking to improve test coverage in large-scale repositories, the most effective strategy is to use IDE plugins to automate context gathering and boilerplate generation. By providing the AI with a structurally sound template, developers can eliminate compilation errors and significantly reduce the time spent on manual refinement, ensuring that even complex edge cases are covered with minimal effort.