Techlist.io - Korean Tech Blog Curator

meta4 min readCurated summary

Ranking Engineer Agent (REA): The Autonomous AI Agent Accelerating Meta’s Ads Ranking Innovation

Meta’s Ranking Engineer Agent (REA) autonomously manages much of the ads-ranking ML experimentation lifecycle, from generating hypotheses and launching training jobs to debugging failures and analyzing results. Unlike session-based AI assistants, REA maintains context across workflows lasting days or weeks, while engineers retain oversight at strategic checkpoints. In its first production rollout, REA doubled average model accuracy across six models and helped three engineers produce launch proposals for eight models—about five times the historical engineering output. ## The Bottleneck in Traditional ML Experimentation - Meta’s advertising systems rely on large, complex ML models serving billions of users across Facebook, Instagram, Messenger, and WhatsApp. - Improving these models traditionally requires engineers to: - Form hypotheses - Design experiments - Launch training jobs - Debug failures - Analyze results - Iterate on promising approaches - Each cycle can take days or weeks, and mature models make meaningful improvements increasingly difficult to find. - The sequential, hands-on process became a bottleneck to experimentation and innovation. ## REA as an Autonomous ML Agent - Existing ML AI tools generally assist with isolated tasks such as drafting hypotheses, writing configurations, or interpreting logs. - REA instead coordinates the full experimentation process and advances it without continuous prompting. - Its design addresses three central challenges: - **Long-running workflows:** Persistent state and memory allow REA to manage multiday or multiweek experiments. - **Hypothesis quality:** It combines historical experiment data with current ML research. - **Operational resilience:** It handles failures and compute limits within engineer-approved safeguards. ## Hibernate-and-Wake Workflow Management - Training jobs may run for hours or days, so REA delegates waiting to a background system. - It hibernates to conserve resources and automatically wakes when jobs finish. - This lets it preserve context and continue experiments without constant human supervision. - REA is built on Meta’s Confucius agent framework, which provides: - Code-generation capabilities - Integration with job schedulers - Experiment tracking - Codebase navigation tools ## Dual-Source Hypothesis Generation - REA draws ideas from two systems: - **Historical Insights Database:** A repository of previous experiments, successes, and failures used for pattern recognition and in-context learning. - **ML Research Agent:** A research component that examines baseline configurations and proposes new optimization strategies. - Combining these sources produces configurations that may not emerge from either source alone. - Some of REA’s strongest improvements resulted from combining model architecture changes with training-efficiency techniques. ## Three-Phase Experiment Planning - Before running experiments, REA proposes an exploration plan, estimates GPU costs, and obtains engineer confirmation. - Its typical strategy includes: - **Validation:** Test individual hypotheses in parallel to establish baselines. - **Combination:** Combine promising ideas to identify synergistic effects. - **Exploitation:** Intensively optimize the strongest candidates within the approved compute budget. ## Autonomous Failure Handling and Safeguards - REA adapts to infrastructure problems, unexpected errors, poor results, and compute constraints without escalating every issue to an engineer. - It uses runbooks and diagnostic reasoning to: - Exclude jobs with clear out-of-memory failures - Detect training instability, such as exploding losses - Debug preliminary infrastructure failures - Reprioritize experiments when results are weak - Its autonomy is constrained by: - Access limited to Meta’s ads-ranking codebase - Explicit engineer approval through preflight reviews - Confirmed GPU budgets - Automatic pausing or stopping when thresholds are reached ## Production Results - Across six models, REA-driven iterations achieved approximately **2× the average model accuracy** compared with baseline. - Three engineers produced proposals to launch improvements for eight models. - Historically, that volume of work would have required roughly two engineers per model, resulting in about **5× greater engineering output** with REA. REA demonstrates that autonomous agents can improve ML experimentation by handling long-running execution, generating broader hypotheses, and recovering from routine failures. The most effective deployment model combines substantial agent autonomy with explicit compute limits, codebase restrictions, and human review at major strategic decisions.

Read original(opens in new tab)
dropbox3 min readCurated summary

How we optimized Dash's relevance judge with DSPy

Dropbox Dash needed a relevance judge that could score query–document pairs accurately, cheaply, and reliably at scale. Its original judge used OpenAI’s o3, but the cost made it impractical for large-scale labeling, while its prompt performed poorly when moved to the cheaper gpt-oss-120b model. Dropbox used DSPy’s GEPA optimizer to turn prompt tuning into a measurable feedback loop, improving alignment with human judgments while preserving production-ready output formatting. ## Measuring Agreement with Human Reviewers - The judge rates each query–document pair on a 1–5 relevance scale: - **5** means a perfect match. - **1** means no meaningful connection to the query or user intent. - Human annotators provide both: - A relevance score. - A short explanation for their judgment. - Dropbox evaluates the model with normalized mean squared error (NMSE): - It measures the squared difference between model and human ratings. - Scores are normalized to a 0–100 scale. - **0** represents perfect agreement; higher values indicate worse performance. - Invalid JSON or incorrectly structured responses are treated as fully incorrect because they cannot be consumed reliably by downstream systems. - The optimization objective is therefore twofold: - Minimize disagreement with human ratings. - Ensure consistently parseable, production-ready outputs. ## Moving from o3 to a Lower-Cost Model - The original judge used OpenAI’s o3 because it delivered strong agreement with human ratings. - Running o3 across orders of magnitude more query–document pairs was too expensive. - Dropbox selected **gpt-oss-120b**, an open-weight model offering a better cost-performance balance. - The carefully tuned o3 prompt did not transfer directly: - Relevance quality declined under the NMSE metric. - Manual prompt rewriting would have required extensive iteration and regression testing. ## DSPy and GEPA-Based Prompt Optimization - Dropbox defined the optimization problem using: - A fixed relevance-rating task. - Human-annotated examples. - NMSE as the evaluation metric. - DSPy’s **GEPA optimizer** iteratively improves prompts for a specific target model. - Instead of relying only on an aggregate score, GEPA analyzes individual disagreements and generates structured feedback. - Feedback combines: - The difference and direction between predicted and human ratings. - The human annotator’s explanation. - The model’s reasoning. - DSPy then uses a reflection loop: - Evaluate the current prompt. - Identify recurring failure modes. - Revise the prompt with generalizable rules. - Repeat the process against the human-alignment metric. - This approach can address systematic errors such as: - Overvaluing keyword overlap. - Undervaluing document recency. - Misinterpreting user intent. - The feedback explicitly discourages overfitting to individual examples and preserves core task constraints, including the 1–5 rating range. Dropbox’s experience suggests that relevance judges should be optimized systematically rather than tuned manually. Defining a clear human-alignment metric, including structural validity, allows DSPy to adapt prompts across models while reducing cost and limiting regressions.

Read original(opens in new tab)
github3 min readCurated summary

Investing in the people shaping open source and securing the future together

Open source security depends on supporting the maintainers who sustain critical software, not merely hosting their code. GitHub argues that funding, education, practical security tools, and AI assistance can reduce maintainer burnout while improving the broader software supply chain. Its new commitments focus on making security work more manageable as AI accelerates both vulnerability discovery and attacks. ## A $12.5 Million Open Source Security Commitment - GitHub is joining Anthropic, AWS, Google, and OpenAI in committing $12.5 million to the Linux Foundation’s Alpha-Omega initiative. - The funding will help integrate emerging AI security capabilities into existing open source workflows. - The effort builds on GitHub’s broader role as a provider of security tools, education, and long-term maintainer support. ## Expanding Maintainer Resources - More than 280,000 GitHub maintainers are eligible for free access to: - Core GitHub services - GitHub Copilot Pro - GitHub Actions - Code scanning and Autofix - Secret scanning and push protection - Dependency alerts - GitHub’s Secure Open Source Fund is adding $5.5 million in Azure credits and funding for training, expertise, community support, and new partners such as Datadog, Open WebUI, the Atlantic Council, and OWASP. - GitHub Security Lab is improving security advisories and Private Vulnerability Reporting to reduce low-quality reports and ease the burden on maintainers. ## Results from Security-Focused Funding - Previous Secure Open Source Fund programs supported 138 projects and more than 200 maintainers across 38 countries. - Participating projects produced: - 191 new CVEs - More than 250 prevented secret leaks - More than 600 detected and resolved leaked secrets - These projects collectively affect billions of monthly software downloads. - GitHub concludes that security improves when maintainers receive dedicated time, funding, education, and tools that fit naturally into their workflows. ## Using AI to Reduce Maintainer Burden - AI has increased the speed and scale of vulnerability discovery for both attackers and defenders. - Maintainers are facing more automated pull requests and security reports, often with poor signal-to-noise ratios, contributing to burnout. - GitHub’s goal is to use AI for triage, pull request review, vulnerability identification, and remediation—not simply to generate more findings. - GitHub has open sourced an AI-powered security research framework so maintainers, rather than only specialized security teams, can benefit from it. - Copilot Pro provides eligible maintainers with AI-assisted code review, agentic security remediation workflows, and access to multiple leading models. GitHub’s overall recommendation is to treat AI as a force multiplier and pair it with sustained funding, education, and workflow-integrated security tools. Supporting maintainers directly is presented as the most effective way to protect the wider software ecosystem.

Read original(opens in new tab)
line3 min readCurated summary

Integration of LINE App’s Multi-party Chat Features

LINE is consolidating its two multi-person chat types—temporary “Rooms” and long-term “Groups”—into a single Group Chat model. The change aims to simplify the user experience, make all chat features available everywhere, and reduce duplicated server and client resources. A gradual migration strategy is being used to avoid disruption. ## Two Original Chat Models - **Rooms** were designed for temporary conversations: - No room name was required. - Invited friends joined immediately without approval. - Features such as albums and notes were unavailable. - **Groups** were designed for long-term communities: - They had names and supported features such as group albums and notes. - Invitees had to accept or reject invitations before joining. - Users often created Rooms without realizing their limitations, then created a new Group later when they needed additional features. ## Reasons for Unification - Users found the distinction between Rooms and Groups difficult to understand. - Existing conversations could not be converted from Rooms into Groups, forcing users to abandon their conversation history. - Users frequently created multiple chats with the same members, causing: - Cluttered conversation lists. - Unnecessary data accumulation on servers. - Increased client and server resource usage. - The unified model standardizes behavior and features while retaining flexibility in how invitations work. ## Migrating Groups to Group Chats - LINE introduced new Group Chat APIs and used **dual reads** to maintain compatibility with existing Group APIs and storage. - The migration proceeded gradually: 1. The new API initially read Group data through a routing layer. 2. The number of Group Chats was progressively increased. 3. Eventually, only Group Chats were created. - Batch processing migrated all existing Group data. - After migration, LINE stopped dual reads and relied exclusively on the Group Chat model. ## Differences Between Rooms and Groups ### Invitation Mechanisms - Groups required invitees to explicitly accept or reject an invitation. - Rooms added people immediately when they were invited. - The unified creation flow lets users choose whether invitees should join immediately or confirm participation first. ### Feature Availability - Rooms lacked many Group features because they were intended to be temporary. - The new model is based on the Group architecture, so all newly created conversations support the full feature set, including future features. ## Improving Conversation Discovery - Users often created a new chat with the same participants instead of finding an older, inactive conversation in a long chat list. - The new creation workflow displays a hint when an equivalent existing conversation is found. - Users can then return to the existing conversation, reducing duplicate rooms and improving navigation. ## Migration Plans for Existing Rooms - Conversations created in current LINE versions are already Group Chats. - Groups created with older app versions are being converted server-side. - The remaining objective is to migrate existing Rooms so their participants can use the complete set of Group Chat features. The project is a long-term effort designed to minimize disruption while improving consistency and efficiency. Duplicate conversations with identical participants fell from 15% for Rooms to 0.78% for invitation-free Group Chats, demonstrating the practical impact of the consolidation.

Read original(opens in new tab)
discord3 min readCurated summary

How ROOST is Advancing Online Safety

Discord argues that online safety should be built through shared, open-source infrastructure rather than isolated corporate systems. Its donated rules engine, Osprey, lets platforms detect suspicious behavior and harmful activity in real time, while ROOST develops and maintains tools for broad industry adoption. Early adoption, including by Bluesky, suggests this model can raise baseline safety standards across the internet. ## The Need for Shared Safety Tools - Nearly 100 million people use Discord daily, generating hundreds of millions of events that must be evaluated for threats. - Generative AI has increased the scale and sophistication of phishing, deepfakes, and coordinated abuse. - Smaller platforms often lack the resources to build effective trust-and-safety systems from scratch. - ROOST aims to make proven safety technologies open, shared, and auditable. ## How Osprey Works - Osprey is a real-time rules engine for event processing and behavioral analysis. - It can evaluate logins, messages, account creation, content posts, and platform-specific actions. - Safety teams write rules in a simple language and deploy them without engineering dependencies. - The engine produces transparent decisions indicating whether activity is safe, suspicious, or malicious. - Discord runs thousands of rules across hundreds of action types. - Investigation findings feed new rules, while enforcement generates additional signals for future detection. - The open-source release is based on Discord’s production system rather than a reduced version; improvements from ROOST were later reintegrated into Discord. ## ROOST’s Collaborative Model - ROOST builds on earlier cross-industry efforts such as image hashing for child-safety work, the Tech Coalition’s Lantern program, GIFCT incident response, and shared ISO safety standards. - Unlike organizations that primarily steward open-source projects, ROOST also develops and maintains a suite of public-interest safety tools. - Its projects include Osprey and Coop, a comprehensive review tool. - Open-source tools can raise the minimum level of protection available to smaller platforms and reduce the spread of threats across services. - The model also enables companies to build managed services around free tools, similar to businesses built around Linux. ## Adoption and Industry Impact - Musubi announced a managed Coop offering, while Zentropi integrated its labeling engine with Coop. - Osprey v1 was introduced at FOSDEM, prompting collaboration among engineers from multiple organizations and protocols. - Platforms such as Bluesky are already using Osprey. - More than 360 million users across participating platforms are now covered by open-source safety tooling. - ROOST continues development through public contributor and adopter working-group meetings held every two weeks. ROOST’s approach suggests that open, production-grade safety infrastructure can help platforms respond faster to emerging threats while creating a shared foundation for industry-wide improvement.

Read original(opens in new tab)
toss3 min readCurated summary

Embracing the Software 3.0 Era

Software 3.0 replaces hand-written rules with natural-language instructions to LLMs, but models alone cannot reliably perform real-world work. The missing piece is the harness: tools, context, and environments that connect an LLM to codebases, commands, databases, and users. Claude Code illustrates how familiar Software 1.0 architecture can guide agent design while adding a new capability—asking humans for judgment when uncertainty arises. ## From Software 1.0 to Software 3.0 - **Software 1.0:** Developers explicitly write logic using languages such as Python, Java, or C++. - **Software 2.0:** Data and training produce neural-network weights that function as the program. - **Software 3.0:** Prompts and natural-language instructions direct LLM behavior. - Karpathy’s central claim is that Software 3.0 is increasingly absorbing both traditional code and trained models. ## Harnesses Make LLMs Useful - A raw LLM cannot independently read a codebase, execute commands, modify files, or access databases. - A **harness** supplies the tools and environment needed to turn model capability into practical work. - Claude Code is presented as a harness for Claude: it transforms a language model into an agent capable of completing and shipping tasks. ## Mapping Agent Concepts to Layered Architecture The terminology of agent systems can be understood through familiar Software 1.0 design patterns: - **Slash commands → Controllers** - They serve as entry points for user requests, such as `/review` or `/refactor`. - **Sub-agents → Service layer** - They coordinate multiple skills to complete a workflow. - Each sub-agent has an independent context and acts as a self-contained unit of work. - **Skills → Domain components** - Each skill should have one focused responsibility, such as reviewing code, generating tests, or writing documentation. - **MCP → Infrastructure or adapters** - MCP provides abstraction boundaries for external systems such as APIs and databases. - **CLAUDE.md → Project constitution** - It records stable project information: technology choices, conventions, and build commands. - Frequently changing task details should be provided through the conversation or injected into an agent’s context instead. ## Agent Design Has Familiar Anti-Patterns Traditional code smells also apply to agent systems: - **Feature Envy:** A skill relies excessively on another skill’s data. - **Duplication:** Prompts are copied across multiple skills. - **Long Method:** A single sub-agent performs an overly long sequence of many skills. - Clear boundaries, single responsibility, and limited coupling remain valuable. ## The Difference: Agents Can Ask Humans Layered architecture generally requires every failure and edge case to be handled through predefined exceptions, policies, or branches. - Traditional code must decide what to do when an unusual case occurs. - An agent using human-in-the-loop interaction can pause and ask the user for clarification. - In this model, exceptions become questions, allowing the agent to continue after receiving a decision. Agents should ask when: - An action is difficult to reverse, such as deletion or deployment. - Several valid options exist without a clear best choice. - The decision has significant consequences. They should proceed automatically when: - The operation is safely repeatable. - Existing conventions provide a clear answer. - The action is easy to undo. ## What Carries Forward into Software 3.0 The new paradigm does not make established engineering practices irrelevant. - Move away from explicitly coding every possible rule and edge case. - Do not reduce LLMs to simple autocomplete tools. - Preserve layered design, single responsibility, abstraction, dependency management, and interface design. - Continue emphasizing testability, debugging, code review, and iterative improvement. The practical approach is to combine Software 3.0’s flexible reasoning with Software 1.0’s architecture and engineering discipline, while giving agents a clear way to involve humans when decisions require judgment.

Read original(opens in new tab)
figma2 min readCurated summary

How We Rebuilt the Foundations of Component Instances | Figma Blog

Figma replaced its decade-old Instance Updater with a reactive foundation called Materializer. The change separates component resolution from layout and variable evaluation, while enabling granular updates instead of recalculating entire instance trees. Figma reports that common operations in large design systems are now up to 50% faster, and the architecture can support other dynamic features beyond components. ## Why the Original Architecture No Longer Scaled - Instance Updater was introduced around 2016 to resolve component properties, manage instance structure, and synchronize instances with their main components. - Over time, Figma added features such as: - Auto layout in 2019 - Variants in 2020 - Component properties in 2022 - Variables in 2023 - Slots, introduced as an open beta feature in 2025 - Modern instances can combine variants, variable bindings, auto layout, nested instances, variable modes, and extensive overrides. - A small edit can therefore propagate through deeply nested trees and trigger widespread recalculation. - Instance Updater accumulated specialized logic for layout, variables, and other integrations, making it increasingly fragile and slow. - In extreme cases, systems repeatedly invalidated one another, causing actions such as instance swaps or property changes to take seconds. ## The Goals of the Rewrite - Figma concluded that incremental optimization would not solve the underlying architectural problems. - The new design aims to separate responsibilities: - The instance system resolves which properties and children an instance should have. - Layout systems handle layout calculations. - Variable systems handle variable evaluation. - Updates should be granular, affecting only the portions of a tree that actually changed rather than rebuilding entire instances. - Figma also wanted reusable infrastructure for dependency tracking, reactive updates, and efficient invalidation across the editor. ## Materializer and Derived Subtrees - Figma built **Materializer**, a generic system that operates on the document tree. - Its purpose is to create and maintain **derived subtrees**—structures whose properties and hierarchy are computed from other sources of truth. - Component instances are one use case, but the same model can support features such as rich text nodes whose content is synchronized with an external CMS. - This broader abstraction allows teams to build dynamic features without embedding bespoke update logic into a single specialized runtime. The rewrite turns component instances from a monolithic synchronization problem into one application of a more general reactive system. By isolating responsibilities and invalidating only affected subtrees, Figma improves performance for complex design systems while creating a foundation for future dynamic features.

Read original(opens in new tab)
figma2 min readCurated summary

Design’s Influence Is Expanding, and Here’s Why That Feels Hard | Figma Blog

Design is expanding into more products, interactions, and strategic decisions, especially as AI introduces new software categories and interfaces. Although AI makes design work faster, it also increases output, expectations, and workload rather than reducing effort. This leaves designers divided: the field is growing, but many are unsure whether it is improving. ## Design’s Expanding Influence - Each technological shift—from graphical interfaces to the web and mobile apps—has increased design’s scope. - AI is creating new categories such as agent orchestration systems and answer engines. - Existing products are gaining generative, conversational, and predictive features. - Users now interact through prompts, speech, and image uploads, creating new design challenges: - Translating ambiguous input into clear intent - Making automated experiences understandable and human - Designing beyond traditional screen-by-screen navigation - Survey results show mixed sentiment: - 36% of designers think the profession has improved - 35% think it has worsened - 29% see no change - Meanwhile, 82% of hiring managers say demand for designers has increased or remained steady, though only 20% believe the industry itself is improving. ## AI Expands the Work - AI helps teams address new design problems more quickly, but it does not necessarily reduce the amount of work. - Product builders reported a 17.5% year-over-year increase in the number of tasks they perform. - Research from UC Berkeley found that AI users work faster while also taking on more tasks and working longer hours. - Workers often feel more productive without feeling less busy. ## The Jevons Paradox in Design - As AI makes creation cheaper and easier, teams produce more designs, explore more options, and iterate more deeply. - This follows the Jevons Paradox: efficiency increases can lead to greater overall consumption rather than reduced consumption. - Software development experienced a similar pattern when cloud infrastructure made releases easier, resulting in more frequent releases and redesigns. - AI has changed the rhythm and volume of design work rather than eliminating it. Designers should view AI as a force multiplier, not a shortcut to less work. Its benefits will depend on managing rising expectations and workload while developing clearer approaches to complex, automated interactions.

Read original(opens in new tab)
google3 min readCurated summary

Improving breast cancer screening workflows with machine learning

Google Research’s AIMS studies evaluated whether machine learning could support the UK’s mammography double-reading workflow. Across five NHS screening services, the AI system improved cancer detection sensitivity without reducing specificity, detected some cancers missed by human readers, and processed cases far faster. The studies also showed that safe deployment requires local calibration, monitoring for distribution shifts, and evaluation of how clinicians interact with AI results. ## NHS Screening Challenges - The UK NHS uses two human readers for each mammogram, with arbitration when their assessments require review. - A projected shortage of clinical radiologists—currently around 30% and expected to reach 40% by 2028—threatens the sustainability of this model. - AI could help increase detection while reducing pressure on radiology services. ## Study 1: Standalone Performance - The retrospective evaluation included mammograms from approximately 116,000 women screened across five NHS services. - The services represented three different double-reading and arbitration workflows. - AI thresholds were calibrated separately for each service to account for local populations and procedures. - Performance was measured against the original first reader using a 39-month follow-up period, including interval and subsequent-round cancers. - Researchers also assessed: - Comparisons with second and consensus readers - Lesion-level localization - Performance across demographic groups ## Study 1: Results - Cancer detection increased from 7.54 to 9.33 cases per 1,000 women. - The AI system achieved significantly higher sensitivity than the original first reader without compromising specificity. - It detected 25% of interval cancers missed by the original double-reading process. - Performance was especially strong for invasive cancers and women attending their first screening. - The study found no notable systematic disparities by age, ethnicity, breast density, or socioeconomic status. ## Prospective Technical Deployment - The system was deployed non-interventionally at 12 sites across two London screening services. - It processed 9,266 cases over roughly two months per service. - Mammograms were pseudonymized and sent to a secure Google Cloud-based system. - Median AI processing time was 17.7 minutes, compared with more than two days for the first human read. - The deployment detected a distribution shift between historical training data and current clinical data. - Researchers adjusted operating points during deployment to maintain safe and appropriate recall rates for local workflows. ## Study 2: AI in the Double-Reading Workflow - The second study examined how human readers performed when using AI as part of arbitration, rather than evaluating AI in isolation. - Twenty-two readers reviewed thousands of cases using real screening-service rules. - Two workflows were compared: - **Standard care:** decisions from the historical first and second human readers - **AI-enabled care:** the historical first-reader decision paired with the AI decision - This design aimed to assess the practical effects of replacing the second human read with an AI reader. The findings support AI as a potential second reader in breast cancer screening, but broader prospective clinical validation is still needed. Successful adoption should include phased deployment, local calibration, continuous monitoring, and careful evaluation of human-AI decision-making.

Read original(opens in new tab)
google3 min readCurated summary

Google Research at The Check Up: from healthcare innovation to real-world care settings

Google Research argues that AI is entering a new phase in healthcare: moving beyond isolated tools toward personalized care, clinical collaboration, public-health planning, and scientific discovery. The post highlights research partnerships, open models, and real-world deployments designed to make healthcare more accurate, accessible, and proactive. Google emphasizes that these advances must be developed responsibly through clinical validation, peer review, and collaboration with healthcare institutions. ## AI for Personalized Healthcare - A Fitbit collaboration studied how AI could support preventative care across the United States. - The research found that a Personal Health Agent (PHA) modeled on a collaborative health team could provide more effective long-term support than single-purpose fitness or tracking apps. - The PHA combines: - Data analysis - Medical and domain expertise - Health coaching - Large multimodal models can transform wearable data into personalized guidance about sleep, fitness, and overall health. ## AI as a Clinical Collaborator - Google’s breast cancer research with Imperial College London and the UK’s NHS used diverse datasets and expert-validated ground truth data. - The experimental system identified 25% of “interval cancers”—cancers missed during screening and later detected after symptoms appeared. - Integrated into clinical workflows, the system could reduce radiologists’ workload while maintaining safe detection performance. - Google’s diabetic retinopathy screening model has been deployed through partnerships with medical institutions in India, Thailand, and Australia. - It has supported more than one million screenings. - Patients can receive results in roughly two minutes. - AMIE, a multi-agent medical AI system, can reason across medical histories, laboratory results, and medical images to identify overlooked patterns. - Google is testing AMIE with Beth Israel Deaconess Medical Center to assist with pre-visit history-taking and flag urgent symptoms. - An IRB-approved national study with Included Health will evaluate AI-supported telehealth care. ## Open Models for Healthcare Developers - Google’s Health AI Developer Foundations (HAI-DEF) provides free open-weight models and open-source tools for building healthcare applications. - MedGemma supports: - Medical text and image interpretation - High-dimensional 3D imaging - Medical-specific speech recognition - The All India Institute of Medical Sciences is using MedGemma for outpatient triage and dermatology screening. - Singapore’s Ministry of Health is adapting the model for locally relevant primary- and specialty-care applications. - The MedGemma Impact Challenge received more than 850 submissions aimed at turning AI research into practical, human-centered healthcare tools. ## AI for Public Health - Google Earth AI combines geospatial models and datasets to study connections between environmental conditions, population behavior, and health outcomes. - Researchers at Mount Sinai and Boston Children’s Hospital/Harvard used Google data and surveys to estimate childhood MMR vaccination coverage at ZIP-code resolution. - The resulting “super-resolution” maps identified pockets of under-vaccination that corresponded with recent measles outbreaks. - Such analysis could help public-health officials target outreach and prevention efforts more effectively. ## AI for Biomedical Discovery - Co-Scientist and Gemini Deep Think are being used to generate scientific hypotheses and support research across fields including single-cell analysis, public health, and neuroscience. - Google is also exploring evolutionary coding agents that run scientific-computing experiments in parallel. - DeepSomatic, a genomic analysis tool, is designed to improve the detection of cancer-related genetic mutations across multiple cancer types. Google’s broader recommendation is to treat AI as a validated collaborator and infrastructure layer rather than a replacement for clinicians or researchers. Continued clinical testing, expert oversight, transparent publication, and open developer access will be essential to translating these systems into safe, practical benefits.

Read original(opens in new tab)
cloudflare3 min readCurated summary

Standing up for the open Internet- why we appealed Italy’s Piracy Shield fine

Cloudflare argues that Italy’s “Piracy Shield” undermines the open Internet by allowing private media companies to order broad website and IP-address blocks without judicial oversight, transparency, or due process. After refusing to register with the system, Cloudflare was fined €14 million by Italy’s communications regulator, AGCOM, and appealed the penalty on March 8. The company is also challenging Piracy Shield’s legality under EU law. ### What Piracy Shield Does - Piracy Shield requires registered service providers to block websites and IP addresses within 30 minutes of a submission. - Blocking decisions are made through an electronic portal by an unidentified group of private Italian media companies. - The system provides: - No judicial oversight. - Little or no transparency about who requested a block or why. - No opportunity for website owners to challenge blocks beforehand. - No effective redress for wrongful blocking. - Cloudflare says the system’s design favors major rightsholders, including Italy’s Serie A soccer league. - The system was reportedly donated to the government by SP Tech, an arm of a law firm representing major beneficiaries. ### Problems with IP Blocking - Shared IP addresses can host thousands of unrelated websites, making accidental overblocking unavoidable. - Piracy Shield has caused outages affecting: - Ukrainian government, education, and research websites. - European small businesses and NGOs supporting women and children. - Google Drive, which was inaccessible to many Italian users for more than 12 hours. - A September 2025 University of Twente study found that legitimate sites were routinely blocked for months. - Despite these failures, AGCOM expanded the scheme to global DNS providers and VPNs, services closely linked to privacy and free expression. ### Cloudflare’s Legal Challenge - Cloudflare warned AGCOM in 2024 that Piracy Shield threatened the Internet’s technical architecture and proposed alternative approaches. - It later challenged AGCOM’s attempt to force Cloudflare to register with the system in Italian administrative courts. - Cloudflare and the CCIA also complained to the European Commission, arguing that Piracy Shield conflicts with the EU Digital Services Act. - The DSA requires content restrictions to be proportionate and governed by meaningful procedural safeguards. - The European Commission criticized Piracy Shield’s lack of oversight in a June 2025 letter. - In December 2025, an Italian administrative court ordered AGCOM to provide Cloudflare with records supporting its blocking orders, though Cloudflare says it has not yet received them. ### The €14 Million Fine - AGCOM issued the fine on December 29, 2025, shortly after being ordered to disclose Piracy Shield records. - Cloudflare argues that the penalty’s timing and calculation raise serious concerns. - The article ends while beginning to discuss Italian legal limits on fines and the alleged flaws in AGCOM’s calculation. Cloudflare’s position is that copyright enforcement cannot justify an opaque, privately operated blocking system with no effective safeguards. It recommends defending the legal and technical principles of an open Internet rather than expanding Piracy Shield.

Read original(opens in new tab)
aws3 min readCurated summary

AWS Weekly Roundup: Amazon S3 turns 20, Amazon Route 53 Global Resolver general availability, and more (March 16, 2026) | Amazon Web Services

Amazon S3 marked its 20th anniversary with major milestones in scale, performance, and cost reduction, while AWS introduced account regional namespaces for improved bucket-name control. The week’s featured launch was the general availability of Amazon Route 53 Global Resolver, providing secure, globally accessible DNS resolution across 30 AWS Regions. Other updates covered stateful AI agent infrastructure, Windows Server 2025, simplified AWS identity access, and reusable Redshift ingestion templates. ## Amazon S3 Reaches 20 Years - Launched publicly on March 14, 2006, S3 has grown from object storage into a foundational cloud service. - As of March 2026, it stores: - More than 500 trillion objects - Hundreds of exabytes of data - Over 200 million requests per second globally - Storage prices have fallen by approximately 85% since launch, to just over $0.02 per gigabyte. - New account regional namespaces let organizations reserve bucket names within their own account namespace by adding an account-specific suffix. - Adoption can be enforced with IAM and AWS Organizations service control policies using the `s3:x-amz-bucket-namespace` condition key. ## Route 53 Global Resolver Becomes Generally Available - Amazon Route 53 Global Resolver is an internet-reachable, anycast DNS resolver for authorized clients anywhere. - It is available across 30 AWS Regions and supports IPv4 and IPv6 DNS queries. - It resolves: - Public internet domains - Private domains associated with Route 53 private hosted zones - Security features include filtering for malicious, unsafe, DNS tunneling, and Domain Generation Algorithm (DGA) domains. - General availability adds protection against Dictionary DGA threats. - Centralized DNS query logging is also included. ## Additional AWS Service Updates - **Bedrock AgentCore Runtime** - Adds stateful MCP server support through the `Mcp-Session-Id` header. - Dedicated microVMs isolate each user session and preserve context across interactions. - MCP servers can use elicitation, sampling, and progress notifications in addition to resources, prompts, and tools. - **Amazon WorkSpaces** - Adds Windows Server 2025 bundles for WorkSpaces Personal and WorkSpaces Core. - Security features include TPM 2.0, UEFI Secure Boot, Credential Guard, HVCI, Secured-core server, and DNS-over-HTTPS. - Existing Windows Server 2016, 2019, and 2022 bundles remain supported. - **AWS Builder ID** - Adds GitHub and Amazon as sign-in options alongside Google and Apple. - Users can access AWS Builder Center, Training and Certification, and Kiro without maintaining separate credentials. - **Amazon Redshift** - Introduces reusable templates for `COPY` operations. - Templates centralize frequently used parameters, improve consistency, and automatically apply future updates to subsequent loads. - The feature is available in commercial and AWS GovCloud Regions. ## Upcoming AWS Events - AWS Summits are scheduled for Paris, London, and Bengaluru. - AWS Community Days are planned in Pune, San Francisco, and Romania. - AWS will participate in NVIDIA GTC 2026 in San Jose. - AWS Community GameDay Europe will offer hands-on troubleshooting challenges across more than 50 cities. For practitioners, the most significant developments are Route 53 Global Resolver for centralized global DNS security, S3 namespaces for organizational naming governance, and AgentCore’s stateful MCP support for more capable AI applications.

Read original(opens in new tab)
github3 min readCurated summary

GitHub for Beginners: Getting started with GitHub Actions

GitHub Actions is GitHub’s built-in platform for automating CI/CD and repetitive repository tasks. Workflows are YAML files triggered by events such as pushes, pull requests, schedules, or newly opened issues, then executed as jobs on hosted or self-hosted runners. The post guides beginners through creating a workflow that automatically labels new issues. ## What GitHub Actions Provides - GitHub Actions supports: - Continuous integration and delivery - Automated tests and vulnerability scans - Release creation - Team reminders and other repetitive tasks - Workflows are stored in the repository and run automatically when configured events occur. - Jobs execute in virtual machines called runners, provided by GitHub or managed by the user. ## How Workflows Operate - **Events** trigger workflows, such as: - Pushing code - Opening or merging pull requests - Creating issues - Scheduled times - **Runners** are virtual machines that execute workflow jobs. GitHub offers Ubuntu, Windows, and macOS hosted runners, while teams can also use self-hosted runners. - **Jobs** contain groups of steps executed on the same runner. - **Steps** can either run shell commands or invoke reusable Marketplace actions. ## Workflow Structure Workflow files use YAML and live in `.github/workflows`. The three main sections are: - **`name`**: Describes the workflow. - **`on`**: Specifies the event or events that trigger it. - **`jobs`**: Defines the work performed after triggering. The post recommends descriptive filenames such as `build-and-test.yml`, `security-scanner.yml`, or `label-new-issue.yml`. ## Creating an Issue-Labeling Workflow The example workflow automatically adds a `triage` label whenever a new issue is opened. - It is named `Label New Issues`. - Its trigger is configured as: ```yaml on: issues: types: [opened] ``` - The `label-issues` job runs on `ubuntu-latest`. - Permissions are explicitly granted: - `issues: write` allows the workflow to add labels. - `contents: read` allows it to access repository content. ## Using Actions and Shell Commands The workflow contains two steps: - `actions/checkout@v6` uses a prebuilt Marketplace action to check out the repository code. - A shell command uses the GitHub CLI to add the label: ```bash gh issue edit "$ISSUE_NUMBER" --add-label "$LABEL" ``` Environment variables provide the command with: - `GITHUB_TOKEN` for authentication - The issue number from `github.event.issue.number` - The label name, `triage` The `uses` keyword invokes reusable actions, while `run` executes a shell command directly. Start with a small workflow in `.github/workflows`, define its trigger and required permissions carefully, and build from reusable actions plus simple commands. The post also recommends practicing with GitHub’s “Hello GitHub Actions” exercise to become familiar with workflow creation.

Read original(opens in new tab)
cloudflare4 min readCurated summary

From legacy architecture to Cloudflare One

Moving from fragmented VPNs to Cloudflare One is presented as a gradual modernization effort rather than a risky “big bang” cutover. Cloudflare and CDW recommend a tiered, application-aware migration that combines Zero Trust controls with careful dependency analysis and staged deployment. The central conclusion is that legacy applications can gain modern security protections without immediate code rewrites or major downtime. ## Reducing Big-Bang Migration Risk - Large organizations may need to transition hundreds or thousands of applications and users from legacy VPNs. - A single firewall error, dependency failure, or session timeout can disrupt essential services. - These risks often prevent organizations from adopting Zero Trust despite vulnerable, aging infrastructure. - CDW applies lessons from failed deployments to create a risk-aware migration roadmap. - Applications are categorized by complexity, with simpler systems migrated first and legacy systems handled later under tighter controls. - A public-sector migration of 500 applications caused widespread disruption because more than 4,000 applications had not been prioritized or tiered. ## Treating Migration as Application Modernization - Traditional migrations often treat networks as basic connectivity infrastructure and overlook application ecosystems. - CDW analyzes: - Backend databases and APIs - Identity and authentication dependencies - Hidden service-to-service calls - Legacy session behavior - Security requirements are incorporated into the architecture from the beginning rather than added after connectivity is restored. - The migration becomes an application modernization program instead of a simple VPN replacement. ## Protecting Legacy Applications with Cloudflare Access - Cloudflare Access replaces broad network-level VPN access with request-by-request Zero Trust authorization. - Each request can be evaluated using: - User identity - Device posture - Hardware-based MFA - Other contextual signals - This limits lateral movement and reduces the attack surface. - Legacy applications can be “wrapped” with modern security controls without rewriting their code. - Cloudflare Tunnel provides: - An outbound-only connection - SSO and MFA integration - No public IP exposure for the application - Access policies can require endpoint MFA and a device health check before traffic reaches the server. - This approach allows organizations to modernize security incrementally while legacy applications continue operating. ## Pre-Migration Audit ### Architectural and Identity Assessment - Identify whether applications use a federated identity provider such as Okta or legacy local directories. - Map database, API, and backend dependencies. - Verify that hidden API calls and service-token-based Tunnel connections will continue functioning after migration. - Assess whether applying least-privilege controls could break application behavior. ### Establishing a Strategy and Implementation Firebreak - Create separate groups for: - Security strategy and standards - Deployment and operational implementation - This separation prevents deployment speed from overriding requirements designed to limit lateral movement. ### Testing Persistent Sessions - Identify applications that depend on persistent sessions, particularly for users switching between cellular towers. - Cloudflare’s edge architecture and Dynamic Path MTU Discovery (PMTUD) help maintain sessions even when client IP addresses change. - This assessment can identify opportunities to replace rigid legacy hardware with a modern single-pass architecture. ### Categorizing Applications and Setting Timelines - **Tier 0: Modern SaaS applications** - Native SAML/OIDC support - Cloudflare can act as a clientless identity-provider proxy - Estimated effort: 1–3 hours per application - **Tier 1: Internal web applications** - Support identity headers and modern web protocols - Use a clientless reverse proxy with Cloudflare Tunnel - Estimated effort: 3–6 hours per application - **Tier 2: Non-web client-server applications** - Require specific port/protocol support or thick-client configurations - Use both Cloudflare One Client and Cloudflare Tunnel - Estimated effort: 4–8 hours per application A phased migration built around application dependencies, identity readiness, session behavior, and technical complexity offers a safer path to Cloudflare One. Organizations should begin with an audit and pilot, secure legacy applications using Access and Tunnel, and expand tier by tier rather than attempting a single cutover.

Read original(opens in new tab)
google3 min readCurated summary

Testing LLMs on superconductivity research questions

LLMs may help physicists navigate complex research, but their reliability depends heavily on the quality and curation of their sources. In a high-temperature superconductivity study, systems grounded in expert-selected literature—especially NotebookLM and a custom retrieval-augmented generation system—outperformed general web-access models. The results suggest that trustworthy scientific AI requires balanced reasoning, strong evidence, and carefully controlled reference collections. ## Evaluating LLMs on Superconductivity - Researchers from Google Research and Cornell University tested whether LLMs could answer expert-level questions in condensed matter physics. - The study focused on cuprate high-temperature superconductors, whose underlying mechanism remains unresolved despite decades of research. - Understanding superconductivity in these materials could help scientists discover compounds that work at higher temperatures. - The field contains thousands of experimental and theoretical papers and competing explanations, making it difficult for researchers—especially newcomers—to establish a balanced view. ## Study Design and Sources - Six systems were evaluated: - GPT-4o - Perplexity - Claude 3.5 - Gemini Advanced Pro 1.5 - Google NotebookLM - A custom retrieval-augmented generation (RAG) system - Four models had broad web access, including 765 open-access experimental papers and 1,553 theoretical papers. - NotebookLM and the custom RAG system used a curated database: - Twelve superconductivity experts selected 15 review articles. - Those reviews contained approximately 3,300 references. - A final collection of 1,726 experimental papers and reviews was assembled. - Experts created 67 difficult questions, including questions about doping levels and evidence for quantum criticality in cuprates. ## Evaluation Criteria Experts used masked reviews and scored responses from 0 to 2 on: - Balance between competing scientific perspectives - Comprehensiveness and factual depth - Conciseness and clarity - Evidence and links to sources - Relevance of supplied images - Qualitative comments ## Results - NotebookLM achieved the strongest overall performance. - The custom RAG system ranked second overall, showing the value of retrieval from the same expert-curated sources. - NotebookLM, Gemini, and the custom RAG system performed best at presenting balanced and comprehensive answers. - NotebookLM provided the strongest evidence and citations but was less concise than the other systems. - Image quality was generally weaker; the custom RAG system performed best among the models that regularly supplied images. - All systems showed areas needing improvement, particularly when addressing nuanced, unresolved research questions. ## Practical Implication For scientific research, LLMs should be paired with expert-curated, quality-controlled literature rather than relying solely on unrestricted web searches. Such systems can serve as research tutors or thought partners, but their answers still require expert verification, especially in fields with competing theories and rapidly evolving evidence.

Read original(opens in new tab)