Techlist.io - Korean Tech Blog Curator

googleOriginal article

Sensible Agent: A framework for unobtrusive interaction with proactive AR agents (opens in new tab)

Sensible Agent is a research prototype designed to move AR agents beyond explicit voice commands toward proactive, context-aware assistance. By leveraging real-time multimodal sensing of a user's environment and physical state, the framework ensures digital help is delivered unobtrusively through the most appropriate interaction modalities. This approach fundamentally reshapes human-computer interaction by anticipating user needs while minimizing cognitive and social disruption. ## Contextual Understanding via Multimodal Parsing The framework begins by analyzing the user's immediate surroundings to establish a baseline for assistance. * A Vision-Language Model (VLM) processes egocentric camera feeds from the AR headset to identify high-level activities and locations. * YAMNet, a pre-trained audio event classifier, monitors environmental noise levels to determine if audio feedback is appropriate. * The system synthesizes these inputs into a parsed context that accounts for situational impairments, such as when a user’s hands are occupied. ## Reasoning with Proactive Query Generation Once the context is established, the system determines the specific type of assistance required through a sophisticated reasoning process. * The framework uses chain-of-thought (CoT) reasoning to decompose complex problems into intermediate logical steps. * Few-shot learning, guided by examples from data collection studies, helps the model decide between actions like providing translations or displaying a grocery list. * The generator outputs a structured suggestion that includes the specific action, the query format (e.g., binary choice or icons), and the presentation modality (visual, audio, or both). ## Dynamic Modality and Interaction Management The final stage of the framework manages how the agent communicates with the user and how the user can respond without breaking their current flow. * The prototype, built on Android XR and WebXR, utilizes a UI Manager to render visual panels or generate text-to-speech (TTS) prompts based on the agent's decision. * An Input Modality Manager activates the most discreet response methods available, such as head gestures (nods), hand gestures (thumbs up), or gaze tracking. * This adaptive selection ensures that if a user is in a noisy room or a social setting, the agent can switch from verbal interaction to subtle visual cues and gesture-based confirmations. By prioritizing social awareness and context-sensitivity, Sensible Agent provides a blueprint for AR systems that feel like helpful companions rather than intrusive tools. Implementing such frameworks is essential for making proactive digital assistants practical and acceptable for long-term, everyday use in public and private spaces.

figma3 min readCurated summary

Figma Rendering: Powered by WebGPU | Figma Blog

Figma replaced its WebGL-based renderer with a WebGPU backend to unlock GPU compute, clearer resource management, and better error handling. The migration required more than swapping APIs: Figma redesigned its graphics interface, supported both WebGL and WebGPU, and built tooling to translate existing shaders. The project also improved the existing WebGL renderer by making rendering inputs explicit and reducing opportunities for state-related bugs. ## Why Figma Moved Beyond WebGL - Figma originally chose WebGL to deliver a smooth, browser-based infinite canvas when most design tools were native applications. - WebGPU, supported by Chromium since 2023, enables: - Compute shaders that move parallelizable work from the CPU to the GPU. - Less reliance on WebGL’s bug-prone global state. - More capable and understandable error handling. - The transition had to preserve WebGL compatibility and avoid performance regressions or disruptions during rollout. ## Making Draw Calls Explicit - The previous interface mirrored WebGL’s global-state model: - Buffers, textures, materials, and framebuffers were bound separately. - Resources remained bound after a draw call. - Developers could accidentally reuse stale state. - Figma redesigned `draw()` so that required resources are passed directly as arguments. - The WebGL implementation lazily updates bindings only when necessary, preserving efficiency while making dependencies explicit. - This redesign fixed several WebGL bugs before WebGPU support was introduced. ## Supporting GLSL and WGSL Shaders - WebGL uses GLSL, while WebGPU uses WGSL. - Maintaining separate GLSL and WGSL versions of every shader would have created excessive duplication and maintenance work. - Figma built a custom shader processor that: - Parses existing WebGL 1–compatible GLSL. - Translates it into a newer GLSL structure. - Uses the open-source `naga` tool to convert it to WGSL. - Generates both GLSL and WGSL outputs. - Extracts shader metadata such as input types and data layouts. - Supports file includes for shader reuse and modularity. - This allowed engineers to continue writing and maintaining one primary shader source while supporting both rendering backends. ## Adapting Uniform Data - Uniforms provide shader inputs such as colors and transformations. - WebGL allows uniforms to be updated individually through calls such as `uniform1f` and `uniformMatrix3fv`. - Figma’s original graphics interface followed this model with methods such as `setUniform1f`. - WebGPU requires uniforms to be grouped into uniform buffers and uploaded together. - Consequently, simply switching APIs could have reduced performance; Figma needed to redesign uniform handling carefully rather than directly reproducing WebGL behavior. ## Practical Outcome Figma’s WebGPU migration was an architectural modernization as much as a graphics API upgrade. The recommended approach is to introduce an explicit, backend-independent rendering interface, automate shader translation, and optimize data layouts carefully so WebGPU’s capabilities improve performance without sacrificing compatibility.

Read original(opens in new tab)
googleOriginal article

Making LLMs more accurate by using all of their layers (opens in new tab)

Self Logits Evolution Decoding (SLED) is a novel decoding strategy designed to reduce hallucinations and improve the factual accuracy of large language models without requiring external data or fine-tuning. By leveraging the internal representations of all model layers rather than just the final output, SLED aligns generation with the model’s intrinsic knowledge more effectively. Research shows that this approach consistently enhances performance across diverse tasks, including complex reasoning, multiple-choice questions, and open-ended generation. ## Limitations of Standard Decoding * Standard LLMs typically generate text by relying solely on the "logits" (prediction scores) of the final layer to determine the next token. * This process often leads to hallucinations because the final layer may prioritize "popular" or common patterns from training data over factual accuracy. * While techniques like Retrieval Augmented Generation (RAG) provide external context, they increase system complexity and do not address the model's internal tendency to ignore subtle contextual cues during the final projection. ## The Technical Mechanism of SLED * SLED utilizes "early exit" logits from every intermediate layer of the Transformer architecture, rather than just the final one. * The strategy reuses the model's final projection matrix on these intermediate layers to create multiple probability distributions across the same set of potential tokens. * By calculating a weighted average of the distributions from all layers, SLED refines the prediction to better reflect the model's latent knowledge. * This multi-layer approach allows the model to catch nuances—such as specific math constraints or geographic facts—that might be "smoothed over" by the final layer’s preference for high-probability sequences. ## Practical Performance and Reasoning * In chain-of-thought tasks, SLED helps the model maintain logic; for example, it can correctly identify when a discount should be applied in a math problem by favoring intermediate layers that recognize the "if/then" logic over a simple arithmetic pattern. * The method is model-agnostic and has shown consistent accuracy gains across various LLM scales and configurations. * SLED is highly flexible and can be integrated with existing factuality decoding methods or speculative decoding to further reduce hallucinations without the need for additional training data. For developers and researchers seeking to boost the reliability of LLMs, SLED offers a computationally efficient alternative to fine-tuning. By simply adjusting the decoding strategy to incorporate the rich information available in intermediate layers, models can achieve higher factuality and more robust reasoning capabilities in real-world applications.

googleOriginal article

Learn Your Way: Reimagining textbooks with generative AI (opens in new tab)

Google Research has introduced Learn Your Way, an AI-driven educational experiment that reimagines traditional textbooks as personalized, multimodal learning journeys. By leveraging the LearnLM family of models integrated into Gemini 2.5 Pro, the system transforms static source material into tailored content based on a student’s specific grade level and interests. Early efficacy studies demonstrate that this approach significantly enhances retention, with students scoring 11 percentage points higher than those using standard digital readers. ### Pedagogical Foundations and Dual Coding The research is built on the "dual coding theory," which suggests that forming mental connections between different representations of information strengthens conceptual understanding. * The system moves away from a "one-size-fits-all" model toward a student-driven experience where learners can choose and intermix formats. * Personalization is used as a tool to enhance situational interest and motivation by adapting content to specific student attributes. * The framework incorporates active learning through real-time quizzing and feedback to address knowledge gaps as they arise. ### The Personalization Pipeline The technical architecture begins with a layered pipeline that processes source material, such as a textbook PDF, to create a foundational text for all other formats. * The original material is first "re-leveled" to match the learner’s reported grade level while maintaining the integrity and scope of the curriculum. * Generic examples within the text are strategically replaced with personalized examples based on user interests, such as sports, music, or food. * This personalized base text serves as the primary input for generating all subsequent multimodal representations, ensuring consistency across formats. ### Multimodal Content Generation To produce a wide variety of educational assets, the system utilizes a combination of large language models and specialized AI agents. * **Agentic Workflows:** While tools like mind maps and timelines are generated directly by Gemini, complex assets like narrated slides use multi-step agentic workflows to ensure pedagogical effectiveness. * **Custom Visuals:** Because general-purpose image models often struggle with educational accuracy, the researchers fine-tuned a dedicated model specifically for generating educational illustrations. * **Diverse Representations:** The interface provides "immersive text" with embedded questions, audio lessons for auditory learning, and interactive slides that mimic recorded classroom sessions. ### Research Outcomes and Future Application The project’s effectiveness was validated through a study comparing the GenAI approach against standard digital reading materials. * Students using the personalized AI tools showed a significant improvement in retention test scores. * Beyond retention, the system aims to transform passive reading into an active, multimodal experience that follows established learning science principles. * The "Learn Your Way" experiment is currently available on Google Labs, providing a practical look at how adaptive, learner-centric materials might replace static textbooks in future K-12 and higher education settings.

airbnb4 min readCurated summary

Viaduct, Five Years On: Modernizing the Data-Oriented Service Mesh

Viaduct, Airbnb’s data-oriented service mesh, has evolved substantially over five years while retaining its core model: a central schema, hosted business logic, and re-entrant composition through GraphQL. Its usage has grown eightfold, supporting more than 130 teams and over 1.5 million lines of production code, without increasing operational overhead. Viaduct Modern now aims to simplify its developer API and establish stronger architectural boundaries, alongside the project’s release as open source. ## Adoption and Evolution - Viaduct traffic has increased by a factor of eight since 2020. - More than 130 teams now host code in Viaduct, supported by hundreds of weekly active developers. - The hosted codebase has grown to over 1.5 million lines, with roughly the same amount of test code. - Operational overhead has remained constant, incident-minutes have been cut in half, and costs have grown linearly with QPS. - Viaduct is now available as open-source software. ## Core Principles That Remain - **Central schema:** Viaduct provides one integrated schema connecting domains across Airbnb. - More than 75% of requests are internal. - The schema is developed by many teams but exposed as a connected graph. - **Hosted business logic:** Teams run business logic directly in Viaduct rather than maintaining separate microservices. - This reduces operational overhead and can allow standalone services to be retired. - Viaduct provides a serverless environment so developers can focus on application logic. - **Re-entrancy:** Hosted logic composes with other hosted logic through GraphQL fragments and queries. - This supports modularity. - It helps avoid the tightly coupled structure and maintenance problems associated with traditional monoliths. ## Problems with the Earlier Design - Viaduct’s APIs evolved reactively in response to individual use cases. - Multiple mechanisms emerged for accomplishing similar tasks, creating confusion for developers. - Some capabilities were well supported while others were not. - The framework’s layers had loose, inconsistent interfaces. - The boundary between Viaduct and hosted application code was weak. - These issues made framework improvements increasingly risky because changes could disrupt existing users. ## Simplifying the Tenant API - Viaduct Modern overhauls the developer-facing API and execution engine. - The new Tenant API reduces the implementation choices to two mechanisms: - **Node resolvers** - **Field resolvers** - The choice is determined by the schema rather than by ad hoc behavioral distinctions. - Resolver APIs have been unified wherever possible. - The goal is a smaller, more consistent surface that preserves successful ideas from the old API while removing unnecessary alternatives. ## Tenant Modules and Re-Entrant Composition - Viaduct uses modules and re-entrancy to provide boundaries similar to service definitions and RPC APIs in microservice architectures. - A tenant module combines: - Schema owned by a team - The code implementing that schema - Modules can create rich connections in the shared graph, but direct code dependencies between teams are discouraged. - Instead, teams declare their data requirements through GraphQL fragments and queries. ### Example: Extending the `User` Type - A Core User team owns the base `User` type and resolves fields such as `firstName` and `lastName`. - A Messaging team can extend `User` with a `displayName` field. - Its resolver declares that it needs `firstName` and `lastName`. - Messaging does not depend directly on Core User’s implementation or need to know where those fields originate. - This declarative model lets teams collaborate through the schema while preserving ownership and modularity. ## Framework Modularity - Viaduct Modern also restructures the framework itself. - The system consists of: - The GraphQL execution engine - The Tenant API - Hosted application code - Historically, the interfaces between these layers were weak, making performance and reliability improvements difficult to introduce safely. - The redesign focuses on stronger abstraction boundaries so the framework can evolve independently of application code. Viaduct’s modernization is intended to preserve its centralized, data-oriented model while making development simpler and framework evolution safer. The open-source release provides an opportunity for other organizations to evaluate or adopt this approach to schema-driven, modular service composition.

Read original(opens in new tab)
figma3 min readCurated summary

Free Association: Production Designer Jeremy Hindle on Building Severance | Figma Blog

Jeremy Hindle’s production design for *Severance* began with a minimal script description—“four desks, large room”—and developed into Lumon Industries’ eerie, emotionally charged world. Drawing on modernist architecture, industrial design, cinema, and television, Hindle prioritized how spaces feel over what they communicate intellectually. His conclusion is that the show’s unsettling power comes from the tension between sterile order, vast scale, and profound loneliness. ## Designing Emotion - Hindle describes his goal as creating an atmosphere that is “stunning and lonely.” - His experience designing hundreds of commercial offices informed Lumon’s corporate environments. - Rather than focusing only on narrative information, he designs for viewers’ “kinetic” responses—the physical and emotional sensations produced by a space. - The sparse script gave him room to build a complete visual language around emptiness, repetition, and institutional control. ## *Playtime* and Architectural Absurdity - Jacques Tati’s *Playtime* was a central reference for *Severance*. - Tati’s film satirizes modern corporate architecture through vast, impersonal spaces and carefully choreographed movement. - Its mixture of comedy, alienation, and architectural scale helped shape Lumon’s uncanny office environment. ## John Deere Headquarters and Lumon’s Monumental Scale - Kevin Roche and Eero Saarinen’s 1964 John Deere headquarters provided a blueprint for Lumon’s institutional architecture. - Its monumental structure influenced the show’s sense of corporate permanence and authority. - The building also inspired the iconic desks used by the Macrodata Refinement team. - Hindle translated large-scale architectural references into environments that feel both imposing and strangely empty. ## Dieter Rams and Controlled Modernism - Lumon’s largely fabricated sets include selected real design classics, including Dieter Rams’s Braun TS 45 and TG 60 wall console. - Rams’s restrained, functional aesthetic reinforces the company’s carefully controlled visual identity. - The inclusion of recognizable industrial design creates a believable mid-century modern world while making Lumon’s environment feel more curated and artificial. ## Lessons from Film Production - Hindle’s work on *Zero Dark Thirty*, his first feature film, strengthened his commitment to constructing large-scale physical worlds. - He used full-scale models to help the entire crew understand how environments would function and remain visually readable. - This experience supported the detailed spatial planning required for *Severance*’s complex office interiors. ## Negative Space and the Scale of *Fargo* - A single still from the Coen brothers’ *Fargo* influenced Hindle’s vision for the scale of *Severance*. - The image demonstrated how negative space can make characters appear isolated and vulnerable. - Hindle showed the reference to director Ben Stiller, using it to communicate the show’s desired balance of composition, distance, and loneliness. ## The Emotional World of *Her* - Spike Jonze’s *Her*, with production design by K.K. Barrett and costumes by Casey Storm, served as another important reference. - Hindle admired the film’s ability to create a complete emotional world through design. - Its coordinated use of architecture, interiors, color, and clothing helped inform *Severance*’s approach to character and atmosphere. ## *Twin Peaks* and Television with Staying Power - David Lynch’s *Twin Peaks* represented the kind of distinctive, enduring television Hindle hoped *Severance* could become. - Although he had not previously been strongly interested in television, the project offered an opportunity to build a world with lasting cultural and visual identity. - Like *Twin Peaks*, *Severance* combines recognizable everyday settings with surrealism, mystery, and an unsettling sense of place. Hindle’s work shows how production design can make an abstract corporate dystopia physically memorable. By combining modernist architecture, precise industrial objects, cinematic negative space, and emotionally charged emptiness, *Severance* turns Lumon into one of the story’s most powerful characters.

Read original(opens in new tab)
discord3 min readCurated summary

Discord Patch Notes: September 3, 2025

Discord’s September 3, 2025 patch focuses on scaling, performance, mobile improvements, permissions, and a large collection of cross-platform bug fixes. Discord increased the default server capacity from 2.5 million to 25 million members while improving update processing and asynchronous operations for very large communities. The Android app also moved to React Native’s new architecture, with early results showing smoother performance, particularly on lower-end devices. ## Large-Server Improvements - Raised the default server member limit to **25 million**, up from 2.5 million. - Batched certain server updates and shifted more work to asynchronous processes to improve stability. - Added further improvements for large servers and those using Community features. - Continued work based on administrator feedback, including changes described in Discord’s Community Server Cleanup Report. ## Android and Mobile Updates - Upgraded Android to React Native’s new architecture across most of the app. - Early performance data shows: - Smoother app operation. - Fewer low-frame-rate problems. - Better behavior on lower-end devices. - Improved the mobile Shop with: - Faster loading. - Gifting. - Nameplate purchases. - A redesigned layout closer to the desktop experience. - Added camera-enable and camera-disable sound effects, especially useful when using keyboard shortcuts. - Improved mobile rendering and behavior for Markdown, member statuses, thread lists, voice-region menus, browser settings, and rotated layouts. ## Server Management and Permissions - Increased the pinned-message limit from **50 to 250**. - Added a dedicated permission for pinning and unpinning messages, separating it from **Manage Messages**. - Fixed issues involving: - Role assignment during Server Onboarding. - Reordering roles on iOS. - Server subscriptions and subscription modals. - Community server icons. - Server Insights announcement-channel charts. - Twitch integration’s Force Sync button. - Nitro boosts and localized Japanese formatting. ## General Usability Fixes - Mobile member lists now correctly display “Can’t Wait For” and “Obsessed With” status messages. - Fixed profile views that could become impossible to close after rapidly opening mutual-friend profiles. - Corrected poll notification placeholders on Android. - Restored proper Markdown rendering on mobile Event pages. - Fixed Quick Switcher search prefixes such as `!`, `@`, `#`, and `*`. - Improved DM context-menu hitboxes and fixed several spacing, padding, alignment, and tooltip issues. - Corrected invalid invite embeds, owner crown alignment, role and app badge positioning, and Friend Suggestions layout. - Fixed desktop pop-out windows incorrectly displaying a menu bar. ## Notifications, Themes, and Accessibility - Automod incident notifications now follow users’ existing push-notification settings. - Fixed randomly changing client theme colors. - Improved visibility and formatting for light-theme Student Hub search placeholders. - Corrected localization issues involving Nitro Home and other interface elements. - Fixed several unreadable, misaligned, or incorrectly sized buttons and controls. ## Platform-Specific Fixes - **iOS:** Fixed thread-list flickering, browser selection, voice-region menu backgrounds, role reordering, phone-related interface behavior, and channel details after rotation. - **Android:** Fixed phone country-code controls, app badge alignment, poll notifications, and mobile layout issues. - **Windows:** Corrected navigation from Action Center notifications. - **macOS:** Fixed an unclickable title bar. - **Desktop/Web:** Resolved issues with custom keybind warnings, pop-out windows, search, tooltips, subscription flows, and various menus. Discord says all listed fixes have been merged, though they may still be rolling out gradually across platforms.

Read original(opens in new tab)
lineOriginal article

Code Quality Improvement Techniques Part (opens in new tab)

When implementing resource management patterns similar to Kotlin's `use` or Java's try-with-resources, developers often face the challenge of handling exceptions that occur during both primary execution and resource cleanup. Simply wrapping these multiple failures in a custom exception container can inadvertently break the calling code's error-handling logic by masking the original exception type. To maintain code quality, developers should prioritize the primary execution exception and utilize the `addSuppressed` mechanism to preserve secondary errors without disrupting the expected flow. ### The Risks of Custom Exception Wrapping Creating a new exception class to consolidate multiple errors during resource management can lead to significant issues for the caller. * Wrapping an expected exception, such as an `IOException`, inside a custom `DisposableException` prevents specific `catch` blocks from identifying and handling the original error. * This pattern often results in unhandled exceptions or the loss of specific error context, especially when the wrapper is hidden inside utility functions. * While this approach aims to be "neat" by capturing all possible failures, it forces the caller to understand the internal wrapping logic of the utility rather than the business logic errors. ### Prioritizing Primary Logic over Cleanup When errors occur in both the main execution block and the cleanup (e.g., `dispose()` or `close()`), it is critical to determine which exception takes precedence. * The exception from the main execution block is typically the "primary" failure that reflects a business logic or IO error, whereas a cleanup failure is often secondary. * Throwing a cleanup exception while discarding the primary error makes debugging difficult, as the root cause of the initial failure is lost. * In a typical `try-finally` block, if the `finally` block throws an exception, it naturally suppresses any exception thrown in the `try` block unless handled manually. ### Implementing Better Suppression Logic A more robust implementation mimics the behavior of Kotlin’s `Closeable.use` by ensuring the most relevant error is thrown while keeping others accessible for debugging. * Instead of creating a wrapper class, use `Throwable.addSuppressed()` to attach the cleanup exception to the primary exception. * If only the primary block fails, throw that exception directly to satisfy the caller's `catch` requirements. * If both the primary block and the cleanup fail, throw the primary exception and add the cleanup exception as a suppressed error. * If only the cleanup fails, it is then appropriate to throw the cleanup exception as the standalone failure. ### Considerations for Checked and Unchecked Exceptions The impact of exception handling varies by language, particularly in Java where checked exceptions are enforced by the compiler. * Converting a checked exception into an unchecked `RuntimeException` inside a wrapper can cause the compiler to miss necessary error-handling requirements. * If exceptions have parent-child relationships, such as `IOException` and `Exception`, wrapping can cause a specific handler to be bypassed in favor of a more generic one. * It is generally recommended to only wrap checked exceptions in `RuntimeException` when the error is truly unrecoverable and the caller is not expected to handle it. When designing custom resource management utilities, always evaluate which exception is most critical for the caller to see. Prioritize the primary execution error and use suppression for auxiliary cleanup failures to ensure that your error-handling remains transparent and predictable for the rest of the application.

googleOriginal article

VaultGemma: The world's most capable differentially private LLM (opens in new tab)

VaultGemma represents a significant milestone in privacy-preserving AI as the most capable large language model trained from scratch using differential privacy (DP). By establishing new scaling laws specifically for DP training, researchers have optimized the complex trade-offs between compute, privacy budgets, and model utility. The resulting 1-billion-parameter model demonstrates that high-performance generative AI can be achieved while maintaining rigorous mathematical guarantees against data memorization. ## Scaling Laws for Differentially Private Training * Performance in DP-trained models is primarily governed by the "noise-batch ratio," which measures the amount of random privacy noise relative to the size of the training data groups. * Research suggests that for any given compute and privacy budget, there exists an optimal training configuration that balances model size, iterations, and batch size to achieve the lowest possible training loss. * A critical finding indicates that DP training requires a departure from standard scaling practices, favoring significantly larger batch sizes and smaller model architectures than traditional non-DP training. ## Synergies in Privacy, Compute, and Data * Increasing the privacy budget (epsilon) in isolation leads to diminishing returns unless it is paired with a proportional increase in compute (FLOPs) or data (tokens). * Visualizations of the scaling laws show that different model sizes can provide similar utility if the number of training iterations and batch sizes are correctly adjusted. * The optimal configuration shifts between investing in larger models versus more iterations depending on the specific constraints of the data and privacy budgets. ## Training at Scale with Algorithmic Advancements * VaultGemma is built on the Gemma 2 architecture and utilizes a 1B parameter setup optimized for the unique constraints of DP. * To overcome hardware limitations when processing the massive batch sizes required for DP training, the team developed a "Virtual Batch" technique in JAX to aggregate gradients across multiple steps. * Training from scratch allows the model to outperform traditional DP-finetuned models, which often struggle to balance utility with the noise introduced during the fine-tuning process. ## Performance and Evaluation * VaultGemma achieves competitive results against standard 1B parameter models while providing formal privacy protections. * The model demonstrates superior privacy-utility trade-offs, proving that carefully scaled DP models can retain high levels of reasoning and language capability. * The release includes the model weights and a comprehensive technical report to assist the community in developing the next generation of private-by-design AI. VaultGemma provides a practical blueprint for developers who need to balance the power of large language models with strict data confidentiality requirements. By leveraging the provided scaling insights, organizations can now train models that are mathematically resistant to data leakage without sacrificing significant performance.

figma3 min readCurated summary

Is the App Layer Where AI Proves Its Value? | Figma Blog

AI’s next breakthrough may come less from larger models than from the application layer that makes them useful and accessible. Like graphical interfaces made personal computers mainstream, well-designed AI products can translate complex capabilities into intuitive, context-specific experiences. The products that succeed will combine reliable infrastructure with thoughtful interaction design and emotional resonance. ## From MS-DOS to the App Layer - Today’s prompt-driven AI resembles the MS-DOS era: powerful, but requiring users to know how to issue precise commands. - Existing models have a “capabilities overhang,” meaning much of their potential remains difficult to access. - Personal computers became mainstream through graphical user interfaces, not MS-DOS itself. - Similarly, browsers, search engines, smartphone apps, and services such as Uber and Instagram transformed underlying technology into everyday tools. ## Design Makes Technology Adoptable - Building an app layer is not enough; adoption depends on the quality of the interactions surrounding the technology. - Successful products combine functionality with intuitive design: - Pinch-to-zoom and inertial scrolling on smartphones - Live maps in Uber - Simple navigation in browsers and search engines - AI products will need new interaction patterns that make model capabilities feel natural rather than like conversations with a raw chatbot. ## AI Products Must Be Context-Specific - Most people will use AI through specialized products rather than directly interacting with language models. - Effective AI applications will adapt their content, tone, interface, and responses to particular audiences and situations. - The Good Inside parenting app illustrates this approach: - It uses a chatbot trained on Dr. Becky’s parenting guidance. - Vague prompts receive empathetic, actionable advice. - Simple cards, a calm color palette, readable typography, and subtle animations create a reassuring experience. - The same principle applies to products for lawyers, doctors, designers, artists, and other professional or consumer groups. ## The Interface Can Matter More Than the Model - User reactions to GPT-5’s simplified model picker showed that interface changes can provoke stronger responses than improvements to model capability. - This does not make the underlying models unimportant, but users primarily experience AI through how its capabilities are packaged and presented. - Atlassian’s acquisition of The Browser Company suggests that even browsers may evolve into active AI interfaces that help applications work together, rather than merely displaying tabs. ## Design as a Competitive Advantage - AI products will compete on the feelings and confidence they create: - Support for parents - Inspiration for artists - Confidence for lawyers - Product teams must choose interactions that present AI outputs seamlessly while maintaining reliable, scalable systems. - Many new AI applications will emerge, but the strongest may distinguish themselves through design and become as transformative as graphical user interfaces were for computing. The practical opportunity for AI builders is to focus not only on model performance, but on designing specialized, emotionally resonant products that turn raw capability into useful everyday experiences.

Read original(opens in new tab)
googleOriginal article

Smarter nucleic acid design with NucleoBench and AdaBeam (opens in new tab)

Google Research and Move37 Labs have introduced NucleoBench, a comprehensive open-source benchmark for nucleic acid design, alongside AdaBeam, a high-performing new optimization algorithm. While AI models have become highly proficient at predicting the biological properties of DNA and RNA, generating optimal sequences within massive search spaces—such as the $2 \times 10^{120}$ possible variations for a 5' UTR—remains a significant hurdle. By standardizing evaluation across 16 distinct biological tasks, this research identifies AdaBeam as a superior method that scales effectively to the large-scale models required for modern drug discovery. ## Standardizing the Optimization Pipeline The process of computational nucleic acid design typically follows a five-step workflow: data collection, training a predictive model, generating candidate sequences (the design step), wet-lab validation, and iterative retraining. NucleoBench focuses specifically on the design step, which has historically lacked standardized evaluation. * Most existing benchmarks rely on decades-old methods like simulated annealing or vanilla genetic algorithms. * Traditional algorithms often treat predictive models as "black boxes," failing to leverage internal model data to guide the search. * The vastness of genomic search spaces makes brute-force optimization impossible, necessitating more intelligent, model-aware generation strategies. ## The NucleoBench Framework NucleoBench is the first large-scale benchmark designed to compare gradient-free and gradient-based design algorithms under identical conditions. The framework encompasses over 400,000 experiments to ensure statistical rigor across diverse biological challenges. * **Algorithm Categories**: It compares gradient-free methods (like directed evolution), which are simple but ignore model internals, against gradient-based methods (like FastSeqProp), which use the model’s internal "direction of steepest improvement" to find better sequences. * **Task Diversity**: The 16 tasks include controlling gene expression in specific cell types (liver or neuronal), maximizing transcription factor binding, and improving chromatin accessibility. * **Scale**: The benchmark includes long-range DNA sequence challenges using large-scale models like Enformer, which are computationally demanding but critical for understanding complex genomic interactions. ## AdaBeam’s Hybrid Optimization Performance Drawing on insights from the NucleoBench evaluation, the researchers developed AdaBeam, a hybrid algorithm that combines the strengths of various optimization strategies. * **Success Rate**: AdaBeam outperformed existing algorithms on 11 of the 16 tasks in the benchmark. * **Efficiency and Scaling**: Unlike many gradient-based methods that struggle with computational overhead, AdaBeam demonstrates superior scaling properties as sequences become longer and predictive models grow in complexity. * **Methodology**: It functions as a hybrid approach, using sophisticated search techniques to navigate the sequence space more effectively than "vanilla" algorithms developed before the era of deep learning. The researchers have made AdaBeam and the NucleoBench repository freely available to the scientific community. By providing a standardized environment for testing, they aim to accelerate the development of next-generation treatments, including more stable mRNA vaccines and precise CRISPR gene therapies.

googleOriginal article

Speculative cascades — A hybrid approach for smarter, faster LLM inference (opens in new tab)

Speculative cascades represent a hybrid inference method that integrates the cost-efficiency of model cascades with the latency-reducing benefits of speculative decoding. By utilizing a smaller drafter model to generate token sequences that are verified in parallel by a larger expert model, this approach allows for high-speed generation while maintaining flexible quality standards. The result is a system that achieves superior cost-quality trade-offs and higher speed-ups than either traditional cascading or standard speculative decoding alone. ### Limitations of Cascades and Speculative Decoding * **Sequential Bottlenecks in Cascades:** Traditional cascades use a deferral rule to decide if a small model can handle a prompt. If the small model is not confident, the system waits for it to finish before starting the large model from scratch, wasting significant time. * **Strict Matching in Speculative Decoding:** This method requires the large model to verify the small model’s tokens. Even if the small model produces a factually correct and high-quality response, the large model will reject the entire draft if the tokens do not match its own preferred output exactly. * **Trade-off Divergence:** Cascades prioritize reducing computational costs but suffer from latency when deferring, while speculative decoding prioritizes speed but often performs redundant work because it mandates identical output to the larger model. ### The Speculative Cascades Mechanism * **Parallel Verification with Deferral:** Speculative cascades use the parallel processing of speculative decoding but introduce a flexible decision rule. The system can choose to accept the smaller model’s draft even if it differs from the larger model’s prediction, provided it meets a confidence threshold. * **Flexible Token Matching:** Unlike standard speculative decoding, which often relies on strict token-by-token matching, speculative cascades allow for "probabilistic matches" or quality-based acceptance to prevent unnecessary rejections. * **Resource Optimization:** By strategically deferring to the smaller model for certain segments of the generation, the system reduces the total work required from the expensive expert model without losing the speed of parallel execution. ### Empirical Results and Performance * **Model Testing:** The approach was validated using Gemma and T5 models across diverse language tasks, including reasoning, coding, translation, and question answering. * **Superior Trade-offs:** Testing showed that speculative cascades consistently outperformed baselines in cost-quality metrics, providing faster inference without the strict "all-or-nothing" quality constraints of speculative decoding. * **Task Versatility:** The hybrid method proved effective across both creative tasks (like summarization) and factual tasks (like math or coding), where different levels of "correctness" are acceptable. Speculative cascades offer a practical path for scaling LLM deployments by balancing the high cost of large models with the need for low-latency user experiences. Developers looking to optimize inference should consider this hybrid approach to capture the efficiency of small models while retaining the oversight of larger, more capable ones.

figma2 min readCurated summary

Issue No.12: New Roles, New Rules | Figma Blog

Figma’s “New roles, new rules” highlights how AI and faster iteration are blurring traditional boundaries between product roles. Product managers, designers, and developers are increasingly working across disciplines, using prototypes and shared principles to collaborate directly. The issue argues that effective teams are replacing rigid handoffs with experimentation, co-creation, and better systems for guiding AI. ## Shifting Roles - Research found that: - 64% of product builders identify with two or more roles. - 56% of non-designers perform design-related work. - Faster development cycles and AI tools are enabling people to contribute further outside their formal specialties. - As responsibilities expand, teams are also reconsidering how they manage time, ownership, and collaboration. ## Prototyping to Create Shared Understanding - Figma product designer Natasha Tenggoro struggled to explain how video playback should work in Figma Buzz. - Instead of relying on verbal descriptions, she used Figma Make to build prototypes herself. - The prototypes helped the team reach three “aha” moments and provided a clearer basis for discussion. - The example demonstrates how building interactive artifacts can replace lengthy explanations and accelerate alignment. ## Music-Inspired Design in Figma Draw - Figma Draw’s new scatter brushes were shaped by musical concepts such as tempo, texture, and volume. - Designers translated the character of genres including Honky-tonk, Screamo, Doo-wop, and Vaporwave into brush behavior. - The release adds 10 scatter brushes, giving designers more control over attributes such as gap, wiggle, and jitter. ## Duolingo’s Collaborative Method - Duolingo’s Math team is changing the traditional design-to-engineering handoff for its math games. - Designers and engineers work through co-creation, scrappy prototypes, and continuous experimentation. - Shared principles—including “show, don’t tell” and distinguishing between a v1 and an MVP—help the team move quickly while refining ideas. - Collaboration is treated as an ongoing product practice rather than a stage that ends when design is handed off. ## Broader Changes in Creative Work - Agencies and freelancers are also abandoning rigid client boundaries and involving clients throughout the creative process. - Design systems can improve AI-generated code by giving agents structured, relevant, and brand-consistent input. - MCP servers are presented as an important connection between design systems and AI-powered workflows, helping agents produce more useful output. Overall, the issue recommends embracing broader roles and replacing formal handoffs with shared prototypes, collaborative iteration, and well-structured design systems—especially as AI becomes more involved in product development.

Read original(opens in new tab)
figma3 min readCurated summary

Are Roles and Responsibilities a Thing of the Past? | Figma Blog

Product teams are moving away from rigid job boundaries toward more fluid, overlapping responsibilities. Figma’s research, based on 51 interviews and a survey of 1,199 product professionals, found that 64% of respondents identify with at least two roles, while more than a third span three or more. This shift can improve collaboration and speed, but it also creates friction through tool overload and requires teams to coordinate more deliberately. ## Roles Are Becoming More Fluid - Product development responsibilities increasingly overlap across design, product management, engineering, research, data, and marketing. - 56% of non-designers say they participate heavily in at least one design-related task. - PMs are prototyping ideas, engineers are contributing to early design decisions, and marketers and content specialists are commenting directly in design files. - Figma conducted the research with Factworks and Fusion Hill through 51 qualitative interviews and a survey of 1,199 participants. ## Design Isn’t Just for Designers - Non-designer participation in design tasks such as mockups and brand exploration rose by 10% in the past year. - 70% of product managers create low-fidelity mockups or wireframes, and 59% create interactive prototypes. - One in four product builders has recently adopted a new design tool; 42% of those who have not yet adopted one plan to do so within a year. - Cross-functional design work can: - Help teams align on ideas earlier. - Expose technical feasibility problems before implementation. - Reduce back-and-forth and save time. - Designers can use the time saved for research, strategy, design-system maintenance, and craft. - The goal is not to eliminate design expertise, but to encourage shared visual communication and earlier collaboration. - Teams can support this transition by building stronger relationships, establishing a shared vision early, and welcoming contributions from colleagues who are still learning design tools. ## Tool Overload Is Causing Friction - 72% of respondents identify AI tools as the main force changing their roles. - Tools such as Figma Make, Claude Code, and GitHub Copilot make prototyping and code generation more accessible. - Notion, Figma Buzz, and Airtable similarly support rapid marketing production and project management. - Role expansion has led 71% of respondents to use more tools and software. - Participants use tools across an average of 7.6 out of 16 categories, including spreadsheets, project management, graphic design, and coding assistants. - Although new tools enable more people to contribute, the growing number of platforms can make workflows harder to manage and reduce the effectiveness of individual tools. Teams should treat overlapping roles as an opportunity for earlier alignment and broader participation, while actively simplifying their tool ecosystems. Clear collaboration practices—not rigid ownership boundaries—will be essential as product work continues to evolve.

Read original(opens in new tab)
googleOriginal article

Accelerating scientific discovery with AI-powered empirical software (opens in new tab)

Google Research has introduced an AI-powered system designed to accelerate scientific discovery by automating the creation and optimization of "empirical software." By leveraging the Gemini model and tree search optimization, the system can propose, implement, and iteratively improve code for complex multidisciplinary challenges, achieving results that match or exceed human expert performance. This approach transforms scientific hypothesis evaluation from a months-long manual coding process into an automated search that can be completed in hours or days. ### The Concept of Empirical Software and Scorable Tasks * The system shifts focus from traditional functional correctness to "empirical software," where the primary objective is to maximize a predefined quality score. * It targets "scorable tasks," which are defined by a problem description, a specific scoring metric, and a dataset for training and validation. * This framework addresses the research bottleneck where scientists must manually test hundreds of models or parameters to achieve a breakthrough. ### System Architecture and Optimization Strategy * The engine takes a task description and optional context—such as ideas from scientific literature—as input to generate novel methodological concepts. * It utilizes a tree search strategy inspired by AlphaZero, employing an upper confidence bound to navigate and prioritize thousands of potential code variants. * The LLM acts as an iterative rewriter, refining executable code within a sandbox to continuously improve the performance score. * Outputs are designed to be fully verifiable, interpretable, and reproducible, providing scientists with the specific coded solutions used to reach a result. ### Demonstrated Performance Across Scientific Domains * The system was tested on six diverse benchmarks, including genomics, public health, geospatial analysis, neuroscience, and time-series forecasting. * In genomics, the system tackled the "batch integration" of single-cell RNA sequencing (scRNA-seq) data, a complex problem involving the removal of noise while preserving biological signals. * The AI discovered 40 novel methods that outperformed top expert-developed tools within the OpenProblems V2.0.0 batch integration benchmark. * Evaluation focused on advanced capabilities such as zero-shot generalization, high-dimensional signal processing, and uncertainty quantification. This system represents a significant shift toward "research engines" that participate actively in the scientific method through iterative experimentation. Scientists can utilize these tools to explore a much broader range of hypotheses than manual coding allows, potentially leading to faster breakthroughs in data-heavy fields like genomics and climate modeling.