Generative AI

125 posts

figma3 min readCurated summary

Is the App Layer Where AI Proves Its Value? | Figma Blog

AI’s next breakthrough may come less from larger models than from the application layer that makes them useful and accessible. Like graphical interfaces made personal computers mainstream, well-designed AI products can translate complex capabilities into intuitive, context-specific experiences. The products that succeed will combine reliable infrastructure with thoughtful interaction design and emotional resonance. ## From MS-DOS to the App Layer - Today’s prompt-driven AI resembles the MS-DOS era: powerful, but requiring users to know how to issue precise commands. - Existing models have a “capabilities overhang,” meaning much of their potential remains difficult to access. - Personal computers became mainstream through graphical user interfaces, not MS-DOS itself. - Similarly, browsers, search engines, smartphone apps, and services such as Uber and Instagram transformed underlying technology into everyday tools. ## Design Makes Technology Adoptable - Building an app layer is not enough; adoption depends on the quality of the interactions surrounding the technology. - Successful products combine functionality with intuitive design: - Pinch-to-zoom and inertial scrolling on smartphones - Live maps in Uber - Simple navigation in browsers and search engines - AI products will need new interaction patterns that make model capabilities feel natural rather than like conversations with a raw chatbot. ## AI Products Must Be Context-Specific - Most people will use AI through specialized products rather than directly interacting with language models. - Effective AI applications will adapt their content, tone, interface, and responses to particular audiences and situations. - The Good Inside parenting app illustrates this approach: - It uses a chatbot trained on Dr. Becky’s parenting guidance. - Vague prompts receive empathetic, actionable advice. - Simple cards, a calm color palette, readable typography, and subtle animations create a reassuring experience. - The same principle applies to products for lawyers, doctors, designers, artists, and other professional or consumer groups. ## The Interface Can Matter More Than the Model - User reactions to GPT-5’s simplified model picker showed that interface changes can provoke stronger responses than improvements to model capability. - This does not make the underlying models unimportant, but users primarily experience AI through how its capabilities are packaged and presented. - Atlassian’s acquisition of The Browser Company suggests that even browsers may evolve into active AI interfaces that help applications work together, rather than merely displaying tabs. ## Design as a Competitive Advantage - AI products will compete on the feelings and confidence they create: - Support for parents - Inspiration for artists - Confidence for lawyers - Product teams must choose interactions that present AI outputs seamlessly while maintaining reliable, scalable systems. - Many new AI applications will emerge, but the strongest may distinguish themselves through design and become as transformative as graphical user interfaces were for computing. The practical opportunity for AI builders is to focus not only on model performance, but on designing specialized, emotionally resonant products that turn raw capability into useful everyday experiences.

Read original(opens in new tab)
figma2 min readCurated summary

How to Harness Skills That AI Can’t Automate | Figma Blog

Great products require more than AI-generated functionality; they depend on human craft, including curiosity, intuition, taste, and intention. AI accelerates prototyping and iteration, but people must decide what questions to ask, when to expand the scope, and whether an experience truly resonates. The article argues that these uniquely human skills are essential for steering AI toward thoughtful, high-quality outcomes. ## Curiosity: Starting from First Principles - Curiosity means asking “why” and “what if,” exploring multiple directions, and validating ideas through iteration. - Tools such as Figma Make, Claude, and ChatGPT make it faster to generate prototypes, campaign concepts, and alternative approaches. - AI can expand possibilities within a prompt, but it cannot independently challenge the original scope, follow an unexpected hunch, or recognize that a tangent deserves deeper investigation. - Early exploration improves confidence and reduces rework: - Figma designer Natasha Tenggoro used Figma Make to test video-playback concepts for Figma Buzz. - Prototyping exposed edge cases, clarified feasibility, and helped engineers understand the feature before implementation. - Curiosity can also lead to valuable scope expansion: - While designing icons for four new Figma products, Tim Van Damme realized the existing icon family needed a broader redesign. - He explored hundreds of variations and redesigned the suite for greater consistency, distinctiveness, scalability, and a unified Figma identity. ## Intuition: Following a Feeling - AI can help bring products to market quickly, but it cannot determine what will emotionally or practically resonate with customers. - Human designers can recognize when an interaction feels confusing even if it technically works. - Intuition may also identify qualitative improvements—such as adding white space to a crowded layout—that are difficult to express as strict requirements. - Eliel Johnson of CVS Health describes design as an active practice centered on asking whether an experience “feels good,” not merely whether it satisfies functional constraints. The practical lesson is to use AI for speed and breadth while reserving human judgment for direction, exploration, emotional resonance, and quality. Human craft is what turns technically workable output into a product people genuinely value.

Read original(opens in new tab)
figma2 min readCurated summary

IDC Study Says the Global Workforce Engaged in Software Design Is Expanding | Figma Blog

IDC forecasts that the global workforce involved in software design and development will grow by more than 30%, reaching 144 million people by 2029. This expansion reflects the growing role of design as a competitive differentiator and the increasing demand for design talent. Generative AI is expected to accelerate product development while raising expectations for usability and visual quality. ## Workforce Growth in Software Design - The population of knowledge workers and developers designing digital products and interfaces is projected to grow from **107 million in 2025 to 144 million in 2029**. - The research covers multiple industries, company sizes, and roles involved in software design and product development. - IDC based its model on quantitative research involving more than 1,000 IT and business leaders. ## UX Design Leads Growth - UX professionals are expected to grow at a **7.6% compound annual growth rate** between 2025 and 2029. - This growth is expected to outpace other knowledge-worker categories. - IDC links the trend to rising demand for design-intensive digital products and solutions. ## Generative AI Raises the Bar - Generative AI will increase the speed and volume of software development. - Faster development may give users more product choices and enable better-designed digital experiences. - Designers and developers will face greater pressure to make AI-powered features highly usable, appealing, and intuitive. Overall, the study suggests that companies should treat design as a strategic capability and invest in UX talent as software development becomes faster and more AI-driven.

Read original(opens in new tab)
googleOriginal article

How Google’s AI can help transform health professions education (opens in new tab)

To address a projected global deficit of 11 million healthcare workers by 2030, Google Research is exploring how generative AI can provide personalized, competency-based education for medical professionals. By combining qualitative user-centered design with quantitative benchmarking of the pedagogically fine-tuned LearnLM model, researchers have demonstrated that AI can effectively mimic the behaviors of high-quality human tutors. The studies conclude that specialized models, now integrated into Gemini 2.5 Pro, can significantly enhance clinical reasoning and adapt to the individual learning styles of medical students. ## Learner-Centered Design and Participatory Research * Researchers conducted interdisciplinary co-design workshops featuring medical students, clinicians, and AI researchers to identify specific educational needs. * The team developed a rapid prototype of an AI tutor designed to guide learners through clinical reasoning exercises anchored in synthetic clinical vignettes. * Qualitative feedback from medical residents and students highlighted a demand for "preceptor-like" behaviors, such as the ability to manage cognitive load, provide constructive feedback, and encourage active reflection. * Analysis revealed that learners specifically value AI tools that can identify and bridge individual knowledge gaps rather than providing generic information. ## Quantitative Benchmarking via LearnLM * The study utilized LearnLM, a version of Gemini fine-tuned specifically for educational pedagogy, and compared its performance against Gemini 1.5 Pro. * Evaluations were conducted using 50 synthetic scenarios covering a spectrum of medical education, ranging from preclinical topics like platelet activation to clinical subjects such as neonatal jaundice. * Medical students engaged in 290 role-playing conversations, which were then evaluated based on four primary metrics: overall experience, meeting learning needs, enjoyability, and understandability. * Physician educators performed blinded reviews of conversation transcripts to assess whether the AI adhered to medical education standards and core competencies. ## Pedagogical Performance and Expert Evaluation * LearnLM was consistently rated higher than the base model by both students and educators, with experts noting it behaved "more like a very good human tutor." * The fine-tuned model demonstrated a superior ability to maintain a conversation plan and use grounding materials to provide accurate, context-aware instruction. * Findings suggest that pedagogical fine-tuning is essential for AI to move beyond simple fact-delivery and toward true interactive tutoring. * These specialized learning capabilities have been transitioned from the research phase into Gemini 2.5 Pro to support broader educational applications. By integrating these specialized AI behaviors into medical training pipelines, institutions can provide scalable, individualized support to students. The transition of LearnLM’s pedagogical features into Gemini 2.5 Pro provides a practical framework for developers to create tools that not only provide medical information but actively foster the critical thinking skills required for clinical practice.

googleOriginal article

From massive models to mobile magic: The tech behind YouTube real-time generative AI effects (opens in new tab)

YouTube has successfully deployed over 20 real-time generative AI effects by distilling the capabilities of massive cloud-based models into compact, mobile-ready architectures. By utilizing a "teacher-student" training paradigm, the system overcomes the computational bottlenecks of high-fidelity generative AI while ensuring the output remains responsive on mobile hardware. This approach allows for complex transformations, such as cartoon style transfer and makeup application, to run frame-by-frame on-device without sacrificing the user’s identity. ### Data Curation and Diversity * The foundation of the effects pipeline relies on high-quality, properly licensed face datasets. * Datasets are meticulously filtered to ensure a uniform distribution across different ages, genders, and skin tones. * The Monk Skin Tone Scale is used as a benchmark to ensure the effects work equitably for all users. ### The Teacher-Student Framework * **The Teacher:** A large, powerful pre-trained model (initially StyleGAN2 with StyleCLIP, later transitioning to Google DeepMind’s Imagen) acts as the "expert" that generates high-fidelity visual effects. * **The Student:** A lightweight UNet-based architecture designed for mobile efficiency. It utilizes a MobileNet backbone for both the encoder and decoder to ensure fast frame-by-frame processing. * The distillation process narrows the scope of the massive teacher model into a student model focused on a single, specific task. ### Iterative Distillation and Training * **Data Generation:** The teacher model processes thousands of images to create "before and after" pairs. These are augmented with synthetic elements like AR glasses, sunglasses, and hand occlusions to improve real-world robustness. * **Optimization:** The student model is trained using a sophisticated combination of loss functions, including L1, LPIPS, Adaptive, and Adversarial loss, to balance numerical accuracy with aesthetic quality. * **Architecture Search:** Neural architecture search is employed to tune "depth" and "width" multipliers, identifying the most efficient model structure for different mobile hardware constraints. ### Addressing the Inversion Problem * A major challenge in real-time effects is the "inversion problem," where the model struggles to represent a real face in latent space, leading to a loss of the user's identity (e.g., changes in skin tone or clothing). * YouTube uses Pivotal Tuning Inversion (PTI) to ensure that the user's specific features are preserved during the generative process. * By editing images in the latent space—a compressed numerical representation—the system can apply stylistic changes while maintaining the core characteristics of the original video stream. By combining advanced model distillation with on-device optimization via MediaPipe, YouTube demonstrates a practical path for bringing heavy generative AI research into consumer-facing mobile applications.

googleOriginal article

Enabling physician-centered oversight for AMIE (opens in new tab)

Guardrailed-AMIE (g-AMIE) is a diagnostic AI framework designed to perform patient history-taking while strictly adhering to safety guardrails that prevent it from providing direct medical advice. By decoupling data collection from clinical decision-making, the system enables an asynchronous oversight model where primary care physicians (PCPs) review and finalize AI-generated medical summaries. In virtual clinical trials, g-AMIE’s diagnostic outputs and patient communications were preferred by overseeing physicians and patient actors over human-led control groups. ## Multi-Agent Architecture and Guardrails * The system utilizes a multi-agent setup powered by Gemini 2.0 Flash, consisting of a dialogue agent, a guardrail agent, and a SOAP note agent. * The dialogue agent conducts history-taking in three distinct phases: general information gathering, targeted validation of a differential diagnosis, and a conclusion phase for patient questions. * A dedicated guardrail agent monitors and rephrases responses in real-time to ensure the AI abstains from sharing individualized diagnoses or treatment plans directly with the patient. * The SOAP note agent employs sequential multi-step generation to separate summarization tasks (Subjective and Objective) from more complex inferential tasks (Assessment and Plan). ## The Clinician Cockpit and Asynchronous Oversight * To facilitate human review, researchers developed the "clinician cockpit," a web interface co-designed with outpatient physicians through semi-structured interviews. * The interface is structured around the standard SOAP note format, presenting the patient’s perspective, measurable data, differential diagnosis, and proposed management strategy. * This framework allows overseeing PCPs to review cases asynchronously, editing the AI’s proposed differential diagnoses and management plans before sharing a final message with the patient. * The separation of history-taking from decision-making ensures that licensed medical professionals retain ultimate accountability for patient care. ## Performance Evaluation via Virtual OSCE * The system was evaluated in a randomized, blinded virtual Objective Structured Clinical Examination (OSCE) involving 60 case scenarios. * g-AMIE’s performance was compared against primary care physicians, nurse practitioners, and physician assistants who were required to operate under the same restrictive guardrails. * Overseeing PCPs and independent physician raters preferred g-AMIE’s diagnostic accuracy and management plans over those of the human control groups. * Patient actors reported a preference for the messages generated by g-AMIE compared to those drafted by human clinicians in the study. While g-AMIE demonstrates high potential for human-AI collaboration in diagnostics, the researchers emphasize that results should be interpreted with caution. The workflow was specifically optimized for AI characteristics, and human clinicians may require specialized training to perform effectively within such highly regulated guardrail frameworks.

lineOriginal article

Hey, won't you become a (opens in new tab)

Hack Day 2025 serves as a cornerstone of LY Corporation’s engineering culture, bringing together diverse global teams to innovate beyond their daily operational scopes. By fostering a high-intensity environment focused on creative freedom, the event facilitates technical growth and strengthens interpersonal bonds across international branches. This 19th edition demonstrated how rapid prototyping and cross-functional collaboration can transform abstract ideas into functional AI-driven prototypes within a strict 24-hour window. ### Structure and Participation Dynamics * The hackathon follows a "9 to 9" format, providing exactly 24 hours of development time followed by a day for presentations and awards. * Participation is inclusive of all roles, including developers, designers, planners, and HR staff, allowing for holistic product development. * Teams can be "General Teams" from the same legal entity or "Global Mixed Teams" comprising members from different regions like Korea, Japan, Taiwan, and Vietnam. * The Developer Relations (DevRel) team facilitates team building for remote employees using digital collaboration tools like Zoom and Miro. ### AI-Powered Personality Analysis Project * The author's team developed a "Scouter" program inspired by Dragon Ball, designed to measure professional "combat power" based on communication history. * The system utilizes Slack bots and AI models to analyze message logs and map them to the Big 5 Personality traits (Openness, Conscientiousness, Extraversion, Agreeableness, and Neuroticism). * Professional metrics are visualized as game-like character statistics to make personality insights engaging and less intimidating. * While the original plan involved using AI to generate and print physical character cards, hardware failures with photo printers forced a technical pivot to digital file downloads. ### High-Pressure Presentation and Networking * Every team is allotted a strict 90-second window to pitch their product and demonstrate a live demo. * The "90-second rule" includes a mandatory microphone cutoff to maintain momentum and keep the large-scale event engaging for all attendees. * Dedicated booth sessions follow the presentations, allowing participants to provide hands-on experiences to colleagues and judges. * The event emphasizes "Perfect the Details," a core company value, by encouraging teams to utilize all available resources—from whiteboards to AI image generators—within the time limit. ### Environmental Support and Culture * The event occupies an entire office floor, providing a high-density yet comfortable environment designed to minimize distractions during the "Hack Time." * Cultural exchange is encouraged through "humanity snacks," where participants from different global offices share local treats in dedicated rest areas. * Strategic scheduling, such as "Travel Days" for international participants, ensures that teams can focus entirely on technical execution once the event begins. Participating in internal hackathons provides a vital platform for testing new technologies—like LLMs and personality modeling—that may not fit into immediate product roadmaps. For organizations with hybrid work models, these intensive in-person events are highly recommended to bridge the communication gap and build lasting trust between global teammates.

figma3 min readCurated summary

Figma’s IPO: Design Is Everyone’s Business | Figma Blog

Figma’s IPO marks a new phase for the company, but not a change in its founding mission: narrowing the gap between imagination and reality. CEO Dylan Field argues that AI will make design more accessible and important, while emphasizing that Figma will prioritize decades-long growth over short-term efficiency or share-price performance. He sees Figma as a collaborative platform where more people can shape products and ideas. ## Figma’s IPO and Long-Term Mission - Figma went public to improve corporate governance, increase brand awareness, provide liquidity, strengthen its acquisition currency, and access capital markets. - Field especially values public ownership because it allows the broader Figma community to share in the company’s success. - He cautions investors that public-market performance is unpredictable and does not promise share-price growth. - Figma will prioritize supporting designers’ evolving needs and pursuing long-term growth over maximizing quarterly efficiency. - The company expects to take significant risks, including large platform investments and mergers and acquisitions. ## AI as a Strategic Investment - Figma is investing heavily in AI and plans to increase that investment further. - This spending may reduce efficiency for several years, but Field considers AI central to the future of design workflows. - Existing capabilities, including Figma Make and other AI features, are presented as only the beginning. - AI could help designers work more effectively and bring more people into the design process. ## Design as a Competitive Advantage - Creating a minimum viable product is easier than ever, making design, craft, and distinctive perspective more important differentiators. - Design is no longer an afterthought focused only on form and function; it can determine whether a product succeeds or fails. - As design becomes more central, companies will involve more types of contributors and encourage greater experimentation and creativity. - Figma must balance accessibility with professional power while supporting collaboration and decision-making across large organizations. ## The Evolution of AI Interfaces - Field compares current AI interfaces, which rely heavily on prompts, to the MS-DOS era of computing. - He expects new, domain-specific design patterns to make AI capabilities easier and more intuitive to use. - Just as graphical user interfaces expanded access to computers, well-designed AI interfaces could make advanced capabilities available to everyday users. ## Figma’s Role in the Future - Field describes Figma as a “peaceful garden” where individuals and teams can develop ideas together. - He believes tools alone do not change the world; people use them to create meaningful change. - Figma’s long-term ambition is to help more people participate in design and turn ideas into reality. Figma’s direction is therefore one of patient, ambitious investment: accept near-term inefficiency, expand access to design, and build AI-powered tools for a much broader creative community.

Read original(opens in new tab)
lineOriginal article

LY's Tech Conference, ' (opens in new tab)

LY Corporation’s Tech-Verse 2025 conference highlighted the company's strategic pivot toward becoming an AI-centric organization through the "Catalyst One Platform" initiative. By integrating the disparate infrastructures of LINE and Yahoo! JAPAN into a unified private cloud, the company aims to achieve massive cost efficiencies while accelerating the deployment of AI agents across its entire service ecosystem. This transformation focuses on empowering engineers with AI-driven development tools to foster rapid innovation and deliver a seamless, "WOW" experience for global users. ### Infrastructure Integration and the Catalyst One Platform To address the redundancies following the merger of LINE and Yahoo! JAPAN, LY Corporation is consolidating its technical foundations into a single internal ecosystem known as the Catalyst One Platform. * **Private Cloud Advantage:** The company maintains its own private cloud to achieve a four-fold cost reduction compared to public cloud alternatives, managed by a lean team of 700 people supporting 500,000 servers. * **Unified Architecture:** The integration spans several layers, including Infrastructure (Project "DC-Hub"), Cloud (Project "Flava"), and specialized Data and AI platforms. * **Next-Generation Cloud "Flava":** This platform integrates existing services to enhance VM specifications, VPC networking, and high-performance object storage (Ceph and Dragon). * **Information Security:** A dedicated "SafeOps" framework is being implemented to provide governance and security across all integrated services, ensuring a safer environment for user data. ### AI Strategy and Service Agentization A core pillar of LY’s strategy is the "AI Agentization" of all its services, moving beyond simple features to proactive, personalized assistance. * **Scaling GenAI:** Generative AI has already been integrated into 44 different services within the group. * **Personalized Agents:** The company is developing the capacity to generate millions of specialized agents that can be linked together to support the unique needs of individual users. * **Agent Ecosystem:** The goal is to move from a standard platform model to one where every user interaction is mediated by an intelligent agent. ### AI-Driven Development Transformation Beyond user-facing services, LY is fundamentally changing how its engineers work by deploying internal AI development solutions to all staff starting in July. * **Code and Test Automation:** Proof of Concept (PoC) results showed a 96% accuracy rate for "Code Assist" and a 97% reduction in time for "Auto Test" procedures. * **RAG Integration:** The system utilizes Retrieval-Augmented Generation (RAG) to leverage internal company knowledge and guidelines, ensuring high-quality, context-aware development support. * **Efficiency Gains:** By automating repetitive tasks, the company intends for engineers to shift their focus from maintenance to creative service improvement and innovation. The successful integration of these platforms and the aggressive adoption of AI-driven development tools suggest that LY Corporation is positioning itself to be a leader in the "AI-agent" era. For technical organizations, LY's model serves as a case study in how large-scale mergers can leverage private cloud infrastructure to fund and accelerate a company-wide AI transition.

microsoft3 min readCurated summary

Enhancing Code Quality at Scale with AI-Powered Code Reviews

Microsoft developed an AI-powered pull request reviewer to reduce routine review work, catch defects earlier, and help developers merge code faster. What began as an internal experiment now supports more than 90% of Microsoft’s PRs—over 600,000 per month—and has influenced GitHub’s Copilot for Pull Request Reviews. The central lesson is that AI works best as a human-in-the-loop assistant embedded directly into existing workflows. ## Addressing PR Review Bottlenecks - Human reviewers often spend time on style issues and minor bugs while overlooking architectural or security concerns. - Large, multi-file PRs can lack sufficient context and may wait days or weeks for review. - The AI reviewer automatically joins new PRs and handles repetitive or easily missed checks, allowing humans to focus on higher-level decisions. ## AI-Powered Review Features - **Automated comments:** Flags issues such as missing null checks, error-handling problems, sensitive-data risks, inefficient algorithms, and style inconsistencies. - **Suggested fixes:** Provides corrected snippets or alternative implementations, but authors must explicitly review and apply changes. AI does not commit changes automatically. - **PR summaries:** Generates descriptions of the change and highlights key modifications across the diff. - **Interactive Q&A:** Reviewers can ask questions about parameters, code behavior, or the impact on other modules directly in the PR discussion. - **Workflow integration:** The assistant behaves like a normal reviewer, requiring no separate tools or interfaces and optionally engaging as soon as a PR is opened. ## Effects on Quality and Development Speed - AI-assisted reviews reduced median PR completion times by 10–20% in early studies across 5,000 repositories. - Early feedback reduces waiting time, back-and-forth cycles, and the chance that minor issues delay approval. - The system has identified bugs such as missing null checks and incorrectly ordered API calls before they reached production. - Developers, particularly new hires, can use the explanations as continuous guidance on coding standards and best practices. ## Team-Specific Customization - Teams can configure repository-specific review guidelines. - Custom prompts support specialized checks, including regression detection based on historical crash patterns and validation of deployment or change gates. - This extensibility allows the reviewer to address concerns beyond generic code quality rules. ## Feedback Between Internal and External Products - Microsoft’s internal deployment provided early feedback on review quality, usability, and developer trust. - Internal experiments helped shape features such as inline suggestions and human-controlled change application. - These lessons contributed to GitHub Copilot for Pull Request Reviews, which reached general availability in April 2025. - Microsoft also uses learnings from GitHub’s broader external adoption to improve its internal development practices, creating an ongoing feedback loop between first-party and third-party products. Overall, the post recommends treating AI review as an always-available first pass—not a replacement for human judgment. Its greatest value comes from seamless integration, strong customization, and keeping authors and reviewers accountable for final decisions.

Read original(opens in new tab)
googleOriginal article

MedGemma: Our most capable open models for health AI development (opens in new tab)

Google Research has expanded its Health AI Developer Foundations (HAI-DEF) collection with the release of MedGemma and MedSigLIP, a series of open, multimodal models designed specifically for medical research and application development. These models offer a high-performance, privacy-preserving alternative to closed systems, allowing developers to maintain full control over their infrastructure while leveraging state-of-the-art medical reasoning. By providing both 4B and 27B parameter versions, the collection balances computational efficiency with complex longitudinal data interpretation, even enabling deployment on single GPUs or mobile hardware. ## MedGemma Multimodal Variants The MedGemma collection utilizes the Gemma 3 architecture to process both image and text inputs, providing robust generative capabilities for healthcare tasks. * **MedGemma 27B Multimodal:** This model is designed for complex tasks such as interpreting longitudinal electronic health records (EHR) and achieves an 87.7% score on the MedQA benchmark, performing within 3 points of DeepSeek R1 at approximately one-tenth the inference cost. * **MedGemma 4B Multimodal:** A lightweight version that scores 64.4% on MedQA, outperforming most open models under 8B parameters; it is optimized for mobile hardware and specific tasks like chest X-ray report generation. * **Clinical Accuracy:** In unblinded studies, 81% of chest X-ray reports generated by the 4B model were judged by board-certified radiologists to be sufficient for patient management, achieving a RadGraph F1 score of 30.3. * **Versatility:** The models retain general-purpose capabilities from the original Gemma base, ensuring they remain effective at instruction-following and non-English language tasks while handling specialized medical data. ## MedSigLIP Specialized Image Encoding MedSigLIP serves as the underlying vision component for the MedGemma suite, but it is also available as a standalone 400M parameter encoder for structured data tasks. * **Architecture:** Based on the Sigmoid loss for Language Image Pre-training (SigLIP) framework, it bridges the gap between medical imagery and text through a shared embedding space. * **Diverse Modalities:** The encoder was fine-tuned on a wide variety of medical data, including fundus photography, dermatology images, histopathology patches, and chest X-rays. * **Functional Use Cases:** It is specifically recommended for tasks involving classification, retrieval, and search, where structured outputs are preferred over free-text generation. * **Data Retention:** Training protocols ensured the model retained its ability to process natural images, maintaining its utility for hybrid tasks that mix medical and non-medical visual information. ## Technical Implementation and Accessibility Google has prioritized accessibility for developers by ensuring these models can run on consumer-grade or limited hardware environments. * **Hardware Compatibility:** Both the 4B and 27B models are designed to run on a single GPU, while the 4B and MedSigLIP versions are adaptable for edge computing and mobile devices. * **Open Resources:** To support the community, Google has released the technical reports, model weights on Hugging Face, and implementation code on GitHub. * **Developer Flexibility:** Because these are open models, researchers can fine-tune them on proprietary datasets without compromising data privacy or being locked into specific cloud providers. For medical AI development, the choice of model should depend on the specific output requirement: MedGemma is the optimal starting point for generative tasks like visual question answering or report drafting, while MedSigLIP is the preferred tool for building high-speed classification and image retrieval systems.

lineOriginal article

Hosting the Tech Conference Tech- (opens in new tab)

LY Corporation is hosting its global technology conference, Tech-Verse 2025, on June 30 and July 1 to showcase the engineering expertise of its international teams. The event features 127 sessions centered on core themes of AI and security, offering a deep dive into how the group's developers, designers, and product managers solve large-scale technical challenges. Interested participants can register for free on the official website to access the online live-streamed sessions, which include real-time interpretation in English, Korean, and Japanese. ### Conference Overview and Access * The event runs for two days, from 10:00 AM to 6:00 PM (KST), and is primarily delivered via online streaming. * Registration is open to the public at no cost through the Tech-Verse 2025 official website. * The conference brings together technical talent from across the LY Corporation Group, including LINE Plus, LINE Taiwan, and LINE Vietnam. ### Multi-Disciplinary Technical Tracks * The agenda is divided into 12 distinct categories to cover the full spectrum of software development and product lifecycle. * Day 1 focuses on foundational technologies: AI, Security, Server-side development, Private Cloud, Infrastructure, and Data Platforms. * Day 2 explores application and management layers: AI Use Cases, Frontend, Mobile Applications, Design, Product Management, and Engineering Management. ### Key Engineering Case Studies and Sessions * **AI and Data Automation:** Sessions explore the evolution of development processes using AI, the shift from "Vibe Coding" to professional AI-assisted engineering, and the use of Generative AI to automate data pipelines. * **Infrastructure and Scaling:** Presentations include how the "Central Dogma Control Plane" connects thousands of services within LY Corporation and methods for improving video playback quality for LINE Call. * **Framework Migration:** A featured case study details the strategic transition of the "Demae-can" service from React Native to Flutter. * **Product Insights:** Deep dives into user experience design and data-driven insights gathered from LINE Talk's global user base. Tech-Verse 2025 provides a valuable opportunity for developers to learn from real-world deployments of AI and large-scale infrastructure. Given the breadth of the 127 sessions and the availability of real-time translation, tech professionals should review the timetable in advance to prioritize tracks relevant to their specific engineering interests.

lineOriginal article

AI and Writer's Partnership (opens in new tab)

LY Corporation is addressing the chronic shortage of high-quality technical documentation by treating the problem as an engineering challenge rather than a training issue. By utilizing Generative AI to automate the creation of API references, the Document Engineering team has transitioned from a "manual craftsmanship" approach to an "industrialized production" model. While the system significantly improves efficiency and maintains internal context better than generic tools, the team concludes that human verification remains essential due to the high stakes of API accuracy. ### Contextual Challenges with Generic AI Standard coding assistants like GitHub Copilot often fail to meet the specific documentation needs of a large organization. * Generic tools do not adhere to internal company style guides or maintain consistent terminology across projects. * Standard AI lacks awareness of internal technical contexts; for example, generic AI might mistake a company-specific identifier like "MID" for "Member ID," whereas the internal tool understands its specific function within the LY ecosystem. * Fragmented deployment processes across different teams make it difficult for developers to find a single source of truth for API documentation. ### Multi-Stage Prompt Engineering To ensure high-quality output without overwhelming the LLM's "memory," the team refined a complex set of instructions into a streamlined three-stage workflow. * **Language Recognition:** The system first identifies the programming language and specific framework being used. * **Contextual Analysis:** It analyzes the API's logic to generate relevant usage examples and supplemental technical information. * **Detail Generation:** Finally, it writes the core API descriptions, parameter definitions, and response value explanations based on the internal style guide. ### Transitioning to Model Context Protocol (MCP) While the prototype began as a VS Code extension, the team shifted to using the Model Context Protocol (MCP) to ensure the tool was accessible across various development environments. * Moving to MCP allows the tool to support multiple IDEs, including IntelliJ, which was a high-priority request from the developer community. * The MCP architecture decouples the user interface from the core logic, allowing the "host" (like the IDE) to handle UI interactions and parameter inputs. * This transition reduced the maintenance burden on the Document Engineering team by removing the need to build and update custom UI components for every IDE. ### Performance and the Accuracy Gap Evaluation of the AI-generated documentation showed strong results, though it highlighted the unique risks of documenting APIs compared to other forms of writing. * Approximately 88% of the AI-generated comments met the team's internal evaluation criteria. * The specialized generator outperformed GitHub Copilot in 78% of cases regarding style and contextual relevance. * The team noted that while a 99% accuracy rate is excellent for a blog post, a single error in a short API reference can render the entire document useless for a developer. To successfully implement AI-driven documentation, organizations should focus on building tools that understand internal business logic while maintaining a strict "human-in-the-loop" workflow. Developers should use these tools to generate the bulk of the content but must perform a final technical audit to ensure the precision that only a human author can currently guarantee.

figma3 min readCurated summary

8 Essential Tips for Using Figma Make | Figma Blog

Figma Make works best when users provide clear context, prepare clean design files, and refine complex projects incrementally. Detailed initial prompts reduce revisions, while organized Figma layers and Auto Layout help designs translate into functional prototypes. For ambitious builds, breaking work into focused prompts and separate code folders improves control, maintainability, and debugging. ## Provide Detailed Initial Prompts - Include: - The task Figma Make should perform - Product or flow context - Essential design elements - Expected interactions and behaviors - Device, layout, and visual constraints - Front-loading requirements helps produce a stronger first version with fewer follow-up prompts. - Use precise, measurable instructions instead of vague requests: - “Move this element down 20 pixels” - “Add 16px of space between these buttons” - If repeated adjustments are not working, restart with a new file and use lessons from the first attempt. - Effective project prompts can include an overview, platform, purpose, features, visual direction, technical details, and an explicit first implementation step. ## Clean Up Figma Files Before Importing Them - Figma Make can either create new designs or turn existing Figma frames into interactive prototypes. - Before copying a frame into Figma Make: - Organize the file - Apply appropriate constraints - Use Auto Layout correctly - Name layers according to their purpose - Figma tools such as Suggest Auto Layout and Rename Layers with AI, along with plugins like Clean Document, can help prepare files. - If the result is too large or not responsive, use prompts such as: - “Scale this to the size of my screen and make it responsive.” - “Keep this mobile-sized.” - A well-structured Auto Layout design can enable complex interactions from a single prompt, such as making a CD spin when a music player starts. ## Build Complex Projects Incrementally - Use a detailed first prompt to establish the overall foundation, then make smaller, focused changes. - Smaller requests allow the model to respond more precisely and reduce the risk of unwanted changes elsewhere. - Incremental prompting is useful for: - Building complex interfaces - Creating multi-page flows - Adding individual features - Maintaining the intended visual direction - Ask Figma Make to place separate elements in separate code folders to improve organization, maintainability, and error isolation. - Large projects may require many prompts; one financial dashboard and onboarding flow took more than 150 focused iterations. - Example follow-ups included adding journal post-its, inserting a detailed finance table, and adding a currency-selection checkbox. - Separating 3D landmarks into individual coded files similarly makes it easier to refine components without affecting the wider environment. Use Figma Make as an iterative design-and-development tool: prepare the source file carefully, describe the desired result precisely, and make complex changes one manageable step at a time.

Read original(opens in new tab)
googleOriginal article

Zooming in: Efficient regional environmental risk assessment with generative AI (opens in new tab)

Google Research has introduced a dynamical-generative downscaling method that combines physics-based climate modeling with probabilistic diffusion models to produce high-resolution regional environmental risk assessments. By bridging the resolution gap between global Earth system models and city-level data needs, this approach provides a computationally efficient way to quantify climate uncertainties at a 10 km scale. This hybrid technique significantly reduces error rates compared to traditional statistical methods while remaining far less computationally expensive than full-scale dynamical simulations. ## The Resolution Gap in Climate Modeling * Traditional Earth system models typically operate at a resolution of ~100 km, which is too coarse for city-level planning regarding floods, heatwaves, and wildfires. * Existing "dynamical downscaling" uses regional climate models (RCMs) to provide physically realistic 10 km projections, but the computational cost is too high to apply to large ensembles of climate data. * Statistical downscaling offers a faster alternative but often fails to capture complex local weather patterns or extreme events, and it struggles to generalize to unprecedented future climate conditions. ## A Hybrid Dynamical-Generative Framework * The process begins with a "physics-based first pass," where an RCM downscales global data to an intermediate resolution of 50 km to establish a common physical representation. * A generative AI system called "R2D2" (Regional Residual Diffusion-based Downscaling) then adds fine-scale details, such as the effects of complex topography, to reach the target 10 km resolution. * R2D2 specifically learns the "residual"—the difference between intermediate and high-resolution fields—which simplifies the learning task and improves the model's ability to generalize to unseen environmental conditions. ## Efficiency and Accuracy in Risk Assessment * The model was trained and validated using the Western United States Dynamically Downscaled Dataset (WUS-D3), which utilizes the "gold standard" WRF model. * The dynamical-generative approach reduced fine-scale errors by over 40% compared to popular statistical methods like BCSD and STAR-ESDM. * A key advantage of this method is its scalability; the AI requires training on only one dynamically downscaled model to effectively process outputs from various other Earth system models, allowing for the rapid assessment of large climate ensembles. By combining the physical grounding of traditional regional models with the speed of diffusion-based AI, researchers can now produce granular risk assessments that were previously cost-prohibitive. This method allows for a more robust exploration of future climate scenarios, providing essential data for farming, water management, and community protection.