Ai Inference

5 posts

aws2 min readCurated summary

Announcing Amazon EC2 G7 instances accelerated by NVIDIA RTX PRO 4500 Blackwell Server Edition GPUs | Amazon Web Services

Amazon EC2 G7 instances are now generally available with NVIDIA RTX PRO 4500 Blackwell Server Edition GPUs and custom sixth-generation Intel Xeon processors. Compared with G6 instances, they provide up to 4.6× higher AI inference performance and 2.1× better graphics performance. AWS positions them for AI inference, rendering, video, virtual desktops, spatial computing, and GPU-accelerated analytics. ## GPU and Performance Improvements - Each GPU provides 32 GB of memory, with up to 256 GB across eight GPUs. - GPU memory capacity is 1.33× higher and bandwidth is 2.45× higher than G6. - GPUs include fifth-generation Tensor Cores and fourth-generation RT Cores. - G7 instances accelerate analytics workloads running on Amazon EMR with Amazon EKS. ## Networking and Storage - Up to 700 Gbps of EFA-enabled networking—seven times the G6 throughput. - Up to 7.6 TB of local NVMe SSD storage keeps large models and datasets close to the GPUs. - Support for NVIDIA GPUDirect P2P and GPUDirect RDMA with EFA enables low-latency GPU communication across GPUs, nodes, and FSx for Lustre. ## Video Processing - Ninth-generation NVENC and sixth-generation NVDEC engines support 4:2:2 encoding and decoding. - They deliver up to 1.5× more concurrent video streams than G6 instances. ## Instance Configurations - Seven instance sizes are available. - Configurations offer up to: - 8 NVIDIA RTX PRO 4500 GPUs - 192 vCPUs - 768 GiB of system memory - 700 Gbps network bandwidth - 7.6 TB local NVMe storage - Detailed instance specifications were listed as “coming soon” in the announcement. ## Software and Availability - AWS provides Deep Learning AMIs and NVIDIA Workstation AMIs with preinstalled drivers. - Amazon EKS users should build AMIs with NVIDIA driver version R595. - Supported operating systems include Amazon Linux, Ubuntu, RHEL, and Windows Server. - NVIDIA integration supports DirectX, Vulkan, and OpenGL. - G7 instances are initially available in US East (Ohio) and US West (Oregon). - Purchasing options include On-Demand, Savings Plans, Spot Instances, and Dedicated Instances for selected sizes. G7 instances are a strong option for GPU-intensive workloads requiring higher inference, graphics, networking, and video performance. Organizations can launch them through the EC2 console and evaluate pricing across the available purchasing models.

Read original(opens in new tab)
cloudflare2 min readCurated summary

Introducing Custom Regions for precision data control

Cloudflare is expanding Regional Services with three new managed regions—Turkey, the UAE, and IRAP—and introducing Custom Regions. The feature lets customers define exactly where TLS termination and Layer 7 processing occur while still using Cloudflare’s global network for traffic ingestion and DDoS protection. This combines local data-sovereignty control with globally distributed security and performance. ## Global Security with Local Compliance - Traffic enters through the nearest Cloudflare data center, where Cloudflare applies global-scale Layer 3 and Layer 4 DDoS protection. - Before decryption, request metadata is inspected and traffic is routed over Cloudflare’s private backbone to a data center inside the customer’s designated region. - TLS termination and Layer 7 services—including WAF, Bot Management, and Workers—run only within that region. - After processing, traffic is re-encrypted and sent securely to the origin. - This design localizes data inspection without forcing customers to sacrifice Cloudflare’s global attack-mitigation capacity. ## Expanded Cloudflare Managed Regions - Regional Services originally supported the EU, UK, and U.S. - Cloudflare now offers 35 predefined regions. - Newly added options include: - Turkey - United Arab Emirates - IRAP, supporting Australian compliance - ISMAP, supporting Japanese compliance ## Custom Regions - Customers can define their own geographical boundaries instead of selecting only Cloudflare-managed regions. - Custom Regions support: - Individual countries - Arbitrary combinations of countries - Regions that exclude specified countries - Early-access use cases include: - Keeping AI prompts and responses within selected countries - Running country-specific promotions - Meeting government contractual requirements - Aligning regions with corporate structures such as EMEA, MENA, or APAC - Example definitions include North America, everywhere except North America, or countries where Fahrenheit is commonly used. ## How Region Enforcement Works - A region represents a set of Cloudflare data centers. - Managed regions use Cloudflare-defined membership. - Custom Regions use expressions, commonly based on the data center’s ISO country code: - `country_code == "TR"` selects Turkey. - `country_code in ["DE", "FR", "NL"]` selects Germany, France, and the Netherlands. - Negated expressions can exclude specific countries. - The same Regional Services architecture applies regardless of who defines the boundary; Custom Regions simply give the customer control over the membership rules. Custom Regions are best suited to organizations with precise sovereignty, compliance, performance, or organizational requirements. They provide fine-grained control over where sensitive traffic is decrypted and processed while preserving Cloudflare’s global network protections.

Read original(opens in new tab)
aws2 min readCurated summary

Amazon EC2 C8id, M8id, and R8id instances with up to 22.8 TB local NVMe storage are generally available | Amazon Web Services

Amazon has generally released the EC2 C8id, M8id, and R8id instances, combining custom Intel Xeon 6 processors with up to 22.8 TB of local NVMe SSD storage. Compared with prior sixth-generation instances, they provide up to 43% more compute performance, 3.3× higher memory bandwidth, and significant gains for I/O-intensive databases and analytics. The instances target compute-heavy, balanced, and memory-intensive workloads respectively. ## Instance Families and Workloads - **C8id:** Designed for compute-intensive applications such as video encoding, image processing, and media workloads requiring fast local storage. - **M8id:** Balances compute and memory for data logging, media processing, and medium-sized data stores. - **R8id:** Targets memory-intensive workloads, including large SQL/NoSQL databases, in-memory databases, analytics, and AI inference. ## Capacity and Performance - Scale up to **96xlarge**, compared with 32xlarge in the previous generation. - Offer up to: - **384 vCPUs** - **3 TiB memory** - **22.8 TB local NVMe storage** - Available in **metal-48xl** and **metal-96xl** configurations for workloads needing direct physical-resource access. - Deliver up to: - **43% higher compute performance** - **3.3× greater memory bandwidth** - **46% better I/O-intensive database performance** - **30% faster I/O-intensive real-time analytics queries** ## Networking, Storage, and Compatibility - Support **Instance Bandwidth Configuration**, allowing network or EBS bandwidth to be increased by up to 25% depending on workload needs. - Use sixth-generation AWS Nitro cards to offload virtualization, storage, and networking operations. - Require AMIs with **ENA and NVMe drivers**; current AWS Windows and Linux AMIs include the NVMe driver by default. - Local NVMe devices appear automatically after boot and do not require block device mappings. - Storage is hardware-encrypted with **XTS-AES-256** and unique keys. - Local NVMe storage is temporary: it is destroyed when the instance is stopped or terminated. ## Availability and Purchasing - Available in US East (N. Virginia), US East (Ohio), and US West (Oregon). - R8id instances are also available in Europe (Frankfurt). - Offered as On-Demand, Savings Plans, Spot Instances, Dedicated Instances, and Dedicated Hosts. These instances are best suited to applications that can exploit high-performance, ephemeral local NVMe storage; persistent data should remain on services such as Amazon EBS.

Read original(opens in new tab)
awsOriginal article

Announcing Amazon EC2 G7e instances accelerated by NVIDIA RTX PRO 6000 Blackwell Server Edition GPUs (opens in new tab)

Amazon has announced the general availability of EC2 G7e instances, a new hardware tier powered by NVIDIA RTX PRO 6000 Blackwell Server Edition GPUs designed for generative AI and high-end graphics. These instances deliver up to 2.3 times the inference performance of their G6e predecessors while providing significant upgrades to memory and bandwidth. This launch aims to provide a cost-effective solution for running medium-sized AI models and complex spatial computing workloads at scale. **Blackwell GPU and Memory Advancements** * The G7e instances feature NVIDIA RTX PRO 6000 Blackwell GPUs, which provide twice the memory and 1.85 times the memory bandwidth of the G6e generation. * Each GPU provides 96 GB of memory, allowing users to run medium-sized models—such as those with up to 70 billion parameters—on a single GPU using FP8 precision. * The architecture is optimized for both spatial computing and scientific workloads, offering the highest graphics performance currently available in the EC2 portfolio. **High-Speed Connectivity and Multi-GPU Scaling** * To support large-scale models, G7e instances utilize NVIDIA GPUDirect P2P, enabling direct communication between GPUs over PCIe interconnects with minimal latency. * These instances offer four times the inter-GPU bandwidth compared to the L40s GPUs found in G6e instances, facilitating more efficient data transfer in multi-GPU configurations. * Total GPU memory can scale up to 768 GB within a single node, supporting massive inference tasks across eight interconnected GPUs. **Networking and Storage Performance** * G7e instances provide up to 1,600 Gbps of network bandwidth, a four-fold increase over previous generations, making them suitable for small-scale multi-node clusters. * Support for NVIDIA GPUDirect Remote Direct Memory Access (RDMA) via Elastic Fabric Adapter (EFA) reduces latency for remote GPU-to-GPU communication. * The instances support GPUDirect Storage with Amazon FSx for Lustre, achieving throughput speeds up to 1.2 Tbps to ensure rapid model loading and data processing. **System Specifications and Configurations** * Under the hood, G7e instances are powered by Intel Emerald Rapids processors and support up to 192 vCPUs and 2,048 GiB of system memory. * Local storage options include up to 15.2 TB of NVMe SSD capacity to handle high-speed data caching and local processing. * The instance family ranges from the g7e.2xlarge (1 GPU, 8 vCPUs) to the g7e.48xlarge (8 GPUs, 192 vCPUs). For developers ready to transition to Blackwell-based architecture, these instances are accessible through AWS Deep Learning AMIs (DLAMI). They represent a major step forward for organizations needing to balance the high memory requirements of modern LLMs with the cost efficiencies of the G-series instance family.

awsOriginal article

Amazon EC2 X8i instances powered by custom Intel Xeon 6 processors are generally available for memory-intensive workloads (opens in new tab)

Amazon has announced the general availability of EC2 X8i instances, specifically engineered for memory-intensive workloads such as SAP HANA, large-scale databases, and data analytics. Powered by custom Intel Xeon 6 processors with a 3.9 GHz all-core turbo frequency, these instances provide a significant performance leap over the previous X2i generation. By offering up to 6 TB of memory and substantial improvements in throughput, X8i instances represent the highest-performing Intel-based memory-optimized option in the AWS cloud. ### Performance Enhancements and Processor Architecture * **Custom Silicon:** The instances utilize custom Intel Xeon 6 processors available exclusively on AWS, delivering the fastest memory bandwidth among comparable Intel cloud processors. * **Memory and Bandwidth:** X8i provides 1.5 times more memory capacity (up to 6 TB) and 3.4 times more memory bandwidth compared to previous-generation X2i instances. * **Workload Benchmarks:** Real-world performance gains include a 50% increase in SAP Application Performance Standard (SAPS), 47% faster PostgreSQL performance, 88% faster Memcached performance, and a 46% boost in AI inference. ### Scalable Instance Sizes and Throughput * **Flexible Sizing:** The instances are available in 14 sizes, including new larger formats such as the 48xlarge, 64xlarge, and 96xlarge. * **Bare Metal Options:** Two bare metal sizes (metal-48xl and metal-96xl) are available for workloads requiring direct access to physical hardware resources. * **Networking and Storage:** The architecture supports up to 100 Gbps of network bandwidth with Elastic Fabric Adapter (EFA) support and up to 80 Gbps of Amazon EBS throughput. * **Bandwidth Control:** Support for Instance Bandwidth Configuration (IBC) allows users to customize the allocation of performance between networking and EBS to suit specific application needs. ### Cost Efficiency and Use Cases * **Licensing Optimization:** In preview testing, customers like Orion reduced SQL Server licensing costs by 50% by maintaining performance thresholds with fewer active cores compared to older instance types. * **Enterprise Applications:** The instances are SAP-certified, making them ideal for RISE with SAP and other high-demand ERP environments. * **Broad Utility:** Beyond databases, the instances are optimized for Electronic Design Automation (EDA) and complex data analytics that require massive memory footprints. For organizations managing massive datasets or expensive licensed database software, migrating to X8i instances offers a clear path to both performance optimization and infrastructure cost reduction. These instances are currently available in the US East (N. Virginia), US West (Oregon), and Europe (Ireland) regions through On-Demand, Spot, and Reserved purchasing models.