Cloud Monitoring

5 posts

aws3 min readCurated summary

AWS Interconnect is now generally available, with a new option to simplify last-mile connectivity | Amazon Web Services

AWS Interconnect is a managed service for private, high-speed connectivity between AWS and other clouds or on-premises networks. Its multicloud capability, now generally available, initially connects AWS with Google Cloud while Microsoft Azure support is planned for later in 2026. The service aims to replace complex VPN, colocation, and third-party networking setups with a turnkey, resilient configuration managed through AWS. ## AWS Interconnect Capabilities - **Interconnect – multicloud** connects an AWS VPC privately to VPCs on other cloud providers. - **Interconnect – last mile** simplifies connectivity from branch offices, data centers, and remote sites through existing network providers. - Both capabilities provide: - Dedicated bandwidth - Private connectivity - Managed provisioning - Reduced infrastructure and configuration overhead - Connections can be configured through the AWS Console by selecting the location or provider, AWS Region, and bandwidth. ## Multicloud Connectivity - The service provides a managed **Layer 3 connection** between AWS and another cloud provider. - Traffic uses the AWS global backbone and the partner’s private network rather than the public internet. - This improves: - Latency predictability - Throughput consistency - Isolation from internet congestion - Google Cloud is supported at launch; Microsoft Azure is expected later in 2026. ## Security, Resilience, and Monitoring - Physical links between AWS and partner routers use **IEEE 802.1AE MACsec encryption** by default. - Each cloud provider handles encryption on its own backbone, so customers must verify that the resulting deployment satisfies compliance requirements. - Connections use multiple logical links across at least two physical facilities to protect against device or facility failures. - Amazon CloudWatch integration includes: - A Network Synthetic Monitor for round-trip latency and packet loss - Bandwidth utilization metrics for capacity planning ## Open Partner Specification - AWS has published the underlying Interconnect specification on GitHub under the **Apache 2.0 license**. - Other cloud providers can become partners by implementing the specification and meeting AWS requirements for: - Resiliency - Support - Service-level agreements - Operational readiness ## Provisioning an AWS–Google Cloud Connection - The demonstration connects a single AWS VPC to a Google Cloud VPC using a Direct Connect Gateway. - In the AWS Direct Connect console, the user: - Selects Google Cloud as the provider - Chooses AWS Region `eu-central-1` - Chooses Google Cloud Region `europe-west3` - Specifies bandwidth - Selects a Direct Connect Gateway - Enters the Google Cloud project ID - AWS then generates an activation key for use on the Google Cloud side. ## Configuring Google Cloud - Because a Google Cloud web console option was unavailable at the time, the example uses the `gcloud` CLI. - The user creates a transport resource with: - The AWS activation key - The Google Cloud region - The target VPC network - Advertised AWS routes - After the transport reaches the appropriate state, the user creates a VPC peering connection between the Google Cloud VPC and the generated transport network. - Custom routes are imported and exported through the peering configuration. ## Completing the AWS Configuration - Once the Google Cloud transport and peering are configured: - The AWS Interconnect status can be checked in the Interconnect console. - The Direct Connect Gateway shows the new attachment. - The final AWS-side step is associating the gateway with the appropriate Virtual Private Gateway. - The Virtual Private Gateway must be in the same AWS Region as the Interconnect. - AWS routing still requires a final route entry so workloads can reach the remote Google Cloud network. AWS Interconnect is best suited to organizations operating hybrid or multicloud environments that want private, resilient connectivity without managing physical links or complex third-party networking. The managed provisioning process can reduce setup time to minutes, but teams should still validate routing, encryption responsibilities, regional constraints, and compliance requirements.

Read original(opens in new tab)
awsOriginal article

Amazon S3 Storage Lens adds performance metrics, support for billions of prefixes, and export to S3 Tables (opens in new tab)

Amazon S3 Storage Lens has introduced three significant updates designed to provide deeper visibility into storage performance and usage patterns at scale. By adding dedicated performance metrics, support for billions of prefixes, and direct export capabilities to Amazon S3 Tables, AWS enables organizations to better optimize application latency and storage costs. These enhancements allow for more granular data-driven decisions across entire AWS organizations or specific high-performance workloads. ## Enhanced Performance Metric Categories The update introduces eight new performance-related metric categories available through the S3 Storage Lens advanced tier. These metrics are designed to pinpoint specific architectural bottlenecks that could impact application speed. * **Request and Storage Distributions:** New metrics track the distribution of read/write request sizes and object sizes, helping identify small-object patterns that might be better suited for Amazon S3 Express One Zone. * **Error and Latency Tracking:** Users can now monitor concurrent PUT 503 errors to identify throttling and analyze FirstByteLatency and TotalRequestLatency to measure end-to-end request performance. * **Data Transfer Efficiency:** Metrics for cross-Region data transfer help identify high-cost or high-latency data access patterns, suggesting where compute resources should be co-located with storage. * **Access Patterns:** Tracking unique objects accessed per day identifies "hot" datasets that could benefit from higher-performance storage tiers or caching solutions. ## Support for Billions of Prefixes S3 Storage Lens has expanded its analytical scale to support the monitoring of billions of prefixes. This allows organizations with massive, complex data structures to maintain granular visibility without sacrificing performance or detail. * **Granular Visibility:** Users can drill down into massive datasets to find specific prefixes causing performance degradation or cost spikes. * **Scalable Analysis:** This expansion ensures that even the largest data lakes can be monitored at a level of detail previously limited to smaller buckets. ## Integration with Amazon S3 Tables The service now supports direct export of storage metrics to Amazon S3 Tables, a feature optimized for high-performance analytics. This integration streamlines the workflow for administrators who need to perform complex queries on their storage metadata. * **Analytical Readiness:** Exporting to S3 Tables makes it easier to use SQL-based tools to query storage trends and performance over time. * **Automation:** This capability allows for the creation of automated reporting pipelines that can handle the massive volume of data generated by prefix-level monitoring. To take full advantage of these features, users should enable the S3 Storage Lens advanced tier and configure prefix-level monitoring for buckets containing mission-critical or high-throughput data. Organizations experiencing latency issues should specifically review the new request size distribution metrics to determine if batching objects or migrating to S3 Express One Zone would improve performance.

awsOriginal article

New and enhanced AWS Support plans add AI capabilities to expert guidance (opens in new tab)

AWS has announced a major transformation of its support plans, moving from a reactive model to a proactive, AI-driven approach for issue prevention and workload optimization. By integrating AI-powered capabilities with deep technical expertise, these enhanced plans aim to help organizations identify potential operational risks before they impact business performance. This new tier-based structure provides businesses with varying levels of contextual assistance, ranging from intelligent automated recommendations to direct access to specialized engineering teams. ### Business Support+ * Introduces intelligent, AI-powered assistance designed to provide contextual recommendations for developers, startups, and small businesses. * Features a seamless transition from AI tools to human experts, with critical case response times reduced to 30 minutes—twice as fast as previous standards. * Provides personalized workload optimization suggestions based on the user's specific environment via a low-cost monthly subscription. ### Enterprise Support * Assigns a designated Technical Account Manager (TAM) who utilizes data-driven insights and AI tools to mitigate risks and identify optimization opportunities. * Grants access to the AWS Security Incident Response service at no additional fee, centralizing the tracking, monitoring, and investigation of security events. * Guarantees a 15-minute response time for production-critical issues, with support engineers receiving AI-generated context to ensure faster, more personalized resolution. * Includes access to hands-on workshops and interactive programs to foster continuous technical growth within the organization. ### Unified Operations Support * Provides the highest level of context-aware assistance through a dedicated core team including a TAM, a Domain Engineer, and a Senior Billing and Account Specialist. * Delivers industry-leading 5-minute response times for critical incidents, supported by around-the-clock monitoring and AI-powered proactive risk identification. * Offers on-demand access to specialized experts in migration, incident management, and security through the customer’s preferred collaboration channels. These updates reflect AWS’s commitment to using generative AI to shorten resolution times and provide more personalized architectural guidance. Organizations should evaluate their operational complexity and required response times to select the plan that best aligns with their mission-critical cloud needs.

datadog3 min readCurated summary

Engineering spotlight: Jeromy Carriere

Jeromy Carriere, Datadog’s SVP of Product Engineering, describes engineering leadership as balancing strategy, execution, people development, and organizational processes. His career across Google, Facebook, and Datadog shaped his passion for observability and taught him to lead through mistakes, autonomy, and accountability. He argues that sustained innovation requires both intentional direction and space for teams and individuals to grow. ## The Responsibilities of Engineering Leadership - Carriere’s work shifts with organizational cycles: - Quarterly planning focused on connecting initiatives and increasing collaboration. - Execution periods focused on removing resource and decision-making blockers. - Ongoing performance management across the broader engineering organization. - He works on how Engineering operates, including: - Process definition and improvement. - Hiring and performance management. - Reviewing design documents and code. - Observing incidents and postmortems. - His central challenge is balancing attention across strategy, execution, people, and technical quality. ## From Cloud Monitoring to Datadog - At Google in 2014, Carriere helped create a cloud monitoring offering because Google Cloud lacked capabilities comparable to Datadog. - He later worked on observability at Facebook and developed a strong interest in improving developer and engineer productivity. - He returned to Datadog after seeing its ability to innovate and deliver products with sustained velocity. - He emphasizes that velocity means more than moving quickly: it requires direction, strategy, and consistency over time. ## Learning Through Mistakes - Carriere believes the most valuable lessons come from making and owning mistakes. - Earlier in his career, he was sometimes too directive, limiting team creativity and ownership. - At other times, he was too distant and failed to provide enough support. - His current leadership approach aims to: - Give teams substantial autonomy. - Provide support when needed. - Hold teams accountable for agreed-upon outcomes. - He also learned that people may not have a clear five-year career plan. Leaders should help them identify the work that provides satisfaction and enables them to perform at their best. ## Creating Space for Career Decisions - People often become focused on the immediate task and overlook other possibilities. - Carriere recommends deliberately stepping back to observe: - What activities feel satisfying. - What opportunities exist nearby. - What kinds of work could better match an individual’s strengths and interests. - This reflection requires intentional time and freedom rather than waiting for clarity to emerge automatically. ## The Value of Co-op and Internship Programs - Carriere credits the University of Waterloo’s co-op program with giving him early experience as a professional software developer. - The combination of strong academic training and repeated, high-quality industry placements helped connect theory with real work. - He sees a similar benefit in Datadog’s internship program, where interns are trusted with meaningful projects and often produce some of the company’s strongest work. Datadog’s engineering approach, as described by Carriere, combines strategic product velocity with thoughtful organizational support. For both leaders and individual contributors, the practical recommendation is to learn from mistakes, create room for reflection, and build environments where people have autonomy, meaningful work, and accountability.

Read original(opens in new tab)
datadog2 min readCurated summary

Rethinking UX for AI-driven alerting | Datadog

Datadog’s page announces that the company was named a Leader in the 2026 Gartner Magic Quadrant for Observability Platforms. The supplied content, however, primarily contains site navigation rather than the referenced blog post, so it does not provide details about the article’s argument concerning AI-driven alerting. ## Gartner Recognition - Datadog highlights its recognition as a Leader in Gartner’s Magic Quadrant for Observability Platforms. - The announcement is presented as a promotional resource linked from the Datadog website. ## Datadog’s Product Portfolio The navigation emphasizes Datadog’s broad observability and security platform, including: - **Infrastructure:** infrastructure, container, network, serverless, GPU, storage, and cloud-cost monitoring. - **Applications:** APM, service monitoring, profiling, dynamic instrumentation, and agent observability. - **Data and logs:** database, data-stream, job, quality, log, sensitive-data, and pipeline monitoring. - **Security:** code, cloud, vulnerability, workload, application, API, and SIEM security tools. - **Digital experience:** browser and mobile RUM, session replay, synthetic monitoring, product analytics, and error tracking. - **Software delivery:** CI visibility, test optimization, code coverage, feature flags, and developer portals. - **Service management:** incident response, SLOs, event management, workflows, and case management. - **AI:** Bits AI agents, investigations, chat, security analysis, agent observability, and MCP integrations. The provided text does not include enough of the actual “Rethinking UX for AI-Driven Alerting” article to summarize its technical concepts or conclusions.

Read original(opens in new tab)