Google Cloud

Why AI apps fail in production (And how Google solved it)

AI applications that work perfectly in local development environments fail when deployed to production in enterprise settings due to infrastructure constraints, cascading errors, and organizational governance barriers.

ml-systems observability
5 min
AWS

Eclipse Dataspace Components on AWS: Cost optimization strategies

Predicting and controlling infrastructure costs when deploying Eclipse Dataspace Components connectors on AWS without clear cost benchmarks.

observability general
4 min
AWS

Unlocking the future of video data: March Networks cloud storage on AWS

Storing and managing petabytes of distributed video surveillance data across thousands of locations while meeting growing retention requirements and enabling operational insights extraction.

storage-systems distributed-systems
5 min
Meta

Modernizing the Meta Ads Service With an Open-Source Kernel Scheduler

Linux kernel upgrades risked introducing latency regressions across Meta's ad serving fleet, which operates at scale where milliseconds of latency degradation significantly impact ads performance.

observability distributed-systems
5 min
AWS

S&P Global’s innovative disaster recovery strategy using Amazon FSx for NetApp ONTAP snapshots

S&P Global needed to implement a disaster recovery strategy for their Capital IQ platform that could achieve rapid failover to a secondary region while maintaining data consistency for mission-critical financial operations.

storage-systems databases
5 min
AWS

Lessons learned from scaling to 1 million Lambda functions

Managing operational stability and resource constraints when scaling a serverless SaaS platform from thousands to over 1 million concurrent Lambda functions.

distributed-systems rate-limiting
5 min
Cloudflare

Content Independence Day, one year on: building the business model for the agentic Internet

How to build infrastructure that enables direct monetization of content creators in an AI-agent-driven internet where traditional search referral models are becoming obsolete.

api-design security
4 min
Meta

10 Years of Meta’s Commitment to Python

Meta needed to maintain and advance Python as a critical language across its engineering infrastructure while supporting the broader open-source community that sustains the language.

general
5 min
Meta

How Meta Engineered Ultra-Narrow Batteries for AI Glasses

How to design and manufacture ultra-narrow batteries that fit within the temple arms of smart glasses while providing sufficient power for cameras, speakers, AI workloads, and displays.

general
5 min
Cloudflare

Build your own vulnerability harness

Cloudflare needed to automatically discover, triage, and manage security vulnerabilities at scale while minimizing false positives and handling the computational constraints of large language models.

ml-systems security
4 min
Google Cloud

How I learned Go in a Day with Antigravity 2.0 and How You Can Do the Same

Reducing runtime overhead and resource consumption by migrating from a resource-intensive Node.js/NPM stack to a compiled, single-binary solution for a CLI tool managing Agent Skills.

general
5 min
Google Cloud

How customer collaboration is shaping the future of GenAI security with Model Armor

Securing generative AI models and deployments in production environments while maintaining usability for enterprise customers in regulated industries like telecommunications.

security ml-systems
5 min
Google Cloud

Scaling the Next Generation of Global Innovation: How Google Supports Top Startups Around the World

How to provide startups with the technical infrastructure, architectural guidance, and cloud platform support necessary to scale from prototype to market-defining global businesses.

distributed-systems general
5 min
AWS

Introducing the Snowflake and AWS Custom Lens for the AWS Well-Architected Framework

Organizations needed a unified framework to architect solutions that effectively integrate AWS cloud services with Snowflake's data platform while following best practices for both.

general databases
5 min
Google Cloud

10 Indispensable Prompts Our Team Refuses to Build Without

Engineers waste time creating ad-hoc, inconsistent prompts for AI tools instead of having a reusable, refined set of prompts that consistently produce high-quality outputs.

general
5 min
AWS

Align your architecture backlog with Tech Roadmap Prioritization (TRP)

Engineering teams struggle to prioritize competing architectural initiatives and backlog items in a way that aligns technical decisions with business impact and resource constraints.

general
4 min
AWS

Scaling oncology patient support: How New York Cancer and Blood Specialists transformed customer experience with AWS and Pronetx, now part of Caylent

NYCBS needed to modernize their patient engagement and contact center infrastructure to improve patient enrollment and streamline communication with oncology patients.

general api-design
4 min
Cloudflare

VoidZero is joining Cloudflare

Building performant JavaScript tooling and compilers that can scale across modern web development workflows.

general
3 min
Spotify

Coding Is No Longer the Constraint: Scaling Developer Experience to Teams and Agents at Spotify

Scaling developer productivity and experience when coding is no longer the primary bottleneck, requiring infrastructure and tooling that enable both human teams and AI agents to work effectively.

observability microservices
4 min
Dropbox

Beyond code generation: rethinking engineering productivity in the age of AI agents

How to transition from code-generation AI tools that only assist engineers to autonomous agentic systems capable of executing complete, scoped engineering tasks independently.

microservices api-design
3 min
Google Cloud

Developer's guide to Gemini Enterprise and A2UI integration

Conversational AI agents lack a standard way to render rich UI components (date pickers, maps, multi-select lists) within chat interfaces, forcing agents to rely only on text or markdown responses.

api-design general
5 min
Google Cloud

From keynote to the terminal: Join our Next ‘26 developer livestreams

Google Cloud needed to bridge the gap between high-level keynote announcements and practical implementation details that developers could immediately apply.

general observability
5 min
Google Cloud

Introducing the Builders Hub from the Google Developer Program

Developers lose productivity navigating fragmented tooling across multiple consoles, documentation sites, and services to manage their projects and stay informed.

api-design general
5 min
Google Cloud

Migrating to Google Cloud’s Application Load Balancer: A practical guide

Migrating business-critical load balancer configurations from on-premises hardware solutions to Google Cloud while preserving existing traffic manipulation logic.

load-balancing distributed-systems
5 min
Google Cloud

Next '26 Hands-On: 10 Codelabs to Build Featured Tech

How to help developers transition from understanding AI concepts to building and maintaining production agentic systems in cloud environments.

observability microservices
5 min
Google Cloud

Pioneering AI-assisted code migration: How Google achieved 6x faster migration from TensorFlow to JAX

Google needed to accelerate large-scale codebase migrations (TensorFlow to JAX) that are too complex and interconnected for manual developer effort or standard AI coding tools to handle efficiently.

ml-systems general
5 min
Google Cloud

Ship code within minutes with the Gemini CLI DevOps Extension

Developers avoid deploying applications because the deployment process (containerization, CI/CD, IAM configuration) is time-consuming and interrupts the fast inner development loop.

devops general
5 min
AWS

Choosing between single or multiple organizations in AWS Organizations

Organizations must determine whether to operate under a single AWS organization or split into multiple organizations based on their operational, security, and scaling requirements.

security distributed-systems
4 min
Cloudflare

When "idle" isn't idle: how a Linux kernel optimization became a QUIC bug

CUBIC congestion control algorithm's congestion window was becoming pinned at minimum values in QUIC, causing severe performance degradation due to incorrect idle period detection.

networking security
4 min
Netflix

Scaling ArchUnit with Nebula ArchRules

Netflix needed a way to enforce consistent architectural patterns and build standards across tens of thousands of Java repositories in their polyrepo strategy.

microservices general
5 min
Spotify

Let’s Talk Agentic Development: Spotify x Anthropic Live

How can software engineers leverage AI agents to improve development workflows and productivity at scale?

general
4 min
Cloudflare

Introducing Dynamic Workflows: durable execution that follows the tenant

Enable multi-tenant platforms to execute millions of unique, durable workflows without incurring significant idle infrastructure costs.

distributed-systems microservices
4 min
Cloudflare

Building the agentic cloud: everything we launched during Agents Week 2026

How to enable developers to build and deploy AI agents at scale across a distributed edge computing network while maintaining security and providing necessary infrastructure tools.

distributed-systems security
4 min
Cloudflare

Making Rust Workers reliable: panic and abort recovery in wasm‑bindgen

Rust panics in Cloudflare Workers were fatal and poisoned the entire worker instance, making applications unreliable when unhandled errors occurred.

security observability
4 min
Cloudflare

Shared Dictionaries: compression that keeps up with the agentic web

Web pages are growing larger and slower to load due to increased dynamic content, requiring better compression techniques that can adapt to modern agentic web patterns.

caching api-design
3 min
Meta

Capacity Efficiency at Meta: How Unified AI Agents Optimize Performance at Hyperscale

Meta needed to automatically identify and remediate performance inefficiencies across their massive infrastructure to reduce power consumption and free up engineering capacity.

observability distributed-systems
5 min
Cloudflare

500 Tbps of capacity: 16 years of scaling our global network

How to scale a global content delivery and DDoS mitigation network to handle massive throughput (500 Tbps) while maintaining capacity to protect against record-breaking attacks.

load-balancing distributed-systems
3 min
Cloudflare

Welcome to Agents Week

How to enable AI agents to operate effectively at the edge of the internet with the security, performance, and reliability characteristics of Cloudflare's existing infrastructure.

distributed-systems security
4 min
Meta

Escaping the Fork: How Meta Modernized WebRTC Across 50+ Use Cases

Meta needed to modernize WebRTC across 50+ use cases while maintaining synchronization with upstream open-source development, avoiding the drift that typically occurs when large projects fork internally.

distributed-systems real-time-systems
5 min
Meta

Trust But Canary: Configuration Safety at Scale

Safely deploying configuration changes at scale while minimizing the risk of widespread failures caused by faulty configurations.

observability distributed-systems
5 min
Dropbox

Reducing our monorepo size to improve developer velocity

Monorepo growth was causing increased build times, slower dependency resolution, and reduced developer velocity as the codebase expanded.

general observability
3 min
Meta

AI for American-Produced Cement and Concrete

Designing high-quality, sustainable concrete mixes that are produced in the United States while optimizing for performance characteristics.

ml-systems general
5 min
Meta

KernelEvolve: How Meta’s Ranking Engineer Agent Optimizes AI Infrastructure

Meta needed to automatically optimize low-level infrastructure and kernel-level parameters for AI ranking models to improve performance without manual tuning.

ml-systems distributed-systems
5 min
LinkedIn

Announcing Our LinkedIn-Cornell 2024 Grant Recipients

Advancing AI research requires collaboration between industry and academia, but funding and partnership models need structured programs.

ml-systems general
3 min
LinkedIn

Career stories: Influencing engineering growth at LinkedIn

Growing engineering teams at scale requires clear career frameworks and mentorship to help engineers develop technical leadership skills.

general
3 min
LinkedIn

Career stories: The math-music connection in data science

Data science teams need diverse skill sets that blend mathematical rigor with creative problem-solving to build effective ML systems.

ml-systems general
3 min