Agenda

AI Engineer · all times Europe/London

Tuesday, 10 November 2026

09:00
Gray Failure: The Achilles’ Heel of Cloud-Scale Systems

Tuesday, 10 November 2026 · 09:00 to 09:45 · Main hall

Tvorimir Luchakivskiy · Principal Engineer at Latticework Systems

Cloud scale provides the vast resources necessary to repair and fix failed components, but this is useful only if those failures can be detected. For this reason, the major availability breakdowns and performance anomalies we see in cloud environments tend to be caused by subtle underlying faults, i.e., gray failure rather than fail-stop failure. In this talk, we discuss our experiences with gray failure in Microsoft Azure to show its broad scope and consequences with several case studies. We also argue that a key feature of gray failure is differential observability: that the system’s failure detectors may not notice problems even when applications are afflicted by them. We will show how Microsoft Azure applied the differential observability in practice and bridged the gap between differeShow more
OpsTalk (45 min)
Sign in to star this24
Real Talk: What We Think We Know — That Just Ain’t So

Tuesday, 10 November 2026 · 09:00 to 09:45 · Studio

Marcus Mccoy · Head of Platform at Kestrel Labs

Dr. David Woods once said to me “We cannot call it a scientific field unless we can admit we’ve gotten things wrong in the past.” Do we, in this community, do that? Well-formed critique is critical for any field — including SRE — to progress. I’d like to talk about a few ideas, assumptions, and concepts often talked about in this community but whose validity is rarely questioned or explored. John Allspaw has worked in software systems engineering and operations for over twenty years in many different environments. John’s publications include the books *The Art of Capacity Planning* (2009) and *Web Operations* (2010) as well as the forward to *The DevOps Handbook.* His 2009 Velocity talk with Paul Hammond, *10+ Deploys Per Day: Dev and Ops Cooperation* helped start the DevOps movement. JohnShow more
PracticeTalk (25 min)
Sign in to star this5
Product Reliability for Google Maps

Tuesday, 10 November 2026 · 09:00 to 09:45 · Workshop room

Nora Leroy · Distinguished Engineer at Fathom Analytics Group

As our organization has gotten very good at protecting server SLOs with reliability best practices like scaling globally distributed at-scale architectures, toil mitigation, and continuous reliability improvements we noticed that a majority of incidents impacting our end-users were not showing up as an SLO miss. In many cases these outages were not even observable from the server side - for example, the rollout of a new version of the consumer mobile application (that our services powers) to an app store could break one or more critical feature(s) due to bugs in client code. This reality has led to a change in the way we approach reliability - we’re shifting our focus from server reliability to product reliability. We’re not yet finished with the transition, but we’re starting to see very Show more
ResearchTalk (25 min)
Sign in to star this4
10:00
Using Generative AI Patterns for Better Observability

Tuesday, 10 November 2026 · 10:00 to 10:45 · Main hall

Sidónio Moreira · Research Scientist at Clearwater Storage

What happens when you've got a team of SREs who get downsized and smacked with an increasing workload? This is the story of my small team's wild ride of revamping our in-house incident tool. Spoiler: It turned into a "make or buy" debate real quick! Picture this: a team of 10 becomes a merry band of 3, drowning in a sea of issues, requests, and feedback. Unexpectedly, our team changed like a game of musical chairs, skills shuffling all over. In our pursuit to be the heroes, we hit the brakes and pitched buying a ready-made tool for a change. It wasn't just a business decision; it was a cultural rollercoaster for everyone involved! So, let's take a stroll down memory lane, exploring how we switched gears to evaluate third-party tools to pave the ground for improving the reliability of our pShow more
SystemsTalk (25 min)
Sign in to star this5
Strengthening Apache Pinot's Query Processing Engine with Adaptive Server Selection and Runtime ## Query Killing

Tuesday, 10 November 2026 · 10:00 to 10:45 · Studio

Irmingard Frese · Staff Engineer at Kestrel Labs

Apache Pinot, a real-time, distributed OLAP database, encountered resiliency challenges in its large production deployment due to server failures, slowness, and unpredictable query patterns. To ensure adherence to strict SLAs, enhancements were made to its Query Processing framework through two key features: Intelligent Query Routing and Runtime Query Killing. Intelligent Query Routing replaces the traditional round-robin server query distribution with a dynamic approach based on continuous server performance metrics. This adaptive method proactively routes queries to more efficient servers, reducing the need for manual intervention and preserving availability during transient server issues. Runtime Query Killing addresses the issue of complex queries causing Out of Memory errors and availShow more
PracticeTalk (45 min)
Sign in to star this5
Optimizing Resilience and Availability by Migrating from JupyterHub to the Kubeflow Notebook ## Controller

Tuesday, 10 November 2026 · 10:00 to 10:45 · Workshop room

Andrea Johnson · Head of Platform at Meridian Payments

This presentation details our transition from JupyterHub to the Kubeflow Notebook Controller. JupyterHub was architected in a backend agnostic way that "supports" Kubernetes but isn't truly Kubernetes-native. As a result, it has significant shortcomings with respect to resilience and high availability. In particular, the core component, the hub API, can only have one replica at any given time. In contrast, The Kubeflow Notebook controller is built from the ground up to be Kubernetes native using the operator pattern. There's far less complexity, fewer components, less brittleness, and improved resilience and high availability. As a result, our platform has been able to scale to four times as many users, including ten times as many concurrent executions. Our users are happier and there's leShow more
ResearchTalk (45 min)
Sign in to star this5
11:00
Cross-System Interaction Failures: Don't Fail through the Cracks

Tuesday, 10 November 2026 · 11:00 to 11:45 · Main hall

Larisa Kladochniy · Research Scientist at Jane Street

Modern cloud systems are orchestrated by many independent and interacting subsystems, each specializing in important services such as scheduling, data processing and storage, resource management, etc. Hence, overall cloud system reliability is determined not only by the reliability of each individual subsystem, but also by the interactions between them. With recent practices of microservice and serverless architectures, each individual subsystem becomes simpler and smaller, while their interactions grow in complexity and diversity. We observe that many recent production incidents of large-scale cloud systems are manifested through failures of the interactions across system boundaries, which we term "cross-system interaction failures", or "CSI failures". However, understanding and addressinShow more
Talk (45 min)
Sign in to star this
Storytelling as an Incident Management Skill

Tuesday, 10 November 2026 · 11:00 to 11:45 · Studio

Yashodha Kaur · Observability Engineer at Snowflake

Managing incidents well requires a number of skills — debugging and systems understanding, strong communication, and high-speed project management — but we rarely talk about the power of storytelling in our incident management loop. From oncall preparation, to incident handling, to postmortem creation, skill at storytelling can support and even improve our engineering skills. Establishing a setting, building a coherent narrative, and understanding your characters builds memorable and compelling communication that can level up all stages of your incident management. Laura de Vesine is a 20+ year software industry veteran. She has spent the last 8 years in SRE working in incident analysis and prevention, chaos engineering, and the intersection of technology and organizational culture. Laura Show more
Talk (45 min)
Sign in to star this
OIDC and CICD: Why Your CI Pipeline Is Your Greatest Security Threat

Tuesday, 10 November 2026 · 11:00 to 11:45 · Workshop room

Noemí Marrero · Engineering Manager at Fathom Analytics Group

Your CI/CD Process is chock full of credentials, and almost anyone in your company has access to it. Configuring your CI correctly is vital to supply chain security. We discuss how to reduce that attack surface by enforcing proper branch permissions and using OIDC to reduce long-lived credentials and tie branches to roles.Show more
PracticeLightning talk (10 min)
Sign in to star this5
12:00
Lunch
13:30
Synthesizing Sanity with, and in Spite of, Synthetic Monitoring

Tuesday, 10 November 2026 · 13:30 to 14:15 · Main hall

Madeleine Skotnes · Site Reliability Engineer at Halyard Cloud

Synthetic monitoring, particularly browser-based monitoring, is hard to do well. When tests pass, synthetic monitoring provides a uniquely intuitive kind of psychological safety - human-like, verified confidence, compared to other forms of monitoring. When tests fail, synthetic monitoring is often blamed as flaky, misconfigured, or unreliable. If not properly implemented, it can not only be financially, mentally and organisationally draining, but damaging to real customer experience. This talk is a conceptual and technical story of 4 years working with Atlassian’s in-house synthetic monitoring solution, being the owning developer for a tool actively used by 30-40 internal teams to build and manage synthetic monitoring for Jira. How can we make synthetic monitoring better serve its purpose Show more
ResearchWorkshop (90 min)
Sign in to star this4
Measuring Reliability Culture to Optimize Tradeoffs: Perspectives from an Anthropologist

Tuesday, 10 November 2026 · 13:30 to 14:15 · Studio

Carolina Concepción · Site Reliability Engineer at Fathom Analytics Group

Understanding how culture influences engineering practices is often a black box. At Meta, we focused on transforming our mantra of “move fast and break things,” to “move fast with stable infrastructure” by giving attention to the cultural elements of doing reliability work. This talk describes this process and decodes how to systematically measure the on-the-ground perspectives of engaging with reliability work so that we can have an informed perspective on how to best optimize the right degree of reliability. The audience will take away actionable practices that can be used to understand how to evaluate their underlying reliability culture, take data-driven approaches to measuring reliability sentiment and barriers and facilitators to performing the work, and identify practices that allowShow more
PracticePoster / ePoster
Sign in to star this4
The Sins of High Cardinality

Tuesday, 10 November 2026 · 13:30 to 14:15 · Workshop room

Milja Joki · Director of Engineering at Clearwater Storage

High cardinality is a sin in observability metrics collection. Cardinality in the form of granular labels or other metadata can cause exponential growth in your observability time-series storage and compute resources; costing money and slowing down queries. It's not always economical nor practical to collect all possible metrics with all possible labels and then worry about how to extract value using queries after the fact. Sure, we want as much granularity as possible in our observability data, but it's a trade-off, we need to be strategic in using metric cardinality to get the granularity needed to discover and remediate problems. This talk will focus on presenting different strategies to constrain the impact of high metrics cardinality referencing applicable open source Prometheus metriShow more
SystemsPoster / ePoster
Sign in to star this5

Wednesday, 11 November 2026

09:00
Scam or Savings? A Cloud vs. On-Prem Economic Slapfight

Wednesday, 11 November 2026 · 09:00 to 09:45 · Main hall

Harry Lee · Staff Engineer at Northbound Data

Since not giving a single crap about money turned out to be a purely zero-interest-rate phenomenon, the "cloud is overpriced vs. no it is not" argument has reignited—usually championed on either side by somebody with something to sell you. The speaker has nothing to sell you, and many truths to tell. It's time to find out where the truth is hiding.Show more
OpsTalk (25 min)
Sign in to star this4
Is It Already Time To Version Observability? (Signs Point To Yes.)

Wednesday, 11 November 2026 · 09:00 to 09:45 · Studio

Aline Yildirim · Site Reliability Engineer at Meridian Payments

Pillars, cardinality, metrics, dashboards ... the definition of observability has been debated to death, and I'm done with it. Let's just say that observability is a property of complex systems, just like reliability or performance. This definition feels both useful and true, and I am 100% behind it. However, there has recently been a generational sea change in data types, usability, workflows, and cost models, along with what users report is a massive, discontinuous leap in value. In the parlance of semantic versioning, it is a breaking, backwards-incompatible change. Which means it’s time to bump the major version number. Observability 1.0, meet Observability 2.0. In this presentation, we will outline the technical and sociotechnical characteristics of each generation of tooling and descShow more
OpsLightning talk (10 min)
Sign in to star this5
Automating Disaster Recovery: The Ultimate Reliability Challenge

Wednesday, 11 November 2026 · 09:00 to 09:45 · Workshop room

Jazzy Van Olderen · Director of Engineering at Ironwood Software

Here's how I explain my job to non-techies: if a meteor struck our servers, it's on my team to fix it. But what if it did? Realistically, what would happen if a meteor struck your datacenter? Here's the story of a vision, one to fully automate disaster recovery away, how I pushed back on it claiming it was impossible, and how we still executed on it to great success.Show more
OpsWorkshop (90 min)
Sign in to star this5
10:00
Compliance & Regulatory Standards Are NOT Incompatible with Modern Development Best ## Practices

Wednesday, 11 November 2026 · 10:00 to 10:45 · Main hall

Esther Giménez · Database Engineer at Latticework Systems

Modern software development is all about fast feedback loops, with best practices like testing in production, continuous delivery, observability driven development, and feature flags. Yet I often hear people complaining that only startups can get away with doing these things; real, grown-up companies are subject to regulatory oversight, which prevents engineers from deploying their own code due to separation of concerns, requires managers to sign off on changes, etc.Show more
PracticeTalk (45 min)
Sign in to star this4
What Can You See from Here?

Wednesday, 11 November 2026 · 10:00 to 10:45 · Studio

Marcus Mccoy · Head of Platform at Kestrel Labs

What you can see depends on where you stand. Your vantage point plays a big part in how you work, what you think is important, and how you interpret what's going on around you. Even your job title or team name can influence–and limit!–how you see the world. It's easy to find yourself stuck: no longer learning, polishing a service nobody else cares about, or wondering why it's so hard to get anything done outside your own team. A change in perspective can change what you think is important, how you influence the decisions that you care about, and even what you think is possible for yourself. In this talk, we'll look at how to get a broader view.Show more
OpsPoster / ePoster
Sign in to star this5
11:00
Coffee and posters