Agenda
AI Engineer · all times Europe/London
Tuesday, 10 November 2026
Tuesday, 10 November 2026 · 09:00 to 09:45 · Main hall
Tvorimir Luchakivskiy · Principal Engineer at Latticework Systems
Cloud scale provides the vast resources necessary to repair and fix failed components, but this is useful only if those failures can be detected. For this reason, the major availability breakdowns and performance anomalies we see in cloud environments tend to be caused by subtle underlying faults, i.e., gray failure rather than fail-stop failure. In this talk, we discuss our experiences with gray failure in Microsoft Azure to show its broad scope and consequences with several case studies. We also argue that a key feature of gray failure is differential observability: that the system’s failure detectors may not notice problems even when applications are afflicted by them. We will show how Microsoft Azure applied the differential observability in practice and bridged the gap between differeShow moreShow less
Tuesday, 10 November 2026 · 09:00 to 09:45 · Studio
Marcus Mccoy · Head of Platform at Kestrel Labs
Dr. David Woods once said to me “We cannot call it a scientific field unless we can admit we’ve gotten things wrong in the past.” Do we, in this community, do that? Well-formed critique is critical for any field — including SRE — to progress. I’d like to talk about a few ideas, assumptions, and concepts often talked about in this community but whose validity is rarely questioned or explored. John Allspaw has worked in software systems engineering and operations for over twenty years in many different environments. John’s publications include the books *The Art of Capacity Planning* (2009) and *Web Operations* (2010) as well as the forward to *The DevOps Handbook.* His 2009 Velocity talk with Paul Hammond, *10+ Deploys Per Day: Dev and Ops Cooperation* helped start the DevOps movement. JohnShow moreShow less
Tuesday, 10 November 2026 · 09:00 to 09:45 · Workshop room
Nora Leroy · Distinguished Engineer at Fathom Analytics Group
As our organization has gotten very good at protecting server SLOs with reliability best practices like scaling globally distributed at-scale architectures, toil mitigation, and continuous reliability improvements we noticed that a majority of incidents impacting our end-users were not showing up as an SLO miss. In many cases these outages were not even observable from the server side - for example, the rollout of a new version of the consumer mobile application (that our services powers) to an app store could break one or more critical feature(s) due to bugs in client code. This reality has led to a change in the way we approach reliability - we’re shifting our focus from server reliability to product reliability. We’re not yet finished with the transition, but we’re starting to see very Show moreShow less
Tuesday, 10 November 2026 · 10:00 to 10:45 · Main hall
Sidónio Moreira · Research Scientist at Clearwater Storage
What happens when you've got a team of SREs who get downsized and smacked with an increasing workload? This is the story of my small team's wild ride of revamping our in-house incident tool. Spoiler: It turned into a "make or buy" debate real quick! Picture this: a team of 10 becomes a merry band of 3, drowning in a sea of issues, requests, and feedback. Unexpectedly, our team changed like a game of musical chairs, skills shuffling all over. In our pursuit to be the heroes, we hit the brakes and pitched buying a ready-made tool for a change. It wasn't just a business decision; it was a cultural rollercoaster for everyone involved! So, let's take a stroll down memory lane, exploring how we switched gears to evaluate third-party tools to pave the ground for improving the reliability of our pShow moreShow less
Tuesday, 10 November 2026 · 10:00 to 10:45 · Studio
Irmingard Frese · Staff Engineer at Kestrel Labs
Apache Pinot, a real-time, distributed OLAP database, encountered resiliency challenges in its large production deployment due to server failures, slowness, and unpredictable query patterns. To ensure adherence to strict SLAs, enhancements were made to its Query Processing framework through two key features: Intelligent Query Routing and Runtime Query Killing. Intelligent Query Routing replaces the traditional round-robin server query distribution with a dynamic approach based on continuous server performance metrics. This adaptive method proactively routes queries to more efficient servers, reducing the need for manual intervention and preserving availability during transient server issues. Runtime Query Killing addresses the issue of complex queries causing Out of Memory errors and availShow moreShow less
Tuesday, 10 November 2026 · 10:00 to 10:45 · Workshop room
Andrea Johnson · Head of Platform at Meridian Payments
This presentation details our transition from JupyterHub to the Kubeflow Notebook Controller. JupyterHub was architected in a backend agnostic way that "supports" Kubernetes but isn't truly Kubernetes-native. As a result, it has significant shortcomings with respect to resilience and high availability. In particular, the core component, the hub API, can only have one replica at any given time. In contrast, The Kubeflow Notebook controller is built from the ground up to be Kubernetes native using the operator pattern. There's far less complexity, fewer components, less brittleness, and improved resilience and high availability. As a result, our platform has been able to scale to four times as many users, including ten times as many concurrent executions. Our users are happier and there's leShow moreShow less
Tuesday, 10 November 2026 · 11:00 to 11:45 · Main hall
Larisa Kladochniy · Research Scientist at Jane Street
Modern cloud systems are orchestrated by many independent and interacting subsystems, each specializing in important services such as scheduling, data processing and storage, resource management, etc. Hence, overall cloud system reliability is determined not only by the reliability of each individual subsystem, but also by the interactions between them. With recent practices of microservice and serverless architectures, each individual subsystem becomes simpler and smaller, while their interactions grow in complexity and diversity. We observe that many recent production incidents of large-scale cloud systems are manifested through failures of the interactions across system boundaries, which we term "cross-system interaction failures", or "CSI failures". However, understanding and addressinShow moreShow less
Tuesday, 10 November 2026 · 11:00 to 11:45 · Studio
Yashodha Kaur · Observability Engineer at Snowflake
Managing incidents well requires a number of skills — debugging and systems understanding, strong communication, and high-speed project management — but we rarely talk about the power of storytelling in our incident management loop. From oncall preparation, to incident handling, to postmortem creation, skill at storytelling can support and even improve our engineering skills. Establishing a setting, building a coherent narrative, and understanding your characters builds memorable and compelling communication that can level up all stages of your incident management. Laura de Vesine is a 20+ year software industry veteran. She has spent the last 8 years in SRE working in incident analysis and prevention, chaos engineering, and the intersection of technology and organizational culture. Laura Show moreShow less
Tuesday, 10 November 2026 · 11:00 to 11:45 · Workshop room
Noemí Marrero · Engineering Manager at Fathom Analytics Group
Your CI/CD Process is chock full of credentials, and almost anyone in your company has access to it. Configuring your CI correctly is vital to supply chain security. We discuss how to reduce that attack surface by enforcing proper branch permissions and using OIDC to reduce long-lived credentials and tie branches to roles.Show moreShow less
Tuesday, 10 November 2026 · 13:30 to 14:15 · Main hall
Madeleine Skotnes · Site Reliability Engineer at Halyard Cloud
Synthetic monitoring, particularly browser-based monitoring, is hard to do well. When tests pass, synthetic monitoring provides a uniquely intuitive kind of psychological safety - human-like, verified confidence, compared to other forms of monitoring. When tests fail, synthetic monitoring is often blamed as flaky, misconfigured, or unreliable. If not properly implemented, it can not only be financially, mentally and organisationally draining, but damaging to real customer experience. This talk is a conceptual and technical story of 4 years working with Atlassian’s in-house synthetic monitoring solution, being the owning developer for a tool actively used by 30-40 internal teams to build and manage synthetic monitoring for Jira. How can we make synthetic monitoring better serve its purpose Show moreShow less
Tuesday, 10 November 2026 · 13:30 to 14:15 · Studio
Carolina Concepción · Site Reliability Engineer at Fathom Analytics Group
Understanding how culture influences engineering practices is often a black box. At Meta, we focused on transforming our mantra of “move fast and break things,” to “move fast with stable infrastructure” by giving attention to the cultural elements of doing reliability work. This talk describes this process and decodes how to systematically measure the on-the-ground perspectives of engaging with reliability work so that we can have an informed perspective on how to best optimize the right degree of reliability. The audience will take away actionable practices that can be used to understand how to evaluate their underlying reliability culture, take data-driven approaches to measuring reliability sentiment and barriers and facilitators to performing the work, and identify practices that allowShow moreShow less
Tuesday, 10 November 2026 · 13:30 to 14:15 · Workshop room
Milja Joki · Director of Engineering at Clearwater Storage
High cardinality is a sin in observability metrics collection. Cardinality in the form of granular labels or other metadata can cause exponential growth in your observability time-series storage and compute resources; costing money and slowing down queries. It's not always economical nor practical to collect all possible metrics with all possible labels and then worry about how to extract value using queries after the fact. Sure, we want as much granularity as possible in our observability data, but it's a trade-off, we need to be strategic in using metric cardinality to get the granularity needed to discover and remediate problems. This talk will focus on presenting different strategies to constrain the impact of high metrics cardinality referencing applicable open source Prometheus metriShow moreShow less
Wednesday, 11 November 2026
Wednesday, 11 November 2026 · 09:00 to 09:45 · Main hall
Harry Lee · Staff Engineer at Northbound Data
Since not giving a single crap about money turned out to be a purely zero-interest-rate phenomenon, the "cloud is overpriced vs. no it is not" argument has reignited—usually championed on either side by somebody with something to sell you. The speaker has nothing to sell you, and many truths to tell. It's time to find out where the truth is hiding.Show moreShow less
Wednesday, 11 November 2026 · 09:00 to 09:45 · Studio
Aline Yildirim · Site Reliability Engineer at Meridian Payments
Pillars, cardinality, metrics, dashboards ... the definition of observability has been debated to death, and I'm done with it. Let's just say that observability is a property of complex systems, just like reliability or performance. This definition feels both useful and true, and I am 100% behind it. However, there has recently been a generational sea change in data types, usability, workflows, and cost models, along with what users report is a massive, discontinuous leap in value. In the parlance of semantic versioning, it is a breaking, backwards-incompatible change. Which means it’s time to bump the major version number. Observability 1.0, meet Observability 2.0. In this presentation, we will outline the technical and sociotechnical characteristics of each generation of tooling and descShow moreShow less
Wednesday, 11 November 2026 · 09:00 to 09:45 · Workshop room
Jazzy Van Olderen · Director of Engineering at Ironwood Software
Here's how I explain my job to non-techies: if a meteor struck our servers, it's on my team to fix it. But what if it did? Realistically, what would happen if a meteor struck your datacenter? Here's the story of a vision, one to fully automate disaster recovery away, how I pushed back on it claiming it was impossible, and how we still executed on it to great success.Show moreShow less
Wednesday, 11 November 2026 · 10:00 to 10:45 · Main hall
Esther Giménez · Database Engineer at Latticework Systems
Modern software development is all about fast feedback loops, with best practices like testing in production, continuous delivery, observability driven development, and feature flags. Yet I often hear people complaining that only startups can get away with doing these things; real, grown-up companies are subject to regulatory oversight, which prevents engineers from deploying their own code due to separation of concerns, requires managers to sign off on changes, etc.Show moreShow less
Wednesday, 11 November 2026 · 10:00 to 10:45 · Studio
Marcus Mccoy · Head of Platform at Kestrel Labs