Course · Training · Workshop

Site Reliability Engineering

Two-day SRE and testing workshop with a shift-left approach: track SLA-relevant metrics and SLOs, build tests from unit to end-to-end and analyse deployment failures.

Transform your team's approach with our concise SRE and testing workshop. Learn to identify key questions, focus on meaningful testing and track SLA-relevant metrics. Our "shift left" approach to SRE will guide your team to prioritize effectively in complex environments. The workshop also fosters a fun and engaging culture around all types of testing, from unit to end-to-end testing and deployment failure analysis. Improve your team's SRE and testing skills in a fun, collaborative environment.

What participants say

These customers booked courses in the same topic cluster.More customers →

Content

Our SRE and Testing Workshop explores the key elements of Site Reliability Engineering, combined with a hands-on approach to testing and failure analysis. Via a blend of theory and exercises, this course covers both the basics and advanced techniques in the following areas:

  • Introduction to Site Reliability Engineering
  • Determining Key Performance Indicators (KPIs) and Service Level Agreements (SLAs)
  • Developing effective testing strategies, from unit to end-to-end tests
  • Analyzing and preventing deployment failures
  • Implementing a ‘shift left’ approach to enhance reliability early in the development process
  • Practice-based case studies and group exercises to reinforce learning
  • Creating and managing a collaborative testing culture within the team

The goal is to equip participants with the necessary tools and techniques to make their applications and systems more reliable and their development processes more efficient.

The actual course content may differ from the above depending on the trainer, delivery, duration and the composition of participants.

Request this course in-house

By submitting you accept our Privacy Policy.

Request a public date

No suitable public date? Register without obligation — once there is enough interest we schedule a new public date and let you know first.

Number of participants (approx.)

More than 3 participants? Best to request a dedicated in-house date directly.

By submitting you accept our Privacy Policy.

More about Site Reliability Engineering (SRE)

Site Reliability Engineering (SRE) is an approach that applies software engineering principles to infrastructure and operations problems to create scalable and highly reliable software systems. Developed by Google, the concept merges aspects of traditional IT operations with agile software development and defines clear objectives such as Service Level Objectives (SLOs) and error budgets. SRE promotes a culture of shared ownership between development and operations teams, enabling faster and safer deployments.

Further resources:

History

Site Reliability Engineering was developed at Google in the mid-2000s under the leadership of Ben Treynor Sloss. When Google faced the challenge of reliably operating its rapidly growing infrastructure in 2003, Treynor Sloss founded the first SRE team with the goal of deploying software engineers for operations work, bringing automation and technical excellence to the forefront. The core principles – including error budgets, toil reduction, and blameless postmortems – became landmark concepts for the entire industry.

In 2016, Google published the book "Site Reliability Engineering", making the principles and practices of SRE accessible to the entire software industry. Since then, SRE has gained worldwide adoption and significantly influences modern DevOps practices, platform engineering, and the way organizations balance reliability with the speed of innovation.