Top AI Repos — open-source AI, indexed and scored
Top AI Repos tracks AI repositories on GitHub and answers two different questions about each one: is it moving right now, and would you bet a product on it.
Top AI Repos tracks AI repositories on GitHub and answers two different questions about each one: is it moving right now, and would you bet a product on it.
A curated list of Site Reliability and Production Engineering resources.
| Date | Stars |
|---|---|
| 2026-07-24 | 13398 |
| 2026-07-25 | 13400 |
| 2026-07-28 | 13400 |
| 2026-07-30 | 13400 |
| 2026-08-06 | 13400 |
Today
— stars today
This week
— stars this week
This month
— stars this month
Momentum
0.0
growth rate 0.00%/day
# Awesome Site Reliability Engineering [](https://github.com/sindresorhus/awesome) [<img src="awesome-sre-logo.svg" align="right" width="100">](https://dastergon.gr/awesome-sre) A curated list of awesome [Site Reliability](https://www.usenix.org/conference/srecon14/technical-sessions/presentation/keys-sre) and [Production](https://www.usenix.org/conference/srecon15/program/presentation/canahuati) Engineering resources. #### What is Site Reliability Engineering? > "Fundamentally, it's what happens when you ask a software engineer to design an operations function." - Ben Treynor Sloss, VP Google Engineering, founder of Google SRE ## Contributing Please take a look at the [contribution guidelines](CONTRIBUTING.md) first. Contributions are always welcome! ## Contents - [Culture](#culture) - [Education](#education) - [Books](#books) - [Hiring](#hiring) - [Reliability](#reliability) - [Monitoring & Observability & Alerting](#monitoring--observability--alerting) - [On-Call](#on-call) - [Post-Mortem](#post-mortem) - [Capacity Planning](#capacity-planning) - [Service Level Agreement](#service-level-agreement) - [Performance](#performance) - [Programming](#programming) - [Misc Articles](#misc-articles) - [Real-time Messaging](#real-time-messaging) - [Blogs](#blogs) - [Newsletters](#newsletters) - [Conferences & Meetups](#conferences-meetups) - [Twitter](#twitter) - [SRE Tools](#sre-tools) - [SRE Podcasts](#podcasts) ## Culture * [What is Site Reliability Engineering?](https://landing.google.com/sre/interview/ben-treynor.html) * [Keys To SRE by Ben Treynor](https://www.usenix.org/conference/srecon14/technical-sessions/presentation/keys-sre) * [Google SRE Resources](https://landing.google.com/sre/resources.html) * [Notes from Production Engineering by Pedro Canahuati](https://www.usenix.org/conference/srecon15/program/presentation/canahuati) * [PostOps: Recovery from Operations](https://www.usenix.org/conference/srecon15europe/program/presentation/underwood) * [Love DevOps? Wait 'till you meet SRE](https://www.atlassian.com/it-service/site-reliability-engineering-sre) [[video]](https://youtu.be/fsTpRx8Pt-k) * [How Google Does Planet-Scale Engineering for Planet-Scale Infra](https://www.youtube.com/watch?v=H4vMcD7zKM0) * [Site Reliability Engineering at Facebook](https://www.facebook.com/notes/facebook-engineering/site-reliability-engineering-at-facebook/291616313919/) * [A History of Site Reliability Engineering at Uber](https://www.youtube.com/watch?v=qJnS-EfIIIE&nohtml5=False) * [Case Study: Adopting SRE Principles at StackOverflow](https://www.usenix.org/conference/srecon15/program/presentation/limoncelli) * [Site Reliability Engineering at Dropbox](https://www.youtube.com/watch?v=ggizCjUCCqE) * [Site Reliability Engineers — Keeping Google up and running 24/7](https://www.youtube.com/watch?v=yXI7r0_J29M) * [Site Reliability Engineering at Salesforce](https://www.salesforce.com/video/193050/) * From Sys Admin to Netflix SRE - [video](https://www.youtube.com/watch?v=lZI51YzIgVE) and [slides](https://www.socallinuxexpo.org/sites/default/files/presentations/Scale%20x14%20Slides.pdf) * [SRE@Google: Thousands of DevOps Since 2004](https://www.youtube.com/watch?v=iIuTnhdTzK0) * [Transactional System Administration Is Killing Us and Must be Stopped](https://www.usenix.org/conference/lisa15/conference-program/presentation/limoncelli) * [A hierarchy of SRE needs](https://web.archive.org/web/20190401220948/https://plus.google.com/+lizthegrey/posts/MLAJFVyEb2f) * [PostOps: A Non-Surgical Tale of Software, Fragility, and Reliability](https://www.usenix.org/conference/lisa13/technical-sessions/plenary/underwood) * [SRE: An incomplete guide to cultural Narnia](https://web.archive.org/web/20180820235243/http://anthonycaiafa.com/2016/04/10/sre-cultural-narnia/) - [[Video]](https://www.youtube.com/watch?v=__wypEhdcrQ&t=0s) * [Putting T
Excerpt of 61,505 characters
Read on GitHubPavlos Ratis · Apple · United Kingdom
525
5
5
5
4
3
3
Raghu Chinnannan · India
3
Denny Schäfer · AVIV Germany
3
2
2
2
Peter Thaleikis · @bringyourownideas
2
2
Lasantha Kularatne · United States
2
2
2
b murphy · @Grafana
2
2
2
Would you bet a product on this? Bounded 0–100 and slow moving.
matched fp:31be11fae6a79dd9, topic:awesome, topic:awesome-list, desc:curated list
matched fp:31be11fae6a79dd9, topic:monitoring, readme:observability