User:SGrabarczuk (WMF)/Site Reliability Engineering
Notify us of emergencies with Klaxon.
About us
Would be nice to add something here about SRE as a whole.
Leadership
Teams
Collaboration Services
[edit]We are responsible for building and maintaining the infrastructure aspects of the source code management, CI and CD, task and ticket management systems as well as hosting non-MediaWiki websites and other collaboration services.
Data Center Operations
[edit]We are responsible for all of Wikimedia's data center deployments and logistics as well as maintaining our presence in locations across the world. We perform on-site work and maintain the full 5-year life cycle (specs, purchasing, physical install, break/fix and decommissioning) for all hardware.
Data Persistence
[edit]We focuse on Wikimedia's persistent data storage and retrieval systems, including (No)SQL databases, (distributed) object storage, file storage and backup systems.
#wikimedia-data-persistence connect
Infrastructure Foundations
[edit]We focus on building and maintaining our base platform (“metal cloud”) that forms the foundations upon which nearly everything else in our infrastructure builds upon. On top of our bare metal deployments, their responsibilities include (but are not limited to) configuration management systems, infrastructure automation, orchestration tooling, infrastructure security and network operations.
#wikimedia-sre-foundations connect
Observability
[edit]We provide teams with diagnostic tools, platforms, and insights into how systems and services perform. We leverage technologies such as Grafana, Kibana/Logstash, OpenSearch, Prometheus, AlertManager and more.
#wikimedia-observability connect #wikimedia-serviceops connect
Service Operations
[edit]We take care of public and “user-visible” services in close collaboration with the P&T teams. This includes our MediaWiki platform and the SOA service infrastructure based on Kubernetes.
Tools Infrastructure
[edit]We provide support and infrastructure for Toolforge and the Tools Platform team. We also maintain several other cloud services, including cloud-vps.
Traffic
[edit]We are responsible for the critical first layer of high-traffic infrastructure which now spans much of the globe, including our TLS termination and caching layers (ATS, Varnish), load balancing, DNS and our own network.
Contact
- See at SRE/SRE Team requests.
- #wikimedia-sre connect
- Additional documentation related to our infrastructure and the team's work can be found on Wikitech.