Context & constraints
At Striveworks, I owned production reliability, observability, and operational strategy for a multi-region, GPU-backed MLOps platform. Reliability work spanned infrastructure and application components.
My contribution
As incident commander for high-severity production incidents, I coordinated cross-functional diagnosis and service restoration, followed by post-incident analysis and remediation.
I defined and implemented SRE practices including SLIs and SLOs, automated alerting, incident workflows, and operational readiness reviews.
I worked with product and engineering teams on capacity, production behavior, and release readiness.
Engineering focus
- Incident coordination
- Connect the people diagnosing the failure with the work needed to restore service and address what happened.
- Operational visibility
- Use observability, service objectives, and automated alerting to support operational decisions.
- Release readiness
- Bring production behavior and reliability requirements into engineering planning.