All work

Professional work / Striveworks

Production incident leadership

Coordinating diagnosis, service restoration, and follow-through for a multi-region, GPU-backed MLOps platform.

Senior Site Reliability Engineer · 2023–2024

Context & constraints

At Striveworks, I owned production reliability, observability, and operational strategy for a multi-region, GPU-backed MLOps platform. Reliability work spanned infrastructure and application components.

My contribution

As incident commander for high-severity production incidents, I coordinated cross-functional diagnosis and service restoration, followed by post-incident analysis and remediation.

I defined and implemented SRE practices including SLIs and SLOs, automated alerting, incident workflows, and operational readiness reviews.

I worked with product and engineering teams on capacity, production behavior, and release readiness.

Engineering focus

Incident coordination
Connect the people diagnosing the failure with the work needed to restore service and address what happened.
Operational visibility
Use observability, service objectives, and automated alerting to support operational decisions.
Release readiness
Bring production behavior and reliability requirements into engineering planning.