About Justin Bailey

Software, platforms,
and the people using them.

I'm a software and platform engineer, technical lead, and U.S. military veteran. My work connects the software teams build with the infrastructure and operating practices that keep it dependable.

How I work

Platforms are products.

I care about the interface as much as the implementation: what a developer needs to express, what the platform should own, and how an operator knows it is working.

My experience spans Kubernetes operators and APIs, self-service infrastructure, production reliability, container maintenance, and technical delivery. I stay close to the implementation while connecting engineering decisions to the people using the system.

Beyond professional work, I operate a home environment that keeps me learning. The portfolio explores that work in more detail.

Professional experience

Roles & contributions.

2023–Present

BrainGu

Platform Development Technical Lead

Build platform software, Kubernetes capabilities, and self-service infrastructure while leading delivery and customer-facing technical work.

  • Developed a Kubernetes operator covering tenant onboarding, RBAC, network policy, secrets management, and upgrade safety gates across EKS and bare-metal environments.
  • Developed Crossplane-based capabilities and self-service workflows for cloud provisioning.
  • Built a self-service infrastructure catalog using OpenTofu and Terragrunt.
  • Partnered with developers, product teams, security engineers, and customers to translate recurring requirements into reusable platform APIs and deployment patterns.
  • Automated container maintenance workflows for Platform One Iron Bank and internal products.
  • Go
  • Kubernetes
  • Crossplane
  • OpenTofu
  • Terragrunt
  • GitOps

2023–2024

Striveworks

Senior Site Reliability Engineer

Owned production reliability for a multi-region, GPU-backed MLOps platform.

  • Acted as incident commander for high-severity production incidents, coordinating diagnosis, service restoration, post-incident analysis, and remediation.
  • Automated recurring operational work and established SRE practices.
  • Observability
  • SRE
  • SLIs & SLOs
  • Incident response

2021–2023

NetApp

DataOps and Observability Engineer

Established a DataOps and observability function for Azure NetApp Files, connecting operational telemetry with engineering decisions.

  • Connected metrics, logs, and telemetry to troubleshooting and operational decisions for distributed storage.
  • Partnered with SRE, product, and engineering teams to define observability requirements and address systemic reliability issues.
  • Azure NetApp Files
  • Observability
  • DataOps
  • Metrics & logs

2019–2022

GovCIO (formerly Salient CRGT)

DevOps Engineer

Built enterprise and edge infrastructure automation for mission-critical systems.

  • Automated system configuration with Ansible and Nautobot.
  • Built tooling for repeatable bare-metal deployment.
  • Mentored DevOps engineers on Linux automation and operational best practices.
  • Ansible
  • Nautobot
  • Infrastructure automation

Foundations

Service & systems.

  • ARMA Global / General DynamicsSenior System Administrator · 2018–2019
  • United States Army ReserveIntelligence Analyst · 2018–2022
  • United States ArmyFire Direction Control Specialist · 2014–2018

Education

Norwich University

Bachelor of Science — Vulnerability and Computer Forensics

Graduated in 2022.