Principal Site Reliability Engineer / Platform Engineer
I build and operate production Kubernetes platforms for AI, data and real-time voice systems: cloud networking, GitOps delivery, secrets, observability, databases and disaster recovery.
10+ years · GKE, AWS, Yandex Cloud, Hetzner · Production systems
- Kubernetes platform engineering: GKE, EKS, Yandex Managed Kubernetes, Talos, Helm, operators, KEDA and Knative.
- Cloud & Infrastructure as Code: Terraform, Terragrunt, Ansible, Packer, AWS, GCP, Yandex Cloud and Hetzner.
- Security & secrets: OpenBao, Vault, Cloud KMS, Workload Identity, External Secrets, TLS and vulnerability scanning.
- Observability & reliability: Prometheus, Grafana, VictoriaMetrics, OpenTelemetry, Sentry, SLOs, incidents, backups and recovery tests.
- AI & data platforms: LLM gateways, embeddings, RAG, pgvector, Milvus, Kubeflow, MLflow, NiFi, Spark and Trino.
- Voice & networking: OpenSIPS, SIP, RTP, WebRTC, LiveKit, IPsec, Cloud NAT and hybrid connectivity.
- Delivery & automation: GitLab, GitLab Runner, Artifact Registry, Argo CD, CI/CD, Python and Bash automation.
- Making production systems predictable, observable and boring.
- Debugging hard incidents across Kubernetes, networks, databases and queues.
- Automating repetitive operational pain.
- Designing safer migrations, reversible rollouts and useful runbooks.
- Keeping dashboards, alerts and documentation low-noise and actionable.
Kubernetes operators, GitOps, OpenBao, workload identity, Gateway API, OpenTelemetry, eBPF, RAG, real-time voice infrastructure and production AI platforms.




