Senior DevOps / Site Reliability Engineer β 8+ years designing, automating, and operating production infrastructure, from bare-metal datacenters to Kubernetes.
Currently I run the platform behind a B2B search product serving ~3M API requests/day: a bare-metal fleet with virtualization and distributed storage, a 50+ node full-text search cluster, and five Kubernetes clusters.
- Kubernetes at production scale β deploy and operate clusters; migrate stateful systems from bare metal into K8s with zero customer-facing downtime
- Observability that people trust β metrics and logs pipelines, data-driven alert thresholds, and a runbook behind every alert
- Incidents end to end β root cause on live systems, fix, then verify with measurements before calling it done
- Backup & disaster recovery β HA backup infrastructure covering 60+ hosts, tested restores, monitored liveness
- Internal tooling β Go, Python, and TypeScript: from a Kubernetes admission controller to a full alert-management console used as the team's on-call cockpit
A growing part of my work is making AI agents do real infrastructure tasks safely: MCP servers that expose production tooling to agents with controlled scopes, and autonomous agents embedded into daily operations. The latest one is a production agent that compiles daily infrastructure digests from Git, Jira, and Slack β built on a least-privilege credential model with no direct access to the infrastructure it reports on.
| Project | What it is |
|---|---|
| KENT | Kubernetes events exporter in Go: ships cluster events to VictoriaLogs or stdout β namespace filters, multi-tenant streams, Helm chart, health probes. Born from a real production monitoring need |
| gocd-mcp | MCP server that lets AI agents operate GoCD CI/CD pipelines β inspect, trigger, and manage runs through a controlled interface |
| beaver | Go event pipeline: live Wikimedia edit stream β Kafka β ClickHouse (event log) + PostgreSQL (materialized state), with at-least-once delivery and Prometheus metrics |
Claims get verified against reality. Documentation is fact-checked against live systems, fixes are confirmed with numbers, and every alert ships with a runbook.
π Certified Kubernetes Administrator (CKA) β The Linux Foundation



