SRE / DevOps

King's Choice

Location
Remote
Employment
Full-time
Level
Mid-level
Category
DevOps & SRE
Posted

Description

We're looking for an SRE / DevOps engineer to build the reliability layer for a platform that commands live power plants. No unified log pipeline, no dashboard anyone watches, no rehearsed restore - the app runs in production, the layer underneath it is a set of habits. You decide what the observability stack looks like, what an alert may wake a person up for, and how failover is designed and tested.

Second part is physical: with the Infra Lead, dispatch control room - workstations, UPS, local network, secured remote access to plants. Grafana in one tab, a rack in the other.

Must have

4+ years SRE / DevOps / systems engineering, Linux daily: services, systemd, storage, performance diagnostics

Real experience with physical or VPS servers, not cloud consoles only - UPS, RAID, remote management. Cloud-only profile is an automatic no

Cloud VMs: provisioning, patching, hardening, capacity planning, TLS, nginx

Docker / Docker Compose in production: dev/staging/prod layouts, registry, zero-downtime deploys

Observability you built yourself: Prometheus + Grafana or equivalent, log pipeline (fluent-bit / ELK class), and an alerting philosophy you can defend

Redundancy you tested: PostgreSQL backup/restore and replication, Redis failover, measured recovery. You can describe a backup you personally restored

Networking: WireGuard or similar, firewalls, DNS, TLS, diagnosing cloud ↔ office ↔ plant links

CI/CD as code (GitHub Actions or similar), promotion through environments

Bash + Python for automation

English B2

Apply at the source

This role was published by King's Choice and listed via Djinni. Applications are handled there, not on this site.

View & apply on djinni.co ↗