DevOps Engineer

whitelark.io

Location
Remote
Employment
Full-time
Level
Specialist
Category
DevOps & SRE
Posted

Description

Requirements:

Must have

5+ years in DevOps, SRE, or Platform Engineering.

Production AWS experience (EKS, IAM, VPC, RDS).

Terraform — modules, remote state, multi-environment infrastructure.

Kubernetes — production operations, Helm.

CI/CD — building, maintaining, and improving pipelines.

Linux, networking, scripting (Bash + Python or Go).

Secrets management and security fundamentals (IAM, least privilege).

Observability — metrics, logs, and alerting (Prometheus/Grafana or equivalent).

AI fluency — actively uses AI tools (Claude Code, Cursor, Codex, or similar) in day-to-day engineering work. Not just "aware of AI" — uses them to multiply productivity while applying engineering judgment to the output.

Proactivity — identifies problems before they surface, proposes solutions, and drives them to completion. Doesn't wait to be told what to do.

English: confident reading and listening is a must.

What you'll do:

Infrastructure as Code: Design and evolve Terraform modules for AWS (composite modules, accounts as code), and build a multi-environment foundation across non-prod / pre-prod / prod.

Kubernetes: Operate and evolve EKS clusters — workloads, namespaces, resources, autoscaling, and network policies.

CI/CD: Build, maintain, and improve GitHub Actions pipelines — faster, more reliable delivery with fewer manual steps.

GitOps / Delivery: Manage Helm charts and ArgoCD for declarative Kubernetes deployments.

Secrets & Access: Own secrets management, rotation, least-privilege access, and keep secrets out of code.

Observability: Build metrics, dashboards, and alerting (Amazon Managed Prometheus + Amazon Managed Grafana), define SLOs, and improve incident response.

Security & Networking: Manage IAM, environment separation, and secure access with Okta and Twingate.

Databases (Operations): Operate PostgreSQL (RDS) and Redshift — configuration, backups, secret rotation, and operational reliability.

Reliability: Drive production readiness, capacity planning, cost optimization, HA where it matters, and operational runbooks.

Stack:

Cloud: AWS (EKS, RDS, IAM, VPC, etc.)

IaC: Terraform

Containers: Docker, Kubernetes (EKS), Helm

GitOps: ArgoCD

CI/CD: GitHub Actions

Observability: Amazon Managed Prometheus (AMP), Amazon Managed Grafana (AMG)

Database: PostgreSQL (RDS), AWS Redshift (OLAP)

Auth / Access: Okta, Twingate

AI: AI-first — active use of AI tools and AI agents in daily engineering work.

Nice to have

FinTech, payments, or banking infrastructure experience.

Experience building toward or maintaining PCI DSS compliance.

ArgoCD / GitOps.

Amazon Managed Prometheus / Grafana.

AWS Redshift.

Multi-account AWS environments with clear non-prod / pre-prod / prod separation.

Apply at the source

This role was published by whitelark.io and listed via Djinni. Applications are handled there, not on this site.

View & apply on djinni.co ↗