Senior DevOps Engineer — Relocation to Dubai

UAPP.AI

Location
Remote
Employment
Full-time
Level
Mid-level
Category
DevOps & SRE
Posted

Description

We are a product company building AI-powered mobile applications used by over 10 million people in 150 countries. We own the entire product lifecycle — from the initial idea to large-scale global implementation. We thrive in a fast-paced environment, making data-driven decisions without unnecessary bureaucracy. If you value autonomy, rapid releases, and seeing the direct impact of your work, you will fit right in.

The Role

We are looking for a Senior DevOps Engineer to take full ownership and leadership of the infrastructure for our ambitious project in the Cloud Gaming & Streaming space.

In this role, you will architect, scale, and maintain a high-performance, fault-tolerant system provisioning and streaming content across a server fleet growing from hundreds toward thousands of machines. You will lead infrastructure decisions, set technical standards, build key systems (like observability and bare-metal orchestration) from scratch, and mentor a Mid DevOps Engineer on the team.

Work Format & Relocation

Location: Dubai, UAE (Office).

Trial Period: First 2–3 months are fully remote for a smooth onboarding.

Relocation Support: Full corporate visa sponsorship and relocation package provided after the trial period.

Key Responsibilities

Fleet Architecture & Provisioning: Architect and automate end-to-end bare-metal and virtualized host provisioning — storage, GPU passthrough (vfio), and VM lifecycle across a large, heterogeneous server fleet.

Orchestration & Scale: Design and maintain fleet-scale orchestration systems: parallel deployments, health checks, and self-healing routines ensuring consistency across thousands of machines.

Infrastructure as Code & Configuration: Lead IaC practices using Terraform and Ansible to maintain reproducible host configurations, deterministic automation, and VM images.

Advanced Virtualization & OS Administration: Deeply manage Linux/KVM virtualization (libvirt, QEMU) and Windows guest systems, including complex GPU passthrough and low-latency virtual networking.

Greenfield Observability: Architect and implement the observability stack from scratch (Prometheus, Grafana, Loki/ELK) across the entire fleet for deep performance monitoring and alerting.

Network & Traffic Management: Design and optimize network configurations at scale — NAT, port forwarding, routing, and traffic paths for thousands of concurrent, low-latency streaming clients.

CI/CD & Deterministic Automation: Build robust CI/CD pipelines (GitLab CI / GitHub Actions) for infrastructure code, images, and deterministic scripting (Bash, Python, PowerShell).

Security & Storage: Implement enterprise-grade secrets management (Vault / SOPS) and manage large-scale content distribution via S3-compatible object storage (e.g., Cloudflare R2).

Leadership & Incident Response: Set reliability engineering standards, drive incident response/on-call processes, and collaborate closely with core C++/C# infrastructure engineers.

Requirements

Experience: 4+ years in a DevOps / Infrastructure role with proven experience building and managing large-scale, high-load systems or large server fleets (dozens/hundreds+ of nodes).

OS & Virtualization Expertise: Deep administration skills in both Linux and Windows; practical expertise with KVM/libvirt, QEMU, and GPU passthrough / vfio.

IaC & Configuration Management: Strong hands-on experience with Terraform and Ansible for host consistency, image management, and fleet-wide automation.

Scripting Mastery: High proficiency in Python, Bash, or PowerShell with the ability to write deterministic, idempotent, and restart-safe automation scripts across heterogeneous hardware.

Networking & Observability: Strong grasp of Linux networking (NAT, iptables, routing, traffic shaping) and experience standing up monitoring/logging stacks (Prometheus, Grafana, Loki/ELK) from scratch.

CI/CD & Security: Experience automating infrastructure via GitLab CI/GitHub Actions and managing secrets at scale (Vault, SOPS).

Language: English (Upper-Intermediate / B2 or higher — required for technical collaboration).

Nice-to-have

Direct experience in Cloud Gaming, real-time data streaming, or high-load gaming environments.

Hands-on experience with GPU-heavy workloads and large-file distribution via object storage (S3 / Cloudflare R2).

Basic knowledge of C++ or C# to collaborate effectively with core streaming infrastructure developers.

What We Offer

Architectural Greenfield & Autonomy: Full ownership to design the fleet observability, orchestration, and automation architecture from zero.

Leadership & Real Impact: Direct influence on the tech stack of a Cloud Gaming platform targeting a global, multi-million user audience.

Relocation Package: Full financial and organizational support for your move to Dubai (corporate visa, flight tickets, initial housing) after the trial period.

Modern Office: Work from a vibrant, tech-driven workspace in Dubai alongside a strong engineering team.

Fast-Paced Culture: A product-driven environment with quick decision-making, high engineering freedom, and zero bureaucracy.

Compensation: Highly competitive USD-pegged salary matching your senior technical background.

Apply at the source

This role was published by UAPP.AI and listed via Djinni. Applications are handled there, not on this site.

View & apply on djinni.co ↗