Senior MLOps / ML Platform Engineer

Sigma Software

Location
Remote
Employment
Full-time
Level
Specialist
Category
Data Science & ML
Posted

Description

We are looking for a Senior MLOps Engineer to join Sigma Software and help build a production-grade ML platform for a large-scale AdTech ecosystem. You will work on infrastructure powering predictive decision-making systems that process hundreds of millions of auction requests daily.

As part of a dedicated engineering team, you will contribute to scalable ML orchestration, model lifecycle automation, observability, and real-time optimization workflows. This role is ideal for engineers with strong production experience who enjoy solving complex platform and operational challenges.

We at Sigma Software offer the opportunity to work on cutting-edge ML infrastructure projects, collaborate with experienced engineers, and influence architecture decisions in a long-term strategic engagement.

CUSTOMER

Our Customer is a technology company operating supply-side infrastructure within the programmatic advertising ecosystem. The company manages a high-load ad exchange platform handling hundreds of millions of auction requests every day and is investing in advanced predictive decision-making capabilities to improve advertiser outcomes and real-time optimization processes.

PROJECT

Sigma Software is building a predictive modeling and optimization platform integrated with a live ad exchange environment. The solution enables real-time supply scoring and filtering, audience look-alike generation, contextual performance estimation, and multi-objective optimization under operational constraints.

The project combines large-scale ML infrastructure, automated model lifecycle management, multi-tenant architecture, and advanced observability practices. The team focuses on delivering reliable, reproducible, and scalable ML systems ready for long-term Customer ownership.

Key Technologies: Python, Kubernetes, Docker, GCP, Vertex AI, MLflow, Airflow, Kubeflow, Argo Workflows, Terraform

Job Description

Build and maintain ML training orchestration pipelines across hourly, daily, and weekly schedules

Implement retries, backfills, and idempotent execution mechanisms

Design and support model registry workflows including versioning, lineage, evaluation gates, and promotion processes

Develop isolated per-advertiser model environments with namespace and configuration separation

Build scalable refresh pipelines and publishing workflows for serving infrastructure

Implement shadow mode and champion/challenger deployment strategies

Develop monitoring and alerting for ML-specific metrics including feature drift, prediction drift, train/serve skew, and calibration decay

Ensure reproducibility of ML workflows using containerized environments, pinned dependencies, and data snapshots

Monitor training and scoring costs across tenants

Collaborate with DevOps and SRE engineers on CI/CD and infrastructure automation

Prepare operational documentation and platform handover materials

Qualifications

5+ years of experience in MLOps, ML platform engineering, or infrastructure engineering supporting production ML systems

Strong Python skills and experience building platform-level tooling and automation

Hands-on experience with Kubernetes and Docker

Experience building CI/CD pipelines for ML workloads

Hands-on production experience with MLflow, Kubeflow, Airflow, Argo Workflows, Vertex Pipelines, or similar orchestration and ML lifecycle platforms

Experience with ML platforms and model lifecycle tools such as Vertex AI, MLflow, or Kubeflow

Strong understanding of ML observability including drift detection, train/serve skew monitoring, and incident response

Experience designing or supporting multi-tenant ML systems and isolated model environments

Experience working with cloud platforms, preferably GCP

Experience with infrastructure-as-code tools such as Terraform

Experience with Linux environments

Understanding of the ML lifecycle and productionization processes

Upper-Intermediate English level or higher

WILL BE A PLUS

Experience with feature stores and feature consistency management

Experience with large-scale batch scoring systems operating under freshness SLAs

Familiarity with experiment tracking platforms and evaluation gates

Experience with on-premises Kubernetes or bare-metal Linux infrastructure

Knowledge of DVC, lakeFS, or other data versioning tools

Experience with Bigtable, Redis, Aerospike, or similar low-latency serving databases

GPU scheduling and training cost optimization experience

Familiarity with SOC 2, ISO 27001, or GDPR-related compliance requirements

PERSONAL PROFILE

Strong ownership mindset and focus on operational reliability

Ability to work independently in complex distributed systems environments

Strong collaboration and communication skills

Analytical thinking with attention to scalability and maintainability

Comfortable working in fast-paced product-oriented environments

Apply at the source

This role was published by Sigma Software and listed via Djinni. Applications are handled there, not on this site.

View & apply on djinni.co ↗