Senior Data Engineer
- Location
- Remote
- Employment
- Full-time
- Level
- Specialist
- Category
- Data & Analytics
- Posted
Description
Join Sigma Software to build large-scale data infrastructure powering a real-time AdTech platform processing hundreds of millions of auction requests daily. We are looking for a Senior Data Engineer who enjoys solving complex distributed data challenges and building production-grade ML-oriented data systems.
You will become part of a dedicated Sigma Software team developing predictive modeling and optimization capabilities for a live advertising ecosystem. The role combines large-scale event processing, streaming and batch pipelines, experimentation infrastructure, and high-throughput data engineering in a cloud-native environment.
We as a company offer the opportunity to work on impactful global products, collaborate with experienced engineers, and contribute to architecture decisions while growing your expertise in large-scale distributed systems and modern data platforms.
Customer
Our Customer is a technology company operating supply-side infrastructure within the programmatic advertising ecosystem. The company manages a large-scale ad exchange handling hundreds of millions of auction requests per day and is actively investing in predictive decisioning technologies to optimize advertising outcomes in real time.
Project
The project focuses on building a predictive modeling and optimization platform on top of a live ad exchange environment. The platform performs real-time supply scoring and filtering, contextual performance estimation, look-alike audience generation, and multi-objective optimization under business constraints.
The solution processes massive-scale event and auction datasets and includes feature engineering pipelines, streaming and batch ingestion, experimentation infrastructure, point-in-time-correct training data generation, and ML-oriented data services with strict operational reliability and compliance requirements.
Requirements
5+ years of experience in Data Engineering
At least 2 years of experience working with production ML or large-scale analytics pipelines
Expert-level SQL skills including window functions and incremental processing patterns
Strong Python skills for production-grade pipeline development
Hands-on experience with Spark or PySpark
Experience designing ETL / ELT pipelines with Airflow, Cloud Composer, Dagster, or similar tools
Experience working with cloud data warehouses at scale, preferably BigQuery
Strong understanding of data modeling and point-in-time correctness
Experience working with event-driven or clickstream datasets at very large scale
Experience supporting business-critical production pipelines
Upper-Intermediate English level or higher
WILL BE A PLUS
Experience with GCP services including Dataflow, Pub/Sub, GCS, and Beam
Experience building streaming or near-real-time ingestion systems
Understanding of feature stores, train/serve skew, and label leakage prevention
Experience in AdTech or auction-based environments
Experience handling delayed or incomplete labels in ML systems
Experience with dbt or similar transformation frameworks
Experience delivering solutions into Customer-owned infrastructure
Knowledge of GDPR/CCPA-related privacy engineering practices
Experience with experimentation infrastructure and statistical validation pipelines
Experience working in hybrid cloud/on-prem Linux environments
Terraform and Kubernetes experience
Experience optimizing warehouse cost and performance
Personal Profile
Strong analytical and problem-solving skills
Ownership-oriented mindset
Ability to work independently in a client-facing environment
Strong communication and documentation skills
Comfortable working in a fast-paced engineering environment
Collaborative and proactive attitude
Responsibilities
Write and defend diagnostic SQL queries against large-scale production datasets
Build and maintain ingestion pipelines for bid, win, and impression logs into BigQuery
Harmonize fields across independently designed datasets and maintain versioned field mappings
Develop point-in-time-correct feature tables and aggregation pipelines
Design and maintain conversion and labeling pipelines with delayed label handling
Own the data serving write path, schema contracts, publishing flows, and freshness SLOs
Build experimentation infrastructure including traffic splitting and reporting pipelines
Perform large-scale historical backfills and safe reprocessing after mapping changes
Implement data isolation and safe-aggregation controls for advertiser data protection
Develop automated data quality validation frameworks
Collaborate closely with Customer engineers and prepare operational documentation
Contribute to architecture discussions and platform scalability improvements
Apply at the source
This role was published by Sigma Software and listed via Djinni. Applications are handled there, not on this site.
Original posting: https://djinni.co/jobs/848309-senior-data-engineer/