Senior Data Engineer (Content Analytics)

N-iX

Location
Remote
Employment
Full-time
Level
Specialist
Category
Data & Analytics
Posted
Deadline

Description

N-iX is looking for a Senior Data Engineer to support our client's Data Licensing team, which develops training datasets for frontier AI labs. You will own the onboarding and enrichment of new multimedia content. Extend scalable event-driven and batch-processing systems that produce reliable metadata for curation and customer-facing applications.

Responsibilities:

  • Partner with the Principal Engineer, senior technical leaders and platform teams to implement and evolve the architectural direction for metadata-processing systems; independently own the design and delivery of assigned components.
  • Evaluate and evolve orchestration for enrichment and backfill workloads, either by extending existing ingestion capabilities or integrating with new cloud-native services, while maintaining consistent contracts and lifecycle handling.
  • Evolve shared frameworks so data scientists can package, test, validate and run production enrichment workloads with minimal data engineering support.
  • Define and maintain metadata contracts, lineage, validation and quality controls.
  • Contribute to and apply shared engineering standards, and provide technical guidance through design reviews, documentation and collaboration with data engineering and data science teams.
  • Estimate and refine the cost, capacity and completion time of enrichment and backfill workloads.
  • Partner with platform teams responsible for operating the underlying cloud infrastructure.
  • Own time-sensitive customer and partner deliverables through production release.

Requirements:

  • 8+ years of experience in data engineering, backend software engineering or distributed systems, including hands-on ownership of large-scale production data workloads.
  • Demonstrated experience designing and operating large-scale ingestion or distributed data-processing systems.
  • Strong Python software engineering and distributed processing experience.
  • Practical knowledge of high-volume batch and event-driven architectures, including idempotency, replay, backpressure, retries and failure recovery.
  • Experience enabling data science or ML workloads to move from prototype to production through reusable frameworks and engineering guardrails.
  • Production experience with AWS and GCP, with the ability to design workflows spanning multiple cloud environments.
  • Experience managing evolving data contracts, schemas, lineage and data-quality expectations.
  • Ability to benchmark distributed workloads and translate results into credible capacity, execution-time and cloud-cost projections.
  • Experience setting engineering standards, reviewing system designs and mentoring engineers or data scientists.
  • Track record of delivering critical, time-sensitive systems into production.

Nice to have:

  • Containerized CPU/GPU batch and ML workloads using managed compute services such as SageMaker, AWS Batch, ECS, Cloud Batch or GKE/EKS.
  • Event-driven orchestration and messaging using services such as Step Functions, Lambda, EventBridge, Kinesis, Workflows or Pub/Sub.
  • Snowflake, ClickHouse or Apache Iceberg.
  • Production-scale image, video or audio processing

Apply at the source

This role was published by N-iX and listed via Djinni. Applications are handled there, not on this site.

View & apply on djinni.co ↗