Site Reliability Engineer (SRE) - Migration Operations (IRC305279)
- Location
- Remote
- Employment
- Full-time
- Level
- Mid-level
- Category
- DevOps & SRE
- Posted
- Deadline
Description
Preferred qualifications & skills
Education & experience: Bachelor of Science in Computer Science, Systems Engineering, or a related field, plus heavy experience in production system administration and incident response.
System Operations: Expertise in Linux system administration, storage mounting, network routing, and process monitoring.
Monitoring & Tooling: Experience using Prometheus, Grafana, and log aggregators to monitor high-throughput data operations.
Automation: Proficiency in Python, Bash, and SaltStack.
Nice to have skills:
Experience managing live data center maintenance windows and fleet migrations
Job Responsibilities
Monitor and maintain host stability, network bandwidth, and storage health during active fleet migrations.
Develop monitoring dashboards and alert triggers to detect migration degradation or hardware failures.
Execute operational runbooks and coordinate technical mitigation during migration maintenance windows.
Automate infrastructure recovery steps for failed VM migration attempts.
Department/Project Description
We are seeking a Senior SRE to join our Fleet Operations Team. The team ensures uninterrupted service reliability and system health during major hardware and software platform maintenance cycles. You will be responsible for operational readiness, performance monitoring, and incident management while migrating workloads to new clusters.
Eager to know more? If you are proactive, bring new ideas and suggestions, don’t waste any second and apply!
Apply at the source
This role was published by GlobalLogic and listed via Djinni. Applications are handled there, not on this site.
Original posting: https://djinni.co/jobs/850822-site-reliability-engineer-sre-migration-opera/