We're partnering with a large US financial services company to source a Principal DevOps Engineer to join them on a contract basis.
We are seeking a highly motivated and experienced Principal DevOps Engineer to lead the evolution of cloud, platform engineering, and operational excellence capabilities.
In this role, you will provide technical leadership in the design, implementation, and adoption of modern DevOps, Site Reliability Engineering (SRE), GitOps, and AI-powered engineering practices.
You will help shape the future of software delivery by driving intelligent automation, resilience engineering, observability, and platform standardization across our engineering ecosystem.
You will partner closely with development, architecture, infrastructure, and security teams to build highly scalable, resilient, and observable platforms while improving developer productivity through self-service capabilities, Infrastructure as Code (IaC), and AI-driven operational excellence.
Reliability is engineered into everything through automation, observability, resilience testing, chaos engineering, and continuous improvement.
The Expertise You Have
- Bachelor's degree or higher in Computer Science, Engineering, or a related technical field; master’s degree preferred.
- 9+ years of hands-on experience designing, implementing, and supporting enterprise-scale CI/CD platforms and distributed systems.
- 3+ years of experience building and operating cloud-native solutions in AWS, including infrastructure modernization and migration initiatives.
- 3+ years of software development experience utilizing Python, Java, or Node.js, with a strong focus on automation and software engineering best practices.
- 2+ years of hands-on GitOps experience leveraging platforms such as Argo CD, Flux CD, or equivalent solutions.
- Strong experience with Kubernetes and containerized application platforms.
- Experience designing and operating highly available, distributed, service-oriented architectures.
- Experience building and operating Internal Developer Platforms (IDPs) and enabling self-service engineering capabilities.
- Experience utilizing AI-assisted engineering tools such as GitHub Copilot, Amazon Q, ChatGPT Enterprise, Cursor, or equivalent technologies.
- Exposure to AIOps and Agentic AI frameworks supporting operational intelligence, automation, incident response, and workflow orchestration.
The Skills You Bring
- Strong automation skills using Python, Shell scripting, and related technologies.
- Deep expertise in CI/CD practices, software configuration management, automated testing, code quality analysis, and deployment automation utilizing tools such as Jenkins, Ansible, and Docker.
- Hands-on experience with Infrastructure as Code frameworks including Terraform, AWS CloudFormation, and Ansible.
- Strong understanding of cloud-native architectures, DevOps principles, SRE practices, and platform engineering.
- Advanced Kubernetes administration and application delivery experience.
- Experience with GitOps methodologies, including declarative infrastructure, automated reconciliation, progressive delivery, and version-controlled operations.
- Hands-on experience with Argo CD, Flux CD, Argo Rollouts, or similar GitOps technologies.
- Hands-on experience with modern observability platforms including Datadog, Prometheus, Grafana, ELK/OpenSearch, Splunk, and OpenTelemetry.
- Strong expertise in monitoring, logging, alerting, instrumentation, and performance engineering for large-scale distributed systems.
- Demonstrated ability to improve platform scalability, reliability, performance, and resiliency.
- Strong analytical and problem-solving skills, including incident management, root cause analysis, and troubleshooting under pressure.
- Excellent verbal and written communication skills with the ability to influence both technical and executive stakeholders.
- Proven ability to collaborate effectively across diverse engineering teams and organizational boundaries.
- Demonstrated expertise supporting mission-critical production environments with a focus on availability, reliability, and operational excellence.
- Participation in an on-call rotation may be required.
Other preferred skills – Nice to have:
- Experience applying AI and AIOps capabilities across software delivery, automation, monitoring, and operational workflows.
- Familiarity with Agentic AI concepts, autonomous workflows, and AI-driven operational automation.
- Understanding AI governance, security, compliance, and responsible AI practices.
- Experience leveraging data analytics and visualization tools to drive operational insights.
- Passion for continuous learning and adopting emerging technologies and best practices.
The Value You Deliver
- Lead the adoption of GitOps and Platform Engineering practices to improve deployment reliability, scalability, governance, and developer experience.
- Drive AI-enabled DevOps initiatives that reduce operational toil and accelerate engineering productivity through intelligent automation.
- Identify opportunities to leverage Agentic AI capabilities for incident management, deployment validation, operational workflows, and production support within established governance frameworks.
- Partner with engineering, architecture, infrastructure, and security teams to establish next-generation DevOps capabilities incorporating GitOps, AIOps, and AI-driven operations.
- Accelerate software delivery while improving reliability through predictive insights, intelligent automation, and autonomous operational capabilities.
- Support development squads with modern DevOps practices, cloud adoption, and engineering excellence initiatives.
- Deliver high-quality, maintainable, secure, and cost-effective solutions that meet both functional and non-functional business requirements.
- Define and execute enterprise-grade reliability and observability strategies that ensure critical systems remain highly available and performant.
- Leverage technical, operational, and financial data to drive efficiency, reduce complexity, and eliminate manual effort.
- Establish and promote engineering standards, platform consistency, and operational excellence across development and SRE organizations.
- Lead troubleshooting and resolution of complex, cross-functional issues spanning applications, infrastructure, networks, and cloud platforms.
- Mentor and coach engineers and technical leaders in reliability engineering, observability, automation, and cloud-native best practices.
Click 'Apply Now' to submit your resume and start your application. Alternatively, you can send your resume directly to [email protected]
