Cloud Platforms builds and runs the internal platform that PEI’s engineering and data teams build on: the cloud accounts, the reusable infrastructure code, the delivery pipelines, the observability, and the security and cost guardrails that wrap around them. We treat that platform as a product with internal customers – the measure of our work is whether those teams can ship quickly and safely without needing us in the room: self-service by default, guardrails rather than gates, and shared components everyone can use instead of bespoke infrastructure rebuilt per project.
As a Cloud Platforms Engineer, you will help develop and build that platform. Ownership of it sits with the team rather than with individuals, so you will contribute across all of it – the components other teams consume, the estate they run on, and the path between someone writing a change and that change being live.
A good share of the work sits around the Databricks data platform running on Azure. Prior Databricks experience is not expected – a genuine appetite for learning an unfamiliar platform matters far more.
Responsibilities and Duties
- Build reusable building blocks. Write and maintain Terraform modules and GitLab CI/CD (continuous integration and delivery) components that other teams consume to stand up services safely, so that the common case needs no bespoke work from us at all.
- Run the cloud estate. Design, implement and operate infrastructure across multiple accounts and environments on Amazon Web Services (AWS), on Microsoft Azure, where our Databricks data platform sits, and on a smaller Google Cloud Platform footprint.
- Make shipping fast and predictable. Improve how PEI’s teams get code into production: pipeline design, deployment safety, environment parity, and shortening the gap between a change being written and it being live.
- Make the platform observable by default. Ship monitoring, alerting and dashboards as part of the platform rather than as an afterthought, and keep tightening the loop between something going wrong and someone useful knowing about it.
- Prefer guardrails to gates. Bake security and compliance into those shared components – least-privilege access, policy as code, automated scanning in the pipeline – so that the secure path is also the easiest path.
- Own identity and access. Manage who can reach what across AWS, Azure, Databricks and GitLab – joiners and leavers, group-based permissions, service principals and machine credentials – and keep moving that towards something automated rather than a queue of requests.
- Keep production healthy. Help diagnose and resolve production issues when they arise, bring a clear view of what good incident management looks like, and fix the class of problem rather than just the instance.
- Make sure we can recover. Look after backup and disaster recovery across the estate, and make sure recovery is something we have tested and proven rather than assumed.
- Make spend visible. Help keep cloud cost attributable to the teams and products that drive it, and act on what that shows.
- Support the teams who depend on us. Be a responsive point of contact for teams across the business – failing pipelines, blocked releases, access requests, alerts that need a human. Then treat anything that comes up twice as a candidate for automation, documentation or self-service, rather than accepting it as a permanent part of the job.