Requirements
We are looking for a DevOps/MLOps Engineer with strong hands-on experience in CI/CD, AWS infrastructure management, and MLOps across the full model lifecycle. Key requirements include expertise in source-control-native CI/CD tooling (e.g., GitHub Actions), strong AWS skills (IAM, cross-account architecture, OIDC federation, ECR, container and serverless compute), and Infrastructure as Code (Terraform or AWS CDK). Experience with container scanning, vulnerability gating, and building containerized deployment pipelines is essential.
Description
DevOps / MLOps Engineer | AWS | CI/CD | AI/ML InfrastructureAbout the roleWe are looking for a strong DevOps / MLOps Engineer to own and evolve the infrastructure, automation and deployment practices supporting AI/ML models and applications.This role sits at the intersection of DevOps, cloud infrastructure and MLOps. The ideal candidate has deep hands-on experience building robust CI/CD systems, managing AWS infrastructure and taking machine learning models from development through controlled production deployment and monitoring.What you'll doAssess, improve and extend existing CI/CD pipelines, increasing reliability, consistency and coverageEnhance automation frameworks and repository templates used to build, validate and deploy packages, applications and ML modelsOwn controlled promotion across development and production environments, including approval gates and rollback strategiesManage cloud services underpinning the AI/ML and application estate, including:Identity and access architectureFederated authentication from source controlArtifact and container registriesCompute and model-serving infrastructureMaintain build infrastructure, including self-hosted runners with access to accelerated computeOwn the ML model lifecycle, including:Model registry and versioningPackaging models for inferencePromotion criteriaProduction deploymentMonitoring of deployed modelsDefine container image build standards and embed image scanning and vulnerability gates into CI/CD pipelinesMaintain and improve Infrastructure as Code, ensuring environment configuration remains consistent and in parityContribute to data management and governance standards supporting AI/ML readinessWhat we're looking forStrong, hands-on DevOps and CI/CD experienceDeep experience with source-control-native CI/CD tooling, particularly GitHub Actions or equivalentExperience designing reusable, modular and templated workflowsStrong AWS experience, including:IAMCross-account architectureOIDC federationECRContainer computeServerless computeStrong Infrastructure as Code experience using Terraform or AWS CDKProven MLOps experience across the full model lifecycle, from model registry through production serving and monitoringExperience building and maintaining containerized deployment pipelinesStrong understanding of software supply-chain security, including container scanning and vulnerability gatingAbility to work independently and take ownership of infrastructure end-to-endNice to haveExperience connecting on-premises AI/ML training infrastructure with cloud-based model servingExperience supporting AI/ML platforms in a regulated or enterprise engineering environmentExperience with cloud cost management and capacity governanceExperience working with accelerated or GPU-based compute infrastructure