As a Principal Core Infrastructure Engineer (AI/ML Forward Deployed Infrastructure Engineer) at Oracle, you will be responsible for designing, implementing, and maintaining the infrastructure that supports customer AI and machine learning initiatives. This role involves working closely with customer technical teams and internal cloud services to ensure efficient, secure, and scalable AI/ML solution deployments. You will troubleshoot issues for proof of concept and production deployments, optimize infrastructure for performance, reliability, and cost-effectiveness, and ensure security and compliance standards are met. The position requires strong technical skills in areas like scripting, automation, containerization, Linux administration, and networking, along with excellent problem-solving and collaboration abilities. Experience with AI/ML or HPC workloads, programming languages, and DevOps practices are preferred.
Medical, dental, and vision insurance, short term and long term disability, life insurance, supplemental life insurance, Health care and dependent care Flexible Spending Accounts, Pre-tax commuter and parking benefits, 401(k) Savings and Investment Plan with company match, Flexible Vacation, 11 paid holidays, Paid sick leave, Paid parental leave, Adoption assistance, Employee Stock Purchase Plan, Financial planning and group legal, Voluntary benefits including auto, homeowner and pet insurance.
Requirements
The ideal candidate will have experience in designing, implementing, and managing infrastructure for AI/ML or HPC workloads, including expertise in scripting and automation (Ansible, Terraform, Python), containerization (Docker, Kubernetes), and Linux system administration. Strong problem-solving, communication, and collaboration skills are essential. Familiarity with machine learning frameworks and DevOps practices is preferred.
Description
As a Principal Core Infrastructure Engineer (AI/ML Forward Deployed Infrastructure Engineer), you will play a critical role in designing, implementing, and maintaining the infrastructure that supports our customers AI and machine learning initiatives.
You will work closely with customer technical teams (data scientists, software engineers, and Infra/IT professionals) and internal cloud services team to ensure their AI/ML solutions are deployed efficiently, securely, and at scale.
Your expertise will be crucial in optimizing our infrastructure for performance, reliability, and cost-effectiveness. In this hands-on technical role you’ll troubleshoot issues for proof of concept (POC) and production deployments.