Principal Core Infrastructure Engineer

🏢 Oracle
📍 Nashville, United StatesOn-site
📅 Posted: 1w ago🔄 Updated: 1w ago
CV%
✨ AI Summary
As a Principal Core Infrastructure Engineer (AI/ML Forward Deployed Infrastructure Engineer), you will design, implement, and maintain the infrastructure for customer AI and machine learning initiatives. This role involves working with customer technical teams and internal cloud services to ensure efficient, secure, and scalable AI/ML solutions. You will be responsible for optimizing infrastructure for performance, reliability, and cost-effectiveness, and troubleshooting issues for proof of concept and production deployments. The ideal candidate will have experience in scripting, automation, containerization, orchestration, networking, security, and Linux system administration, with a strong understanding of AI/ML infrastructure and HPC workloads.
Required Skills
Other
AnsibleSlurmPBSRHELCentOSUbuntuDebiantechnical oversightinclusivity
Information Technology
TerraformPythonKubernetesDockerCore NetworkingCloud SecurityLinuxSystems AdministrationShell ScriptingPerformance OptimizationTechnical Documentation
Business, Sales & Management
HR ManagementMentoringProcess Improvement
Soft Skills & Professional Competencies
Problem SolvingCommunicationCollaborationStrategic ThinkingPlanningExecutionDelegationPrioritizationContinuous Learning
Nice to have:
Information Technology
RustGoJavaScalabilityTensorFlowPyTorchScikit-learnJenkinsCI/CDPrometheus
🎁 Benefits & Perks
Medical, dental, and vision insurance, including expert medical opinion; Short term disability and long term disability; Life insurance and AD&D; Supplemental life insurance (Employee/Spouse/Child); Health care and dependent care Flexible Spending Accounts; Pre-tax commuter and parking benefits; 401(k) Savings and Investment Plan with company match; Paid time off: Flexible Vacation is provided to all eligible employees assigned to a salaried (non-overtime eligible) position. Accrued Vacation is provided to all other employees eligible for vacation benefits. For employees working at least 35 hours per week, the vacation accrual rate is 13 days annually for the first three years of employment and 18 days annually for subsequent years of employment. Vacation accrual is prorated for employees working between 20 and 34 hours per week. Employees working fewer than 20 hours per week are not eligible for vacation.; 11 paid holidays; Paid sick leave: 72 hours of paid sick leave upon date of hire. Refreshes each calendar year. Unused balance will carry over each year up to a maximum cap of 112 hours.; Paid parental leave; Adoption assistance; Employee Stock Purchase Plan; Financial planning and group legal; Voluntary benefits including auto, homeowner and pet insurance.
Requirements
Experience in scripting and automation using tools like Ansible, Terraform, Python and/or Kubernetes. Experience with containerization technologies (e.g., Docker, Kubernetes) and orchestration tools (like Slurm, PBS, etc.) for managing distributed systems. Solid understanding of networking concepts, security principles, and best practices. Excellent problem-solving skills, with the ability to troubleshoot complex issues and drive resolution in a fast-paced environment. Strong communication and collaboration skills, with the ability to work effectively in cross-functional teams and convey technical concepts to non-technical stakeholders. Strong documentation skills. Strong Linux skills with hands-on experience in Oracle Linux/RHEL/CentOS, Ubuntu, and Debian distributions, including system administration, package management, shell scripting, and performance optimization.
Description

As a Principal Core Infrastructure Engineer (AI/ML Forward Deployed Infrastructure Engineer), you will play a critical role in designing, implementing, and maintaining the infrastructure that supports our customers AI and machine learning initiatives. 

You will work closely with customer technical teams (data scientists, software engineers, and Infra/IT professionals) and internal cloud services team to ensure their AI/ML solutions are deployed efficiently, securely, and at scale. 

Your expertise will be crucial in optimizing our infrastructure for performance, reliability, and cost-effectiveness. In this hands-on technical role you’ll troubleshoot issues for proof of concept (POC) and production deployments. 

 

 

#LI-ES2

✨ Premium Match Details
Deep-dive CV analysis, customized Cover Letters, and Interview prep!
📊 Match Analysis
Insights against your active CV
📊
Personalized Match Analysis
Upload your CV to see exact matching percentages, detailed skills mapping, and gap analysis for this role.
🎯 Overalli74%
⚡ Skillsi85%
View Breakdown
Ontology Match: 85.0
Matched:✓ Requirements Matching✓ Ontology Skills Mapping
📜 Eligibilityi49%
View Breakdown
Local: 19600%
🏗️ Career Fiti91%
View Breakdown
Seniority: 91.0
📋 Requirementsi67%
View Breakdown
Domain: 67.0
🔥 Motivationi78%
View Breakdown
Title Fit: 78.00