✨ AI Summary
Oracle Cloud Infrastructure (OCI) is seeking an Architect for AI Infrastructure to play a pivotal role in ensuring their AI infrastructure meets the demands of Enterprise and AI/ML customers. This role involves designing and developing architectural changes for GPU delivery, health monitoring, triage automation, and diagnostic services. The ideal candidate will ensure reliability and customer satisfaction through proactive issue management, recurring pattern resolution, and hands-on debugging and log analysis. Collaboration with internal teams and leaders is crucial for maintaining uptime, performance, and customer growth, as well as implementing monitoring and optimization frameworks for AI workloads, including latency, throughput, and cost efficiency. The role also includes leading incident response and post-mortem processes for infrastructure issues impacting AI services.
🎁 Benefits & Perks
Medical, dental, and vision insurance, including expert medical opinion; Short term disability and long term disability; Life insurance and AD&D; Supplemental life insurance (Employee/Spouse/Child); Health care and dependent care Flexible Spending Accounts; Pre-tax commuter and parking benefits; 401(k) Savings and Investment Plan with company match; Paid time off (Flexible Vacation/Accrued Vacation); 11 paid holidays; Paid sick leave; Paid parental leave; Adoption assistance; Employee Stock Purchase Plan; Financial planning and group legal; Voluntary benefits including auto, homeowner and pet insurance.
Requirements
The role requires designing and developing fundamental architectural changes for GPU delivery, health monitoring, triage automation, and diagnostic services for AI infrastructure. Responsibilities include ensuring reliability, proactive issue management, driving efficiency, and hands-on debugging and log analysis for AI/ML/HPC workloads.
Description
Here at OCI we’re building the world’s largest AI clusters and we’re the fastest at bringing them to market. OCI (Oracle Cloud Infrastructure) AI Infrastructure is at the forefront of building a cutting-edge, ultra-high-performance GPU platform designed to support AI/ML/HPC workloads. This is your chance to be part of the AI revolution, working with systems that allow customers to scale from tens to thousands of GPUs without compromising performance.
Our team is responsible for designing and developing fundamental architectural changes for GPU delivery, health monitoring, triage automation, and diagnostic services. You will have the opportunity to work with cutting-edge technologies and make a significant impact on our organization's success.