Forward Deployed Engineer - LLMOps

🏢 Systems Limited
📍 Saudi ArabiaFull-timeOn-site
📅 Posted: Yesterday🔄 Updated: Yesterday
CV%
✨ AI Summary
The Forward Deployed Engineer - LLMOps role at Systems Limited focuses on owning the production operations for LLM and agentic workloads, including serving, cost management, and observability. Key responsibilities involve managing production serving and scaling for LLM/agentic workloads, controlling inference costs, building observability for LLM-specific failure modes, and managing model/version rollout strategies. The role also includes incident response for LLM/agent production issues and partnering with GenAI Engineers on production-readiness. This position requires 4-6 years of platform/MLOps engineering experience with hands-on LLM/GenAI production experience. A deep understanding of LLM inference economics, experience with LLM observability tooling, and familiarity with various model hosting platforms (Microsoft Azure AI Foundry, AWS Bedrock, Google Vertex AI, and self-hosted options like vLLM, TGI) are essential. The candidate must be comfortable with the unpredictability of agentic workloads and possess strong communication skills to explain cost dynamics to stakeholders.
Required Skills
Other
AWS BedrocktracingLLM inference economicsTGIGoogle Vertex AILLM GenAI production experiencevLLMeval pipelines
Information Technology
AzureObservability
Business, Sales & Management
HR Management
Requirements
Requires 4-6 years of platform/MLOps engineering experience with hands-on LLM/GenAI production experience. Must have a deep understanding of LLM inference economics, experience with LLM observability tooling, and familiarity with multiple model hosting platforms (Microsoft Azure AI Foundry, AWS Bedrock, Google Vertex AI, and self-hosted options like vLLM, TGI). Experience building canary/rollback strategies for probabilistic systems and comfort with agentic workloads are also essential.
Description
ABOUTOwns production operations for LLM and agentic workloads — serving, cost, and observability for a fundamentally less predictable class of system than classical ML.KEY RESPONSIBILITIESOwn production serving and scaling for LLM/agentic workloads (inference infra, load balancing, caching)Monitor and control inference cost — token usage, retry/loop cost, model routing decisionsBuild observability for LLM-specific failure modes: hallucination rate, latency spikes, prompt driftManage model/version rollout strategy (canary releases, fallback models, A/B testing)Own incident response for LLM/agent production issuesPartner with GenAI Engineers and Agentic AI Architects on production-readiness reviewsExplain token-cost dynamics to client finance/business stakeholdersCollaborate closely with GenAI Engineers without needing a hard line between build and runSupport the practice in setting cost governance policy for LLM workloadsREQUIREMENTS & SKILLS4–6 yrs platform/MLOps engineering with hands-on LLM/GenAI production experienceDeep understanding of LLM inference economics — token costs, batching, caching, model routingExperience with LLM observability tooling (tracing, eval pipelines, prompt/version management)Familiarity with multiple model hosting platforms and their cost/performance tradeoffs — Microsoft Azure AI Foundry, AWS Bedrock, and Google Vertex AI, plus self-hosted open-source options (vLLM, TGI) as a good-to-haveExperience building canary/rollback strategies for probabilistic systemsComfortable with the higher unpredictability of agentic workloads vs. classical ML servingCost-conscious communicator — can explain a token-cost blowup to a client's finance stakeholderCollaborates closely with GenAI Engineers without needing a hard line between build and runCalm under pressure during live incidents affecting client-facing systems
✨ Premium Match Details
Deep-dive CV analysis, customized Cover Letters, and Interview prep!
📊 Match Analysis
Insights against your active CV
📊
Personalized Match Analysis
Upload your CV to see exact matching percentages, detailed skills mapping, and gap analysis for this role.
🎯 Overalli74%
⚡ Skillsi85%
View Breakdown
Ontology Match: 85.0
Matched:✓ Requirements Matching✓ Ontology Skills Mapping
📜 Eligibilityi49%
View Breakdown
Local: 19600%
🏗️ Career Fiti91%
View Breakdown
Seniority: 91.0
📋 Requirementsi67%
View Breakdown
Domain: 67.0
🔥 Motivationi78%
View Breakdown
Title Fit: 78.00