Requirements
3+ years of hands-on ML/AI engineering with demonstrated end-to-end system ownershipProduction experience building LLM-powered applications (not just API consumption)Hands-on experience with agent orchestration: LangGraph, CrewAI, or AutoGen in productionProduction RAG experience with evaluation metrics, hybrid search, and re-ranking strategiesExperience building ML models: churn, propensity, LTV, segmentation, recommendation systemsHands-on experience with data pipelines: Spark for batch, Flink or Kafka Streams for real-timeStrong Python proficiency: production code structure, async, multiprocessing, profiling, optimizationExperience with vector databases at scale: OpenSearch k-NN, Qdrant, or MilvusProduction MLOps experience: MLflow, experiment tracking, model registry, drift monitoringReal-time ML inference experience at 1,000+ QPSGood to Have:Experience at AI-first companies or building AI/ML platforms from scratchTelco or enterprise data platform backgroundExperience with LLM fine-tuning: LoRA, QLoRA, PEFT techniquesExperience with embedding models: sentence-transformers, fine-tuning for domainKubernetes for ML workload orchestration and GPU schedulingKnowledge of PII detection (Presidio) and LLM guardrails (NeMo Guardrails)
Description
Own ML/AI systems end-to-end: data pipelines, model training, serving infrastructure, monitoring, and iterationBuild LLM-powered applications with custom pipelines, prompt management, evaluation, and optimizationImplement multi-agent orchestration systems using LangGraph, CrewAI, or AutoGen for autonomous workflowsBuild and optimize RAG pipelines using LlamaIndex with chunking strategies, embedding selection, re-ranking, and evaluationDeploy and manage LLM inference infrastructure using vLLM or Ollama for on-premise sovereign deploymentsBuild traditional ML scoring models: churn prediction, propensity scoring, LTV estimation, next-best-actionDesign and build feature pipelines using Apache Flink (streaming) and Spark (batch) for real-time and batch MLImplement MLOps practices: model versioning, registry, drift monitoring, A/B testing, and staged rolloutsDesign and implement AI operators for visual low-code canvas (LLM Gateway, RAG Pipeline, Intent Classifier)Optimize ML inference for latency and throughput at scale (10K+ QPS)Collaborate with Data Engineering and Platform teams to integrate ML systems with data infrastructure