Oversee the proper implementation of event management and monitoring guidelines across the Technology department
Lead the proactive event management through automation and predictive analytics.
Design, configure, and optimize observability platforms (e.g., Datadog, Dynatrace, Prometheus) for comprehensive visibility across IT and network services
Oversee the Configuration and maintaining of monitoring thresholds, alerting rules, and automated remediation scripts
Collaborate with cross-functional teams to conceptualize and implement monitoring requirements for new services and technologies
Ensure rapid escalation and resolution of incidents, minimizing service downtime and customer impact.
Analyze incident and performance data to identify trends, drive process improvements, and recommend automation and proactiveness opportunities.
Work closely with internal teams and external vendors to ensure comprehensive monitoring coverage and rapid incident response.
Design and Implement machine learning models for anomaly detection, root cause analysis, and predictive incident management.
Develop and maintain dashboards for real-time service and customer impact monitoring
Ensure relevant MMC stakeholders are constantly and timely aware of technology faults
Ensure SLAs related to management of exception conditions in the technology ecosystem are respected by the suppliers
Ensure the routine operational tasks are carried out by Technology Operations Control, which is staffed by shifts of operators who provide centralized monitoring and control activities
Advocate for a culture of proactivity, automation, and data-driven decision-making within the team
Provide regular reports and insights to management on system health, incident trends, and improvement initiatives