NOC Lead - Enterprise Operations

🏢 k20s - kinetic technologies private limited
📍 Sharjah, United Arab EmiratesFull-timeRemote
📅 Posted: Today🔄 Updated: Today
CV%
✨ AI Summary
NOC Lead responsible for 24x7 monitoring and operational support of enterprise systems including websites, applications, networks, servers, and cloud services. This role involves leading shift governance, incident response, major incident coordination, vendor escalations, ITIL process compliance, SLA reporting, and continuous service improvement. Key responsibilities include managing enterprise monitoring tools like Site24x7, leading incident management bridges, supervising NOC engineers, coordinating with vendors, enforcing ITIL processes via BMC Service Management, and generating operational reports. The position requires strong technical troubleshooting, disciplined leadership, and the ability to work rotating shifts in a highly available environment.
Required Skills
Other
Site24x7ManageEngine OpManager
Information Technology
DashboardsMonitoring
Hospitality, Retail & Customer Service
Customer Service Management
Requirements
The role requires a Bachelor's degree in Computer Science, Information Technology, Computer Engineering, Electronics, or a related field, with 5-8 years of experience in enterprise NOC, infrastructure operations, application monitoring, or technical support. Candidates must have at least 2 years of experience in team leadership, shift supervision, or operational coordination within a 24x7 NOC environment. Proficiency in enterprise monitoring platforms like Site24x7 and engine-based products, ITSM tools like BMC Service Management, and core network/server fundamentals are essential. Experience with ITIL practices, vendor coordination, SLA management, and working in highly available environments is also required. The ability to work rotating shifts and lead teams under pressure is crucial.
Description
Role SummaryNOC Lead will be responsible for leading enterprise-wide 24x7 monitoring and operational support across websites, applications, networks, Windows servers, cloud services, and end-user computing platforms. The role owns shift governance, alert response, major incident coordination, vendor escalations, ITIL process compliance, BMC ticket governance, SLA and KPI reporting, and continuous service improvement. The position requires hands-on command of engine-based monitoring products and Site24x7, strong technical troubleshooting capability, and disciplined leadership to maintain service availability, performance, reliability, and operational excellence across the enterprise. Key Responsibilities 2.1 Enterprise Monitoring & Command Centre OperationsLead 24x7 monitoring of enterprise websites, applications, APIs, network devices, Windows servers, cloud services, and other business-critical platforms.Take operational ownership of engine-based products and monitoring tools, including hands-on administration, troubleshooting, health checks, and service restoration support.Operate and optimize Site24x7 monitoring for website and application availability, response time, transaction performance, infrastructure health, and alerting.Ensure monitoring thresholds, probes, synthetic checks, dashboards, notification rules, and escalation paths remain accurate and aligned with business criticality.Review monitoring coverage for new and changed services and ensure there are no unmanaged assets, blind spots, or unsupported alerts.Reduce false positives and alert noise through regular tuning, correlation, suppression, and improvement of monitoring logic.2.2 Incident, Event & Major Incident ManagementEnsure all alerts are validated, prioritized, acknowledged, recorded, and actioned within defined operational targets.Lead the technical and operational response for Priority 1 and Priority 2 incidents, including bridge coordination, task allocation, escalation, and service restoration tracking.Maintain clear communication with business stakeholders, technical teams, service owners, and management throughout critical incidents.Ensure incidents are linked to related problem, change, vendor, and known-error records where applicable.Drive post-incident reviews, root cause analysis, corrective actions, and preventive measures for recurring or high-impact failures.Verify that shift handovers include all active alerts, open incidents, pending vendor actions, planned changes, risks, and follow-up items.2.3 Team Leadership & Shift GovernanceSupervise, mentor, coach, and schedule NOC Engineers to ensure complete and effective coverage across all shifts, including nights, weekends, and public holidays.Prepare shift rosters, manage leave coverage, distribute workloads, and maintain adequate staffing for operational and business requirements.Set clear expectations for punctuality, ownership, ticket quality, communication, escalation discipline, and a solution-oriented can-do attitude.Conduct shift briefings, knowledge-sharing sessions, technical coaching, performance reviews, and competency development activities.Maintain and enforce standard operating procedures, runbooks, escalation matrices, checklists, and shift handover standards.Review team performance and take timely corrective action for process gaps, missed alerts, delayed escalations, or recurring quality issues.2.4 Vendor Coordination & SLA ManagementAct as the primary operational liaison with third-party vendors, managed service providers, telecom providers, application partners, and support contractors.Raise and track vendor cases, provide required evidence and diagnostics, coordinate troubleshooting sessions, and escalate delays or service risks.Monitor vendor response and resolution performance against contractual SLAs, operational level agreements, and service commitments.Conduct regular vendor service reviews and follow up on chronic issues, pending root cause reports, recurring incidents, and improvement actions.Maintain vendor contact details, support entitlements, contract references, escalation paths, and service coverage information.2.5 ITIL Process & BMC Service ManagementImplement and enforce ITIL-based incident, problem, change, event, service request, and knowledge management processes within NOC operations.Use BMC Service Management / BMC Helix / BMC Remedy for ticket creation, categorization, assignment, escalation tracking, work notes, resolution, closure, and reporting.Ensure ticket records contain accurate timestamps, impact and urgency, troubleshooting evidence, actions taken, ownership, customer communication, and closure details.Review aging, breached, reopened, misclassified, and unassigned tickets and drive timely corrective action.Participate in change planning, change advisory discussions, maintenance windows, implementation monitoring, validation, and rollback coordination.Develop and maintain knowledge articles, troubleshooting guides, templates, and known-error documentation.2.6 Reporting, Dashboards & Service ImprovementCompile and share daily, weekly, and monthly reports covering incident trends, alert volumes, service availability, SLA attainment, MTTA, MTTR, ticket aging, recurring issues, and team performance.Maintain operational dashboards for enterprise services, critical infrastructure, website and application health, shift status, and management visibility.Analyze trends and recurring patterns to identify risks, capacity concerns, monitoring gaps, automation opportunities, and service improvement priorities.Present concise operational updates, incident summaries, service review inputs, and executive-level metrics to management and stakeholders.Track corrective and preventive actions to completion and demonstrate measurable service improvements.2.7 Infrastructure & End-User Support CoordinationMaintain working knowledge of TCP/IP, DNS, DHCP, routing and switching fundamentals, VPN connectivity, and common network failure indicators.Coordinate first-line checks for Windows Server services, event logs, CPU, memory, disk, processes, scheduled tasks, and basic operating system issues.Support basic troubleshooting for Windows 10/11 desktops, connectivity, authentication, endpoint services, and standard enterprise applications.Engage network, server, database, cloud, cybersecurity, application, and end-user support teams based on alert type and troubleshooting evidence.Ensure the NOC gathers accurate diagnostics before escalation and validates service restoration after technical teams complete corrective actions.2.8 Operational Governance, Risk & ContinuityMaintain disciplined 24x7 operational control, including shift readiness, attendance, access availability, communication channels, and escalation coverage.Support business continuity, disaster recovery, failover, high-availability, and crisis-management exercises from an operational monitoring perspective.Identify operational risks, single points of failure, repeat service interruptions, and control weaknesses and escalate them for remediation.Ensure compliance with enterprise security policies, data handling requirements, audit controls, and access governance.Provide audit evidence, incident records, monitoring reports, change records, SOPs, and control documentation when requested.2.9 Projects, Automation & Monitoring EnhancementRepresent NOC operations in infrastructure, application, cloud, migration, and technology upgrade projects.Define monitoring, alerting, ticketing, support ownership, escalation, documentation, and operational acceptance requirements before go-live.Promote automation for repetitive checks, ticket enrichment, reporting, service validation, alert correlation, and dashboard generation.Coordinate user acceptance and operational readiness testing for new monitoring capabilities and support procedures.Maintain a continuous improvement roadmap for tools, processes, team capability, service visibility, and response effectiveness. Technical Skills 3.1 Enterprise Monitoring PlatformsSite24x7 - website, application, server, network, cloud, synthetic, and availability monitoring.Engine-based monitoring products and monitoring collectors, agents, probes, dashboards, event engines, and alert workflows.ManageEngine OpManager, and AppDynamics or equivalent enterprise monitoring platforms.Threshold configuration, service dependency mapping, alert correlation, notification policies, escalation rules, and monitoring health checks.3.2 ITSM & ITIL OperationsBMC Service Management / BMC Helix / BMC Remedy ticketing and workflow management.ITIL incident, problem, change, event, service request, knowledge, and continual improvement practices.Major incident management, SLA tracking, escalation governance, ticket quality control, and service review reporting.3.3 Network FundamentalsTCP/IP, DNS, DHCP, NAT, VPN, ports and protocols, subnetting basics, and network connectivity validation.Basic routing and switching concepts, LAN/WAN fundamentals, interface status, latency, packet loss, and bandwidth indicators.Common network troubleshooting utilities such as ping, traceroute, nslookup, ipconfig, netstat, and telnet/test-netconnection.3.4 Server & End-User PlatformsWindows Server operating system fundamentals, services, event logs, performance counters, storage, processes, and scheduled tasks.Windows 10/11 desktop fundamentals, endpoint connectivity, authentication, services, applications, and remote support basics.Basic awareness of Active Directory, virtualization, cloud platforms, backup systems, storage, and endpoint security tools.3.5 Website, Application & Service MonitoringHTTP/HTTPS, DNS resolution, SSL/TLS certificate validity, ports, URLs, APIs, response codes, response time, and content checks.Website and application availability, synthetic transactions, user journey monitoring, dependency awareness, and performance indicators.Basic interpretation of application, web server, operating system, and monitoring logs for first-line diagnosis.3.6 Reporting & AnalyticsOperational dashboards, SLA and KPI reporting, trend analysis, incident analytics, availability reporting, and performance scorecards.Microsoft Excel, PowerPoint, and reporting or visualization tools such as Power BI or equivalent platforms.Data quality validation, management summaries, action tracking, and service improvement measurement.3.7 Automation & ScriptingWorking familiarity with PowerShell, Python, SQL, APIs, or similar tools for operational automation and reporting.Ability to identify repetitive NOC activities suitable for scripts, workflows, templates, or automated validation.Basic understanding of integration between monitoring tools, ticketing systems, dashboards, and notification platforms.3.8 Project & Operational SkillsRoot cause analysis, risk analysis, change management, capacity awareness, documentation, and continual service improvement.Shift planning, workload management, vendor management, stakeholder communication, and operational governance.Ability to work calmly under pressure and coordinate multiple teams during critical incidents.Skills: management,itil,troubleshooting Educational QualificationsBachelor's degree in Computer Science, Information Technology, Computer Engineering, Electronics, or a related field. Preferred CertificationsITIL Foundation certification.CompTIA Network+ or Cisco CCNA certification.BMC Helix / BMC Remedy administration or user certification.Site24x7, ManageEngine, AppDynamics, or equivalent monitoring certification. Experience Requirements5-8 years of experience in enterprise NOC, infrastructure operations, application monitoring, or technical support environments.At least 2 years of team leadership, shift supervision, or operational coordination experience in a 24x7 NOC or support function.Hands-on experience managing website and application monitoring on a continuous 24x7 basis.Practical experience with engine-based monitoring products, Site24x7, and enterprise monitoring dashboards.Demonstrated experience using BMC Service Management or a comparable ITSM ticketing platform.Experience coordinating vendors, managing escalations, and driving SLA compliance.Experience applying ITIL practices in incident, problem, change, and service request management.Experience preparing service reports, operational metrics, trend analysis, and management dashboards.Ability and willingness to work or provide leadership coverage across rotating shifts, nights, weekends, and public holidays.Experience in highly available, mission-critical environments such as banking, healthcare, government, telecom, retail, or large enterprises. Core CompetenciesLeadership, coaching, mentoring, and shift management.Strong troubleshooting, analytical, and decision-making skills.Customer-focused service delivery and ownership mindset.Excellent verbal, written, and incident communication skills.Vendor, stakeholder, and cross-functional team management.Strong documentation, reporting, and presentation practices.Reliability, punctuality, responsibility, and a can-do attitude.Ability to remain calm and decisive during critical incidents.Attention to detail, process discipline, and continuous improvement.Continuous learning and adoption of emerging monitoring and operations technologies.
✨ Premium Match Details
Deep-dive CV analysis, customized Cover Letters, and Interview prep!
📊 Match Analysis
Insights against your active CV
📊
Personalized Match Analysis
Upload your CV to see exact matching percentages, detailed skills mapping, and gap analysis for this role.
🎯 Overalli74%
⚡ Skillsi85%
View Breakdown
Ontology Match: 85.0
Matched:✓ Requirements Matching✓ Ontology Skills Mapping
📜 Eligibilityi49%
View Breakdown
Local: 19600%
🏗️ Career Fiti91%
View Breakdown
Seniority: 91.0
📋 Requirementsi67%
View Breakdown
Domain: 67.0
🔥 Motivationi78%
View Breakdown
Title Fit: 78.00