Data Engineer (Apache Spark)

🏢 Aptivamena Tech
📍 Cairo, EgyptFull-timeOn-site
📅 Posted: 3w ago🔄 Updated: 3w ago
CV%
✨ AI Summary
The Data Engineer role focuses on building and operating high-scale batch and near-real-time data pipelines using Apache Spark as the core processing engine. This position involves designing and managing analytical data systems end-to-end on self-managed infrastructure. The ideal candidate will possess extensive experience in data engineering, software development, and deep hands-on production knowledge of Apache Spark, including performance tuning and deployment on platforms like YARN or Kubernetes. Proficiency in related technologies such as Kafka, distributed query engines, ETL tools, SQL, and real-time analytical stores is crucial.
Required Skills
Information Technology
JavaScalabilityPythonApache SparkKafkaInformaticaApachedbtSQLData ModelingData WarehouseHadoopDatabricksDockerKubernetesApache AirflowData Engineering & AnalyticsClaudeSoftware EngineeringPySpark
Other
Spark Structured StreamingTrino/PrestoDatastageClickHouseClouderaDataHubCodex
Requirements
7+ years of experience in data engineering and software developmentAbility to write high-quality code in Java/Scala, Python or equivalent languagesDeep, hands-on production experience with Apache Spark - batch and Spark Structured Streaming (core requirement)Demonstrated Spark performance tuning: partitioning, caching and persistence, broadcast joins, shuffle reduction, data-skew handling and Adaptive Query ExecutionExperience operating Spark on self-managed clusters (YARN, Kubernetes or standalone) - executor sizing, resource allocation and multi-tenant workloadsPractical experience with Kafka (or equivalent messaging systems) as a Spark source and sink for high-volume workloads, including offset and checkpoint managementPractical experience with distributed query engines (e.g. Trino/Presto or similar)Practical experience with ETL / data integration tools (e.g. Datastage, Informatica, Apache NiFi) and SQL-based transformation frameworks (e.g. dbt)Strong SQL skills and understanding of data modeling and data warehousing for analytical workloadsHands-on experience with real-time / low-latency analytical stores (columnar or OLAP engines, e.g. Apache Pinot/ClickHouse or similar)Practical experience with big-data platforms and distributions (e.g. Cloudera, Hadoop ecosystem, Databricks or similar)Practical experience containerizing and operating data workloads (Docker; Kubernetes a plus) and workflow orchestration tools (e.g. Airflow)Familiarity with data lake table formats (e.g. Apache Iceberg, Delta Lake or similar), including schema evolution and compactionFamiliarity with data governance / cataloging tools (e.g. DataHub) and lakehouse management systems (e.g. Apache Amoro)Familiarity using AI tools for development and debugging (Claude, Cursor, Codex)
Description
Build high-scale batch and near-real-time data pipelines deployed on infrastructure we run ourselves (on-prem), not managed cloud services. You will design and operate high-volume analytical data systems end to end, with Apache Spark as the core processing engine for both batch and streaming workloads.
✨ Premium Match Details
Deep-dive CV analysis, customized Cover Letters, and Interview prep!
📊 Match Analysis
Insights against your active CV
📊
Personalized Match Analysis
Upload your CV to see exact matching percentages, detailed skills mapping, and gap analysis for this role.
🎯 Overalli74%
⚡ Skillsi85%
View Breakdown
Ontology Match: 85.0
Matched:✓ Requirements Matching✓ Ontology Skills Mapping
📜 Eligibilityi49%
View Breakdown
Local: 19600%
🏗️ Career Fiti91%
View Breakdown
Seniority: 91.0
📋 Requirementsi67%
View Breakdown
Domain: 67.0
🔥 Motivationi78%
View Breakdown
Title Fit: 78.00