Data Architect

🏢 inception42
📍 Abu Dhabi, United Arab EmiratesFull-timeOn-site
📅 Posted: 2d ago🔄 Updated: 2d ago
CV%
✨ AI Summary
Inception42 is looking for a Data Architect in Abu Dhabi to design and build the data foundations for AI, agentic, and analytics products. This hands-on role involves end-to-end data architecture across ingestion, transformation, storage, modelling, retrieval, APIs, governance, and production operations for both structured and unstructured data. You will collaborate with various teams to ensure scalability, performance, security, and maintainability. Key responsibilities include designing data architectures, defining domain boundaries and data contracts, architecting lakehouse and distributed-data platforms, building production-grade pipelines in SQL and Python, and designing secure data services and APIs. You will also focus on building retrieval and knowledge pipelines, implementing access controls, establishing data quality controls, and driving production readiness through CI/CD and automation. The ideal candidate possesses strong data architecture and engineering fundamentals, fluency in SQL and Python, and experience with modern data architecture patterns and AI product data systems.
Required Skills
Other
lakehousedistributed-compute
Information Technology
Data LineageData PlatformsCI/CDAI Integration
Engineering, Construction & Trades
Automation
Requirements
We are seeking a Data Architect with strong data architecture and engineering fundamentals, including hands-on experience with SQL and Python for building production-grade ETL/ELT pipelines. The ideal candidate will have practical depth in data modelling, experience designing cloud-scale data platforms, and a working knowledge of retrieval-augmented generation and vector search. A strong understanding of API and integration patterns, security, governance, observability, and CI/CD is also required.
Description
Data ArchitectLocation: Abu Dhabi, UAEInception42, a G42 company, is the region's leading innovator of AI-powered domain-specific as well as industry-agnostic products, built on a rich heritage of research and development. Within the G42 ecosystem, Inception42 functions as the core intelligence layer – transforming data and compute infrastructure into real-world, applied AI solutions. Beyond its commercial endeavors, Inception42 is committed to creating positive societal impact. For more information, please visit www.inceptionai.aiOverviewWe are looking for a Data Architect to design and build the data foundations behind AI, agentic, and analytics products. This is a hands-on architecture role spanning ingestion, transformation, storage, modelling, retrieval, APIs, governance, and production operations across structured and unstructured data. You will work closely with data engineering, AI, platform, security, and product teams in an environment where technical decisions must hold up under real-world requirements for scale, performance, security, and maintainability.What You'll OwnDesign and implement end-to-end data architectures across ingestion, processing, storage, modelling, serving, and consumption layers.Define domain boundaries, canonical models, data contracts, integration patterns, and schema-evolution strategies for structured and unstructured data.Architect scalable lakehouse and distributed-data platforms using open storage formats, resilient processing patterns, and clear separation of compute, storage, and serving concerns.Build reference implementations and production-grade ETL/ELT pipelines in SQL and Python, setting a practical standard for testability, performance, and operational readiness.Design secure data services and APIs using contract-first patterns across REST, GraphQL, event-driven integration, and API-management layers.Build retrieval and knowledge pipelines covering parsing, chunking, enrichment, embedding, indexing, reranking, grounding, and citation, with lifecycle controls for refresh, deletion, versioning, evaluation, retention, and agent memory.Embed RBAC and ABAC controls, identity integration, encryption, auditability, retention, and policy enforcement across data products, APIs, and knowledge resources.Establish data-quality controls, validation gates, lineage, metadata capture, cataloguing, pipeline service levels, failure handling, and observability.Drive production readiness through CI/CD, infrastructure as code, containerisation, automated testing, release controls, and disciplined environment management.Collaborate closely with AI, product, engineering, and platform teams to turn use cases into scalable designs, resolve trade-offs, guide delivery through production, and produce clear architecture decisions, data-flow diagrams, reusable patterns, and technical documentation.What We're Looking ForStrong data-architecture and engineering fundamentals, with a track record of taking systems from design through implementation, testing, and production operation.Hands-on fluency in SQL and Python and experience building reliable, production-grade ETL/ELT pipelines and reusable data components.Practical depth in data modelling, including dimensional, relational, and domain-oriented approaches, data contracts, and safe schema evolution.Experience designing cloud-scale data platforms using lakehouse, distributed-compute, streaming, or comparable modern data architecture patterns.Working knowledge of retrieval-augmented generation, vector search, indexing, evaluation, and knowledge-system design for AI products.Strong understanding of API and integration patterns, identity and access management, secure data access, and enterprise governance controls.Experience with observability, data quality, lineage, CI/CD, containers, infrastructure automation, and the operational concerns of production data systems.Strong systems thinking, debugging, and communication skills, with the ability to work through ambiguity, make pragmatic trade-offs, and align engineering, AI, product, platform, and security stakeholders around executable designs.Nice to HaveExperience with Databricks, Delta Lake, Unity Catalog, Microsoft Fabric, Azure data services, or equivalent enterprise data platforms.Exposure to Kubernetes, Azure DevOps, API Management, Entra ID, Keycloak, Terraform, Bicep, or comparable cloud-native platform capabilities.Experience designing agent data systems, Model Context Protocol integrations, knowledge graphs, semantic layers, or long-term memory patterns.Familiarity with Power BI semantic models, governed analytics consumption, and the downstream implications of architecture decisions for reporting and self-service data products.Inception42 is building AI systems designed for practical deployment at scale — across industries, infrastructure, and national-level initiatives.If you want to work close to both advanced AI capability and real-world execution, we'd like to hear from you.
✨ Premium Match Details
Deep-dive CV analysis, customized Cover Letters, and Interview prep!
📊 Match Analysis
Insights against your active CV
📊
Personalized Match Analysis
Upload your CV to see exact matching percentages, detailed skills mapping, and gap analysis for this role.
🎯 Overalli74%
⚡ Skillsi85%
View Breakdown
Ontology Match: 85.0
Matched:✓ Requirements Matching✓ Ontology Skills Mapping
📜 Eligibilityi49%
View Breakdown
Local: 19600%
🏗️ Career Fiti91%
View Breakdown
Seniority: 91.0
📋 Requirementsi67%
View Breakdown
Domain: 67.0
🔥 Motivationi78%
View Breakdown
Title Fit: 78.00