Careers

Lead Data Engineer – Architecture & Accelerator R&D [Remote] (LDE [29926])

Key Responsibilities:

1. Data Engineering Accelerator R&D

* Understand existing accelerators, validate their capabilities and outputs, recommend improvements, and deliver client demos.
* Work with the founder to prioritize new accelerator capabilities across the data engineering lifecycle.
* Translate data engineering challenges into clear requirements, workflows, and acceptance criteria for AI engineers.
* Provide domain guidance on architecture, database development, data modeling, pipelines, and modernization.
* Define accelerator inputs, expected outputs, engineering rules, and validation checks.
* Create realistic test scenarios and review generated designs, models, and code for correctness, completeness, and performance.
* Work hands-on with AI engineers to resolve issues and improve output quality.
* Capture client feedback and document reusable patterns, standards, and lessons learned.

2. Technical Leadership for Client Data Programs

* Support client proposals with solution approaches, scope, effort estimates, and delivery plans.
* Guide engineering teams and review architectures, designs, and code.
* Attend client calls to clarify requirements, explain recommendations, and resolve technical concerns.
* Contribute hands-on to advisory, solution design, development, and troubleshooting as needed.
* Manage technical risks and dependencies and help teams resolve delivery blockers.
* Apply relevant accelerators to client programs and use delivery feedback to improve them.

Required Skills and Experience

* Strong database fundamentals across OLTP, OLAP, NoSQL, data warehouses, lakehouses, and analytical systems.
* Proven hands-on database development experience, including advanced SQL, stored procedures, indexing, and performance tuning.
* End-to-end delivery experience in both greenfield data platforms and brownfield modernization, from requirements and architecture through deployment and operations.
* Strong Python skills for data processing, integration, and automation.
* Ability to translate business requirements into practical designs and explain architecture trade-offs.
* Strong technical leadership, client communication, problem-solving, and ownership.
* Ability to guide AI engineers and critically validate AI-generated technical outputs.

Required Technical Coverage – Full Data Engineering Stack

Hands-on experience across a complete data engineering stack is mandatory, from source integration and storage through transformation, modeling, reporting, and operations. Candidates must have delivered solutions on at least one modern data platform. Experience with every listed tool is not required; equivalent technologies are acceptable.

* Data platforms: Microsoft Fabric, Databricks, Snowflake, Amazon Redshift, or Google BigQuery.
* Databases: SQL Server, Oracle, PostgreSQL, MySQL, or equivalent, with an understanding of NoSQL use cases.
* Pipelines and orchestration: Azure Data Factory (ADF), Fabric Data Factory, Airflow, or equivalent.
* Data processing: SQL, Python, and Spark/PySpark or equivalent distributed processing tools.
* ETL/ELT design: Source-to-target mapping, transformations, incremental loading, CDC, error handling, and reconciliation.
* Data modeling: Conceptual, logical, physical, and dimensional models, including star schemas and slowly changing dimensions.
* Reporting and analytics: Power BI or equivalent, semantic models, business metrics, and analytics-ready datasets.
* Quality and governance: Profiling, validation, metadata, lineage, access controls, and sensitive-data handling.
* Engineering and operations: Git, CI/CD, automated testing, monitoring, and performance and cost optimization.

Apply for this Job