Empresa: Mars
Job Description: At Mars Data Platform Services (DPS) , we are transforming our enterprise data platform across Azure Databricks and Unity Catalog into an automated, self-service product. Instead of hand-building custom pipelines for every new dataset, we are engineering an intelligent, metadata-driven ingestion engine that configures and scales pipelines automatically. We are looking for a DataOps Engineer to lead the technical evolution of our data inflow and landing zone ecosystem. In this role, you will build the central framework that turns pipeline creation, data onboarding, storage provisioning, and data archiving into automated, configuration-driven processes. What Will Be Your Key Responsibilities? 1. Build & Scale the Metadata-Driven Ingestion Engine Design and enhance our central metadata framework so onboarding new data sources requires simple configuration updates rather than writing new pipeline code. Ensure the engine dynamically orchestrates data ingestion, schema evolution, and merge logic reliably across the platform. 2. Lead Landing Zone Reliability & Data Quality Safeguards Design and maintain high-throughput landing zones across cloud storage and Unity Catalog. Build automated safeguards that isolate schema errors or corrupted data into quarantine areas and dead-letter queues before they can impact production data tables. 3. Automate Unity Catalog Assets & Long-Term Archiving Automate the programmatic creation and setup of storage locations, schemas, and tables in Unity Catalog. Implement automated lifecycle policies that handle time-based data retention and migrate older data into low-cost archival storage seamlessly. 4. Champion DataOps Standards & Operational Enablement Drive modern software engineering standards into our data workflows through automated CI/CD testing and deployment. Create reusable deployment templates, monitoring dashboards, and clear operational playbooks that empower our 24/7 support teams to monitor and maintain pipelines with confidence. What Are We Looking For? Essential Requirements: • Core Data Engineering Background: 4+ years of hands-on data engineering experience building reliable batch and streaming data pipelines using Python/PySpark, SQL , and cloud data warehouses or lakehouses (Azure Databricks / Delta Lake). • Metadata-Driven Mindset: Proven track record designing or maintaining config-driven/metadata-driven frameworks where pipeline generation, schema creation, Delta merge logic, and scheduling are driven dynamically by central control tables rather than hand-coded per dataset. • Programmatic Unity Catalog & Landing Zone Architecture: Practical expertise managing and provisioning Databricks Unity Catalog storage assets (Volumes, External Locations, Managed/External Tables) programmatically via code/APIs, alongside high-throughput Azure Data Lake Storage (ADLS Gen2) landing zones. • DataOps & Automation Practices: Experience treating data pipelines like software products—applying automated CI/CD deployment, version control (Git), automated testing, and modular packaging (such as Databricks Asset Bundles or Terraform). Nice-to-Haves : • Experience with declarative pipeline tools (such as Delta Live Tables or dbt). • Practical experience designing automated data retention schedules and long-term cloud archival storage. • Familiarity with supporting or partnering with operational support teams and Managed Service Providers (MSPs #TBdigital