Senior Data Engineer
Straive- Posted an hour ago
- Be among the first 10 applicants
Job Description
Job Title: Senior Data Engineer (Azure, Databricks & Microsoft Fabric)
Location: Straive Locations
Experience Level: 5–8 Years
Job Summary
We are seeking a highly skilled Senior Data Engineer to design, build, and optimize our enterprise
data platform and analytics solutions. In this role, you will serve as a key hands-on engineering track
lead responsible for transforming complex business requirements into scalable, production-grade data
pipelines and Lakehouse architectures.
You will work heavily across Azure, Databricks, Microsoft Fabric, Python, PySpark, and SQL to
execute robust Medallion Lakehouse implementations (Bronze, Silver, Gold), build resilient
API/database integrations, and engineer clean data structures. Additionally, you will play an active
role in preparing data pipelines for modern AI/ML workloads, ensuring our data platform is fully
optimized for downstream analytics, LLMs, and intelligent applications.
Key Responsibilities
Data Pipeline & Lakehouse Engineering: Design, implement, and maintain high-performance
batch and real-time streaming data pipelines using Python, PySpark, and Spark SQL on
Azure Databricks and Microsoft Fabric. Build multi-tier Medallion architectures (Bronze, Silver,
Gold) using Delta Lake principles. Good experience on Genie.
Microsoft Fabric Platform Delivery: Leverage Fabric capabilities (OneLake, Direct Lake,
Lakehouses, and Data Factory Gen2) to streamline zero-copy ingestion, optimize semantic
models, and eliminate redundant ETL pipelines.
API & Systems Integration: Build resilient API connectors, CDC workflows, and data
integration pipelines to ingest data from heterogeneous sources (relational databases, flat
files, third-party REST APIs, and enterprise cloud applications).
AI & Analytics Enabling: Structure, clean, and optimize lakehouse data layers for downstream
Power BI reporting as well as AI/ML applications, ensuring data is clean, validated, and
structured for AI grounding.
Data Quality & Governance: Implement robust data validation, access controls, data lineage,
and metadata management using Unity Catalog, Fabric Governance, and automated testing
frameworks.
Technical Mentorship & CI/CD: Perform thorough code reviews, enforce engineering best
practices, optimize query performance, and manage automated deployment pipelines using
Azure DevOps or GitHub Actions.
Technical Skills & Qualifications
Primary Requirements (Must-Have)
1. Core Development Languages: Advanced hands-on proficiency in Python, PySpark, and SQL
(Spark SQL / T-SQL) for large-scale data manipulation, performance tuning, and complex
transformations.
2. Databricks Platform: Strong hands-on experience with Azure Databricks (Delta Lake, Unity
Catalog, Delta Live Tables, Auto Loader) and Genie.
3. Microsoft Fabric: Demonstrated experience building pipelines and workspace items within
Microsoft Fabric (OneLake, Fabric Lakehouses/Warehouses, Data Factory Gen2, Direct Lake
semantic models).
4. Cloud Infrastructure (Azure): Solid expertise across Azure Data Lake Storage Gen2 (ADLS
Gen2), Azure Data Factory (ADF), Azure Key Vault, and Azure SQL/Synapse.
5. AI Data Concepts & AI Readiness: Strong conceptual and practical understanding of data
engineering requirements for AI/ML workloads—including vector embeddings, RAG
(Retrieval-Augmented Generation) data ingestion, feature stores, and structuring
unstructured/semi-structured data for LLMs.
6. Data Integration & Modeling: Proven experience building Source-to-Target Mappings (STTM),
handling schema drift, parsing complex JSON/REST API payloads, and modeling Star
Schema / Dimensional Gold layers.
Preferred Skills (Good-to-Have)
Agentic AI Concepts & Frameworks: Familiarity with Agentic AI architecture patterns (e.g.,
structuring data tools/APIs for AI Agents, function calling data schemas, LangChain /
LangGraph, AutoGen, or vector stores like Azure AI Search / Pinecone).
Data Transformation & CI/CD Tools: Working knowledge of dbt, Infrastructure-as-Code
(Terraform), and automated CI/CD pipelines (Azure DevOps / GitHub Actions).
BI Optimization: Hands-on experience optimizing Direct Lake semantic models and Delta
tables for fast Power BI reporting.
Education &- Certifications
Education: Bachelor's or Master's degree in Computer Science, Information Technology, or a
related quantitative field.
Preferred Certifications (Added Advantage):
o Microsoft Certified: Fabric Data Engineer Associate (DP-700)
o Databricks Certified Data Engineer Associate / Professional
More Info
Key Skills
OneLake
Dimensional Gold layers
Azure Data Lake Storage Gen2
Azure Key Vault
Azure SQL Synapse
Delta Live Tables
GitHub Actions
Source-to-Target Mappings
AI ML workloads
Data Factory Gen2
schema drift
Auto Loader
Unity Catalog
Delta Lake
Microsoft Fabric




