Position Overview
We are looking for a Data Engineer with 2–3 years of hands-on experience in developing and supporting enterprise data pipelines using Databricks, Apache Spark, SQL and Python/PySpark. The candidate will work on data engineering initiatives for a leading banking client and should be capable of independently developing, troubleshooting and optimizing data pipelines while following enterprise security, data-quality and governance standards.
Key Responsibilities
- Develop and maintain ETL/ELT data pipelines using Databricks for large-volume enterprise data.
- Perform data ingestion, transformation, cleansing, validation and enrichment using PySpark and SQL.
- Develop and maintain Apache Spark DataFrame-based processing and transformation logic.
- Work extensively with Delta Lake/Delta Tables, including MERGE, schema evolution and incremental processing.
- Develop complex SQL queries using joins, CTEs, subqueries, aggregations and window functions.
- Implement incremental and batch data processing and handle structured and semi-structured data.
- Work with Bronze, Silver and Gold data layers and understand data movement across the Medallion architecture.
- Monitor and troubleshoot Databricks Jobs/Workflows and production pipeline failures.
- Implement data-quality checks and investigate data discrepancies and pipeline issues.
- Maintain technical documentation, coding standards and deployment practices.
- Follow banking-domain requirements related to data security, confidentiality, auditability and governance.
Mandatory Skills
- Databricks – Notebooks, Jobs/Workflows, clusters and basic platform administration concepts
- Apache Spark / PySpark – DataFrames, transformations, joins, aggregations and performance fundamentals
- SQL – Advanced joins, CTEs, subqueries, window functions, aggregations and query optimization
- Python – Good programming fundamentals, functions, exception handling and data processing
- Delta Lake / Delta Tables – MERGE, ACID transactions, schema evolution and incremental processing
- ETL/ELT – Data ingestion, transformation, validation, error handling and incremental loads
- Data Lake / Data Warehouse – Understanding of data modeling and enterprise data architecture
- Data Quality – Validation, reconciliation, duplicate handling and basic data-quality controls
Good to Have
- Experience with Azure/AWS cloud data services
- Exposure to Azure Data Factory or similar orchestration tools
- Git and basic CI/CD knowledge
- Experience with REST/API or file-based data ingestion
- Exposure to banking/financial data and related security/governance practices
Candidate Profile
- Capable of independently developing and supporting end-to-end data pipelines.
- Strong analytical and problem-solving skills with good debugging ability.
- Should be able to explain at least one real-world data engineering project end-to-end, including data sources, transformations, processing logic, data quality and deployment.
- Good communication and teamwork skills.
- Strong focus on data accuracy, security, confidentiality and reliability.
Education & Experience
- Experience: 2–3 years of relevant hands-on Data Engineering experience
- Qualification: B.E./B.Tech/M.Tech/MCA/M.Sc. in Computer Science, IT, Data Science, Data Engineering or related disciplines.
Key Skills: Databricks | Apache Spark | PySpark | SQL | Python | ETL/ELT | Delta Lake | Data Engineering | Data Quality

