Senior Databricks Data Engineer - PySpark - Bangalore
Bangalore, Karnataka / Pune, Maharashtra / Remote (Global), Global
Remote, On-site
4-7 Yrs
INR 15-25 LPA
Permanent
10 positions
Job Description
Hiring for: A technology company delivering enterprise data engineering and analytics solutions, with a strong focus on modern cloud data platforms and Databricks.
Role: Senior Databricks Data Engineer - PySpark - Bangalore/ Pune
Positions: 10
Experience: 4 to 7 years
Location(s): Remote (Bangalore, Pune)
Type: On-site, Remote / Permanent
Salary: Up to INR 25 LPA
Notice Period: Immediate to 30 days
Must Have: Databricks Development, Pyspark, Python, SQL, Unity Catalogs,
Job Summary
We are seeking an experienced PySpark ETL Developer with 4+ years of experience in designing, developing, and optimizing enterprise ETL pipelines using PySpark. The ideal candidate should have strong expertise in Python, Apache Spark, Databricks, and Snowflake, along with hands-on experience in processing large-scale data and building scalable data engineering solutions. You will work closely with business stakeholders and cross-functional teams to develop high-performance data pipelines that support business intelligence and analytics initiatives.
Key Responsibilities
- Design, develop, and maintain scalable ETL pipelines using PySpark and Spark SQL.
- Collaborate with stakeholders to gather business requirements and translate them into efficient data engineering solutions.
- Extract data from multiple sources, including databases, APIs, data lakes, files, and streaming platforms.
- Transform, cleanse, and validate data using PySpark to ensure high data quality and consistency.
- Develop and optimize Spark jobs for performance, scalability, and efficient resource utilization.
- Build and maintain batch and streaming data pipelines to support real-time and near real-time processing.
- Load transformed data into data lakes, data warehouses, and analytical platforms.
- Implement robust error handling, logging, monitoring, and troubleshooting mechanisms for ETL workflows.
- Document ETL processes, data lineage, transformation logic, and technical specifications.
- Develop unit, integration, and performance tests to ensure reliable and accurate data processing.
- Collaborate with cross-functional teams to deliver scalable, secure, and high-quality data solutions.
Required Skills
- 4+ years of experience as a Data Engineer or ETL Developer.
- Strong hands-on experience with PySpark, Apache Spark, Spark SQL, and Python.
- Expertise in Databricks.
- Strong understanding of ETL design, data transformation, and data integration.
- Experience with Data Lakes, Data Warehouses, and Big Data technologies.
- Knowledge of distributed computing, parallel processing, and data partitioning concepts.
- Strong SQL skills with experience in performance tuning and query optimization.
- Experience working with batch and streaming data processing.
- Excellent analytical, debugging, and problem-solving skills.
- Strong verbal and written communication skills.
Preferred Skills
- Experience with cloud platforms such as AWS or Azure.
- Knowledge of Hadoop ecosystem and modern data engineering technologies.
- Experience with API integration and data ingestion from multiple data sources.
- Familiarity with CI/CD pipelines and Agile development methodologies.
- Exposure to enterprise-scale data engineering projects.
Screening Questions
Please note that you will be asked to answer these questions during the application process:
- 1.How many years of hands-on Databricks implementation experience do you have?
- 2.Current location
- 3.Current CTC (Lakhs per annum)
- 4.Expected CTC (Lakhs per annum)
- 5.Notice period
- 6.Are you currently serving notice period?
- 7.How many years of building Databricks pipelines end-to-end from source ingestion to final data layer do you have?
- 8.How many years of experience you have working with Medallion Architecture (Bronze/Silver/Gold)?
- 9.Mention orchestration tools have you used hands-on
- 10.Which sources have you personally integrated - APIs, DBs, files, streaming?
- 11.How many years of DEDICATED experience you have working on a Databricks project where you onboarded new data sources and built the pipeline from scratch?
- 12.How many years of experience you have working on Databricks/Spark performance optimization?
- 13.How many years of hands-on experience do you have in Python development?
- 14.How many years of hands-on experience do you have working with PySpark on complex, large-scale data engineering systems?
- 15.How many years of hands-on experience do you have working with SQL for complex data engineering projects, including query optimization and performance tuning?
Skills & Technologies
Apply for this Role
Submit your details to apply for this role.