Actively Hiring

Senior Databricks Data Engineer - PySpark - Bangalore

Location

Bangalore, Karnataka / Pune, Maharashtra / Remote (Global), Global

Work Mode

Remote, On-site

Experience

4-7 Yrs

Salary Range

INR 15-25 LPA

Job Type

Permanent

Openings

10 positions

Job Description

Hiring for: A technology company delivering enterprise data engineering and analytics solutions, with a strong focus on modern cloud data platforms and Databricks.

Role: Senior Databricks Data Engineer - PySpark - Bangalore/ Pune

Positions: 10

Experience: 4 to 7 years

Location(s): Remote (Bangalore, Pune)

Type: On-site, Remote / Permanent

Salary: Up to INR 25 LPA

Notice Period: Immediate to 30 days


Must Have: Databricks Development, Pyspark, Python, SQL, Unity Catalogs,


Job Summary

We are seeking an experienced PySpark ETL Developer with 4+ years of experience in designing, developing, and optimizing enterprise ETL pipelines using PySpark. The ideal candidate should have strong expertise in Python, Apache Spark, Databricks, and Snowflake, along with hands-on experience in processing large-scale data and building scalable data engineering solutions. You will work closely with business stakeholders and cross-functional teams to develop high-performance data pipelines that support business intelligence and analytics initiatives.


Key Responsibilities

  • Design, develop, and maintain scalable ETL pipelines using PySpark and Spark SQL.
  • Collaborate with stakeholders to gather business requirements and translate them into efficient data engineering solutions.
  • Extract data from multiple sources, including databases, APIs, data lakes, files, and streaming platforms.
  • Transform, cleanse, and validate data using PySpark to ensure high data quality and consistency.
  • Develop and optimize Spark jobs for performance, scalability, and efficient resource utilization.
  • Build and maintain batch and streaming data pipelines to support real-time and near real-time processing.
  • Load transformed data into data lakes, data warehouses, and analytical platforms.
  • Implement robust error handling, logging, monitoring, and troubleshooting mechanisms for ETL workflows.
  • Document ETL processes, data lineage, transformation logic, and technical specifications.
  • Develop unit, integration, and performance tests to ensure reliable and accurate data processing.
  • Collaborate with cross-functional teams to deliver scalable, secure, and high-quality data solutions.


Required Skills

  • 4+ years of experience as a Data Engineer or ETL Developer.
  • Strong hands-on experience with PySpark, Apache Spark, Spark SQL, and Python.
  • Expertise in Databricks.
  • Strong understanding of ETL design, data transformation, and data integration.
  • Experience with Data Lakes, Data Warehouses, and Big Data technologies.
  • Knowledge of distributed computing, parallel processing, and data partitioning concepts.
  • Strong SQL skills with experience in performance tuning and query optimization.
  • Experience working with batch and streaming data processing.
  • Excellent analytical, debugging, and problem-solving skills.
  • Strong verbal and written communication skills.


Preferred Skills

  • Experience with cloud platforms such as AWS or Azure.
  • Knowledge of Hadoop ecosystem and modern data engineering technologies.
  • Experience with API integration and data ingestion from multiple data sources.
  • Familiarity with CI/CD pipelines and Agile development methodologies.
  • Exposure to enterprise-scale data engineering projects.


Screening Questions

Please note that you will be asked to answer these questions during the application process:

  • 1.How many years of hands-on Databricks implementation experience do you have?
  • 2.Current location
  • 3.Current CTC (Lakhs per annum)
  • 4.Expected CTC (Lakhs per annum)
  • 5.Notice period
  • 6.Are you currently serving notice period?
  • 7.How many years of building Databricks pipelines end-to-end from source ingestion to final data layer do you have?
  • 8.How many years of experience you have working with Medallion Architecture (Bronze/Silver/Gold)?
  • 9.Mention orchestration tools have you used hands-on
  • 10.Which sources have you personally integrated - APIs, DBs, files, streaming?
  • 11.How many years of DEDICATED experience you have working on a Databricks project where you onboarded new data sources and built the pipeline from scratch?
  • 12.How many years of experience you have working on Databricks/Spark performance optimization?
  • 13.How many years of hands-on experience do you have in Python development?
  • 14.How many years of hands-on experience do you have working with PySpark on complex, large-scale data engineering systems?
  • 15.How many years of hands-on experience do you have working with SQL for complex data engineering projects, including query optimization and performance tuning?

Skills & Technologies

Required Skills
Apache HadoopApache Spark Performance OptimizationBig Data ProcessingCI/CDCloud Platforms (AWS/Azure/GCP)Data EngineeringData IngestionData LakeData WarehousingDatabricksDatabricks WorkflowsDistributed ComputingETL Pipeline DevelopmentMedallion ArchitecturePipeline OrchestrationPySparkPythonSnowflakeSQL
EmailVerifyUploadDetails

Apply for this Role

Submit your details to apply for this role.