Apply with hirly
PySpark Developer
Infosys · Bangalore, India
Upload your resume to see how well you match this job — free, in seconds, no account needed.
Your resume is used only to score it against this job. If you don't create an account, it is deleted within 24 hours.
Already have an account? Sign in to see your saved application
We are looking for a skilled PySpark Developer with up to 3 years of experience in building and maintaining large-scale data processing pipelines. The ideal candidate should have hands-on experience in PySpark, Python, SQL, and Big Data technologies and be capable of developing efficient ETL workflows for enterprise data platforms. Responsibilities Design, develop, and maintain data pipelines using PySpark. Develop ETL/ELT processes for ingesting, transforming, and loading large volumes of data. Write optimized PySpark code for data processing and transformation. Work with structured and semi-structured data from multiple sources. Develop and optimize SQL queries for data extraction and validation. Troubleshoot data quality and performance issues. Collaborate with Data Engineers, Analysts, and Business teams to understand requirements. Participate in code reviews and follow data engineering best practices. Monitor and support production data pipelines. Technical requirements Strong experience in PySpark. Good programming knowledge of Python. Hands-on experience with SQL. Understanding of ETL/ELT concepts and data warehousing. Experience working with large datasets and distributed processing. Knowledge of Spark SQL, DataFrames, and Spark transformations. Familiarity with Linux/Unix environment. Additional responsibilities Exposure to cloud platforms such as AWS, Azure, or GCP. Knowledge of Databricks. Experience with workflow orchestration tools such as Airflow. Understanding of CI/CD concepts. Education Bachelor Of Comp. Applications,Bachelor Of Computer Science,Bachelor Of Science,Bachelor of Engineering,Bachelor Of Technology