Apply with hirly
AI Data Engineer
Infosys · Bangalore, India
Upload your resume to see how well you match this job — free, in seconds, no account needed.
Your resume is used only to score it against this job. If you don't create an account, it is deleted within 24 hours.
About the job: Join a team where data engineering meets intelligent automation to power next-generation analytics and AI experiences. In this role, you’ll help build reliable, scalable data pipelines and enable GenAI-ready datasets that support experimentation, model development, and business decision-making. You’ll collaborate closely with analysts, data scientists, and platform teams to turn raw, fragmented data into trusted, well-modeled, and accessible assets. If you enjoy solving complex data challenges, improving performance, and bringing structure to fast-moving AI initiatives, this is a great opportunity to grow your impact. You’ll work in a culture that values ownership, continuous learning, and practical innovation—where clean data, strong engineering, and thoughtful collaboration come together to deliver real outcomes. Responsibilities Key Responsibilities:
- Design, build, and maintain robust ETL/ELT pipelines to ingest, transform, and curate data from multiple sources.
- Develop and optimize data models and curated datasets to support analytics, reporting, and AI/ML workloads.
- Implement data quality checks, validation rules, and monitoring to ensure accuracy, completeness, and reliability.
- Enable GenAI initiatives by preparing high-quality datasets for downstream consumption (e.g., feature-ready and retrieval-ready data).
- Collaborate with cross-functional teams to gather requirements, define data contracts, and deliver reusable data assets.
- Troubleshoot pipeline failures and performance bottlenecks; improve scalability, latency, and cost efficiency.
- Maintain documentation for pipelines, transformations, lineage, and operational runbooks to support maintainability. Minimum Qualifications:
- BTECH, MTECH, MCA, MSC or equivalent education.
- 3–5 years of experience in data engineering with hands-on ownership of production-grade pipelines.
- Strong experience in ETL processes including extraction, transformation, orchestration, and scheduling.
- Working exposure to GenAI-oriented data preparation needs and supporting AI/ML data workflows.
- Ability to collaborate with stakeholders to translate requirements into scalable data solutions. Technical requirements Good to have skills: SQL, Python, Apache Spark, Airflow, Data Modeling, data engineering, ai Additional responsibilities Preferred Qualifications:
- Experience designing scalable data architectures and implementing reusable data frameworks for multiple use cases.
- Familiarity with building datasets for GenAI use cases such as retrieval workflows and knowledge augmentation patterns.
- Proven ability to improve pipeline reliability through automation, alerting, and proactive monitoring.
- Strong problem-solving skills with a track record of optimizing transformations and reducing end-to-end processing time.
- Experience working in agile teams and contributing to code reviews, documentation, and engineering best practices. Education MCA,MSc,MTech,Bachelor of Engineering,BTech