Join a fast‑moving, innovation‑driven team building enterprise‑scale real‑time data pipelines. You’ll design, develop, and optimise streaming applications that power Genpact’s AI solutions for global clients. What You'll Do Design and maintain high‑performance Apache Flink streaming applications. Implement transformation logic using Flink, Kafka Streams, KSQLDB, and SMTs. Build scalable mapping frameworks to harmonise source data into enterprise models. Optimise pipelines for low latency, high throughput, and efficient resource use. Handle late‑arriving events with watermarks, checkpointing, and state management. Integrate streams with Kafka, Schema Registry, APIs, and downstream analytics. Develop CI/CD pipelines and automated tests for streaming jobs. What You Need 6–10 years of data engineering experience focusing on real‑time streaming. Deep hands‑on expertise with Apache Flink and Apache Kafka ecosystem. Strong Java or Scala programming skills; Python a plus. Proficiency with event‑time processing, watermarks, windowing, and stateful streams. Experience using Avro, Protobuf, JSON and Schema Registry for serialization. Advanced SQL skills on large structured and semi‑structured datasets. Familiarity with cloud platforms (Azure, AWS, or GCP). Good to Have Experience with Confluent Platform and enterprise Kafka deployments. Knowledge of modern lakehouse solutions like Databricks, Snowflake, or Delta Lake. Exposure to CDC patterns and API integration. Agile/Scrum delivery background. The Opportunity Genpact’s AI Gigafactory accelerates advanced technology solutions, letting you work on cutting‑edge AI projects that drive measurable business outcomes. You’ll collaborate with architects and business stakeholders in a values‑driven environment that rewards curiosity and impact.
Sr. Data Engineer – Data Engineering 4B
Full Time
