Join a leading creative‑technology company to shape its data platform, bridging traditional big‑data pipelines with cutting‑edge generative AI infrastructure. You’ll build robust pipelines on Databricks and deliver high‑quality datasets for AI models.
What You’ll Do
- Design scalable distributed data processing systems.
- Build ETL/ELT pipelines for complex datasets on Databricks.
- Engineer high‑quality data sets for generative AI models.
- Implement retrieval architectures and vector‑database solutions.
- Develop backend services (FastAPI/Flask/Node.js) to expose data products.
- Optimize code performance, data quality, and storage efficiency.
What You Need
- Expert in Apache Spark, Hadoop, Kafka, and Databricks/Delta Lake.
- Proficient with Python ecosystem: PySpark, Pandas, NumPy, and performance tuning.
- Advanced experience deploying on AWS or Azure cloud platforms.
- Deep knowledge of relational and non‑relational data modeling.
- Minimum 7 years of professional software engineering experience.
- Demonstrated ability to ship production‑grade data pipelines.
Good to Have
- Experience with LangChain, LlamaIndex, and vector databases such as Pinecone or Milvus.
- Building APIs using FastAPI, Flask, or Node.js.
- Containerization with Docker, Kubernetes, and CI/CD pipelines.
The Opportunity
You’ll work at Adobe, a global leader with 30,000 employees, building AI‑powered products like Creative Cloud and Firefly. The role offers direct impact on Adobe’s next‑generation data platform.
