Join a fast‑growing frontier tech team building backend services and AI agents that power end‑to‑end retail operations. You’ll design, ship, and own highly available systems that enable autonomous incident handling and root‑cause analysis. What You'll Do Build scalable backend services and event pipelines meeting latency, throughput, and availability targets. Optimize performance via profiling, efficient data access, concurrency control, and load testing. Engineer resilience with replication, failover, backpressure, graceful degradation, and disaster‑recovery validation. Create incident‑correlation and root‑cause systems that reconstruct timelines and attach evidential links. Develop AI agents that investigate incidents, execute permitted actions, and maintain audit trails. Own production health through observability, on‑call rotation, postmortems, and preventive improvements. What You Need Proven ownership of high‑traffic production services with clear performance and availability metrics. Strong Python and distributed‑systems skills: APIs, async processing, data modelling, consistency, idempotency. Hands‑on experience with replication, failover, backup validation, and meeting RTO/RPO objectives. Deep production debugging using logs, metrics, tracing, profiling, and query analysis. Delivered LLM‑based applications or agents with tool‑calling, structured outputs, and performance monitoring. Ability to make sound operational judgments on evidence, permissions, rollbacks, and human intervention. Good to Have Experience in commerce, fulfilment, logistics, or payments workflow automation. Familiarity with Kubernetes/GCP, MongoDB, and React ecosystems. Built operational attribution or observability systems using Prometheus and Grafana. The Opportunity Fynd offers limitless growth with a culture that encourages ownership and continuous learning, and the team works from the office five days a week to foster collaboration.
SDE – 2/3 | Backend & Agentic AI | Granary
Full Time
