
Job Description & Key Responsibilities
Amazon is the worldβs largest eβcommerce and cloudβcomputing company, operating in more than 20 countries and serving millions of customers daily. In India, Amazon has built a robust ecosystem that includes retail, logistics, digital services, and a thriving marketplace for millions of sellers. The companyβs culture of customer obsession, innovation, and operational excellence has made it a top employer for technology talent across the globe.
The Selling Partner Insights and Analytics (SPIA) team is a critical part of Amazonβs sellerβcentric strategy. It powers Paragon, the secondβlargest HumanβinβtheβLoop platform at Amazon, handling over 500β―million cases a year for more than 200 internal teams and 70β―000+ users. The team focuses on turning massive, heterogeneous data sets into actionable insights that improve seller experience, reduce friction, and enable AIβdriven decision making across Amazonβs marketplace.
As a Data Engineer on the SPIA team, you will design, build, and operate scalable data pipelines and infrastructure on native AWS services. You will work closely with product owners, data scientists, and ML engineers to curate data for reporting, analytics, and largeβlanguageβmodel (LLM) training. The role demands a blend of strong SQL skills, handsβon experience with bigβdata technologies, and a passion for building reliable, costβeffective data platforms that can grow with Amazonβs expanding marketplace footprint.
Key Responsibilities:
1. Design and maintain costβeffective, highly available data pipelines on AWS (S3, Glue, EMR, Redshift, etc.).
2. Develop logical and physical data models that support Paragonβs reporting and ML workloads.
3. Build and optimize ETL jobs using Spark, Hive, and Python/Scala scripts.
4. Collaborate with business stakeholders to gather requirements and translate them into scalable data solutions.
5. Implement data governance, access controls, and compliance standards for sensitive datasets.
6. Automate monitoring, alerting, and remediation to achieve BestβAtβAmazon (BAA) operational metrics.
7. Enable selfβservice data exploration for analysts through curated data marts and catalogues.
8. Participate in code reviews, performance tuning, and capacity planning.
9. Contribute to documentation, knowledge sharing, and mentorship of junior engineers.
10. Stay updated with emerging AWS services and industry best practices to continuously improve the data platform.
Tech Stack: AWS (S3, Redshift, Glue, EMR, Lambda), Hadoop ecosystem (Hive, Spark), SQL, PL/SQL, SparkSQL, Python, Scala, data modeling tools, ETL frameworks (Informatica/SSIS alternatives), Linux/KornShell scripting.
Growth Path: Starting as a Data Engineer, you can progress to Senior Data Engineer, Lead Data Engineer, and eventually Data Architect or Manager β Data Engineering, with opportunities to work on highβimpact projects across Amazonβs global marketplace and AI initiatives.
Why Join Amazon? You will work on one of the largest data platforms in the world, influence decisions that affect millions of sellers and buyers, and be part of a culture that rewards invention, ownership, and relentless customer focus. The role offers exposure to cuttingβedge AI/LLM projects, a collaborative environment, and clear career advancement pathways.