Data pipeline execution
Worked on Big Data Technologies' orchestration platform for data pipelines and analytics jobs, with a focus on scheduling, retries, distributed state, and fault recovery.
I work on Big Data Technologies systems that help Amazon data engineers schedule, run, monitor, and recover data pipelines across Redshift and EMR. My work is mostly in orchestration, dependency management, lineage, and operational tooling.
A few examples of the backend platform work I have done across orchestration, lineage, and operations.
Worked on Big Data Technologies' orchestration platform for data pipelines and analytics jobs, with a focus on scheduling, retries, distributed state, and fault recovery.
Built lineage and completeness capabilities for dependency evaluation, blocking-dataset detection, workflow recovery, redrive, and debugging across analytics workflows.
Built tools for operational investigation and root-cause analysis, including an Amazon Bedrock RAG-based project that helped summarize failures and suggest next steps.
10+ years building backend platforms, distributed systems, and developer infrastructure.
Working primarily on Big Data Technologies' orchestration systems for data pipelines and analytics jobs, including design, implementation, deployment, operations, and technical direction for platform initiatives.
Built dependency management, data completeness, lineage, and recovery systems for analytics workflows at Amazon scale.
Developed backend services and distributed ingestion workflows for Amazon Appstore application onboarding and publishing.
Languages, platforms, and systems I use to build reliable data infrastructure.