Design and build shared data curation infrastructure, including reusable components, libraries, and workflows
Build scalable LLM-driven data transformation pipelines using raw sources such as the Netflix catalog and metadata
Use large-scale batch inference with attention to quality and token cost
Develop sampling strategies covering coverage, diversity, difficulty, and balance across content and member segments
Develop filtering and quality-control methods, including LLM-as-judge, evaluation-model-based scoring, deduplication, and validation
Partner with researchers to measure how curation choices affect model performance
Make curated datasets discoverable artifacts with versioning, explicit lineage, and reproducibility
Drive adoption of shared curation practices across MEDC and partner modeling teams
Requirements:
Strong software engineering in Python, with experience building reusable infrastructure, libraries, or frameworks used by other engineers and researchers
Experience building LLM-driven data generation or transformation pipelines, including synthetic data, structured outputs, or batch inference at scale
Hands-on experience with data quality methods: sampling strategies, filtering, deduplication, and model-based quality scoring such as LLM-as-judge
Understanding of how data choices affect model behavior and ability to design experiments that measure it
Experience with distributed data processing such as Spark, Ray, or similar
Excellent collaboration skills, particularly with researchers, data scientists, and platform teams
Experience with LLM evaluation systems (must-have for L6)
Technical leadership across data and evaluation infrastructure; experience setting technical direction for a multi-engineer effort (must-have for L6)
Experience with dataset versioning, lineage, and artifact management
Experience optimizing cost and throughput for large-scale LLM inference
Experience with human annotation workflows and calibrating LLM judges against human raters
Experience with pipeline orchestration frameworks such as Metaflow, Airflow, or similar
Background in recommendation systems, personalization, search, or working with content catalog and metadata
Benefits:
Health Plans
Mental Health support
401(k) Retirement Plan with employer match
Stock Option Program
Disability Programs
Health Savings and Flexible Spending Accounts
Family-forming benefits
Life and Serious Injury Benefits
Paid leave of absence programs
Full-time hourly employees accrue 35 days annually for paid time off to be used for vacation, holidays, and sick paid time off
Full-time salaried employees are immediately entitled to flexible time off
Tags
Carta de presentación
Inicia sesión para generar una carta de presentación para esta vacante.