Tasks:
- * Build data cleaning and annotation quality tools
- * Build data services and visualizations
- * Build high-throughput, low-latency distributed processing systems
- * Coordinate with model and engineering teams to deliver solutions
- * Design data models
- * Design data pipeline core workflows
- * Develop data management platform from collection to training
- * Develop data mining tools for model error analysis
- * Enable unified data access and collaboration
- * Implement data versioning, lineage, metadata, and retrieval
- * Optimize end-to-end data pipeline performance
- * Process batch and real-time data
- * Support autonomous driving, smart cockpit, and robotics data workflows
- * Support logging and connected-vehicle data collection
- * Synchronize, clean, and standardize data
Perks/Benefits:
Skills/Tech stack required:
[Apache Iceberg] [Batch Processing] [Columnar Storage] [Data cleaning] [Data Lake] [Data Lineage] [Data Pipelines] [Data Versioning] [Data Warehouse] [Distributed Systems] [Docker] [ETL] [Go] [Java] [Kafka] [Kubernetes] [Lance] [Metadata Management] [MongoDB] [MySQL] [Performance Tuning] [PostgreSQL] [Pulsar] [Python] [RabbitMQ] [Real Time] [Real-time Processing] [Redis] [Stream processing] [Time processing]
Educational requirements:
[Bachelor's Degree]
Role(s):
[Data Engineer] [Data Pipeline Engineer] [Engineer] [Pipeline Engineer] [Senior Data Engineer]