
Software Engineer, Data Infrastructure
DatologyAIover 1 year ago
Base Salary
$180k - $300k/yr
Responsibilities
- Design, build, and maintain highly scalable data processing solutions with strong scalability, reliability, and security.
- Architect, build, and deploy backend systems and services powering the data curation platform.
- Partner with researchers and engineers to deliver new product features and research capabilities to customers.
- Operate reliable and secure systems that earn customer trust.
- Lead development of the core product and data platform and contribute to business-critical technical decisions.
Requirements
- Meaningful experience leading and building production data systems for major product initiatives.
- Experience building and managing highly scalable data processing solutions, data lakes or warehouses, distributed storage systems, workflow management systems, and supporting infrastructure.
- Proficiency in at least one Data Engineering programming language, such as Python, Scala, or Java.
- Expertise with ETL schedulers such as Airflow, Dagster, or similar frameworks.
- Experience maintaining high standards for system design, correctness, and testing.
- Experience serving as the technical lead of a Data Engineering, Platform, or Infrastructure team.
- Experience building machine learning or deep learning systems and/or data infrastructure supporting training of large machine learning models.
- Ability to own problems end-to-end and learn missing knowledge as needed.
- Collaborative, humble attitude and willingness to support team success.
Benefits
- Role is based in Redwood City, California, with four days per week in the office.
- Visa sponsorship is available for selected candidates.
- 100% covered medical, vision, and dental benefits.
- 401(k) plan with a 4% company match.
- Unlimited paid time off.
- Annual wellness and learning and development stipends.
- Daily office lunches and snacks.
- Relocation assistance for employees moving to the Bay Area.
Tech Stack
Categories
BackendData Engineering
About DatologyAI
DatologyAI builds tools to automatically select the best data on which to train deep learning models. Our tools leverage cutting-edge research—much of which we perform ourselves—to identify redundant, noisy, or otherwise harmful data points. The algorithms that power our tools are modality-agnostic—they’re not limited to text or images—and don’t require labels, making them ideal for realizing the next generation of large deep learning models. Our products allow customers in nearly any vertical to train better models for cheaper.