3 months ago
Base Salary
$190k - $240k/yr
Responsibilities
- Own the reliability and availability of data platform infrastructure across all environments.
- Enforce and improve disciplined promotion from lower environments to production.
- Define and uphold standard operating procedures for deployments, maintenance windows, and change management.
- Instrument and monitor platform health and build meaningful alerting.
- Participate in architecture and deployment discussions and challenge deployments that are not ready.
- Collaborate with data scientists, engineers, and product managers on infrastructure needs.
- Identify and remediate reliability risks before they become incidents.
- Support customer-facing and internal systems with a stability-first approach.
Requirements
- At least 3 years of experience in cloud infrastructure, SRE, or platform engineering.
- Experience with high-availability architecture, including blue/green deployments, data replication, and load balancing.
- Experience with workflow orchestration and DAG-based schedulers or large-scale job scheduling.
- Strong Linux fundamentals and scripting experience with Bash, Python, or similar tools.
- Experience with distributed data processing or managing clusters that run distributed processing frameworks.
- Experience with containerization and orchestration.
- Experience with data ingestion, ETL, streaming systems, message queues, or pipelines.
- Experience with infrastructure-as-code and provisioning.
- Experience operating OLAP and OLTP databases, including query patterns, indexing, and operational care.
- Experience with monitoring, logging, and observability tooling.
- Experience administering and scaling managed data platforms such as Databricks.
- Understanding of network infrastructure, including load balancers, DNS, auto-scaling, multi-region topologies, and proxies.
- Understanding of least-privilege security, secrets management, and controls for data systems.
- MLOps concepts or tooling are a plus.
- AWS experience is preferred; GCP or Azure experience translates.
- Strong operational discipline, communication, ownership, independence, humility, and methodical execution are valued.
Benefits
- Equity compensation
- Health insurance coverage for the employee and dependents
- 401(k), FSA, and commuter benefits
- $150 monthly spending account
- $1,000 annual continued education benefit
- $500 Newbie Productivity Perk
- Unlimited PTO and sick days
- Monthly company wellness day off
- Snacks, drinks, and catered lunches at the office
- Team-building events
- Hybrid schedule with two days per week in the office
Tech Stack
Amazon RedshiftApache AirflowApache FlinkApache KafkaApache SparkAWSAzureBashClickHouseDatabricksDatadogDockerGoogle Cloud PlatformHelmKibanaKubernetesLinuxPostgreSQLPrometheusPythonTerraform
Categories
DevOpsSite Reliability
About Tatari
Tatari is building the infrastructure to modernize TV advertising for Brands, Agencies, and Publishers. Clients include Made In, Daily Harvest, Wpromote, and Fubo. Recognized by Business Insider as one of the Hottest Ad Tech Companies of 2023 and by Digiday as the Best CTV Platform of 2024, Tatari is headquartered in San Francisco with further offices in Los Angeles and New York. For additional information, please visit tatari.tv
