1 month ago
Base Salary
$129k - $193k/yr
Responsibilities
- Design and implement monitoring and alerting systems for data platform stability, reliability, and performance.
- Respond to and resolve failures affecting data pipelines, storage layers, and backend data platforms.
- Develop automation tools and scripts for deployment, monitoring, backup, recovery, and disaster recovery.
- Optimize data storage, query performance, data flows, latency, and processing speed.
- Perform incident troubleshooting, coordinate cross-team resolution efforts, and maintain high availability.
- Forecast capacity requirements and scale infrastructure to accommodate data growth.
- Document data platform architecture, configurations, and operational procedures, and provide team training.
- Ensure data platforms meet security and compliance requirements and prevent unauthorized access or data misuse.
- Collaborate with data engineering, data science, product, and development teams on data product implementation and reliability issues.
Requirements
- At least 8 years of experience as an SRE, DevOps, or Data Operations Engineer.
- Experience with AWS, GCP, or Azure cloud platforms.
- Familiarity with Kafka, Hadoop, Spark, Cassandra, HDFS, and AWS S3 or comparable big data and distributed storage technologies.
- Extensive database management experience with NoSQL databases, MySQL, and PostgreSQL.
- Proficiency with Ansible, Terraform, Kubernetes, and Docker for automating data system deployment and maintenance.
- Familiarity with CI/CD pipelines.
- Programming proficiency in Python, Go, Java, or Scala for scripting and automation tool development.
- Experience with Prometheus, Grafana, and ELK Stack for monitoring and log management.
- Strong production troubleshooting and debugging skills.
- Bachelor’s degree or higher in Computer Science, Software Engineering, or a related field, with consideration for equivalent coursework, experience, or extensive related professional experience.
- Preferred experience with Aerospike, Kafka, Snowflake, large-scale distributed systems, data quality management, data governance, or ETL pipelines.
- Preferred familiarity with containerization, microservices architecture, and Kubernetes.
Benefits
- Comcast provides benefits and personalized support options for eligible employees, including physical, financial, and emotional wellbeing resources.
- The role is based in Virginia and offers a stated base-pay range; eligible employees may also receive bonuses and other benefits.
Tech Stack
AnsibleApache CassandraApache HadoopApache KafkaApache SparkAWSAzureDockerGoGoogle Cloud PlatformGrafanaJavaKubernetesMySQLPostgreSQLPrometheusPythonScalaSnowflakeTerraform
Categories
Site Reliability
About Comcast
Welcome to Comcast. From the connectivity and platforms we provide to the content and experiences we create, we bring people together, globally. Our people think the world of our work, and that’s why our work is the best in the world.
