
SRE Monitoring Platform Software Engineer (Early Career / Temporary)
BitDeer Technologies Group11 days ago
Remote, United States or San Jose, CA, USAEntry Level
Responsibilities
- Build collection, ingestion, query, storage, alerting, correlation, SLO, topology, cluster-health, remediation, workflow, inspection, and job-scheduling components for the SRE platform.
- Write clean production code and unit, integration, contract, chaos, and soak tests.
- Instrument services with metrics, logs, and traces using OpenTelemetry and build observability dashboards.
- Write operational runbooks and participate in on-call as a shadow before taking primary responsibility.
- Ship components through the GitOps and CI/CD release pipeline while meeting declared SLOs and maintaining drift-free systems.
Requirements
- Have 0–2 years of software engineering experience; strong projects or internships are accepted for new graduates.
- Be proficient in one programming language, preferably Go, or alternatively Python, Java, or Rust.
- Understand data structures, algorithms, concurrency, TCP/HTTP networking, operating-system concepts, and distributed-systems fundamentals.
- Have some hands-on exposure to monitoring or observability, including Prometheus, Grafana, Loki, PromQL, or service instrumentation.
- Be familiar with Linux, shell usage, system logs, debugging tools, Kubernetes basics, Git, pull requests, and CI pipelines.
- Practice writing unit and integration tests and communicate clearly in written and spoken English.
- Preferred qualifications include monitoring or observability internships or projects, GPU/AI infrastructure exposure, AIOps or anomaly-detection experience, and open-source cloud-native contributions.
Benefits
- Temporary early-career role with senior and principal engineer mentorship.
- Guided progression from shadowed on-call participation toward independent ownership within 12 months.
Categories
Site Reliability