1 day ago
Remote, United States +4 moreSenior
Responsibilities
- Implement scalable Infrastructure-as-Code patterns with Terraform for standardized cloud provisioning.
- Own and evolve the Kubernetes platform, including EKS or self-managed Kubernetes environments.
- Optimize CI/CD pipelines and improve deployment frequency, lead time, and release confidence.
- Design secure networking, IAM, and secrets management strategies across environments.
- Improve observability through metrics, logs, and tracing with DataDog.
- Optimize cloud costs through rightsizing, autoscaling, and architectural improvements.
- Implement disaster recovery, backup, and multi-region resilience strategies.
- Modernize manually managed infrastructure into automated, testable, and reproducible systems.
- Introduce infrastructure tooling and architectural changes, and drive adoption through documentation, workshops, and hands-on support.
- Partner with engineering teams to reduce friction in CI/CD, deployments, and cloud environments.
- Lead incident response processes and improve postmortem outcomes.
- Maintain reliability, infrastructure stability, operational efficiency, and cost-optimization targets.
Requirements
- 8+ years of experience in DevOps, SRE, or infrastructure engineering.
- Proven experience designing and operating production Kubernetes environments at scale.
- Deep hands-on experience with AWS infrastructure and cloud networking.
- Strong experience building and maintaining Terraform modules across large cloud environments.
- Experience owning CI/CD systems and improving DORA metrics.
- Experience leading incident response and driving meaningful postmortem outcomes.
- Strong understanding of distributed systems, event-driven architectures, Kafka, and PostgreSQL performance.
- Experience modernizing legacy infrastructure and eliminating manual operational toil.
- Ability to take infrastructure projects from ambiguity to production without daily direction.
- Ability to build trust across teams while improving reliability.
- Nice to have experience coding applications, operating high-throughput Kafka clusters, tuning PostgreSQL and Redis, implementing autoscaling, using service mesh technologies, building internal developer platforms, applying zero-trust networking and policy-as-code, operating multi-region systems, and introducing SLO, error-budget, or chaos-testing frameworks.
Benefits
- Medical, dental, and vision health care plan.
- 401(k) retirement plan.
- Life insurance.
- Flexible paid time off.
- Nine paid holidays.
- Family leave.
- Remote work arrangement.
- Orlando-based associates may use hybrid work arrangements.
- Free food and snacks in Orlando.
- Wellness resources.
- Remote hires must travel to Orlando, Florida at least twice per year for Town Halls and team collaboration, in addition to Orlando orientation.
Tech Stack
Apache KafkaAWSBashDatadogGitHub ActionsJavaScriptKubernetesPostgreSQLPythonRedisTerraformTypeScript
Categories
About Worth AI
Worth AI builds a SaaS platform that automates onboarding and underwriting for business credit, using AI-driven risk models, real-time SMB data, and continuous portfolio monitoring. It sells to banks, lenders, and fintechs that need faster credit decisions and portfolio visibility. Founded in 2023 and headquartered in Winter Park, Florida, the company is privately held.
